Voice Control for Non-Voice Applications via Screen Data Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users of computing devices often rely on tactile input methods like remote controls, mice, and keyboards to interact with displayed content, but there is a need for additional input means, particularly for third-party applications that are not optimized for voice control.
Innovation Solution
The system enables voice control of computing devices by capturing audio commands, performing automatic speech recognition, and using natural language understanding to determine user intents, even for applications not originally designed for voice interaction, by sending directive data to the device to perform actions on displayed content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If voice control is implemented for third-party applications, then ease of operation is improved, but device complexity increases
Solution Approach 1:
The patent introduces a remote system as an intermediary that handles voice command processing, speech recognition, and natural language understanding. This external mediator processes voice inputs and translates them into actionable directives, allowing the user device to implement voice control without significant internal complexity increases. The remote system acts as the mediator between the user's voice commands and the application execution environment.
Solution Approach 2:
The patent creates a universal voice control framework that can interface with multiple third-party applications across different platforms. The system uses standardized protocols to enable voice control functionality in applications that were not originally designed for voice interaction, making the voice control capability universally applicable rather than requiring application-specific implementations.
2Adaptability or versatility
If voice control is added to applications not designed for it, then adaptability is improved, but manufacturing precision deteriorates
Solution Approach 1:
The patent implements a dynamic adaptation mechanism where the system adjusts its voice control approach based on the specific application and context. The framework dynamically determines which portions of displayed content can be selected and interacted with via voice commands, adapting to different application types and interfaces rather than applying a rigid, pre-defined voice control scheme.
Solution Approach 2:
The system changes operational parameters such as confidence thresholds, selection criteria, and interaction models based on the application context. By dynamically adjusting these parameters, the system maintains appropriate precision and reliability when enabling voice control in applications not originally designed for it, rather than using fixed, imprecise parameters for all scenarios.
3Speed
If voice commands are processed locally, then speed is improved, but use of energy increases
Solution Approach 1:
The patent segments the voice processing workload between the user device and the remote system. The user device performs local audio capture and basic preprocessing, while the computationally intensive speech recognition and natural language understanding tasks are performed remotely. This segmentation allows the device to maintain responsiveness for immediate feedback while offloading energy-intensive processing to the remote system.
Data Source
AI summary
Systems and methods for voice control of computing devices are disclosed. Applications may be downloaded and/or accessed by a device having a display, and content associated with the applications may be displayed. Many applications do not allow for voice commands to be utilized to interact with the displayed content. Improvements described herein allow for non-voice-enabled applications to utilize voice commands to interact with displayed content by determining screen data displayed by the device and utilizing the screen data to determine an intent associated with the application. Directive data to perform an action corresponding to the intent may be sent to the device and may be utilized to perform the action on an object associated with the displayed content.


