Voice Control for Non-Voice Applications via Screen Data Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users of computing devices often rely on tactile input methods like remote controls, mice, and keyboards to interact with displayed content, but there is a need for additional input means, particularly for third-party applications that are not optimized for voice control.

Innovation Solution

The system enables voice control of computing devices by capturing audio commands, performing automatic speech recognition, and using natural language understanding to determine user intents, even for applications not originally designed for voice interaction, by sending directive data to the device to perform actions on displayed content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice control is implemented for third-party applications, then ease of operation is improved, but device complexity increases

Engineering Contradiction:
Improveease of operationVSAvoiddevice complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent introduces a remote system as an intermediary that handles voice command processing, speech recognition, and natural language understanding. This external mediator processes voice inputs and translates them into actionable directives, allowing the user device to implement voice control without significant internal complexity increases. The remote system acts as the mediator between the user's voice commands and the application execution environment.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a universal voice control framework that can interface with multiple third-party applications across different platforms. The system uses standardized protocols to enable voice control functionality in applications that were not originally designed for voice interaction, making the voice control capability universally applicable rather than requiring application-specific implementations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If voice control is added to applications not designed for it, then adaptability is improved, but manufacturing precision deteriorates

Engineering Contradiction:
ImproveadaptabilityVSAvoidmanufacturing precision
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent implements a dynamic adaptation mechanism where the system adjusts its voice control approach based on the specific application and context. The framework dynamically determines which portions of displayed content can be selected and interacted with via voice commands, adapting to different application types and interfaces rather than applying a rigid, pre-defined voice control scheme.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes operational parameters such as confidence thresholds, selection criteria, and interaction models based on the application context. By dynamically adjusting these parameters, the system maintains appropriate precision and reliability when enabling voice control in applications not originally designed for it, rather than using fixed, imprecise parameters for all scenarios.

Inventive Principle:
Principle #35Parameter changes

3Speed

If voice commands are processed locally, then speed is improved, but use of energy increases

Engineering Contradiction:
ImprovespeedVSAvoiduse of energy
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The patent segments the voice processing workload between the user device and the remote system. The user device performs local audio capture and basic preprocessing, while the computationally intensive speech recognition and natural language understanding tasks are performed remotely. This segmentation allows the device to maintain responsiveness for immediate feedback while offloading energy-intensive processing to the remote system.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10884701B2Voice enabling applications
Publication Date: 2021.01.05 AMAZON TECH INC
  • US10884701B2 patent drawing
  • US10884701B2 patent drawing
  • US10884701B2 patent drawing

AI summary

Systems and methods for voice control of computing devices are disclosed. Applications may be downloaded and/or accessed by a device having a display, and content associated with the applications may be displayed. Many applications do not allow for voice commands to be utilized to interact with the displayed content. Improvements described herein allow for non-voice-enabled applications to utilize voice commands to interact with displayed content by determining screen data displayed by the device and utilizing the screen data to determine an intent associated with the application. Directive data to perform an action corresponding to the intent may be sent to the device and may be utilized to perform the action on an object associated with the displayed content.