Voice Assistant Control of Unconfigured Apps Using Synonym Biasing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automated assistants face limitations when interacting with applications that are not pre-configured for their functionality or are rendered in a non-native language, leading to inefficient resource usage and misrecognition in speech processing.
Innovation Solution
The automated assistant processes application interface data to identify synonymous terms and biases speech processing towards these terms, allowing it to accurately interpret user commands and control applications that are not pre-configured or in a different language.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the automated assistant uses standard speech processing to interact with applications not pre-configured for its functionality, then the assistant can attempt to control any application, but this leads to misrecognition and inefficient resource usage
Solution Approach 1:
The system performs preliminary analysis of the application interface by capturing screenshots and extracting visual elements before the user provides speech commands. This preliminary action involves identifying buttons, links, and interface elements, then creating a mapping between these visual elements and their functional equivalents in the assistant's command system. This preparation enables accurate speech recognition without requiring pre-configured application integration.
2Adaptability or versatility
If the automated assistant processes content in a non-native language interface, then the assistant can interact with applications in different languages, but this increases complexity of language processing
Solution Approach 1:
The system introduces a visual intermediary layer between the user's speech commands and the application interface. Instead of directly processing text in the application's native language, the system captures the visual representation of the interface (screenshots), identifies elements visually, and creates a translation layer that maps visual elements to the assistant's command language. This intermediary approach enables multi-language support without requiring the assistant to process multiple language texts directly.
3Ease of operation
If the user frequently switches between automated assistant and manual touch interaction, then the user can control applications, but this wastes computing resources
Solution Approach 1:
The system creates a universal interface layer that handles both visual display and speech control functions through a single integrated system. The same visual element identification and mapping infrastructure that displays application content also processes speech commands, eliminating the need for separate control mechanisms. This multi-functional approach enables seamless voice-based control without requiring additional hardware or separate processing systems.
Data Source
AI summary
Implementations set forth herein relate to an automated assistant that can interact with applications that may not have been pre-configured for interfacing with the automated assistant. The automated assistant can identify content of an application interface of the application to determine synonymous terms that a user may speak when commanding the automated assistant to perform certain tasks. Speech processing operations employed by the automated assistant can be biased towards these synonymous terms when the user is accessing an application interface of the application. In some implementations, the synonymous terms can be identified in a responsive language of the automated assistant when the content of the application interface is being rendered in a different language. This can allow the automated assistant to operate as an interface between the user and certain applications that may not be rendering content in a native language of the user.


