Voice Assistant Control of Unconfigured Apps Using Synonym Biasing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automated assistants face limitations when interacting with applications that are not pre-configured for their functionality or are rendered in a non-native language, leading to inefficient resource usage and misrecognition in speech processing.

Innovation Solution

The automated assistant processes application interface data to identify synonymous terms and biases speech processing towards these terms, allowing it to accurately interpret user commands and control applications that are not pre-configured or in a different language.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the automated assistant uses standard speech processing to interact with applications not pre-configured for its functionality, then the assistant can attempt to control any application, but this leads to misrecognition and inefficient resource usage

Engineering Contradiction:
Improveability to control applicationsVSAvoidspeech recognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system performs preliminary analysis of the application interface by capturing screenshots and extracting visual elements before the user provides speech commands. This preliminary action involves identifying buttons, links, and interface elements, then creating a mapping between these visual elements and their functional equivalents in the assistant's command system. This preparation enables accurate speech recognition without requiring pre-configured application integration.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If the automated assistant processes content in a non-native language interface, then the assistant can interact with applications in different languages, but this increases complexity of language processing

Engineering Contradiction:
Improvelanguage compatibilityVSAvoidspeech processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system introduces a visual intermediary layer between the user's speech commands and the application interface. Instead of directly processing text in the application's native language, the system captures the visual representation of the interface (screenshots), identifies elements visually, and creates a translation layer that maps visual elements to the assistant's command language. This intermediary approach enables multi-language support without requiring the assistant to process multiple language texts directly.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If the user frequently switches between automated assistant and manual touch interaction, then the user can control applications, but this wastes computing resources

Engineering Contradiction:
Improveuser interaction capabilityVSAvoidcomputing resource consumption
Core Design Contradiction:
Ease of operationVSLoss of energy

Solution Approach 1:

The system creates a universal interface layer that handles both visual display and speech control functions through a single integrated system. The same visual element identification and mapping infrastructure that displays application content also processes speech commands, eliminating the need for separate control mechanisms. This multi-functional approach enables seamless voice-based control without requiring additional hardware or separate processing systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12603092B2Automated assistant control of non-assistant applications via identification of synonymous term and/or speech processing biasing
Publication Date: 2026.04.14 GOOGLE LLC
  • US12603092B2 patent drawing
  • US12603092B2 patent drawing
  • US12603092B2 patent drawing

AI summary

Implementations set forth herein relate to an automated assistant that can interact with applications that may not have been pre-configured for interfacing with the automated assistant. The automated assistant can identify content of an application interface of the application to determine synonymous terms that a user may speak when commanding the automated assistant to perform certain tasks. Speech processing operations employed by the automated assistant can be biased towards these synonymous terms when the user is accessing an application interface of the application. In some implementations, the synonymous terms can be identified in a responsive language of the automated assistant when the content of the application interface is being rendered in a different language. This can allow the automated assistant to operate as an interface between the user and certain applications that may not be rendering content in a native language of the user.