Contextual Assistants Using Spatial Cues to Disambiguate Voice Queries

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Digital assistants struggle to provide accurate answers to context-driven queries when the spoken query is ambiguous, lacking sufficient context to identify the object of interest.

Innovation Solution

A computer-implemented method that combines speech recognition with spatial input cues, using a graphical user interface to disambiguate queries by identifying the object on the screen through mouse pointing or touch inputs, enabling accurate information retrieval.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speech recognition is used to process queries, then the assistant can understand spoken commands, but the query may be ambiguous and lack sufficient context to identify the object of interest

Engineering Contradiction:
Improvevoice-based query inputVSAvoidcontext information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent combines speech recognition with spatial input processing to create a hybrid query interpretation system. The speech transcript is merged with spatial location data from the graphical user interface to disambiguate references and identify the intended object, resolving the contradiction between ease of voice input and loss of contextual information.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The spatial input indication acts as an intermediary between the ambiguous speech query and the object identification process. It provides additional contextual cues that bridge the gap between the spoken words and the intended target object, enabling accurate query interpretation without requiring fully unambiguous speech.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If the assistant processes queries without spatial context, then the system remains simple, but the assistant cannot accurately identify which object the user is referring to when multiple objects are displayed

Engineering Contradiction:
Improvequery processing systemVSAvoidobject identification accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The query processing system is segmented into distinct functional components: speech recognition module, spatial input processing module, and query interpretation module. Each component handles specific aspects of the query, with the spatial input module separately processing location data from the graphical user interface to identify the referenced object among multiple displayed objects.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system adds a spatial dimension to the query processing by incorporating graphical user interface location data. This dimensional addition transforms the one-dimensional speech transcript into a multi-dimensional query that includes both linguistic and spatial information, enabling precise object identification without significantly increasing overall system complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Reliability

If the assistant requires unambiguous queries to identify objects, then the answer accuracy improves, but the user must provide additional context which increases interaction complexity

Engineering Contradiction:
Improvequery answer accuracyVSAvoiduser interaction simplicity
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system performs self-service by automatically extracting spatial context from the graphical user interface without requiring explicit user instructions. The query interpretation module autonomously processes the combination of speech and spatial data to disambiguate object references, maintaining high answer accuracy while keeping the user interaction simple and intuitive.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250306853A1Contextual Assistant Using Mouse Pointing or Touch Cues
Publication Date: 2025.10.02 GOOGLE LLC
  • US20250306853A1 patent drawing
  • US20250306853A1 patent drawing
  • US20250306853A1 patent drawing

AI summary

A method for a contextual assistant to use mouse pointing or touch cues includes receiving audio data corresponding to a query spoken by a user, receiving, in a graphical user interface displayed on a screen, a user input indication indicating a spatial input applied at a first location on the screen, and processing the audio data to determine a transcription of the query. The method also includes performing query interpretation on the transcription to determine that the query is referring to an object displayed on the screen without uniquely identifying the object, and requesting information about the object. The method further includes disambiguating, using the user input indication indicating the spatial input applied at the first location on the screen, the query to uniquely identify the object that the query is referring to, obtaining the information about the object requested by the query, and providing a response to the query.