Contextual Assistant Spatial Cues for Ambiguous Screen Queries

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Digital assistants struggle to provide accurate answers to ambiguous queries, as they require additional context to disambiguate the query and identify the relevant object on a screen.

Innovation Solution

A computer-implemented method that receives audio data corresponding to a query spoken by a user and captures spatial input indications on a screen, using a speech recognition model to transcribe the query and perform query interpretation to identify the object on the screen, and disambiguates the query using spatial input to uniquely identify the object.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a digital assistant processes only audio data for query interpretation, then the system complexity remains low, but the accuracy of answering ambiguous queries deteriorates

Engineering Contradiction:
Improvequery interpretation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple input modalities (audio data, spatial input data, and screen context data) into a unified query interpretation process. The system merges these different data types to disambiguate queries and identify objects on the screen, thereby improving interpretation accuracy while managing system complexity through integrated processing.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces spatial input data as an intermediary element that bridges audio queries and visual context on the screen. This spatial input acts as a mediator that helps the system connect the user's spoken query with the relevant object location, improving query disambiguation without requiring direct complex interaction between audio and visual systems.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the digital assistant requests additional context for ambiguous queries, then the answer accuracy improves, but the response time increases

Engineering Contradiction:
Improveanswer accuracyVSAvoidresponse time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent captures spatial input data and screen context information in advance, before the query is fully processed. By having this contextual information readily available when an ambiguous query is detected, the system can immediately disambiguate the query without requiring additional back-and-forth interactions, thus maintaining fast response times while improving answer accuracy.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If the system captures spatial input data from the screen, then the ability to disambiguate queries improves, but the ease of operation deteriorates

Engineering Contradiction:
Improvequery disambiguation accuracyVSAvoiduser interaction simplicity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent makes the screen itself serve as the input interface by capturing spatial coordinates directly from touch or cursor movements. The system automatically processes these spatial inputs without requiring users to learn new interaction patterns or perform additional actions, thereby maintaining ease of operation while improving query disambiguation accuracy.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12346634B2Contextual assistant using mouse pointing or touch cues
Publication Date: 2025.07.01 GOOGLE LLC
  • US12346634B2 patent drawing
  • US12346634B2 patent drawing
  • US12346634B2 patent drawing

AI summary

A method for a contextual assistant to use mouse pointing or touch cues includes receiving audio data corresponding to a query spoken by a user, receiving, in a graphical user interface displayed on a screen, a user input indication indicating a spatial input applied at a first location on the screen, and processing the audio data to determine a transcription of the query. The method also includes performing query interpretation on the transcription to determine that the query is referring to an object displayed on the screen without uniquely identifying the object, and requesting information about the object. The method further includes disambiguating, using the user input indication indicating the spatial input applied at the first location on the screen, the query to uniquely identify the object that the query is referring to, obtaining the information about the object requested by the query, and providing a response to the query.