Contextual Assistants Using Spatial Cues to Disambiguate Voice Queries
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Digital assistants struggle to provide accurate answers to context-driven queries when the spoken query is ambiguous, lacking sufficient context to identify the object of interest.
Innovation Solution
A computer-implemented method that combines speech recognition with spatial input cues, using a graphical user interface to disambiguate queries by identifying the object on the screen through mouse pointing or touch inputs, enabling accurate information retrieval.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If speech recognition is used to process queries, then the assistant can understand spoken commands, but the query may be ambiguous and lack sufficient context to identify the object of interest
Solution Approach 1:
The patent combines speech recognition with spatial input processing to create a hybrid query interpretation system. The speech transcript is merged with spatial location data from the graphical user interface to disambiguate references and identify the intended object, resolving the contradiction between ease of voice input and loss of contextual information.
Solution Approach 2:
The spatial input indication acts as an intermediary between the ambiguous speech query and the object identification process. It provides additional contextual cues that bridge the gap between the spoken words and the intended target object, enabling accurate query interpretation without requiring fully unambiguous speech.
2Device complexity
If the assistant processes queries without spatial context, then the system remains simple, but the assistant cannot accurately identify which object the user is referring to when multiple objects are displayed
Solution Approach 1:
The query processing system is segmented into distinct functional components: speech recognition module, spatial input processing module, and query interpretation module. Each component handles specific aspects of the query, with the spatial input module separately processing location data from the graphical user interface to identify the referenced object among multiple displayed objects.
Solution Approach 2:
The system adds a spatial dimension to the query processing by incorporating graphical user interface location data. This dimensional addition transforms the one-dimensional speech transcript into a multi-dimensional query that includes both linguistic and spatial information, enabling precise object identification without significantly increasing overall system complexity.
3Reliability
If the assistant requires unambiguous queries to identify objects, then the answer accuracy improves, but the user must provide additional context which increases interaction complexity
Solution Approach 1:
The system performs self-service by automatically extracting spatial context from the graphical user interface without requiring explicit user instructions. The query interpretation module autonomously processes the combination of speech and spatial data to disambiguate object references, maintaining high answer accuracy while keeping the user interaction simple and intuitive.
Data Source
AI summary
A method for a contextual assistant to use mouse pointing or touch cues includes receiving audio data corresponding to a query spoken by a user, receiving, in a graphical user interface displayed on a screen, a user input indication indicating a spatial input applied at a first location on the screen, and processing the audio data to determine a transcription of the query. The method also includes performing query interpretation on the transcription to determine that the query is referring to an object displayed on the screen without uniquely identifying the object, and requesting information about the object. The method further includes disambiguating, using the user input indication indicating the spatial input applied at the first location on the screen, the query to uniquely identify the object that the query is referring to, obtaining the information about the object requested by the query, and providing a response to the query.


