Spoken Item Selection Using Visual Feedback to Cut Interaction Latency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computer-based approaches for selecting a subset of items, such as vehicle features or event tickets, suffer from high latency due to the need for thorough review of options and sequential rendering of audible synthesized responses, which can be impractical for users with limited dexterity or in noisy environments.
Innovation Solution
A system that enables users to select items using only spoken input, with visual output guiding the selection process, minimizing latency by rendering visual outputs simultaneously and leveraging a semantic parser and display-dependent parser to process spoken inputs for quick item selection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If turn-based audible dialogs with synthesized spoken responses are used, then users can provide spoken input, but latency increases due to the time required to render audible responses
Solution Approach 1:
The patent extracts the audible synthesized spoken response rendering from the interaction loop, replacing it with visual outputs only. This removes the time-consuming speech synthesis and audio rendering process while maintaining the core functionality of guiding users through item selection via spoken input and visual feedback.
Solution Approach 2:
The patent substitutes the acoustic field-based communication (synthesized speech) with visual field-based communication (display outputs). This replacement eliminates the sequential nature of audio rendering, allowing parallel processing and immediate visual feedback, thereby reducing latency significantly.
2Ease of operation
If touchscreens with hierarchical navigation are used, then users can select items, but latency increases due to thorough review requirements and inadvertent navigation
Solution Approach 1:
The patent implements immediate visual feedback that updates the display in real-time as users provide spoken input. The system processes spoken commands and instantly reflects the current selection state, allowing users to see their progress and make corrections without navigating through hierarchical levels, thereby reducing latency and preventing inadvertent navigation errors.
Solution Approach 2:
The patent segments the item selection process into discrete, speakable units that are presented visually. Instead of requiring users to navigate through nested hierarchical structures, the system breaks down selections into atomic spoken commands that directly manipulate the selection state, reducing the cognitive load and time required for thorough review.
3Loss of time
If visual outputs are rendered simultaneously, then latency is reduced, but device complexity increases
Solution Approach 1:
The patent implements dynamic rendering where visual outputs are updated in real-time based on the current processing state. The display continuously reflects the latest interpretation of spoken input and current selection state, allowing parallel processing of multiple items while maintaining a coherent, up-to-date visual representation that adapts to user input without requiring complex sequential control.
Data Source
AI summary
Mitigating latency in guiding a user, during an interaction between the user and a computing system, in selecting a subset of item(s), from a superset of candidate items, and causing performance of further action(s) based on the selected subset of item(s). In guiding a user in selecting the subset of items, various implementations enable the user to provide only spoken input(s) in selecting the subset of item(s), and provide visual output(s) that are responsive to the spoken input(s) and that guide the user in selecting the item(s). In some of those various implementations, there is not any (or there is only de minimis) audible spoken synthesized spoken output rendered by the computing system in guiding the user in selecting the subset of item(s).


