Voice Input Routing via Contextual Disambiguation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic devices face difficulties in determining the appropriate target application for voice inputs when multiple voice-enabled tasks coincide, requiring users to manually switch between applications, as current systems lack the ability to coordinate voice inputs in parallel.
Innovation Solution
The system employs multi-modal inputs such as eye tracking, gesture recognition, situational, and contextual data to select and route voice inputs to the appropriate active voice-enabled application, allowing multiple applications to operate in parallel without forcing users to explicitly switch between them.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If multiple voice enabled applications operate in parallel, then the system can handle multiple voice inputs simultaneously, but it becomes difficult to determine the appropriate target application for each voice input
Solution Approach 1:
The patent introduces an intermediary system that sits between the audio receiver and multiple voice-enabled applications. This intermediary analyzes contextual information, application states, and voice input characteristics to determine which application should receive each voice input, thereby resolving the ambiguity of target identification when multiple applications are active simultaneously.
Solution Approach 2:
The system dynamically changes parameters such as application focus state, voice input routing priorities, and contextual relevance weights based on current system conditions. By adjusting these parameters in real-time, the system can accurately route voice inputs to the appropriate application even when multiple applications are running in parallel.
2Measurement precision
If users manually switch between voice enabled applications, then the system can determine the correct target application, but this increases user effort and reduces operational efficiency
Solution Approach 1:
The system implements self-service by automatically determining the correct target application without requiring user intervention. The intermediary system autonomously analyzes the context and routes voice inputs to the appropriate application, eliminating the need for users to manually switch between applications while maintaining high accuracy in target identification.
Solution Approach 2:
The system uses feedback mechanisms to continuously monitor application states and user interactions. By analyzing this feedback data, the intermediary system learns user preferences and contextual patterns, improving its ability to automatically identify the correct target application and reducing the need for manual switching over time.
3Device complexity
If the system processes voice inputs serially in currently active applications, then target application determination is straightforward, but this reduces productivity when multiple voice enabled tasks coincide
Solution Approach 1:
The patent segments the voice input processing system into distinct components: an audio receiver, an intermediary routing system, and multiple independent voice-enabled applications. This segmentation allows each component to operate independently and simultaneously, enabling parallel processing of multiple voice inputs while maintaining clear boundaries and coordination mechanisms between components.
Solution Approach 2:
The intermediary system serves multiple functions: it receives voice inputs from the audio receiver, analyzes contextual information, determines the appropriate target application, and routes the voice input accordingly. This multi-functional intermediary enables the system to handle multiple voice-enabled tasks in parallel while maintaining straightforward processing coordination through a centralized routing mechanism.
Data Source
AI summary
One embodiment provides a method, including: receiving, at an audio receiver of a device, a voice input; selecting, using a processor of a device, an active target voice enabled resource for the voice input from among a plurality of active target voice enabled resources; and providing, using a processor of the device, the voice input to the active target voice enabled resource selected. Other aspects are described and claimed.


