Non-Linguistic Voice Input for Faster VUI Actions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice user interfaces (VUIs) face challenges in providing quick responses and accurately conveying user intent and emotion, particularly in time-critical scenarios and when using linguistic inputs alone is insufficient.
Innovation Solution
A voice user interface system that processes both linguistic and non-linguistic inputs, such as paralinguistic and prosodic inputs, to enable faster device actions and emotion recognition, utilizing microphones, sensors, and neural networks for real-time processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If natural language processing (NLP) systems are used for voice input, then the system can process linguistic information, but the response time is delayed and cannot achieve real-time performance
Solution Approach 1:
The patent segments voice input processing into two independent pathways: linguistic processing (for semantic meaning) and non-linguistic processing (for tone, pitch, and emotional content). This segmentation allows the non-linguistic pathway to operate independently and trigger actions without waiting for the slower NLP processing, thereby reducing response time while preserving linguistic processing capability for complex tasks.
Solution Approach 2:
The system performs preliminary analysis of non-linguistic voice characteristics (tone, pitch, volume) in real-time as the voice signal is received. This preliminary action enables the system to detect emotional states and intent markers immediately, allowing it to prepare for or even execute time-critical actions before the full NLP processing completes, thus reducing overall response latency.
2Measurement precision
If only linguistic inputs are processed, then the system structure remains simple, but the system cannot recognize user emotion and intent accurately
Solution Approach 1:
The patent divides the voice processing system into separate linguistic and non-linguistic processing modules. The non-linguistic module specifically analyzes tone, pitch, volume, and speech patterns to detect emotional states and intent markers. This segmentation enables accurate emotion recognition by dedicating specific processing resources to acoustic features without requiring complete system redesign.
Solution Approach 2:
The system implements a multi-functional processing architecture where the non-linguistic analysis module serves multiple purposes: detecting emotion, determining user intent, and triggering actions. This universal module handles various emotional states (joy, sadness, anger, frustration) and intent types (humor, sarcasm, urgency) through a single integrated processing pathway, reducing overall system complexity while improving measurement precision.
3Speed
If the system waits for complete voice command processing before acting, then processing accuracy is maintained, but time-critical actions cannot be executed
Solution Approach 1:
The system performs preliminary analysis of non-linguistic voice characteristics to detect intent markers and emotional states that indicate time-critical situations. When such markers are detected (e.g., urgent tone, frustrated speech patterns), the system can immediately execute appropriate actions without waiting for complete NLP processing, while still maintaining processing accuracy through subsequent verification of the detected intent.
Solution Approach 2:
The patent introduces an intermediary non-linguistic analysis module that acts as a mediator between the voice input and the action execution system. This intermediary rapidly analyzes acoustic features and emotional content, and when time-critical conditions are detected, it can directly trigger actions or prioritize processing, thereby increasing action execution speed while maintaining reliability through the intermediary's verification of intent markers.
Data Source
AI summary
A voice user interface (VUI) and methods for operating the VUI are disclosed. In some embodiments, the VUI configured to receive and process linguistic and non-linguistic inputs. For example, the VUI receives an audio signal, and the VUI determines whether the audio input comprises a linguistic and/or a non-linguistic input. In accordance with a determination that the audio signal comprises a non-linguistic input, the VUI causes a system to perform an action associated with the non-linguistic input.


