Non-Linguistic Voice Input for Faster VUI Actions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice user interfaces (VUIs) face challenges in providing quick responses and accurately conveying user intent and emotion, particularly in time-critical scenarios and when using linguistic inputs alone is insufficient.

Innovation Solution

A voice user interface system that processes both linguistic and non-linguistic inputs, such as paralinguistic and prosodic inputs, to enable faster device actions and emotion recognition, utilizing microphones, sensors, and neural networks for real-time processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If natural language processing (NLP) systems are used for voice input, then the system can process linguistic information, but the response time is delayed and cannot achieve real-time performance

Engineering Contradiction:
Improveresponse timeVSAvoidlinguistic processing capability
Core Design Contradiction:
Loss of timeVSLoss of information

Solution Approach 1:

The patent segments voice input processing into two independent pathways: linguistic processing (for semantic meaning) and non-linguistic processing (for tone, pitch, and emotional content). This segmentation allows the non-linguistic pathway to operate independently and trigger actions without waiting for the slower NLP processing, thereby reducing response time while preserving linguistic processing capability for complex tasks.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary analysis of non-linguistic voice characteristics (tone, pitch, volume) in real-time as the voice signal is received. This preliminary action enables the system to detect emotional states and intent markers immediately, allowing it to prepare for or even execute time-critical actions before the full NLP processing completes, thus reducing overall response latency.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If only linguistic inputs are processed, then the system structure remains simple, but the system cannot recognize user emotion and intent accurately

Engineering Contradiction:
Improveemotion recognition accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the voice processing system into separate linguistic and non-linguistic processing modules. The non-linguistic module specifically analyzes tone, pitch, volume, and speech patterns to detect emotional states and intent markers. This segmentation enables accurate emotion recognition by dedicating specific processing resources to acoustic features without requiring complete system redesign.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system implements a multi-functional processing architecture where the non-linguistic analysis module serves multiple purposes: detecting emotion, determining user intent, and triggering actions. This universal module handles various emotional states (joy, sadness, anger, frustration) and intent types (humor, sarcasm, urgency) through a single integrated processing pathway, reducing overall system complexity while improving measurement precision.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Speed

If the system waits for complete voice command processing before acting, then processing accuracy is maintained, but time-critical actions cannot be executed

Engineering Contradiction:
Improveaction execution speedVSAvoidprocessing accuracy
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The system performs preliminary analysis of non-linguistic voice characteristics to detect intent markers and emotional states that indicate time-critical situations. When such markers are detected (e.g., urgent tone, frustrated speech patterns), the system can immediately execute appropriate actions without waiting for complete NLP processing, while still maintaining processing accuracy through subsequent verification of the detected intent.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary non-linguistic analysis module that acts as a mediator between the voice input and the action execution system. This intermediary rapidly analyzes acoustic features and emotional content, and when time-critical conditions are detected, it can directly trigger actions or prioritize processing, thereby increasing action execution speed while maintaining reliability through the intermediary's verification of intent markers.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12417766B2Voice user interface using non-linguistic input
Publication Date: 2025.09.16 MAGIC LEAP INC
  • US12417766B2 patent drawing
  • US12417766B2 patent drawing
  • US12417766B2 patent drawing

AI summary

A voice user interface (VUI) and methods for operating the VUI are disclosed. In some embodiments, the VUI configured to receive and process linguistic and non-linguistic inputs. For example, the VUI receives an audio signal, and the VUI determines whether the audio input comprises a linguistic and/or a non-linguistic input. In accordance with a determination that the audio signal comprises a non-linguistic input, the VUI causes a system to perform an action associated with the non-linguistic input.