Communication Fusion App Interprets Non-Verbal Cues

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional chat-based applications fail to effectively interpret spoken user input due to the lack of consideration for non-verbal cues such as intonation, gestures, and emotional tone, leading to inappropriate responses and reduced user experience.

Innovation Solution

A computer-implemented method that generates a predicted context based on non-verbal cues, integrating inputs from speech-to-text, sound-to-cues, personality prediction, and video-to-cues models to enhance the interpretation of spoken user input, allowing chat-based applications to account for emotional states and intentions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If speech-to-text models are used to translate spoken input to text input, then the translation process is automated and efficient, but non-verbal cues such as intonation, gestures, and emotional tone are not taken into account, leading to inaccurate interpretation

Engineering Contradiction:
Improveautomation of speech-to-text translationVSAvoidaccuracy of spoken input interpretation
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent combines multiple models (speech-to-text, sound-to-cues, personality prediction, video-to-cues) into a unified system that processes both verbal and non-verbal cues simultaneously. The communication fusion application merges outputs from these models to generate a comprehensive predicted context that includes emotional states, intentions, and personality traits, thereby resolving the contradiction between automated translation efficiency and interpretative accuracy.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent adds new dimensions of analysis by incorporating sound-to-cues models that detect emotional states and personality traits from audio characteristics, and video-to-cues models that analyze facial expressions and gestures. This dimensional expansion from purely textual translation to multi-modal analysis enables accurate interpretation of non-verbal cues while maintaining automated processing efficiency.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If chat-based applications respond based solely on text input, then the response generation is simple and fast, but the applications cannot properly interpret emotional states or intentions, resulting in inappropriate responses

Engineering Contradiction:
Improvespeed of response generationVSAvoidappropriateness of responses
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary analysis of non-verbal cues through sound-to-cues and video-to-cues models before generating responses. By pre-determining emotional states, intentions, and personality traits from audio and visual inputs, the system prepares comprehensive context information that guides subsequent response generation, ensuring both speed and appropriateness of responses.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The communication fusion application continuously integrates feedback from multiple models (speech-to-text, sound-to-cues, personality prediction, video-to-cues) to dynamically adjust response generation. This multi-source feedback mechanism ensures that responses are both rapid and contextually appropriate, reflecting the user's emotional state and intentions.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If the system incorporates multiple models for analyzing non-verbal cues, then the interpretation accuracy improves, but the system complexity increases

Engineering Contradiction:
Improveaccuracy of non-verbal cue detectionVSAvoidnumber of models and processing components
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex analysis task into separate specialized models: speech-to-text for verbal translation, sound-to-cues for emotional state detection from audio, personality prediction for trait identification, and video-to-cues for facial expression and gesture analysis. Each model handles a specific aspect of non-verbal cue detection, improving overall accuracy while maintaining modular architecture that manages complexity through functional decomposition.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11887600B2Techniques for interpreting spoken input using non-verbal cues
Publication Date: 2024.01.30 DISNEY ENTERPRISES INC
  • US11887600B2 patent drawing
  • US11887600B2 patent drawing
  • US11887600B2 patent drawing

AI summary

In various embodiments, a communication fusion application enables other software application(s) to interpret spoken user input. In operation, a communication fusion application determines that a prediction is relevant to a text input derived from a spoken input received from a user. Subsequently, the communication fusion application generates a predicted context based on the prediction. The communication fusion application then transmits the predicted context and the text input to the other software application(s). The other software application(s) perform additional action(s) based on the text input and the predicted context. Advantageously, by providing additional, relevant information to the software application(s), the communication fusion application increases the level of understanding during interactions with the user and the overall user experience is improved.