Multimodal Conversation Interpretation With Confidence Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional technologies face challenges in accurately interpreting the meaning and context of words based on voice and facial expression data, leading to unreliable interpretation results.

Innovation Solution

A system comprising a voice collection unit, facial expression collection unit, analysis unit, confidence evaluation unit, and presentation unit, which uses AI to analyze voice and facial expression data to generate interpretation candidates with confidence scores and provide guidance via audio, enabling accurate understanding of conversation partners.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional voice and facial expression analysis methods are used, then the system is simple, but the interpretation accuracy and reliability are insufficient

Engineering Contradiction:
Improveinterpretation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the interpretation process into distinct functional modules: voice collection unit, facial expression collection unit, analysis unit, confidence evaluation unit, and presentation unit. Each module handles a specific aspect of the interpretation task, allowing for specialized processing while maintaining overall system organization and manageability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The analysis unit serves as an intermediary between the data collection units and the presentation unit. It processes raw voice and facial expression data, generates interpretation candidates, and passes them to the confidence evaluation unit, which acts as another intermediary to filter and rank results before final presentation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If multiple data sources (voice and facial expression) are analyzed, then the reliability of interpretation results improves, but the device complexity increases

Engineering Contradiction:
Improveinterpretation reliabilityVSAvoiddata processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system merges voice data and facial expression data into a unified analysis framework. The analysis unit simultaneously processes both data types and integrates their information to generate comprehensive interpretation candidates, leveraging the complementary nature of auditory and visual cues for more reliable results.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The analysis unit is designed with multi-functionality to handle both voice and facial expression data using the same underlying AI model architecture. This universal processing approach allows the system to accommodate multiple data sources without requiring entirely separate processing pipelines for each modality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If AI-based analysis is implemented, then the interpretation quality improves, but the computational resources and processing time increase

Engineering Contradiction:
Improveinterpretation qualityVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system generates multiple interpretation candidates beyond what would be minimally required, allowing the confidence evaluation unit to select the most reliable results. This approach of producing excess candidates ensures high-quality interpretations while maintaining reasonable processing times by filtering out lower-confidence results.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The confidence evaluation unit provides feedback on the quality of interpretation candidates generated by the analysis unit. This feedback mechanism allows the system to adjust its processing focus, prioritizing high-confidence interpretations and reducing computational effort on low-confidence cases, thereby optimizing the balance between quality and processing time.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20260072640A1system
Publication Date: 2026.03.12 SOFTBANK GROUP CORP
  • US20260072640A1 patent drawing
  • US20260072640A1 patent drawing
  • US20260072640A1 patent drawing

AI summary

The system according to the embodiment comprises a voice collection unit, a facial expression collection unit, an analysis unit, an interpretation unit, a confidence evaluation unit, a presentation unit, and a voice guidance unit. The voice collection unit collects voice data. The facial expression collection unit collects facial expression data. The analysis unit analyzes data collected by the voice collection unit and the facial expression collection unit. The interpretation unit interprets the meaning and context of words based on the data analyzed by the analysis unit. The confidence evaluation unit evaluates the confidence of the interpretation results obtained by the interpretation unit. The presentation unit presents the interpretation results evaluated by the confidence evaluation unit to a smartphone. The voice guidance unit explains the interpretation results evaluated by the confidence evaluation unit via audio through earphones.