Multimodal Intention Prediction via Segmented Predictors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multimodal interface methods face challenges in efficiently and accurately deducing user intentions due to high complexity in signal-level integration and difficulty in perceiving associations between modalities when analyzing meanings individually.
Innovation Solution
A user intention deduction apparatus and method that utilizes a first predictor to analyze motion information and a second predictor to interpret multimodal information, generating control signals for operations, allowing for efficient prediction and execution of user intentions across various modalities, such as voice and gesture inputs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multimodal inputs are combined at signal level, then the system can process simultaneously generated signals, but the characteristic space becomes very large and the model complexity increases
Solution Approach 1:
The patent segments the intention deduction process into two independent predictors: a first predictor that processes motion information and a second predictor that processes multimodal sensor information. This segmentation divides the complex characteristic space into smaller, more manageable subspaces, reducing the overall model complexity while maintaining processing accuracy for simultaneously generated signals.
2Reliability
If multimodal inputs are combined at signal level, then simultaneous signals can be processed, but the amount of learning required for the model becomes high
Solution Approach 1:
By segmenting the learning task into two separate predictors with specialized functions, the patent reduces the learning burden on each individual model. The first predictor learns motion patterns while the second predictor learns multimodal sensor patterns, significantly reducing the total learning time compared to training one large model to handle all signals simultaneously.
3Adaptability or versatility
If modality inputs are analyzed individually at meaning level, then learning and expansion become easier, but associations between modalities become difficult to find
Solution Approach 1:
The patent merges the outputs of the first predictor (motion intention) with multimodal sensor information in the second predictor. This merging occurs at the intention level rather than the raw signal level, allowing the system to maintain modality independence for easier learning while still capturing associations between different modalities through their combined contribution to the final intention prediction.
4Ease of operation
If modality inputs are analyzed individually at meaning level, then independencies between modalities are maintained, but the associations between modalities that users exploit become difficult to perceive
Solution Approach 1:
The patent segments the analysis into two stages: first analyzing motion information independently to predict partial intention, then analyzing multimodal sensor information independently while incorporating the motion-based prediction. This segmented approach maintains the ease of learning individual modalities while the sequential combination recovers the associations between modalities that users naturally exploit.
Data Source
AI summary
Disclosed are an apparatus and method of deducing a user's intention using multimodal information. The user's intention deduction apparatus includes a first predictor to predict a part of a user's intention using at least one piece of motion information, and a second predictor to predict the user's intention using the predicted part of the user's intention and multimodal information received from at least one multimodal sensor.


