Adaptive Inference System for Multi-Modal User Intention

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multi-modal inference systems face challenges in accurately inferring user intentions due to the complexity of integrating and analyzing various modalities, such as visual, voice, and text information, and lack of personalization based on user history and context.

Innovation Solution

An adaptive inference system that collects multi-modal information including visual, voice, and text data, uses recognition techniques like object, face, and emotion recognition, and incorporates user history and personal information to infer intentions, enabling more accurate and personalized inferences by integrating these insights.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multi-modal information (visual, voice, text) is collected and integrated for user intention inference, then inference accuracy is improved, but system complexity increases

Engineering Contradiction:
Improveinference accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the complex multi-modal inference process into distinct modules: a collection unit that gathers multi-modal information (visual, voice, text), a storage unit that maintains user history and context, and an inference unit that processes the data. This segmentation allows each module to handle specific tasks independently, improving overall inference accuracy while managing system complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If user history information and personal information are integrated into the inference process, then personalization accuracy is improved, but information processing complexity increases

Engineering Contradiction:
Improvepersonalization accuracyVSAvoidinformation processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary action by pre-collecting and storing user history information and personal information in a dedicated storage unit before the actual inference process. This allows the inference unit to access pre-processed user context data without performing complex real-time analysis, thereby improving personalization accuracy while reducing the computational burden during active inference operations.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If multiple recognition techniques (object recognition, face recognition, emotion recognition, voice recognition) are applied, then recognition precision is improved, but processing time increases

Engineering Contradiction:
Improverecognition precisionVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system merges multiple recognition techniques (object recognition, face recognition, emotion recognition, voice recognition) into a unified inference framework. The collection unit simultaneously captures multi-modal information and the inference unit processes these diverse data types together, leveraging complementary information from different modalities to improve recognition precision while optimizing processing efficiency through integrated analysis.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11455837B2Adaptive inference system and operation method therefor
Publication Date: 2022.09.27 KOREA ELECTRONICS TECH INST
  • US11455837B2 patent drawing
  • US11455837B2 patent drawing
  • US11455837B2 patent drawing

AI summary

This application relates to an adaptive inference system and an operation method therefor. In one aspect, the system includes a user terminal for collecting multi-modal information including at least visual information, voice information and text information. The system may also include an inference support device for receiving the multi-modal information from the user terminal, and inferring the intention of a user on the basis of pre-stored history information related to the user terminal, individualized information and the multi-modal information.