Multimodal Unsupervised Meta-Learning Encoder for Cross-Modal Task Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current unsupervised meta-learning technologies are limited to single-modal data, making it difficult to implement artificial intelligence that can handle multimodal data like humans do, which is essential for flexible thinking and accurate task performance in real-world applications.

Innovation Solution

A multimodal unsupervised meta-learning apparatus that extracts conceptual features from a massive multimodal dataset without labels, generates a source task similar to a target task, and derives a learning method to improve learning efficiency and performance, using a combination of neural networks and feature extraction techniques to train models with a small number of target datasets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If unsupervised meta-learning is applied to single-modal data, then learning efficiency is improved, but the ability to handle multimodal data like humans is lost

Engineering Contradiction:
Improvelearning efficiencyVSAvoidmultimodal data handling capability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent applies universality by designing a unified multimodal unsupervised meta-learning framework that can handle multiple data modalities (image, audio, text) simultaneously. The system uses a single encoder that processes different modalities through their respective embedding layers, enabling one model to perform multiple tasks across different data types without requiring separate single-modal models for each modality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent merges multiple single-modal processing streams into a unified multimodal framework. Different modalities are processed through separate embedding layers (image embedding layer, audio embedding layer, text embedding layer) and then combined through a shared encoder structure, allowing the system to integrate information from multiple sources while maintaining the efficiency of unsupervised learning.

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If massive labeled data is used for training, then model accuracy is improved, but data acquisition cost and time increase significantly

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata acquisition time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts labels from the data through unsupervised learning techniques. Instead of requiring pre-labeled data, the system automatically discovers patterns and structures in the data by training the encoder to reconstruct or classify data representations, thereby extracting meaningful labels and concepts without human annotation.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs self-labeling through unsupervised meta-learning. The encoder learns to represent data in a way that automatically organizes it into meaningful categories and concepts, allowing the model to serve its own labeling needs without external intervention or pre-labeled data.

Inventive Principle:
Principle #25Self-service

3Loss of information

If multimodal data is processed using traditional deep learning, then comprehensive information capture is improved, but computational resource requirements increase

Engineering Contradiction:
Improveinformation capture completenessVSAvoidcomputational resource consumption
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The patent segments the processing of different modalities into separate embedding layers, allowing each modality to be processed independently through its own specialized layer. This segmentation enables efficient processing by avoiding the computational burden of processing all modalities simultaneously through a single monolithic model, while still capturing comprehensive information through the integrated encoder.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240232648A1Multimodal unsupervised meta-learning method and apparatus
Publication Date: 2024.07.11 ELECTRONICS & TELECOMM RES INST
  • US20240232648A1 patent drawing
  • US20240232648A1 patent drawing
  • US20240232648A1 patent drawing

AI summary

Disclosed herein are a multimodal unsupervised meta-learning method and apparatus. The multimodal unsupervised meta-learning method includes training, by a multimodal unsupervised feature representation learning unit, an encoder configured to extract features of individual single-modal signals from a source multimodal dataset, generating, by a multimodal unsupervised task generation unit, a source task based on the features of individual single-modal signals, deriving, by a multimodal unsupervised learning method derivation unit, a learning method from the source task using the encoder, and training, by a target task performance unit, a model based on the learning method and features extracted from a small number of target datasets by the encoder, thus performing the target task.