Multi-Modal Machine Learning Model for Entity Interpretation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning models struggle to interpret natural language text and images effectively, particularly in multi-modal datasets, leading to incomplete and non-interpretable outputs.

Innovation Solution

A fully connected multi-modal machine learning model with text- and image-based layers that generate, pass, augment, and reinterpret intermediate representations to create holistic outputs, combining contextual insights from both textual and image-based information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple individually trained models are used to process each data type in multi-modal datasets, then each model can be optimized for its specific data type, but the system complexity increases and the outputs remain individually derived without holistic integration

Engineering Contradiction:
Improveinterpretation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple individually trained models into a unified multi-modal machine learning model that processes text, images, and other data types through integrated layers. The model merges text-based layers, image-based layers, and audio-based layers into a single architecture that generates holistic outputs rather than separate individual outputs, resolving the contradiction between optimization precision and system complexity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The multi-modal machine learning model is designed with universal functionality to handle multiple data types (text, images, audio) within a single model architecture. The model can process different modalities through specialized layers while maintaining a unified structure, allowing one model to perform multiple functions that previously required separate models.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If deep learning models are used to capture broad representations of textual features, then complex relationships between concepts can be discerned, but the outputs become difficult to interpret by human experts

Engineering Contradiction:
Improveconcept relationship detectionVSAvoidoutput interpretability
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent segments the deep learning model into distinct interpretable layers including text-based layers, image-based layers, and audio-based layers, each with specific functions. By segmenting the model architecture and output components, the system maintains the ability to detect complex relationships while providing human-interpretable outputs through structured layer representations and separate processing paths for different modalities.

Inventive Principle:
Principle #1Segmentation

3Productivity

If algorithmic NLP techniques are used to identify and distill targeted concepts from text, then human-interpretable outputs can be generated quickly, but the system struggles with complex concepts that span multiple documents

Engineering Contradiction:
Improveprocessing speedVSAvoidcomplex concept handling
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent merges algorithmic NLP processing with deep learning capabilities in a unified multi-modal model. The combination allows the system to maintain the speed and interpretability of algorithmic approaches while incorporating the complex relationship detection abilities of deep learning, enabling handling of complex concepts that span multiple documents through integrated text, image, and audio processing layers.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250131196A1Deep model integration techniques for machine learning entity interpretation
Publication Date: 2025.04.24 OPTUM INC
  • US20250131196A1 patent drawing
  • US20250131196A1 patent drawing
  • US20250131196A1 patent drawing

AI summary

Various embodiments of the present disclosure provide machine learning training techniques for implementing a multi-modal interpretation process to generate holistic outputs for an event. The techniques may include generating, using first layers of a multi-modal machine learning model, text-based intermediate representations for an entity based on textual input data. The techniques include generating, using second layers of the multi-modal machine learning model, image-based intermediate representations for the entity based on the text-based intermediate representations and input images for the entity. The techniques include generating, using one or more third layers of the multi-modal machine learning model, an entity representation summary based on the image-based intermediate representations and an image narrative summary for the input images. The techniques include initiating the performance of a prediction-based action based on the entity representation summary.