Multimodal Classification System for Distributed Resource Transfer

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems lack an efficient method for intelligent multimodal classification in distributed technical environments, which affects the accuracy of resource transfer executions in multimodal communications, as they do not effectively weigh and utilize various communication modalities.

Innovation Solution

A system that retrieves multimodal communications, applies feature extraction algorithms to extract relevant features, generates training datasets, and uses machine learning to classify communications into class labels, determining a multimodal metric for each modality to assess communication efficiency and initiate appropriate actions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional communication systems are used without intelligent classification, then the system complexity is low, but the accuracy of resource transfer execution deteriorates

Engineering Contradiction:
Improveaccuracy of resource transfer executionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the classification process into distinct functional modules: feature extraction module that processes raw communication data, classification module that applies machine learning models, and execution module that acts on classified results. This segmentation allows each module to be optimized independently while maintaining overall system accuracy for resource transfer execution.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary feature extraction and classification before resource transfer execution. By pre-processing communication data to extract relevant features and classify communication effectiveness in advance, the system prepares classification results that guide subsequent resource transfer decisions, thereby improving execution accuracy without adding complex real-time processing requirements.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If multiple communication modalities are processed without feature extraction, then the processing speed is fast, but the classification accuracy deteriorates

Engineering Contradiction:
Improveclassification accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSSpeed

Solution Approach 1:

The patent extracts only the most relevant features from multiple communication modalities (text, audio, video) using dedicated feature extraction algorithms. Instead of processing all raw data, the system identifies and extracts key features such as text sentiment, audio tone, and video engagement metrics. This selective extraction maintains classification accuracy while significantly reducing processing time compared to analyzing complete raw datasets.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies different feature extraction strategies tailored to each communication modality's specific characteristics. Text modalities receive NLP-based feature extraction, audio modalities receive acoustic feature extraction, and video modalities receive visual feature extraction. This localized approach ensures each modality contributes its most relevant features efficiently, balancing accuracy and processing speed.

Inventive Principle:
Principle #3Local quality

3Measurement precision

If all communication modalities are treated equally, then the system operation is simple, but the communication effectiveness assessment deteriorates

Engineering Contradiction:
Improvecommunication effectiveness assessmentVSAvoidsystem operation simplicity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent assigns different weights to different communication modalities based on their relative effectiveness for resource transfer communication. The system dynamically determines weight coefficients for text, audio, and video modalities according to the specific communication context and effectiveness metrics. This weighted approach allows the system to operate automatically without manual configuration while accurately assessing which modalities contribute most to communication effectiveness.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system changes the parameter of modality importance by dynamically adjusting weight coefficients based on communication context, historical effectiveness data, and real-time performance metrics. This parameter adjustment allows the system to adapt to different communication scenarios automatically, improving assessment accuracy without requiring complex manual intervention to reconfigure system operation.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11528248B2System for intelligent multi-modal classification in a distributed technical environment
Publication Date: 2022.12.13 BANK OF AMERICA CORP
  • US11528248B2 patent drawing
  • US11528248B2 patent drawing
  • US11528248B2 patent drawing

AI summary

Systems, computer program products, and methods are described herein for intelligent multimodal classification in a distributed technical environment. The present invention is configured to retrieve one or more multimodal communications from a data repository; initiate one or more feature extraction algorithms on the one or more communication modalities to extract one or more features; generate a training dataset based on at least the one or more features extracted from the one or more communication modalities; initiate one or more machine learning algorithms on the training dataset to generate a first set of parameters; receive an unseen multimodal communication; generate an unseen dataset based on at least the unseen multimodal communication; classify, using the first set of parameters, the unseen multimodal communication into one or more class labels; and initiate an execution of one or more actions on the unseen multimodal communication based on at least the classification.