Multi-task Attention Text Encoder for NLP Training Efficiency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing natural language processing solutions face inefficiencies in training and resource utilization due to the need for large datasets and extensive computational resources, particularly in generating semantically rich inputs for downstream models.

Innovation Solution

An attention-based text encoder machine learning model is trained using a multi-task training routine that combines language modeling, similarity determination, and document classification tasks to generate word-wise and document-wide embedded representations, optimizing parameter values through sequential and concurrent learning loss models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional single-task training is used for NLP models, then the model can be trained with simpler training procedures, but the model requires large datasets and extensive computational resources to achieve satisfactory performance

Engineering Contradiction:
Improvetraining efficiencyVSAvoidtraining samples required
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent combines multiple training tasks (language modeling, similarity determination, and document classification) into a single multi-task training framework. This merging allows the model to learn multiple objectives simultaneously, improving training efficiency and reducing the amount of data needed compared to traditional single-task training approaches.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The trained attention-based text encoder model is designed to perform multiple functions: generating word-wise embeddings, generating document-wide embeddings, and supporting various downstream NLP tasks. This multi-functionality allows a single model to replace what would traditionally require multiple specialized models, reducing overall computational resources and data requirements.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Loss of energy

If traditional training approaches are used, then the training process is simpler, but storage and computational resources are consumed at higher rates

Engineering Contradiction:
Improvecomputational resource consumptionVSAvoidtraining routine complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The patent merges multiple training objectives into a unified loss function that combines language modeling loss, similarity determination loss, and document classification loss. This consolidation allows the model to learn multiple tasks simultaneously, reducing the need for separate training processes and thereby reducing overall computational resource consumption.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The model performs preliminary learning of language structures and semantic relationships through the language modeling and similarity determination tasks before being applied to specific downstream tasks. This preliminary action pre-trains the model with general linguistic knowledge, reducing the computational resources needed for subsequent task-specific training.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If multi-task training is implemented, then training efficiency improves and fewer samples are needed, but the training routine becomes more complex

Engineering Contradiction:
Improvetraining speedVSAvoidtraining routine complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent merges multiple training tasks into a unified framework with a single model architecture and a combined loss function. This merging enables simultaneous optimization of multiple objectives, improving training speed and sample efficiency while managing complexity through a cohesive training routine rather than multiple separate training processes.

Inventive Principle:
Principle #5Merging (Combining)

4Reliability

If conventional embedding methods are used, then the implementation is simpler, but the semantic richness of the generated inputs is insufficient

Engineering Contradiction:
Improvesemantic qualityVSAvoidmodel architecture complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements local quality by generating both word-wise embeddings (local level) and document-wide embeddings (global level) using the same attention-based model. This dual-level embedding approach ensures that both fine-grained word-level semantics and coarse-grained document-level semantics are captured, improving the overall semantic quality of the generated representations.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12112132B2Natural language processing machine learning frameworks trained using multi-task training routines
Publication Date: 2024.10.08 OPTUM SERVICES IRELAND LTD
  • US12112132B2 patent drawing
  • US12112132B2 patent drawing
  • US12112132B2 patent drawing

AI summary

Various embodiments of the present invention provide methods, apparatus, systems, computing devices, computing entities, and/or the like for performing natural language processing operations using an attention-based text encoder machine learning model that is trained using a multi-task training routine that is associated with two or more training tasks (e.g., a multi-task training routine that is associated with two or more sequential training tasks, a multi-training routine that is associated with two or more concurrent training tasks, and/or the like).