Shared Machine Learning Model for Text Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Training multiple machine learning models for natural language processing tasks, especially those in niche applications, often requires large labeled datasets, leading to suboptimal performance due to limited training samples and differing industry-specific expressions and intents.

Innovation Solution

A shared machine learning model is trained using diverse training data from multiple machine learning models, performing text embedding tasks and adjusting weights to generate representations that enable accurate intent determination across different classification tasks, thereby optimizing performance for various industry-specific applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple machine learning models are trained separately for different text classification tasks, then each model can be optimized for its specific task, but the models require large labeled datasets for each niche application, leading to suboptimal performance due to limited training samples

Engineering Contradiction:
Improveclassification accuracyVSAvoidtraining data volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent combines multiple task-specific machine learning models into a single shared model that performs multiple text classification tasks simultaneously. The shared model learns common representations from diverse training data across different domains (e.g., healthcare, finance, retail), allowing each task to benefit from the collective information in the combined dataset rather than requiring separate large datasets for each niche application.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The shared machine learning model is designed to be universal and multi-functional, capable of performing multiple text classification tasks across different industries and domains. Instead of training separate specialized models for each task, the shared model learns domain-agnostic features and representations that can be applied universally across various text classification problems, reducing the need for task-specific large labeled datasets.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If multiple machine learning models are trained for different industries, then each model can capture industry-specific expressions, but the models may overfit to their specific datasets and fail to generalize across different industries

Engineering Contradiction:
Improveindustry-specific performanceVSAvoidgeneralization capability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent merges training data from multiple industries into a unified shared model, allowing the model to learn both industry-specific patterns and common linguistic features. By exposing the shared model to diverse expressions and contexts across healthcare, finance, retail, and other domains simultaneously, it learns to generalize better rather than overfitting to any single industry's data characteristics.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The shared model incorporates task-specific adaptation layers or fine-tuning mechanisms that allow it to capture local industry-specific expressions while maintaining a shared representation layer that learns universal patterns. This enables the model to adapt to industry-specific nuances without sacrificing its ability to generalize across different domains.

Inventive Principle:
Principle #3Local quality

3Adaptability or versatility

If separate machine learning models are trained for each text classification task, then each model can be specialized, but the overall system complexity increases and training resources are duplicated

Engineering Contradiction:
Improvetask specializationVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges multiple separate machine learning models into a single shared model architecture that handles multiple text classification tasks. Instead of maintaining separate models for healthcare text classification, finance text classification, retail text classification, etc., the shared model unifies these tasks under one framework, reducing system complexity while maintaining task-specific performance through shared representation learning.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The shared machine learning model is designed as a universal framework that can perform multiple text classification tasks simultaneously. It uses a common architecture and shared parameters that are trained on multi-task data, eliminating the need for multiple separate model deployments and reducing overall system complexity while maintaining adaptability to different tasks.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20230169362A1Shared network learning for machine learning enabled text classification
Publication Date: 2023.06.01 SAP FRANCE
  • US20230169362A1 patent drawing
  • US20230169362A1 patent drawing
  • US20230169362A1 patent drawing

AI summary

A method may include training a first machine learning model to perform a question generation task and a second machine learning model to perform a question answering task. The first machine learning model and the second machine learning model may be subj ected to a collaborative training in which a first plurality of weights applied by the first machine learning model generating one or more questions are adjusted to minimize an error in an output of the second machine learning model answering the one or more questions. The first machine learning model and the second machine learning model may be deployed to perform a natural language processing task that requires the first machine learning model to generate a question and/or the second machine learning model to answer a question. Related methods and articles of manufacture are also disclosed.