Type-Specific MHC Class II Augmentation for Rare-Class Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing artificial intelligence-based predictive models for MHC class II binding and immunogenicity require improved data augmentation methods to enhance prediction confidence and efficiency, particularly in handling data imbalance and rare class detection.

Innovation Solution

A data augmentation method that selectively augments MHC class II binding and immunogenicity predictive models by categorizing data into first and second types, applying type-specific augmentation and labeling strategies, and optimizing peptide sequences to improve data quality and reduce computational resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data augmentation is performed using conventional methods, then the quantity of training data increases, but the quality and relevance of augmented data deteriorates due to data imbalance and rare class detection issues

Engineering Contradiction:
Improvequantity of training dataVSAvoidquality of augmented data
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent segments the training data into multiple types (e.g., first-type data and second-type data) based on specific characteristics such as peptide length and IC50 values. This segmentation allows different augmentation strategies to be applied to different data segments, thereby maintaining data quality while increasing quantity. The segmented approach prevents uniform augmentation from degrading overall data quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by implementing type-specific augmentation conditions and labeling strategies for different data types. Each data type receives customized augmentation parameters and labeling rules tailored to its characteristics, ensuring that augmented data maintains high relevance and quality for its specific category rather than applying a one-size-fits-all approach.

Inventive Principle:
Principle #3Local quality

2Reliability

If the number of input data is increased to improve prediction confidence, then the prediction confidence improves, but the computational resources and processing time increase

Engineering Contradiction:
Improveprediction confidenceVSAvoidcomputational resources
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary action by pre-categorizing data into types and pre-defining augmentation conditions and labeling rules for each type before the actual augmentation process. This preliminary organization enables more efficient processing during training, as the system doesn't need to make complex decisions during data generation, thereby reducing computational overhead while still achieving high prediction confidence through increased data quantity.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If uniform augmentation conditions are applied to all data, then the processing simplicity is maintained, but the adaptability to different data characteristics deteriorates

Engineering Contradiction:
Improveprocessing simplicityVSAvoidadaptability to data characteristics
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamics by making the augmentation conditions adaptive rather than static. The system dynamically selects augmentation parameters and labeling rules based on the identified data type, allowing the processing approach to change according to data characteristics. This dynamic adaptation maintains ease of operation through automated type-based selection while achieving high adaptability to different data types.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250299781A1Data augmentation methods, devices and programs for major histocompatibility complex class ii binding and immunogenicity predictive models
Publication Date: 2025.09.25 LG MANAGEMENT DEV INST CO LTD
  • US20250299781A1 patent drawing
  • US20250299781A1 patent drawing
  • US20250299781A1 patent drawing

AI summary

Data augmentation methods, devices, and programs for an MHC class II binding and immunogenicity predictive models may select a plurality of augmentation target data including first-type data and second-type data from original data according to a predetermined selection condition, to generate a plurality of augmentation data by augmenting the plurality of selected augmentation target data according to a predetermined augmentation condition, wherein the plurality of selected augmentation target data is augmented according to each of an augmentation condition of the first-type data and an augmentation condition of the second-type data, and to modify labeling of the plurality of augmentation data, wherein labels are modified according to different labeling conditions for each of the first-type data and the second-type data.