Segment-wise Input Feature Representation for Genomics Data Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing predictive data analysis systems face challenges in efficiently processing large and data-intensive datasets, particularly in genomics, due to the massive size of genetic data and hardware/software limitations, making it difficult to represent genetic variants consistently for ingestion by Deep Learning algorithms.

Innovation Solution

The approach involves converting large datasets into input feature representation super-segments and segments, using a transformer-based language model to generate multi-segment input feature representations, and performing predictive data analysis by segmenting input spaces hierarchically, enabling faster and less resource-intensive processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If large genomics datasets are processed using traditional Deep Learning algorithms, then predictive data analysis can be performed, but hardware resources and processing time become excessively large

Engineering Contradiction:
Improvepredictive data analysis capabilityVSAvoidhardware resource consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent segments the input feature representation into multiple segments (e.g., chromosomal segments) and processes each segment independently through a shared embedding model. This segmentation reduces the computational burden on hardware by breaking down large datasets into manageable units, thereby resolving the contradiction between performing predictive analysis on large genomics data and avoiding excessive hardware resource consumption.

Inventive Principle:
Principle #1Segmentation

2Productivity

If large genomics datasets are processed using traditional Deep Learning algorithms, then predictive data analysis can be performed, but processing time becomes excessively large

Engineering Contradiction:
Improvepredictive data analysis capabilityVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

By dividing the input feature representation into chromosomal segments and processing them in parallel through the shared embedding model, the patent significantly reduces processing time. The segmentation enables concurrent computation on multiple data portions, thereby resolving the contradiction between performing predictive analysis and minimizing processing time.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The shared embedding model serves multiple functions by processing different chromosomal segments through a single unified model architecture. This multi-functionality reduces the need for multiple separate model instances, thereby decreasing processing time while maintaining predictive analysis capability on large genomics datasets.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If genetic variants are represented consistently for Deep Learning ingestion, then predictive analysis can be performed, but device complexity increases

Engineering Contradiction:
Improveconsistent genetic variant representationVSAvoiddata processing system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the genetic data into chromosomal segments, each processed by the shared embedding model to generate consistent representations. This segmentation approach simplifies the overall processing system by using a single model architecture for all segments, thereby resolving the contradiction between achieving consistent genetic variant representation and avoiding increased device complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The shared embedding model acts as an intermediary that transforms diverse genetic variant data into consistent representations. By using a single intermediary model for all chromosomal segments, the patent maintains representation consistency while avoiding the complexity of multiple specialized models, thereby resolving the contradiction between reliable representation and system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20230089140A1Machine learning techniques using segment-wise representations of input feature representation segments
Publication Date: 2023.03.23 OPTUM SERVICES IRELAND LTD
  • US20230089140A1 patent drawing
  • US20230089140A1 patent drawing
  • US20230089140A1 patent drawing

AI summary

Various embodiments of the present invention provide methods, apparatus, systems, computing devices, computing entities, and/or the like for performing health-related predictive data analysis. Certain embodiments of the present invention utilize systems, methods, and computer program products that perform predictive data analysis by using at least one of shared segment embedding machine learning models or transformer-based machine learning models.