Machine Learning-Based Genome Segmentation for Precise Element Annotation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current genomic analysis methods struggle to accurately identify and characterize various protein coding and regulatory elements within DNA sequences, hindering a comprehensive understanding of disease mechanisms and therapy development.

Innovation Solution

Employing machine-learning technologies, particularly language models, to segment and annotate nucleotide sequences at high resolution, leveraging unsupervised training on unlabeled data and multi-task models to enhance accuracy and efficiency in genomic element localization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine learning models are used to annotate genomic elements, then measurement precision of genomic element locations is improved, but device complexity increases

Engineering Contradiction:
Improvegenomic element location precisionVSAvoidmodel architecture complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the complex annotation task into multiple specialized output channels, each dedicated to predicting specific genomic element types (e.g., promoters, enhancers, gene bodies). This modular approach improves measurement precision for each element type while managing model complexity through functional decomposition of the prediction task.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a universal machine learning model that performs multiple annotation functions simultaneously through a single multi-task framework. The model generates predictions for various genomic elements (promoters, enhancers, insulators, gene bodies) in one unified process, improving efficiency and consistency while reducing the need for separate specialized models.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If multiple output channels are used for different genomic elements, then measurement precision is improved, but device complexity increases

Engineering Contradiction:
Improvegenomic element classification accuracyVSAvoidnumber of output channels
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The model architecture segments the classification task into distinct output channels, each specialized for a specific genomic element type. This segmentation enables the model to apply element-specific prediction strategies and improves classification accuracy for each category while maintaining a unified model structure.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from a single-output classification problem to a multi-dimensional output space where each dimension corresponds to a different genomic element type. This dimensional expansion allows simultaneous prediction of multiple element types, improving overall annotation precision while organizing complexity into a structured multi-dimensional framework.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If language models are used for unsupervised training, then productivity of genomic analysis is improved, but loss of information increases

Engineering Contradiction:
Improvegenomic sequence processing speedVSAvoidlabeled data requirements
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent applies preliminary action through unsupervised pre-training of language models on large corpora of unlabeled genomic sequences. This pre-training phase enables the model to learn general genomic patterns and representations beforehand, allowing rapid processing of new sequences with minimal labeled data required for fine-tuning, thus improving productivity while reducing labeled data needs.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The language model performs self-service learning by automatically extracting patterns and features from unlabeled genomic data through unsupervised training. The model serves its own feature extraction needs without requiring extensive manual labeling, reducing information loss and enabling scalable genomic analysis with limited annotated datasets.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250335785A1Systems and methods for machine learning-based genome annotation
Publication Date: 2025.10.30 INSTADEEP LTD
  • US20250335785A1 patent drawing
  • US20250335785A1 patent drawing
  • US20250335785A1 patent drawing

AI summary

The present disclosure, among other things, provides machine-learning technologies for identifying and localizing particular genomic elements (e.g., gene elements and/or regulatory elements) within nucleotide sequences, such as DNA and/or RNA sequences. In certain embodiments, similar to the manner in which image processing methods can be used to localize particular objects in images at pixel level resolution, referred to as “segmentation,” systems and methods of the present disclosure predict presence and locations of certain genomic elements within nucleotide sequences, thereby “segmenting” nucleotide sequences. Accordingly, genomic element segmentation technologies described herein may be used to generate annotations that identify and label portions of nucleotide sequences according to their predicted (e.g., via machine learning models described herein) function—e.g., as protein-coding genes, untranslated regions, splice sites, promotors, enhancers, etc. Among other things, these genomic annotations may be used to inform underlying biological processes driving diseases and facilitate development of new therapies.