Graph-Based Genomic Reference Model for Read Alignment Bias

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current genomic analysis methods face challenges such as reference allele bias, inefficient data storage, and computational limitations in processing large-scale genomic data, particularly in representing polymorphisms and structural variations, which hinder accurate and efficient DNA sequence analysis.

Innovation Solution

A graph-based reference genome framework, known as the GNOmics Graph Model (GGM), is introduced, which represents all known polymorphisms, including SNPs, indels, and structural changes, using nodes and edges to enable efficient data compression and parallel computing, allowing for accurate read alignment and reduced storage requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a linear reference genome is used for read mapping, then the mapping process is simple and straightforward, but reference allele bias occurs and polymorphic variants are not accurately represented

Engineering Contradiction:
Improvesimplicity of read mappingVSAvoidaccuracy of variant detection
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The linear reference genome is segmented into multiple alternative alleles at polymorphic positions, creating a graph structure where each node represents a possible variant. This segmentation allows reads to be mapped against multiple alternative sequences simultaneously, eliminating reference allele bias while maintaining manageable complexity through localized branching at variant sites.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from a one-dimensional linear reference genome to a multi-dimensional graph structure where additional dimensions represent alternative alleles and polymorphic variants. This dimensional expansion allows the system to represent genetic diversity without proportionally increasing computational complexity, as the graph structure organizes alternatives in a navigable framework.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If multiple alternative reference genomes are created to represent population diversity, then polymorphism representation improves, but data storage requirements and computational complexity increase significantly

Engineering Contradiction:
Improverepresentation of genetic variationVSAvoidcomputational and storage complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

Multiple alternative reference genomes are merged into a single graph-based reference structure where shared sequences are represented once and alternative alleles are connected through graph edges. This merging eliminates redundant storage of identical sequences while maintaining the ability to represent population diversity, reducing both storage requirements and computational complexity compared to maintaining separate reference genomes.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The graph-based reference genome serves multiple functions simultaneously: it represents the canonical reference sequence, incorporates population-specific variants, and provides alternative alleles for diverse populations. This universal structure eliminates the need for separate reference genomes for different populations, reducing overall system complexity while enhancing adaptability to various genetic backgrounds.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If a graph-based reference genome is used to represent all polymorphisms, then reference allele bias is eliminated and variant detection accuracy improves, but the complexity of the reference structure increases

Engineering Contradiction:
Improveaccuracy of read alignmentVSAvoidcomplexity of reference genome structure
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The graph-based reference genome is designed to be dynamically adaptable, allowing efficient navigation and traversal algorithms that adjust the effective complexity based on the query sequence. Reads are mapped by dynamically selecting relevant paths through the graph based on sequence similarity, which maintains computational efficiency despite the underlying complex structure. The system adapts to each mapping query rather than requiring exhaustive processing of all possible paths.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10600217B2Methods for the graphical representation of genomic sequence data
Publication Date: 2020.03.24 THE UNIV OF BRITISH COLUMBIA
  • US10600217B2 patent drawing
  • US10600217B2 patent drawing
  • US10600217B2 patent drawing

AI summary

This disclosure provides a computational framework with related methods and systems to enhance the analysis of genomic information. More specifically, the disclosure provides for a graph-based reference genome framework, referred to as a GNOmics Graph Model (GGM), which represents genomic sequence information in edges with nodes representing transitions between edges. The disclosed GGM framework can represent all known polymorphisms simultaneously, including, SNPs, indels, and various rearrangements, in a data-efficient manner. The edges can contain weights to reflect the likelihood of a path within the GGM incorporating any particular edge. The disclosure also provides for systems and methods for using the GGM as a reference model for the rapid assembly of short sequence reads and analysis of DNA sequence variation with enhanced computational efficiency.