Genome Graphs for Accurate Epigenetic Methylation Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for studying patterns of epigenetic modification in the human genome are limited, particularly in accurately detecting methylated cytosines due to the conversion of unmethylated cytosines to uracils during bisulfite sequencing, leading to inaccurate base calls.
Innovation Solution
The use of a directed acyclic graph (DAG) to represent the genome, where divergent paths account for potential modifications, combined with bisulfite sequencing (Methyl-Seq) to profile and align segments, allowing for sensitive and accurate determination of methylation patterns.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If bisulfite sequencing is used to detect methylated cytosines, then the method can identify epigenetic modifications, but unmethylated cytosines are converted to uracils leading to inaccurate base calls
Solution Approach 1:
The patent introduces an intermediary reference sequence that represents the expected DNA sequence without methylation. This reference acts as a mediator between the bisulfite-converted sequence (where unmethylated cytosines become uracils) and the original genomic sequence. By comparing the converted sequence against this intermediary reference, the system can accurately distinguish between originally methylated cytosines (which remain as cytosines after bisulfite treatment) and originally unmethylated cytosines (which become uracils), thereby resolving the information loss problem while maintaining detection accuracy.
2Device complexity
If a linear reference sequence is used for alignment, then the alignment process is simple, but it cannot accurately represent regions with potential modifications
Solution Approach 1:
The patent segments the reference sequence into multiple paths, creating a graph structure where each path represents a possible sequence variant. This segmentation allows the reference to accommodate regions with potential modifications (such as SNPs or epigenetic variations) without increasing overall complexity. The graph is divided into nodes and edges, where nodes represent sequence positions and edges represent possible transitions, enabling accurate alignment even when the actual sequence deviates from the linear reference.
Solution Approach 2:
The patent transitions from a one-dimensional linear reference sequence to a multi-dimensional graph structure. This dimensional change allows the reference to represent not only the canonical sequence but also alternative sequences at each position. The graph adds a vertical dimension of sequence variation while maintaining the horizontal progression through the sequence, enabling precise localization of modifications without overwhelming complexity.
3Measurement precision
If the genome is represented as a graph with divergent paths, then accurate alignment of modified regions is enabled, but the alignment process becomes more complex
Solution Approach 1:
The patent performs preliminary actions by pre-processing the bisulfite-converted sequence to identify and mark potential modification sites before alignment. It also pre-computes the graph structure with all possible paths representing known variations. This preliminary preparation simplifies the actual alignment process, as the system only needs to navigate pre-defined paths rather than constructing the graph dynamically during alignment, thereby reducing computational complexity while maintaining high precision.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables precise identification of methylated cytosines across substantial genomic regions, providing insights into organism development and potential early warnings of cancer or other clinical issues.
Implementation Method 1
Treatment of DNA with bisulfite converts un-methylated cytosines to uracil without affecting methylated cytosines
Data Source
AI summary
The invention provides systems and methods for determining patterns of modification to a genome of a subject by representing the genome using a graph, such as a directed acyclic graph (DAG) with divergent paths for regions that are potentially subject to modification, profiling segments of the genome for evidence of epigenetic modification, and aligning the profiled segments to the DAG to determine locations and patterns of the epigenetic modification within the genome.


