Reduced Representation Library for Target DNA Genome Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current DNA analysis methods face challenges with massive data generation from next-generation sequencing (NGS), including difficulties in genome assembly due to short read lengths, data storage and transfer issues, ambiguities in repeat DNA areas, and limitations in handling low amounts of sample material, leading to incomplete or incorrect data.
Innovation Solution
The method involves creating a reduced representation library (RRL) of the target DNA genome using predetermined sequences, such as restriction enzyme recognition sites, for efficient sequencing and analysis, allowing for clustering of non-overlapping segments with similar metrics to provide master segments, which include inferred boundaries, read counts, and ancestral probabilities for enhanced interpretation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If whole genome sequencing is performed using NGS, then comprehensive genome data is obtained, but data storage and processing requirements become excessively large
Solution Approach 1:
The patent extracts and sequences only specific genomic regions of interest (e.g., disease-associated genes, exons, or targeted loci) rather than the entire genome. This is achieved through hybridization capture or PCR amplification of selected regions, reducing data volume by orders of magnitude while maintaining clinical relevance.
Solution Approach 2:
The genome is divided into specific segments or regions of interest for sequencing. The patent focuses on sequencing only the necessary portions of the genome (e.g., 0.3% of the genome for exome sequencing), thereby reducing the overall data burden while preserving critical genetic information.
2Productivity
If short read lengths are used in NGS, then sequencing coverage is increased, but assembly and alignment accuracy deteriorate
Solution Approach 1:
The patent segments the genome into specific target regions and sequences each region with sufficient depth. By focusing on discrete genomic locations rather than attempting to assemble the entire genome, the method achieves high coverage of clinically relevant areas without the assembly challenges of short reads.
Solution Approach 2:
The patent uses reference genomes and alignment algorithms as intermediaries to map short reads to known genomic positions. This approach bypasses the need for de novo assembly, allowing short reads to be accurately placed in the genome based on their similarity to reference sequences.
3Measurement precision
If massive amounts of NGS data are generated, then sequencing depth is improved, but data transfer and storage challenges increase
Solution Approach 1:
The patent extracts only the essential genomic information needed for clinical diagnosis, storing results in compact formats such as variant calls (VCF files) or structured reports rather than raw sequence data. This reduces storage requirements from terabytes to megabytes while preserving diagnostic capability.
Solution Approach 2:
The patent discards redundant or low-value sequence data during processing, retaining only clinically significant variants and essential quality metrics. This selective data retention maintains diagnostic accuracy while minimizing storage requirements.
4Ease of operation
If discrete base calls are used as input for analysis, then data processing is simplified, but information loss occurs during downstream analysis
Solution Approach 1:
The patent maintains continuous retention of raw sequence data, base quality scores, and alignment information throughout the analysis pipeline. Rather than converting to discrete calls early, the system preserves the full spectrum of data types, allowing flexible re-analysis and correction as new insights emerge.
Solution Approach 2:
The patent performs preliminary quality control and filtering steps that preserve maximum information, rather than making premature discrete calls. Quality metrics and confidence scores are calculated and stored in advance, enabling later refinement of variant calls without losing underlying data.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach reduces the amount of DNA needed for sequencing, decreases NGS run time, and increases computational and storage efficiency, enabling reliable genome-wide analysis, especially with limited sample material, and improves the detection of genetic variations and risk alleles.
Implementation Method 1
a reduced representation library (RRL) of the target DNA genome using predetermined sequences, such as restriction enzyme recognition sites
Data Source
Figure 1
Figure 2
Figure 3A
AI summary
A method of target DNA genome analysis is provided. The method comprises the steps of : - obtaining non-overlapping segments of target DNA stretches with segment boundaries defined by the presence of particular restriction enzyme recognition sites, whereby the assembly of said non-overlapping segments compose a reduced representation library of said target DNA genome; - obtaining for said segments, raw metrics from a sequencing process applied on said reduced representation library; - clustering non-overlapping, nearby segments with similar raw metrics to provide master segments; - providing metrics describing the master segments, - making a final discrete DNA call based on the master segments and its metrics.