Merging DNA and RNA Sequencing Data for Variant Calling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for analyzing multi-omic data fail to integrate functional relationships between DNA and RNA sequencing data, leading to suboptimal variant calling and characterization, particularly in understanding the pathological impacts of genetic variants.
Innovation Solution
A system and method that merge aligned RNA and DNA sequencing data to identify and characterize variants, assigning RNA editing and expression status using allele-specific expression categorizations, thereby improving variant calling accuracy and providing insights into transcriptional abundance and functional impacts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If DNA and RNA sequencing data are analyzed separately by different bioinformatics pipelines, then the analysis process is simple and straightforward, but the functional relationships between the data are not utilized and variant characterization is incomplete
Solution Approach 1:
The patent merges aligned DNA sequencing data and aligned RNA sequencing data into a single merged alignment file, enabling integrated analysis of both data types. This combining approach allows the system to capture functional relationships between DNA variants and RNA expression while maintaining a unified analysis framework rather than separate silos.
Solution Approach 2:
The merged alignment system serves multiple functions: it enables variant calling from both DNA and RNA data, performs expression analysis, characterizes variant impact on transcription, and identifies RNA editing events. This multi-functional approach eliminates the need for separate analysis pipelines while comprehensive information extraction.
2Measurement precision
If variant calling is performed using only DNA sequencing data, then the variant detection process is straightforward, but the accuracy and sensitivity of variant calls are limited
Solution Approach 1:
The system uses RNA sequencing data to cross-validate and improve DNA-based variant calls. By comparing variants detected in both DNA and RNA data, the system can confirm true positives and filter false positives, thereby improving variant call accuracy through feedback from the RNA dataset.
Solution Approach 2:
The patent creates a composite analysis approach by integrating DNA and RNA sequencing data into a unified framework. This composite methodology combines the strengths of both data types—DNA for comprehensive variant detection and RNA for expression context and validation—to achieve superior variant calling accuracy.
3Loss of information
If multiple types of -omic data are generated for the same sample, then better understanding of biological systems is achieved, but the data analysis becomes more complex and computationally intensive
Solution Approach 1:
The patent merges DNA and RNA sequencing data into a single aligned framework, enabling simultaneous analysis of genomic and transcriptomic information. This unified approach captures comprehensive biological system information including variant effects on transcription, expression levels, and RNA editing while avoiding the complexity of entirely separate analysis pipelines.
Data Source
AI summary
A method (100) for characterizing variant expression status for variants identified from a genomic sample, comprising: (i) obtaining (110) DNA sequencing data for the genomic sample; (ii) obtaining (110) RNA sequencing data for the genomic sample, wherein the obtained RNA sequencing data further comprises expression data for each variant; (iii) merging (130) the aligned DNA and RNA sequencing data into a merged alignment; (iv) identifying (140) a plurality of variants relative to the reference genome to generate a set of variants; (v) characterizing (150) an RNA-editing and/or expression status for each of at least a plurality of variants, wherein the expression status comprises one of a plurality of allele-specific expression categorizations comprising expression information for an alternative allele of the variant and expression information for a reference allele of the variant if there is one; and (vi) generating (160) a report comprising the characterized expression status for the variants.


