Polymorphism-Based Sample Matching for Sequencing Libraries
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Diagnostic testing for pathogens in patient samples faces challenges in accurately matching RNA and DNA sequencing libraries due to the possibility of sample mis-assignment, despite being prepared by highly trained technologists following standard procedures.
Innovation Solution
The method involves identifying and matching polymorphisms in RNA and DNA sequencing libraries using shared markers such as single nucleotide polymorphisms (SNPs) and mitochondrial DNA haplogroups, and assigning them a random index to ensure correct sample association, utilizing techniques like reverse transcription and next-generation sequencing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If RNA and DNA sequencing libraries are prepared separately by highly trained technologists following standard operating procedures, then the quality and reliability of library preparation is maintained, but there remains a possibility of sample mis-assignment where the RNA library is not from the same patient sample as the DNA library
Solution Approach 1:
The patent applies preliminary action by extracting and storing genetic markers (SNPs and mitochondrial haplogroups) from patient samples during the library preparation process, before sequencing occurs. These markers are embedded in the sequencing data as identity tags, enabling later verification of sample matching without interfering with the primary sequencing objectives
Solution Approach 2:
The patent implements feedback by comparing genetic markers between RNA and DNA sequencing libraries after sequencing is complete. The system provides feedback on whether the libraries derive from the same patient sample by matching SNP profiles and mitochondrial haplogroups, allowing detection and correction of any sample mis-assignment errors
2Reliability
If genetic markers such as SNPs and mitochondrial haplogroups are used to identify and match sequencing libraries, then sample mis-assignment is detected and prevented, but sensitive patient information including genotypes and haplogroups is exposed
Solution Approach 1:
The patent applies the taking out principle by extracting only the essential genetic marker information (SNP presence/absence patterns and mitochondrial haplogroup classifications) needed for sample identification, while excluding detailed genomic sequences and other sensitive patient information. This selective extraction maintains identification capability while minimizing privacy exposure
Solution Approach 2:
The patent applies parameter changes by transforming detailed genetic sequence data into simplified categorical markers - converting continuous genomic information into discrete SNP allele patterns and haplogroup classifications. This parameter transformation maintains the discriminatory power needed for sample matching while reducing the sensitivity and privacy risk of the stored information
3Productivity
If a small subset of highly polymorphic SNP loci is selected for genotyping to match samples, then the matching process becomes efficient and practical, but the precision and comprehensiveness of sample identification may be reduced compared to using the full genome
Solution Approach 1:
The patent applies local quality by selecting specific highly polymorphic SNP loci distributed across different genomic regions, particularly focusing on regions with known high variability. This targeted approach concentrates identification power in strategically chosen locations rather than uniformly sampling the entire genome, achieving efficient matching with adequate precision
Solution Approach 2:
The patent applies universality by selecting SNP markers that are highly polymorphic across diverse human populations, making the marker set universally applicable for sample matching across different ethnicities and patient groups. This universal marker set maintains high discriminatory power across diverse samples while keeping the genotyping process efficient and standardized
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach effectively identifies and verifies that RNA and DNA sequencing libraries derive from the same patient sample, enhancing diagnostic accuracy and reducing the risk of sample mis-assignment, while also protecting sensitive patient information by obfuscating genotypes and haplogroups.
Implementation Method 1
preparing (e.g., independently preparing) sequencing libraries for both the RNA (e.g., RNA converted to complementary DNA (cDNA)) and DNA molecules
Data Source
AI summary
The present disclosure provides methods and systems for processing samples including nucleic acid molecules. The methods may comprise identifying polymorphisms in a plurality of sequencing libraries and using the polymorphisms to identify the plurality of sequencing libraries as being associated with the same sample.


