Cross-Platform SNP Matching for Sequencing Sample Identity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current short-read sequencing technologies in clinical settings are prone to missing important target regions and SNPs due to misaligned sequence overlap, while nanopore long-read sequencing offers a promising alternative but requires methods to ensure sample integrity and prevent mix-ups between assays.
Innovation Solution
A method is developed to match samples sequenced using disparate techniques by identifying short-read SNPs, determining target regions for long-read sequencing, and comparing short-read and long-read sequence data to confirm sample identity, leveraging existing sequencing data without additional wet or dry lab protocols.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If short-read sequencing is used to identify variants, then the sequencing can be performed in clinical settings with established technology, but important target regions and SNPs may be missed due to misaligned sequence overlap
Solution Approach 1:
The patent segments the genome into specific target regions containing SNPs of interest. By focusing sequencing efforts on these predefined segments rather than attempting to sequence the entire genome continuously, the method ensures comprehensive coverage of critical regions while avoiding the misalignment issues that plague whole-genome short-read sequencing.
Solution Approach 2:
The patent performs preliminary identification of target regions and SNPs using short-read sequencing data before conducting long-read sequencing. This preliminary action allows the methodology to focus subsequent long-read sequencing efforts on specific regions of interest, ensuring that no important SNPs are missed while optimizing resource utilization.
2Reliability
If nanopore long-read sequencing is used as an alternative method, then structural variant breakpoint detection is improved and direct native DNA sequencing is possible, but sample mix-ups and identity confirmation between assays become critical challenges
Solution Approach 1:
The patent employs SNPs as universal identifiers that serve multiple functions: they characterize samples for disease variant detection, enable sample identity verification, and facilitate cross-platform data integration. This multi-functionality resolves the complexity of sample verification while maintaining the advantages of long-read sequencing.
Solution Approach 2:
The patent uses SNP profiles as an intermediary mechanism to bridge short-read and long-read sequencing data. By comparing SNP characteristics across both sequencing platforms, the method enables sample identity confirmation without requiring direct integration of the disparate sequencing technologies, thus reducing operational complexity.
3Reliability
If short-read and long-read sequencing data are integrated, then comprehensive variant detection is achieved, but sample matching and data correlation between platforms become necessary
Solution Approach 1:
The patent implements self-service sample matching by utilizing the SNP profiles inherently present in both short-read and long-read sequencing data. The system automatically extracts and compares SNP characteristics without requiring manual sample matching procedures, thus maintaining ease of operation while achieving comprehensive variant detection across platforms.
Data Source
AI summary
Disclosed herein are methods of determining sample matches. An example method includes receiving short-read sequence data, wherein the short-read sequence data is generated from a first sample using a short-read sequencing technique. The method further includes identifying a plurality of short-read single nucleotide polymorphisms (SNPs) from the short-read sequence data and selecting one or more SNPs from the plurality of short-read SNPs. The method further includes receiving long-read sequence data, wherein the long-read sequence data is generated from a second sample using a long-read sequencing technique, wherein the long-read sequence data comprises sequence data for at least the one or more SNPs. The method further includes determining whether the first sample and the second sample match based on the short-read sequence data and the long-read sequence data.


