Single Cell Genome Reconstruction via Transposase Fragmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current single-cell DNA sequencing technologies fail to accurately determine discrete gene copy numbers due to their reliance on secondary analysis methodologies developed for bulk sequencing, leading to variability and inaccuracy in copy number determination, especially in cancer cells where gene copy numbers are heterogeneous.
Innovation Solution
A method involving the use of a transposase to fragment genomic DNA, labeling fragments with identical transposon ends, and reconstructing the genome by identifying overlapping sequences to determine the phase and ploidy of genomic regions, allowing for precise counting of DNA molecules and discrete gene copy numbers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If secondary analysis methodologies developed for bulk sequencing are used to determine copy numbers in single-cell sequencing, then the analysis can be performed using existing bulk sequencing pipelines, but the copy number determination becomes inaccurate and variable, especially in cancer cells with heterogeneous gene copy numbers
Solution Approach 1:
The patent segments the genome into individual molecules by using transposase to insert unique molecular identifiers (UMIs) at random positions across each DNA molecule. This segmentation allows each molecule to be tracked independently through sequencing, enabling accurate copy number determination by counting individual molecules rather than relying on bulk statistical methods. The segmentation principle directly resolves the contradiction by maintaining ease of analysis through standardized pipelines while achieving precise single-molecule resolution.
Solution Approach 2:
The patent uses transposase to create copies of unique molecular identifiers (UMIs) at the ends of each DNA fragment. These copied UMI sequences serve as molecular barcodes that allow reconstruction of original DNA molecules from fragmented sequencing reads. This copying mechanism enables accurate tracking of individual genome copies through the sequencing process, resolving the accuracy issue while maintaining compatibility with existing sequencing workflows.
2Adaptability or versatility
If statistical methods with binning and segmentation are used to infer copy numbers from mapped reads, then the analysis can handle bulk sequencing data, but the inferred copy numbers vary when different bin sizes or segmentation options are selected
Solution Approach 1:
The patent performs preliminary action by inserting unique molecular identifiers (UMIs) and transposon ends into each DNA molecule before fragmentation and sequencing. This pre-labeling allows direct tracking of individual molecules through the entire sequencing process, eliminating the need for post-sequencing statistical inference. The preliminary tagging action ensures that copy number determination relies on direct molecular counting rather than indirect statistical estimation, resolving the reliability issue while maintaining adaptability to existing pipelines.
3Quantity of substance
If current single-cell DNA sequencing technologies are used, then single-cell analysis can be performed, but the discrete nature of gene copy numbers is not captured, resulting in decimal approximations rather than true integer values
Solution Approach 1:
The patent employs self-service by using the DNA molecules themselves to carry unique molecular identifiers (UMIs) that automatically track their identity and origin. Each DNA molecule serves its own identification purpose through the embedded UMI, eliminating the need for external statistical inference methods. This self-identification mechanism enables direct counting of discrete molecular copies, achieving true integer copy number values while maintaining single-cell analysis capability.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enables accurate reconstruction of single-cell genomes and precise counting of DNA molecules, overcoming the limitations of existing technologies by providing a robust method for determining discrete gene copy numbers in single cells, particularly in cancer cells where heterogeneity is a challenge.
Implementation Method 1
contacting and fragmenting the genomic DNA using a transposase loaded with two identical transposon ends to form a plurality of genomic DNA fragments each labeled with an identical transposon end at its 5′ and 3′ ends
Implementation Method 2
extending a complementary strand of each fragment using a universal primer comprising a sequence complementary to the identical transposon end to generate one or more extension products
Data Source
AI summary
Single-cell sequencing provides a new level of granularity in studying the heterogeneous nature of cancer cells. For some cancers, this heterogeneity is the result of copy number changes of genes within the cellular genomes. The ability to accurately determine such copy number changes is critical in tracing and understanding tumorigenesis. Current single-cell genome sequencing methodologies infer copy numbers based on statistical approaches followed by rounding decimal numbers to integer values. Such methodologies are sample dependent, have varying calling sensitivities which heavily depend on the sample's ploidy and are sensitive to noise in sequencing data. Described herein are novel methods for reconstructing the genome of a single cell. The methods comprise fragmenting the genome using a loaded transposase, linking together fragments based on the overlapping 8-10 nucleotide genomic sequence immediately next to the transposon end to restore the order of the fragments as originally present in the genome, and reconstructing the genome by disregarding fragments that result from a defective transposase reaction and therefore cannot be linked with a neighboring fragment.


