Combinatorial Barcode Sets for Single-Cell Genomic Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current genomic and transcriptomic profiling methods face inefficiencies due to Poisson-based distribution of nucleic acid barcodes, leading to low throughput and high loss of rare cells, especially when trying to achieve 95% of containers receiving exactly one cell and one barcode, resulting in more than 90% of containers being unproductive.
Innovation Solution
Introducing multiple barcodes that form a distinguishable barcode set within each reaction container, allowing all barcodes to be associated with each other, enabling the identification of the cell of origin for genomic or transcriptomic targets without requiring a majority of containers to be empty, thus overcoming the inefficiencies of Poisson distributions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If Poisson distribution is used to distribute cells and barcodes to containers, then the distribution follows statistical laws, but more than 90% of containers remain unproductive
Solution Approach 1:
The patent divides the barcode assignment problem into segments by allowing multiple barcodes per container rather than requiring exactly one. This segmentation of the barcode-to-container relationship enables containers to receive varying numbers of barcodes (0, 1, 2, or more) while maintaining identifiability through combinatorial code sets, thereby increasing the percentage of productive containers from less than 0.24% to potentially much higher values.
Solution Approach 2:
The patent changes the parameter of barcode concentration and distribution from a strict Poisson distribution with lambda=0.05 (designed for exactly one barcode per container) to a more flexible distribution where multiple barcodes can be present. By adjusting the expected number of barcodes per container (lambda) and allowing combinatorial codes, the system transforms the unproductive waste into useful information, as multiple barcodes can still identify a single cell through their combination.
2Measurement precision
If lambda is set to 0.05 for both cells and barcodes to achieve 95% single-occupancy, then fidelity is maintained, but fewer than 0.24% of containers are productive
Solution Approach 1:
The patent moves from a one-dimensional barcode assignment (one barcode per container) to a multi-dimensional combinatorial code space where multiple barcodes can coexist in the same container. By using combinations of barcodes (e.g., if barcodes A and B are present, they together identify a unique cell), the system adds a dimensional layer to the identification process that preserves fidelity while dramatically increasing productivity.
Solution Approach 2:
The patent employs multiple copies of barcodes in each container rather than a single unique barcode. Instead of requiring exactly one barcode per container, the system places multiple barcodes (potentially including duplicates or complementary codes) that together form a combinatorial identifier. This copying approach maintains measurement precision because the combination of barcodes uniquely identifies the cell, while increasing the likelihood that at least one barcode will be present in each productive container.
3Productivity
If multiple barcodes are introduced per container, then throughput increases, but barcode assignment complexity increases
Solution Approach 1:
The patent implements feedback mechanisms in the data analysis pipeline to handle multiple barcodes per container. By using algorithms that can deconvolute combinatorial barcode sets and assign them to parent cells, the system provides feedback that resolves the apparent complexity. The computational methods can identify which combinations of barcodes belong together and assign them to the appropriate cell profiles, thereby managing the increased complexity through systematic feedback-based resolution.
Solution Approach 2:
The patent creates a universal combinatorial barcode system that can handle multiple barcodes per container through a unified mathematical framework. The combinatorial code approach provides a universal method for assigning identity to cells regardless of how many barcodes are present, whether 1, 2, or more. This multi-functional system can accommodate varying barcode counts while maintaining a consistent assignment methodology, thereby managing complexity through universality rather than requiring separate handling for each scenario.
4Reliability
If single barcode per container is required, then unambiguous profiles are obtained, but greater than 95% of cells are lost from analysis
Solution Approach 1:
The patent applies preliminary combinatorial barcode assignment before cell profiling, where multiple barcodes are pre-assigned to potential cell identities. By establishing a combinatorial code framework in advance, the system prepares for multiple barcodes per container rather than requiring strict single-occupancy. This preliminary action of creating combinatorial code sets enables the system to maintain profile unambiguity even when multiple barcodes are present, as the combinatorial relationships were pre-established and can be resolved during data analysis.
Data Source
Figure 1~2B
Figure 3
Figure 4
AI summary
Provided herein are methods of identifying the origin of a nucleic acid sample. The methods include forming a reaction mixture comprising a nucleic acid sample comprising nucleic acid molecules from a single cell and a set of barcodes, incorporating the set of barcodes into the nucleic acid molecules of the sample, and identifying the set of barcodes incorporated into the nucleic acid molecules of the single cell thereby identifying the origin of the nucleic acid sample.