Multiplexed Immune Receptor Pairing via Sample Pooling and UMIs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for high-throughput sequencing of adaptive immune receptor nucleic acids are limited in their ability to efficiently pair adaptive immune receptor sequences from multiple biological samples simultaneously, requiring extensive resources and being challenging for assessing infiltrating immune cells in tissues or solid tumors.
Innovation Solution
A method involving multiplex PCR and high-throughput sequencing to determine rearranged nucleic acid sequences encoding adaptive immune receptor heterodimers, followed by pooling samples and statistical analysis to assign cognate pairs, allowing for simultaneous pairing of adaptive immune receptor sequences across multiple samples.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If current methods are used for high-throughput sequencing of adaptive immune receptor nucleic acids from multiple samples, then sequencing coverage is achieved, but time and resource input are excessive
Solution Approach 1:
The method segments the sequencing process by assigning unique molecular identifiers (UMIs) and sample-specific barcodes to individual nucleic acid molecules before pooling samples. This allows parallel processing of multiple samples through multiplexed sequencing while maintaining the ability to computationally separate and analyze each sample's data independently, thereby achieving high throughput without sacrificing analysis accuracy
Solution Approach 2:
The protocol uses universal primers and reagents that can process multiple sample types and immune receptor classes simultaneously. The same sequencing library preparation workflow handles diverse nucleic acid inputs from different biological samples, reducing the need for sample-specific optimization and decreasing overall processing time
2Productivity
If current methods are used for high-throughput sequencing of adaptive immune receptor nucleic acids from multiple samples, then sequencing coverage is achieved, but resource input is excessive
Solution Approach 1:
Multiple biological samples are combined into a single pooled library for simultaneous sequencing. By merging samples with unique barcodes and processing them through a single sequencing run, the method achieves multiplexed throughput while reducing the total amount of reagents, sequencing capacity, and infrastructure resources required compared to running separate sequencing experiments for each sample
Solution Approach 2:
The method discards low-quality or non-informative sequencing reads through computational filtering while recovering and retaining high-quality reads that contain valid immune receptor sequence information. This selective recovery approach maximizes the utility of sequencing resources by focusing analysis on productive data and eliminating waste from poor-quality reads
3Measurement precision
If pairing methods are used for adaptive immune receptor sequences, then pairing accuracy is achieved, but false discovery rates increase
Solution Approach 1:
Unique molecular identifiers (UMIs) serve as intermediary markers that track individual nucleic acid molecules through the sequencing process. These UMIs enable computational methods to distinguish true cognate pairings from random associations by verifying that paired sequences originate from the same original molecule, thereby reducing false discoveries while maintaining pairing accuracy
Solution Approach 2:
The method incorporates iterative computational validation where initial pairing predictions are tested against multiple criteria including UMI consistency, sequence quality metrics, and statistical significance thresholds. Results feed back into refined pairing assignments, with false positives identified and corrected through multiple rounds of validation, thereby reducing false discovery rates while preserving true pairings
4Measurement precision
If single-sample processing is used for adaptive immune receptor sequencing, then sample-specific analysis is achieved, but throughput is limited
Solution Approach 1:
The method segments each original nucleic acid molecule with a unique barcode identifier before pooling samples. This molecular-level segmentation allows the sequencer to process thousands of samples in parallel while computational methods can subsequently segment and analyze each sample's data independently, achieving both high throughput and sample-specific resolution simultaneously
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach significantly reduces time and resource input while accurately pairing adaptive immune receptor sequences, enabling efficient analysis of large numbers of biological samples with improved throughput and reduced false discovery rates.
Implementation Method 1
determining the first rearranged nucleic acid sequences may include: for each source sample, amplifying rearranged nucleic acid molecules extracted from the source sample in a single multiplex polymerase chain reaction (PCR) using a plurality of V-segment primers and a plurality of J-segment primers
Implementation Method 2
sequencing said plurality of rearranged nucleic acid amplicons to determine sequences of the first rearranged nucleic acid sequences in each source sample
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The invention is directed to methods for highly-multiplexed simultaneous detection of nucleic acids encoding paired adaptive immune heterodimers from a large number of biological samples containing lymphocytes of interest. Methods of the invention comprise performing a single pairing assay on a pool of source samples to determine nucleic acids encoding paired cognate receptor heterodimer chains in the combined pool. Separately, single-locus, high-throughput sequencing is performed on each individual sample to determine the plurality of sequences encoding one of the two receptor polypeptide chains. Pairs of cognate sequences determined in the pooled sample may then be mapped back to a single source sample by comparing the expression patterns of the single-locus sequences.