Paired immune receptor sequencing from 3' barcoding based rna-seq methods

WO2025165960A3PCT designated stage Publication Date: 2025-12-11REGENERON PHARMACEUTICALS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/013737
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-31
Filing Date
2025-01-30
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing high-throughput single cell RNA sequencing methods using 3' barcoding chemistries are unable to effectively capture full-length immune receptor sequences, particularly TCR or BCR information, due to the distance between the 3' barcodes and the 5' VDJ regions, leading to incomplete sequencing results.

Method used

A method involving the generation of a 3'-barcoded cDNA library with immune receptor sequences, amplification, probe linking, capture, and long-read sequencing to produce full-length immune receptor sequences, followed by computational processing to annotate and correct the variable regions using a processor.

Benefits of technology

Enables complete capture of full-length immune receptor sequences and accurate annotation of allele-level variable regions, improving the understanding of immune receptor diversity and clonality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025013737_11122025_PF_FP_ABST
    Figure US2025013737_11122025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for long-read sequencing of immune receptor are disclosed. A method in accordance with the present disclosure comprises generating a 3'-barcoded cDNA library comprising an immune receptor target sequence, wherein the immune receptor target sequence comprises a variable region and a barcoded region; amplifying the immune receptor target sequence; linking the variable region of the immune receptor target sequence to a probe to generate a labeled immune receptor target sequence; capturing the labeled immune receptor target sequence; amplifying the captured immune receptor target sequence; passing the captured immune receptor target sequence through a nanopore sequencing device; and generating, by way of the nanopore sequencing device, a full-length immune receptor target sequence.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] PAIRED IMMUNE RECEPTOR SEQUENCING FROM 3' BARCODING BASED RNA-SEQ

[0002] METHODS

[0003] Field

[0004] The present disclosure generally relates to paired immune receptor sequencing and, more particularly, to systems and methods for long-read sequencing of immune receptors using 3' single cell RNA sequencing (scRNA-seq) methods.

[0005] Background

[0006] High-throughput (HT) single cell RNA sequencing (scRNA-seq) allows for the analysis of the gene expression profiles of individual cells within a complex biological sample. The technique allows researchers to uncover rare cell types, identify cell subpopulations, and characterize cellular diversity within a tissue or sample. The "high throughput" aspect refers to the ability to process a large number of individual cells in parallel. Modern HT scRNA-seq platforms can handle thousands to millions of cells in a single experiment, making it feasible to analyze complex and heterogeneous samples comprehensively.

[0007] HT scRNA-seq involves several steps, including cell isolation, cell lysis, reverse transcription of RNA into cDNA, and library preparation. During reverse transcription, unique molecular identifiers (UMIs) and / or barcodes are often introduced to each cDNA molecule from a single cell. UMIs help in accurate estimation of gene expression as it corrects for amplification bias by uniquely tagging the original RNA molecule. Barcodes help distinguish individual cells' gene expression profiles in sequencing data. After library preparation, the cDNA libraries from individual cells are pooled and subjected to HT DNA sequencing. Sequence data for each cDNA molecule is generated, including information regarding which genes are expressed in each cell and at what levels.

[0008] VDJ-Seq, also known as immune repertoire sequencing or TCR / BCR sequencing, is a technique used to analyze the diversity and composition of T-cell receptor (TCR) or B-cell receptor (BCR) repertoires. The immune repertoire refers to the collection of different TCR or BCR sequences expressed by lymphocytes in an individual. VDJ-Seq involves sequencing the variable regions of TCR or BCR genes to capture the unique receptor sequences present in a sample. More specifically, VDJ-seq allows for profiling of the immune repertoire by sequencing the rearranged V (variable), D (diversity), and J (joining) gene segments that make up the TCR or BCR chains. This technique provides information about the clonality, diversity, and frequency of TCR or BCR sequences within a sample.

[0009] Various HT cell or spatial barcoding RNA-seq methods, with or without VDJ-seq capabilities, are listed in Table 1 below:

[0010] Table 1 lOx Immune Profiling (lOx Genomics) is a platform for studying the immune system at single-cell resolution. The platform utilizes a combination of single-cell RNA-sequencing (scRNA-seq) and immune-specific 5' barcoding chemistry to analyze gene expression in individual immune cells. lOx 3' GEX is designed to analyze the 3' end of messenger RNA (mRNA) molecules in individual cells, providing insights into gene expression patterns and cellular heterogeneity. The platform utilizes a combination of microfluidics, barcoding and next-generation sequencing (NGS) to capture and sequence the 3' end of mRNA molecules from thousands to millions of individual cells simultaneously. The workflow of lOx 3' GEX involves encapsulating individual cells into droplets along with barcoded oligonucleotides that tag and capture the mRNA from each cell. Within each droplet, reverse transcription of the captured mRNA occurs, followed by amplification after the droplets are broken. Barcoded complementary DNA (cDNA) libraries for each cell. The resulting cDNA libraries can then be sequenced using the NGS platforms to obtain gene expression data. BD Rhapsody VDJ is a technology platform developed for the analysis of immune repertoire diversity and clonality. The platform is configured to allow for study of the variable regions of immunoglobulins (Ig) and T cell receptors (TCR) at the single-cell level. The workflow involves isolating and capturing individual immune cells into droplets, where the cell's RNA is barcoded and amplified. Subsequently, the variable regions of Ig and TCR genes are specifically targeted and sequenced. lOx Visium is a platform that enables the study of gene expression patterns and spatial organization within intact tissue samples. The platform involves the preparation of tissue sections mounted on a slide or a capture area. The tissue is permeabilized to allow for the capture of RNA, and then the capture oligonucleotides are printed on the slide. Each capture oligonucleotide carries a spatial barcode, which marks the specific location of the tissue section on the slide. Only RNA in that spot will produce cDNA with the specific spatial barcode. The captured RNA is then reverse transcribed into complementary DNA (cDNA) libraries that retain the spatial information of their origin. After library preparation, the cDNA libraries are sequenced using next-generation sequencing (NGS) technologies. The resulting data provides a spatially resolved transcriptome, where the expression levels of genes can be mapped to their respective locations within the tissue section.

[0011] CurioSeeker enables whole-transcriptome spatial mapping of fresh frozen tissues at single-cell resolution. The CurioSeeker platform utilizes spatially indexed beads to capture mRNA from tissue sections, followed by reverse transcription and next-generation sequencing (NGS). The resulting data may be processed using a bioinformatics pipeline to create detailed spatial gene expression maps. lOx Multiome enables the simultaneous profiling of gene expression and chromatin accessibility or DNA sequence variations at the single-cell level. The lOx Multiome platform allows for the integration of single-cell RNA sequencing with either single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) or single-nucleus variant calling (SNV) analysis. The workflow involves encapsulation of individual cells or nuclei into droplets containing unique barcodes, followed by the simultaneous processing of these cells for RNA-seq and either ATAC-seq or SNV analysis. As shown in Table 1 above, VDJ-Seq of TCRs and BCRs using HT single cell RNA-seq methods generally requires 5' barcoding chemistry. Furthermore, VDJ-Seq with 5' barcoding requires target enrichment using constant (C) gene primers, as illustrated in FIG. 1. C gene primers refer to primers designed to target and amplify the constant (C) region of immune cell receptors. The same target enrichment, however, is incompatible with VDJ-Seq of TCRs and / or BCRs using 3' barcoding chemistries, as illustrated in FIG. 2.

[0012] In 3' barcoding, the cell barcode and unique molecular identifiers (UMIs) are at the 3' end and are captured first, but the VDJ regions, which are towards the 5’ end, may not be sequenced using short read sequencing due to the distance from the 3’ end. This can result in incomplete information about the immune receptor sequences. In contrast, with 5’ barcoding, both the cell barcode and the VDJ regions are near the start of the sequencing process. This allows for the simultaneous capture of the cell's identity (through the barcode) and its immune receptor sequences (through VDJ sequencing), providing a more complete picture of the cell's characteristics and function, even with short read sequencing.

[0013] As such, 3' barcoding chemistries cannot be effectively used to obtain TCR or BCR information. Alternative methods of obtaining TCR and BCR information with 3' barcoding chemistries are needed.

[0014] Summary

[0015] The present disclosure provides systems and methods for long-read sequencing of immune cell receptors. In some aspects, a method of long-read sequencing of an immune cell receptor is provided. The method may include generating a 3'-barcoded cDNA library comprising an immune receptor target sequence, wherein the immune receptor target sequence comprises a variable region and a barcoded region; amplifying the immune receptor target sequence; linkingthe variable region of the immune receptor target sequence to a probe to generate a labeled immune receptor target sequence; capturing the labeled immune receptor target sequence; amplifying the captured immune receptor target sequence; passing the captured immune receptor target sequence through a long-read sequencing device; and generating, by way of the long-read sequencing device, a full-length immune receptor target sequence. The method may include additional, less, or alternate functionality, including that discussed elsewhere herein.

[0016] The present disclosure also provides a computer-implemented method for annotating the variable region of a full-length immune receptor target sequence. The method may be implemented using a system including a computing device including a processor communicatively coupled to a memory device. Additionally, or alternatively, the computer- implemented method may be implemented via one or more local or remote processors, servers, transceivers, memory units, mobile devices, wearables, smart watches, smart contact lenses, smart glasses, augmented reality glasses, virtual reality headsets, mixed or extended reality glasses or headsets, voice or chat bots, ChatGPT bots, and / or other electronic or electrical components, which may be in wired or wireless communication with one another. The computer-implemented method comprises receiving, by way of a processor, a full-length immune receptor target sequence generated from a 3'-barcoded cDNA comprising an immune receptor target sequence comprising a variable region and a barcoded region; removing, by way of the processor, unwanted sequences from the immune receptor target sequence; identifying, by way of the processor, the barcode of the full-length immune receptor target sequence; comparing byway of the processor, the identified barcode with a reference barcode; wherein the comparing comprises: aligning the identified barcode with the reference barcode; and determining whetherthe identified barcode originated from the reference barcode. The computer-implemented method further includes annotating, by way of the processor, the variable region of the full-length immune receptor target sequence; and correcting, by way of the processor, the variable region of the full-length immune receptor target sequence. The computer-implemented method may include additional, less, or alternate functionality, including those discussed elsewhere herein.

[0017] The present disclosure provides a system comprising at least one memory and at least one processor in communication with the at least one memory. The processor may be programmed to: receive, by way of the processor, a full-length immune receptor target sequence generated from a 3' barcoded cDNA comprising an immune receptor target sequence comprising a variable region and a barcoded region; remove, by way of the processor, one or more unwanted sequences from the full-length immune receptor target sequence; identify, by way of the processor, the barcode of the full-length TCR target sequence; compare, by way of the processor, the identified barcode with a reference barcode; wherein the processor is further programmed to align the identified barcode with the reference barcode; and determine whether the identified barcode originated from the reference barcode; annotate, by way of the processor, the variable region of the full-length immune receptor target sequence; and correct, by way of the processor, the variable region of the full-length immune receptor target sequence. The system may include additional, less, or alternate actions, including those discussed elsewhere herein.

[0018] The present disclosure provides at least one non-transitory computer-readable storage media having computer-executable instructions embodied thereon. The computerexecutable instructions, when executed by at least one processor, cause the at least one processorto: receive, byway of the processor, a full-length immune receptor target sequence generated from a 3' barcoded cDNA comprising an immune receptor target sequence comprising a variable region and a barcoded region; remove, by way of the processor, one or more unwanted sequences from the full-length immune receptor target sequence; identify, by way of the processor, the barcode of the full-length immune receptor target sequence; compare, by way of the processor, the identified barcode with a reference barcode; wherein the processor is further programmed to align the identified barcode with the reference barcode; and determine whether the identified barcode originated from the reference barcode; annotate, by way of the processor, the variable region of the full-length immune receptor target sequence; and correct, by way of the processor, the variable region of the full-length immune receptor target sequence. The storage medium may include additional, less, or alternate actions, including those discussed elsewhere herein.

[0019] Advantages will become more apparent to those skilled in the art from the following description of the preferred embodiments which have been shown and described by way of illustration. As will be realized, the present embodiments may be capable of other and different embodiments, and their details are capable of modification in various respects. Accordingly, the drawings and description are to be regarded as illustrative in nature and not as restrictive. Brief Description Of The Drawings

[0020] Those of skill in the art will understand that the figures, described below, are for illustrative purposes only. The figures are not intended to limit the scope of the present disclosure in any way.

[0021] FIG. 1 illustrates a conventional VDJ-seq method with 5' barcoding wherein target enrichment using C gene primers is used.

[0022] FIG. 2 illustrates a VDJ-seq method with 3' barcoding wherein target enrichment is incompatible.

[0023] FIG. 3 illustrates an embodiment of the long-read sequencing method of the present disclosure.

[0024] FIG. 4 illustrates an example nanopore sequencing process in accordance with an embodiment of the present disclosure.

[0025] FIG. 5 illustrates an example full-length TCR target sequence of the present disclosure generated using the nanopore sequencing process of FIG. 4.

[0026] FIG. 6 illustrates an example embodiment of the variable region annotation method of the present disclosure.

[0027] FIG. 7 illustrates an embodiment of a comparative T-cell receptor enrichment to link clonotypes by sequencing (TREKseq) method.

[0028] FIG. 8 illustrates another embodiment of the comparative TREKseq method.

[0029] FIG. 9 illustrates an embodiment of the comparative TREKseq data processing pipeline.

[0030] FIG. IDA illustrates an example 3' TCR read structure generated by the comparative TREKseq method.

[0031] FIG. 10B illustrates an example 3' TCR read structure generated by the nanopore sequencing method of the present disclosure.

[0032] FIG. 11 illustrates an embodiment of the cell barcode identification step of FIG. 6.

[0033] FIG. 12A illustrates a first step of an embodiment of the cell barcode correction step of FIG. 6.

[0034] FIG. 12B illustrates a second step of an embodiment of the cell barcode correction step of FIG. 6. FIG. 12C illustrates a third step of an embodiment of the cell barcode correction step of FIG. 6.

[0035] FIG. 13 illustrates an embodiment of the consensus building and contig correction step of FIG. 6.

[0036] FIG. 14(a)-(b) illustrate variable region capture achieved by an embodiment of the variable region annotation method of the present disclosure as compared to the comparative TREKseq method.

[0037] FIG. 15 illustrates the top 20 V gene annotations generated by the comparative TREKseq method.

[0038] FIG. 16 illustrates barcode region capture achieved by the variable region annotation method of the disclosure, as compared to the comparative TREKseq method.

[0039] FIG. 17 illustrates the recovery rate of a / (3 T-cell droplets by the variable region annotation method of the present disclosure when compared to a GEX library.

[0040] FIG. 18 illustrates full-length TCR sequencing of Curio Seeker and lOx Visium cDNA with 3' spatial barcodes.

[0041] FIG. 19A illustrates an example 3' TCR read structure generated by the nanopore sequencing method of the present disclosure.

[0042] FIG. 19B illustrates an example embodiment of the variable region annotation method of the present disclosure applied to the 3' TCR read structure of FIG. 19A.

[0043] FIG. 20 illustrates detection levels of productive TCRs achieved by the variable region annotation method of FIG. 19B.

[0044] FIG. 21 illustrates detected TCRs co-localized with T-cell marker genes.

[0045] FIG. 22 illustrates raw reads from lOx 3' GEX cDNA that do not include identifiable barcodes.

[0046] FIG. 23 illustrates raw reads from Curio Seeker cDNA that do not include identifiable barcodes.

[0047] FIG. 24 illustrates removal of TSO-TSO (Template Switch Oligos) artefacts after TCR capture in an embodiment of the method of the disclosure.

[0048] The Figures depict preferred embodiments for purposes of illustration only. One skilled in the art will readily recognize from the following discussion that alternative embodiments of the systems and methods illustrated herein may be employed without departing from the principles of the invention described herein.

[0049] Description Of Embodiments

[0050] The present embodiments may relate to, inter alia, systems and methods for long- read sequencing of immune cell receptors. For example, in some embodiments, the method comprises generating a 3'-barcoded cDNA library comprising an immune receptor target sequence, wherein the immune receptor target sequence comprises a variable region and a barcoded region; amplifying the immune receptor target sequence; linking the variable region of the immune receptor target sequence to a probe to generate a labeled immune receptor target sequence; capturing the labeled immune receptor target sequence; amplifying the captured immune receptor target sequence; passing the captured immune receptor target sequence through a long-read sequencing device; and generating, by way of the long-read sequencing device, a full-length immune receptor target sequence.

[0051] In other embodiments, the method further comprises a computer-implemented method of annotating a variable region of an immune receptor, comprising receiving, by way of a processor, the full-length immune receptor target sequence generated by the long-read sequencing method of the disclosure; removing, by way of the processor, unwanted sequences from the full-length immune receptor target sequence; identifying, by way of the processor, the barcode of the full-length immune receptor target sequence; comparing, by way of the processor, the identified barcode with a reference barcode; wherein the comparing comprises aligning the identified barcode with the reference barcode; and determining whether the identified barcode originated from the reference barcode; annotating, by way of the processor, the variable region of the full-length immune receptor target sequence; and correcting, by way of the processor, the variable region of the full-length immune receptor target sequence. The systems and methods described herein may include additional, less, or alternate functionality, including that discussed elsewhere herein.

[0052] "Long-read sequencing," as used herein, may refer to a DNA sequencing technique that allows for the generation of longer DNA sequence reads compared to short-read sequencing methods, such as Illumina. In various aspects, long-read sequencing may produce DNA sequence reads of greater than 300 base pairs, or greater than 1000 base pairs, or greater than 10000 base pairs.

[0053] "3'-barcoding," or "3'-barcoded," as used herein, may refer to the addition of a unique barcode sequence to the 3' end of an RNA molecule or a cDNA molecule within a cell before sequencing using a single-cell RNA-sequencing (scRNA-seq) or other high-throughput sequencing method. In various aspects, the barcode may include a molecular tag.

[0054] "Immune receptor", as used herein, may refer to any immune receptor comprising a variable region. In some aspects, the immune receptor may be a B cell receptor (BCR) or a T cell receptor (TCR). "Variable region," as used herein, refers to a specific region of a BCR or TCR that exhibits variability in its amino acid or nucleotide sequence.

[0055] "Target sequence," as used herein, may refer to any sequence on an immune receptor that includes a variable region.

[0056] "Amplifying," as used herein, may refer to any method by which at least a part of a target nucleic acid sequence, is reproduced, typically in a template-dependent manner, including without limitation, a broad range of techniques for amplifying nucleic acid sequences, either linearly or exponentially. Example means for performing an amplifying step include ligase chain reaction (LCR), ligase detection reaction (LDR), ligation followed by Q- replicase amplification, PCR, primer extension, strand displacement amplification (SDA), hyperbranched strand displacement amplification, multiple displacement amplification (MDA), nucleic acid strand-based amplification (NASBA), two-step multiplexed amplifications, rolling circle amplification (RCA) and the like, including multiplex versions or combinations thereof, for example but not limited to, OLA / PCR, PCR / OLA, LDR / PCR, PCR / PCR / LDR, PCR / LDR, LCR / PCR, PCR / LCR (also known as combined chain reaction-CCR), and the like.

[0057] "Probe," as used herein, may refer generally to a capture probe such as a biotin probe and the like.

[0058] "Nanopore sequencing," as used herein, may refer to a method that determines the sequence of a polynucleotide with the aid of a nanopore. In some embodiments, the sequence of the polynucleotide is determined in a template-dependent manner.

[0059] "Full-length" immune receptor target sequences, as used herein, may refer to the target sequence produced after passing a captured immune receptor target sequence through a long-read sequencing device of the disclosure. "Full-length" immune receptor target sequences, may also refer to "full-length" paired immune receptor target sequences. "Full-length" paired immune receptor target sequences may include paired TCR sequences including the complete genetic or nucleotide sequences of both the alpha (a) and beta (P) chains of a TCR or the complete genetic or nucleotide sequences of both the gamma (y) and delta (5) chains of a TCR. "Full-length" paired immune receptor sequences may also include paired BCR sequences including the complete genetic or nucleotide sequences of both the heavy chain and light chain of a BCR.

[0060] "Annotating," as used herein, refers to the process of adding descriptive information or metadata to biological sequences, such as DNA, RNA or protein sequences.

[0061] The technical effect of the systems and methods described herein may be achieved by performing at least one of the following steps: generating a 3'-barcoded cDNA library comprising an immune receptor target sequence, wherein the immune receptor target sequence comprises a variable region and a barcoded region, amplifyin the immune receptor target sequence; linking the variable region of the immune receptor target sequence to a probe to generate a labeled immune receptor target sequence; capturing the labeled immune receptor target sequence; amplifying the captured immune receptor target sequence; passing the captured immune receptor target sequence through a long-read sequencing device; generating, by way of the long-read sequencing device, a full-length immune receptor target sequence; receiving, by way of a processor, the full-length immune receptor target sequence; removing by way of the processor, unwanted sequences from the full-length immune receptor target sequence; identifying, by way of the processor, the barcode of the full-length immune receptor target sequence; comparing, by way of the processor, the identified barcode with a reference barcode; wherein the comparing comprises aligning the identified barcode with the reference barcode, and determining whether the identified barcode originated from the reference barcode; annotating, by way of the processor, the variable region of the full-length immune receptor target sequence; and correcting, by way of the processor, the variable region of the full-length immune receptor target sequence.

[0062] At least one of the technical problems addressed by the systems and methods disclosed herein may include: (i) challenges in obtaining TCR or BCR information from 3' barcoding chemistries; (ii) incomplete capture of full-length immune receptor sequences; (iii) inability to capture allele level variable region annotation.

[0063] The resulting technical effects may include, for example: (i) single cell immune receptor sequencing of samples that are outside of the lOx Genomics 5' VDJ chemistry; (ii) substantially complete capture of full-length immune receptor sequences; (iii) ability to capture allele level variable region annotation.

[0064] In various aspects, a method of long-read sequencing of an immune receptor is disclosed. In various aspects, the method comprises generating a 3'-barcoded cDNA library comprising an immune receptor target sequence, wherein the immune receptor target sequence comprises a variable region and a barcoded region, amplifyingthe immune receptor target sequence; linking the variable region of the immune receptor target sequence to a probe to generate a labeled immune receptor target sequence; capturing the labeled immune receptor target sequence; amplifying the captured immune receptor target sequence; passing the captured immune receptor target sequence through a long-read sequencing device; generating, by way of the long-read sequencing device, a full-length immune receptor target sequence.

[0065] In various aspects, the method may comprise generating a 3'-barcoded cDNA library comprising an immune receptor target sequence using any suitable method known in the art. In some aspects, the 3' barcoded cDNA library may be generated from a 3' barcoded messenger RNA (mRNA) coding for one or more immune receptors. In some aspects, the 3' barcoded cDNA library may be generated from a 3' barcoded mRNA coding for one or more of an alpha (a) chain of a TCR, a beta (3) chain of a TCR, a gamma (y) chain of a TCR, or a delta (5) chain of a TCR, or combinations thereof. In other aspects, the 3' barcoded cDNA library may be generated from 3' barcoded mRNA coding for a BCR.

[0066] In various aspects, the immune receptor target sequence may include, for example, a TCR target sequence or a BCR target sequence. In various aspects, the TCR target sequence may include an aP TCR target sequence, or a y5 TCR target sequence. In various aspects, the target sequence may include a variable region, a barcoded region, and / or a constant region. In various aspects, the barcoded region may further include a molecular identifier, such as a unique molecular identifier (UMI). "Unique molecular identifier", as used herein, refers to a molecular barcode that provides error correction and increased accuracy during sequencing.

[0067] In various aspects, the method may comprise amplifying the immune receptor target sequence using any suitable method known in the art. In various aspects, the method comprises amplifying the immune receptor target sequence and linking the variable region of the immune receptor target sequence to a probe to generate a labeled immune receptor target sequence. In various aspects, suitable probes may include any probe that may be linked to the variable region of an immune receptor to generate a labeled immune receptor target sequence. In various aspects, suitable probes may include biotin probes, antibodies, fluorescent probes, quantitative PCR (aPCR) probes, next-generation sequencing (NGS) probes, peptide probes, immunoglobulin-binding proteins, and the like.

[0068] In various aspects, the method may comprise capturingthe labeled immune receptor target sequence with any suitable capture probe. Examples of capture probes may include antibodies, DNA capture probes, protein A / G capture probes, streptavidin capture probes, phage display peptide probes, and immobilized nucleic acid capture probes.

[0069] In various aspects, the method may comprise amplifying the captured immune receptor target sequence and passing the captured immune receptor target sequence through a long-read sequence device. Suitable long-read sequencing devices may include devices utilizing nanopore sequencing technology (MinlON, GridlON, and PromethlON; Oxford Nanopore Technologies (ONT)) or utilizing single-molecule, real-time (SMRT) sequencing technology (PacBio Sequel II; Pacific Biosciences).

[0070] In various aspects, the method may further comprise, generating, by way of the long- read sequencing device, a full-length immune receptor target sequence, such as a full-length TCR target sequence, or a full-length BCR target sequence. In some aspects, the full-length immune receptor target sequence may be a full-length paired immune receptor target sequence, such as a full-length paired TCR target sequence including both the alpha (a) and beta (3) chains of the TCR or both the gamma (y) and delta (5) chains of the TCR; or a full- length paired BCR target sequence including both the heavy chain and light chain of the BCR.

[0071] In some aspects, the method may comprise the sequential steps of: (i) generating a 3'-barcoded cDNA library comprising a TCR target sequence, wherein the TCR target sequence comprises a variable region and a barcoded region; (ii) amplifying the TCR target sequence; (iii) linking the barcoded region of the TCR target sequence to a probe to generate a barcode region-labeled TCR target sequence; (iv) capturing the barcode region- labeled TCR target sequence; (v) amplifying the captured TCR target sequence of step (iv); (vi) linking the variable region of the amplified TCR target sequence of step (v) to a probe to generate a variable region-labeled TCR target sequence; (vii) capturing the variable region- labeled TCR target sequence of step (vi); (viii) amplifying the captured TCR target sequence of step (vii); (ix) passing the amplified TCR target sequence of step (viii) through a long-read sequencing device; and (x) generating, by way of the long-read sequencing device, a full- length TCR target sequence.

[0072] In some aspects, the method may comprise the sequential steps of: (i) generating a 3'-barcoded cDNA library comprising a TCR target sequence, wherein the TCR target sequence comprises a variable region and a barcoded region; (ii) amplifying the TCR target sequence; (iii) linking the variable region of the amplified TCR target sequence to a probe to generate a variable region-labeled TCR target sequence; (iv) capturing the variable region- labeled TCR target sequence; (v) amplifying the captured TCR target sequence of step (iv); (vi) linking the barcoded region of the TCR target sequence of step (v) to a probe to generate a barcode region-labeled TCR target sequence; (vii) capturing the barcode region- labeled TCR target sequence of step (vi); (viii) amplifying the captured TCR target sequence of step (vii); (ix) passing the amplified TCR target sequence of step (viii) through a long-read sequencing device; and (x) generating, by way of the long-read sequencing device, a full- length TCR target sequence.

[0073] In various aspects, a computer-implemented method for annotating a variable region of a full-length immune receptor target sequence is disclosed. The computer- implemented method comprises receiving, by way of a processor, a full-length immune receptor target sequence generated from a 3'-barcoded cDNA comprising an immune receptor target sequence comprising a variable region and a barcoded region, such as those generated by the methods provided herein; removing, by way of the processor, one or more unwanted sequences from the full-length immune receptor target sequence; identifying, by way of the processor, the barcode of the full-length immune receptor target sequence; comparing, by way of the processor, the identified barcode with a reference barcode, wherein the comparing comprises aligning the identified barcode with the reference barcode, and determining whether the identified barcode originated from the reference barcode; annotating, by way of the processor, the variable region of the full-length immune receptor target sequence; and correcting, by way of the processor, the variable region of the full-length immune receptor target sequence.

[0074] In various aspects, the one or more unwanted sequences may be removed using any suitable trimming tool known in the art, such as Porechop, Cutadapt, and the like. In various aspects, unwanted sequences may include any sequence that is not needed for variable region annotation. In various aspects, unwanted sequences may include adapters, primers, untranslated regions (UTRs) and leader regions.

[0075] In various aspects, methods of identifying the barcode of the full-length immune receptor target sequence may include retrieving barcodes based on R1 matching and / or anchoring, for example, retrieving barcodes based on partial R1 matches. In some aspects, comparing the identified barcode with a reference barcode, may comprise computing the posterior probability that the observed barcode originated from a reference barcode. In various aspects, suitable reference barcodes include any barcode suitable for alignment with an identified barcode of the full-length immune receptor sequence disclosed herein. In some aspects, the reference barcode may include a whitelist barcode.

[0076] In various aspects, the variable region of the full-length immune receptor target sequence may be annotated using any suitable variable region annotation tool known in the art, such as IgBLAST. In some aspects, annotating the variable region of the full-length immune receptor target sequence may comprise including information about the CDR3 (complementarity-determining region 3) length, mutations, and other sequence characteristics.

[0077] In various aspects, correcting the variable region of the full-length immune receptor target sequence may comprise consensus building using any suitable alignment tool, such as IgBLAST, ClustalW and MAFFT; and contig correction using any suitable correction tool, such as Quiver, RACON, Redundans, GapFiller, SSPACE, HaploMerger2, LoRDEC, CORTEX, MEDAKA, or the like. In various aspects, a system comprising at least one memory and at least one processor in communication with the at least one memory is disclosed. In some aspects, the at least one processor is programmed to receive, by way of the processor, a full-length immune receptor target sequence generated from a 3' barcoded cDNA comprising an immune receptor target sequence comprising a variable region and a barcoded region, such as those generated by the methods provided herein; remove, by way of the processor, one or more unwanted sequences from the full-length immune receptor target sequence; identify, by way of the processor, the barcode of the full-length immune receptor target sequence; compare, by way of the processor, the identified barcode with a reference barcode, wherein the processor is programmed to align the identified barcode with the reference barcode, and determine whether the identified barcode originated from the reference barcode; annotate, by way of the processor, the variable region of the full-length immune receptor target sequence; and correct, by way of the processor, the variable region of the full-length immune receptor target sequence.

[0078] In various aspects, the processor may be further programmed to correct errors from sequencing, using any suitable correction tool, such as RACON, LoRDEC, Quiver, Coral, Fiona, HiTEC, and the like. In various aspects, the processor may be further programmed to correct errors from amplification using a molecular identifier using any suitable correction tool, such as UMI-tools, Alevin, Cell Ranger, scumi, and the like.

[0079] In various aspects, at least one non-transitory computer-readable storage media having computer-executable instructions embodied thereon is disclosed. In various aspects, the computer-executable instructions, when executed by at least one processor, cause the at least one processor to receive, by way of the processor, a full-length immune receptor target sequence generated from a 3' barcoded cDNA comprising an immune receptor target sequence comprising a variable region and a barcoded region, such as those generated by the methods provided herein; remove, by way of the processor, one or more unwanted sequences from the full-length immune receptor target sequence; identify, by way of the processor, the barcode of the full-length immune receptor target sequence; compare, by way of the processor, the identified barcode with a reference barcode, wherein the processor is programmed to align the identified barcode with the reference barcode, and determine whether the identified barcode originated from the reference barcode; annotate, by way of the processor, the variable region of the full-length immune receptor target sequence; and correct, by way of the processor, the variable region of the full-length immune receptor target sequence.

[0080] As used herein, a processor may include any programmable system including systems using micro-controllers, reduced instruction set circuits (RISC), application specific integrated circuits (ASICs), logic circuits, and any other circuit or processor capable of executing the functions described herein. The above examples are example only, and are thus not intended to limit in any way the definition and / or meaning of the term "processor."

[0081] In some aspects, a computer program is provided, and the program is embodied on a computer readable medium. In an example embodiment, the system is executed on a single computer system, without requiring a connection to a server computer. In a further embodiment, the system is being run in a Windows® environment (Windows is a registered trademark of Microsoft Corporation, Redmond, Washington). In yet another embodiment, the system is run on a mainframe environment and a UNIX® server environment (UNIX is a registered trademark of X / Open Company Limited located in Reading, Berkshire, United Kingdom). The application is flexible and designed to run in various different environments without compromising any major functionality.

[0082] Additionally, the computer systems discussed herein may include additional, less, or alternate functionality, including that discussed elsewhere herein. The computer systems discussed herein may include or be implemented via computer-executable instructions stored on non-transitory computer-readable media or medium.

[0083] In some aspects, the system includes multiple components distributed among a plurality of computing devices. One or more components may be in the form of computerexecutable instructions embodied in a computer-readable medium. The systems and processes are not limited to the specific embodiments described herein. In addition, components of each system and each process can be practiced independent and separate from other components and processes described herein. Each component and process can also be used in combination with other assembly packages and processes. The present embodiments may enhance the functionality and functioning of computers and / or computer systems.

[0084] Examples

[0085] Nanopore sequencing of a single cell TCR

[0086] FIG. 3 illustrates an example nanopore sequencing of a single cell TCR. 3' barcoded cDNA were first generated from mRNA coding for the a and P chain of a TCR and then amplified to produce full-length TCRa and TCRp 3'-barcoded cDNA. Each of the TCRa and TCRp 3'-barcoded cDNA included a variable region (VJ), a constant region (Ca, Cp) and a barcoded region (BC). The variable regions of the TCRa and TCRp 3'-barcoded cDNA were linked to biotin probes to produce biotinylated TCRaand TCRp 3'-barcoded cDNA. The biotinylated TCRaand TCRp 3'-barcoded cDNA were captured using streptavidin beads in a target enrichment step and the enriched TCRaand TCRp 3'-barcoded cDNA were amplified using PCR. The amplified TCRa and TCRp 3'-barcoded cDNA were then passed through a nanopore sequencing device (Oxford Nanopore) utilizing a neural network based base caller (Guppy or Dorado), as illustrated in FIGS. 4 and 6, to form the full-length TCR target sequence (hereinafter, the ONT Read Structure) as shown in FIG. 5. More specifically, the cDNA was passed through a protein nanopore. As each base passed through the nanopore, it caused a change in an electrical current, which was detected and recorded. The base caller, e.g., Guppy, then used a neural network to interpret these changes in current and determine the sequence of bases in the cDNA to form the ONT read structure. The ONT read structure was then further processed using the computational steps set forth in FIG. 6.

[0087] More specifically, as shown in FIG. 6, the ONT read structure of FIG. 5 was subjected to a trimming step conducted using Porechop to find and remove adapters. Cutadapt was then used to find and remove unwanted sequences, such as the R2, UTR and leader regions shown in FIG. 5, in order to match anchor reads.

[0088] Cell barcode identification was conducted by first retrieving barcodes based on partial R1 (13bp) matches. Cell barcode identification steps are illustrated in FIG. 11, wherein the raw reads with adapters trimmed (~5.4 M reads) were first retrieved. Barcode identification involved retrieving barcodes based on partial R1 reads (13 bp) with at most one mismatch used as an anchor (~1.35 M reads). Stringent partial R1 match was used as anchor to ensure good quality reads. Cell barcode (16bp) and UMI (12 bp) were retrieved based on the anchored read (Rl) locus as depicted in FIGS. 12A-12C. Barcodes overlapping with a whitelist were then retrieved (~16K), resulting in ~1.16 M reads and barcodes with less than 3 reads (~8K) were filtered, resulting in ~1.05 M reads.

[0089] Cell barcode correction was then conducted by computing the posterior probability that the observed barcode originated from a whitelist barcode, as substantially illustrated in FIGS. 12A-12C. More specifically, in an initial step, the observed frequency in the dataset of every barcode on the whitelist was counted. For every observed barcode in the dataset that was not on the whitelist, but had a corresponding whitelist sequence that was 1 Hamming distance away, the posterior probability that the observed barcode originated from the whitelist barcode with a sequencing error at the differing base (based on the base Q score) was computed. The observed barcode was replaced with the whitelist barcode having the highest posterior probability that exceeds 0.9 to generate a corrected ONT read structure.

[0090] As shown in FIG. 6, consensus building was then conducted on the corrected ONT read structure using IgBlast. Additional correction steps were then conducted after consensus building using RACON and MEDAKA.

[0091] More specifically, as shown in FIG. 13, before correction, clone read groups to identify representative TCR chains within each droplet barcode were defined. Clones were defined based on variable gene (V-gene) and CDR3 amino acid sequences within each cell barcode. Those with the same UMI reads were collapsed and the longest read with valid annotations was selected. The read counts were summed within each clone to produce a consensus count and identify unique sequences. Sequences with a consensus count equal to 1 were filtered out to produce an output of representative sequences for each cell barcode. The representative sequences were then corrected by subjecting the sequences to two to four rounds of RACON. RACON is a sequence-graph based method wherein inputs in the form of three files: contigs in FASTA / FASTQ format, reads in FASTA / FASTQ format and overlaps / alignments in MHAP / PAF / SAM format are taken in and a set of polished contigs are outputted. After correction, a consensus TCR chain was built using Medaka, which uses neural networks to build a consensus from the cluster representative sequences. Clones were then defined among the consensus TCR chains within each cell barcode to merge matching sequences post-correction to produce an output of consensus sequences

[0092] As further shown in FIG. 6, variable region annotation was conducted after contig correction using IgBlast.

[0093] TREK-seq Validation

[0094] TREKseq (T-cell receptor enrichment to link clonotypes by sequencing) is a technique used to identify and analyze T-cell receptor sequences and their clonal relationships. TREKseq is described in more detail in Tu et al., Nat. Immunol. (2019) and Miller et al. Nat. Biotech. (2022), which are incorporated by reference herein in their entirety.

[0095] FIGS. 7 and 8 illustrate an example TREKseq method. As shown in FIGS. 7 and 8, 3' barcoded cDNA are first generated from mRNA coding for the a and 3 chain of a TCR and then amplified to produce full-length TCRaand TCRp 3'-barcoded cDNA. Each of the TCRaand TCRp 3'-barcoded cDNA include a variable region (VJ), a constant region (Ca, Cp) and a barcoded region (BC). The constant regions of the TCRaand TCRp 3'-barcoded cDNA are linked to biotin probes to produce biotinylated TCRaand TCRp 3'-barcoded cDNA. The biotinylated TCRaand TCRp 3'-barcoded cDNA are captured using streptavidin beads in a target enrichment step and the enriched TCRaand TCRp 3'-barcoded cDNA are subjected to primer extension using a pool of a or P V-gene primers (36a + 363) and further amplified using SI PCR (site-specific PCR). The amplified TCRaand TCRp 3'-barcoded cDNA are combined and sequenced using miSeq / NextSeq (Illumina).

[0096] FIG. 9 illustrates the TREKseq data processing pipeline, performed using Presto and IgBlast, which involves, in a first step, compiling the raw reads (~200 M) and filtering the low quality reads (average quality >20 and min length 120), which results in 7.5 M reads. In a subsequent step, cell barcodes are extracted (16 bp) and matched with a white list, which results in ~700k reads. The UMI (12 bp) is extracted for each barcode and the reads with the same UMI are collapsed into one consensus read to produce unique sequences. The sequences with greater than or equal to two reads supporting them are maintained and the variable regions of these sequences are annotated.

[0097] FIG. 10A provides an illustration of a TREKseq read structure, while FIG. 10B provides an illustration of a read structure of the disclosure. As can be seen in these Figures, the methods of the disclosure are able to provide a read structure having a full-length VDJ region, while TREKseq provides a read structure having only a partial VDJ read. FIG. 14(a)-(b) provides an illustration of the full-length TCR capture by the methods of the disclosure, whereas TREKseq only captures the FWR3, CDR3, and J regions of the TCR. As further illustrated in FIG.

[0098] 5 15, TREKseq reads are not able to provide allele level variable region annotation.

[0099] Additionally, as shown in FIG. 16, as well as Table 2 below, the full-length sequencing methods of the disclosure exhibit improved sensitivity in capturing barcodes with productive a / 3 TCRs compared to TREKseq. 0 Table 2. Chain pairing of ONT and TREKseq droplets

[0100] Greater than 90% consensus was also found between the ONT and TREKseq VDJ annotations, as shown in Table 3 below: 5 Table 3. Comparison between ONT and TREKseq VDJ annotations

[0101]

[0102] The number of droplets with a productive receptor sequence included 4669 in the ONT dataset and 2478 in the TREKseq dataset. 2325 common barcodes (93%) were uncovered. As shown in FIG. 17 and Table 4 below, ONT recovered ~ 80% of a / 3 T cell droplets when compared to a GEX library. GEX data included 4369 a / (3 T cells which were compared to 4669 ONT droplets. The percentage of common droplets between GEX and ONT was 78% (3418). Table 4. Chain pairing of matched and unmatched ONT droplets with GEX library

[0103] In contrast, common droplets between GEX and TREKseq numbered only 50% (2251).

[0104] Curio Seeker (Spatial) ONT TCR Full-length TCR sequencing of cDNA with 3'spatial barcodes is shown in FIG. 18 and was observed to allow spatial tracking of T cell clones. The pipeline of FIG. 19B, when applied to the 3' TCR read structure of FIG. 19A retrieved 17% (848k / 4.9M) reads with valid cell barcodes. As shown in FIG. 20 and Table 5 below, about 2500 out of 16000 beads had productive TCRs. TCR beta chains were primarily detected.

[0105] Table 5. Chain Pairing of ONT Curio Seeker beads

[0106] * Multichains can mean multiple T cells on the bead Detected TCRs were observed to co-localize with T cell marker genes as shown in FIG. 21. Protocol to remove TSO-TSO artefacts

[0107] As substantially illustrated in FIG. 22, it was found that ~70% of ONT raw reads from lOx 3' GEX cDNA did not have identifiable barcodes. FIG. 23 shows that CurioSeeker cDNA also contained similar artefacts (no barcode). It was further found that TSO-TSO and UPS-UPS artefacts made up a large fraction of discarded reads, as shown in Table 6 below: Table 6

[0108] An example protocol to remove the TSO-TSO artefacts is illustrated in FIG. 24. As shown in FIG. 24, during generation of the cDNA, barcoded cDNA and TSO-TSO artefacts were generated. Hybrid capture for separate TRAV and TRBV enrichment was conducted and PCR was used to amplify the enriched cDNA. Amplified barcoded cDNA was linked to a biotin probe and pulled down. ONT adapters were then added by ligation to the PCR amplified pulled-down barcoded cDNA. As used herein, an element or step recited in the singular and preceded by the word "a" or "an" should be understood as not excluding plural elements or steps, unless such exclusion is explicitly recited. Furthermore, references to "example embodiment" or "one embodiment" of the present disclosure are not intended to be interpreted as excluding the existence of additional embodiments that also incorporate the recited features.

[0109] The patent claims at the end of this document are not intended to be construed under 35 U.S.C. § 112(f) unless traditional means-plus-function language is expressly recited, such as "means for" or "step for" language being expressly recited in the claim(s). This written description uses examples to disclose the disclosure, including the best mode, and also to enable any person skilled in the art to practice the disclosure, including making and using any devices or systems and performing any incorporated methods. The patentable scope of the disclosure is defined by the claims, and may include other examples that occur to those skilled in the art. Such other examples are intended to be within the scope of the claims if they have structural elements that do not differ from the literal language of the claims, or if they include equivalent structural elements with insubstantial differences from the literal language of the claims.

Claims

What Is Claimed Is:

1. A method of long-read sequencing of a T-cell receptor (TCR) comprising: generating a 3'-barcoded cDNA library comprising a TCR target sequence, wherein the TCR target sequence comprises a variable region and a barcoded region; amplifying the TCR target sequence; linking the variable region of the TCR target sequence to a probe to generate a labeled TCR target sequence; capturing the labeled TCR target sequence; amplifying the captured TCR target sequence; passing the captured TCR target sequence through a nanopore sequencing device; and generating, by way of the nanopore sequencing device, a full-length TCR target sequence.

2. The method of claim 1, wherein the TCR target sequence further comprises a constant region.

3. The method of claim 1 or 2, wherein the barcoded region further comprises a molecular identifier.

4. The method of any one of claims 1-3, wherein the TCR comprises an a£ TCR.

5. The method of any one of claims 1-4, wherein the method comprises linking the variable region of the TCR target sequence to a capture probe to generate the labeled TCR target sequence.

6. The method of claim 5, wherein the method comprises capturing the labeled TCR target sequence using one or more capture probes.

7. The method of any one of claims 1-6, wherein the method comprises linking the variable region of the TCR target sequence to a biotin probe to generate a biotin labeled TCR target sequence.

8. The method of claim 7, wherein the method comprises capturing the biotin labeled TCR target sequence using one or more streptavidin beads.

9. The method of claim 1, wherein the method comprises the sequential steps of:(i) generating a 3'-barcoded cDNA library comprising a TCR target sequence, wherein the TCR target sequence comprises a variable region and a barcoded region;(ii) amplifying the TCR target sequence;(iii) linking the barcoded region of the TCR target sequence to a probe to generate a barcode region-labeled TCR target sequence;(iv) capturing the barcode region-labeled TCR target sequence;(v) amplifying the captured TCR target sequence of step (iv);(vi) linking the variable region of the amplified TCR target sequence of step (v) to a probe to generate a variable region-labeled TCR target sequence;(vii) capturing the variable region-labeled TCR target sequence of step (vi);(viii) amplifying the captured TCR target sequence of step (vii);(ix) passing the amplified TCR target sequence of step (viii) through a nanopore sequencing device; and(x) generating, by way of the nanopore sequencing device, a full-length TCR target sequence.

10. The method of claim 1, wherein the method comprises the sequential steps of: (i) generating a 3'-barcoded cDNA library comprising a TCR target sequence, wherein the TCR target sequence comprises a variable region and a barcoded region;(ii) amplifying the TCR target sequence;(iii) linking the variable region of the amplified TCR target sequence to a probe to generate a variable region-labeled TCR target sequence;(iv) capturing the variable region-labeled TCR target sequence;(v) amplifying the captured TCR target sequence of step (iv);(vi) linking the barcoded region of the TCR target sequence of step (v) to a probe to generate a barcode region-labeled TCR target sequence;(vii) capturing the barcode region-labeled TCR target sequence of step (vi);(viii) amplifying the captured TCR target sequence of step (vii);(ix) passing the amplified TCR target sequence of step (viii) through a nanopore sequencing device; and(x) generating, by way of the nanopore sequencing device, a full-length TCR target sequence.

11. A system comprising at least one memory and at least one processor in communication with the at least one memory, wherein the at least one processor is programmed to: receive, by way of the processor, the full-length TCR target sequence of any one of claims 1-10; remove, by way of the processor, one or more unwanted sequences from the full- length TCR target sequence; identify, by way of the processor, the barcode of the full-length TCR target sequence; compare, by way of the processor, the identified barcode with a reference barcode, wherein the processor is programmed to align the identified barcode with the reference barcode, and determine whether the identified barcode originated from the reference barcode; annotate, by way of the processor, the variable region of the full-length TCR target sequence; andcorrect, by way of the processor, the variable region of the full-length TCR target sequence.

12. The system of claim 11, wherein the one or more unwanted sequences comprise adapters, primers, untranslated regions (UTRs) and leader regions.

13. The system of claim 11 or 12, wherein the reference barcode is a whitelist barcode.

14. The system of any one of claims 11-13, wherein the processor is further programmed to correct errors from sequencing.

15. The system of any one of claims 11-14, wherein the processor is further programmed to correct errors from amplification using a molecular identifier.

16. A computer-implemented method of annotating the variable region of a full- length TCR target sequence comprising: receiving, by way of a processor, the full-length TCR target sequence of any one of claims 1-10; removing, by way of the processor, one or more unwanted sequences from the full- length TCR target sequence; identifying, by way of the processor, the barcode of the full-length TCR target sequence; comparing, by way of the processor, the identified barcode with a reference barcode, wherein the comparing comprises aligning the identified barcode with the reference barcode, and determining whether the identified barcode originated from the reference barcode; annotating, by way of the processor, the variable region of the full-length TCR target sequence; andcorrecting, by way of the processor, the variable region of the full-length TCR target sequence.

17. The computer-implemented method of claim 16, wherein the one or more unwanted sequences comprise adapters, primers, untranslated regions (UTRs) and leader regions.

18. The computer-implemented method of claim 16 or 17, wherein the reference barcode is a whitelist barcode.

19. The computer-implemented method of any one of claims 16-18, wherein the method comprises correcting errors from sequencing.

20. The computer-implemented method of any one of claims 16-19, wherein the method comprises correcting errors from amplification using a molecular identifier.

21. At least one non-transitory computer-readable storage media having computer-executable instructions embodied thereon, wherein when executed by at least one processor, the computer-executable instructions cause the at least one processor to: receive, by way of the processor, the full-length TCR target sequence of any one of claims 1-10; remove, by way of the processor, one or more unwanted sequences from the full- length TCR target sequence; identify, by way of the processor, the barcode of the full-length TCR target sequence; compare, by way of the processor, the identified barcode with a reference barcode, wherein the processor is programmed to align the identified barcode with the reference barcode, and determine whether the identified barcode originated from the reference barcode;annotate, by way of the processor, the variable region of the full-length TCR target sequence; and correct, by way of the processor, the variable region of the full-length TCR target sequence.

22. A computer-implemented method of annotating the variable region of a full- length TCR target sequence comprising: receiving, by way of a processor, a full-length TCR target sequence generated from a 3'-barcoded cDNA comprising a TCR target sequence comprising a variable region and a barcoded region; removing, by way of the processor, one or more unwanted sequences from the full- length TCR target sequence; identifying, by way of the processor, the barcode of the full-length TCR target sequence; comparing, by way of the processor, the identified barcode with a reference barcode, wherein the comparing comprises aligning the identified barcode with the reference barcode, and determining whether the identified barcode originated from the reference barcode; annotating, by way of the processor, the variable region of the full-length TCR target sequence; and correcting, by way of the processor, the variable region of the full-length TCR target sequence.

23. The computer-implemented method of claim 22, wherein the full-length TCR target sequence is generated by generating a 3'-barcoded cDNA library comprising a TCR target sequence, wherein the TCR target sequence comprises a variable region and a barcoded region; amplifying the 3'-barcoded cDNA comprising the TCR target sequence;linking the variable region of the TCR target sequence to a probe to generate a labeled TCR target sequence; capturing the labeled TCR target sequence; amplifying the captured TCR target sequence; passing the captured TCR target sequence through a nanopore sequencing device; and generating, by way of the nanopore sequencing device, a full-length TCR target sequence.

24. The computer-implemented method of claim 22 or 23, wherein the TCR target sequence further comprises a constant region.

25. The computer-implemented method of any one of claim 23 or 24, wherein the method comprises linking the variable region of the TCR target sequence to a capture probe to generate the labeled TCR target sequence.

26. The computer-implemented method of claim 23, wherein the method comprises capturing the labeled TCR target sequence using one or more capture probes.

27. The computer-implemented method of claim 26, wherein the method comprises linking the variable region of the TCR target sequence to a biotin probe to generate a biotin labeled TCR target sequence.

28. The computer-implemented method of claim 27 , wherein the method comprises capturing the biotin labeled TCR target sequence using one or more streptavidin beads.

29. The computer-implemented method of claim 23, wherein the full-length TCR target sequence is generated by(i) generating a 3'-barcoded cDNA library comprising a TCR target sequence, wherein the TCR target sequence comprises a variable region and a barcoded region;(ii) amplifying the TCR target sequence;(iii) linking the barcoded region of the TCR target sequence to a probe to generate a barcode region-labeled TCR target sequence;(iv) capturing the barcode region-labeled TCR target sequence;(v) amplifying the captured TCR target sequence of step (iv);(vi) linking the variable region of the amplified TCR target sequence of step (v) to a probe to generate a variable region-labeled TCR target sequence;(vii) capturing the variable region-labeled TCR target sequence of step (vi);(viii) amplifying the captured TCR target sequence of step (vii);(ix) passing the amplified TCR target sequence of step (viii) through a nanopore sequencing device; and(x) generating, by way of the nanopore sequencing device, a full-length TCR target sequence.

30. The computer-implemented method of claim 23, wherein the full-length TCR target sequence is generated by(i) generating a 3'-barcoded cDNA library comprising a TCR target sequence, wherein the TCR target sequence comprises a variable region and a barcoded region;(ii) amplifying the TCR target sequence;(iii) linking the variable region of the amplified TCR target sequence to a probe to generate a variable region-labeled TCR target sequence;(iv) capturing the variable region-labeled TCR target sequence;(v) amplifying the captured TCR target sequence of step (iv);(vi) linking the barcoded region of the TCR target sequence of step (v) to a probe to generate a barcode region-labeled TCR target sequence;(vii) capturing the barcode region-labeled TCR target sequence of step (vi);(viii) amplifying the captured TCR target sequence of step (vii);(ix) passing the amplified TCR target sequence of step (viii) through a nanopore sequencing device; and(x) generating, by way of the nanopore sequencing device, a full-length TCR target sequence.

31. A system comprising at least one memory and at least one processor in communication with the at least one memory, wherein the at least one processor is programmed to: receive, by way of the processor, a full-length TCR target sequence generated from a 3'-barcoded cDNA comprising a TCR target sequence comprising a variable region and a barcoded region; remove, by way of the processor, one or more unwanted sequences from the full- length TCR target sequence; identify, by way of the processor, the barcode of the full-length TCR target sequence; compare, by way of the processor, the identified barcode with a reference barcode, wherein the processor is programmed to align the identified barcode with the reference barcode, and determine whether the identified barcode originated from the reference barcode; annotate, by way of the processor, the variable region of the full-length TCR target sequence; and correct, by way of the processor, the variable region of the full-length TCR target sequence.

32. The system of claim 31, wherein the full-length TCR target sequence is generated by generating a 3'-barcoded cDNA library comprising a TCR target sequence, wherein the TCR target sequence comprises a variable region and a barcoded region amplifying the TCR target sequence;linking the variable region of the TCR target sequence to a probe to generate a labeled TCR target sequence; capturing the labeled TCR target sequence; amplifying the captured TCR target sequence; passing the captured TCR target sequence through a nanopore sequencing device; and generating, by way of the nanopore sequencing device, a full-length TCR target sequence.

33. The system of claim 31 or 32, wherein the TCR target sequence further comprises a constant region.

34. The system of claim 32, wherein the linking comprises binding the variable region of the TCR target sequence to a capture probe to generate a labeled TCR target sequence.

35. The system of claim 34, wherein the capturing comprises binding the labeled TCR target sequence using one or more capture probes.

36. The system of claim 34, wherein the linking comprises binding the variable region of the TCR target sequence to a biotin probe to generate a biotin labeled TCR target sequence.

37. The system of claim 36, wherein the capturing comprises binding the biotin labeled TCR target sequence using one or more streptavidin beads.

38. The system of claim 32, wherein the full-length TCR target sequence is generated by(i) generating a 3'-barcoded cDNA library comprising a TCR target sequence, wherein the TCR target sequence comprises a variable region and a barcoded region;(ii) amplifying the TCR target sequence;(iii) linking the barcoded region of the TCR target sequence to a probe to generate a barcode region-labeled TCR target sequence;(iv) capturing the barcode region-labeled TCR target sequence;(v) amplifying the captured TCR target sequence of step (iv);(vi) linking the variable region of the amplified TCR target sequence of step (v) to a probe to generate a variable region-labeled TCR target sequence;(vii) capturing the variable region-labeled TCR target sequence of step (vi);(viii) amplifying the captured TCR target sequence of step (vii);(ix) passing the amplified TCR target sequence of step (viii) through a nanopore sequencing device; and(x) generating, by way of the nanopore sequencing device, a full-length TCR target sequence.

39. The system of claim 32, wherein the full-length TCR target sequence is generated by(i) generating a 3'-barcoded cDNA library comprising a TCR target sequence, wherein the TCR target sequence comprises a variable region and a barcoded region;(ii) amplifying the TCR target sequence;(iii) linking the variable region of the amplified TCR target sequence to a probe to generate a variable region-labeled TCR target sequence;(iv) capturing the variable region-labeled TCR target sequence;(v) amplifying the captured TCR target sequence of step (iv);(vi) linking the barcoded region of the TCR target sequence of step (v) to a probe to generate a barcode region-labeled TCR target sequence;(vii) capturing the barcode region-labeled TCR target sequence of step (vi);(viii) amplifying the captured TCR target sequence of step (vii);(ix) passing the amplified TCR target sequence of step (viii) through a nanopore sequencing device; and(x) generating, by way of the nanopore sequencing device, a full-length TCR target sequence.

40. At least one non-transitory computer-readable storage media having computer-executable instructions embodied thereon, wherein when executed by at least one processor, the computer-executable instructions cause the at least one processor to: receive, by way of the processor, a full-length TCR target sequence generated from a 3'-barcoded cDNA comprising a TCR target sequence comprising a variable region and a barcoded region; remove, by way of the processor, one or more unwanted sequences from the full- length TCR target sequence; identify, by way of the processor, the barcode of the full-length TCR target sequence; compare, by way of the processor, the identified barcode with a reference barcode, wherein the processor is programmed to align the identified barcode with the reference barcode, and determine whether the identified barcode originated from the reference barcode; annotate, by way of the processor, the variable region of the full-length TCR target sequence; and correct, by way of the processor, the variable region of the full-length TCR target sequence.

41. The non-transitory computer-readable storage media of claim 40, wherein the full-length TCR target sequence is generated by generating a 3'-barcoded cDNA library comprising a TCR target sequence, wherein the TCR target sequence comprises a variable region and a barcoded region; amplifying the TCR target sequence; linking the variable region of the TCR target sequence to a probe to generate a labeled TCR target sequence;capturing the labeled TCR target sequence; a m lifyi ng the captured TCR target sequence; passing the captured TCR target sequence through a nanopore sequencing device; and generating, by way of the nanopore sequencing device, a full-length TCR target sequence.

Citation Information

Patent Citations

  • Using random priming to obtain full-length v(d)j information for immune repertoire sequencing

    US20210139970A1

  • Phenotypic and molecular characterisation of single cells

    US20210317522A1

  • Single cell sequencing libraries of genomic transcript regions of interest in proximity to barcodes, and genotyping of said libraries

    US20230348899A1