TCR / BCR profiling
The method of RNA isolation and hybrid capture probe enrichment for TCR/BCR genes in sequencing addresses the challenge of profiling TCR/BCR repertoires, enabling precise immune response analysis and targeted treatments.
Patent Information
- Application Number
- JP2022564450
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-04-08
- Filing Date
- 2021-04-21
- Publication Date
- 2025-08-15
AI Technical Summary
Existing methods are inadequate for accurately profiling T cell receptor (TCR) and B cell receptor (BCR) repertoires, which limits effective immune response profiling and treatment strategies, particularly in conditions like infectious diseases and cancers.
A method involving RNA isolation, hybrid capture probe enrichment for TCR/BCR genes, and whole-transcriptome sequencing to determine TCR/BCR profiles, allowing for disease diagnosis, treatment, and predicting treatment outcomes by analyzing TCR/BCR clones and sequences.
Enables precise profiling of TCR/BCR repertoires, facilitating disease diagnosis, treatment, and predicting treatment responses, including identifying SARS-CoV-2 exposure and lymphocytic infiltration, and enabling therapies like CAR-T cell treatments.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] Cross-reference to related patent applications This application claims the benefit of priority under 35 U.S.C. §119(e) of U.S. Provisional Application No. 63 / 013,130, filed April 21, 2020, U.S. Provisional Application No. 63 / 084,459, filed September 27, 2020, and U.S. Provisional Application No. 63 / 201,020, filed April 8, 2021. The contents of each provisional application are incorporated herein by reference in their entirety.
[0002] The present disclosure relates to systems, methods, and compositions useful for profiling T cell receptor (TCR) and B cell receptor (BCR) repertoires using next-generation sequencing (NGS) methods. The present disclosure also relates to systems and methods for diagnosing, treating, or predicting infection, disease, condition, or treatment outcome or effectiveness based on TCR / BCR profile data of a subject in need thereof. In some embodiments, the method includes detecting SARS-CoV-2 exposure. [Background technology]
[0003] The vertebrate immune system consists of two major parts: the innate part and the adaptive part. The innate part of the immune system has evolved to respond quickly and effectively to foreign antigens or danger signals. However, in many cases, the innate immune response is insufficient to provide sterilizing immunity. Additionally, the adaptive part of the immune system lacks the capacity for "memory" and is unable to mount a further effective response against a pathogen upon subsequent challenge with the same or similar pathogen. Thus, the innate part of the immune system (and / or non-immune cells, e.g., infected cells) present antigens to the adaptive immune system, which can then initiate the process of selection of antigen-specific immune cells, T lymphocytes (T cells) and B lymphocytes (B cells). This process is facilitated by the presence of a large variety of antigen-specific cells that are more likely to respond to antigenic challenge. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] U.S. Patent Application No. 16 / 657,804 [Patent Document 2] U.S. Patent Application No. 17 / 112,877 [Patent Document 3] Published Patent Application No. 16 / 533,676 [Patent Document 4] U.S. Patent Application Serial No. 16 / 927,976 [Patent Document 5] U.S. Patent Application Serial No. 16 / 789,288 [Patent Document 6] U.S. Patent Application Serial No. 15 / 930,234 [Patent Document 7] U.S. Patent Application No. 17 / 706,704 [Patent Document 8] U.S. Patent Application No. 16 / 581,706 [Patent Document 9] U.S. Patent Application Serial No. 16 / 732,229 [Patent Document 10] PCT / US19 / 69161 [Patent Document 11] U.S. Patent Application No. 17 / 074,984 [Patent Document 12] U.S. Patent Application No. 16 / 789,413 [Patent Document 13] U.S. Patent Application No. 16 / 888,357 [Patent Document 14] U.S. Patent Application Serial No. 16 / 830,186 [Patent Document 15] U.S. Patent Application No. 16 / 789,363 [Patent Document 16] PCT US20 / 18002 issue [Patent Document 17] U.S. Patent Application No. 16 / 994,315 [Patent Document 18] U.S. Patent Application No. 16 / 533,676 [Patent Document 19] U.S. Patent Application No. 62 / 804,509 [Patent Document 20] U.S. Patent Application Serial No. 16 / 653,868 [Patent Document 21] U.S. Patent Application No. 16 / 945,588 [Patent Document 22] U.S. Patent Application Serial No. 16 / 732,168 [Patent Document 23] PCT / US19 / 69149 [Patent Document 24] U.S. Patent Application Serial No. 16 / 693,117 [Patent Document 25] PCT / US20 / 56930 [Patent Document 26] U.S. Patent Application No. 17 / 114,386 [Patent Document 27] U.S. Provisional Patent Application No. 62 / 924,515 [Patent Document 28] U.S. Provisional Patent Application No. 16 / 802,126 [Patent Document 29] PCT / US21 / 18619 [Non-patent literature]
[0005] [Non-Patent Document 1] BMC Medical Genomics 12, 195 pages (2019) [Non-patent document 2] Kalendar et al., 2009, Genes, Genomes, and Genomics, Issue 3 (Special Issue 1), pp. 1-14. [Non-patent document 3] Kivioja et al., 2011, Nat. Methods 9(1), pp. 72-74 [Non-patent document 4] Islam et al., 2014, Nat. Methods 11(2), pp. 163-66 [Non-Patent Document 5] http: / / www.imgt.org [Non-patent document 6] "Landscape of tumor-infiltrating T cell repertoire of human cancers" (Li et al., 2016, Nat. Genet., 48(7), pp. 725-732) [Non-Patent Document 7] "Landscape of B cell immunity and related immune evasion in human cancers" (Hu et al., 2019, Nat. Genet., 51(3), pp. 560-567) [Non-patent document 8] "BASIC: BCR assembly from single cells" (Canzar et al., 2017, Bioinformatics, 33(3), pp. 425-427) [Non-Patent Document 9] "Simultaneously inferring T cell fate and clonality from single cell transcriptomes" (Stubbington et al., 2015, BioRxiv https: / / doi.org / 10.1101 / 025676) [Non-Patent Document 10] "Antigen receptor repertoire profiling from RNA-seq data" (Bolotin et al., 2017, Nat. Biotech., 35(10), pp. 908-911) [Non-Patent Document 11] Nicolas L Bray, Harold Pimentel, Pall Melsted and Lior Pachter, "Near-optimal probabilistic RNA-seq quantification", Nature Biotechnology, 34, pp. 525-527 (2016), doi:10.1038 / nbt.3519 [Non-Patent Document 12] https: / / pachterlab.github.io / kallisto / (カリフォルニアState Technology Research Institute (カリフォルニアPasadena))
Non-patent document 13
Non-patent document 14
Non-patent document 15
Non-patent document 16
Non-patent document 17
Non-patent document 18
Non-patent document 19
Non-patented document 25
Non-patent document 26
Non-patent document 27
Non-patent document 28
Non-patent document 29
Non-patent document 30
Non-patent document 31
[0006] Disclosed herein are methods for determining a patient's TCR / BCR profile. In some embodiments, the method includes: a) isolating RNA from a patient sample; b) enriching the isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes and enriching for a targeted whole-transcriptome panel using a collection of transcriptome hybrid capture probes; c) sequencing the RNA of (b) to generate sequencing data; and e) analyzing the sequencing data to determine the patient's TCR / BCR profile. In some embodiments, the collection of TCR / BCR hybrid capture probes includes a first collection of TCR / BCR hybrid capture probes that include a first collection of TCR / BCR constant region probes that include a BCR constant region probe. Pool , the second containing a BCR non-constant region probe Pool , a third containing a TCR constant region probe Pool , the fourth containing a TCR non-constant region probe Pool , and a fifth containing transcriptome hybrid capture probes. Pool Includes. [Means for solving the problem]
[0007] In some embodiments, the first Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool In some embodiments, the ratio of the first Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool , vs. 5th Pool The ratio is 1:2.5:100:100:10. In some embodiments, 2% or less of the reads in the sequencing data map to TCR / BCR genes. In some embodiments, the sample is a blood sample.
[0008] In some embodiments, step (d) comprises identifying multiple TCR / BCR clones in the sample, and / or identifying the most abundant TCR / BCR clone in the sample, and / or identifying the most abundant non-constant region sequence in the sample.
[0009] In some embodiments, step (c) comprises whole transcriptome sequencing or short read sequencing.
[0010] In some embodiments, the patient's BCR / TCR profile is compared to a control TCR / BCR profile, and based on the comparison, the patient is identified as having a disease or medical condition. In some embodiments, the disease or condition is an infectious disease, cancer, an autoimmune disease, or an allergy. In some embodiments, the cancer or infectious disease is one or more of those listed in embodiment 114. In some embodiments, the infectious disease includes contact with SARS-CoV-2. In some embodiments, the subject is suspected of having been diagnosed with COVID-19. In some embodiments, the disease is cancer. In some embodiments, the analysis includes determining the presence or extent of lymphocytic infiltration of the tumor. In some embodiments, the method includes treating the patient with a therapeutic agent. In some embodiments, the therapeutic agent includes an immunotherapeutic agent. In some embodiments, the immunotherapeutic agent is a vaccine. In some embodiments, the immunotherapeutic agent is a chimeric antigen receptor (CAR) T cell.
[0011] In some embodiments, methods of treating a disease or condition in a patient are provided, the methods comprising: a) isolating RNA from a patient sample; b) enriching the isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes and enriching for a targeted whole transcriptome panel using a collection of transcriptome hybrid capture probes; c) sequencing the RNA of (b) to generate sequencing data; and d) analyzing the sequencing data, wherein the analysis determines the most abundant TCR / BCR clones in the sample and optionally determines a TCR / BCR profile for the patient, wherein the collection of TCR / BCR hybrid capture probes comprises a first collection of TCR / BCR hybrid capture probes that include a first collection of TCR / BCR constant region probes that include a BCR constant region probe. Pool , the second containing a BCR non-constant region probe Pool , a third containing a TCR constant region probe Pool , the fourth containing a TCR non-constant region probe Pool , and a fifth containing transcriptome hybrid capture probes. Pool and e) treating the patient.
[0012] In some embodiments, the first Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool In some embodiments, the ratio of the first Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool , vs. 5th Pool The ratio is 1:2.5:100:100:10. In some embodiments, 2% or less of the reads in the sequencing data map to TCR / BCR genes. In some embodiments, the sample is a blood sample.
[0013] In some embodiments, treatment involves expanding the most abundant TCR / BCR clones in vitro and administering the expanded clones to the patient. In some embodiments, a single most abundant clone is expanded. In some embodiments, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, or 50 most abundant clones are expanded.
[0014] In some embodiments, step (d) comprises identifying the most abundant TCR non-constant region sequences in the sample, and the treatment administered in step (e) comprises administering a CAR-T cell therapy, wherein the CAR-T cells comprise at least one of the most abundant TCR non-constant region sequences.
[0015] In some embodiments, a method for characterizing the effect of a treatment on a patient's TCR / BCR profile is provided. In some embodiments, the method includes: a) at a first time point, i) isolating RNA from a patient sample, ii) enriching the isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes and enriching for a targeted whole transcriptome panel using a collection of transcriptome hybrid capture probes, iii) sequencing the RNA of (b) to generate sequencing data, and iv) analyzing the sequencing data to determine the patient's TCR / BCR profile; b) at a second time point, i) isolating RNA from the patient sample, ii) enriching the isolated RNA for TCR / BCR genes using a collection of transcriptome hybrid capture probes and enriching for a targeted whole transcriptome panel using a collection of transcriptome hybrid capture probes, iii) sequencing the RNA of (b) to generate sequencing data, and iv) analyzing the sequencing data to determine the patient's TCR / BCR profile; iii) sequencing the RNA of (b) to generate sequencing data; and iv) analyzing the sequencing data to determine the patient's TCR / BCR profile. And c) comparing the TCR / BCR profile determined in step (a) with the TCR / BCR profile determined in step (b) to characterize the effect of treatment on the patient's TCR / BCR profile, wherein the set of hybrid capture probes includes a first set of hybrid capture probes that includes TCR constant region probes. Pool , the second containing a TCR non-constant region probe Pool , a third containing a BCR constant region probe Pool , the fourth containing a BCR non-constant region probe Pool , and a fifth containing transcriptome hybrid capture probes. Pool Includes.
[0016] In some embodiments, the first time point is before the therapeutic agent is administered and the second time point is after the therapeutic agent is administered. In some embodiments, the first time point comprises a first time during a first course of treatment and the second time point is a second time during treatment with the first therapy or after the course of treatment with the first therapy has ended. In some embodiments, a third, fourth, fifth, or Nth time point is analyzed. One or more of the Nth time points may be time points used during a longitudinal study, e.g., a course of treatment, a clinical trial, or before, during, or after multiple treatments.
[0017] In some embodiments, the first Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool In some embodiments, the ratio of the first Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool , vs. 5th Pool The ratio is 1:2.5:100:100:10. In some embodiments, 2% or less of the reads in the sequencing data map to TCR / BCR genes. In some embodiments, the sample is a blood sample.
[0018] In some embodiments of any of the above methods, the sample comprises a blood sample or a solid tumor sample. In some embodiments of any of the above methods, step (c) comprises whole transcriptome sequencing or short read sequencing.
[0019] In some embodiments, methods are provided for determining the TCR / BCR profile of a patient with COVID-19 or another disease. In some embodiments, the method includes: a) isolating RNA from a patient sample; b) enriching the isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes and enriching for a targeted whole transcriptome panel using a collection of whole transcriptome hybrid capture probes; c) sequencing the RNA of (b) to generate sequencing data; and d) analyzing the sequencing data to determine the TCR / BCR profile of the patient, wherein the collection of TCR / BCR hybrid capture probes includes a first collection of TCR constant region probes. Pool , the second containing a TCR non-constant region probe Pool , a third containing a BCR constant region probe Pool , the fourth containing a BCR non-constant region probe Pool , and a fifth containing transcriptome hybrid capture probes. Pool Includes:
[0020] In some embodiments, the first Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool In some embodiments, the ratio of the first Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool , vs. 5th PoolThe ratio of is 1:2.5:100:100:10. In some embodiments, 2% or less of the reads in the sequencing data map to TCR / BCR genes. In some embodiments, the sample is a blood sample. In some embodiments, the patient's TCR / BCR profile is compared to a SARS-CoV-2 TCR / BCR positive control profile, and in some embodiments, it is determined whether the patient has been exposed to SARS-CoV-2. In some embodiments, if the determination suggests exposure to SARS-CoV-2, the subject is treated.
[0021] In some embodiments, a method for determining SARS-CoV-2 exposure in a patient is provided. In some embodiments, the method includes: a) isolating RNA from a patient sample; b) enriching the isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes and enriching for a targeted whole-transcriptome panel using a collection of transcriptome hybrid capture probes; c) sequencing the RNA of (b) to generate sequencing data; d) analyzing the sequencing data to determine a TCR / BCR profile of the patient; and e) comparing the TCR / BCR profile of the patient to a positive control to determine SARS-CoV-2 exposure, wherein the collection of TCR / BCR hybrid capture probes includes a first collection of TCR constant region probes. Pool , the second containing a TCR non-constant region probe Pool , a third containing a BCR constant region probe Pool , and a fourth containing a BCR non-constant region probe. Pool , as well as a fifth containing transcriptome hybrid capture probes. Pool Includes:
[0022] In some embodiments, the first Pool , versus , the second Pool , versus , the third Pool , versus 4th PoolIn some embodiments, the ratio of the first Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool , vs. 5th Pool is 1:2.5:100:100:10. In some embodiments, 2% or less of the reads in the sequencing data map to TCR / BCR genes. In some embodiments, the sample is a blood sample. In some embodiments, the patient has been in contact with or is suspected of having been in contact with SARS-CoV-2. In some embodiments, the patient is experiencing flu-like symptoms or symptoms associated with respiratory illness. In some embodiments, the method includes treating the patient for SARS-CoV-2 exposure once it is determined that the patient has been in contact with SARS-CoV-2.
[0023] In some embodiments of any of the above methods, step (c) comprises whole transcriptome sequencing or short read sequencing.
[0024] In some embodiments, methods are provided for identifying TCR / BCR non-constant region sequences enriched in a cohort of patients with SARS-CoV-2. In some embodiments, the methods include: a) isolating RNA from a sample of each patient in the cohort; b) enriching the isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes and enriching for a targeted whole-transcriptome panel using a collection of transcriptome hybrid capture probes; c) sequencing the RNA of (b) to generate sequencing data; d) analyzing the sequencing data to determine the TCR / BCR profiles of patients in the cohort; and e) identifying TCR / BCR non-constant region sequences enriched in the cohort compared to a control group without the disease or condition, wherein the collection of hybrid capture probes is a first collection of TCR constant region probes. Pool, the second of the TCR non-constant region probes Pool , the third of the BCR constant region probe Pool , and the fourth BCR non-constant region probe Pool , as well as the fifth transcriptome hybrid capture probe Pool Includes:
[0025] In some embodiments, the first Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool In some embodiments, the ratio of the first Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool , vs. 5th Pool The ratio is 1:2.5:100:100:10. In some embodiments, 2% or less of the reads in the sequencing data map to TCR / BCR genes. In some embodiments, the sample is a blood sample. [Brief explanation of the drawings]
[0026] [Figure 1A] FIG. 1 shows an exemplary TCR / BCR immune repertoire display (report) showing additional or alternative fields for review by a physician. [Figure 1B] FIG. 1 shows an example of a TCR / BCR immune repertoire display (report) showing the clonality of a patient after analysis by a novel hybrid capture approach, in this case related to BCR clonality. [Figure 2]This diagram outlines our novel hybrid capture approach to immune profiling. 1) Tumor sampling RNA is isolated from formalin-fixed, paraffin-embedded primary tumor samples. Samples exhibit widespread lymphocytic infiltration, which is highly dependent on the tissue origin of the sample. 2) TCR / BCR transcript enrichment - Specially designed and optimized hybrid capture probes pool target genes for seven lymphocyte receptors (TCR-α, TCR-β, TCR-γ, TCR-δ, Ig-heavy chain, Ig-κ, and Ig-λ) to enrich for immune receptors in the RNA-seq output without compromising downstream transcriptome analysis. 3) RNA Sequencing - Transcriptome analysis of tumor samples is performed using state-of-the-art RNA-seq platforms (for examples of RNA-seq platforms, see U.S. Patent Application No. 16 / 657,804, filed October 18, 2019, entitled "Data Based Cancer Research and Treatment Systems and Methods," and U.S. Patent Application No. 17 / 112,877, filed December 4, 2020, entitled "Systems and Methods for Automating RNA Expression Calls in a Cancer Prediction Pipeline"). Application of rep-seq probes enriches TCR / BCR reads, which do not exceed 2% of total reads for 95% of RNA-seq runs. 4) Repertoire Sequencing Analysis - RNA-seq data are processed using a rep-seq bioinformatics pipeline (in one example, the rep-seq bioinformatics pipeline includes the open-source rep-seq software TRUST4, see https: / / github.com / liulab-dfci / TRUST4). Candidate TCR / BCR reads are aligned to IMGT reference allele sequences and hypervariable complementarity-determining region 3 (CDR3) sequence clonotypes.The assignment and relative abundance of the CDR3 genes are quantified. [Figure 3] 1 is a table showing an exemplary embodiment of a novel hybrid capture approach for immune profiling by number of individual probes (right column) per generic target (left column). [Figure 4] FIG. 1 is a schematic diagram showing a probe deployment approach for enriching TCR and BCR sequences in a novel hybrid capture approach for immune profiling. [Figure 5] 1 is a histogram showing the distribution of TCR / BCR read frequencies as a proportion of total unfiltered reads in a sequencing run by using a novel hybrid capture approach for immune profiling. [Figure 6A] Figure 1 shows samples that underwent enriched RNA-based rep-seq and high-sensitivity TCR-β receptor DNA sequencing assays. Only TCR-β CDR3 nucleotide sequences were quantified and compared across runs on this basis. [Figure 6B] Figure 6 shows separate RNA-based rep-seq runs. The x-axis shows the abundance of each CDR3 nucleotide sequence in one run of data from the RNA-based rep-seq assay. The y-axis shows the abundance of each CDR3 nucleotide sequence in a second run of data from either a highly sensitive TCR-β receptor DNA sequencing assay (A) or an RNA-based rep-seq assay (B). In each example, enriched RNA-based rep-seq is considered less sensitive than the stand-alone DNA-based assay, but the RNA-based rep-seq method detects and reproduces the relative abundance of the most frequent clonotypes, even in the relatively small TCR-β repertoire shown in Figure 6. There is also high consistency for highly abundant clonotypes across assays. [Figure 7]Scatter plot showing the relationship between productive clonotypes and CDR3-supported read fragments for 501 human cancer samples analyzed using a novel hybrid capture immunoprofiling approach. Each data point represents data for one sample, and the cancer type associated with the sample is represented by a specific color / shape combination (see legend). The x-axis represents the number of productive clonotypes detected in each sample, and the y-axis represents the number of CDR3-supported sequence read fragments (sequence read fragments with portions that map to the CDR3 locus) detected. [Figure 8] Repertoires generated from 501 tumor transcriptomes demonstrate a broad distribution of clonotype richness: all productive clonotypes for BCR (Ig heavy chain, Ig-κ, and Ig-λ) and TCR (TCR-α, TCR-β, TCR-γ, and TCR-δ) (excluding CDR3 sequences with partial alignments, frameshifts, and internal stop codons). [Figure 9] Figure 1 shows that repertoires generated from 501 tumor transcriptomes demonstrate a broad distribution of clonotype richness. Gene expression-based estimates for B cells (published patent application Ser. No. 16 / 533,676, incorporated herein by reference, and PMID: 30864330) (y-axis) are correlated with clonotype yield (reads supporting productive CDR3s, x-axis) for each receptor (one-tailed Pearson - 95% CI). Samples with an infiltrate estimate of 0.001 or less are indicated by that value. [Figure 10] Figure 1 shows that repertoires generated from 501 tumor transcriptomes demonstrate a broad distribution of clonotype richness. Gene expression-based estimates for CD4 / CD8 T cells (published patent application Ser. No. 16 / 533,676, incorporated herein by reference, and PMID: 30864330) (y-axis) are correlated with clonotype yield (reads supporting productive CDR3s, x-axis) for each receptor (one-tailed Pearson - 95% CI). Samples with an infiltrate estimate of 0.001 or less are indicated by that value. [Figure 11]Scatter plots (left) show the relationship between the number of TCR beta productive clonotypes and normalized Shannon entropy within each immune profile of 501 human cancer samples sequenced using the novel hybrid capture approach disclosed herein. Increased normalized Shannon entropy correlates with increased clonotypic diversity within the sample. Nine repertoires were selected (indicated by color-coded stars representing the sample's cancer type, from bottom to top: red for acute lymphocytic leukemia, orange for T-cell lymphoma, yellow for T-cell lymphoma, cyan for clear cell renal cell carcinoma, deep blue for pancreatic cancer, purple for ovarian cancer, light green for non-small cell lung cancer, green for non-small cell lung cancer, and black for breast cancer). A close-up of the top 10 clonotypes in the selected TRB repertoire (right) shows the productive receptor frequencies for the top 10 clonotypes (each color represents one of the top 10 clonotypes, with the remaining repertoire shown in gray). [Figure 12] 1 is a bar graph showing the frequency of the top 10 clonotypes in individuals with B-cell lymphoma previously treated with anti-CD19 CAR. Yellow stars indicate clonotypes representing reads aligned to the heavy chain of a chimeric antigen receptor. [Figure 13] This is a bar graph showing the productive frequencies for the top 10 clonotypes analyzed from SARS-CoV-2-infected individuals using a novel hybrid capture method. The data was then matched against a database of putative SARS-CoV-2-reactive TCR B clonotypes. Yellow and purple stars indicate clonotypes that were consistent with the MIRA assay data, suggesting that these clonotypes are specific for SARS-CoV-2. [Figure 14] FIG. 1 shows the number of genes in each class of IG(BCR) or TCR genes with 1, 2, 3, 4 or 5+ alleles and depicts the allelic variation of these genes. [Figure 15] FIG. 1 shows examples of aligned TCR reference sequences. [Figure 16] FIG. 1 shows the cumulative distribution of the number of mismatched base pairs (bp) and the proportion of mismatched bp (number of mismatches relative to gene length). [Figure 17] Figure 1 shows the difference in total desired coverage length (in base pairs) when using the complete set (upper limit) of IG and TCR allele sequences (Table 1) versus gene-level consensus sequences (Table 2). DETAILED DESCRIPTION OF THE INVENTION
[0027] Various aspects of the subject disclosure will now be described with reference to the drawings, in which like reference numerals correspond to like elements throughout the several views. However, these drawings and the following detailed description regarding them are not intended to limit the claimed subject matter to the particular forms disclosed. Rather, the intention is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the claimed subject matter.
[0028] In the following detailed description, reference is made to the accompanying drawings which form a part hereof, and in which is shown, by way of illustration, specific embodiments in which the present disclosure may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the disclosure. It should be understood, however, that the detailed description and specific examples, while indicating examples of embodiments of the present disclosure, are for purposes of illustration only and are not intended to be limiting. From this disclosure, various substitutions, modifications, additions, rearrangements, or combinations thereof may be made within the scope of the present disclosure and will be apparent to those skilled in the art.
[0029] According to common practice, various features illustrated in the drawings may not be drawn to scale. The figures shown herein are not intended to be actual views of particular methods, devices, or systems, but are merely idealized representations employed to illustrate various embodiments of the present disclosure. Accordingly, dimensions of various features may be arbitrarily increased or decreased for clarity. Also, some of the drawings may be simplified for clarity. Thus, the drawings may not depict all of the components of a given apparatus (e.g., device) or method. Also, like reference numerals may be used throughout the specification and drawings to denote like features.
[0030] The information and signals described herein may be represented using any of a variety of different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referenced throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof. Some figures may show signals as single signals for ease of representation and explanation. A signal may also represent a bus of signals, which may have various bit widths, and those skilled in the art will understand that the present disclosure may be implemented for any number of data signals, including a single data signal.
[0031] Various illustrative logical blocks, modules, circuits, and algorithmic operations associated with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or a combination thereof. To clearly illustrate the interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and operations are generally described in terms of their functionality. Whether such functionality is implemented in hardware or software depends on the specific application and design constraints imposed on the overall system. Those skilled in the art may implement the disclosed functions in various ways depending on each specific application, but such implementation decisions should not be interpreted as causing a departure from the scope of the embodiments of the disclosure described herein.
[0032] It is also noted that embodiments may be described in terms of processes that are depicted as flowcharts, flow diagrams, structure diagrams, or block diagrams. While a flowchart may describe operational acts as a sequential process, many of these acts may be performed in parallel or substantially simultaneously in separate processes. The order of acts may also be rearranged. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. Additionally, methods disclosed herein may be implemented in hardware, software, or both. If implemented in software, the functions may be stored on or transmitted over as one or more instructions or code on a computer-readable medium. Computer-readable media includes computer storage media and communication media, including any medium that facilitates transfer of a computer program from one place to another.
[0033] References to elements herein using phrases such as "first" and "second" should be understood not to limit the quantity or order of those elements, unless a limit is explicitly stated. Rather, these phrases may be used herein as a convenient method of distinguishing between two or more elements or instances of a single element. Thus, references to a first and a second element do not imply that only two elements may be employed or that the first element must precede the second element. Also, unless otherwise stated, a collection of elements may include one or more elements.
[0034] definition As used in the specification and claims, the singular forms "a," "an," and "the" include the plural forms unless the context clearly indicates otherwise. For example, the term "polypeptide fragment" should be interpreted to mean "one or more polypeptide fragments" unless the context clearly indicates otherwise. As used herein, the term "plurality" means "two or more."
[0035] As used herein, the terms "about," "approximately," "substantially," and "significantly" will be understood by those of ordinary skill in the art and will vary to some extent depending on the context in which they are used. If a term is used that is not clear to those of ordinary skill in the art, taking into account the context in which it is used, "about" and "approximately" will mean up to plus or minus 10% of the specific term, and "substantially" and "significantly" will mean plus or minus more than 10% of the specific term.
[0036] As used herein, the terms "include" and "including" have similar meanings to the terms "comprise" and "comprising." The terms "comprise" and "comprising" should be interpreted as "open" transitional terms that allow for the inclusion of additional elements beyond those recited in the claims. The terms "consist" and "consisting of" should be interpreted as "closed" transitional terms that do not allow for the inclusion of additional elements beyond those recited in the claims. The term "consist essentially of" should be interpreted as partially closed, allowing for the inclusion of only additional elements that do not fundamentally change the nature of the claimed subject matter.
[0037] As used herein, the term "subject" may be used interchangeably with "patient" or "individual," and may include "animals," and particularly "mammals." Mammalian subjects may include humans and other primates, as well as domestic animals, livestock, and pets, such as dogs, cats, guinea pigs, rabbits, rats, mice, horses, cattle, and dairy cows.
[0038] As used herein, a "subject sample" or a "biological sample" from a subject refers to a sample obtained from a subject, such as, without limitation, a tissue sample (e.g., fat, muscle, skin, nerve, tumor, biopsy (e.g., solid tumor biopsy), lymph node, etc.), or a fluid sample (e.g., saliva, mucus, blood, serum, plasma, lymph, urine, stool, cerebrospinal fluid, etc.), and / or a cell, cultured cell (e.g., organoid), or subcellular structure such as vesicles and exosomes.
[0039] "BCR" or "B cell receptor," depending on the context in which it is used herein, refers to an immunoglobulin molecule that forms a receptor protein that is normally located on the outer surface of a type of lymphocyte known as a "B cell." Depending on the context, the term BCR or b cell receptor refers to at least a portion of the region of the genome involved in the development of the B cell receptor.
[0040] As used herein, a "comprehensive genomic profiling panel" refers to a genomic profiling panel that includes more than 10 genes.
[0041] "Contig" refers to a collection of overlapping DNA segments that together represent a common region of DNA.
[0042] "IgM" refers to immunoglobulin M antibodies and their isotypes.
[0043] "IgD" refers to immunoglobulin D antibodies and their isotypes.
[0044] "IgG" refers to immunoglobulin G antibodies and their isotypes.
[0045] "IgA" refers to immunoglobulin A antibodies and their isotypes.
[0046] "IgE" refers to immunoglobulin E antibodies and their isotypes.
[0047] "NGS" refers to next-generation sequencing technology.
[0048] "Profiling" refers to any one of a variety of methods that can be used to learn about genes in a person or a specific cell type, and / or the way those genes interact with each other and / or the environment.
[0049] As used herein, "RNAseq" or "rna-seq" refers to an abbreviation for "RNA sequencing," a sequencing technology that uses NGS to reveal the presence and quantity of RNA in a biological sample. RNA-seq can be used for whole-transcriptome analysis, whole-exome analysis, targeted panel analysis, and combinations thereof.
[0050] As used herein, "clonal" refers to a population of cells derived from a single cell. For example, a single T cell undergoes several successive mitotic divisions to generate many T cells with the same T cell receptor. This population of T cells is considered clonal.
[0051] As used herein, "oligoclonal" refers to a population of cells derived from two or more cells but fewer than many single cells. For example, a population of T cells derived by mitotic expansion of 2, 3, 4, 5, 6, 7, 8, 9, or 10 distinct T cell clones is considered oligoclonal.
[0052] As used herein, "polyclonal" refers to a population of cells derived from many single clones. For example, a population of T cells derived by mitotic expansion of 11, 20, 50, or 100 or more different T cell clones is considered polyclonal.
[0053] "TCR" or "t cell receptor," depending on the context in which it is used herein, refers to a protein complex found on the surface of T cells or T lymphocytes that is involved in recognizing fragments of antigens as peptides bound to major histocompatibility complex (MHC) molecules. In some contexts, the term TCR or t cell receptor refers to at least a portion of a region of the genome involved in the development of the t cell receptor.
[0054] As used herein, the term "repertoire" refers to the entire information (including, without limitation, presence, absence, expression level, and variants) obtained by nucleic acid sequencing, such as NGS, relating to a particular class of molecules, such as receptors, e.g., B cell receptors and / or T cell receptors, or a collection of molecules in a system, such as an immune repertoire, which is believed to contain information about B cell receptor information, T cell receptor information, and other immune-related genes, such as MHC genes.
[0055] As used herein, the term "TCR / BCR repertoire" refers to the entire information obtained by nucleic acid sequencing (e.g., by NGS) of a sample isolated from a subject using hybrid capture probes comprising a TCR / BCR probe set (probe panel), alone or in combination with a targeted exome panel. When a targeted whole transcriptome or targeted whole exome panel is used, the TCR / BCR repertoire does not include the remaining (non-TCR / BCR) transcriptome data. The TCR / BCR repertoire includes information, e.g., in the form of sequence data, about the gene targets in the TCR / BCR panel. This information can be analyzed by methods known in the art to determine the receptor type (e.g., TCR type α:β or γ:δ, BCR type IgD, IgM, IgA, IgG, or IgE), receptor identity (e.g., based on non-constant region sequences), and the abundance of specific receptors to obtain a TCR / BCR profile.
[0056] As used herein, a "TCR / BCR profile" refers to a subset of information in the TCR / BCR repertoire that allows prediction or identification of a condition or status observation, such as a medical state, disease, treatment response, tumor infiltration, etc., of a subject or cohort. In some embodiments, the TCR / BCR profile comprises a clinically actionable observation. By way of example, a TCR / BCR patient profile is typically obtained by analysis, such as statistical analysis of NGS sequencing data (e.g., of the TCR / BCR repertoire or immune repertoire), and the results of such analysis may be presented or output in any form, such as a report or other visual representation, summary, list, display, etc. Exemplary information in a TCR / BCR profile includes, but is not limited to, one or more of multiple TCR / BCR receptor sequences (clones), receptor abundance (e.g., indicating clonal abundance), the most abundant receptor, the abundance of one or more specific receptors, the degree of receptor diversity (clonality), receptor type, abundance of non-constant regions, and any combination of the above. By way of example only, TCR / BCR profiles may include the top 10 most abundant receptors (see, e.g., Example 3), clonality within the repertoire in various cancers (see, e.g., Example 7), and identification of receptors common to cohort databases (see, e.g., Example 9).
[0057] "TCR / BCR profiling" refers to the profiling of at least a portion of the region of the genome involved in the development of the T cell receptor or the B cell receptor.
[0058] "V(D)J recombination" refers to the near-random rearrangement of variable (V), joining (J), and sometimes diversity (D) gene segments that results in a wide variety of amino acid sequences in the antigen-binding regions of immunoglobulins and TCRs that enable recognition of antigens from pathogens, including bacteria, viruses, fungi, parasites and helminths, as well as certain cancer cells.
[0059] "V region" refers to a BCR or TCR variable gene segment or its gene product.
[0060] "D region" refers to a BCR or TCR diversity gene segment, or its gene product.
[0061] "J region" refers to a BCR or TCR joining gene segment or its gene product.
[0062] "C region" refers to a BCR or TCR constant gene segment, or its gene product.
[0063] As used herein, "transcriptome" refers to the entire range of messenger RNA molecules expressed by an organism, a particular tissue, or a particular cell. A transcriptome can be defined at a particular time point, e.g., a particular developmental stage, a particular disease stage, etc.
[0064] "Whole transcriptome" refers to coding and non-coding RNA expressed in cells, tissues, organs, and / or the entire body.
[0065] "Whole transcriptome sequencing" or "whole transcriptome profile" refers to the measurement of the full complement of transcripts in a sample at a given time. Whole transcriptome sequencing captures both coding transcripts (mRNA) and non-coding transcripts (e.g., miRNA, tRNA, rRNA, etc., when rRNA is of interest), providing a "snapshot" of expression levels, exons, introns, and variants. In some embodiments, whole transcriptome sequencing begins by removing rRNA from the sample (rRNA typically accounts for the majority of sequencing reads). In some embodiments, whole transcriptome sequencing includes a transcriptome enrichment step using a targeted panel to enrich for specific RNA sequences and / or remove or reduce the presence of other sequences (e.g., by removing high-abundance RNA species using species-specific rRNA probes). By way of example, a whole transcriptome targeted panel for enrichment for specific RNA sequences can include probes that enrich for 5,000, 10,000, or 20,000 or more RNA targets. In some embodiments, a whole transcriptome targeted panel comprising transcriptome hybrid capture probes can include a whole exome panel, for example, the Integrated DNA Technologies xGen Exome Research Panel v2.
[0066] As used herein, "exome" refers to the portion of the genome composed of exons, i.e., sequences that, when transcribed, remain in the mature RNA after introns are removed by RNA splicing and contribute to the final protein product encoded by that gene.
[0067] As used herein, "whole exome sequencing" refers to sequencing the protein-coding region of the genome, typically using NGS sequencing methods. Because the human exome comprises less than 2% of the genome but contains up to 85% of known disease-associated variants, this method is a cost-effective alternative to whole genome sequencing. In some embodiments, whole exome sequencing is performed including an exome enrichment step using an exome-targeted panel for enrichment of exome sequences (and, e.g., removal of non-coding sequences). Such panels are commercially available and typically include probes that enrich for 5,000, 10,000, or 20,000 or more genes. As an example, a non-limiting exome panel is the Integrated DNA Technologies xGen Exome Research Panel v2.
[0068] As used herein, the terms "targeted panel" and "targeted gene sequencing panel" or "targeted panel" are used interchangeably to refer to a set of probes directed to a select set of genes or gene regions of interest. Targeted panels are useful tools for detecting a set of specific sequences in a given sample. In some embodiments, targeted panels generate smaller, more manageable data sets (e.g., TCR / BCR profiles) compared to more common approaches such as whole genome sequencing. In some embodiments, targeted panels include whole exome panels or whole transcriptome panels, covering 5,000, 10,000, or 20,000 or more targets. In some embodiments, targeted panels include hybrid capture probes.
[0069] As used herein, a "hybridization capture probe" or "hybrid capture probe" refers to a biotinylated oligonucleotide that contains a region complementary to a nucleic acid sequence of interest sufficient to bind to (hybridize to) the nucleic acid sequence of interest and provide a means for enriching it using streptavidin-conjugated capture moieties bound to a solid support structure, e.g., a bead. In various embodiments, other capture moieties may be used in place of streptavidin and biotinylation. Examples of binding moieties include, but are not limited to, biotin:streptavidin, biotin:avidin, biotin:avidin:streptavidin, antibody:antigen, antibody:antibody, and covalent chemical bonds (e.g., click chemistry).
[0070] As used herein, "probe" means a Pool The terms "probe collection" and "panel" refer to a collection of probes useful for enriching nucleic acid targets prior to sequencing. In some embodiments, probe collection and probe panel are used interchangeably. In some embodiments, a panel is a collection of probes. Pool In some embodiments, additional probes may be used in combination with the TCR / BCR panel to enrich for additional target genes or sequences of interest. Pool By way of example, probes directed to cancer-specific sequences (e.g., sequences that function as diagnostic, prognostic, and / or therapeutic biomarkers) are provided. Pool may be included with a BCR / TCR panel instead of or in addition to a whole transcriptome or whole exome panel.
[0071] The terms "polynucleotide," "nucleic acid," and "nucleic acid molecule" are used interchangeably and refer to a covalently linked sequence of nucleotides (i.e., ribonucleotides in RNA and deoxyribonucleotides in DNA) in which the 3' position of the pentose of one nucleotide is linked to the 5' position of the pentose of the next nucleotide by a phosphodiester group. The sequenced nucleotides may be nucleotides of any form of nucleic acid, including, without limitation, RNA, DNA, and cfDNA molecules. These terms also refer to complementary DNA (cDNA), which is DNA synthesized from a single-stranded RNA (e.g., messenger RNA (mRNA) or microRNA (miRNA)) template in a reaction catalyzed by reverse transcriptase. The term "polynucleotide" includes, without limitation, single-stranded and double-stranded polynucleotides.
[0072] As used herein, the term "gene" refers to a nucleic acid sequence that encodes a gene product, which may be a polypeptide or a functional RNA molecule. The term "gene" is used broadly herein to encompass both the genomic DNA form of a gene (i.e., a specific portion of a specific chromosome) and the mRNA and cDNA forms of the gene produced therefrom. During gene expression, genomic DNA is transcribed into RNA, which may be immediately functional or translated into a functional polypeptide. In addition to the coding region (i.e., the sequence encoding the gene product), a gene also contains "non-coding regions." Non-coding regions may be immediately adjacent to the coding region (e.g., 5' and 3' non-coding regions flanking the coding region) or may be separated from the coding region (e.g., many kilobases upstream or downstream). Some non-coding regions, including "introns" (i.e., regions removed by RNA splicing before translation) and translational regulatory elements (e.g., ribosome binding sites, terminators, and start and stop codons), are transcribed into RNA but not translated. Other non-coding regions, including basal transcriptional regulatory regions, are not transcribed. A gene requires a "promoter," a sequence recognized and bound by proteins (i.e., transcription factors) that recruit RNA polymerase and facilitate its binding and initiation of transcription. A gene can have more than one promoter, resulting in messenger RNA (mRNA) of varying sizes extending to the 5' end. As used herein, a gene may also contain more distally located transcriptional regulatory elements (i.e., "enhancers" and "silencers") that can loop in close proximity to the promoter, allowing proteins (i.e., "transcription factors") bound to distant regulatory sites to affect transcription. For example, an "enhancer" increases transcription by binding to activating proteins that facilitate the recruitment of RNA polymerase or the initiation of transcription. Conversely, a "silencer" binds to repressor proteins that make DNA inaccessible to RNA polymerase or inhibit transcription. A gene may also contain "insulator" elements that protect the promoter from inappropriate regulation.Insulators can function by blocking interactions with enhancers or silencers, or by acting as a barrier to prevent the spread of condensed chromatin. Although enhancers and silencers are not generally considered to be part of the gene itself (because a single enhancer or silencer can regulate the expression of multiple genes), the term gene as used herein encompasses distant elements that affect its expression.
[0073] As used herein, the term "promoter" refers to a DNA sequence capable of controlling the expression of a coding sequence or functional RNA. Generally, the coding sequence is located 3' to the promoter sequence. A promoter may be derived entirely from a native gene, may be composed of different elements derived from different naturally occurring promoters, or may even contain synthetic DNA segments. Those skilled in the art will understand that different promoters can cause expression of a gene in different tissues or cell types, or at different developmental stages, or in response to different environmental conditions. Artificial promoters that cause gene expression in many cell types at many times are generally referred to as "constitutive promoters." Artificial promoters that allow selective expression of a gene in many cell types are referred to as "inducible promoters."
[0074] The term "gene sequence" or "sequence" refers to a series of nucleotides present in a DNA, RNA, or cDNA molecule. In the context of the present invention, a sequence is determined by sequencing nucleic acids present in a biological sample.
[0075] The term "read" refers to a DNA sequence of sufficient length (e.g., about 30 bp) that it can be used to identify a larger sequence or region by aligning it with a chromosome, e.g., a genomic region or gene. A read can be paired-end or single-end.
[0076] As used herein, the term "reference genome" refers to any specific known genome sequence, whether partial or complete, of any organism or virus that can be used to reference identifying sequences from a subject. Many reference genomes are listed on the National Center for Biotechnology Information website (www.ncbi.nlm.nih.gov). "Genome" refers to the complete genetic information of an organism or virus that is expressed in a nucleic acid sequence.
[0077] As used herein, the terms "aligned," "alignment," or "aligning" refer to a process used to identify regions of similarity. In the context of the present invention, alignment refers to matching sequences with locations in a reference genome based on the order of nucleotides in those sequences. Alignment can be performed manually or by a computer algorithm, such as the Efficient Local Alignment of Nucleotide Data (ELAND) computer program sold as part of the Illumina Genomics analysis pipeline. Alignment can refer to a 100% sequence match or a match less than 100% (an incomplete match). In various examples, alignment includes a false alignment.
[0078] The terms "library" and "sequencing library" are used herein to refer to a sequence of adaptor-ligated DNA fragments. Pool Adapters are generally designed to interact with the surface of a particular sequencing platform, such as a flow cell (Illumina) or bead (Ion Torrent), to facilitate the sequencing reaction.
[0079] The terms "sequencing probe" or "sequencing primer" are used herein to refer to short oligonucleotides used to sequence nucleic acids (i.e., cDNA or DNA). Sequencing probes may hybridize to a target sequence within the nucleic acid, or to an adapter sequence attached to the nucleic acid to allow for non-specific amplification and sequencing.
[0080] The term "RNA read count" is used herein to refer to the number of sequencing reads generated by a genetic analyzer. The term "RNA read count" is often used to refer to the number of reads that overlap with a given feature (e.g., a gene or chromosome).
[0081] The term "genetic profile" is used herein to refer to information about specific genes in an individual or a particular type of tissue. This information may include genetic variations (e.g., single nucleotide polymorphisms), gene expression data, and other genetic or epigenetic characteristics (e.g., DNA methylation patterns) determined, for example, by analysis of next-generation sequencing data.
[0082] The term "variant" is used to refer to a difference in a gene sequence or gene profile compared to a reference genome or reference gene profile.
[0083] The term "expression level" is used herein to refer to the number of copies of a particular RNA or protein molecule produced by a gene or other gene regulatory region (long non-coding RNA, enhancer) that can be defined by chromosomal location or other genetic mapping marker, which may or may not be normalized using standard methods (e.g., counts per million, taking the base 10 logarithm of the original read count).
[0084] The term "gene product" is used herein to mean a protein or RNA molecule produced by expression (ie, transcription, translation, post-translational modification, etc.) of a gene or other gene regulatory region.
[0085] The terms "extracted," "recovered," "isolated," and "separated" refer to a compound (e.g., protein, cell, nucleic acid, or amino acid) that has been removed from at least one component with which it is naturally associated and occurs in nature.
[0086] As used herein in conjunction with nucleic acid sample preparation, e.g., for NGS sequencing methods, the term "enriched" or "enrichment" refers to a process that increases the amount of one or more nucleic acid species in a sample. Exemplary enrichment methods can include chemical and / or mechanical means, and can also include amplifying nucleic acids contained in a sample. By way of example, enrichment may include the use of hybrid capture probes and polymerase chain reaction (PCR). Enrichment can be sequence-specific (e.g., the use of hybrid capture probes or target-specific PCR primers) or non-specific (i.e., involving any of the nucleic acids present in the sample). As used herein in reference to the concentration or amount of one or more biomolecules in a sample, "enriched" refers to an increase in the concentration or amount of one or more biomolecules, such as nucleic acids or proteins, compared to a control concentration or (in relative amount) other biomolecules in the sample. In the context of data science, "enrichment" refers to statistical enrichment.
[0087] As used herein, the term "cancer" refers to any one or more of a wide range of benign or malignant tumors, including those capable of invasive growth and metastasis in humans or animals or portions thereof, e.g., via the lymphatic system and / or bloodstream. As used herein, "tumor" includes both benign and malignant tumors, as well as solid growths. Exemplary cancers include, but are not limited to, carcinomas, lymphomas, or sarcomas, such as ovarian cancer, colon cancer, breast cancer, pancreatic cancer, lung cancer, prostate cancer, urethral cancer, uterine cancer, acute lymphocytic leukemia, Hodgkin's disease, small cell carcinoma of the lung, melanoma, neuroblastoma, glioblastoma, and sarcomas of human soft tissue.
[0088] For the purposes of this disclosure, the term biomarker shall be taken to mean any genetic variant or molecule or set of molecules, or molecular feature (e.g., location, expression level, etc.) that is indicative of or correlates with a trait of interest, e.g., the presence of an infection, medical condition, or disease such as cancer, or a patient's susceptibility to an infection, condition, or disease; the probability that the infection, medical condition, or disease is of some subtype; the likelihood that the patient will or will not respond to a particular treatment or treatments; the degree of positive response (e.g., survival and / or progression-free survival quantifiable as a time interval) expected to a treatment or treatment group, whether the patient responds to the treatment or not; or the probability that the infection, medical condition, or disease has or will progress, or has progressed beyond its primary site (metastasized). In some embodiments, a biomarker comprises a TCR / BCR profile.
[0089] Terms such as "treatment" and "treating" are used herein generally to mean obtaining a desired pharmacological and / or physiological effect. The effect may be prophylactic, in terms of completely or partially preventing a disease or its symptoms, and / or therapeutic, in terms of partially or completely curing a disease and / or the deleterious effects resulting from the disease. As used herein, "treatment" encompasses any treatment of a disease in a mammal, including (a) preventing the onset of a disease in a subject believed to be susceptible to the disease but not yet diagnosed as having the disease; (b) suppressing the disease, i.e., arresting its development; or (c) alleviating the disease, i.e., causing the disease to regress. A therapeutic agent may be administered before, during, or after the onset of a disease or injury. Treatment of an ongoing disease, in which treatment stabilizes or alleviates undesirable clinical symptoms in the patient, is of particular interest. Administration of a therapeutic agent to a subject is desirably during the symptomatic stage of the disease, and optionally after the symptomatic stage of the disease.
[0090] The term "effective amount" refers to an amount of an active agent sufficient to exert a detectable therapeutic effect when used in accordance with the disclosed methods without undue adverse side effects (such as toxicity, irritation, and allergic reactions), consistent with a reasonable benefit-risk ratio. The effective amount for a patient will depend on the type of patient, the patient's size and health, the nature and severity of the condition being treated, the method of administration, the duration of treatment, the nature of concomitant therapy (if any), and the specific formulation employed. Thus, it is not possible to specify an exact effective amount in advance. However, the effective amount for a given situation can be determined by one of ordinary skill in the art through routine experimentation based on knowledge in the art and the information presented herein. The optimal administration regimen can be determined by one of ordinary skill in the art without undue experimentation.
[0091] overview The diversity of lymphocyte receptors (T cell receptors, "TCR" and B cell receptors, "BCR") is increased by recombination, i.e., approximately 10 18This is achieved by generating a theoretical diversity of unique receptors for each lymphocyte. Thus, upon antigen challenge and presentation by innate immune cells, lymphocytes bearing high-affinity receptors for antigens or antigens bound to major histocompatibility complexes I or II (MHCI or II) are activated and clonally expanded. Importantly, lymphocytes that have already been activated and differentiated into "effector" cells persist within an individual for some time. Furthermore, "memory" cells may persist throughout the host's lifetime, even after resolution of an infection or disappearance of a pathogen. The specific types of lymphocyte receptors present in an individual subject are referred to as a "repertoire." Thus, an individual subject's lymphocyte receptor repertoire contains a signature of their response to antigen challenge, including diseases such as cancer, and provides a record of the pathogens they have encountered. For example, in the case of blood cancers, the repertoire can be monitored to identify "tumors" or cancer cells, and in the case of solid tumors, the repertoire can be used to investigate why the immune system is not recognizing and eliminating cancer cells as non-self. For at least these reasons, there is great interest in the lymphocyte receptor profile of individual subjects.
[0092] The field of immune profiling utilizes a technology known as next-generation sequencing (NGS) to accurately sequence the immune receptor repertoire in individual subjects. NGS can generate millions of sequencing “reads” that are subsequently aligned to a reference genome or transcriptome to obtain a relatively complete picture of an individual's genome or sample's transcriptome. However, significant technical and computational challenges exist in assembling and analyzing the reads to detect and evaluate T and B cell receptors. Quality sequencing data depends on two factors known as sequencing breadth and depth. Sequencing breadth refers to the number or percentage of genomic bases covered by sequencing, while sequencing depth refers to the approximate number of times a particular base or region is covered by a sequencing run. However, in any given sample, the presence of specific transcripts encoding lymphocyte receptors, or possibly lymphocytes as a whole, can be very limited. Therefore, obtaining deep sequencing results that accurately represent an individual subject's T and B cell repertoire without directly enriching or selecting for the non-constant regions of the TCR and BCR may be challenging.
[0093] T and B cell receptors consist of discrete genes that rearrange to form a large repertoire present in an individual. Therefore, any strategy for selectively enriching T and B cell receptor transcripts must be adapted to address low-abundance transcripts encoding lymphocyte receptors and the diverse genes assembled to encode recombinant antigen receptors. Additionally, mapping of sequencing reads to TCRs and BCRs can be unbalanced, potentially biasing the detection of either TCR or BCR clones in a given sample. Furthermore, the most important information in immune profiling resides in hypervariable (non-constant) regions, not constant regions. Therefore, sequences directed toward hypervariable regions may require enrichment. In some cases, sample quantity and / or quality may be limited because the sample is derived from a biopsy or other rare sample. Therefore, there is a need in the art for a method that can not only extract high-quality RNA sequencing data from a single sample or sequencing run, but also provide deep and accurate immune profiling at scale. The disclosed methods and systems address this need in the art.
[0094] Targeted sequencing of T and B cells can be a powerful tool for mapping the immune system (the "immunome") in cancer and other conditions such as autoimmune diseases, infectious diseases, and transplantation. Individual non-clonal T and B cells are unique at the DNA level, differing in the T cell receptor (TCR) or B cell receptor (BCR) genes that determine the pathogens or antigens to which the cells respond. As disclosed herein, assembling (determining) the sequences of TCR and BCR genes via RNA sequencing allows for more precise mapping of the immune system and generates new compartmentalized immune-specific signatures that can provide the information needed, for example, to predict immune responses, diagnose or confirm diseases, conditions, or pathogen exposures, determine disease severity, measure or confirm treatment efficacy and effectiveness, determine minimal residual disease (MRD), and create specific therapies such as chimeric antigen receptor (CAR) T cells (CAR-T cells), NK cells (CAR-NK cells), macrophages (CAR-M cells), or other cell types engineered to express CARs, anti-cancer immune mobilizing monoclonal T cell receptors (ImmTACs), other adoptive cell therapies, and vaccines.
[0095] Determination of TCR / BCR profile Disclosed herein are methods, systems, and compositions for determining a TCR / BCR profile of a subject. In some embodiments, the method includes (a) isolating RNA from a patient sample, (b) enriching the isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes, (c) sequencing the RNA of (b) to generate sequencing data, and (d) analyzing the sequencing data to determine the TCR / BCR profile of the patient. In some embodiments, the collection of hybrid capture probes includes a first collection of probes that contain BCR constant region probes. Pool , the second containing a BCR non-constant region probe Pool , a third containing a TCR constant region probe Pool , and a fourth containing a TCR non-constant region probe. PoolIn some embodiments, a TCR / BCR probe set obtained according to the method of Example 1 is used.
[0096] In some embodiments, TCR / BCR profiling may be performed as a standalone assay. In some embodiments, the TCR / BCR profiling methods and systems disclosed herein may be configured for use within a more extensive RNAseq whole-transcriptome or whole-exome RNA panel, thereby preserving rare patient samples, shortening time to diagnosis or treatment recommendations, and providing a novel and valuable method for obtaining specific information about a subject's immune profile in addition to gene expression data and associated genetic data (such as, but not limited to, alternative splicing events, fusions, and genetic variants). When incorporated into an RNAseq platform, TCR / BCR profiling is considered to have expanded utility and an even more unique and valuable resource for generating insights. By way of example, but not limitation, TCR / BCR profiling can be used to profile and track a variety of disease states and associated immune responses, including cancer, infectious diseases, transplants, allergic diseases (induced by airway, food, or other allergens), and autoimmunity. Allergic diseases include contact dermatitis, asthma, anaphylaxis, non-IgE-mediated food allergies associated with atopic dermatitis, etc. Autoimmune diseases include type 1 diabetes, rheumatoid arthritis, lupus, celiac disease, Sjogren's syndrome, multiple sclerosis, polymyalgia rheumatica, ankylosing spondylitis, alopecia areata, vasculitis, temporal arteritis, etc. TCR / BCR profiling, which may include allele typing, can be used for biomarker discovery to predict immune response, health outcomes, and / or disease severity.
[0097] Thus, in some embodiments, the TCR / BCR profiling methods disclosed herein are performed using sequencing techniques such as next-generation sequencing. In some embodiments, bulk (multi-cell) sequencing using short-read RNA sequencing can be used. The resulting reads can then be used, for example, to assemble contigs from TCR / BCR gene regions. In some embodiments, for example, the fifth of a set of probes comprising an exon-targeting panel can be used to assemble contigs from the TCR / BCR gene region. Pool is provided.
[0098] As noted above, one exemplary benefit of the methods, compositions, and systems of the present invention is that TCR / BCR profiling can be included within the scope of a more comprehensive RNAseq whole transcriptome panel. In oncology or more general profiling settings, TCR / BCR profiling can be added to other analyses of other parts of the transcriptome, such as cytokine expression, immune cell composition, potential viral / bacterial signals, and inflammatory signatures. Whole transcriptome analysis allows for the capture of that data while also providing a TCR / BCR snapshot.
[0099] Exemplary sample preparation methods useful for TCR / BCR profiling, NGS, and related methods are described below. The techniques of the present invention are not intended to be limited to sample preparation methods, and one of skill in the art will understand that substitutions, alternative reagents, and alternative processing steps may be used.
[0100] RNA extraction Transcriptome analysis, i.e., the study of the complete set of RNA transcripts (i.e., transcriptome) produced by cells, and exome analysis, i.e., the study of RNA encoding protein products, provide promising means for identifying genetic variants correlated with disease states and disease progression. For example, to identify genetic variants associated with cancer, transcriptome and / or exome analysis can be performed on samples collected from patients with cancer cells. Suitable patient samples include tissue samples, tumors (e.g., solid tumors), biopsies, lymph nodes, and body fluids (e.g., blood, serum, plasma, lymph, sputum, lavage fluid, cerebrospinal fluid, urine, semen, sweat, tears, saliva). Alternatively, transcriptome and / or exome analysis can be performed on organoids (i.e., "tumor organoids") generated from human cancer specimens. Sequencing can be performed on single-cell specimens or multicellular specimens.
[0101] While RNA sequencing (RNA-seq) can be performed on any patient sample containing RNA, those skilled in the art understand that sequencing protocols should be adapted to the specific sample being used. For example, RNA tends to be significantly degraded in tissue samples that have already been processed for histology (e.g., formalin-fixed, paraffin-embedded (FFPE) tissue sections). Therefore, researchers modify several key steps in RNA-seq protocols to reduce sequencing artifacts (see, for example, BMC Medical Genomics 12, 195 (2019)).
[0102] Currently, transcriptome and exome analyses are overwhelmingly performed using high-throughput RNA sequencing (RNA-Seq), which uses next-generation sequencers to detect RNA transcripts in a sample. The first step in performing RNA-Seq is to extract RNA from a sample.
[0103] Cell lysis The first step in RNA extraction from a sample is often lysing the cells present in the sample. Cells are commonly lysed using several physical disruption methods, including mechanical disruption (e.g., using a blender or tissue homogenizer), liquid homogenization (e.g., using a dounce or French press), high-frequency sonication (e.g., using a sonicator), freeze-thaw cycles, heating, manual grinding (e.g., using a mortar and pestle), and bead-beating (e.g., using a BioSpec Mini-beadbeater-96). Cells are also commonly lysed using detergent-containing reagents, many of which are commercially available (e.g., QIAzol Lysis Reagent from QIAGEN, FastBreak™ Cell Lysis Reagent from Promega). Physical disruption methods are often performed in a "homogenization buffer" containing a lysis reagent, such as a detergent to enhance lysis efficiency or a protease (e.g., proteinase K). The homogenization buffer may also contain an antifoaming agent and / or an RNase inhibitor to protect RNA from degradation. Those skilled in the art will appreciate that different cell lysis techniques may be required to obtain the highest possible yield from different tissues, with techniques that minimize degradation of the released RNA and that avoid release of nuclear chromatin being preferred.
[0104] RNA isolation After cells are lysed, RNA can be separated from other cellular components. Total RNA is typically isolated using guanidine thiocyanate-phenol-chloroform extraction (e.g., using TRIzol) or by trichloroacetic acid / acetone precipitation followed by phenol extraction. However, a number of column-based systems for RNA extraction are also commercially available (e.g., Invitrogen's PureLink RNA Mini Kit and Zymo Research's Direct-zol Miniprep Kit).
[0105] Ideally, the isolated RNA is largely free of DNA and enzymatic contamination. To this end, isolation methods may utilize agents that eliminate DNA (e.g., TURBO DNase-I) and / or remove enzymatic proteins from the sample (e.g., Agencourt® RNAClean® XP beads from Beckman Coulter).
[0106] Whole-transcriptome sequencing is sometimes used to analyze all transcripts present in a cell, including messenger RNA (mRNA) and all non-coding RNAs. Focusing on the entire transcriptome allows researchers to map exons and introns and identify splice variants. Notably, most whole-transcriptome library preparation protocols include a step to remove ribosomal RNA (rRNA), which would otherwise be expected to capture the majority of sequencing reads. rRNA depletion is commonly achieved using kits such as Illumina's Ribo-Zero Plus rRNA Depletion Kit and Zymo's Seq RiboFree Total RNA Library Kit.
[0107] In other instances, more targeted RNA-Seq protocols are used to focus on specific types of RNA. For example, mRNA-seq is typically used to selectively study the "coding" portions of the genome, which typically account for only 1-2% of the total transcriptome. Enriching samples for mRNA increases the sequencing depth achieved for coding genes, allowing for the identification of rare transcripts and variants. Polyadenylated mRNA is typically enriched when using oligo-dT beads (e.g., Invitrogen's Dynabeads™). This enrichment step can be performed on either isolated total RNA or raw cell lysates.
[0108] Targeted approaches have also been developed for the analysis of microRNAs (miRNAs) and small interfering RNAs (siRNAs), which are typically isolated using kits designed for efficient recovery of small RNAs (e.g., Invitrogen's mirVana™ miRNA isolation kit).
[0109] Library preparation After RNA is extracted from a sample, the next major step is to convert the RNA into a form suitable for next-generation sequencing (NGS). Through a series of steps, the RNA is converted into a collection of DNA fragments known as a "sequencing library." After the library is sequenced, the resulting sequencing "reads" are aligned against a reference genome or transcriptome to determine the expression profile of the analyzed cell.
[0110] In some embodiments, library preparation is automated, allowing for higher sample throughput, minimizing errors, and reducing hands-on time. For example, fully automated library preparation can be performed using a liquid handling robot (e.g., PerkinElmer's SciClone® NGSx).
[0111] Reverse transcription / cDNA preparation After RNA is extracted from a sample, the next major step is to convert the RNA into a form suitable for next-generation sequencing (NGS). Through a series of steps, the RNA is converted into a collection of DNA fragments known as a "sequencing library." After the library is sequenced, the resulting sequencing "reads" are aligned against a reference genome or transcriptome to determine the expression profile of the analyzed cell.
[0112] In some cases, library preparation is automated, allowing for higher sample throughput, minimizing errors, and reducing hands-on time. For example, fully automated library preparation can be performed using liquid handling robots (e.g., PerkinElmer's SciClone® NGSx).
[0113] For sequencing, RNA is converted into more stable, double-stranded complementary DNA (cDNA) using reverse transcription (RT). In some cases, reverse transcription is performed directly on the sample lysate prior to RNA isolation. In other instances, reverse transcription is performed on isolated RNA.
[0114] Reverse transcription is catalyzed by reverse transcriptase, an enzyme that synthesizes a complementary strand of cDNA using an RNA template and a short primer complementary to the 3' end of the RNA. This first strand of cDNA is then made double-stranded either by subjecting it to PCR or by using a combination of DNA polymerase I and DNA ligase. In the latter method, an RNase (e.g., RNase H) is typically used to digest the RNA strand, allowing the first cDNA strand to serve as a template for synthesizing the second cDNA strand.
[0115] Many reverse transcriptases are commercially available, including avian myeloblastosis virus (AMV) reverse transcriptase (e.g., AMV Reverse Transcriptase from New England BioLabs) and Moloney murine leukemia virus (M-MuLV, MMLV) reverse transcriptase (e.g., SMARTscribe™ from Clontech, SuperScript II™ from Life Technologies, and Maxima H Minus™ from Thermo Scientific). Notably, many of the available reverse transcriptases have been engineered to improve thermal stability or efficiency (e.g., by eliminating 3' to 5' exonuclease activity or reducing RNase H activity).
[0116] The primer that serves as the starting point for synthesizing the new strand can be a random primer (i.e., directed to RT of any RNA), an oligo dT primer (i.e., directed to RT of mRNA), or a gene-specific primer (i.e., directed to RT of a specific target RNA).
[0117] Following reverse transcription, an exonuclease (eg, exonuclease I) can be added to the sample to degrade any primers remaining after the reaction and prevent them from interfering with the subsequent amplification step.
[0118] concentrated For some applications, sequencing the entire transcriptome of a sample is not necessary. Instead, "targeted sequencing" can be used to study a selected set of genes or specific genomic elements. Libraries enriched for target sequences are generally prepared using hybridization-based methods (i.e., hybridization capture-based target enrichment). Hybridization can be performed on a solid surface (microarray) or in solution. Solution-based methods use biotinylated oligonucleotide probes that specifically hybridize with the genes or genomic elements of interest. Pool are added to the library. The probes are then captured and purified using streptavidin-coated magnetic beads, after which the sequences hybridized to these probes are amplified and sequenced. Numerous probe panels for library enrichment are commercially available, including those from IDT (e.g., xGen Exome Research Panel v1.0 and v2.0 probes) and Roche (e.g., SeqCap® probes). The customizability of many available probe panels allows researchers to design a collection of capture probes precisely tailored to a particular application. Numerous kits (e.g., Roche's SeqCap EZ MedExome Targeted Exome Enrichment Kit) and hybridization mixes (e.g., IDT's xGen Lockdown) that facilitate target enrichment are also commercially available.
[0119] In some cases, it may be advantageous to treat the library with a reagent that reduces off-target capture before performing target enrichment. For example, libraries are commonly treated with oligonucleotides that bind to adapter sequences (e.g., xGen blocking oligos) or repeat sequences (e.g., human Cot DNA) to reduce non-specific binding to capture probes.
[0120] A detailed discussion of enrichment and an exemplary enrichment scheme for the TCR / BCR gene region is provided below.
[0121] Library amplification While not required for all sequencing applications, library preparation typically includes at least one amplification step to enrich for sequencing-compatible DNA fragments (i.e., fragments with adapter-ligated ends) and generate sufficient amounts of library material for downstream processing. Amplification may be performed using standard polymerase chain reaction (PCR) techniques. However, care should be taken, whenever possible, to minimize amplification bias and limit the introduction of sequencing artifacts. This is achieved through the selection of appropriate enzymes and protocol parameters. To this end, several companies offer high-fidelity DNA polymerases (e.g., Roche's KAPA HiFi DNA polymerase) that have been shown to generate more accurate sequencing data. These DNA polymerases are often purchased as part of a PCR master mix (e.g., New England BioLabs' NEBNext® High-Fidelity 2X PCR Master Mix) or as part of a kit (e.g., Roche's KAPA HiFi Library Amplification Kit).
[0122] Those skilled in the art will appreciate that even with highly optimized PCR protocols, PCR conditions must be fine-tuned for each sequencing experiment. For example, depending on the initial concentration of DNA in the library and the requirements of the sequencer being used, it may be desirable to subject the library to anywhere from 4 to 14 cycles of PCR.
[0123] In some cases, the library preparation protocol includes multiple rounds of library amplification, for example, in some cases, an additional amplification is performed after the library is accumulated, followed by a PCR cleanup.
[0124] Spike-in control Because cells from different experimental conditions may not yield identical amounts of RNA, sequencing data may be normalized to accurately identify changes across various experimental conditions. Normalization may be useful, for example, to address global transcriptional changes between different experimental conditions. For normalization, "spike-in controls" may be added to the sequencing library. In some embodiments, spike-in controls comprise, for example, DNA sequences added at a known ratio to the sample. Control DNA can be any DNA that can be easily distinguished from the experimental cDNA during data analysis. For example, control libraries typically contain synthetic DNA or DNA from an organism other than the organism of interest (e.g., PhiX spike-in controls may be added to a human-derived library).
[0125] Fragmentation and Size Selection For sequencing techniques that cannot easily analyze long DNA strands, DNA is generally fragmented into uniform fragments prior to sequencing. The optimal fragment length depends on both the sample type and the sequencing platform used. For example, whole genome sequencing typically works best with DNA fragments 350 bp or less in length, while targeted sequencing using hybridization capture (see Section 2G) works best with DNA fragments 200 bp or less in length.
[0126] In some cases, fragmentation is performed after reverse transcription (i.e., on the cDNA). Suitable DNA fragmentation methods include physical methods (e.g., using sonication, sound, spray, centrifugal force, needles, or hydrodynamics), enzymatic methods (e.g., using NEBNext dsDNA Fragmentase from New England BioLabs), and tagmentation (e.g., using the Nextera™ system from Illumina).
[0127] In other instances, fragmentation is performed prior to reverse transcription (i.e., on the RNA). In addition to fragmentation methods suitable for DNA, RNA may also be fragmented using heat and magnesium (e.g., using the KAPA Hyper Prep kit from Roche).
[0128] A size selection step may then be performed to enrich the library for fragments of an optimal length or length range. Traditionally, size selection was achieved by separating fragments of different sizes using agarose gel electrophoresis, excising fragments of the desired size, and performing gel extraction (e.g., using Qiagen's MinElute Gel Extraction Kit™). However, size selection is now commonly achieved using magnetic bead-based systems (e.g., Beckman Coulter's AMPure XP™, Promega's ProNex® Size-Selective Purification System).
[0129] Sequencing adapter ligation Prior to sequencing, cDNA fragments are ligated to sequencing adaptors. Sequencing adaptors are short DNA oligonucleotides that contain (1) sequences required for amplifying the cDNA fragments during the sequencing reaction and (2) sequences that interact with the NGS platform (e.g., the surface of an Illumina flow cell or Ion Torrent beads). Therefore, adaptors must be selected based on the sequencing platform to be used.
[0130] Libraries from multiple samples are typically pooled and analyzed during a single sequencing run (see "Polling," below). To track the origin of each cDNA in the pooled samples, adapters ligated to the cDNA fragments in each library contain a unique molecular barcode (or a combination of multiple barcodes). During the sequencing reaction, a sequencer reads this barcode sequence in addition to the biological sequence of the cDNA. The barcodes are then used during data analysis to assign each cDNA to its original sample, a process called "demultiplexing."
[0131] The indexing strategy used for sequencing reactions should be selected based on the number of samples stored and the desired level of accuracy. For example, to ensure libraries are demultiplexed with high accuracy, unique double indexing, which adds unique identifiers to both ends of cDNA, is commonly used. Adapters may also contain unique molecular identifiers (UMIs), short sequences often containing degenerate bases that incorporate a unique barcode into each molecule within any given sample library. UMIs reduce false-positive variant calls and increase variant detection sensitivity by allowing true variants to be distinguished from errors introduced during library preparation, target enrichment, or sequencing. Numerous sets of index sequences and adapters are commercially available, including Roche's SeqCap double-end adapters, IDT's xGen double-index UMI adapters, and Illumina's TruSeq UD index.
[0132] Library Cleanup Following PCR, the amplified DNA is typically purified to remove residual enzymes, nucleotides, primers, and buffer components from the reaction. Purification is generally achieved using phenol-chloroform extraction followed by ethanol precipitation, or using spin columns containing a silica matrix to which DNA selectively binds in the presence of chaotropic salts. Numerous column-based PCR cleanup kits are commercially available, including those from Qiagen (e.g., MinElute PCR Purification Kit), Zymo Research™ (DNA Clean & Concentrator™-5), and Invitrogen (e.g., PureLink™ PCR Purification Kit). Alternatively, purification can be achieved using paramagnetic beads (e.g., Axygen™ AxyPrep Mag™ PCR Cleanup Kit).
[0133] Accumulation To keep sequencing cost-effective, clinical laboratory technicians or researchers often store diverse libraries together, each with a unique barcode (see "Sequencing Adapter Ligation" above), allowing them to be sequenced en masse. The sequencer used and the desired sequencing depth will dictate the number of samples stored. For example, in some applications, it may be advantageous to store fewer than 12 libraries to achieve greater sequencing depth, while in other applications it may be prudent to store more than 100 libraries.
[0134] When sequencing multiple libraries in a single run, care should be taken to ensure roughly equal sequencing coverage for each library. To this end, equal amounts of each library (based on molar concentration) should be accumulated. Furthermore, the total molar concentration of the accumulated libraries must be compatible with the sequencer. Therefore, it is important to accurately quantify the DNA in the libraries (e.g., using the methods described in "Quality Control" below) and perform the necessary calculations before accumulating the libraries. In some cases, it may be necessary to concentrate the accumulated libraries, for example, using a vacuum concentrator, to achieve a suitable total molar concentration.
[0135] In various embodiments, accumulation is performed twice: in some embodiments, sequencer adaptor ligation and accumulation (e.g., accumulation of about 5-10 samples) is performed before enrichment / library amplification, and a second accumulation step is performed after library cleanup.
[0136] Quality control (cDNA library integrity, fragment size) Prior to sequencing, libraries may be assessed to ensure they contain sufficient DNA quantity and quality to produce useful sequencing results. DNA may be quantified to verify that the library concentration is sufficient for loading onto the sequencer. Commonly used DNA quantification methods include gel electrophoresis, UV spectrophotometry (e.g., NanoDrop®), fluorometry (e.g., Qubit™, Picofluor™), real-time PCR (also known as quantitative PCR), or droplet digital emulsion PCR (ddPCR). DNA quantification is often aided by a wide selection of commercially available dyes and stains (e.g., ethidium bromide, SYBR Green, RiboGreen®). Given the very narrow recommended input range for NGS, it is preferable to use a highly accurate quantification method to verify that the final library concentration is appropriate.
[0137] The fragment size distribution of the final library should also be assessed to verify that the fragment lengths are suitable for sequencing. Traditionally, fragment size distribution was determined by spreading the sample on an agarose gel. However, more advanced capillary electrophoresis methods (e.g., Bioanalyzer®, TapeStation®, Fragment Analyzer™, all manufactured by Agilent) that require smaller sample inputs are now more commonly employed. Conveniently, these methods can be used to analyze both DNA fragment size and concentration.
[0138] Clonal amplification When sequencing a library, it is applied to a device, typically a flow cell (Illumina) or chip (Ion Torrent), where the sequencing chemistry occurs. The cDNAs in the library can be attached to these devices by decorating them with short oligonucleotides complementary to adapter sequences. Prior to sequencing, the device undergoes clonal amplification (e.g., by cluster generation (Illumina) or microemulsion PCR (Ion Torrent)), which generates clusters of multiple copies of each cDNA on its surface, thereby amplifying the signal generated by each cDNA during the sequencing reaction. Clonal amplification is often performed using commercially available kits (e.g., Illumina's paired-end cluster kit). After clonal amplification, the library is ready for sequencing.
[0139] Exemplary enrichment of TCR / BCR gene regions In some embodiments, a plurality of nucleic acid probes (e.g., a hybrid capture probe set) is used to enrich for one or more target sequences in a nucleic acid sample (e.g., an isolated nucleic acid sample or a nucleic acid sequencing library), for example, where the one or more target sequences are informative for TCR / BCR profiling. The probes may be designed and generated according to methods known in the art. In some embodiments, the TCR / BCR probe set is obtained according to the method of Example 1. In some embodiments, the probe set includes probes that target one or more genetic loci, such as exonic or intronic loci. In some embodiments, the probe set includes probes that target one or more non-protein-coding genetic loci, such as regulatory loci, miRNA loci, and other non-coding genetic loci, for example, one or more genetic loci recognized to be associated with a particular disease or medical condition (e.g., cancer). In some embodiments, the plurality of loci comprises at least 25, 50, 100, 150, 200, 250, 300, 350, 400, 500, 750, 1000, 2500, or 5000 or more human genomic loci.
[0140] Generally, probes for enriching nucleic acids (e.g., complementary DNA, cDNA, generated from nucleic acids extracted or isolated from a biological specimen, including extracted or isolated RNA) comprise DNA, RNA, or modified nucleic acid structures having a base sequence complementary to a locus of interest. For example, a probe designed to hybridize to a locus within a cDNA molecule can comprise a sequence complementary to either strand, since the cDNA molecule can be double-stranded. In some embodiments, each probe in the plurality of probes comprises a nucleic acid sequence identical to or complementary to at least 10, at least 11, at least 12, at least 13, at least 14, or at least 15 consecutive bases of a locus of interest. In some embodiments, each probe in the plurality of probes comprises a nucleic acid sequence identical to or complementary to at least 20, 25, 30, 40, 50, 75, 100, 150, or 200 or more consecutive bases of a locus of interest.
[0141] By way of example, but not limitation, probe sequences may be selected according to the methods described in FastPCR software for PCR primer and probe design and iterative searching (Kalendar et al., 2009, Genes, Genomes, and Genomics, Vol. 3 (Special Issue 1), pp. 1-14), which is incorporated herein by reference.
[0142] Targeted panels offer several benefits for nucleic acid sequencing. In one example, a panel targeting genes (including TCR and BCR genes) with high variability in individual subjects, humans, or even cells within a subject or human body can facilitate the bioinformatics process of determining the sequences of those genes. For example, if a "whole exome" or targeted sequencing panel does not generate a sufficient number of sequencing reads mapping to highly variable genes, probes targeting highly variable genes can be added to the whole exome or targeted sequence panel probes to increase the number of reads mapping to highly variable genes.
[0143] In some embodiments, the gene panel is a whole exome panel that analyzes the exome of a biological sample. In some embodiments, the gene panel is a whole genome panel that analyzes the genome of a specimen. In some embodiments, the gene panel is a whole transcriptome panel that analyzes the transcriptome of a specimen. In some embodiments, the gene panel is a targeted whole transcriptome panel that analyzes the transcriptome of a specimen. In some embodiments, the gene panel is used in conjunction with a TCR / BCR gene panel (e.g., to support clinical decisions related to immunological profile or immunome).
[0144] In some embodiments, the probes of the panel contain additional nucleic acid sequences that do not share homology to the loci of interest. For example, in some embodiments, the probes also contain nucleic acid sequences that include an identifier sequence, such as a unique molecular identifier (UMI), that is unique to a particular sample or subject. Examples of identifier sequences are described, for example, in Kivioja et al., 2011, Nat. Methods 9(1), pp. 72-74, and Islam et al., 2014, Nat. Methods 11(2), pp. 163-66, which are incorporated herein by reference. Similarly, in some embodiments, the probes also contain primer nucleic acid sequences, useful for amplifying nucleic acid molecules of interest, for example, using polymerase chain reaction (PCR). In some embodiments, the probes also contain a capture sequence designed to hybridize to an anti-capture sequence to recover nucleic acid molecules of interest from the sample.
[0145] Similarly, in some embodiments, each probe comprises a non-nucleic acid affinity moiety covalently attached to a nucleic acid molecule complementary to a locus of interest to recover a nucleic acid molecule of interest. Non-limiting examples of non-nucleic acid affinity moieties include biotin, digoxigenin, and dinitrophenol. In some embodiments, the probes are attached to a solid surface or particle, such as a dipstick or magnetic bead, to recover a nucleic acid of interest. In some embodiments, the methods described herein include amplifying nucleic acids that bind to the probe population prior to further analysis, such as sequencing. Nucleic acid amplification methods, such as PCR, are known in the art.
[0146] An enriched probe set for a TCR / BCR gene region (TCR / BCR gene panel) may include probes targeting one or more TCR and / or BCR genes or gene regions. The probes may target TCR and BCR gene segments located in the V, D, J, and constant regions. The probes may target gene segments responsible for the TCR alpha, beta, gamma, and delta chains. The probes may target gene segments responsible for the BCR kappa, lambda, and heavy chains, as well as various B cell receptor constant region isotype variants (such as IgM, IgG, IgA, IgD, and IgE).
[0147] The target in the constant region may be adjacent to the V / D / J recombination site. For example, the target may exclude a 1200 bp region downstream of the VDJ region. For example, the probe design may be configured to exclude all but the most proximal probe covering each of the constant regions, e.g., all but the two, three, four, or five most proximal probes. In some embodiments, the probe design is configured to exclude all but the three most proximal probes. This configuration provides sufficient signal to capture RNA fragments containing the VDJ fork and to identify and distinguish different constant regions from one another to determine whether the region is associated with IgG, IgM, or IgA. Annotated sequences are known in the art. See, for example, the IGMT database at http: / / www.imgt.org.
[0148] In some embodiments, a TCR / BCR gene panel (e.g., a panel of hybrid capture probes directed to portions of one or more TCR genes and one or more BCR genes) is used. PoolThe target genes for TCR / BCR enrichment via the panel including IGKV1OR1-1, IGKV2-18, IGKV3OR2-268, IGKC, IGKJ5, IGKJ4, IGKJ3, IGKJ2, IGKJ1, IGKV4-1, IGKV5-2, IGKV7-3, IGKV2-4, IGKV1-5, IGKV1-6, IGKV3-7, IGKV1-8, IGKV1-9, IGKV3-11, IGKV1-12, IGKV1-13, IGKV3-15, IGKV1-16, IGKV1-17, IGKV3-20, IGKV6-21, IGKV6-22, IGKV6-23, IGKV6-24, IGKV6-25, IGKV6-26, IGKV6-27, IGKV6-28, IGKV6-29, IGKV6-30, IGKV6-31, IGKV6-32, IGKV6-33, IGKV6-34, IGKV6-35, IGKV6-36, IGKV6-37, IGKV6-38, IGKV6-39, IGKV6-40, IGKV6-41, IGKV6-42, IGKV6-43, IGKV6-44, IGKV6-45, IGKV6-46, IGKV6-47, IGKV6-48, IGKV6-49, IGKV6-50, IGKV6-51, IGKV6-52, IGKV6-5 V2-24, IGKV1-27, IGKV2-28, IGKV2-29, IGKV2-30, IGKV1-33, IGKV1-37, IGKV1-39, IGKV2-40, IGKV2D-40, IGKV1D-39, IGKV1D-37, IGKV1D-33, IGKV2D -30,IGKV2D-29,IGKV2D-28,IGKV2D-26,IGKV2D-24,IGKV6D-21,IGKV3D-20,IGKV2D-18,IGKV6D-41,IGKV1D-17,IGKV1D-16,IGKV3D-15,IGKV1D-13,I GKV1D-12, IGKV3D-11, IGKV1D-42, IGKV1D-43, IGKV1D-8, IGKV3D-7, IGKV1OR2-118, IGKV1OR2-1, IGKV1OR2-2, IGKV1OR2-3, IGKV1OR2-9, IGKV2OR2-7 D, IGKV1OR2-11, IGKV1OR2-108, TRGC2, TRGJ2, TRGJP2, TRGC1, TRGJP, TRGJP1, TRGV11, TRGV10, TRGV9, TRGVA, TRGV8, TRGV5P, TRGV5, TRGV4, TRGV3, TRG V2, TRGV1, TRBV1, TRBV2, TRBV3-1, TRBV4-1, TRBV5-1, TRBV6-1, TRBV7-1, TRBV4-2, TRBV6-2, TRBV7-2, TRBV6-4, TRBV7-3, TRBV5-3, TRBV9, TRBV10-1, TRBV11-1, TRBV12-1, TRBV10-2, TRBV11-2, TRBV12-2, TRBV6-5, TRBV7-4, TRBV5-4, TRBV6-6, TRBV5-5, TRBV6-7, TRBV7-6, TRBV5-6, TRBV6-8, TRBV7-7,TRBV5-7、TRBV7-9、TRBV13、TRBV10-3、TRBV11-3、TRBV12-3、TRBV12-4、TRBV12-5、TRBV14、TRBV15、TRBV16、TRBV17、TRBV18、TRBV19、TRBV20-1、TRBV21-1、TRBV23-1、TRBV24-1、TRBV25-1、TRBV26、TRBV27、TRBV28、TRBV29-1、TRBD1、TRBJ1-1、TRBJ1-2、TRBJ1-3、TRBJ1-4、TRBJ1-5、TRBJ1-6、TRBC1、TRBJ2-1、TRBJ2-2、TRBJ2-2P、TRBJ2-3、TRBJ2-4、TRBJ2-5、TRBJ2-6、TRBJ2-7、TRBC2、TRBV30、IGLV8OR8-1、TRBV20OR9-2、TRBV21OR9-2、TRBV23OR9-2、TRBV24OR9-2、TRBV26OR9-2、TRBV29OR9-2、IGKV1OR9-2、IGKV1OR-2、IGKV1OR9-1、IGKV1OR-3、IGKV1OR10-1、IGHG2、TRAV1-1、TRAV1-2、TRAV2、TRAV3、TRAV4、TRAV5、TRAV6、TRAV7、TRAV8-1、TRAV9-1、TRAV10、TRAV11、TRAV12-1、TRAV8-2、TRAV8-3、TRAV13-1、TRAV12-2、TRAV8-4、TRAV13-2、TRAV14DV4、TRAV9-2、TRAV12-3、TRAV8-6、TRAV16、TRAV17、TRAV18、TRAV19、TRAV20、TRAV21、TRAV22、TRAV23DV6、TRDV1、TRAV24、TRAV25、TRAV26-1、TRAV8-7、TRAV27、TRAV29DV5、TRAV30、TRAV26-2、TRAV34、TRAV35、TRAV36DV7、TRAV38-1、TRAV38-2DV8、TRAV39、TRAV40、TRAV41、TRDV2、TRDD1、TRDD2、TRDD3、TRDJ1、TRDJ4、TRDJ2、TRDJ3、TRDC、TRDV3、TRAJ61、TRAJ60、TRAJ59、TRAJ58、TRAJ57、TRAJ56、TRAJ55、TRAJ54、TRAJ53、TRAJ52、TRAJ51、TRAJ50、TRAJ49、TRAJ48、<h2 style=";text-align:left;direction:ltr">TRAJ47,TRAJ46,TRAJ45,TRAJ44,TRAJ43,TRAJ42,TRAJ41,TRAJ40,TRAJ39,TRAJ38,TRAJ37,TRAJ36,TRAJ35,TRAJ34,TRAJ33,TRAJ32,TRAJ31,TRAJ3 0、TRAJ29、TRAJ28、TRAJ27、TRAJ26、TRAJ25、TRAJ24、TRAJ23、TRAJ22、TRAJ 21、TRAJ20、TRAJ19、TRAJ18、TRAJ17、TRAJ16、TRAJ14、TRAJ13、TRAJ12、TRAJ 11、TRAJ10、TRAJ9、TRAJ8、TRAJ7、TRAJ6、TRAJ5、TRAJ4、TRAJ3、TRAJ2、TRAJ 1、TRAC、IGHA2、IGHE、IGHG4、IGHA1、IGHG1、IGHG2、IGHG3、IGHD、IGHM、IGHJ6 IGHJ3P, IGHJ5, IGHJ4, IGHJ3, IGHJ2P, IGHJ2, IGHJ1, IGHHD7-27, IGHJ1P, IGHHD1-26, IGHHD6-25, IGHHD5-24, IGHHD4-23, IGHHD3-22, IGHHD2-21, IGHHD1-20, I GHD6-19、IGD5-18、IGHD4-17、IGD3-16、IGD2-15、IGD1-14、IGD6-13、 IGHD5-12、IGHD4-11、IGHD3-10、IGHD3-9、IGHD2-8、IGHD1-7、IGHD6-6、IGH D4-4, IGHHD3-3, IGHHD2-2, IGHHD1-1, IGHV6-1, IGHV1-2, IGHV1-3, IGHV4-4, IGHV7-4-1, IGHV2-5, IGHV3-7, IGHV3-64D, IGHV5-10-1, IGHV3-11, IGHV3-13 、IGHV3-15、IGHV3-16、IGHV1-18、IGHV3-19、IGHV3-20、IGHV3-21、IGHV3-2 2、IGHV3-23、IGHV1-24、IGHV3-25、IGHV2-26、IGHV4-28、IGHV3-32、IGHV3-3 0、IGHV3-30-2、IGHV4-31、IGHV3-29、IGHV3-33、IGHV3-33-2、IGHV4-34、IG HV7-34-1、IGHV3-35、IGHV3-38、IGHV4-39、IGHV7-40、IGHV3-43、IGHV1-45、IGHV1-46、IGHV3-47、IGHV3-48、IGHV3-49、IGHV5-51、IGHV3-52、IGHV3-53、IGHV3-54、IGHV4-55、IGHV1-58、IGHV4-59、IGHV4-61、IGHV3-62、IGHV3-63、IGHV3-64、IGHV3-66、IGHV1-68、IGHV1-69、IGHV2-70D、IGHV3-69-1、IGHV1-69-2、IGHV1-69D、IGHV2-70、IGHV3-71、IGHV3-72、IGHV3-73、IGHV3-74、IGHV5-78、IGHV7-81、IGHV1OR15-9、IGHV1OR15-2、IGHV3OR15-7、IGHV1OR15-1、IGHV1OR15-3、IGHV4OR15-8、IGHV1OR15-4、IGHV3OR16-9、IGHV2OR16-5、IGHV3OR16-15、IGHV3OR16-6、IGHV3OR16-10、IGHV3OR16-8、IGHV3OR16-12、IGHV3OR16-13、IGHV3OR16-16、IGHV1OR21-1、IGKV1OR22-5、IGKV2OR22-4、IGLV4-69、IGLV10-54、IGLV1-62、IGLV8-61、IGLV4-60、IGLV6-57、IGLV11-55、IGLV5-52、IGLV1-51、IGLV1-50、IGLV9-49、IGLV5-48、IGLV1-47、IGLV7-46、IGLV5-45、IGLV1-44、IGLV7-43、IGLV1-41、IGLV1-40、IGLV5-37、IGLV1-36、IGLV2-34、IGLV2-33、IGLV3-32、IGLV3-31、IGLV3-27、IGLV3-25、IGLV2-23、IGLV3-22、IGLV3-21、IGLV3-19、IGLV2-18、IGLV3-16、IGLV2-14、IGLV3-13、IGLV3-12、IGLV2-11、IGLV3-10、IGLV3-9、IGLV2-8、IGLV2-5、IGLV4-3、IGLV3-1、IGLJ1、IGLC1、IGLJ2、IGLC2、IGLJ3、IGLC3、IGLJ4、IGLJ5、IGLJ6、IGLC6、IGLJ7、IGLC7、TRBV3-2、TRBV4-3、TRBV6-9、TRBV7-8、and TRBV5-8.
[0149] Each probe may be designed to cover only the TCR and / or BCR region, or to cover both the TCR and / or BCR region and the non-TCR / BCR region.
[0150] In some embodiments, gene regions can be represented by gene name or Ensembl ID. Ensembl ID can be represented as ENSG or ENST. For example, the gene IGKV3OR2-268 would map to an Ensembl ID of ENSG00000233999-ENSG00000233999 or ENST00000421835-ENST00000421835.
[0151] In some embodiments, a probe set for TCR / BCR profiling according to the systems and methods disclosed herein is obtained according to the method of Example 1.
[0152] In some embodiments, the probes are Pool In some embodiments, Pool In some embodiments, the probe set is: 1) a BCR constant region group, 2) a BCR non-constant region group (VDJ), 3) a TCR constant region group, and 4) a TCR non-constant region group (VDJ). Pool In some embodiments, the probe set comprises at least one probe from each Pool 1 to 5 probes from each Pool Five to ten probes from each Pool 10 to 50 probes from each Pool In some embodiments, each probe comprises 100 to 200 probes from PoolThe number of probes in each group varies. For example, in one embodiment, the probes in each group are as follows: the TCR non-constant region group includes about 100-1000 probes, the TCR constant region group includes about 10-50 probes, the BCR non-constant region group includes about 500-2000 probes, and the BCR constant region group includes about 20-100 probes. In some embodiments, the probe set includes a TCR non-constant region group including about 650 probes, a TCR constant region group including about 18 probes, a BCR non-constant region group including about 894 probes, and a BCR constant region group including about 45 probes.
[0153] Probe concentration In some embodiments, the TCR / BCR hybrid capture probes may be included as part of a comprehensive genomic profiling panel. Examples include whole-exome / whole-transcriptome RNA-seq panels, targeted enrichment sequencing panels, whole-exome panels, whole-genome panels, whole-transcriptome panels, etc. In some embodiments, the probes may be included as part of a variety of Pool In some embodiments, Pool are: 1) BCR constant, 2) BCR non-constant (VDJ), 3) TCR constant, 4) TCR non-constant (VDJ). In some embodiments, a fifth panel or other panel including transcriptome targeting is included.
[0154] In some embodiments, the resulting probe population is Pool By way of example, for TCR / BCR enrichment and / or profiling, probes may be grouped or categorized as TCR non-constant region groups, TCR constant region groups, BCR non-constant region groups, and BCR constant region groups. PoolIn some embodiments, the number of probes in each group is the same, while in other embodiments, the number of probes in each group is different. In some embodiments, the number of probes in two or more groups is the same. By way of example only, in one embodiment, the probes in each group are as follows: the TCR non-constant region group has about 100-1000 probes, the TCR constant region group has about 10-50 probes, the BCR non-constant region group has about 500-2000 probes, and the BCR constant region group has about 20-100 probes. In some embodiments, the TCR non-constant region group has 650 probes, the TCR constant region group has 18 probes, the BCR non-constant region group has 894 probes, and the BCR constant region group has 45 probes.
[0155] In some embodiments, each probe used in a genomic profiling panel Pool The amount of the first probe containing the BCR constant region probe is characterized as a ratio. Pool , the second containing a BCR non-constant region probe Pool , a third containing a TCR constant region probe Pool , and a fourth containing a TCR non-constant region probe. Pool may be provided in a ratio of about 0.1-10:0.25-25:10-1000:10-1000, or about 0.5-5:1.25-12.5:50-500:50-500, or about 0.7-1.3:1.7-7.5:75-125:75-125, or about 1:2.5:100:100. In some embodiments, a fifth panel comprising an exome-targeted panel Pool In some embodiments, the first Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool , vs. 5th Pool The ratio is about 0.1-10:0.25-25:10-1000:10-1000:1-100, or about 0.5-5:1.25-12.5:50-500:50-500:5-50, or about 0.7-1.3:1.7-7.5:75-125:75-125:7-12, or about 1:2.5:100:100:10.
[0156] In some embodiments, the probes in the genomic profiling panel Pool Concentrations are characterized by concentrations in attomoles per probe per capture (i.e., per reaction well). By way of example and without limitation, in some embodiments, BCR stationary probes are about 0.25-25, 0.5-12.5, about 1-10, 1-5, or about 2.5 attomoles per probe per capture; BCR non-stationary probes are about 0.6-62.5, about 1-30, about 2-20, about 5-15, or about 6.25 attomoles per probe per capture; TCR stationary probes are about 25-500, or about 30-300, or about 100-300, or about 200-300, or about 250 attomoles per probe per capture; and TCR non-stationary probes are about 25-500, or about 30-300, or about 100-300, or about 200-300, or 250 attomoles per probe per capture.
[0157] By way of example and without limitation, in some embodiments, the BCR constant probe is about 2.5 attomoles / probe / capture, the BCR non-constant probe is about 6.25 attomoles / probe / capture, the TCR constant probe is about 250 attomoles / probe / capture, and the TCR non-constant probe is about 250 attomoles / probe / capture. In some embodiments, an exome probe is additionally used. In some embodiments, the exome probe is provided at a concentration of 25 attomoles / probe / capture, and thus in some embodiments, the TCR and BCR probes Pool exome Pool are used at concentrations of 0.1, 0.25, 10, and 10 times higher than those of the corresponding 200 mg / kg bw.
[0158] Lead Processing and Analysis The sequenced reads can be processed for further analysis, which in some embodiments may include one or more of alignment, assembly, annotation, and quantification steps.
[0159] In one example, the systems and methods disclosed herein receive an RNA-seq FASTQ file comprising the raw output from an NGS sequencing pipeline, which includes a list of all reads generated by the sequencer and the quality information associated with each read.
[0160] In an optional filtering step, the systems and methods may filter out amplification duplicates (e.g., PCR duplicates, two or more reads derived from the same source template, or the same nucleic acid molecule). In one example, the systems and methods may utilize unique molecular identifiers (UMIs) to filter out amplification duplicates. The systems and methods may filter out low-quality reads, or reads with a quality score below a selected threshold.
[0161] During the TCR / BCR gene sequence assembly process, the system and method may provide the RNA-seq FASTQ (forward and reverse read files) to a dedicated repertoire sequencing (rep-seq) aligner and / or gene sequence assembler, specifically a dedicated aligner designed for immune receptor quantification.
[0162] Examples of TCR and / or BCR gene sequence assembly methods are described in, for example, “Landscape of tumor-infiltrating T cell repertoire of human cancers” (Li et al., 2016, Nat. Genet., 48(7), pp. 725-732), “Landscape of B cell immunity and related immune evasion in human cancers” (Hu et al., 2019, Nat. Genet., 51(3), pp. 560-567), “BASIC: BCR assembly from single cells” (Canzar et al., 2017, Bioinformatics, 33(3), pp. 425-427), “Simultaneously inferring T cell fate and clonality from single cell transcriptomes” (Stubbington et al., 2015, BioRxiv https: / / doi.org / 10.1101 / 025676), and “Antigen receptor repertoire” (Allen et al., 2016, BioRxiv https: / / doi.org / 10.1101 / 025676), all of which are incorporated herein by reference. This is described in (Bolotin et al., 2017, Nat. Biotech., 35(10), 908-911)).
[0163] For example, using identified anchor reads with a first paired end that aligns to a TCR / BCR gene and a second paired end that does not, paired-end reads can be aligned to a predetermined immunological receptor gene sequence or to the entire genome. Alignment tools such as STAR or Kallisto can be used to align RNA-seq data in fastq files to references such as hg19 and GRCh37. See, for example, Nicolas L. Bray, Harold Pimentel, Pall Melsted, and Lior Pachter, "Near-optimal probabilistic RNA-seq quantification," Nature Biotechnology, 34, 525-527 (2016), doi:10.1038 / nbt.3519, incorporated herein by reference. See also https: / / pachterlab.github.io / kallisto / (California Institute of Technology, Pasadena, CA). For example, STAR can be used to generate RNA-seq data for deconvolution of immune CDR3 sequences. See Dobin et al., STAR: ultrafast universal RNA-seq aligner, Bioinformatics, January 2013, 29(1), pp. 15-21, which is incorporated herein by reference.
[0164] Reads can be filtered to those that map to BCR or TCR regions. In one example, there are three TCR regions, and in the case of the hg19 reference genome, coordinates can include TCRα (chr. 14:22,090,057-23,021,075), TCRβ (chr. 7:141,998,851-142,510,972), and TCRγ (chr. 7:38,279,625-38,407,656). Because the TCRδ gene region (chr. 14:22,891,537-22,935,569) is embedded in the TCRα region, reads for this region can be obtained together with TCRα reads. Counts of reads mapping to the non-constant regions of the TCRα and TCRβ chains can be used to estimate the usage of different genes and PCAs. In one example, there are three BCR regions, and in the case of the hg19 reference genome, for example, the coordinates may include IGH (chr. 14:106,032,614-107,288,051), IGK (chr. 2:89,890,568-90,274,235), and IGL (chr. 22:22,380,474-23,265,085).
[0165] Among all the mapped reads extracted in the above step, for a subset that has unmapped mates that may have been generated from the CDR3 region and could not be aligned to the reference genome, the reads in the BAM file can be screened to search for mates for each such read until all mapped reads in the TCR region are paired. Unmapped reads found in this step and thought to be associated with the CDR3 region may be used for CDR3 de novo assembly.
[0166] As another example, each read may be aligned to a predetermined immunological receptor gene sequence. A plurality of anchor windows may be identified, each window associated with a plurality of reads exceeding a threshold. Anchor sequences may be generated from the anchor reads using reads that align to a region of the anchor window, referred to as "anchor reads." The anchor window, anchor sequences, and unaligned reads may be provided to an assembly process to generate contig sequences. Each contig sequence may be annotated or otherwise associated with at least one immunological gene region segment selected from one of V, D, J, and C. Optionally, portions of each contig sequence located outside the CDR3 region may be removed. The number of contig sequences annotated and / or associated with each segment may be quantified.
[0167] In some embodiments, at least one read that aligns to a given immunological receptor gene sequence may be to a CDR3 region, e.g., to aid in recognition of one or more antigens. At least one read that aligns to a given immunological receptor gene sequence may be to a CDR3 adjacent region.
[0168] In one example, the systems and methods include an assembler that outputs unique receptor sequence non-constant regions, including immune receptor clonotypes. In various embodiments, the assemblers disclosed herein output a list of CDR3 sequences at their nucleotide level, along with read counts (e.g., the number of reads associated with each CDR3 sequence). In alternative embodiments, the output sequences may correspond to all or part of the CDR, CDR2, and / or CDR3 portions of TCR and / or BCR genes. In various embodiments, the assembler may output data associated with zero to tens of thousands of different CDR3 sequences assembled from sequencing reads. In one example, complementarity determining region 3 (CDR3) is the region where immune genes recombine during the VDJ recombination process.
[0169] Specimens may have different numbers of immune cells. Some specimens, such as glioblastoma specimens, may have little or no immune infiltrate. For such specimens, the number of CDR3 sequences output will be very low. Other specimens or tumors will have large numbers of immune cells. In some cases, specimens include immune cell cancers or specimens from patients with an active adaptive immune response to an infection or other disease state. As another example, consider a large immune cell population derived from a lymph node that has a highly diverse collection of CDR3 sequences (e.g., 5,000, 10,000, 20,000, 30,000, etc. sequences).
[0170] In various embodiments, the output of the assembler is in the form of a table where each row represents a sequence, which may be several hundred nucleotides or even less than 100. For each sequence, the systems and methods can also output a confidence or quality metric.
[0171] In one example, each sequence can be associated with a read quantity, which in various embodiments may reflect the number of reads and / or the proportion of reads that align to that sequence out of all reads detected in a specimen.
[0172] The systems and methods return genetic segment identifiers that include CDR3 sequences, e.g., a list of genetic segments that most likely recombined to form a particular CDR3 sequence during VDJ recombination. In one embodiment, for each genetic segment (V segment, J segment, and, if applicable, D segment), the systems and methods return a list of multiple possible genetic segment identifiers, e.g., identifiers representing the top three most likely genetic segments.
[0173] In various embodiments, the systems and methods may filter sequences to remove sequences predicted to be non-productive, which may include sequences with detected frameshift mutations, premature stop codons, or partially assembled clonotypes.
[0174] The systems and methods can calculate secondary statistics based on the read counts associated with the sequences listed in the output table or filtered output table. In various embodiments, secondary statistical categories can include abundance (e.g., the number of distinct clonotypes or unique sequences detected in the sample and / or represented in the table) and evenness (e.g., whether the read counts for all clonotypes are roughly equal or to what extent the distribution of read counts is skewed toward one or a few clones). Example statistics include Shannon entropy, Simpson's index, GINI index, and the like. For examples of statistical methods that can be applied to output table data, see Bolotin et al., Nat Biotechnol 35, 908-911 (2017), https: / / doi.org / 10.1038 / nbt.3979, incorporated herein by reference in its entirety.
[0175] The selection of a statistical calculation may be based on various criteria. In one example, criteria may include the distribution of values calculated for multiple specimens in a database, the range of possible output values, and / or the reproducibility or similarity of statistics for technical or biological replicates. For example, a small distribution in the values of a secondary statistic across multiple specimens may make it difficult to distinguish one specimen from another. The distribution may be measured by various statistical methods. Regarding the range of possible values, Shannon entropy is not limited to the range of 0 to 1, which may be advantageous in various embodiments. In other embodiments, a statistic limited to a certain range (e.g., 0 to 1) may be desirable. Technical replicates may be multiple NGS runs on the same specimen, and biological replicates may be multiple slices from the same biopsy, and reproducibility may be calculated by comparing the values of the statistic calculated for each replicate. Comparisons may include standard deviation, standard error of the mean, etc.
[0176] The systems and methods can also determine the protein structure associated with each sequence. The systems and methods can cluster sequences according to the similarity of their associated protein structures. The systems and methods can also analyze protein structures, including antigens predicted to bind to TCRs or BCRs and / or human leukocyte antigen (HLA) / major histocompatibility complex (MHC) molecules, particularly antigens associated with a patient's disease state (e.g., antigens produced by specific pathogens during infection, neoantigens produced by cancer cells, antigens or allergens that cause allergic reactions, or antigens that cause autoimmune diseases). The analysis can include combining multiple sequences or predicting pairing of two or more sequences. Sequences can be predicted to be paired sequences from the same heterodimeric protein, e.g., heavy and light chain sequences, alpha and beta chain sequences, gamma and delta chain sequences, etc. In various embodiments, two sequences can be predicted to be associated with the same heterodimeric protein if the number of reads associated with each sequence is roughly equivalent. For example, if 30% of the detected heavy chain reads are sequence A and 28% of the detected light chain reads are sequence alpha, then sequences A and alpha can be predicted to be paired. These predicted pairings can be confirmed using single-cell sequencing or other methods of analyzing TCR or BCR gene and / or protein sequences. For examples of analysis of TCR or BCR protein structure, see Glanville et al., Nature 547, 94-98 (2017), https: / / doi.org / 10.1038 / nature22976, incorporated herein by reference in its entirety.
[0177] The systems and methods may include storing the TCR / BCR sequencing results in a database, and may associate the TCR and / or BCR sequences with additional molecular data (e.g., HLA sequences, genomic, transcriptomic, epigenomic, proteomic, metabolic, etc. data) and / or clinical data (e.g., demographic information, diagnostic data, disease severity, immune response, phenotypic, treatment response data, etc.). The systems and methods may access similar databases to determine whether the TCR / BCR sequences are associated with particular molecular or clinical data features (e.g., the presence of a variant in genomic data, or a particular response to a therapy or therapy category such as immunotherapy).
[0178] The system and method can be used to identify patients whose disease is likely to be therapeutically effective as reference information for the development of therapies based on antibodies, vaccines, CAR-Ts, CAR-NK, ImmTAC, etc. Pool This may involve discovering TCR or BCR sequences (individual sequences or groups of sequences) from a genome or from individual patient data.
[0179] The systems and methods may include designing experiments to test therapeutic responses associated with one or more detected sequences, which may be biochemical assays, organoid experiments, t-cell and organoid co-culture experiments, etc.
[0180] The system and method can be used to determine differential gene expression. One application of RNA-seq data, including data obtained using a TCR / BCR enrichment panel, is to identify genes that are expressed differently between two or more experimental groups. For example, RNA sequencing data can be used to identify genes that are significantly more or less expressed in patients (e.g., patients with cancer, autoimmune disease, infectious disease, allergy, and / or requiring transplantation) compared to healthy individuals. This can be achieved by performing statistical analysis to compare the normalized read counts of each gene across different experimental groups. The aim of this analysis is to determine whether the observed difference in read counts is significant, i.e., whether it is greater than the expected (enriched) difference compared to the difference due to natural random variation.
[0181] Several data processing steps may be performed to prepare raw sequencing data for analysis. Sequencing data are typically provided in FASTQ format, where each sequencing read is associated with a quality score. First, the data is processed to remove sequencing artifacts, such as adapter sequences and low-complexity reads. Based on the read quality scores, sequencing errors are identified and removed or corrected. Publicly available tools such as TagDust, SeqTrim, and Quake can be used to perform these "data grooming" steps.
[0182] In the next stage of data processing, the reads are aligned to a reference genome using an alignment tool. Several publicly available tools can be used for this step, including TopHat, Cufflinks, and Scripture. These programs can be used to reconstruct transcripts, identify variants, and quantify the expression levels of each transcript and gene.
[0183] After the reads are aligned and quantified, differential expression analysis can be performed. Commonly used statistical methods for differential expression analysis include those based on the negative binomial distribution (e.g., edgeR and DESeq) and Bayesian approaches based on the negative binomial model (e.g., baySeq and EBSeq).
[0184] report 1A and 1B are diagrams showing examples of reports.
[0185] The results of TCR / BCR profiling can be presented to the requesting clinician or other individual. Results can be provided in a number of formats, such as by gene or by segment. Results can also be aggregated. Results may include the clonality of the immune repertoire, such as the putative clonality of BCR or TCR sequences in the specimen. An example of a result is shown in Figures 1A and 1B in the form of a report excerpt. This excerpt is from a report on a cancer specimen, but similar reports can be generated for other disease states, infections, or medical conditions.
[0186] Summary tab This section may include multiple data fields and / or conclusions related to those data fields based on the TCR / BCR sequencing data, including estimated tumor purity (burden), estimated immune cell composition (proportion of immune cells in the sample, such as B cells, macrophages, T cells, CD4 T cells, CD8 T cells, CD8 T cell subtypes, NK cells, etc.), estimated immune infiltrate fraction, and immune receptor clonality fraction.
[0187] Immune Repertoire Tab This section contains immune repertoire profiles generated by utilizing TCR / BCR sequencing data. For hematologic malignancies, this profile can potentially be used to define and track dominant clonotypes, and this information or associated conclusions can be included in the report. For patients receiving CAR therapy, this profile can also be used to track the abundance of the CAR product over time, and this information or associated conclusions can be included in the report.
[0188] Bar graph: If most of the sequences detected in a specimen are clonal or from a single clone, it may suggest the expansion of a particular V(D)J combination. In various embodiments, a threshold value can be used to classify each CDR3 sequence as clonal, oligoclonal, or polyclonal. For example, a CDR3 sequence associated with fewer than 25 reads can be classified as polyclonal, a CDR3 sequence associated with 25-99 reads can be classified as oligoclonal, and a CDR3 sequence associated with 100 or more reads can be classified as clonal. In the bar graph, the percentages associated with each category refer to the proportion of reads associated with the CDR3 sequence that fall into that category. In another example, if one dominant clone is expected for a certain disease state, the CDR3 sequence of that clone would be the only sequence classified as clonal. CDR3 sequence listing: The report may include a table or list showing the most common VDJ or VJ combinations (in one example, there may be one / a few V(D)J combinations that account for the majority of sequences detected in a specimen).
[0189] In this example, the most common heavy chain sequence accounts for approximately 40% of the detected heavy chain sequences, and the most common light chain sequence accounts for approximately 35% of the detected light chain sequences. These two sequences account for an equal proportion of the total reads for each chain type, and based on this equal proportion, they can be predicted to pair in the same protein heterodimer.
[0190] The report may include any statistical interpretations calculated from the read counts. For example, if a sequence skews homogeneity such that the proportion of reads exceeds a read threshold (e.g., 10%, 20%, or 50% or more of the reads), this may indicate an expanded immune cell population. Sequences associated with read counts above the read threshold may suggest TCRs or BCRs that recognize or bind to infectious pathogens (or antigens derived from a pathogen), allergens, neoantigens, or cancer cells. The report may also include antigens and / or HLA sequences predicted to bind to a given TCR or BCR sequence or sequence combination, and may further include any associations between antigens and genomic data associated with the specimen.
[0191] The report may include treatments and / or clinical trials matched to the patient (or organoid) based on the TCR / BCR profile. For example, suitable treatments and / or clinical trials include adoptive cell therapy, cancer vaccines, immuno-oncology drugs, immunotherapy, checkpoint blockade, immune checkpoint inhibitors, chemotherapy, cancer-specific treatments, vaccines, antivirals, antibiotics, antiparasitics, antifungals, one or more antibodies (which may be monoclonal, polyclonal, etc., and may be isolated from another patient after recovery from an infection), antihistamines, nasal sprays, anti-leukotrienes, leukotriene modifiers, leukotriene receptor antagonists, allergy injections or other methods of inducing an isotype that switches allergic IgE to a more tolerable IgG, anti-inflammatory treatments, steroids, oral corticosteroids, prednisone, anti-rheumatic drugs (DMARDS), biologics targeting common anti-inflammatory pathways, TNF pathway antagonists (including Remicade), B cell depletion (including Rituxan), immunosuppressants, insulin, bone marrow transplant, anti-inflammatory diets, physical therapy, surgery, topical medications, and / or topical scalp medications.
[0192] The report may include a conclusion related to CAR-T cell, CAR-NK cell, CAR-M cell, another CAR cell, or ImmTAC cell monitoring (e.g., whether CAR cells are present in large numbers) in the patient based on the sequences detected. The report may include a conclusion related to hemo-cancer (e.g., lymphoid or myeloid cancer, lymphoma, etc.) and / or minimal or measurable residual disease (MRD) status based on the expanded immune cells detected in the patient.
[0193] The report may include predicted therapeutic responses in relation to the TCR or BCR sequences detected in the specimen, for example, predicted immunotherapeutic responses based on infiltrating lymphocytes detected or predicted to be present in the tumor specimen.
[0194] A report may exclude sequences for a variety of reasons, for example, if a sequence is found to be of no relevance to the patient's disease state, the sequence may not be included in the report.
[0195] In one example, a patient may have a genomic alteration that is a documented antigen or neoantigen. In one example, a TCR / BCR profile can be generated for a patient with colon cancer that is a KRAS P12D alteration and an HLA C08.02 allele known to represent this altered KRAS peptide. The TCR / BCR profile can be analyzed for CDR3 sequences predicted to recognize the altered KRAS peptide (see the World Wide Web and the NCBI NLM database at nih.gov / pmc / articles / PMC5178827 / ).
[0196] In one example, a TCR / BCR profile can be generated for a patient with multiple myelomas and a RAS mutation, and the TCR / BCR profile can be analyzed for CDR3 sequences predicted to recognize altered RAS peptides.
[0197] In one example, a patient's TCR / BCR profile suggests that the patient's repertoire is skewed (e.g., 90% of sequencing reads are associated with the top clone). This patient can be monitored with longitudinal studies to determine whether the clones are consistent over time. In one example, the top clone is associated with 50% of sequencing reads at time x and only 20% of sequencing reads at a later time point, which may indicate that the patient's current therapy is having some efficacy. If the clones are below the detection limit of the systems and methods disclosed herein, the patient's report may suggest a sensitive follow-up MRD assay and may include information about confounding factors (including biopsy site, sample variation, etc.).
[0198] The reports may include various data visualizations, particularly of repertoire sequencing (rep-seq) data, immunological profiling data, and / or TCR or BCR sequencing data. Examples include Circos plots, heat maps, or histograms / distribution charts (e.g., number of reads associated with each V, D, or J gene family, number of instances of amino acids in the primary protein structure predicted from the TCR or BCR sequence, rearrangement rate versus CDR3 length, IgG / IgM / etc. subdivisions, etc.), box plots (e.g., showing diversity scores or mutation frequencies for various samples or groups of samples), transition tables showing the frequency of each possible base (nucleotide) change, plots showing the genetic location of mutations (base changes), etc. For data visualization related to rep-seq data, see, e.g., IJ Speeert et al., J Immunol, 2017, 198:4156-4165, doi:10.4049 / jimmunol.1601921, and Ni Q, Zhang J, Zheng Z, Chen G, Christian L, Gronholm J, Yu H, Zhou D, Zhuang Y, Li QJ, and Wan Y, (2020), "VisTCR: An Interactive Software for T Cell Repertoire Sequencing Data Analysis," Front. Genet. 11:771, doi:10.3389 / fgene.2020.00771, the contents of each of which are incorporated by reference in their entirety for all purposes.
[0199] The report may include predicted antigens or epitopes recognized by the TCR or BCR sequences contained in the report. These predicted antigens or epitopes can be used in vaccine development. For example, the most prevalent antigens or epitopes can be included as part of a vaccine that would further include an adjuvant.
[0200] For example, coronavirus epitopes (antigens recognized by BCR or TCR) may include those listed in Table 1.
[0201] [Table 1]
[0202] In this table, each row represents a SARS-CoV-2 peptide and the corresponding SARS-CoV-1 peptide that can be recognized by a T cell receptor. The table includes information on the T cell type (CD4 or CD8) of the TCR that recognizes the peptide and the protein / amino acid position of the peptide's origin within the viral protein. See Le Bert, N., Tan, A.T., Kunasegaran, K., et al., "SARS-CoV-2-specific T cell immunity in cases of COVID-19 and SARS, and uninfected controls," Nature 584, 457-462 (2020), https: / / doi.org / 10.1038 / s41586-020-2550-z, the contents of which are incorporated herein by reference in their entirety for all purposes.
[0203] Table 2 includes human coronavirus peptides and the corresponding amino acid position and source viral protein of each peptide. Peptides listed in the same row are homologous peptides from different coronaviruses. See Mateus et al., DOI: 10.1126 / science.abd3871, the entire contents of which are incorporated herein by reference for all purposes. In this table, "VP" in column 1 refers to the viral protein, and "1st AA" in column 2 refers to the first amino acid position.
[0204] [Table 2]
[0205] Table 3 includes SARS-CoV-2 peptides, the corresponding amino acid position and source viral protein for each peptide, and the HLA match for each peptide. See Sekine et al., "Robust T cell immunity in convalescent individuals with asymptomatic or mild COVID-19," Cell (2020), doi: https: / / doi.org / 10.1016 / j.cell.2020.08.017, the contents of which are incorporated herein by reference in their entirety for all purposes.
[0206] [Table 3]
[0207] Additional coronavirus peptides are described in the scientific literature. See, e.g., Dijkstra JM and Hashimoto K., "Expected immune recognition of COVID-19 virus by memory from earlier infections with common coronaviruses in a large part of the world population [2nd ed., peer-reviewed 2 accepted]," F1000Research 2020, 9:285, https: / / doi.org / 10.12688 / f1000research.23458.2, and Peng et al., "Broad and strong memory CD4+ and CD8+ T cells induced by SARS-CoV-2 in UK convalescent COVID-19 patients," bioRxiv2020.06.05.134551, (2020), Pmid:32577665, the contents of which are incorporated herein by reference in their entirety for all purposes.
[0208] The above-described methods and systems can be utilized in conjunction with or as part of a digital laboratory healthcare platform for medical and research purposes generally. It should be understood that numerous uses of the above-described methods and systems are possible in conjunction with such a platform. One example of such a platform is described in U.S. Patent Application No. 16 / 657,804, filed October 18, 2019, entitled "Data Based Cancer Research and Treatment Systems and Methods," which is incorporated herein by reference in its entirety for all purposes.
[0209] For example, an implementation of one or more embodiments of the above-described methods and systems may include microservices that comprise a digital laboratory healthcare platform supporting TCR / BCR profiling. Embodiments may include a single microservice for performing and providing TCR / BCR profiling, or multiple microservices, each with a specific role that collectively performs one or more of the above-described embodiments. In one example, a first microservice may perform TCR / BCR profiling so that the profile results can be provided to a second microservice for reporting.
[0210] When the above embodiments are implemented in one or more microservices, or in combination with or as part of a digital laboratory healthcare platform, one or more of the microservices can be part of an order management system that coordinates the mechanisms of events as necessary at the appropriate times and in the appropriate order necessary to instantiate the above embodiments. A microservices-based order management system is disclosed, for example, in U.S. Patent Application No. 16 / 927,976, filed July 13, 2020, entitled "Adaptive Order Fulfillment and Tracking Methods and Systems," which is incorporated herein by reference in its entirety for all purposes.
[0211] For example, while continuing with the first and second microservices described above, the request management system can receive a request for RNA sequencing and notify the first microservice that it is ready for processing. When the first microservice is ready to provide the RNA sequencing to the second microservice, it can perform the RNA sequencing and notify the request management system. The request management system can also verify that the execution parameters (requirements) for the second microservice are met, including that the first microservice is completed, and notify the second microservice that it may continue processing the request to provide a completed RNA report according to the embodiments described above.
[0212] When the digital laboratory medical platform further includes a genetic analyzer system, the genetic analyzer system may include a targeted panel and / or sequencing probes. Examples of targeted panels are disclosed in U.S. Patent Application Nos. 16 / 789,288 and 15 / 930,234, filed February 12, 2020, and May 12, 2020, respectively, which are incorporated by reference in their entireties for all purposes. In one example, the targeted panel may provide next-generation sequencing results, according to the above embodiments, for genes, including immunological genes (e.g., TCR and BCR genes), that exhibit a high degree of sequence variability between individuals and / or between cells of an individual. An example of next-generation sequencing probe design is disclosed in U.S. Patent Application No. 17 / 706,704, filed October 21, 2020, entitled "Systems and Methods for Next Generation Sequencing Uniform Probe Design," which is incorporated herein by reference in its entirety for all purposes.
[0213] If the digital laboratory medicine platform further includes a bioinformatics pipeline, the above-described methods and systems can be utilized after the systems and methods utilized in the bioinformatics pipeline are completed or substantially completed. For example, the bioinformatics pipeline can receive next-generation gene sequencing results and return a series of binary files, such as one or more BAM files, reflecting DNA and / or RNA read counts aligned to a reference genome. The above-described methods and systems can be utilized, for example, to ingest DNA and / or RNA read counts and generate TCR / BCR sequence profiling results.
[0214] If the digital laboratory medicine platform further includes an RNA data normalizer, any RNA read counts can be normalized before processing in the above embodiments. An example of an RNA data normalizer is disclosed, for example, in U.S. Patent Application No. 16 / 581,706, filed September 24, 2019, entitled "Methods of Normalizing and Correcting RNA Expression Data," which is incorporated herein by reference in its entirety for all purposes.
[0215] When the digital laboratory medical platform further includes a genetic data deconvolutor, the systems and methods for deconvolution can be utilized to analyze genetic data associated with a specimen having two or more biological components to determine the contribution of each component to the genetic data and / or to determine what genetic data would be expected to be associated with any component of the specimen if that component were purified. Examples of genetic data deconvolutors are disclosed, for example, in U.S. Patent Application Nos. 16 / 732,229 and PCT / US19 / 69161, both filed December 31, 2019, entitled "Transcriptome Deconvolution of Metastatic Tissue Samples," and U.S. Patent Application No. 17 / 074,984, filed October 20, 2020, entitled "Calculating Cell-type RNA Profiles for Diagnosis and Treatment," which are incorporated by reference in their entireties for all purposes.
[0216] If the digital laboratory medicine platform further includes an automated RNA expression caller, the RNA expression levels can be adjusted to be expressed relative to a reference expression level, often to prepare multiple RNA expression datasets for analysis to avoid artifacts when the datasets are not generated using the same methods, equipment, and / or reagents. An example of an automated RNA expression caller is disclosed, for example, in U.S. Patent Application No. 17 / 112,877, filed December 4, 2020, and entitled "Systems and Methods for Automating RNA Expression Calls in a Cancer Prediction Pipeline," which is incorporated herein by reference in its entirety for all purposes.
[0217] The digital laboratory medicine platform may further include one or more insight engines that provide information, characteristics, or decisions related to disease states, which may be based on genetic and / or clinical data associated with a patient and / or specimen. Exemplary insight engines include a tumor of unknown origin engine, a human leukocyte antigen (HLA) loss of homozygosity (LOH) engine, a tumor mutation burden engine, a PD-L1 status engine, a homologous recombination deletion engine, a cellular pathway activation reporting engine, an immune infiltration engine, a microsatellite instability engine, a pathogen infection status engine, and the like. An example of a tumor of unknown origin engine is disclosed in U.S. Patent Application No. 15 / 930,234, filed May 12, 2020, and entitled "Systems and Methods for Multi-Label Cancer Classification," which is incorporated herein by reference in its entirety for all purposes. An example of an HLA LOH engine is disclosed, for example, in U.S. Patent Application No. 16 / 789,413, filed February 12, 2020, entitled "Detection of Human Leukocyte Antigen Class I Loss of Heterozygosity in Solid Tumor Types by NGS DNA Sequencing," which is incorporated herein by reference in its entirety for all purposes. An example of a tumor mutation burden (TMB) engine is disclosed, for example, in U.S. Patent Application No. 16 / 789,288, filed February 12, 2020, entitled "Targeted-Panel Tumor Mutational Burden Calculation Systems and Methods," which is incorporated herein by reference in its entirety for all purposes.An example of a PD-L1 status engine is disclosed, for example, in U.S. patent application Ser. No. 16 / 888,357, entitled "A Pan-Cancer Model to Predict The PD-L1 Status of a Cancer Cell Sample Using RNA Expression Data and Other Patient Data," filed May 29, 2020, which is incorporated by reference in its entirety for all purposes. An additional example of a PD-L1 status engine is disclosed, for example, in U.S. patent application Ser. No. 16 / 830,186, entitled "Determining Biomarkers from Histopathology Slide Images," filed March 25, 2020, which is incorporated by reference in its entirety for all purposes. An example of a homologous recombination deficiency engine is disclosed, for example, in U.S. Patent Application No. 16 / 789,363 and PCT Application No. PCT US20 / 18002, both filed February 12, 2020, entitled "An Integrative Machine-Learning Framework to Predict Homologous Recombination Deficiency," which are incorporated herein by reference in their entireties for all purposes. An example of a cellular pathway activation reporting engine is disclosed in U.S. Patent Application No. 16 / 994,315, filed August 14, 2020, entitled "Systems And Methods For Detecting Cellular Pathway Dysregulation In Cancer Specimens," which are incorporated herein by reference in their entireties for all purposes.An example of an immune infiltration engine is disclosed, for example, in U.S. Patent Application No. 16 / 533,676, filed August 6, 2019, entitled "A Multi-Modal Approach to Predicting Immune Infiltration Based on Integrated RNA Expression and Imaging Features," which is incorporated herein by reference in its entirety for all purposes. An additional example of an immune infiltration engine is disclosed, for example, in U.S. Patent Application No. 62 / 804,509, filed February 12, 2019, entitled "Comprehensive Evaluation of RNA Immune System for the Identification of Patients with an Immunologically Active Tumor Microenvironment," which is incorporated herein by reference in its entirety for all purposes. An example of an MSI engine is disclosed, for example, in U.S. patent application Ser. No. 16 / 653,868, filed October 15, 2019, entitled "Microsatellite Instability Determination System and Related Methods," which is incorporated herein by reference in its entirety for all purposes. An additional example of an MSI engine is disclosed in U.S. patent application Ser. No. 16 / 945,588, filed July 31, 2020, entitled "Systems and Methods for Detecting Microsatellite Instability of a Cancer Using a Liquid Biopsy," which is incorporated herein by reference in its entirety for all purposes.
[0218] If the digital laboratory medical platform further includes a report generation engine, the above methods and systems can be utilized to generate a summary report of the patient's genetic profile and the results of one or more insight engines for presentation to a physician. For example, the report may provide the physician with information regarding the extent to which a sequenced specimen contains tumor or normal tissue from a first organ, a second organ, a third organ, etc. For example, the report may provide a genetic profile for each tissue type, tumor, or organ in the specimen. The genetic profile may represent the gene sequences present in the tissue type, tumor, or organ, and may include information about variants, expression levels, gene products, or other information obtained from the genetic analysis of the tissue, tumor, or organ. The report may include appropriate therapies and / or clinical trials based on some or all of the genetic profile or insight engine findings and summaries. For example, clinical trials may be adapted according to the systems and methods disclosed in U.S. Patent Application No. 16 / 889,779, filed June 1, 2020, entitled "Systems and Methods of Clinical Trial Evaluation," which is incorporated herein by reference in its entirety for all purposes.
[0219] The report may include a comparison of the results against a database of results from multiple specimens. An example of a method and system for comparing results against a database of results is disclosed in U.S. Patent Application No. 16 / 732,168 and PCT / US19 / 69149, both filed December 31, 2019, entitled "A Method and Process for Predicting and Analyzing Patient Cohort Response, Progression and Survival," which are incorporated herein by reference in their entireties for all purposes. This information, optionally in conjunction with similar information and / or clinical response information from additional specimens, can lead to biomarker discovery or clinical trial design.
[0220] If the digital laboratory medicine platform further includes applying one or more embodiments described herein to organoids developed in connection with the platform, the methods and systems can be used to further evaluate genetic sequencing data obtained from the organoids to provide information regarding the extent to which the sequenced organoids contain a first cell type, a second cell type, a third cell type, etc. For example, the report can provide a genetic profile of each cell type in the specimen. The genetic profile can represent the gene sequences present in any cell type, or can include information about variants, expression levels, gene products, or other information obtained from genetic analysis of the cells. The report can also include tailored therapies based on some or all of the deconvoluted information. These therapies can be tested on the organoids, their derivatives, and / or similar organoids to determine the organoids' susceptibility to those therapies. For example, organoids can be cultured and tested according to the systems and methods disclosed in U.S. patent application Ser. No. 16 / 693,117, entitled "Tumor Organoid Culture Compositions, Systems, and Methods," filed November 22, 2019; PCT / US20 / 56930, entitled "Systems and Methods for Predicting Therapeutic Sensitivity," filed October 22, 2020; and U.S. patent application Ser. No. 17 / 114,386, entitled "Large Scale Phenotypic Organoid Analysis," filed December 7, 2020, which are incorporated by reference in their entireties for all purposes.
[0221] Where the digital laboratory medicine platform further includes applying one or more of the above in combination with or as part of a medical device or laboratory-developed test for medical and research purposes generally, the results of the laboratory-developed test or medical device can be enhanced and personalized through the use of artificial intelligence. One example of a laboratory-developed test, particularly a test that can be enhanced by artificial intelligence, is disclosed, for example, in U.S. Provisional Patent Application No. 62 / 924,515, filed October 22, 2019, entitled "Artificial Intelligence Assisted Precision Medicine Enhancements to Standardized Laboratory Diagnostic Testing," which is incorporated herein by reference in its entirety for all purposes.
[0222] It should be understood that the above examples are illustrative and not limiting of the use of the systems and methods described herein in combination with a digital laboratory healthcare platform.
[0223] application The present disclosure provides methods for analyzing the clonal number and distribution of T cell receptors (TCRs) and B cell receptors (BCRs). The sequences encoding TCRs and BCRs contain a variety of information useful for clinical and research applications. For example, immune profiling can be performed to determine the clonality of T and B cell repertoires.
[0224] In one example, T and B cells specific for the pathogen SARS-CoV-2 are activated and expanded after infection without apparent symptoms. Immune profiling of individuals in such cases is expected to reveal an expansion of SARS-CoV-2-specific lymphocytes. Furthermore, the humoral immune response to SARS-CoV-2 has been shown to wane over time, resulting in fewer SARS-CoV-2-specific antibodies remaining in the circulation (Self WH et al., MMWR Morb Mortal Wkly Rep 2020, 69:1762-1766). This feature of SARS-CoV-2 infection (COVID-19) potentially reduces the validity of tests for SARS-CoV-2 exposure based on virus-specific antibody titers.
[0225] For example, immune profiling is also important in detecting T and B cell lymphomas, as these cancers typically have dominant clones that emerge and expand as the cancer progresses. The presence of dominant T or B cell clones can be assessed in individuals to help determine the extent or severity of the disease.
[0226] However, what is lacking in conventional immune profiling assays is the ability to assess the molecular phenotype of cells of interest. Therefore, the technology of the present application combines complete next-generation sequencing of DNA or RNA-based samples with immune profiling. Previously, performing RNA or exome sequencing and immune profiling on a sample required splitting the sample material into two separate assays and integrating the data after sequencing. The method of the present application allows for analysis of both genomic / transcriptomic data and immune profiling in a single assay, without compromising the quality of the data from either component. Therefore, the method of the present application possesses exceptional efficiency that translates into precision medicine at a scale that is viable for routine use by medical practitioners for a variety of potential applications.
[0227] The disclosed method utilizes hybrid capture probes to enrich for sequences most relevant to understanding the T and B cell repertoire in an individual subject. Novel probes are designed to span the constant and non-constant regions of TCR and BCR sequences. Probe sets are designed to provide deep sequencing in critical areas of TCR and BCR sequences, enabling the development of complete immune profiles with fewer reads than conventional assays. Additionally, probe sets are formulated to provide productive sequences that cover both TCRs and BCRs. Furthermore, using the disclosed method, probe sets can be further fine-tuned for individual applications to provide maximum coverage of the immune repertoire. This novel hybrid capture approach allows immune profiling to be achieved while dedicating less than 2% of the reads to TCR / BCR profiling in any given sequencing run. As a result, over 98% of sequencing reads can be utilized for other applications, enabling high-quality, deep sequencing to be achieved simultaneously with immune profiling.
[0228] General In any given tumor, tissue, or blood sample, there are hundreds, thousands, tens of thousands, or even millions of different TCR and BCR sequences. These sequences can be used to predict, for example, past infection and potentially which T cells are killing tumor cells. While standard RNAseq can estimate the proportion of T cells in a tumor (infiltration), TCR sequencing can reveal whether the majority of T cells in a tumor are specific for a single neoantigen or are diverse. Pool Tracking TCRs and BCRs across patient cohorts can identify specific receptors that recur in patients with the same alterations, generating information that can be used to direct TCR-based / CAR cell therapies.
[0229] In the context of past or current infection, TCR and BCR sequencing results may be useful for characterizing the infection and / or immune response to the infection. When TCR / BCR sequencing is performed as part of a whole-exome RNA-seq assay, the RNA sequences and expression levels of various immune genes (e.g., cytokines, checkpoint molecules, innate immune genes) may also contribute to that characterization.
[0230] By way of example, but not limitation, TCR / BCR profiling results may be used to determine whether an individual has been in contact with one or more infectious pathogens, to detect whether an individual has TCR or BCR sequences associated with sterilizing immunity and / or neutralizing antibodies to a group of infectious pathogens or to a specific infectious pathogen, to identify adaptive immune responses to specific pathogens or antigens, to analyze and improve infectious disease treatment protocols for general patient populations or subpopulations, to identify associations between disease severity and immune profiles, and to identify and evaluate the efficacy and safety of TCR / BCR profiling. They can classify or predict the severity of a person's disease (see, e.g., SchultheiB et al., 2020, Immunity, https: / / doi.org / 10.1016 / j.immuni.2020.06.024, incorporated herein by reference in its entirety), assist physicians in selecting treatment protocols, adapting treatment protocols depending on an individual's immune response, developing and / or evaluating the efficacy of therapeutic or prophylactic treatments (e.g., vaccines), designing clinical trials or defining patient cohorts, and / or obtaining additional relevant information.
[0231] In some instances, serology tests can determine whether an individual has developed antibodies in response to infectious agents and / or antigens. However, it is known that not all infectious agents elicit a strong antibody (B cell) response or seroconversion in individual infections. These infections may be caused by pathogens whose life cycle occurs mostly within host cells (e.g., Listeria monocytogenes), or by viruses that do not cause viremia or are not found in high concentrations in an individual's blood. Examples of viruses that generally do not cause viremia include coronaviruses such as SARS, MERS, and SARS-CoV-2.
[0232] In some cases, infections that do not elicit a strong B cell response can still be controlled and resolved by individuals, and one hypothesized mechanism for this control in the absence of a B cell response is T cell responses (see Gallais et al., 2020, MedRxiv, https: / / doi.org / 10.1101 / 2020.06.21.20132449). Numerous assays (e.g., ELISpot, Fluorospot, ELISA, etc.) can be used to analyze an individual's T cell responses and / or memory B cells specific to a particular pathogen and / or antigen. However, these assays often require cell culture techniques and / or incubation periods, limiting the number of tests that can be performed each day. In various instances, TCR / BCR sequencing can be made more amenable to large-scale testing, allowing for multiple samples to be processed daily.
[0233] The T cell, B cell, and antibody assays described above detect TCRs and BCRs that respond to antigens included in the assay, but may not detect TCRs or BCRs that respond to antigens generated during infection, cancer, or other disease states not included in the assay. Also, these assays do not automatically provide the genetic sequence (and consequently, protein structure) of the BCR or TCR molecule, which is another advantage of the TCR / BCR sequencing methods disclosed herein.
[0234] Additional exemplary, non-limiting applications of the techniques of the present invention are set forth below.
[0235] Applications related to direct analysis of patient TCR / BCR profiles Disease testing - Cancer diagnosis and severity measurement / confirmation In some embodiments, patient samples, including blood or tumor samples, are collected and the severity of disease is assessed. By way of example, but not limitation, the severity of hematological malignancies, including T and B cell lymphomas, can be assessed by performing TCR / BCR hybrid capture and sequencing to generate an immune profile for the patient. The immune profile provides information about the clonality of normal and malignant cells. This information can be used by medical practitioners to better understand the tumor burden in patients and to help guide treatment decisions.
[0236] In some embodiments, therapy is recommended or tailored based on the TCR / BCR profile. The TCR / BCR profile provides information about the predominant clones that comprise a malignant tumor. Therefore, the TCR / BCR profile may serve as a reference for medical practitioners in making treatment decisions. By way of example, but not limitation, recommended therapies after TCR / BCR profiling include adoptive cell therapy / ACT, CAR-T cell therapy, chimeric antigen receptor macrophage (CAR-M) therapy, or other categories of cells engineered to express chimeric antigen receptors (CARs). Additional therapies include, but are not limited to, cancer vaccines, immuno-oncology drugs, immunotherapy, checkpoint inhibition, immune checkpoint inhibitors, chemotherapy, cancer-specific treatments, vaccines, antivirals, antibiotics, antiparasitics, antifungals, one or more antibodies (which may be monoclonal, polyclonal, etc., and may be isolated from another patient after recovery from an infection), antihistamines, nasal sprays, anti-leukotrienes, leukotriene modifiers, leukotriene receptor antagonists, allergy injections or other methods of inducing an isotype that switches allergic IgE to a more tolerated IgG, anti-inflammatory treatments, steroids, oral corticosteroids, prednisone, anti-rheumatic drugs (DMARDS), biologics targeting common anti-inflammatory pathways, TNF pathway antagonists (including Remicade), B cell depletion (including Rituxan), immunosuppressants, insulin, bone marrow transplant, anti-inflammatory diets, physical therapy, surgery, topical medications, and / or topical scalp medications.
[0237] In some embodiments, the technology of the present invention is used to simultaneously perform one or more of the following functions: assess the presence and degree of lymphocytic infiltration in solid tumor samples, measure / confirm disease severity, or detect infiltration biomarkers. TCR / BCR profiling of patient samples obtained from solid tumors provides information on the frequency and clonality of tumor-infiltrating lymphocytes (TILs). In some embodiments, therapy is recommended based on TIL analysis performed using TCR / BCR profiling. By way of example, but not limitation, treatments that may be recommended after TCR / BCR profiling include ACT, CAR-T, and / or other immuno-oncology (IO) therapies.
[0238] In some embodiments, TCR / BCR data are combined with other infiltration predictors (engines) and / or combined as a single feature to refine these predictive models. One example of an immune infiltration engine is disclosed, for example, in U.S. Patent Application No. 16 / 533,676, filed August 6, 2019, entitled "A Multi-Modal Approach to Predicting Immune Infiltration Based on Integrated RNA Expression and Imaging Features," which is incorporated herein by reference in its entirety for all purposes. An additional example of an immune infiltration engine is disclosed, for example, in U.S. Patent Application No. 62 / 804,509, filed February 12, 2019, entitled "Comprehensive Evaluation of RNA Immune System for the Identification of Patients with an Immunologically Active Tumor Microenvironment," which is incorporated by reference in its entirety for all purposes.
[0239] In one example, a TCR / BCR profile can be generated for a patient with non-small cell lung cancer (NSCLC) and an EGFR mutation. This TCR / BCR profile can be analyzed with or without output from an invasion predictor used to analyze the patient's data. These results can be used to match therapy (e.g., immunotherapy, checkpoint inhibition, etc.) to the patient.
[0240] Therapeutic Efficacy Testing In some embodiments, the technology of the present invention is used to identify whether therapeutic immune cells have infiltrated a target tumor. By way of example, and not limitation, cells detected using TCR / BCR profiling include CAR-T cells or cells provided through adoptive cell transfer (ACT) therapy. By way of example, and not limitation, additional probes can be added to specifically target sequences unique to particular therapeutic modalities.
[0241] In some embodiments, the technology of the present invention can be used to conduct longitudinal studies before and after the administration of immune-affecting therapies, such as chemotherapy. Chemotherapeutic drugs often have adverse effects on a patient's immune system. By way of example, and without limitation, chemotherapeutic drugs include anthracyclines such as doxorubicin and epirubicin, taxanes such as paclitaxel and docetaxel, 5-fluorouracil, cyclophosphamide, or carboplatin, or other drugs used to treat cell proliferation disorders. In some embodiments, the technology of the present invention can be used to monitor the extent and nature of adverse events or off-target effects of the immune repertoire. By way of example, and without limitation, TCR / BCR profiling can be used to understand whether a subject is more susceptible to infections or other diseases associated with an immunocompromised state. In some embodiments, TCR / BCR profiling can enable physicians to provide additional, directed therapy in response to TCR / BCR profiling analysis. By way of example, but not limitation, TCR / BCR profiling may lead to the administration of cytokines known to have a positive effect on the immune system, and in some embodiments, long-term TCR / BCR analysis provides data to determine the degree and nature of immune repertoire improvement and to modify treatment as necessary.
[0242] Sequencing, analysis and listing of cell and / or receptor sequences in patient samples In some embodiments, the techniques of the present invention are used to determine the TCR / BCR profile of a patient, or of a patient sample obtained, for example, from malignant tissue. In some embodiments, the methods are useful for determining the presence and degree of lymphocytic infiltration in tumor samples and for identifying the most abundant clones in tumor samples. By way of example, but not limitation, the profile may enable the selection of highly representative clones to be expanded for patient-specific adoptive cell transfer and / or used to identify highly expressed receptor non-constant regions to generate patient-specific chimeric receptors. This information provides the basis for developing personalized medical treatment approaches, for example, using CAR-T cell therapy.
[0243] In some embodiments, the technology of the present invention is used to determine the TCR / BCR profile of a patient suffering from or suspected of being infected with a pathogen. In some embodiments, the most abundant clone in a patient sample (e.g., a blood sample) is identified to be expanded for patient-specific adoptive cell transfer. In some embodiments, receptor non-constant regions for this patient sample are identified to generate patient-specific chimeric receptors for use in CAR cell therapy.
[0244] Diagnosis or confirmatory diagnosis of patients based on TCR / BCR analysis In some embodiments, TCR / BCR data can be used in combination with other pathogen detection or prediction methods and / or used as a feature to refine those prediction models. One example of a pathogen detection or prediction method is disclosed, for example, in U.S. Patent Application No. 16 / 802,126, filed February 26, 2020, which is incorporated by reference in its entirety for all purposes. Another example of a pathogen detection or prediction method is disclosed, for example, in PCT / US21 / 18619, filed February 18, 2021, which is incorporated by reference in its entirety for all purposes. In some embodiments, the data can be used to predict whether a patient will be protected from future infection. By way of example, but not limitation, this information may be beneficial for patients after vaccination or natural viral infection. In some embodiments, analyses can be performed longitudinally to characterize the immune response to pathogenic organisms. In the case of oncogenic pathogens, immune profiling can be used to predict whether a patient is already / has been infected, which can guide treatment decisions made by medical practitioners.
[0245] In some embodiments, a list of receptor sequences associated with a pathogen is provided (see Applications of the Technology above). In some embodiments, large datasets containing positive and negative controls are provided, for example, to determine which TCR / BCR sequences are associated with a given disease, pathogen, or antigen (see, e.g., Example 2). By way of example, but not limitation, the following is an example of utilizing the TCR / BCR profiling methods of the present invention for diagnostic or confirmatory diagnostic applications.
[0246] SARS-CoV-2 diagnosis In various embodiments, BCR sequences that recognize SARS-CoV-2 antigens may include those listed in Table 4, which provides examples of positive control data for SARS-CoV-2 contact and / or infection.
[0247] [Table 4]
[0248] Each row in Table 4 represents a BCR sequence, including the heavy chain V, D, and J family classifications and the heavy chain CDR3 amino acid sequence, and the light chain V and J family classifications and the light chain CDR3 amino acid sequence. The amino acid sequence may represent a consensus sequence of amino acids present in multiple CDR3 sequences when aligned and compared, and a tilde (~) may indicate a position in the sequence that does not share the same amino acid in the aligned CDR3 sequences. In various embodiments, two BCR sequences may be paired. For example, rows 3 and 4 may represent two alleles expressed by the same cell to create a heterodimeric protein BCR structure. Similarly, rows 5 and 6 may be paired sequences, rows 7 and 8 may be paired, rows 9 and 10 may be paired, and so on.
[0249] [Table 5A]
[0250] [Table 5B]
[0251] [Table 5C]
[0252] [Table 5D]
[0253] [Table 5E]
[0254] Table 5F
[0255] Table 5G
[0256]
Table 5H
[0257]
Table 5I
[0258]
Table 5J
[0259] Table 5K
[0260]
Table 5L
[0261] Table 5M
[0262]
Table 5N
[0263] Each row in Table 5 represents a BCR sequence, including the nucleotide and amino acid sequences of the heavy chain VDJ region and the light chain VJ region (see Robbiani et al., doi: https: / / doi.org / 10.1101 / 2020.05.13.092619, the contents of which are incorporated by reference in their entirety for all purposes).
[0264] The characteristics and phenotypes of SARS-CoV-2-specific T cells and the viral antigens (epitopes) they recognize have been described in the scientific literature, see, e.g., Weiskopf et al., "Phenotype and kinetics of SARS-CoV-2-specific T cells in COVID-19 patients with acute respiratory distress syndrome," Sci. Immunol. 5, eabd2071 (2020), doi:10.1126 / sciimmunol.abd2071pmid:32591408, the contents of which are incorporated herein by reference in their entirety for all purposes.
[0265] TCR and / or BCR sequences that recognize coronaviruses may further include sequences contained in the iReceptor database (https: / / gateway.ireceptor.org / login). For an example of a database containing data such as metadata, clinical data, coronavirus infection status, over 135,000 TCR sequences, coronavirus peptide sequences, and peptide-TCR binding pair data associated with 1,414 specimens, see Nolan et al., DOI: 10.21203 / rs.3.rs-51964 / v1, the contents of which are incorporated herein by reference in their entirety for all purposes.
[0266] Therefore, a diagnosis or confirmatory diagnosis of COVID or SARS-CoV-2 exposure can be provided based on TCR / BCR analysis.
[0267] While the above examples are directed to SARS-CoV-2, the techniques of the present invention are not so limited and can be used to diagnose, confirm, predict, or identify other diseases, infections, or conditions. By way of example, and not limitation, in some embodiments, exposure to one or more specific cancer types, infectious diseases such as influenza A, HIV, EBV, CMV, SARS-CoV-2, Lyme disease, allergies, and autoimmune diseases (such as diabetes, celiac disease, and psoriasis) is diagnosed or confirmed by determining the TCR / BCR profile of a subject sample.
[0268] Minimal Residual Disease (MRD) Testing In some embodiments, the techniques of the present invention are used to detect small numbers of residual tumor cells in patients undergoing treatment, for example, for hematological malignancies. In some embodiments, the methods of the present invention detect a single malignant T / B cell clone in 1,000, 10,000, or 100,000 cells. By way of example and not limitation, detection of a single cell in 1,000, 10,000, or 100,000 cells may be useful in helping medical practitioners decide whether to restart therapy or select a second-line therapy to treat the disease.
[0269] Applications of TCR / BCR profiling related to comparison of patient data with known cohort data Disease biomarkers and therapy development In some embodiments, multiple TCR / BCR profiles are generated and analyzed from a cohort of patients suffering from a particular disease, infection, or medical condition. In some embodiments, TCR / BCR profiles are provided from a cohort of patients suffering from a particular disease, infection, or medical condition, including positive controls, i.e., patients known to have been diagnosed with the particular disease, infection, or medical condition, and negative controls, i.e., patients known not to have the particular disease, infection, or medical condition. By way of example, in some embodiments, TCR / BCR profiles are generated from a cohort of people suffering from a single type of tumor or cancer, a particular infectious disease, a particular autoimmune disease, a particular allergy, or a particular medical condition.
[0270] In some embodiments, receptor chain pairs are identified. In some embodiments, TCR reactivity is verified by major histocompatibility complex (MHC) tetramer assay testing.
[0271] In some embodiments, the identified consensus TCR / BCR sequences are used in immunotherapy generation, such as CAR-T cell generation, for various disease states, including but not limited to, infectious diseases, autoimmune diseases, allergies, and cancer.
[0272] In some embodiments, the common TCR / BCR sequences identified for a disease cohort are used for antigen prediction (e.g., vaccine development) for diseases or conditions, including, but not limited to, infectious diseases, autoimmune diseases, allergies, and cancer. In some embodiments, patients are stratified by HLA type to facilitate prediction of antigens corresponding to common receptor sequences. In some embodiments, machine learning is used to predict one or more antigens. In some embodiments, antigens are validated by wet-lab experiments, including, but not limited to, multiplex identification of antigen-specific T cells (MIRA) and Biacore or other similar assays, based on surface plasmon resonance to detect binding energy and molecular interactions.
[0273] Predictive testing In some embodiments, models generated from TCR / BCR profiles derived from disease cohort analysis can also be applied to individual patient data.
[0274] In some embodiments, longitudinal studies of TCR / BCR profiles are performed during treatment or clinical trials to determine the efficacy / efficacy of any therapeutic approach. Similarly, longitudinal studies of TCR / BCR profiles can be performed in conjunction with vaccination, for example, in the context of a cancer vaccine, to determine efficacy and estimated time to progression or estimated time to remission, disease progression (with or without therapy), and / or outcome or efficacy of treatment. In some embodiments, longitudinal studies of TCR / BCR profiles are performed in conjunction with immuno-oncology (IO) therapy to provide detailed and accurate information regarding the efficacy of immuno-oncology treatment modalities.
[0275] In some embodiments, a single sampling point may be sufficient to assess a patient's response or efficacy in an IO treatment modality.
[0276] In some embodiments, the TCR / BCR profile data is integrated with other immunotherapy response predictors to accurately assess a patient's response to immunotherapy. In some embodiments, the TCR / BCR profile can be used as additional data to refine other existing or yet to be conceived predictive models.
[0277] In some embodiments, the large cohort data of people suffering from a particular disease or disorder includes TCR / BCR profile data and treatment outcomes.
[0278] In some embodiments, TCR / BCR analysis predicts protective or sterilizing immunity after natural infection / pathogen exposure or vaccination. In some embodiments, large datasets with positive and negative controls and receptor sequence enrichment level distributions are used to identify threshold levels of specific receptor sequence enrichment associated with pathogen infection / exposure.
[0279] In some embodiments, TCR / BCR analysis is used for HLA typing. In some embodiments, TCR / BCR data can be used in combination with other HLA typing methods and / or used as a feature to refine their predictive models. An example of an HLA typing method is disclosed, for example, in U.S. Patent Application No. 16 / 789,413, filed August 20, 2019, which is incorporated herein by reference in its entirety for all purposes. Exemplary Embodiments
[0280] Several non-limiting, exemplary embodiments of the present technology are disclosed herein.
[0281] Embodiment 1. In a first embodiment, a method for determining a patient's TCR / BCR profile is provided. In some embodiments, the method includes (a) isolating RNA from a patient sample; (b) enriching the isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes; (c) sequencing the RNA of (b) to generate sequencing data; and (d) analyzing the sequencing data to determine the patient's TCR / BCR profile. In some embodiments, the collection of TCR / BCR hybrid capture probes includes a first set of BCR constant region probes. Pool , the second of the BCR non-constant region probe Pool , the third of the TCR constant region probe Pool , and the fourth TCR non-constant region probe Pool Includes:
[0282] Embodiment 2. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool 2. The method of embodiment 1, wherein the ratio of
[0283] Embodiment 3. Step (b) comprises: (1) a targeted whole transcriptome panel; (2) a targeted whole exome panel; (3) a targeted panel directed to at least 10 target sequences of interest; or (4) any combination of 1-3, wherein step (b) comprises: Pool 2. The method of embodiment 1, further comprising concentrating using
[0284] Embodiment 4. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool , vs. 5th Pool The method of embodiment 3, wherein the ratio of
[0285] Embodiment 5. The method of embodiment 4, wherein 2% or less of the reads in the sequencing data map to TCR / BCR genes.
[0286] Embodiment 6. Step (b) comprises determining the fifth of the hybrid capture probes for the targeted whole exome panel. Pool 2. The method of embodiment 1, further comprising concentrating using
[0287] Embodiment 7. The method of embodiment 1, wherein step (d) comprises identifying multiple TCR / BCR clones in the sample.
[0288] Embodiment 8. The method of embodiment 1, wherein step (d) comprises identifying the most abundant TCR / BCR clones in the sample.
[0289] Embodiment 9. The method of embodiment 1, wherein step (d) comprises identifying the most abundant non-constant region sequences in the sample.
[0290] Embodiment 10. The method of embodiment 1, wherein the sample is a blood sample or a solid tumor sample.
[0291] Embodiment 11. The method of any of the above embodiments, further comprising diagnosing the patient with a disease or condition based on the TCR / BCR profile.
[0292] Embodiment 12. The method of embodiment 11, wherein the disease or condition comprises one or more of cancer, an infectious disease, an autoimmune condition, an allergy, or graft-versus-host disease.
[0293] Embodiment 13. The method of embodiment 12, wherein the cancer or infectious disease is one or more of those listed in embodiment 114.
[0294] Embodiment 14. The method of embodiment 11, wherein diagnosing comprises comparing the TCR / BCR profile of the subject to a control, and if the BCR / TCR profile of the subject is similar to the control (e.g., the abundance, identity, and / or clonality of one or more BCR / TCR receptors is comparable or identical to that of the control), then the subject is diagnosed with the disease or condition.
[0295] Embodiment 15. In some methods of any of the above embodiments, a control TCR / BCR panel for a disease (such as cancer or an infectious disease) or medical condition is provided.
[0296] Embodiment 16. In some embodiments, methods are provided for diagnosing a patient having a disease or condition based on the patient's TCR / BCR profile. In some embodiments, the method includes: a) isolating RNA from a patient sample; b) enriching the isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes; c) sequencing the RNA of (b) to generate sequencing data; d) analyzing the sequencing data to determine the patient's TCR / BCR profile; and e) comparing the patient's TCR / BCR profile to a standard collection to diagnose the patient having the disease or condition, wherein the collection of TCR / BCR hybrid capture probes includes a first TCR constant region probe and a second TCR constant region probe. Pool , the second of the TCR non-constant region probes Pool , the third of the BCR constant region probe Pool , and the fourth BCR non-constant region probe Pool Includes:
[0297] Embodiment 17. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool 17. The method of embodiment 16, wherein the ratio of
[0298] Embodiment 18. Step (b) comprises determining the fifth of the hybrid capture probes for the targeted whole transcriptome panel. Pool 17. The method of embodiment 16, further comprising concentrating using
[0299] Embodiment 19. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool , vs. 5th Pool 19. The method of embodiment 18, wherein the ratio of
[0300] Embodiment 20. The method of embodiment 19, wherein 2% or less of the reads in the sequencing data map to TCR / BCR genes.
[0301] Embodiment 21. Step (b) comprises determining the fifth of the hybrid capture probes for the targeted whole exome panel. Pool 17. The method of embodiment 16, further comprising concentrating using
[0302] Embodiment 22. The method of embodiment 16, wherein the disease or condition is an infectious disease, cancer, an autoimmune disease, an allergy, or graft-versus-host disease.
[0303] Embodiment 23. The method of embodiment 22, wherein the cancer or infectious disease is one or more of those listed in embodiment 114.
[0304] Embodiment 24. The method of embodiment 23, wherein diagnosing comprises comparing the TCR / BCR profile of the subject to a control, and if the BCR / TCR profile of the subject is similar to the control (e.g., the abundance, identity, and / or clonality of one or more BCR / TCR receptors is comparable or identical to that of the control), then the subject is diagnosed with the disease or condition.
[0305] Embodiment 25. In some embodiments, a control TCR / BCR panel for a disease (such as cancer or an infectious disease) or medical condition is provided.
[0306] Embodiment 26. In some embodiments, a method of assessing the severity or progression of a disease or condition based on a patient's TCR / BCR profile is provided. In some embodiments, the method includes: a) isolating RNA from a patient sample; b) enriching the isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes; c) sequencing the RNA of (b) to generate sequencing data; d) analyzing the sequencing data to determine the patient's TCR / BCR profile; and e) comparing the patient's TCR / BCR profile to a standard collection to characterize the severity or progression of the disease, wherein the collection of TCR / BCR hybrid capture probes includes a first set of TCR constant region probes. Pool , the second of the TCR non-constant region probes Pool , the third of the BCR constant region probe Pool , and the fourth BCR non-constant region probe Pool Includes:
[0307] Embodiment 27. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool 27. The method of embodiment 26, wherein the ratio of
[0308] Embodiment 28. Step (b) comprises determining the fifth of the hybrid capture probes for the targeted whole transcriptome panel. Pool 27. The method of embodiment 26, further comprising concentrating using
[0309] Embodiment 29. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool , vs. 5th Pool 29. The method of embodiment 28, wherein the ratio of
[0310] Embodiment 30. The method of embodiment 29, wherein 2% or less of the reads in the sequencing data map to TCR / BCR genes.
[0311] Embodiment 31. Step (b) comprises determining the fifth of the hybrid capture probes for the targeted whole exome panel. Pool 27. The method of embodiment 26, further comprising concentrating using
[0312] Embodiment 32. The method of embodiment 26, wherein the disease is an infectious disease, cancer, an autoimmune disease, or an allergy.
[0313] Embodiment 33. The method of embodiment 29, wherein the sample is a solid tumor sample.
[0314] Embodiment 34. The method of embodiment 30, wherein step (e) comprises determining the presence or extent of lymphocytic infiltration of the tumor.
[0315] Embodiment 35. The method of embodiment 32, wherein the cancer or infectious disease is one or more of those listed in embodiment 114.
[0316] Embodiment 36. The method of embodiment 35, wherein diagnosing comprises comparing the TCR / BCR profile of the subject to a control, and if the BCR / TCR profile of the subject is similar to the control (e.g., the abundance, identity, and / or clonality of one or more BCR / TCR receptors is comparable or identical to that of the control), then the subject is diagnosed with the disease or condition.
[0317] Embodiment 37. In some embodiments, a control TCR / BCR panel for a disease (such as cancer or an infectious disease) or medical condition is provided.
[0318] Embodiment 38. In some embodiments, methods are provided for treating a disease or condition in a patient based on the patient's TCR / BCR profile. In some embodiments, the method includes: a) isolating RNA from a patient sample; b) enriching the isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes; c) sequencing the RNA of (b) to generate sequencing data; d) analyzing the sequencing data to determine the patient's TCR / BCR profile; and e) administering treatment based on the patient's TCR / BCR profile, wherein the collection of TCR / BCR hybrid capture probes includes a first TCR constant region probe and a second TCR constant region probe. Pool , the second of the TCR non-constant region probes Pool , the third of the BCR constant region probe Pool , and the fourth BCR non-constant region probe Pool Includes:
[0319] Embodiment 39. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool 39. The method of embodiment 38, wherein the ratio of
[0320] Embodiment 40. Step (b) comprises determining a fifth of the hybrid capture probes for the targeted whole transcriptome panel. Pool 39. The method of embodiment 38, further comprising performing the concentration using
[0321] Embodiment 41. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool , vs. 5th Pool 41. The method of embodiment 40, wherein the ratio of
[0322] Embodiment 42. The method of embodiment 41, wherein 2% or less of the reads in the sequencing data map to TCR / BCR genes.
[0323] Embodiment 43. Step (b) comprises determining a fifth hybrid capture probe for a targeted whole exome panel. Pool 39. The method of embodiment 38, further comprising performing the concentration using
[0324] Embodiment 44. The method of embodiment 38, wherein step (d) comprises identifying the most abundant TCR / BCR clone in the sample, and wherein the treatment performed in step (e) comprises expanding the most abundant clone in vitro and readministering the expanded cells to the patient.
[0325] Embodiment 45. The method of embodiment 38, wherein step (d) comprises identifying the most abundant TCR non-constant region sequences in the sample, and wherein the treatment performed in step (e) comprises administering a CAR-T cell therapy comprising at least one of the most abundant TCR non-constant region sequences.
[0326] Embodiment 46. The method of embodiment 38, wherein the disease or condition is an infectious disease, cancer, an autoimmune disease, or an allergy.
[0327] Embodiment 47. The method of embodiment 46, wherein the cancer or infectious disease is one or more of those listed in embodiment 114.
[0328] Embodiment 48. In some embodiments, a method is provided for characterizing the effect of a treatment on a patient's TCR / BCR profile. In some embodiments, the method includes: (a) at a first time point, prior to administering the treatment, (i) isolating RNA from a patient sample; (ii) isolating RNA from a patient sample; (iii) sequencing the RNA of (ii) to generate sequencing data, and (iv) analyzing the sequencing data to determine the patient's TCR / BCR profile; (b) at a second time point, after administering the treatment, (i) isolating RNA from the patient sample, (ii) using the collection of hybrid capture probes to enrich the isolated RNA for TCR / BCR genes, (iii) sequencing the RNA of (ii) to generate sequencing data, and (iv) analyzing the sequencing data to determine the patient's TCR / BCR profile; and (c) comparing the TCR / BCR profile determined in step (a) with the TCR / BCR profile determined in step (b) to characterize the effect of the treatment on the patient's TCR / BCR profile, wherein the collection of TCR / BCR hybrid capture probes comprises a first set of TCR constant region probes. Pool , the second of the TCR non-constant region probes Pool , the third of the BCR constant region probe Pool , and the fourth BCR non-constant region probe Pool Includes:
[0329] Embodiment 49. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool 49. The method of embodiment 48, wherein the ratio of
[0330] Embodiment 50. Step (b) comprises determining a fifth of the hybrid capture probes for the targeted whole transcriptome panel. Pool 49. The method of embodiment 48, further comprising performing the concentration using
[0331] Embodiment 51. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool , vs. 5th Pool 51. The method of embodiment 50, wherein the ratio of is 1:2.5:100:100:10.
[0332] Embodiment 52. The method of embodiment 51, wherein 2% or less of the reads in the sequencing data map to TCR / BCR genes.
[0333] Embodiment 53. Step (b) comprises determining a fifth hybrid capture probe for a targeted whole exome panel. Pool 49. The method of embodiment 48, further comprising performing the concentration using
[0334] Embodiment 54. The method of embodiment 48, wherein the therapeutic agent is an immunotherapeutic agent.
[0335] Embodiment 55. The method of embodiment 54, wherein the immunotherapeutic is a vaccine.
[0336] Embodiment 56. The method of embodiment 54, wherein the immunotherapeutic is a chimeric antigen receptor (CAR) T cell.
[0337] Embodiment 57. The method of any one of embodiments 48-57, further comprising modifying the treatment prescribed to the patient based on the observed effect.
[0338] Embodiment 58. In some embodiments, methods are provided for identifying TCR / BCR non-constant region sequences enriched in a cohort of patients with a particular disease or condition. In some embodiments, the methods include: a) isolating RNA from a sample of each patient in the cohort; b) isolating a first sequence of TCR constant region probes; Pool , the second of the TCR non-constant region probes Pool , the third of the BCR constant region probe Pool , and the fourth BCR non-constant region probe Pool(b) enriching the isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes comprising: (a) a TCR / BCR hybrid capture probe set comprising: (b) a TCR / BCR hybrid capture probe set ...c) enriching the isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes comprising: (a) a TCR / BCR hybrid capture probe set comprising: (b) a TCR / BCR hybrid capture probe set comprising: (c) a TCR / BCR hybrid capture probe set comprising: (b) a TCR / BCR hybrid capture probe set comprising: (a) a TCR / BCR hybrid capture probe set comprising: (b) a TCR / BCR hybrid capture probe set comprising: (b) a TCR / BCR hybrid capture probe set comprising:
[0339] Embodiment 59. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool 59. The method of embodiment 58, wherein the ratio of
[0340] Embodiment 60. The collection of hybrid capture probes comprises a fifth of the probes comprising a targeted whole exome panel. Pool 59. The method of embodiment 58, further comprising:
[0341] Embodiment 61. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool , vs. 5th Pool 61. The method of embodiment 60, wherein the ratio of
[0342] Embodiment 62. The method of embodiment 61, wherein 2% or less of the reads in the sequencing data map to TCR / BCR genes.
[0343] Embodiment 63. The collection of hybrid capture probes comprises a fifth of the probes comprising a targeted whole exome panel. Pool 59. The method of embodiment 58, further comprising:
[0344] Embodiment 64. The method of embodiment 58, wherein the disease or condition is an infectious disease, an autoimmune disease, an allergy, or a cancer.
[0345] Embodiment 65. The method of embodiment 58, further comprising identifying disease-specific antigens using the enriched TCR / BCR non-constant region sequences.
[0346] Embodiment 66. The method of embodiment 65, further comprising producing a vaccine comprising the disease-specific antigen.
[0347] Embodiment 67. The method of embodiment 65 or 66, wherein the disease-specific antigen is a tumor antigen.
[0348] Embodiment 68. The method of embodiment 64, wherein the cancer or infectious disease is one or more of those listed in embodiment 114.
[0349] Embodiment 69. In some embodiments, a kit for determining a patient's TCR / BCR profile is provided. In some embodiments, the kit comprises a collection of TCR / BCR hybrid capture probes.
[0350] Embodiment 70. In some embodiments, a method of determining a patient's TCR / BCR profile is provided. In some embodiments, the method includes: a) isolating RNA from a patient sample; b) enriching the isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes and enriching for a targeted whole transcriptome panel using a collection of hybrid capture probes; c) sequencing the RNA of (b) to generate sequencing data; and d) analyzing the sequencing data to determine the patient's TCR / BCR profile, wherein the collection of TCR / BCR hybrid capture probes includes a first set of BCR constant region probes. Pool , the second of the BCR non-constant region probe Pool , the third of the TCR constant region probe Pool , and the fourth TCR non-constant region probe Pool a whole transcriptome targeted panel in the set, versus a first Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool The ratio of TCR to BCR is 10:1:2.5:100:100, and less than 2% of the reads in the sequencing data map to TCR / BCR genes.
[0351] Embodiment 71. In some embodiments, a method of determining a TCR / BCR profile of a patient is provided. In some embodiments, the method comprises: a) isolating RNA from a patient sample; b) enriching the isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes; c) sequencing the RNA of (b) to generate sequencing data; and d) analyzing the sequencing data to determine the TCR / BCR profile of the patient, wherein the patient has been exposed to or is suspected of being exposed to SARS-CoV-2, and the collection of TCR / BCR hybrid capture probes comprises a first group of TCR constant region probes. Pool , the second of the TCR non-constant region probes Pool , the third of the BCR constant region probe Pool , and the fourth BCR non-constant region probe Pool Includes:
[0352] Embodiment 72. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool 72. The method of embodiment 71, wherein the ratio of
[0353] Embodiment 73. Step (b) comprises determining the fifth of the hybrid capture probes for the targeted whole transcriptome panel. Pool 72. The method of embodiment 71, further comprising performing the concentration using
[0354] Embodiment 74. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool , vs. 5th Pool74. The method of embodiment 73, wherein the ratio of
[0355] Embodiment 75. The method of embodiment 74, wherein 2% or less of the reads in the sequencing data map to TCR / BCR genes.
[0356] Embodiment 76. Step (b) comprises determining the fifth hybrid capture probe for the whole exome targeting panel. Pool 72. The method of embodiment 71, further comprising performing the concentration using
[0357] Embodiment 77. The method of embodiment 71, further comprising identifying multiple TCR / BCR clones in the sample.
[0358] Embodiment 78. The method of embodiment 71, further comprising identifying the most abundant TCR / BCR clone in the sample.
[0359] Embodiment 79. The method of embodiment 71, further comprising identifying the most abundant non-constant region sequences in the sample.
[0360] Embodiment 80. The method of embodiment 71, wherein the sample is a blood sample or a solid tumor sample.
[0361] Embodiment 81. In some embodiments, a method for assessing the severity or progression of COVID-19 based on a patient's TCR / BCR profile is provided. In some embodiments, the method includes: a) isolating RNA from a patient sample; b) enriching the isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes; c) sequencing the RNA of (b) to generate sequencing data; d) analyzing the sequencing data to determine the patient's TCR / BCR profile; and e) comparing the patient's TCR / BCR profile to a standard collection to characterize the severity or progression of the disease, wherein the collection of TCR / BCR hybrid capture probes includes a first set of TCR constant region probes. Pool , the second of the TCR non-constant region probes Pool , the third of the BCR constant region probe Pool , and the fourth BCR non-constant region probe Pool Includes:
[0362] Embodiment 82. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool 82. The method of embodiment 81, wherein the ratio of
[0363] Embodiment 83. Step (b) comprises determining a fifth of the hybrid capture probes for the targeted transcriptome panel. Pool 82. The method of embodiment 81, further comprising performing the concentration using
[0364] Embodiment 84. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool , vs. 5th Pool 84. The method of embodiment 83, wherein the ratio of
[0365] Embodiment 85. The method of embodiment 84, wherein 2% or less of the reads in the sequencing data map to TCR / BCR genes.
[0366] Embodiment 86. Step (b) comprises determining a fifth hybrid capture probe for a targeted whole exome panel. Pool 82. The method of embodiment 81, further comprising performing the concentration using
[0367] Embodiment 87. In some embodiments, a method of treating COVID-19 based on a patient's TCR / BCR profile is provided. In some embodiments, the method includes: a) isolating RNA from a patient sample; b) enriching the isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes; c) sequencing the RNA of (b) to generate sequencing data; d) analyzing the sequencing data to determine the patient's TCR / BCR profile; and e) administering treatment based on the patient's TCR / BCR profile, wherein the collection of TCR / BCR hybrid capture probes includes a first TCR constant region probe and a second TCR constant region probe. Pool , the second of the TCR non-constant region probes Pool , the third of the BCR constant region probe Pool , and the fourth BCR non-constant region probe Pool Includes:
[0368] Embodiment 88. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool 88. The method of embodiment 87, wherein the ratio of is 1:2.5:100:100.
[0369] Embodiment 89. Step (b) comprises determining the fifth of the hybrid capture probes for the targeted whole transcriptome panel. Pool 88. The method of embodiment 87, further comprising performing the concentration using
[0370] Embodiment 90. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool , vs. 5th Pool 90. The method of embodiment 89, wherein the ratio of is 1:2.5:100:100:10.
[0371] Embodiment 91. The method of embodiment 90, wherein 2% or less of the reads in the sequencing data map to TCR / BCR genes.
[0372] Embodiment 92. Step (b) comprises determining a fifth of the hybrid capture probes for the targeted whole transcriptome panel. Pool 88. The method of embodiment 87, further comprising performing the concentration using
[0373] Embodiment 93. The method of embodiment 87, wherein step (d) comprises identifying the most abundant TCR / BCR clone in the sample, and wherein the treatment performed in step (e) comprises expanding the most abundant clone in vitro and readministering the expanded cells to the patient.
[0374] Embodiment 94. The method of embodiment 87, wherein step (d) comprises identifying the most abundant TCR non-constant region sequences in the sample, and wherein the treatment performed in step (e) comprises administering a CAR-T cell therapy comprising at least one of the most abundant TCR non-constant region sequences.
[0375] Embodiment 95. In some embodiments, a method is provided for characterizing the effect of a COVID-19 treatment on a patient's TCR / BCR profile. In some embodiments, the method includes: a) at a first time point, before administering the treatment, i. isolating RNA from a patient sample, ii. using a collection of TCR / BCR hybrid capture probes to enrich the isolated RNA for TCR / BCR genes, iii. sequencing the RNA of (aii) to generate sequencing data, and iv) analyzing the sequencing data to determine the patient's TCR / BCR profile; b) at a second time point, after administering the treatment, i) isolating RNA from the patient sample, ii). iii) sequencing the RNA of (bii) to generate sequencing data; and iv) analyzing the sequencing data to determine a TCR / BCR profile of the patient; and c) comparing the TCR / BCR profile determined in step (a) with the TCR / BCR profile determined in step (b) to characterize the effect of treatment on the TCR / BCR profile of the patient, wherein the set of TCR / BCR hybrid capture probes comprises a first set of TCR constant region probes. Pool , the second of the TCR non-constant region probes Pool , the third of the BCR constant region probe Pool , and the fourth BCR non-constant region probe Pool Includes:
[0376] Embodiment 96. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool 96. The method of embodiment 95, wherein the ratio of is 1:2.5:100:100.
[0377] Embodiment 97. Steps (aii) and (bii) comprise determining a fifth of the hybrid capture probes for a targeted whole transcriptome panel. Pool96. The method of embodiment 95, further comprising performing the concentration using
[0378] Embodiment 98. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool , vs. 5th Pool 98. The method of embodiment 97, wherein the ratio of is 1:2.5:100:100:10.
[0379] Embodiment 99. The method of embodiment 98, wherein 2% or less of the reads in the sequencing data map to TCR / BCR genes.
[0380] Embodiment 100. Steps (aii) and (bii) comprise: Pool 96. The method of embodiment 95, further comprising performing the concentration using
[0381] Embodiment 101. The method of embodiment 95, wherein the therapeutic agent is an immunotherapeutic agent.
[0382] Embodiment 102. The method of embodiment 101, wherein the immunotherapeutic is a vaccine.
[0383] Embodiment 103. The method of embodiment 101, wherein the immunotherapeutic is a chimeric antigen receptor (CAR) T cell.
[0384] Embodiment 104. The method of any one of embodiments 95-104, further comprising modifying the treatment prescribed to the patient based on the observed effect.
[0385] Embodiment 105. In some embodiments, a method is provided for identifying TCR / BCR non-constant region sequences enriched in a cohort of patients with SARS-CoV-2. In some embodiments, the method includes: a) isolating RNA from a sample of each patient in the cohort; b) enriching the isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes; c) sequencing the RNA of (b) to generate sequencing data; d) analyzing the sequencing data to determine a TCR / BCR profile of patients in the cohort; and e) identifying TCR / BCR non-constant region sequences enriched in the cohort compared to a control group without the disease or condition, wherein the collection of hybrid capture probes comprises a first collection of TCR constant region probes. Pool , the second of the TCR non-constant region probes Pool , the third of the BCR constant region probe Pool , and the fourth BCR non-constant region probe Pool Includes:
[0386] Embodiment 106. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool 106. The method of embodiment 105, wherein the ratio of
[0387] Embodiment 107. The collection of hybrid capture probes comprises a fifth of the probes comprising a whole transcriptome targeting panel. Pool 106. The method of embodiment 105, further comprising:
[0388] Embodiment 108. The first in the set Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool , vs. 5th Pool 108. The method of embodiment 107, wherein the ratio of is 1:2.5:100:100:10.
[0389]
[0062] Embodiment 109. The method of embodiment 108, wherein 2% or less of the reads in the sequencing data map to TCR / BCR genes.
[0390] Embodiment 110. The collection of hybrid capture probes comprises a fifth of the probes comprising a whole exome targeting panel. Pool 106. The method of embodiment 105, further comprising:
[0391] Embodiment 111. The method of embodiment 105, further comprising identifying SARS-CoV-2-specific antigens using the enriched TCR / BCR non-constant region sequences.
[0392] Embodiment 112. The method of embodiment 108, further comprising producing a vaccine comprising a SARS-CoV-2-specific antigen.
[0393] Embodiment 113. In some embodiments, a kit for determining the TCR / BCR profile of a COVID-19 patient is provided. In some embodiments, the kit comprises a collection of TCR / BCR hybrid capture probes. In some embodiments, the collection of probes comprises a first TCR constant region probe Pool , the second of the TCR non-constant region probes Pool , the third of the BCR constant region probe Pool , and the fourth BCR non-constant region probe Pool Four separate Pool In some embodiments, the first Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool In some embodiments, the set of probes is (1) a whole transcriptome targeted panel, (2) a whole exome targeted panel, or (3) a fifth set of probes. Pool The first targeting panel is used in combination with one of 10,000 to 20,000 targets of interest. Pool , versus , the second Pool , versus , the third Pool , versus 4th Pool , vs. 5th Pool The ratio of TCR / BCR genes is 1:2.5:100:100:10. When used in sequencing reactions such as RNA-seq, the TCR / BCR panel is configured so that 2% or less of the reads in the sequencing data map to TCR / BCR genes.
[0394] Embodiment 114. In some of the above embodiments, (a) a subject or cohort is diagnosed as having, suspected of having, or afflicted with a disease or medical condition, such as cancer or an infectious disease (infectious disease), or (b) a method of diagnosing a disease or medical condition, such as cancer or an infectious disease (infectious disease), is provided. By way of example, but not limitation, in any of the above embodiments, the cancer may be selected from the group consisting of chondrosarcoma, Ewing's sarcoma, malignant fibrous histiocytoma / osteosarcoma of bone, osteosarcoma, rhabdomyosarcoma, leiomyosarcoma, myxosarcoma, astrocytoma, brainstem glioma, pilocytic astrocytoma, ependymoma, primitive neuroectodermal tumor, cerebellar astrocytoma, cerebral astrocytoma, glioblastoma, glioma, medulloblastoma, neuroblastoma, oligodendroglioma, pineal astrocytoma, pituitary adenoma, breast cancer, invasive lobular carcinoma, tubular adenocarcinoma, invasive cribriform carcinoma, medullary carcinoma. , male breast cancer, phyllodes tumor, inflammatory breast cancer, adrenocortical carcinoma, islet cell carcinoma (endocrine pancreatic cancer), multiple endocrine neoplasia syndrome, parathyroid cancer, pheochromocytoma, thyroid cancer, Merkel cell carcinoma, uveal melanoma, retinoblastoma, anal cancer, appendix cancer, bile duct cancer, carcinoid tumor, gastrointestinal cancer, colon cancer, extrahepatic bile duct cancer, gallbladder cancer, stomach cancer, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor, hepatocellular carcinoma, pancreatic cancer, islet cell carcinoma, rectal cancer, bladder cancer, cervical cancer, endometrial cancer, extragonadal germ cell Tumors, ovarian cancer, ovarian epithelial cancer (epidermal stromal tumor), ovarian germ cell tumor, penile cancer, renal cell cancer, renal pelvis and ureter cancer, transitional cell carcinoma, prostate cancer, testicular cancer, gestational trophoblastic tumor, ureter and renal pelvis cancer, transitional cell carcinoma, urethral cancer, uterine sarcoma, vaginal cancer, vulvar cancer, Wilms' tumor, esophageal cancer, head and neck cancer, nasopharyngeal cancer, oral cancer, oropharyngeal cancer, paranasal sinus and nasal cavity cancer, pharyngeal cancer, salivary gland cancer, hypopharyngeal cancer, acute mixed lineage leukemia, acute eosinophilic leukemia, acute lymphocytic leukemia, acute myeloid leukemia, acute Myeloid dendritic cell leukemia, AIDS-related lymphoma, anaplastic large cell lymphoma, angioimmunoblastic T-cell lymphoma, B-cell prolymphocytic leukemia, Burkitt's lymphoma, chronic lymphocytic leukemia, chronic myeloid leukemia, cutaneous T-cell lymphoma, diffuse large B-cell lymphoma, follicular lymphoma, hairy cell leukemia, hepatosplenic T-cell lymphoma, Hodgkin's lymphoma, hairy cell leukemia, intravascular large B-cell lymphoma, large granular lymphocyte leukemia, lymphoplasmacytic lymphoma, lymphomatoid granulomatosis,Mantle cell lymphoma, marginal zone B-cell lymphoma, mast cell leukemia, mediastinal large B-cell lymphoma, multiple myeloma / plasmacytoma, myelodysplastic syndrome, mucosa-associated lymphoid tissue lymphoma, mycosis fungoides, nodal marginal zone B-cell lymphoma, non-Hodgkin's lymphoma, precursor B-lymphoblastic leukemia, primary central nervous system lymphoma, primary cutaneous follicular lymphoma, primary cutaneous immunocytoma, primary effusion lymphoma, plasmablastic lymphoma, Sézary syndrome, splenic marginal zone lymphoma The tumor may be one or more of: lymphoma, T-cell prolymphocytic leukemia, basal cell carcinoma, squamous cell carcinoma, skin adnexal tumor (e.g., sebaceous gland carcinoma), melanoma, primary sarcoma of cutaneous origin (e.g., dermatofibrosarcoma protuberans), primary lymphoma of cutaneous origin, bronchial adenoma / carcinoid, small cell lung cancer, mesothelioma, non-small cell lung cancer, pleuropulmonary blastoma, laryngeal carcinoma, thymic carcinoma, Kaposi's sarcoma, epithelioid hemangioendothelioma (EHE), desmoplastic small round cell tumor, or liposarcoma. In any of the above embodiments, the infectious disease (infectious disease) is selected from the group consisting of Acinetobacter infection, actinomycosis, African sleeping sickness, AIDS (acquired immune deficiency syndrome), amebiasis, anaplasmosis, angiostrongyliasis, anisakiasis, anthrax, hemolytic alkanobacterial infection, Argentine hemorrhagic fever, ascariasis, aspergillosis, astrovirus infection, babesiosis, Bacillus cereus infection, bacterial meningitis, bacterial pneumonia, bacterial vaginosis, Bacteroides infection, balantidiosis, bartonellosis, baylisascaris infection, BK virus infection, black sand mites, blastocystosis, blastomycosis, Bolivian hemorrhagic fever, botulism (and infant botulism). toxicosis), Brazilian hemorrhagic fever, brucellosis, bubonic plague, Burkholderia infection, Buruli ulcer, Calicivirus infection (norovirus and sapovirus), Campylobacteriosis, Candidiasis (moniliasis, thrush), Capillariasis, Carrion disease, Cat scratch disease, Cellulitis, Chagas disease (American trypanosomiasis), Chancroid, Chickenpox, Chikungunya fever, Chlamydia infection, Chlamydia pneumoniae infection (Taiwan acute respiratory agent or TWAR), Cholera, Fungal infection, Chytridiomycosis, Clostridium difficile enteritis, Coccidioidomycosis, Colorado tick fever (CTF), Cold (acute viral rhinopharyngitis, acute coryza), Coronavirus disease 2019 (COVID-19), Creutzfeldt-Jakob disease (CJD),Crimean-Congo hemorrhagic fever (CCHF), cryptococcosis, cryptosporidiosis, cutaneous larva migrans (CLM), cyclosporiasis, neurocysticercosis, cytomegalovirus infection, dengue fever, desmodesmus infection, dientamebiasis, diphtheria, diphyllobothriasis, dracunculiasis, Ebola hemorrhagic fever, echinococcosis, ehrlichiosis, enterobiasis, enterococcus infection, enterovirus infection, epidemic typhus, erythema infectiosum (5th disease), exanthema subitum (6th disease), fascioliasis, fascioliasis, fatal familial insomnia (FFI), filariasis, clostridium velutipes Dium perfringens food poisoning, free-living amoeba infection, Fusobacterium infection, gas gangrene (clostridial myonecrosis), geotrichosis, Gerstmann-Straussler-Scheinker syndrome (GSS), Giardiasis, glanders, gnathostomiasis, gonorrhea, inguinal granuloma, group A streptococcal infection, group B streptococcal infection, Haemophilus influenzae infection, hand, foot and mouth disease (HFMD), hantavirus pulmonary syndrome (HPS), heartland virus disease, Helicobacter pylori infection, hemolytic uremic syndrome (HUS), hemorrhagic fever with renal syndrome (HFRS), Hendra virus infection, Hepatitis A, Hepatitis B, Hepatitis C, Hepatitis D, Hepatitis E, Herpes simplex, Histoplasmosis, Hookworm disease, Human bocavirus infection, Human ehrlichiosis, Human granulocytic anaplasmosis (HGA), Human metapneumovirus infection, Human monocytic ehrlichiosis, Human papillomavirus (HPV) infection, Human parainfluenza virus infection, Membranous taeniasis, Epstein-Barr virus infectious mononucleosis, Influenza, Isosporosis, Kawasaki disease, Keratitis, Kingella kingae infection, Kuru, Lassa fever, Legionnaires' disease, Pontiac fever, Leishmaniasis , leprosy, leptospirosis, listeriosis, Lyme disease (Lyme borreliosis), lymphatic filariasis (elephantiasis), lymphocytic choriomeningitis, malaria, Marburg hemorrhagic fever (MHF), measles, Middle East respiratory syndrome (MERS), melioidosis (Whitmore disease), meningitis, meningococcal disease, metagonism, microsporidiosis, molluscum contagiosum (MC), monkeypox, mumps, typhus (epidemic typhus), mycoplasma pneumonia, Mycoplasma genitalium infection, mycetoma, myiasis, neonatal conjunctivitis (ophthalmia neonatorum), Nipah virus infection, norovirus,(New) Variant Creutzfeldt-Jakob disease (vCJD, nvCJD), nocardiosis, onchocerciasis (river blindness), opisthorchiasis, paracoccidioidomycosis (South American blastomycosis), paragonimiasis, pasteurellosis, head lice, body lice, pubic lice, pelvic inflammatory disease (PID), whooping cough, bubonic plague, pneumococcal infection, Pneumocystis pneumonia (PCP), pneumonia, poliovirus infection, Prevotella infection, primary amebic meningoencephalitis (PAM), progressive multifocal leukoencephalopathy, psittacosis, Q fever, rabies, relapsing fever, respiratory syncytial virus infection, rhinosporidiosis, rhinovirus infection, rickettsial infection, rickettsial smallpox, Rift Valley fever (RVF), Rocky Mountain spotted fever (RMSF), rotavirus infection, rubella, salmonellosis, severe acute respiratory syndrome (SARS), scabies, scarlet fever, Schistosomiasis, septicemia, shigellosis, shingles, smallpox, sporotrichosis, staphylococcal food poisoning, staphylococcal infection, strongyloidiasis, subacute sclerosing panencephalitis, bejel, syphilis, yaws, taeniasis, tetanus, tinea barbae, tinea capitis, tinea corporis, tinea rotundifolia, tinea manus, tinea nigra, tinea unguium, tinea versicolor, toxic shock syndrome (TSS), toxocariasis (ocular larva migrans (OLM)), toxocariasis (visceral larva migrans (VLM)), The illness may be one or more of the following: toxoplasmosis, trachoma, trichinosis, trichomoniasis, trichuriasis, tuberculosis, tularemia, typhoid fever, typhus, Ureaplasma urealyticum infection, valley fever, Venezuelan equine encephalitis, Venezuelan hemorrhagic fever, Vibrio vulnificus infection, Vibrio parahaemolyticus, viral pneumonia, West Nile fever, white sand flu, Mycobacterium pseudotuberculosis infection, yersiniosis, yellow fever, zeaspora, Zika fever, and zygomycosis.
[0395] Embodiment 115. A method of sequencing at least one TCR or BCR region of a specimen using a plurality of probes, wherein the probes are a first TCR constant region probe. Pool , the second of the TCR non-constant region probes Pool , the third of the BCR constant region probe Pool , and the fourth BCR non-constant region probe Pool The first Pool has a first density level and a second Poolhas a second density level and a third Pool has a third concentration level and the fourth pool has a fourth concentration level.
[0396] Embodiment 116. The method of embodiment 115, wherein the first concentration level, the second concentration level, the third concentration level, and the fourth concentration level are different from each other.
[0397] Embodiment 117. The first Pool The concentration level of the probe in the second Pool The concentration level of the probe in the third Pool The concentration level of the probe in the fourth Pool 117. The method of embodiment 115 or 116, wherein the concentration level of the probe in the first Pool The concentration level of the probe in the second Pool The concentration level of the probe in the third Pool The concentration level of the probe in the fourth Pool Lower than the concentration level of the probe in
[0398] Embodiment 118. The first and third Pool The concentration levels of the probes in the second and fourth Pool 118. The method of any one of embodiments 115-117, wherein the concentration of the probe in the third and fourth samples is at least 2-fold less, at least about 5-fold less, at least about 10-fold less, at least about 15-fold less, at least about 20-fold less, at least about 30-fold less, at least about 40-fold less, or at least about 50-fold less. Pool The concentration levels of the probes in the first and second Pool The concentration level of the probe in the sample is at least 2-fold less, at least about 5-fold less, at least about 10-fold less, at least about 15-fold less, at least about 20-fold less, at least about 30-fold less, at least about 40-fold less, or at least about 50-fold less.
[0399] Embodiment 119. Selecting a plurality of probes from a set of probes Pool 1. A method for sequencing at least one TCR or BCR region of a specimen, comprising forming a Pool wherein the plurality of probes are selected to exclude at least a portion of the constant region of at least one TCR or BCR region.
[0400] Embodiment 120. Pool 120. The method of embodiment 119, wherein said method comprises a probe for sequencing at least a portion of the constant region of the TCR or BCR.
[0401] Embodiment 121. The method of embodiment 119 or 120, wherein the sequencing is whole transcriptome sequencing.
[0402] Embodiment 122. The method of any of embodiments 117-121, wherein the sequencing is short-read sequencing.
[0403] Embodiment 123. The method of embodiment 115, wherein sequencing is performed on a sample taken from the patient, and the results are used to predict the patient's susceptibility to the disease.
[0404] Embodiment 124. The method of embodiment 115, wherein the TCR or BCR region is associated with a viral infection and the specimen is taken prior to administering to the patient a vaccine designed to protect against viral infection.
[0405] Embodiment 125. The method of embodiment 115, wherein sequencing is performed on a specimen collected from a patient, and the patient has been in contact with an infectious agent prior to specimen collection.
[0406] Embodiment 126. The method of embodiment 125, wherein the patient has generated antibodies against the infectious pathogen.
[0407] Embodiment 127. The method of embodiment 125, wherein the patient has not produced a substantial amount of antibodies to the infectious pathogen.
[0408] Embodiment 128. The method of embodiment 125, wherein the infectious agent did not cause seroconversion.
[0409] Embodiment 129. The method of embodiment 125, wherein high concentrations of infectious pathogens were not detectable in the patient's blood.
[0410] Embodiment 130. The method of embodiment 125, wherein the infectious pathogen is SARS-CoV-2.
[0411] Embodiment 131. The method of embodiment 115, wherein sequencing is performed on a specimen collected from a patient, and the patient is experiencing symptoms associated with a respiratory disease.
[0412] Embodiment 132. The method of embodiment 115, wherein sequencing is performed on a specimen collected from a patient, and the patient is experiencing influenza-like symptoms.
[0413] Embodiment 133. The method of embodiment 115, wherein the specimen is a tissue specimen.
[0414] Embodiment 134. The method of embodiment 115, wherein the specimen is a tumor specimen.
[0415] Embodiment 135. The method of embodiment 115, wherein the specimen is a blood specimen.
[0416] Embodiment 136. The method of embodiment 115, wherein the specimen is a saliva specimen.
[0417] Embodiment 137. The method of embodiment 115, wherein the specimen is a mucosal specimen.
[0418] Embodiment 138. The method of embodiment 115, wherein the specimen is a cerebrospinal fluid specimen.
[0419] Embodiment 139. The method of embodiment 115, wherein the sequencing is performed by whole transcriptome sequencing.
[0420] Embodiment 140. A method for sequencing an RNA transcriptome, comprising the method of embodiment 115.
[0421] Embodiment 141. The method of embodiment 115, further comprising identifying multiple TCR clones in the specimen.
[0422] Embodiment 142. The method of embodiment 141, further comprising determining the proportion of at least one TCR clone in the plurality of TCR clones in the specimen.
[0423] Embodiment 143. The method of embodiment 115, further comprising identifying multiple BCR clones in the specimen.
[0424] Embodiment 144. The method of embodiment 143, further comprising identifying a proportion of at least one BCR clone in the plurality of BCR clones in the specimen.
[0425] Embodiment 145. The assembly comprises a TCR constant region. Pool , TCR non-constant region Pool , BCR constant region Pool , and BCR non-constant region Pool 142. The method of any of embodiments 115-141, comprising at least one oligonucleotide from
[0426] Embodiment 146. The method of any of the above embodiments, wherein the population of TCR / BCR probes is obtained as described in Example 1. [Example]
[0427] The following examples are illustrative and should not be construed as limiting the claimed subject matter.
[0428] Example 1 TCR / BCR profiling probe and assay development A. Methods for Selecting or Designing Hybrid Capture Probe Sequences For example, probes can be designed to enrich for nucleic acids associated with TCR / BCR genes in a sequencing library within an RNA-seq assay.
[0429] Step 1 is to generate a list of reference target gene sequences that are located within the desired target gene.
[0430] Step 1 may include collecting a complete set of reference sequences for these genes and corresponding alleles from a database of potential probe designs. In another embodiment, a set of reference sequences for these genes can be collected to generate a list. In one embodiment, the reference sequence for a gene may include all exons and all introns associated with the gene, only a portion of the exons, only a portion of the introns, or no introns at all. In one example, for each gene, a segment (portion) of the gene may be selected as the target gene sequence. In one embodiment, each target gene sequence has a length of about 400 bp. Each gene may have multiple alleles, and each allele may have a unique reference sequence.
[0431] In one embodiment, step 1 involves collecting a complete set of IG and TCR gene sequences from a gene sequence database, such as the IMGT database. In one example, the database contains 296 IG (BCR) and 222 TCR genes, which generally range in length from about 100 to 1000 bp. As can be seen in Figure 14, approximately half of these genes have multiple annotated alleles.
[0432] The IG and TCR loci may contain hundreds of genes with substantial homology and allelic variation. Figure 14 ("Gene and Allele Counts") shows the gene counts (y-axis) for each class of IG (BCR) or TCR genes with 1, 2, 3, 4, or 5+ alleles (see legend for color scheme), demonstrating the allelic variation of these genes. Each class of gene is represented along the x-axis (IGHC, IGHD, IGHJ, IGHV, IGKC, IGKJ, IGKV, IGLC, IGLJ, IGLV, TRAC, TRAJ, TRAV, TRBC, TRBD, TRBJ, TRBV, TRDC, TRDD, TRDJ, TRDV, TRGC, TRGJ, TRGV, etc.).
[0433] Step 2 is an optional step of determining a gene consensus sequence across multiple alleles. This step may involve comparing allele sequences of a gene to determine a consensus sequence.
[0434] A probe set covering all alleles may ensure complete coverage while eliminating a significant amount of redundancy due to high sequence similarity between alleles. At a basic level, a representative (consensus) target sequence at the gene level may result in a probe panel that covers the majority of allelic variation.
[0435] Comparison of allelic sequences may include filling in gaps in the reference sequence prior to comparison. In this example, IMGT provides reference sequences in a curated alignment format (IMGT Gapped Fasta). Unfortunately, many of these IMGT reference sequences are incompletely sequenced at the 5' or 3' end. As a result, in addition to single nucleotide variations, truncations are often present in the raw IMGT allelic sequences. This issue is illustrated in Figure 15, with truncations in various alleles at both the 5' and 3' ends, as shown in example TRAV 8-4. Figure 15 shows an example of aligned TCR reference sequences.
[0436] In one example, filling gaps in the reference sequence (converting the raw IMGT reference sequence to a full allele reference sequence) is done by determining a consensus sequence based on a curated alignment of IMGTs (the most frequent nucleotide at each position) and using that consensus sequence to fill in (substitute) the truncated or missing segments in each allele. In this example, the processed set of filled in reference sequences contains the set of target sequences that the probes should cover. (See Figure 15, "IMGT Sequence Processing").
[0437] Step 3 is an optional step that assesses sequence similarity across multiple alleles.
[0438] This step can use the processed IMGT reference sequences (filled in, if applicable, and / or portions of the sequences if missing as described above) to determine whether a gene-level consensus sequence can cover the potential allelic variation.
[0439] This step may involve comparing each allele sequence to its corresponding gene consensus sequence. Figure 16 shows the cumulative distribution of the number of mismatched base pairs (bp) and the percentage of mismatched bp (number of mismatches relative to gene length).
[0440] FIG. 16 ("Allele Sequence Similarity") shows that most alleles are highly similar to their gene consensus sequences, according to the empirical cumulative distribution function (CDF).
[0441] In this example, 98.6% of all alleles have mismatches of less than 15 bp, and 98.2% of all alleles have at least 95% identity compared to the gene consensus sequence. For a small number of alleles (fewer than 20) with low consensus sequence identity, it may be appropriate to cover those sequence differences separately in order to design a set of probes that covers all alleles.
[0442] Step 4 is an optional step of filtering the list of genes, alleles, and / or target segments. The filtering strategies described herein can be used individually, in conjunction with other probe design list filtering strategies known in the art, or in any combination thereof.
[0443] The list may be filtered with the goal of reducing sequencing reads from less desirable targets. In one example, constant region targets are less desirable than non-constant region targets. In this example, constant region targets may be filtered and removed from the list if they are located more than a specified distance threshold (in bp) from the non-constant region. In another example, constant region targets may be filtered and removed from the list if they are not within the two to five targets located closest to the non-constant region of a gene.
[0444] The list may be filtered with the goal of reducing redundancy of targets and duplicate or substantially equivalent probe sequences designed based on those targets.
[0445] In one example, allele reference sequences can be removed from the list or replaced with a gene consensus sequence if the gene consensus sequence has at least 95% sequence identity (e.g., at least 95% of the sequences have the same nucleotide as the consensus sequence for positions in the consensus sequence and corresponding positions in the allele sequence). For the 19 alleles for which 95% sequence identity is not achieved, the original allele sequence is retained. In this example, all allele sequences that are at least 95% identical to this gene consensus set are likely to be covered by the final probe set.
[0446] Step 5 is an optional step that calculates an estimate of the total desired probe coverage for these loci.
[0447] The table in Figure 17 shows the difference in total desired coverage length (in base pairs) using the full set of IG and TCR allele sequences (upper bound, unfiltered target list) (Table 1 in Figure 17) versus the gene-level consensus sequence (filtered target list) (Table 2 in Figure 17). Using the gene-level consensus strategy reduces the number of gene sequences in the set from 1098 all-allele sequences to 532 all-gene consensus sequences, reducing the total coverage length from 325 kb to 125 kb. This reduced set of sequences is expected to correspondingly reduce the number of probes required for coverage. In this example, using the gene-level consensus strategy reduces the number of possible target sequences / different 120-mers (example probe lengths) in the IG / TCR sequences from 115,920 to 68,746. Results may vary depending on the probe length selected.
[0448] Step 6 is an optional step of providing the target list to a probe design specialist. The list may be the filtered list of targets generated in step 4. The probe design specialist may be a commercial vendor that designs and / or manufactures sequencing probes and / or primers. One example of such a commercial vendor is IDT.
[0449] Step 7 is to select (design) probe sequences based on the list (for example, using probe design software). Probe sequence selection may be performed by a probe design specialist.
[0450] By way of example, but not limitation, probe sequences may be selected or designed according to the methods described in FastPCR software for PCR primer and probe design and iterative searching (Kalendar et al., 2009, Genes, Genomes, and Genomics, Vol. 3 (Special Issue 1), pp. 1-14), which is incorporated herein by reference.
[0451] B. TCR / BCR assay development using probes obtained by the method of step A.
[0452] This example illustrates the development of one embodiment of a TCR / BCR profiling assay. In this embodiment, TCR / BCR sequencing is performed in combination with RNA sequencing. In the embodiment described here, seven receptors are tiled: IGH, IGK, IGL, TRA, TRB, TRG, and TRD. Thus, the repertoire data includes annotated CDR3 hypervariable sequence quantification for IgH, IgK, IgL, TCR-alpha, TCR-beta, TCR-delta, and TCR-gamma receptors. (See, e.g., Figure 4.)
[0453] The capture method of this embodiment is optimized to produce an RNAseq output in which 2% or less of all unfiltered read pairs map to TCR and BCR sequences in 95% of samples. This capture rate maintains transcriptome integrity for downstream analysis while still capturing sufficient depth to adequately identify receptor clonotypes from the most abundant infiltrating lymphocyte clones. (See Figures 3 and 5.)
[0454] In some early attempts to perform this analysis, some of the least informative regions (constant regions) accounted for the majority of reads. Also, BCR region coverage overwhelmingly exceeded TCR region coverage. To address this issue, in some embodiments, useful constant region coverage was identified and retained, while many constant region probes that were unlikely to generate informative TCR / BCR reads were removed. Additionally, probes were assigned to different TCR regions to allow for independent fine-tuning of signals from the TCR and BCR regions. Pool and BCR Pool In some embodiments, the probes are divided into TCR transient, TCR constant, BCR transient, and BCR constant probe concentrations. PoolBy independently fine-tuning the TCR transient, TCR constant, BCR transient, and BCR constant probe concentrations, more informative TCR / BCR information was obtained with far fewer reads. It also ensured that the information from TCR / BCR profiling was more evenly balanced between TCR and BCR.
[0455] Several experiments were performed that led to the construction of TCR / BCR probes that could be successfully used in RNAseq assays, a brief description of which follows and a schematic diagram of the method is shown in Figure 2.
[0456] First, two sets of TCR / BCR probes were designed. Design 1 contained all TCR / BCR probes in a single tube. Several attempts were made to optimize this configuration, with specific probes being included. Pool and the concentration of whole-exome panel probes relative to Pool Design 2 involved dividing the probes into four groups: TCR transient, TCR constant, BCR transient, and BCR constant. Probes were then selected for each group, resulting in the final composition of each group as follows:
[0457] I. "BCR Steady" - 45 probes
[0458] II. "BCR Non-Stationary" - 893 probes
[0459] III. "TCR Stationary" - 18 probes
[0460] IV. "TCR transient" - 650 probes
[0461] The TCR / BCR probe concentrations relative to each other and relative to the exome probe were also evaluated. The exome probe was tested at 25 attomoles / probe / capture. The ratios below refer to relative concentrations compared to the exome. For example, a 10-fold spike means a final amount of 250 attomoles / probe / capture (the exome in this case is 25 attomoles / probe / capture). In various embodiments, probes designed to target the entire human exome can be hybridized with DNA or RNA molecules. When the probes are hybridized with RNA molecules, the molecules in the library can be referred to as a human transcriptome.
[0462] [Table 6]
[0463] Leveraging specially designed probes to integrate repertoire sequencing (rep-seq) into a large-scale RNA-seq workflow, the method disclosed herein can capture a snapshot of the immune receptor repertoire without compromising transcriptome analysis.
[0464] Example 2 Sequencing results In this example, TCR and BCR sequences in blood samples taken from patients with B cell lymphoma were analyzed according to the systems and methods disclosed herein.
[0465] method Sample Preparation (including enrichment with TCR / BCR hybrid capture probes obtained according to the method of Example 1)
[0466] RNA was quantified using the Quant-it Ribogreen RNA Assay (ThermoFisher Scientific, Part Number R11490) and authenticated using the Fragment Analyzer High Sens RNA Analysis Kit (Agilent Technologies, Part Number DNF-472-1000). RNA was normalized to 10 ng / uL in a 10-uL starting volume and then subjected to thermal and chemical fragmentation, with variable parameters to obtain fragments of comparable size from RNA inputs with different starting size distributions. Library preparation was performed using a commercially available kit (KAPA RNA HyperPrep Kit for Illumina, Part Number KK8544) in conjunction with IDT unique double-indexed (UDI) unique molecular identifier (UMI) adapters. This involved first-strand synthesis using reverse transcriptase (RT) to create first-strand cDNA, followed by treatment with RNAse to degrade RNA and DNA polymerase to achieve second-strand synthesis, which creates double-stranded cDNA. IDT UDI-UMI adapters were ligated to the cDNA, and the adapter-ligated libraries were cleaned using a magnetic bead-based method (Roche Diagnostics, part number KK8002). The libraries were amplified by high-fidelity, low-bias PCR using primers complementary to the adapter sequences. The amplified libraries were then subjected to magnetic bead-based cleanup (Axygen, part number MAG-PCR-CL-250) to remove unused primers and assess quantity. Prior to hybridization, samples were normalized by library mass and captured. Pool Each sample consists of 6 to 8 samples. PoolThe libraries were multiplexed into a single library. The xGen Exome Research Panel v2 probe set, which contains complementary custom-designed probes including the TCR / BCR probes obtained by the method in Step A, was used in conjunction with the xGen Universal Blocker (Integrated DNA Technologies, Part Number 1075475) and the xGen Hybridization / Wash Kit (Integrated DNA Technologies, Part Number 1080584) for library hybridization and capture. The enriched targets were amplified using KAPA HiFi HotStart ReadyMix and primers (Roche Diagnostics, Part Number KK2621) and subjected to additional magnetic bead-based cleanup. The quantity and quality of the final libraries were assessed, and success was determined based on a molar concentration calculation that incorporated both quantification and quality qualification measurements.
[0467] Sequencing: The amplified target-capture library was sequenced to an average of 50 million reads on an Illumina NovaSeq 6000 system using patterned flow cell technology.
[0468] Analysis: Repertoire sequencing analysis was performed using TRUST4 v1.0.0 software on RNA sequencing data in the form of FASTQ files containing read pairs. TRUST4 v1.0.0 was run according to the developer's instructions, without modification, using the FASTQ files containing read pairs as input and the human IMGT reference sequence file provided with the software to generate quantitative data related to the TCR and BCR clonotypes (productive, non-productive, and partial) identified by the software. The tabular clonotype report generated by TRUST4 was used to calculate the Shannon entropy of productive clonotypes in each immune receptor chain (IGH, IGK, IGL, TRA, TRB, TRG, and TRD). The rows of the TRUST4 report and additional non-statistical annotations were combined to form the final data table, with column descriptions listed below.
[0469] result A total of 1,957 clonotypes of expanded B and T cells were detected in the specimens, and 1,074 of the detected clonotypes were determined to be productive sequences (e.g., they did not contain stop codons, were not out-of-frame, were not partial sequences, etc.).
[0470] Table 6 (Table 7) shows the top 10 most abundant sequences (e.g., sequences associated with the most supporting sequence reads). Each column represents a clone. The columns closer to the left have more raw abundance (e.g., detected supporting sequence reads) associated with that clone. Clones showing greater gains are closer to the left. In this example, the first (leftmost) IGH CDR3 is an IGH-productive clonotype with an abundance of up to 25% (see the "receptor_productive_frequency" row). Complete results are included as appendices in Appendix I of U.S. Provisional Patent Applications Nos. 63 / 013,130, 63 / 084,459, and 63 / 201,020. The most frequent clones likely correspond to expanded populations of B or T cells. In this case, the expanded population of B cells can be analyzed to track B-cell lymphoma and detect progression, response to treatment, MRD, etc. In one embodiment, the TCR / BCR sequencing methods disclosed herein are utilized on multiple specimens taken from a patient at different time points to track the disease over time.
[0471] Below are row names and descriptions in the appendix as an example of the various data that may comprise the TCR / BCR immune repertoire sequencing data associated with each CDR3 sequence or clonotype. a. Count - the integer number of read fragments supporting clonotypes (e.g., the number of read fragments that align to a given clonotype reference sequence) b. Frequency - Clonotype frequency within the BCR or TCR c. CDR3nt - CDR3 nucleotide sequence d. CDR3aa - CDR3 amino acid sequence (if the sequence is non-productive, "_" means a stop codon, "out_of_frame" means a frameshift mutation, or "partial" means a partial sequence) e. V - Called V gene clonotype assignment {formatted as gene*allele} ("null" means no gene is called) (may include V gene family, V gene, and / or V allele) f. D - Called D gene clonotype assignment {gene * allele} (null if no gene is called or does not apply to the receptor) (may include D gene family, D gene, and / or D allele) g. J - J gene clonotype assignment called {gene * allele} (null if no gene called) (may include J gene family, J gene, and / or J allele) h. C - C gene clonotype assignment {gene} to be called (no allele information is returned for C genes) (null if no gene is called) i. Receptor [type] - {IGH, IGK, IGL, TRA, TRB, TRG, TRD, mixed} (in some instances, mixed may be alpha / delta TCR) j. productive_status - {"in", "partial", "out_of_frame", "internal_stop"} (in means in-frame / productive, partial means partial sequence, out_of_frame means the sequence is frameshifted and not expected to be productive, internal_stop means the sequence has a stop codon and is not expected to be productive) k. receptor_frequency - receptor-specific clonotype frequency (for that receptor) l. receptor_productive_frequency - frequency of productive receptor clonotypes within all productive receptor clonotypes m. V_gene_family - (e.g., IGLV3-25*03 → IGLV3) n. V_gene - (e.g., IGLV3-25*03 → IGLV3-25) o. V_allele - (e.g. IGLV3-25*03→03) p. D_gene_family q. D_gene r. D_allele s. J_gene_family t. J_gene u.J_allele v. IGH_isotype - null if not called / not applicable: {"A1", "A2", "D", "E", "G1", "G2", "G3", "G4", "M"} otherwise w. has_CDR3nt_twin - "True" is entered if a duplicate of the nt sequence of this clonotype exists in the repertoire x. has_CDR3aa_twin - "True" is entered if a duplicate of this clonal aa sequence exists in this repertoire
[0472] In various embodiments, clonotype frequencies and / or genetic clonotype assignments can be determined by TCR or BCR sequence assembly algorithms included in the systems and methods described herein.
[0473] [Table 7A]
[0474] [Table 7B]
[0475] Example 3 TCR / BCR sequence database and its uses In this example, a reference dataset may be generated or an existing reference dataset may be selected. The data may be de-identified. The data may be free of protected health information (PHI). The reference dataset may include TCR / BCR sequencing data linked to annotated clinical documentation, as well as additional NGS-based outputs, including but not limited to, patient HLA-typed or matched NGS DNA / RNA sequencing, viral / pathogen sequencing, whole exome or targeted panel sequencing of patient specimens. Clinical documentation may include: disease characterization and duration, severity of symptoms or disease (e.g., diseases associated with infection by a pathogen), symptom description and / or severity grading, one or more therapies (e.g., cancer therapies such as immunotherapy or vaccines) and duration and outcome, time from disease onset and / or end to sample collection, sample collection site / specimen information (e.g., saliva, blood, mucosa, nasal / anterior nares swab, nasopharyngeal swab, nylon flocked swab, spun polyester swab, nasopharyngeal aspirate, bronchoalveolar lavage, Mawls tube, or Longhorn Primerstore) These may include specimens collected in MTM tubes, nasopharyngeal / nasal / nares or other specimens collected in viral transport media / VTM, feces, etc.), infection status of one or more pathogens determined by one or more diagnostic assays (e.g., PCR-based, isothermal nucleic acid amplification-based, NGS-based, serology-based, array / microarray / array card / open array plate / FilmArray, etc., ELISA, ELISpot, FluoroSpot, antigen-based, rapid antigen test, or other molecular assay), etc. HLA data can be used to further annotate or contextualize the TCR sequence data. For example, certain combinations of TCR sequences and HLA types may be incompatible, meaning that a particular TCR sequence would be expected in the context of a particular HLA type. For example, if a patient lacks a TCR sequence that is expected in the context of exposure to a particular pathogen, the absence of that sequence would be expected if the patient does not have an HLA type that matches that TCR sequence.
[0476] This reference dataset can be mined to identify TCR or BCR sequences enriched in patients responding to or recently recovered from a disease caused by a particular pathogen or combination of pathogens. See, e.g., Emerson, R., DeWitt, W., Vignali, M. et al., "Immunosequencing identifies signatures of cytomegalovirus exposure history and HLA-mediated effects on the T cell repertoire," Nat Genet 49, 659-665 (2017), https: / / doi.org / 10.1038 / ng.3822, the contents of which are incorporated herein by reference in their entirety for all purposes.
[0477] Mining may involve the use of machine learning clustering techniques on TCR / BCR sequence databases. An example method involves the detection of pathogen-associated TCR or BCR sequences in data from patients with specific types of cancer. These sequences can be used as biomarkers, i.e., indicators for predicting response to checkpoint inhibitors or IO. Cross-reactivity means that sequences may be present in a higher proportion of patients, especially when infections with a pathogen are prevalent, making them more likely to be the first sequences found in common across multiple patients.
[0478] The systems and methods can be used to detect receptor sequences that are produced in response to one disease state but are cross-reactive and can be used as adoptive cellular therapy for another disease state. For example, influenza infection or vaccine, or SARS-CoV-2 infection or vaccine, can produce receptor sequences that subsequently attack cancer cells (see https: / / onlinelibrary.wiley.com / doi / 10.1111 / bjh.17116).
[0479] In one example, a patient has non-small cell lung cancer (NSCLC) and a virus-associated TCR sequence. The TCR beta chains are divided into affinity groups (based on similar amino acid structures). Certain virus-associated TCRs cross-react with cancer antigens (in patients with the same HLA). (See the Cell Press immunology article, "Global analysis of shared T cell specificities in human non-small cell lung cancer enables HLA inference and antigen discovery," Chiou et al., https: / / www.sciencedirect.com / science / article / pii / S1074761321000 807.)
[0480] Subsequent observations of these TCR and BCR sequences in the patient can then be analyzed by a predictive model developed based on the reference dataset or a subset of the reference dataset (e.g., only records with data deemed relevant for prediction, including data associated with known negative or positive status or a numerical score associated with a predictive target or category) to calculate the likelihood that the patient has an infection status, contact history, and / or potential protection or resistance to the infection associated with either the TCR and / or BCR sequences, which may be based on associations or trends captured in the reference dataset.
[0481] The reference dataset can also be analyzed to find associations between disease severity and various genetic, immunological, or clinical factors or characteristics, such as alleles or variants associated with ABO blood group genes, genes located on chromosome 9q34.2, immunological genes, genes located on chromosomes 3 or 6, HLA genes, etc., immunological characteristics, clinical data / conditions (age, heart disease, history of diabetes, blood glucose levels, high blood pressure, obesity, asthma, COPD, etc.), and / or the presence of specific TCR and / or BCR sequences.
[0482] The TCR and / or BCR sequences may have been generated in response to the disease-causing pathogen or in response to another pathogen, for example, if the disease is COVID-19, SARS, or MERS, the TCR or BCR may have been generated in response to OC43, HKU1, 229E, and NL63 coronaviruses, and may cross-react with SARS-CoV-2, SARS-CoV-1, MERS, etc.
[0483] By way of example, immunological features associated with severe COVID-19 disease include large populations of activated CD4 T cells, few or no circulating follicular helper T cells (cTfh), activated and / or exhausted CD8 T cell populations, TEMRA-like cells, T-bet+ cells (including plasmablasts), Ki67+ cells (including plasmablasts), few or no memory B cells, a strong or Tbetbright effector-like CD8 T cell response, a weak CD4 T cell response, a diminished lymphocyte response, a strong plasmablast response in the absence of large populations of activated cTfh, or impaired T or B cell responses. (See Mathew et al., Science, September 4, 2020, Vol. 369, No. 6508, the contents of which are incorporated herein by reference in their entirety for all purposes.)
[0484] By way of example, immunological features associated with mild or asymptomatic COVID-19 include SARS-CoV-2-specific T cell responses (e.g., targeting internal viral proteins, viral surface proteins, viral nucleocapsid proteins, viral membranes, or viral spike proteins), persistently functional memory T cell responses, T cells expressing CD38, HLA-DR, Ki-67, PD-1 (or other inhibitory receptors), CCR7, CD127, CD45RA, and / or TCF1, SARS-CoV-2-specific IgG, and inflammatory markers (e.g., in patient plasma). (See Sekine et al., Robust T cell immunity in convalescent individuals with asymptomatic or mild COVID-19, Cell, (2020), doi: https: / / doi.org / 10.1016 / j.cell.2020.08.017, the entire contents of which are incorporated herein by reference for all purposes.)
[0485] Thus, in some embodiments, reference TCR / BCR sequence datasets can be analyzed to find associations between immunological and / or clinical features (including diagnoses) associated with, for example, severe, mild, or asymptomatic COVID-19. This information can then be used, for example, to provide or confirm a diagnosis, predict disease severity, and predict treatment efficacy.
[0486] Example 4 Coronavirus-specific TCR / BCR sequences In this example, a coronavirus cohort (e.g., data collected from a group of patients infected with or exposed to coronavirus and / or a negative control group known not to have been exposed to or affected by coronavirus for a specified period immediately prior to specimen collection) can be analyzed. The coronavirus cohort may be assembled from a reference dataset, such as the reference dataset described in Example 2, particularly a subset of patients from the dataset who are known or suspected to have been exposed to coronavirus. The coronavirus may include individual coronaviruses that infect humans (e.g., SARS-CoV-1, SARS-CoV-2, MERS-CoV, coronavirus HKU1, coronavirus NL63, coronavirus 229E, coronavirus OC43, etc.), or combinations thereof.
[0487] A cohort or reference dataset can be used to determine which TCR and / or BCR sequences are associated with coronavirus infection or exposure. For coronavirus-associated TCR sequences, the system and method can match HLA types as described in Example 1. For example, the method assumes that the top 5 or 10 HLA alleles in a population of interest are present in all patients, unless the patient's HLA type is known.
[0488] The quantity and presence of these TCR and / or BCR sequences in a patient can be used to predict the patient's exposure to coronavirus (including, for example, using a predictive model developed based on the reference dataset in Example 1, a subset of that dataset, or another dataset), predict the patient's likelihood of having mild or severe symptoms if exposed to coronavirus, and / or gauge the patient's response to a vaccine. It is expected that it will be possible to determine whether the presence of those TCRs and / or BCRs confers protection from (or susceptibility to) severe COVID symptoms, which can then be used to stratify patients based on their potential risk of complications following SARS-CoV-2 (or other coronavirus) infection.
[0489] This analysis can also be used to predict which TCR and / or BCR sequences associated with coronavirus exposure are cross-protective against SARS-CoV-2 (or other coronaviruses), measure the abundance of specific sequences across populations, or determine whether TCR-HLA combinations are protective or associated with particular symptom severity for design trials.
[0490] Example 5 Predicting patient susceptibility to pathogens and / or diseases In this example, a specimen from a patient can be analyzed for TCR and / or BCR sequences using the methods disclosed above. The patient may have a respiratory infection and / or flu-like symptoms but may not have received a specific diagnosis regarding which respiratory infection is causing the symptoms. The detected TCR and / or BCR sequences can be analyzed using a predictive model refined using a reference dataset, such as the dataset described in Example 2, to predict the pathogen most likely causing the symptoms and / or the likelihood that the patient will suffer from mild or severe disease (e.g., the patient's susceptibility to disease). In various embodiments, the patient's COVID-19 status (whether or not infected with SARS-CoV-2) is unknown. In another embodiment, the patient has been diagnosed with COVID-19 and / or previously had a positive result on a SARS-CoV-2 diagnostic assay.
[0491] Specimens from patients can also be analyzed to identify evidence of the presence of the pathogen (eg, by the assays listed in Example 1). Pathogens that may be screened by the assay include those commonly associated with respiratory infections and / or influenza-like symptoms, such as SARS-CoV-1, SARS-CoV-2, MERS-CoV, coronavirus HKU1, coronavirus NL63, coronavirus 229E, coronavirus OC43, influenza A, influenza A H1, influenza A H1-2009, influenza A H1N1, influenza A H3, influenza B, influenza C, parainfluenza virus 1, parainfluenza virus 2, parainfluenza virus 3, parainfluenza virus 4, rhinovirus / enterovirus, adenovirus, respiratory syncytial virus, respiratory syncytial virus A, respiratory syncytial virus B, human metapneumovirus, bocavirus, human bocavirus, chlamydia pneumonia, mycoplasma pneumonia, legionella pneumonia, Bordetella, Bordetella holmesii, holmesii, Bordetella pertussis, Streptococcus pneumoniae, Coxiella burnetii, Staphylococcus aureus, Klebsiella pneumoniae, Moraxella catarrhalis, Haemophilus influenzae, Pneumocystis jirovecii, enterovirus D68, Epstein-Barr virus (EBV), mumps, measles, cytomegalovirus, human herpesvirus (HHV-6), varicella-zoster virus (VZV), and parechovirus.
[0492] Example 6 Treatment response assessment – vaccines In this example, TCR and / or BCR sequences according to the methods disclosed above can be analyzed in specimens from vaccine trial subjects or from patients vaccinated in other settings to assess the patient's response to the vaccine and their susceptibility to the disease associated with the vaccine (e.g., the disease that the vaccine is designed to prevent, reduce, or alleviate) or another disease.
[0493] The reference dataset described in Example 2 can be used to determine which TCR and / or BCR sequences are associated with response to the vaccine. The presence or absence of these TCR / BCR sequences, and additional statistics or outcomes associated with the TCR / BCR sequences in a patient, can be used to predict the patient's degree of disease susceptibility. Additional clinical or molecular data as described in the previous examples can also be taken into account in predicting disease susceptibility.
[0494] The report may include the TCR / BCR sequences detected, the predicted disease susceptibility, the basis for the predicted disease susceptibility, and other relevant information.
[0495] In one embodiment, the sample is collected after administering the vaccine to the subject. In another embodiment, multiple samples are collected from the patient, including a first sample collected before administering the vaccine to the patient and a second sample collected after administering the vaccine to the patient. In another embodiment, the vaccine is administered to the patient multiple times, and samples can be collected after each administration of the vaccine.
[0496] Example 7 Clonal expansion of T cells in various cancers and estimation of clonality within the repertoire preface The degree of immune cell infiltration into tumors is influenced by various factors, such as the tumor's immunogenicity, the type of tissue from which the tumor arises, the degree to which immune cells can physically migrate through tumor interstitial material, and the metabolic repression and hypoxia of solid tumors. However, the level of immune filtration can be characteristic of similar tumor types, with brain tumors characteristically exhibiting low immune filtration and lung and skin malignancies exhibiting high infiltration. Even tumors of different origins may have more or fewer distinct clones of lymphocytes present within the tumor microenvironment. Lymphomas and leukemias present unique cases, as malignant T or B cells may arise from a single clone. Thus, samples from patients with lymphoma may contain only a few T cell receptor (TCR) or B cell receptor (BCR) clones in their immune profiles. In summary, sequencing various tumor samples of different cellular origins and developing immune profiles can provide valuable insights into the efficacy of specific immune profiling assays. The development of immune profiles from a variety of malignancies will enable unbiased evaluation of the application of the novel hybrid capture approach described herein to real-world diagnostic challenges.
[0497] method Sample preparation and sequencing was performed as described in Example 2 above.
[0498] Analysis: Analysis was performed as described above for Example 2 (gene expression RUO invasion analysis is known in the art and described, for example, at https: / / www.worldscientific.com / doi / abs / 10.1142 / 9789813279827_0026 and in published patent application Ser. No. 16 / 533,676, which is incorporated herein by reference).
[0499] result We hypothesized that the novel hybrid capture approach disclosed herein would provide accurate and efficient sampling of immune profiles. To test this hypothesis, we obtained RNA samples from 501 individuals, prepared cDNA libraries, and isolated sequences using the novel hybrid capture probe approach described herein. Sequencing generated an average of approximately 20,000 reads for each of the 501 sequenced samples. Sequencing data were graphed to demonstrate the average number of clonotypes present in tumor samples, with productive clonotypes on the X-axis and CDR3-supported read fragments on the Y-axis (Figure 7). Consistent with other reports, brain tumor samples (blue circles) were found to be largely devoid of productive clonotypes, while lung tumor samples showed significant infiltration of diverse lymphoma clonotypes. Furthermore, immune cell profiles demonstrated broad clonal abundance (Figure 8). Furthermore, productive T and B cell reads correlated with putative T and B cell infiltration in tumors (Figures 9 and 10). These data support the hypothesis that novel hybrid capture approaches can be used to efficiently collect informative immune profiling data from a variety of tumor types, especially when the repertoire yield recapitulates tissue-specific expectations of lymphocyte infiltration.
[0500] Next, we further analyzed immune profiles from hematological malignancies. Hematological cancers often consist of very few clonal types due to the clonal nature of tumor cell proliferation. Therefore, we hypothesized that our novel hybrid capture approach would reveal very few clones in samples from leukemia and lymphoma, while other tumor types would reveal numerous clones. We normalized Shannon entropy to the theoretical maximum homogeneity at any given repertoire size, representing the clonal distribution of each receptor. This provided a measure for assessing the expected versus observed clonal diversity in each sample. Samples from T-cell lymphoma or leukemia exhibited low normalized Shannon entropy, while samples from, for example, melanoma, breast cancer, or oropharyngeal carcinoma exhibited high normalized Shannon entropy (Figure 11, left panel). Further decomposition of the relative frequencies of the top 10 most productive TCR beta (TRB) clonotypes is shown in the right panel of Figure 11. Thus, these data support the hypothesis that the novel hybrid capture assay described herein can efficiently isolate and amplify sequences highly relevant to clinical diagnostic decisions. Indeed, application of the novel hybrid capture assay described herein to a diverse collection of tumor samples demonstrated that the assay is effective in capturing diverse immune infiltrates and repertoire differences that reflect known biological trends.
[0501] Example 8 B-cell lymphoma case study: Demonstration of anti-CD19 CAR detection in patients preface Chimeric antigen receptor (CAR) technology utilizes a comprehensive biology approach to treat diseases in an antigen-specific manner. CARs are engineered to act as "surrogate" T cell receptors directed against antigens of choice. Thus, T cells can be isolated from patients, transduced with the CAR, and administered as autologous transplants. CARs are constructed from an antibody's antigen-recognition domain (scFv) and various intercellular signaling domains. When the CAR binds to a cognate antigen, the intercellular signaling domain transduces signals similar to those of the natural T cell receptor, along with costimulatory signals, activating the effector function of the CAR T cell. Successful, durable CAR T cell treatment relies on the engraftment and persistence of the patient's CAR T cells. Therefore, technologies that can longitudinally detect the presence and status of a patient's CAR T cells will aid medical practitioners' ability to make informed treatment decisions.
[0502] method Subject sample preparation and sequencing were performed as described in Example 2.
[0503] Analysis was performed as described above (no special analysis was required to prepare the data for detection of the CDR3 amino acid sequence of this CAR in the rep-seq data).
[0504] Subject history The patient was diagnosed with B-cell lymphoma in 2015. Initial treatment with rituximab led to a complete remission of the disease. The disease recurred in 2017, prompting a second treatment with rituximab, which also led to a complete remission. The patient experienced another disease recurrence in 2019 and was treated with anti-CD19 CAR (axicabtagene-ciloleucel), which led to a complete remission. The patient experienced another recurrence in 2020 and was treated with rituximab and an anti-CD-79b monoclonal antibody. The patient's sample was collected one year after treatment with the anti-CD19 CAR. At this time, the malignant cells were CD19+, CD20- by Cytometry.
[0505] result We hypothesized that the novel hybrid capture and sequencing approach disclosed herein would efficiently and accurately detect CAR T cell engraftment. Therefore, to test this hypothesis, we generated immune profiles of subjects who had successfully received anti-CD19 CAR treatment but subsequently experienced relapse. The anti-CD19 CAR axicabtagene ciloreucel utilizes the FMC63 scFv as the extracellular domain to detect CD19. In various embodiments, any CAR sequence could be detected using the systems and methods described herein. While the IGHJ4 heavy chain is of murine origin, the high sequence homology between humans and mice in this region should allow detection using the novel hybrid capture approach, in this example, designed for human use. In various embodiments, additional probes specific for the desired CAR sequence could be added to the systems and methods disclosed herein to increase the sequencing reads corresponding to the CAR sequence. Immune profiling revealed that the overall repertoire was 60% or less the size of the comparable B-cell lymphoma repertoire. Additionally, T cells represented a high proportion of the subject's repertoire (Figure 12). One interpretation of this result is that extensive rituximab treatment reduced the B cell repertoire. Notably, this approach detected 20 / 164 IGH-aligned reads that mapped to anti-CD19 CAR (Figure 12, yellow asterisk). Thus, these data demonstrate that the novel hybrid capture approach can efficiently and accurately detect anti-CD19 CAR scFv sequences in subjects.
[0506] Example 9 COVID-19 Case Study: Demonstrating Compatibility of rep-seq Data with External Data - Detection of Putative SARS-CoV-2-Specific TCRs in COVID-19 Patients preface Antigen-specific immune cell repertoires contain potentially important information about a subject's immunological history. For example, a subject's T and B cell repertoire reflects their exposure to pathogens. Circulating antibodies directed against pathogens may also fade over time. In contrast, pathogen-specific T cells may persist indefinitely. Therefore, sequencing a subject's immune repertoire and generating an immune profile may be a more efficient and accurate indicator of exposure than, for example, serology testing.
[0507] The pandemic caused by SARS-CoV-2 has led to incredible loss of life and human suffering worldwide. However, in response to this pandemic, unprecedented efforts and resources have been invested in immunological research. Through these efforts, immune repertoires have been obtained and published from individuals infected with SARS-CoV-2. Therefore, immune profiles generated using the novel hybrid capture approach disclosed herein can be generated from individuals infected with SARS-CoV-2 and validated against externally generated data.
[0508] method Sample preparation, sequencing and analysis were performed as described above.
[0509] result The novel hybrid capture approach disclosed herein efficiently generates high-quality immune profiles from patient samples. To test the hypothesis that the novel hybrid capture approach can detect pathogen-specific TCR sequences, immune profiles were generated from individuals with a history of SARS-CoV-2 infection. Immune profiling revealed 47 TCR beta and 56 TCR alpha clonotypes. The TCR beta clonotypes were then compared to a public database of SARS-CoV-2-specific TCR beta clonotypes developed using multiplexed identification of T cell receptor antigen assays (MIRA) (PMID: 32793896). This repertoire contains 160,000 TCR beta clonotypes with affinity for SARS-CoV-2 peptides and can be considered a positive control panel for SARS-CoV-2 exposure and / or infection. Interestingly, four TCR beta clonotypes matched the clonotypes discovered by the MIRA assay (Figure 13). The four CDR3 reads that matched the MIRA assay data were CASSIGVNTEAFF (11 reads found in 509 COVID-19+ repertoires, purple asterisk in Figure 13), CASSLSGGPYNEQFF (7 reads found in 30 COVID-19+ repertoires, yellow asterisk in Figure 13), CASSSGIQPQHF (not detected in 500 COVID-19-validation samples), and CASSVSYEQYF (not detected in 500 COVID-19-validation samples). These data support the hypothesis that the novel hybrid capture approach described herein can efficiently and accurately identify immune profiles (TCR / BCR profiles) in people infected with pathogens such as SARS-CoV-2.
[0510] Example 10 In one example, a TCR / BCR profile can be generated for a patient suffering from colorectal cancer. The patient is believed to have a KRAS P12D mutation and also an HLA C08.02 allele known to be present in this altered KRAS peptide. The TCR / BCR profile can be analyzed for CDR3 sequences predicted to recognize the altered KRAS peptide.
[0511] Subject sample preparation and sequencing can be performed as described in Example 2. Analysis can be performed as described in Example 2 above.
[0512] The novel hybrid capture approach described herein can be used to generate immune profiles of people with colorectal cancer. In some embodiments, the patient carries a KRAS P12D mutation. In further embodiments, the patient carries an HLA C08.02 allele known to present as a mutated KRAS peptide as a neoantigen. Thus, in some embodiments, immune profiles generated from individuals with the above KRAS mutation and the specific HLA C08.02 allele include TCR clonotypes that recognize KRAS neoantigens generated by the P12D mutation in KRAS. In some embodiments, these clonotype sequences can be used to select lymphocyte clones for generating patient-specific precision medicine therapies. In some embodiments, putative neoantigen-specific clonotypes obtained from a patient's immune profile can be used to assemble a database of such clonotypes. In some embodiments, the database of neoantigen-associated clonotypes can be further used for comparison with patients with unknown KRAS mutation status. In some embodiments, a database of neoantigen-associated clonotypes can be used to diagnose patients with KRAS P12D mutations based on the patient's immune profile, in addition to or without the need for a biopsy and additional sequencing steps.
[0513] In some embodiments, an immune profile generated from a colorectal cancer patient can be used to select a suitable therapy for treating the tumor. In further embodiments, the therapy selected for colorectal cancer treatment based on the immune profile is selected from cytotoxic chemotherapy, targeted therapy, such as a Janus kinase inhibitor, or immunotherapy. In some embodiments, the selected immunotherapy is selected from the group consisting of checkpoint inhibitor therapy, CAR T cell therapy, CAR M therapy, cancer vaccines, or other immuno-oncology treatment modalities.
[0514] In some embodiments, the TCR / BCR profile of a colorectal cancer patient can be used to generate a report detailing the most abundant T and B cell clones. In some embodiments, the top 10 most frequent clones can be displayed. The most frequent clones are believed to represent expanded populations of B or T cells. In one embodiment, the TCR / BCR sequencing methods disclosed herein are utilized on multiple samples taken from a patient at different time points to track the disease over time.
[0515] Example 11 In one example, a TCR / BCR profile can be generated for a patient with non-small cell lung cancer (NSCLC) and an EGFR mutation, and the TCR / BCR profile can be analyzed for CDR3 sequences predicted to recognize peptides encoded by the mutant EGFR gene.
[0516] Subject sample preparation and sequencing can be performed as described in Example 2. Analysis can be performed as described in Example 2 above.
[0517] The novel hybrid capture approach described herein can be used in conjunction with sequencing to generate immune profiles of people with non-small cell lung cancer. In some embodiments, the patients harbor EGFR mutations. In some embodiments, immune profiles generated from people harboring EGFR mutations include TCR clonotypes that recognize EGFR neoantigens. In some embodiments, these clonotype sequences can be used to select lymphocyte clones for generating patient-specific precision medicine therapies. In some embodiments, putative neoantigen-specific clonotypes obtained from a patient's immune profile can be used to assemble a database of such clonotypes. In some embodiments, the database of neoantigen-associated clonotypes can be further used for comparison with patients with unknown EGFR mutation status. In some embodiments, the database of neoantigen-associated clonotypes can be used to diagnose patients with EGFR mutations based on their immune profile, in addition to or without the need for biopsy and additional sequencing steps.
[0518] In some embodiments, immune profiles generated from non-small cell lung cancer patients can be used to select a suitable therapy for tumor treatment. In further embodiments, the therapy selected for colorectal cancer treatment based on the immune profile is selected from cytotoxic chemotherapy, targeted therapy, such as a Janus kinase inhibitor, or immunotherapy. In some embodiments, the selected immunotherapy is selected from the group consisting of checkpoint inhibitor therapy, CAR T cell therapy, CAR M therapy, cancer vaccines, or other immuno-oncology treatment modalities.
[0519] In some embodiments, the TCR / BCR profile of an NSCLC patient can be used to generate a report detailing the most abundant T and B cell clones. In some embodiments, the top 10 most abundant clones can be displayed. The most abundant clones are believed to represent expanded populations of B or T cells. In one embodiment, the TCR / BCR sequencing methods disclosed herein are utilized on multiple samples taken from a patient at different time points to track the disease over time.
Claims
1. a) isolating RNA from a patient sample; b) Enriching isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes and enriching for a targeted whole transcriptome panel using a collection of transcriptome hybrid capture probes; c) sequencing the RNA of (b) to generate sequencing data; and d) Analyzing the sequencing data to determine the patient's TCR / BCR profile.
1. A method for determining a TCR / BCR profile of a patient, comprising: The collection of TCR / BCR hybrid capture probes includes a first pool containing BCR constant region probes, a second pool containing BCR non-constant region probes, a third pool containing TCR constant region probes, and a fourth pool containing TCR non-constant region probes, and the collection of TCR / BCR hybrid capture probes and the collection of transcriptome hybrid capture probes are contained in one reaction well.
2. 2. The method of claim 1, wherein the ratio of the first pool to the second pool to the third pool to the fourth pool within the population is 1:2.5:100:
100.
3. The method described in claim 1, wherein the ratio of the first pool to the second pool to the third pool to the fourth pool to the set of transcriptome hybrid capture probes is 1:2.5:100:100:
100.
4. The method of claim 3, wherein 2% or less of the reads in the sequencing data map to TCR / BCR genes.
5. The method of claim 1, wherein step (c) comprises whole transcriptome sequencing or short read sequencing.
6. The method of claim 1, wherein step (d) comprises identifying multiple TCR / BCR clones in the sample.
7. The method of claim 1, wherein step (d) comprises identifying the most abundant TCR / BCR clone in the sample.
8. 2. The method of claim 1, wherein step (d) comprises identifying the most abundant non-constant region sequences in the sample.
9. 10. The method of claim 1, wherein the sample is a blood sample or a solid tumor sample.
10. a) isolating RNA from a patient sample; b) Enriching isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes and enriching for a targeted whole transcriptome panel using a collection of transcriptome hybrid capture probes; c) sequencing the RNA of (b) to generate sequencing data; and d) Analyzing the sequencing data to determine the patient's TCR / BCR profile.
1. A method for determining a TCR / BCR profile of a patient, comprising: a first pool containing BCR constant region probes, a second pool containing BCR non-constant region probes, a third pool containing TCR constant region probes, and a fourth pool containing TCR non-constant region probes; wherein the ratio of transcriptome hybrid capture probes to the first pool, to the second pool, to the third pool, and to the fourth pool is 10:1:2.5:100:100; wherein 2% or less of the reads in the sequencing data map to TCR / BCR genes; and wherein the set of TCR / BCR hybrid capture probes and the set of transcriptome hybrid capture probes are contained in one reaction well.
11. The method of any one of claims 1 to 10, wherein step (c) comprises whole transcriptome sequencing or short read sequencing.
12. 12. A composition comprising a collection of TCR / BCR hybrid capture probes for use in the method of any one of claims 1 to 11, comprising comparing a patient's BCR / TCR profile with a control TCR / BCR profile, and identifying the patient as having a disease or medical condition based on the comparison.
13. 13. The composition of claim 12, wherein the disease or condition is an infectious disease, cancer, an autoimmune disease, or an allergy.
14. 14. The composition of claim 13, wherein the cancer or infectious disease is one or more of those listed in embodiment 114.
15. The composition of any one of claims 1 to 11, wherein the analysis comprises determining the presence or degree of lymphocytic infiltration of the tumor.
16. The composition of any one of claims 12 to 14, wherein the method further comprises treating the patient with a therapeutic agent.
17. 17. The composition of claim 16, wherein the therapeutic agent is an immunotherapeutic agent.
18. 18. The composition of claim 17, wherein the immunotherapeutic is a vaccine.
19. 18. The composition of claim 17, wherein the immunotherapeutic is a chimeric antigen receptor (CAR) T cell.
20. a) isolating RNA from a patient sample; b) Enriching isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes and enriching for a targeted whole transcriptome panel using a collection of hybrid capture probes; c) sequencing the RNA of (b) to generate sequencing data; and d) analyzing the sequencing data, wherein the analysis includes determining the most abundant TCR / BCR clones in the sample and optionally determining the TCR / BCR profile of the patient; the collection of TCR / BCR hybrid capture probes comprises a first pool comprising BCR constant region probes, a second pool comprising BCR non-constant region probes, a third pool comprising TCR constant region probes, and a fourth pool comprising TCR non-constant region probes; and (e) Treating patients Including, A set of TCR / BCR hybrid capture probes and a set of transcriptome hybrid capture probes are contained in one reaction well; A composition comprising a collection of TCR / BCR hybrid capture probes for use in a method of treating a disease or condition in a patient.
21. 21. The composition of claim 20, wherein the ratio of the first pool to the second pool to the third pool to the fourth pool within the population is 1:2.5:100:
100.
22. 21. The composition of claim 20, wherein the ratio of the first pool to the second pool to the third pool to the fourth pool to the transcriptome hybrid capture probes within the collection is 1:2.5:100:100:
100.
23. 23. The composition of claim 22, wherein 2% or less of the reads in the sequencing data map to TCR / BCR genes.
24. 21. The composition of claim 20, wherein the treatment comprises expanding the most abundant TCR / BCR clones in vitro and administering the expanded clones to the patient.
25. 21. The composition of claim 20, wherein step (d) comprises identifying the most abundant TCR non-constant region sequences in the sample, and wherein the treatment administered in step (e) comprises administering a CAR-T cell therapy, and wherein the CAR-T cells comprise at least one of the most abundant TCR non-constant region sequences.
26. a) at a first time point prior to administration of the therapeutic agent; i) isolating RNA from a patient sample; ii) Enriching isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes and enriching for a targeted whole transcriptome panel using a collection of transcriptome hybrid capture probes; iii) sequencing the RNA of (b) to generate sequencing data; and iv) analyzing the sequencing data to determine the patient's TCR / BCR profile; b) at a second time point after administration of the therapeutic agent; i) isolating RNA from a patient sample; ii) Enriching isolated RNA for TCR / BCR genes using a collection of TCR / BCR hybrid capture probes and enriching for a targeted whole transcriptome panel using a collection of transcriptome hybrid capture probes; iii) sequencing the RNA of (b) to generate sequencing data; and iv) analyzing the sequencing data to determine the patient's TCR / BCR profile; and c) comparing the TCR / BCR profile determined in step (a) with the TCR / BCR profile determined in step (b) to characterize the effect of the treatment on the patient's TCR / BCR profile.
1. A composition comprising a collection of TCR / BCR hybrid capture probes for use in a method for characterizing the effect of a treatment on a patient's TCR / BCR profile, comprising: The collection of hybrid capture probes includes a first pool containing TCR constant region probes, a second pool containing TCR non-constant region probes, a third pool containing BCR constant region probes, and a fourth pool containing BCR non-constant region probes, and the collection of TCR / BCR hybrid capture probes and the collection of transcriptome hybrid capture probes are contained in one reaction well.
27. 27. The composition of claim 26, wherein the ratio of the first pool to the second pool to the third pool to the fourth pool within the population is 1:2.5:100:
100.
28. The composition described in claim 27, wherein the ratio of the first pool to the second pool to the third pool to the fourth pool to the transcriptome hybrid capture probe is 1:2.5:100:100:
100.
29. 29. The composition of claim 28, wherein 2% or less of the reads in the sequencing data map to TCR / BCR genes.
Citation Information
Patent Citations
JP1101025676A
JP1142978981A
JP16/533
PCT/US19/69149
PCT/US19/69161