Determination of the amount of lymphocytes in a mixed sample

JP2024525963A5Pending Publication Date: 2025-07-24THE FRANCIS CRICK INST LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024504149
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-03-11
Filing Date
2022-07-22
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Current methods for estimating the amount of tumor-infiltrating lymphocytes (TILs) in mixed samples, such as tumor samples, are not universally advanced and require high sequencing depth, are not suitable for addressing aneuploidy in non-T cell populations, and lack accuracy in predicting responses to cancer immunotherapy.

Method used

A method using whole-exome sequencing (WES) to calculate lymphocyte fractions by analyzing signals from T cell receptor excision circles (TRECs) during VDJ recombination, particularly in the TCRA gene, to estimate T and B cell percentages in samples, which correlates with orthogonal immune infiltration scores and provides predictive value for immunotherapy response.

Benefits of technology

The method accurately estimates lymphocyte fractions, correlating with RNA sequencing and histopathological scores, and predicts response to cancer immunotherapy, offering superior predictive value over mutation burden, applicable to both tumor and blood-derived samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method is described for determining the lymphocyte fraction in a mixed sample containing genomic material from multiple cell types. The method includes obtaining a read depth profile for a sample along a predetermined genomic region of interest that includes at least a portion of a genomic locus that undergoes VDJ recombination, obtaining a plurality of read depth ratios (ri) by normalizing the read depth with reference to a baseline read depth derived from a subset of the regions of interest, obtaining a summarized read depth ratio value (rVDJ) for the subset of the regions of interest that are likely to be deleted by VDJ recombination, and determining a lymphocyte fraction (f) in the sample as a function of the summarized read depth ratio value (rVDJ). Methods of providing a diagnosis or prognosis based on lymphocyte fraction, and related systems and products are also described.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Field of Disclosure The present disclosure relates to methods for estimating lymphocyte fractions, such as T lymphocytes or B lymphocytes, in a mixed sample containing multiple cell types, such as a tumor sample or a blood sample, based on a read depth profile derived from the sample, as well as systems and related products for carrying out such methods. Also described is a method for providing a prognosis for a cancer patient based at least in part on an estimated tumor infiltrating lymphocyte fraction from a tumor sample from the patient. [Background technology]

[0002] background Lymphocytes are an important part of the human body machinery that plays a key role in human health and disease. For example, in the case of tumors, the immune system is widely recognized as one of the key factors influencing both tumor formation and subsequent cancer progression. T cells in particular play a key role in the cancer immune microenvironment by eliminating neoantigen-bearing cancer cells, and therefore the number of T cells in tumor samples is an important clinical factor. In fact, the status of the tumor immune microenvironment in clinical presentations can be a prognostic and response predictor for immunotherapy. In recent years, immunotherapy, especially checkpoint inhibitors (CPIs), has emerged as an innovative treatment for many cancer types. Indeed, malignant melanoma and non-small cell lung cancer (NSCLC) have shown remarkable improvements in both disease-free and overall survival rates after anti-CTL4 or anti-PDL1 therapy (Robert et al., 2011; Schadendorf et al., 2015; Topalian et al., 2012). However, responses to immunotherapy are not universal, and clinically beneficial responses occur only in a fraction of patients (Goodman et al., 2017). Therefore, determining which patients will benefit from CPIs is of paramount importance.

[0003] Current data suggest that response to CPIs is primarily governed by two features: immune stimuli, such as the presence of neoantigens, which can be predicted from next-generation sequencing (NGS); and the presence of T cells that can respond to that stimuli. The presence of infiltrating T cells has long been associated with improved survival in cancer types such as ovarian cancer (Zhang et al., 2003) and colorectal cancer (Galon et al., 2006). More recent studies have focused on the predictive potential of neoantigens, with increased tumor mutation burden (TMB) emerging as one of the best predictors of response to immunotherapy (Samstein et al., 2019). A recent study by Litchfield et al. (2020) systematically investigated how well various potential biomarkers of CPI response perform in a pan-cancer CPI-treated dataset. The two best predictive features identified were clonal TMB, reflecting mutations present in all cancer cells, and RNA-sequencing-derived transcriptomic signatures of T cells (e.g., CD8A RNA expression), indicating the presence of tumor-infiltrating lymphocytes (TILs).

[0004] Therefore, estimating both the presence and quantity of TILs is of great importance in cancer treatment. However, at present, there is no universal advanced method to estimate TILs. Microarray signatures such as CIBERSORT (Newman et al., 2015), and CIBERSORTx for RNA-seq (Newman et al., 2019), ESTIMATE (Yoshihara et al., 2013), and RNA-seq signatures such as those compiled by Danaher et al. have been used to quantify the transcriptomic signatures of TILs in tumor samples. Alternatively, the presence of immune infiltration can be determined based on hematoxylin and eosin (H&E) histopathological slides or, alternatively, by appropriate cell-specific staining, most commonly immunohistochemical characterization. If more phenotypic knowledge of T cells is required, sequencing of T cell receptors (TCRs) can be performed after enrichment for sequences encoding the highly variable VDJ recombined CDR3 regions of the TCR (Bolotin et al., 2012). All these methods require acquisition of dedicated data with the aim of identifying immune cells in mixed samples, such as tumor samples. Recently, in McGrath et al. (2017), an approach was proposed to characterize the TCR repertoire in a sample using WES data by isolating rearranged reads of T cells (reads that span the hypervariable CDr3 and contain both V and J segments). However, this approach requires a very high sequencing depth, cannot address the presence of aneuploidy in non-T cell populations (which is common, for example, in the case of cancer), and is only suitable for characterizing the TCR repertoire. Summary of the Invention [Problem to be solved by the invention]

[0005] Thus, there is a need for new methods for determining the amount of lymphocytes in a mixed sample that alleviate one or more of the shortcomings of existing methods. [Means for solving the problem]

[0006] Brief description of the invention Broadly speaking, the inventors recognized that whole exome sequencing (WES) and even whole genome sequencing (WGS) are frequently performed on tumor samples to calculate tumor mutation burden and identify actionable mutations for targeted therapy. Similarly, this type of data is increasingly being collected in other disease contexts, such as neonatal care. However, at present, WES cannot be used to accurately infer immune infiltration levels. Therefore, the inventors develop a method to estimate T cell (or B cell) fractions from WES / WGS or similar data using signals derived from VDJ recombination, e.g., T cell receptor excision circles (TRECs) during VDJ recombination of the T cell receptor alpha (TCRA) gene. This score directly measures the proportion of T cells (and sometimes B cells) in the sample. The inventors further demonstrated that this score significantly correlates with an orthogonal immune infiltration score based on RNA sequencing and histopathological TIL scoring in the TRACERx100 cohort. Using a meta-analysis of patient cohorts undergoing immunotherapy, we confirmed that this score can predict response to cancer immunotherapy, providing predictive value that surpasses mutational burden. We further demonstrated the broad applicability of the method by demonstrating that T cell fraction can also be calculated from blood-derived germline samples. Thus, the method broadens the possibilities for studying the immune microenvironment of WES (or WGS) tumor samples, which is particularly useful when orthogonal data are not available.

[0007] Accordingly, in a first aspect, the present disclosure provides a method for determining a lymphocyte fraction in a mixed sample comprising genomic material from a plurality of cell types, comprising obtaining a read depth profile of the sample comprising read depths along predetermined genomic regions of interest comprising at least a portion of a genomic locus undergoing VDJ recombination, and determining a plurality of read depth ratios (r) by normalizing the read depths along the predetermined regions of interest with reference to a baseline read depth derived from a subset of the regions of interest. i) and a summarized read depth ratio value (r VDJ ) and the fraction of lymphocytes in the sample (f) was calculated based on the summarized read depth ratio value (r VDJ and determining the temperature as a function of

[0008] The method may have any one or more of the following optional features.

[0009] A mixed sample may be a cell or tissue sample (where the sample may contain multiple cell types, each containing their respective genomic material), or a sample of nucleic acid derived therefrom.

[0010] The genomic locus undergoing VDJ recombination may be selected from the TCRA, TCRB, TCRG, TCRD, IGH, IGL, or IGK loci. Thus, the lymphocyte fraction may be a T lymphocyte fraction, a B lymphocyte fraction, an αβT lymphocyte fraction, a γδT lymphocyte fraction, a B or T cell fraction comprising any gene segment at the TCRA, TCRB, TCRG, TCRD, IGH, IGL, or IGK loci, or a B or T cell fraction comprising any gene segment listed in Table 8 or a corresponding gene segment in another species (the coordinates listed in Table 8 are merely indicative, and in particular, corresponding coordinates in another reference sequence may be used). The genomic locus undergoing VDJ recombination may be selected from the TCRA, TCRB, TCRG, or IGH loci. The predetermined region of interest may include one or more exons from the TCRA, TCRB, TCRG, TCRD, IGH, IGL, or IGK loci. The predetermined region of interest may include one or more exons from the TCRA, TCRB, TCRG, or IGH locus. The predetermined region of interest may include a plurality of exons from the TCRA, TCRB, TCRG, TCRD, IGH, IGL, or IGK locus, the plurality of exons including at least one exon corresponding to each of the V, D, and C segments of the respective loci. The plurality of exons may include at least one exon corresponding to each of the V, D, J, and C segments of the respective loci. The plurality of exons may include exons corresponding to multiple V segments of the respective loci. The predetermined region of interest may include a region encoding multiple V, D, J, and / or C segments from the TCRA, TCRB, TCRG, TCRD, IGH, IGL, or IGK locus. The predetermined region of interest may include one or more regions in multiple V, D, J, and / or C segments from the TCRA, TCRB, TCRG, TCRD, IGH, IGL, or IGK locus. A given region of interest may include all exons from the TCRA, TCRB, TCRG, TCRD, IGH, IGL, or IGK locus.A given region of interest may include all of the TCRA, TCRB, TCRG, TCRD, IGH, IGL, or IGK locus (i.e., including both exons and introns).

[0011] The predetermined region of interest may include a subset of exons from the TCRA, TCRB, TCRG, TCRD, IGH, IGL, or IGK loci. The predetermined region of interest may include a subset of the TCRA, TCRB, TCRG, TCRD, IGH, IGL, or IGK loci. For example, the predetermined region may exclude any exons or regions determined to be associated with systematic bias, such as, for example, bias due to the platform used to collect the read data, bias associated with GC content at the exon or region, or bias associated with coverage at the exon or region. For example, some exon capture kits have been shown to be associated with bias that results in coverage at certain exons that deviates from expectations. Similarly, certain regions may have lower or higher coverage than expected with whole genome sequencing, for example, due to artifacts in the sequencing or alignment process. Such exons or regions may be identified by comparing read coverage between datasets that are not expected to be subject to the same systematic bias, such as those acquired using different platforms (e.g., different capture kits). For example, the average logR ratios of multiple candidate exons or regions can be calculated for a first sample set and a second sample set (where the two sample sets are not expected to be subject to the same systematic bias), and the median across each sample set can be compared for each of the multiple candidate exons or regions to identify exons or regions that differ significantly between the two sample sets. For example, a difference between the median logR values ​​that exceeds a threshold (e.g., 0.5) can be used as a criterion to eliminate exons / regions that are likely to be subject to bias. Regions with lower or higher than expected coverage can be identified by fitting a model, e.g., a generalized linear model, to the average read depth ratio data across the multiple samples and removing any regions of a predetermined size (e.g., 100bp, 200bp, 500bp, 1000bp) that are associated with a summarized value (e.g., the mean or median) that is not within a predetermined distance from the fitted model.Alternatively or additionally, the regions with lower or higher coverage than expected can be identified by determining lymphocyte fractions for a plurality of samples, and performing a genome-wide association study to identify single nucleotide polymorphisms associated with the determined lymphocyte fractions, where the presence of association indicates a bias in coverage in a region of a predetermined size that includes the single nucleotide polymorphism. Thus, the method can include determining lymphocyte fractions for a plurality of samples, performing a genome-wide association study to identify single nucleotide polymorphisms associated with the determined lymphocyte fractions, selecting samples that include the single nucleotide polymorphisms, fitting a model to the average read depth ratio data across the selected samples, and removing regions of a predetermined size (e.g., 100bp, 200bp, 500bp, 1000bp) associated with summarized values ​​(e.g., mean or median) that are not within a predetermined distance from the fitted model, thereby identifying regions with lower or higher coverage than expected in whole genome sequencing.

[0012] The read depth or read depth ratio may be processed to correct for biases associated with the read depth data acquisition platform used. For example, the read depth / read depth ratio may be GC corrected. GC correction of read depth data may be performed using a linear model that captures the effect of GC content on read depth at every position, with the residuals of the model representing the normalized read depth value. The effect of GC content on read depth may be calculated at the exon or region level (GC exon / GC region ) and a term reflecting GC content (where the region is, for example, 1000 bp (in this case, GC region "GC 1000bp ) that reflects the GC content at the level of the entire region of interest (e.g., using a 1000 bp window along the region of interest), and a term that reflects the GC content obtained from a model fitted to the smoothed data, e.g., a linear model). macro ) can be expressed as: lm(basepair~GC exon +GC 2 exon +GCmacro +GC 2 macro ), or equivalently lm(basepair~GC region +GC 2 region +GC macro +GC 2 macro ) (especially lm(basepair~GC 1000bp +GC 2 1000bp +GC macro +GC 2 macro A linear model of the form r = 0.05, where "basepair" is the read depth at every position can be used. Residuals of such a model can be used as corrected coverage values. The corrected (e.g., GC corrected) read depth values ​​can be used to calculate baseline read depths derived from a subset of regions of interest, and summarized read depth ratio values ​​(r ) for the subset of regions of interest that are likely to be deleted by VDJ recombination. VDJ) can be determined. For example, the median of the corrected values ​​(or other summarized criteria, e.g., a statistical measure of centrality) for each subset of regions of interest can be used. If the subset of regions of interest includes multiple genomic locations (e.g., the beginning and end of a region likely to undergo VDJ recombination), one of the multiple medians can be used, e.g., the highest value. The read depth profile can be obtained from sequencing data, preferably whole exome sequencing, whole genome sequencing (including low-pass WGS, also called shallow genome sequencing), or panel sequencing data. When the read depth profile is obtained from panel sequencing data, the panel sequencing data includes data regarding a given genomic region of interest. For example, the data may have been obtained using capture probes targeting multiple genomic regions including the given genomic region of interest. The read depth profile can be obtained from sequencing data having a sequencing depth of at least 0.1x, at least 0.5x, at least 2x, at least 5x, at least 10x, preferably at least 15x, preferably at least 20x, more preferably at least 30x. The read depth profile may be obtained from whole exome sequencing data having a sequencing depth of at least 10x, preferably at least 15x, preferably at least 20x, more preferably at least 30x. The read depth profile may be obtained from whole genome sequencing data having a sequencing depth of at least 0.1x, at least 0.5x, at least 2x, at least 5x, or at least 10x, preferably at least 2x. The read depth profile may be obtained from sequencing data that has been filtered to remove exons with a sequencing depth below a threshold, such as 10x, preferably 15x. The read depth profile may be obtained from sequencing data that has been filtered to remove regions (defined using a fixed size window, such as 1000bp) with a sequencing depth below a threshold, such as 2x.Sequencing data that contains less than a predetermined number of exons in a given genomic region of interest with a sequencing depth above a threshold (e.g., 10x or 15x) may be discarded. For example, data that contains less than 30 exons in a given region of interest (e.g., the TCRA locus) with a sequencing depth above a threshold (e.g., 15x) may be discarded. In such cases, new sequence data may be obtained for the sample, or the sequence data may be used to investigate another given genomic region of interest for which the data has sufficient coverage.

[0013] The summarized read depth value may be a summarized log read depth ratio. The log may be a base 2 logarithm. The fraction of lymphocytes in the sample (f) may be multiplied by the summarized read depth ratio value (r VDJ ) is used to determine the fraction of lymphocytes (f) in a sample as a function of the summarized read depth ratio value (r VDJ ), the fraction of abnormal cells that are aneuploid within a given region of interest in a mixed sample (p), and the copy number of abnormal cells within a given region of interest (Ψ T ) as a function of the fraction of abnormal cells in the mixed sample (p) and the copy number of abnormal cells within a given region of interest (Ψ T ) can be determined by obtaining an abnormal cell-specific copy number profile using methods known in the art, such as, for example, the ASCAT method described in Van Loo et al., or the pureCN method described in Riester et al. (2016) and available at http: / / bioconductor.org / packages / PureCN / . Determining the lymphocyte fraction (f) in the sample can be used to determine the degree of abnormal cell-specific copy number profile.

number

number

[0014] The subset of regions of interest likely to be deleted by VDJ recombination may include the region between the end of the last V segment of the genomic locus undergoing VDJ recombination and the start of the J segment of the genomic locus undergoing VDJ recombination. The subset of regions of interest likely to be deleted by VDJ recombination may include one or more D segments of the genomic locus undergoing VDJ recombination, such as all D segments. The subset of regions of interest likely to be deleted by VDJ recombination preferably does not include a V or C segment. Alternatively or additionally, the subset of regions of interest likely to be deleted by VDJ recombination does not include a J segment. The subset of regions of interest likely to be deleted by VDJ recombination may include, consist of, or be included in the gap between the last V gene segment and the first segment encoding a J segment of the genomic locus undergoing VDJ recombination, or in the case of TCRA, the first segment encoding a portion of the TCR δ chain, or any corresponding segment in another predetermined region of interest. For example, looking at the TCRA locus, we found it appropriate to define the location of maximal VDJ recombination as the gap between the last TCRA V segment and the first segment of the TCR delta gene (having a size of 80000 bp and rounded to start at position 22800000 (resulting in the region chr14:22800000-22880000). Other definitions are possible and small changes in the size of this region are not expected to adversely affect the performance of the method. Regions of maximal VDJ recombination can be defined using a similar approach for other VDJ loci, i.e. by defining a region that is at least partially located around or within the region between the last V gene segment and the first J gene segment. Regions of maximal VDJ recombination can be defined for each VDJ locus using the genomic coordinates in Table 7 or the corresponding coordinates for other species or reference genomes.The subset of regions of interest likely to be deleted by VDJ recombination can include regions corresponding to any V or J segment of a genomic locus that undergoes VDJ recombination, which can be used when the fraction of lymphocytes is the fraction of B or T cells that contain any gene segment at the TCRA, TCRB, TCRG, TCRD, IGH, IGL, or IGK locus, or the fraction of B or T cells that contain any gene segment listed in Table 8 or a corresponding gene segment in another species.

[0015] The method may further include fitting the model to a plurality of read depth ratios and generating a summarized read depth ratio value (r ) for a subset of the regions of interest that are likely to be deleted by VDJ recombination. VDJ) may include obtaining a summarized read depth ratio value based on the values ​​of the model in a subset of the regions of interest that are likely to be deleted by VDJ recombination. The model may capture the log read depth ratio as a function of genomic location along a given genomic region of interest. The multiple read depth ratios may be referred to herein as a "read depth ratio profile." As will be appreciated by those skilled in the art, the multiple read depth ratios are mutually organized along genomic coordinates, and thus, the reference to fitting a model to the multiple values ​​refers to fitting the model to a data series that includes the multiple read depth ratios as a function of genomic coordinates. The model may be a generalized linear model, such as a generalized additive model, a piecewise constant model, or a Bayesian model. The model may be a piecewise constrained linear model with multiple cut points selected to correspond to the positions of multiple V and J genes in the region of interest. For example, the model may be a piecewise constrained linear model with multiple cut points selected from the start and end of the genes listed in Table 8, from the genomic locations listed in Table 8, or from corresponding genes or locations in other species or reference genomes. The model may be a generalized linear model, and the sequence data may have been obtained using a panel capture method such as whole exome sequencing. The model may be a piecewise constrained linear model, and the sequence data may have been obtained using whole genome sequencing. The piecewise constrained linear model may have one or more (or all) of the following additional constraints: the model may have a value of 0 before the first V segment in the given genomic region of interest; the model may have a value of 0 after the last J segment in the given genomic region of interest; the model may be such that the value of each subsequent V segment in the given genomic region of interest has a lower value along the genomic coordinate than the value of the preceding segment; the model may be such that the value of each subsequent J segment in the given genomic region of interest has a higher value along the genomic coordinate than the value of the preceding segment. Any approach that fits a smooth curve (whether continuous constant or piecewise constant) to the data can be used in the context of this disclosure.As another example, the model may be a Bayesian model that models segment usage from read depth and a prior distribution of usage of each of the multiple V and J genes in the region of interest. The prior distribution of usage of each of the multiple V and J genes in the region of interest may be a uniform distribution. Such a distribution may define equal likelihood that each gene is used. The prior distribution of usage of each of the multiple V and J genes in the region of interest may be a distribution that uses prior knowledge of V and J gene (segment) usage in a reference population. For example, T cell receptor sequencing data obtained from a reference population may be used to determine the distribution of V and J gene usage in the reference population. The reference population may be a human population. The reference population may be a human population with one or more specific characteristics selected from age, sex, and HLA type. The model may be a Bayesian model, and the sequence data may have been obtained using whole genome sequencing. The Bayesian model may be implemented using a gradient-based Markov Chain Monte Carlo (MCMC) method to calculate a probability distribution for segment usage of each V and J gene in the region of interest.

[0016] In general, the summarized value may be the median, the mean, the trimmed mean, or any other known statistical criterion of centrality of multiple values. The summarized read depth ratio value for a region may be a statistical criterion of centrality for multiple read depth ratio values ​​within the region. In an embodiment where a model (e.g., a generalized additive model) is fitted to the read depth ratio profile, the summarized read depth ratio value for a region may be obtained as a statistical criterion of centrality for the values ​​of the fitted model within the region. For example, the average of the fitted model over the entire region may be used. In an embodiment where a piecewise constant model is fitted to the read depth ratio profile, the summarized read depth ratio value for a region may be obtained as the value of the fitted model for the segment corresponding to the region or multiple segments corresponding to the region, or the value of the model for the segments within the region. For example, using a piecewise constant model constrained as described above, the summarized read depth ratio value for a subset of regions of interest that are likely to be deleted by VDJ recombination may be selected as the minimum value of the model (which is also equal to the maximum deviation of the model from 0).

[0017] When the fraction of lymphocytes is the fraction of B or T cells that contain any gene segment at the TCRA, TCRB, TCRG, TCRD, IGH, IGL, or IGK locus, or the fraction of B or T cells that contain any gene segment listed in Table 8 or the corresponding gene segment in another species, the subset of regions of interest that are likely to be deleted by VDJ recombination may correspond to the gene segments. The summarized read depth ratio value for a gene segment may be the difference between the summarized read depth ratio value for the gene segment and the summarized read depth ratio value for the preceding gene segment, or the absolute value of the difference. The summarized read depth ratio value may be a summarized log read depth ratio. The summarized read depth ratio may be obtained from the value of the model fitted to the read depth ratio. Any logarithm may be a base 2 logarithm.

[0018] The method may further include obtaining a likelihood for the model fitted to the plurality of read depth ratios, and a likelihood below a predetermined threshold indicates excessive noise in the sample. The method may further include obtaining a plurality of candidate models and one or more statistical criteria (such as likelihood and / or confidence interval) associated with the fitting of the plurality of candidate models, and selecting the candidate model based on the statistical criteria. For example, the model with the highest likelihood with a confidence interval within a predetermined range may be selected. Obtaining a plurality of candidate models may include obtaining a predetermined number of models, such as 10, 30, 50, or 100 models.

[0019] The method may further include obtaining a baseline read depth from a subset of regions of interest. The baseline read depth may be a summarized read depth value across all subjects of the region of interest. For example, the baseline read depth may be a median read depth across a given subset of regions of interest. The subset of regions of interest used to obtain the baseline read depth may be a subset of regions of interest that are unlikely to be deleted by VDJ recombination. The subset of regions of interest that are unlikely to be deleted by VDJ recombination may include the first n V segments of a genomic locus that undergoes VDJ recombination. The subset of regions of interest that are unlikely to be deleted by VDJ recombination may include a C segment of a genomic locus that undergoes VDJ recombination. Preferably, the subset of regions of interest used to obtain the baseline read depth does not include a D or J segment of a genomic locus that undergoes VDJ recombination. Indeed, according to the methods disclosed herein, it is possible to use the beginning, the end, or both the beginning and the end (but not both) of the VDJ locus under investigation, or other adjacent regions that can be assumed to have copy number values ​​that are not affected by VDJ recombination. A subset of regions of interest used to obtain baseline read depth can be defined for each VDJ locus using the genomic coordinates in Table 7, or corresponding coordinates for other species or reference genomes.

[0020] Without wishing to be bound by theory, it is believed that the use of longer normalization regions (e.g., using both leading and trailing regions and / or increasing the length of the leading region) advantageously compensates for the noise typical in read depth data. For example, the leading region can be extended to include a larger number of V segments that are less likely to be frequently lost, and also to include regions outside the gene when data are available (e.g., when using whole genome sequencing data). Furthermore, since signals associated with VDJ recombination are closer to the end of the gene than to the beginning of the gene (as J segments are clustered near the end of the gene, whereas V segments are more widely spread), it is believed that the use of signals at the beginning of the gene is beneficial. Furthermore, it is also possible to use regions outside the VDJ locus under investigation, for example, before or after the locus. However, as the distance from the locus of interest increases, it becomes more likely that the signals seen correspond to different copy number events in the cancer cells and therefore do not reflect the tumor copy number at the VDJ locus under investigation. The parameter n can be determined by using appropriate training data to select a value of n that results in the most accurate estimate of lymphocyte fraction across the training data when compared to, for example, an orthogonal criterion for lymphocyte fraction. The orthogonal criterion for lymphocyte fraction can be obtained, for example, based on transcriptomic and / or histopathological data. The parameter n can be between 8 and 16, between 10 and 14, for example 10, 11, or 12.

[0021] Obtaining the read depth profile may include obtaining a plurality of raw read depth values ​​over a predefined region of interest and smoothing the read depth values. Smoothing the read depth values ​​may include replacing the raw read depth values ​​by corresponding rolling median values ​​over a window of predefined width. The predefined width of the window may be selected between about 20 bp and about 200 bp, or between about 50 bp and about 150 bp. For example, the read depth profile may include values ​​obtained by calculating rolling median values ​​in a 50 bp window along the predefined region of interest. The appropriate size of the window can be empirically determined for a particular data set, type of data, or situation by comparing the estimate of lymphocyte fraction obtained using the present method to various candidate window sizes for an orthogonal measure of lymphocyte fraction, such as the Danaher score.

[0022] The mixed sample may be a blood sample or a tumor sample. The method may further include obtaining from the subject a mixed sample comprising genomic material from a plurality of cell types. The method may further include obtaining from the subject read depth data from the mixed sample comprising genomic material from a plurality of cell types by one or more in vitro steps. The one or more in vitro steps may include nucleic acid extraction, library preparation, target sequence capture, sequencing, and / or hybridization to a copy number array. As will be appreciated by those skilled in the art, once data indicating the presence of nucleic acids in the sample is obtained, for example by sequencing, additional steps may be applied, such as, for example, quality control steps, reference genome alignment steps (also referred to as mapping), normalization steps, etc., as known in the art or as described herein. The method may further include repeating the method of any embodiment of this aspect on one or more additional mixed samples, where the initial mixed sample and the additional mixed sample are obtained from the same subject, and the initial mixed sample and the additional mixed sample are tumor samples. The method may further include obtaining one or more summarized lymphocyte fractions based on the lymphocyte fractions (f) for the first mixed sample and the further mixed sample. The summarized lymphocyte fractions may be obtained as the average fraction across the first mixed sample and the further mixed sample, the minimum fraction across the first mixed sample and the further mixed sample, the maximum fraction across the first mixed sample and the further mixed sample, and / or the fold difference between the minimum fraction across the first mixed sample and the further mixed sample and the maximum fraction across the first mixed sample and the further mixed sample. When multiple mixed samples are available for a subject, one or more of the summarized lymphocyte fractions may be used to provide a prognosis, as described further below, instead of the individual lymphocyte fractions for a single sample. For example, the minimum fraction across multiple samples, the minimum and maximum fraction across multiple mixed samples, and / or the fold difference between the minimum fraction across the first mixed sample and the further mixed sample and the maximum fraction across the first mixed sample and the further mixed sample may be used to provide a prognosis for the subject.Preferably, a prognosis is provided using the minimum fraction across multiple samples.

[0023] The method may further include providing the determined lymphocyte differential fraction and / or any value derived therefrom to a user, for example via a user interface. Values ​​derived therefrom may include diagnostic or prognostic indicators.

[0024] As will be appreciated by those skilled in the art, the complexity of the operations described herein (due to at least the amount of data typically generated by genomic DNA sequencing) is beyond the scope of human mental activity, and therefore, unless the context indicates otherwise (such as when sample preparation or acquisition steps are described), all steps of the methods described herein are computer-implemented.

[0025] The method of the present disclosure can be used in many clinical situations where different diagnostic and / or prognostic predictions may be associated with different levels of lymphocytes in samples from subjects. Thus, the present disclosure also relates to a method for stratifying a population of subjects according to lymphocyte differential status (e.g., tumor infiltrating lymphocyte (TIL) status), comprising performing the method of the first aspect of the present disclosure on one or more samples (e.g., tumor samples) from each subject in the population. For example, subjects can be divided into "immune cold" and "immune hot" groups depending on the lymphocyte differential estimated for one or more tumor samples of the patient. This can be particularly applied to identifying patients for personalized medicine and / or clinical trials. Patients classified as immune cold are expected to be less responsive to immunotherapy, such as the use of checkpoint inhibitors, compared to patients classified as immune hot. Thus, immune hot may be more likely to benefit from immunotherapy and immune cold may be more likely to benefit from alternative treatment options. Alternatively or additionally, classified patients may have a worse survival prognosis than patients classified as immune hot. Thus, the methods of the present disclosure can also be used to classify subjects diagnosed with cancer into at least two groups with different prognostic and / or sensitivity predictions to one or more immunotherapies.

[0026] Thus, according to a further aspect, the present disclosure provides a method of providing a prognosis for a subject diagnosed with cancer, comprising determining lymphocyte fraction in one or more tumor samples from the subject using the method of any embodiment of the first aspect. Also described is a method of monitoring a subject diagnosed with cancer, comprising determining lymphocyte fraction in one or more tumor samples from the subject obtained at a first time point and one or more tumor samples from the subject obtained at a further time point using the method of any embodiment of the first aspect. For example, the first time point may be before treatment and the further time point may be after treatment. Alternatively, both the first time point and the further time point may be after treatment. The treatment may involve any anti-cancer therapy, including immunotherapy. The method of providing a prognosis or monitoring a subject may further comprise classifying the sample as immune hot or immune cold depending on the lymphocyte fraction in the sample (such as, for example, T lymphocyte fraction). For example, samples with lymphocyte fraction above a threshold (e.g., about 0.1, i.e., about 10%) may be classified as immune hot, and samples with lymphocyte fraction below the threshold may be classified as immune cold. The method may further include classifying subjects into good and bad prognosis groups depending on the number of samples from the patient classified as immune cold. For example, when analyzing samples from different tumor regions, a number of samples classified as immune cold equal to or greater than a threshold (e.g., 2) may classify the subject into a bad prognosis group. The bad prognosis group may be associated with a reduced recurrence-free survival and / or reduced overall survival compared to the good prognosis group. The threshold for lymphocyte fraction and / or number of immune cold samples may be determined using a suitable training cohort, for example, by evaluating the performance of various thresholds that can classify patients between groups associated with significantly different prognosis predictions. Also described is a method of determining whether a subject diagnosed with cancer is likely to respond to immunotherapy, comprising determining the lymphocyte differential in one or more tumor samples from the subject using the method of any embodiment of the first aspect.The method may further include administering immunotherapy to the subject diagnosed as likely to respond to immunotherapy. The method may include recommending treatment with immunotherapy to the subject diagnosed as likely to respond to immunotherapy. The method may include administering an alternative therapy (such as conventional chemotherapy or radiation therapy) and / or recommending treatment with the alternative therapy to the subject if the subject is diagnosed as unlikely to respond to immunotherapy.

[0027] The method may further include classifying the subject into a group likely to respond to immunotherapy and a group unlikely to respond to immunotherapy. Indeed, the inventors have demonstrated that the T cell fraction in the tumor, and even in the blood sample, determined according to the present disclosure, is a predictor of response to immunotherapy. For example, the method may further include classifying the sample as immune hot or immune cold depending on the lymphocyte fraction in the sample (such as T lymphocyte fraction). For example, a sample having a lymphocyte fraction above a threshold (e.g., about 0.1, i.e. about 10%) may be classified as immune hot, and a sample having a lymphocyte fraction below the threshold may be classified as cold. Then, if one or more samples from the subject are immune cold, the subject may be classified into a group unlikely to respond to immunotherapy. Alternatively, if a sample from the subject has a lymphocyte fraction below the threshold, the subject may be classified into a group unlikely to respond to immunotherapy, and otherwise into a group likely to respond to immunotherapy.

[0028] The immunotherapy may be any therapy that modulates the function of the immune system and / or utilizes existing immune function to treat cancer. The immunotherapy may be selected from T cell transfer therapy (tumor infiltrating lymphocyte (or TIL) therapy and CAR-T cell therapy), therapeutic antibodies, cancer treatment vaccines, checkpoint inhibitor therapy, and immune system regulators (e.g., interferons and interleukins). The method may include classifying subjects into groups that are likely to respond to immunotherapy (e.g., CPI therapy) and groups that are unlikely to respond to immunotherapy (e.g., CPI therapy) based on lymphocyte differentials in one or more samples from the subject and the estimated tumor mutation burden for the subject. Such classification may be obtained using a multivariate classification model (e.g., a classification method using a trained generalized linear model). The parameters of the classification model (e.g., thresholds, parameters of the multivariate model) may be determined using an appropriate training cohort, for example, using a group of subjects with known immunotherapy response status.

[0029] Also described is a method of determining whether a subject has T-cell lymphopenia, comprising determining lymphocyte fraction in one or more samples from the subject using the method of any embodiment of the first aspect. The one or more samples may be blood samples. The subject may be a neonatal subject. The method of determining whether a subject has T-cell lymphopenia may be performed in the context of a method for screening for severe combined immunodeficiency. The method may further comprise treating the subject for lymphopenia or recommending treatment for lymphopenia to the subject.

[0030] Also described is a method of diagnosing a subject as having a disease, disorder, or condition associated with abnormal lymphocyte counts, comprising determining lymphocyte differentials in one or more samples from the subject using the method of any embodiment of the first aspect. Also described is a method of monitoring a subject diagnosed as having a disease, disorder, or condition associated with abnormal lymphocyte counts, comprising determining lymphocyte differentials in one or more samples from the subject using the method of any embodiment of the first aspect. The disease, disorder, or condition associated with abnormal lymphocyte counts may be an autoimmune disease, a bone marrow disease, or an infectious disease.

[0031] According to a further aspect, the present disclosure provides a method for producing a method for manufacturing a semiconductor device comprising: A system comprising at least one processor and at least one non-transitory computer readable medium comprising instructions, the instructions, when executed by the at least one processor, include operations for obtaining a read depth profile of the sample comprising read depths along predetermined genomic regions of interest comprising at least a portion of a genomic locus undergoing VDJ recombination; and generating a plurality of read depth ratios (r) by normalizing the read depths along the predetermined regions of interest with reference to a baseline read depth derived from a subset of the regions of interest. i ) and summarized read depth ratio values ​​(r VDJ ) and the fraction of lymphocytes in the sample (f) are calculated based on the summarized read depth ratio value (r VDJ and determining, by at least one processor, the magnitude of the difference between the first and second inputs of the signal as a function of the signal strength. The system according to this aspect may comprise any of the features mentioned in relation to the first aspect.

[0032] A system according to this aspect may be configured to perform the method of any embodiment of the previous aspect. In particular, at least one non-transitory computer-readable medium may include instructions that, when executed by at least one processor, cause the at least one processor to perform operations including any operations described in connection with the first aspect. The system may further include one or more of the following in operative connection with the processor: a user interface, where the instructions cause the processor to provide at least the estimated amount of lymphocytes or a value derived therefrom to the user interface for output to a user; one or more sequence read depth data acquisition devices (e.g., a sequencing device, etc.); one or more data stores, such as a sequence read depth data store.

[0033] According to a further aspect, a non-transitory computer-readable medium is provided that includes instructions that, when executed by at least one processor, cause the at least one processor to perform a method of any embodiment of any aspect described herein.

[0034] According to a further aspect, there is provided a computer program comprising code which, when the code is executed on a computer, causes the computer to perform the method of any embodiment of any aspect described herein.

[0035] The present invention includes the described aspects and combinations of preferred features except where expressly not permissible or explicitly avoided. These and further aspects and embodiments of the invention are described in further detail below with reference to the accompanying examples and figures. [Brief description of the drawings]

[0036] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1] 1 is a flow chart that generally illustrates a method for determining lymphocyte differentials in a mixed sample. [Diagram 2] FIG. 1 illustrates one embodiment of a system for determining lymphocyte fraction in a mixed sample. [Figure 3A] FIG. 1 shows a schematic of the biological concept behind the determination of lymphocyte fraction in a mixed sample using VDJ recombination signals as described herein, and how the read depth ratio signal is used to detect gain or loss events of somatic copy number alterations in a standard tumor and matched germline analysis, where the cells consist of three different cell types: tumor cells, T cells, and all other stromal cells. [Figure 3B] Figure 1 shows a schematic of the biological concept behind the determination of lymphocyte fraction in mixed samples using VDJ recombination signals as described herein, and how this same process operates when focusing on the TCRA gene in the context of VDJ recombination and removal of TRECs. The bottom left panel shows an increase in the number of breakpoints detected in the TRACERx100 dataset within the TCRA gene compared to the surrounding region of 14q, suggesting that the TREC signal is captured. [Figure 4A] 1 is a plot showing an example of read depth ratios in two TRACERx100 regions showing high levels of T cell content in blood compared to matched tumors. The VDV segments represent variable segments of both the TCR alpha and TCR delta loci. [Figure 4B] 1 is a plot showing an example of read depth ratios in two TRACERx100 regions showing high levels of T cell content in tumors compared to matched blood. The VDV segments represent variable segments of both the TCR alpha and TCR delta loci. [Diagram 5] FIG. 4 is a schematic illustrating how the signal shown in FIG. 3 can be converted into an estimate of T cell fraction according to an embodiment of the present disclosure. [Figure 6A] FIG. 1 shows theoretical values ​​for naive TCRA T cell fraction scores for various regional copy number and tumor purity values. The points and lines are colored by the actual T cell fraction, and the solid horizontal lines indicate where the points should be. [Figure 6B]FIG. 1 shows the distribution of regional copy numbers of the TCRA gene across the TRACERx100 cohort. [Figure 6C] FIG. 1 shows the distribution of tumor purity across the TRACERx100 cohort. [Figure 6D] Figure 1 shows the exact and naive scores by copy number for the TRACERx100 cohort. As can be seen, the naive score is overestimated when the copy number is 1, and underestimated for other copy numbers above 2. [Figure 7] FIG. 1 shows the results of validation of the approach described herein using cell line data. Pie charts of calculated TCRA T cell fractions from WES of either T cell-derived or non-T cell-derived cell lines are shown. [Figure 8A] FIG. 1 shows the results of validating the approach described herein using simulated data, showing calculated versus actual T cell fractions in a simulated dataset with a background TCRA copy number of 2 and T cell fraction values ​​ranging from 0.01 to 0.99. [Figure 8B] FIG. 1 shows the results of validating the approach described herein using simulated data, showing simulated log read depth ratios from a sample consisting of 24% T cells and 75% tumor (TCRA copy number=1). [Figure 8C] FIG. 1 shows the results of validating the approach described herein using simulated data, showing the difference between calculated and actual naive T cell fractions for various tumor copy numbers and purities. [Figure 8D] FIG. 1 shows the results of validating the approach described herein using simulated data, showing the difference between TCRA T cell fractions and actual fractions for various tumor copy numbers and purities. [Figure 9A]FIG. 1 shows the results of validating the approach described herein by comparison to orthogonal measures of immune infiltration in TRACERx100 data, showing the association between estimated T cell fractions and Danaher RNA-based scores for CD8+, T cells, and total TILs. [Figure 9B] FIG. 1 shows validation of the approach described herein by comparison to orthogonal measures of immune infiltration in TRACERx100 data, showing associations between pathological TIL scores and measures of CD8+ T cell content based on RNA-seq (Danaher, Davoli, EPIC, TIMER, CIBERSORT, and xCell) or DNA (TIL ExTRECT and CDR3 VDJ scores). [Figure 9C] FIG. 1 shows the results of validating the approach described herein by comparison with orthogonal measures of immune infiltration in TRACERx100 data, showing the association between estimated TCRA T cell fractions and RNA-based scores for various immune-related cell types from Danaher, Davoli, EPIC, TIMER, CIBERSORT, and xCell methods. Scores are ordered according to the strength of association with TCRA T cell fractions as assessed by Spearman's rho coefficient. [Figure 9D] FIG. 1 shows the validation of the approach described herein by comparison to orthogonal measures of immune infiltration in TRACERx100 data, showing the association between iDNA-based CDR3 VDJ read scores and TCRA T cell fractions. [Figure 10A] FIG. 1 shows the results of validating the approach described herein by comparison with orthogonal measures of immune infiltration, showing lymphocyte fractions of TCGA LUAD and LUSC samples in TRACERx100 data calculated from methylation data compared with TCRA T cell fractions calculated with CIBERSORT (A). [Figure 10B]FIG. 1 shows the results of validating the approach described herein by comparison with orthogonal measures of immune infiltration, comparing lymphocyte fractions of TCGA LUAD and LUSC samples in TRACERx100 data calculated from methylation data with CD8+ T cell fractions calculated by CIBERSORT. [Figure 11A] Figure 1 shows the results of validating the approach described herein regarding the impact of read coverage (A-B) or potential FFPE-related changes (C) on estimated TCRA T cell fractions, showing downsampling of five TRACERx100 regions to different depths. [Figure 11B] FIG. 1 shows the results of validating the approach described herein on the impact of read coverage (A-B) or potential FFPE-related changes (C) on estimated TCRA T cell fractions, showing simulated downsampling of data to various depth levels. [Figure 11C] Figure 1 shows the results of validating the approach described herein regarding the impact of read coverage (A-B) or potential FFPE-related changes (C) on estimated TCRA T cell fraction. TCRA T cell fraction (non-GC corrected) values ​​are shown for FFPE and fresh frozen samples for bladder tumor and melanoma within the CPI1000+ cohort. [Figure 11D] Figure 1 shows the results of validating the approach described herein on the impact of read coverage (A-B) or potential FFPE-related changes (C) on estimated TCRA T cell fractions, downsampling the five TRACERx100 regions with the highest CDR3 read counts to various depths and the resulting CDR3 read counts. [Figure 12A] FIG. 1 shows the results of using TCRA T cell fraction to derive a prognostic index within the TRACERx100 cohort. Kaplan-Meier curves showing the difference in survival between high-risk and low-risk patients, defined by the number of areas with <10% T cell content, as inferred from the TCRA T cell fraction estimator. [Figure 12B]FIG. 1 shows the results of using TCRA T cell fraction to obtain a prognostic indicator within the TRACERx100 cohort. Kaplan-Meier curves for LUAD and LUSC histological subtypes in TRACERx100 are shown. [Figure 12C] FIG. 1 shows the results of using TCRA T cell fraction to obtain a prognostic indicator within the TRACERx100 cohort. Kaplan-Meier curves for LUAD and LUSC histological subtypes in TRACERx100 are shown. [Figure 13A] Figure 1 shows the results of the analysis of TCRA T cell fraction within the TCGA cohort. Kaplan-Meier curves showing overall survival in TCGA LUAD and LUSC, with patients separated by the median TCRA T cell fraction within each group. [Figure 13B] Figure 1 shows the results of the analysis of TCRA T cell fraction within the TCGA cohort. Kaplan-Meier curves showing overall survival with LUAD alone, with patients separated by the median TCRA T cell fraction within each group. [Figure 13C] Figure 1 shows the results of the analysis of TCRA T cell fraction within the TCGA cohort. Kaplan-Meier curves showing overall survival in LUSC alone, with patients separated by the median TCRA T cell fraction within each group. [Figure 14] Figure 1 shows the results of an analysis of TCRA T cell fraction within the TRACERx100 cohort. Box plots show the difference in TCRA T cell fraction estimates between blood and tumor in TRACERx100. [Figure 15] Figure 1 shows the results of an analysis of TCRA T cell fractions within the TRACERx100 cohort. Comparison of blood and tumor T cell fractions separated by histological features. Areas from the same patient are connected with lines, the type of line depending on whether all tumor areas have high (long dashed line), low (solid line), or heterogeneous T cell fractions (short dashed line) compared to matched blood T cell fractions in TRACERx100. [Figure 16A]FIG. 1 shows the results of an analysis of TCRA T cell fraction within the TRACERx100 cohort, with box plots showing the difference in TCRA T cell fraction between blood and tumor samples within the TRACERx100 cohort. [Figure 16B] FIG. 1 shows the results of an analysis of the TCRA T cell fraction within the TRACERx100 cohort, showing predictors of blood TCRA T cell fraction within the TRACERx100 cohort. [Figure 16C] Figure 1 shows the results of an analysis of TCRA T cell fractions within the TRACERx100 cohort. Kaplan-Meier curves are shown for the multi-region TRACERx100 cohort, divided by LUAD (top) and LUSC (bottom) divided by the number of cold regions present in the tumor (increasing from left to right). Hot and cold regions were defined by using the mean of all tumor regions (0.08095) as the threshold. In each Kaplan-Meier curve, included patients were limited to those with a total region greater than the number of cold regions used to define the threshold. [Figure 16D] Figure 1 shows the results of an analysis of TCRA T-cell fraction within the TRACERx100 cohort, showing hazard ratios for separate Cox regression models relating disease-free survival to various multi-domain measures associated with TCRA T-cell fraction for the entire TRACERx100 cohort, and for LUAD and LUSC patients individually. Relative domain immune evasion score is defined as the maximum domain TCRA score divided by the minimum domain (upper 95% CI) TCRA score. [Figure 16E] FIG. 1 shows the results of an analysis of TCRA T cell fraction within the TRACERx100 cohort. Cox proportional hazards model for the entire TRACERx100 cohort including minimum and maximum area TCRA T cell fraction, tumor stage, age, and sex. [Figure 16F]Figure 1 shows the results of an analysis of TCRA T cell fractions within the TRACERx100 cohort. Kaplan-Meier curves are shown for the multi-region TRACERx100 cohort, divided by LUAD (top) and LUSC (bottom) divided by the number of cold regions present in the tumor (increasing from left to right). Hot and cold regions were defined by using the median of all tumor regions (0.0736) as the threshold. In each Kaplan-Meier curve, included patients were limited to those with a total region greater than the number of cold regions used to define the threshold. [Figure 17A] FIG. 1 shows the results of an analysis of TCRA T cell fraction within the TCGA cohort. Box plot showing the difference in TCRA T cell fraction between blood and tumor samples for LUAD within the TCGA cohort. [Figure 17B] FIG. 1 shows the results of an analysis of TCRA T cell fraction within the TCGA cohort. Box plot showing the difference in TCRA T cell fraction between blood and tumor samples for LUSC within the TCGA cohort. [Figure 17C] FIG. 1 shows the results of an analysis of TCRA T cell fraction within the TCGA cohort, and is a box plot showing the difference in TCRA T cell fraction between LUAD and LUSC in blood within the TCGA cohort. [Figure 17D] FIG. 1 shows the results of an analysis of TCRA T cell fraction within the TCGA cohort. Box plot showing the difference in TCRA T cell fraction between LUAD and LUSC in tumor samples. [Figure 17E] FIG. 1 shows the results of an analysis of TCRA T cell fraction within the TCGA cohort, and predictors of blood TCRA T cell fraction within the TCGA LUAD and LUSC cohorts. [Figure 17F] Figure 1 shows the results of an analysis of TCRA T cell fraction within the TCGA cohort. Kaplan-Meier curves of overall survival and progression-free survival in the TCGA LUAD cohort are shown, where the mean of the TCGA LUAD cohort (0.109) was used as the threshold to divide the cohort into immune hot and cold groups. [Figure 17G]FIG. 1 shows the results of an analysis of TCRA T-cell fraction within the TCGA cohort, showing hazard ratios from separate Cox regression models for TCRA T-cell fraction for the TCGA LUAD and LUSC cohorts for both overall survival (OS) and progression-free survival (PFS). [Figure 17H] Figure 2 shows the results of an analysis of TCRA T cell fraction within the TCGA cohort. Log2 (hazard ratios) from Kaplan-Meier plots for TCGA, where tumor samples were divided into hot and cold based on different thresholds from 0 to 0.16 (in increments of 0.0025) for overall survival and progression-free survival. [Figure 17I] Figure 1 shows the results of an analysis of TCRA T cell fraction within the TCGA cohort. Kaplan-Meier curves for overall survival and progression-free survival for the TCGA LUSC and TCGA LUAD & LUSC cohorts are shown, using the mean of the TCGA LUAD cohort (0.109) as the threshold for distinguishing hot from cold tumors. [Figure 18A] FIG. 13 is a box plot showing gender differences in TCRA T cell fractions in blood samples in the TRACERx100 cohort. [Figure 18B] Box plot showing sex differences in TCRA T cell fractions in blood samples in the TCGA cohort. [Figure 18C] FIG. 11 is a box plot showing the association between age and TCRA T cell fraction for germline blood samples in the RACERx100 cohort. [Figure 18D] Box plot showing the association between age and TCRA T cell fraction for germline blood samples in the TCGA cohort. [Figure 19] FIG. 1 shows an overview of cohorts in the CPI1000+ dataset. [Figure 20] Figure 1 shows a violin plot showing significant differences in calculated tumor TCRA T cell fractions for non-responders and responders across the CPI1000+meta cohort. The dotted black line at 0.1 separates tumors classified as either immune hot or cold. [Figure 21] Figure 1 shows a plot showing tumor TCRA T cell fractions versus clonal TMB in CPI1000+ data. The dashed lines separate the cohort into four quadrants with high / low mutational burden and hot / cold tumors. The inset pie chart shows the proportion of patients who responded to CPI therapy. [Figure 22] Figure 1. Results of univariate meta-analysis of predictors of immune response across multiple cohorts containing at least 10 patients of each cancer type and both DNA and RNA-seq data. Left panel: Forest plot showing OR values ​​and associated p-values ​​of different clinical factors with respect to predictive value of response. Right panel: Heat map of OR values ​​across individual studies in the CPI1000+ dataset. Focus on cohorts with both RNA-seq and TCRA T-cell fraction available. [Figure 23] FIG. 1 shows ROC plots for GLM models for predicting CPI response in the CPI1000+ dataset (1: clonal TMB, 3: clonal TMB+TCRA T cell fraction, 2: clonal TMB+CD8A expression, 4: clonal TMB+TCRA T cell fraction+CD8A expression). [Figure 24A] Figure 1 shows a cohort overview of the CPI lung dataset. The red line in the top panel reflects the median TCRA T cell fraction in patients with responses (0.10) or no responses (0.0070) to CPI. Note that tumor TCRA T cell fractions are often zero. [Figure 24B] FIG. 1 shows the results of a univariate meta-analysis across three NSCLC CPI datasets that included DNA sequencing data but not RNA sequencing data. [Diagram 25]TCRA T cell fraction for all regions in the TRACERx100 cohort as a function of the number of TCRA V-segments used in read depth normalization (see step 2 in FIG. 5 ) (top), the percentage of regions in the TRACERx100 cohort that are non-zero as a function of the number of segments used (middle), and the -log(p-value) of correlation with the Danaher T cell gene signature as a function of the number of V-segments used for normalization (bottom). [Figure 26A] FIG. 1 shows an investigation of the impact of bias in the exon capture process used in TCGA data. Plots showing examples of log read depth ratios for exons in a single sample from the TCGA LUAD cohort. Certain segments, particularly the J gene exons, show a bimodal distribution indicating coverage bias. [Figure 26B] Figure 1 shows an investigation of the impact of bias in the exon capture process used in TCGA data, showing the median log ratio values ​​for a selected range of exons in the TRACERx100 (blue) or TCGA LUAD (red) cohorts. Many exons in the TCGA LUAD cohort have much lower coverage and therefore lower log ratio values. [Figure 26C] FIG. 1 shows an investigation of the impact of bias in the exon capture process used in TCGA data. TCRA T cell fractions using a reduced exon set for the TCGA cohort before GC correction. [Figure 26D] FIG. 1 shows an investigation of the impact of bias in the exon capture process used in TCGA data. TCRA T cell fractions using a reduced exon set for the TCGA cohort after GC correction. [Figure 27A] FIG. 1 shows an investigation of infection-associated determinants of TCRA T cell fraction in blood and malignant tissues. Association of microbial leads from KRAKEN with TCRA T cell fraction in blood samples. [Figure 27B]FIG. 1 shows an investigation of infection-associated determinants of TCRA T cell fraction in blood and malignant tissues. Association of microbial leads from KRAKEN with TCRA T cell fraction in tumor samples. [Figure 27C] FIG. 1 shows an investigation of infection-associated determinants of TCRA T cell fraction in blood and malignant tissues. LUAD shows microbial species for which normalized read counts are significantly associated with TCRA T cell fraction in tumors. [Figure 27D] FIG. 1 shows an investigation of infection-associated determinants of TCRA T cell fraction in blood and malignant tissues. Normalized read counts show microbial species significantly associated with intratumoral TCRA T cell fraction in LUSC. [Figure 28A] FIG. 1 shows the investigation of determinants of TCRA T cell fraction in blood and malignant tissues using multi-region data. Summary plots of the PNE cohort including multi-region microdissected tissue paired with normal blood samples are shown. [Figure 28B] FIG. 1 shows investigation of determinants of TCRA T cell fraction in blood and malignant tissues using multi-domain data. Predictors of blood TCRA T cell fraction in the PNE cohort. [Figure 28C] FIG. 1 shows investigation of determinants of TCRA T cell fraction in blood and malignant tissues using multi-domain data. Predictors of blood TCRA T cell fraction in the PNE cohort. [Figure 28D] FIG. 1 shows the investigation of determinants of TCRA T cell fraction in blood and malignant tissues using multi-region data. The effect of tumor sample location on TCRA T cell fraction in the Messaoudene et al. breast cancer cohort. [Figure 28E] FIG. 1 shows the investigation of determinants of TCRA T cell fraction in blood and malignant tissues using multi-region data. The effect of tumor sample location on TCRA T cell fraction in the TRACERx100 cohort of Messaoudene et al. [Figure 28F]FIG. 1 shows the investigation of determinants of TCRA T cell fraction in blood and malignant tissues using multi-domain data. Linear model for TCRA T cell fraction in PNE samples from genomic factors. [Figure 28G] (a) Examination of determinants of TCRA T cell fraction in blood and malignant tissues using multi-domain data. Log10 p-values ​​for 59 microbial species tested for association with TCRA T cell fraction in blood and tumor samples in LUAD and LUSC. Red line represents significance threshold at P=0.000423. [Fig. 28H] FIG. 1 shows the investigation of determinants of TCRA T cell fraction in blood and malignant tissues using multi-region data. Summary of mean TCRA T cell fraction in the PNE cohort. [Figure 29A] Figure 1. Investigation of the association between subclonal SCNAs and immune heterogeneity within a multi-sample pan-cancer cohort. Shows an overview of immune heterogeneity across a multi-sample pan-cancer cohort. The dashed red line is the TCRA T cell fraction of 0.1. [Figure 29B] FIG. 13. Investigation of the association between subclonal SCNA and immune heterogeneity within a multi-sample pan-cancer cohort, showing the proportion of patients in one of three categories: uniformly immune hot (defined herein as all regions having a TCRA T cell fraction of ≥0.1), uniformly immune cold (defined herein as all regions having a TCRA T cell fraction of <0.1), or heterogeneous, including both hot and cold regions. [Figure 29C]Figure 1. Examination of the association of subclonal SCNA with immune heterogeneity within a multi-sample pan-cancer cohort. Shown is a subset of patients (n=58) from a multi-sample with heterogeneous immune infiltrates, defined as having both a pair of regions with very similar TCRA T cell fractions (pairwise difference <0.1) and a pair of regions with very different TCRA T cell fractions (pairwise difference ≥0.1), plotted against pairwise SCNA heterogeneity, defined as the total proportion of the genome with unique SCNA alterations in either region in the comparison. Data are averaged per tumor to account for bias from tumors with a large number of regions. [Figure 29D] Figure 1. Investigation of the association of subclonal SCNA with immune heterogeneity within a multi-sample pan-cancer cohort. The bottom panel shows the number of tumors in the pan-cancer multi-sample cohort with subclonal gain (above the dark red-0 horizontal line) or loss (below dark blue-0) across the genome. The horizontal line represents regions with more than 30 tumors with subclonal gain or loss (see Methods). The top panel shows the -log10(p-value) of 160 cytoband regions tested for association with TCRA T-cell fraction and subclonal gain (dark red dots) or loss (dark blue dots). The red horizontal line indicates the significance threshold, with only one region being significant, the loss event at chromosome 12q24.31-32. [Figure 29E] Figure 1. Investigation of the association between subclonal SCNA and immune heterogeneity within a multi-sample pan-cancer cohort. Changes in TCRA T cell fractions between regions with and without subclonal loss events at chromosome 12q24.31-32. [Figure 29F]Figure showing the investigation of the association between subclonal SCNA and immune heterogeneity within a multi-sample pan-cancer cohort, presenting a volcano plot from Limma-Voom analysis of TRACERx100 tumors with subclonal loss at 12q24.31-32, performed on regions with loss and regions without loss. After adjustment for multiple hypotheses, only 8 genes were significant, 3 of which (SPPL3, ABCB9, and OGFOD2) were in the loss region and are labeled. [Diagram 30] Figure showing the implementation of the method of Figure 5 for WGS data. Two alternatives are shown: a GAM model using a bin-based approach for GC correction and quality control, and a segment model using the known positions of V and J segments. The segment model has the following constraints: 1. Model the log read depth ratio as a linear model consisting of fixed segments aligned to the positions of V(D)J segments. 2. The model has a value of 0 before the first V segment. 3. V segments are monotonically decreasing in the order they appear in the genome. For example, TRVA2 < TRVA1. 4. J segments are monotonically increasing in the order they appear in the genome. 5. The model has a value of 0 after the last segment. [Figure 31-1] Figure showing an exemplary output of the GAM model of Figure 30. [Figure 31-2] Figure showing an exemplary output of the GAM model of Figure 30. [Figure 32A] Figure showing the results of benchmarking the WGS T cell ExTRECT model using the TRACERx100 cohort, presenting a comparison of the TCRA T cell fraction scores of WGS and WES. [Figure 32B] Figure showing the results of benchmarking the WGS T cell ExTRECT model using the TRACERx100 cohort, presenting the WGS T cell fraction score compared to the RNAseq Danaher score for T cells. [Figure 32C]FIG. 1 shows the results of benchmarking the WGS T cell ExTRECT model using the TRACERx100 cohort, showing the WGS B cell fraction score compared to the RNAseq Danaher score for B cells using the entire TRACERx100 cohort. [Fig. 32D] FIG. 1 shows the results of benchmarking the WGS T cell ExTRECT model using the TRACERx100 cohort, showing the WGS B cell fraction score compared to the RNAseq Danaher score for B cells using a cohort that includes samples from which samples with low segment model log-likelihood samples were removed. [Figure 32E] FIG. 1 shows the results of benchmarking the WGS T cell ExTRECT model with the TRACERx100 cohort, showing a comparison of T cell fraction scores from the WGS GAM model. [Fig. 32F] FIG. 1 shows the results of benchmarking the WGS T cell ExTRECT model with the TRACERx100 cohort, showing a comparison of T cell fraction scores from the WGS segment models. [Fig. 32G] FIG. 1 shows the results of benchmarking the WGS T cell ExTRECT model using the TRACERx100 cohort, comparing TRDV1 segment fraction from the segment model with TPM (transcripts per million) RNA scores from Kallisto for the TCRD gene segment. [Diagram 33] Figure 31 shows the results of an analysis of the robustness of the method of Figure 30 with reduced sequencing depth. T cell ExTRECT scores were obtained on downsampled bam using either the GAM model or the segment model. [Figure 34A] FIG. 31 shows the use of the segment model in FIG. 30 to investigate TCR and BCR diversity. An exemplary V-segment usage plot for patient CRUK0085 (for TCRA) is shown. [Figure 34B]FIG. 31 shows the use of the segment model in FIG. 30 to investigate TCR and BCR diversity. An exemplary V-segment usage plot for patient CRUK0085 (with respect to TCRB) is shown. [Figure 34C] FIG. 31 shows the use of the segment model in FIG. 30 to investigate TCR and BCR diversity. An exemplary V-segment usage plot for patient CRUK0085 (with respect to TCRG) is shown. [Fig. 34D] FIG. 31 shows the use of the segment model in FIG. 30 to investigate TCR and BCR diversity. An exemplary V-segment usage plot for patient CRUK0045 (with respect to IGH) is shown. [Figure 34E] FIG. 31 shows the use of the segment model in FIG. 30 to explore TCR and BCR diversity. A heat map of the proportion of TRACERx100 samples utilizing different TCR V-segments is shown. [Fig. 34F] FIG. 31 shows the use of the segment model in FIG. 30 to investigate TCR and BCR diversity, showing a comparison of the number of V-segments called with MiXCR from matched RNAseq data and those called from the T cell ExTRECT segment model. [Figure 34G] FIG. 31 shows the use of the segment model in FIG. 30 to investigate TCR and BCR diversity, showing a comparison of the number of V-segments called with MiXCR from matched RNAseq data and those called from the T cell ExTRECT segment model. [Fig. 34H] FIG. 31 shows the use of the segment model in FIG. 30 to investigate TCR and BCR diversity, showing a comparison of the number of V-segments called with MiXCR from matched RNAseq data and those called from the T cell ExTRECT segment model. [Fig. 34I]FIG. 31 shows the use of the segment model in FIG. 30 to investigate TCR and BCR diversity, showing a comparison of the number of V-segments called with MiXCR from matched RNAseq data and those called from the T cell ExTRECT segment model. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0037] Detailed Description In describing the present invention, the following terms will be employed and are intended to be defined as set forth below.

[0038] "And / or," as used herein, shall be construed as a specific disclosure of each of the two specified features or components, without regard to the presence or absence of the other. For example, "A and / or B" shall be construed as a specific disclosure of (i) A, (ii) B, and (iii) each of A and B, as if each were individually set forth herein.

[0039] A "sample", as used herein, may be a cell or tissue sample (e.g., by biopsy), a biological fluid, an extract (e.g., a DNA extract obtained from a subject) from which genomic material can be obtained for genomic analysis, such as genomic sequencing (whole genome sequencing, whole exome sequencing, targeted sequencing (also called "panel" sequencing)) or copy number array profiling. A sample may be a cell, tissue, or biological fluid sample obtained from a subject (e.g., by biopsy). Such a sample may be referred to as a "subject sample". In particular, a sample may be a blood sample or a tumor sample, or a sample derived therefrom. A sample may be freshly obtained from a subject or may be processed and / or stored (e.g., frozen, fixed, or subjected to one or more purification, enrichment, or extraction steps) prior to genomic analysis. In particular, a sample may be a cell or tissue culture sample. Thus, a sample as described herein may refer to any type of sample that includes cells or genomic material derived therefrom, and may be from a biological sample obtained from a subject, or from a sample obtained, for example, from a cell line. The sample is preferably from a jawed vertebrate (such as a jawed vertebrate cell sample or a sample from a jawed vertebrate subject), suitably a mammal (such as a mammalian cell sample or a sample from a mammalian subject, including model animals such as mice and rats, among others), preferably a human (such as a human cell sample or a sample from a human subject). In some embodiments, the sample is a sample obtained from a subject, such as a human subject. Furthermore, the sample may be transported and / or stored, collection may occur at a location remote from the site of genomic sequence data acquisition (e.g., sequencing), and / or the computer-implemented method steps may occur at a location remote from the site of sample collection and / or the site of genomic sequence data acquisition (e.g., sequencing) (e.g., the computer-implemented method steps may be performed by a networked computer, such as a "cloud" provider).

[0040] A "mixed sample" refers to a sample that is assumed to contain multiple cell types, or genetic material derived from multiple cell types. A sample obtained from a subject, such as a tumor sample, is typically a mixed sample (unless it has been subjected to one or more purification and / or separation steps). Preferably, a mixed sample is a sample that contains lymphocytes, or is assumed (expected) to contain lymphocytes. Suitably, the sample contains lymphocytes and at least one other cell type. For example, the sample may be a tumor sample. A "tumor sample" refers to a sample that is derived from or obtained from a tumor. Such a sample may contain tumor cells, immune cells (such as lymphocytes), and other normal (non-tumor) cells. In the context of a tumor sample, the term "purity" refers to the proportion of cells in a sample that are tumor cells (sometimes referred to as the "cancer cell fraction" or "tumor fraction"), or the equivalent proportion of cells in the case of a sample that contains genetic material derived from cells. In the context of a sample containing genetic material, the tumor fraction can be estimated using sequence analysis processes that attempt to deconvolute the tumor genome and the germline genome, such as, for example, ASCAT (Van Loo et al., 2010), ABSOLUTE (Carter et al., 2012), or ichorCNA (Adalsteinsson et al., 2017). In the context of a tumor sample, lymphocytes in the sample may be referred to as "tumor infiltrating lymphocytes" (TILs). The tumor sample may be a primary tumor sample, a tumor-associated lymph node sample, or a sample from a metastatic site from a subject. The sample containing tumor cells or genetic material derived from tumor cells may be a bodily fluid sample. Thus, the genetic material derived from tumor cells may be circulating tumor DNA or tumor DNA within exosomes. Alternatively or additionally, the sample may contain circulating tumor cells. The mixed sample may be a sample of cells, tissues, or bodily fluids that have been processed to extract genetic material. Methods for extracting genetic material from biological samples are known in the art. The lymphocytes may be T lymphocytes, B lymphocytes, or a mixture thereof. In particular, the T lymphocytes may include αβT lymphocytes and / or γδ lymphocytes. The B lymphocytes may include B cells with IGL light chains and / or B cells with IGK light chains.

[0041] The term "lymphocyte fraction" refers to the percentage of DNA-containing cells in a mixed sample that are lymphocytes or are a specific type of lymphocyte (e.g., T lymphocytes, B lymphocytes, αβT lymphocytes, T lymphocytes or B lymphocytes containing any selected gene segment in the TCR or BCR gene, etc.). The term "T lymphocyte fraction" refers to the percentage of DNA-containing cells in a mixed sample that are T lymphocytes. The term "B lymphocyte fraction" refers to the percentage of DNA-containing cells in a mixed sample that are B lymphocytes. In the context of the present disclosure, lymphocyte fraction is estimated based on signals derived from genomic regions that undergo VDJ recombination in the specific lymphocyte population being quantified. Thus, lymphocyte fraction may more specifically refer to the percentage of DNA-containing cells in a mixed sample that are lymphocytes characterized by having a given genomic region of interest that has undergone VDJ recombination. The predetermined genomic region of interest may be a region of the TCRA / TCRD, TCRB, TCRG, IGH, IGL, or IGK gene (also referred to herein as the TCRA / TCRD, TCRB, TCRG, IGH, IGL, and IGK loci). Depending on the species under investigation, other less common loci may be used, such as the loci of certain TCRs present in marsupials and sharks (e.g., TCRμ). Depending on the selection of the predetermined region of interest (genomic locus) used in the method, the lymphocyte fraction may capture different subpopulations of lymphocytes. For example, the use of the TCRB locus may allow quantification of the αβ T lymphocyte fraction. As another example, the use of the TCRA locus may allow quantification of the combined αβ T lymphocyte and γδ T lymphocyte fraction, since the TCRD locus is believed to be completely contained within the TCRA locus. Furthermore, the use of one or more segments corresponding to TCRD in the TCRA locus may allow for quantification of the fraction of αβ and γδ T lymphocytes, which are gamma delta T lymphocytes. The selection of a particular segment corresponding to TCRD may be made depending on the correlation between the fraction estimated using the segment and an orthogonal criterion indicating the fraction of TCRD. As another example, the use of the TCRG or TCRD locus may allow for quantification of the fraction of gamma delta T lymphocytes.The use of the TCRG locus may be preferred for quantifying the γδ T lymphocyte fraction, since the TCRD locus is very small and therefore not expected to contain many informative signals. Conversely, the use of both the TCRA locus and the TCRB locus or the TCRG locus may allow for separate quantification of the αβ T lymphocyte fraction and the γδ T lymphocyte fraction. As yet another example, the use of the IGH locus may allow for the quantification of B cells (regardless of whether they have an IGL or IGK light chain), while the use of the IGL or IGK locus may allow for the quantification of B cells with either an IGL or IGK light chain. Furthermore, separate lymphocyte fractions can be quantified for multiple subpopulations of lymphocytes. These can be added together to obtain lymphocyte fractions representing multiple or all subpopulations of lymphocytes. Furthermore, some mixed samples may be assumed to contain primarily or exclusively a single lymphocyte population. Therefore, it can be assumed that the lymphocyte fraction estimated using a single predefined genomic region of interest represents the (overall) lymphocyte fraction for the sample. For example, it can be assumed that the sample contains mainly T lymphocytes (or conversely, it can be assumed that it contains a relatively small proportion of B lymphocytes compared to the proportion of T lymphocytes), the majority of which can be expected to be αβ T lymphocytes. In such cases, the lymphocyte fraction is therefore estimated using a single predefined genomic region of interest corresponding to the TCRA or TCRB locus. Similarly, some mixed samples may be assumed to contain mainly or exclusively a single T lymphocyte population. Therefore, it can be assumed that the lymphocyte fraction estimated using a single predefined genomic region of interest represents the (overall) T lymphocyte fraction for the sample. For example, it can be assumed that αβ T lymphocytes are the main type of T lymphocyte present in the sample, and thus the T lymphocyte fraction can be estimated using a single predefined genomic region of interest corresponding to the TCRA or TCRB locus.As another example, the TCRD locus is thought to be completely contained within the TCRA locus (together forming a locus called TCRA / TCRD), such that the lymphocyte fraction estimated using this single, predefined genomic region of interest can be assumed to represent the (total) T lymphocyte fraction for the sample (including specifically αβ and γδ T lymphocytes).

[0042] The TCRA gene refers to the T cell receptor alpha locus with gene ID 6955 (HGNC symbol: TRA) in Homo sapiens, or any homologous region in other jawed vertebrates. In Homo sapiens, this gene is located on chromosome 14 at position NC_000014.9 (21621904..22552132) (assembly GRCh38.p13), which is also referred to as position NC_000014.8 (22090057..23021075) (assembly GRCh37.p13). Examples of homologous loci include the mouse tcra locus (Gene ID: 21473, chromosome 14, assembly GRCm39-NC_000080.7 (52665424..54461655), assembly GRCm38.p6-NC_000080.6 (52427967..54224198)). Example coordinates that can be used as the region of maximum V(D)J recombination and as the region used to obtain baseline read depth for the TCRA gene (and any genes therein, such as the TCRD gene) are provided in Table 7.

[0043] The TCRB gene refers to the T cell receptor beta locus with gene ID 6957 (HGNC symbol: TRB) in Homo sapiens, or any homologous region in other jawed vertebrates, in which the gene is located on chromosome 7 at position NC_000007.14 (142299011..142813287) (assembly GRCh38.p13) or NC_000007.13 (141998851..142510972) (assembly GRCh37.p13). An example of a homologous locus is the mouse tcrb locus (Gene ID: 21577, chromosome 6, assembly GRCm39-NC_000072.7 (40868230..41535305) or assembly GRCm38.p6-NC_000072.6 (40891296..41558371)). Example coordinates that can be used as the region of maximum V(D)J recombination and as the region used to obtain baseline read depth for the TCRB gene are provided in Table 7.

[0044] The TCRG gene refers to the T cell receptor gamma locus with gene ID 6965 (HGNC symbol: TRG) in Homo sapiens, or any homologous region in other jawed vertebrates, in which the gene is located on chromosome 7 at position NC_000007.14 (38240024..38368055, complement) (GRCh38.p13) or NC_000007.13 (38279625..38407656, complement) (GRCh37.p13). Examples of homologous loci include the mouse tcrg locus (Gene ID: 110067, chromosome 13, assembly GRCm39-NC_000079.7 (19362212..19540646), assembly GRCm38.p6-NC_000079.6 (19178042..19356476)). Example coordinates that can be used as regions of maximum V(D)J recombination and as regions used to obtain baseline read depth for the TCRG gene are provided in Table 7.

[0045] The TCRD gene refers to the T cell receptor delta locus with gene ID 6964 (HGNC symbol: TRD) in Homo sapiens, or any homologous region in other jawed vertebrates, in which the gene is located on chromosome 14 at position NC_000014.9 (22422546..22466577) (GRCh38.p13) or NC_000014.8 (22891537..22935569) (GRCh37.p13). Examples of homologous loci include the mouse tcrd locus (Gene ID: 110066, chromosome 14, assembly GRCm39-NC_000080.7 (54183530..54396655), assembly GRCm38.p6-NC_000080.6 (53946073..54159198)). Examples of coordinates that can be used as the region of maximum V(D)J recombination and as the region used to obtain baseline read depth for the TCRD gene are provided in Table 7 (with reference to the region used in TCRA) and Table 8 (with reference to the TCRD segment within the TCRA locus).

[0046] The IGH gene refers to the immunoglobulin heavy locus with gene ID 3492 (HGNC symbol: IGH) in Homo sapiens, or any homologous region in other jawed vertebrates, in which the gene is located on chromosome 14 at position NC_000014.9 (105586437..106879844, complement) (GRCh38.p13) or NC_000014.8 (106032614..107288051, complement) (GRCh37.p13). Examples of homologous loci include the mouse Igh locus (Gene ID: 11507, chromosome 12, assembly GRCm39-NC_000078.7 (113222388..115973574, complement), assembly GRCm38.p6-NC_000078.6 (113258768..116009954, complement). Example coordinates that can be used as the region of maximum V(D)J recombination and as the region used to obtain baseline read depth for the IGH gene are provided in Table 7.

[0047] The IGL gene refers to the immunoglobulin lambda locus with gene ID 3535 (HGNC symbol: IGL) in Homo sapiens, or any homologous region in other jawed vertebrates, where the gene is located on chromosome 22 at position NC_000022.11 (22026076..22922913) (GRCh38.p13) or NC_000022.10 (22380474..23265085) (GRCh37.p13). Examples of homologous loci include the mouse Igl locus (Gene ID: 111519, chromosome 16, assembly GRCm39-NC_000082.7 (18845608..19079594, complement), assembly GRCm38.p6-NC_000082.6 (19026858..19260844, complement)).

[0048] The IGK gene refers to the immunoglobulin kappa locus with gene ID 50802 (HGNC symbol: IGK) in Homo sapiens, or any homologous region in other jawed vertebrates, in which the gene is located on chromosome 2 at position NC_000002.12 (88857361..90235368) (GRCh38.p13) or NC_000002.11 (89890568..90274235), (89156874..89630436, complement) (GRCh37.p13). An example of a homologous locus is the mouse Igk locus (Gene ID: 243469, chromosome 6, assembly GRCm39-NC_000072.7 (67532620..70703738), assembly GRCm38.p6-NC_000072.6 (67555636..70726754)).

[0049] The term "sequence data" refers to information indicating the presence and preferably also the amount of genomic material in a sample having a particular sequence. Such information can be obtained using sequencing technologies, such as next generation sequencing (NGS, such as whole exome sequencing (WES), whole genome sequencing (WGS), or sequencing of captured genomic loci (targeted or panel sequencing)), or using array technologies such as copy number variation arrays or other molecular counting assays. When NGS technologies are used, the sequence data may include a count of the number of sequencing reads having a particular sequence. The sequence data may be mapped to a reference sequence, such as a reference genome, using methods known in the art, such as Bowtie (Langmead et al., 2009). Thus, the number of sequencing reads or equivalent non-digital signals may be associated with a particular genomic location ("genomic location" refers to the location in the reference genome to which the sequence data is mapped). The term "read depth" refers to a signal indicating the amount of genomic material in a sample that maps to a particular genomic location. Such signals can be obtained using sequencing techniques, for example using next generation sequencing (NGS, e.g. WES, WGS, or sequencing of captured genomic loci), or using array techniques, for example copy number variation arrays. When NGS techniques are used, the read depth can be a read depth within the general meaning of the term, i.e. a count of the number of sequencing reads that map to a genomic location. When array techniques are used, the read depth can be an intensity value associated with a particular genomic location, which can be compared to a control to provide an indication of the amount of genomic material that maps to a particular location. The term "read depth profile" refers to a collection of read depth values ​​for multiple genomic locations. For example, the read depth at a particular genomic location i can represent the read depth at the base at location i in a reference genome, and the read depth profile can represent the read depth for multiple locations i in one or more regions of interest.

[0050] As used herein, "treatment" refers to the reduction, alleviation, or elimination of one or more symptoms of the disease being treated compared to the symptoms prior to treatment. "Prevention" (or prophylaxis) refers to the delay or prevention of the onset of symptoms of the disease. Prevention may be absolute (the disease does not occur) or may be effective only in some individuals or for a limited period of time.

[0051] As used herein, the term "computer system" includes hardware, software, and data storage devices for implementing the system or for executing the method according to the above-mentioned embodiments. For example, a computer system may comprise a processing unit (such as a central processing unit (CPU) and / or a graphic processing unit (GPU)), input means, output means, and data storage, which may be embodied as one or more connected computing devices. Preferably, the computer system has a display or comprises a computing device with a display to provide a visual output display (e.g. in the design of a business process). The data storage may include RAM, disk drives, or other computer-readable media. A computer system may include multiple computing devices connected by a network and capable of communicating with each other via the network. It is expressly envisioned that a computer system may also consist of or comprise a cloud computer.

[0052] As used herein, the term "computer-readable medium" includes, but is not limited to, any non-transitory medium that can be read and accessed directly by a computer or computer system. Media can include, but are not limited to, magnetic storage media, such as floppy disks, hard disk storage media, and magnetic tape; optical storage media, such as optical disks and CD-ROMs; electrical storage media, such as memory, including RAM, ROM, and flash memory; and hybrids and combinations of the above media, such as magnetic / optical storage media.

[0053] Determination of lymphocyte fraction in mixed samples The present disclosure provides a method for determining lymphocyte fraction in a mixed sample using read depth data from a sample including read depth data for at least a predetermined genomic region of interest. An exemplary method is described with reference to FIG. 1. In optional step 10, a mixed sample including genomic material from multiple cell types can be obtained from a subject. In optional step 12, sequence read / read depth data can be obtained from the mixed sample, for example, by sequencing the genomic material in the sample using one of whole exome sequencing, whole genome sequencing, or panel sequencing. In step 14, a read depth profile for the sample including read depth along the predetermined genomic region of interest is obtained. This may include step 14A of obtaining a plurality of raw read depth values ​​across the predetermined region of interest, and step 14B of smoothing and / or correcting (e.g., GC correction) the read depth values.

[0054] In step 16, a baseline read depth is obtained from the subset of regions of interest. This may include an optional step 16A of identifying a subset of regions of interest that can be used as the subset of regions of interest that are unlikely to be deleted by VDJ recombination to obtain the baseline read depth. For example, the regions of interest that are unlikely to be deleted by VDJ recombination include the first n V segments of a genomic locus that undergoes VDJ recombination. The parameter n can be determined by using suitable training data to select a value of n that results in the most accurate estimate of lymphocyte fraction across the training data when compared to, for example, an orthogonal criterion of lymphocyte fraction. The orthogonal criterion of lymphocyte fraction can be obtained, for example, based on transcriptome data and / or histopathology data.

[0055] In step 18, a plurality of read depth ratios (r i ) is obtained. This is the result of (multiple read depth ratios (r i ) to obtain a smoothed read depth profile. In step 20, summarized read depth ratio values ​​(r VDJ ) is obtained (optionally based on the value of the model fitted in step 18A in a subset of regions of interest that are likely to be deleted by VDJ recombination).

[0056] In step 22, the fraction of abnormal cells (p) in the mixed sample (where abnormal cells are cells that are aneuploid within the given region of interest) and the copy number of abnormal cells (Ψ T ) can be optionally determined by obtaining copy number profiles specific to abnormal cells using methods known in the art. In step 24, the fraction of lymphocytes in the sample (f) is calculated based on the summarized read depth ratio values ​​(rVDJ ) is determined as a function of the summarized read depth ratio value (r VDJ ), the fraction of lymphocytes in the sample as a function of the α-threshold (f), the fraction of abnormal cells in the mixed sample (p), and the copy number of abnormal cells in a given region of interest (Ψ T ).

[0057] At optional step 26, the determined lymphocyte differential and / or any values ​​derived therefrom may be provided to a user, for example via a user interface. Values ​​derived therefrom may include prognostic and / or diagnostic information, as described further below.

[0058] Purpose The above method has applications in various clinical settings. For example, in the context of cancer, determining the lymphocyte fraction in a tumor sample, or preferably multiple samples from a tumor, can provide an indication of the immune status of the tumor. This has been shown to have prognostic value for various cancers, including non-small cell lung cancer (demonstrated herein), ovarian cancer (see, e.g., Zhang et al., 2003), colorectal cancer (see, e.g., Galon et al., 2006), breast cancer (see, e.g., Dieci et al., 2018), and melanoma. The inventors have demonstrated that T and / or B cell fractions in blood and / or tumor samples are indicative of prognosis in at least some cancer types (see Example 6, Figure 35).

[0059] Thus, also described herein is a method of determining the immune status of a tumor from a subject, the method comprising determining the lymphocyte differential in one or more tumor samples from the subject.

[0060] Also described herein is a method of providing a prognosis for a subject diagnosed with cancer, comprising determining a lymphocyte fraction in one or more tumor samples from the subject. The method may further comprise classifying the sample as immune hot or immune cold depending on the lymphocyte fraction in the sample (such as, for example, T lymphocyte fraction). For example, a sample having a lymphocyte fraction above a threshold (e.g., about 0.1, i.e., about 10%) may be classified as immune hot, and a sample having a lymphocyte fraction below the threshold may be classified as cold. The method may further comprise classifying the subject into a good prognosis group and a bad prognosis group depending on the number of samples from the patient classified as immune cold. For example, when analyzing samples from different tumor regions, a number of samples classified as immune cold equal to or greater than a threshold (e.g., 2) may classify the subject into a bad prognosis group. The bad prognosis group may be associated with a decreased recurrence-free survival rate and / or decreased overall survival rate compared to the good prognosis group. Threshold values ​​for lymphocyte differentials and / or number of immune cold samples can be determined using an appropriate training cohort, for example by assessing the performance of various threshold values ​​that can classify patients between groups associated with significantly different prognoses.

[0061] Furthermore, continuing in the context of cancer, the presence of TILs has been shown to be associated with an increased likelihood of response to immunotherapy, including checkpoint inhibitors (CPIs) in particular (see, e.g., Litchfield et al., 2020), but also T cell transfer therapy (tumor infiltrating lymphocyte (or TIL) therapy and CAR-T cell therapy), therapeutic antibodies, cancer treatment vaccines, and immune system regulators (e.g., interferons and interleukins). Thus, also described herein is a method for determining whether a subject diagnosed with cancer is likely to benefit from treatment with an immunotherapy, comprising determining a lymphocyte fraction in one or more tumor samples from the subject. The immunotherapy may be any therapy that modulates the function of the immune system to treat cancer. Indeed, any such therapy is likely to be influenced by the presence or absence of immune cells in the tumor. In some embodiments, the immunotherapy is a CPI therapy. The CPI therapy includes, for example, treatment with an anti-CTL4 drug or an anti-PDL1 drug. The method may further include classifying the subject into a group that is likely to respond to immunotherapy and a group that is unlikely to respond to immunotherapy. For example, the method may further include classifying the sample as immune hot or immune cold according to lymphocyte fraction (such as T lymphocyte fraction) in the sample. Then, if one or more samples from the subject are immune cold, the subject can be classified into a group that is unlikely to respond to immunotherapy (such as CPI therapy). Alternatively, if the sample from the subject has a lymphocyte fraction below a threshold, the subject may be classified into a group that is unlikely to respond to immunotherapy (such as CPI therapy), otherwise, the subject may be classified into a group that is likely to respond to immunotherapy (such as CPI therapy). Alternatively, the method may further include classifying the subject into a group that is likely to respond to CPI therapy and a group that is unlikely to respond to CPI therapy based on the lymphocyte fraction in one or more samples from the subject and the estimated tumor mutation load for the subject. Such a classification may be obtained using a multivariate model. The parameters of the classification model (e.g., thresholds, parameters of the multivariate model) may be determined using an appropriate training cohort.

[0062] The lymphocyte differentials described herein can also be used to determine whether a subject has severe combined immunodeficiency (SCID) and / or T-cell lymphopenia. This is particularly used in the context of neonatal care. Indeed, prior to the present method, SCID was usually identified using a complete blood count. This could, in theory, be confirmed using sequencing and quantification of the TRECs themselves, but this is rarely done in actual clinical practice. Thus, the ability to use data from WES to make a diagnosis regarding SCID, which is increasingly being performed as part of genomic screening of newborns to identify potentially harmful germline mutations, is highly valuable. Thus, the present disclosure also relates to a method of identifying T-cell lymphopenia and / or SCID in a subject, particularly a neonatal subject, comprising determining the lymphocyte differentials described herein in one or more samples from the subject. The sample may, for example, be a blood sample.

[0063] The method can also be used in any clinical situation where the presence or absence or amount of immune cells or a particular type of immune cell in a sample is diagnostic or prognostic. For example, some autoimmune diseases are associated with reduced lymphocyte counts (e.g., lupus, rheumatoid arthritis, etc.), as are some bone marrow diseases (e.g., aplastic anemia, etc.) and infectious diseases (e.g., HIV, sepsis, influenza, malaria, viral hepatitis, tuberculosis, typhoid, etc.). Thus, also described herein is a method for diagnosing a subject as having a disease, disorder, or condition associated with abnormal lymphocyte counts, comprising determining lymphocyte differentials in one or more samples from the subject.

[0064] The lymphocyte fraction described herein can also be used to determine, for example, by GWAS analysis, whether one or more germline mutations, such as single nucleotide polymorphisms (SNPs), are associated with an increased or decreased T / B cell fraction in, for example, a blood sample. This can help identify variants (germline mutations) that may be responsible for the increased / decreased T cell fraction.

[0065] system FIG. 2 illustrates an embodiment of a system for determining lymphocyte fraction in a mixed sample and / or providing a prognosis or treatment recommendation based at least in part on the lymphocyte fraction according to the present disclosure. The system comprises a computing device 1 comprising a processor 101 and a computer readable memory 102. In the illustrated embodiment, the computing device 1 also comprises a user interface 103, which is shown as a screen but may include any other means of conveying information to a user, for example by an auditory or visual signal. The computing device 1 is communicatively connected, for example via a network 6, to a read depth data acquisition means 3, such as a sequencing machine, and / or one or more databases 2 that store read depth data. The one or more databases may further store other types of information that can be used by the computing device 1, such as, for example, reference sequences and parameters. The computing device may be a smartphone, tablet, personal computer, or other computing device. The computing device is configured to perform a method for determining lymphocyte fraction in a mixed sample as described herein. In an alternative embodiment, the computing device 1 is configured to communicate with a remote computing device (not shown), which itself is configured to execute the method for determining lymphocyte differential in a mixed sample as described herein. In such a case, the remote computing device may be configured to transmit the results of the method for determining lymphocyte differential to the computing device. The communication between the computing device 1 and the remote computing device may be via a wired or wireless connection, and may be via a local or public network, for example via the public Internet or WiFi. The lead depth data acquisition means may be wired to the computing device 1, or may be capable of communicating via a wireless connection, for example WiFi, as shown.The connection between the computing device 1 and the read depth data acquisition means 3 may be direct or indirect (e.g., via a remote computer). The read depth data acquisition means 3 is configured to acquire read depth data from a nucleic acid sample, e.g., a genomic DNA sample extracted from a cell and / or tissue sample. In some embodiments, the sample may be subjected to one or more pre-processing steps, such as DNA purification, fragmentation, library preparation, target sequence capture (e.g., exon capture and / or panel sequence capture). Preferably, the sample has not been subjected to amplification, or if it has been subjected to amplification, the amplification has been performed in the presence of amplification bias control means, e.g., the use of a unique molecular identifier. Any sample preparation process suitable for use in determining a genome copy number profile (whether whole genome or sequence specific) may be used within the context of the present disclosure. The read depth data acquisition means is preferably a next generation sequencer. The sequence data acquisition means 3 may be directly or indirectly connected to one or more databases 2 in which (raw or partially processed) sequence data may be stored.

[0066] The following are offered by way of example and should not be construed as limiting the scope of the claims. EXAMPLES

[0067] Working Example The immune microenvironment influences tumor progression and can be a prognostic and response predictor for immunotherapy. However, measurement of tumor infiltrating lymphocytes (TILs) is limited by a lack of appropriate data. Whole exome sequencing (WES) of DNA is often performed to calculate tumor mutation burden and identify actionable mutations. In these examples, we develop a method to estimate T cell fraction from WES samples using a signal from T cell receptor excision circle (TREC) loss during VDJ recombination of the T cell receptor alpha (TCRA) gene. The method is applicable to various types of clinically important samples, including cancer samples, for example to measure TILs. In these examples, we show that this score significantly correlates with the orthogonal TIL estimator and can be calculated from fresh frozen or formalin-fixed paraffin-embedded samples. Blood TCRA T cell fraction correlated with immune infiltration in tumors, the presence of bacterial sequencing reads, and was higher in women. Tumor TCRA T-cell fraction is a prognostic predictor in lung adenocarcinoma, and using a meta-analysis of tumors treated with immunotherapy, we found that the score is a predictor of response to immunotherapy, providing value beyond mutational burden. Application of the score to a multi-sample pan-cancer cohort revealed extensive diversity of immune infiltration within tumors. Subclonal loss of 12q24.31-32, including SPPL3, was associated with reduced TCRA T-cell fraction. The method described herein, T-cell ExTRECT (T-cell exome TREC tool), uncovers T-cell infiltration in WES samples.

[0068] Example 1 - Estimation of T cell fraction in a sample from WES data introduction Determining the level of TILs in tumors is crucial both for understanding the tumor immune microenvironment and for predicting the response to immunotherapy. However, at present, there is no universal advanced method to estimate TILs. Microarray signatures such as CIBERSORT (Newman et al., 2015), and CIBERSORTx for RNA-seq (Newman et al., 2019), ESTIMATE (Yoshihara et al., 2013), and RNA-seq signatures such as those compiled by Danaher et al. have been used to quantify the transcriptomic signatures of TILs in tumor samples. Alternatively, the presence of immune infiltration can be determined based on hematoxylin and eosin (H&E) histopathological slides or, alternatively, by appropriate cell-specific staining, most commonly immunohistochemical characterization. If more phenotypic knowledge of T cells is required, sequencing of the T cell receptor (TCR) can be performed after enrichment for sequences encoding the highly diverse VDJ recombined CDR3 regions of the TCR ( Bolotin et al., 2012 ).

[0069] A significant drawback of most existing T cell quantification methods is that they require additional materials, time, and expertise in addition to WES. Thus, although these methods provide additional knowledge that can help predict immune responses, the application of these methods in clinical practice requires increased both time and expenditure, adding to the already high costs of immunotherapy.

[0070] In this example, we propose a method to estimate T cell fraction within a sample directly from WES data, utilizing somatic copy number-based signals from VDJ recombination and excision of T cell receptor excision circles (TRECs).

[0071] The method can be computed from any type of NGS platform, including both whole genome sequencing and targeted panel-based methods, and has the potential to be extended to all genes that undergo VDJ recombination.

[0072] result T cell diversity is the product of VDJ recombination, in which various gene segments within the T cell receptor gene are recombined. This results in a large number of unselected gene segments being excised from the TCRA gene as TRECs. This process results in copy number differences between T cells and other cells, and T cells essentially undergo deletion events within the TCRA gene.

[0073] Standard somatic copy number alteration estimation tools (e.g., ASCAT (Van Loo et al., 2010), Sequenza (Favero et al., 2015), FACETS (Shen & Seshan, 2015), or ABSOLUTE (Carter et al., 2012)) in principle rely on two related signals to obtain a somatic copy number alteration profile of cancer: the B allele frequency, which reflects the relative frequency of heterozygous SNPs in the tumor sample; and the read depth ratio, which reflects the logarithm of the ratio of reads between the tumor sample and its matched germline (often the buffy coat from a blood sample). Any deviation in the read depth ratio is assumed to reflect copy number alterations specifically in the tumor (as shown in Figure 3A). However, in the context of TCRA genes, this assumption does not hold. Rather than reflecting only somatic copy number changes in the tumor, deviations in read depth ratios overlapping the TCRA gene also reflect detection of deletion events in the TCRA gene of T cells present in the tumor sample. Whether this is a loss or gain event depends on the difference in the proportion of T cells present in the blood compared to the tumor sample (as shown in Figure 3B). Specifically, if the T cell content in the sequenced tumor sample is low compared to the blood sample, there will be more reads from the TCRA gene identified in the tumor sample, resulting in an amplified signal (log read depth ratio > 0). Conversely, if the T cell content in the tumor is high compared to the blood, there will be fewer TCRA reads and the somatic copy number change estimation tool will infer a deletion or loss event (read depth ratio < 0).

[0074] Examining the ASCAT-derived somatic copy number alteration profiles of the TRACERx100 cohort of non-small cell lung cancer (NSCLC) tumors undergoing multi-region sequencing (Jamal-Hanjani et al., 2017), we noted that somatic copy number alteration segments with both breakpoints within the TCRA gene occurred in 165 / 327 of the tumor regions undergoing WES (as shown in Figure 1B). Examining two such cases in the TRACERx100 cohort, we found that the read depth ratio peaked at the genomic location of the D genomic segment, which is the location most frequently contained within TRECs (as shown in Figure 4). Thus, these data suggest that signals in TRECs can be identified from WES data.

[0075] To explicitly utilize this signal to calculate the T cell fraction in individual samples, we developed TIL ExTRECT (T cell Exome T-cell Receptor Excision Circle Tool). Briefly, TIL ExTRECT relies on the analysis of the read depth ratio in the TCRA gene to directly measure T cell infiltration in WES samples. First, we calculate the read depth at each position in the TCRA gene using genomic segments at the beginning and end of TCRA that are not affected by VDJ recombination as controls, and calculate a corrected log read depth ratio between the original read depth of the coverage value and the median of the values ​​in the control genomic segments. Finally, given that both the fraction of tumor cells in the sequenced sample (tumor purity) and the local somatic copy number around TCRA are known, an accurate estimate of the T cell fraction can be calculated based on the magnitude of deviation of the log (base 2) read depth ratio. In cases where knowledge of tumor purity and local somatic copy number is not available, a naive estimate can also be made that we demonstrate to have similar predictive value (described in detail in the Methods section, shown in Figure 5). Notably, unlike the RNA-seq score, the TCRA score represents a direct measurement of the proportion of T cells, hereafter referred to as the "TCRA T cell fraction."

[0076] A detailed description of TIL ExTRECT and optimization of parameters such as selection of segments for normalization and quality control of selected exons are provided in Methods below. The accuracy of the TCRA T cell fraction was verified using four orthogonal approaches described in Example 2 below. Its clinical utility was then demonstrated using two separate approaches in Examples 3 and 4. Finally, its use in assessing key determinants of T cell immune infiltration in various sample types was explored in Example 5.

[0077] method Defining the local copy number around the TCRA locus in different cells within a tumor sample In T cells and other stromal cells known to be diploid, there are two copies of the TCRA gene, and one can say that the copy number around the TCRA locus is equal to 2. Assuming that in T cells there is a region within the TCRA locus that is constantly lost during VDJ recombination, this would result in a copy number of 0 in this small genomic segment (as shown in FIG. 3B). In tumor cells, there can be both copy number gains and losses across chromosome 14 (where the TCRA gene is located), resulting in different copy number states of the TCRA gene that can be different from 2. This method assumes that there are no breakpoints within the TCRA gene that result from somatic alterations in cancer cells (however, in tumor cells the region as a whole may not be diploid).

[0078] As described in more detail in the following two subsections, the average copy number at the beginning of the TCRA gene (chr14:22090057(hg19)) was inferred from all cancer cells in the sample using ASCAT and is referred to as the local tumor copy number at the TCRA gene. This is used as a term in the SCNA-recognizing TCRA T cell fraction provided by Equation (7) below. Additionally, when tumor somatic copy number alteration information is not available, the SCNA-naive TCRA T cell fraction (Equation (3)) can be used, which is referred to as the naive TCRA T cell fraction.

[0079] Calculation of TCRA T cell fraction from WES data The ASCAT approach (Allele-Specific Copy Number Analysis of Tumors) described in Van Loo et al. (2010) provides a method for estimating allele-specific copy numbers from DNA extracted from samples containing a mixture of tumor and non-abnormal cells. The following formula is the formula used by ASCAT (Van Loo et al., 2010) to estimate tumor ploidy at genomic position i, combined with a second formula for the B allele frequency at that position (see Van Loo et al.):

[0080]

number

[0081] where r i is the logR at position i (a measure of the total signal intensity that quantifies the total copy number at a genomic locus), γ is a constant that depends on the technology used (e.g., for Illumina Hiseq a value of γ=1 can be used), p is the tumor purity (fraction of tumor cells among the cell population in the sample), and n Ai and n Bi is the allele-specific copy number, the factor 2 represents the diploid copy number assumed for normal tissue, and Ψ S is the average ploidy of the tumor samples. When using ASCAT with sequencing data, r i is the logarithm (base 2) of the read depth ratio at genomic position i, where the ratio is the ratio between the coverage for the tumor and germline samples at the same genomic position.

[0082] In this study, the aim was to determine the fraction of T cells among the cell population in a sample. Furthermore, this sample may contain abnormal cells (e.g., tumor cells) that may have unknown ploidy at genomic locus i. By analogy with the mixed population investigated in Van Loo et al., iThe value represents the ratio of coverage of samples containing T cells among other (non-T) cells to samples without T cells. For example, tumor biopsy samples and matched blood samples are often available for tumor patients. In practice, it is often possible to use non-tumor samples to obtain r i In contrast to the tumor situation, where it is possible to quantify the logR, there are no known samples from patients with a clear lack of T cells. Indeed, both tumor and matched blood samples are likely to contain T cells. We therefore devised a new formula and method to estimate logR from a single sample (as described in detail in the following section). Briefly, this is calculated by using the median coverage of the genomic regions at the beginning and end of the TCRA gene (chr14:22090057-22298223 in hg19, chr14:23016447-23221076 in hg19) as the normal background rate, which is assumed to have copy number values ​​unaffected by VDJ recombination. We then calculate the logR over the entire TCRA gene (chr14:22090057-23221076) by dividing the coverage over the TCRA genomic region by this median. T We make an estimator for . Therefore, this case can be expressed as:

[0083]

number

[0084] where f is the T cell fraction of the sample and n T is the T cell copy number at the TCRA locus (or other VDJ locus under investigation, depending on the type of cell to be quantified), γ is a constant that depends on the technology used (e.g., for Illumina Hiseq a value of γ=1 may be used), and r i is logR calculated from a single sample as described below, and Ψ Si is the copy number for the T cell-excluded sample at the TCRA locus (or any other VDJ locus under investigation), and Ψ Sis the copy number at this locus for the entire sample without taking into account changes due to VDJ recombination. VDJ ), n T ≈ 0. Therefore, at this position, equation (2) simplifies to equation (2b).

[0085]

number

[0086] At this point, two different situations can exist: (a) for normal samples, for example where the average ploidy is always 2 (e.g. in the case of blood samples), or for samples where the average ploidy may not be 2 but where a small fraction of T cells are present (e.g. cancer samples with low T cell content), Ψ S ≒Ψ Si (b) the above assumption is not true, for example in tumor samples where the T cell content is not small. Case (a) helps simplify Equation 2 to Equation (3).

[0087]

number

[0088] This is called the formula for the "naive TCRA T cell fraction." If the assumptions in (a) do not apply (case (b)), Ψ S and Ψ Si The following formula for can be used: Psi Si =2(1-p')+p'Ψ T (4) Psi S =2(1-p)+pΨ T (5)

[0089]

number

[0090] Here, ΨT is the local tumor copy number at the TCRA locus (calculated as the average copy number at the start of the TCRA gene (chr14:22090057(hg19)) estimated from all cancer cells in the sample using ASCAT), and p' is the adjusted tumor purity that takes into account the absence of T cell DNA at the position of maximum VDJ recombination. T is calculated as the average copy number from all cancer cells (i.e., across all subclones, each of which may have a different copy number at that locus) at the start of the TCRA locus (or any other VDJ locus under investigation). This is because the beginning of this gene is expected to show no signal from VDJ recombination. Therefore, the end of the gene can also be used instead of the beginning in any embodiment without affecting the results, since the end of the gene is also expected to show no signal from VDJ recombination. Without wishing to be bound by theory, it is believed that the use of signals at the beginning of the gene is beneficial because signals associated with VDJ recombination are closer to the end of the gene than to the beginning of the gene (because J segments are clustered near the end of the gene, whereas V segments are more widely spread). Furthermore, in any embodiment, it is also possible to use the average copy number from all cancer cells in the region outside the VDJ locus under investigation, e.g., before or after the locus (if such data is available, e.g., when whole genome sequencing is used instead of whole exome sequencing). However, as the distance from the locus of interest increases, it becomes more likely that the signal seen corresponds to different copy number events within the cancer cell and therefore does not reflect the tumor copy number at the VDJ locus under investigation.

[0091] Substituting equations (4), (5), and (6) into equation (2), we obtain equation (7).

[0092]

number

[0093] Equation (7) represents a "more complete" version of equation (3) (i.e., a solution of equation (2) without all the assumptions made to obtain equation (3)). It is therefore possible to recover equation (3) by starting from equation (7) and applying the same assumptions. In particular, equation (7) immediately simplifies to equation (3) if these are not cancer cells (not cells with a ploidy different from normal 2; p=0). Similarly, if there exists a tumor cell with the same peripheral local copy number of TCRA as in normal cells (Ψ T = 2), this formula also simplifies to Equation (3). Equation (3) can therefore be used to calculate the TCRA T cell fraction in samples from normal tissue, or can be used as a naive estimate when tumor purity or copy number status is unknown.

[0094] In other words, the non-cancerous components of the sample (1-p) can be rewritten by dividing them into T cell and non-T cell subsets. (1-p)=(1-p)f'+(1-p)(1-f') Here, f′ represents the fraction of non-tumor cells that are T cells and is related to the total T cell fraction of the sample that are T cells (f) according to the following formula:

[0095]

number

[0096] Ignoring allele-specific copy numbers, let Ψ be the total copy number in the tumor at genomic locus i. T and the copy number of the T cell compartment is n T and the copy number in the non-T cell normal cell compartment is 2 at all genomic loci, we obtain the following formula:

[0097]

number

[0098] Here, we consider the genomic location of a VDJ recombination event at the TCRA gene in T cells. T →0, and further at this position, r VDJ Assume that is estimated directly from the read depth ratio. S is calculated as the average copy number of the sample without taking into account VDJ recombination, i.e., Ψ S =2(1-p)+pΨ T Therefore, at the genomic locus of VDJ recombination, the following formula holds:

[0099]

number

[0100] Substituting the equation for f' into this equation, and transforming it into the equation for the fraction of T cells in the sample, is shown to be provided by equation (7) above. VDJ Note that if the value of is >0, the resulting fraction will be negative. In such cases, it can be assumed that no TILs are present in the sample and only sample noise is being measured. Thus, f can be set to 0. Conversely, if the tumor purity p is known (and / or it can be assumed that the calculation of tumor purity is error-free), then a value of f higher than 1-p is unlikely (as the tumor purity+T cell fraction of a sample cannot be greater than 1), and f can be truncated to this value if a higher value is obtained.

[0101] Calculation of log read depth ratio from a single sample in WES To calculate the T cell fraction from equation (3) (for a naive estimate of T cell content) or equation (7) (when tumor copy number and purity are known), the log read depth ratio r i needs to be estimated from the raw coverage data (i.e., from a single sample, not from matched samples that each contain or do not contain the cell population to be quantified).

[0102] Briefly, this logR is calculated by using the median coverage of the genomic regions at the beginning and end of the TCRA gene (or any other VDJ locus under investigation, depending on the type of cells to be quantified) as the normal background rate, which is assumed to have a copy number value that is not affected by VDJ recombination. As explained above, in any embodiment, it is also possible to use the beginning or end (but not both) of the VDJ locus under investigation, or other (preferably nearby) regions that can be assumed to have a copy number value that is not affected by VDJ recombination. Without wishing to be bound by theory, it is believed that the use of a longer normalization region (e.g., using both the beginning and the end regions and / or increasing the length of the beginning region) advantageously compensates for the noise that is typical in read depth data. For example, the beginning region can be extended to include a large number of V segments that are unlikely to be frequently lost, and also to include regions outside the gene when data are available (e.g., when using whole genome sequencing data). In this case, the following regions were used: chr14:22090057-22298223 and chr14:23016447-23221076 (hg19). The coverage across the entire TCRA gene (in this case, chr14:22090057-23221076 (hg19)) was then calculated by dividing the coverage across the TCRA genomic region (or any other VDJ locus under investigation) by this median. i We obtain an estimate for . The idea behind this is illustrated in Figure 5.

[0103] More specifically, from the aligned BAM files to either hg19 (based on GRCh37) or hg38 (GRCh28) (depending on the dataset, e.g. TCGA data was obtained pre-aligned to hg38, while TRACERx and other data were aligned in house to hg19), coverage at the individual base level is extracted using the depth function (with parameters, -q20-Q20) of samtools (version 1.3.1). After this is done, only bases within known exons defined by SureSelect Human All-Exon Probes (version 5) are used in TRACERx, and then within each exon, reads are normalized with a rolling median in windows of 50 base pairs (i.e. the median within each 50 bp window is taken as the new value). Note that the size of the windows is not fixed and in some embodiments windows larger or smaller than 50 bp can be used, e.g. between 20-200 bp or between 50-150 bp. By comparing the estimates of lymphocyte differential fraction obtained using the present method to various candidate window sizes for orthogonal measures of lymphocyte differential fraction, such as the Danaher score, the appropriate size of the window can be empirically determined for a particular data set, type of data, or situation.

[0104] The median baseline coverage value is calculated within the region of TCRA assumed to be unaffected by VDJ recombination using the regions at the beginning and end of the gene (chr14:22090057-22298223 and chr14:23016447-23221076 (hg19), corresponding to the first 12 V and C segments of TCRA). All coverage values ​​are then divided by this value and the logarithm is then taken to calculate the "single sample" log ratio. r VDJ To complete the calculation, we use the average value of the model at the position of maximum VDJ recombination, r VDJA generalized additive model is fitted to the data, taking x, bs, cs, y, bs, s, and bs as values ​​for x, bs ... The position of maximum VDJ recombination is defined as the gap between the last TCR V segment and the first segment of the TCR delta gene, rounded to have a size of 80000 bp (80 kbp), starting at position 22800000 (resulting in the region chr14:22800000-22880000). This is explained below. Other definitions are possible and it is expected that small changes in the size of this region will not make a big difference. The region of maximum VDJ recombination can be defined using a similar approach for other VDJ loci, i.e. by defining a region located at least partially around or within the region between the last V gene segment and the first J gene segment, for example a region corresponding to the gap between the last V gene segment and the first J gene segment, optionally rounded to the nearest kbp or nearest 10 kbp, etc. for ease of manipulation.

[0105] In summary, for many tumor samples, equation (7) can be used, which requires knowledge of both the tumor purity and the tumor copy number at the TCRA gene. In the special case where the copy number at TCRA is exactly 2, as is the case in all blood-derived samples and many tumors, equation (3) can be used to calculate the T cell fraction, and is accurate. Furthermore, when the tumor purity and tumor copy number at TCRA are unknown, equation (3) can serve as a naive estimate for the T cell fraction (however, this naive estimate may not be accurate depending on whether the above-mentioned assumptions hold). This is discussed in the next section.

[0106] Naïve TCRA T cell fraction and accurate estimate of T cell fraction FIG. 6A shows the (theoretical) difference between the estimated naive value for T cells and various actual T cell fraction values ​​for various local copy number values ​​and tumor purity, based on theoretical values ​​derived from equations (2), (3), and (7). For low T cell fractions and local copy number values ​​close to 2, the estimate is very accurate. FIG. 6B-C shows, using real data from the TRACERx100 cohort (more information on this data is provided in Example 2), that the distribution of local TCRA copy numbers has mode 2, and tumor purity values ​​are typically low. FIG. 6D shows the correlation between the naive estimate and the calculated accurate T cell fraction for cases where the local TCRA copy number is not 2.

[0107] Log ratio (r VDJ Optimizing the segments used to estimate r VDJ Calculation of requires the selection of (i) the segment of maximal VDJ recombination, and (ii) the segment that should be used to calculate a "normal baseline" against which coverage across the TCRA genomic region is compared (by calculating the ratio as explained above).

[0108] For the focal segment representing the location of maximum VDJ recombination, the gap between the last V gene segment and the first segment encoding part of the TCR delta chain was used (hg19, chr14:22800000-22880000). This region of the genome is theoretically likely to be within the T cell receptor excision circle (TREC) and overlaps with a region previously used for PCR primers measuring TREC (Kuss et al., 2005). As explained above, any region between the last V segment and the first J segment of the VDJ locus of interest can be used.

[0109] In calculating the normal baseline using V gene segments that are less likely to be commonly lost in VDJ recombination, it is desirable to make the region as large as possible due to sequencing noise (see Calculating log ratios from a single sample in WES). However, the more V gene segments are selected, the more likely it is that T cell clonotypes reside in locations where these gene segments have been excised by TREC (and therefore coverage in that region will be affected by VDJ recombination).

[0110] For the local segments, the first n V gene segments and the last C segment were selected using an optimization scheme outlined below. The first n V segments were taken in order and the score (TCRA T cell fraction, f, non-GC corrected) was calculated across the entire TRACERx100 cohort (see Example 2 for further details on this cohort). The number of V segments was selected based on which n had the most non-zero values ​​of f and maximized the association between the calculated T cell fraction and the Danaher Transcriptome Score for T cell infiltration (a comparative measure of T cell infiltration, see Example 2). It was found that the length of the segments used reduced the level of noise (i.e. the more segments used the better the normalization in terms of noise), but in some cases if they were too long they would overlap with segments that were missing from a particular T cell clonotype (i.e. longer segments are more likely to include one or more segments that were lost in one or more T cell clonotypes). Therefore, as a conservative compromise between high correlation with Danaher score and as short a segment as possible, the first 12 V-segments were chosen for use (shown in FIG. 25, where the top plot shows the TCRA T-cell fraction as a function of the number of V-segments used for normalization in the TRACERx100 data, the middle plot shows the corresponding fractions for non-zero samples, and the bottom plot shows the significance of the correlation with the Danaher T-cell signature assessed using Spearman's method). Note that other numbers of segments could also be used, such as the first 11 V-segments, without significantly affecting the results.

[0111] GC normalization method used within TIL ExTRECT GC content is known to bias sequence coverage values. Therefore, TIL ExTRECT can normalize coverage values ​​by GC content before calculating T cell fraction. To do this, GC content is calculated by subtracting the GC content (GC exon) is calculated at the local exon level and at a larger macro-scale (GC macro ) at two scales: at the larger scale, GC content is calculated in 1000 bp windows across the entire TCRA gene. A generalized additive model is then fitted to this data to create a smoothed value for macro-scale GC content across the TCRA gene. To normalize the GC content, the coverage values ​​of every base pair within all TCRA exons are put into the following linear model:

[0112]

number

[0113] The residuals of this model are then taken as the new GC-normalized coverage values.Unless otherwise stated, all TCRA T cell fractions used herein are GC-corrected.

[0114] After GC correction, the baseline of the score is rescaled by taking the median at the beginning and end of the TCRA gene (chr14:22090057-22298223 and chr14:23016447-23221076 (hg19)) and the median of the generalized additive model (GAM) at the position of maximum VDJ recombination (chr14:22800000-22880000). Since in theory the T cell fraction cannot be negative, which necessarily requires that the log read depth ratio at chr14:22800000-22880000 must be ≦0, the maximum of the two medians is taken as the new baseline and set to 0.

[0115] Calculating confidence intervals within TIL ExTRECT 95% confidence intervals for the TCRA T-cell fraction were calculated taking into account two components: 1) the noise in the final baseline adjustment after GC correction; 2) the uncertainty in the fit of the GAM model used in the final calculation of the TCRA T-cell fraction.

[0116] To account for noise in the baseline adjustment, 95% confidence intervals on the adjusted values ​​were calculated as 1.96 times the standard deviation of the coverage values ​​in the region used for normalization. For the GAM model fits, simultaneous confidence intervals were calculated (using the confint function) using the R package "gratia" (v0.5.1). Finally, these two sources of uncertainty were combined to generate a 95% confidence interval.

[0117] Overestimated tumor TCRA copy number biases TCRA T cell fraction Overestimated values ​​of TCRA tumor copy number lead to inflated values ​​for TCRA T-cell fraction. This may result from a poor quality solution selected from copy number segmentation algorithms such as ASCAT. One indication that this is the case is whether the calculated TCRA T-cell fraction from TIL ExTRECT exceeds a value of 1-tumor purity. In these cases, the given copy number solution is considered unreliable and a naive estimate for T-cell fraction that instead assumes the TCRA tumor copy number is 2 is used as an alternative solution.

[0118] Example 2 - Validation of WES-derived TCRA T cell fraction introduction To assess the accuracy of the TCRA T cell fraction criteria described in Example 1, we used four orthogonal approaches.

[0119] First, we used WES data from cell lines as positive and negative controls for T cell content. Second, we used simulated next-generation sequencing (NGS) data that included various T cell fraction values. Simulated data was also used to investigate the impact of local tumor TCRA copy number and tumor purity on the accuracy of the score. Third, we examined how well the score compares to orthogonal immune-related data in multiple NSCLC cohorts. Fourth, we calculated how well the score is similar to alternative DNA-based methods for inferring immune content.

[0120] result As a first validation of the approach described in Example 1, we utilized WES data from cell lines consisting of 14 independent samples derived from the HCT116 cell line, including tetraploid and diploid clones with different degrees of genomic complexity (Lopez et al. 2020), and three samples (JURKAT, PEER, and HPB-ALL) derived from cell lines derived from T cell lymphomas from the cancer cell line encyclopaedia (Gandi et al. 2019). In theory, NGS from the HCT116 cell line should have a TCRA T cell fraction of zero, while cell lines derived from T cell lymphomas (which should all have undergone VDJ recombination) should have a TCRA T cell fraction of 1 (reflecting 100% of cells as T cells). As can be seen, all HCT116 cell lines had a calculated fraction of 0, regardless of their diploid or tetraploid status (as shown in Figure 7). Conversely, three T cell-derived cell lines had scores calculated in the naive estimation that were close to 1 (~0.95 to ~0.96) (see Figure 7). This small difference from an estimated fraction of exactly 100% is likely explained by technical factors (e.g., misaligned reads) or biological factors.

[0121] As a second validation approach, to further evaluate the accuracy of the TCRA T cell fraction over binary (presence / absence), simulated NGS data containing various T cell fraction values ​​were acquired for various T cell fractions (see Methods). As can be seen in Figure 8A, there was a highly significant, nearly perfect relationship between the simulated and calculated T cell fractions in samples with a background TCRA copy number of 2 (ρ=1, p<2.2e-16). Further investigation of the impact of local tumor TCRA copy number (see Methods) and tumor purity on the accuracy of the score was investigated using simulated data, which also demonstrated a significant relationship between the simulated and estimated TCRA T cell fractions (as shown in Figure 8C-D).

[0122] As a third validation approach, we examined how well the scores compared to orthogonal immune-related data within the TRACERx100 cohort and within the lung adenocarcinoma (LUAD) TCGA (Cancer Genome Atlas Research Network, 2014) and lung squamous cell carcinoma (LUSC) (Cancer Genome Atlas Research Network, 2012) cohorts. These NSCLC cohorts contained estimators for T cell content from a variety of different data types, including RNA-seq, histopathology slides, and methylation data.

[0123] The TRACERx100 cohort consists of 327 tumor regions that underwent WES, derived from 100 tumors from 100 patients, each with a matched germline blood sample that also underwent WES (see Table 1). Additionally, 189 of these tumor regions had matched RNA-seq data from which a transcriptome-based immune score could be calculated.

[0124] [Table 1]

[0125] In particular, we examined how well the scores compared to orthogonal non-WES immune-related data within the TRACERx100 cohort. Immune-related signature scores for various cell types from Danaher et al., Davoli et al., xCell, TIMER, EPIC, and CIBERSORT were calculated for 189 TRACERx100 regions using RNA-seq data (Rosenthal et al., 2019). TCRA T cell fraction from WES was found to have a significant positive relationship with multiple immune scores derived from various methods. The top three items with the strongest associations were all T cell related (Danaher Th1: ρ = 0.68, P = 9.0e-23, xCell CD8 T cell memory ρ = 0.67, P = 2.3e-22, and Danaher T cell: ρ = 0.67, P = 2.94e-22). The Davoli score for NK cells (ρ = 0.67, P = 3.92e-22) and other non-T cell scores such as activated dendritic cells identified by xCell (ρ = 0.61, P = 1.26e-17) were also ranked very highly, indicating the complexity of the tumor immune microenvironment and that T cell infiltration is strongly correlated with infiltration from other immune cells. TCRA T cell fraction from WES was found to have a significant positive relationship with Danaher transcriptome signatures corresponding to CD8+ cells (ρ = 0.62, P = 9.6e-20), T cells (ρ = 0.63, P = 2.3e-20), and total TILs (ρ = 0.59, P = 1.4e-17) (shown in Figure 9A, 9C).

[0126] To further validate the TCRA T cell fraction, we evaluated the association between the TCRA T cell fraction and the TIL score based on the available histopathological H&E slides for 88 patients, which were manually inspected and scored by a pathologist. As can be seen, the TCRA T cell fraction significantly correlated with the pathology TIL fraction estimate. Notably, the pathology TIL fraction estimate does not exclusively include CD8+ and CD4+ T cells, and therefore a perfect correlation between the estimate and the TCRA T cell fraction is not expected.

[0127] As a fourth approach, we calculated how similar the score was to an alternative DNA-based method {Levy et al} that infers immune content based on the number of reads at each WES aligning to the CDR3 region after VDJ recombination versus total coverage (CDR3 VDJ score, see Methods for details). In the TRACERx100 cohort, we observed a significant positive correlation between TCRA T cell fraction and CDR3 VDJ score (Figure 9D, ρ=0.36, P=1.4e-13). However, we noted that despite the high coverage in the TRACEERx100 cohort, the CDR3 VDJ score was significantly constrained by sequencing depth. The number of reads aligning to the CDR3 region was typically very low (Q1=0, median=2, mean=2.335, Q3=3, max=14). To explicitly quantify how robust the T cell ExTRECT-derived TCRA T cell fractions and CDR3 VDJ scores are to differences in sequencing coverage, we employed two complementary methods (described below).

[0128] Next, selecting a subset of tumor regions (147 regions) that contained both RNA-seq data and pathology-derived TIL scores, we were also able to assess how well TCRA T cell fraction, CDR3 VDJ score, and six RNA-seq-based immune measures of CD8+ cells (Danaher, Davoli, xCell, TIMER, CIBERSORT, and EPIC) compared with TIL scores (shown in Figure 9B). The Danaher CD8+ score had the strongest association with pathology TILs (ρ=0.49), followed by our TCRA T cell fraction (ρ=0.41), Davoli (ρ=0.4), xCell (ρ=0.36), CIBERSORT (ρ=0.23), TIMER (ρ=0.2), CDR3 VDJ score (ρ=0.2), and EPIC (ρ=0.082), with all but EPIC being significantly associated. Thus, TCRA T cell fraction estimates represent a significant improvement over alternative DNA-based measures and are comparable to advanced RNA-seq-based immune signatures in revealing immune cell tumor infiltrates, but unlike many existing methods, they provide accurate point estimates of T cell fraction.

[0129] Additionally, we further evaluated TIL ExTRECT in an independent NSCLC dataset with lower coverage from TCGA (median coverage: LUAD=84, LUSC=88.2 compared to median depth 426X in TRACERx). Samples comprising the LUAD and LUSC cohorts are listed in Table 2. We correlated TCRA T cell fraction with immune-related data from Thorsson et al. (Thorsson et al., 2018). In Thorsson et al., a measure of total TILs in a sample, called "leukocyte fraction", was calculated from DNA methylation and RNA-seq CIBERSORT-based T cell CD8 fraction.

[0130] [Table 2]

[0131] Compared to the CIBERSORT-based CD8+ T cell estimator, the TCRA T cell fraction had a stronger correlation with the methylation-derived leukocyte fraction used by Thorrson et al. (ρ=0.29, P=1.2e-18 vs. ρ=0.1, P=0.00097, as shown in Figure 10), indicating both that the TCRA T cell fraction correlates well with a methylation-based measure of total TILs and that this is potentially superior to specific RNA measures of T cell content.

[0132] Additionally, we sought to test how robust the TCRA T cell fraction from TIL ExTRECT was to differences in sequencing coverage. Specifically, to ascertain the minimum sequencing depth required to obtain accurate T cell fraction estimates, we employed two complementary methods. First, we selected five NSCLC tumor regions from TRACERx100. The TCRA T cell fraction from TIL ExTRECT ranged from 0 to 0.35. These regions were then individually downsampled 10 times to 5, 10, 20, 30, 40, 50, 75, 100, and 200-fold coverage. We found that the T cell fraction estimates remained consistent at coverages of 30-fold or more (ρ=0.84, P=1.4e-14) (shown in Figure 11A). Consistent with this, simulated data showed a reliable correspondence between 30x coverage and known T cell fractions (shown in Figure 11B, ρ=0.96, P=1.9e-14). We observed that while useful information may be obtainable even with 10x coverage (as shown in Figure 11A), TCRA T cell fractions begin to lose fidelity at coverages below 20x. This is likely due to increased noise, especially for lower T cell fractions that begin to fall below the detection limit. Thus, coverages below 30x (including 10x and 20x) may still be useful in distinguishing between high and low T cell fractions, but may not be sufficient to obtain reliable estimates of very low T cell fractions. As T cell fractions below 0.1 are common, the use of coverages above 30x may be advantageous in providing additional certainty and broader applicability. In contrast, we found that the CDR3 method for determining T cell infiltration was heavily skewed by low sequencing coverage: when we selected the five samples with the highest CDR3 reads and downsampled them to 50x coverage, one sample had three reads detected, while the remaining sample had only one CDR3 read (Figure 11D).Consistent with this, it is noteworthy that the TCGA BRCA analysis by Levy et al. also identified no CDR3 reads in 56% of tumors.

[0133] Finally, we identified no systematically significant differences in T cell fraction depending on whether the sample was fresh frozen or formalin-fixed paraffin embedded (FFPE) across multiple different datasets (see Methods and Figure 11C), suggesting that the method is not sensitive to DNA alterations inherent to FFPE.

[0134] Thus, the TCRA T cell fraction estimation method described herein improves upon alternative DNA-based measures and rivals advanced RNA-seq-based immune signatures in revealing immune cell tumor infiltration, but unlike many existing methods, provides an accurate point estimate of T cell fraction. Furthermore, T cell ExTRECT can be applied to any sample that has undergone WES, thus enabling analysis of T cell fraction in both tumor and blood samples.

[0135] method statistics All statistical tests in this example and all other statistical tests were performed in R3.6.1. No statistical methods were used to predetermine sample sizes. Tests involving correlations were performed using “stat_cor” from the R package ggpubr (v0.4.0) with Spearman’s method. Tests involving comparison of distributions were performed using “stat_compare_means” with the unpaired option or with “wilcox.test” unless otherwise stated. Effect sizes for the corresponding Wilcoxon tests were measured using the “wilcox_effsize” function from the rstatix ​​package (v0.6.0). Hazard ratios and p-values ​​were calculated using the “survival” package (v3.2-3) for both Kaplan-Meier curves and Cox proportional hazards models. For all statistical tests, the number of included data points is plotted or annotated in the corresponding figures.

[0136] Defining immune hot and cold tumor regions Throughout these examples, immune hot regions were defined as regions containing more than 10% T cells as measured by TIL ExTRECT, and immune cold regions were defined as regions containing less than 10% T cells. This threshold was defined based on expert knowledge, and other thresholds could be used depending on the context and purpose of the analysis.

[0137] Fresh frozen vs. FFPE samples To test that TCRA T cell fractions were reliable and consistent for both fresh-frozen and FFPE samples, we calculated non-GC-corrected TCRA T cell fractions for six different studies in the CPI responder cohort. Three of these studies utilized WES from FFPE tissues (n=460), while the other three utilized WES samples from fresh-frozen tissues (n=357). We fitted linear models to predict TCRA T cell fractions by histology and FFPE status (Table 6), revealing that cancer type was the main driver of this significance, while FFPE status was not. Furthermore, no significant differences were found for melanoma and bladder tumors with FFPE and fresh-frozen WES samples (Figure 11C). Therefore, it can be concluded that whether the WES samples were derived from fresh-frozen or FFPE tissues does not significantly affect the TCRA T cell fraction values ​​calculated by TIL ExTRECT.

[0138] [Table 3]

[0139] Simulated Data Simulated data for testing the non-GC-corrected TCRA T-cell score outcomes were generated using ART Illumina (version 2.5.8) (Huang, Li, Myers, & Marth, 2012) in combination with tools from the MASCoTE (Zaccaria & Raphael, 2020) tumor simulation method. Briefly, ART was used to generate simulated paired-end FASTQ files for human chromosome 14 from a HiSeq 2500 DNA sequencer with coverage 30 and read length 150 (parameters -p-ss HS25-f 30-na-1 150-m 200-s 10). where -ss represents the sequencing platform, -f represents coverage, -l represents read length, -m represents the average size of DNA fragments, -s represents the standard deviation of fragment sizes, and -na indicates that no output alignment file can be provided (see https: / / manpages.debian.org / stretch / art-nextgen-simulation-tools / art_illumina.1 for details). These FASTQ files were then aligned to the human genome hg19 using bwa mem (v0.7.15). The resulting BAM files were then cleaned, sorted, and indexed using Picard tools (version 1.107). As only coverage at the TCRA gene is important for testing the simulated data, a BAM file was created using samtools (version 1.3.1) view of all reads mapping to chr14:20000000-24000000. This BAM file served as a template for normal and tumor cells. A T cell template was created from this BAM file by extracting only reads that mapped to chr14:20000000-22500000 or chr14:23000000-24000000 and merging these results into a single BAM file using samtools merge.Thus, this T cell BAM file contains an artificially created 100% deletion between positions chr14:22500000-23000000. Reads that partially mapped to either of the regions on either side of this gap were included (i.e., were not deleted). Since the read size is 150 bp, the presence or absence of these reads is considered to be insignificant compared to the size of the deletion. This T cell and normal / cancer BAM file was used to generate various mixed BAM files using the MixBam.py module used in the MASCoTE method. This creates mixed BAMs by sampling normal / cancer and T cell BAMs according to both the genomic length and copy number of the mixture. In this way, simulated data could be created that contained different proportions of normal, T cell, and cancer cell populations with different cancer somatic copy number change values.

[0140] Simulated data using the method was generated for all possible combinations of tumor purity = [0.25, 0.5, 0.75], tumor copy number = [1, 3, 4, 5], and T cell fraction = [0.01-0.25, in 0.01 increments], and for tumor copy number 2 with T cell fractions ranging from 0.01-0.99 in 0.01 increments. After generating this simulated data, the T cell fraction was calculated using only reads from positions within the exons of the TCRA gene, as is done in the standard TIL ExTRECT method (described in Example 1) applied to WES. An example of output for a single simulated run with 24% T cells and 75% tumor, with a local copy number of 1 around TCRA, is shown in Figure 8B. Figures 8C and 8D show the difference between simulated and calculated values ​​at various local tumor copy number and purity values. The naive score overestimates the T cell fraction when the copy number is less than 2 and underestimates it when the copy number is greater than 2. However, the exact score follows the line y=x, showing that it is indeed accurate.

[0141] Investigating the impact of sequencing depth on TCRA T cell fraction estimation To evaluate the impact of lower coverage, five TRACERx100 regions were selected based on the range of TCRA T cell fractions they represent (high = 0.35, medium = 0.156, low = 0.053, very low = 0.010, none = 0). For simplicity, all regions were assumed to have a local TCRA tumor copy number of 2. In the TRACERx100 as study, the median coverage was 430.91x. Although the samples had various depths, the aligned BAM files were downsampled to a specific depth using the samtools view with the -s option, taking into account the original depth of the region. In this way, 10 downsampled depths created from different random seeds were created, forming BAM files with depths of coverage of 200, 100, 75, 50, 40, 30, 20, 10, and 5x. After downsampling to generate these BAMs, non-GC-corrected TCRA T cell fractions were estimated using the methods described above.

[0142] Exon bias detection for capture kits and quality control of exons used in TCRA T cell fraction calculations As a general quality control step, all exons were checked for sufficient coverage before being used in the calculation of TCRA T cell fraction: exons with a median coverage below 15x were removed, and if more than 30 exons were below this threshold, the sample was marked as failing due to low coverage.

[0143] In the course of analyzing the TCGA dataset, bias was observed due to the capture kit used (an extensive custom set based on Agilent SureSelect v2). This bias was previously described by Wang et al. (Wang, Kim, & Chuang, 2018) and affects 4833 genes, including TCRA. This bias caused coverage of certain exons to be much lower than expected, thus interfering with the score calculation. An example of this bias can be seen in Figure 26A. To address this issue, we calculated the average logR ratios in individual exons (i.e., the average log ratio of read coverage in an exon to the median read coverage in the first 12 V and C segments, as explained above) within the TRACERx100 and TCGA LUAD cohorts (see Figure 26B). The TCGA LUSC cohort was not used for this analysis, but we do not expect the results from this cohort to be significantly different from those of the TCGA LUAD cohort. Exons with median differences greater than 0.5 (for each exon, calculate the median across all samples in each cohort, then examine the median difference between cohorts) were flagged and removed from the TCGA analysis. This resulted in the exclusion of 99 / 192 exons. Note that in the following analysis, these exons were removed from the TRACERx100 data (however, this data is not expected to be subject to the same bias as the TCGA data, as the bias is likely related to the capture kit used), to demonstrate that the difference in using this reduced set in the absence of bias is minimal (compared to the large effects that can be observed when bias is present). Other cohorts, particularly within the CPI dataset (Rizvi et al., 2015; Shim et al., 2020), using the Agilent SureSelect v2 capture kit, were also observed to have the same bias (i.e., bias affecting the same exons).This is shown by calculating the median values ​​for all exons as described above and correlating these with the median values ​​of exons from a dataset affected by bias (i.e., in this case, the TCGA LUAD cohort) with the score calculated using this reduced set of exons. Furthermore, after GC correction, bias was observed in a single exon when using the Agilent SureSelect v2 capture kit in one of the TCRδ gene segments. As this segment is within the focal region used to calculate the TCRA fraction, and as the number of exons in the reduced set is smaller, a small change in this exon could produce a large effect on the calculated score. After GC correction, this turned out to be much more pronounced, and therefore, in cohorts such as TCGA where this effect is large, this exon was also removed from the analysis.

[0144] We applied the reduced exon set to the TCGA cohort, in which additional TCRδ gene segments were also removed, and compared the TCRA T cell scores, observing a large difference between the TCGA and TRACERx cohorts for tumor samples from both LUAD and LUSC (Figure 26C: Wilcoxon test, LUAD: P<2.2e-16LUSC: P<2.2e-16). However, this difference is entirely due to the GC bias within this reduced exon set. When the GC correction method is used, the significant difference in the TCRA T cell fraction with LUAD patients is only marginally significant, likely due to the difference in the cohort population (Figure 26D, Wilcoxon test, LUAD: P=0.028, LUSC: P=0.97).

[0145] There were no other biases other than those correctable by GC correction detected with any other capture kit used in any cohort herein. However, the cohort using the Nimblegen capture kit was found to be sufficiently different from the cohort of exons defined from the Agilent capture kit. All samples from the Nimblegen capture kit were calculated using exons defined directly from the Nimblegen exon capture regions. In addition, some capture kits, such as Nextera Rapid Capture and IDT xGen Exome Research Panel, were identified to have extremely low coverage within the TCRA gene. Manual checking of the exome capture regions of these kits revealed that regions within the TCRA gene were not included in their design. For datasets using these kits, TIL ExTRECT cannot be used to calculate TCRA T cell fractions.

[0146] Calculation of CDR3 VDJ score CDR3 VDJ scores were calculated following the procedure outlined in Levy et al.. Initial reads aligning to TCRB (hg19:chr7:142000817-142510993) and unaligned reads were extracted with samtools, the resulting bam was converted to fastq using bedtools, and the tool IMSEQ (v1.1.0) was then used on the resulting output to identify VDJ recombined reads aligning to the CDR3 region, and the number of aligned reads was normalized by the total number of reads in the original bam file (measured with samtools flagstat) to generate a CDR3 VDJ score.

[0147] TRACERx100 patients The first 100 patients prospectively analyzed by the NSCLC TRACERx study (https: / / clinicaltrials.gov / ct2 / show / NCT01888601, approved by an independent research ethics committee, 13 / LO / 1546) were used in this study. This is the same cohort of 100 patients originally described in Jamal-Hanjani et al. (Jamal-Hanjani et al., 2017). Briefly, informed consent was a mandatory requirement for participation in the TRACERx study. This NSCLC cohort consisted of 68 male and 32 female patients with a median age of 68 years. Finally, this cohort consisted mainly of early stage tumors (Ia(26), Ib(36), IIa(13), IIb(11), IIIa(13), and IIIb(1)), with 28 patients also receiving adjuvant therapy. Both WES (aligned to hg19) and RNA-seq samples were obtained from the TRACERx study of the first 100 patients. Methods for processing these samples were as previously described (Jamal-Hanjani et al., 2017). For WES samples specifically, exome capture was performed using a custom version of the Agilent Human All Exome V5 kit, following the manufacturer's instructions.

[0148] TCGA LUAD and LUSC cohorts Aligned BAM files (hg38) from the TCGA LUAD and LUSC cohorts were downloaded from the Genome Data Commons (dataset ID: phs000178.v10.p8). Sample purity and ploidy calls were generated using ASCAT (v2.4.2) and were obtained from a previous analysis of TCGA data (Middleton et al., 2020). Briefly, Affymetrix SNP6 profiles from tumor-normal paired samples (dataset ID: phs000178.v10.p8) were processed through the PennCNV library (K. Wang et al., 2007) to obtain GC-corrected BAF and log ratios before processing with ASCAT. Immune-related data, including leukocyte differential and CD8+ differential, were obtained from Thorsson et al. 2018.

[0149] Cancer cell line data The non-T cell derived cell line HCT116 was sequenced on an Illumina HiSeq 2500 and aligned to bwa mem using hg19 as described in Lopez et al., 2020. The T cell derived cell lines were from the dataset described in Ghandi et al., 2019, downloaded from the Sequence Read Archive (SRA) under the accession number PRJNA523380. The T cell derived cell lines were selected by ensuring that cell lines derived from precursor T cell acute lymphoblastic leukemia were excluded as they have not undergone VDJ recombination. This process resulted in the selection of WES data from three cell lines: JURKAT, HPB-ALL, and PEER. Naïve TCRA T cell fractions were used for all cell line work, as it is difficult to perform ASCAT without matched germline samples.

[0150] Orthogonal immune measures: 1. RNA-seq signatures We used the method of Danaher et al. (Danaher et al., 2018) as our primary method to estimate T cell content from RNA-seq measures, as it has previously been demonstrated to correlate most strongly with TIL scores calculated with TRACERx (Rosenthal et al., 2019). Other RNA-seq signatures tested against TCRA T cell fractions were the Davoli method (Davoli, Uno, Wooten, & Elledge, 2017), xCell (Aran, Hu, & Butte, 2017), TIMER (Li et al., 2017), and EPIC (Racle, de Jonge, Baumgaertner, Speiser, & Gfeller, 2017).

[0151] Orthogonal immune scales: 2. Pathological TIL score TILs were estimated from histopathology slides using internationally established guidelines developed by the International Immuno-Oncology Biomarker Working Group (Hendry et al., 2017) as previously described in Rosenthal et al. (Rosenthal et al., 2019). Briefly, the relative percentage of stromal area to tumor area was determined from pathology slides for a given tumor area. TILs were reported for the stromal compartment (=percent stromal TILs). The denominator used to determine the percent of stromal TILs was the area of ​​stromal tissue (i.e., area occupied by mononuclear inflammatory cells relative to the total interstitial area in tumors) rather than the number of stromal cells (i.e., fraction of total stromal nuclei representing mononuclear inflammatory cell nuclei). The method has been demonstrated to be reproducible between trained pathologists (Denkert et al., 2016). Inter-individual concordance was performed, which showed high reproducibility. The International Immuno-Oncology Biomarker Working Group is developing a freely available training tool to train pathologists on optimal TIL assessment on hematoxylin-eosin slides (www.tilsincancer.org).

[0152] Example 3 - Evaluation of the prognostic value of WES-derived TCRA T cell fraction in TRACERx100 and TCGA NSCLC cohorts introduction To explore the potential clinical utility of TIL ExTRECT, we first considered whether TCRA T cell fraction is a prognostic predictor in the TRACERx100 NSCLC cohort. TIL levels inferred from histopathological H&E slides from samples within the TRACERx100 cohort have previously been associated with disease-free survival (AbdulJabbar et al., 2020). Therefore, we investigated whether a similar association could be confirmed using the novel TCRA T cell fraction described herein. We also performed a similar investigation using data from TCGA.

[0153] Finally, we used a unique approach to quantify TCRA T cell fractions in matched blood samples within these cohorts.

[0154] result When patients in the TRACERx100 NSCLC cohort were divided into two groups according to the number of immune cold areas in the patient's tumor (≥2 immune cold areas, where immune cold areas are defined as areas with a TCRA T cell fraction ≤0.1), significant differences emerged: the presence of ≥2 observed immune cold areas was found to be associated with a reduced recurrence-free survival (see Figure 12A, log-rank test P=0.0068, HR=2.3). The threshold of ≥2 immune cold areas was chosen based on expert knowledge (see, for example, AbdulJabbar et al., 2020). Dividing the cohort by histological characteristics revealed that this significance mainly came from LUAD patients, while the recurrence-free survival of LUSC patients was not significant (see Figure 12B, LUAD: P=0.0037, HR=4.1; Figure 12C, LUSC: P=ns, HR=1.12).

[0155] This result in the TRACERx100 cohort suggests that TCRA T cell fraction has a significant association with recurrence-free survival in LUAD patients. However, when we evaluated whether overall survival in the TCGA LUAD dataset (Table 2) was also related to TCRA T cell fraction, no significant association was observed when patients were divided into two categories based on the definition of immune cold regions with TCRA T cell fraction <0.1 (see Figure 13, log-rank test, P = ns). The TCGA dataset is single-region, subject to sampling bias, and cannot distinguish patients with tumors that are uniformly cold regions from those with a single cold region. The lack of significance within the TCGA cohort compared to TRACERx100 suggests the importance of multi-region immune data in predicting survival.

[0156] In addition to estimating TCRA T cell infiltration in tumor samples, TIL ExTRECT can also be applied to matched blood samples from the TRACERx100 cohort. We observed that matched blood samples had significantly higher TCRA T cell fractions than their corresponding paired primary tumor samples (see Figure 14, Wilcoxon test P = 0.0012, effect size = 0.16). However, it should be noted that this fraction reflects the fraction of cells that contain DNA, and thus this fraction of T cells in blood is much lower when red blood cells are considered. Notably, tumor TCRA T cell fractions were not consistently lower in tumor regions compared to matched blood. In fact, substantially less than half of tumor regions (139 / 339) contained higher TCRA T cell fractions in the tumor regions than matched blood. To summarize this at the patient level rather than the region level, 44 patients scored lower than their matched blood samples in all regions, 24 patients scored higher than their matched blood samples in all regions, and the remaining 32 were heterogeneous. We also observed that tumor TCRA T cell fraction was significantly associated with blood TCRA T cell fraction in both LUAD and LUSC (LUAD: ρ=0.39, P=1.2e-07, LUSC: ρ=0.45, P=7.8e-07, see Figure 15) (using TRACERx100 data; one blood sample is matched to multiple tumor regions due to the multi-regional nature of the data). This suggests that TCRA T cell fraction in blood is associated with tumors from the same patient.

[0157] We divided the TRACERx100 cohort by histological characteristics and found that the circulating TCRA T cell fraction was significantly higher in LUAD patients compared to LUSC (see FIG. 16A, left plot; Wilcoxon test P=0.0053, effect size=0.29), and primary tumor samples also had a higher TCRA T cell fraction identified in LUAD compared to LUSC tumors (see FIG. 16A, right plot; Wilcoxon test, P=2.2e-12, effect size=0.42). Similar results were observed in TCGA tumors (see FIG. 17A-D), and we also found significant differences according to sample type, with a higher TCRA T cell fraction found in blood than in primary tumors (see FIG. 17A and B; Wilcoxon test, LUAD: P<2.2e-16, effect size=0.28, LUSC: P<2.2e-16, effect size=0.38). These results suggest that tumors that stimulate strong immune responses may result in higher levels of T cells circulating in the blood, or potentially, elevated levels of T cells in the blood may result in a stronger immune response within the tumor.

[0158] To further explore factors that may explain the changes in TCRA T cell fraction in blood samples, we considered the clinical characteristics of patients, including sex and age. Indeed, there is considerable sexual dimorphism in the human immune system during aging (Marquez et al., 2020), and we noted that a decrease in CD8+ T cell content was associated with both aging and males. For both the TRACERx100 and TCGA cohorts, there was a significant difference between male and female patients in terms of TCRA T cell fraction in blood samples, with increased levels in women (TRACERx100: Figure 18A, Wilcoxon test P = 0.027, effect size = 0.22; TCGA: Figure 18B, Wilcoxon test, P = 1.2e-5, effect size = 0.12). The effect of age was relatively weak. The TRACERx100 cohort showed a small, non-significant negative trend (see FIG. 18C: ρ=-0.091, P=0.37), while the TCGA cohort showed a weak but significant association (see FIG. 18D: ρ=-0.095, P=0.013).

[0159] In further analysis, we analyzed LUAD and LUSC separately, dividing patients into either immune hot or cold groups based on the number of cold regions present in the tumor. Here, we used a simple threshold based on the average TCRA T cell fraction in all tumor regions (0.081) and limited the analysis to tumors that exceeded the number of cold regions used at each threshold. The results in Figure 16C show that for LUAD, as the number of cold regions used in the threshold increases, greater significance is observed in survival between immune hot and cold groups (LUAD: ≧2 cold regions, HR=3.1, P=0.0063 log-rank test; LUAD: ≧3 cold regions, HR=7.3, P=0.00024 log-rank test). In contrast, for LUSC, there was no significant difference in survival between immune cold and immune hot groups for any of the different thresholds (Figure 16C). Similar results were obtained using the median (0.074) instead of the mean as the threshold for immune hot or cold regions (Figure 16F).

[0160] Also, in the single-region TCGA LUAD cohort, we found significance in survival between hot and cold tumors using the mean (0.11) as the threshold between hot and cold tumors (Overall survival (OS): HR=0.61, P=0.0043, Progression-free survival (PFS): HR=0.67, P=0.016; see FIG. 17F). For both OS and PFS, there were various possible thresholds that yielded similar results (FIG. 17H). Using the same thresholds to define immune hot and cold tumors, no significant association with survival was found for either OS or PFS in the TCGA LUSC cohort (FIG. 17I).

[0161] We then used a Cox regression model to determine whether sequential analyses using different multi-region scores related to TCRA T cell fractions were similarly prognostic for survival. We selected four scores for each patient to study: 1) mean TCRA T cell fraction across all regions (indicative of overall T cell infiltration in the tumor), 2) minimum TCRA T cell fraction across all regions, 3) maximum TCRA T cell fraction across all regions, and 4) divergence of TCRA T cell fractions, defined as the maximum region's TCRA T cell fraction divided by the upper 95% confidence interval (to avoid division by 0 or very small numbers) of the minimum region's TCRA T cell fraction. This score represents the fold difference between the regions with the maximum and minimum TCRA T cell fractions, and when high, it may indicate T cell divergence / heterogeneity and possibly subclones that have undergone immune evasion. Since the relative regional immune evasion score is determined by the minimum and maximum scores, we also constructed a Cox model that included both of these scores. The results of this analysis show that the lowest TCRA scores are summarized in Figures 16D and 17G. 32 Consistent with the importance of tumor regions with TCRA, the minimum TCRA fraction across tumor regions was prognostic in the entire TRACERx cohort (HR=0.52, P=0.048). However, the mean and maximum TCRA T cell fractions were not significant in either group. The TCRA T cell fraction divergence score was significant in LUAD (HR=2.2, P=0.023 log-rank test). A model including both the minimum and maximum TCRA scores was significant in both the entire cohort (minimum: HR=0.5, P=0.005; maximum: HR=1.5, P=0.061) and the LUAD subset (minimum: HR=0.36, P=0.016; maximum: HR=2.52, P=0.029) (see Figure 16D). This suggests additional predictability when considering heterogeneity in the TCRA T cell fraction. There were no significant TCRA-related scores in LUSC, and the maximum TCRA T-cell fraction, in combination with the minimum, was only significant in LUAD.

[0162] When the minimum and maximum TCRA T-cell fractions were combined with other clinical phenotypes such as tumor stage, sex, and age, the minimum TCRA T-cell fraction remained significant (HR=0.52, P=0.022, see FIG. 16E). Examining the single-region TCGA survival data for both overall survival (OS) and progression-free survival (PFS), in univariate Cox models, TCRA T-cell fraction was not significant in any comparison, with the closest association being TCGA LUAD in the OS model (HR=0.85, P=0.069).

[0163] Taken together, we believe that this data suggests that TCRA T cells are a prognostic predictor of survival in LUAD patients. Notably, we noted a stronger overall signal in the TRACERx100 cohort despite its small size, indicating the importance of intratumoral immune heterogeneity revealed by multiregional data.

[0164] method See Examples 1-2.

[0165] Multi-sample tumor cohort of patients A multi-sample pan-cancer cohort (see Table 5) was created by combining the TRACERx cohort with a subset of the cohort recently presented by Watkins et al. Tumors were included if they had at least two regions sequenced within the primary tumor where TIL ExTRECT could be used to calculate TCRA T cell fractions. Thus, the final cohort consisted of the multi-region primary tumor dataset with the addition of any metastatic samples that were also sequenced for these patients.

[0166] [Table 4]

[0167] In addition to TRACERx100, the following datasets were combined into the final multi-sample pan-cancer cohort: 1. Brastianos et al. - A cohort focused on the study of brain metastases originating from different histological types. From this cohort, only tumors with multiregional primary samples were included. 2. Gerlinger et al. - A multisample primary cohort of patients with clear cell renal cell carcinoma (KIRC). 3.Harbst et al.- A multi-regional primary cohort of patients with cutaneous malignant melanoma (SKCM). 4. Lamy et al.- A multi-regional primary cohort of patients with bladder cancer (BLCA). 5. Savas et al. - A multiple sample cohort of ER+ and triple negative breast cancer patients (BRCA ER+ and TNBC). 6. Suzuki et al.- A multiregional primary cohort of gliomas. 7. Turajlic et al.- A multi-regional primary cohort of patients with clear cell renal carcinoma (KIRC), papillary renal carcinoma (KIRP), and chromophobe renal carcinoma (KICH). 8. Messaudene et al.- A multidisciplinary primary cohort of patients with HER2+ and ER+ breast cancer.

[0168] Selection of subregions for multi-region sequencing in various datasets In all multi-region cohorts, regions were selected by various methods (see related literature) with two main criteria in mind. The first criterion was to maximize tumor content at the expense of stroma to ensure high-quality mutational and copy number analysis for the primary genomic objective, and the second criterion was that each region represented a physically separate and discrete part of the tumor. In cases where these were not in distinct locations, different measures were used. For example, in the TRACERx100 cohort, sequenced regions were a minimum of 3 mm apart.

[0169] Example 4 - TCRA T cell fraction can predict response to immune checkpoint blockade introduction To further explore areas of research where TCRA T cell fraction may be clinically useful, we used data from the CPI1000+ cohort (Litchfield et al., 2020) to evaluate whether TCRA T cell fraction could be utilized to predict response to immune checkpoint blockade.

[0170] result The CPI1000+ cohort (Litchfield et al., 2020) multi-study dataset consists of 1070 post-CPI (checkpoint inhibitor) treated tumors that received either anti-CTL4 or anti-PDL1 therapy across eight major cancer types (see Figure 19 and Table 3). Response was defined as patients showing a radiological response with complete response (CR) or partial response (PR), while non-responders were defined as stable disease (SD) or progressive disease (PD) by RECIST criteria. Table 3 details the percentage of tumors that contained both WES and RNA-seq data.

[0171] [Table 5]

[0172] Consistent with the importance of tumor TCRA T cell fraction in predicting response to CPIs, we observed a significant difference in TCRA T cell fraction between responders and non-responders across the cohort (see FIG. 20, P=2.3e-7, effect size=0.17). The median TCRA T cell fraction in tumor DNA samples for responders was 0.053 (Q1=0, Q3=0.17), whereas in non-responders it was 0.000268 (Q1=0, Q3=0.084). For immune cold tumors (here defined as tumors with a TCRA T cell fraction of <0.1), it was highly significantly higher in the non-responders (FIG. 21, immune cold tumors are tumors below the dotted line, 78% vs. 63%, Fisher's exact test, odds ratio (OR)=0.47, P=2.25e-06).

[0173] When we separate cohorts by both mutational burden (high: total clonal TMB > 68; low: clonal TMB < 68) and immune microenvironment (hot: TCRA T cell fraction > 0.016; cold: TCRA T cell fraction < 0.016), we find that the association between TCRA T cell fraction and response is independent of mutational burden (see Figure 21). In patients with low mutational burden, immune hot tumors have a response rate of 21% compared to an 8% response rate for immune cold tumors. In tumors with high mutational burden, there is a 45% response rate when the tumor is immune hot and a 30% response rate when it is cold.

[0174] To further evaluate the utility of TCRA T cell fraction compared to RNA-seq-based measurements, all studies with more than 10 samples from individual cancer types with both RNA-seq and TCRA T cell fraction were selected for univariate meta-analysis (see Figure 22; 557 patients across 7 studies and 5 cancer types). As expected, TCRA T cell fraction (odds ratio (OR) = 1.39, P = 0.00858), clonal TMB (OR = 1.59, P = 6.021e-05), and CD8A expression (OR = 1.45, P = 0.0004479) were all found to be significantly associated with response. Neither TCRA T cell fraction from blood nor tumor purity was found to be significantly associated with response, with the difference between tumor and blood TCRA T cell fractions trending toward significance (OR = 1.27, P = 0.016).

[0175] We then evaluated whether tumor TCRA T cell fraction provides additional clinical utility over TMB (tumor mutation burden) and whether it improves response prediction more than RNA-seq measures such as CD8A expression. We created four generalized linear models (GLMs) to predict whether a patient was a responder or non-responder. The first model consisted of clonal TMB alone, while the second and third models used clonal TMB in combination with CD8A RNA-seq expression or TCRA T cell fraction. The fourth model used clonal TMB in combination with both CD8A expression and TCRA T cell fraction. Using either CD8A expression or TCRA improved the predictive value of the models (see Figure 23; AUC=0.66 for CD8A expression and AUC=0.70 for TCRA T cell fraction compared to 0.62 for clonal TMB alone). However, of these, only the model for clonal TMB+TCRA T cell fraction (ROC test, P=0.0028, GLM:clonal TMB+TCRA, AUC=0.68, GLM:clonal TMB, AUC=0.62) was significant compared to clonal TMB alone. Combining TCRA T cell fraction with CD8A expression did not significantly increase the predictive value of the model (see Figure 23, AUC=0.68, ROC test, versus clonal TMB+TCRA model P=0.72). However, when examining the significance of variables in all four models, TCRA T cell fraction was more significant than CD8A (GLM: clonal TMB + TCRA, P = 4.62e-05; GLM: clonal TMB + CD8A, P = 0.000431), and when combined into a single model, TCRA T cell fraction remained significant, but CD8A expression was not (TCRA, P = 0.00601; CD8A, P = 0.06246).

[0176] Collectively, these results suggest that tumor TCRA T-cell fraction can be used as a surrogate for RNA-seq measures of CD8+ infiltration and further suggest that WES-based TCRA T-cell fraction estimates add prognostic value to TMB estimators.

[0177] Given that RNA-seq is often not available, we next evaluated the predictive potential of TCRA T cell fraction in a combined NSCLC CPI cohort (see Table 4). Figure 24A provides an overview of this cohort, including 266 patients assayed with WES, and importantly, does not include RNA-seq or other orthogonal immune measures. Analysis of this lung CPI cohort (see Figure 24B), in which univariate analysis was performed, showed that TCRA T cell fraction (OR=1.44, P=0.005) and blood TCRA T cell fraction (OR=1.39, P=0.0015) were significantly associated with response to CPI. For individual studies, TCRA T cell fraction had an OR>1 in both the Hellman and Shim cohorts, but not in the Rizvi cohort; in contrast, blood TCRA had an OR>1 in all three individual studies. These results demonstrate that TCRA T cell fractions can be calculated from WES data alone and that such estimates from NSCLC and matched blood are predictive of response to immunotherapy.

[0178] [Table 6]

[0179] method See Examples 1-2.

[0180] Meta-analysis of the CPI1000+ cohort The CPI1000+ cohort is described in detail in Litchfield et al. (2020) and includes the following datasets: 1. Snyder et al. (Snyder et al., 2014), Cohort of patients with advanced melanoma following anti-CTLA-4 treatment. 2. Van Allen et al. (Van Allen et al., 2016), Cohort of patients with advanced melanoma following anti-CTLA-4 treatment. 3. Hugo et al. (Hugo et al., 2016), Cohort of patients with advanced melanoma following anti-PD-1 treatment. 4. Riaz et al. (Riaz et al., 2017), Cohort of patients with advanced melanoma following anti-PD-1 treatment. 5. Cristescu et al. (Cristescu et al., 2018), Cohort of patients with advanced melanoma following anti-PD-1 treatment. 6. Cristescu et al. (Cristescu et al., 2018), Cohort of patients with advanced head and neck cancer following anti-PD-1 treatment. 7. Cristescu et al. (Cristescu et al., 2018), “all other tumor types” cohort following treatment with anti-PD-1 (from the KEYNOTE-028 and KEYNOTE-012 studies). 8. Snyder et al. (Snyder et al., 2017), Cohort of metastatic urothelial carcinoma following anti-PD-L1 treatment. 9. Mariathasan et al. (Mariathasan et al., 2018), Cohort of metastatic urothelial carcinoma following anti-PD-L1 treatment. 10. McDermot et al. (McDermott et al., 2018), Cohort of metastatic renal cell carcinoma following anti-PD-L1 treatment. 11. Rizvi et al. (Rizvi et al., 2015), Cohort of patients with non-small cell lung cancer following anti-PD-1 treatment. 12. Cohort of post-anti-PD-1 treatment non-small cell lung cancer samples used by Hellman et al., Litchfield et al., 2020. 13. Le et al. (Le et al., 2015), Colorectal cancer cohort following treatment with anti-PD-1 therapy.

[0181] Of these studies, Snyder et al. (Snyder et al., 2017) was excluded from the analysis due to very poor coverage within the TCRA gene. All samples were aligned to hg19 using bwa mem (v0.7.15) with purity and copy number data calculated from ASCAT as described in Litchfield et al. (2020). Notably, 1008 / 1125 (90%) samples had WES data and 941 / 1125 (83%) samples had sufficient purity and coverage to allow copy number calculation, allowing the calculation of TCRA T cell fraction. Of these samples, 833 / 1125 (74%) had matched RNA-seq data, allowing orthogonal assessment of T cell abundance. Regarding the expansion of this dataset (Shim et al. (Shim et al., 2020)), a cohort of NSCLC post-anti-PD-1 treatment was added for specific NSCLC analysis.

[0182] Univariate and multivariate models for CPI response For univariate models, we followed an adapted procedure from Litchfield et al. (2020). The main difference was that only samples with complete data (RNA-seq for CD8A, clonal TMB, and TCRA T-cell fractions) were included. Meta-analysis of univariate models was performed using the R package “meta” (version 4.13-0). Multivariate models were created with generalized linear models using the function “glm” from the “stats” R package using default values. For ROC curve analysis we used the R package “ROCR” (version 1.0-11).

[0183] Example 5 - Determinants of TCRA T cell fraction in blood, normal and malignant tissues While previous analyses have necessarily focused on T cell infiltration within tumor tissue, the TIL ExTRECT method described herein provides the opportunity to examine T cell infiltration in any sample that has undergone WES. Thus, in this example, we performed analyses to assess the key determinants of T cell immune infiltration in various sample types.

[0184] result We first examined the extent of T cell infiltration in blood collected as normal controls for tumor analysis. Within the TRACERx100 cohort, blood TCRA T cell fractions were significantly higher in women compared to men, and a significant positive relationship was observed between tumor sample TCRA T cell fractions and blood TCRA T cell fractions (Figure 16B, histology: P = 0.066, ES = 0.19; sex: P = 0.0057, ES = 0.28; mean tumor infiltration: ρ = 0.42, P = 1.7e-05). These data suggest that tumor immune infiltration may affect circulating blood T cell levels. Sex and mean tumor infiltration were significant univariate predictors of TCRA T cell fractions and remained significant in linear models predicting blood TCRA T cell fractions that accounted for smoking status, age, and histology. In contrast, no significant differences in T cell infiltration in blood were observed between LUAD and LUSC patients in multivariate models. We also analyzed matched blood samples from TCGA LUAD and LUSC cohorts and found similar results, with gender and tumor histology showing strong associations in linear models predicting TCRA T cell fraction in blood (Figure 17E).

[0185] We then evaluated T cell ExTRECT in an independent NSCLC dataset with lower coverage from TCGA, as described above and in Table 2. Broadly consistent results were observed in blood samples from LUAD and LUSC TCGA patients (Figure 17E).

[0186] Another factor that may affect T cell infiltration in blood is the presence of a viral or bacterial infection elsewhere in the body. To examine this hypothesis, we used the data presented in Poore et al. and quantified the number of microbiome reads from WGS blood samples and RNA-seq tumor samples from the LUAD and LUSC TCGA cohorts using the bioinformatics tool KRAKEN. We observed that blood samples with a high number of microbial reads (greater than the median of 6.81) had significantly higher blood TCRA T cell fractions than those with a low number (Figure 27A, P=0.00092, ES=0.31, Wilcoxon test). In contrast, no association was observed between the total number of microbial reads and tumor TCRA T cell fractions. This suggests that the presence of viruses or bacteria does not elevate the levels of T cell fractions seen in cancer (Figure 27B, P=ns). We then tested whether there were viruses or bacteria associated with TCRA T cell fraction at the individual species level (see Methods). There was no significant association after adjusting for multiple hypotheses for blood samples. For tumor samples, there was one hit for bacteria of the genus Williamsia and genus Paeniclostridium for LUAD and LUSC tumors, respectively (Figures 27C and 28G, Williamsia in LUAD: ρ=-0.17, P=0.00011, FDR P=0.013; Figures 27D and 28G, Paeniclostridium in LUSC: ρ=-0.2, P=0.00013, FDR P=0.015). However, both of these were associated with higher normalized log-cpm values ​​when the TCRA T cell fraction was lower. This suggests that the presence of these bacterial species may not result in increased T cell infiltration and may be opportunistic species taking advantage of the immune-cold tumor microenvironment.

[0187] To further understand the key determinants of blood TCRA T cell fraction, we investigated WES sequencing samples derived from both blood and physiologically normal esophageal epithelium (PNE) tissue. In particular, we examined a dataset containing microdissected WES sequencing samples derived from physiologically normal esophageal epithelium (PNE) tissue as described in Yokoyama et al. Here, a wide range of TCRA T cell fractions was calculated in blood samples, whereas the majority of PNE tissues had no detectable T cell infiltration (Figure 28A, Figure 28H). Dividing participants according to the presence or absence of T cell infiltration in PNE samples revealed a significant association with blood TCRA T cell fraction (Figure 28B, P = 0.021, ES = 0.29), again indicating that, similar to tumor samples, high levels of T cell infiltration in normal tissues influence the TCRA T cell fraction found in blood. In a linear model predicting T cell fraction in blood, only the level of infiltration in normal tissues was a significant independent predictor (Figure 28C). In separate linear models, genomic factors such as normal tissue mutation burden or cancer driver mutation status were found to be not predictive of T cell infiltration in PNE tissues (Figure 28F). This suggests that the detected T cell infiltration may be due to the presence of microbial infection rather than ongoing immune surveillance.

[0188] To further evaluate factors influencing T cell infiltration in tumor tissue, we took advantage of a recently published pan-cancer cohort of multi-sample data (Watkins et al.) to investigate the extent to which different regions of the same tumor may exhibit different levels of immune infiltration and whether the genomic basis for this difference could be identified. In total, we were able to evaluate T cell infiltration in 739 tumor regions from 182 tumors across 14 cancer types (see Table 5).

[0189] A variety of T cell infiltrations were observed across the cohorts (Figure 29A, range 0-58%). Interestingly, the third highest score, comprising 51% of sequenced cells, was found to represent T cells within a single region of the KIRC tumor from patient RMH002. Notably, patient RMH002 had been treated with the antiangiogenic drug sunitinib for 14 weeks prior to nephrectomy, and sunitinib treatment has been shown to increase T cell infiltration in KIRC tumors (Haywood et al.). To quantify the degree of diversity of T cell infiltration within tumors, we classified each tumor into one of three categories based on whether the calculated TCRA T cell fraction was uniformly hot across all regions (defined herein as the mean TCRA T cell fraction within the cohort being ≥ 0.11 across all regions), uniformly cold (defined herein as < 0.11 across all regions), or heterogeneous in T cell fraction. There were significant differences in the proportion of tumors with uniformly hot, uniformly cold, and heterogeneous immune infiltration assessed by cancer type (Figure 29B, Chi-squared test: P = 1.62e-07), with BRCA ER+ tumors being the most heterogeneous (83%) and LUSC tumors being the least heterogeneous (22%). There were also clear differences in the degree of immune infiltration and the proportion of different immune group categories "uniformly hot," "uniformly cold," or "heterogeneous" across cancer types. For example, although BLCA and LUSC contained a similar number of heterogeneous tumors (36% vs. 37%), approximately 64% of BLCA tumors were uniformly hot and 0% were uniformly cold. In contrast, 37% of tumors were uniformly cold and 25% were uniformly hot in LUAD. This suggests that highly localized immune infiltration exists for certain cancer types, particularly BRCA ER+, BRCA HER+, LUAD, and KIRC, which may be subject to significant sampling bias.

[0190] Next, we examined whether there was a relationship between genomic diversity and immune diversity, and investigated subclonal SCNA heterogeneity as a potential mechanism that may lead to heterogeneous immune responses within a single tumor. We first limited the analysis to tumors that contained at least three samples and a heterogeneous mix of T cell fractions (see Methods). Pairwise SCNA heterogeneity between any two regions was calculated as the sum of the proportion of the genome that had unique SCNAs in either region. Figure 29C shows that pairwise analysis shows that pairs of regions with larger differences in TCRA T cell fractions within a tumor (average of all pairwise distances ≧0.065) were more likely to have higher levels of pairwise SCNA heterogeneity (all events: P=0.00021, effect size=0.352; gain events: P=0.0045, effect size=0.324; loss or LOH events: P=0.024, effect size=0.257, n=77).

[0191] Next, to examine whether any particular subclonal SCNA events were associated with immune loss or activation, we identified cytoband regions that underwent subclonal loss or gain in over 30 tumors across a pan-cancer multi-sample cohort (Figure 29D) and tested whether these loss or gain events were associated with changes in TCRA T cell fraction. Only one cytoband-level event, subclonal loss of 12q24.31-32, was found to be significantly associated with a decrease in TCRA T cell fraction (Figure 29E: P=1.3e-05, effect size=0.735).

[0192] To determine whether this effect was due to a single gene, we selected tumors with subclonal loss of 12q24.31-32 within the TRACERx100 cohort and performed differential gene expression analysis on the associated RNA-seq data. After testing 16168 genes, only eight remained significant after multiple testing correction: SPPL3, C12orf76, LYRM9, CIT, UBE3B, ABCB9, OGFOD2, and USP30 (Figure 29F). In particular, SPPL3 was the most significant gene and is located at 12q24.31 together with ABCB9 and OGFOD2. SPPL3 has recently been found to enhance B3GNT5 enzyme activity, which upregulates cell surface glycosphingolipids, thereby interfering with class I HLA function and reducing CD8+ T cell activation (Jongsma et al.). Thus, these data suggest that subclonal loss of 12q24.31 may be selected for during tumor progression across cancer types (occurring in 18.7% or 34 / 182 of tumors within the cohort) as a mechanism of immune evasion.

[0193] method See Examples 1-4.

[0194] KRAKEN TCGA Analysis Preprocessed microbiome data output from the KRAKEN (Wood & Salzberg) analysis performed by Poore et al. were downloaded from ftp: / / ftp.microbio.me / pub / cancer_microbiome_analysis / .

[0195] To create high and low KRAKEN microbiome groups for both blood and tumor samples, we downloaded the file Kraken-TCGA-Voom-SNM-Most-Stringent-Filtering-Data.csv, which contains the normalized log-cpm values, and added the rows for each sample to obtain an overall "microbiome" score. Samples were then split into high and low groups based on the median of this score.

[0196] To investigate the role of any individual microbial species affecting the TCRA T cell fraction, a reduced list of species from the Kraken-TCGA-Voom-SNM-Most-Stringent-Filtering-Data.csv file was selected by removing all species with fewer than 1000 raw reads in total in the TCGA LUAD and LUSC cohorts called from the raw data file Kraken-TCGA-Raw-Data-17625-Samples.csv. This left a total of 59 microbial species, which were individually tested for association with the TCRA T cell fraction using Spearman's correlation for both blood and tumor samples in both LUAD and LUSC.

[0197] Identification of gain, loss, and LOH events in a pan-cancer multi-sample cohort Analysis of whole-exome sequencing was performed as previously described by Jamal-Hanjani et al. Copy number segmentation, tumor purity, and ploidy for each sample were estimated using ASCAT (Van Loo et al.) as previously described (Jamal-Hanjani et al.). These data were used as input to a multi-sample SCNA estimation approach to generate genome-wide estimates for the presence of loss of heterozygosity and for loss, neutral, gain, and amplified copy number states relative to the ploidy of the sample. Log ratio values present in each copy number segment with a log ratio value of ≥5 in all samples of the tumor were examined by comparison to three log ratio thresholds adjusted for sample ploidy using a one-sided t-test with a threshold of P < 0.01. These log ratio thresholds corresponded to <log2[1.5 / 2] for loss and >log2[2.5 / 2] for gain in diploid tumors. Segments not classified as loss or gain were classified as neutral. For each segment, these values relative to the definition of ploidy were combined with loss of heterozygosity detection across all samples from a single tumor.

[0198] Pairwise subclonal SCNA score To calculate the pairwise subclonal SCNA measure, we used the classification outlined in the previous Methods section to create three groups of pairwise subclonal SCNA scores. First, we considered segments affected by either gains or losses over ploidy or LOH as abnormal, and compared each pair of regions from a single patient's disease, classifying the abnormal region as clonal if it was abnormal in both samples, and subclonal if it was abnormal in only one sample. This same process was repeated only for gains over ploidy, and then losses over ploidy and LOH were considered together.

[0199] Site band level SCNA analysis Segments were mapped to hg19 cytobands to allow comparisons between tumors. When multiple segments mapped to a cytoband, the SCNA status (gain or loss relative to ploidy) of the segment with the greatest overlap with the cytoband was selected. For SCNA gain and loss analysis, cytoband-level events were selected if they occurred subclonally more than 30 times across the entire cohort. Bands passing this threshold within the same region (e.g., all cytobands on 1p36) were then grouped together. A paired Wilcoxon test was used to assess whether tumor regions within a single patient with subclonal SCNA events had significant differences in TCRA T cell fractions versus regions without events.

[0200] Selection of multiple sample tumors with heterogeneous immune infiltration To be included, tumors had to have at least three regions sequenced and fulfill two requirements: 1) have a pair of regions with a major change in immune infiltration defined as a difference in TCRA T cell fraction of ≥ 0.1, and 2) have a pair of regions with a minor or no change in immune infiltration defined as a difference in TCRA T cell fraction of < 0.1. An example of a tumor that meets this requirement is a tumor with three regions R1, R2, and R3 with TCRA T cell fractions of 0.01, 0.01, and 0.2, respectively. The R1-R2 pair has a difference in TCRA T cell fraction of 0, and both the R1-R3 and R2-R3 pairs have a major difference of 0.19. Within the multi-sample tumor cohort, 58 patients met these criteria.

[0201] RNA-seq differential gene expression analysis of patients with subclonal 12q24.31-32 losses Differential gene expression analysis was performed on TRACERx100 RNA-seq patients with subclonal 12q24.31-32 loss. Using R 4.0.0, first the edgeRR package (version 3.32.1) was used for sample-specific TMM (trimmed mean of M-values) normalization, then standard edgeR filtering methods were used to filter out genes with low expression, after which the Limma-Voom method from the limmaR package (version 3.46.0) was used to calculate Voom fits and obtain p-values ​​for gene expression differences. Comparisons were weighted for patient and histological characteristics as inhibitors and p-values ​​were FDR corrected for multiple testing. Results were then visualized using the R EnhancedVolcano package (version 1.8.0).

[0202] Example 6 - Estimation of T cell fraction, B cell fraction, and immune clonotypes in samples from WGS data introduction The approach described in Example 1 and exemplified in Examples 2-5 can be used to explore additional situations, such as when whole genome sequence data are available, when immune cell compartments other than αβ T cells are of interest, and / or when the use of specific V and J segments, for example to investigate clonotypes, is of interest. In this example, we demonstrate the implementation of the general approach described in Example 1 with respect to WGS data. By utilizing a subset of the TRACERx100 cohort that includes both WES and WGS data, we have validated the WGS version of the method, which shows high concordance between its use with WES data and its use with WGS data. Furthermore, we take advantage of the greater coverage within the WGS samples to provide an alternative model for modeling the read depth ratio (the segment model) that takes into account the nature of V(D)J recombination. Using orthogonal RNA-seq data, we show that the score accurately depicts the immune microenvironment, and using downsampling, we show that the method is accurate up to two-fold depth. We then show that this approach can also be used to investigate γδ T cell fractions and TCR and BCR diversity. Finally, we also demonstrate the use of this method to study survival in a large pan-cancer cohort (data not shown).

[0203] result Unlike whole exome sequencing (WES), WGS includes coverage across the entire genome. In the context of estimating copy number changes, this brings several important advances: 1) uniform coverage allows precise localization of copy number breakpoints, and 2) less bias, such as GC content bias or reference allele bias. Therefore, we have developed the T-cell ExTRECT method described in Example 1 applied to WGS. 3We reasoned that the method may be able to estimate T cell fractions with greater accuracy at lower depths. Furthermore, six other genes (in addition to TCRA as mentioned above), TCRD, TCRG, TCRB, IGH, IGK, and IGL, may also undergo V(D)J recombination and be similarly quantifiable, potentially allowing estimation of αβ and γδ T cell compartments as well as B cells within a sample. This is particularly beneficial since many WGS datasets lack accompanying immune data. The method described herein can provide a comprehensive immune profile of a sample of any WGS sample without the need for additional data.

[0204] The method described above was adapted for WGS by applying GC correction and quality control to 1000bp windows instead of exon segments from the WES capture kit, and applying additional quality control to identify any 100bp genomic windows that showed evidence of being outliers in that they had much higher or lower read depth than nearby regions (see Methods for details). An overview of this adapted method (called "adapted GAM") is given in Figure 30. Furthermore, we were able to apply the same WGS method to TCRB, TCRG, and IGH by choosing different genomic loci for focal and normalization regions (see Table 7 for genomic locations used; Figure 31 shows a plot of CRUK0085T1-R3 as an example). This particular dataset was not suitable for illustrating the method for IGL and IGK, because the quality of sequence alignments within these loci is highly questionable, and likely due to their high complexity and the quality of the reference sequences, it is difficult to map short read sequences in a uniform and unbiased manner. Nevertheless, the same method could be applied to these loci if higher quality alignments were available, for example by utilizing higher quality reference sequences. Similarly, due to the small size of the TCRD locus within TCRA, the method was not directly applied to TCRD. However, the segment usage-based approach described below was able to estimate the fraction of T cells containing TCRD segments.

[0205] Benchmarking applications to WGS data To take advantage of the uniform coverage in the WGS data, an alternative model to the smooth generalized linear model (GAM) was designed. It is a piecewise constrained linear model in which breakpoints are forced to be located at known positions of the V and J segments within the V(D)J locus. This "segment model", described in Methods and Figure 30, has the following theoretical advantages over the GAM model: 1) it is based on actual V(D)J recombination biology; 2) individual fractions for individual V and J gene segments can be calculated, providing insight into segment usage and T or B cell receptor diversity; and 3) fractions of TCRD-specific segments can be used as an indicator of γδ T cells. This model can be used with any sequencing data that has sufficient coverage over the region of interest. In certain WES data, this approach (although theoretically possible) was less advantageous because the particular exon capture kit used has large gaps in coverage in the region of interest. However, this may not be the case in other capture data sets.

[0206] We tested the accuracy of both the GAM and segment models applied to WGS data by determining the fractional fractions of T and B cells in a set of 322 TRACERx100 samples for which matched WES and WGS were available, and 126 WGS samples with orthogonal RNAseq data. Using the GAM model to compare WES and WGS TCRA T cell fractions, we found significant correlations in both blood samples (ρ=0.71, P=4.7e-16, FIG. 32A) and tumor samples (ρ=0.72, P<2.2e-16, FIG. 32A). Using the segment model, we found a similar significant correlation with GAM-based WES scores (blood: ρ=0.69, P=6.7e-15; tumor: ρ=0.7, P<2.2e-16, FIG. 32A). However, for both models, we found that WGS values ​​were higher at lower T cell fractions. This may be due to the increased accuracy and sensitivity possible with WGS data. In support of this, when examining the orthogonal TRACERx100 RNAseq data, a stronger correlation with Danaher T cell score was seen for WGS scores than for matched WES scores (WES: ρ=0.64, P=1.6e-15, (Danaher T cell score)=)1.9+8.3 (ExTRECT T cell fraction (WES)); WGS(GAM): ρ=0.82, P<2.2e-16; WGS(Segment model): ρ=0.81, P<2.2e-16; Figure 32B).

[0207] Validation of the approach for other loci Danaher scores calculated from matched RNAseq data were used to orthogonally validate the accuracy of T cell fractions calculated from TCRB and TCRG and B cell fractions calculated from IGH. Significant correlations with T cell Danaher scores were found for TCRB (GAM: ρ = 0.66, P < 2.2e-16; segment model: ρ = 0.81, P < 2.2e-16, Figure 32B) and TCRG (GAM: ρ = 0.64, P = 3.1e-16; segment model: ρ = 0.8, P < 2.2e-16, Figure 32C), and significant correlations with B cell Danaher scores were found for IGH (GAM: ρ = 0.47, P = 2.3e-8; segment model: ρ = 0.44, P = 2.1e-07, Figure 32C). For the B cell fraction scores from the segment model, there were some obvious outlier points with high B cell fraction but low B cell Danaher scores. These scores tended to have very low log likelihood values ​​from the segment model, indicating very noisy samples and poor fit. Removing all samples with log likelihood <0 improved the correlation with B cell Danaher scores (GAM: ρ = 0.41, P = 4.3e-06; segment model: ρ = 0.6, P = 5.1e-13). Overall, these results indicate that the segment model produces scores that have a stronger association with Danaher scores than the corresponding values ​​from the GAM model.

[0208] Comparing the different T cell fractions from TCRA, TCRB, and TCRG, the TCRB T cell fraction was highly correlated with the TCRA T cell fraction, but the values ​​were about half in the GAM model (ρ=0.69, P<2.2e-16, y=0.00059+0.49x, Fig. 32E) and slightly more than half in the segment model (ρ=0.91, P<2.2e-16, y=0.033+0.57x, Fig. 32F), likely due to the exclusion of alleles where only one of the TCRB alleles frequently undergoes V(D)J recombination. In contrast, we found that the TCRG T cell fraction was unexpectedly higher than the 1–5% of CD3+ T cells typically characterized as γδ T cells, and was also highly correlated with the TCRA T cell fraction (GAM: ρ = 0.71, P < 2.2e-16, y = -0.028 + 0.71, Figure 32E and segmented model: ρ = 0.94, P < 2.2e-16, y = 0.005 + x). Although there were significantly more samples with TCRG fractions close to 0 in the GAM model compared to TCRA (1.7% TCRA T cell fraction < 1e-5 vs. 32.8% TCRG T cell fraction < 1e-5), the high fraction in the majority of samples, and the near exact correlation between the two in the segmented model, are widely considered biologically implausible if this signal is entirely due to γδ T cells. Instead, αβ T cells are known to first undergo rearrangement of the TCRG locus before committing to their final lineage, and results from the segment model suggest that possessing a rearranged TCRG locus is a near-universal feature of αβ T cells and that the TCRG locus can provide an accurate measure of T cell fraction.

[0209] The TCRD locus is located entirely within the TCRA locus and is too small in size to accurately generate a score for the fraction of γδ T cells using a GAM model with the available data due to a large amount of noise. However, a segment model can be used to calculate the fraction of T cells that contain segments derived from TCRD gene segments and therefore the fraction of T cells that arise from γδ T cells. The segment-based score for TRDV1 was found to be significantly correlated with TPM RNAseq values ​​for TCRD segments TRDV1 (ρ=0.25, P=0.0011, Figure 32G), TRDJ1+TRDJ2+TRDJ3+TRDJ4 (ρ=0.15, P=0.047, Figure 32G), and TRDC (ρ=0.23, P=0.0021, Figure 32G) in the TRACERx100 cohort.

[0210] Investigating Depth Requirements A subset of TRACERx100 samples (n=19), representing various T and B cell fractions from either the TCRA, TCRB, TCRG, or IGH loci, was selected for downsampling analysis to ascertain the lowest depth at which either the GAM or segment model could summarise the complete depth score. Using samtools, each sample was randomly downsampled four times to depths of 60x, 30x, 10x, 5x, 2x, 1x, 0.5x or 0.1x. The results of this analysis are given in Figure 33 and show high fidelity to total depth values ​​at 2x in all cases for both the GAM and segment models (TCRA GAM: ρ = 0.83, P < 2.2e-16, TCRA seg: ρ = 0.75, P = 7.5e-15, TCRB GAM: ρ = 0.46, P = 2.7e-05, TCRB seg: ρ = 0.53, P = 8.4e-07, TCRG GAM: ρ = 0.5, P = 4.2e-06, TCRG seg: ρ = 0.61 P = 5e-09, T-cell average GAM: ρ = 0.74, P = 1.9e-14, T-cell average seg: ρ = 0.76 P = 1.4e-15, IGH GAM: ρ = 0.37, P = 0.0011, IGH seg: ρ=0.56, P=1.3e-07, Figure 33A-J). However, at this low depth, values ​​are often scaled down (e.g., the T-cell average segment model at 1x has ρ=0.59, but points on the line y=-0.028+1.6x when compared to full depth values). The observed high fidelity of scores at these low depth downsampled files allows the T-cell ExTRECT method to be designed to perform very well on low pass WGS samples (e.g., above 0.1x, between 0.1x and 0.5x, between 0.1x and 2x, etc.) using either the GAM model or the segment model by calculating coverage in bins (e.g., bins from about 100bp to about 10000bp, specifically values ​​between 100bp and 1000bp are likely to be particularly useful. Approximate values ​​can be selected using a benchmarking approach as described above).

[0211] The ExTRECT approach can be used to investigate immune cell repertoire diversity To gain insight into more general TCR and BCR diversity, we investigated the use of the approach described above. This was done within samples by examining the V and J segment usage predicted from the best fit of the segment model. Figures 34A-C show the proportion of V segments used for TCRA, TCRB, and TCRG across different tumor regions and germline blood samples in TRACERx patient CRUK0085, and Figure 34D shows the V segment usage for IGH in TRACERx patient CRUK0045. These can be obtained using the same formula for f as described above based on the position of maximum recombination, but instead, by using r VDJThese clonotype plots show the potential of the method to identify samples with broader clonotypes, such as in the case of CRUK0085, where more than half of the T cells are predicted to use the V-segment TRAV18 at R4, with a high proportion of TRBV29_1. It is possible to plot the change in the proportion of segment usage across the cohort, as shown in Figure 34E for the TRAV segment. In this case, despite the formation of clear clusters, there was no clear relationship to histological features, clinical sex, or whether the data was from tumor or blood samples. Using the bioinformatics software MiXCR (Bolotin et al., 2015) to compare the V segments called with those from the matched RNAseq data, we identified a significant correlation with the number of V segments called from the T-cell ExTRECT segment model (TCRA: ρ = 0.54, P = 1.5e-12, TCRB: ρ = 0.39, P = 6.6e-07, TCRG: ρ = 0.29, P = 0.00025, IGH: ρ = 0.42, P = 5.9e-08, Figure 34F-I). However, the number of V segments called with MiXCR was significantly higher than the T-cell ExTRECT model in all cases, suggesting that the fitted models are oversimplified and do not capture the full TCR or BCR diversity. This can be addressed by extending the V and J segments selected from the model by obtaining a sample of best-fitting solutions rather than simply the "optimal" solutions. For example, the top 10, 20, 30, 50, or 100 solutions (in terms of fit to the data) can be selected and examined, e.g., based on their confidence intervals, to select a model that captures the full TCR or BCR diversity. These results are particularly important because several features of TCR or BCR diversity have been shown to be indicative of cancer prognosis and / or response to therapy.For example, B cell gene expression signatures associated with various types of diversity have been found to be prognostic in breast and ovarian cancer (Iglesia et al., 2014), TCR repertoire analysis, including determination of TCR baseline clonality and post-treatment expansion, has been shown to be a source of biomarkers for response to checkpoint inhibitor therapy (Aversa et al., 2020; Hopkins et al., 2018), and TCR diversity in TILs at baseline was found to be a prognostic predictor in various cancers (Valpione et al., 2021).

[0212] Associations of the cancer immune environment with survival, clinical determinants, and cancer type in a pan-cancer cohort The T cell ExTRECT segment model was applied to a large pan-cancer WGS cohort of over 15,000 participants across 20 different cancer types. The analysis provided T cell fractions from TCRA, TCRB, and TCRG, B cell fractions from IGH, and fractions for individual V and J gene segments. The majority of these participants had matched blood germline WGS samples, and the majority of the remaining hematological cancer patients had matched germline samples from saliva or other normal tissues. This provided the opportunity to investigate differences in immune fractions in patient blood at the time of sampling for the first time in a large pan-cancer cohort.

[0213] This analysis revealed extensive diversity in the range of immune landscapes (TCRA T cell fraction, TCR diversity measured by the number of TRAV segments called by the model with fractions above 0.01) within and between cancer types in blood and tumor samples. Linear models can also be fitted for both TCRA T cell fraction and the number of TRAV segments called by the segment model to ascertain the dependence of these values ​​on sex, age, and ancestry. These models can be fitted for blood TCRA T cell fraction and the number of predicted TRAV segments to investigate, for example, whether age is associated with a decrease in T cell fraction and diversity in pan-cancer cohorts and whether this is observed to different degrees across various cancer subtypes.

[0214] Such analyses can also be performed on IGH B cell fraction scores from T cell ExTRECT to investigate IGH B cell fractions in blood and tumors across various cancer types, as well as clinical determinants of IGH B cell fraction and diversity (e.g., gender, age, ancestry).

[0215] To ascertain the contribution of germline and ancestry to the determination of T cell fraction in blood, a GWAS analysis can be performed using PLINK and a set of unrelated patients in a cancer or pan-cancer cohort (see Methods). To improve performance, GWAS can be run separately for blood TCRA, TCRB, and TCRG T cell fraction values. These values ​​are calculated independently from various regions of the genome, and by looking at hits above the recommended significance level (P<1e-5) in all three, one can be confident that they are not the result of noise or artifacts in the data arising from SNPs in the V(D)J genes that are associated with poor mapping quality and coverage values ​​(rather than actual T cell fraction signal). For any SNPs identified in one or all three of these GWAS, one can then look at the enrichment of these SNPs in various cancer types.

[0216] Finally, survival data can be analyzed from the latest records in hospital episode statistics, e.g., utilization data on the date of cancer diagnosis, death records, and latest follow-up time. Kaplan-Meier curves can then be obtained for e.g., TCRA T cell and IGH B cell fractions in blood and tumor samples, in each case dividing the cohort into high or low fractions by the median. This analysis can be used to investigate whether the TCRA T cell fraction in blood / tumor or the IGH B cell fraction in blood / tumor is significantly associated with survival.

[0217] Finally, to better understand the relationship of immune cell fraction to other clinical features and its heterogeneity across various cancer types, a Cox proportional model can be fitted consisting of the TCRA T cell fraction in blood and tumor, the IGH B cell fraction in blood and tumor, and the TRDV1 T cell fraction as a surrogate for γδ T cells in blood and tumor. The Cox model can also include age and sex, and whether the patient received chemotherapy before surgery, to account for possible confounding associations with the T or B cell fraction. This last factor can be included as a surrogate for how aggressive or late stage the cancer was clinically. Alternative criteria such as cancer stage can also be used. The results of such an analysis can be inspected by looking at a heat map of z-scores from the Cox models run on any particular cohort. This can reveal whether the TCRA T cell fraction in blood / tumor or the IGH B cell fraction in blood / tumor is a significant risk in the pan-cancer cohort and in any individual cancer type.

[0218] method The overall process for calculating the T cell fraction from WGS closely followed the method in Example 1. This included normalizing coverage using the regions within the TCRA locus at the beginning and end previously defined, calculating the single sample read depth ratio using γ=1 for Illumina Hiseq, and formulas in equations (3) or (7) for the T cell fraction. The only difference was that r VDJ (deviation from 0 of the read depth ratio at the position of maximum V(D)J recombination). In particular, the GC correction method described in Example 1 was adapted to work on WGS rather than exons as described by the capture kit. GC content was calculated as the number of exons within a 1000 bp window (GC 1000bp ) and a smoothed version (GC macro The coverage values ​​were then normalized with respect to GC content by fitting and employing the residuals from the following linear model:

[0219]

number

[0220] This is because the local GC content is at the level of individual exons (GC exon ) but at the level of 1000bp windows (GC 1000bp ) are the same as those described in Example 1.

[0221] As a final step, a robust quality control process identified any 100 bp segments within any V(D) gene that were outliers in terms of coverage, either 1) across the entire cohort, or 2) associated with a subset of the cohort and linked to a known germline genomic variant. For the first step, the average GC-corrected ratio of all V(D)J genes was calculated for the set of all GEL lung cohorts. This revealed segments with extremely low or extremely high coverage compared to the surrounding regions. By fitting a GAM model to the average GC-corrected ratio, it is possible to filter out those above and below a certain threshold (±0.25) from the fitted line. This worked well for TCRA, but for other genes such as TCRB, there were large clustered segments that prevented the fitting of the GAM model. For these, some obvious outlier regions were first manually removed, e.g., all regions with an average GC ratio above 0.25 for TCRB. After removal of the segments thus identified, additional segments with bias were identified after GWAS analysis using the PLINK software (see Renteria et al., 2013). In particular, GWAS analysis was performed to identify any SNPs associated with T or B cell fraction. Genetic regions with coverage bias in specific exons were identified and flagged for removal. The scores were recalculated and the GWAS process was repeated to check for removal of bias. The removed segments were either under- or over-covered compared to the surrounding regions associated with a specific germline genotype, making a specific variant in a gene strongly associated with T or B cell fraction purely due to this artifact. In these cases, specific samples with the relevant genotype were selected and the average GC-corrected ratio was calculated. Additional outlier segments were then identified and flagged for removal using the same procedure as described above, and the PLINK / GWAS analysis was rerun to ensure that no further artifactual segments were present.

[0222] The method was extended to other T cell receptor genes, TCRB and TCRG, which undergo V(D)J recombination, and the B cell receptor IGH, using the above method, but redefining the positions used for normalization of the read depth ratio and the positions where the maximum deviation is expected. In general, the maximum deviation can be expected to be located between the last V segment and the first J segment, and the region was selected accordingly. This is given in Table 7.

[0223] [Table 7]

[0224] Similarly, the method was also applied to examine segment usage, using the method described above, but with r VDJ For the value of , the estimate from the segment model for the particular segment corresponding to the particular gene under investigation (specifically the absolute value of the difference between the estimate from the segment model for the particular segment and the preceding segment, i.e., decrease or increase depending on whether the V or J segment is considered at the start of the segment under investigation, as shown in FIG. 31). Thus, the normalization regions in Table 7 are used with the segments corresponding to the "maximum V(D)J recombination region" for each gene in Table 8 below (provided as gene hgnc symbol-start_hg38-end_hg38). For the estimation of the fraction of γδ T cells, any of TRDV1, TRDV2, TRDJ1, TRDJ2, TRDJ3, and TRDJ4 can be used (all located within the TCRA locus). We found that the V segment TRDV1 correlates most strongly with the RNAseq values ​​associated with γδ T cells. Therefore, in this example, this segment was used for the estimation of the fraction of γδ T cells. In this example, a segment model is used when examining segment usage, and the estimate from the model for the corresponding segment is r VDJHowever, the same regions provided in Tables 7 and 8 can be used in combination with GAM, preferably when coverage allows for the study of these regions, to estimate summary values ​​for each of multiple segments to be investigated and to estimate the r for the second of a pair of consecutive segments. VDJ Determine the difference between consecutive segments as:

[0225] [Table 8]

[0226] [Table 9]

[0227] [Table 10]

[0228] [Table 11]

[0229] [Table 12]

[0230] Creation of a new T cell ExTRECTV(D)J segment model The method described in Example 1 uses a GAM model to VDJ, and therefore the T cell fraction in the sample. The GAM model simplifies the modeling of the process of V(D)J recombination as a smoothed process. This is useful when using WES data due to the unevenness of coverage in large segments of the genome within the TCRA locus that are not sequenced by any exome capture kit. This comes at the cost of increased vulnerability to noise in the data. Furthermore, this is not constrained to match the known biological reality of the V(D)J recombination process. For WES data, the benefits outweigh the costs when the coverage within the locus under investigation is not uniform. However, using WGS data (or any other data that does not suffer from this bias due to the specific regions captured), it is possible to use a model that is directly based on the biological process of V(D)J recombination. This is implemented by modeling the read depth ratio as a series of constant piecewise segments, where the positions of the breakpoints represent the possible beginnings of the deletion sites after V(D)J recombination, which are known from the positions of the V and J gene segments and can therefore be provided as constraints on the model. Therefore, we created a model in which segments are preselected and aligned with the positions of V and J segments. Furthermore, the model may be constrained by knowledge from V(D)J recombination that the read depth ratio should start from 0 and decrease monotonically from there until the point of maximum V(D)J recombination. For example, in TCRAV(D)J recombination, only some TCR chains select the first TCRV-1 segment, and all other V segments are deleted, but after the last V segment, all TCR chains have deletions. Similarly for J segments, they should increase monotonically until they reach 0.

[0231] To fit this model, each possible segment with known breakpoints corresponding to the V and J genes (see Table 8) was converted into a vector of ones and zeros, with ones inside the region and zeros outside the region. These vectors were then fitted to the normalized read ratio data using a constrained linear model (using functions from the R package restriktor v0.3), with inequality constraints chosen for each of the n V and m J segments as follows: V1<0; V2<0. <V1;…;V n <V n-1 ;J1>V n ;J2>J1;.…;J m >J m-1 ;J m <0. Using these fitted values, the total T or B cell fraction can be calculated from the maximum deviation of the model, e.g., from the last V segment, and the individual fractions from individual segments, e.g., segments associated with either TRDV1 or TRDV2 in the TCRA locus representing γδ T cells.

[0232] GWAS analysis with PLINK PLINK was run on a cohort of unrelated participants with WGS germline samples from blood, controlling for the first 20 genetic PCs, covariates for sex, age, and disease type. PLINK input was preprocessed to have a MAF>0.001 across the cohort, passed QC filters including missingness and sufficient depth, underwent variant normalization, made all multi-allelic variants biallelic, and was variant aligned to ensure parsimoniousness. In addition to these pre-set filters and QC before running PLINK, we performed LD pruning in 500 kb windows with an R2 threshold of 0.2 to ensure that all variants had a MAF>0.001 and genotype missingness ≤0.2 within the cohort. Finally, a Hardy-Weinberg equilibrium test was performed with a threshold of 0.000001 before running PLINK.

[0233] Example 7 - Discussion The immune system is widely recognized as one of the key factors influencing both tumor formation and subsequent cancer progression. T cells play a key role in the cancer immune microenvironment by eliminating neoantigen-bearing cancer cells, and therefore the number of T cells in a tumor sample is an important clinical factor that may determine the course of the disease. In the context of cancer, bulk DNA sequencing of tumor tissues has been used primarily to characterize the somatic changes that drive tumor development. However, the present inventors demonstrate that DNA sequencing can also be used to study the immune microenvironment of a sample with the method presented herein.

[0234] TRECs have been previously clinically examined in the context of screening for severe combined immunodeficiency (SCID) in newborns (van der Spek, Groenwold, van der Burg, & van Montfrans, 2015), where the absence of TRECs is used to infer T cell lymphopenia. However, these screening methods are based on sequencing and quantification of TRECs themselves, not WES data. A recent study of breast cancer single-cell genomic sequencing identified deletion events in T cells within the TCRA gene (Baslan et al., 2020). Here, we build on this study and explicitly use TRECs to estimate T cell fractions in WES samples, as detailed in Example 1.

[0235] The method described herein (TIL ExTRECT) provides accurate estimators of immune infiltration, which demonstrate clinical utility. We found that TCRA T cell fraction estimators are strongly correlated with orthogonal immune measures, and cancer cell line and simulated WES data support their reliability (Example 2). Furthermore, we demonstrate that the inferred TCRA T cell fraction is a prognostic predictor in LUAD, validating this finding in the TCGA LUAD cohort (Example 3). In this context, we showed that TCRA T cell fraction is associated with response to CPIs in a pan-cancer cohort, improving the predictive value of clonal TMB alone (Example 4). We further demonstrated that TCRA T cell fraction in blood and tumor samples, like IGH B cell fraction, is a predictor of survival in a pan-cancer cohort and in many specific cancer types (Example 6). The fraction of γδ T cells estimated based on the fraction of TRDV1 T cells was found to be a predictor of survival in certain cancer types.

[0236] Furthermore, TIL ExTRECT allows the calculation of T cell (or B cell) fractions in the dataset, which was not possible before. Using this, we demonstrate that T cell fractions in blood are heterogeneous, significantly higher in women than men (consistent with current findings) and are also associated with microbial infections as expected (Example 5). It has been shown that T cell ExTRECT can be used for datasets without previous immune annotations, rather than just summarizing known immune associations. In a pan-cancer multi-sample cohort without RNA-seq, we observe striking variations in the degree of heterogeneity of T cell infiltration across different cancer types (Example 5). Much of this heterogeneity appears to be driven by subclonal SCNAs, and we identify that subclonal loss of 12q24.31-32 is associated with a significant reduction in T cell infiltration. This may be related to reduced expression of SPPL3, which causes upregulation of cell surface glycosphingolipids, thereby interfering with the function of class I HLA molecules. This loss may be an immune evasion mechanism in multiple cancer types and is subclonal and therefore frequently occurs late in the tumor progression trajectory.

[0237] In Example 2, we demonstrate that TCRA T cell fraction is strongly correlated with orthogonal immune measures from different platforms and across multiple different datasets. Cancer cell lines and simulated WES data confirm that this score provides a reliable estimate of the fraction of T cells in sequenced samples. Simulation and downsampling of high-depth TRACERx100 samples shows that 30x coverage provides sufficient signal to calculate a reliable T cell fraction (although lower coverage can be used to distinguish between high and low T cell content, which may be useful as a cruder biomarker). This relatively low coverage means that it is applicable to most DNA sequencing datasets. In Example 6, we demonstrate that the method can be applied to WGS data, and that signals with higher resolution than WES can be obtained outside the TCRA region, even at lower depths. Using orthogonal RNA-seq data, we show that the score accurately depicts the immune microenvironment, and using downsampling, we show that the method is accurate to 2-fold depth. In Example 6, we further demonstrate the potential of the method to investigate both γδ T cell fraction and TCR and BCR diversity. Besides WES and WGS, any NGS method that targets TCRA genes (or any gene undergoing V(D)J recombination under investigation) may be compatible with this approach. Without wishing to be bound by theory, it is believed that sequencing data is preferable to, for example, SNP array data, because SNP array data usually have a rather low resolution in the main region of the VDJ locus of interest. Furthermore, a targeted panel-based sequencing approach can be used that specifically targets TCRA genes (or genes undergoing V(D)J recombination under investigation) and provides TCRA T cell fraction (or TCRD, TCRG, TCRB, IGH, IGK, or IGL cell fraction). Indeed, one can easily envision combining such an approach with current panel-based sequencing approaches to estimate TMB to facilitate its use as a diagnostic tool.The sequencing platform specific constant (γ in equations (2), (3), and (7), where a value of 1 is appropriate for WES or WGS data) can be adjusted depending on the sequencing platform used. Additionally, platform specific biases such as GC content bias can be assessed and taken into account.

[0238] For these reasons, TIL ExTRECT may be useful in the context of reanalysis of previously published NGS cancer datasets that do not normally have unbiased estimators of T cell content. TCRA T cell fraction may also be of value as a simple biomarker in the clinic due to its association with response to immunotherapy as demonstrated in Example 4, and with survival as demonstrated in Example 6.

[0239] While there are other methods that provide immune-related data from samples, such as imaging slides, which require FFPE blocks to be acquired, cut, and stained before being evaluated manually or digitally, these are often unavailable or logistically cumbersome. On the other hand, T cell ExTRECT can be used in conjunction with DNA sequencing platforms, which are increasingly commonly used in research and in the clinic.

[0240] Although the method was developed within the framework of cancer research, it is worth considering its application in a broader clinical setting. The measure can be calculated from any NGS analysis obtained from any healthy or diseased tissue or blood. Analysis of matched blood samples within the cancer cohort demonstrated a significant relationship between circulating TCRA T cell fraction and the cancer type and gender of the patient, demonstrating the potential of this approach (see also Examples 3 and 6). Furthermore, analysis of only blood samples in the lung CPI cohort showed that the inferred TCRA T cell fraction was a predictor of response to immunotherapy without analysis of the tumor sample itself (Figure 24B). Moreover, non-NGS polymerase chain reaction (PCR) quantification of TREC levels in blood samples has already been used to monitor T cell generation in the thymus in immunotherapy trials in breast cancer (Page et al.) and prostate cancer (Madan et al.). Furthermore, T cell ExTRECT can be applied outside of oncology. One future clinical use of blood TCRA T cell fraction could be in the context of SCID, a rare congenital syndrome that often results in severe T cell deficiency requiring rapid hematopoietic stem cell transplantation for treatment, where TIL ExTRECT could be combined with newborn genomic screening to simultaneously identify potentially damaging germline mutations and reduced T cell content. Unfortunately, at the time of writing, SCID is not subject to routine newborn screening in most countries, despite affecting 1 in 50,000 live births, and the majority of cases are diagnosed only after severe opportunistic infections, often when treatment is too late (Chan et al.). TIL ExTRECT could potentially identify SCID by including TCRA within newborn screening gene panels for inherited diseases.

[0241] Since T cells and VDJ recombination are defining features of the adaptive immune system in jawed vertebrates, the general method presented herein is species agnostic. Although the TIL ExTRECT method described in the above examples is optimized for the human genome, the method can be extended to other species, including many research model organisms (including mouse, chicken, ferret, etc.) (mainly by defining the corresponding specific genomic regions to be used and evaluating and addressing any region-specific biases). It is also worth noting that VDJ recombination supports this novel approach to TCRA T cell fraction estimation, but is not unique to genes encoding TCR-α receptors. Other TCR genes, namely β, γ, and δ, also undergo VDJ recombination, as do the BCR immunoglobulin genes IGH, IGL, IGK. Thus, the method can also be used to calculate fractions of both αβ, γδ T cells and B cells with either IGL or IGK light chains, as demonstrated in Example 6. These can use appropriate chromosomal regions, such as (hg19):TCRB chr7:141998851-142510972;TCRG chr7:38279625-38407656;IGH chr14:106032614-107288051;IGL chr22:22380474-23265085;IGK chr2:89156674-90274235, or regions in Tables 7 and 8, to assess and address region-specific bias. These other immune cell types are generally less prevalent, so some of these may not perform as well depending on noise in the data and fraction of cells of interest in the sample, but this extended method may provide a detailed description of the immune microenvironment of cancer samples utilizing DNA sequencing data alone (demonstrated in Example 6).

[0242] In summary, the approach described herein, i.e., TIL ExTRECT, may have important applications in both basic and translational research by providing a cost-effective technique to characterize immune infiltrates along with somatic alterations in both human and model system datasets without the need for RNA sequencing.

[0243] References AbdulJabbar et al. (2020). Geospatial immune variability illuminates differential evolution of lung adenocarcinoma. Nature Medicine, 26(7), 1054-1062. Aran, D., Hu, Z., & Butte, A. J. (2017). xCell: Digitally portraying the tissue cellular heterogeneity landscape. Genome Biology, 18(1), 1-14. Aversa I, Malanga D, Fiume G, Palmieri C. Molecular T-Cell Repertoire Analysis as Source of Prognostic and Predictive Biomarkers for Checkpoint Blockade Immunotherapy. Int J Mol Sci. 2020 Mar 30;21(7):2378. Baslan et al. (2020). Novel insights into breast cancer copy number genetic heterogeneity revealed by single-cell genome sequencing. ELife, 9, 1-21. Bolotin et al. (2012). Next generation sequencing for TCR repertoire profiling: Platform-specific features and correction algorithms. European Journal of Immunology, 42(11), 3073-3083. Bolotin, D. A. et al. MiXCR: software for comprehensive adaptive immunity profiling. Nat. Methods 12, 380-381 (2015). Brastianos, P. K. et al. Genomic characterization of brain metastases reveals branched evolution and potential therapeutic targets. Cancer Discov. 5, 1164-1177 (2015). Cancer Genome Atlas Research Network. (2012). Comprehensive genomic characterization of squamous cell lung cancers. Nature, 489(7417), 519-525. Cancer Genome Atlas Research Network. (2014). Comprehensive molecular profiling of lung adenocarcinoma. Nature, 511(7511), 543-550. Carter et al. (2012). Absolute quantification of somatic DNA alterations in human cancer. Nature Biotechnology, 30(5), 413-421. Chan, K. & Puck, J. M. Development of population-based newborn screening for severe combined immunodeficiency. J. Allergy Clin. Immunol. 115, 391-398 (2005). Cristescu et al. (2018). Pan-tumor genomic biomarkers for PD-1 checkpoint blockade-based immunotherapy. Science, 362(6411). Danaher et al.(2018). Pan-cancer adaptive immune resistance as defined by the Tumor Inflammation Signature (TIS): Results from The Cancer Genome Atlas (TCGA). Journal for ImmunoTherapy of Cancer, 6(1), 1-17. Davoli, T., Uno, H., Wooten, E. C., & Elledge, S. J. (2017). Tumor aneuploidy correlates with markers of immune evasion and with reduced response to immunotherapy. Science, 355(6322). Denkert et al. (2016). Standardized evaluation of tumor-infiltrating lymphocytes in breast cancer: Results of the ring studies of the international immuno-oncology biomarker working group. Modern Pathology, 29(10), 1155-1164. Dieci et al. (2018) Update on tumor-infiltrating lymphocytes (TILs) in breast cancer, including recommendations to assess TILs in residual disease after neoadjuvant therapy and in carcinoma in situ: A report of the International Immuno-Oncology Biomarker Working Group on Breast Cancer. Seminars in Cancer Biology, Volume 52, Part 2, 2018, Pages 16-25. Favero et al. (2015). Sequenza: Allele-specific copy number and mutation profiles from tumor sequencing data. Annals of Oncology, 26(1), 64-70. Galon et al. (2006). Type, density, and location of immune cells within human colorectal tumors predict clinical outcome. Science, 313(5795), 1960-1964. Gerlinger, M. et al. Intratumor Heterogeneity and Branched Evolution Revealed by Multiregion Sequencing. N. Engl. J. Med. 366, 883-892 (2012). Gerlinger, M. et al. Genomic architecture and evolution of clear cell renal cell carcinomas defined by multiregion sequencing. Nat. Genet. 46, 225-233 (2014). Ghandi et al. (2019). Next-generation characterization of the Cancer Cell Line Encyclopedia. Nature, 569(7757), 503-508. Goodman et al. (2017). Tumor mutational burden as an independent predictor of response to immunotherapy in diverse cancers. Molecular Cancer Therapeutics, 16(11), 2598-2608. Harbst, K. et al. Multiregion whole-exome sequencing uncovers the genetic evolution and mutational heterogeneity of early-stage metastatic melanoma. Cancer Res. 76, 4765-4774 (2016). Haywood et al. Sunitinib’s effect on tumor infiltration of CD8 T cells in renal cell carcinoma (RCC) and modulation of their function by altering VEGF-induced upregulation of PD1 expression. Journal of Clinical Oncology, 34(2),591 (2016). Hendry, S., Salgado, R., Gevaert, T., Russell, P. A., John, T., Thapa, B., … Fox, S. B. (2017). Assessing Tumor-infiltrating Lymphocytes in Solid Tumors. Advances In Anatomic Pathology (Vol. 24). Hopkins et al. T cell receptor repertoire features associated with survival in immunotherapy-treated pancreatic ductal adenocarcinoma. JCI Insight. 2018 Jul 12;3(13):e122092. Huang, W., Li, L., Myers, J. R., & Marth, G. T. (2012). ART: A next-generation sequencing read simulator. Bioinformatics, 28(4), 593-594. Hugo et al. (2016). Genomic and Transcriptomic Features of Response to Anti-PD-1 Therapy in Metastatic Melanoma. Cell, 165(1), 35-44. Iglesia et al. Prognostic B-cell signatures using mRNA-seq in patients with subtype-specific breast and ovarian cancer. Clin Cancer Res. 2014 Jul 15;20(14):3818-29. Jamal-Hanjani et al.(2017). Tracking the Evolution of Non-Small-Cell Lung Cancer. New England Journal of Medicine, 376(22), 2109-2121. Jongsma et al. The SPPL3-Defined Glycosphingolipid Repertoire Orchestrates HLA Class I-Mediated Immune Responses. Immunity. 2021 Jan 12;54(1):132-150.e9. doi: 10.1016 / j.immuni.2020.11.003. Epub 2020 Dec 2. Erratum in: Immunity. 2021 Feb 9;54(2):387. PMID: 33271119. Lamy, P. et al. Paired exome analysis reveals clonal evolution and potential therapeutic targets in urothelial carcinoma. Cancer Res. 76, 5894-5906 (2016). Le, D. T., Uram, J. N., Wang, H., Bartlett, B. R., Kemberling, H., Eyring, A. D., … Diaz, L. A. (2015). PD-1 Blockade in Tumors with Mismatch-Repair Deficiency. New England Journal of Medicine, 372(26), 2509-2520. Levy, E., Marty, R., Garate Calderon, V.et al.Immune DNA signature of T-cell infiltration in breast tumor exomes.Sci Rep6,30064 (2016). Li et al. (2017). TIMER: A web server for comprehensive analysis of tumor-infiltrating immune cells. Cancer Research, 77(21), e108-e110. Litchfield et al. Meta-analysis of tumor and T cell-intrinsic mechanisms of sensitization to checkpoint inhibition. Cell 184.3 (2021): 596-614. Lopez et al. (2020). Interplay between whole-genome doubling and the accumulation of deleterious alterations in cancer evolution. Nature Genetics, 52(3), 283-293. Madan, R. A. et al. Clinical and immunologic impact of short-course enzalutamide alone and with immunotherapy in non-metastatic castration sensitive prostate cancer. J. Immunother. Cancer 9, (2021). Mariathasan et al. (2018). TGFβ attenuates tumour response to PD-L1 blockade by contributing to exclusion of T cells. Nature, 554(7693), 544-548. Marquez et al. (2020). Sexual-dimorphism in human immune system aging. Nature Communications, 11(1). McDermott et al. (2018). Clinical activity and molecular correlates of response to atezolizumab alone or in combination with bevacizumab versus sunitinib in renal cell carcinoma. Nature Medicine, 24(6), 749-757. McGrath et al. Detecting T cell receptor rearrangements in silico from non-targeted DNA-sequencing (WGS / WES). bioRxiv preprint doi.org / 10.1101 / 201947; October 13, 2017. Messaoudene, M. et al. T-cell bispecific antibodies in node-positive breast cancer: Novel therapeutic avenue for MHC class i loss variants. Ann. Oncol. 30, 934-944 (2019). Middleton et al. (2020). The National Lung Matrix Trial of personalized therapy in lung cancer. Nature, 583(7818), 807-812. Newman et al. (2015). Robust enumeration of cell subsets from tissue expression profiles. Nature Methods, 12(5), 453-457. Newman et al. (2019). Determining cell type abundance and expression from bulk tissues with digital cytometry. Nature Biotechnology, 37(7), 773-782. Page, D. B. et al. A phase II study of dual immune checkpoint blockade (ICB) plus androgen receptor (AR) blockade to enhance thymic T-cell production and immunotherapy response in metastatic breast cancer (MBC). J. Clin. Oncol. 37, TPS1106-TPS1106 (2019). Poore GD, Kopylova E, Zhu Q, et al. Microbiome analyses of blood and tissues suggest cancer diagnostic approach. Nature. 2020;579(7800):567-574. Racle et al. (2017). Simultaneous enumeration of cancer and immune cell types from bulk tumor gene expression data. ELife, 6, 1-25. Riaz et al.(2017). Tumor and Microenvironment Evolution during Immunotherapy with Nivolumab. Cell, 171(4), 934-949.e15. Renteria ME, Cortes A, Medland SE. Using PLINK for Genome-Wide Association Studies (GWAS) and data analysis. Methods Mol Biol. 2013;1019:193-213. Riester, M., Singh, A.P., Brannon, A.R.et al.PureCN: copy number calling and SNV classification using targeted short read sequencing.Source Code Biol Med11,13 (2016). Rizvi et al. (2015). Mutational landscape determines sensitivity to PD-1 blockade in non-small cell lung cancer. Science, 348(6230), 124-128. Robert et al. (2011). Ipilimumab plus Dacarbazine for Previously Untreated Metastatic Melanoma. New England Journal of Medicine, 364(26), 2517-2526. Rosenthal et al. (2019). Neoantigen-directed immune escape in lung cancer evolution. Nature, 567(7749), 479-485. Samstein et al. (2019). Tumor mutational load predicts survival after immunotherapy across multiple cancer types. Nature Genetics, 51(2), 202-206. Savas, P. et al. The Subclonal Architecture of Metastatic Breast Cancer: Results from a Prospective Community-Based Rapid Autopsy Program “CASCADE”. PLoS Med. 13, 1-25 (2016). Schadendorf et al. (2015). Pooled analysis of long-term survival data from phase II and phase III trials of ipilimumab in unresectable or metastatic melanoma. Journal of Clinical Oncology, 33(17), 1889-1894. Shen, R., & Seshan, V. (2015). FACETS : Fraction and Allele-Specific Copy Number Estimates from Tumor Sequencing. Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series., 1, 50. Retrieved from http: / / biostats.bepress.com / mskccbiostat / paper29 Shim et al. (2020). HLA-corrected tumor mutation burden and homologous recombination deficiency for the prediction of response to PD-(L)1 blockade in advanced non-small-cell lung cancer patients. Annals of Oncology, 31(7), 902-911. Snyder et al. (2014). Genetic Basis for Clinical Response to CTLA-4 Blockade in Melanoma. New England Journal of Medicine, 371(23), 2189-2199. Snyder et al. (2017). Contribution of systemic and somatic factors to clinical response and resistance to PD-L1 blockade in urothelial cancer: An exploratory multi-omic analysis. PLoS Medicine, 14(5), 1-24. Suzuki, H. et al. Mutational landscape and clonal architecture in grade II and III gliomas. Nat. Genet. 47, 458-468 (2015). Thorsson et al. (2018). The Immune Landscape of Cancer. Immunity, 48(4), 812-830.e14. Topalian et al. (2012). Safety, Activity, and Immune Correlates of Anti-PD-1 Antibody in Cancer. New England Journal of Medicine, 366(26), 2443-2454. Turajlic, S. et al. Deterministic Evolutionary Trajectories Influence Primary Tumor Growth: TRACERx Renal. Cell 173, 595-610.e11 (2018). Valpione, S., Mundra, P.A., Galvani, E.et al.The T cell receptor repertoire of tumor infiltrating T cells is predictive and prognostic for cancer survival.Nat Commun12,4098 (2021). Van Allen et al. (2016). Erratum for the report “genomic correlates of response to CTLA-4 blockade in metastatic melanoma.” Science, 352(6283), 207-212. van der Spek, J., Groenwold, R. H. H., van der Burg, M., & van Montfrans, J. M. (2015). TREC Based Newborn Screening for Severe Combined Immunodeficiency Disease: A Systematic Review. Journal of Clinical Immunology, 35(4), 416-430. Van Loo et al. (2010). Allele-specific copy number analysis of tumors. Proceedings of the National Academy of Sciences of the United States of America, 107(39), 16910-16915. Wang et al. (2007). PennCNV: An integrated hidden Markov model designed for high-resolution copy number variation detection in whole-genome SNP genotyping data. Genome Research, 17(11), 1665-1674. Wang, V. G., Kim, H., & Chuang, J. H. (2018). Whole-exome sequencing capture kit biases yield false negative mutation calls in TCGA cohorts. PLoS ONE, 13(10), 1-14. Watkins, T. B. K. et al. Pervasive chromosomal instability and karyotype order in tumour evolution. Nature 587, 126-132 (2020). Yokoyama, A., Kakiuchi, N., Yoshizato, T. et al. Age-related remodelling of oesophageal epithelia by mutated cancer drivers. Nature 565, 312-317 (2019). Yoshihara et al. (2013). Inferring tumour purity and stromal and immune cell admixture from expression data. Nature Communications, 4. Zaccaria, S., & Raphael, B. J. (2020). Accurate quantification of copy-number aberrations and whole-genome duplications in multi-sample tumor sequencing data. Nature Communications, 11(1). Zhang et al. (2003). Intratumoral T Cells, Recurrence, and Survival in Epithelial Ovarian Cancer. New England Journal of Medicine, 348(3), 203-213. Wood, DE & Salzberg, SL Kraken: Ultrafast metagenomic sequence classification using exact alignments. Genome Biol. 15, (2014).

[0244] All references cited in this specification are incorporated by reference in their entirety for all purposes to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference in its entirety.

[0245] This application claims priority to UK Patent Application No. 2110555.6, filed July 22, 2021, and UK Patent Application No. 2203451.6, filed March 11, 2022. The entire contents of both priority applications are incorporated herein by reference for all purposes.

[0246] The specific embodiments described herein are provided by way of example and not by way of limitation. Various modifications and variations of the compositions, methods, and uses of the techniques described will be apparent to those skilled in the art without departing from the scope and spirit of the techniques described. Subheadings herein are included for convenience only and should not be construed as limiting the disclosure in any way.

[0247] The methods of any embodiment described herein may also be provided as a computer program, or as a computer program product or computer readable medium including a computer program configured to perform the above-described method when executed on a computer.

[0248] Unless the context dictates otherwise, the above feature descriptions and definitions are not limited to any particular aspect or embodiment of the invention, but apply equally to all aspects and embodiments described.

[0249] Throughout this specification and the claims, the following terms have the meanings expressly associated therewith, unless the context clearly dictates otherwise. The phrase "in one embodiment" as used herein may refer to the same embodiment, but not necessarily to the same embodiment. Furthermore, the phrase "in another embodiment" as used herein may refer to different embodiments, but not necessarily to different embodiments. Thus, as described below, various embodiments of the invention can be readily combined without departing from the scope or spirit of the invention.

[0250] It should be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from "about" one particular value and / or to "about" another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values ​​are expressed as approximations, by use of the antecedent "about," it will be understood that the particular value forms another embodiment. The term "about" in reference to numerical values ​​is optional and may mean, for example, ±10%.

[0251] Throughout this specification, including the claims which follow, unless the context otherwise requires, the words "comprises" and "including," and variations such as "comprising," "comprising," and "including," will be understood to imply the inclusion of a specified integer, step, or group of integers or steps, and not the exclusion of any other integer, step, or group of integers or steps.

[0252] Other aspects and embodiments of the present invention provide those aspects and embodiments described above, except where the context indicates otherwise, replacing the word "comprising" with the word "consisting of" or "consisting essentially of."

[0253] The features disclosed in the above description, the appended claims, or the accompanying drawings may be appropriately expressed in their specific form or in terms of means for performing a disclosed function or a method or process for obtaining a disclosed result, but such features may be used individually or in any combination to realize the invention in various of its forms.

Claims

**Claim 1** A computer-implemented method for determining the fraction of lymphocytes in a mixed sample containing genomic material from multiple cell types, comprising: obtaining a read depth profile of the sample, including read depths along a predetermined genomic region of interest that includes at least a portion of a genomic locus undergoing VDJ recombination, from sequencing data; Normalizing the lead depth along the predetermined region of interest with reference to a baseline lead depth derived from a lead depth profile in a subset of the region of interest, to obtain a plurality of lead depth ratios (r i ), and Obtaining a summarized read depth ratio value (r VDJ ) for a subset of the regions of interest that are likely to be deleted by VDJ recombination, wherein the subset of the regions of interest that are likely to be deleted by VDJ recombination includes a region between the end position of the last V segment of the genomic locus undergoing VDJ recombination and the start position of the J segment of the genomic locus undergoing VDJ recombination, Determining the fraction ratio (f) of lymphocytes in the sample as a function of the summarized read depth ratio value (r VDJ ), and determining the fraction ratio (f) of lymphocytes in the sample as a function of the summarized read depth ratio value (rVDJ) can be 【Number 1】 determining a value of the fraction (f) of lymphocytes in the sample, where γ is a constant, p is the fraction of abnormal cells with aneuploidy in the predetermined region of interest in the mixed sample, and ΨT is the copy number of the abnormal cells within the predetermined region of interest; A method comprising the steps of: **Claim 2** The method according to claim 1, wherein the genomic locus undergoing VDJ recombination is selected from the TCR A, TCR B, TCR G, TCR D, IGH, IGL, or IGK loci, and / or the fraction of lymphocytes is the fraction of T lymphocytes, the fraction of B lymphocytes, the fraction of αβ T lymphocytes, the fraction of γδ T lymphocytes, the fraction of B cells or T cells containing any gene segment of the TCR A, TCR B, TCR G, TCR D, IGH, IGL, or IGK loci, or the fraction of B cells or T cells containing any gene segment listed in Table 8. **Claim 3** The method according to claim 2, wherein the predetermined region of interest includes one or more exons from the TCR A, TCR B, TCR G, TCR D, IGH, IGL, or IGK loci, or includes a plurality of exons from the TCR A, TCR B, TCR G, TCR D, IGH, IGL, or IGK loci, and the plurality of exons includes at least one exon corresponding to each of the V, D, and C segments of each of the respective loci. The method according to claim 2. **Claim 4** The method according to claim 3, wherein the plurality of exons includes exons corresponding to a plurality of V segments of each of the respective loci. **Claim 5** The method according to claim 1, wherein the read depth profile is obtained from whole exome sequencing, whole genome sequencing, or panel sequencing data. **Claim 6** Determining the fraction (f) of lymphocytes in the sample comprises: 【Number 2】 Determining a value of f, where γ is a constant, according to the method of claim 1. **Claim 7** The method of claim 1, wherein the subset of the regions of interest that are likely to be deleted by VDJ recombination includes one or more D segments (optionally all D segments) of the genomic locus undergoing VDJ recombination.

8. further comprising applying the model to the plurality of read depth ratios, and obtaining a summarized read depth ratio value (r VDJ ), which is based on the value of the model within the subset of the regions of interest that are likely to be deleted by VDJ recombination, and optionally, the model is a generalized linear model, a piecewise constrained linear model having a plurality of breakpoints selected to correspond to the positions of a plurality of V and J genes within the region of interest, or a Bayesian model that models the use of segments from the read depth and the previous distribution of the use of each of the plurality of V and J genes within the region of interest. The method according to claim 1.

9. The method of claim 1, further comprising obtaining a baseline read depth from the subset of the regions of interest, wherein the baseline read depth is a summarized read depth value across the entire subset of the regions of interest, and / or the subset of the regions of interest used to obtain the baseline read depth is a subset of the regions of interest that are less likely to be deleted by VDJ recombination.

10. The method of claim 1, wherein the baseline read depth is obtained from a subset of the regions of interest that are less likely to be deleted by VDJ recombination, the subset including the first n V segments of the genomic locus undergoing VDJ recombination and / or the C segment of the genomic locus undergoing VDJ recombination.

11. Obtaining a read depth profile includes obtaining a plurality of raw read depth values across the predetermined region of interest and smoothing the read depth values. Optionally, smoothing the read depth values includes replacing the raw read depth values with corresponding rolling median values over a window of a predetermined width. The method of claim 1.

12. The method of claim 1, wherein the mixed sample is a blood sample or a tumor sample.

13. A computer-implemented method of providing a prognosis prediction for a subject diagnosed with cancer and / or determining whether the subject is likely to respond to immunotherapy, comprising Determining the lymphocyte fraction in one or more tumor samples from the subject using the method according to claim 1, and classifying the sample as "immunocold" or "immunohot" according to the estimated lymphocyte fraction in one or more of the tumor samples from the patient, wherein a sample having a lymphocyte fraction exceeding a threshold is classified as immunohot, and a sample having a lymphocyte fraction below the threshold is classified as immunocold, (a) a patient classified as immunocold has a worse survival prognosis than a patient classified as immunohot, and / or (b) a patient classified as immunohot is more likely to benefit from immunotherapy than a patient classified as immunocold, a computer-implemented method comprising the above.

14. A method for determining whether a subject has T cell lymphocytopenia, the method comprising determining the lymphocyte fraction in one or more samples from the subject using the method according to claim 1, a computer-implemented method.

15. A processor, A computer-readable medium including instructions that, when executed by the processor, cause the processor to execute the steps of the method according to any one of claims 1 to 14, A system comprising the above.