Systems and methods for deconvolution of expression data
Nonlinear regression models are employed to analyze RNA expression data, enhancing the accuracy and efficiency of determining cellular composition ratios in tumor tissues, particularly for cell types like B cells, CD4+ T cells, and endothelial cells, improving tumor microenvironment analysis.
Patent Information
- Application Number
- JP2024135878
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-30
- Filing Date
- 2024-08-16
- Publication Date
- 2026-02-20
- Estimated Expiration
- 2041-03-12
AI Technical Summary
Existing methods for analyzing cell composition in tumor tissues are limited in accuracy and efficiency, particularly in distinguishing between different cell types and their proportions, which is crucial for understanding tumor microenvironments.
A method utilizing nonlinear regression models, specifically nonlinear regression models, to analyze RNA expression data from biological samples, allowing for precise determination of cellular composition ratios by processing expression data through these models to estimate the proportions of various cell types, including B cells, CD4+ T cells, CD8+ T cells, endothelial cells, fibroblasts, lymphocytes, and macrophages, among others.
Enhances the accuracy and efficiency of determining cellular composition ratios in tumor tissues, providing a detailed understanding of the tumor microenvironment by accurately estimating the proportions of different cell types, thereby improving diagnostic and therapeutic strategies.
Smart Images

Figure 0007818662000048 
Figure 0007818662000049 
Figure 0007818662000050
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a division of "SYSTEMS AND METHODS FOR DECONVOLUTION OF No. 63 / 108,262, filed on March 12, 2020, entitled "GENE EXPRESSION DATA" and U.S. Provisional Patent Application No. 63 / 108,262, filed on March 12, 2020, entitled "GENE EXPRESSION DATA" "MACHINE LEARNING SYSTEMS AND METHODS FOR DECONVOLUTION OF GENE E" filed on 35 U.S.C. § 119(e) of U.S. Provisional Patent Application No. 62 / 988,700 entitled "EXPRESSION DATA" and US Provisional Patent Application No. 2006 / 0119994, each of which is incorporated herein by reference in its entirety. be absorbed. [Background technology]
[0002] Generally, a tumor mass (or other diseased tissue) contains a population of malignant cells (e.g., cancer cells) and a population of malignant cells (e.g., cancer cells), e.g., The microenvironment may contain immune cells, fibroblasts, and extracellular matrix proteins. It is done. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Newman et al., “Robust enumeration of cell subsets from tissue expression profiles,” Nat. Methods 12, pp. 453–457 (2015) [Non-patent document 2] Newman et al., “Determining cell type abundance and expression from bulk tissues with digital cytometry,” Nat Biotechnol 37, pp. 773–782 (2019) [Non-patent document 3] Finotello et al., “Molecular and pharmacological modulators of the tumor immune contexture revealed by deconvolution of RNA-seq data,” Genome Med 11, p. 34 (2019) [Non-patent document 4] Hao et al., “Fast and Robust Deconvolution of Tumor Infiltrating Lymphocytes from Expression Profiles using Least Trimmed Squares”, bioRxiv p. 358366; doi: https: / / doi.org / 10.1101 / 358366 [Non-Patent Document 5] Aran et al., “xCell: digitally portraying the tissue cellular heterogeneity landscape,” Genome Biol. 18, p. 220 (2017) [Non-patent document 6] Monaco et al., “RNA-Seq signatures normalized by mRNA abundance allow absolute deconvolution of human immune cell types,” Cell Rep. 26, pp. 1627–1640, e1627 (2019) [Non-Patent Document 7] Vaught et al., “Biospecimens and biorepositories: from afterthought to science,” Cancer Epidemiol Biomarkers Prev. 2012 Feb;21(2):253–5. [Non-patent document 8] Vaught and Henderson, “Biological sample collection, processing, storage and information management”, IARC Sci Publ. 2011;(163):23-42. [Non-Patent Document 9] Li et al., JCO Precis Oncol. 2018; 2: PO.17.00091 [Non-Patent Document 10] Consea et al., "A survey of best practices for RNA-seq data analysis," Genome Biology 201617:13 [Non-Patent Document 11] Pereira and Rueda (bioinformatics-core-shared-training.github.io / cruk-bioinf-sschool / Day2 / rnaSeq_align.pdf) [Non-Patent Document 12] Wagner et al., Theory Biosci. (2012) 131:281-285 Summary of the Invention [Means for solving the problem]
[0004] Some embodiments utilize at least one computer hardware processor obtaining expression data for a biological sample, the biological sample having been previously obtained from the subject; and the expression data comprises a first expression data set associated with a first set of genes associated with a first cell type. and performing one or more nonlinear regressions including the expression data and a first nonlinear regression model. determining a first cellular composition ratio for a first cell type using the regression model; The first cellular proportion indicates an estimated proportion of cells of a first cell type in the biological sample, and the first cells The step of determining a first cellular composition ratio for the type includes: subjecting the first expression data to a first nonlinear regression model; processing the sample with a data set to determine a first cellular composition ratio for the first cell type; and The present invention provides a method for analyzing a cell composition ratio of a sample, the method comprising the steps of: (a) performing the steps (a) and (b); and (c) outputting the cell composition ratio of the sample.
[0005] Some embodiments include at least one hardware processor and at least one hardware When executed by a hardware processor, it is executed by at least one hardware processor. obtaining expression data for a biological sample, wherein the biological sample has been previously obtained from the subject; a first expression data set associated with a first set of genes associated with a first cell type; and performing one or more nonlinear regressions including the expression data and a first nonlinear regression model. determining a first cellular composition ratio for a first cell type using the regression model; The first cellular composition ratio indicates an estimated proportion of cells of a first cell type in a biological sample, The step of determining a first cell composition ratio for a cell type includes: subjecting the first expression data to a first nonlinear regression; processing the data through the model to determine a first cellular composition ratio for the first cell type; and and outputting the first cellular composition ratio. and at least one non-transitory computer-readable storage medium storing the Provides systems.
[0006] Some embodiments, when executed by at least one hardware processor, obtaining expression data for the biological sample on at least one hardware processor; wherein the biological sample has previously been obtained from the subject and the expression data relates to the first cell type. a step of obtaining first expression data relating to a first set of genes; and and one or more nonlinear regression models, including a linear regression model, for the first cell type. A step of determining a first cellular constituent ratio, wherein the first cellular constituent ratio is a first cellular constituent ratio in a biological sample. Determining a first cell type cell composition ratio by indicating an estimated proportion of cells of one cell type The step processes the first expression data through a first nonlinear regression model to generate a first expression profile for the first cell type. determining a first cellular constituent ratio based on the determined cellular constituent ratio; and outputting the first cellular constituent ratio. at least one non-transitory computer storing processor-executable instructions for performing the steps of A computer-readable storage medium is provided.
[0007] In some embodiments, the subject has, is suspected of having, or has cancer. There is a risk of this happening.
[0008] In some embodiments, the expression data is RNA expression data.
[0009] In some embodiments, processing the first expression data with a first nonlinear regression model provides first expression data as input to a first nonlinear regression model to generate a first expression vector from a first cell type; obtaining a corresponding output representing an estimated proportion of RNA from the first cell type; and determining a first cellular composition ratio for the first cell type based on the above.
[0010] In some embodiments, the expression data relates to a first set of genes associated with a first cell type. and a first nonlinear regression model using the first expression data as an input. and generating a first value for an estimated proportion of RNA from a first cell type using the The first submodel is compared with the second expression data and the estimated proportion of RNA from the first cell type. The first value and are used as inputs to generate a second value for the estimated proportion of RNA from the first cell type. and a second sub-model configured to generate a
[0011] In some embodiments, the expression data is a second cell type associated with a second cell type that is different from the first cell type. and one or more nonlinear regression models, wherein the second expression data relates to the set of two genes. comprises a second nonlinear regression model. Some embodiments are based, at least in part, on the second cell The second expression data is subjected to a second nonlinear regression model to determine a second cellular proportion for the type. determining a second cell composition ratio for a second cell type by treating the cells with a It further includes a process.
[0012] In some embodiments, the first cell type is a B cell, a CD4+ T cell, a CD8+ T cell, an endothelial cell, a thymocyte, or a leukocyte. A group consisting of fibroblasts, lymphocytes, macrophages, monocytes, NK cells, neutrophils, and T cells are selected.
[0013] In some embodiments, the first expression data is for a first cell type in Table 2 The present invention includes expression data for at least 10 genes selected from the group of genes of
[0014] In some embodiments, the expression data comprises a plurality of gene sets associated with each of the plurality of cell types. the plurality of gene sets includes expression data associated with a first gene set and a first cell; and a plurality of cell types, including a plurality of cell types, wherein the one or more nonlinear regression models are Some embodiments use expression data associated with multiple gene sets to generate a model. The method further comprises determining a plurality of cell composition ratios for a plurality of cell types, In some embodiments, the step of determining a plurality of cellular constituent ratios comprises a first cellular constituent ratio. The process includes, for each cell type of the plurality of cell types, determining, at least in part, a cellular composition for the cell type. Using each of the plurality of nonlinear regression models to determine a ratio, By processing expression data related to sets of genes associated with cell types, The method includes a step of determining the composition ratio of each cell type.
[0015] In some embodiments, the genes in the set of genes are selected from the set of genes in Table 2. The gene comprises at least 25 genes selected from a group of genes, and determines the composition ratio of multiple cells. The steps include processing expression data for at least 25 genes.
[0016] In some embodiments, the genes in the set of genes are selected from the set of genes in Table 2. The gene comprises at least 35 genes selected from a group of genes, and determines the ratio of multiple cell constituents. The steps include processing expression data for at least 35 genes.
[0017] In some embodiments, the genes in the set of genes are selected from the set of genes in Table 2. The gene comprises at least 50 genes selected from a group of genes, and determines the ratio of multiple cell constituents. The steps include processing expression data for at least 50 genes.
[0018] In some embodiments, the genes in the set of genes are selected from the set of genes in Table 2. The gene comprises at least 75 genes selected from a group of genes, and determines the ratio of multiple cell constituents. The steps include processing expression data for at least 75 genes.
[0019] In some embodiments, the genes in the set of genes are selected from the set of genes in Table 2. A gene encoding a gene for determining a plurality of cell constituent ratios, comprising at least 100 genes selected from a group of genes. The step of processing comprises processing expression data for at least 100 genes.
[0020] In some embodiments, the one or more nonlinear regression models are based on one or more random for Includes rest regression models.
[0021] In some embodiments, the one or more nonlinear regression models are based on one or more neural networks. Includes network regression models.
[0022] In some embodiments, the one or more nonlinear regression models are based on one or more support vectors. Includes terminology regression models.
[0023] In some embodiments, the first nonlinear regression model simulates, at least in part, obtaining simulated expression data; and performing a first nonlinear and training a regression model.
[0024] Some embodiments include obtaining simulated expression data and and training a first nonlinear regression model using the expression data.
[0025] In some embodiments, obtaining the simulated expression data comprises generating simulated expression data, the generating the simulated expression data comprising: Obtaining a set of RNA expression data from one or more biological samples, The set includes microenvironment cell expression data and malignant cell expression data; generating simulated microenvironment cell expression data using the cell expression data; and generating simulated malignant cell expression data using the malignant cell expression data. The process involves comparing the simulated microenvironment cell expression data with the simulated malignant cell expression data. and combining the expression data to generate at least a portion of the simulated expression data. Includes:
[0026] Some embodiments provide an expression profile for a first cell type and a determining a malignant tumor expression profile using the first cellular composition ratio of .
[0027] In some embodiments, the first nonlinear regression model is based on the simulated RNA expression data. obtaining training data comprising simulated RNA expression data for a first cell type; and collecting first RNA expression data for a first set of genes associated with the first cell. training a first nonlinear regression model to estimate the proportion of RNA from a cell type, The training step comprises training a first cell using a first nonlinear regression model and first RNA expression data. generating an estimated ratio of RNA from the first cell type, as well as using the estimated ratio of RNA from the second cell type; and updating parameters of the first nonlinear regression model. So they are trained.
[0028] Some embodiments utilize at least one computer hardware processor obtaining RNA expression data for a biological sample, the biological sample having cancer; previously obtained from a subject suspected of having or at risk of having cancer and A. Expression data includes first RNA expression data associated with a first set of genes associated with a first cell type. The first RNA expression data includes the gene expression data for the first cell type in Table 2. a first cell comprising expression data for at least 10 genes selected from the group of cells; The types are B cells, CD4+ T cells, CD8+ T cells, endothelial cells, fibroblasts, lymphocytes, and macrophages. a first RNA selected from the group consisting of a first RNAi gene, ... determining a first cellular proportion for a first cell type using the expression data; The first cellular composition ratio indicates an estimated ratio of cells of a first cell type in the biological sample, and the first The step of determining a first cellular composition ratio for a cell type includes inputting the first cellular composition ratio into a first nonlinear regression model. to provide first RNA expression data and generate corresponding output data representing estimated proportions of RNA from the first cell type. obtaining a titer for the first cell type based on the estimated ratio of RNA from the first cell type; and a step of determining a first cellular composition ratio. .
[0029] Some embodiments include at least one hardware processor and at least one hardware When executed by a hardware processor, it is executed by at least one hardware processor. obtaining RNA expression data for a biological sample, the biological sample having cancer; previously obtained from a subject suspected of having or at risk of having cancer and A. Expression data includes first RNA expression data associated with a first set of genes associated with a first cell type. The first RNA expression data includes the gene expression data for the first cell type in Table 2. a first cell comprising expression data for at least 10 genes selected from the group of cells; The types are B cells, CD4+ T cells, CD8+ T cells, endothelial cells, fibroblasts, lymphocytes, and macrophages. a first RNA selected from the group consisting of a first RNAi gene, ... determining a first cellular proportion for a first cell type using the expression data; The first cellular composition ratio indicates an estimated ratio of cells of a first cell type in the biological sample, and the first The step of determining a first cellular composition ratio for a cell type includes inputting the first cellular composition ratio into a first nonlinear regression model. to provide first RNA expression data and generate corresponding output data representing estimated proportions of RNA from the first cell type. obtaining a titer for the first cell type based on the estimated ratio of RNA from the first cell type; determining a first cellular composition ratio; and at least one non-transitory computer-readable storage medium storing the Provides systems.
[0030] Some embodiments, when executed by at least one hardware processor, obtaining RNA expression data for the biological sample on at least one hardware processor; wherein the biological sample is from a subject having, suspected of having, or at risk of having cancer. The RNA expression data is previously obtained from a subject, and the RNA expression data is from a first cell type associated with the first cell type. The first RNA expression data includes first RNA expression data associated with a set of genes, the first RNA expression data being listed in Table 2 ( At least 10 genes selected from the group of genes for the first cell type in Table 2 The first cell type includes expression data for B cells, CD4+ T cells, CD8+ T cells, endothelial cells, and Consists of cells, fibroblasts, lymphocytes, macrophages, monocytes, NK cells, neutrophils, and T cells. and using the first RNA expression data to generate a first RNA expression profile for the first cell type. determining a first cellular composition ratio of a first cellular composition ratio in a biological sample; Determining a first cellular composition ratio for a first cell type, indicating an estimated proportion of cells of the cell type. providing the first RNA expression data as input to a first nonlinear regression model to estimate the first cell type obtaining a corresponding output representing an estimated proportion of RNA from the first cell type; and an estimate of RNA from the first cell type. determining a first cellular composition ratio for the first cell type based on the ratio; at least one non-transitory computer-implemented program that stores processor-executable instructions for causing the program to execute the program; A data-readable storage medium is provided.
[0031] In some embodiments, the RNA expression data comprises a first set of genes associated with a first cell type. and a first nonlinear regression model for the first RNA expression data. as input to generate a first value for the estimated proportion of RNA from a first cell type. The first submodel is constructed as follows: For the estimated proportion of RNA from the first cell type, use the first values for σ and σ as inputs: and a second sub-model configured to generate a second value.
[0032] In some embodiments, the RNA expression data comprises a second set of genes associated with a second cell type. The second RNA expression data includes second RNA expression data related to the first RNA expression data in Table 2. Expression of at least 10 genes selected from a group of genes for two cell types The second cell type includes B cells, CD4+ T cells, CD8+ T cells, endothelial cells, and fibroblasts. , lymphocytes, macrophages, monocytes, NK cells, neutrophils, and T cells. In some embodiments, determining the second cellular proportion for the second cell type comprises: The second RNA expression data is used to determine a second cellular proportion for a second cell type. Processing with a nonlinear regression model is included.
[0033] In some embodiments, the RNA expression data includes a plurality of genes associated with each of the plurality of cell types. the plurality of gene sets includes RNA expression data associated with the first gene set and the second gene set; and a plurality of cell types, including one cell type. Some embodiments include a plurality of gene sets associated with the plurality of gene sets. Determining multiple cellular constituent ratios for multiple cell types using RNA expression data In some embodiments, the method further comprises the steps of: The step of determining the composition ratio of a plurality of cells includes determining at least one composition ratio for each of the plurality of cell types. In this section, we used each nonlinear regression model to determine the cellular composition ratio for each cell type. by processing RNA expression data related to sets of genes associated with cell types. The method includes determining the proportion of each cell type in the cell population.
[0034] In some embodiments, the first nonlinear regression model comprises a random forest regression model. nothing.
[0035] In some embodiments, the first nonlinear regression model is a neural network regression model Includes:
[0036] In some embodiments, the first nonlinear regression model is a support vector machine regression model Includes:
[0037] In some embodiments, the first nonlinear regression model simulates, at least in part, The method is trained by generating training data including RNA expression data. In embodiments, generating training data comprises collecting RNA expression data from one or more biological samples. obtaining a set of RNA expression data, the set of RNA expression data comprising microenvironment cell RNA expression data and and malignant cell RNA expression data, and using the microenvironment cell RNA expression data, staining generating simulated microenvironment cell RNA expression data; and extracting the malignant cell RNA expression data. generating simulated malignant cell RNA expression data using the simulated We combined the RNA expression data of simulated microenvironment cells with the RNA expression data of simulated malignant cells. and generating at least a portion of the simulated RNA expression data. .
[0038] Some embodiments provide an RNA expression profile for the first cell type and a method for detecting the presence of a nucleotide sequence in the first cell type. determining a malignant tumor expression profile using the first cellular composition ratio in the include.
[0039] In some embodiments, the first RNA expression data is selected from the group of genes in Table 2. The data includes expression data for at least 25 genes selected.
[0040] In some embodiments, the first RNA expression data is selected from the group of genes in Table 2. The data includes expression data for at least 50 selected genes.
[0041] In some embodiments, the first RNA expression data is selected from the group of genes in Table 2. The data includes expression data for at least 100 genes selected.
[0042] In some embodiments, the first nonlinear regression model is based on the simulated RNA expression data. obtaining training data comprising simulated RNA expression data for a first cell type; and second RNA expression data for a first set of genes associated with the first cell. training a first nonlinear regression model to estimate the proportion of RNA from a cell type, The training step comprises training a first nonlinear regression model using the first nonlinear regression model and the second RNA expression data. generating an estimated ratio of RNA from the first cell type, as well as using the estimated ratio of RNA from the second cell type; and updating parameters of the first nonlinear regression model. So they are trained.
[0043] Some embodiments utilize at least one computer hardware processor obtaining training data comprising simulated RNA expression data, The collected RNA expression data includes a first RNA expression vector for a first gene associated with a first cell type. the data and a second RN for a second gene associated with a second cell type different from the first cell type. A method for estimating the proportion of RNA from each of one or more cell types, comprising: training a plurality of nonlinear regression models to calculate the first A nonlinear regression model to estimate the proportion of RNA from one cell type and from a second cell type Multiple nonlinear regression models were used, including a second nonlinear regression model to estimate the RNA ratio. The training step comprises, at least in part, training a first nonlinear regression model and a first RNA expression data. generating an estimated ratio of RNA from the first cell type using Update the parameters of the first nonlinear regression model using the estimated proportion of RNA in training a first nonlinear regression model by the steps of: and outputting the trained plurality of nonlinear regression models, including the first nonlinear regression model and the second nonlinear regression model. and a method comprising the steps of:
[0044] Some embodiments include at least one computer hardware processor and at least one and when executed by one computer hardware processor, Training a computer hardware processor with simulated RNA expression data obtaining data, the simulated RNA expression data being associated with the first cell type; and a second cell type different from the first cell type. and one or more respective RNA expression data for related second genes. training multiple nonlinear regression models to estimate the proportions of RNA from different cell types; Thus, multiple nonlinear regression models are used to estimate the proportion of RNA from a first cell type. A nonlinear regression model and a second nonlinear regression model to estimate the proportion of RNA from a second cell type. and training a plurality of nonlinear regression models includes, at least in part, Using the linear regression model and the first RNA expression data, estimate the proportion of RNA from the first cell type. generating a first nonlinear regression model using the estimated proportions of RNA from the first cell type; training a first nonlinear regression model by updating the parameters of the model; a trained nonlinear regression model including a first nonlinear regression model and a second nonlinear regression model; and outputting the plurality of nonlinear regression models. and at least one non-transitory computer-readable storage medium for provide.
[0045] Some embodiments are implemented by at least one computer hardware processor. When executed, the simulated obtaining training data comprising simulated RNA expression data, The data includes first RNA expression data for a first gene associated with a first cell type and a first cell type. and second RNA expression data for a second gene associated with a second cell type distinct from the cell type. , and performing multiple nonlinear iterations to estimate the proportion of RNA from each of one or more cell types. training a regression model, wherein the plurality of nonlinear regression models are A first nonlinear regression model for estimating the proportion of RNA from a second cell type was used. and training a plurality of nonlinear regression models includes a second nonlinear regression model for training the plurality of nonlinear regression models. A first cell is analyzed using, at least in part, the first nonlinear regression model and the first RNA expression data. generating an estimated proportion of RNA from the first cell type, and using the estimated proportion of RNA from the first cell type updating parameters of the first nonlinear regression model using the training a nonlinear regression model; and and outputting a plurality of trained nonlinear regression models, including a linear regression model. at least one non-transitory computer-readable memory device storing processor-executable instructions; Provide a storage medium.
[0046] In some embodiments, obtaining training data comprises obtaining RNA expression data from one or more biological samples. obtaining a set of RNA expression data, the set of RNA expression data being microenvironment cell RNA expression data; and based on the microenvironment cell RNA expression data, Obtaining simulated microenvironment cell RNA expression data and analyzing the malignant cell RNA expression data. obtaining simulated malignant cell RNA expression data based on the simulated malignant cell RNA expression data; We combined the RNA expression data of simulated malignant cells with that of the microenvironment cells. and generating at least a portion of the simulated RNA expression data. generating at least a portion of the simulated RNA expression data.
[0047] Some embodiments may further comprise: The method further includes adding noise to the expression data.
[0048] In some embodiments, the noise is at least one of Poisson noise or Gaussian noise. Includes one.
[0049] In some embodiments, generating simulated microenvironment cellular RNA expression data Using the first portion of the microenvironment cell RNA expression data for the first microenvironment cell type, , generating a first RNA expression profile.
[0050] In some embodiments, the first portion of the microenvironment cell RNA expression data comprises a first microenvironment cell It contains RNA expression data from multiple subtypes of the disease.
[0051] In some embodiments, generating a first RNA expression profile comprises generating a first microenvironment cell Resample the first portion of the microenvironment cell RNA expression data using multiple subtypes of cell types The method includes a ringing step.
[0052] In some embodiments, the first portion of the microenvironment cellular RNA expression data comprises RNA expression data from a plurality of samples. A Expression data included.
[0053] In some embodiments, generating a first RNA expression profile comprises generating a first RNA expression profile from a plurality of samples. The first part of the microenvironment cellular RNA expression data is generated by taking as input several samples. The method includes a step of resampling.
[0054] In some embodiments, generating simulated microenvironment cellular RNA expression data for a second microenvironment cell type, using the second portion of the microenvironment cell RNA expression data generating a second RNA expression profile; and comparing the first RNA expression profile with the second RNA expression profile. The small amount of simulated microenvironment cellular RNA expression data was combined with the current profiles. and generating at least some of the
[0055] In some embodiments, the first and second RNA expression profiles are combined. Generate at least some of the RNA expression data for the simulated microenvironment cells determining a weighted sum of the first RNA expression profile and the second RNA expression profile. The process includes:
[0056] In some embodiments, the malignant cell RNA expression data comprises RNA expression data from a plurality of malignant cell samples. Includes data.
[0057] In some embodiments, generating the simulated malignant cell RNA expression data comprises: The method includes combining RNA expression data from multiple malignant cell samples.
[0058] In some embodiments, generating the simulated malignant cell RNA expression data comprises: Adding noise to the simulated malignant cell RNA expression data.
[0059] In some embodiments, the coefficients of the weighted sum are determined using the output of a previously trained nonlinear regression model. It is determined using
[0060] In some embodiments, the first RNA expression data is for a first cell type in Table 2. The present invention includes expression data for at least 10 genes selected from the group of genes in the present invention.
[0061] In some embodiments, the second RNA expression data is for a second cell type in Table 2. The present invention includes expression data for at least 10 genes selected from the group of genes in the present invention.
[0062] In some embodiments, the first cell type and the second cell type are B cells, CD4+ T cells, CD8+ T cells, or the like. alveoli, endothelial cells, fibroblasts, lymphocytes, macrophages, monocytes, NK cells, neutrophils, and T The cell is selected from the group consisting of:
[0063] In some embodiments, the simulated RNA expression data is a first RNA expression data set associated with a first cell type. The second contains RNA expression data for one gene.
[0064] In some embodiments, a first nonlinear regression model of the one or more nonlinear regression models is a Using the RNA expression data of one cell type as input, calculate the estimated proportion of RNA from the first cell type. a first sub-model configured to generate values of Calculate the R from the first cell type using the first values for the estimated ratio of RNA from the first cell type and R as inputs. and a second sub-model configured to generate a second value for the estimated proportion of NA.
[0065] In some embodiments, the second sub-model comprises a plurality of cell types other than the first cell type. For the estimated proportion of RNA from the first cell type, use the estimated proportion of RNA from The second value is further configured to generate a second value of
[0066] Some embodiments utilize at least one computer hardware processor obtaining expression data for a biological sample, the biological sample comprising: previously obtained from a subject suspected of having or at risk of having cancer, obtaining a plurality of expression profiles for a corresponding plurality of cell types, Each of the files contains one or more genes associated with each cell type from a plurality of cell types. and, at least in part, combining the expression data and the plurality of expression data. A piecewise continuous error function between the current profile and the plurality of cells is optimized. and determining a plurality of cellular composition ratios for the type.
[0067] Some embodiments include at least one computer hardware processor and at least one and when executed by one computer hardware processor, obtaining expression data for the biological sample in a computer hardware processor; Therefore, the biological sample may be collected from individuals who have, are suspected of having, or are at risk of having cancer. The process involves multiple expression profiles for multiple cell types that have been previously obtained from the subject. obtaining a file, each of the expression profiles comprising a respective one of the expression profiles from a plurality of cell types; and including respective expression data from one or more genes associated with the cell type; At least in part, a piecewise continuous error function between the expression data and multiple expression profiles By optimizing the ratio of the cell types, multiple cell composition ratios can be determined. A process of determining at least one computer-readable medium storing processor-executable instructions for causing the and a storage medium capable of storing the data.
[0068] Some embodiments are implemented by at least one computer hardware processor. When executed, the at least one computer hardware processor is configured to perform a and obtaining expression data from the biological sample, the biological sample being selected from a group of individuals having or suspected of having cancer. or a corresponding plurality of samples previously obtained from a subject at risk of having cancer. Obtaining a plurality of expression profiles for the cell type, each of the expression profiles comprising: This involves the identification of each gene from one or more genes associated with each cell type from multiple cell types. and, at least in part, the expression data and the plurality of expression profiles. A piecewise continuous error function between determining a cellular composition ratio; Another computer-readable storage medium is provided.
[0069] In some embodiments, the expression data is RNA expression data and the plurality of expression profiles is NA expression profile.
[0070] In some embodiments, determining a plurality of cellular composition ratios for a plurality of cell types includes determining an error. determining a weighted sum of the difference values, the error values being determined using a piecewise continuous error function; can be.
[0071] In some embodiments, determining a plurality of cellular composition ratios for a plurality of cell types includes determining an error. Minimizing a weighted sum of the difference values.
[0072] In some embodiments, the one or more genes are selected from fewer than 5000 genes in Table 2. It is a gene that consists of at least two genes.
[0073] Some embodiments provide a plurality of expression profiles for a corresponding plurality of cell types and a plurality of The method further comprises determining a malignant tumor expression profile using the cellular constituent ratios of the above. [Brief explanation of the drawings]
[0074] [Figure 1A] FIG. 1A is a diagram depicting a system for determining cellular composition ratios based on expression data, according to some embodiments of the technology described herein. [Figure 1B] FIG. 1B is a diagrammatic example for determining various cellular constituent ratios of various cell types and cell subtypes using nonlinear regression models of each respective cell type and cell subtype, according to some embodiments of the technology described herein. [Figure 1C] FIG. 1C shows a t-SNE visualization depicting an exemplary cell population including malignant and microenvironment cells, according to some embodiments of the techniques described herein. [Figure 1D] FIG. 1D shows a t-SNE visualization depicting an exemplary malignant cell population in accordance with some embodiments of the techniques described herein. [Figure 1E] FIG. 1E is a chart depicting exemplary gene expression in various cells according to some embodiments of the technology described herein. [Figure 1F] FIG. 1F is a chart depicting exemplary correlations between genes and selected cell proportions in a sample mixture of diverse cell types, according to some embodiments of the technology described herein. [Figure 1G-1] FIG. 1G is a chart depicting exemplary gene expression of tumor cell lines, according to some embodiments of the technology described herein. [Figure 1G-2] FIG. 1G is a chart depicting exemplary gene expression of tumor cell lines, according to some embodiments of the technology described herein. [Figure 1G-3]FIG. 1G is a chart depicting exemplary gene expression of tumor cell lines, according to some embodiments of the technology described herein. [Figure 1G-4] FIG. 1G is a chart depicting exemplary gene expression of tumor cell lines, according to some embodiments of the technology described herein. [Figure 2A] FIG. 2A is a flowchart depicting an exemplary non-linear method for determining cellular constituent ratios based on expression data, according to some embodiments of the technology described herein. [Figure 2B] FIG. 2B is a flowchart illustrating an example implementation of a method 200 for determining cellular composition ratios based on expression data, according to some embodiments of the technology described herein. [Figure 2C] FIG. 2C is a flowchart illustrating an example implementation of act 216a of method 200 according to some embodiments of the technology described herein. [Figure 3A] FIG. 3A depicts the use of machine learning methods to determine RNA ratios based on RNA expression data, according to some embodiments of the technology described herein. [Figure 3B] FIG. 3B is a diagram depicting the use of a nonlinear regression model including sub-models to determine RNA ratios based on RNA expression data, according to some embodiments of the technology described herein. [Figure 3C] FIG. 3C depicts a method for determining cellular constituent ratios based on RNA ratios, according to some embodiments of the technology described herein. [Figure 3D] FIG. 3D depicts an example method for determining a malignant tumor expression profile based on cellular composition, according to some embodiments of the technology described herein. [Figure 4] FIG. 4 is a flowchart depicting an exemplary method for training one or more nonlinear regression models for determining cellular constituent proportions based on RNA expression data, according to some embodiments of the technology described herein. [Figure 5A]FIG. 5A is a diagram depicting an example method for training one or more machine learning models including validation and multi-stage training, in accordance with some embodiments of the techniques described herein. [Figure 5B-1] FIG. 5B is a diagram depicting an example method for training one or more machine learning models including validation and multi-stage training, in accordance with some embodiments of the techniques described herein. [Figure 5B-2] FIG. 5B is a diagram depicting an example method for training one or more machine learning models including validation and multi-stage training, in accordance with some embodiments of the techniques described herein. [Figure 6A] FIG. 6A depicts an exemplary method for training one or more nonlinear regression models that includes generating simulated RNA expression data, according to some embodiments of the technology described herein. [Figure 6B] FIG. 6B is an exemplary diagram for generating artificial mixtures of RNA expression data to mimic real tissues, according to some embodiments of the techniques described herein. [Figure 6C] FIG. 6C is an illustrative diagram for generating and using artificial mixtures to train cell-type models, according to some embodiments of the techniques described herein. [Figure 6D] FIG. 6D is an exemplary illustration for generating specific artificial mixtures for training particular cell type / subtype models, according to some embodiments of the techniques described herein. [Figure 6E] FIG. 6E is an exemplary illustration for generating specific artificial mixtures for training specific cell type / subtype models, according to some embodiments of the techniques described herein. [Figure 6F] FIG. 6F is an exemplary diagram illustrating a technique for processing a dataset and generating an artificial mixture, according to some embodiments of the techniques described herein. [Figure 7A]FIG. 7A is a chart comparing simulated RNA expression data with RNA expression data from a biological sample, according to some embodiments of the technology described herein. [Figure 7B] FIG. 7B is a chart depicting exemplary cellular constituent ratios predicted according to a deconvolution technique developed by the inventors and the corresponding true cellular constituent ratios, according to some embodiments of the techniques described herein. [Figure 7C] FIG. 7C is a chart comparing an exemplary predictive accuracy of a deconvolution technique developed by the inventors against the predictive accuracy of an alternative algorithm, according to some embodiments of the techniques described herein. [Figure 7D] FIG. 7D is a chart comparing an exemplary predictive accuracy of a deconvolution technique developed by the inventors against the predictive accuracy of an alternative algorithm, according to some embodiments of the techniques described herein. [Figure 7E] FIG. 7E is a diagram depicting the expression of four selected genes in normal tissues, immune cell types, and cancerous tissues, according to some embodiments of the technology described herein. [Figure 7F] FIG. 7F is a chart depicting exemplary predictive specificity of a deconvolution technique developed by the present inventors, in accordance with some embodiments of the techniques described herein. [Figure 7G] FIG. 7G is a chart comparing exemplary non-specificity scores of a deconvolution technique developed by the inventors against non-specificity scores of alternative algorithms, according to some embodiments of the techniques described herein. [Figure 8] FIG. 8 is a flowchart depicting an exemplary linear method for determining cellular constituent ratios based on RNA expression data, according to some embodiments of the technology described herein. [Figure 9A-1] FIG. 9A depicts exemplary RNA expression profiles and global RNA expression data according to some embodiments of the technology described herein. [Figure 9A-2] FIG. 9A depicts exemplary RNA expression profiles and global RNA expression data according to some embodiments of the technology described herein. [Figure 9A-3] FIG. 9A depicts exemplary RNA expression profiles and global RNA expression data according to some embodiments of the technology described herein. [Figure 9A-4] FIG. 9A depicts exemplary RNA expression profiles and global RNA expression data according to some embodiments of the technology described herein. [Figure 9B] FIG. 9B is a diagram depicting an example piecewise continuous error function in accordance with some embodiments of the techniques described herein. [Figure 10] FIG. 10 depicts an illustrative implementation of a computer system that may be used in connection with some embodiments of the technology described herein. [Figure 11] FIG. 11 is a block diagram of an illustrative environment in which one or more embodiments of the techniques described herein may be implemented. [Figure 12A] FIG. 12A is a chart and graphs representing the analysis and results from an experiment to establish RNA transcript normalization and analyze sequencing technical noise, as described in connection with Example 1. [Figure 12B] FIG. 12B is a chart and graphs representing the analysis and results from an experiment establishing RNA transcript normalization and analyzing sequencing technical noise, as described in connection with Example 1. [Figure 12C] FIG. 12C is a chart and graphs representing the analysis and results from an experiment to establish RNA transcript normalization and analyze sequencing technical noise, as described in connection with Example 1. [Figure 12D] FIG. 12D is a chart and graphs representing the analysis and results from an experiment establishing RNA transcript normalization and analyzing sequencing technical noise, as described in connection with Example 1. [Figure 12E]FIG. 12E is a chart and graphs representing the analysis and results from an experiment to establish RNA transcript normalization and analyze sequencing technical noise, as described in connection with Example 1. [Figure 12F] FIG. 12F is a chart and graphs representing the analysis and results from an experiment to establish RNA transcript normalization and analyze sequencing technical noise, as described in connection with Example 1. [Figure 12G] FIG. 12G is a chart and graphs representing the analysis and results from an experiment to establish RNA transcript normalization and analyze sequencing technical noise, as described in connection with Example 1. [Figure 12H] FIG. 12H is a chart and graphs representing the analysis and results from an experiment to establish RNA transcript normalization and analyze sequencing technical noise, as described in connection with Example 1. [Figure 12I] FIG. 12I is a chart and graphs representing the analysis and results from an experiment to establish RNA transcript normalization and analyze sequencing technical noise, as described in connection with Example 1. [Figure 12J] FIG. 12J is a chart and graphs representing the analysis and results from an experiment establishing RNA transcript normalization and analyzing sequencing technical noise, as described in connection with Example 1. [Figure 12K] FIG. 12K is a chart and graphs representing the analysis and results from an experiment establishing RNA transcript normalization and analyzing sequencing technical noise, as described in connection with Example 1. [Figure 13A] FIG. 13A is a chart and graphs representing the analysis and results from an experiment to deconvolute RNA-seq of multiple normal and cancer tissues, as described in connection with Example 2. [Figure 13B] FIG. 13B is a chart and graphs representing the analysis and results from an experiment to deconvolute RNA-seq of multiple normal and cancer tissues, as described in connection with Example 2. [Figure 13C]FIG. 13C is a chart and graphs representing the analysis and results from an experiment to deconvolute RNA-seq of multiple normal and cancer tissues, as described in connection with Example 2. [Figure 13D-1] FIG. 13D is a chart and graphs representing the analysis and results from an experiment to deconvolute RNA-seq of multiple normal and cancer tissues, as described in connection with Example 2. [Figure 13D-2] FIG. 13D is a chart and graphs representing the analysis and results from an experiment to deconvolute RNA-seq of multiple normal and cancer tissues, as described in connection with Example 2. [Figure 13E-1] FIG. 13E is a chart and graphs representing the analysis and results from an experiment to deconvolute RNA-seq of multiple normal and cancer tissues, as described in connection with Example 2. [Figure 13E-2] FIG. 13E is a chart and graphs representing the analysis and results from an experiment to deconvolute RNA-seq of multiple normal and cancer tissues, as described in connection with Example 2. [Figure 13F] FIG. 13F is a chart and graphs representing the analysis and results from an experiment to deconvolute RNA-seq of multiple normal and cancer tissues, as described in connection with Example 2. [Figure 13G] FIG. 13G is a chart and graphs representing the analysis and results from an experiment to deconvolute RNA-seq of multiple normal and cancer tissues, as described in connection with Example 2. [Figure 13H] FIG. 13H is a chart and graphs depicting the analysis and results from an experiment to deconvolute RNA-seq of multiple normal and cancer tissues, as described in connection with Example 2. [Figure 13I] FIG. 13I is a chart and graphs representing the analysis and results from an experiment to deconvolute RNA-seq of multiple normal and cancer tissues, as described in connection with Example 2. [Figure 13J]FIG. 13J is a chart and graphs representing the analysis and results from an experiment to deconvolute RNA-seq of multiple normal and cancer tissues, as described in connection with Example 2. [Figure 14A] FIG. 14A is a chart and graphs representing the analysis and results from an experiment to deconvolve single-cell RNA-seq data and bulk RNA-seq of blood, as described in connection with Example 3. [Figure 14B] FIG. 14B is a chart and graphs representing the analysis and results from an experiment to deconvolute single-cell RNA-seq data and bulk RNA-seq of blood, as described in connection with Example 3. [Figure 14C] FIG. 14C is a chart and graphs representing the analysis and results from an experiment to deconvolve single-cell RNA-seq data and bulk RNA-seq of blood, as described in connection with Example 3. [Figure 14D] FIG. 14D is a chart and graphs representing the analysis and results from an experiment to deconvolute single-cell RNA-seq data and bulk RNA-seq of blood, as described in connection with Example 3. [Figure 14E] FIG. 14E is a chart and graphs representing the analysis and results from an experiment to deconvolve single-cell RNA-seq data and bulk RNA-seq of blood, as described in connection with Example 3. [Figure 14F] FIG. 14F is a chart and graphs representing the analysis and results from an experiment to deconvolve single-cell RNA-seq data and bulk RNA-seq of blood, as described in connection with Example 3. [Figure 14G] FIG. 14G is a chart and graphs representing the analysis and results from an experiment to deconvolve single-cell RNA-seq data and bulk RNA-seq of blood, as described in connection with Example 3. [Figure 15A]FIG. 15A is a chart and graphs representing the analysis and results from an experiment deconvolving several different cancer tissues, as described in connection with Example 4. [Figure 15B] FIG. 15B is a chart and graphs representing the analysis and results from an experiment deconvolving several different cancer tissues, as described in connection with Example 4. [Figure 15C] FIG. 15C is a chart and graphs representing the analysis and results from an experiment deconvolving several different cancer tissues, as described in connection with Example 4. [Figure 15D] FIG. 15D is a chart and graphs representing the analysis and results from an experiment deconvolving several different cancer tissues, as described in connection with Example 4. [Figure 15E] FIG. 15E is a chart and graphs representing the analysis and results from an experiment deconvolving several different cancer tissues, as described in connection with Example 4. [Figure 15F] FIG. 15F is a chart and graphs representing the analysis and results from an experiment deconvolving several different cancer tissues, as described in connection with Example 4. [Figure 15G] FIG. 15G is a chart and graphs representing the analysis and results from an experiment deconvolving several different cancer tissues, as described in connection with Example 4. [Figure 15H] FIG. 15H is a chart and graphs representing the analysis and results from an experiment deconvolving several different cancer tissues, as described in connection with Example 4. [Figure 15I] FIG. 15I is a chart and graphs representing the analysis and results from an experiment deconvolving several different cancer tissues, as described in connection with Example 4. DETAILED DESCRIPTION OF THE INVENTION
[0075] The present inventors have investigated the RNA expression data (e.g., biological samples) using sequencing techniques, e.g., Based on the data collected by RNA sequencing, The cellular composition (e.g., specific components) of a sample (e.g., a sample from a tumor or other diseased tissue) In some embodiments, we have developed machine learning methods to determine the proportion of each type of cell. The step of determining the cell composition ratio for one or more cell types may include one or more nonlinear loops. The method may include using a regression model to estimate the respective cellular constituent proportions for the cell types. Linear regression models can be used in the methods described herein, e.g., to assess the efficacy and safety of various malignant and / or minimally invasive cancers. Combining RNA expression data for environmental cell types and / or The method includes the steps of: It can be trained using simulated RNA expression data, which can be generated by .
[0076] The inventors have demonstrated that the tumor microenvironment (TME) plays a key role in disease progression (e.g., whether the tumor is eradicated or metastasized). Recognize and understand that HIV can play an important role in determining whether or not a patient is resistant to treatment and in determining treatment response / resistance. For example, as recognized and understood by the inventors, the immune and non-immune components of the TME The components are involved in cell-cell contact and a variety of different molecular signals, such as growth factors and cytokines. Furthermore, the present inventors have investigated the role of TME in the survival, maintenance, growth and development of tumors. mediates tumor survival by controlling the host immune system, resulting in tumor immunosurveillance. Therefore, the inventors have developed an understanding of the quantity and functionality of TME components. is essential to cancer research and is important for understanding treatment and its clinical impact. However, despite the importance of understanding TME components, existing cancer research This study focuses on a limited set of cellular components of the TME due to limitations of conventional methods for analyzing TME components. We have focused only on elements such as immunohistochemistry, flow cytometry, and CyTO. Techniques such as F rely on the availability of target-specific antibodies and unique tags, e.g., fluorescent dyes. There are limits to what can be done.
[0077] The present inventors further demonstrated that the method can simultaneously provide information on tens of thousands of genes in a biological sample. Bulk RNA sequencing (RNA-seq) allows for signal capture that represents the combined contributions of multiple cell types. However, the present inventors recognize and understand that this method allows for the detection of all RNAs of this type. A. Expression data do not provide information about the origin of individual RNA molecules and therefore bulk RNA- Many challenges remain in determining the cellular composition (e.g., cellular composition ratios) of the TME from seq. We recognize that the process of determining cellular constituent ratios from RNA expression data is described herein. This is sometimes called "deconvolution."
[0078] We believe that one of the key issues in cellular deconvolution is the identification of tumors and their microenvironments. Many genes can be expressed simultaneously by several types of cells present in the environment This is true for closely related cell types (e.g., T cells). specific cell types that may be considered subtypes (e.g., CD4+ and CD8+ T cell subtypes, etc.) This poses a particular challenge to identification, as genetic markers between closely related cell types are often lacking. In some embodiments, the cells A type can be thought of as a population of cells with a distinct expression profile. For example, CD4+ T cells CD8+ T cells, and NK cells were analyzed for metabolic, signaling, and surface markers. The expression of a significant number of structural and regulatory genes, including It appears that globulin is specifically expressed by mature dendritic cells and macrophages. Therefore, we believe that RNA expression data are consistent. Recognize and understand that the present invention may include both specific marker genes and genes associated with cell lineage. In addition, the ratio of markers to lineage-specific gene expression provides information about cell subtypes. The inventors recognize that the information may or may not be available. (For example, the CD4 / CD3D gene ratio may be a marker for CD4+ T cells, but CD3D is a marker for hemoglobin.) (It is not a specific marker of any subtype of T cell.) Different types of cells may be involved, even if they Even if these factors are closely related, their effects on tumor pathogenesis may differ significantly. However, we have found that it is not possible to distinguish cell populations, even among closely related cell types. I still think it is important.
[0079] Another challenge of cellular deconvolution that we recognize is the number of cells and their The difficulty in distinguishing between conditions, e.g., cell type-specific or semi-specific genetic Expression of the gene may vary depending on the activation state of the cell type, or may be expressed by subtypes of the cell type. Multiple studies can determine the sequence of similar cell subtypes. However, they may be captured in different biological states. The authors argue that variation in biological conditions is important for deriving accurate estimates of cellular composition. recognize and understand that they can play a vital role in
[0080] Furthermore, the inventors have demonstrated that the tumor microenvironment accounts for a relatively small proportion of the tumor as a whole. We recognize and understand that small cell populations from bulk RNA-seq data may be difficult to extract. Identification of groups can be particularly challenging due to low signal-to-noise ratios. They argue that even small cell populations can have a significant impact on response to therapy. We acknowledge that identifying changes in small cell populations (e.g., NK cells) remains important. Furthermore, the value of RNA expression of a gene varies depending on the specific measurement technique, library, and - Preparation protocols and RNA enrichment methods (e.g., total RNA-seq (REF), polyA enrichment (REF), exon enrichment (EXO)) The inventors recognize and understand that the use of 3' scRNA-seq (REF) can be highly dependent on 3' scRNA capture. Even with the use of techniques such as single-cell RNA-seq (scRNA-seq), the coverage of such techniques is limited. In this case, it is generally not possible to extract marker genes useful for identifying cell types.
[0081] Therefore, the present inventors have developed an accurate and robust cellular We recognize the need for deconvolution techniques. We have applied machine learning techniques to estimate cellular composition based on data (e.g., RNA expression data). In some embodiments, a biological sample from a subject is collected and analyzed using a novel system and method. obtaining expression data (e.g., bulk RNA-seq data) for one or more cells; types (e.g., B cells, CD4+ T cells, CD8+ T cells, endothelial cells, fibroblasts, lymphocytes, macrophages) The method includes determining the cell composition ratios of the following cells: phage, monocytes, NK cells, neutrophils, and T cells. The cellular composition ratio is determined by the ratio of each specific cellular component in a biological sample. According to some embodiments, the estimated proportion of cells of a particular cell type may be displayed. The step of determining the proportion of components is performed by determining the set of genes associated with that cell type (e.g., one or more marker genes, which may be type-specific or semi-specific genes obtaining expression data for the specific cell types and determining the cellular composition of the specific cell types; According to some embodiments, the method may include processing the current data with a non-linear regression model. This process can be repeated multiple times to achieve deconvolution across multiple cell types. Each of the cell types (which may include subtypes of cell types, as described herein) 7. As described herein, these techniques represent a significant improvement over the prior art.
[0082] In some embodiments, the machine learning method used to determine the cellular composition ratios Multiple nonlinear models, each trained to determine the cell composition ratio for each specific cell type, are used. In some embodiments, the non-linear regression model may include using a multi-dimensional regression model. parameters (e.g., thousands, tens of thousands, hundreds of thousands, at least a million, millions, tens of millions, or The nonlinear regression model may have a number of parameters (hundreds of millions of parameters), and the step of training the nonlinear regression model may include Such parameters can be calculated by computer from simulated expression data. In some embodiments, the simulated training data may include estimating the value of The generating step may involve generating multiple training sets (e.g., For example, at least 25,000, at least 50,000, at least 100,000, at least 150 ,000, at least 200,000, at least 500,000, etc. In some embodiments, multiple nonlinear regression models are used to identify multiple cell types (e.g., at least five, at least 10, at least 20, at least 30, at least 40, etc.) This can be trained.
[0083] The techniques described herein and developed by the inventors use machine learning techniques This allows for a more robust computational approach to determining cell composition than conventional methods. This approach also provides significant improvements in performance, accuracy, and efficiency. For example, Figures 7C and 7D show the performance improvement over the conventional approach. Compared with the method, the nonlinear deconvolution method developed by the inventors (e.g. In the "Kassandra" model, even in the presence of overexpression noise in cancer cells, , which has been shown to provide more accurate predictions of cellular composition for different cell types ( For example, as shown in Figure 7D). As a result, the techniques described herein allow for the This results in a general improvement in cellular structure and, specifically, the techniques described herein. This study provides an improved method for determining the composition ratios of cells (e.g., specifically for cell populations within the tumor microenvironment). This provides improvements to support clinical decision making and understanding of tumor pathogenesis.
[0084] For example, unlike conventional approaches, the machine learning approach described herein can Expression data relating to genes associated with the type (e.g., specific and / or semi-specific) By using it as input to a nonlinear regression model trained specifically for that subtype. This allows for successful identification of dependencies and interrelationships between genes in phenotypically closely related cell types. This allows accurate detection of cell subtypes even with similar expression patterns. (Figures 7A and 7B) We used training data that mimics the cellular complexity and diversity of tumor biopsies. By utilizing the unique characteristics of expression profiles and cell population markers, Thus, the nonlinear deconvolution techniques described herein are superior to previous algorithms. It is more robust and shows more consistent accuracy across different cell types / subtypes, provides significantly more accurate results than conventional methods for complex and noisy data (Fig. 7C, 7D, 13F, 15G). In the context of the tumor microenvironment (e.g., in the clinical setting of patients), These more accurate results could lead to improved cancer diagnosis and prognosis, as well as personalized treatment for patients. enabling more tailored treatment options.
[0085] One of the factors that contribute to the accuracy and robustness of the approach developed by the inventors is This aspect is particularly relevant for each respective cell type to determine the corresponding cellular composition ratio. For example, for a given cell type, the expression data may be: It may include expression data related to specific genes associated with that given cell type. In the form of at least as described herein with respect to Figures 1D-1E and Table 2 The expression data may include expression data relating to genes for a given cell type. As described herein, the step of identifying genes associated with a particular cell type can be carried out by identifying a particular Identifying genes that are expressed only or predominantly in cell types or subtypes to obtain data from multiple databases and / or using various sequencing techniques. This may include processing expression data from multiple samples, which may allow for any particular associated with a particular cell type, regardless of how the gene for that cell type is determined The use of expression data associated with specific genes is a novel method for detecting and analyzing cell deaths. Cellular deconvolution techniques allow us to determine which genes are expressed by which cell types. This allows for the exploitation of domain-specific knowledge about Contribute to success.
[0086] Another aspect of the approach developed by the inventors that contributes to its performance is the The architecture employed for both the training described in the paper and the use of nonlinear deconvolution techniques is For example, as described herein, in some embodiments, Estimate the cellular composition ratio for each cell type and / or subtype analyzed in the sample. To do this, separate nonlinear regression models are trained and used (e.g., at least Figure 3A (including as described herein with respect to the detection of cellular components in biological samples). This may allow for more accurate identification of cell types and / or subtypes (e.g., Figure 7A 7G). Furthermore, in some embodiments, the model architecture Hierarchical structures that may be used as part of the training and / or use of the machine learning techniques described herein. structure (e.g., as described herein, including with respect to at least FIG. 5A). For example, this model architecture includes multiple sub-models corresponding to multiple stages. In this case, the output of one or more previous sub-models (e.g., one or more The initial prediction of one or more cellular proportions for a cell type is used for subsequent sub-models. This allows the model to be used as part of the input for the model (e.g., in the second, third, etc. stages of training and / or use of By improving initial predictions (e.g., from the first stage of training and / or using the model), According to some embodiments, multiple cell types and / or or the output from the first sub-model across multiple models for a subtype is The hierarchical structure of the model can be used to provide input to subsequent submodels. For example, the first submodel prediction of cellular composition ratios for all cell types (e.g., provide as input to a second sub-model (for another cell type and / or subtype) This allows subsequent submodels (e.g., second submodels) to be modeled on cell types and / or or subtype interdependencies can be taken into account, thereby allowing for the analysis of various cell types and The present invention can provide a more accurate prediction of cellular composition across tumor types and / or subtypes.
[0087] Another advantage of the approach developed by the inventors is that, in some embodiments, The model is trained on data representing an artificial mixture of cell types, The training process is conducted by physically sampling and analyzing tumor samples. malignant cells and microenvironmental cells across a much larger range of diverse sample compositions than would otherwise be possible Consider the diverse and tissue-specific expression of This allows for the cellular deconvolution The effort and computational resources associated with training a nonlinear regression model for this purpose are significantly reduced. The artificial mixtures described in the specification are intended to mimic technical noise and to account for the wide biological variability. The data is obtained in a way that captures the data, and a machine learning model trained on this data is This improves our ability to identify biologically meaningful signals in the presence of such noise and variability. For example, as described herein, quantification of technical noise can be performed. A practical noise model has been developed and may be applied to artificial mixtures. The RNA expression data used to develop these artificial mixtures represent various biological states. These artificial mixtures are derived from multiple different samples across multiple cell populations. The objective is to determine the cellular composition ratio across various cell types in real tumor samples using a nonlinear regression model. Improve your ability to effectively estimate
[0088] As described herein below, including with respect to Figures 8 and 9A-9B, the present invention The method developed by
[14] includes an improved linear method for cellular deconvolution. As described herein, one aspect that contributes to the success of linear approaches is , the use of an error function developed by the inventors. As described herein, including Compared to traditional methods, such as squared distance, the piecewise continuous error function This takes into account genes that are strongly expressed in tumor cells. The accuracy of the deconvolution can be improved by using such an error function. The method developed by the present inventors reduces the error associated with the predicted cell composition ratio. can be accurately modeled (e.g., as described herein, including with respect to Figures 8 and 9A). (as published in [Publication ID]) provides improved results over conventional techniques.
[0089] The following describes the cellular deconvolution system and method developed by the present inventors. A more detailed description of various concepts and embodiments of the method described herein is provided. It should be appreciated that various aspects can be implemented in any of numerous ways. Examples of typical implementations are provided herein for illustrative purposes only. The various aspects described in the embodiments can be used alone or in any combination and are not to be construed as limiting the scope of the present invention. It is not limited to the combinations explicitly stated.
[0090] FIG. 1A depicts a system 100 for determining cellular composition 110. The illustrated system may be used in clinical or practical applications as described herein, including with respect to 11. It can be implemented in a laboratory setting.
[0091] As shown, the system 100 includes a biological sample 102, which may be, for example, a sample from a subject (e.g., , subjects with, suspected of having, or at risk of having cancer) The subject may have a genetic predisposition to cancer (e.g., a known genetic predisposition to cancer) or a tumor biopsy sample obtained by a genetic test. If you have one or more genetic mutations, or if you may have been exposed to a cancer-causing agent, In some cases, there may be a risk of having cancer. The biological sample 102 may be obtained by performing a biopsy on a patient. This can be obtained by obtaining a blood sample, saliva sample, or any other suitable biological sample from The biological sample 102 may have been previously obtained from the subject. Any process applied to the method (e.g., obtaining expression data from a biological sample) may be performed in vitro. The biological sample 102 may include diseased tissue (e.g., tumor) and / or healthy tissue. In some embodiments, the biological sample may be provided by a doctor, hospital, clinic, or other healthcare provider. In some embodiments, the source or preparation method of the biological sample may include " Some embodiments may include any of the embodiments described in the "Biological Sample" section. In this aspect, the subject matter may include any of the embodiments described in the "Subject Matter" section.
[0092] The system 100 further includes a sequencing platform 104 that can generate sequence information 106. In some embodiments, the sequencing platform 104 may include next-generation sequencing. sensing platforms (e.g., Illumina™, Roche™, Ion Torrent™) ), or any high-throughput or massively parallel sequencing platform. In some embodiments, the sequencing platform 104 may be any suitable sequence Sequencing device and / or any sequencing system including one or more devices In some embodiments, these methods may be automated, and some implementations may involve In some embodiments, there may be manual intervention. It may also be the result of next generation sequencing (e.g., Sanger sequencing). In embodiments, sample preparation may follow the manufacturer's protocol. , sample preparation may be performed using custom protocols or for research, diagnostic, prognostic, and / or It may also be other protocols for clinical purposes. In some embodiments, the protocol is an experimental In some embodiments, the origin or preparation of the sequence information may be unknown. good.
[0093] The sequence information 106 includes sequence data (e.g., , next-generation sequencing, Sanger sequencing, etc. In addition to the sequence of nucleotides, the information contained therein (e.g., information indicating origin, tissue type, etc.) and consider them as information that can be inferred or determined from sequence data. For example, in some embodiments, the nucleic acid may be primarily polyadenylated. In some embodiments, the RNA sequence information can be analyzed to determine whether The column information 106 may include information contained in a FASTA file, a description and / or a sequence of sequences contained in a FASTQ file. The quality score, the aligned position contained in the BAM file, and / or any preferred It may also include any other suitable information obtained from a suitable file.
[0094] In some embodiments, the sequence information 106 is generated using nucleic acids from a sample derived from the subject. A reference to a nucleic acid can refer to one or more nucleic acid molecules (e.g., multiple nucleic acid molecules). In some embodiments, the sequence information can be used to identify individuals who have or are suspected of having a disease. DNA and / or DNA from previously obtained biological samples of subjects with or at risk of having the disease may be sequence data representing the nucleotide sequence of the RNA. In some embodiments, the nucleic acid is a deoxyribonucleic acid (DNA). In some embodiments, the nucleic acid is a nucleic acid in which the entire genome is contained within the nucleic acid. In some embodiments, the nucleic acid is prepared to be present in a protein-coding region of the genome. The genome is processed so that only the region (e.g., exome) remains. When nucleic acids are prepared for sequencing, this is called whole exome sequencing (WES). Various methods for isolating exomes for sequencing are known in the art. For example, in solution-based isolation, tagged probes are used to identify target regions (e.g., exons). The hybridization of the nucleotide sequence from the nucleotide sequence of the target site (e.g., a non-binding oligonucleotide) is then performed. These tagged fragments can then be prepared and sequenced. It is possible.
[0095] In some embodiments, the nucleic acid is ribonucleic acid (RNA). The RNA sampled includes both coding and non-coding transcribed RNA found in the sample. When such RNA is used for sequencing, sequencing is performed using "total RNA." This is said to be done by sequencing the entire genome, and can also be called whole transcriptome sequencing. , nucleic acids so that coding RNA (e.g., mRNA) can be isolated and used for sequencing. This can be prepared by any means known in the art, for example, by by isolating or screening RNA for adenylated sequences This is sometimes called mRNA-Seq.
[0096] In some embodiments, the sequence information 106 may include raw DNA or RNA sequence data, DNA exome sequences, Data (e.g., whole exome sequencing (WES), DNA genome sequence data (e.g., whole genome sequencing (WES)) (from genome sequencing (WGS)), RNA expression data, gene expression data, bias-corrected gene expression data obtained from a sequencing platform 104. and / or data derived from data obtained from the sequencing platform 104. In some embodiments, the sequence data may include any other suitable type of sequence data, including sequence data. The origin or preparation of information 106 can be classified as "expression data," "obtaining RNA expression data," "alignment," or " "Sample and annotation", "Removal of non-coding transcripts", and "Conversion to TPM and gene assembly". " section.
[0097] Regardless of the sequence data obtained, the sequence information 106 is used to determine the cell composition ratio 110. The sequence information 106 can be processed using a computing device 108. For example, 1 running on a computing device 108 (e.g., as described herein with respect to FIG. 10 ). For example, sequence information can be processed by one or more software programs. 106 in combination with the machine learning-based approach of FIGS. 2A-2C or the Any other method described herein (e.g., at least the method shown in FIGS. 2A-2C and 3A-3C) The nonlinear deconvolution method described with respect to at least FIGS. 8 and 9A-9B The data can be processed according to the linear deconvolution method described above. In some embodiments, the computing device 108 may be a physician, clinician, researcher, patient, or other For example, the user may input the sequence information 106 into a computer. can be provided as input to the computer device 108 (e.g., upload a file) by sequencing), and / or using sequence information to specify processing or other methods to be performed. User input can be provided to
[0098] Regardless of how the sequence information 106 is processed, the result is one or more cellular constituents. As described herein, each cellular constituent ratio can be a percentage 110 of the total number of cells in the biological sample 102. In some embodiments, the biological sample may represent an estimated proportion of each specific type of cell present in the sample. The cell composition ratios are normalized to represent 100%. Cell types include, for example, B cells, plasma plasma B cells, non-plasma B cells, T cells, CD4+ T cells, CD8+ T cells, regulatory T cells, helper T cells Cells, CD8+PD1-high, CD8+PD1-low, NK cells, monocytes, macrophages, resting tumor-associated macrophages phages (TAM), M1-like or activated macrophages, neutrophils, endothelial cells, fibroblasts, and and / or any other suitable cell type. According to some embodiments, the cell type is one or T cells may include multiple subtypes, for example, CD4+ T cells, CD8+ T cells, and regulatory T cells. The cell composition ratio 110 can be used to determine the cell subtype. In addition to all ratios, it may also include ratios for cell types that are not subtypes of any other cell type. According to some embodiments, the cellular composition ratios include ratios for "other" cell types. This allows for the analysis of cells not considered in other cell populations (e.g., cells not explicitly included in the analysis). The estimated proportion of one or more types of cells not included in the total number of cells in the sera can be expressed.
[0099] FIG. 1B illustrates a method for detecting each of the individual cell types and cells according to some embodiments of the technology described herein. A nonlinear regression model for cell subtypes was used to identify different cell types and cell subtypes. FIG. 1 is an exemplary diagram for determining different cellular composition ratios for a sample.
[0100] As shown in this example, the first nonlinear regression model, Model A 126, is used to The cell composition ratio 128 for cell type A 122 is calculated using sequence information 124 related to cell type A 122. A second nonlinear regression model, Model B136, can be used to estimate the cell type. The cellular composition ratio 138 for B 132 is estimated using sequence information 134 associated with cell type B 136. It is possible.
[0101] For this example, cell type A 122 and cell type B 132 are different cell types. 122 may include B cells, while cell type B 132 may include T cells. As the described embodiment of the method is not limited in that respect, cell type A and / or cell type B may be It may be of any suitable cell type.
[0102] In some embodiments, sequence information 124 and sequence information 134 are from cell type A 122 and cell type B 134, respectively. In some embodiments, sequence information can be obtained for cell type B 132. For example, sequence information 1 24 can be linked to a first set of genes specific to cell type A122, while sequence information The information 134 can be linked to a second set of genes specific to cell type B 132. Methods for identifying cell type and / or subtype specific and / or semi-specific genes includes any of the embodiments described in the "Gene Selection and Specificity" section. obtain.
[0103] As shown in Figure 1B, different nonlinear regression models predicted the cellular structure for different cell types. For example, model A126 is used to determine the cell type A122. The model B136 was used to estimate the cell fraction for cell type B132. In some embodiments, the cell composition ratio 138 is used to estimate the cell composition ratio 138. Each of the models, including those described herein, is used for a particular cell type. It can be trained to estimate the cellular composition of
[0104] In some embodiments, the different cell types can include cell subtypes. As such, cell subtypes of close origin may be differentiated (e.g., from each other and / or from the cells into which they differentiate). As shown in Figure 1B, cell type B 1 32 includes subtype A 142 and subtype B 162. For example, cell type B 132 is a T cell , and subtype A 142 and subtype B 162 are T cell subtypes (e.g. , CD4+ T cells and CD8+ T cells).
[0105] In some embodiments, a third nonlinear regression model, model C146, is used to determine the subtypes. The cellular composition 148 for group A 142 can be estimated using the sequence information 144. A fourth nonlinear regression model, model D156, was used to estimate the cytoarchitecture for subtype B162. The proportion of copies 158 can be estimated using sequence information 164 .
[0106] In some embodiments, sequence information 144 and sequence information 164 are subtype A 142 and subtype B 164, respectively. In some embodiments, this can be obtained for subtype B 162. Obtain sequence information related to gene sets containing type-specific and / or semi-specific genes For example, the sequence information 144 may include a first set of genes specific to subtype A 142. The sequence information 164 may be related to the subtype B 144, while the sequence information 164 may be related to the second subtype B 144. May be associated with a set of genes specific to a particular cell type and / or subtype Methods for identifying specific and / or semi-specific genes include "gene selection and specificity" methods. Any of the embodiments described in relation to the section may be included.
[0107] FIG. 1C shows the expression of multiple genes for exemplary cell populations, including malignant cells and microenvironment cells. A t-SNE plot depicting the current data. As indicated in the legend, the t-SNE plot Cell types and / or subtypes depicted in include macrophages, M1 macrophages, , M2 macrophages, B cells, B cells (non-plasma), plasma B cells, T cells, CD8+ T cells, PD1+ CD8+ T cells, PD1- CD8+ T cells, CD4+ T cells, regulatory T cells, helper T cells, endothelial cells These include cells, monocytes, NK cells, fibroblasts, neutrophils, and tumor cells (e.g., cancer cells). Infectious cells include tumor cells or any other cells associated with disease and / or diseased tissue. Microenvironmental cells may be, for example, immune cells, skin cells, or tumor cells. Any non-tumor cells may be included, including any other cells.
[0108] The t-SNE plot in Figure 1C shows the t-SNE results obtained via any of the sequencing approaches described herein. A large number (e.g., at least 1000, at least 5000, at least 10000) of the antibodies may be collected from a biological sample. Delineating cell types / subtypes across RNA-seq samples (at least 10,000). In this study, we combined RNA-seq datasets, uniformly annotated them, and analyzed bioinformatics data. Recalculate the expression values using a bioinformatics method (e.g., convert the expression values into By using the method (recalculated using the RT-PCR method), accurate and comparable measurements of transcript expression can be obtained. For the example, RNA-seq data was collected from 12,450 selected samples (e.g., flow cytometry). (sorted by magnetic-assisted sorting of cells using beads) This was possible, and we were able to subdivide it into 19 cell populations of interest. After removal and quality inspection, the selected samples were classified into 10 major subpopulations as shown in Table 1 below. They were distributed among alveolar types and 19 cell subpopulations.
[0109] In the illustrated example, the t-SNE plot 140 shows the cell types / subtypes listed before quality control. The plot depicts RNA-seq samples (n=12450) from each type, while the t-SNE plot 150 depicts the quality control data. RNA-seq samples from listed cell types / subtypes after removing samples that did not pass the analysis The quality control methodology includes "data collection, analysis, and preprocessing." This may include any of the embodiments described in the section, or any other suitable quality control method. For example, in some embodiments, data from cells having an abnormal physiological state may be collected simultaneously. can be excluded (e.g., based on annotations provided with the data). For example, in some embodiments, phorbol myristate acetate / ionomycin All T cell samples with IL-1 activation and / or induced pluripotent stem cell-derived samples were excluded. In some embodiments, low isolation purity, sequencing quality parameters, or other organisms (e.g., Also remove samples with high contamination (e.g., organisms other than the primary organism under consideration) and / or low coverage. was done.
[0110] [Table 1]
[0111] As shown in plot 150, the cell population can include tumor cells 152. 152 is a t-SNE plot of cancer cell lines (n = 2166) color-coded by cancer type. Figure 1D As shown, the types of cancer include breast, colorectal, Cancer, head and neck cancer, kidney cancer, lung cancer, melanoma, pancreatic cancer, prostate cancer, stomach cancer, and / or Or any other type of cancer.
[0112] According to some embodiments, a portion of the sample of RNA expression data plotted in Figures 1C and 1D or all of the above, as described herein, including with respect to at least FIG. 1E. Use cell type / subtype specific and / or semi-specific genes as part of the selection In some embodiments, at least one of the methods described herein with respect to FIG. , some or all of the depicted samples of RNA expression data into artificial mixtures of RNA expression data. In some embodiments, the processes shown in FIGS. 1C and 1D can be used as part of generating The RNA expression data included in the data set and the RNA expression data plotted in Figures 1C and 1D are Data similar to the present data are derived from public datasets and included in the Gene Expression Omnibus (GE 0) and open source databases such as ArrayExpress. In an embodiment, RNA expression data similar to the RNA expression data plotted in Figures 1C and 1D are For example, a dataset containing the following can be used, each of which is shown in Table 1: The cells represented in Table 1 are represented by multiple samples from multiple datasets. A similar data set containing some or all of the types can be used.
[0113] FIG. 1E is a heat map depicting the expression 170 of exemplary genes for cell types 160. As shown, the vertical axis represents cell types160 and the horizontal axis represents gene expression170 per million. Each row in the heatmap represents one RNA-seq sample. As described in the literature, some genes appear to be specific to certain cell types. For example, as shown in the heatmap in Figure 1F, the selected genes 190 The ratio of RNAs in the corresponding sorted cell population 180 may be correlated with the ratio of RNAs in the corresponding sorted cell population 180. For example, As shown in the heatmap in Figure 1G, the 192 selected genes were significantly different from the 182 tumor cell lines. These genes may have limited or no expression.
[0114] As shown below, Table 2 shows the results for each of several cell types. Decombination techniques that are believed to be specific or semi-specific and / or that are described herein The paper specifies a set of genes that can be used in the solution approach.
[0115] Gene selection and specificity In some embodiments, the cellular deconvolution method developed by the present inventors uses specific gene expression data to determine the cellular composition of a particular cell type. For example, in some embodiments, the method may involve using only the As described herein, including those relating to specific and / or semi-specific antibodies to particular cell types, In some embodiments, only differential gene expression data can be used. Cell type (e.g., non-malignant cell type) specific and / or semi-specific genes are specifically expressed. In particular, genes that are highly expressed in malignant cells (e.g., cancer cell lines) (e.g., tumor cell-specific genes) In some embodiments, specific and / or specific to a particular cell type may be excluded. The step of selecting semi-specific genes includes the steps of performing any or all of the following techniques: May include: literature analysis, statistical Kruskal-Wallis test (similar to non-parametric ANOVA) Number change analysis, Conover-Iman test (nonparametric pairwise test for multiple comparisons), and / or correlation analysis using RNA-seq data from Figures 1C-1D.
[0116] In some embodiments, a set of genes (e.g., for a particular cell type) is obtained from various sources. In some embodiments, only genes with known function are used. Some genes may be similar to the labels used in CYTOF. Some may be derived from literature data (indicating specificity of a particular gene) (Some genes may be present in the genome) and / or some genes may be absent in existing RNA-seq samples of sorted cells. (e.g., experimental conditions, sequencing quality, and expression quality) Searching for genes in samples can be done in several ways: Correlation of gene expression with proportion of cells in artificial mixtures using differential gene expression (e.g., as described herein, including with respect to at least FIG. 6A ) , gene expression and TCGA (The Cancer Genome Atlas) samples or selected cell samples. Correlation with several marker cell genes (such as CD3 for T cells) in pooled TCGA samples (e.g., to add a larger proportion of cells to the sample, increase the number of read counts. and to reduce the correlation between the presence of various cells within the tumor), artificial mixtures Use linear regression (e.g., with L1 regularization) on We use some key metrics (e.g., SHAP or gradient boosting tree gain) or by using some genetic algorithms to generate artificial cells with known cell composition. and / or the legacy that gives the best quality of predictions of machine learning methods on actual independent data. Select a combination of genes, or use any combination or linkage of these described methods. Use.
[0117] If a gene is expressed only in a particular cell type or cell subtype, then that gene is Genes can be considered "specific" for a particular cell type or subtype. can be considered "semi-specific" for a particular cell type or subtype if: (1) It is a cell type or subtype that is associated with both one or more other cell types or subtypes. (2) if it is expressed in a particular cell type or subtype and not in other cell types or subtypes; For example, a gene is expressed more abundantly than a specific cell type or subtype. The average expression of a gene is less than the average expression of the same gene in other cell types or subtypes. At least a threshold percentage (e.g., 50%, 100%, 200%, 500%, 1000%, etc.) or a threshold factor (e.g., 2 If the coefficient (e.g., 5, 10, 15, 20, etc.) is high, the gene is semi-specific to a particular cell type or subtype. In one example, the cell type or subtype of a gene may be considered to be heterogeneous. The average expression in one cell type is less than the average expression of the gene in other cell types or subtypes. If both are 10 times larger, the gene is semi-specific for a particular cell type or subtype. For example, between macrophages and monocytes, between CD4+ T cells and CD8+ T cells, There may be common genes between NK cells and CD8+ T cells. In this context, common genes can be considered semi-specific to cell type and / or subtype. In some embodiments, the antibody can be semi-specific for both CD4+ T cells and CD8+ T cells. The present study identifies genes whose expression is significantly lower or absent in malignant cell (e.g., tumor) lines. In some embodiments, as described above, multiple data Specificity criteria can be assessed when evaluated against the combined expression data from the dataset. In some embodiments, if several types of cells are present in the same dataset, For each such dataset, we consider similar idiosyncrasies within the dataset to reduce batch effects. Gender analysis can also be performed.
[0118] In some embodiments, for each set of genes, the genes are compared to TCGA (The Cancer To determine how the gene is expressed in the Genome Atlas for the desired tumor type For example, for a given cell type, the average expression of TCGA can be calculated. In other words, it may be desirable for the ratio of the actual specific or semi-specific genes (e.g., in a set of specific or semi-specific genes) The mean expression of is 70% of the mean expression in the sample of sorted cells, while this set If the expression of other genes is around 5%, the specific or semi-specific gene is likely to be involved in tumor development. It is likely that this gene is expressed by other cells or that cells within the tumor do not express this gene. The manifestations of the children are very different.
[0119] Additionally or alternatively, expression of genes from the same set may be associated with this type of tumor (e.g., If correlation between TCGA samples for the desired tumor type is desired, For this, we can analyze the average correlation with other genes from the set. The expression signatures of the genes considered in TCGA LUAD may be low (e.g., 10 TPM Therefore, the correlation between these genes may also be low (e.g., In some cases, NK cell and neutrophil genes Expression may be particularly low.
[0120] The inventors have demonstrated that cells with a common origin and function may often express the same genes. For example, hematopoietic immune cells express CD45 (PTPRC) and HCLS1. Due to their development, immune cells can be divided into lymphoid and myeloid cells. In the current study, lymphocytes can be divided into T cells, B cells, and NK cells, and among T cells, CD4+ T cells However, some of these cells are involved in tumor development and There are also subtypes that may play an important role in both the treatment and treatment process. As described in the literature, determining the cellular composition of a particular cell subtype However, the present inventors have found that the specific features expressed in cell subtypes There may be few specific and / or semi-specific genes, and such genes in the tumor microenvironment may be The number of cells that are viable may be smaller than the combined cell population, so we It is recognized and understood that isolating cell subtypes in vivo can be difficult.
[0121] The present inventors have proposed one method to improve the accuracy of determining both cell type and subtype. The method involves determining the cell composition ratio for a cell subtype by using a combination of cells (e.g., specific and / or semi-specific for cell types and subtypes that share common genes We have discovered that using information about gene expression can be a common Genes can be used, for example, to determine the cellular composition of individual cell types and subtypes. Another method using genes common to a group of cell subtypes is described elsewhere herein. As described in [2], first, the cell composition ratio for the combination group is calculated, and then the individual cells in the group are analyzed. It may be possible to refine the calculation to determine the cellular composition ratios for different cell types. do.
[0122] [Table 2A]
[0123] [Table 2B]
[0124] [Table 2C]
[0125] [Table 2D]
[0126] [Table 2E]
[0127] [Table 2F]
[0128] [Table 2G]
[0129] [Table 2H]
[0130] [Table 2I]
[0131] FIG. 2A depicts a method 200 for determining cellular composition for at least one cell type. In some embodiments, the method 200 is 10) For example, a computing device may include at least one processor and at least one and a processor and processor-executable instructions that, when executed, perform the operations of method 200. and at least one non-transitory storage medium for storing the data. In a system such as system 100 (which may include, for example, a clinical setting or a laboratory setting) , by one or more computing devices, for example, by computing device 108 This can be done.
[0132] In operation 202, the method 200 begins with obtaining expression data for a biological sample from a subject. In some embodiments, obtaining expression data comprises obtaining expression data from a subject using any suitable technique, In some embodiments, the expression data may include obtaining expression data from a previously obtained biological sample. Obtaining the data can include obtaining previously obtained expression data from a biological sample (e.g., In some embodiments, the expression data may include accessing a database to obtain expression data. The data is RNA expression data. Examples of RNA expression data are provided herein. In some embodiments, the subject has, is suspected of having, or has cancer. As described herein, including with respect to FIG. 1A, Body samples include biopsy samples (e.g., from a subject's tumor or other diseased tissue), also referred to as "biological samples." any of the embodiments described herein, including those relating to In some embodiments, the origin or preparation of the expression data may include any of a variety of biological samples. Any of the embodiments described with respect to the sections "Expression Data" and "Obtaining RNA Expression Data" For example, the expression data may include RNA expression data extracted using any suitable technique. As another example, the expression data obtained in act 202 may include expression data measured by TPM. The data may include RNA expression data collected.
[0133] In some embodiments, the expression data is stored in at least one storage medium and For example, expression data can be stored in one or more files, or accessed as part of a In some embodiments, the RNA expression At least one storage medium storing the current data is local to the computing device. (e.g., stored on the same at least one non-transitory storage medium), It may be external to the computing device (e.g., a remote database or cloud storage) The expression data may be stored in a single storage medium or in multiple The data may be distributed across multiple storage media.
[0134] In some embodiments, the expression data of operation 202 is a first cell type (e.g., a first cell type in a biological sample). a first set of genes associated with the cell type and / or subtype being analyzed; In some embodiments, the first set of genes may include first expression data associated with at least one of the first set of genes. 1E, at least as described herein with respect to FIG. 1E. For example, for endothelial cell types, the set of genes may include genes such as ANGPT2, AP LN, CDH5, CLEC14A, ECSCR, EMCN, ENG, ESAM, ESM1, FLT1, HHIP, KDR, MMRN1, MMRN2, These may include NOS3, PECAM1, PTPRB, RASIP1, ROBO4, SELE, TEK, TIE1, and / or VWF. In some embodiments, the first set of genes includes at least those described in Figures 4-6. Train a corresponding nonlinear regression model for that cell type as described herein The same as the set of genes or subset of the set of genes used as part of the process. It is also possible.
[0135] At operation 204, the method 200 determines a first cellular composition ratio for at least a first cell type. Proceed to step 1003. As shown, determine a first cell composition ratio for a first cell type. The method comprises: generating first expression data relating to a first set of genes for a first cell type; processing the data with a nonlinear regression model (e.g., one or more nonlinear regression models) to generate a first The method may include determining a first cellular composition ratio for a cell type. For example, the first expression data may be provided as an input to the first nonlinear regression model. Other information may be provided as part of the input to the regression model, for example, the median expression data may be included as part of the input to the non-linear regression model. Any other suitable information may additionally or alternatively be provided as part of the input (e.g., The mean of the expression data, the median or mean of a subset of the expression data, or the mean of the expression data or any other suitable statistic derived from or otherwise related to expression data.
[0136] In some embodiments, part of operation 204 may be performed for each cell type and / or subtype being analyzed. For example, the expression data may be analyzed in a multi-step manner. The subsets are then compared to each nonlinear regression model for each cell type and / or subtype. may be provided as input.
[0137] In some embodiments, the output of the nonlinear regression model is an RNA fragment from a first cell type in the sample. 2C and 3C. The estimated proportion of RNA from the first cell type is used to determine the proportion of RNA from the first cell type, as described in the literature. In some embodiments, the corresponding cellular composition can be calculated using at least one of the following: The techniques described herein, including those relating to 3C, involve processing a nonlinear regression model. In this case, the output of the nonlinear regression model is the estimated ratio of RNA The ratio may be an estimated cellular composition ratio for the first cell type, rather than a ratio.
[0138] In some embodiments, the process 200 then performs a dynamic analysis to output the first cellular composition ratio. Proceed to step 206. Regardless of the architecture or input, the output of one or more nonlinear regression models can be used in the method 200. It can be combined as part of, stored, or otherwise post-processed. For example, the cellular composition ratio for each cell type can be determined by the computer used to perform method 200. Some of the information may be stored locally on the computer device (e.g., on a non-transitory storage medium). In embodiments, the cellular composition ratios are stored in one or more external storage media (e.g., a remote database). or cloud storage environment, etc.).
[0139] FIG. 2B is an example implementation of a method 200 for determining cellular composition based on expression data. In some embodiments, the steps for implementing method 200 may be the same as those included in the exemplary flowchart of FIG. 2B. In some embodiments, the steps for implementing method 200 may include any suitable combination of the steps included in the method. The steps may include additional or alternative steps not shown in FIG. 2B. For example, method 200 may be The steps of performing may include all the operations included in the exemplary flowchart. Method 200 may include only a subset of the operations included in the exemplary flowchart (e.g., , operation 212 and operation 216, operations 212, 214, 216 and 218, operations 212, 216 and 220, etc.).
[0140] In some embodiments, implementation 220 begins with operation 212, where a biological sample from a subject is The step of obtaining expression data for a biological sample from a subject is shown in FIG. 2A, including in connection with operation 202.
[0141] In some embodiments, operation 212 includes obtaining first expression data and second expression data. The first expression data may be associated with a first set of genes associated with the first cell type. and the second expression data is associated with a second set of genes associated with the second cell type. For example, the first expression data may be a set of first genes associated with B cells. and the second expression data can be associated with a second set of genes associated with T cells. Additionally or alternatively, the first expression data can be associated with a first cell subpopulation. The second expression data can be associated with a first set of genes associated with a second type. The gene can be associated with a second set of genes associated with a cell subtype, e.g., the first expression data can be associated with a first set of genes associated with CD4+ cells; The second expression data can relate to a second set of genes associated with CD8+ cells. Techniques for identifying genes associated with different cell types and / or subtypes are known as "genetic This disclosure is described herein, including with respect to the "Child Selection and Specificity" section.
[0142] In some embodiments, the exemplary method 220 proceeds to operation 214, where the expression data is preprocessed. In some embodiments, preprocessing involves using one or more nonlinear regression models. For example, expression data can be generated that is suitable for processing by filtering. , combining, batching, filtering, or any other In some embodiments, the expression data can be pre-processed by any suitable method. These methods include "alignment and annotation," "removal of non-coding transcripts," and Any of the embodiments described in the "Conversion to TPM and Gene Assembly" section are included. It is possible.
[0143] After the expression data is preprocessed, the exemplary method 220 proceeds to operation 216, where the expression data and one or more nonlinear regression models (e.g., at least 5, at least 10, at least Using the 15 models, multiple cellular proportions can be determined for multiple cell types. In some embodiments, the nonlinear regression models include at least those related to FIGS. The neural network may be trained first according to the techniques described herein.
[0144] In some embodiments, a separate nonlinear regression model is used to analyze each cell type and / or subtype. For example, operation 216 may include operations 216a and 216b. and operation 216b, which may include operation 216c, which may include operation 216e, which may include operation 216f, which may include operation 216f, which may include operation 216g, which may include operation 216h, which may include operation 216i, which may include operation 216j, which may include operation 216j, which may include operation 216i ... A separate nonlinear regression model is used to determine the cellular composition of the Act 216a includes a step of: Operation 216b includes determining a first cellular composition ratio for the first cell type. A second cellular composition is determined for the second cell type using the expression data and a second nonlinear regression model. In some embodiments, operation 216 includes determining a ratio. In some embodiments, operation 216 may include only one of the following: one or more cellular components to determine the cellular composition of the cell (e.g., a third cell type or subtype) An example implementation of operation 216a may include using an additional nonlinear regression model of These and other related arts are described herein.
[0145] In some embodiments, the exemplary method 220 includes operation 2 for outputting the plurality of cellular composition ratios. In some embodiments, the plurality of cellular constituent ratios can be displayed using a graphical user interface. output via an interface, stored in memory, and transmitted to one or more other computer devices. The image data may be transmitted to a device and / or output in any other suitable manner.
[0146] In some embodiments, the output of the plurality of cellular composition ratios in operation 218 and / or the data obtained in operation 212 Techniques for post-processing the collected expression data can be used. Thus, the post-processing method uses the cellular composition and expression data to perform a post-processing analysis of the biological sample in operation 220. The method may include determining a malignant tumor expression profile for the sample. The file may contain information indicating the occurrence of malignant cells contained in the biological sample. In some embodiments, the expression of a plurality of different genes associated with malignant cells is The step of determining a malignant tumor expression profile includes (a) determining expression profiles for TME cells in the biological sample; (b) estimating the current expression profile of the biological sample (e.g., bulk expression data, The method may include subtracting expression of TME cells from the expression data (e.g., the expression data obtained in act 212). Exemplary methods for determining tumor expression profiles are described herein, including with respect to FIG. 3D. It is written in the book.
[0147] Figure 2C shows the results for the first cell type using the first expression data and the first nonlinear regression model. 10 shows an example implementation of operation 216a for determining a first cellular composition ratio using the Thus, in some embodiments, the first nonlinear regression model is a regression model based on the first expression data (e.g., FIG. 3C The method may include a first sub-model and / or a second sub-model for processing the first and second sub-models (as shown in
[0148] In some embodiments, the first expression data comprises a first set of genes associated with a first cell type. In addition to the first expression data associated with the first cell type, a second set of genes associated with the first cell type is also obtained. The expression data may include second expression data associated with the gene.
[0149] In some embodiments, the implementation uses the first sub-model to analyze RNA from a first cell type. In some embodiments, the method begins with operation 232 for predicting a first value for the estimated ratio of a first set of genes and / or any other input information; It can be provided as input to the first sub-model of the linear regression model, the output of which is the first It can be a predicted ratio of RNA from a cell type.
[0150] In some embodiments, after predicting the first value, the implementation may use the second sub-model to The method then proceeds to operation 234 for predicting a second value for the estimated proportion of RNA from the first cell type. In some embodiments, second expression data relating to a second set of genes is provided to the first sub-model. In addition to the predictions from the model and / or any other input information provided in the first sub-model, The second sub-model of the gene expression model may be provided as an input to the second sub-model of the gene expression model. In general, first expression data associated with a first set of genes is used as input to a second sub-model. According to some embodiments, predictions from multiple nonlinear regression models can be provided. The measurements (e.g., the output of the first sub-model of each nonlinear regression model for each cell type) are It can be provided as input to the second submodel of the nonlinear regression model for cell type. The output of the second submodel of a nonlinear regression model is independent of the input to the second submodel. The output of the second sub-model can be the estimated proportion of RNA from the first cell type in the sample. The power, in some embodiments, may include the output of a nonlinear regression model for the first cell type. .
[0151] In some embodiments, the nonlinear regression model may include more than two sub-models. For example, the second submodel can be repeated any number of times, each time using one or more of the previous submodels. The predictions from the submodel are included as input.
[0152] In some embodiments, the implementation then performs a second analysis of the estimated proportion of RNA from the first cell type. Using the value of 2, proceed to operation 236 to determine the cell composition ratio for the first cell type. In some embodiments, determining the estimated proportion of RNA from a first cell type comprises: (a) extracting a biological sample; and (b) estimating the number of cells of the first type contained in the biological sample. estimating the number of cells of the first type (e.g., using formula 350). is the estimated ratio of RNA (e.g., R in Equation 350) cell ) into the RNA factor per cell (e.g., A in Equation 350) c ell The step of estimating the total number of cells may include a step of comparing the number of cells of each cell type with the number of cells of each cell type. The method for estimating the cellular composition ratio may include the steps of: As described herein, including with respect to 3C.
[0153] FIG. 3A shows an exemplary use of a machine learning method for determining RNA ratios based on RNA expression data. In the illustrated example, the primary tumors available in the TCGA database RNA expression data from tumor sample 302 was compared with corresponding data for T cells, CD4+ T cells, and CD8+ T cells. To arrive at the estimated RNA ratio 306, the present specification, including at least those relating to Figures 2A-2C, The data is processed according to the machine learning techniques described in the book.
[0154] In the illustrated example, the RNA expression data for tumor sample 302 is From a line database (e.g., in this example, The Cancer Genome Atlas (TCGA) database) In some embodiments, the RNA expression data is obtained from one or more databases, such as TCGA. It can be obtained from any suitable source, including a database, or directly from a biological sample. (eg, as described herein, including with respect to at least FIG. 1A).
[0155] Regardless of how the RNA expression data is obtained from the tumor sample 302, the RNA expression data can be processed using a non-linear regression model 304. According to some embodiments, The nonlinear regression model 304 is a nonlinear regression model as described herein, including with respect to at least FIGS. Implement it using a gradient boosting technique (e.g., as implemented in XGBoost) According to some embodiments, the method described herein, including with respect to Figures 2A-2C, can be used. As described above, the nonlinear regression model 304 generates a separate nonlinear regression model for each of the multiple cell types. In the illustrated example, the nonlinear regression model 304 is a T cell A nonlinear regression model for CD4+ T cells, a nonlinear regression model for CD8+ T cells As shown, in some embodiments, one or provides additional nonlinear regression models for multiple additional cell types and / or subtypes It is possible.
[0156] In some embodiments, the inputs to the nonlinear regression models 304 are: For example, the RNA expression data of Figures 2A-2C may include a selected subset of the RNA expression data of Figures 2A-2C. As described herein, including the inputs to the nonlinear regression model for a particular cell type, The data may include RNA expression data for genes specific and / or semi-specific to that cell type. For example, in the illustrated example, the nonlinear regression model for T cells is The input RNA expression data included CAMK4, CBLB, CD2, CD226, CD3D, CD3E, CD3G, CD48, and CD 5, CD6, CD7, FLT3LG, ITK, KCNA3, KLRB1, LAG3, LAT, LCK, LTA, SIRPG, SIT1, SLA2, TBX21, TCF7, TESPA1, TRAC, TRAF3IP3, TRAT1, TRBC2, TRDC, TRGC1, TRGC2, UBASH3A, In some embodiments, other information related to the RNA expression data (e.g., For example, the median of the RNA expression data, or any other suitable statistical value, is used as input to the nonlinear regression model. This may additionally or alternatively be provided as a force.
[0157] In some embodiments, the output of the nonlinear regression model 304 is a set of data for each cell type and / or subtype. For example, the nonlinear regression model for T cells may be As an output, we generate a predicted proportion of RNA from T cells in the input RNA expression data. Similarly, the nonlinear regression model for CD4 T cells can be used to calculate the RNA content from CD4 T cells. The predicted ratio was generated as the output, and the nonlinear regression model for CD8 T cells was A predicted proportion of RNA can be generated as output. As such, the predicted proportions of RNA are used to identify a portion or subtypes of the cell type and / or subtype being analyzed. The corresponding cellular composition ratios can be calculated for all.
[0158] In the illustrated example, the prediction for T cells is compared with the prediction for CD4 T cells + CD8 T cells. A plot comparing the predicted values for the subtypes is shown. may or may not be equal to the prediction for types containing those subtypes. For example, the sum of the predictions for CD4 T cells and CD8 T cells gives the prediction for T cells. Sometimes the sum of the predictions for CD4 T cells and CD8 T cells is greater than the prediction for T cells. In some embodiments, the sum of the subtype predictions may be greater than the total type prediction. May be equal and / or normalize subtype predictions to be equal to the total of type predictions Or it can be adjusted.
[0159] FIG. 3B illustrates a first sub-model 326, 327 for determining RNA ratios based on RNA expression data. 8, 330 and the use of nonlinear regression models 320, 322, 324 including second submodels 338, 340, 342. This is a diagram depicting the above.
[0160] As shown in the exemplary embodiment of FIG. 3B, different nonlinear regression models 320, 322, 32 4 was used to identify genes associated with each of the cell types: cell type A 308, cell type B 310, and cell type C 312. In some embodiments, the expression data 314, 316, 318 for each exemplary non-linear The shape regression model generates first values 332, 334, 336 for the estimated proportion of RNA from each cell type. the first submodel for the estimated proportion of RNA from each cell type; 344, 346, 348. Second sub-models 338, 340, 342 are included to generate the models 344, 346, 348.
[0161] Non-limiting examples for using a non-linear regression model that includes one or more sub-models include: 322, a nonlinear regression model trained to estimate RNA ratios for cell type B 310. In some embodiments, expression data 316 from a set of genes associated with cell type B 310 is collected. can be obtained and used as input to the nonlinear regression model 322. For example, 10 can include immune cells, and the expression data 316 can include genes ADAP2, ADGRE3, ADGRG3, C1QA , C1QC, and C3AR1 (e.g., immune cell-related gene sets listed in Table 2) In some embodiments, the expression data 316 may include expression data from at least Some (e.g., expression data related to a subset of genes, some related to all genes) The expression data for gene A is used as input to the first submodel 328. A subset of expression data 316 containing expression data for DAP2, ADGRE3, and ADGRG3 was input. The first submodel then uses the input expression data as a Processing can be performed to determine a first value 334 of an estimated proportion of RNA from cell type B 310 .
[0162] In some embodiments, the exemplary nonlinear regression model 322 estimates RNA from cell type B 310. A second sub-model 340 may be included to generate a second value 346 of the ratio. The second sub-model 340 may use one or more inputs to generate a second value 340. For example, in some embodiments, at least a portion of the expression data 316 can be used as input. In some embodiments, the expression data can be used in the same way to the first sub-model 328. expression data input (e.g., expression data for genes ADAP2, ADGRE3, and ADGRG3) In some embodiments, the expression data is input to the first sub-model using the same expression data. may include additional expression data in addition to the genes ADAP2, ADGRE3, ADGRG3, C1QA, and In some embodiments, the expression data is based on a first sub-model. may contain expression data different from the expression data input to Expression data for 3AR1).
[0163] Additionally or alternatively, in some embodiments, the second sub-model 340 may include other cell types 308, 3 Estimation of RNA output by the first submodels 326, 330 of the nonlinear regression models 320, 324 for 12 The ratio can be taken as an input. As shown, the second ratio for cell type B 310 The submodel 340 generates a first value 332 for the estimated proportion of RNA from cell type A 308 and a second value 332 for the estimated proportion of RNA from cell type C 3 It takes as input a first value 336 for the estimated proportion of RNA from 12. This type of input can be used in another Determine the ratio of RNA from cell types associated with the same gene or set of genes. For example, cell type B 310 may be differentiated from cell type C 312. When related to the same gene, gene X, the expression data obtained for gene X is Which of the two cell types is present in a biological sample may not be highly informative. This is because it may be unclear which cell type generated the expression data. However, if the first sub-model 330 determines a first estimated ratio of RNAs determined for cell type C, Consider a scenario where the output is 0% as value 336. This is because there are no cells of cell type C 312 in the biological sample. As a result, all expression data obtained for gene X The second submonomer must be expressed by cell type B 310. In some embodiments, the second submonomer The delta 340 can use the first values 332, 336 to make such an inference.
[0164] In some embodiments, the output of the second sub-model 340 is an estimated proportion of RNA from cell type B 310 As described herein, including with respect to FIG. 3D , The estimated RNA ratios can be processed to determine the cellular constituent ratios for each cell type. do.
[0165] FIG. 3C depicts a method for determining cellular composition ratios 370 based on RNA ratios 360. For example, the method of FIG. 3C may involve the analysis of a portion or portions of the cell types and / or subtypes. To arrive at a prediction of cellular composition for all, including those related to Figures 2 and 3A, and can be applied to predicted RNA ratios according to the techniques described herein.
[0166] As shown in the figure, the process of obtaining cell composition ratios based on RNA ratios is In some embodiments, formula 350 can be applied to the ratio of RNA for , may be applied to each RNA ratio individually (e.g., sequentially), or in some embodiments may be applied to some or all of the RNA ratios together (e.g., in parallel) In some embodiments, formula 350 calculates RNA ratios for cell types that are not subsets of each other. In some embodiments, the formula 350 can then be applied to the first It can be applied to RNA ratios for cell types that are subtypes of one or more cell types. In some embodiments, the calculation of the cell composition ratio for a cell subtype is performed by first calculating For example, in some embodiments, the ratio of the cellular components of the cells can be adjusted based on the ratio of the cellular components of the cells. The cell composition ratios for the cell subtypes calculated in the first place are summed to give the total cell type. So that the cell composition ratio of the body (i.e., the first calculated cell type that is a subtype) It may be normalized or otherwise adjusted.
[0167] For a given cell type, cell, equation 350 is:
[0168]
number
[0169] wherein C cell is the cell composition ratio for that cell type, and R cell is the cell type is the RNA ratio for A cell is the RNA coefficient per cell. As shown in Equation 350, The denominator may include the sum of all cell types and / or subtypes analyzed. Therefore, the formula
[0170]
number
[0171] is first calculated for all cell types and / or subtypes, and then for each cell type and / or subtype or individual C for subtypes cell can be used to calculate the value.
[0172] According to some embodiments, the RNA ratio for a cell type can be expressed as a fraction or decimal. (e.g., for purposes of calculation by Equation 350). In some embodiments, RNA ratios can sum to 1 (e.g., Σ cells R cell =1). In some embodiments, RN If the sum of the A ratios is less than 1, then R other We can introduce the formula, which is 1-Σ cell s R cell In some embodiments, when the sum of the RNA ratios is greater than 1, , R other = 0, and the RNA ratios can be normalized so that they sum to 1.
[0173] In some embodiments, formula 350 calculates the RNA factor A per cell. cell per cell The inventors have demonstrated that the RNA abundance per cell is related to the size and / or abundance of the cell. We recognize and understand that the results may depend on other factors. This can lead to different amounts of RNA for the bulk sample. can be used to convert RNA ratios to corresponding cellular constituent ratios. So, the RNA factor per cell, A cell can be determined as part of the model training process. can be used (e.g., simulated or artificial cells in which the ratios of different cell types are known). data). In some embodiments, the RNA factor A per cell cell Some or all of the details The cell type can be experimentally determined. For example, the RNA factor per cell is determined for each cell. Access data on RNA expression for the type (e.g., available scientific literature, e.g., From PMID:29130882, PMID:30726743, or average or nonlinearly transformed UM by cell type (estimated from single-cell data using I counts) and use that data to Determine the corresponding RNA factor per cell (e.g., purity and / or histological TCGA phosphorylation). This can be obtained by analyzing the plasma cell data.
[0174] In some embodiments, the RNA factor per cell may be tissue specific and may be analyzed It may vary based on disease (e.g., cancer). In some embodiments, the RN per cell The A coefficient may be tissue independent and may not differ depending on the disease being analyzed (this For example, even among different cancers, tissues, or diseases, non-malignant microenvironment cells remain the same. or may be represented by a substantially similar cellular phenotype). We combine data from multiple cancer types, tissues, diseases, etc. to calculate the RNA coefficient for each For example, in some embodiments, the RNA factor per cell for a cell type can be calculated. More than 10,000 different cancer tissue samples from TCGA were analyzed as part of the process to determine The present inventors have demonstrated that the non-malignant cell composition ratio is determined by histology and WES analysis. We recognize and understand that certain embodiments may be In some embodiments, determining the RNA coefficient per cell may include deriving a coefficient for RNA per cell type. To obtain this information, the non-malignant cell composition ratio obtained from RNA was compared with the cell composition ratio obtained from DNA. An aligning step may be included.
[0175] The techniques described herein are not limited to application to RNA-seq data only. For example, some embodiments of the techniques described herein may be implemented in a multi-threaded manner. For this purpose, expression values can be calculated using the 1 for RNA-seq. Normalize the expression to be in a similar range to the transcripts per million (TPM) values (e.g., total expression is 1 million), a linear scale can optionally be used.
[0176] FIG. 3D illustrates a method for analyzing a cell population based on cellular composition in accordance with some embodiments of the technology described herein. 3 is a diagram depicting an exemplary method 380 for determining a malignant tumor expression profile using This involves obtaining a biological sample (e.g., a biopsy sample) and identifying malignant cells contained in the biological sample. In some embodiments, the method may include determining the expression of a gene (e.g., the expression of an individual gene). This allows for the differentiation of TME cell expression from the overall expression of a biological sample (e.g., expression in a bulk biopsy sample). The method may include a step of removing the
[0177] As shown, this exemplary method includes three steps. The first step 382 is to In some embodiments, this includes determining the average expression profile of non-malignant cell types. may include using expression data from the sorted cell types. For example, this may involve cells, B cells, macrophages, fibroblasts, and any other cells that may be contained in the TME In some embodiments, the method may include obtaining and using RNA-seq data from a suitable cell type. The cell types can exclude tumor (e.g., malignant) cells. The average expression profile includes: , the average expression of a set of genes for each cell type.
[0178] This exemplary method then uses cellular deconvolution techniques to separate the cellular components. The cellular composition of the biological sample (e.g., a biopsy sample) can be analyzed by the second step 384 to predict the presence or absence of a cellular component. As shown, this includes the proportion of each cell type in the cell composition. A step of generating a ratio vector may be included. The process may be performed in accordance with any of the embodiments described herein, including those relating to Figures 1-3C. It may include.
[0179] Average expression profiles of multiple different cell types contained in the TME (e.g., first step 382) and and the proportion of each of those cell types in the biological sample (e.g., second step 384). As shown, the expression of each cell type in a biological sample can be estimated. Step 386 includes determining the product of the expression profile matrix and the cell proportion vector. The resulting vector can be used to identify putative expression profiles of TME cells in a biological sample. It is a file.
[0180] In some embodiments, determining the tumor expression profile comprises determining the tumor expression profile by bulk expression profiling of the biological sample. This may include subtracting the TME expression profile from (e.g., expression in a bulk biopsy sample). As shown, this includes expression profiles of TME cells from bulk expression vectors. This involves subtracting the vector generated for
[0181] FIG. 4 illustrates the use of one or more nonlinear 4 is a flowchart illustrating a method 400 for training a shape regression model. As such, the method 400 may include one or more nonlinear regression models (e.g., at least five, but not more than five). A set of nonlinear regression models (at least 10, but not less than 15) is trained to identify the corresponding phenotypes in the biological samples. In some embodiments, the method may include estimating the cellular composition ratio for one or more cell types. In this case, each nonlinear regression model predicts the cellular composition of a particular cell type in a biological sample. A separate nonlinear model for each cell type and / or subtype is trained to estimate A regression model can be trained.
[0182] In some embodiments, the method 400 may be performed on a computing device (e.g., , as described herein, including at least with respect to FIG. 10). The computer device includes at least one processor and a processor that, when executed, performs the operations of method 400. and at least one non-transitory storage medium storing processor-executable instructions for .
[0183] At operation 402, method 400 includes obtaining training data including simulated RNA expression data. In some embodiments, "simulated" RNA expression data can be used to The data may include RNA expression data that is partially generated in silico. The RNA expression data collected includes data from multiple expression datasets from purified cell type samples. In some embodiments, the data may include data obtained by sampling the reads. In the illustrated example, the RNA expression data may include expression data measured by TPM. the RNA expression data includes first RNA expression data for a first gene associated with a first cell type; and second RNA expression data for a second gene associated with the second cell type. The gene may be, for example, a first cell type specific and / or semi-specific gene, while Alternatively, the second gene may be a gene specific and / or semi-specific to the second cell type. In some embodiments, the training data is for each cell type and / or subtype being analyzed, and The data may include RNA expression data of genes associated with the serotype and / or other cell types.
[0184] In some embodiments, training data may be generated as part of operation 402. As described herein, including at least with respect to FIG. 6A, in some embodiments: RNA expression data from malignant cells (e.g., cancer cells) and microenvironmental cells (e.g., immune cells, dermal cells) We combine RNA expression data from various cells (e.g., human skin cells) with multiple simulated samples for training. A process for generating an RNA mixture (which may be referred to herein as an "artificial mixture" or "mixture") The process can generate simulated RNA expression data. In some embodiments, at least 1,000, at least 10,000, at least 100,000, or at least 1,000,000. The mixture may be generated and / or accessed as part of operation 402.
[0185] The training data may be obtained in any suitable manner in operation 402. For example, the training data may be The data is stored on at least one storage medium (e.g., in one or more files or in a database). In some embodiments, the training data may be stored in at least one Another storage medium may be local to the computing device (e.g., stored on at least one non-transitory storage medium) or on a computer device It may be external (e.g., stored in a remote database or cloud storage environment) The training data may be stored on one storage medium or distributed across multiple storage media. It may be scattered.
[0186] In some embodiments, operation 402 includes pre-processing the training data in any suitable manner. It may also include, for example, filtering, combining, and organizing the training data into batches. The image may be pre-processed by filtering, filtering, or any other suitable method. Preprocessing can be done, for example, using one or more nonlinear regression models. In some embodiments, training data suitable for performing the method according to at least one embodiment of FIG. As described herein, including dividing the training data into separate training, validation, and holdout The data can be divided into separate datasets.
[0187] In operations 404 through 408, the method 400 uses the training data to generate one or more The process can proceed to train the nonlinear regression model. In particular, acts 404 through 408 , a first model of the nonlinear regression model is trained to estimate the cell composition for the corresponding first cell type. The steps of estimating the ratio are described. Operations 404 and 406 are referred to herein as the training steps. According to some embodiments, each model of the nonlinear regression model may include at least one Alternatively, each nonlinear regression model can be trained separately for each cell type (e.g., (using corresponding different input data and different learned parameters). In some embodiments, each nonlinear regression model of the one or more nonlinear regression models is 406 to 408, as modified as necessary, according to the techniques described herein. , and trained and / or saved according to operation 408.
[0188] In operation 404, training a first one of the nonlinear regression models includes: Proceeding to generate an estimated ratio of RNA for the first cell type using the RNA expression data As described herein, the first RNA expression data can be related to a first cell type. a first gene associated with the first cell type (e.g., only genes specific and / or semi-specific to the first cell type) In some embodiments, the first RNA expression data is provided as input to the first model. In some embodiments, other inputs may additionally or alternatively be provided to the first model. For example, median, mean, or any other arbitrary value relating to some or all of the RNA expression data may be used. Any suitable information may be provided as part of the input to the first model.
[0189] In operation 406, training a first one of the nonlinear regression models includes training a first one of the nonlinear regression models based on RNA from a first cell type. We can then proceed to update the parameters using the estimated ratio of In an embodiment, as part of operation 406, the estimated proportion of RNA from the first cell type is calculated based on the proportion of RNA from the first cell type. The ratio of RNA from the serotype can be compared to known values. For example, the ratio of RNA from the serotype can be compared to known values. A loss function can be applied to the estimated and known values to determine the loss. In this embodiment, the loss can be used to update the parameters of the model. For example, to update the model parameters to minimize the loss, Gradient descent, or any other suitable optimization technique, can be applied to .
[0190] The first model may be generated by any suitable method, including non-linear regression methods, as described herein. In some embodiments, the first model can be ,Gradient boosting machine learning techniques can be used. For example, the first model is Use an ensemble of weak predictive models such as decision trees or gradient boosting algorithms The present invention may include any other suitable predictive models that can be combined in an iterative manner. In this embodiment, a gradient boosting framework such as XGBoost or LightGBM is used as the first It can be used as part of the process of training a model. The forest model can be used as part of the process of training the first model. .
[0191] In some embodiments, operations 404 through 406 are performed multiple times for a given nonlinear regression model. (e.g., at least 100 times, at least 1,000 times, at least 10,000 times, at least 100,000 times, In some embodiments, operations 404 through 406 may be repeated (at least 1 million times). , may be repeated for a set number of iterations, or may be repeated until a threshold is exceeded (e.g., (e.g., until the loss falls below a threshold). In some embodiments, including at least those relating to FIG. 5A. As described herein, a nonlinear regression model is trained in two or more stages. It is possible.
[0192] At operation 408, the method 400 continues by training a nonlinear regression model including the first nonlinear regression model and the second nonlinear regression model. In some embodiments, the method may proceed to output the generated nonlinear regression models. , the step of outputting the trained plurality of nonlinear regression models may include: and storing one or more of the following on at least one non-transitory computer for later access. storing the model in a readable storage medium (e.g., memory); providing the model to a recipient; (e.g., using any suitable communication network or other means to communicate with the model) transmitting the associated data to the recipient; Graphical user interface and / or any other method that outputs the trained model. displaying the information to the user in a suitable manner, which is an embodiment of the technology described herein. The method is not limited in this respect.
[0193] FIG. 5A shows a graph of the correlation coefficients obtained by using one or more nonlinear regression models according to the method developed by the present inventors. 5 shows an example method 500 for training a model. The illustrated technique involves at least In combination with any of the other techniques described herein, including those relating to FIGS. 2 and 4 It can be used as such.
[0194] As shown, the method 500 includes preparing one or more datasets for training. In some embodiments, the data set may be , may be generated as part of operation 502 (e.g., the present invention, including at least those relating to FIG. 6A). according to the techniques described herein) and / or may be accessed (e.g., one or more database), as described in more detail herein, including with respect to FIG. The dataset may contain multiple artificial mixtures of RNA expression data, including species RNA expression data from various malignant (e.g., tumor) and / or microenvironment cells may be included. In embodiments, the dataset is at least 1000, at least 10,000, at least 10 It may contain tens of thousands, or at least a million artificial mixtures.
[0195] In some embodiments, the dataset may be divided into a training dataset and a holdout dataset. For example, in some embodiments, the data sets can be used for training and holding, respectively. The set proportion of data sets to be used is randomly assigned to the training and holdout data sets. For example, in the illustrated example, 80% of the data set is the training data set. The remaining 20% is left as a reserve dataset.
[0196] As shown in the figure, the holdout dataset is used to develop quality metrics. 7B). In some embodiments, the reserved data set is used so that the entire data set can be used for training. As shown in the diagram of operation 502, the training dataset is Further, one or more (e.g., 10) sets of data, each containing a respective training set and validation set. According to some embodiments, the training dataset can be subdivided into lane folds. In some embodiments, crossforks are used as part of training. It is possible to perform field verification.
[0197] Regardless of how the data set is prepared in act 502, method 500 In operation 510 and operation 520, a step of training a plurality of nonlinear regression models using the dataset is performed. As described herein, including at least with respect to FIG. Each nonlinear regression model is then applied to the corresponding RNA expression data from a particular cell type based on the input RNA expression data. It can be trained to estimate the proportion of RNA that is present in the sample. First, the nonlinear regression model is trained in a first stage corresponding to training a first sub-model of the nonlinear regression model. The training is done in two stages: a first stage, which corresponds to training the second submodel of the nonlinear regression model, and a second stage, which corresponds to training the second submodel of the nonlinear regression model. You can train.
[0198] In the first stage, in operation 510, a first sub-model of each nonlinear regression model is trained to An initial prediction of the proportion of RNA from each cell type can be generated. For each first submodel, the inputs are corresponding cell type-specific and / or semi-specific genetic In some embodiments, the RNA expression data may include cell type-specific and / or semi-specific RNA expression data. In some embodiments, only RNA expression data for genes that are not required can be provided as input. can provide other information, for example, median expression data. Regardless of the inputs used, the output of the first stage is an initial estimate of the proportion of RNA from each cell type. Each first sub-model of each nonlinear regression model may be predictive, and each first sub-model of each nonlinear regression model may be predictive for its respective cell type. Provides predictions about
[0199] In a second stage, in operation 520, a second sub-model of each nonlinear regression model is trained to A second prediction can be generated for the proportion of RNA from each cell type. For each second submodel of the model, the inputs are specific and / or semi-specific for the corresponding cell type. In some embodiments, the method may include RNA expression data of target genes as well as predictions from the first stage. Is the RNA expression data used in the second stage different from the RNA expression data used in the first stage? For example, in some embodiments, a nonlinear regression model is trained in the second stage. For purposes of this, some or all of the training data may be regenerated (e.g., as shown in Figures 5B and 5C). (According to the techniques described herein, including with respect to FIG. 6). The training data for the first and second stages are parallelized so that the training data for each stage is different. In addition to the RNA expression data, the first The predictions from the stage may be provided as input to the second stage. An initial prediction for all cell types may be provided as input to the second stage. This allows the second stage to effectively correct the predictions from the first stage, resulting in a final model. This may improve the consistency and / or accuracy of the rules.
[0200] Regardless of the input provided in the second stage, the output of the second stage is the RNA from each cell type. and a second prediction of the ratio of A prediction for each cell type may be provided. In some embodiments, the second prediction is , which may be the final output of a nonlinear regression model (e.g., the present invention, including those relating to Figures 2 and 4). As described herein. In some embodiments, additional training stages (e.g., additional sub- (e.g., third stage, fourth stage, etc.), each stage taking as input It takes new training data (e.g., RNA expression data) and predictions from the previous stage.
[0201] The process of providing predictions from a previous stage as part of the input to the next stage is performed in a manner that is consistent with the specific cell type. The model for the cell types uses information about the estimated proportions of other cell types to This allows for adaptation (e.g., if the total number of T cells is equal to 10 and the number of CD4+ T cells is 8 (Knowing that there are 1, the number of CD8+ T cells cannot exceed 2.) As shown, a multi-stage training procedure allows the model to take this into account. This procedure allows information from multiple different cell types and subtypes to be obtained from each individual cell. It may be possible to use it for type models.
[0202] FIG. 5B illustrates a block diagram of a machine learning model trained in accordance with some embodiments of the techniques described herein. 2 is an exemplary, non-limiting illustration of a method for detecting a smeared image. and in combination with any of the other approaches described herein, including those relating to FIG. It can be used.
[0203] As shown, the diagram 530 is a representation of the method described herein, including with respect to FIG. 5A. This diagram illustrates the division of a dataset into one or more folds, for example: Randomly split the dataset into three folds and train each of the three folds. In some embodiments, the data set may be further divided into a validation data set and a validation data set. Using the dataset as described herein, including with respect to FIG. 6A, A target mixture can be produced.
[0204] In some embodiments, folds are then used to: One or more models for a given set of parameters (e.g., parameters 550) The parameters can be trained in the set of predefined ranges shown in Table 3. In some embodiments, the number of occurrences may be generated (e.g., randomly) based on the number of occurrences. Each cell type model is trained individually using at least some (e.g., all) of the fields. In some embodiments, a validation mix can then be used to test each parameter. A set of metrics can be evaluated to generate associated evaluation data. In an embodiment, the parameters may be: It can be updated at each stage of training and / or used as input to subsequent training stages. For example, the first fold can be used as input for the first stage of training. The first set of parameters can then be generated by training the second fold. The updated set of parameters is used as the input for the second stage of Table 4 and Table 5 show the first and second stages of training, respectively. and exemplary parameters for one or more cell type models after the first and second stages. There are.
[0205] [Table 3]
[0206] [Table 4]
[0207] [Table 5]
[0208] FIG. 6A illustrates a method for generating simulated RNA expression data using one or more nonlinear FIG. 6 depicts an example method 600 for training a regression model (e.g., at least All of the data used as training data as described herein, including those relating to FIGS. In some embodiments, the simulated RNA expression data is generated in branch 6 of method 600. 10 and 620, malignant cells (e.g., cancer cells) and microenvironmental cells (e.g., generated by combining samples of RNA expression data from various cells (e.g., immune cells, stromal cells, etc.) An exemplary process for generating an artificial mixture of RNA expression data can be: As described below with respect to Figure 6A.
[0209] FIG. 6B illustrates a method for mimicking real tissue in accordance with some embodiments of the techniques described herein. 1 is a diagram depicting an example of a process for generating an artificial mixture of RNA expression data for In some embodiments, the RNA expression data is collected from one or more live samples, as shown in branch 630. One or more selected genes that represent a biological state (e.g., positive gene regulation, negative gene regulation, etc.) In some embodiments, the nuclei are derived from the same cell type / subtype as shown at branches 640 and 650. As such, one or more cell types / subtypes can be mixed in various proportions to create artificial mixtures. do.
[0210] FIG. 6C illustrates a method for training a cell type model according to some embodiments of the techniques described herein. 1 is an exemplary diagram for generating and using an artificial mixture to In this example, the data set is a fall as described herein, including with respect to FIG. In some embodiments, the resulting dataset is used to Then, in some embodiments, the artificial mixture is used to create a or train one or more nonlinear regression models specific to multiple cell types / subtypes, respectively. In some embodiments, the folds are mixed and validated as described with respect to FIG. The resulting models from each of these can be considered jointly or independently.
[0211] 6D and 6E illustrate the use of specific cell types in accordance with some embodiments of the technology described herein. An exemplary method for generating specific artificial mixtures for training subtype models In some embodiments, the compounds described herein, including those related to Table 6, To train a specific cell type / subtype model, one or more data Sets can be excluded.
[0212] FIG. 6F illustrates a process for processing a dataset in accordance with some embodiments of the techniques described herein. 1 is an exemplary diagram showing an approach for generating artificial mixtures using
[0213] As shown, operation 602 involves rebalancing (e.g., avoiding overtraining the model). Before the process of resampling large datasets to identify cell types, In some embodiments, the present specification, including those relating to FIG. The data sets are rebalanced 604 and combined as described below: It may be the entire set of samples for a particular cell type. The samples are then randomly selected in operation 608 and averaged in operation 612. In some embodiments, in accordance with the techniques described herein, As mentioned above, overexpression noise can be added to the expression of cell types.
[0214] Data collection, analysis and pre-processing According to some embodiments, the sample of RNA expression data is at least as related to Figures 1C-1D. For example, selected malignant cells can be obtained as described herein, including Constructing an artificial mixture of RNA expression data using multiple samples of cells and microenvironment In some embodiments, the number of samples can be as large as the number of samples included in Table 1. In some embodiments, the number of samples may be at least 5,000, at least 1 0,000, at least 15,000, at least 20,000, at least 30,000, at least 5 The number of samples may be 0,000, at least 100,000, or any number of suitable samples. In this paper, we will introduce open-source datasets, such as Gene Expression Omnibus (GEO) and Arra. In some embodiments, the dataset used is: Selected to meet the following criteria: human (homo sapiens) only, read length > 31 bp Standard RNA-seq (no polyA depletion, targeted panel, etc.). In some embodiments, artificial mixtures are constructed. To construct a marker, only cell types relevant to the particular disease being analyzed (e.g., a particular type of tumor) are selected. In contrast, the methods described herein, including at least those relating to FIG. 1E, can be used. As described in the previous section, the analysis of gene expression specificity involves data for all cell types. It can also be used for
[0215] In some embodiments, the selection of the dataset is based on biological parameters and bioinformatics. The parameters can be based on both physiological and physiological parameters. A dataset of samples cultured under similar conditions may be used. CD overstimulated by ball 12-myristate 13-acetate and ionomycin activation Datasets of 4+ T cells or macrophages co-cultured with excessive bacterial cultures As with the datasets, datasets with abnormal stimulation were excluded. Only samples with coding read counts of at least 4 million were used.
[0216] In some embodiments, quality control is performed on the RNA expression data prior to construction of the artificial mixture. can be implemented (e.g., to exclude strange or unreliable data sets). For example, some samples of CD4+ T cells lack expression of the CD45, CD4, or CD3 genes. In some embodiments, other The same can be done for cell types. For example, samples for some cell types can be , if they significantly express genes that are not typical for that type of cell (e.g., T cells). In the samples with cells, CD19, CD33, MS4A1, etc. were expressed in significant amounts, and most of the other In some embodiments, these T cell samples may be excluded if their expression is low. In this condition, if a sample of CD4+ T cells expresses significant amounts of the CD8 gene, they are depleted. In some embodiments, a method such as t-SNE or PCA with different sets of genes can be used. Several methods of expression analysis can be used to delineate similarities and differences between datasets. (e.g., as shown in Figures 1C and 1D). If a particular cell type in a dataset fails to cluster with the same cell type in other datasets, If so (e.g., in t-SNE, PCA, or other plots), quality control the dataset. further analysis as part of the data set and excluding some or all of the data from that data set. This can be done.
[0217] Building the mixture According to some embodiments, a sample prepared as described herein above is used. to generate various artificial mixtures of RNA expression data (e.g., representing simulated tumor tissues) The artificial mixture can be constructed using the sample expression and calculated as TPM (transcripts per million). The gene expression for the whole sample is generated by the expression of each individual cell from that sample. In some embodiments, the eigenvalues described herein may be formed as a linear combination. As described below, RNA expression data from samples of various cell types were mixed in predetermined ratios. As shown in Figure 6A, the simulated RNase activity for malignant cells A. Expression data (e.g., generated as shown in branch 610) is used to generate stains for microenvironment cells. Combine with simulated RNA expression data (e.g., as generated as shown in branch 620) It is possible.
[0218] Referring now to branch 620, a method for generating simulated microenvironment cellular RNA expression data is provided. In the illustrated example, a sample of each cell type (e.g., For example, a sample of RNA expression data for genes GSE1, GSE2, GSE3, or GSE4, as indicated. , per dataset (e.g., reducing the weight of datasets with a large number of samples) and sub- It is possible to rebalance by type (e.g., changing the proportion of a sample subtype). The rebalancing method can be divided into two categories: "rebalancing by data set" and The following are described herein, including with respect to the "Rebalancing by Subtype" section: For each cell type, multiple samples can then be randomly selected and averaged. Then, rebalanced / balanced samples are prepared for some or all of the cell types used. The equalized samples can be mixed together in a specific ratio (e.g., the actual tumor to simulate the tumor microenvironment).
[0219] Referring now to branch 610, a method for generating simulated malignant cell RNA expression data is In the illustrated example, cancer cells (e.g., NSCLC, A random sample of 1000 serogroups (e.g., ccRCC, Mel, HNCK, etc.) can then be selected. Adding high expression noise to the RNA expression data to consider abnormal gene expression in malignant cells For example, tumor cells may express genes that are not normally present in the parent cell type. Specific, semi-specific, or inflammatory cytokines associated with immune or stromal cells within the TME may be present. Where this is the case for a marker gene, the highly expressed gene is one of the genes described herein. Contains high expression noise that may interfere with the deconvolution methods described Whether or not the branch 610 results in simulated malignant cell RNA expression data. It is possible.
[0220] As shown in the figure, simulated RNA expression data for malignant cells (e.g., For example, the generated RNs are shown in branch 610) and the simulated RNs for the microenvironment cells. A) and expression data (e.g., generated as shown in branch 620) to generate an artificial mixture. (referred to as "expression mixture" in FIG. 6A). Simulated RNA expression data for cells and simulated RNA expression data for microenvironment cells The RNA expression data collected was randomly sorted based on a given distribution for cancer cells. In some embodiments, due to technical noise and biological variability, Noise can then be added to the mixture to mimic the noise that Each type can be designated according to one or more suitable distributions. For example, in FIG. As shown, technical noise is specified by a Poisson distribution, and biological variability is The noise due to the noise may be specified to follow a normal distribution. The technical noise may have multiple components that may be specified by other distributions. For example, Another component of technical noise may be specified by a non-Poisson distribution. Regardless of how the product is produced, in some embodiments, the artificial mixture is It is possible to represent an artificial tumor containing a tumor microenvironment (TME).
[0221] When creating artificial mixtures, the inventors have found that different cells of the same type from different samples I recognize and understand that there are cases where it is desirable to use a few By using only one sample per cell type, This may result in poorer performance for tumor samples (e.g., due to the cell state and its development). This is due to the limited number of read counts for differential expression, as well as the variability of expression. noise, alignment errors, and other sources of technical noise). In creating the artificial mixture, we use as many available cell samples as possible. We recognize that it may be desirable to use
[0222] Therefore, for this example, a large number of RNA-seq samples from various cell types (e.g., at least at least 100, at least 500, at least 1000, at least 2000, or at least 5000 In some embodiments, malignant cells (e.g., pure cultures for various diagnostic purposes) were collected. Several datasets of cancer cells (cancer cells, cancer cell lines, or cancer cells selected from tumors) were For each cell type, several corresponding data sets from different datasets can be collected. Table 7 shows the percentage of samples remaining after quality control for some cell types. The quantities of samples are listed.
[0223] In some embodiments, artificial The mixture is used as a training data set for training one or more nonlinear regression models. In some embodiments, a nonlinear regression model can be used to model the cell type / subtype. Thus, in some embodiments, for each specific cell type model To train a model using the As shown in Figures 6D and 6E, the set of mixtures used for each model may contain specific datasets that allow differentiation between specific cell types / subtypes. For example, to train a model of CD4+ T cells, To avoid uncertainty about the proportion of CD4+ T cells in the dataset, we used unspecified T cells. As an example, Table 6 shows a dataset containing one or The mixture of multiple corresponding cell types / subtypes used to train the model is specified. are.
[0224] [Table 6]
[0225] Sample averaging In some embodiments, multiple samples for each cell type are averaged in any suitable manner. (e.g., to improve the quality of the sample before adding artificial noise). In some embodiments, the averaging step can be performed in groups of two, e.g., 400 An averaged sample of 10,000 reads may contain information from 8 million reads. In this case, the step of averaging across multiple samples may be affected by technical factors during sequencing. This can reduce the expression noise caused by
[0226] In some embodiments, for each cell type, num av Select samples and average their expression do(num av The values of are shown in the parameter table, Table 9. Any subtype sample can be used at this stage as a type sample. For example, in some embodiments, regulatory T cells can be treated together with T cells. Although this approach allows for greater subtype diversity in artificial samples, it is not enough to Averaging multiple samples reduces biological variability in gene expression within cell types or subtypes. The degree of averaging employed may affect the results of the training. For this reason, the number of samples for averaging is indicated as a parameter and can be adjusted accordingly during training. can be selected along with other parameters (e.g., to enhance or maximize quality). (To do so).
[0227] Sample Rebalancing The number of cell samples available varies greatly for different datasets and cell subtypes. Because of this possibility, in some embodiments the number of samples can be rebalanced. As described herein below, in one example, samples are rebalanced for each data set. The cells can be sorted and then rebalanced by subtype. From the rebalanced number of samples, num av A sample can be selected.
[0228] Rebalancing per dataset In some embodiments, the number of sorted cell samples in a dataset can range from one to several hundred (e.g., For example, at least 5, at least 10, at least 50, or at least 100 samples Typically, each dataset is selected and sequenced in the same way. Cell samples within the same dataset may include samples of one or two cell types for which A specific marker set for selection or a specific condition, such as a specific disease of the patient from whom the cells were collected Datasets with a large number of samples may also have conditions. This can lead to overtraining of the model. Therefore, samples from all datasets are rebalanced for each dataset. It is sampled.
[0229] For example, in some embodiments, for each data set, the number of samples is set to a number N dataset,new to Replace and resample.
[0230]
number
[0231] In the formula, N max is the number of samples in the largest data set (e.g., for a particular cell type) Ri, N dataset,old is the number of original samples in the data set. The parameter is a value in the range [0,1], where 0 means no change in the number of samples. where 1 means there are the same number of samples for each data set. In this case, the rebalancing parameters can be selected during training.
[0232] Rebalancing by cell subtype For some cell types, samples of this type as well as more specific subtypes are available. The number of available subtype samples may be limited by the number of these subtypes. The ratio of the mixture to the specified ratio may not be consistent with that specified for the cell type. When creating a mixture of these subtypes, the samples can be rebalanced. do.
[0233] For example, in some embodiments, a CD4+ T cell (and helper T cells along with regulatory T cells) sample may be available in significantly greater numbers than CD8+ T cells. In this case, the average T cell The proportion of CD4+ and CD8+ T cell samples was determined prior to random selection of samples to form the cell sample. For example, the percentages can be changed in the TCGA or PBMC samples for these cell types. The ratio of RNA fractions can be chosen to be similar to the average ratio of RNA fractions expected for some experiments. In an embodiment, prediction is based on one or more linear models trained with a mixture of equal cell proportions. can be obtained using the del.
[0234] The subtype rebalancing algorithm can be as follows: For a given type To rebalance each subtype, P subtype *msize / min P +1 Resampling is performed with replacement for a number of samples equal to . In the formula, P subtype is a number reflecting the proportion of a given subtype (e.g., for a given type, The proportion of this subtype among all subtypes in the population. (This can be expressed as the number of all samples divided by the total number of samples of this type) , msize is the maximum number of samples among all subtypes of a given type, and min P is, all P between all subtypes subtypeAccording to some embodiments, rebalancin The tagging operation will handle all nested subtypes (e.g., types that themselves have subtypes). This can be done recursively on all subtypes (including the subtypes).
[0235] Generation of microenvironment cell ratios According to some embodiments, the resulting samples of the plurality of different cell types are subjected to a simulated The RNA expression data of the microenvironment cells were randomly mixed with each other to generate the RNA expression data. For example, a first fraction of an artificial mixture can be prepared using a random proportion of each cell type. You can generate a set,
[0236]
number
[0237] In the formula, R cell is a random number uniformly distributed between 0 and 1, and K cell The relationship between the specific cell type It's a number.
[0238] According to some embodiments, the coefficient K in the above equation cell is the most likely site of cellular mRNA The ratios are close to those observed in TCGA or PBMC samples. You can choose to These approximate ratios were obtained using a model trained without such ratios, and were compared with the TCGA or PBMC samples. For example, the approximate percentage for a given type of tissue can be calculated from the You can use a vector of numbers that reflects the vector. Multiply each number in the vector by a random number between 0 and 1. The resulting coefficients are normalized to the sum and used in a linear combination. In this embodiment, K cellFor each of the multiple cell types, we performed the following experiments based on tumor tissue and blood (PBMC). The most likely proportions of cell types based on the phenotype can be selected from Table 7. do.
[0239] We believe that the deconvolution algorithm works across the entire range of cells. We recognize and appreciate that in some cases it may be desirable to use a cell suspension from a tumor sample. Preparation of a suspension can lead to a dramatic increase in the proportion of lymphocytes, and the syphilis of such a suspension It may be desirable for the algorithm to operate on sequencing data. The authors found that the formation of cell ratios by the described method can be achieved by targeting certain cell types, such as NK cells. A sample with a high percentage (e.g., 70-100%) of It is recognized and understood that, in some embodiments, a parameter is An additional mixture is created in which the proportions are generated from a Dirichlet distribution with a parameter 1 / number_of_types. This parameter is selected along with other parameters to create the mixture. The number of samples in the data set formed in this way can be determined by the parameter This parameter can be controlled by the dilute sample ratio (Table 9). can also be selected as a parameter for creating a mixture. Thus, in the final data set, each cell type can be found in a proportion between 0 and 100%. Here, most of the characteristic quantities likely reflect cell populations that mimic actual tumors. There is.
[0240] In some embodiments, expression of the artificial tissue is controlled by expression vectors for each cell type and the cells. For example, the RNA may be generated based on a randomly selected proportion of the RNA. Expression vectors were then transfected with random vectors reflecting the RNA content of those cells, as described in Add to the coefficient.
[0241]
number
[0242] where α is a random coefficient reflecting the random proportion of cellular RNA for each cell type. can be,
[0243]
number
[0244] represents the RNA expression data of a particular gene for that cell,
[0245]
number
[0246] represents the RNA expression data of a particular gene for that mixture.
[0247] [Table 7]
[0248] Noise Generation As shown in Figure 6A, after the artificial mixture was generated, the RNA expression data showed no significant differences between the neutrophils and the neutrophils. Adding noise (e.g., technical noise, uniform noise, or any suitable type of noise) For example, noise can be generated and RNA expression data can be generated according to the process described below. The data is added to the
[0249]
number
[0250] In some embodiments, the expression of each gene can contribute noise to the overall tissue expression. For example, a single gene
[0251]
number
[0252] The expression of
[0253]
number
[0254] can be expressed as In the formula, μ Ti represents the true expression of the gene,
[0255]
number
[0256] represents the Poisson technical noise, and N prepi is derived from sequencing library preparation represents normally distributed noise, N bioi represents the variable biological noise.
[0257] In some embodiments, the relative standard deviation (δ Pi ) and normally distributed noise Relative standard deviation (δ Ni ) to calculate the quantitative relative standard deviation.
[0258]
number
[0259] Technical variability is due to differences in sample and library preparation (non-Poisson noise), as well as colorimetric analysis. Random transcript selection (Poisson's method) in the sequencing pathway due to limited coverage Many cell types in the microenvironment are typically present in tumor samples. Therefore, we have decided to use the It is important to consider the different levels of variability or noise depending on the level of expression. For example, in some embodiments, technical noise (Poisson's A mathematical noise model based on TPM is provided, taking into account both Poisson and non-Poisson noise. In an embodiment of the present invention, a nonlinear regression model is generated to train the nonlinear regression model as described herein. A model of this variability can be added to the artificial mixture created. The technical non-Poisson noise is assumed to be normally distributed. Account for variations in manufacturing, alignment, or human handling of different samples. In contrast, Poisson noise is a type of technical noise and is May be related to the coverage or read count of the target gene and may not be normally distributed. The resulting dependence of technical noise on coverage and gene expression is given by the formula :
[0260]
number
[0261] can be expressed by In the formula, l i is the effective gene length, Tj is the average TPM in technical replicates, and R is the read where α is the estimated proportionality coefficient. According to this formula, the lower the coverage, the more dispersion there is. According to this formula, genes with low expression levels have a high level of Poisson noise. It is thought that this will become more common.
[0262] As described herein below with respect to Example 1, this model uses purified cells As a result of the expression levels and coverage as shown using technical replicates of the population As shown, this allows accurate representation of gene expression variability (Figure 12I). In this case, the detection limit for gene expression ranged from 1 TPM at a coverage of 20 million total reads to 1000 samples. The coverage ranged from 12 TPM at 1 million reads per 10000 reads. The ability to replicate as a function of read count can be affected by the amount of material available. Calculate the noise factor (α) for Poisson noise by plotting the values of By calculating this coefficient, the technique for each sample and each gene can be calculated (Figure 12K). The technical noise can be estimated according to an estimation formula.
[0263] In addition to technical noise, biological noise that may be related to different activation states of cells also affects RNA- In some embodiments, artificial mixtures may contribute to the overall variance of the chromatin-seq samples. It may not be necessary to add biological noise, because this noise is a reflection of the biological state. Already existing through the use of RNA-seq data derived from cell subsets that represent variability As described herein below with respect to Example 1, this The overall variance can be calculated by comparing data for the same cell type obtained in different experiments. This can be evaluated by plotting the Poisson and non-Poisson data (Figure 12J). Both technical noise as well as biological variability depend on the average sequencing coverage. An example of this is presented in Figure 12J. In this example, for a particular cell type, , there is an average increase of 10% to 26% in noise from technical to biological replicates ( Figure 12J, right).
[0264] In some embodiments, the noise contribution from single gene expression is calculated as described herein. The analysis is applied to simulate technical and biological noise in artificial mixtures. For example, noise can be added to the total gene expression as two summands: It is possible.
[0265]
number
[0266] In the formula, ξ P ,ξ N ~N(0,1), β is the coefficient of the Poisson noise level coefficient, and γ is the uniform level are the coefficients of non-Poisson noise (Table 9).
[0267] As described herein below with respect to Example 1, the above approach can be used in a non-technical manner. The results are obtained by subtracting technical Poisson noise from biological noise. In the example of Figures 12L to 12M, an average variance of about 16% was obtained, which was then In this example, after technical correction, the noise was Technical non-Poisson and biological variability are independent of the measurement method. This is to be expected, since
[0268] The noise models described herein are used to analyze artificial mixtures using technical (Poisson and This allows for the addition of variability (both Poisson and non-Poisson) to more accurately represent the actual tissue. A good imitation artificial mixture is obtained. The improved artificial mixture is then used to Train the deconvolution algorithm (e.g., see Figures 4-6). (as described in the specification) Stability can be ensured.
[0269] Hyperparameter estimation As shown in Figure 6A, a nonlinear regression model was used according to the method developed by the present inventors. In some embodiments, training the model includes estimating and and / or updating the model as described herein. In addition to the weights learned for the model, the parameters of Several parameters, called parameters, may be included (e.g., at least those related to FIG. 4). (including as described herein) such hyperparameters and An exemplary list of these values is shown in Table 9.
[0270] In some embodiments, each time a nonlinear regression model is trained, the hyperparameters For example, some or all of the hyperparameters can be estimated using the training data. can be updated based on one or more validation sets of data (e.g., model In some embodiments, the hyperparameters are calculated based on the TCGA data. For example, the Results are checked for consistency against TCGA data to ensure TCGA model agreement is achieved. For example, in the illustrated example, for a given cell type (e.g., lymphocytes), The sum of the results across cell subtypes (e.g., T cells, B cells, and NK cells) is used to calculate the percentage of the cell. It has been confirmed that the overall results are equal (or close) to those for the cell types.
[0271] In some embodiments, as part of the hyperparameter estimation, a parameter search can be implemented using random search, grid search, or genetic algorithms. Any suitable parameter search technique can be used. For example, parameter testing can be performed using Bayesian optimization, gradient-based optimization, or evolutionary optimization. In some embodiments, a parametric search can be used to Select one or more hyperparameter values from a given range associated with the parameter. You can choose.
[0272] Table 8 and Table 9 list examples of hyperparameters: Averaging The number of samples (Nav), uniform noise level (γ), Dirichlet sample proportion (Dp), rebalancing The r parameter, the high expression ratio (Hf), and the maximum high expression level (Mhl).
[0273] For each cell type, the "Na" v" samples are selected and their expression is averaged.
[0274] As described above, including with respect to the "Generation of Microenvironment Cell Fractions" section, We can create several artificial mixtures "Dp" whose proportions are generated from the distribution.
[0275] As described above, including with respect to the "Per Dataset Rebalancing" section, The rebalancing parameter "r" is used in the formula to determine the number of new samples in the data set. As stated, "r" is a value in the range [0,1], where 0 means no change in the number of samples, 1 means there are the same number of samples for each dataset In some embodiments, the rebalancing parameters are selected during training. It is possible.
[0276] As described above, including with respect to the "Constructing the Mixture" section, To mimic the over-expression of the gene, we added high-expression noise to each of the artificial mixtures. In some embodiments, the selected tumor sample residues can be used to create each mixture. A random value is added to the gene expression with a small probability. For example, the probability of "Hf" is between 0 and A uniformly distributed random number ranging from "Mhl" to "Mhl" can be added to the expression of each gene.
[0277] [Table 8]
[0278] [Table 9]
[0279] Computational Complexity The machine learning models described herein may have tens of thousands, hundreds of thousands, or even millions of parameters. For example, the nonlinear regression model 304 may include at least As described herein, including with respect to FIGS. 2-6, at least 10,000 parameters computer, at least 100,000 parameters, or at least 1 million parameters Therefore, processing the data with a machine learning model such as a nonlinear regression model 304 is The process of creating them requires millions of calculations, even after they have been trained, and is However, it is practically impossible to do this mentally without a computer.
[0280] Algorithms for training machine learning models such as those described herein are also widely used. This may require a large amount of computational resources, since such models may be required in the tens of thousands, This is because they are trained using hundreds of thousands or millions of artificial mixtures (e.g., at least (See also FIG. 6A and described herein). In one specific example, a two-stage Three million artificial mixtures were generated to train a nonlinear regression model over time. (eg, as described herein, including with respect to at least FIG. 5A). Without the computational resources, neither the training algorithm nor the use of the trained model can be implemented. I can't come.
[0281] result The following description describes various methods that have been achieved using the techniques developed by the present inventors. The results are described with respect to Figures 7A-7G. The method developed here substantially outperforms previous methods for cellular deconvolution. In the figure, the cellular deconvolution developed by the present inventors is shown. This technique is sometimes called "Cassandra."
[0282] FIG. 7A illustrates simulated RNA expression data 702 (e.g., generated according to the methodology of FIG. 6A). RNA expression data 704 from multiple biological samples (e.g., tumors) are compared with multiple artificial mixtures of RNA. In the illustrated example, the RNA expression data 702 includes the data related to FIG. The results were obtained from 500 artificial lung cancer samples derived using the techniques described herein. In comparison, RNA expression data 704 is from 500 non-small cell lung cancer RNA-seq data from TCGA. As shown in the illustrated example, the data includes gene expression patterns from artificial mixtures. The gene expression patterns for the IL-16 and actual tumors are quite similar. As a result, the correlation between real and artificial tumors reached 0.9 (p=0.001).
[0283] Figure 7B shows the predicted values according to the deconvolution method developed by the present inventors. 1 is a chart depicting exemplary cellular composition ratios and corresponding true cellular composition ratios. In the illustrated example, the performance of the deconvolution method developed by the inventors is , measured as the Pearson correlation for the artificial mixture of holds (e.g., for Figure 5A ). As shown, the correlation is The correlations for types are greater than 0.94, and for multiple cell types are greater than 0.98 (p=0).
[0284] Figures 7C and 7D show the predicted artificial mixture values versus the true artificial mixture values for different cell types. 1 is an exemplary chart showing the Pearson correlation between mixture values (e.g., prediction accuracy). These graphs are illustrative of the deconvolution method developed by the inventors. Figure 7C shows the predicted accuracy of the cancer cell lineage. Figure 7D shows the prediction accuracy without high expression noise. The prediction accuracy in some cases is presented.
[0285] Random overexpression, as described herein, including with respect to at least FIG. 6A Noise can be added to the artificial mixture (e.g., the data developed by the inventors). The convolution approach allows aberrant expression from malignant cells in the sample to be ignored. To generate accurate high-expression noise, we used TCGA data from four different cancer types. Four exemplary gene markers in tumors: CD14 in bladder cancer, F in cutaneous melanoma We analyzed CRLA, STAP1 in clear cell renal cell carcinoma, and PADl2 in squamous cell carcinoma of the lung. Each of these markers was found to be highly expressed in the corresponding type of cancer. These markers are not expressed in the corresponding normal tissues, but are expressed in immune cells. It has been found to be expressed (Fig. 7E).
[0286] As a result, the deconvolution method developed by the inventors is able to extract the As shown in Figures 7C and 7D, the present inventors The method developed by this study allows for the detection of chromatin fragments across multiple cell types, even in the presence of high expression noise. This produces accurate predictions (Figure 7D). Furthermore, Figure 7D shows that the performance of alternative algorithms is significantly higher. The presence of noise significantly reduces the detection accuracy. This shows that the correlation score remains high in the evidence dataset.
[0287] Alternative algorithms include CIBERSORT, CIBERSORTx, QuanTISeq, FARDEEP, Xcell, and AB These include IS, EPIC, MCP-Counter, Scaden, and MuSiC. Newman et al. ("Robust enumera tion of cell subsets from tissue expression profiles,” Nat. Methods 12, 453–45. 7 (2015)) describes CIBERSORT. Newman et al. ("Determining cell type ab undance and expression from bulk tissues with digital cytometry”, Nat Biotechno 37, pp. 773-782 (2019)) describes CIBERSORTx. ar and pharmacological modulators of the tumor immune context revealed by dec "Convolution of RNA-seq data," Genome Med 11, 34 (2019) describes QuanTIseq. Hao et al. ("Fast and Robust Deconvolution of Tumor Infiltrating Lymphocytic Cells") te from Expression Profiles using Least Trimmed Squares”, bioRxiv page 358366; doi : https: / / doi.org / 10.1101 / 358366) describes FARDEEP. Aran et al. ("xCell : digitally portraying the tissue cellular heterogeneity landscape”, Genome Bio I. 18, pp. 220 (2017)) describes the X cell. Monaco et al. ("RNA-Seq signatures" normalized by mRNA abundance allow absolute deconvolution of human immune cells types,” Cell Rep. 26, pp. 1627–1640. e1627 (2019)) describes ABIS.
[0288] Figure 7F shows the different fine grained images for the deconvolution method developed by the inventors. Predicted artificial mixture values versus true artificial mixture values for cell types (e.g., prediction accuracy) This is a heatmap showing the Pearson correlation between the For data from selected samples, predicted cell proportions for different cell types are shown. As shown, the deconvolution method developed by the present inventors , achieving high prediction accuracy scores across multiple cell types, including closely related cell types.
[0289] FIG. 7G shows an exemplary deconvolution method developed by the inventors. A chart comparing the non-specificity scores with those of alternative algorithms In the illustrated example, the non-specificity scores of 11 alternative algorithms are shown. The chart values in Figure 7G are relative to specific (true positive) predictions for several different cell types. The nonspecific score represents the proportion of nonspecific (false positive) predictions that A lower ratio of predictions indicates a more specific model (e.g., a pure The detection of signals for each cell type in the population was assessed, including B cells, T cells, and myeloid cells. Further subdivision of the macrophages revealed that each subclass was clearly distinct from the others. was done.
[0290] Linear methods for deconvolution According to some embodiments of the techniques developed by the present inventors, cellular deconvolution An exemplary linear deconvolution technique is shown in FIG. 8 and 9A-9C herein below.
[0291] FIG. 8 illustrates a method for determining cellular composition ratios based on expression data (e.g., RNA expression data). 8 is a flowchart depicting an exemplary linear method 800. , method 800 generates an expression profile (e.g., RNA expression as shown in FIG. 9A, and / or expression profiles) for one or more cell types in the biological sample. The method may include a step of estimating the proportion of all cells in the sample.
[0292] In some embodiments, the method 800 may be performed on a computing device (e.g., , as described herein, including at least with respect to FIG. 10). The computer device includes at least one processor and a processor that, when executed, performs the operations of method 800. and at least one non-transitory storage medium storing processor-executable instructions for Method 800 may be implemented, for example, in a system such as system 100 (including, for example, in a clinical setting or may include a laboratory setting), by one or more computing devices, e.g. This can be done by the computing device 108.
[0293] At operation 802, the method 800 begins by obtaining RNA expression data of a biological sample from a subject. In some embodiments, operation 802 may include analyzing RNA expression data previously obtained from a biological sample. As described herein, including with respect to FIG. 1A , The biological sample may include a biopsy sample (e.g., of a subject's tumor or other diseased tissue) or any other Any suitable type of biological sample can be included, and expression data can be obtained using any suitable technique. The expression obtained in operation 802 includes RNA expression data measured by TPM. In some embodiments, the source or preparation of the biological sample may include the method described in the "Biological Sample" section. In some embodiments, the expression data may include any of the embodiments described above. The origin or method of generation of the data is described in the sections "Expression Data" and "Obtaining RNA Expression Data." Any of the described embodiments may be included.
[0294] In some embodiments, the expression data is stored in at least one storage medium and For example, expression data can be stored in one or more files, or accessed as part of a can be stored in a database and then retrieved as part of operation 802. In embodiments, the at least one storage medium storing the expression data is a computer device. may be local to the storage device (e.g., stored on the same or at least one non-transitory storage medium) stored on the computer device) or may be external to the computer device (e.g., in a remote database) (The expression data is stored in a single storage medium.) The data may be stored in one storage medium or may be distributed across multiple storage media.
[0295] In some embodiments, operation 802 includes preprocessing the expression data in any suitable manner. For example, the expression data may be sorted, combined, or organized into batches. ing, filtering, or pre-processing in any other suitable manner. By preprocessing, the expression data can be processed as described in this specification, including those related to operations 804 to 806. The linear regression technique described in
[0046] may be suitable for processing. In some embodiments, pre-processing of RNA includes "alignment and annotation," "non-coordinated sequencing," and "non-coordinated sequencing." "Removal of gene transcripts," and "Conversion to TPM and gene assembly" sections are described. It may include any of the embodiments.
[0296] As described herein with respect to operations 804 through 806, method 800 may include: 4. Analyze RNA expression using linear regression techniques to determine the proportion of one or more cellular components that correspond to the RNA expression You can then proceed to process the data.
[0297] At operation 804, the method 800 generates a plurality of expression profiles (e.g., For example, one can proceed to obtain a If method 800 is used to analyze CD4+ T cells, NK cells, and CD8+ T cells, operation 802 Expression profiles for NK cells, and CD8+ T cells Expression profiles (e.g., RNA expression profiles) can be obtained for Each of the files (files) contains one or more genes associated with each cell type from a plurality of cell types. In some embodiments, each of the offspring may include expression data (e.g., RNA expression data). The genes associated with each cell type are specific and / or semi-specific genes for that cell type. For example, the genes associated with each cell type are listed in Table 2. In some embodiments, the corresponding genes include those included in Table 2. At least 2 genes, At least 4 genes, At least 6 genes, At least 8 genes, at least 10 genes, at least 12 genes, at least 14 genes In some embodiments, the corresponding genes may include at least 16 genes. is less than 10,000 units, less than 5,000 units, less than 2,000 units, less than 1,000 units, less than 500 units, less than 250 units, or Fewer than 100 genes may be included.
[0298] The expression profile can be obtained in any suitable manner. For example, 8. Save the file in one or more files or in a database as part of operation 804. In some embodiments, the expression profile can be stored in a small area. At least one storage medium may be local to the computing device (e.g., , stored on the same or at least one non-transitory storage medium), or on a computer device It may be external to the database (e.g., stored in a remote database or cloud storage environment). The expression profiles may be stored in a single storage medium or in multiple storage media. It may be dispersed throughout the medium.
[0299] At operation 806, the method 800 proceeds, at least in part, by combining the expression data and the plurality of expression profiles. , and a piecewise continuous error function between (e.g., the exemplary piecewise continuous error function described with respect to FIG. 9A ). By optimizing a continuous error function, multiple cell composition ratios for multiple cell types are obtained. Operation 806 can proceed to determining the ratio across multiple cell types simultaneously or It may be performed iteratively, and in some embodiments (e.g., for a set number of iterations). The method may be repeated until the error measurement is below a threshold.
[0300] According to some embodiments, operation 806 comprises: In some embodiments, the method may include performing a linear regression using a piecewise continuous error function. In some embodiments, this may involve optimizing a piecewise continuous error function. The process of optimizing a piecewise continuous error function is a global optimization of the piecewise continuous error function. It is not limited to finding maximum or minimum values, but also to finding local maximums or minimums within a threshold distance of the global maximum or minimum. For example, operation 806 may involve finding a particular maximum or minimum value for the expression data. and a minimum error or error below a threshold (e.g., measured using a piecewise continuous error function). determining a combination (e.g., a weighted sum) of expression profiles with a given error This may include the process.
[0301] For a particular cell type, operation 806 performs a piecewise linear regression for each gene associated with that cell type. determining the corresponding output of a continuous error function (e.g., the error function of FIG. 9C) A piecewise continuous error function is a practical measurement from real data (e.g., RNA-seq data). Using the expression values obtained and the expression of genes in the expression profile for that cell type, It may be useful to compare the predicted expression values calculated using the quantification method (e.g., obtained in operation 804) with the predicted expression values calculated using the quantification method (e.g., obtained in operation 804). For example, predicted expression values are calculated based on the expression of genes in the expression profile and the corresponding expression values for that cell type. It can be calculated as the product of the coefficient α.
[0302] For a given gene and cell type, the input to the error function is a coefficient α, the The expression of genes in the cell type (g) and the expression of genes in the expression profile for that cell type (p) The error function can be updated as part of operation 806, as shown in FIG. The coefficients a, b, and k may be as described herein, including those related to the According to an embodiment, operation 806 is performed iteratively or in parallel for some or all of the genes. For example, a piecewise continuous error function can be minimized or below a threshold. Operation 806 is repeated across multiple cell types until a coefficient α is found for each cell type such that According to some embodiments, for a given cell type, the coefficient The value of α is the weighted error sum across all genes (e.g., The calculated piecewise error function is described herein, including with respect to operation 806 and FIG. 9C. The coefficients can be determined by finding the coefficient values that minimize
[0303] In some embodiments, the coefficient α may represent a cellular composition ratio for the corresponding cell type. (For example, α defines the weight of each expression profile in the weighted sum of expression data.) For example, the step of determining multiple cell composition ratios for multiple cell types may involve multiple processing the coefficients to obtain corresponding cell composition ratios for each of the cell types, e.g. For example, it may include a normalization step.
[0304] FIG. 9A is a diagram depicting exemplary RNA expression profiles and global RNA expression data. The illustrated examples include known RNA expression profiles for CD4+ T cells, NK cells, and CD8+ T cells. The current profiles are shown. Each RNA expression profile has genes on the horizontal axis and are shown as bar graphs representing the expression of those genes. , each RNA expression profile can be unique for a given cell type.
[0305] As shown in the illustrated example, the overall expression observed for a biological sample is The expression profiles of the various cell types that make up the chromosome can be considered as the sum of the expression profiles of the various cell types that make up the chromosome. Although not specified, each RNA expression profile can be weighted by a coefficient α, so that the biological sample is According to some embodiments, this sum can be considered as a weighted sum of the NA expression profiles. The expression level can further include terms for unknown expression in other cell types. It can represent expression data that is not explained by the weighted sum of the current profile (e.g., As shown in grey in the expression observed for the sample).
[0306] FIG. 9B depicts an exemplary piecewise continuous error function for use in the method of FIG. 8. As shown in the plot, the error function f is piecewise and the coefficients a and b are related. The coefficient k affects the shape of the rightmost interval of the error function. For the intervals, the error can be calculated according to the illustrated expression.
[0307] biological samples Any method, system, or other claimed element uses a biological sample from a subject. In some embodiments, the biomolecules can be used to measure or analyze the biomolecules. trial The fee is for subjects who have cancer, are suspected of having cancer, or are at risk of having cancer. The biological sample may be obtained from a subject, for example, a biological sample of a bodily fluid (e.g., blood, urine, or brain). spinal fluid), one or more cells (e.g., oral mucosal specimen or tracheal brushing) scraping or brushing), tissue fragments (cheek tissue, muscle tissue, lung tissue, heart tissue, brain tissue, or skin tissue), or organs (brain, lung, liver, bladder, kidney, pancreas, intestine, or Any type of biological sample containing part or all of a tissue (e.g., muscle) or any other type of biological sample (e.g., For example, feces or hair).
[0308] In some embodiments, the biological sample is a tumor sample from the subject. In some embodiments, the biological sample is a blood sample from the subject. A tissue sample.
[0309] A tumor sample, in some embodiments, refers to a sample containing cells from a tumor. In some embodiments, the tumor sample includes cells from a benign tumor, e.g., non-cancerous cells. In some embodiments, the tumor sample includes cells from a precancerous tumor, e.g., a precancerous tumor. In this embodiment, the tumor sample comprises cells from a malignant tumor, e.g., cancer cells.
[0310] Examples of tumors include adenomas, fibromas, hemangiomas, lipomas, cervical dysplasia, pulmonary metaplasia, and leukoplakia. Cancers include, but are not limited to, cancers, carcinomas, sarcomas, germ cell tumors, and blastomas.
[0311] In some embodiments, the sample of blood comprises cells, e.g., a sample comprising cells from a blood sample. In some embodiments, the blood sample comprises non-cancerous cells. The fluid sample contains precancerous cells. In some embodiments, the blood sample contains cancer cells. In some embodiments, the sample of blood comprises blood cells. In some embodiments, the sample of blood comprises white blood cells. The fluid sample includes platelets. Examples of cancerous blood cells include leukemia, lymphoma, and myeloma. In some embodiments, the blood sample includes, but is not limited to, The sample is harvested to obtain cell-free nucleic acid (e.g., cell-free DNA).
[0312] The blood sample may be a whole blood sample or a fractionated blood sample. The sample of blood comprises whole blood. In some embodiments, the sample of blood comprises fractionated blood. In embodiments, the sample of blood comprises a buffy coat. In some embodiments, the sample of blood comprises serum. In some embodiments, the sample of blood comprises plasma. In some embodiments, the sample of blood comprises a blood clot. Includes:
[0313] A tissue sample, in some embodiments, refers to a sample containing cells from a tissue. In some embodiments, the tumor sample comprises non-cancerous cells from the tissue. The material includes precancerous cells from the tissue.
[0314] The disclosed methods can be used to treat muscle tissue, brain tissue, lung tissue, liver tissue, epithelial tissue, connective tissue, and nerve tissue. Various tissues, including organ and non-organ tissues, including but not limited to organ tissues. In some embodiments, the tissue may be normal tissue, diseased tissue, or In some embodiments, the tissue may be a tissue section or a whole tissue. In some embodiments, the tissue is animal or human tissue. Animal tissues may include rodent (e.g., rat or mouse), primate (e.g., monkey) Examples of tissues that may be used include, but are not limited to, tissues obtained from animals such as cats, dogs, cats, and farm animals.
[0315] Biological samples include, but are not limited to, any bodily fluid, such as blood (e.g., whole blood, serum, or (or plasma), saliva, tears, synovial fluid, cerebrospinal fluid, pleural fluid, pericardial fluid, peritoneal fluid, and / or urine], Hair, skin (including the epidermis, dermis, and / or hypodermis), oropharynx, laryngopharynx, esophagus, stomach , bronchi, salivary glands, tongue, oral cavity, nasal cavity, vaginal cavity, anal cavity, bones, bone marrow, brain, thymus, spleen, small intestine, Appendix, colon, rectum, anus, liver, biliary tract, pancreas, kidneys, ureters, bladder, urethra, uterus, vagina, vulva the genitals, ovaries, cervix, scrotum, penis, prostate, testicles, seminal vesicles, and / or any type of tissue ( any tissue within the subject's body, including, for example, muscle tissue, epithelial tissue, connective tissue, or nervous tissue It may be from a supply source.
[0316] Any of the biological samples described herein can be obtained from a subject using any known technique. For example, the collection, processing, and storage of biological samples can each be entirely See the following publications, which are incorporated herein in their entirety: Vaught et al., "Bi ospecimens and biorepositories: from afterthought to science” (Cancer Epidemiol) Biomarkers Prev. 2012 Feb;21(2):253-5), and Vaught and Henderson, "Biol ological sample collection, processing, storage and information management” (IARC Sci Publ. 2011;(163):23~42).
[0317] In some embodiments, the biological sample is obtained from a surgical procedure (e.g., laparoscopic surgery, microsurgical surgery, or endoscopic surgery), bone marrow biopsy, punch biopsy, endoscopic biopsy, or needle biopsy (e.g., fine needle aspiration) Biopsy can be obtained by a variety of techniques, including core needle biopsy, vacuum-assisted biopsy, or image-guided biopsy.
[0318] In some embodiments, one or more cells (i.e., a cellular biological sample) are obtained by scraping or brushing. Cellular biological samples can be obtained from subjects using methods such as the cervical, esophageal, stomach, any area in or from the body of a subject, including one or more areas of the bronchial, or oral cavity In some embodiments, one or more tissue fragments (e.g., tissue fragments) from a subject can be obtained from the tissue fragments. For example, a tissue biopsy sample can be used. In certain embodiments, the tissue biopsy sample can be , known to have cancerous cells or suspected to have cancerous cells, One or more (e.g., 2, 3, 4, 5, 6, 7) tumors or tissues , 8, 9, 10, or more than 10 biological samples.
[0319] Any of the biological samples from a subject described herein may be prepared using any suitable method to maintain the stability of the biological sample. In some embodiments, the biological sample may be stored using any method that maintains its stability. is a component of a biological sample (e.g., DNA, RNA, protein, or tissue structure or morphology) , so that the measured value represents the state of the sample when it was obtained from the subject. In some embodiments, the biological sample is The components of the biological sample (e.g., DNA, RNA, proteins, or tissue structures or The composition is stored in a manner that can protect the product (or its form) from deterioration. When used, degradation occurs at a rate such that the original form is no longer detectable at the same level as before degradation. It is the transformation of one component into another.
[0320] In some embodiments, the biological sample (e.g., tissue sample) is fixed. When a sample is "fixed," it is meant to prevent or mitigate decay or deterioration, such as autolysis or putrefaction of the sample. It relates to a sample that has been treated with one or more agents or processes to reduce Examples of fixation processes include, but are not limited to, heat fixation, immersion fixation, and perfusion. In some embodiments, the sample to be fixed is treated with one or more fixatives. Examples of fixing agents include cross-linking agents (e.g., formaldehyde, formalin, glutaraldehyde, etc.) aldehydes such as aldehydes), precipitants (e.g., ethanol, methanol, acetone, xylene, etc.), Alcohol), mercury (e.g., B-5, Zenker's fixative, etc.), picric acid, Hepes-glutamine These include, but are not limited to, acid buffer-mediated organic solvent protective effect (HOPE) fixatives. In some embodiments, the biological sample (e.g., a tissue sample) is treated with a cross-linking agent. In embodiments, the cross-linking agent comprises formalin. In some embodiments, formalin-fixed The biological sample is embedded in a solid substrate, for example, paraffin wax. The biological samples are formalin-fixed, paraffin-embedded (FFPE) samples. The method is known and is described, for example, by Li et al., JCO Precis Oncol. 2018; 2: PO.17.00091. It is listed.
[0321] In some embodiments, the biological sample is stored using cryopreservation. Typical examples include, but are not limited to, step-down freezing, quick freezing, direct plunge freezing, and These include nap freezing, slow freezing using a programmable freezer, and vitrification. In some embodiments, the biological sample is stored using lyophilization. After collection of the biological sample from the subject, the biological sample may be treated with a preservative (e.g., a preservative for preserving RNA). placed in a container already containing RNALater and then frozen (e.g., by snap freezing) In some embodiments, such storage in a frozen state is carried out immediately after collection of the biological sample. In some embodiments, the biological sample is frozen in or without a preservative. in a buffer solution containing HCl for a period of time (e.g., up to 1 hour, up to 8 hours, or up to 1 day, or The solution can be kept at either room temperature or 4°C for several days.
[0322] Non-limiting examples of preservatives include formalin solution, formaldehyde solution, RNALater or other equivalent solution, TriZol or other equivalent solution, DNA / RNA Shield or equivalent solution, EDTA (e.g. For example, Buffer AE (10 mM Tris-Cl, 0.5 mM EDTA, pH 9.0) and other coagulants, as well as Citric Acids te Dextronse (e.g., for blood samples).
[0323] In some embodiments, specialized containers are used to collect and / or store biological samples. For example, a vacutainer can be used to store blood. In some embodiments, the vacutainer may contain a preservative (e.g., a coagulant or anticoagulant). In some embodiments, the container in which the biological sample is stored may be modified for better preservation or to prevent contamination. It may be contained in a secondary container to avoid
[0324] Any biological sample from a subject described herein may be prepared using any suitable method that maintains the stability of the biological sample. In some embodiments, the biological sample may be stored under any conditions that maintain the stability of the biological sample. In some embodiments, samples are stored at room temperature (e.g., 25°C). In some embodiments, the sample is stored under refrigeration (e.g., at 4°C). In some embodiments, the sample is stored under refrigerated conditions (e.g., at -20°C). In some embodiments, the sample is stored under cryogenic conditions (e.g., −50° C. to −800° C.). In some embodiments, the biological sample is stored at temperatures between -60°C and -1700°C. -80°C (e.g., -70°C) for up to 5 years (e.g., up to 1 month, up to 2 months, up to 3 months) , up to 4 months, up to 5 months, up to 6 months, up to 7 months, up to 8 months, up to 9 months, Up to 10 months, up to 11 months, up to 1 year, up to 2 years, up to 3 years, up to 4 years, or up to In some embodiments, the biological sample is stored for up to 5 years. For up to 20 years (e.g., up to 5 years, up to 10 years, up to It is stored for up to 15 years, or up to 20 years.
[0325] The methods of the present disclosure include obtaining one or more biological samples from a subject for analysis. In some embodiments, one biological sample is obtained from the subject for analysis. Now, multiple (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13) 1, 14, 15, 16, 17, 18, 19, 20 or more biological samples are In some embodiments, a single biological sample from a subject is analyzed. In some embodiments, multiple (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more A biological sample is analyzed. If multiple biological samples from a subject are analyzed, the biological samples are analyzed simultaneously. The biological samples may be procured separately (e.g., multiple biological samples may be collected in the same procedure), or the biological samples may be At different time points (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 days after the first procedure, 1, 2 , 3, 4, 5, 6, 7, 8, 9, 10 weeks after procedure, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 months after procedure , 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 years after the procedure, or 10, 20, 30, 40, 50, 60, 70, 80, 9 can be collected at different times, including 0, 100, or 200 years later.
[0326] A second or subsequent biological sample may be taken from the same area (e.g., from the same tumor or tissue area). Or they may be taken or obtained from different regions (eg, containing different tumors). The second or subsequent biological sample may be taken from the subject after one or more treatments, or may be obtained and may be taken from the same or different regions. The two or subsequent biosamples may be compared to determine whether the cancer in each biosample has different characteristics (e.g., , in the case of biological samples taken from two physically separate tumors in a patient), or whether the patient responded to one or more treatments (e.g., the same tumor or lesion before and after treatment) In the case of two or more biological samples taken from different tumors, In some embodiments, each of the at least one biological sample is a body fluid sample, a cell sample, or a , or a tissue biopsy sample.
[0327] In some embodiments, one or more biological specimens are combined (e.g., For example, the first tumor from the subject may be placed in the same container for storage. The sample can be combined with a second sample of a second tumor obtained from the subject, and the first and The second tumor may or may not be the same tumor. In some embodiments, the first tumor and the second tumor The tumors may be similar but not identical (e.g., two tumors in a subject's brain). In an embodiment, the first biological sample and the second biological sample from the subject are samples from different tumor types. (e.g., tumors in muscle tissue and tumors in brain tissue).
[0328] In some embodiments, the sample from which RNA and / or DNA is extracted (e.g., a tumor sample, or blood A sample of at least 2 μg (e.g., at least 2 μg, at least 2.5 μg, at least The size of the sample should be large enough that at least 3 μg, at least 3.5 μg or more of RNA can be extracted. In some embodiments, the sample from which RNA and / or DNA is extracted is peripheral blood mononuclear cells (PBMCs). In some embodiments, the sample from which RNA and / or DNA is extracted may be any type of In some embodiments, the sample from which RNA and / or DNA is extracted (e.g., For example, a tumor sample or a blood sample should be prepared so that at least 1.8 μg of RNA can be extracted from it. In some embodiments, the amount is at least 50 mg (e.g., at least 1 mg). , at least 2 mg, at least 3 mg, at least 4 mg, at least 5 mg, at least 10 mg, At least 12 mg, at least 15 mg, at least 18 mg, at least 20 mg, at least 22 mg, at least 25 mg, at least 30 mg, at least 35 mg, at least 40 mg, at least 45 mg, or A tissue sample (of at least 50 mg) is taken from which RNA and / or DNA is extracted. In this embodiment, a tissue sample of at least 20 mg is taken, from which RNA and / or DNA is extracted. In some embodiments, at least 30 mg of tissue sample is collected. In some cases, at least 10 to 50 mg (e.g., 10 to 50 mg, 10 to 15 mg, 10 to 30 mg, 10 to 40 mg, 20 to 30 mg) A tissue sample (0 mg, 20-40 mg, 20-50 mg, or 30-50 mg) is collected from which RNA and / or DNA is extracted. In some embodiments, at least 30 mg of tissue sample is taken. In this embodiment, at least 20-30 mg of tissue sample is collected and RNA and / or DNA are extracted therefrom. In some embodiments, the sample from which RNA and / or DNA is extracted (e.g., tumor The sample, or blood sample) should contain at least 0.2 μg (e.g., at least 200 ng, at least 300 ng) , at least 400ng, at least 500ng, at least 600ng, at least 700ng, at least 800ng, at least 900ng, at least 1μg, at least 1.1μg, at least 1.2μg, at least at least 1.3μg, at least 1.4μg, at least 1.5μg, at least 1.6μg, at least 1. 7 μg, at least 1.8 μg, at least 1.9 μg, or at least 2 μg) of RNA is extracted therefrom In some embodiments, the RNA and / or DNA can be extracted. The sample (e.g., tumor sample or blood sample) to be analyzed should contain at least 0.1 μg (e.g., at least 1 00ng, at least 200ng, at least 300ng, at least 400ng, at least 500ng, less At least 600ng, at least 700ng, at least 800ng, at least 900ng, at least 1μg, at least 1.1 μg, at least 1.2 μg, at least 1.3 μg, at least 1.4 μg, at least 1.5 μg, at least 1.6 μg, at least 1.7 μg, at least 1.8 μg, at least 1.9 μg, or at least 2 μg) of RNA can be extracted.
[0329] subject Aspects of the present disclosure relate to a biological sample obtained from a subject. In some embodiments, the subject , mammals (e.g., humans, mice, cats, dogs, horses, hamsters, cows, pigs, or other In some embodiments, the subject is a human. In some embodiments, the subject is a In some embodiments, the subject is an adult (e.g., 18 years of age or older). In some embodiments, the subject is a child (e.g., 18 years of age or older). In some embodiments, the human subject has at least one form of cancer. or a person who has been diagnosed with at least one form of cancer. The cancer from which the subject is suffering is carcinoma, sarcoma, myeloma, leukemia, lymphoma, or Carcinoma is a mixed type of cancer that includes more than one of the following: myeloma, leukemia, and lymphoma. Sarcomas are malignant neoplasms of epithelial origin or cancers of the lining or outer membranes of the body. It refers to cancer that originates in the supportive and connective tissues, such as muscle and fat. Myeloma is a type of cancer that develops in the bone marrow. Leukemia ("liquid cancer" or "blood cancer") is a cancer that originates in the bone marrow (blood cell-producing cells). Lymphoma is a cancer of the white blood cells, or lymphocytes, which cleanse the body's fluids and fight infection. The glands of the lymphatic system, a network of blood vessels, nodes, and organs (especially the spleen, tonsils, and thymus) that produce lymphatic fluid Occurring in nodules. Non-limiting examples of mixed cancers include adenosquamous carcinoma, mixed mesodermal tumor, In some embodiments, the subject has a tumor. The tumor may be: In some embodiments, the cancer is selected from the group consisting of skin cancer, lung cancer, breast cancer, and pre-cancerous carcinoma. Any one of the following cancers: prostate cancer, colon cancer, rectal cancer, cervical cancer, and uterine cancer In some embodiments, the subject is diagnosed with a genetic disorder, e.g., the subject has one or more genetic risk factors. or exposure to one or more carcinogens (e.g., tobacco smoke or chewing tobacco). Are at risk of developing cancer because they have been exposed to or have been exposed to do.
[0330] Expression data Expression data (e.g., indicative of expression levels) of a plurality of genes can be analyzed by the methods described herein. The number of genes that can be examined can be determined by the number of genes of interest. In some embodiments, the number of genes of interest may be up to an inclusive number. As a non-limiting example, four or more, five or more genes can be used to examine the expression level. or more, 6 or more, 7 or more, 8 or more , 9 or more, 10 or more, 11 or more, 12 or or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more Above, 20 or more, 21 or more, 22 or more, 23 or less or more, 24 or more, 25 or more, 26 or more, 27 10 or more, 28 or more, 29 or more, 30 or more or more, 35 or more, 40 or more, 50 or more, 60 or less or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or less or more, 200 or more, 225 or more, 250 or more , 275 or more, or 300 or more genes, as described herein. As another set of non-limiting examples, expression data can be used to assess any The data includes, for each cell type listed in Table 2, the At least 5, at least 10, at least 1 gene selected from a group of genes for cell type 5, at least 20, at least 25, at least 35, at least 50, at least 7 Expression data for at least 100 genes can be included.
[0331] To obtain expression data (e.g., indicative of expression levels) for a plurality of genes, Any method can be used on samples taken from. A non-limiting set of examples are: The expression data may be RNA expression data, DNA expression data, or protein expression data. obtain.
[0332] In some embodiments, the DNA expression data refers to the level of DNA in a sample from a subject. The level of DNA in a sample from a subject having cancer, e.g., If the level is elevated compared to the level in DNA from a subject who does not have the gene duplication, The level of DNA in a sample from a subject with cancer may be a measure of the level of cancer, e.g., cancer patient reduced levels of DNA in a sample from a subject without the gene depletion compared to levels in a sample from a subject without the gene depletion This may be the case.
[0333] In some embodiments, the DNA expression data is about the expressed DNA (or genes) in the sample. data, e.g., sequencing data for expressed genes in patient samples. Such data may, in some embodiments, indicate that a patient has one or more of the following characteristics associated with a particular cancer: can be useful in determining whether a gene has multiple mutations.
[0334] RNA expression data can be obtained from any of a variety of methods known in the art, including, but not limited to: It can be obtained using any method: whole transcriptome sequencing, Total RNA sequencing, mRNA sequencing, targeted RNA sequencing, small RNA sequencing sequencing, ribosome profiling, RNA exome capture sequencing, and / or deep RNA sequencing. DNA expression data may be obtained from any publicly available DNA sequencing The protein can be obtained using any method known in the art, including known methods. For example, DNA sequencing can be used to identify one or more mutations in a subject's DNA. Any technique used in the art for sequencing DNA is referred to herein as As a non-limiting set of examples, the methods and compositions described in A is single molecule real-time sequencing, Ion Torrent sequencing, Pyrosequencing sequencing, sequencing by synthesis, sequencing by ligation (SOLiD) sequencing), nanopore sequencing, or Sanger sequencing (chain transcription) Proteins can be sequenced by cleavage or cleavage. Expression data can be obtained from any of the methods known in the art, including, but not limited to: It can be obtained using the following methods: N-terminal amino acid analysis, C-terminal amino acid analysis, Doman degradation (including by using machines such as protein sequencing machines), or Quantitative analysis method.
[0335] In some embodiments, the expression data is obtained by bulk RNA sequencing. Bulk RNA sequencing is a method for sequencing RNA extracted from multiple input cell populations. The method can include obtaining expression levels of a plurality of genes, the population including a plurality of different cells. In some embodiments, the expression data may be obtained by single-cell sequencing (e.g., Single-cell sequencing involves sequencing individual cells. This may include:
[0336] In some embodiments, the expression data comprises whole exome sequencing (WES) data In some embodiments, the expression data comprises whole genome sequencing (WGS) data. In some embodiments, the expression data comprises next generation sequencing (NGS) data. In an embodiment, the expression data comprises microarray data.
[0337] Obtaining RNA expression data In some embodiments, RNA expression data (e.g., data obtained from RNA sequencing ( The method for processing RNA-seq data (also referred to herein as RNA-seq data) involves analyzing a subject (e.g., a person with cancer) The method includes obtaining RNA expression data for a subject (or a subject diagnosed with cancer). In some embodiments, obtaining RNA expression data includes obtaining a biological sample and analyzing the RNA sequence as described herein. 2. Processing the RNA for sequencing using one of the following sequencing methods: In some embodiments, the RNA expression data is obtained by performing an experiment to obtain the RNA expression data. Some samples are obtained from the laboratory or facility where they were performed (e.g., the laboratory or facility where they performed RNA-seq). In an embodiment, the laboratory or facility is a medical laboratory or medical facility.
[0338] In some embodiments, the RNA expression data is stored in a computer storage medium (e.g., a hard disk drive) on which the data resides. In some embodiments, the RNA expression data is obtained by obtaining the RNA expression data. are obtained via a secure server (e.g., an SFTP server or Illumina BaseSpace) In some embodiments, the data is stored in a text-based file (e.g., a FASTQ file). In some embodiments, the file in which the sequencing data is stored is The template also includes a quality score for the sequencing data. The file in which the sequencing data is stored also contains sequencing identifier information.
[0339] Alignment and annotation In some embodiments, RNA expression data (e.g., data obtained from RNA sequencing ( The method for processing the RNA expression data (also referred to herein as RNA-seq data) involves identifying genes in the RNA expression data. The annotation is performed by aligning the sequence with the known sequence of the human genome. The method includes obtaining segmented RNA expression data.
[0340] In some embodiments, alignment of RNA expression data aligns the data to a specific species of interest. Known assembled genome (e.g., human genome) or transcriptome data A variety of sequence alignment software programs are available. available and the data are compiled into assembled genome or transcriptome databases. Alignment software can be used to align a sequence to a target sequence. Non-limiting examples of aligners include short (unspliced) aligners (e.g., BLAT; BFAS T, Bowtie, Burrows-Wheeler Aligner, Short Oligonucleotide Analysis package, or Mosaik), spliced aligners, aligners based on known splice junctions (e.g. e.g., Errange, IsoformEx, Splice Seq), or de novo splice aligners (e.g., A In some embodiments, data alignment is performed using a Any suitable tool can be used to perform the analysis and annotation, for example Kallisto (github.com / pachterlab / kallisto) provides data alignment and annotation. In some embodiments, the known genome is referred to as a reference genome. A reference genome (also called a reference assembly) is a representative set of genes for a species. In some embodiments, the present invention provides a digital nucleic acid sequence database. The human and mouse reference genomes used in any one of the methods described herein are available from Genome The Human Reference Release is maintained and improved by the Global Reference Consortium (GRC). Non-limiting examples include GRCh38, GRCh37, NCBI Build 36.1, NCBI Build 35, and NCBI B uild 34. Non-limiting examples of transcriptome databases include the Examples include transcriptome shotgun assembly (TSA).
[0341] In some embodiments, annotation of RNA expression data is performed based on the data to be processed. The location of genes and / or coding regions in the assembled genome or Identifying the transcriptome by comparing it with a transcriptome database. Non-limiting examples of data sources for the application include GENCODE (www.gencodegenes.org), R efSeq (see, e.g., www.ncbi.nlm.nih.gov / refseq / ), and Ensembl. In this embodiment, the annotation of genes in the RNA expression data is based on the GENCODE database. The annotations are based on the GENCODE V23 annotations (e.g., www.gencodegenes.org).
[0342] Consea et al. (A survey of best practices for RNA-seq data analysis; Genome Biology2 01617:13) provides best practices for analyzing RNA-seq data, which is The present invention is applicable to any one of the methods described herein, and is incorporated by reference in its entirety. This document is incorporated herein by reference. Pereira and Rueda (bioinformatics-core-shared-training.github. io / cruk-bioinf-sschool / Day2 / rnaSeq_align.pdf) can also be used in conjunction with any of the methods described herein. We describe a method for analyzing RNA sequencing data that is applicable to one of the following: is incorporated herein by reference in its entirety.
[0343] Removal of non-coding transcripts In some embodiments, RNA expression data (e.g., data obtained from RNA sequencing ( The method for processing annotated RNA expression data (also referred to herein as RNA-seq data) is and removing non-coding transcripts from the data. The annotation allows for the identification of coding and non-coding reads. In this embodiment, the analytical effort is focused on identifying proteins (e.g., proteins potentially involved in cancer pathology). Remove non-coding reads from transcripts to focus on the expression of some In embodiments, removing non-coding transcript reads from the data can reduce the likelihood of mismatching, e.g., or the variance of data in replicate tests of similar samples (e.g., nucleic acids from the same cell or cell type). In some embodiments, non-limiting examples of expression data that are removed include: one or more non-coding transcripts belonging to one or more gene groups selected from the list consisting of (For example, 10-50, 50-100, 100-1,000, 1,000-2,500, 2,500-5,000 or more) These include: pseudogenes, polymorphic pseudogenes, and processing pseudogenes (more than 100 non-coding transcripts). , transcribed processed pseudogenes, unitary pseudogenes, non-processed pseudogenes, transcribed Transcribed unitary pseudogenes, constant chain immunoglobulin (IGC) pseudogene, and joining chain immunoglobulin lin (IG J) pseudogene, variable chain immunoglobulin (IG V) pseudogene, transcribed and unprocessed Pseudogene, translated non-processing pseudogene, T cell receptor binding chain (TRJ) pseudogene, possible mutant T cell receptor (TR V) pseudogene, small nuclear RNA (snRNA), small nucleolar RNA (snRNA), microRNA (miRNA), ribozyme, ribosomal RNA (rRNA), mitochondrial tRNA (Mt tRNA), Mitochondrial rRNA (Mt rRNA), Cajal body-specific small RNA (scaRNA), residual introns , sense intron RNA, sense overlap RNA, nonsense-mediated decay RNA, nonstop min Sense RNA, antisense RNA, long intervening non-coding RNA (lincRNA), macro-long non-coding RNA (macro lncRNAs), processed transcripts, 3' overlapping non-coding RNAs (3' overlapping ncrnas), small RNAs (sRNAs), and Other RNAs (misc RNAs), vault RNAs, and TEC RNAs.
[0344] In some embodiments, one or more of these types of transcripts may be used. Information (e.g., sequence information) of a number of transcripts is available from nucleic acid databases (e.g., the Gencode database). Gencode V23, the Genbank database, the EMBL database, or other databases. In some embodiments, the non-coding sequences described herein can be obtained at transcripts, histone-encoding genes, mitochondrial genes, and interleukin-encoding genes. genes encoding the IL-1 receptor, genes encoding collagen, and / or genes encoding T cell receptors Part of the child (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, or 99.5% or more) of the RNA expression data were aligned and annotated. is removed from the data.
[0345] Conversion to TPM and gene assembly In some embodiments, RNA expression data (e.g., data obtained from RNA sequencing ( The method for processing RNA-seq data (also referred to herein as RNA-seq data) involves determining the length of the transcripts read. Convert RNA expression data into (e.g., transcripts per kilobase per million (TPM) format) In some embodiments, the normalized RNA expression per transcript length is First, align and annotate the current data. Convert the data to TPM. This allows for the presentation of expression in the form of concentrations rather than counts, thereby providing a comprehensive This allows for comparison of samples with different read counts and read lengths.
[0346] In some embodiments, the RNA expression data is normalized per transcript read length. By analyzing the data, gene expression data (expression data for each gene) is obtained. Gene assembly is the sum of the reads for all isoform transcripts of a gene. The method includes combining expression data from the gene sequences to obtain expression data for the gene. In some embodiments, gene aggregation to obtain gene expression data is performed after TPM normalization. In some embodiments, the data are analyzed using a method that involves identifying genes that introduce bias. Gene assembly occurs before the conversion of TA to TPM.
[0347] Wagner et al. (Theory Biosci. (2012) 131:281-285) explain how TPM is calculated. provides a description of the method, which is incorporated herein by reference in its entirety. In some embodiments, the following formula is used to calculate the TPM:
[0348]
number
[0349] Computer implementation and sample processing environment Any of the embodiments of the techniques described herein (e.g., the methods of FIGS. 2, 4, and 6) An example implementation of a computer system 1000 that may be used with the computer system 1000, such as the computer system 1000 shown in FIG. The computer system 1000 includes one or more processors 1010 and a non-transitory computer readable medium. The system includes a removable storage medium (e.g., memory 1020 and one or more non-volatile storage media 1030). The processor 1010 includes one or more articles of manufacture including a memory 1020 and a non-volatile storage device. The writing and reading of data to device 1030 can be controlled in any suitable manner. However, this is because aspects of the technology described herein are not limited in this respect. To perform any of the functions described herein, the processor 1010 a non-transitory computer that stores processor-executable instructions for execution by processor 1010; One or more non-transitory computers that function as readable storage media (e.g., memory 1020). One or more processor-executable instructions stored on a computer-readable storage medium can be executed.
[0350] The computing device 1000 may also be configured to communicate with other computing devices. A network input / output (I / O) interface that can communicate with (e.g., over a network) a computer device providing output to a user, One or more user I / O interfaces that can receive input from the user The user I / O interface can also include a keyboard, mouse, microphone, display device (e.g., monitor or touch screen), speakers, camera, and and / or various other types of I / O devices.
[0351] In some embodiments, the techniques described herein may be implemented in the exemplary environment shown in FIG. 11, within the exemplary environment 1100, One or more biological samples from 1180 can be provided to the laboratory 1170. The laboratory 1170 The body sample is processed to obtain expression data (e.g., DNA, RNA, and / or protein expression data) and / or Alternatively, sequence information may be obtained and transmitted via the network 1110 to a subject (e.g., a patient) 1180. can be provided to at least one database 1160 storing the
[0352] The network 1110 may be a wide area network (e.g., the Internet), a local area network (LON), or a network (e.g., a corporate intranet), and / or any other suitable type of network. Any of the devices shown in FIG. 11 may be a workpiece. using a wire link, one or more wireless links, and / or any suitable combination thereof 1110.
[0353] In the illustrated embodiment of FIG. 11, at least one database 1120 stores information about the object (e.g., expression data or sequence information about a subject (e.g., a patient), medical history data about a subject (e.g., a patient), data, test result data about the subject (e.g., patient), and / or other assignments related to the subject 1180. Any suitable information may be stored. Test result data stored for a subject (e.g., a patient) Examples of data include biopsy results, imaging results (e.g., MRI results), and blood test results. The information stored in the at least one database 1120 may be stored in any suitable database. The data may be stored in any suitable format and / or structure, including any of the formats described herein. The present invention is not limited in this respect because the present invention is not limited in this respect. 120 may be implemented in any suitable manner (e.g., one or more databases, one or more files, ) can store data. At least one database 1120 can store a single data It may be a single database or multiple databases.
[0354] As shown in FIG. 11, the exemplary environment 1100 may include information about patients other than patient 1180. and one or more external databases 1120 in which the data can be stored. The database 1160 stores expression data and / or sequence information (e.g., , imaging test results, biopsy results, blood test results), medical history data of one or more patients, or multiple patient test result data, one or more patient demographic and / or Personal or biographical information, and / or any other suitable type of information may be stored. In some embodiments, the external database 1160 includes the Cancer Genome Atlas (TCGA), clinical trials, and one or more databases of experimental information and / or commercial sequencing suppliers. one or more publicly accessible databases, such as one or more databases maintained by The information available can be stored in a number of databases. 60 may store such information in any suitable manner using any hardware. Although aspects of the technology described herein are not limited in this respect, This is because.
[0355] In some embodiments, at least one of the databases 1120 and the external database 1160 , may be the same database or part of the same database system, or may be physically co-located, although this is not intended to be a substitute for the techniques described herein. This is because there is no limitation in this respect.
[0356] For example, in some embodiments, server 1140 may store data in databases 1120 and / or 1160. access information stored in the database and use this information to identify one of the biological samples and / or sequence information; or to determine a number of characteristics (e.g., determining its cellular composition), The process can be carried out as follows.
[0357] In some embodiments, server 1140 may include one or more computing devices. If the server 1140 includes multiple computing devices, the devices may be installed in the same physical location. They may be located in a single location (e.g., a single room) or may be distributed across multiple physical locations. In some embodiments, the server 1140 may be a cloud computing infrastructure. In some embodiments, one or more servers 1140 may be part of a , in the same place within a facility operated by an entity (e.g., a hospital, a research institute) to which the doctor 1150 belongs. In such an embodiment, the server 1140 may be located at a location where the patient 1180 It is likely that accessing all personal medical data will be easier.
[0358] As shown in FIG. 11, in some embodiments, the distribution performed by the server 640 The results of the analysis are then transmitted to a computing device 1130 (which may be a laptop or a smartphone). portable computing devices such as a MacBook, or fixed computers such as a desktop computer The results may be provided to the physician 1150 via a computer device. The information may be provided in writing, by email, through a graphical user interface, and / or through any other suitable means. In the embodiment of FIG. 11, the results are provided to a physician 1150. In other embodiments, the results of the analysis may be communicated to the patient 1180 or a medical professional, such as a caregiver, nurse, or other caregiver for the patient 1180. It should be understood that the information may be provided to a medical provider or clinical trial participant.
[0359] In some embodiments, the results are presented to a physician 1150 via a computing device 1130. In some embodiments, the display may be part of a graphical user interface (GUI) that is provided by the display. The GUI is displayed by a web browser running on the computing device 1130. In some embodiments, the GUI may be presented to the user as part of a web page. application programs (such as a web browser) running on the computing device 1130; For example, in some embodiments, The computing device 1130 may be a mobile device (e.g., a smartphone). Rather, a GUI is a graphical user interface (GUI) that is implemented by an application program (e.g., an application program) running on a mobile device. The user may be presented with the information via the "Application" ("Application"). [Example]
[0360] Example 1 Establishing RNA transcript normalization and analyzing sequencing technical noise To establish an exemplary process for RNA transcript normalization as described herein, Experiments were performed to analyze the sequencing technical noise.
[0361] Figure 12A shows purified B cells (as an example of a cell type) sequenced in various laboratories. The transcripts per million covering different biological types were calculated in different samples of The percentage of transcripts per million (TPM) is shown. The GEO and ArrayExpress IDs of the samples are shown as labels on the X-axis. are shown in the legend (according to GENCODE annotations, version 23). In addition, the total expression variability attributed to short RNA transcripts is reduced by the variability due to normalization of the length of the short transcripts. This strongly distorts the TPM distribution of the gene of interest by increasing the number of non-coding transcripts. As mentioned above, including with regard to the "Considerations for the analysis of non-coding transcripts" section, This may reduce the variance of the data.
[0362] Figure 12B shows the sequence of the reference human transcriptome (GENCODE, v23), as indicated in the legend. The transcript distribution by transcript biotype and length is shown. The percentages of transcripts of various lengths in the biotypes are shown (all retained in Figure 12C). (with further categories of all removed transcripts). Additionally, a significant amount of noise appears to correspond to V, D, or J regions in the transcriptome. These were derived from short transcripts of the annotated TCR and BCR coding genes. Cells make long transcripts after VDJ recombination, and these short transcripts are never synthesized. Therefore, without specific rearrangement, various TCR and BCR variants (TCR and BCR repertoires) Finally, we used a filter to measure short non-coding RNA sequences. In addition to removing these TCR and BCR protein-encoding transcripts, TPM normalization was performed. The exclusion of non-coding transcripts and TCR- and BCR-transcripts was shown in Figure 12. As shown in B, this may reduce the variance of the data.
[0363] Figure 12C is a schematic diagram of an exemplary process for expression quantification and TPM renormalization. Expression was calculated by Kallisto (Bray et al. 2016). Next, non-coding transcripts, short V, D or is a transcript encoding the TCR / BCR associated with the J segment and its biological properties and quality / Filter other transcripts that follow the evidence information. Finally, aggregate transcripts by gene. and normalized to 1 million TPM.
[0364] [Table 10A]
[0365] [Table 10B]
[0366] Figures 12D-12E show the results before (red) and after (blue) transcript filtration and TPM renormalization. The expression of 3515 housekeeping genes (Eisenberg and Levanon 2013) in various cell types The data are from total RNA-seq (Figure 12D) or polynucleotides. A. Grouped based on type of library preparation using either RNA-seq (Figure 12E). The P values shown are calculated by two-tailed Wilcoxon test. Median and rank biserial correlation coefficients are shown.
[0367] Figure 12F shows the total R before (left) and after (right) the proposed transcript filtration and renormalization. Selected results from experiments using either NA-seq (green) or polyA RNA-seq (red). 1 is a PCA projection of RNA expression in B cells. After the normalization procedure, there is an unwanted batch effect between the expression profiles. The techniques for TPM normalization are described herein, including with respect to the "Transformation and Gene Assembly" section. It will be published.
[0368] Figure 12G shows the dependence of the relative standard deviation of technical replicates on gene expression levels (TPM). RNA-se with total coverage of 0 million (pink), 5 million (yellow), and 10 million (green) read counts q Experiments are presented.
[0369] Figure 12H (left) shows the mean standard deviation of gene expression across all read counts in RNA-seq. The graphs shown show the dependence of noise levels on the amplitude: Technical Poisson noise only (blue), all technical noise (yellow) and technical and biological Figure 12H (right) shows the results calculated in samples with different types of noise. 6 is a violin plot showing the distribution of the same standard deviation of gene expression. As mentioned above, the technical noise component can be characterized by a Poisson distribution. Another component of technical noise may be identified by non-Poisson noise. ,Biological noise may be specified by a normal distribution.
[0370] Figure 12I shows the results of technical replicates of an RNA-seq experiment with different total read count coverage. 1 is a plot showing the measured Poisson noise coefficient. Poisson noise is a measure of the noise in the RNA-seq data. is inversely proportional to the square root of the total read count coverage of the data.
[0371] Figure 12J (left) shows the mean standard deviation of gene expression across all read counts in RNA-seq. The graph shown shows the dependence of gene expression on the level of the imputed Poisson noise (green). Figure 12J (right) shows data for the same sample with (yellow) and all technical noise. The dependence of the mean standard deviation of gene expression on the total coverage of read counts in RNA-seq. The graph shown is the same as that presented in the left graph after subtracting the imputed Poisson noise. The same data as shown is shown, showing the non-Poisson addition to the technical noise. The technical noise does not show any dependence on sequencing coverage.
[0372] Figure 12K (left) shows the mean standard deviation of gene expression across all read counts in RNA-seq. The graphs shown show the dependence of the densities on the number of cells in a single cell across various laboratories and experiments. The gene expression of the strains is shown, accounting for both biological and technical noise. The calculated imputed Poisson technical noise is shown in green. Figure 12K (right) shows the gene expression Figure 1 shows the dependence of the mean standard deviation of the current read counts on the total coverage in RNA-seq. The graph shown is the gene expression pattern shown on the left after subtracting the imputed Poisson noise. This indicates pure biological noise in the sample, which is due to the sequencing coverage. It did not depend on the
[0373] Example 2 Deconvolution of the microenvironment from RNA-seq of multiple normal and cancer tissues The RNA-seq data from multiple normal and cancer tissues was used to Experiments were carried out to perform cellular deconvolution according to the techniques described in this paper. The cellular deconvolution technique developed by the authors is called "Kassandra" Specifically, they are cell type and / or subtype specific and and / or semi-specific genes are selected, artificial mixtures are generated, and multiple nonlinear regression models are trained. We then developed a method to determine multiple cellular constituent ratios of multiple cell types and used a trained nonlinear regression model. The techniques for determining cellular composition, as well as other pretreatments and methods described herein, are used to and post-processing techniques.
[0374] Figure 13A shows a schematic of the validation experiment for deconvolution based on TCGA data. Figure 1. Hematoxylin and eosin (H&E) slides and whole exome sequencing ( Data on cell numbers obtained from other methods (WES) are used.
[0375] Figure 13B shows the B cell, CD4+, CD8+, and macrophage counts in 10,489 tumor biopsies from TCGA. Deconvolution of phage, fibroblasts, and endothelial cells as described herein Cellular composition estimated using a technique (e.g., using a trained nonlinear regression model) 1 is a violin plot showing the distribution of tumor rates. As shown, the tumor tissue are divided by type of cancer.
[0376] Figure 13C shows the TCGA and GTEX assays calculated based on the deconvoluted cell fractions. t-SNE plot showing the correlation coefficients.
[0377] Figure 13D shows the TCGA RNA-seq data predicted by the techniques described herein. , and lymphoid tissue predicted by mechanical analysis of histological TCGA data by (Saltz et al. 2018). 10 is a graph showing the Pearson correlation between sphere proportions.
[0378] FIG. 13E shows predicted genomic DNA sequences of malignant cells from RNA-seq using the techniques described herein. Plot showing correlation between ratio and tumor purity estimated from WES for 11 TCGA cancer types is.
[0379] Figure 13F shows the correlation between tumor purity and the predicted proportion of malignant cells based on RNA-seq data. Graph showing correlation between tumor size and tumor progression. Tumor data was derived from TCGA. The graph is Pearson correlation of predictions by the techniques described in
[14] as well as predictions by various alternative algorithms. The Pearson correlation is shown. The nonlinearity of the proposed algorithm is compared with other algorithms. Shape deconvolution technology more accurately predicts the proportion of malignant cells and outperforms conventional techniques Improvement was demonstrated.
[0380] FIG. 13G shows the LUSC of T cell RNA ratios predicted by the techniques described herein. Pearson correlation between MiXCR and T cell receptor (CDR3 region of TCR) reads in TCGA data 1 is a graph showing
[0381] FIG. 13H shows plasma B cell RNA ratios predicted by the techniques described herein. In the LUSC TCGA data, peer analysis of B cell receptor (IgH CDR3 region) leads by MiXCR was performed. 10 is a graph showing the correlation between the
[0382] Figure 13I shows predicted T cell RNA ratios for various cancer types from TCGA data. 1 is a graph showing Pearson correlation values with cell receptor (CDR3 region of TCR) leads. Predictions are shown according to the techniques described herein and according to various alternative algorithms. Each data point corresponds to a different cancer type (COAD, KIRC, LUAD, LUSC, READ, SKCM, TNBC). do.
[0383] Figure 13J shows predicted plasma B cell RNA ratios in various cancer types from TCGA. 2 is a graph showing Pearson correlation values with B cell receptor (IgH CDR3 region) leads. Predictions are shown for the technique described in
[1999] and for various alternative algorithms. Each data point corresponds to a different cancer type (COAD, KIRC, LUAD, LUSC, READ, SKCM, TNBC). Respond.
[0384] In this experiment, we analyzed the cellular composition of TCGA samples of various tumor types and healthy tissues. B cells, CD4+ T cells, CD8+ T cells, macrophages, fibroblasts, and endothelium were analyzed (Figure 13B). Five major cell populations, including IL-16 cells, were quantified (Fig. 13C). These values are consistent with those reported. For example, the DLBC RNA-seq data showed a strong enrichment of B cells. Tumor purity values predicted by the techniques described herein and other deconvolution techniques are The correlation between the quantification algorithms was compared using an established purity algorithm (Figure 13E (Figure 13F). This analysis was performed using the techniques described herein to refine the analysis from bulk RNAseq data. This supported the ability to accurately predict cell populations.
[0385] In this example, we analyzed the expressed T cell receptor (TCR) and IgH / L (B cell) The proportion of T cells or plasma B cells actively producing immunoglobulins is MIXCR was used to realign the sequences and identify various T and plasma B cell clones. The abundance and diversity of CDR3 transcripts associated with the nucleotide sequence were measured. The techniques described herein are the only ones that allow for the predicted T cell proportions to be accurately predicted within a sample. Provides a strong correlation between the IgH / L transcript fraction and the number of TCRs found and with the plasma B cell ratio (Fig. 13G to Fig. 13J).
[0386] Example 3 Deconvolution of single-cell and bulk RNA-seq data from blood Herein, we use single-cell RNA-seq data and bulk RNA-seq data from blood. Experiments were carried out to perform cellular deconvolution according to the techniques described in
[14] . The cellular deconvolution technique developed by the present inventors is called "Cassandra" Specifically, artificial mixtures are generated and cell types and / or subtypes are identified. Select specific and / or semi-specific genes for the group and train multiple nonlinear regression models. Determine multiple cellular constituent ratios of multiple cell types using a trained nonlinear regression model and other pretreatments and techniques described herein to determine the cellular composition ratio. Post-processing techniques.
[0387] Figure 14A shows the validation for deconvolution using scRNA-seq samples from PBMCs. Schematic of the recognition experiment. The scRNA-seq data were artificially mixed to create a bulk RNA-seq data set. We created a mutant strain.
[0388] Figure 14B shows the results across nine single-cell PBMC datasets provided by 10x Genomics. t-SNE plots of cell phenotyping. The concatenated plots are normalized by SCTransform, The Seurat pipeline (Butler et al. 2018; Stuart et al. 2019) included batch correction and upfront PCA. As shown, various cell types and / or subtypes have been identified using the methods that distinguish them. The cells express important cellular markers (e.g., specific and / or semi-specific genes) that contribute to the development of the cell.
[0389] Figure 14C shows the true cell ratios from scRNA-seq of PBMCs compared with the true ratios for the bulk RNA-seq mixture. 1 is a graph showing correlations between predictions made using the techniques described herein.
[0390] Figure 14D shows the true ratios from scRNA-seq of PBMCs compared with the ratios described herein for eight cell subtypes. using techniques described in the literature (e.g., nonlinear regression to determine cellular composition). 1 is a plot showing the correlation of predictions made using the model.
[0391] Figure 14E shows the validity of deconvolution using bulk RNA-seq of PBMCs or whole blood. FIG. 1 is a schematic diagram of a sex confirmation experiment and FACS measurement of the same sample.
[0392] Figures 14F-1 and 14F-2 show the expression of various cell types (CD4+ T cells, CD8+ T cells, NK cells, B cells, monocytes, and Prediction of neutrophils and leukocytes from bulk RNA-seq using the techniques described herein The correlation between the estimated cell ratio and the actual cell ratio obtained by flow cytometry was The data set used for comparison is GSE107572 (Finotello et al. 201 9), GSE115823 (Altman et al. 2019), GSE60424 (Linsley et al. 2014), SDY67 (Zimmermann et al. 20 16), GSE127813 (Newman et al. 2019), GSE53655 (Shin et al. 2014), and GSE64655 (Hoek et al. 2015). Pearson correlations are shown for all cell types combined.
[0393] In this experiment, we applied the techniques described herein to peripheral blood mononuclear cells (PBMs). C) was applied to artificial bulk RNA-seq constructed from scRNA-seq datasets derived from the seq library (Figure 1 4A to 14B). High correlation was observed when the true scRNA-seq ratio was aligned with the predicted RNA-seq ratio. Correlation values were obtained (Figure 14C). In this example, when the correlations for each cell type were graphed separately, , cell types present in high numbers have the most significant correlation between true and predicted values ( Figure 14D).
[0394] FACS analysis was then performed on the blood samples that were available using the techniques described herein. Bulk RNA-seq of the PBMCs was analyzed (Figure 14E). Eight different PBMC samples were analyzed, and for each sample, F The ACS analysis was compared to the cellular organization predicted by the techniques described herein. As shown, all analyses presented correlation coefficients ranging from 0.900 to 0.984 (Figure 14F-1 and and Figure 14F-2).
[0395] Example 4 Deconvolution of the microenvironment from various cancer tissues We used scRNA-seq data from several tumor tissues, including melanoma, head and neck cancer, and lung cancer. Cellular deconvolution is performed according to the techniques described herein using The figure shows the results of the cellular deconvolution technique developed by the present inventors. The technique is sometimes called "Cassandra." Specifically, it creates an artificial mixture and Select genes specific and / or semi-specific for the type and / or subtype, and A shape regression model is trained to determine multiple cell composition ratios of multiple cell types, and the trained nonlinear The technique for determining cellular composition ratios using a shape regression model, as well as the method described herein Other pre- and post-processing techniques that may be used.
[0396] Figure 15A shows, from left to right, melanoma (GSE72056) (Tirosh et al. 2016), lung cancer (E-MTAB-6149 and EM TAB-6653) (Lambrechts et al. 2018) and head and neck cancer (HNC) (GSE103322) (Puram et al. 2017) single t-SNE plot of cell phenotyping in a cell dataset. t-SNE plot of lung cancer. The Seurat pipeline (Butler et al. 2014) includes SCTransform normalization, batch correction, and upfront PCA. 2018; Stuart et al. 2019). Melanoma and head and neck cancer t-SNE plots were obtained using cell type specificity. The log TPM expression values of differential genes were obtained by t-SNE transformation.
[0397] Figure 15B is a schematic diagram of a validation experiment using scRNA-seq data derived from cancer tissues. The scRNA-seq data were artificially mixed to create a bulk RNA-seq dataset.
[0398] Figures 15C, 15D, 15E, and 15F show the true cell proportion values derived from scRNA-seq data (Figure 15A). and decombination using the techniques described herein from artificial bulk RNA-seq data. Figure 15C shows the correlation between the solution predictions and melanoma (n=19) and lung cancer (n=19). (n=12), HNC (Fig. 15E) (n=22), and B-cell lymphoma (Fig. 15F) (n=12) in various cell subpopulations. A correlation has been shown.
[0399] Figures 15G and 15H show predictions from artificial bulk RNA-seq data for melanoma, lung cancer, and HNC. The average Pearson correlation values between the estimated values and the true values derived from the scRNA-seq data (Figure 15G) and 15A and 15B are heat maps showing the average MAE (mean absolute error) scores (FIG. 15H). The results obtained from the techniques described herein are compared with those obtained from alternative algorithms. In particular, when compared with conventional techniques for deconvolution, The nonlinear regression techniques developed in this study can, on average, more accurately predict the cellular composition of various cell types. The results are shown to have a lower mean absolute error.
[0400] FIG. 15I shows the cell ratios predicted by the techniques described herein and the data set. FAC of lymphocytes, fibroblasts, and lung adenocarcinoma cell lines from the lab GSE121127 (Wang et al. 2018) (top). S and bone marrow CYTOF from dataset GSE120444 (Oetjen et al. 2018) (bottom). The Pearson correlation value (r) indicates the correlation between the actual cell proportions measured and all the combined The correlation values of cell types are shown.
[0401] In this experiment, cells obtained from scRNA-seq were manually annotated (Figure 15A) to identify the markers for each cell type. Specific ratios of the nucleotides were mixed to resemble the bulk RNA-seq sample (e.g., at least as related to Figure 6A). These cell ratios are then compared to the The results were compared with the values predicted by the techniques described in the literature. The ability of the techniques to reconstitute the cellular composition of each cell type was measured (Fig. 15C-F). The median correlation of the reconstruction reached approximately 0.97, the highest among the other methods.
[0402] The techniques described herein demonstrate the ability to detect and quantify chromatin fragments in mixed samples derived from scRNA-seq data. When compared to alternative techniques in its ability to estimate cell number, The described technique yielded the most accurate results with the highest correlation score (Figure 15G) and lowest mean average error (MAE) (Figure 15H). Many cell types were achieved. Only the techniques described here allow for the identification of CD4+ T cells and control It is accurate in reconstituting T cells, providing average Pearson correlation values of up to 0.87 and 0.95. (Fig. 15G). Thus, although these cell types have a high number of overlapping genes, The techniques developed by the inventors have successfully produced more accurate results than alternative algorithms. Drop.
[0403] [Table 11A]
[0404] Table 11B
[0405] Table 11C
[0406]
Table 11D
[0407] Table 11E
[0408] Table 11F
[0409]
Table 11G
[0410] Table 11H
[0411]
Table 11I
[0412]
Table 11J
[0413]
Table 11K
[0414]
Table 11L
[0415] [Table 11M]
[0416] Thus, several aspects and embodiments of the techniques described in this disclosure have been described. It should be understood that various changes, modifications, and improvements will readily occur to those skilled in the art. Such changes, modifications, and improvements are intended to be within the spirit and scope of the technology described herein. For example, it is contemplated that one skilled in the art would be able to perform the functions described herein and and / or various other means to obtain the result and / or one or more benefits. and each of such variations and / or modifications is within the scope of the present invention. Those skilled in the art will appreciate that this is within the scope of the embodiments described herein. It is understood that many equivalents to the specific embodiments described herein may be recognized using no more than experimentation. It will be appreciated that the above embodiments are provided by way of example only. and within the scope of the appended claims and their equivalents, embodiments of the present invention are specifically described. It should be understood that the present invention may be practiced differently from that described. Any two or more of the features, systems, articles, materials, kits and / or methods described Any combination of such features, systems, articles, materials, kits and / or methods may be used in conjunction with one another. Where not inconsistent, it is within the scope of this disclosure.
[0417] The above-described embodiments can be implemented in any of numerous ways, including by performing a process or method. One or more aspects and embodiments of the present disclosure, including a device (e.g., a computer, a processor, a process or device using program instructions executable by the device (such as a processor or other device) In this respect, various inventive concepts can be implemented or controlled. When implemented on one or more computers or other processors, the various implementations described above One or more programs implementing a method for implementing one or more of the aspects A computer-readable storage medium (or a plurality of computer-readable storage media) (e.g., computer memory, one or more floppy disks, compact Disks, optical disks, magnetic tapes, flash memories, field programmable gates Circuitry in a chip array or other semiconductor device, or other tangible computer The computer-readable medium(s) may be embodied as a storage medium. The program(s) to be executed may be executed on one or more different computers or other processes. The software may be transportable so that it can be loaded into a computer and implemented in various ways. In some embodiments, the computer-readable medium may be a non-transitory medium.
[0418] The terms "program" and "software" are used in their general sense in this specification. and may be used to configure a computer or other processor to implement various aspects as described above. Any kind of computer code that can be used to program or It should further be appreciated that, according to one aspect, One or more computer programs that, when executed, perform the methods of the present disclosure may be implemented solely by It need not reside on a single computer or processor, and implements various aspects of the present disclosure. Therefore, it can be distributed modularly among several different computers or processors.
[0419] Computer-executable instructions may be used, for example, to execute one or more computers or other devices. The software may take many forms, such as a program module executed by a Program modules include programs that perform particular tasks or implement particular abstract data types. This includes routines, programs, objects, components, data structures, etc. Typically, the functionality of the program modules may be combined or combined as desired in various embodiments. It can be dispersed.
[0420] Also, data structures may be stored in any suitable form on computer-readable media. For simplicity of illustration, the data structure is organized as follows: Such a relationship may be indicated by the field's storage being By assigning a location in a computer-readable medium that conveys the relationship between However, pointers, tags, or other mechanisms that establish relationships between data elements may be used. The information in the fields of the data structure may be stored using any suitable mechanism, including by using Establish relationships between information sources.
[0421] When implemented in software, the software code may be provided to a single computer. Any suitable processor, whether centralized or distributed among multiple computers. Or it can be run on a collection of processors.
[0422] Furthermore, it should be understood that computers may be used in any of several forms, e.g. Non-limiting examples include rack-mounted computers, desktop computers, The present invention can be embodied in a laptop computer or a tablet computer. on devices that are not generally considered computers but have suitable processing power. This can be implemented in a personal digital assistant (PDA), smartphone, tablet, or other Suitable portable or stationary electronic devices are included.
[0423] A computer may also have one or more input and output devices. The device can be used, among other things, to present a user interface. Visual presentation of output is an example of an output device that can be used to provide an interface. a printer or display screen for the or other sound generating devices. Examples of input devices that can be used include a keyboard and a pointing device, such as a mouse. Examples include touchpads and discretized tablets. Input information can be received in a cognitive or other audible form.
[0424] Such computers may be connected to a local area network or a wireless network, such as a business network. Area Networks and Intelligent Networks (IN) or Internet The networks may be interconnected by one or more networks of any suitable form, including Such networks may be based on any suitable technology and may use any suitable protocol. It can operate according to the standard and can be used over a wireless network, a wired network, or a fiber optic network. It may contain a work.
[0425] Also, as described, some aspects can be embodied as one or more methods. The acts carried out as part of the Act may be ordered in any suitable way. Therefore, embodiments can be constructed in which the actions are performed in a different order than illustrated, which may include: Some acts may be performed simultaneously even though they are shown as sequential acts in the exemplary embodiment. and
[0426] All definitions as defined and used herein are dictionary definitions, by reference. Definitions in incorporated documents and / or ordinary meanings of defined terms are governed by the It should be understood.
[0427] As used herein and in the claims, indefinite coronary The words "a" and "an" are understood to mean "at least one" unless expressly indicated to the contrary. It should be understood.
[0428] As used herein and in the claims, "and" means The term "or" means "either or both" of the elements so conjoined, i.e., It is understood to mean elements which are conjunctive in some cases and disjunctive in other cases. Multiple elements listed with "and / or" are treated in the same way, i.e. "and / or" should be interpreted as "one or more" of the elements so conjoined. Any elements other than those specifically identified by the clause shall be construed as including those specifically identified elements. It may be present in any form, whether related or unrelated. For example, a reference to "A and / or B" may be used in conjunction with open-ended language such as "including." When used, in one embodiment, only A (optionally including elements other than B) is used, and in another embodiment, In some embodiments, only B (optionally including elements other than A) is used, and in other embodiments, both A and B are used. It may refer to a method (optionally including other elements), etc.
[0429] As used herein and in the claims, one or more With respect to a list of multiple elements, the phrase "at least one" refers to any one of the elements in the list. should be understood to mean at least one element selected from one or more elements of but not necessarily at least one of every element specifically listed in the list of elements It is not intended to be inclusive or exclusive of any combination of elements in the list of elements. This definition also includes all elements, whether related or not to those elements specifically identified. and the elements specifically identified in the list of elements referred to by the phrase "at least one." It is therefore possible that other elements may be optionally present. and "at least one of A and B" (or equivalently, "at least one of A or B" or equivalently In one embodiment, "at least one of A and / or B" refers to at least one, optionally , including more than one A and no B (optionally including elements other than B), as another embodiment In this case, there is at least one, and optionally more than one, B, and no A (and optionally and optionally including elements other than A), and in yet another embodiment at least one, optionally more than one contains more than one A and at least one, optionally more than one B (and optionally , including other elements).
[0430] In the claims and in the above specification, "comprising" and "including," "carrying," "having," "containing" ), "involving", "holding", "composed of" All transitional phrases, such as ", ... "consisting of" and "essentially consisting of" should be understood to mean Only the transitional phrase "consisting essentially of" is used in closed or semi-closed sentences, respectively. This is a transitional phrase.
[0431] The terms "approximately," "substantially," and "about" refer in some embodiments to within ± of a target value. within 20%, in some embodiments within ±10% of the target value, and in some embodiments within ±5% of the target value "Approximately" may be used to mean within ±2% of a target value in some embodiments. The terms "substantially" and "about" may be inclusive of the target value. [Explanation of symbols]
[0432] 100 systems 102 Biological samples 104 Sequencing Platform 106 Sequence Information 108 Computer Devices 110 Cell composition ratio 122 Cell type A 124 Sequence Information 126 Model A 128 Cell composition ratio 132 Cell type B 134 Sequence Information 136 Model B 138 Cell composition ratio 140 t-SNE plots 142 Subtype A 144 Sequence Information 146 Model C 148 Cell composition ratio 150 t-SNE plots 152 tumor cells 156 Model D 158 Cell composition ratio 160 cell types 162 Subtype B 164 Sequence Information 170 Gene Expression 180 cell populations 182 tumor cell lines 190 genes 192 genes 200 Methods and Processes 202 Operation 204 Operation 206 Operation 212 operation 214 operation 216 operation 216a operation 216b operation 218 operation 220 Operation, Implementation Examples, Methods 232 operation 234 operation 236 operation 302 primary tumor samples 304 Nonlinear Regression Models 306 RNA ratio 308 Cell type A 310 Cell type B 312 Cell type C 314 Expression Data 316 Expression Data 318 Expression Data 320 Nonlinear Regression Models 322 Nonlinear Regression Models 324 Nonlinear Regression Models 326 First Submodel 328 First Submodel 330 First Submodel 332 First Value 334 First Value 336 First Value 338 Second Submodel 340 Second Submodel 342 Second submodel 344 Second Value 346 Second Value 348 Second Value 350 formula 360 RNA ratio 370 Cell composition ratio 380 method 382 First Step 384 Second Process 386 Third Step 400 ways 402 operation 404 operation 406 operation 408 operation 500 ways 502 operation 510 operation 520 operation 530 Diagram 540 diagram 550 parameters 600 ways 602 operation 604 Rebalancing 608 operation 610 Branch 612 operation 614 operation 620 Branch 630 Branch 640 Branch, Server 650 Branch 702 RNA expression data 704 RNA expression data 800 ways 802 operation 804 operation 806 operation 1000 Computer Systems 1010 processor 1020 memory 1030 Non-volatile storage media 1040 Network Input / Output (I / O) Interface 1050 User I / O Interface 1100 Environment 1110 Network 1120 Database 1130 Computer Devices 1140 Server 1150 Doctors 1160 databases 1170 Laboratory 1180 targets
Claims
1. Using at least one computer hardware processor, obtaining RNA expression data for a biological sample, the biological sample having been previously obtained from a subject having, suspected of having, or at risk of having cancer; the RNA expression data includes first RNA expression data associated with a first set of genes associated with a first cell type; the first cell type is selected from the group consisting of B cells, CD4+ T cells, CD8+ T cells, endothelial cells, fibroblasts, lymphocytes, macrophages, monocytes, NK cells, neutrophils, and T cells; The process and determining a first cellular constituent ratio for the first cell type using the first RNA expression data, wherein the first cellular constituent ratio indicates an estimated proportion of cells of the first cell type in the biological sample, and determining the first cellular constituent ratio for the first cell type comprises: providing the first RNA expression data as input to a first nonlinear regression model to obtain a corresponding output representing an estimated proportion of RNA from the first cell type; and determining the first cellular constituent ratio for the first cell type based on the estimated ratio of RNA from the first cell type, wherein a first nonlinear regression model is trained to estimate cellular constituent ratios for a single cell type; and and The first RNA expression data is related to the following genes: ADAP2, ADGRE3, ADGRG3, ADORA3,AIF1, AOAH, APOBEC3D, ARHGAP15, ARHGAP30, ARHGAP9, ARHGDIB, BANK1, BLK, C1QA, C1QC, C3AR1, C5AR1, CAMK4, CBLB, CCDC69, CCL5, CCL7, CCR1, CCR2, CCR3, CD14, CD160, CD163, CD19, CD1D, CD2, CD22, CD226, CD244, CD247, CD27, CD300A, CD300C, CD300E, CD300LB, CD302, CD33, CD37, CD3D, CD3E, CD3G, CD4, CD48, CD5, CD53, CD6, CD68, CD69, CD7, CD79A, CD79B, CD86, CEACAM8, CECR1, CELF2, CLDND2, CLEC17A, CLEC2D, CLEC5A, CLEC7A, CMKLR1, CORO1A, CPNE5, CR2, CSF1R, CSF2RA, CSF3R, CTSS, CTSW, CXCR1, CXCR2, CXCR5, CYBB, CYFIP2, CYTH4, CYTIP, DENND1C, DERL3, DOCK2, EAF2, ELF1, ELMO1, EVI2B, FAM129C, FAM78A, FCER1G, FCGR1A, FCGR1B, FCGR2A, FCGR3B, FCMR, FCN1, FCRL1, FCRL2, FCRL3, FCRL5, FCRLA, FERMT3, FFAR2, FGR, FKBP11, FLT3LG, FMNL1, FNBP1, FPR1, FPR2, FPR3, GLCCI1, GLT1D1, GPR174, GZMM, HCK, HCLS1, HLA-DOB, HMHA1, ICAM3, IFI30, IFITM2, IGFLR1, IGHG1, IGHG3, IGHM, IGKC, IGLL5, IKZF1, IKZF3, IL10, IL16, IL2RB, IL2RG, IL4I1, INPP5D, IRF5, ITGAL, ITGAX,ITGB2, ITGB7, ITK, KCNA3, KCNAB2, KCNJ15, KIR2DL1, KIR2DL2, KIR2DL3, KIR2DL4, KIR2DS2, KIR3DL1, KIR3DL2, KLRB1, KLRC2, KLRC3, KLRD1, KLRF1, KLRK1, LAG3, LAIR1, LAPTM5, LAT, LAX1, LCK, LCP1, LIM2, LRRC25, LSP1, LTA, LY9, MAP4K1, MEFV, MMP25, MNDA, MRC1, MS4A1, MS4A4A, MS4A6A, MSR1, MYO1F, MYO1G, MZB1, NCAM1, NCF2, NCKAP1L, NCR1, NCR3, NFATC2, NKG7, NLRC3, NMUR1, P2RY10, P2RY13, P2RY8, PADI2, PADI4, PARVG, PAX5, PGLYRP1, PHOSPHO1, PIK3AP1, PILRA, PLA2G7, PLCB2, POU2AF1, PPP1R16B, PRF1, PRKCB, PTGDR, PTPN22, PTPN6, PTPRC, PTPRCAP, PVRIG, PYHIN1, RAB7B, RAC2, RASGRP1, RASGRP2, RASGRP4, RASSF5, RCSD1, RHOH, RLTPR, S1PR5, SAMD3, SAMSN1, SASH3, SEC11C, SH2D1B, SIGLEC1, SIGLEC5, SIGLEC7, SIGLEC9, SIRPB2, SIRPG, SIT1, SLA2, SLAMF6, SNX20, SP140, SPI1, SPIB, SPN, SSR4, STAP1, STAT5A, STK4, TAGAP, TBC1D10C, TBX21, TCF7, TESPA1, TLR2, TMC8, TMIGD2, TNFAIP8, TNFAIP8L2, TNFRSF10C, TNFRSF13B, TNFRSF13C, TNFRSF17, TRAC, TRAF3IP3, TRAT1, TRBC2, TRDC, TREM2, TRGC1, TRGC2, TXNDC11, TXNDC5, TYROBP,A group of immune cell-related genes, including UBASH3A, VAV1, VNN2, VNN3, VPREB3, VSIG4, WAS, XCL2, and ZBED2; and the following genes: BANK1, BLK, CD19, CD22, CD37, CD79A, CD79B, CLEC17A, CPNE5, CR2, CXCR5, DERL3, EAF2, FAM129C, FCRL1, FCRL2, FCRL3, FCRL5, FCRLA, FKBP11, GLCCI1, HLA-DOB, IGHG1, IGHG3, IGHM, IGKC, IGLL5, MS4A1, MZB1, PAX5, POU2AF1, SEC11C, SPIB, SSR4, STAP1, TNFRSF13B, TNFRSF13C, and TNFRSF17. A group of B cell-related genes, including TXNDC11, TXNDC5, and VPREB3; A group of plasma B cell-related genes, including the following genes: BANK1, BLK, CD19, CD22, CD37, CD79A, CD79B, CLEC17A, CPNE5, CR2, DERL3, EAF2, FAM129C, FCRL1, FCRL2, FCRL3, FCRL5, FCRLA, FKBP11, GLCCI1, HLA-DOB, IGHG1, IGHG3, IGHM, IGKC, IGLL5, MZB1, POU2AF1, SEC11C, SPIB, SSR4, STAP1, TNFRSF13B, TNFRSF13C, TNFRSF17, TXNDC11, and TXNDC5; A group of genes associated with non-plasma B cells, including the following genes: ADAM28, BANK1, BCL11A, BLK, CD19, CD22, CD37, CD72, CD79A, CD79B, CLEC17A, CPNE5, CR2, CXCR5, FAM129C, FCER2, FCRL1, FCRL2, FCRL3, FCRL5, FCRLA, HLA-DOB, MS4A1, PAX5, POU2AF1, RALGPS2, SPIB, STAP1, TNFRSF13B, TNFRSF13C, and VPREB3; A group of T cell-related genes, including the following genes: CAMK4, CBLB, CD2, CD226, CD3D, CD3E, CD3G, CD48, CD5, CD6, CD7, FLT3LG, ITK, KCNA3, KLRB1, LAG3, LAT, LCK, LTA, SIRPG, SIT1, SLA2, TBX21, TCF7, TESPA1, TRAC, TRAF3IP3, TRAT1, TRBC2, TRDC, TRGC1, TRGC2, UBASH3A, and ZBED2; A group of CD4 T cell-related genes, including the following genes: ANKRD55, CCR4, CD2, CD27, CD28, CD3D, CD3E, CD3G, CD4, CD40LG, CD5, CD6, FHIT, FLT3LG, ICOS, IKZF1, IL2RA, IL9, IRF4, ITK, LCK, LEF1, LTA, TESPA1, TNFRSF4, TRAC, TRAT1, TRBC2, and UBASH3A; A group of genes associated with regulatory T cells, including the following genes: CCR4, CCR8, CD2, CD27, CD4, CTLA4, ENTPD1, FOXP3, HAVCR2, IKZF2, IKZF4, IL21R, IL2RA, IL2RB, IL2RG, ITGAE, ITK, LAG3, LTB, SIRPG, TIGIT, TNFRSF18, TNFRSF4, TNFRSF8, TNFRSF9, and TRAC; A group of genes related to helper T cells, including the following genes: ANKRD55, CD2, CD28, CD40LG, CD5, CD6, FHIT, FLT3LG, IL7R, ITK, ITM2A, KLRB1, LCK, LEF1, LRRN3, NELL2, P2RY8, TCF7, TESPA1, THEMIS, TRAF3IP3, and TRAT1; A group of CD8 T cell-related genes, including the following genes: CCL5, CD2, CD3D, CD3E, CD3G, CD6, CD7, CD8A, CD8B, CD96, CRTAM, CXCR3, EOMES, FCRL6, FLT3LG, GZMA, GZMB, GZMH, GZMK, ITK, KLRC2, KLRC4, KLRK1, PRF1, PRKCQ, PTGDR, PVRIG, SH2D1A, TBX21, TCF7, THEMIS, TIGIT, TRAC, TRAT1, TRBC2, UBASH3A, XCL2, ZAP70, and ZBED2; A group of genes associated with CD8 PD1 low T cells, including the following genes: CCR7, CD160, CD28, CD5, CD8A, CD8B, CRTAM, EOMES, FCRL6, FGFBP2, GZMK, GZMM, IL7R, KCNA3, KLRF1, KLRG1, KLRK1, PRKCQ, PTGDR, PVRIG, S1PR5, SH2D1A, TCF7, and ZAP70; A group of genes associated with CD8 PD1 high T cells, including the following genes: CBLB, CD2, CD226, CD244, CD27, CD38, CD8A, CD8B, CRTAM, CTLA4, ENTPD1, FASLG, HAVCR2, ICOS, IL2RA, IL2RB, IRF4, ITGAE, KLRC1, KLRK1, LAG3, LTA, PDCD1, PRDM1, PRKCQ, PVRIG, SH2D1A, SIRPG, TIGIT, TMIGD2, and TNFRSF9; A group of NK cell-related genes, including the following genes: CCL5, CD160, CD244, CD247, CD7, CLDND2, CTSW, GZMM, IL2RB, KIR2DL1, KIR2DL2, KIR2DL3, KIR2DL4, KIR2DS2, KIR3DL1, KIR3DL2, KLRB1, KLRC2, KLRC3, KLRD1, KLRF1, KLRK1, LIM2, NCAM1, NCR1, NCR3, NKG7, NMUR1, PRF1, PTGDR, PYHIN1, S1PR5, SAMD3, SH2D1B, TMIGD2, and XCL2; A group of monocyte-related genes, including the following genes: AOAH, CCR1, CCR2, CD1D, CD300C, CD300E, CD300LB, CD302, CD33, CECR1, CSF1R, CTSS, CYBB, FCN1, IRF5, MEFV, MS4A6A, PADI4; A group of macrophage-related genes, including the following genes: ADAP2, ADORA3, C1QA, C1QC, C3AR1, C5AR1, CCL7, CCR1, CD14, CD163, CD33, CD4, CD68, CLEC5A, CMKLR1, CSF1R, CYBB, FPR3, IL10, IL4I1, MRC1, MS4A4A, MS4A7, MSR1, PLA2G7, RAB7B, SIGLEC1, TREM2, and VSIG4; A group of genes associated with M1 macrophages, including the following genes: C15orf48, C1QC, C3AR1, CCL3, CCL3L3, CCL4L2, CCL7, CD14, CD68, CLEC5A, CSF1R, CXCL3, CYBB, GADD45G, GRAMD1A, IL10, IL12B, IL15RA, IL1RN, IL27, IL4I1, LILRB4, MMP19, PFKFB3, PLA2G7, SIGLEC1, SLAMF7, SOCS3, SOD2, SPHK1, TNF, TNFAIP6, TNIP3, and VSIG4; A group of genes associated with M2 macrophages, including the following genes: ADAP2, C1QC, CCR1, CD14, CD163, CD209, CD4, CD68, CLEC5A, CMKLR1, CSF1R, CYBB, FKBP15, FPR3, GPNMB, LACC1, LIPA, MRC1, MS4A4A, MSR1, NPL, PLA2G7, RAB42, SIGLEC1, SLC38A6, STAB1, TREM2, and VSIG4; A group of neutrophil-related genes, including the following genes: ADGRE3, ADGRG3, C5AR1, CCR3, CEACAM8, CLEC7A, CSF3R, CXCR1, CXCR2, EVI2B, FCGR2A, FCGR3B, FFAR2, FPR1, FPR2, GLT1D1, IFITM2, KCNJ15, LILRB3, MEFV, MMP25, MNDA, P2RY13, PADI2, PADI4, PGLYRP1, PHOSPHO1, RASGRP4, SIGLEC5, TNFRSF10C, VNN2, VNN3, and WAS; A group of fibroblast-related genes, including the following genes: ACTA2, ADAMTS2, CD248, COL16A1, COL1A1, COL1A2, COL3A1, COL4A1, COL5A1, COL6A1, COL6A2, COL6A3, FAP, FBLN2, FBN1, FGF2, LOXL1, MFAP5, PCOLCE, PDGFRA, PDGFRB, TAGLN, THBS2, THY1, and VEGFC; A group of endothelial cell-related genes, including: ANGPT2, APLN, CDH5, CLEC14A, ECSCR, EMCN, ENG, ESAM, ESM1, FLT1, HHIP, KDR, MMRN1, MMRN2, NOS3, PECAM1, PTPRB, RASIP1, ROBO4, SELE, TEK, TIE1, and VWF; The following genes: ACRBP, ADAP2, ADGRE2, ADGRE3, ADGRG3, ADORA3, AIF1, AOAH, C1QA, C1QC, C3AR1, C5AR1, CCL7, CCR1, CCR2, CCR3, CD14, CD163, CD1D, CD300A, CD300C, CD300E, CD300LB, CD302, CD33, CD4, CD68, CD86, CEACAM8, CECR1, CLEC5A, CLEC7A, CMKLR1, CSF1R, CSF2RA, CSF3R, CTSS, CXCR1, CXCR2, CYBB, EMILIN2, EVI2B, FCER1G, FCGR1A, FCGR1B, FCGR2A, FCGR3B, FCN1, FFAR2, FGL2, FPR1, FPR2, FPR3, GLT1D1, HCK, HK3, IFI30, IFITM2, IGSF6, IL10, IL4I1, IRF5, ITGAM, ITGAX, KCNJ15, LILRA3, LILRA5, LILRA6, LILRB2, LRRC25, LYN, LYZ, MAFB, MEFV, MMP25, MNDA, MPP1, MRC1, MS4A4A, MS4A6A, MSR1, NCF2, NINJ1, OSCAR, P2RX1, P2RY13, PADI2, PADI4, PGLYRP1, PHOSPHO1, PILRA, PLA2G7, PLEK, PRKCD, PSAP, RAB7B, RASGRP4, RNASE6, RP2, SIGLEC1, SIGLEC14, SIGLEC5, a group of myeloid cell-related genes, including SIGLEC9, SIRPB2, SPI1, STX11, TLR2, TNFRSF10C, TNFSF13, TREM2, TYROBP, VNN2, VNN3, VSIG4, and WAS; The following genes were identified: ACAP1, ANXA2R, APOBEC3D, APOBEC3G, BANK1, BLK, CAMK4, CARD11, CBLB, CCL5, CD160, CD19, CD2, CD22, CD226, CD244, CD247, CD27, CD37, CD3D, CD3E, CD3G, CD48, CD5, CD6, CD69, CD7, CD79A, CD79B, CLDND2, CLEC17A, CLEC2D, CPNE5, CR2, CTSW, CXCR5, CYFIP2, DEF6, DERL3, EAF2, ETS1, EVL, FAM129C, FCMR, FCRL1, FCRL2, FCRL3, FCRL5, FCRLA, FKBP11, FLT3LG, GLCCI1, GPR174, GPR18, GRAP2, GZMM, HLA-DOB, IGHG1, IGHG3, IGHM, IGKC, IGLL5, IKZF1, IKZF3, IL16, IL2RB, IL2RG, ITGB7, ITK, KCNA3, KIR2DL1, KIR2DL2, KIR2DL3, KIR2DL4, KIR2DS2, KIR3DL1, KIR3DL2, KLRB1, KLRC2, KLRC3, KLRD1, KLRF1, KLRK1, LAG3, LAT, LAX1, LCK, LIM2, LTA, LY9, MAP4K1, MS4A1, MZB1, NCAM1, NCR1, NCR3, NFATC2, NKG7, NLRC3, NMUR1, P2RY10, P2RY8, PARP15, PAX5, PIK3IP1, POU2AF1, PPP1R16B, PPP3CC, PRF1, PTGDR, PTPRCAP, PVRIG, PYHIN1, RASAL3, RASGRP1, RASGRP2, RHOH, RLTPR, S1PR5, SAMD3, SEC11C, SH2D1B, SIRPG, SIT1, SKAP1, SLA2, SLAMF6, SP140, SPIB, SSR4, STAP1, TBC1D10C, TBX21, TCF7, TESPA1, TMC6, TMC8, TMIGD2, TNFRSF13B, TNFRSF13C,A group of lymphocyte-related genes, including TNFRSF17, TRAC, TRAF3IP3, TRAT1, TRBC2, TRDC, TRGC1, TRGC2, TXNDC11, TXNDC5, UBASH3A, VPREB3, XCL2, ZBED2, and ZNF101. comprising expression data for at least 10 genes selected from method.
2. the RNA expression data includes second RNA expression data associated with the first set of genes associated with the first cell type; The first nonlinear regression model is a first sub-model configured to use the first RNA expression data as input to generate a first value for the estimated proportion of RNA from the first cell type; and a second sub-model configured to use the second RNA expression data and the first value for the estimated proportion of RNA from the first cell type as inputs to generate a second value for the estimated proportion of RNA from the first cell type.
3. The method of claim 1 or 2, wherein the first nonlinear regression model comprises a random forest regression model, a neural network regression model, and / or a support vector machine regression model.
4. The first nonlinear regression model is obtaining simulated expression data; training a first nonlinear regression model using the simulated expression data; trained at least in part by obtaining the simulated data includes obtaining a plurality of artificial mixtures by combining RNA expression data from samples of a plurality of cell types in predetermined proportions; 4. The method of claim 1, wherein the simulated expression data optionally comprises simulated malignant cell RNA expression data, and wherein obtaining the simulated data comprises adding random overexpression noise to the obtained expression data.
5. The first nonlinear regression model is trained, at least in part, by generating training data comprising simulated RNA expression data, the generating training data comprising: obtaining a set of RNA expression data from one or more biological samples, wherein the set of RNA expression data includes microenvironment cell RNA expression data and malignant cell RNA expression data; generating simulated microenvironment cellular RNA expression data using said microenvironment cellular RNA expression data; generating simulated malignant cell RNA expression data using said malignant cell RNA expression data; and combining the simulated microenvironment cellular RNA expression data with the simulated malignant cellular RNA expression data to generate at least a portion of the simulated RNA expression data.
6. generating training data further comprises adding noise to the simulated RNA expression data before training the first nonlinear regression model; Optionally, the noise comprises at least one of Poisson noise or Gaussian noise, and / or generating the simulated malignant cell RNA expression data comprises adding noise to the simulated malignant cell RNA expression data, and / or the malignant cell RNA expression data comprises RNA expression data from a plurality of malignant cell samples, and generating the simulated malignant cell RNA expression data comprises combining the RNA expression data from the plurality of malignant cell samples.
7. The step of generating the simulated microenvironment cellular RNA expression data includes:
7. The method of claim 5 or 6, comprising generating a first RNA expression profile for a first microenvironment cell type using a first portion of the microenvironment cellular RNA expression data, optionally wherein the first portion of the microenvironment cellular RNA expression data comprises RNA expression data from multiple subtypes of the first microenvironment cell type.
8. 8. The method of claim 7, wherein generating the first RNA expression profile comprises resampling the first portion of the microenvironment cellular RNA expression data using the plurality of subtypes of the first microenvironment cell type, and wherein the first portion of the microenvironment cellular RNA expression data comprises RNA expression data from a plurality of samples.
9. 9. The method of claim 8, wherein generating the first RNA expression profile comprises resampling a first portion of the microenvironment cellular RNA expression data using samples from the plurality of samples as input.
10. The step of generating the simulated microenvironment cellular RNA expression data includes: generating a second RNA expression profile for a second microenvironment cell type using a second portion of the microenvironment cell RNA expression data; combining the first RNA expression profile and the second RNA expression profile to generate at least some of the simulated microenvironment cellular RNA expression data; Further comprising: Optionally, combining the first RNA expression profile and the second RNA expression profile to generate the at least some of the simulated microenvironment cellular RNA expression data comprises: determining a weighted sum of said first RNA expression profile and said second RNA expression profile; 8. The method of claim 7, wherein optionally, the coefficients of the weighted sum are determined using the output of a pre-trained non-linear regression model.
11. determining a malignant tumor expression profile using the RNA expression profile for the first cell type and the first cellular constituent ratio for the first cell type; The method according to any one of claims 1 to 10, further comprising Law.
12. 12. The method of claim 1, wherein the first RNA expression data comprises expression data for at least 25 genes, at least 50 genes, or at least 100 genes selected from the group of genes in claim 1.
13. The first nonlinear regression model is obtaining training data comprising simulated RNA expression data, the simulated RNA expression data comprising second RNA expression data for the first set of genes associated with the first cell type; training the first nonlinear regression model to estimate the proportion of RNA from the first cell type, said training comprising: generating an estimated proportion of RNA from the first cell type using the first nonlinear regression model and the second RNA expression data; and updating parameters of the first nonlinear regression model using the estimated proportion of RNA from the first cell type. and The method according to any one of claims 1 to 12, wherein the method is trained by
14. at least one hardware processor; at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by said at least one hardware processor, cause said at least one hardware processor to perform the method of any one of claims 1 to 13; A system including:
15. At least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one hardware processor, cause the at least one hardware processor to perform the method of any one of claims 1 to 13.
Citation Information
Patent Citations
JP1101358366A
Methods and systems for determining ratios of different cell subsets
JP2018512071A
Systems and methods for analyzing mixed cell populations
WO2019018684A1
Learning method, mixing ratio prediction method and learning device
WO2020004575A1