Machine learning techniques for gene expression analysis
By employing gene rankings in statistical models, the analysis of gene expression data can be standardized across various sequencing platforms, addressing the challenges of data variability and platform compatibility.
Patent Information
- Application Number
- JP2022533583
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-08-03
- Filing Date
- 2020-12-05
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-12-05
AI Technical Summary
Existing techniques struggle to analyze gene expression data across different sequencing platforms, leading to variations in expression level values and requiring separate statistical models for each platform, which is impractical and inefficient.
The use of gene rankings instead of specific expression level values allows for the development of statistical models that can operate independently of the sequencing platform, enabling analysis of gene expression data from various sources using a common data processing pipeline.
This approach enables consistent analysis of gene expression data across different sequencing platforms, improving the practicality and efficiency of bioinformatics analysis by allowing a larger sample size and rationalizing the handling of different data formats.
Smart Images

Figure 0007684303000035 
Figure 0007684303000036 
Figure 0007684303000037
Abstract
Description
Technical Field
[0001] Related Applications This application claims the benefit of, and is a continuation of, U.S. Provisional Patent Application No. 62 / 943,976, filed Dec. 5, 2019, entitled “MACHINE LEARNING TECHNIQUES FOR GENE EXPRESSION ANALYSIS,” and U.S. Provisional Patent Application No. 63 / 060,512, filed Aug. 3, 2020, entitled “MACHINE LEARNING TECHNIQUES FOR DETERMINING PERIPHERAL T-CELL LYMPHOMA (PTCL) SUBTYPE USING GENE EXPRESSION DATA,” the entire contents of each of which are incorporated herein by reference.
[0002] Aspects of the technology described herein relate to determining characteristics of a biological sample obtained from a subject known or suspected to have or be at risk of having cancer by sequencing a biological sample using one or more sequencing platforms and analyzing the resulting gene expression data using machine learning techniques. In particular, the technology described herein involves using gene expression data from one or more sequencing platforms to determine characteristics of a biological sample such as the tissue of origin and cancer grade.
Background Art
[0003] The characteristics of living cells can be related to the expression levels of several genes. For example, cancer cells can have some genes that are upregulated and other genes that are downregulated relative to normal and healthy cells. This relationship between cell characteristics and gene expression levels can be utilized when analyzing gene expression data for living cells, such as by using gene expression microarrays or by analyzing data obtained by performing next-generation sequencing, to determine the characteristics of living cells.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Non-Patent Documents
[0005]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Non - Patent Document 8
Non - Patent Document 9
Non - Patent Document 10
Summary of the Invention
Means for Solving the Problems
[0006] Some embodiments are computer-implemented methods that include obtaining expression data, at least in part, by sequencing a biological sample of a subject having, suspected of having, or at risk of having cancer, using at least one computer hardware processor, the expression data comprising expression levels for a plurality of genes, the plurality of genes comprising a set of genes; ranking at least some of the genes in the set of genes based on their expression levels in the expression data to obtain a gene ranking; and determining at least one characteristic of the biological sample using the gene ranking and a statistical model trained using training data representing a plurality of gene rankings for at least some of the genes in the obtained set of genes, wherein each of the plurality of gene rankings is obtained based on respective expression levels for at least some of the genes in the set of genes.
[0007] At least one characteristic can be selected from cancer grade (e.g., breast cancer grade, kidney clear cell cancer grade, lung adenocarcinoma grade) for cells in a biological sample, origin tissue (e.g., lung, pancreas, stomach, colon, liver, bladder, kidney, thyroid, lymph node, adrenal gland, skin, breast, ovary, prostate, or cell of origin in tissues such as germinal center B cell (GCB) or activated B cell (ABC)) for cells in a biological sample, histological information (e.g., tissue types such as adenocarcinoma, squamous cell carcinoma, carcinoma, cystadenocarcinoma, sarcoma, and glioma) for cells in a biological sample, and cancer subtype (e.g., PTCL subtypes such as anaplastic large cell lymphoma (ALCL), angioimmunoblastic T cell lymphoma (AITL), natural killer / T cell lymphoma (NKTCL), and adult T cell leukemia / lymphoma (ATLL)) for cells in a biological sample, and viral status (e.g., HPV status such as HPV positive or HPV negative for head and neck squamous cell carcinoma).
[0008] In some embodiments, at least one characteristic of the biological sample is a physiological characteristic of the cells in the biological sample, or the tissue from which the cells originate. In some embodiments, at least one characteristic is selected from cancer grade for cells in the biological sample, origin tissue for cells in the biological sample, tissue type for cells in the biological sample, and cancer subtype for cells in the biological sample.
[0009] In some embodiments, the method further includes performing sequencing of the biological sample using a gene expression microarray prior to the step of obtaining expression data. In some embodiments, the method further includes performing next-generation sequencing of the biological sample prior to the step of obtaining expression data.
[0010] In some embodiments, at least one characteristic includes the cancer grade for the cells in the biological sample. In some embodiments, at least one characteristic includes the tissue of origin for the cells in the biological sample.
[0011] In some embodiments, the subject has, is suspected of having, or is at risk of having breast cancer. In some embodiments, the set of genes is selected from the group of genes described in Table 1. In some embodiments, the set of genes comprises at least 3 genes selected from the group of genes described in Table 1. In some embodiments, the set of genes comprises at least 5 genes selected from the group of genes described in Table 1. In some embodiments, the set of genes comprises at least 10 genes selected from the group of genes described in Table 1. In some embodiments, the set of genes comprises at least 20 genes selected from the group of genes described in Table 1.
[0012] In some embodiments, the subject has, is suspected of having, or is at risk of having kidney cancer. In some embodiments, the subject has, is suspected of having, or is at risk of having clear cell kidney cancer. In some embodiments, the set of genes is selected from the group of genes described in Table 2. In some embodiments, the set of genes comprises at least 3 genes selected from the group of genes described in Table 2. In some embodiments, the set of genes comprises at least 5 genes selected from the group of genes described in Table 2. In some embodiments, the set of genes comprises at least 10 genes selected from the group of genes described in Table 2. In some embodiments, the set of genes comprises at least 20 genes selected from the group of genes described in Table 2.
[0013] In some embodiments, the subject has, is suspected of having, or is at risk of having lymphoma. In some embodiments, the set of genes is selected from the group of genes set forth in Table 3. In some embodiments, the set of genes comprises at least 3 genes selected from the group of genes set forth in Table 3. In some embodiments, the set of genes comprises at least 5 genes selected from the group of genes set forth in Table 3. In some embodiments, the set of genes comprises at least 10 genes selected from the group of genes set forth in Table 3. In some embodiments, the set of genes comprises at least 20 genes selected from the group of genes set forth in Table 3.
[0014] In some embodiments, the subject has, is suspected of having, or is at risk of having squamous cell carcinoma of the head and neck. In some embodiments, the set of genes is selected from the group of genes set forth in Table 8. In some embodiments, the set of genes comprises at least 10 genes selected from the group of genes set forth in Table 8.
[0015] In some embodiments, at least one characteristic includes the human papillomavirus status for cells in a biological sample. In some embodiments, the set of genes is selected from the group of genes set forth in Table 8. In some embodiments, the set of genes comprises at least 10 genes selected from the group of genes set forth in Table 8.
[0016] In some embodiments, the method further includes ranking at least some genes in a second set of genes based on their expression levels in expression data to obtain a second gene ranking, and using the second gene ranking and a second statistical model trained using second training data indicating a plurality of rankings for at least some of the genes in the second set of genes to determine at least one second characteristic of a biological sample.
[0017] In some embodiments, the at least one second characteristic includes a cancer grade for cells in the biological sample. In some embodiments, the at least one second characteristic includes a tissue of origin for cells in the biological sample.
[0018] In some embodiments, the step of determining the gene ranking includes determining a relative ranking for each gene in the set of genes based on the expression levels. In some embodiments, the step of determining the at least one characteristic further includes providing the gene ranking as an input to the statistical model and obtaining an output indicative of the at least one characteristic. In some embodiments, the statistical model comprises a gradient boosted decision tree classifier. In some embodiments, the statistical model comprises a classifier selected from the group consisting of a gradient boosted decision tree classifier, a decision tree classifier, a gradient boosted classifier, a random forest classifier, a clustering-based classifier, a Bayesian classifier, a Bayesian network classifier, a neural network classifier, a kernel-based classifier, and a support vector machine classifier.
[0019] In some embodiments, the set of genes comprises at least 5 genes. In some embodiments, the set of genes consists of 5 to 50 genes. In some embodiments, the set of genes consists of 5 to 300 genes.
[0020] In some embodiments, the method further comprises presenting to the user an indication of at least one characteristic. In some embodiments, presenting an indication of at least one characteristic further comprises displaying to the user at least one characteristic in a graphical user interface (GUI).
[0021] In some embodiments, at least one characteristic includes a cancer grade for cells in a biological sample, and the cancer grade is selected from the group consisting of grade 1, grade 2, grade 3, grade 4, and grade 5. In some embodiments, at least one characteristic includes the origin tissue for cells in a biological sample, and the origin tissue is selected from the group consisting of lung tissue, pancreatic tissue, stomach tissue, colon tissue, liver tissue, bladder tissue, kidney tissue, thyroid tissue, lymph node tissue, adrenal tissue, skin tissue, breast tissue, ovarian tissue, prostate tissue, urothelial tissue, cervical tissue, esophageal tissue, brain tissue, soft tissue, connective tissue, head tissue, and neck tissue. In some embodiments, at least one characteristic includes the tissue type for cells in a biological sample, and the tissue type is selected from the group consisting of adenocarcinoma, squamous cell carcinoma, carcinoma, cystadenocarcinoma, sarcoma, and glioma.
[0022] In some embodiments, at least one characteristic comprises the human papillomavirus (HPV) status for cells in a biological sample, and the set of genes comprises at least 5 genes selected from the group of genes listed in Table 8 (Table 8). In some embodiments, at least one characteristic comprises the subtype of peripheral T cell lymphoma (PTCL) for cells in a biological sample, and the set of genes comprises at least 5 genes selected from the group of genes listed in Table 10 (Table 10). In some embodiments, the subtype of PTCL is selected from the group consisting of anaplastic large cell lymphoma (ALCL), angioimmunoblastic T cell lymphoma (AITL), natural killer / T cell lymphoma (NKTCL), and adult T cell leukemia / lymphoma (ATLL).
[0023] In some embodiments, the subject has, is suspected of having, or is at risk of having breast cancer, and the set of genes comprises at least 5 genes selected from the group of genes listed in Table 1 (Table 1). In some embodiments, the set of genes comprises at least 10 genes selected from the group of genes listed in Table 1 (Table 1). In some embodiments, the subject has, is suspected of having, or is at risk of having kidney cancer, and the set of genes comprises at least 5 genes selected from the group of genes listed in Table 2 (Table 2). In some embodiments, the subject has, is suspected of having, or is at risk of having lymphoma, and the set of genes comprises at least 5 genes selected from the group of genes listed in Table 3 (Table 3). In some embodiments, the subject has, is suspected of having, or is at risk of having diffuse large B-cell lymphoma (DLBCL), and the set of genes comprises at least 10 genes selected from the group of genes listed in Table 3 (Table 3), and at least one characteristic is an origin cell selected from the group consisting of germinal center B cells (GCB) and activated B cells (ABC). In some embodiments, the subject has, is suspected of having, or is at risk of having lung adenocarcinoma, and the set of genes comprises at least 5 genes selected from the group of genes listed in Table 6 (Table 6).
[0024] In some embodiments, at least one characteristic is selected from the group consisting of cancer grade for cells in a biological sample, origin tissue for cells in a biological sample, tissue type for cells in a biological sample, and cancer subtype for cells in a biological sample.
[0025] In some embodiments, the step of determining at least one characteristic further includes providing gene ranking and obtaining an output indicative of the at least one characteristic as inputs to a statistical model. In some embodiments, the at least one characteristic is selected from the group consisting of cancer grade for cells in a biological sample, origin tissue for cells in a biological sample, tissue type for cells in a biological sample, and cancer subtype for cells in a biological sample.
[0026] In some embodiments, the subject has, is suspected of having, or is at risk of having head and neck squamous cell carcinoma, and the set of genes comprises at least five genes selected from the group of genes described in Table 8. In some embodiments, the set of genes comprises at least ten genes selected from the group of genes described in Table 8.
[0027] In some embodiments, the at least one characteristic includes the human papillomavirus (HPV) status for cells in a biological sample. In some embodiments, the at least one characteristic includes the subtype of peripheral T-cell lymphoma (PTCL) for cells in a biological sample, and the set of genes includes at least five genes selected from the group of genes described in Table 10. In some embodiments, the set of genes comprises at least ten genes selected from the group of genes described in Table 10. In some embodiments, the subtype of PTCL is selected from the group consisting of anaplastic large cell lymphoma (ALCL), angioimmunoblastic T-cell lymphoma (AITL), natural killer / T-cell lymphoma (NKTCL), and adult T-cell leukemia / lymphoma (ATLL).
[0028] Some embodiments are directed to a system comprising at least one hardware processor and at least one non - transitory computer - readable storage medium storing processor - executable instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to execute a method. The method includes obtaining expression data, at least partially obtained by sequencing a biological sample of a subject having, suspected of having, or at risk of having cancer, the expression data comprising expression levels for a plurality of genes, the plurality of genes constituting a set of genes; ranking at least some of the genes in the set of genes based on their expression levels in the expression data to obtain a gene ranking; and determining at least one characteristic of the biological sample using the gene ranking and a statistical model trained using training data showing a plurality of gene rankings for at least some of the genes in the obtained set of genes, wherein each of the plurality of gene rankings is obtained based on respective expression levels for at least some of the genes in the set of genes.
[0029] Some embodiments are directed to at least one non-transitory computer-readable storage medium storing processor-executable instructions, which, when executed by at least one hardware processor, cause the at least one hardware processor to obtain expression data that is at least partially obtained by sequencing a biological sample of a subject having, suspected of having, or at risk of having cancer, wherein the expression data comprises expression levels for a plurality of genes, and the plurality of genes constitute a set of genes, rank at least some of the genes in the set of genes based on their expression levels in the expression data to obtain a gene ranking, and use the gene ranking and a statistical model trained using training data indicating at least some of the genes in the obtained set of genes and a plurality of gene rankings to determine at least one characteristic of the biological sample, wherein each of the plurality of gene rankings is obtained based on respective expression levels for at least some of the genes in the set of genes.
[0030] Some embodiments are methods that use at least one computer hardware processor to obtain expression data for cells in a biological sample from a subject having or suspected of having or at risk of having cancer, rank at least some genes in at least one set of genes based on their expression levels in the expression data to obtain at least one gene ranking, and use at least one statistical model trained using the at least one gene ranking and training data indicating a plurality of rankings for at least some genes in at least one set of genes to determine the tissue of origin for at least some of the cells in the biological sample, wherein each of the plurality of gene rankings is obtained based on respective expression levels of at least some genes in at least one set of genes. The method includes performing the steps of.
[0031] In some embodiments, the expression data is obtained using a gene expression microarray. In some embodiments, the expression data is obtained by performing next-generation sequencing. In some embodiments, the tissue of origin is selected from the group consisting of lung tissue, pancreatic tissue, stomach tissue, colon tissue, liver tissue, bladder tissue, kidney tissue, thyroid tissue, lymph node tissue, adrenal tissue, skin tissue, breast tissue, ovarian tissue, prostate tissue, urothelial tissue, cervical tissue, esophageal tissue, brain tissue, soft tissue, connective tissue, head tissue, and neck tissue.
[0032] In some embodiments, the method further comprises determining a tissue type for at least some of the cells in a biological sample using at least one gene ranking and at least one statistical model. In some embodiments, the tissue type is selected from the group consisting of adenocarcinoma, squamous cell carcinoma, carcinoma, cystadenocarcinoma, sarcoma, and glioma. In some embodiments, the combination of origin tissue and tissue type is selected from the group consisting of lung adenocarcinoma, lung squamous cell carcinoma, melanoma, breast cancer, colorectal adenocarcinoma, ovarian serous cystadenocarcinoma, pheochromocytoma, urothelial carcinoma of the bladder, cervical squamous cell carcinoma, glioblastoma multiforme, head squamous cell carcinoma, neck squamous cell carcinoma, kidney renal clear cell carcinoma, kidney renal papillary cell carcinoma, hepatocellular carcinoma of the liver, pancreatic adenocarcinoma, paraganglioma, prostate adenocarcinoma, sarcoma, gastric adenocarcinoma, thyroid cancer, and uterine corpus endometrial carcinoma.
[0033] In some embodiments, the subject has, is suspected of having, or is at risk of having lymphoma. In some embodiments, the subject has, is suspected of having, or is at risk of having diffuse large B-cell lymphoma (DLBCL). In some embodiments, the origin tissue is an origin cell selected from the group consisting of germinal center B cells (GCB) and activated B cells (ABC). In some embodiments, a set of genes of at least one set of genes is selected from the gene groups listed in Table 3. In some embodiments, a set of genes of at least one set of genes comprises at least three genes selected from the gene groups listed in Table 3. In some embodiments, a set of genes of at least one set of genes comprises at least five genes selected from the gene groups listed in Table 3. In some embodiments, a set of genes of at least one set of genes comprises at least ten genes selected from the gene groups listed in Table 3.
[0034] In some embodiments, a set of genes of at least one set of genes comprises at least five genes. In some embodiments, a set of genes of at least one set of genes consists of 5 to 100 genes. In some embodiments, a set of genes of at least one set of genes consists of 10 to 200 genes. In some embodiments, a set of genes of at least one set of genes consists of 20 to 100 genes. In some embodiments, a set of genes of at least one set of genes consists of 50 to 100 genes.
[0035] In some embodiments, the expression data includes values each representing an expression level for a gene in at least one set of genes, and the step of determining a gene ranking among at least one gene ranking includes determining a relative rank for each gene in one of at least one set of genes based on the values. In some embodiments, the step of determining the origin tissue further includes using at least one gene ranking as an input to at least one statistical model and obtaining an output indicating the origin tissue.
[0036] In some embodiments, at least one statistical model comprises a gradient boosting decision tree classifier. In some embodiments, at least one statistical model comprises at least one classifier selected from the group consisting of a gradient boosting decision tree classifier, a decision tree classifier, a gradient boosting classifier, a random forest classifier, a clustering-based classifier, a Bayesian classifier, a Bayesian network classifier, a neural network classifier, a kernel-based classifier, and a support vector machine classifier.
[0037] In some embodiments, at least one set of genes comprises a first set of genes associated with predicting a first type of tissue and a second set of genes associated with predicting a second type of tissue.
[0038] Some embodiments are directed to a system comprising at least one hardware processor and at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to execute a method. The method includes obtaining expression data for cells in a biological sample of a subject having, suspected of having, or at risk of having cancer; ranking at least some genes in at least one set of genes based on their expression levels in the expression data to obtain at least one gene ranking; and determining an origin tissue for at least some of the cells in the biological sample using the at least one gene ranking and at least one statistical model trained using training data indicating a plurality of rankings for at least some genes in at least one set of genes, wherein each of the plurality of gene rankings is obtained based on respective expression levels of at least some genes in at least one set of genes.
[0039] Some embodiments are directed to at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one hardware processor, cause the at least one hardware processor to obtain expression data for cells in a biological sample of a subject having, suspected of having, or at risk of having cancer; rank at least some genes in at least one set of genes based on their expression levels in the expression data to obtain at least one gene ranking; and use the at least one gene ranking and at least one statistical model trained using training data indicating a plurality of rankings for at least some genes in at least one set of genes to determine an origin tissue for at least some of the cells in the biological sample, wherein each of the plurality of gene rankings is obtained based on respective expression levels for at least some genes in at least one set of genes, to cause the determination to be performed.
[0040] Some embodiments are a method comprising: using at least one computer hardware processor to obtain expression data for cells in a biological sample of a subject having, suspected of having, or at risk of having cancer; ranking at least some genes in a set of genes based on their expression levels in the expression data to obtain a gene ranking; and using the gene ranking and a statistical model trained using training data indicating a plurality of rankings for at least some genes in the set of genes to determine a cancer grade for at least some of the cells in the biological sample, wherein each of the plurality of gene rankings is obtained based on respective expression levels for at least some genes in the set of genes.
[0041] In some embodiments, the expression data was obtained using a gene expression microarray. In some embodiments, the expression data was obtained by performing next-generation sequencing. In some embodiments, the cancer grade is selected from the group consisting of at least grade 1, grade 2, and grade 3. In some embodiments, the cancer grade is selected from the group consisting of at least grade 1, grade 2, grade 3, and grade 4. In some embodiments, the cancer grade is selected from the group consisting of grade 1, grade 2, grade 3, grade 4, and grade 5.
[0042] In some embodiments, the subject has, is suspected of having, or is at risk of having breast cancer. In some embodiments, the set of genes is selected from the group of genes listed in Table 1. In some embodiments, the set of genes comprises at least 3 genes selected from the group of genes listed in Table 1. In some embodiments, the set of genes comprises at least 5 genes selected from the group of genes listed in Table 1. In some embodiments, the set of genes comprises at least 10 genes selected from the group of genes listed in Table 1.
[0043] In some embodiments, the subject has, is suspected of having, or is at risk of having renal cancer. In some embodiments, the subject has, is suspected of having, or is at risk of having clear cell renal cancer. In some embodiments, the set of genes is selected from the group of genes listed in Table 2. In some embodiments, the set of genes comprises at least 3 genes selected from the group of genes listed in Table 2. In some embodiments, the set of genes comprises at least 5 genes selected from the group of genes listed in Table 2. In some embodiments, the set of genes comprises at least 10 genes selected from the group of genes listed in Table 2.
[0044] In some embodiments, the subject has, is suspected of having, or is at risk of having lung adenocarcinoma. In some embodiments, the set of genes is selected from the group of genes listed in Table 6. In some embodiments, the set of genes comprises at least 10 genes selected from the group of genes listed in Table 6. In some embodiments, the set of genes comprises at least 50 genes. In some embodiments, the set of genes consists of 10 to 100 genes. In some embodiments, the set of genes consists of 10 to 30 genes.
[0045] In some embodiments, the expression data includes values representing the expression levels for the genes in the set of genes, and the step of determining a gene ranking includes determining a relative rank for each gene in the set of genes based on the values. In some embodiments, the step of determining a cancer grade further includes using the gene ranking as an input to a statistical model and obtaining an output indicating the cancer grade.
[0046] In some embodiments, the statistical model comprises a gradient boosted decision tree classifier. In some embodiments, the statistical model comprises a classifier selected from the group consisting of a gradient boosted decision tree classifier, a decision tree classifier, a gradient boost classifier, a random forest classifier, a clustering-based classifier, a Bayesian classifier, a Bayesian network classifier, a neural network classifier, a kernel-based classifier, and a support vector machine classifier.
[0047] Some embodiments are directed to a system comprising at least one hardware processor and at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to execute a method. The method includes obtaining expression data for cells in a biological sample of a subject who has or is suspected of having or is at risk of having cancer, ranking at least some genes in a set of genes based on their expression levels in the expression data to obtain a gene ranking, and using the gene ranking and a statistical model trained using training data indicating a plurality of rankings for at least some genes in the set of genes, wherein each of the plurality of gene rankings is obtained based on respective expression levels of at least some genes in the set of genes, to determine a cancer grade for at least some of the cells in the biological sample.
[0048] Some embodiments are directed to at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one hardware processor, cause the at least one hardware processor to obtain expression data for cells in a biological sample of a subject having, suspected of having, or at risk of having cancer, rank at least some of the genes in a set of genes based on their expression levels in the expression data to obtain a gene ranking, and use the gene ranking and a statistical model trained using training data indicating a plurality of rankings for at least some of the genes in the set of genes to determine a cancer grade for at least some of the cells in the biological sample, wherein each of the plurality of gene rankings is obtained based on respective expression levels of at least some of the genes in the set of genes, to perform the determining.
[0049] Some embodiments are methods that include obtaining, using at least one computer hardware processor, expression data for cells in a biological sample of a subject having, suspected of having, or at risk of having cancer, ranking at least some of the genes in at least one set of genes based on their expression levels in the expression data to obtain at least one gene ranking, and using the at least one gene ranking and at least one statistical model to determine a subtype of peripheral T cell lymphoma (PTCL) for at least some of the cells in the biological sample.
[0050] In some embodiments, at least one statistical model is trained using training data that indicates a plurality of rankings of expression levels for at least some genes in at least one set of genes. In some embodiments, each of the plurality of gene rankings is obtained based on respective expression levels for at least some genes in at least one set of genes.
[0051] In some embodiments, expression data is obtained using a gene expression microarray. In some embodiments, expression data is obtained by performing next-generation sequencing. In some embodiments, expression data is obtained using a hybridization-based expression assay.
[0052] In some embodiments, the subtype of PTCL is selected from the group consisting of anaplastic large cell lymphoma (ALCL), angioimmunoblastic T-cell lymphoma (AITL), natural killer / T-cell lymphoma (NKTCL), and adult T-cell leukemia / lymphoma (ATLL). In some embodiments, the subtype of PTCL is selected from the group consisting of peripheral T-cell lymphoma, not otherwise specified (PTCL-NOS), anaplastic large cell lymphoma (ALCL), angioimmunoblastic T-cell lymphoma (AITL), cutaneous T-cell lymphoma (CTCL), natural killer / T-cell lymphoma (NKTCL), Sézary syndrome, adult T-cell leukemia / lymphoma (ATLL), enteropathy-type T-cell lymphoma, nasal NK / T-cell lymphoma, hepatosplenic gamma-delta T-cell lymphoma, T-cell lymphomas of follicular T-cell (TFH) origin, and T-cell lymphomas of the gastrointestinal tract.
[0053] In some embodiments, the set of genes of at least one set of genes is selected from the group of genes described in Table 10. In some embodiments, the set of genes of at least one set of genes comprises at least three genes selected from the group of genes described in Table 10. In some embodiments, the set of genes of at least one set of genes comprises at least five genes selected from the group of genes described in Table 10. In some embodiments, the set of genes of at least one set of genes comprises at least ten genes selected from the group of genes described in Table 10. In some embodiments, the set of genes of at least one set of genes comprises at least fifty genes selected from the group of genes described in Table 10.
[0054] In some embodiments, the set of genes of at least one set of genes comprises at least one gene that is up-regulated in AITL. In some embodiments, the set of genes of at least one set of genes comprises at least one gene that is down-regulated in AITL. In some embodiments, the set of genes of at least one set of genes comprises at least one MF profile gene.
[0055] In some embodiments, the subject has, is suspected of having, or is at risk of having lymphoma. In some embodiments, the subject has, is suspected of having, or is at risk of having peripheral T cell lymphoma (PTCL).
[0056] In some embodiments, a set of genes among at least one set of genes contains at least 5 genes. In some embodiments, a set of genes among at least one set of genes consists of 5 to 100 genes. In some embodiments, a set of genes among at least one set of genes consists of 10 to 200 genes. In some embodiments, a set of genes among at least one set of genes consists of 20 to 100 genes. In some embodiments, a set of genes among at least one set of genes consists of 50 to 100 genes.
[0057] In some embodiments, expression data includes values each representing an expression level for a gene in at least one set of genes, and the step of determining a gene ranking among at least one gene ranking includes determining a relative rank for each gene in one of at least one set of genes based on the values.
[0058] In some embodiments, the step of determining a subtype of PTCL further includes using at least one gene ranking as an input to at least one statistical model and obtaining an output indicating the subtype of PTCL.
[0059] In some embodiments, at least one statistical model comprises a gradient boosted decision tree classifier. In some embodiments, at least one statistical model comprises at least one classifier selected from the group consisting of a gradient boosted decision tree classifier, a decision tree classifier, a gradient boost classifier, a random forest classifier, a clustering-based classifier, a Bayesian classifier, a Bayesian network classifier, a neural network classifier, a kernel-based classifier, and a support vector machine classifier.
[0060] In some embodiments, at least one statistical model includes a multi-class classifier. In some embodiments, the multi-class classifier has at least four outputs, each corresponding to a different subtype of PTCL. In some embodiments, the at least four outputs include a first output corresponding to anaplastic large cell lymphoma (ALCL), a second output corresponding to angioimmunoblastic T cell lymphoma (AITL), a third output corresponding to natural killer / T cell lymphoma (NKTCL), and a fourth output corresponding to adult T cell leukemia / lymphoma (ATLL).
[0061] In some embodiments, at least one statistical model comprises a plurality of classifiers corresponding to different subtypes of PTCL. In some embodiments, the plurality of classifiers includes a first classifier, a second classifier, a third classifier, and a fourth classifier, where the first classifier corresponds to anaplastic large cell lymphoma (ALCL), the second classifier corresponds to angioimmunoblastic T cell lymphoma (AITL), the third classifier corresponds to natural killer / T cell lymphoma (NKTCL), and the fourth classifier corresponds to adult T cell leukemia / lymphoma (ATLL). In some embodiments, at least one set of genes includes a first set of genes associated with the first classifier among the plurality of classifiers and a second set of genes associated with the second classifier among the plurality of classifiers.
[0062] In some embodiments, the subject has, is suspected of having, or is at risk of having lymphoma. In some embodiments, the subject has, is suspected of having, or is at risk of having PTCL.
[0063] In some embodiments, the method further includes presenting to the user an indication of a subtype of PTCL. In some embodiments, the step of presenting an indication of a subtype of PTCL further includes displaying to the user a subtype of PTCL in a graphical user interface (GUI).
[0064] Some embodiments are directed to a system comprising at least one hardware processor and at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to perform a method. The method includes obtaining expression data for cells in a biological sample of a subject having, suspected of having, or at risk of having cancer; ranking at least some genes in at least one set of genes based on their expression levels in the expression data to obtain at least one gene ranking; and using the at least one gene ranking and at least one statistical model to determine a subtype of peripheral T-cell lymphoma (PTCL) for at least some of the cells in the biological sample.
[0065] Some embodiments are directed to at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one hardware processor, cause the at least one hardware processor to obtain expression data for cells in a biological sample of a subject having, suspected of having, or at risk of having cancer, rank at least some genes in at least one set of genes based on their expression levels in the expression data to obtain at least one gene ranking, and use the at least one gene ranking and at least one statistical model to determine a subtype of peripheral T-cell lymphoma (PTCL) for at least some of the cells in the biological sample.
[0066] Various aspects and embodiments are described with reference to the following figures. The figures are not necessarily drawn to scale.
Brief Description of the Drawings
[0067]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 6C
Figure 6D
Figure 7
Figure 8A
Figure 8B
Figure 9A
Figure 9B
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23A
Figure 23B
Figure 23C
Figure 24A
Figure 24B
Figure 24C
Figure 24D
Figure 24E
Figure 25A
Figure 25B
Figure 25C
Figure 25D
Figure 25E
Figure 25F
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
DETAILED DESCRIPTION OF THE INVENTION
[0068] The characteristics of a somatic cell can be related to the expression levels of several genes. For example, cancer cells may have some genes that are upregulated relative to normal, healthy cells and other genes that are downregulated. This relationship between cell characteristics and gene expression levels can be utilized when analyzing gene expression data for somatic cells. In particular, such a relationship can provide several benefits when analyzing the characteristics of somatic cells, which are considered histological characteristics, including origin tissue and cancer grade, that are generally related to the features of the somatic cells visually observed by a person (e.g., a pathologist). In some cases, gene expression data can provide a more consistent assessment of certain cell characteristics than by using histological techniques, which may be subject to variations among pathologists in their assessments.
[0069] Large amounts of gene expression data can be obtained through different platforms, including those by using gene expression microarrays and by performing next generation sequencing, and are currently available or can be generated to characterize living cells. However, the inventors recognize that information derivable from these data is compromised by differences between different gene sequencing platforms, and that such differences can lead to variations in gene expression data generated by those sequencing platforms even when those sequencing platforms are used to sequence the same biological sample. For example, microarray and next generation sequencing (NGS) techniques can result in gene expression data where specific values representing gene expression levels can vary between platforms even when obtained from the same biological sample. This variation in expression values across different sequencing platforms can occur due to the way the expression data is obtained. The processes and devices used to obtain gene expression data using a particular type of sequencing platform (e.g., next generation sequencing, microarray) can affect the specific values for the expression levels obtained. Next, the values for the expression levels depend on which sequencing platform was used to obtain the gene expression data. This variation can occur not only across different types of sequencing platforms, but also when different sequencing platforms are of the same type (e.g., next generation sequencing), and are associated with different systems (e.g., optical systems, detectors) and processes (e.g., biological sample preparation), or even the same device in different locations (e.g., due to differences in calibration, use, environment, etc.).
[0070] The inventor recognizes that such variations in expression level values pose significant challenges when analyzing gene expression data to characterize cells, particularly when using gene expression data obtained using different sequencing platforms. In some cases of expression data, it can be an issue to normalize the expression level values in such a way that expression data obtained using different sequencing platforms can be analyzed using the same or similar techniques.
[0071] Conventional techniques for analyzing expression data are generally applicable only to the analysis of expression data obtained using a single sequencing platform and to only the specific conditions used in sample preparation and sequencing. Such conventional techniques are not applicable to the analysis of expression data obtained from those sequencing platforms, even when the multiple sequencing platforms are of the same type (e.g., next-generation sequencing, microarray). For example, conventional techniques for analyzing gene expression data may involve different data analysis pipelines for expression data obtained using different next-generation sequencing devices. In addition, some conventional techniques involve implementing different data analysis pipelines depending on how the expression data was obtained, even when the same sequencing device was used. For example, conventional techniques for analyzing gene expression data may vary depending on different sequencing conditions or different sample processing methods. As a result, conventional techniques for analyzing expression data cannot be implemented across different sequencing platforms, sample preparation techniques, and sequencing conditions. This significantly affects the usefulness of gene expression data for determining the characteristics of cells.
[0072] One important group of techniques for analyzing expression data includes statistical models (e.g., machine learning models) that receive expression level values (or derivatives thereof) as input and are configured to produce an output of interest, such as a prediction or classification. Examples of such statistical models developed by the inventor are provided herein. Prior to being used, such statistical models are trained in training data comprising input / output pairs. If the training data input comprises expression level values (or derivatives thereof) coming from one type of sequencing platform, a statistical model trained using such data will tend to exhibit inadequate performance when expression level values coming from another type of sequencing platform are provided (in the task for which it was trained). In fact, due to the variation across expression level values from different sequencing platforms, it becomes difficult or impossible to design a single statistical model trained to perform a task using data from any one of a plurality of types of sequencing platforms. Instead, separate statistical models must be trained for each specific sequencing platform using training data obtained for that specific sequencing platform, which is difficult because it requires training multiple models for each platform, thus not only requiring additional computational resources but also simply being infeasible in cases where there is not sufficient training data available for each type of platform.
[0073] Despite differences in the types of expression level data generated by platforms, the inventor recognizes the need for common techniques that can be used to analyze expression data obtained across different sequencing platforms. Such techniques will facilitate the analysis of gene expression data across different subjects, something that conventional gene expression level analysis techniques have not attempted to enable. For example, the techniques described herein for analyzing gene expression data may involve using the same or similar data analysis pipelines (which may include one or more statistical models, examples of which are provided herein) for expression data obtained using the same type of sequencing platform (e.g., next-generation sequencing, microarray) for multiple subjects. Such data analysis pipelines may enable expression data to be analyzed in the same or a similar manner regardless of the sample processing (e.g., DNA extraction, amplification), sequencing conditions (e.g., temperature, pH), and data processing (e.g., data processing for next-generation sequencing, microarray) used to obtain the expression data.
[0074] To address some of the difficulties that arise with conventional techniques for analyzing expression data, the inventor has developed improved techniques for analyzing expression data that are independent of the sequencing platform and data processing used to obtain the expression data. In particular, the inventor has recognized that variations in expression levels between sequencing platforms can be accounted for in subsequent data analysis by using the ranking of sets of genes rather than specific values of expression levels in the data. For example, the inventor has developed various statistical models for determining various characteristics of a biological sample (e.g., the origin tissue for a tissue sample, the cancer grade, the cancer type). Each such statistical model is trained to determine each characteristic of the biological sample using the ranking of each set of genes rather than the expression levels themselves, thereby enabling the statistical model to operate on expression data obtained from different types of sequencing platforms.
[0075] Thus, in some embodiments, a statistical model can be used to predict characteristics of a biological sample based on an input ranking of genes ranked according to their respective expression levels for a sequencing platform. By using the input ranking instead of a specific value for the expression level, the same or a similar data processing pipeline can be used across different expression data, regardless of the specific method by which the expression levels were obtained (e.g., regardless of which sequencing platform, sequencing conditions, sample preparation, data processing, etc. were used to obtain the expression levels). As described herein, the statistical model can be specific to the particular characteristic being determined. Statistical models according to the techniques described herein can be used to predict one or more characteristics, where the characteristics include cancer grade (e.g., breast cancer grade, renal clear cell cancer grade, lung adenocarcinoma grade) for cells in a biological sample, tissue of origin (e.g., lung, pancreas, stomach, colon, liver, bladder, kidney, thyroid, lymph node, adrenal gland, skin, breast, ovary, prostate, or origin cells in a tissue such as germinal center B cells (GCB) or activated B cells (ABC)) for cells in a biological sample, histological information (e.g., tissue types such as adenocarcinoma, squamous cell carcinoma, carcinoma, cystadenocarcinoma, sarcoma, and glioma) for cells in a biological sample, and cancer subtype (e.g., PTCL subtypes such as anaplastic large cell lymphoma (ALCL), angioimmunoblastic T cell lymphoma (AITL), natural killer / T cell lymphoma (NKTCL), and adult T cell leukemia / lymphoma (ATLL)) for cells in a biological sample, viral status (e.g., HPV status such as HPV positive or HPV negative for head and neck squamous cell carcinoma).
[0076] For example, in some embodiments, a ranking of genes based on gene expression levels (in a biological sample) determined by a sequencing platform can be provided as an input to a statistical model trained to predict the origin tissue for the biological sample. As another example, in some embodiments, a ranking of genes based on gene expression levels (in a biological sample) determined by a sequencing platform can be provided as an input to a statistical model trained to predict the cancer grade for the biological sample. In some embodiments, the set of ranked genes depends on a particular biological property of interest. For example, one set of genes can be used to determine the origin tissue and another set of genes can be used to determine the cancer grade.
[0077] Machine learning techniques involving using gene rankings as described herein are an improvement over conventional machine learning techniques because they are superior to conventional machine learning techniques that directly use gene expression values for analyzing gene expression data. For example, due to the benefits provided by using gene rankings when enabling a common statistical model to be implemented regardless of how the expression data is generated, training data obtained using different sequencing platforms can be used when training the statistical models described herein. In contrast, conventional machine learning techniques involving using gene expression values require individual distinct statistical models depending on how the expression data is generated, such as when using different sequencing platforms, sample preparation techniques, etc. Therefore, the machine learning techniques described herein reduce the need to collect training data across different sequencing platforms for training multiple statistical models required for analyzing expression data generated in different ways. Additionally, the statistical models described herein may have better performance as compared to conventional techniques. For example, statistical models according to the techniques described herein can be trained using training data obtained from different sources and thus generally more training data, thereby improving the overall performance of the statistical model being used. In contrast, the sources of training data for conventional machine learning models can be limited to specific sequencing platforms, sample preparation techniques, etc., and the performance can depend on the amount of training data available using the specific method of generating the expression data.
[0078] In addition, having a statistical model that is independent of the sequencing platform, sample preparation, and sequencing conditions used can make the development and use of such statistical models more practical. In actual clinical practice, data from different patients can originate from multiple sources, such as expression data generated using different sample preparation techniques and sequencing platforms. As described above, the techniques described herein enable the ability to uniformly handle patient data originating from these different sources by using a common statistical model. The ability to analyze patient data in this way provides an improvement in bioinformatics technology that depends on the number of patients represented by the patient data, because a larger pool of patients can be analyzed using a common statistical model. These benefits extend to applications in which bioinformatics analysis can be used, including predicting the characteristics of cells in a biological sample, and in that case, it is advantageous to be able to use a larger sample size across a number of patients.
[0079] Furthermore, the machine learning techniques described herein can rationalize the handling of different formats for storing expression data. Different types of sequencing platforms output expression data using different data formats. As described herein, a ranking process is used to generate a gene ranking, and then the gene ranking is input into a common statistical model. The ranking process can enable expression data originating from sources using different data formats to have a similar type of input to the statistical model. This can improve the handling of expression data obtained from different sequencing platforms compared to conventional analysis techniques, which require different data processing pipelines for different input data formats.
[0080] Some embodiments described herein address all of the problems described above that the inventor has recognized with respect to using gene expression data to determine the characteristics of a biological sample. However, not every embodiment described herein addresses every one of these problems, and some embodiments may not address any of them. Thus, it should be understood that the embodiments of the techniques described herein are not limited to addressing all or any of the problems described above with respect to using gene expression data to determine the characteristics of a biological sample.
[0081] Some embodiments involve obtaining gene expression data for a biological sample of interest, ranking genes in a set of genes based on their expression levels in the expression data, and obtaining one or more gene rankings. The one or more gene rankings can be used with a statistical model to determine one or more characteristics of the biological sample, including origin tissue and grade of cancer. The statistical model can be trained using rankings of expression levels for some or all of the genes in the set of genes.
[0082] Gene ranking can be obtained by ranking genes in one or more sets of genes based on their expression levels in expression data. In some embodiments, the expression data includes values each representing an expression level for a gene in a set of genes. Determining a gene ranking can involve determining a relative rank for each gene in the set of genes based on the values. For example, a first gene ranking can be obtained by ranking the genes in a first set of genes based on their expression levels, and a second gene ranking can be obtained by ranking the genes in a second set of genes based on their expression levels. In some embodiments, the first set of genes and the second set of genes can share some or all of the genes. Determining one or more characteristics can involve using the first gene ranking, the second gene ranking, and a statistical model, where the statistical model is trained using training data indicative of the gene rankings of the expression levels for some or all of the genes in the first set of genes and the second set of genes. Different gene sets can correspond to predicting a particular characteristic of a biological sample, and a gene ranking for a particular gene set can be used to determine the characteristic associated with that gene set. For example, if the expression levels for a gene set are associated with predicting a cancer grade, the gene ranking can be used to predict the cancer grade for cells in the biological sample from which the expression data is obtained.
[0083] In some embodiments, the expression data can be obtained for cells in a biological sample, where the subject has or is suspected of having cancer. In the context where the origin tissue is the characteristic being determined, the origin tissue is for the cells in the biological sample. The origin tissue can refer to a specific tissue type from which the cells originate, such as lung, pancreas, stomach, colon, liver, bladder, kidney, thyroid, lymph node, adrenal gland, skin, breast, ovary, and prostate.
[0084] For example, some embodiments involve using a gene set for predicting the tissue of origin, which may include cells of origin, for diffuse large B-cell lymphoma (DLBCL), such as germinal center B cells (GCB) and activated B cells (ABC). Genes in the gene set can be selected from the group consisting of ITPKB, MYBL1, LMO2, BATF, IRF4, LRMP, CCND2, SLA, SP140, PIM1, CSTB, BCL2, TCF4, P2RX5, SPINK2, VCL, PTPN1, REL, FUT8, RPL21, PRKCB1, CSNK1E, GPR18, IGHM, ACP1, SPIB, HLA-DQA1, KRT8, FAM3C, and HLA-DMB.
[0085] In the context where cancer grade is a determined characteristic, the cancer grade is for cells in a biological sample. Cancer grade may refer to the growth and differentiation characteristics of cells in a biological sample, and may refer to numerical grades generally determined by visual observation of cells using microscopy, such as grade 1, grade 2, grade 3, and grade 4. For example, a pathologist may examine tissue biopsied under a microscope and determine the cancer grade for the tissue. Cancer grade generally depends on the amount of cell abnormalities in the tissue and may depend on the cancer type. In grade 1, the tumor cells, and the organization of the tumor tissue, appear similar to normal healthy tissue. Grade 1 tumors tend to grow and spread slowly. In contrast, the cells and tissues of grade 3 and grade 4 tumors do not appear like normal cells and tissues. Grade 3 and grade 4 tumors tend to grow rapidly and spread faster than tumors with lower grades. An exemplary grading system for cancer tissue is described in American Joint Committee on Cancer AJCC Cancer Staging Manual. 7th ed. New York, NY: Springer, 2010, which is incorporated by reference in its entirety. This grading system applies the following definitions, namely, grade X (GX) is an undetermined grade and is applied when the grade of the tissue cannot be assessed, grade 1 (G1) is a low grade and is applied when the cells are well-differentiated, grade 2 (G2) is an intermediate grade and is applied when the cells are moderately differentiated, grade 3 (G3) is a high grade and is applied when the cells are poorly-differentiated, and grade 4 (G4) is a high grade and is applied when the cells are undifferentiated.
[0086] For example, some embodiments involve using a gene set to predict breast cancer grade. The genes in the gene set can be selected from the group consisting of UBE2C, MYBL2, PRAME, LMNB1, CXCL9, KPNA2, TPX2, PLCH1, CCL18, CDK1, MELK, CCNB2, RRM2, CCNB1, NUSAP1, SLC7A5, TYMS, GZMK, SQLE, C1orf106, CDC25B, ATAD2, QPRT, CCNA2, NEK2, IDO1, NDC80, ZWINT, ABCA12, TOP2A, TDO2, S100A8, LAMP3, MMP1, GZMB, BIRC5, TRIP13, RACGAP1, ASPM, ESRP1, MAD2L1, CENPF, CDC20, MCM4, MKI67, PBK, CKS2, KIF2C, MRPL13, TTK, BUB1, TK1, FOXM1, CEP55, EZH2, ECT2, PRC1, CENPU, CCNE2, AURKA, HMGB3, APOBEC3B, LAGE3, CDKN3, DTL, ATP6V1C1, KIAA0101, CD2, KIF11, KIF20A, CDCA8, NCAPG, CENPN, MTFR1, MCM2, DSCC1, WDR19, SEMA3G, KCND3, SETBP1, KIF13B, NR4A2, NAV3, PDZRN3, MAGI2, CACNA1D, STC2, CHAD, PDGFD, ARMCX2, FRY, AGTR1, MARCH8, ANG, ABAT, THBD, RAI2, HSPA2, ERBB4, ECHDC2, FST, EPHX2, FOSB, STARD13, ID4, FAM129A, FCGBP, LAMA2, FGFR2, PTGER3, NME5, LRRC17, OSBPL1A, ADRA2A, LRP2, C1orf115, COL4A5, DIXDC1, KIAA1324, HPN, KLF4, SCUBE2, FMO5, SORBS2, CARD10, CITED2, MUC1, BCL2, RGS5, CYBRD1, OMD, IGFBP4, LAMB2, DUSP4, PDLIM5, IRS2, and CX3CR1.
[0087] As another example, some embodiments involve using a gene set to predict renal clear cell carcinoma grade. The genes in the gene set can be selected from the group consisting of PLTP, C1S, LY96, TSKU, TPST2, SERPINF1, SRPX2, SAA1, CTHRC1, GFPT2, CKAP4, SERPINA3, CFH, PLAU, BASP1, PTTG1, MOCOS, LEF1, SLPI, PRAME, STEAP3, LGALS2, CD44, FLNC, UBE2C, CTSK, SULF2, TMEM45A, FCGR1A, PLOD2, C19orf80, PDGFRL, IGF2BP3, SLC7A5, PRRX1, RARRES1, LHFPL2, KDELR3, TRIB3, IL20RB, FBLN1, KMO, C1R, CYP1B1, KIF2A, PLAUR, CKS2, CDCP1, SFRP4, HAMP, MMP9, SLC3A1, NAT8, FRMD3, NPR3, NAT8B, BBOX1, SLC5A1, GBA3, EMCN, SLC47A1, AQP1, PCK1, UGT2A3, BHMT, FMO1, ACAA2, SLC5A8, SLC16A9, TSPAN18, SLC17A3, STK32B, MAP7, MYLIP, SLC22A12, LRP2, CD34, PODXL, ZBTB42, TEK, FBP1, and BCL2.
[0088] As another example, some embodiments involve using a gene set to predict cancer grade for lung adenocarcinoma. The genes in the gene set can be selected from the group consisting of AADAC, ALDOB, ANXA10, ASPM, BTNL8, CEACAM8, CENPA, CHGB, CHRNA9, COL11A1, CRABP1, F11, GGTLC1, HJURP, IGF2BP3, IHH, KCNE2, KIF14, LRRC31, MYBL2, MYOZ1, PCSK2, PI15, SCTR, SHH, SLC22A3, SLC7A5, SPOCK1, TM4SF4, TRPM8, and YBX2.
[0089] Some embodiments involve predicting the origin cells for diffuse large B cell lymphoma (DLBCL) for a biological sample using the machine learning techniques described herein. Such embodiments can involve using gene sets for predicting origin cells, such as germinal center B cells (GCB) and activated B cells (ABC). The genes in the gene set can be selected from the group consisting of ITPKB, MYBL1, LMO2, BATF, IRF4, LRMP, CCND2, SLA, SP140, PIM1, CSTB, BCL2, TCF4, P2RX5, SPINK2, VCL, PTPN1, REL, FUT8, RPL21, PRKCB1, CSNK1E, GPR18, IGHM, ACP1, SPIB, HLA-DQA1, KRT8, FAM3C, and HLA-DMB.
[0090] Some embodiments involve predicting subtypes of peripheral T-cell lymphoma (PTCL) for a biological sample using the machine learning techniques described herein. Such embodiments may involve using gene sets for predicting PTCL subtypes such as anaplastic large cell lymphoma (ALCL), angioimmunoblastic T-cell lymphoma (AITL), natural killer / T-cell lymphoma (NKTCL), and adult T-cell leukemia / lymphoma (ATLL). Genes in the gene sets may be EFNB2, ROBO1, S1PR3, ANK2, LPAR1, SNAP91, SOX8, RAMP3, TUBB2B, ARHGEF10, NOTCH1, ZBTB17, CCNE1, FGF18, MYCN, PTHLH, SMARCA2, WNK1, NKX2-1, CYP26A1, HPSE, CTLA4, PELI1, PRKCB, SPAST, ALS2, KIF3B, ZFYVE27, GF18, FNTB, REL, DMRT1, SLC19A2, STK3, PERP, TNFRSF8, TMOD1, BATF3, CDC14B, WDFEY3, AGT, ALK, ANXA3, BTBD11, CCNA1, DNER, GAS1, HS6ST2, IL1RAP, PCOLCE2, PDE4DIP, SLC16A3, TIAM2, TUBB6, WNT7B, SMOX, TMEM158, NLRP7, ADRB2, GALNT2, HRASLS, CD244, FASLG, KIR2DL4, LOC100287534, KLRD1, SH2D1B, KLRC2, NCAM1, CXCR5, IL6, ICOS, CD40LG, CD84, IL21, BCL6, MAF, SH2D1A, IL4, PTPN1, PIM1, ENTPD1, IRF4, CCND2, IL16, ETV6, BLNK, SH3BP5, FUT8, CCR4, GATA3, IL5, IL10, IL13, MMEITPKB, MYBL1, LRMP, KIAA0870, LMO2, CR1, LTBR, PDPN, TNFRSF1A, FCER2, ICAM1, FCGR2B, IKZF2, CCR8, TNFRSF18, IKZF4, FOXP3, IL2, TBX21, IFNG, GZMH, GNLY, EOMES, NCR1, GZMB, NKG7, FGFBP2, KLRF1, CD160, KLRK1, CD226, NCR3, TNFRSF8, BATF3, TM It may be selected from the group consisting of OD1, TMEM158, MSC, and POPDC3.
[0091] Some embodiments involve predicting the viral state for a biological sample using the machine learning techniques described herein. In some embodiments, the viral state is the human papillomavirus (HPV) state (e.g., HPV positive state, HPV negative state) for a biological sample. In some embodiments, the HPV state may be determined for a subject having, suspected of having, or at risk of having squamous cell carcinoma of the head and neck. The genes in the gene set may be selected from the group consisting of APOBEC3B, ATAD2, BIRC5, CCL20, CCND1, CDC45, CDC7, CDK1, CDKN2A, CDKN2C, CDKN3, CENPF, CENPN, CXCL14, DCN, DHFR, DKK3, DLGAP5, EPCAM, FANCI, FEN1, GMNN, GPX3, ID4, IGLC1, IL18, IL1R2, KIF18B, KIF20A, KIF4A, KLK13, KLK7, KLK8, KNTC1, KRT19, LAMP3, LMNB1, MCM2, MCM4, MCM5, ME1, MELK, MKI67, MLF1, MMP12, MTHFD2, NDN, NEFH, NEK2, NUP155, NUP210, NUSAP1, PDGFD, PLAGL1, PLOD2, PPP1R3C, PRIM1, PRKDC, PSIP1, RAD51AP1, RASIP1, RFC5, RNASEH2A, RPA2, RPL39L, RSRC1, RYR1, SLC35G2, SMC2, SPARCL1, STMN1, SYCP2, SYNGR3, TIMELESS, TMPO, TPX2, TRIP13, TYMS, UCP2, UPF3B, USP1, ZSCAN18.
[0092] It should be understood that the various aspects and embodiments described herein can be used individually, all together, or in any combination of two or more, because the techniques described herein are not limited in this regard.
[0093] FIG. 1 is a diagram of an exemplary processing pipeline 100 for determining one or more characteristics (e.g., origin tissue, cancer grade, PTCL subtype) of a biological sample based on one or more respective gene rankings of the biological sample according to some embodiments of the technology described herein. The exemplary processing pipeline 100 can include ranking genes based on their gene expression levels and using the rankings and one or more statistical models to determine one or more characteristics. The processing pipeline 100 can be executed on any suitable computing device (e.g., a single computing device, multiple computing devices collocated at a single physical location or located at multiple physical locations remote from each other, one or more computing device portions of a cloud computing system, etc.) because aspects of the technology described herein are not limited in this regard. In some embodiments, the processing pipeline 100 can be executed by a desktop computer, a laptop computer, a mobile computing device. In some embodiments, the processing pipeline 100 can be executed within one or more computing devices that are part of a cloud computing environment.
[0094] As shown in FIG. 1, gene expression data 102 can be obtained for a subject biological sample. The subject can have, be suspected of having, or be at risk of having cancer (e.g., breast cancer, kidney cancer, clear cell renal carcinoma, lymphoma). A subject having, suspected of having, or at risk of having cancer can be a subject showing one or more signs or symptoms of cancer, a subject diagnosed with cancer, a subject having a family history and / or genetic predisposition to cancer, and / or a subject having one or more other risk factors for cancer (e.g., age, exposure to carcinogens, environmental exposure, exposure to viruses associated with a higher likelihood of developing cancer, etc.). The expression data 102 can be obtained using any suitable sequencing platform (e.g., gene expression microarray, next-generation sequencing, hybridization-based expression assay), resulting in expression data (e.g., microarray data, RNAseq data, hybridization-based expression assay data) for the biological sample. Some embodiments may involve performing a sequencing process (e.g., gene expression microarray, next-generation sequencing) on the biological sample prior to obtaining the expression data 102. In some embodiments, obtaining the gene expression data 102 may involve accessing, using a computing device, expression data (e.g., expression data previously obtained from a biological sample) in one or more data stores, receiving expression data from one or more other devices, or obtaining the gene expression data 102 in silico by any other method, etc. In some embodiments, obtaining the gene expression data 102 may involve (in vitro) analyzing the biological sample and accessing the expression data (e.g., by a computing device, by a processor). Further aspects regarding obtaining the expression data are provided in the section entitled "Obtaining Expression Data".
[0095] As shown in FIG. 1, the expression data 102 includes expression level values for N different genes of "Sample 1", "Gene 1", "Gene 2", "Gene 3",... "Gene N". Different sequencing platforms can be used to obtain the expression data 102. In some embodiments, the expression data 102 can be obtained using a gene expression microarray (e.g., by determining the amount of RNA that binds to different probes on the microarray). Gene expression microarrays can detect the expression of thousands of genes at once. The expression data 102 associated with using a gene expression microarray can be associated with 1,000, at least 10,000, or at least 100,000 gene detection events. In some embodiments, the expression data 102 can be obtained by performing next-generation sequencing. Such expression data can be associated with obtaining sequence reads using next-generation sequencing, aligning the sequencing reads to a reference (e.g., by using one or more sequence alignment algorithms), and determining expression level values for some genes based on the alignment. The expression data 102 associated with performing next-generation sequencing can be associated with at least 10,000, at least 100,000, at least 1,000,000, or at least 10,000,000 sequence reads. In some embodiments, the expression data 102 can be obtained by using a hybridization-based expression assay (e.g., a labeled probe for targeting regions of interest in a biological sequence). The expression data 102 associated with using hybridization-based expression can be associated with 1,000, at least 10,000, or at least 100,000 gene detection events.
[0096] In some embodiments, the expression data 102 includes RNA Seq data. In such embodiments, obtaining the expression data 102 can involve obtaining RNA expression levels obtained by performing RNA sequencing. In some embodiments, the expression data 102 is obtained by performing whole-genome sequencing (WGS). In some embodiments, the expression data 102 is obtained by performing whole-exome sequencing (WES). In some embodiments, the expression data 102 includes a combination of RNA Seq data and WGS data. In some embodiments, the expression data 102 includes a combination of RNA Seq data and WES data.
[0097] In some embodiments, the expression data 102 includes values for N different genes, where the values represent the expression levels for a particular gene. For example, a first expression data 102 includes a value of 10.455 that represents the expression level for gene 2 and a value of 0.001 that represents the expression level for gene N, which indicates that gene 2 has a higher expression level than gene N in sample 1. As described above, the sequencing platform used to obtain the expression data 102 can affect the specific values of the expression data and the relative values among genes.
[0098] According to some embodiments, the ranking process 108 may involve ranking genes based on their expression levels in the expression data 102 to obtain a gene ranking 110. The ranking process 108 may involve ranking genes in a set of genes based on the numerical values of their expression levels. In some embodiments, the ranking process 108 may involve ranking some or all of the genes in the expression data 102 to obtain a gene ranking 110. Different gene rankings may be obtained by ranking the expression levels for different gene sets. Determining a gene ranking may involve determining a relative rank for each gene in a set of genes. As shown in FIG. 1, the genes in the expression data 102 may be ranked based on their expression levels using the ranking process 108 for the gene set 1 106a to obtain a first gene ranking 110a. Similarly, the genes in the expression data 102 may be ranked based on their expression levels using the ranking process 108 for the gene set 2 106b to obtain a second gene ranking 110b. The gene ranking 110a and the gene ranking 110b have relative ranks for different genes. As shown in FIG. 1, the gene ranking 110a has relative ranks of 30, N-1, 2, and 1 for genes 1, 2, 3, and N respectively, and the gene ranking 110b has relative ranks of 15, 21, 2, and 1 for genes 1, 2, 3, and N respectively. A gene ranking may include values that specify the relative ranks of the genes in the gene ranking. In some embodiments, the values that specify the relative ranks may include ordinal numbers. In some embodiments, the values that specify the relative ranks may include integers such as those shown in FIG. 1. In some embodiments, the values that specify the relative ranks may be used as inputs (e.g., vectors of relative ranks) to statistical models for predicting characteristics using the techniques described herein. In some embodiments, a gene ranking is a gene by its relative rank It may include a sorted list of genes. In such embodiments, the sorted list of genes may be used as an input (e.g., a vector having the sorted list of genes) to a statistical model for predicting characteristics using the techniques described herein. For example, a gene set may include a gene list A = [x1, x2, x3,... xN-1, xN], and the ranking process 108 may output a sorted list of genes [x2, x15, xN-1... x1, xN] having their corresponding relative ranks as [1, 2, 3,... N-1, N]. The sorted list of genes [x2, x15, xN-1... x1, xN] and their relative ranks [1, 2, 3,... N-1, N] may be used as an input to the statistical model.
[0099] In some embodiments, the ranking process 108 may involve ordering the genes in the gene set from lowest to highest expression level and labeling the list of genes with ranks for the individual genes. For example, the lowest expression level values are ordered first on the list of genes and their corresponding labels are the lowest (e.g., 1, 2, 3, etc.), while the highest expression level values have the corresponding higher labels. In some embodiments, the ranking process 108 may involve ordering the genes in descending order such that the genes in the gene set are ranked from highest to lowest expression level values. In some embodiments, the ranking process 108 may involve one or more preprocessing steps prior to ranking the genes, including binning the gene expression values, rounding the gene expression values. For example, in some embodiments, the gene expression values may be sorted into bins and then ranked. As another example, in some embodiments, the gene expression values may be truncated and then ranked. Other preprocessing steps may be applied to the expression levels and the ranking may be performed on the preprocessed values, because aspects of the techniques described herein are not limited to ranking by simply sorting at the exact gene expression levels obtained.
[0100] In cases where gene groups have equal or substantially similar expression level values, the genes in the group may have a common rank and a label indicating the common rank. In some embodiments, the common rank may be determined as the average of the ranks for the genes in the group. For example, a gene in a gene set may have 30 expression level values, be ranked as 4, and the next genes in the ordered list have expression level values of 35, 35, and 35, respectively, and are ranked as 5, 6, and 7, and then these genes are all ranked as 6 (the average of 5, 6, and 7). In some embodiments, the gene ranking may include two or more genes having a common rank. In some embodiments, the gene ranking of a gene group having a common rank may include consecutive ranking labels (e.g., 1, 2, 2, 2, 3, 4, 5, etc.). In some embodiments, the gene ranking of a gene group having a common rank may include ranking labels that skip one or more values (e.g., 1, 2, 2, 2, 5, 6, 6, 8, etc.). In some embodiments, a gene group having equal or substantially similar expression level values may be ranked according to the minimum rank or the maximum rank in the gene group.
[0101] To determine specific characteristics of a biological sample (e.g., origin tissue, cancer grade, tissue type, such as tissue subtypes like PTCL subtypes, viral states such as HPV status), a selected set of genes can be used in ranking process 108 to obtain gene ranking 110. As shown in FIG. 1, gene set 1 106a is used to obtain gene ranking 110a, and then gene ranking 110a is used to determine characteristic 1 114a. Similarly, gene set 2 106b is used to obtain gene ranking 110b, and then gene ranking 110b is used to determine characteristic 2 114b. For example, a certain set of genes can be used to determine the origin tissue for a biological sample, and another set of genes can be used to determine the cancer grade.
[0102] The number of genes in the gene set can be in the range of 3 to 1,000 genes, 5 to 500 genes, 5 to 200 genes, 5 to 100 genes, 3 to 50 genes, 20 to 100 genes, 50 to 100 genes, 50 to 200 genes, 50 to 300 genes, 100 to 300 genes, and 50 to 500 genes. The gene set can include at least 3 genes, at least 5 genes, at least 10 genes, or at least 20 genes. The gene set can consist of 5 to 50 genes, 5 to 100 genes, 20 to 100 genes, 50 to 100 genes, 5 to 200 genes, 5 to 300 genes, 10 to 200 genes, 50 to 300 genes, 5 to 500 genes, or 50 to 500 genes.
[0103] Gene rankings and statistical models can be used to determine specific characteristics of a biological sample. In particular, a gene ranking can be used as an input to a statistical model, and an output indicative of the characteristic can be obtained. To obtain different characteristics, different gene sets and different statistical models are used, where determining a specific characteristic involves using a specific gene set and a statistical model trained using training data indicating rankings of expression levels for some or all of the genes in the gene set. For example, statistical model 112a is specific to determining characteristic 1114a and is trained using training data indicating rankings of expression levels for some or all of the genes in gene set 1106a. Similarly, statistical model 112b is specific to determining characteristic 2114b and is trained using training data indicating rankings of expression levels for some or all of the genes in gene set 2106b. For example, statistical model 112a and gene set 1106a can be used to determine the cancer grade for cells in a biological sample, and statistical model 112b and gene set 2106b can be used to determine the origin tissue for cells in a biological sample.
[0104] The training data may include rankings of expression levels associated with a plurality of samples, where the samples are associated with characteristics determined using a statistical model. For example, in an embodiment where the statistical model is used to predict cancer grade, the training data may include rankings of expression levels associated with samples of multiple cancer grades (e.g., grade 1, grade 2, grade 3). As another example, in an embodiment where the statistical model is used to predict the tissue of origin, the training data may include rankings of expression levels associated with samples from multiple tissues of origin (e.g., thyroid tissue, lymph node tissue, adrenal tissue, skin tissue, breast tissue, ovarian tissue, prostate tissue, urothelial tissue, cervical tissue, esophageal tissue, brain tissue, soft tissue, connective tissue, head tissue, and neck tissue). As another example, in an embodiment where the statistical model is used to predict the HPV status, the training data may include rankings of expression levels associated with samples from both the HPV positive and HPV negative states. As another example, in an embodiment where the statistical model is used to predict the PTCL subtype, the training data may include rankings of expression levels associated with samples from different PTCL subtypes (e.g., adult T cell leukemia / lymphoma (ATLL), angioimmunoblastic T cell lymphoma (AITL), NK / T cell lymphoma (NKTCL), anaplastic large cell lymphoma (ALCL), and cases belonging to the unspecified type (PTCL-NOS)).
[0105] It should be appreciated that statistical models, such as statistical model 112a and statistical model 112b, can be used to determine one or more characteristics for different biological samples obtained from different subjects. In some cases, the number of subjects for which the same statistical model can be used can be at least 50, 100, 200, 300, 500, 1,000, 2,000, 5,000, 10,000, or more. By using a statistical model for different subjects, the analysis of expression data across different subjects can be facilitated because the same data processing pipeline can be implemented for individual subjects.
[0106] In some embodiments, the ranking process 108 may rank only the genes included in the set of genes such that not all genes in the expression data may obtain a rank or be included during gene ranking. In such embodiments, the ranking is specific to the set of genes and may be used as input to the statistical model 112.
[0107] In some embodiments, the ranking process 108 may involve ranking all of the genes in the expression data 102 such that each gene has a respective rank. In such embodiments, the ranking includes genes outside of the set of genes. In some embodiments, the input to the statistical model may include the ranks determined by the ranking process 108 for the set of genes. In some embodiments, the input to the statistical model may include the ranking obtained by the ranking process 108, and the statistical model may selectively use the ranks for the set of genes in the ranking as part of determining one or more characteristics.
[0108] A statistical model may involve using one or more suitable machine learning algorithms that include one or more classifiers. Examples of classifiers that a statistical model may include are gradient boosted decision tree classifier, decision tree classifier, gradient boosting classifier, random forest classifier, clustering-based classifier, Bayesian classifier, Bayesian network classifier, neural network classifier, kernel-based classifier, and support vector machine classifier. In some embodiments, the statistical model may involve using a gradient boosted decision tree classifier. In some embodiments, the statistical model may involve using a decision tree classifier. In some embodiments, the statistical model may involve using a gradient boosting classifier. In some embodiments, the statistical model may involve using a random forest classifier. In some embodiments, the statistical model may involve using a clustering-based classifier. In some embodiments, the statistical model may involve using a Bayesian classifier. In some embodiments, the statistical model may involve using a Bayesian network classifier. In some embodiments, the statistical model may involve using a neural network classifier. In some embodiments, the statistical model may involve using a kernel-based classifier. In some embodiments, the statistical model may involve using a support vector machine classifier.
[0109] In some embodiments, the statistical model may perform a binary classification of one or more features as an output of the statistical model. For example, such a statistical model may perform a classification of one or more cancer grades (e.g., grade 1, grade 2, grade 3), and the output of the statistical model may include predictions for each of the one or more cancer grades indicating whether a biological sample is categorized as being a particular cancer grade.
[0110] In some embodiments, the statistical model may involve using a machine learning algorithm that implements a gradient boosting framework, such as a gradient boosting decision tree (GBDT) and a gradient boosted regression tree (GBRT). An example of a machine learning algorithm that implements a gradient boosting decision tree is the LightGBM package, which is further described in Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu, LightGBM: A highly efficient gradient boosting decision tree, Advances in Neural Information Processing Systems, pages 3149 - 3157, 2017 (https: / / dl.acm.org / doi / 10.5555 / 3294996.3295074), the entire content of which is incorporated herein by reference. An example of a machine learning algorithm that implements a gradient boosting framework is the XGBoost package, which is further described in Tianqi Chen and Carlos Guestrin, XGBoost: A scalable tree boosting system, In Proceedings of the 22Nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 785 - 794, ACM, 2016 (https: / / dl.acm.org / doi / 10.1145 / 2939672.2939785), the entire content of which is incorporated herein by reference.An example of a machine learning algorithm that implements gradient boosted regression trees is the pGBRT package, which is further described in Stephen Tyree, Kilian Q Weinberger, Kunal Agrawal, and Jennifer Paykin, Parallel boosted regression trees for web search ranking, In Proceedings of the 20th international conference on World wide web, pages 387 - 396, ACM, 2011 (https: / / dl.acm.org / doi / 10.1145 / 1963405.1963461), the entirety of which is incorporated herein by reference.
[0111] The statistical model can be trained using multiple rankings of expression levels for some or all of the genes in a set of genes. The training data may include available expression data obtained through research organizations including the National Cancer Institute (NCI) (e.g., Gene Expression Omnibus (GEO)), the National Center for Biotechnology Information (NCBI) (e.g., Sequence Read Archive (SRA)), The Cancer Genome Atlas Program (TCGA), the ArrayExpress Archive of Functional Genomics Data (by the European Molecular Biology Laboratory), and the International Cancer Genome Consortium.
[0112] For example, a statistical model used to determine cancer grade for breast cancer can be trained using data from the series GSE96058 available through the NCI. As another example, a statistical model used to determine cancer grade for kidney clear cell carcinoma can be trained using data from the Cancer Genome Atlas Kidney Renal Clear Cell Carcinoma (TCGA-KIRC) data collection. As yet another example, a statistical model used to determine the tissue of origin (e.g., ABC, GCB) for DLBCL can be trained using data from one or more of the series GSE117556, Leipzig Lymphoma data set (10.1186 / s13073-019-0637-7), series GSE31312, series GSE10846, series GSE87371, series GSE11318, series GSE32918, series GSE23501, Lymphoma / Leukemia Molecular Profiling Project (LLMPP), and series GSE93984. As another example, a statistical model used to determine the tissue of origin and histological information (e.g., tissue type) for cancer can be trained using data from the Cancer Genome Atlas Program (TCGAP).
[0113] One characteristic that can be determined using the techniques described herein is the cancer grade for cells in a biological sample. Cancer grades can include grade 1, grade 2, grade 3, grade 4, and grade 5. Some cancer grading systems can include any suitable number of grades, or other scores, and it should be understood that the techniques described herein can be used to determine any number of cancer grades regardless of the cancer grading system being implemented. For example, some cancer grading systems can have any number of cancer grades in the range of 1 to 10. Another characteristic is the originating tissue for cells in a biological sample. Originating tissues can include lung tissue, pancreatic tissue, stomach tissue, colon tissue, liver tissue, bladder tissue, kidney tissue, thyroid tissue, lymph node tissue, adrenal tissue, skin tissue, breast tissue, ovarian tissue, prostate tissue, urothelial tissue, cervical tissue, esophageal tissue, brain tissue, soft tissue, connective tissue, head tissue, and neck tissue. In some cases, the originating tissue can refer to the originating cells. For example, if a subject has, is suspected of having, or is at risk of having diffuse large B-cell lymphoma (DLBCL), the originating tissue can be originating cells that include germinal center B cells (GCB) and activated B cells (ABC).
[0114] Another characteristic is the histological information for cells in a biological sample. The histological information can correspond to determinations made by a physician (e.g., a pathologist) using microscopy to visually examine the biological sample. The histological information can include tissue type. Examples of tissue types include adenocarcinoma, squamous cell carcinoma, carcinoma, cystadenocarcinoma, sarcoma, and glioma. In some embodiments, the statistical model can output a combination of the originating tissue and the histological information. The combination of the originating tissue and the histological information can include lung adenocarcinoma, lung squamous cell carcinoma, melanoma, breast cancer, colorectal adenocarcinoma, ovarian serous cystadenocarcinoma, pheochromocytoma, bladder urothelial carcinoma, cervical squamous cell carcinoma, glioblastoma multiforme, head and neck squamous cell carcinoma, renal clear cell carcinoma of the kidney, renal papillary cell carcinoma of the kidney, hepatocellular carcinoma of the liver, pancreatic adenocarcinoma, paraganglioma, prostatic adenocarcinoma, sarcoma, gastric adenocarcinoma, thyroid cancer, and endometrial carcinoma of the corpus uteri.
[0115] Characteristics (e.g., cancer grade, tissue of origin, PTCL subtype) can be output to a user, such as a physician or clinician, by displaying the characteristics to the user in a graphical user interface (GUI), including the characteristics in a report, sending an email to the user, and / or in any other suitable way. The characteristics of the subject can be used for various clinical purposes, including assessing the effectiveness of treatment for cancer, identifying treatment for the subject, performing treatment for the subject, determining the prognosis for the subject, and / or evaluating the suitability of the subject for participation in a clinical trial. In some embodiments, the characteristics of the subject can be used in identifying treatment for the subject. For example, in an embodiment where the tissue of origin is determined for cells in a biological sample, the determined tissue of origin can be used to identify treatment for the subject associated with treating the cancer of the determined tissue of origin. As another example, in an embodiment where the cancer grade is determined for cells in a biological sample, the determined cancer grade can be used to identify treatment for the subject associated with treating the cancer having the determined cancer grade. As another example, in an embodiment where the PTCL subtype is determined for cells in a biological sample, the determined PTCL subtype can be used to identify treatment for the subject suitable for treating the lymphoma of the determined PTCL subtype. Next, the identified treatment can be performed.
[0116] In some embodiments, a characteristic of a subject can be used to effect a treatment for the subject. For example, in embodiments where the originating tissue is determined for cells in a biological sample, a physician can effect a treatment for the subject associated with treating cancer of the determined originating tissue. As another example, in embodiments where the cancer grade is determined for cells in a biological sample, a physician can effect a treatment for the subject associated with treating cancer having the determined cancer grade. As yet another example, in embodiments where the PTCL subtype is determined for cells in a biological sample, a physician can effect a treatment for the subject suitable for treating lymphoma of the determined PTCL subtype. Further examples of characteristics of a biological sample determined using the techniques described herein being used to effect a treatment are provided in the section entitled "Methods of Treatment".
[0117] In some embodiments, a characteristic of a subject can be used when determining a prognosis for the subject. In embodiments where the subject has, is suspected of having, or is at risk of having cancer (e.g., kidney cancer, clear cell kidney cancer, lymphoma, head and neck squamous cell carcinoma, lung adenocarcinoma), the determined characteristic of the subject can be used to determine a prognosis for the subject. For example, in embodiments where the characteristic of the subject is cancer grade, the determined cancer grade (e.g., grade 1, grade 2, grade 3) can be used to determine a prognosis for the subject. Further aspects regarding other applications where characteristics of a biological sample determined using the techniques described herein are used when determining a prognosis are provided in the section entitled "Applications".
[0118] In some embodiments, the determined characteristics of the biological sample can include a cancer grade for cells in the biological sample. In such embodiments, the set of genes used to obtain the gene ranking can include genes associated with biological features, expression pathways, or otherwise associated with determining the cancer grade. Some embodiments involve using a gene set for determining a cancer grade for breast cancer. Examples of genes that can be included in such a gene set are listed in Table 1 below.
[0119]
Table 1A
[0120]
Table 1B
[0121]
Table 1C
[0122]
Table 1D
[0123]
Table 1E
[0124] Some embodiments involve using a gene set for determining a cancer grade for clear cell renal carcinoma. Examples of genes that can be included in such a gene set are listed in Table 2 below.
[0125]
Table 2A
[0126]
Table 2B
[0127]
Table 2C
[0128] In some embodiments, the determined characteristics of a biological sample may include the origin tissue for cells in the biological sample. In such embodiments, the set of genes used to obtain a gene ranking may include genes associated with biological characteristics, expression pathways, or otherwise associated with determining the origin tissue. Some embodiments involve using a gene set to predict the origin tissue for diffuse large B-cell lymphoma (DLBCL), such as germinal center B cells (GCB) and activated B cells (ABC). Examples of genes that may be included in such a gene set are listed in Table 3 below.
[0129]
Table 3A
[0130]
Table 3B
[0131] Some embodiments may involve determining the characteristics of a biological sample by obtaining a characteristic prediction using different gene sets and statistical models corresponding to the different gene sets, where the characteristic prediction is used to determine the characteristics. FIG. 2 is a diagram of an exemplary processing pipeline 200 for determining the characteristics of a biological sample according to some embodiments of the techniques described herein, and the exemplary processing pipeline 200 may include ranking genes based on their gene expression levels, and using the rankings and statistical models to determine the characteristics. The processing pipeline 200 may be executed on any suitable computing device (e.g., a single computing device, multiple computing devices collocated at a single physical location or located at multiple physical locations remote from each other, one or more computing device portions of a cloud computing system, etc.), as the aspects of the techniques described herein are not limited in this regard. In some embodiments, the processing pipeline 200 may be executed by a desktop computer, a laptop computer, a mobile computing device. In some embodiments, the processing pipeline 200 may be executed within one or more computing devices that are part of a cloud computing environment.
[0132] In some embodiments, gene expression data 102 is used to rank genes in different sets of genes based on their expression levels in the gene expression data 102 to obtain a plurality of gene rankings. For example, a gene ranking can be obtained for each gene set, and the gene ranking can be input into a statistical model trained using training data indicating rankings of expression levels for some or all of the genes in the gene set. As shown in FIG. 2, ranking process 108 uses expression data 102 to rank genes in different gene sets including gene set 1 106a, gene set 2 106b, gene set 3 106c, and gene set 4 106d, to obtain gene ranking 1 110a, gene ranking 2 110b, gene ranking 3 110c, and gene ranking 4 110d, respectively. The ranking process 108 can involve ranking the genes in a set of genes based on the numerical values of their expression levels. Different gene rankings can be obtained by ranking the expression levels for different gene sets, and each gene ranking can be input into its respective statistical model to obtain a characteristic prediction. As shown in FIG. 2, gene ranking 1 110a, gene ranking 2 110b, gene ranking 3 110c, and gene ranking 4 110d are provided as inputs to statistical model 1 112a, statistical model 2 112b, statistical model 3 112c, and statistical model 4 112d, respectively.
[0133] In some embodiments, different statistical models and their respective gene sets can correspond to specific characteristics of a biological sample. In such embodiments, each of the statistical models can output a prediction of a biological sample having a specific characteristic. In some cases, the prediction output by the statistical model can include the probability that the biological sample has the characteristic.
[0134] As shown in FIG. 2, the statistical model 1112a outputs the characteristic prediction 1116a, the statistical model 2112b outputs the characteristic prediction 2116b, the statistical model 3112c outputs the characteristic prediction 3116c, and the statistical model 4112d outputs the characteristic prediction 4116d. Predictions output by different statistical models can be analyzed using the prediction analysis process 118 to determine the characteristics 114 for the biological sample. The prediction analysis process 118 can involve aggregating different predictions and selecting specific characteristics for the biological sample from among the different characteristic predictions. In some embodiments, the characteristic prediction can include the probability that the biological sample has a particular characteristic. In such embodiments, the prediction analysis process 118 can involve aggregating the probabilities for different characteristic predictions and selecting a characteristic based on the probabilities. In some embodiments, selecting a characteristic can involve selecting the characteristic with the highest probability as the characteristic 114.
[0135] Although four gene sets and four statistical models are shown in FIG. 2, it should be understood that any suitable number of gene sets and corresponding statistical models can be implemented using the techniques described above in determining characteristic predictions and aggregating the characteristic predictions to obtain characteristics of a biological sample. In some embodiments, the number of gene sets and corresponding statistical models can be in the range of 3 to 100, 3 to 70, 3 to 50, 3 to 40, 3 to 30, 5 to 50, 10 to 60, or 10 to 70.
[0136] In some embodiments, the number of gene sets and corresponding statistical models is less than or equal to the number of classes for the characteristic being predicted using processing pipeline 200. For example, in embodiments where the characteristic being predicted is the originating tissue, the number of classes may correspond to the different types of tissues that can be determined using processing pipeline 200. Such embodiments may involve different gene sets and corresponding statistical models for each type of tissue. For example, gene set 1 106a and statistical model 1 112a can be used to generate a prediction (as characteristic prediction 1 116a) that the biological sample is lung tissue, gene set 2 106b and statistical model 2 112b can be used to generate a prediction (as characteristic prediction 2 116b) that the biological sample is stomach tissue, gene set 3 106c and statistical model 3 112c can be used to generate a prediction (as characteristic prediction 3 116c) that the biological sample is liver tissue, and gene set 4 106d and statistical model 4 112d can be used to generate a prediction (as characteristic prediction 4 116d) that the biological sample is bladder tissue. It should be appreciated that additional gene sets and their corresponding statistical models can be implemented for different tissue types. In some embodiments, there can be 21 gene sets and corresponding statistical models, enabling processing pipeline 200 to predict 21 types of tissues.
[0137] Figure 3 is a flowchart of an exemplary process 300 for determining one or more characteristics of a biological sample using gene ranking and a statistical model, according to some embodiments of the techniques described herein. Process 300 can be executed on any suitable computing device (e.g., a single computing device, multiple computing devices collocated at a single physical location, or multiple computing devices located at multiple physical locations remote from one another, one or more computing device portions of a cloud computing system, etc.), as the aspects of the techniques described herein are not limited in this regard. In some embodiments, the ranking process 108 and the statistical model 112 can execute some or all of process 300 to determine one or more characteristics, such as characteristic 114.
[0138] Process 300 begins at operation 310, where expression data for a target biological sample is obtained. In some embodiments, the expression data may be obtained using a gene expression microarray. In some embodiments, the expression data may be obtained by performing next-generation sequencing. Some embodiments involve performing a sequencing process of the biological sample (e.g., gene expression microarray, next-generation sequencing) prior to obtaining the expression data 102. In some embodiments, obtaining the gene expression data 102 may involve accessing, using a computing device, expression data in one or more data stores (e.g., expression data previously obtained from a biological sample), receiving expression data from one or more other devices, or obtaining the gene expression data 102 in silico by any other method. In some embodiments, obtaining the gene expression data 102 may involve (in vitro) analyzing the biological sample and accessing the expression data (e.g., by a computing device, a processor). Further aspects regarding obtaining expression data are provided in the section entitled "Obtaining Expression Data".
[0139] Next, process 300 proceeds to operation 320, where genes in a set of genes are ranked based on their expression levels in the expression data, such as by using the ranking process 108, to obtain a gene ranking. The expression data may include values each representing an expression level for a gene in a set of genes, and determining the gene ranking may involve determining a relative rank for each gene in the set of genes based on the values.
[0140] In some embodiments, the subject has, is suspected of having, or is at risk of having breast cancer. The set of genes can be selected from the group of genes listed in Table 1. The set of genes can include at least 3, 5, 10, or 20 genes selected from the group of genes listed in Table 1. In some embodiments, the set of genes can include all of the genes listed in Table 1. In some embodiments, the set of genes can include 3 - 100 genes, 5 - 100 genes, 20 - 100 genes, 50 - 100 genes, 80 - 100 genes listed in Table 1. In some embodiments, the set of genes can include 100 or fewer genes, 80 or fewer genes, 50 or fewer genes, 20 or fewer genes listed in Table 1.
[0141] In some embodiments, the subject has, is suspected of having, or is at risk of having clear cell renal cancer. The set of genes can be selected from the group of genes listed in Table 2. The set of genes can include at least 3, 5, 10, or 20 genes selected from the group of genes listed in Table 2. In some embodiments, the set of genes can include all of the genes listed in Table 2. In some embodiments, the set of genes can include 3 - 80 genes, 5 - 80 genes, 20 - 80 genes, 50 - 80 genes, 70 - 80 genes listed in Table 2. In some embodiments, the set of genes can include 80 or fewer genes, 50 or fewer genes, 20 or fewer genes listed in Table 2.
[0142] In some embodiments, the subject has, is suspected of having, or is at risk of having lymphoma. The set of genes can be selected from the group of genes listed in Table 3. The set of genes can include at least 3, 5, 10, or 20 genes selected from the group of genes listed in Table 3. In some embodiments, the set of genes can include all of the genes listed in Table 3. In some embodiments, the set of genes can include 3-25 genes, 5-25 genes, 10-25 genes, 20-25 genes listed in Table 3. In some embodiments, the set of genes can include 25 or fewer genes, 20 or fewer genes, 15 or fewer genes, 10 or fewer genes listed in Table 3.
[0143] Next, process 300 proceeds to operation 330, where one or more characteristics of the biological sample are determined using a statistical model such as gene ranking and statistical model 112. In some embodiments, the characteristics determined by process 300 can include a cancer grade for cells in the biological sample. In some embodiments, the characteristics determined by process 300 can include an origin tissue for cells in the biological sample. The statistical model can be trained using a ranking of expression levels for one or more genes in the set of genes. In some embodiments, the gene ranking can be used as an input to the statistical model to obtain an output indicative of one or more characteristics. In some embodiments, the statistical model comprises a classifier selected from the group consisting of a gradient boosted decision tree classifier, a decision tree classifier, a gradient boost classifier, a random forest classifier, a clustering-based classifier, a Bayesian classifier, a Bayesian network classifier, a neural network classifier, a kernel-based classifier, and a support vector machine classifier.
[0144] In some embodiments, process 300 may include ranking genes in a second set of genes based on their expression levels in the expression data to obtain a second gene ranking. The second gene ranking and the second statistical model may be used to determine one or more second characteristics of the biological sample. The second statistical model may be trained using second training data that indicates rankings of expression levels for some or all of the genes in the second set of genes. The one or more second characteristics of the biological sample may be different from the characteristics determined by operation 330. For example, in some embodiments, the characteristics determined by operation 330 may include a cancer grade for cells in the biological sample, and the second characteristics may include an origin tissue for cells in the biological sample.
[0145] In some embodiments, process 300 may include outputting one or more characteristics to a user (e.g., a physician) in a graphical user interface (GUI), including the one or more characteristics in a report, sending an email to the user, and in any other suitable way.
[0146] In some embodiments, process 300 may include treating a subject based on one or more determined characteristics of the biological sample. For example, in embodiments where an origin tissue is determined for cells in the biological sample, a physician may perform a treatment for the subject associated with treating the cancer of the determined origin tissue. As another example, in embodiments where a cancer grade is determined for cells in the biological sample, a physician may perform a treatment for the subject associated with treating the cancer having the determined cancer grade. Further examples of characteristics of a biological sample determined using the techniques described herein that are used for performing a treatment are provided in the section entitled "Methods of Treatment."
[0147] In some embodiments, process 300 can include identifying a treatment for a subject based on determined characteristics of a biological sample. For example, in embodiments where the origin tissue is determined for cells in the biological sample, the determined origin tissue can be used to identify a treatment for the subject associated with treating cancer of the determined origin tissue. As another example, in embodiments where the cancer grade is determined for cells in the biological sample, the determined cancer grade can be used to identify a treatment for the subject associated with treating cancer having the determined cancer grade.
[0148] In some embodiments, process 300 can include determining a prognosis for a subject based on one or more determined characteristics of a biological sample. For example, in embodiments where the origin tissue is determined for cells in the biological sample, the determined origin tissue can be used to determine a prognosis for the subject associated with treating cancer of the determined origin tissue. As another example, in embodiments where the cancer grade is determined for cells in the biological sample, the determined cancer grade can be used to determine a prognosis for the subject associated with treating cancer having the determined cancer grade. Further aspects regarding additional applications in which characteristics of a biological sample determined using the techniques described herein are used in determining a prognosis are provided in the section entitled "Applications."
[0149] FIG. 4 is a flowchart of an exemplary process 400 for determining an origin tissue for cells in a biological sample, according to some embodiments of the technology described herein. Process 400 can be implemented on any suitable computing device (e.g., a single computing device, multiple computing devices collocated at a single physical location or located at multiple physical locations remote from each other, one or more computing device portions of a cloud computing system, etc.), since aspects of the technology described herein are not limited in this regard. In some embodiments, ranking process 108 and statistical model 112 can perform some or all of process 400 to determine the origin tissue.
[0150] Process 400 begins at operation 410, where expression data for cells in a biological sample of a subject having, suspected of having, or at risk of having cancer is obtained. In some embodiments, the expression data is obtained using a gene expression microarray. In some embodiments, the expression data is obtained by performing next-generation sequencing. Some embodiments involve performing a sequencing process (e.g., gene expression microarray, next-generation sequencing) of the biological sample prior to obtaining expression data 102. In some embodiments, obtaining gene expression data 102 may involve accessing, using a computing device, expression data (e.g., expression data previously obtained from a biological sample) in one or more data stores, receiving expression data from one or more other devices, or obtaining gene expression data 102 in silico by any other means. In some embodiments, obtaining gene expression data 102 may involve (in vitro) analyzing the biological sample and accessing the expression data (e.g., by a computing device, a processor). Further aspects regarding obtaining expression data are provided in the section entitled "Obtaining Expression Data".
[0151] Next, process 400 proceeds to operation 420, where one or more sets of genes are ranked based on their expression levels in the expression data, such as by using ranking process 108, to obtain one or more gene rankings. The expression data may include values each representing an expression level for a gene in one or more sets of genes, and determining a gene ranking may involve determining a relative rank for each gene in the set of genes based on the values.
[0152] In some embodiments, the subject has, is suspected of having, or is at risk of having breast cancer. The set of genes can be selected from the group of genes listed in Table 1. The set of genes can include at least 3, 5, 10, or 20 genes selected from the group of genes listed in Table 1. The set of genes can consist of 5 - 100 genes, 10 - 200 genes, 20 - 100 genes, or 50 - 100 genes. In some embodiments, the set of genes can include all of the genes listed in Table 1. In some embodiments, the set of genes can include 3 - 100 genes, 5 - 100 genes, 20 - 100 genes, 50 - 100 genes, 80 - 100 genes listed in Table 1. In some embodiments, the set of genes can include 100 or fewer genes, 80 or fewer genes, 50 or fewer genes, 20 or fewer genes listed in Table 1.
[0153] Next, process 400 proceeds to operation 430, where the origin tissue for some or all of the cells in the biological sample is determined using one or more gene rankings and one or more statistical models such as statistical model 112. The statistical model can be trained using the ranking of the expression levels for some or all of the genes in the set of genes. Each of the gene rankings can be obtained based on the respective expression levels for one or more genes in the set of genes. In some embodiments, one or more gene rankings can be used as an input to one or more statistical models to obtain an output indicating the origin tissue. The origin tissue can include lung tissue, pancreatic tissue, stomach tissue, colon tissue, liver tissue, bladder tissue, kidney tissue, thyroid tissue, lymph node tissue, adrenal tissue, skin tissue, breast tissue, ovarian tissue, prostate tissue, urothelial tissue, cervical tissue, esophageal tissue, brain tissue, soft tissue, connective tissue, head tissue, and neck tissue.
[0154] In some embodiments, one or more statistical models comprise one or more classifiers selected from the group consisting of gradient boosted decision tree classifiers, decision tree classifiers, gradient boost classifiers, random forest classifiers, clustering-based classifiers, Bayesian classifiers, Bayesian network classifiers, neural network classifiers, kernel-based classifiers, and support vector machine classifiers.
[0155] In some embodiments, process 400 may further include determining histological information (e.g., tissue type) for at least some of the cells in a biological sample using gene ranking and one or more statistical models. The histological information may include adenocarcinoma, squamous cell carcinoma, carcinoma, cystadenocarcinoma, sarcoma, and glioma. The combination of the originating tissue and the histological information may be selected from the group consisting of lung adenocarcinoma, lung squamous cell carcinoma, melanoma, breast cancer, colorectal adenocarcinoma, ovarian serous cystadenocarcinoma, pheochromocytoma, urothelial carcinoma of the bladder, cervical squamous cell carcinoma, glioblastoma multiforme, head squamous cell carcinoma, cervical squamous cell carcinoma, clear cell renal cell carcinoma of the kidney, papillary renal cell carcinoma of the kidney, hepatocellular carcinoma of the liver, pancreatic adenocarcinoma, paraganglioma, prostatic adenocarcinoma, sarcoma, gastric adenocarcinoma, thyroid carcinoma, and endometrial carcinoma of the corpus uteri.
[0156] FIG. 5 is a flowchart of an exemplary process 500 for determining a cancer grade for cells in a biological sample according to some embodiments of the techniques described herein. Process 500 may be executed on any suitable computing device (e.g., a single computing device, multiple computing devices collocated at a single physical location or located at multiple physical locations remote from each other, one or more computing device portions of a cloud computing system, etc.) because aspects of the techniques described herein are not limited in this regard. In some embodiments, ranking process 108 and statistical model 112 may execute some or all of process 500 to determine the cancer grade.
[0157] Process 500 begins at operation 510, where expression data for cells in a biological sample of a subject having cancer, suspected of having cancer, or at risk of having cancer is obtained. In some embodiments, the expression data is obtained using a gene expression microarray. In some embodiments, the expression data is obtained by performing next-generation sequencing. Some embodiments involve performing a sequencing process (e.g., a gene expression microarray, next-generation sequencing) on the biological sample prior to obtaining expression data 102. In some embodiments, obtaining gene expression data 102 may involve accessing expression data (e.g., expression data previously obtained from a biological sample) in one or more data stores using a computing device, receiving expression data from one or more other devices, or obtaining gene expression data 102 in silico by any other means. In some embodiments, obtaining gene expression data 102 may involve (in vitro) analyzing the biological sample and accessing the expression data (e.g., by a computing device, a processor). Further aspects regarding obtaining expression data are provided in the section entitled "Obtaining Expression Data".
[0158] Next, process 500 proceeds to operation 520, where genes in a set of genes are ranked based on their expression levels in the expression data, such as by using ranking process 108, to obtain a gene ranking. The expression data may include values each representing an expression level for a gene in the set of genes, and determining the gene ranking may involve determining a relative rank for each gene in the set of genes based on the values. The set of genes may consist of 5 to 500 genes, 5 to 200 genes, 50 to 500 genes, or 50 to 300 genes.
[0159] In some embodiments, the subject has, is suspected of having, or is at risk of having breast cancer. The set of genes can be selected from the group of genes listed in Table 1. The set of genes can include at least 3, 5, 10, or 20 genes selected from the group of genes listed in Table 1. In some embodiments, the set of genes can include all of the genes listed in Table 1. In some embodiments, the set of genes can include 3 to 100 genes, 5 to 100 genes, 20 to 100 genes, 50 to 100 genes, 80 to 100 genes listed in Table 1. In some embodiments, the set of genes can include 100 or fewer genes, 80 or fewer genes, 50 or fewer genes, 20 or fewer genes listed in Table 1.
[0160] In some embodiments, the subject has, is suspected of having, or is at risk of having clear cell renal carcinoma. The set of genes can be selected from the group of genes listed in Table 2. The set of genes can include at least 3, 5, 10, or 20 genes selected from the group of genes listed in Table 2. In some embodiments, the set of genes can include 3 to 80 genes, 5 to 80 genes, 20 to 80 genes, 50 to 80 genes, 70 to 80 genes listed in Table 2. In some embodiments, the set of genes can include 80 or fewer genes, 50 or fewer genes, 20 or fewer genes listed in Table 2.
[0161] Next, process 500 proceeds to operation 530, where a cancer grade for cells in a biological sample is determined using a gene ranking and a statistical model such as statistical model 112. The statistical model can be trained using a gene ranking of expression levels for one or more genes in a set of genes. Each of the gene rankings can be obtained based on respective expression levels for one or more genes in the set of genes. In some embodiments, the gene ranking can be used as an input to the statistical model to obtain an output indicative of the cancer grade. The cancer grade can include grade 1, grade 2, grade 3, grade 4, and grade 5. In some embodiments, the statistical model comprises a classifier selected from the group consisting of a gradient boosted decision tree classifier, a decision tree classifier, a gradient boost classifier, a random forest classifier, a clustering-based classifier, a Bayesian classifier, a Bayesian network classifier, a neural network classifier, a kernel-based classifier, and a support vector machine classifier.
[0162] An example of how the techniques described herein can be implemented when predicting breast cancer grade is described with respect to FIGS. 6A, 6B, 6C, 6D, and 7. FIG. 6A shows different data sets (data sets that vary in sample preparation, sequencing platform, and data processing used to obtain expression data), the associated clinical cancer grade for samples of the data sets, and the predicted cancer grade obtained using the machine learning techniques described herein, for determining breast cancer grade. In particular, FIG. 6A shows different data sets (upper panel), where each vertical line corresponds to a different sample, and where the shading of the line corresponds to a different data set. FIG. 6A also shows the clinical grade associated with the samples of the data sets, where lighter shading indicates grade 1 ("G1") and darker shading indicates grade 3 ("G3"). The clinical grade can be a determination by a physician (e.g., a pathologist) using microscopy to visually inspect the sample. FIG. 6B shows enrichment signatures for different pathways showing gene expression profiles associated with breast cancer grade 1 and grade 3. Genes in one or more of these pathways can be used to determine breast cancer grade by the techniques described herein. As an example, the HALLMARK_G2M_CHECKPOINT signature is shown in the upper panel and has mostly upregulated genes in the samples in the right portion and mostly downregulated genes in the samples in the left portion. Other examples of pathways associated with cancer grade classification for breast cancer are in Table 4 below. In particular, different pathways enriched in the set of genes upregulated for grade 3 ("G3") and pathways enriched in the set of genes upregulated for grade 1 ("G1") are described in Table 4.
[0163]
Table 4A
[0164]
Table 4B
[0165] Figures 12, 13, 14, 15, 16, and 17 show the relationship between biological features and different grades of breast cancer. In particular, these figures illustrate the biology of molecular grades (grade 1 and grade 3) for breast cancer, where the data shown are for TCGA BRCA, and the predicted breast cancer grades were obtained using the techniques described herein. Figure 12 is the distribution of molecular cancer grades among the PAM50 subtypes. Figure 12 shows that most of the molecular grade 1 samples belong to the luminal subtype. Further comparisons in the breast cancer dataset for Figures 13 - 17 are for the luminal subtype only. Figure 13 shows how the progeny process score corresponds to a given cancer grade and predicted cancer grade in TCGA BRCA. The progeny process score is calculated from expression data. Figure 14 shows a plot comparing different protein expressions for different predicted cancer grades. The protein expression is from RPPA data. Figure 15 is a plot of the cytotoxicity score (CYT) for different predicted cancer grades. Figure 16 is a plot showing the differences in mutations between different cancer grades. Figure 16 shows genes, by WES data, that are significantly differentially mutated between predicted cancer grades. Figure 17 shows segments that are differentially amplified or deleted between predicted cancer grades. The segments shown in Figure 17 are by WES data.
[0166] For purposes of comparison with the computational techniques described herein, FIG. 6A shows prediction grades (lower panel) using expression data and statistical models, according to the techniques described herein. The prediction grades show how different samples are predicted to be, e.g., grade 1 (“G1”) for samples in the left portion and grade 3 (“G3”) for samples in the right portion. This is further shown in the plot of “G3 probability” across different samples below the lower panel of FIG. 6A, where the probability of grade 3 is higher for samples in the right portion than for samples in the left portion. FIGS. 6C and 6D show data similar to that shown in FIGS. 6A and 6B respectively, except that sample and pathway signatures are associated with predicting breast cancer to be grade 1 or grade 3 for grade 2 samples. Here, FIGS. 6C and 6D show how similar the biological features associated with grade 2 are to those associated with grades 1 and 3.
[0167] FIG. 7 is a plot of true positive rate versus false positive rate for several biological samples (shown as solid lines). This plot shows that the predicted cancer grades using the techniques described herein have a high true positive rate while maintaining a low false positive rate.
[0168] As another example, pathways associated with cancer grade classification for renal clear cells are in Table 5 below. In particular, different pathways enriched in sets of genes upregulated for grade 4 (“G4”) and for grade 1 (“G1”) are described in Table 5.
[0169]
Table 5A
[0170]
Table 5B
[0171]
Table 5C
[0172] Figures 18, 19, 20, 21, and 22 show the relationship between biological features and different grades of renal clear cell. In particular, these figures illustrate the biology of the molecular grades (grades 1 and 4) for renal clear cell carcinoma, where the data shown are for TCGA KIRC and the predicted renal clear cell grades were obtained using the techniques described herein. Figure 18 shows how the progeny process score corresponds to a given cancer grade and predicted cancer grade in TCGA KIRC. The progeny process score is calculated from expression data. Figure 19 is a plot showing chromosomal instability (CIN) for different cancer grades. Figure 20 is a plot comparing different protein expressions by RPPA data for different predicted cancer grades. Figures 21 and 22 show genes by WES data that are differentially amplified or deleted between predicted cancer grades.
[0173] Some embodiments involve using the techniques described herein to determine a cancer grade for lung adenocarcinoma. Examples of genes that may be included in a gene set for determining a cancer grade for lung adenocarcinoma are listed in Table 6 below. The gene set may include at least 3, 5, 10, or 20 genes selected from the group of genes listed in Table 6. In some embodiments, the gene set may include all of the genes listed in Table 6. In some embodiments, the gene set may include 3 - 25 genes, 5 - 25 genes, 10 - 25 genes, 20 - 25 genes listed in Table 6. In some embodiments, the gene set may include 25 or fewer genes, 20 or fewer genes, 15 or fewer genes, 10 or fewer genes listed in Table 6.
[0174]
Table 6
[0175] The techniques described in this specification can be implemented when predicting the cancer grade for lung adenocarcinoma, as described with respect to FIGS. 23A, 23B, and 23C. In particular, a cancer grade classifier for lung adenocarcinoma can distinguish between molecular grade 1 (mG1), low grade, and molecular grade 3 (mG3), high grade. Such a classifier can be developed by using samples from TCGA LUAD (from the National Cancer Institute) and CPTAC3 lung adenocarcinoma expression data (from NCBI) as training data. In the classifier described with respect to FIGS. 23A, 23B, and 23C, 117 samples of TCGA LUAD were excluded from the training data set and included as validation data. An initial gene set was formed from genes differentially expressed between grade 1 and grade 3. A genomic grade index (DOI: 10.1093 / jnci / djj052) based on the initial gene set was calculated, and the training data set samples were split into high and low cancer grades based on the survival mode. The number of genes was reduced through the selection of the gene set used for the classifier. For example, in the classifier described with respect to FIGS. 23A, 23B, and 23C, the initial gene set included 321 genes, and the gene set used in the classifier included 31 genes. The validation data set included 117 samples from TCGA LUAD and series GSE68465. After hyperparameter tuning, the performance of the classifier in the validation data set reached a 0.89 AUC score when distinguishing between grade 1 and grade 3. These results demonstrated the ability of the lung molecular grade to be statistically significant when predicting survival.
[0176] Figure 23A shows the enrichment signatures for different pathways, a validation dataset for determining the cancer grade of lung adenocarcinoma, the associated cancer grades reported for samples of the dataset, the predicted cancer grades obtained using the machine learning techniques described herein, and the gene expression profiles associated with grades 1 and 3. The validation dataset shown in Figure 23A varies in sample preparation, sequencing platform, and data processing used to obtain the expression data. Figure 23A shows a dataset (upper panel) where each vertical line corresponds to a different sample, where the shading of the line corresponds to a different dataset. The cancer grades associated with the samples of the dataset are shown, where lighter shading indicates grade 1 and darker shading indicates grade 3. The cancer grade associated with a sample can be a determination by a physician (e.g., a pathologist) using microscopy for visual inspection of the sample. Also shown is the probability of molecular grade 3 predicted using a cancer grade classifier. Figure 23A also shows the enrichment signatures for different pathways that show the gene expression profiles associated with lung adenocarcinoma grades 1 and 3. Genes in one or more of these pathways can be used to determine the lung adenocarcinoma grade by the techniques described herein. As an example, the HALLMARK_G2M_CHECKPOINT signature is shown in the upper panel and has mostly upregulated genes in the samples on the right portion and mostly downregulated genes in the samples on the left portion. Figure 23B shows the results of applying the validation dataset to a lung adenocarcinoma cancer grade classifier. Figure 23C is a plot of true positive rate versus false positive rate for predicting the cancer grades of different biological samples, where the classifier had an AUC score of 0.894.
[0177] FIG. 8A is a flowchart of an exemplary process 800 for selecting a gene set according to some embodiments of the techniques described herein. Process 800 may be executed on any suitable computing device (e.g., a single computing device, multiple computing devices collocated at a single physical location, or multiple computing devices located at multiple physical locations remote from each other, one or more computing device portions of a cloud computing system, etc.), because aspects of the techniques described herein are not limited in this regard. In some embodiments, ranking process 108 and statistical model 112 may implement all or part of process 800 to select a gene set that may be implemented in determining one or more characteristics of a biological sample, such as origin tissue, cancer grade, and PTCL subtype.
[0178] Process 800 begins at operation 810, where expression data is ranked to obtain a gene ranking for genes represented by expression levels in the expression data. Ranking process 108 may be used in ranking the expression data to obtain the gene ranking.
[0179] The expression data used in selecting gene sets may include available expression data obtained through research organizations, including the National Cancer Institute (NCI) (e.g., Gene Expression Omnibus (GEO)), the National Center for Biotechnology Information (NCBI), and The Cancer Imaging Archive (TCIA). For example, a gene set used to predict breast cancer grade may be obtained by using expression data from series GSE2990 available through NCI. As another example, a gene set used to determine cancer grade for clear cell renal carcinoma may be obtained by using expression data from series GSE40435. As another example, a gene set used to determine the origin tissue and histological information (e.g., tissue type) for cancer may be obtained by using expression data from The Cancer Genome Atlas Program (TCGAP). As another example, a gene set used to predict PTCL subtypes may be obtained using the expression data described in Table 9.
[0180] Next, process 800 proceeds to operation 820 where the ranked expression data is input into a statistical model such as statistical model 112. An output indicating one or more desired characteristics may be obtained as a result of inputting the ranked expression data into the statistical model. Process 800 may proceed to operation 830 where a validation quality score is calculated based on the output obtained by inputting the ranked expression data of operation 820 into the statistical model. The validation quality score may be calculated using one or more suitable metrics including negative log loss, AUC, F-score (micro, macro, weighted), accuracy, balanced accuracy, precision, and recall.
[0181] Next, process 800 proceeds to operation 840 where importance values for the different genes included during ranking are calculated. An example of an importance value is the Shapley Additive Explanations (SHAP) value, which is described in "A Unified Approach to Interpreting Model Predictions" by Scott M. Lundberg and Su-In Lee (https: / / arxiv.org / pdf / 1705.07874.pdf), the entire of which is incorporated by reference. Exemplary SHAP values are shown in Table 7 with respect to the germinal center classifier for DLBCL.
[0182] Next, process 800 proceeds to operation 850 where N (e.g., 1, 2, 3, 4) least important genes are excluded based on the importance values. Next, process 800 proceeds to operation 860 where the gene set is updated based on excluding the N least important genes. In some embodiments, at least the genes having the lowest importance values are removed from the gene set.
[0183] Process 800 may be initialized with a larger number of genes (e.g., about 3,000 genes) in the gene set and through subsequent iterations, reduce the number of genes in the set. Process 800 may continue by repeating these operations using the gene set selected in operation 860 of the previous iteration until a desired quality score (e.g., a quality score higher than a threshold) is achieved. In some cases, the initial gene set may be ranked in operation 810 and narrowed by process 800 to achieve a restricted gene set used for the classifier described herein.
[0184] Figure 8B is a flowchart of an exemplary process 900 for selecting a gene set, according to some embodiments of the techniques described herein. Process 900 may be executed on any suitable computing device (e.g., a single computing device, multiple computing devices collocated at a single physical location or located at multiple physical locations remote from each other, one or more computing device portions of a cloud computing system, etc.), as the aspects of the techniques described herein are not limited in this regard. In some embodiments, ranking process 108 and statistical model 112 may execute some or all of process 900 to select a gene set that may be implemented in determining one or more characteristics of a biological sample, such as origin tissue, cancer grade, and PTCL subtype.
[0185] Process 900 begins at operation 910, where an initial gene set is selected. The initial gene set may include a set of genes selected from Table 1, Table 2, Table 3, Table 6, and Table 8. The number of initial genes may be at least 1,000 genes, at least 3,000 genes, or at least 5,000 genes.
[0186] Next, process 900 proceeds to operation 810 as described above with respect to process 800. Next, process 900 proceeds to operation 920, where hyperparameters for the statistical model are selected and the statistical model is fit.
[0187] Next, the process proceeds to operations 840, 850, and 860 as described above with respect to process 800. As described with respect to process 800, the initial set of genes may decrease in number through subsequent iterations of these steps. As a result of these iterative steps, process 900 proceeds to operation 925, where the minimum size of the gene set is reached.
[0188] As part of these iterative steps, process 900 proceeds to operation 930 where a cross-validation score is calculated based on the input of the ranked expression data into the statistical model of operation 820. The cross-validation score can be calculated by performing k-fold cross-validation.
[0189] Process 900 proceeds to operation 940 where a gene set is selected based on the cross-validation score calculated in operation 930. In some embodiments, the selected gene set has the highest cross-validation score from a group of gene sets.
[0190] Next, process 900 proceeds to operation 950 where the expression data is ranked to obtain a gene ranking for genes represented by their expression levels in the expression data. Ranking process 108 can be used to rank the expression data to obtain the gene ranking.
[0191] Next, process 900 proceeds to operation 960 where hyperparameters for the statistical model are selected and fit to the statistical model for the gene set selected in operation 940.
[0192] For example, FIG. 9A is a plot of quality score versus number of genes, showing how the quality score increases by decreasing the number of genes from 30 to 28. FIG. 9B is an exemplary plot of F1 score versus number of genes used in the ranking for ABC / GCB origin tissue prediction according to some embodiments of the techniques described herein.
[0193] Origin cell DLBCL classifier As described herein, some embodiments involve using the techniques described herein to determine the cell of origin for DLBCL. In particular, a cell of origin DLBCL classifier may categorize a sample as being either germinal center B cell (GCB) or activated B cell (ABC). Such a classifier may be developed by using samples from the series GSE117556, the Leipzig Lymphoma dataset (10.1186 / s13073-019-0637-7), the series GSE31312, the series GSE10846, the series GSE87371, the series GSE11318, the series GSE32918, the series GSE23501, the Lymphoma / Leukemia Molecular Profiling Project (LLMPP), and the series GSE93984 as training data. For each dataset, samples were selected such that each dataset had a balanced cell of origin ratio ABC:GCB ratio of 40:60. For example, this may involve selecting samples with cell of origin labeling, followed by rounding of random selection of samples to obtain the desired ABC:GCB ratio. An exemplary cell of origin DLBCL classifier is described with respect to FIGS. 24A, 24B, 24C, 24D, and 24E. In this classifier, the training dataset includes 1,968 samples.
[0194] A suitable dataset can be used to validate the trained progenitor cell DLBCL classifier. Validation of progenitor cell DLBCL may involve using data from series GSE34171 (GPL96 + GPL97), series GSE22898, series GSE64555, series GSE145043, series GSE19246, and the National Cancer Institute Center for Cancer Research (NCICCR) "Genetics and Pathogenesis of Diffuse Large B Cell Lymphoma" dataset. Validation of the classifier described with respect to FIGS. 24A, 24B, 24C, 24D, and 24E involved using a validation dataset of 928 samples.
[0195] The classifier can be further validated using datasets of unknown and unclassified samples. The progenitor cell DLBCL classifier can be validated using data from series GSE69051, series GSE69049, E-TABM-346, series GSE68895, series GSE38202, series GSE2195, the International Cancer Genome Consortium Malignant Lymphoma-DE (ICGC_MALY_DE) dataset (https: / / icgc.org / node / 53049), and the National Cancer Institute Cancer Genome Characterization Initiative (NCICGCI) Non-Hodgkin Lymphoma dataset (https: / / ocg.cancer.gov / programs / cgci / projects / non-hodgkin-lymphoma). For the progenitor cell DLBCL classifier described with respect to FIGS. 24A, 24B, 24C, 24D, and 24E, 1,169 unknown and unclassified samples were used in the validation of the classifier.
[0196] The progenitor cell classifier described with respect to FIGS. 24A, 24B, 24C, 24D, and 24E may involve identifying a gene set, such as by the process 800 shown in FIG. 8. In particular, the initial gene set may be identified from the genes described in Wright G et al., A gene expression-based method to diagnose clinically distinct subgroups of diffuse large B cell lymphoma, PNAS, 2003, 100:9991-9996 (doi:10.1073 / pnas.1732008100), the entirety of which is incorporated herein by reference. The initial gene set was curated to reduce to up to 30 genes that would be used in the classifier. After hyperparameter tuning, the performance of the classifier in the validation dataset reached a 0.93 f1 score and a 0.978 AUC score.
[0197] In this exemplary classifier, binary classification was performed using a gradient booster decision tree classifier in LightGBM. Feature selection was performed by estimating the feature importance in the model using the SHAP package (https: / / github.com / slundberg / shap). Exemplary SHAP importance values calculated for possible genes for inclusion in the progenitor cell classifier for DLBCL are shown in Table 7 below.
[0198] [Table 7]
[0199] FIG. 24A shows a validation data set, an associated origin cell reported for samples of the data set, a predicted origin cell obtained using a machine learning technique described herein, and enrichment signatures for the ABC and GCB subtypes, for determining DLBCL subtypes. FIG. 24B shows a validation data set, an associated origin cell reported for samples of the data set, a predicted origin cell obtained using a machine learning technique described herein, and enrichment signatures for the ABC and GCB subtypes, for determining DLBCL subtypes. The validation data sets shown in FIGS. 24A and 24B vary in sample preparation, sequencing platform, and data processing used to obtain expression data. Both FIGS. 24A and 24B show a data set (upper panel) where each vertical line corresponds to a different sample, where the shading of the line corresponds to a different data set. The origin cell associated with the sample of the data set is shown, where a lighter shade indicates the GCB subtype and a darker shade indicates the ABC subtype. The origin cell associated with the sample can be a determination by a physician (e.g., a pathologist) using microscopy to visually inspect the sample. Enrichment signatures for the ABC and GCB signatures are shown in FIGS. 24A and 24B. The ABC signature generally has mostly upregulated genes in the samples in the right portion, and the GCB signature has mostly upregulated genes in the samples in the left portion. FIGS. 24C and 24D are plots of survival rates for different groups (ABC, GCB). FIG. 24E is a plot of true positive rate versus false positive rate for predicting DLBCL subtypes of different biological samples, where the classifier had a 0.978 AUC score.
[0200] Human Papillomavirus (HPV) Head and Neck Squamous Cell Carcinoma Classifier Some embodiments involve using the techniques described herein to predict the HPV status (HPV positive, HPV negative). Such embodiments can involve determining a sample as having an HPV positive or HPV negative status. In some embodiments, the HPV status can be determined for a subject who has, is suspected of having, or is at risk of having squamous cell carcinoma of the head and neck. Examples of genes that can be included in a gene set for determining the HPV status for squamous cell carcinoma of the head and neck are listed in Table 8 below. The gene set can include at least 3, 5, 10, or 20 genes selected from the group of genes listed in Table 8. In some embodiments, the gene set can include all of the genes listed in Table 8. In some embodiments, the gene set can include 3 to 130 genes, 5 to 130 genes, 20 to 130 genes, 50 to 130 genes, 80 to 130 genes listed in Table 8. In some embodiments, the gene set can include 130 or fewer genes, 100 or fewer genes, 80 or fewer genes, 50 or fewer genes, 20 or fewer genes listed in Table 8.
[0201]
Table 8A
[0202]
Table 8B
[0203]
Table 8C
[0204]
Table 8D
[0205]
Table 8E
[0206] Such a classifier can be developed by using samples from series GSE65858, series GSE41613, E-TABM-302 (from EMBL-EBI), series GSE25727, series GSE3292, series GSE6791, series GSE10300, and the TCGA HNSC dataset (from The Cancer Imaging Archive (TCIA)) as training data. For the classifier described with respect to FIGS. 25A, 25B, 25C, 25D, 25E, and 25F, 60 samples of the TCGA HNSC dataset were excluded from the training data and used in the validation dataset. The validation dataset included 60 samples from the TCGA HNSC dataset and series GSE40774. Series GSE74927 was used as an additional validation dataset representing different strains of the HPV virus, enabling the assessment of the classifier's performance across different HPV strains. The gene set for the classifier was identified from the genes described in Chakravarthy et al., Human Papillomavirus Drives Tumor Development Throughout the Head and Neck: Improved Prognosis Is Associated With an Immune Response Largely Restricted to the Oropharynx, Journal of Clinical Oncology, 34, no. 34 (December 01, 2016) 4132-4141 (DOI:10.1200 / JCO.2016.68.2955), the entire content of which is incorporated herein by reference. The initial gene set was curated to reduce to 82 genes, such as by using process 800 shown in FIG. 8. After hyperparameter tuning, the classifier's performance in the validation dataset with the TCGA HNSC dataset and series GSE40774 reached a 0.975 AUC score and a 0.9 f1 score. The classifier's performance in the validation dataset with series GSE74927 reached a 1.0 AUC score and a 1.0 f1 score.Note that the classifier has successfully recognized several HPV strains, including HPV16, HPV18, HPV33, and HPV55.
[0207] Figure 25A shows enrichment signatures for different pathways showing the validation dataset, the associated HPV status reported for samples of the dataset, the predicted HPV status obtained using the machine learning techniques described herein, and gene expression profiles associated with the HPV status. Figure 25A shows the dataset (upper panel) where each vertical line corresponds to a different sample, where the shading of the line corresponds to a different dataset. The HPV status associated with the samples of the dataset is shown, where lighter shading indicates a negative HPV status and darker shading indicates a positive HPV status. The probability that a sample has a positive HPV status is shown in the middle panel of Figure 25A. Enrichment signatures for different pathways showing gene expression profiles associated with the HPV status are shown in Figure 25A (lower panel). As an example, the HALLMARK_E2F_TARGETS signature is shown in Figure 25A, having mostly upregulated genes in the samples on the right portion and mostly downregulated genes in the samples on the left portion. Figures 25B and 25C are plots of survival rates for different groups of HPV status (positive HPV and negative HPV). Figure 25D is a plot of true positive rate versus false positive rate for predicting the HPV status of different biological samples (from the TCGA HNSC dataset and the series GSE40774 validation data), where the classifier had a 0.975 AUC score. Figure 25E is a plot of true positive rate versus false positive rate for predicting the HPV status of different biological samples (from the series GSE74927 validation data), where the classifier had a 1.0 AUC score. Figure 25F is a plot showing the performance of the classifier for different HPV strains in the series GSE74927 validation data.
[0208] Peripheral T-cell lymphoma (PTCL) classifier Aspects of the present application relate to techniques developed by the inventor for analyzing gene expression data to determine subtypes of peripheral T cell lymphoma (PTCL) for a biological sample. These techniques involve ranking a set of genes based on gene expression levels, as well as using the ranking and one or more statistical models to determine PTCL subtypes. The set of genes can be associated with biological characteristics (e.g., cell morphology, cell migration, cell cycle), expression pathways, or otherwise associated with one or more subtypes of peripheral T cell lymphoma (PTCL).
[0209] Peripheral T cell lymphoma accounts for about 10% of all non-Hodgkin lymphomas. Peripheral T cell lymphoma is a heterogeneous group of diseases that includes more than 20 subtypes, the exact definition of which is limited by modern methods of laboratory diagnosis. Examples of PTCL subtypes include, but are not limited to, peripheral T cell lymphoma, not otherwise specified (PTCL-NOS), anaplastic large cell lymphoma (ALCL), angioimmunoblastic T cell lymphoma (AITL), cutaneous T cell lymphoma (CTCL), natural killer / T cell lymphoma (NKTCL), Sézary syndrome, adult T cell leukemia / lymphoma (ATLL), enteropathy-type T cell lymphoma, nasal NK / T cell lymphoma, hepatosplenic gamma-delta T cell lymphoma, T cell lymphoma of follicular helper T cell (TFH) origin, T cell lymphoma of the gastrointestinal tract (e.g., EATL, MEITL), and the like.
[0210] The most frequent subgroups among PTCLs are adult T-cell leukemia / lymphoma (ATLL), angioimmunoblastic T-cell lymphoma (AITL), NK / T-cell lymphoma (NKTCL), anaplastic large cell lymphoma (ALCL), and cases belonging to the unspecified type (PTCL-NOS), which account for approximately 35% of all PTCL patients. Other PTCL subtypes are rare and are mostly represented by extranodal tumors. More effective annotation of PTCLs is expected to ultimately lead to the design and implementation of individualized therapies. As described herein, the inventors have recognized several benefits by using the ranking of a set of genes as opposed to specific values for gene expression levels. In some embodiments, the techniques described herein involve determining the subtype of peripheral T-cell lymphoma (PTCL) for a biological sample.
[0211] For example, in some embodiments, a ranking of genes based on gene expression levels (in a biological sample) determined by a sequencing platform can be provided as an input to a statistical model trained to predict the PTCL subtype for the biological sample. The statistical model can include a multiclass classifier and can have multiple outputs corresponding to different PTCL subtypes. As another example, in some embodiments, a ranking of genes based on gene expression levels (in a biological sample) determined by a sequencing platform can be provided as an input to multiple statistical models trained to predict different PTCL subtypes. For instance, one statistical model can be trained to predict anaplastic large cell lymphoma (ALCL) for a biological sample, and another statistical model can be trained to predict angioimmunoblastic T-cell lymphoma (AITL) for a biological sample. In such embodiments, the statistical models can be binary classifiers each trained for a different PTCL subtype, or regression-type classifiers that estimate the likelihood of a specific PTCL subtype.
[0212] Different PTCL subtypes may have different molecular signatures. In some embodiments, the set of ranked genes depends on a particular PTCL subtype of interest. In some embodiments, a certain set of genes may be used to determine a group of PTCL subtypes, and another set of genes may be used to determine a different group of PTCL subtypes. For example, a certain set of genes may be used to determine a group of PTCL subtypes including anaplastic large cell lymphoma (ALCL), and another set of genes may be used to determine a different group of PTCL subtypes including angioimmunoblastic T cell lymphoma (AITL), natural killer / T cell lymphoma (NKTCL), and adult T cell leukemia / lymphoma (ATLL). Another set of genes may be used to determine a group of PTCL subtypes including enteropathy-type T cell lymphoma, nasal NK / T cell lymphoma, and hepatosplenic gamma delta T cell lymphoma. As another example, a certain set of genes may be used to determine anaplastic large cell lymphoma (ALCL), and another set of genes may be used to determine natural killer / T cell lymphoma (NKTCL).
[0213] Some embodiments described herein address all of the problems described above that the inventors have recognized with respect to determining the PTCL subtype of a biological sample using gene expression data. However, not every embodiment described herein addresses every one of these problems, and some embodiments may not address any of them. Thus, it should be understood that the embodiments of the techniques described herein are not limited to addressing all or any of the problems described above with respect to determining the PTCL subtype of a biological sample using gene expression data.
[0214] Some embodiments involve obtaining gene expression data for a biological sample of interest, and ranking genes in a set of genes based on their expression levels in the expression data to obtain one or more gene rankings. The one or more gene rankings can be used with one or more statistical models to determine a subtype of PTCL for cells in the biological sample. The statistical model can be trained using rankings of expression levels for some or all of the genes in the set of genes.
[0215] In some embodiments, a gene ranking can be obtained by ranking genes in one or more sets of genes based on their expression levels in the expression data. In some embodiments, the expression data includes values each representing an expression level for a gene in the set of genes. Determining the gene ranking can involve determining a relative rank for each gene in the set of genes based on the values. For example, a first gene ranking can be obtained by ranking genes in a first set of genes based on their expression levels.
[0216] In some embodiments, the expression data can be obtained for cells in a biological sample, where the subject has or is suspected of having cancer. In some embodiments, the expression data can be obtained for cells in a biological sample, where the subject has or is suspected of having lymphoma. In some embodiments, the subject has or is suspected of having PTCL.
[0217] In some embodiments, the processing pipeline 100 shown in FIG. 1 may be used to determine one or more PTCL subtypes. In such embodiments, gene ranking and statistical models may be used to determine one or more PTCL subtypes of a biological sample. In some embodiments, a certain set of genes may be used to determine the PTCL subtype for a biological sample, and another set of genes may be used to determine the origin tissue. For example, statistical model 112a and gene set 1 106a may be used to determine the PTCL subtype for cells in a biological sample, and statistical model 112b and gene set 2 106b may be used to determine the origin tissue for cells in a biological sample. In some embodiments, different gene sets may be used to determine different PTCL subtypes. For example, gene set 1 106a may be used to determine whether a biological sample has an AITL subtype, and gene set 2 106b may be used to determine whether a biological sample has an ATLL subtype.
[0218] In some embodiments, different gene sets and different statistical models may be used to determine different PTCL subtypes. For example, statistical model 112a and gene set 1 106a may be used to determine a certain PTCL subtype (e.g., AITL) for cells in a biological sample, and statistical model 112b and gene set 2 106b may be used for another PTCL subtype (e.g., ATLL) for cells in a biological sample.
[0219] The statistical model used to determine the PTCL subtype can be trained using data from one or more of the series GSE58445, series GSE45712, series GSE1906, series GSE90597, series GSE6338, series GSE36172, series GSE65823, series GSE118238, series GSE78513, series GSE51521, series GSE14317, series GSE80631, series GSE19067, and series GSE20874 available through the GEO database. As another example, the statistical model used to determine the PTCL subtype can be trained using data from one or more of the cohorts described in Table 9 below.
[0220]
Table 9A
[0221]
Table 9B
[0222] In some embodiments, the PTCL subtype can be determined for cells in a biological sample using the techniques described herein. The PTCL subtype can include peripheral T-cell lymphoma, not otherwise specified (PTCL-NOS), anaplastic large cell lymphoma (ALCL), angioimmunoblastic T-cell lymphoma (AITL), cutaneous T-cell lymphoma (CTCL), natural killer / T-cell lymphoma (NKTCL), Sézary syndrome, adult T-cell leukemia / lymphoma (ATLL), enteropathy-type T-cell lymphoma, nasal NK / T-cell lymphoma, hepatosplenic gamma-delta T-cell lymphoma, T-cell lymphoma of follicular T helper (TFH) origin, and T-cell lymphoma of the gastrointestinal tract.
[0223] In some embodiments, the set of genes used to obtain gene rankings can include genes associated with biological characteristics, expression pathways, or otherwise associated with determining one or more PTCL subtypes. Examples of genes that can be included in such a gene set are listed in Table 10 below.
[0224]
Table 10A
[0225]
Table 10B
[0226]
Table 10C
[0227]
Table 10D
[0228]
Table 10E
[0229]
Table 10F
[0230]
Table 10G
[0231] Some embodiments involve using gene sets that include genes associated with the molecular signatures of one or more PTCL subtypes. Examples of genes that may be included in such gene sets are listed in Table 11 below, which shows different genes and their corresponding PTCL subtypes. In some embodiments, one or more genes listed in Table 11 are combined with one or more genes listed in Table 10 to form a gene set that is used to determine PTCL subtypes by the techniques described herein.
[0232] [Table 11]
[0233] For additional examples of genes that may be included in the gene sets used to determine PTCL subtypes by the techniques described herein, see Iqbal J, Wright G, Wang C, et al., Gene expression signatures delineate biological and prognostic subgroups in peripheral T-cell lymphoma, Blood, 2014, 123(19):2915 - 2923 (doi:10.1182 / blood-2013-11-536359), which is hereby incorporated by reference in its entirety.
[0234] Some embodiments may involve using a gene set that includes genes upregulated in angioimmunoblastic T-cell lymphoma (AITL) compared to normal T lymphocytes, which may be referred to herein as "genes upregulated in AITL." For example, one or more genes in the gene set PICCALUGA_ANGIOIMMUNOBLASTIC_LYMPHOMA_UP with the systematic name M12225 in the Gene Set Enrichment Analysis (GSEA) database may be used in determining PTCL subtypes by the techniques described herein. In some embodiments, the gene set includes A2M, ABCC3, ABI3BP, ACKR1, ACTA2, ACVRL1, ADAMDEC1, ADAMTS1, ADAMTS9, ADGRF5, ADGRL4, ADRA2A, ANK2, ANKRD29, ANTXR1, APOC1, APOE, ARHGAP29, ARHGAP42, ARHGEF10, ASPM, ATOX1, C1QA, C1QB, C1QC, C1R, C1S, C2, C3, C4A, C7, CALD1, CARMN, CAV2, CAVIN1, CCDC102B, CCDC80, CCL14, CCL19, CCL2, CCL21, CCN4, CD63, CD93, CDH11, CDH5, CETP, CFB, CFH, CHI3L1, CLMP, CLU, CMKLR1, COL12A1, COL15A1, COL1A1, COL1A2, COL3A1, COL4A1, COL4A2, COL6A1, COL8A2, COX7A1, CP, CSRP2, CTHRC1, CTSC, CTSL, CTTNBP2NL, CXCL10, CXCL12, CXCL9, CYBRD1, CYFIP1, CYP1B1, CYP26B1, CYP27A1, DAB2, DCLK1, DDR2, DEPP1, DHRS7B, DOCK4, DPYSL3, EMCN, EMILIN1, ENG, ENPP2, EPHX1, FAM107A, FAM114A1, FAM20A, FBN1, FCHO2, FERMT2, FLRT2, FN1, FSTL1, FUCA1, GABBR1, GASK1B, GJA1, GJC1, GPNMB, GPRC5B, GUCY1B1, HNMT, HSPB8, HSPG2,It may contain one or more genes selected from the group consisting of IDH1, IFI27, IGFBP5, IGFBP7, IL18, IL33, IRAK3, ITGA9, ITPRIPL2, KCNJ10, KCNMA1, KCTD12, LAMA4, LAMB1, LAMC1, LIFR, LOXL1, LPAR1, LUM, MARCKS, MFAP4, MIR1245A, MIR34AHG, MMP9, MXRA5, MYL9, MYLK, NAGK, NEXN, NFIB, NNMT, NPL, NR1H3, NR2F2, OSMR, P2RY13, PAPSS2, PARVA, PCOLCE, PDGFRA, PDLIM5, PDPN, PGF, PLA2G2D, PLA2G4C, PLD1, PLPP3, PMP22, PPIC, PRRX1, PTGDS, RAB13, RAI14, RARRES2, RASSF4, RBP5, RBPMS, RGL1, RGS5, RHOBTB3, RND3, RPE, RRAS, RSPO3, S1PR3, SAMD9L, SEPTIN10, SERPING1, SERPINH1, SLAMF8, SLC1A3, SLC40A1, SLCO2B1, SMOC2, SPARC, SPARCL1, SPRED1, SULF1, TAGLN, TANC1, TCIM, TDO2, TEAD2, THY1, TJP1, TLR4, TMEM163, TMEM176A, TMEM176B, TNC, TNS1, TNS3, TPM1, TRIM47, VCAM1, VWF, WDFY3, WLS, WWTR1, YAP1, and ZNF226.,
[0235] Some embodiments may involve using a gene set that includes genes that are downregulated in angioimmunoblastic T-cell lymphoma (AITL) compared to normal T lymphocytes, and these genes may be referred to herein as "genes downregulated in AITL". For example, one or more genes in the gene set PICCALUGA_ANGIOIMMUNOBLASTIC_LYMPHOMA_DN with the systematic name M4781 in the gene set enrichment analysis (GSEA) database may be used in determining PTCL subtypes by the techniques described herein.In some embodiments, the gene set may comprise one or more genes selected from the group consisting of AMD1, AREG, ATP2B1-AS1, B3GNT2, BOLA2, BTG1, C16orf72, CBX4, CCDC59, CCNL1, CD6, CD69, CHD1, CLK1, CNOT6L, CNST, COG3, CREM, CSGALNACT2, CSRNP1, DDX3X, DNAJB6, DUSP10, DUSP2, DUSP4, EIF1, EIF4E, EIF4G3, EIF5, EPC1, ETNK1, FBXO33, FBXW7, FOSB, FOSL2, FOXP1, G3BP2, GABARAPL1, GADD45A, GADD45B, GATA3, H2AC18, H3-3B, HAUS3, HECA, HIPK1, ID2, IDS, IER5, IFRD1, IKZF5, ING3, IRF2BP2, IRS2, JMJD1C, JMY, JUN, JUND, KDM3A, KDM6B, KLF10, KLF4, KLF6, LINC-PINT, LINC01578, LY9, MAP3K8, MCL1, MEX3C, MGAT4A, MOAP1, MPZL3, MXD1, MYLIP, NAMPT, NDUFA10, NR4A2, NR4A3, PCIF1, PDE4D, PELI1, PER1, PHF1, PIGA, PMAIP1, PNPLA8, PPP1R15A, PPP1R15B, PRNP, PTGER4, PTP4A1, PTP4A2, RAPGEF6, REL, RGCC, RGS1, RGS2, RNF103, RNF11, RNF139, RSRC2, SARAF, SBDS, SETD2, SIK1, SIK3, SLC2A3, SLC30A1, SMURF2, SNORD22, SNORD3B-1, SON, SRSF5, STK17B, SUCO, THAP2, TIPARP, TMX4, TNFAIP3, TOB1, TP53INP2, TRA2B, TSC22D2, TSC22D3, TSPYL2, TTC7A, TUBB2A, WIPF1, YPEL5, ZBTB10, ZBTB24, ZFAND2A, ZFAND5, ZFC3H1, ZFP36, and ZNF331.
[0236] Some embodiments may involve using a gene set that includes genes associated with a molecular function (MF) profile of a subject, sometimes referred to herein as "MF profile genes." In some embodiments, the genes associated with an MF profile may include genes in one or more modules of the MF profile. Examples of genes associated with MF profiles and modules of MF profiles are described and recited in U.S. Patent No. 10,311,967, entitled "SYSTEMS AND METHODS FOR GENERATING, VISUALIZING AND CLASSIFYING MOLECULAR FUNCTION PROFILES," issued June 4, 2019, the entire contents of which are incorporated herein by reference. In some embodiments, one or more of the genes associated with an MF profile, and one or more of the genes listed in Table 10, may be used in combination as a gene set for determining PTCL subtypes.
[0237] Some embodiments may involve determining the PTCL subtype for cells in a biological sample by using a statistical model that outputs multiple PTCL subtype predictions corresponding to different PTCL subtypes, which are used to determine the PTCL subtype for a biological sample. FIG. 26 is a diagram of an exemplary processing pipeline 2600 for determining the PTCL subtype of a biological sample according to some embodiments of the techniques described herein. The exemplary processing pipeline 2600 may include ranking genes based on their gene expression levels, and using the ranking and a statistical model to determine the PTCL subtype. The processing pipeline 2600 may be executed on any suitable computing device (e.g., a single computing device, multiple computing devices collocated at a single physical location or located at multiple physical locations remote from each other, one or more computing device components of a cloud computing system, etc.), because aspects of the techniques described herein are not limited in this regard. In some embodiments, the processing pipeline 2600 may be executed by a desktop computer, a laptop computer, or a mobile computing device. In some embodiments, the processing pipeline 2600 may be executed within one or more computing devices that are part of a cloud computing environment.
[0238] In some embodiments, gene expression data 102 and a ranking process 108 are used to rank genes based on their expression levels in the gene expression data 102 to obtain a gene ranking 110. The gene ranking 110 may be input into a statistical model 112. The statistical model 112 may be trained using training data that indicates rankings of expression levels for some or all of the genes in a gene set.
[0239] In some embodiments, the statistical model 112 may output a prediction of a biological sample having a specific PTCL subtype. In some cases, the prediction output by the statistical model may include the probability that the biological sample has a PTCL subtype. As shown in FIG. 26, the statistical model 112 outputs a PTCL subtype prediction 1 216a, a PTCL subtype prediction 2 216b, a PTCL subtype prediction 3 216c, and a PTCL subtype prediction 4 216d. The prediction output by the statistical model 112 may be analyzed using the prediction analysis process 118 to determine the PTCL subtype 214 for the biological sample. The prediction analysis process 118 may involve selecting a specific PTCL subtype for the biological sample from among the different PTCL subtype predictions. In some embodiments, the PTCL subtype prediction may include the probability that the biological sample has a specific PTCL subtype. In such embodiments, the prediction analysis process 118 may involve selecting the PTCL subtype based on the probability. In some embodiments, selecting the PTCL subtype may involve selecting the PTCL subtype having the highest probability as the PTCL subtype 214.
[0240] In some embodiments, the statistical model 112 may provide outputs each corresponding to a different PTCL subtype. For example, the PTCL subtype prediction 1 216a may correspond to anaplastic large cell lymphoma (ALCL), the PTCL subtype prediction 2 216b may correspond to angioimmunoblastic T cell lymphoma (AITL), the PTCL subtype prediction 3 216c may correspond to natural killer / T cell lymphoma (NKTCL), and the PTCL subtype prediction 4 216d may correspond to adult T cell leukemia / lymphoma (ATLL). In some embodiments, the statistical model 112 may include a multi-class classifier. In some embodiments, class weights may be implemented for one or more of the classes in the multi-class classifier. Examples of classifiers that the statistical model 112 may include are gradient boosted decision tree classifier, decision tree classifier, gradient boost classifier, random forest classifier, clustering-based classifier, Bayesian classifier, Bayesian network classifier, neural network classifier, kernel-based classifier, and support vector machine classifier.
[0241] Four outputs from the statistical model 112 are shown in FIG. 26, but it should be understood that a statistical model having any suitable number of outputs for PTCL subtype prediction may be implemented using the techniques described above when determining the PTCL subtype of a biological sample. In some embodiments, the outputs may be in the range of 3 to 5, 3 to 10, 3 to 15, or 3 to 20.
[0242] Some embodiments may involve determining a PTCL subtype for cells in a biological sample by using a plurality of statistical models that correspond to different PTCL subtypes and output predictions for those PTCL subtypes, which are used to determine the PTCL subtype for a biological sample. FIG. 27 is a diagram of an exemplary processing pipeline 2700 for determining the PTCL subtype of a biological sample, according to some embodiments of the techniques described herein. The exemplary processing pipeline 2700 may include ranking genes based on their gene expression levels, as well as using the rankings and statistical models to determine the PTCL subtype. The processing pipeline 2700 may be executed on any suitable computing device (e.g., a single computing device, multiple computing devices collocated at a single physical location or located at multiple physical locations remote from each other, one or more computing device portions of a cloud computing system, etc.), because aspects of the techniques described herein are not limited in this regard. In some embodiments, the processing pipeline 2700 may be executed by a desktop computer, a laptop computer, a mobile computing device. In some embodiments, the processing pipeline 2700 may be executed within one or more computing devices that are part of a cloud computing environment.
[0243] In some embodiments, gene expression data 102 and a ranking process 108 are used to rank genes based on their expression levels in the gene expression data 102 to obtain a gene ranking 110. The gene ranking 110 can be input into a statistical model 1 112a, a statistical model 2 112b, a statistical model 3 112c, and a statistical model 4 112d. Each of the statistical model 1 112a, the statistical model 2 112b, the statistical model 3 112c, and the statistical model 4 112d can be trained using training data indicating rankings of expression levels for some or all of the genes in a gene set. The statistical model 1 112a, the statistical model 2 112b, the statistical model 3 112c, and the statistical model 4 112d each correspond to a different PTCL subtype and can output a prediction of a biological sample having that particular PTCL subtype. In some cases, the prediction output by the statistical model can include the probability that the biological sample has a PTCL subtype.
[0244] As shown in Fig. 27, statistical model 1 112a outputs PTCL subtype prediction 1 316a, statistical model 2 112b outputs PTCL subtype prediction 2 316b, statistical model 3 112c outputs PTCL subtype prediction 3 316c, and statistical model 4 112d outputs PTCL subtype prediction 4 316d. Each of statistical model 1 112a, statistical model 2 112b, statistical model 3 112c, and statistical model 4 112d may correspond to a different PTCL subtype. For example, statistical model 1 112a and PTCL subtype prediction 1 316a may correspond to anaplastic large cell lymphoma (ALCL), and statistical model 1 112a may be trained using a ranking of expression levels for one or more genes associated with ALCL, such as those listed in Table 11. As another example, statistical model 2 112b and PTCL subtype prediction 2 316b may correspond to angioimmunoblastic T cell lymphoma (AITL), and statistical model 2 112b may be trained using a ranking of expression levels for one or more genes associated with AITL, such as those listed in Table 11. As yet another example, statistical model 3 112c and PTCL subtype prediction 3 316c may correspond to natural killer / T cell lymphoma (NKTCL), and statistical model 3 112c may be trained using a ranking of expression levels for one or more genes associated with NKTCL, such as those listed in Table 11. As another example, statistical model 4 112d and PTCL subtype prediction 4 316d may correspond to adult T cell leukemia / lymphoma (ATLL), and statistical model 4 112d may be trained using a ranking of expression levels for one or more genes associated with ATLL, such as those listed in Table 11.
[0245] The predictions output by statistical model 1 112a, statistical model 2 112b, statistical model 3 112c, and statistical model 4 112d can be analyzed using prediction analysis process 118 to determine the PTCL subtype 214 for a biological sample. The prediction analysis process 118 may involve selecting a specific PTCL subtype for the biological sample from among different PTCL subtype predictions. In some embodiments, the PTCL subtype prediction may include the probability that the biological sample has a specific PTCL subtype. In such embodiments, the prediction analysis process 118 may involve selecting the PTCL subtype based on the probability. In some embodiments, selecting the PTCL subtype may involve selecting the PTCL subtype with the highest probability as the PTCL subtype 214.
[0246] In some embodiments, one or more of statistical model 1 112a, statistical model 2 112b, statistical model 3 112c, and statistical model 4 112d may include a binary classifier. In some embodiments, each of statistical model 1 112a, statistical model 2 112b, statistical model 3 112c, and statistical model 4 112d includes a binary classifier. In such embodiments, if none of the binary classifiers used is decisive as to which class the biological sample belongs to, the sample may be determined to be unclassified. In some embodiments, statistical model 1 112a, statistical model 2 112b, statistical model 3 112c, and statistical model 4 112d may have a hierarchical classifier configuration.
[0247] Some embodiments may involve a hierarchical configuration of four classifiers in the order of a first classifier for the NKTCL PTCL subtype, a second classifier for the ATLL PTCL subtype, a third classifier for the AITL PTCL subtype, and a fourth classifier for the ALCL PTCL subtype. In some embodiments, each of the first, second, third, and fourth classifiers is a binary classifier.
[0248] Four statistical models and corresponding outputs are shown in FIG. 27. It is to be understood that any number of statistical models can be implemented using the techniques described above when determining the PTCL subtype of a biological sample. In some embodiments, the number of statistical models can be in the range of 3-5, 3-10, 3-15, or 3-20.
[0249] Some embodiments can involve determining the PTCL subtype of a biological sample by obtaining PTCL subtype predictions using different gene sets and statistical models corresponding to the different gene sets, where the PTCL subtype predictions are used to determine the PTCL subtype. FIG. 28 is a diagram of an exemplary processing pipeline 2800 for determining the PTCL subtype of a biological sample according to some embodiments of the techniques described herein. The exemplary processing pipeline 2800 can include ranking genes based on their gene expression levels and using the rankings and statistical models to determine the PTCL subtype. The processing pipeline 2800 can be executed on any suitable computing device (e.g., a single computing device, multiple computing devices collocated at a single physical location or located at multiple physical locations remote from each other, one or more computing device portions of a cloud computing system, etc.) because aspects of the techniques described herein are not limited in this regard. In some embodiments, the processing pipeline 2800 can be executed by a desktop computer, a laptop computer, a mobile computing device. In some embodiments, the processing pipeline 2800 can be executed within one or more computing devices that are part of a cloud computing environment.
[0250] In some embodiments, gene expression data 102 is used to rank genes in different sets of genes based on their expression levels in the gene expression data 102 to obtain a plurality of gene rankings. For example, a gene ranking can be obtained for each gene set, and the gene ranking can be input into a statistical model that is trained using training data indicating the ranking of the expression levels for some or all of the genes in the gene set. As shown in FIG. 28, ranking process 108 can involve using expression data 102 to rank genes in different gene sets, including gene set 1 106a, gene set 2 106b, gene set 3 106c, and gene set 4 106d, to obtain gene ranking 1 110a, gene ranking 2 110b, gene ranking 3 110c, and gene ranking 4 110d, respectively. The ranking process 108 can involve ranking the genes in a set of genes based on the numerical values of their expression levels. Different gene rankings can be obtained by ranking the expression levels for different gene sets, and each gene ranking can be input into its respective statistical model to obtain a PTCL subtype prediction. As shown in FIG. 28, gene ranking 1 110a, gene ranking 2 110b, gene ranking 3 110c, and gene ranking 4 110d are provided as inputs to statistical model 1 112a, statistical model 2 112b, statistical model 3 112c, and statistical model 4 112d, respectively.
[0251] In some embodiments, different statistical models and their respective gene sets can correspond to a particular PTCL subtype of a biological sample. In such embodiments, each of the statistical models can output a prediction of a biological sample having a particular PTCL subtype. In some cases, the prediction output by a statistical model can include the probability that the biological sample has a PTCL subtype.
[0252] As shown in FIG. 28, statistical model 1 112a outputs PTCL subtype prediction 1 416a, statistical model 2 112b outputs PTCL subtype prediction 2 416b, statistical model 3 112c outputs PTCL subtype prediction 3 416c, and statistical model 4 112d outputs PTCL subtype prediction 4 416d. The predictions output by different statistical models can be analyzed using prediction analysis process 118 to determine the PTCL subtype 114 for the biological sample.
[0253] Although four gene sets and four statistical models are shown in FIG. 28, it should be understood that any suitable number of gene sets and corresponding statistical models can be implemented using the techniques described above to determine PTCL subtype predictions in order to obtain the PTCL subtype of a biological sample. In some embodiments, the number of gene sets and corresponding statistical models can be in the range of 3 to 100, 3 to 70, 3 to 50, 3 to 40, 3 to 30, 5 to 50, 10 to 60, or 10 to 70.
[0254] In some embodiments, the number of gene sets and corresponding statistical models is less than or equal to the number of classes for the PTCL subtypes. Such embodiments may involve different gene sets and corresponding statistical models for each PTCL subtype. For example, gene set 1 106a and statistical model 1 112a may be used to generate a prediction that the PTCL subtype (as PTCL subtype prediction 1 416a) is anaplastic large cell lymphoma (ALCL), gene set 2 106b and statistical model 2 112b may be used to generate a prediction that the PTCL subtype (as PTCL subtype prediction 2 416b) is angioimmunoblastic T cell lymphoma (AITL), gene set 3 106c and statistical model 3 112c may be used to generate a prediction that the PTCL subtype (as PTCL subtype prediction 3 416c) is natural killer / T cell lymphoma (NKTCL), and gene set 4 106d and statistical model 4 112d may be used to generate a prediction that the PTCL subtype (as PTCL subtype prediction 4 416d) is adult T cell leukemia / lymphoma (ATLL). It should be understood that additional gene sets and their corresponding statistical models may be implemented for different PTCL subtypes.
[0255] Figure 29 is a flowchart of an exemplary process 2900 for determining a PTCL subtype of a biological sample using gene ranking and a statistical model, according to some embodiments of the techniques described herein. Process 2900 can be executed on any suitable computing device (e.g., a single computing device, multiple computing devices collocated at a single physical location or located at multiple physical locations remote from each other, one or more computing device portions of a cloud computing system, etc.) because aspects of the techniques described herein are not limited in this regard. In some embodiments, ranking process 108 and statistical model 112 can execute some or all of process 2900 to determine a PTCL subtype.
[0256] Process 2900 begins at operation 2910, where expression data for a target biological sample is obtained. In some embodiments, the expression data can be obtained using a gene expression microarray. In some embodiments, the expression data can be obtained by performing next-generation sequencing. In some embodiments, the expression data can be obtained by using a hybridization-based expression assay. Some embodiments involve performing a sequencing process of the biological sample (e.g., gene expression microarray, next-generation sequencing) prior to obtaining the expression data 102. In some embodiments, obtaining the gene expression data 102 can involve accessing, using a computing device, expression data in one or more data stores (e.g., expression data previously obtained from a biological sample), receiving expression data from one or more other devices, or obtaining the gene expression data 102 in silico by any other method, etc. In some embodiments, obtaining the gene expression data 102 can involve (in vitro) analyzing the biological sample and accessing the expression data (e.g., by a computing device, a processor). Further aspects regarding obtaining the expression data are provided in the section entitled "Obtaining Expression Data".
[0257] Next, process 2900 proceeds to operation 2920, where genes in a set of genes are ranked based on their expression levels in the expression data, such as by using the ranking process 108, to obtain a gene ranking. The expression data can include values each representing an expression level for a gene in the set of genes, and determining the gene ranking can involve determining a relative rank for each gene in the set of genes based on the values.
[0258] In some embodiments, the subject has, is suspected of having, or is at risk of having breast cancer. The set of genes can be selected from the group of genes listed in Table 10. The set of genes can include at least 3, 5, 10, or 20 genes selected from the group of genes listed in Table 10. In some embodiments, the set of genes can include all of the genes listed in Table 10. In some embodiments, the set of genes can include 3 to 120 genes, 5 to 120 genes, 20 to 120 genes, 50 to 120 genes, 80 to 120 genes listed in Table 10. In some embodiments, the set of genes can include 120 or fewer genes, 100 or fewer genes, 80 or fewer genes, 50 or fewer genes, 20 or fewer genes listed in Table 10.
[0259] In some embodiments, the subject has, is suspected of having, or is at risk of having lymphoma. In some embodiments, the subject has, is suspected of having, or is at risk of having PTCL.
[0260] Next, process 2900 proceeds to operation 2930 where the PTCL subtype of the biological sample is determined using gene ranking and a statistical model such as statistical model 112. The statistical model can be trained using a ranking of expression levels for one or more genes in a set of genes. In some embodiments, the gene ranking can be used as an input to the statistical model to obtain an output indicating the PTCL subtype. In some embodiments, the statistical model comprises one or more classifiers selected from the group consisting of a gradient boosted decision tree classifier, a decision tree classifier, a gradient boosting classifier, a random forest classifier, a clustering-based classifier, a Bayesian classifier, a Bayesian network classifier, a neural network classifier, a kernel-based classifier, and a support vector machine classifier. In some embodiments, the statistical model can involve using a machine learning algorithm implementing a gradient boosting framework such as gradient boosted decision tree (GBDT) and gradient boosted regression tree (GBRT). Examples of software packages implementing machine learning algorithms that can be used by the techniques described herein include the LightGBM package, the XGBoost package, and the pGBRT package.
[0261] In some embodiments, the statistical model can include a multi-class classifier. The multi-class classifier can provide at least four outputs each corresponding to a different PTCL subtype. For example, the first output can correspond to anaplastic large cell lymphoma (ALCL), the second output can correspond to angioimmunoblastic T cell lymphoma (AITL), the third output can correspond to natural killer / T cell lymphoma (NKTCL), and the fourth output can correspond to adult T cell leukemia / lymphoma (ATLL).
[0262] In some embodiments, the statistical model may include a plurality of classifiers corresponding to different PTCL subtypes. For example, the first classifier may correspond to anaplastic large cell lymphoma (ALCL), the second classifier may correspond to angioimmunoblastic T cell lymphoma (AITL), the third classifier may correspond to natural killer / T cell lymphoma (NKTCL), and the fourth classifier may correspond to adult T cell leukemia / lymphoma (ATLL). In some embodiments, the plurality of classifiers may be binary classifiers. The binary classifier may have a hierarchical classification. For example, the statistical model may include four binary classifiers with a hierarchical structure in the order of a first classifier for the NKTCL PTCL subtype, a second classifier for the ATLL PTCL subtype, a third classifier for the AITL PTCL subtype, and a fourth classifier for the ALCL PTCL subtype.
[0263] In some embodiments, the PTCL subtype is selected from the group consisting of anaplastic large cell lymphoma (ALCL), angioimmunoblastic T cell lymphoma (AITL), natural killer / T cell lymphoma (NKTCL), and adult T cell leukemia / lymphoma (ATLL). In some embodiments, the PTCL subtype is selected from the group consisting of peripheral T cell lymphoma, not otherwise specified (PTCL-NOS), anaplastic large cell lymphoma (ALCL), angioimmunoblastic T cell lymphoma (AITL), cutaneous T cell lymphoma (CTCL), natural killer / T cell lymphoma (NKTCL), Sézary syndrome, adult T cell leukemia / lymphoma (ATLL), enteropathy-type T cell lymphoma, nasal NK / T cell lymphoma, hepatosplenic gamma-delta T cell lymphoma, T cell lymphoma of follicular T helper (TFH) origin, and T cell lymphoma of the gastrointestinal tract.
[0264] In some embodiments, process 2900 may include outputting the PTCL subtype to the user, such as by displaying the PTCL subtype to the user (e.g., a physician) in a graphical user interface (GUI), including the PTCL subtype in a report, sending an email to the user, and any other suitable method.
[0265] In some embodiments, process 2900 may include treating a subject based on the determined PTCL subtype of a biological sample. For example, a physician may perform a treatment for a subject associated with treating lymphoma of the determined PTCL subtype. Further examples of using the PTCL subtype of a biological sample determined using the techniques described herein for performing a treatment are provided in the section entitled "Methods of Treatment."
[0266] In some embodiments, process 2900 may include identifying a treatment for a subject based on the determined PTCL subtype. For example, the determined PTCL subtype may be used to identify a treatment for a subject associated with treating lymphoma of the determined PTCL subtype.
[0267] In some embodiments, process 2900 may include determining a prognosis for a subject based on the determined PTCL subtype. For example, the determined PTCL subtype may be used to determine a prognosis for a subject associated with treating lymphoma of the determined PTCL subtype.
[0268] Further aspects regarding other applications in which the PTCL subtype of a biological sample determined using the techniques described herein is used for determining a prognosis are provided in the section entitled "Applications."
[0269] In some embodiments, the trained statistical model used to determine the PTCL subtype can be evaluated using existing clinical data to determine its performance in identifying the PTCL subtype. As an example, a gene set having the genes listed in Table 10 was used for the ranking process 108, and a multiclass classifier was used to determine whether a sample belongs to the AITL, ATLL, ALCL, NKTCL, or PTCL NOS subtype. The clinical data listed in Table 9 was used for this evaluation process, and Table 12 below shows the PTCL subtypes identified using this process. The statistical model used achieved an 0.84 f1 score. FIG. 30 is a plot of survival rates for different PTCL subtypes (AITL, ATLL, ALCL, NKTCL, and PTCL NOS).
[0270]
Table 12A
[0271]
Table 12B
[0272] In some aspects, the methods for characterizing cancer described herein can be applied to any lymphoma. "Lymphoma" generally refers to cancers (e.g., tumors) that arise from lymph nodes and lymphocytes. Lymphomas are typically classified according to the normal cell type from which the tumor cells arise, such as, for example, T-cell lymphoma, B-cell lymphoma, Hodgkin (lymphocyte) lymphoma, and histiocytic and dendritic cell tumors. Classification of lymphomas is described, for example, by Jiang et al., Expert Rev. Hematol. March 2017, 10(3):239-249. Classification of PTCL lymphomas is described, for example, by Iqbal J, Wright G, Wang C et al., Gene expression signatures delineate biological and prognostic subgroups in peripheral T-cell lymphoma, Blood, 2014, 123(19):2915-2923 (doi:10.1182 / blood-2013-11-536359), which is incorporated herein by reference in its entirety.
[0273] In some embodiments, the lymphoma is a B-cell lymphoma. In some embodiments, the B-cell lymphoma is a diffuse large B-cell lymphoma (DLBCL). Classification of DLBCL is described, for example, by Alizadeh et al., Distinct types of diffuse large B-cell lymphoma identified by gene expression profiling, Nature 403, 503-511 (2000) (doi:10.1038 / 35000501). Examples of DLBCL include, but are not limited to, germinal center B-cell (GCB) subtypes, and activated B-cell (ABC) subtypes.
[0274] In some embodiments, the lymphoma is a T cell lymphoma. In some embodiments, the T cell lymphoma is a mature T cell lymphoma such as peripheral T cell lymphoma (PTCL). More than 25 mature T cell lymphomas have been identified. Examples of PTCL include, but are not limited to, peripheral T cell lymphoma, not otherwise specified (PTCL-NOS), anaplastic large cell lymphoma (ALCL), angioimmunoblastic T cell lymphoma (AITL), cutaneous T cell lymphoma (CTCL), natural killer / T cell lymphoma (NKTCL), Sézary syndrome, adult T cell leukemia / lymphoma (ATLL), enteropathy-type T cell lymphoma, nasal NK / T cell lymphoma, hepatosplenic gamma-delta T cell lymphoma, T cell lymphoma of follicular T cell (TFH) origin, T cell lymphoma of the gastrointestinal tract (e.g., EATL, MEITL), and the like.
[0275] In some embodiments, the lymphoma is anaplastic large cell lymphoma (ALCL). In some embodiments, the ALCL is systemic ALCL. In some embodiments, the ALCL is cutaneous ALCL (e.g., ALCL that affects the skin). In some embodiments, the ALCL is ALK-positive ALCL. In some embodiments, the ALCL is ALK-negative ALCL.
[0276] In some embodiments, the lymphoma is angioimmunoblastic T cell lymphoma (AITL). In some embodiments, the AITL tumor cells express one or more follicular T cell markers, such as CD10 and CD279 (PD-1, PDCD1), CXCL13, BCL6, CD40L, or NFATC1.
[0277] In some embodiments, the lymphoma is adult T cell leukemia / lymphoma (ATLL). In some embodiments, the ATLL results from infection with the HTLV-1 virus.
[0278] In some embodiments, the lymphoma is natural killer / T cell lymphoma (NKTCL). In some embodiments, the NKTCL tumor is located in the palate and / or paranasal sinuses of the subject. In some embodiments, the NKTCL tumor is located in the nasal cavity of the subject.
[0279] Obtaining expression data The expression data described herein (e.g., microarray data, next generation sequencing (NGS) data) can be obtained from a variety of sources. In some embodiments, the expression data can be obtained by analyzing a biological sample of a subject. The biological sample can be analyzed prior to performing the techniques described herein, including techniques for ranking genes based on their expression levels and using the ranking to determine one or more characteristics of the biological sample. In some such embodiments, the data obtained from the biological sample is stored (e.g., in a database) and can be accessed during performance of the techniques described herein. Thus, “obtaining expression data” as described herein can involve obtaining gene expression data in silico by using a computing device to access expression data (e.g., expression data previously obtained from a biological sample) in one or more data stores, receiving expression data from one or more other devices, or any other means, such as (in vitro) analyzing a biological sample, or combinations thereof. Examples of additional techniques regarding how expression data is obtained are described in U.S. Patent No. 10,311,967, titled “SYSTEMS AND METHODS FOR GENERATING, VISUALIZING AND CLASSIFYING MOLECULAR FUNCTION PROFILES,” issued on June 4, 2019, which is hereby incorporated by reference in its entirety.
[0280] In some embodiments, the expression data can include expression levels for all mRNAs in a cell, or for a subset of RNAs in a cell (e.g., for a subset of RNAs expressed from a gene set, or for a subset of genes in one or more of the gene sets described in this application, or including at least some or consisting of some of the genes in those gene sets), for the entire cellular RNA. The RNA levels can be obtained using any suitable technique, including sequencing and / or hybridization-based techniques (e.g., whole exome sequencing data, targeted specificity sequencing data for a subset of RNA, microarray data, etc.).
[0281] Biological sample Any method, system, assay, or other suitable technique can be used to analyze any biological sample from a subject (e.g., a patient). In some embodiments, the biological sample can be any sample from a subject known or suspected to have cancer, including cancer cells or precancerous cells.
[0282] The biological sample can be any type of sample, including, for example, a sample of body fluid, one or more cells, a tissue piece, or part or all of an organ. In some embodiments, the sample can be from a cancerous tissue or organ, or from a tissue or organ suspected of having one or more cancer cells. In some embodiments, the sample can be from a healthy (e.g., non-cancerous) tissue or organ. In some embodiments, a sample from a subject (e.g., a biopsy from a subject) can include both healthy cells and / or tissues and cancer cells and / or tissues. In some embodiments, one sample will be taken from the subject for analysis.
[0283] Any of the biological samples described herein can be obtained from a subject using any known technique. In some embodiments, a biological sample can be obtained from a surgical procedure (e.g., laparoscopy, microsurgery, or endoscopy), bone marrow biopsy, punch biopsy, endoscopic biopsy, or needle biopsy (e.g., fine needle aspiration, core needle biopsy, vacuum assisted biopsy, or image-guided biopsy). In some embodiments, each of the biological samples is a body fluid sample, a cell sample, or a tissue biopsy. In some embodiments, one or more cells (cell sample) are obtained from a subject using a scrape or brush method. The cell sample can be obtained from any area within the subject's body, including, for example, one or more of the following areas: neck, esophagus, stomach, bronchus, or mouth, or from the subject's body. In some embodiments, one or more tissue pieces (e.g., tissue biopsy) from the subject can be used. In some embodiments, the tissue biopsy can comprise one or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10) samples from one or more tumors or tissues known or suspected to have cancer cells.
[0284] Sample Analysis The methods described herein are based at least in part on the identification and characterization of several biological processes and / or molecular and cellular components present within and / or surrounding cancer (e.g., tumor).
[0285] Biological processes within and / or surrounding cancer (e.g., tumor) include, but are not limited to, angiogenesis, metastasis, proliferation, cell activation (e.g., T cell activation), tumor infiltration, immune response, cell signaling (e.g., HER2 signaling), and apoptosis.
[0286] The molecular and cellular composition within and / or surrounding a cancer (e.g., a tumor) includes, but is not limited to, nucleic acids (e.g., DNA and / or RNA), molecules (e.g., hormones), proteins (e.g., wild-type and / or mutant proteins), and cells (e.g., malignant and / or non-malignant cells). As used herein, the cancer microenvironment refers to the molecular and cellular environment in which a cancer (e.g., a tumor) exists, including, but not limited to, blood vessels, immune cells, fibroblasts, bone marrow-derived inflammatory cells, lymphocytes, signaling molecules, and extracellular matrix (ECM) that surround and / or are within the tumor.
[0287] The molecular and cellular composition and biological processes that exist within and / or surrounding a tumor can be directed towards promoting cancer (e.g., tumor) growth and survival (e.g., pro-tumor), and / or inhibiting cancer (e.g., tumor) growth and survival (e.g., anti-tumor).
[0288] The cancer (e.g., tumor) microenvironment can comprise cellular composition and biological processes directed towards promoting cancer (e.g., tumor) growth and survival (e.g., pro-tumor microenvironment), and / or inhibiting cancer (e.g., tumor) growth and survival (e.g., anti-tumor microenvironment). In some embodiments, the cancer (e.g., tumor) microenvironment comprises a pro-cancer (e.g., pro-tumor) microenvironment. In some embodiments, the cancer (e.g., tumor) microenvironment comprises an anti-cancer (e.g., anti-tumor) microenvironment. In some embodiments, the cancer (e.g., tumor) microenvironment comprises both a pro-cancer (e.g., pro-tumor) microenvironment and an anti-cancer (e.g., anti-tumor) microenvironment.
[0289] Any information regarding molecular and cellular compositions and biological processes present within and / or surrounding a cancer (e.g., a tumor) can be used in the methods for characterizing a cancer (e.g., a tumor) described herein. In some embodiments, a cancer (e.g., a tumor) can be characterized based on gene set expression levels (e.g., gene set RNA expression levels). In some embodiments, a cancer (e.g., a tumor) is characterized based on protein expression.
[0290] The methods for characterizing a cancer described herein can be applied to any cancer (e.g., any tumor). Exemplary cancers include, but are not limited to, adrenocortical carcinoma, bladder urothelial carcinoma, breast invasive carcinoma, cervical squamous cell carcinoma, endocervical adenocarcinoma, colon adenocarcinoma, esophageal carcinoma, kidney clear cell carcinoma, kidney papillary cell carcinoma, liver hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, prostate adenocarcinoma, rectal adenocarcinoma, skin cutaneous melanoma, stomach adenocarcinoma, thyroid carcinoma, uterine corpus endometrial carcinoma, and cholangiocarcinoma.
[0291] Expression data Expression data for a plurality of genes (e.g., indicating expression levels) can be used for any of the methods described herein. The number of genes that can be considered can include and be less than all the genes of the subject.
[0292] To obtain expression data for a plurality of genes (e.g., indicating expression levels), any method can be used on a sample from the subject. As a non-limiting set of examples, the expression data can be RNA expression data, DNA expression data, or protein expression data.
[0293] DNA expression data, in some embodiments, refers to the level of DNA in a sample from a subject. For example, the level of DNA in a sample from a subject having cancer, such as gene duplication in a cancer patient sample, can be higher compared to the level of DNA in a sample from a subject not having cancer. For example, the level of DNA in a sample from a subject having cancer, such as gene deletion in a cancer patient sample, can be reduced compared to the level of DNA in a sample from a subject not having cancer.
[0294] DNA expression data, in some embodiments, refers to data about DNA (or genes) expressed in a sample, such as sequencing data about genes expressed in a patient sample. Such data can, in some embodiments, be useful for determining whether a patient has one or more mutations associated with a particular cancer.
[0295] RNA expression data can be obtained using any method known in the art, including but not limited to whole transcriptome sequencing, total RNA sequencing, mRNA sequencing, targeted RNA sequencing, small RNA sequencing, ribosome profiling, RNA exome capture sequencing, and / or deep RNA sequencing. DNA expression data can be obtained using any method known in the art, including any known method of DNA sequencing. For example, DNA sequencing can be used to identify one or more mutations in the DNA of a subject. Any technique used in the art for sequencing DNA can be used in conjunction with the methods described herein. As a non-limiting set of examples, DNA can be sequenced through single molecule real-time sequencing, ion torrent sequencing, pyrosequencing, sequencing by synthesis, sequencing by ligation (SOLiD sequencing), nanopore sequencing, or Sanger sequencing (chain termination sequencing). Protein expression data can be obtained using any method known in the art, including but not limited to N-terminal amino acid analysis, C-terminal amino acid analysis, Edman degradation (including the use of machines such as a protein sequenator), or mass spectrometry.
[0296] In some embodiments, the expression data comprises next-generation sequencing (NGS) data. In some embodiments, the expression data comprises microarray data. In some embodiments, the expression data comprises whole exome sequencing (WES) data. In some embodiments, the expression data comprises whole genome sequencing (WGS) data. In some embodiments, the expression data comprises RNA Seq data (e.g., by performing RNA sequencing). In some embodiments, the expression data comprises a combination of RNA Seq data and WGS data. In some embodiments, the expression data comprises a combination of RNA Seq data and WES data.
[0297] assay Any of the biological samples described herein can be used to obtain expression data using a conventional assay or an assay described herein. The expression data, in some embodiments, includes gene expression levels. The gene expression levels can be detected by detecting the products of gene expression, such as mRNA and / or protein.
[0298] In some embodiments, the gene expression levels are determined by detecting the level of protein in the sample and / or by detecting the level of protein activity in the sample. As used herein, the terms "determining" or "detecting" can include assessing the presence, absence, quantity, and / or amount (which can be an effective amount) of a substance in a sample, which includes deriving qualitative or quantitative concentration levels of such a substance, or alternatively, evaluating the value and / or categorization of such a substance in a sample from a subject.
[0299] Protein levels can be measured using immunoassays. Examples of immunoassays can include (but are not limited to) any known assay, such as the following: immuno-blot assays (e.g., Western blot), immunohistochemical analysis, flow cytometry assays, immunofluorescence assays (IF), enzyme-linked immunosorbent assays (ELISA) (e.g., sandwich ELISA), radioimmunoassays, electrochemiluminescence-based detection assays, magnetic immunoassays, lateral flow assays, and any of the related techniques. Additional suitable immunoassays for detecting the protein levels provided herein will be apparent to those skilled in the art.
[0300] Such immunoassays can involve the use of an agent (e.g., an antibody) that specifically binds to the target protein. An agent such as an antibody that "specifically binds" to a target protein is a term well understood in the art, and methods for determining such specific binding are also well known in the art. An antibody is said to "specifically bind" when it reacts or associates with a particular target protein more frequently, more rapidly, with greater duration, and / or with greater affinity than it does with alternative proteins. Also, by reading this definition, it should be understood that, for example, an antibody that specifically binds to a first target peptide may or may not specifically or selectively bind to a second target peptide. Thus, "specific binding" or "selective binding" does not necessarily require (although it can include) exclusive binding. Generally, although not always, a reference to binding means selective binding. In some examples, an antibody that "specifically binds" to a target peptide or its epitope may not bind to other peptides or other epitopes in the same antigen. In some embodiments, a sample can be contacted with two or more binding agents that bind to different proteins simultaneously or sequentially (e.g., multiplex analysis).
[0301] It will be apparent to those skilled in the art that the present disclosure is not limited to immunoassays. Detection assays that are not antibody-based, such as mass spectrometry, are also useful for the detection and / or quantification of the proteins and / or protein levels provided herein. Assays that rely on chromogenic substrates may also be useful for the detection and / or quantification of the proteins and / or protein levels provided herein.
[0302] Alternatively, the level of nucleic acid encoding a gene in a sample can be measured via conventional methods. In some embodiments, measuring the expression level of a nucleic acid encoding a gene includes measuring mRNA. In some embodiments, the expression level of mRNA encoding a gene can be measured using real-time reverse transcriptase (RT) Q-PCR, or a nucleic acid microarray. Methods for detecting nucleic acid sequences include, but are not limited to, polymerase chain reaction (PCR), reverse transcriptase PCR (RT-PCR), in situ PCR, quantitative PCR (Q-PCR), real-time quantitative PCR (RT Q-PCR), in situ hybridization, Southern blot, Northern blot, sequence analysis, microarray analysis, reporter gene detection, or other DNA / RNA hybridization platforms.
[0303] In some embodiments, the level of nucleic acid encoding a gene in a sample can be measured via a hybridization assay. In some embodiments, the hybridization assay comprises at least one binding partner. In some embodiments, the hybridization assay comprises at least one oligonucleotide binding partner. In some embodiments, the hybridization assay comprises at least one labeled oligonucleotide binding partner. In some embodiments, the hybridization assay comprises at least one pair of oligonucleotide binding partners. In some embodiments, the hybridization assay comprises at least one pair of labeled oligonucleotide binding partners.
[0304] Any binding agent that specifically binds to a desired nucleic acid or protein can be used in the methods and kits described herein to measure the expression level in a sample. In some embodiments, the binding agent is an antibody or aptamer that specifically binds to a desired protein. In other embodiments, the binding agent can be one or more oligonucleotides that are complementary to a nucleic acid or a portion thereof. In some embodiments, a sample can be contacted with two or more binding agents that bind different proteins or different nucleic acids, simultaneously or sequentially (e.g., multiplex analysis).
[0305] To measure the expression level of a protein or nucleic acid, a sample can be contacted with a binding agent under suitable conditions. Generally, the term "contact" refers to exposing the binding agent to the sample or cells collected therefrom for a suitable period sufficient for the formation of a complex between the binding agent and, if any, the target protein or target nucleic acid in the sample. In some embodiments, contacting is performed by capillary action in which the sample is moved across the surface of a support membrane.
[0306] In some embodiments, the assay can be performed on a low-throughput platform that includes a single assay format. In some embodiments, the assay can be performed on a high-throughput platform. Such high-throughput assays can include using a binding agent immobilized on a solid support (e.g., one or more chips). Methods for immobilizing the binding agent can depend on factors such as the nature of the binding agent and the material of the solid support and may require specific buffers. Such methods will be apparent to those skilled in the art.
[0307] gene The various genes described herein are generally named using human gene nomenclature. The various genes are, in some embodiments, described in publicly available resources such as published journal articles. Gene names can be correlated with additional information (including sequence information) by use of databases such as the NCBI GenBank® database available at www.ncbi.nlm.nih.gov, the HUGO (Human Genome Organization) Gene Nomenclature Committee (HGNC) database available at www.genenames.org, and the DAVID Bioinformatics Resource available at www.david.ncifcrf.gov. Gene names can also be correlated with additional information through publications from the above organizations, which are hereby incorporated by reference herein for this purpose. It should be understood that a gene can encompass all variants of that gene. For organisms or subjects other than human subjects, corresponding specific-specific genes can be used. Synonyms, equivalents, and closely related genes (including genes from other organisms) can be identified using similar databases including the NCBI GenBank® database described above.
[0308] Some embodiments involve using a gene set for predicting breast cancer grade that includes the genes listed in Table 1. Some embodiments involve using a gene set for predicting renal clear cell carcinoma grade that includes the genes listed in Table 2. Some embodiments involve using a gene set for predicting the tissue of origin for diffuse large B-cell lymphoma (DLBCL), such as germinal center B cells (GCB) and activated B cells (ABC), that includes the genes listed in Table 3. Some embodiments involve using a gene set for predicting PTCL subtypes that includes the genes listed in Table 10.
[0309] Application Examples The methods for characterizing a biological sample, which may include tumor type characterization, described herein can be used for a variety of clinical purposes including, but not limited to, monitoring cancer progression in a subject, assessing the effectiveness of treatment for cancer, identifying patients suitable for a particular treatment, evaluating patient eligibility to participate in a clinical trial, and / or predicting recurrence in a subject. Thus, what is described herein is a method of diagnosis and prognosis for cancer treatment based on the tumor types described herein.
[0310] The methods described herein can be used to evaluate the effectiveness of cancer treatment, such as that described herein, in view of the correlation between cancer type (e.g., tumor type) and cancer prognosis. For example, multiple biological samples, such as those described herein, can be collected from a subject on whom treatment is being performed, either before and after treatment or during treatment. The cancer type (e.g., tumor type) in the biological sample from the subject can be determined using any of the methods described herein. For example, if the cancer type indicates that the subject has a poor prognosis and, after treatment or over the course of treatment, the cancer type changes to a cancer type indicating a good prognosis, this indicates that the treatment is effective.
[0311] In some embodiments, the cancer type may also be used to identify cancers that may be treatable using specific anti-cancer therapeutics (e.g., chemotherapy). To implement this method, the cancer type in a sample (e.g., a tumor biopsy) collected from a subject having cancer may be determined using the methods described herein. If the cancer type is identified as being susceptible to treatment with a particular anti-cancer therapeutic, the method may further comprise administering an effective amount of that anti-cancer therapeutic to the subject having cancer.
[0312] In some embodiments, the methods for cancer type characterization described herein may be relied upon in the development of new therapies for cancer. In some embodiments, the cancer type may indicate or predict the effectiveness of a new therapy, or the progression of cancer in a subject, before, during, or after the application of the new therapy.
[0313] In some embodiments, the methods for cancer type characterization described herein may be used to evaluate the suitability of patients for participating in clinical trials. In some embodiments, the cancer type may be used to include patients in clinical trials. In some embodiments, patients having a particular cancer grade (e.g., grade 1) are included in clinical trials. In some embodiments, patients having a particular tissue of origin for cancer are included in clinical trials. In some embodiments, the cancer type may be used to exclude patients from clinical trials. In some embodiments, patients having a particular cancer grade (e.g., grade 3) are excluded from clinical trials. In some embodiments, patients having a particular tissue of origin are excluded from clinical trials. In some embodiments, patients having a particular PTCL subtype are excluded from clinical trials.
[0314] In some embodiments, the methods described herein can be used in monitoring the progression of a patient's disease and identifying one or more treatments based on the disease stage determined using the techniques described herein. In some embodiments, the monitoring is performed over a time period during which a first disease stage is identified for a patient for the first time and a second disease stage is identified for the patient for the second time. The second disease stage can be used to identify different types of treatment. For example, in the context of using the techniques described herein to predict cancer grade, monitoring the patient's disease and identifying different treatments based on the disease stage can involve obtaining first expression data obtained by sequencing a first biological sample of a subject (e.g., a subject having kidney cancer), determining a first cancer grade using the first expression data and the statistical models described herein, identifying or recommending a first treatment for the subject based on the first cancer grade, and optionally, performing the first treatment. Monitoring the patient's disease can further involve obtaining second expression data obtained by sequencing a second biological sample of the subject (e.g., a biological sample obtained from the subject at a different time than the first biological sample), determining a second cancer grade using the second expression data, identifying or recommending a second treatment for the subject based on the second cancer grade, and optionally, performing the second treatment. In some embodiments, the first cancer grade is different from the second cancer grade and the first treatment is different from the second treatment. In some embodiments, the monitoring can be performed multiple times (e.g., with multiple medical visits) to evaluate the progression of the treatment, to determine how the patient is responding to a particular treatment, or a combination thereof.
[0315] In some embodiments, the methods described herein can be used in assessing how a subject has responded to a treatment. For example, these techniques described herein can be used in determining whether a subject is responding to a series of treatments, whether the subject is in remission, and whether there is a recurrence of the disease.
[0316] In some embodiments, the characteristics for cells of a biological sample of a subject determined using the techniques described herein can be used in identifying a diagnosis for the subject. In some embodiments, the characteristics can provide information for a physician or other user to determine a diagnosis for the subject. For example, the characteristics alone may be sufficient for a physician to be able to determine a diagnosis. In some embodiments, a combination of the characteristics and other patient medical data can be used by a physician or other user in determining a diagnosis for the subject.
[0317] In some embodiments, the characteristics for cells of a biological sample of a subject determined using the techniques described herein can be used in identifying a prognosis for the subject. In some embodiments, the characteristics can provide information for a physician or other user to determine a prognosis for the subject. For example, the characteristics alone may be sufficient for a physician to be able to determine a prognosis. In some embodiments, a combination of the characteristics and other patient medical data can be used by a physician or other user in determining a prognosis for the subject.
[0318] In some embodiments, a diagnosis or prognosis determined using the techniques described herein can be used in recommending a treatment or therapy for the subject. The therapy can be a drug treatment, radiation, surgery, diet or lifestyle modification, or other therapy. The treatment can be chemotherapy, immunotherapy, hormonal therapy, or other treatment. In some embodiments, recommending a treatment or therapy can include changing a treatment (e.g., a different treatment, an additional treatment, or a different frequency or dosage).
[0319] In some embodiments, the diagnosis or prognosis determined using the techniques described herein can be used in generating recommendations for further analysis of a patient. For example, recommendations for further diagnostic interventions (e.g., more extensive CAT scans, MRIs, more extensive or invasive biopsies, more detailed genetic, proteomic, or histological analysis of one or more tissue samples, etc.).
[0320] In some embodiments, the diagnosis or prognosis determined using the techniques described herein can be used in generating recommendations for changing the frequency of medical examinations during follow - up. For example, recommendations for more frequent medical examinations if the analysis suggests a higher risk, or for less frequent medical examinations if the analysis suggests a lower risk or that the subject is in remission.
[0321] In some embodiments, the characteristics for cells of a biological sample of a subject determined using the techniques described herein can be used in generating a report specific to the subject. For example, the report can be a patient - specific cancer characteristics report. Generating the report can involve generating a file that includes information indicating disease characteristics (e.g., cancer grade, tissue of origin, tissue subtype) determined using the techniques described herein.
[0322] In the context of providing recommendations or other information to a physician or other user, providing such information can involve transmitting electronic information to the physician or other user. In some embodiments, the electronic information can be transmitted to a medical center or to a computer system that hosts patient medical information, and the physician or other user can access the information using a computing device.
[0323] Examples of additional applications of how the characteristics of a biological sample can be used as determined using the techniques described herein are described in U.S. Patent No. 10,311,967, entitled "SYSTEMS AND METHODS FOR GENERATING, VISUALIZING AND CLASSIFYING MOLECULAR FUNCTION PROFILES," issued on June 4, 2019, which is hereby incorporated by reference in its entirety.
[0324] Method of treatment In some of the methods described herein, an effective amount of an anti-cancer therapy as described herein is applied to a subject (e.g., a human) in need of treatment via a suitable route (e.g., intravenous administration) or may be recommended for application.
[0325] Subjects to be treated by the methods described herein can be human patients having cancer, suspected of having cancer, or at risk thereof. Subjects having cancer, suspected of having cancer, or at risk thereof are subjects showing one or more signs or symptoms of cancer, subjects diagnosed with cancer, subjects having a family history and / or genetic predisposition to cancer, and / or subjects having one or more other risk factors for cancer (e.g., age, exposure to carcinogens, environmental exposure, exposure to viruses associated with a higher likelihood of developing cancer, etc.). Examples of cancer include, but are not limited to, melanoma, lung cancer, brain tumor, breast cancer, colorectal cancer, pancreatic cancer, liver cancer, prostate cancer, skin cancer, kidney cancer, or bladder cancer. Subjects to be treated by the methods described herein can be mammals (e.g., can be human). Mammals include, but are not limited to, farm animals (e.g., livestock), sport animals, laboratory animals, pets, primates, horses, dogs, cats, mice, and rats.
[0326] As used herein, an "effective amount" refers to the amount of each active agent necessary to confer a therapeutic effect on a subject, either alone or in combination with one or more other active agents. The effective amount will vary according to factors such as the particular condition being treated, the severity of the condition, and individual patient parameters including age, health, body weight, sex, and weight, as well as the duration of the treatment, the nature of any combination therapy (if any), the specific route of administration, and like factors, within the knowledge and expertise of medical practitioners as recognized by those of ordinary skill in the art. These factors are well known to those of ordinary skill in the art and can be addressed by routine experimentation only. It is generally preferred that the maximum dosage of an individual component or combination thereof, i.e., the most safe dosage by reasonable medical judgment, be used. However, it will be understood by those of ordinary skill in the art that a patient may claim a lower dosage or an acceptable dosage for medical, psychological, or almost any other reason.
[0327] Examples of additional methods of treatment are described in U.S. Patent No. 10,311,967, entitled "SYSTEMS AND METHODS FOR GENERATING, VISUALIZING AND CLASSIFYING MOLECULAR FUNCTION PROFILES," issued on June 4, 2019, which is hereby incorporated by reference in its entirety.
[0328] Quality control analysis In some embodiments, the techniques described herein can be used when performing quality control. One application is quality control analysis in a laboratory setting. For example, a sequencing laboratory may receive a biological sample along with information about the biological sample. In addition to an identifier and / or tracking number, such information may include information about the characteristics of the biological sample (e.g., tissue source, cancer type, cancer grade, etc.). However, due to laboratory errors (e.g., errors such as a patient sample being swapped, being mislabeled, or incorrect information being provided), the biological sample provided may not actually have these characteristics.
[0329] Another application example is for quality control analysis in data analysis settings. For example, patient sequencing data (e.g., reads, aligned reads, expression levels, etc.) can be provided as input to a data processing pipeline. However, if the sequencing data does not correspond to the aligned source (e.g., due to errors and from different patients), the results of the analysis will probably be meaningless.
[0330] In some embodiments, quality control can be performed by comparing the determined characteristics of a biological sample with the predicted characteristics determined using the techniques described herein. When the determined and predicted characteristics match (e.g., are the same or within an acceptable difference), it can be determined that the quality control inspection has been satisfied. On the other hand, if the predicted and determined characteristics do not match, further action may need to be taken. For example, further analysis of the biological sample can be performed, the biological sample can be rejected, the data processing pipeline can be halted or not executed (thereby saving valuable and costly computing resources), and laboratory operators and / or other parties (e.g., clinicians, staff, etc.) can be notified of the potential discrepancy (e.g., by e-mail alerts, messages, reports, entries into log files, etc.).
[0331] For example, a classifier for determining a cancer grade can be used to predict the cancer grade from gene expression data of a sample, and the predicted cancer grade can be compared to the determined cancer grade for the sample. If the predicted cancer grade and the determined cancer grade match, it can be determined that the sample analysis meets the quality control criteria. However, if the predicted cancer grade and the determined cancer grade do not match, further analysis can be performed. As another example, a classifier for determining the tissue of origin can be used to predict the tissue type for a sample, and the predicted tissue type can be compared to the determined tissue type for the sample. If the predicted tissue type and the determined tissue type do not match, further analysis of the biological sample can be performed to identify the tissue type for the sample. Any of the classification techniques described herein can be used in this way, either alone or in combination with each other, to provide multiple quality control checkpoints.
[0332] Examples of additional quality control analysis are described in U.S. Patent Application No. 16 / 920,636, entitled "TECHNIQUES FOR BIAS CORRECTION IN SEQUENCE DATA," filed on July 3, 2020, which is hereby incorporated by reference in its entirety.
[0333] Computing system An exemplary implementation of a computer system 1000 that may be used with any of the embodiments of the technology described herein is shown in FIG. 10. The computer system 1000 includes one or more processors 1010 and one or more manufactured products including non-transitory computer-readable storage media (e.g., memory 1020 and one or more non-volatile storage media 1030). The processor 1010 may control writing data to and reading data from the memory 1020 and the non-volatile storage device 1030 in any suitable manner, because aspects of the technology described herein are not limited in this regard. To execute any of the functions described herein, the processor 1010 may execute one or more processor-executable instructions stored in one or more non-transitory computer-readable storage media (e.g., memory 1020), and the one or more non-transitory computer-readable storage media (e.g., memory 1020) may act as non-transitory computer-readable storage media that store processor-executable instructions for execution by the processor 1010.
[0334] The computing device 1000 may also include a network input / output (I / O) interface 1040 through which the computing device may communicate with other computing devices (e.g., on a network), and one or more user I / O interfaces 1050 through which the computing device may provide output to and receive input from a user. The user I / O interface may include devices such as a keyboard, a mouse, a microphone, a display device (e.g., a monitor or a touch screen), a speaker, a camera, and / or various other types of I / O devices.
[0335] The embodiments described above can be implemented in any of a number of ways. For example, the embodiments can be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed on any suitable processor (e.g., a microprocessor) or set of processors, whether provided on a single computing device or distributed among a plurality of computing devices. It should be understood that any component or set of components that perform the functions described above can generally be regarded as one or more controllers that control the functions described above. The one or more controllers can be implemented in a number of ways, such as using dedicated hardware or using general-purpose hardware (e.g., one or more processors) programmed using microcode or software to perform the functions described above.
[0336] In this regard, it should be understood that one implementation of the embodiments described herein comprises at least one computer-readable storage medium (e.g., RAM, ROM, EEPROM, flash memory, or other memory technology, CD-ROM, digital versatile disk (DVD), or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage, or other magnetic storage device, or other tangible non-transitory computer-readable storage medium) encoded with a computer program (e.g., a plurality of executable instructions) that, when executed on one or more processors, performs the functions described above for one or more of the embodiments. The computer-readable medium can be transferable such that the program stored thereon can be loaded onto any computing device for implementing aspects of the techniques described herein. Additionally, it should be understood that reference to a computer program that, when executed, performs any of the functions described above is not limited to an application program executing on a host computer. Rather, the terms computer program and software are used herein in a general sense to refer to any type of computer code (e.g., application software, firmware, microcode, or any other form of computer instruction) that can be employed to program one or more processors for implementing aspects of the techniques described herein.
[0337] The terms "program" or "software" are used herein in a general sense to refer to any type of computer code, or set of processor-executable instructions, that can be employed to program a computer or other processor to implement the various aspects of the embodiments as described above. Additionally, according to one aspect, one or more computer programs that, when executed, perform the methods of the present disclosure provided herein need not be present on a single computer or processor and may be distributed modularly among different computers or processors to implement the various aspects of the present disclosure provided herein.
[0338] Processor-executable instructions can be in many forms such as program modules that are executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Typically, the functionality of program modules can be combined or distributed as needed in various embodiments.
[0339] Also, data structures can be stored in one or more non-transitory computer-readable storage media in any suitable form. For simplicity of illustration, a data structure can be shown as having associated fields through locations within the data structure. Such relationships can be achieved similarly by allocating storage for the fields along with the locations in the non-transitory computer-readable medium that convey the relationships between the fields. However, any suitable mechanism can be used to establish relationships between the information within the fields of a data structure, including the use of pointers, tags, or other mechanisms that establish relationships between data elements.
[0340] Also, various concepts of the present invention can be implemented as one or more processes, and examples thereof are provided. The operations executed as part of each process can be ordered in any suitable manner. Accordingly, embodiments can be configured in which the operations are executed in an order different from the order in which they are shown, and those embodiments can include performing some operations simultaneously, even though in the exemplary embodiments they are shown as sequential operations.
[0341] Aspects of the technology described herein provide a computer-implemented method for generating, visualizing, and classifying the biological characteristics (e.g., cancer grade, tissue of origin) of a cancer patient.
[0342] In some embodiments, the software program can provide the user with a visual representation of patient characteristics and / or other information regarding the patient's cancer using an interactive graphical user interface (GUI). Such a software program can be executed in any suitable computing environment, including but not limited to a cloud computing environment, a device collocated with the user (e.g., the user's laptop, desktop, smartphone, etc.), one or more devices remote from the user (e.g., one or more servers), and the like.
[0343] For example, in some embodiments, the techniques described herein can be implemented in the exemplary environment 1100 shown in FIG. 11. As shown in FIG. 11, within the exemplary environment 1100, one or more biological samples of patient 1102 can be provided to laboratory 1104. Laboratory 1104 can process the biological sample to obtain expression data (e.g., DNA, RNA, and / or protein expression data) and provide the expression data to at least one database 1106 that stores information about patient 1102 via network 1108.
[0344] Network 1108 can be a wide area network (e.g., the Internet), a local area network (e.g., a corporate intranet), and / or any other suitable type of network. Any of the devices shown in FIG. 11 can be connected to network 1108 using one or more wired links, one or more wireless links, and / or any suitable combination thereof.
[0345] In the illustrated embodiment of FIG. 11, at least one database 1106 can store expression data for a patient, medical history data for a patient, test result data for a patient, and / or any other suitable information about patient 1102. Examples of stored test result data for a patient include biopsy test results, imaging test results (e.g., MRI results), and blood test results. The information stored in at least one database 1106 can be stored in any suitable format and / or using any suitable data structure, because the aspects of the techniques described herein are not limited in this regard. At least one database 1106 can store data in any suitable manner (e.g., one or more databases, one or more files). At least one database 1106 can be a single database or multiple databases.
[0346] As shown in FIG. 11, an exemplary environment 1100 includes one or more external databases 1116 that may store information for patients other than patient 1102. For example, external database 1116 may store expression data (of any suitable type) for one or more patients, medical history data for one or more patients, test result data for one or more patients (e.g., imaging results, biopsy results, blood test results), demographic and / or historical information for one or more patients, and / or any other suitable type of information. In some embodiments, external database 1116 may store information available in one or more publicly accessible databases, such as TCGA (The Cancer Genome Atlas), one or more databases of clinical trial information, and / or one or more databases maintained by commercial sequencing providers. External database 1116 may store such information using any suitable hardware and in any suitable manner, as the aspects of the techniques described herein are not limited in this regard.
[0347] In some embodiments, at least one of database 1106 and external database 1116 may be the same database, may be part of the same database system, or may be physically collocated, as the aspects of the techniques described herein are not limited in this regard.
[0348] For example, in some embodiments, server 1110 may access information stored in database 1106 and / or 1116 and use this information to execute process 300, described with reference to FIG. 3, to determine one or more characteristics of a biological sample.
[0349] As another example, in some embodiments, server 1110 can access information stored in databases 1106 and / or 1116 and use this information to perform process 400, described with reference to FIG. 4, to determine the origin tissue for some or all of the cells in a biological sample.
[0350] As another example, in some embodiments, server 1110 can access information stored in databases 1106 and / or 1116 and use this information to perform process 500, described with reference to FIG. 5, to determine the cancer grade for some or all of the cells in a biological sample.
[0351] As another example, in some embodiments, server 1110 can access information stored in databases 1106 and / or 1116 and use this information to perform process 800, described with reference to FIG. 8A, to select a gene set.
[0352] As another example, in some embodiments, server 1110 can access information stored in databases 1106 and / or 1116 and use this information to perform process 2900, described with reference to FIG. 29, to determine the PTCL subtype of a biological sample. In some embodiments, server 1110 can include one or more computing devices. When server 1110 includes multiple computing devices, those devices can be physically collocated (e.g., in a single room) or dispersed across multiple physical locations. In some embodiments, server 1110 can be part of a cloud computing infrastructure. In some embodiments, one or more servers 1110 can be collocated within a facility operated by an entity (e.g., a hospital, a research institution) with which physician 1114 is associated. In such embodiments, it can be easier to enable server 1110 to access personal medical data for patient 1102.
[0353] As shown in FIG. 11, in some embodiments, the results of the analysis performed by server 1110 can be provided to physician 1114 via computing device 1112, which can be a portable computing device such as a laptop or smartphone, or a fixed computing device such as a desktop computer. The results can be provided in the form of a paper report, an email, a graphical user interface, and / or any other suitable method. In the embodiment of FIG. 11, the results are provided to the physician, but it should be understood that in other embodiments, the results of the analysis can be provided to patient 1102 or a caregiver of patient 1102, a healthcare provider such as a nurse, or a person associated with a clinical trial.
[0354] In some embodiments, the results can be part of a graphical user interface (GUI) presented to physician 1114 via computing device 1112. In some embodiments, the GUI can be presented to the user as part of a web page displayed by a web browser running on computing device 1112. In some embodiments, the GUI can be presented to the user using an application program (different from a web browser) running on computing device 1112. For example, in some embodiments, computing device 1112 can be a mobile device (e.g., a smartphone), and the GUI can be presented to the user via an application program (e.g., an "app") running on the mobile device.
[0355] The GUI presented on computing device 1112 can provide a wide range of oncological data regarding both the patient and the patient's cancer in a new, compact, and highly informative way. Previously, oncological data was obtained from multiple data sources multiple times, and the process of obtaining such information has become costly in terms of both time and expense. Using the techniques and graphical user interfaces described herein, a user can access the same amount of information at once, reducing the burden on the user and the computing resources required to provide such information. Reducing the burden on the user helps to reduce clinician errors associated with exploring various information sources. Reducing the burden on computing resources helps to reduce the processor power, network bandwidth, and memory required to provide a wide range of oncological data, which is an improvement in computing technology. It is to be understood that all definitions as defined and used herein shall prevail over dictionary definitions and / or the ordinary meaning of the defined terms.
[0356] As used in this specification and the claims, the phrase "at least one" with respect to the listing of one or more elements means at least one element selected from any one or more of the elements in the listing of elements, but does not necessarily include at least one of every element specifically listed within the listing of elements, and is understood not to exclude any combination of elements in the listing of elements. This definition also allows for the possibility that elements other than those specifically identified may optionally exist, whether or not they are related to the elements specifically identified within the listing of elements referred to by the phrase "at least one". Thus, by way of non-limiting example, "at least one of A and B" (or equivalently "at least one of A or B", or equivalently "at least one of A and / or B") can, in one embodiment, refer to at least one A where B is absent (and optionally includes elements other than B), optionally including two or more, in another embodiment, refer to at least one B where A is absent (and optionally includes elements other than A), optionally including two or more, and in yet another embodiment, refer to at least one A, optionally including two or more, and at least one B, optionally including two or more (and optionally includes other elements).
[0357] As used herein and in the claims, the phrase "and / or" is to be understood to mean "any or both" of the elements so conjoined, i.e., elements that may exist conjunctively in some cases and disjunctively in others. Multiple elements listed together with "and / or" are to be construed in the same fashion, i.e., as "one or more" of the elements so conjoined. Other elements may optionally exist in addition to those specifically identified by the "and / or" clause, whether related or unrelated to the elements specifically identified thereby. Thus, by way of non-limiting example, a reference to "A and / or B" when used with open-ended terms such as "comprising" may refer, in one embodiment, to only A (optionally including elements other than B), in another embodiment, to only B (optionally including elements other than A), and in yet another embodiment, to both A and B (optionally including other elements).
[0358] The use of ordinal terms such as "first", "second", "third", etc. in the claims to modify a claim element does not by itself imply any priority, precedence, or order of one claim element over another claim element, or the temporal order in which acts of a method are to be performed. Such terms are used merely as labels to distinguish one claim element having a certain name from another element having the same name (except for the use of the ordinal term).
[0359] The syntax and terminology used herein are for the purpose of description and should not be regarded as limiting. The use of "including", "comprising", "having", "containing", "involving", and variations thereof is meant to encompass the items listed thereafter and additional items.
[0360] Although some embodiments of the techniques described in this specification have been described in detail, various modifications and improvements will be readily envisioned by those skilled in the art. Such modifications and improvements are intended to be within the spirit and scope of this disclosure. Accordingly, the above description is by way of example only and not limiting. These techniques are limited only as defined by the following claims and their equivalents.
Explanation of Reference Numerals
[0361] 100, 200, 2600, 2700, 2800 exemplary processing pipelines, processing pipelines 102 gene expression data, expression data, first expression data 106a gene set 1 106b gene set 2 106c gene set 3 106d gene set 4 108 ranking process, rank process 110 gene ranking 110a first gene ranking, gene ranking, gene ranking 1 110b second gene ranking, gene ranking, gene ranking 2 110c gene ranking 3 110d gene ranking 4 112 statistical model 112a statistical model, statistical model 1 112b statistical model, statistical model 2 112c statistical model 3 112d statistical model 4 114 characteristics, PTCL subtype 114a characteristic 1 114b characteristic 2 116a characteristic prediction 1 116b characteristic prediction 2 116c characteristic prediction 3 116d characteristic prediction 4 118 prediction analysis process 214 PTCL subtype 216a, 316a, 416a PTCL Subtype Prediction 1 216b, 316b, 416b PTCL Subtype Prediction 2 216c, 316c, 416c PTCL Subtype Prediction 3 216d, 316d, 416d PTCL Subtype Prediction 4 300, 400, 500, 800, 900, 2900 Exemplary Processes, Processes 1000 Computer System 1010 Processor 1020 Memory 1030 Non-volatile Memory Medium, Non-volatile Memory Device 1040 Network Input / Output (I / O) Interface 1050 User I / O Interface 1100 Exemplary Environment 1102 Patient 1104 Laboratory 1106 Database 1108 Network 1110 Server 1112 Computing Device 1114 Physician 1116 External Database, Database
Claims
1. A computer-implemented method comprising: using at least one computer hardware processor to obtain expression data that is at least partially obtained by sequencing a biological sample of a subject having cancer, suspected of having cancer, or at risk of having cancer, wherein the expression data comprises expression levels for a plurality of genes, and the plurality of genes constitutes a set of genes; ranking at least some of the genes in the set of genes based on their expression levels in the expression data to obtain a gene ranking, wherein the gene ranking includes values that specify relative ranks for the genes in the gene ranking; using the gene ranking and a statistical model trained using training data indicating a plurality of gene rankings for at least some of the genes in the obtained set of genes to determine at least one characteristic of the biological sample, wherein each of the plurality of gene rankings is obtained based on respective expression levels of the at least some of the genes in the set of genes, and the step of determining the at least one characteristic includes providing the gene ranking as an input to the statistical model and obtaining an output indicating the at least one characteristic; performing the steps of a computer-implemented method.
2. The method of claim 1, wherein the at least one characteristic of the biological sample is a physiological characteristic of cells in the biological sample or tissue from which the cells are derived, and / or wherein the at least one characteristic is selected from the group consisting of a cancer grade for cells in the biological sample, an origin tissue for cells in the biological sample, a tissue type for cells in the biological sample, and a cancer subtype for cells in the biological sample.
3. The method of claim 1 or 2, wherein the at least one characteristic includes a cancer grade for cells in the biological sample, and the cancer grade is selected from the group consisting of grade 1, grade 2, grade 3, grade 4, and grade 5.
4. The method according to any one of claims 1 to 3, wherein the subject has breast cancer, is suspected of having breast cancer, or is at risk of having breast cancer, and / or the set of genes is selected from the gene groups described in Table 1.
5. The method according to any one of claims 1 to 4, wherein the subject has clear cell renal carcinoma, is suspected of having clear cell renal carcinoma, or is at risk of having clear cell renal carcinoma, and / or the set of genes is selected from the gene groups described in Table 2.
6. The method according to any one of claims 1 to 5, wherein the subject has lymphoma, is suspected of having lymphoma, or is at risk of having lymphoma, and / or the set of genes is selected from the gene groups described in Table 3.
7. The method according to any one of claims 1 to 6, wherein the subject has lung adenocarcinoma, is suspected of having lung adenocarcinoma, or is at risk of having lung adenocarcinoma, and / or the set of genes is selected from the gene groups described in Table 6.
8. The method according to any one of claims 1 to 7, wherein the at least one characteristic includes the human papillomavirus status for cells in the biological sample, the set of genes includes at least five genes selected from the gene groups described in Table 8, and / or the subject has squamous cell carcinoma of the head and neck, is suspected of having squamous cell carcinoma of the head and neck, or is at risk of having squamous cell carcinoma of the head and neck, and the set of genes is selected from the gene groups described in Table 8.
9. The method according to any one of claims 1 to 8, wherein the statistical model comprises a gradient boosting decision tree classifier.
10. Further comprising the step of presenting an indication of the at least one characteristic to the user, The step of presenting the indication of the at least one characteristic further comprises the step of displaying the at least one characteristic to the user in a graphical user interface (GUI). The method according to claim 1.
11. The at least one characteristic includes the origin tissue for cells in the biological sample, The method according to any one of claims 1 to 10, wherein the origin tissue is selected from the group consisting of lung tissue, pancreatic tissue, gastric tissue, colon tissue, liver tissue, bladder tissue, kidney tissue, thyroid tissue, lymph node tissue, adrenal tissue, skin tissue, breast tissue, ovarian tissue, prostate tissue, urothelial tissue, cervical tissue, esophageal tissue, brain tissue, soft tissue, connective tissue, head tissue, and neck tissue.
12. wherein the at least one characteristic includes an origin tissue for cells in the biological sample and a tissue type for cells in the biological sample, and the combination of the origin tissue and the tissue type is selected from the group consisting of lung adenocarcinoma, lung squamous cell carcinoma, melanoma, breast cancer, colorectal adenocarcinoma, ovarian serous cystadenocarcinoma, pheochromocytoma, bladder urothelial carcinoma, cervical squamous cell carcinoma, glioblastoma multiforme, head squamous cell carcinoma, neck squamous cell carcinoma, renal clear cell carcinoma, renal papillary cell carcinoma, liver hepatocellular carcinoma, pancreatic adenocarcinoma, paraganglioma, prostate adenocarcinoma, sarcoma, gastric adenocarcinoma, thyroid cancer, and uterine corpus endometrial carcinoma, the method according to any one of claims 1 to 11.
13. The method according to any one of claims 1 to 12, wherein the subject has, is suspected of having, or is at risk of having lymphoma, optionally diffuse large B-cell lymphoma (DLBCL), and / or the at least one characteristic includes origin cells selected from the group consisting of germinal center B cells (GCB) and activated B cells (ABC).
14. The method according to any one of claims 1 to 13, wherein the statistical model comprises a plurality of classifiers corresponding to different subtypes of PTCL, and / or the set of genes of at least one set of the genes is selected from the gene groups described in Table 10.
15. The method according to any one of claims 1 to 14, wherein the statistical model comprises a plurality of classifiers corresponding to different subtypes of PTCL, the plurality of classifiers includes a first classifier, a second classifier, a third classifier, and a fourth classifier, the first classifier corresponding to anaplastic large cell lymphoma (ALCL), the second classifier corresponding to angioimmunoblastic T-cell lymphoma (AITL), the third classifier corresponding to natural killer / T-cell lymphoma (NKTCL), and the fourth classifier corresponding to adult T-cell leukemia / lymphoma (ATLL).
16. A system, at least one computer hardware processor, and at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to perform the method according to any one of claims 1 to 15 A system comprising: **Claim 17** At least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to perform the method according to any one of claims 1 to 15.
Citation Information
Patent Citations
JP1038350005A
Methods and devices for identifying patterns in biological systems and methods of use thereof
JP2003529131A
Methods, systems, and arrays for classifying, prognosing, and diagnosing cancers based on the association between p53 status and gene expression profiles
JP2008521383A
US10,311,967
Method for feature selection in a support vector machine using feature ranking
US20080233576A1