Disease detection systems and methods
The method for manufacturing a myelofibrosis-related detection system effectively segments myelofibrosis abnormalities by clustering transcriptome data, enabling the identification of genetic pathways and specific genes associated with each subcategory, thereby facilitating the development of targeted therapeutics.
Patent Information
- Application Number
- PCT/US2024/061140
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-12-19
- Publication Date
- 2025-06-26
AI Technical Summary
Current technologies lack effective methods for reliably segmenting myelofibrosis abnormalities into meaningful subcategories based on transcriptome expression differences, which hinders the identification of genetic pathways and specific genes associated with each subcategory, thereby limiting the development of targeted therapeutics.
A method for manufacturing a myelofibrosis-related detection system that involves obtaining single-nucleus or single-cell transcriptome data for a cohort of subjects, clustering the data to identify distinct subsets of nuclei or cells, and using metadata to determine the myelofibrosis-related state of each subject, thereby enabling the segmentation of myelofibrosis abnormalities into subcategories.
This approach allows for the reliable identification of genetic pathways and specific genes associated with different myelofibrosis subcategories, providing a basis for developing targeted therapeutics that address specific myelofibrosis states.
Smart Images

Figure US2024061140_26062025_PF_FP_ABST
Abstract
Description
DISEASE DETECTION SYSTEMS AND METHODSTECHNICAL FIELD
[0001] The present disclosure is directed to systems, methods, and devices for manufacturing myelofibrosis disease related detection systems.CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to and benefit of U.S. Provisional Patent Application No. 63 / 612,321, filed on December 19, 2023, the content of which is hereby incorporated by reference in its entirety.BACKGROUND
[0003] Myelofibrosis is a rare chronic myeloproliferative neoplasm characterized by progressive bone marrow fibrosis, and inefficient hematopoiesis. The myelofibrosis- associated consequences and medical complications often result in premature death from infection, thrombohemorrhagic events, cardiac or pulmonary failure, and leukemic transformation. The pathogenesis of myelofibrosis is multifactorial, involving multiple cell types, and key mechanisms of the disease and disease progression remains incompletely understood.
[0004] Given the above background, what is needed in the art are myelofibrosis disease related detection systems, based on improved myelofibrosis atlases, that can be used to reliably segment myelofibrosis abnormalities into meaningful subcategories based on relevant and actual differences in transcriptome expression between the different subcategories. Such segmentation can be used to identify genetic pathways and particular genes that are associated with each such subcategory and thus be used as new basis for developing therapeutics that address particular myelofibrosis states.SUMMARY
[0005] The present disclosure addresses the above-identified shortcomings. One aspect of the present disclosure provides a method for manufacturing a myelofibrosis-related detection system. The method comprises, at a computer system having one or more processors, and memory storing one or more programs for execution by the one or more processors, obtaining, in electronic form, first information. The first information comprises single-nucleus or single-cell transcriptome data for a plurality of genes for each nucleus or cell in a first plurality of nuclei or cells. Each nucleus or cell in the first plurality of nuclei or cells is obtained from a sample in a plurality of samples. Each sample in the plurality of samples is from a different subject in a cohort of subjects. The first plurality of nuclei or cells includes a different subset of nuclei or cells from a sample from each subject in the cohort of subjects. In some embodiments, the first plurality of nuclei or cells comprises at least 1000 nuclei or cells. In some embodiments, the first plurality of nuclei or cells comprises 2 or more nuclei or cells.
[0006] The first information further comprises first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has a first myelofibrosis-related state or a second myelofibrosis-related state. At least a first subset of subjects in the cohort of subjects have the first myelofibrosis-related state and a second subset of subjects in the cohort of subjects have the second myelofibrosis-related state.
[0007] The respective single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is optionally barcoded with the subject in the cohort of subjects originating the respective single-nucleus or singlecell transcriptome data.
[0008] The first plurality of nuclei or cells is clustered into a plurality of clusters by (i) computing a plurality of distances using the single-nucleus or single-cell transcriptome data for the plurality of genes for each unique pair of nuclei or cells in the first plurality of nuclei or cells and (ii) evaluating the plurality of distances with a criterion function. The plurality of distances includes a separate distance for each unique pair of nuclei or cells in the first plurality of nuclei or cells. Each respective distance in the plurality of distances represents a different pair of nuclei or cells in the first plurality of nuclei or cells and quantifies a distance between (i) a respective first vector formed by the single-nucleus or single-cell transcriptome data for the plurality of genes for a respective first nucleus or first cell in the different pair of nuclei and (ii) a respective second vector formed by the single-nucleus or single-cell transcriptome data for the plurality of genes for a respective second nucleus or second cell in the different pair of nuclei or cells. Each respective cluster in the plurality of clusters represents a corresponding subset of nuclei or cells of the first plurality of nuclei or cells that are clustered together based on evaluation of distances in the plurality of distances representing different pairs of nuclei or cells within the corresponding subset of nuclei or cells with the criterion function.
[0009] The first metadata is used to identify a first cluster in the plurality of clusters with the first myelofibrosis-related state by determining that the first cluster includes nuclei or cells from subjects in the cohort of subjects that have the first myelofibrosis-related state.
[0010] In some embodiments, the first metadata for each respective subject in the cohort of subjects indicates at least for each respective subject in the cohort of subjects whether the respective subject has the first myelofibrosis-related state or the second myelofibrosis-related state comprises a histologically graded disease status for the respective subject.
[0011] In some embodiments, the first metadata for each respective subject in the cohort of subjects indicates at least for each respective subject in the cohort of subjects whether the respective subject has the first myelofibrosis-related state or the second myelofibrosis-related state comprises a histologically graded disease status for the respective subject determined in accordance with a consistent, verified handling of each sample in the plurality of samples.
[0012] In some embodiments, each sample in the plurality of samples is from a different subject in the cohort of subjects and the plurality of samples comprises 5 samples, 20 samples, 50 samples, or 100 or more samples.
[0013] In some embodiments, the first metadata for each respective subject in the cohort of subjects further indicates one or more features of the subject selected from the group consisting of sex, race, age, and physical condition. In some embodiments the method further comprises using the first cluster to determine an extent to which race is a covariate with respect to the first myelofibrosis-related state. In some embodiments, the method further comprises using the first cluster to determine an extent to which sex is a covariate with respect to the first myelofibrosis-related state. In some embodiments, the method further comprises using the first cluster to determine an extent to which age is a covariate with respect to the first myelofibrosis-related state.
[0014] In some embodiments, sex, race, age, and / or physical condition are each represented in a balanced manner in the cohort of subjects.
[0015] In some embodiments, the method further comprises filtering the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells to remove counts of ambient RNA molecules, doublets and / or empty droplets.
[0016] In some embodiments, the first metadata includes one or more first biomarker annotations for each nucleus or cell in the plurality of nuclei or cells drawn from a first standardized set of biomarkers that each define a unique cell type.
[0017] In some embodiments, the method further comprises using the first cluster to determine an extent to which a first biomarker in the first standardized set of biomarkers is a covariate with respect to the first myelofibrosis-related state.
[0018] In some embodiments, the first metadata includes one or more second biomarker annotations for each nucleus or cell in the plurality of nuclei or cells drawn from a second standardized set of biomarkers that each define a unique cell type at a unique stage.
[0019] In some embodiments, the method further comprises using the first cluster to determine an extent to which a second biomarker in the second standardized set of biomarkers is a covariate with respect to the first myelofibrosis-related state.
[0020] In some embodiments, the method further comprises obtaining genotype data for each subject in the plurality of subjects and overlaying genotype data for each subject represented in the first cluster with a myelofibrosis-related state of each subject in the first cluster.
[0021] In some embodiments, the method further comprises obtaining genotype data for each subject in the plurality of subjects, and determining an extent to which a genotype is a covariate for the first myelofibrosis-related state using the genotype data for each subject represented in the first cluster.
[0022] In some embodiments, the method further comprising using the myelofibrosis abnormality detection system to associate a test subject with the first myelofibrosis-related state by a procedure comprising: obtaining, in electronic form, second information comprising single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in a second plurality of nuclei or cells, wherein each nucleus or cell in the second plurality of nuclei or cells is obtained from a sample obtained from the test subject. The first plurality of nuclei or cells and the second plurality of nuclei or cells are co-clustered into the plurality of clusters. The test subject is identified as having the first myelofibrosis- related state when nuclei or cells from the second plurality of nuclei or cells co-cluster into the first cluster.
[0023] In some embodiments, the first myelofibrosis-related state is myelofibrosis and the second myelofibrosis-related state is absence of myelofibrosis.
[0024] In some embodiments, the first myelofibrosis-related state is absence of myelofibrosis and the second myelofibrosis-related state is myelofibrosis.
[0025] In some embodiments, the first myelofibrosis-related state is a first stage of myelofibrosis and the second myelofibrosis-related state is a second stage of myelofibrosis.
[0026] In some embodiments, the first myelofibrosis-related state is low risk myelofibrosis and the second myelofibrosis-related state is intermediate- 1 risk myelofibrosis.
[0027] In some embodiments, the first myelofibrosis-related state is intermediate- 1 risk myelofibrosis and the second myelofibrosis-related state is intermediate-2 risk myelofibrosis.
[0028] In some embodiments, the first myelofibrosis-related state is intermediate-2 risk myelofibrosis and the second myelofibrosis-related state is high risk myelofibrosis.
[0029] In some embodiments, the first myelofibrosis-related state is low risk or intermediate- 1 risk myelofibrosis and the second myelofibrosis-related state is intermediate-2 risk myelofibrosis or high risk myelofibrosis.
[0030] In some embodiments, the first myelofibrosis-related state is low risk myelofibrosis and the second myelofibrosis-related state is absence of myelofibrosis.
[0031] In some embodiments, the first myelofibrosis-related state is low risk myelofibrosis and the second myelofibrosis-related state is intermediate- 1 risk or intermediate-2 risk myelofibrosis.
[0032] In some embodiments, the first myelofibrosis-related state is a disease-treatable or disease-reversible state and the second myelofibrosis-related state is an undiseased or prediseased state.
[0033] In some embodiments, the first myelofibrosis-related state is a disease-untreatable or disease-irreversible state and the second myelofibrosis-related state is an undiseased or prediseased state.
[0034] In some embodiments, the first myelofibrosis-related state is a first disease-treatable or disease-reversible state and the second myelofibrosis-related state is a second disease- treatable or disease-reversible state.
[0035] In some embodiments, the first myelofibrosis-related state is a disease-untreatable or disease-irreversible state and the second myelofibrosis-related state is a disease-treatable or disease-reversible state.
[0036] In some embodiments, the first myelofibrosis-related state is an abnormal hematopoietic state and the second myelofibrosis-related state is a normal hematopoietic state.
[0037] In some embodiments, the first myelofibrosis-related state is a decreased red blood cell and / or platelet count state and the second myelofibrosis-related state is a normal red blood cell and / or platelet count state.
[0038] In some embodiments, the first myelofibrosis-related state is a bone marrow fibrosis state and the second myelofibrosis-related state is a healthy bone marrow state.
[0039] In some embodiments, the first myelofibrosis-related state is leukemic blood cell state and the second myelofibrosis-related state is healthy blood cell state.
[0040] In some embodiments, the first information further comprises second metadata for each subject in the cohort of subjects comprising age, body mass index, or sex. Further, in such embodiments, the method further comprises using the second information to prune the cohort of subjects based on age, body mass index, or sex thereby causing the cohort of subjects to be free of confounding for age, body mass index, or sex, wherein the pruning causes a subset of subjects to be removed from the cohort of subjects prior to the clustering.
[0041] In some embodiments, the first information further comprises second metadata for each subject in the cohort of subjects comprising age, body mass index, and sex, and the method further comprises using the second information to prune the cohort of subjects based on age, body mass index, and sex thereby causing the cohort of subjects to be free of confounding for age, body mass index, and sex, wherein the pruning causes a subset of subjects to be removed from the cohort of subjects prior to the clustering.
[0042] In some embodiments, the method further comprises using the first metadata to identify a second cluster in the plurality of clusters with the second myelofibrosis-related state by determining that the second cluster includes nuclei or cells from subjects in the cohort that have the second myelofibrosis-related state.
[0043] In some such embodiments, the first myelofibrosis-related state is a first stage of myelofibrosis and the second myelofibrosis-related state is a second stage of myelofibrosis.
[0044] In some such embodiments, the method further determines that the first cluster comprises basophil / mast cells or nuclei of basophil / mast cells and the second cluster comprises cells or nuclei of cells other than basophil / mast cells.
[0045] In some such embodiments, the method further determines that the first cluster comprises monocytes or nuclei of monocytes and the second cluster comprises cells or nuclei of cells other than monocytes.
[0046] In some such embodiments, the method further determines that the first cluster comprises erythroid lineage cells or nuclei of erythroid lineage cells and the second cluster comprises cells or nuclei of cells other than erythroid lineage cells.
[0047] In some such embodiments the method further determines that the first cluster comprises hematopoietic precursor cells or nuclei of hematopoietic precursor cells and the second cluster comprises cells or nuclei of cells other than hematopoietic precursor cells.
[0048] In some such embodiments, the method further determines that the first cluster comprises lymphoid lineage cells or nuclei of lymphoid lineage cells and the second cluster comprises cells or nuclei of cells other than lymphoid lineage cells.
[0049] In some such embodiments, the method further determines that the first cluster comprises megakaryocyte-erythroid progenitor cells or nuclei of megakaryocyte-erythroid progenitor cells and the second cluster comprises cells or nuclei of cells other than megakaryocyte-erythroid progenitor cells.
[0050] In some such embodiments, the method further determines that the first cluster comprises megakaryocyte cells or nuclei of megakaryocyte cells and the second cluster comprises cells or nuclei of cells other than megakaryocyte cells.
[0051] In some such embodiments, the method further determines that the first cluster comprises myeloid lineage cells or nuclei of myeloid lineage cells and the second cluster comprises cells or nuclei of cells other than myeloid lineage cells.
[0052] In some embodiments, the plurality of genes comprises 100 or more genes, 200 or more genes, 500 or more genes, 1000 or more genes, 2000 or more genes, 5000 or more genes, 10000 or more genes, or 15,000 or more genes.
[0053] In some embodiments, the first plurality of nuclei or cells comprises 10,000 or more nuclei or cells, 50,000 or more nuclei or cells, 100,000 or more nuclei or cells, 250,000 or more nuclei or cells, 500,000 or more nuclei or cells, 600,000 or more nuclei or cells, or 1 x 106or more nuclei or cells.
[0054] In some embodiments, the cohort of subjects comprises 25 or more subjects, 50 or more subjects, 75 or more subjects, or 100 or more subjects.
[0055] In some embodiments, the first subset of subjects comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects and the second subset of subjects is other than the first subset of subjects and comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects.
[0056] In some embodiments, the first myelofibrosis-related state or the second myelofibrosis-related state is associated with osteosclerosis, extramedullary hematopoiesis (EMH), inefficient hematopoiesis, inflammation, splenomegaly, cytopenia, portal hypertension, thromboembolism, infection, or acute myeloid leukemia (AML).
[0057] In some embodiments, the method informs a response to a drug compound in a patient or in a plurality of patients.
[0058] In some embodiments, the method informs a response to a dosing amount, duration, and / or frequency of a drug in a patient or in a plurality of patients.
[0059] In some embodiments, the first cluster represents a myelofibrosis disease state and the second cluster represents a healthy state, and the method further comprises identifying a metabolic pathway comprising a set of genes, wherein one or more genes in the set of genes are overexpressed or under-expressed in the first cluster relative to the second cluster.
[0060] In some embodiments, the first cluster represents a myelofibrosis disease state and the second cluster represents a healthy state, and the method further comprises accessing, in electronic form, a plurality of compound-specific differential transcriptional signatures, wherein each respective compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures is a difference between (i) a respective compound-treated transcriptional signature in a plurality of compound-treated transcriptional signatures and (ii) a baseline transcriptional signature data set, wherein the baseline transcriptional signature data set is from a control sample of one or more cells of a cell type, and each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures is from a corresponding compound-treated sample of one or more cells of the cell type separately treated with a different compound in a plurality of at least 10 compounds. A test differential transcription signature is determined by differential comparison of a transcriptional signature of the nuclei or cells of the first cluster and the second cluster. The test differential transcription signature is compared to each respective compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures, thereby identifying a first compound associated with a compound-specific differential transcriptional signature in theplurality of compound-specific differential transcriptional signatures that matches the test differential transcription signature.
[0061] In some such embodiments, the myelofibrosis disease state is selected from low risk, intermediate- 1 risk, intermediate-2 risk, and high risk myelofibrosis.
[0062] In some embodiments, the control sample and each corresponding compound-treated sample is exposed to a solvent, wherein the solvent is the same solvent for the control sample and each corresponding compound-treated sample, optionally wherein the solvent comprises dimethyl sulfoxide (DMSO).
[0063] In some embodiments, the control sample and each corresponding compound-treated sample each comprises a polar aprotic solvent, wherein the polar aprotic solvent is the same polar aprotic solvent for the control sample and each corresponding compound-treated sample.
[0064] In some embodiments, the control sample and each corresponding compound-treated sample comprises DMSO.
[0065] In some embodiments, each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nucleus assay and / or single-cell assay data. Optionally, the single-nucleus assay and / or single-cell assay data is selected from single-cell ribonucleic acid (RNA) sequencing (scRNA-seq) data, single-nucleus RNA sequencing (snRNA-seq) data, scTag-seq data, single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data, CyTOF / SCoP data, E-MS / Abseq data, miRNA-seq data, CITE-seq data, or any combinations thereof.
[0066] In some embodiments, each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises or consists of singlecell RNA sequencing (scRNA-seq) data.
[0067] In some embodiments, each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nuclei or singlecell RNA sequencing (scRNA-seq) data.
[0068] In some embodiments, each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures consists of single-nuclei or single-cell RNA sequencing (scRNA-seq) data.
[0069] In some embodiments, each corresponding compound-treated sample of one or more cells comprises basophils / mast cells, CD14+ cells, monocytes, erythroid lineage cells,hematopoietic precursor cells, lymphoid lineage cells, megakaryocyte-erythroid progenitor cells, megakaryocytes, or myeloid lineage cells.
[0070] In some embodiments, each corresponding compound-treated sample of one or more cells comprises basophils / mast cells.
[0071] In some embodiments, each corresponding compound-treated sample of one or more cells comprises monocytes.
[0072] In some embodiments, each corresponding compound-treated sample of one or more cells consists of basophils / mast cells.
[0073] In some embodiments, each corresponding compound-treated sample of one or more cells consists of monocytes.
[0074] In some embodiments, each corresponding compound-treated sample of one or more cells consists of erythroid lineage cells.
[0075] In some embodiments, each corresponding compound-treated sample of one or more cells comprises hematopoietic precursor cells.
[0076] In some embodiments, each corresponding compound-treated sample of one or more cells comprises lymphoid lineage cells.
[0077] In some embodiments, each corresponding compound-treated sample of one or more cells comprises megakaryocyte-erythroid progenitor cells.
[0078] In some embodiments, each corresponding compound-treated sample of one or more cells consists of macrophages.
[0079] In some embodiments, each corresponding compound-treated sample of one or more cells consists of megakaryocytes.
[0080] In some embodiments, each corresponding compound-treated sample of one or more cells consists of myeloid lineage cells.
[0081] In some embodiments, each corresponding compound-treated sample of one or more cells comprises or consists of one or more cells from an organ, cells from a tissue, stem cells, human cells, cells from umbilical cord blood, cells from peripheral blood, bone marrow cells, cells from a solid tissue, differentiated cells and / or cells from a cell line.
[0082] In some embodiments, each corresponding compound-treated sample of one or more cells is a frozen sample.
[0083] In some embodiments, each respective compound in the plurality of compounds has a molecular weight of less than 2000 Daltons.
[0084] In some embodiments, each respective compound in the plurality of compounds satisfies at least three criteria of the Lipinski rule of five criteria, or optionally each of the Lipinski rule of five criteria.
[0085] Another aspect of the present disclosure provides a method for identifying a compound that transitions a myelofibrosis state to a healthy state. The method in accordance with this aspect of the present disclosure comprises obtaining first information. The first information comprises single-nucleus or single-cell transcriptome data for a plurality of genes for each nucleus or cell in a first plurality of nuclei or cells. Each nucleus or cell in the first plurality of nuclei or cells is obtained from a sample in a plurality of samples from a cohort of subjects. The first information further comprises first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has a myelofibrosis disease state or a healthy state. At least a first subset of subjects in the cohort of subjects have the myelofibrosis disease state and a second subset of subjects in the cohort of subjects have the healthy state. The respective single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is barcoded with the subject in the cohort of subjects originating the respective single-nucleus or single-cell transcriptome data.
[0086] The first plurality of nuclei or cells are clustered into a plurality of clusters based on the single-nucleus or single-cell transcriptome data for the plurality of genes for each pair of nuclei or cells in the first plurality of nuclei or cells. The first information is used to identify a first cluster in the plurality of clusters with the myelofibrosis disease state by determining that the first cluster includes nuclei or cells from subjects in the cohort of subjects that have the myelofibrosis disease state. The first information is used to identify a second cluster in the plurality of clusters with the healthy state by determining that the second cluster includes nuclei or cells from subjects in the cohort that have the healthy state.
[0087] A plurality of compound-specific differential transcriptional signatures is accessed. Each respective compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures is a difference between (i) a respective compound-treated transcriptional signature in a plurality of compound-treated transcriptional signatures and (ii) a baseline transcriptional signature data set. The baseline transcriptional signature data set is from a control sample of one or more cells of a cell type.Each respective compound-treated transcriptional signature in the plurality of compound- treated transcriptional signatures is from a corresponding compound-treated sample of one or more cells of the cell type separately treated with a different compound in a plurality of compounds.
[0088] A test differential transcription signature is determined by differential comparison of a transcriptional signature of the nuclei or cells of the first cluster and the second cluster.
[0089] The test differential transcription signature is compared to each respective compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures, thereby identifying a first compound associated with a compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures that matches the test differential transcription signature.
[0090] In some embodiments, the first metadata for each respective subject in the cohort of subjects, which indicates at least for each respective subject in the cohort of subjects whether the respective subject has the myelofibrosis disease state or the healthy state, comprises a histologically graded disease status for the respective subject.
[0091] In some embodiments, the first metadata for each respective subject in the cohort of subjects, which indicates at least for each respective subject in the cohort of subjects whether the respective subject has the myelofibrosis disease state or the healthy state, comprises a histologically graded disease status for the respective subject determined in accordance with a consistent, verified handling of each sample in the plurality of samples.
[0092] In some embodiments, each sample in the plurality of samples is from a different subject in the cohort of subjects and the plurality of samples comprises 5 samples, 20 samples, 50 samples, or 100 or more samples.
[0093] In some embodiments, the first metadata for each respective subject in the cohort of subjects further indicates one or more features of the subject selected from the group consisting of sex, race, age, and physical condition.
[0094] In some embodiments, the method further comprises filtering the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells to remove counts of ambient RNA molecules, doublets and / or empty droplets.
[0095] In some embodiments, the myelofibrosis state is low risk myelofibrosis intermediate- 1 risk myelofibrosis, intermediate-2 risk myelofibrosis, or high risk myelofibrosis.
[0096] In some embodiments, the method further determines that the first cluster comprises basophil / mast cells or nuclei of basophil / mast cells and the second cluster comprises cells or nuclei of cells other than basophil / mast cells.
[0097] In some embodiments, the method further determines that the first cluster comprises monocytes or nuclei of monocytes and the second cluster comprises cells or nuclei of cells other than monocytes.
[0098] In some embodiments, the method further determines that the first cluster comprises erythroid lineage cells or nuclei of erythroid lineage cells and the second cluster comprises cells or nuclei of cells other than erythroid lineage cells.
[0099] In some embodiments, the method further determines that the first cluster comprises hematopoietic precursor cells or nuclei of hematopoietic precursor cells and the second cluster comprises cells or nuclei of cells other than hematopoietic precursor cells.
[0100] In some embodiments, the method further determines that the first cluster comprises lymphoid lineage cells or nuclei of lymphoid lineage cells and the second cluster comprises cells or nuclei of cells other than lymphoid lineage cells.
[0101] In some embodiments, the method further determines that the first cluster comprises megakaryocyte-erythroid progenitor cells or nuclei of megakaryocyte-erythroid progenitor cells and the second cluster comprises megakaryocyte myeloid lineage cells or nuclei of megakaryocyte myeloid lineage cells.
[0102] In some embodiments, the plurality of genes comprises 100 or more genes, 200 or more genes, 500 or more genes, 1000 or more genes, 2000 or more genes, 5000 or more genes, 10000 or more genes, or 15,000 or more genes.
[0103] In some embodiments, the first plurality of nuclei or cells comprises 10,000 or more nuclei or cells, 50,000 or more nuclei or cells, 100,000 or more nuclei or cells, 250,000 or more nuclei or cells, 500,000 or more nuclei or cells, 600,000 or more nuclei or cells, or 1 x 106or more nuclei or cells.
[0104] In some embodiments, the cohort of subjects comprises 25 or more subjects, 50 or more subjects, 75 or more subjects, or 100 or more subjects.
[0105] In some embodiments, the first subset of subjects comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects and the second subset of subjects is other than the first subset of subjects and comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects.
[0106] In some embodiments, the myelofibrosis disease state is associated with osteosclerosis, extramedullary hematopoiesis (EMH), inefficient hematopoiesis, inflammation, splenomegaly, cytopenia, portal hypertension, thromboembolism, infection, or acute myeloid leukemia (AML).
[0107] In some embodiments, the method informs a response to a drug compound in a patient or in a plurality of patients.
[0108] In some embodiments, the method informs a response to a dosing amount, duration, and / or frequency of a drug in a patient or in a plurality of patients.
[0109] In some embodiments, the method further comprises identifying a metabolic pathway comprising a set of genes, wherein one or more genes in the set of genes are overexpressed or under-expressed in the first cluster relative to the second cluster.
[0110] In some embodiments, the control sample and each corresponding compound-treated sample is exposed to a solvent, wherein the solvent is the same solvent for the control sample and each corresponding compound-treated sample, optionally wherein the solvent comprises dimethyl sulfoxide (DMSO).
[0111] In some embodiments, the control sample and each corresponding compound-treated sample each comprises a polar aprotic solvent, wherein the polar aprotic solvent is the same polar aprotic solvent for the control sample and each corresponding compound-treated sample.
[0112] In some embodiments, the control sample and each corresponding compound-treated sample comprises DMSO.
[0113] In some embodiments, each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nucleus assay and / or single-cell assay data, optionally wherein the single — nucleus-assay and / or single-cell assay data is selected from single-cell ribonucleic acid (RNA) sequencing (scRNA-seq) data, single-nucleus RNA sequencing (snRNA-seq) data, scTag-seq data, single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data, CyTOF / SCoP data, E-MS / Abseq data, miRNA-seq data, CITE-seq data, or any combinations thereof.
[0114] In some embodiments, each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises or consists of single-cell RNA sequencing (scRNA-seq) data.
[0115] In some embodiments, each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nucleus RNA sequencing (scRNA-seq) data.
[0116] In some embodiments, each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures consists of single-nucleus RNA sequencing (scRNA-seq) data.
[0117] In some embodiments, each corresponding compound-treated sample of one or more cells comprises basophils / mast cells, monocytes, CD14+ cells, erythroid lineage cells, hematopoietic precursor cells, lymphoid lineage cells, megakaryocyte-erythroid progenitor cells, megakaryocytes, or myeloid lineage cells.
[0118] In some embodiments, each corresponding compound-treated sample of one or more cells comprises basophils / mast cells.
[0119] In some embodiments, each corresponding compound-treated sample of one or more cells comprises monocytes.
[0120] In some embodiments, each corresponding compound-treated sample of one or more cells consists of basophils / mast cells.
[0121] In some embodiments, each corresponding compound-treated sample of one or more cells consists of monocytes.
[0122] In some embodiments, each corresponding compound-treated sample of one or more cells consists of erythroid lineage cells.
[0123] In some embodiments, each corresponding compound-treated sample of one or more cells comprises hematopoietic precursor cells.
[0124] In some embodiments, each corresponding compound-treated sample of one or more cells comprises lymphoid lineage cells.
[0125] In some embodiments, each corresponding compound-treated sample of one or more cells comprises megakaryocyte-erythroid progenitor cells.
[0126] In some embodiments, each corresponding compound-treated sample of one or more cells consists of macrophages.
[0127] In some embodiments, each corresponding compound-treated sample of one or more cells consists of megakaryocytes.
[0128] In some embodiments, each corresponding compound-treated sample of one or more cells consists of myeloid lineage cells.
[0129] In some embodiments, each corresponding compound-treated sample of one or more cells comprises or consists of one or more cells from an organ, cells from a tissue, stem cells, human cells, cells from umbilical cord blood, cells from peripheral blood, bone marrow cells, cells from a solid tissue, differentiated cells and / or cells from a cell line.
[0130] In some embodiments, each respective compound in the plurality of compounds has a molecular weight of less than 2000 Daltons.
[0131] In some embodiments, each respective compound in the plurality of compounds satisfies at least three criteria of the Lipinski rule of five criteria, or optionally each of the Lipinski rule of five criteria.
[0132] Another aspect of the present disclosure provides a computer system having one or more processors, and memory storing one or more programs for execution by the one or more processors, the one or more programs comprising instructions for performing any of the methods and / or embodiments disclosed herein.
[0133] Another aspect of the present disclosure provides a non-transitory computer readable storage medium storing one or more programs configured for execution by a computer, the one or more programs comprising instructions for carrying out any of the methods and / or embodiments disclosed herein.
[0134] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.BRIEF DESCRIPTION OF THE DRAWINGS
[0135] The embodiments disclosed herein are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings. Like reference numerals refer to corresponding parts throughout the drawings.
[0136] Figure 1 illustrates a block diagram of an exemplary system and computing device for manufacturing a myelofibrosis-related detection system, in accordance with an embodiment of the present disclosure.
[0137] Figures 2A, 2B, 2C, 2D, 2E, 2F, 2G, 2H, 21, 2J, 2K, and 2L collectively provide a flow chart of processes and features of an example method manufacturing a myelofibrosis- related detection system, in accordance with various embodiments of the present disclosure.DETAILED DESCRIPTION
[0138] Given the above background, the present disclosure provides systems and methods for manufacturing a myelofibrosis-related detection system that obtains single-nucleus or single-cell transcriptome data for a plurality of genes for each of a plurality of nuclei or cells. Each nucleus or cell is obtained from a sample. Each sample is from a different subject in a cohort of subjects. Metadata for each subject in the cohort is also obtained, indicating for each respective subject whether they have a first or second myelofibrosis-related state. The transcriptome data for the plurality of genes for each nucleus or cell is barcoded with the corresponding subject in the cohort. The nuclei or cells are clustered into clusters by computing distances with the transcriptome data for the genes for each unique pair of nuclei or cells in the plurality of nuclei or cells and by evaluating the distances with a criterion function. Each distance represents a different pair of nuclei or cells in the plurality of nuclei or cells and quantifies a distance between (i) a vector formed by the transcriptome data for the plurality of genes for a respective first nucleus or cell and (ii) a vector formed by the transcriptome data for the plurality of genes for a second nucleus or cell. Each cluster represents a subset of nuclei or cells of the plurality of nuclei or cells clustered together based on evaluation of distances with the criterion function. The metadata is used to identify a cluster in the plurality of clusters with the first myelofibrosis-related state by determining that the cluster includes nuclei or cells from subjects in the cohort having the first myelofibrosis- related state.
[0139] Reference will now be made in detail to embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerousspecific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to one of ordinary skill in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.
[0140] Plural instances may be provided for components, operations or structures described herein as a single instance. Finally, boundaries between various components, operations, and data stores are somewhat arbitrary, and particular operations are illustrated in the context of specific illustrative configurations. Other forms of functionality are envisioned and may fall within the scope of the implementation(s). In general, structures and functionality presented as separate components in the example configurations may be implemented as a combined structure or component. Similarly, structures and functionality presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the implementation(s).
[0141] It will also be understood that, although the terms “first,” “second,” etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first dataset could be termed a second dataset, and, similarly, a second dataset could be termed a first dataset, without departing from the scope of the present invention. The first dataset and the second dataset are both datasets, but they are not the same dataset.
[0142] The terminology used herein is for the purpose of describing particular implementations only and is not intended to be limiting of the claims. As used in the description of the implementations and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0143] As used herein, the term “if’ may be construed to mean “when” or “upon” or “in response to determining” or “in accordance with a determination” or “in response to detecting,” that a stated condition precedent is true, depending on the context. Similarly, thephrase “if it is determined (that a stated condition precedent is true)” or “if (a stated condition precedent is true)” or “when (a stated condition precedent is true)” may be construed to mean “upon determining” or “in response to determining” or “in accordance with a determination” or “upon detecting” or “in response to detecting” that the stated condition precedent is true, depending on the context.
[0144] Furthermore, when a reference number is given an “zth” denotation, the reference number refers to a generic component, set, or embodiment. For instance, a cellular- component termed “cellular-component z” refers to the zthcellular-component in a plurality of cellular-components.
[0145] In the interest of clarity, not all of the routine features of the implementations described herein are shown and described. It will be appreciated that, in the development of any such actual implementation, numerous implementation-specific decisions are made in order to achieve the designer’s specific goals, such as compliance with use case- and business-related constraints, and that these specific goals will vary from one implementation to another and from one designer to another. Moreover, it will be appreciated that such a design effort might be complex and time-consuming, but nevertheless be a routine undertaking of engineering for those of ordering skill in the art having the benefit of the present disclosure.
[0146] Some portions of this description describe the embodiments of the invention in terms of algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations are commonly used by those skilled in the data processing arts to convey the substance of their work effectively to others skilled in the art. These operations, while described functionally, computationally, or logically, are understood to be implemented by computer programs or equivalent electrical circuits, microcode, or the like.
[0147] The language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the inventive subject matter. It is therefore intended that the scope of the invention be limited not by this detailed description, but rather by any claims that issue on an application based hereon. Accordingly, the disclosure of the embodiments of the invention is intended to be illustrative, but not limiting, of the scope of the invention.
[0148] In general, terms used in the claims and the specification are intended to be construed as having the plain meaning understood by a person of ordinary skill in the art.Certain terms are defined below to provide additional clarity. In case of conflict between the plain meaning and the provided definitions, the provided definitions are to be used.
[0149] Any terms not directly defined herein shall be understood to have the meanings commonly associated with them as understood within the art of the invention. Certain terms are discussed herein to provide additional guidance to the practitioner in describing the compositions, devices, methods and the like of aspects of the invention, and how to make or use them. It will be appreciated that the same thing may be said in more than one way. Consequently, alternative language and synonyms may be used for any one or more of the terms discussed herein. No significance is to be placed upon whether or not a term is elaborated or discussed herein. Some synonyms or substitutable methods, materials and the like are provided. Recital of one or a few synonyms or equivalents does not exclude use of other synonyms or equivalents, unless it is explicitly stated. Use of examples, including examples of terms, is for illustrative purposes only and does not limit the scope and meaning of the aspects of the invention herein.
[0150] Definitions.
[0151] As used herein, the term “about” or “approximately” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, in some embodiments “about” means within 1 or more than 1 standard deviation, per the practice in the art. In some embodiments, “about” means a range of ±20%, ±10%, ±5%, or ±1% of a given value. In some embodiments, the term “about” or “approximately” means within an order of magnitude, within 5-fold, or within 2- fold, of a value. Where particular values are described in the application and claims, unless otherwise stated the term “about” meaning within an acceptable error range for the particular value can be assumed. All numerical values within the detailed description herein are modified by “about” the indicated value, and consider experimental error and variations that would be expected by a person having ordinary skill in the art. The term “about” can have the meaning as commonly understood by one of ordinary skill in the art. In some embodiments, the term “about” refers to ±10%. In some embodiments, the term “about” refers to ±5%.
[0152] As used herein, the terms “abundance,” “abundance level,” or “expression level” refers to an amount of a cellular constituent (e.g., a gene product such as an RNA species, e.g., mRNA or miRNA, or a protein molecule) present in one or more cells, or an averageamount of a cellular constituent present across multiple cells. When referring to mRNA or protein expression, the term generally refers to the amount of any RNA or protein species corresponding to a particular genomic locus, e.g, a particular gene. However, in some embodiments, an abundance can refer to the amount of a particular isoform of an mRNA or protein corresponding to a particular gene that gives rise to multiple mRNA or protein isoforms. The genomic locus can be identified using a gene name, a chromosomal location, or any other genetic mapping metric.
[0153] As used interchangeably herein, a “cell state” or “biological state” refers to a state or phenotype of a cell or a population of cells. For example, a cell state can be healthy or diseased. A cell state can be one of a plurality of diseases. A cell state can be a response to a compound treatment and / or a differentiated cell lineage. A cell state can be characterized by a measure (e.g, an activation, expression, and / or measure of abundance) of one or more cellular constituents, including but not limited to one or more genes, one or more proteins, and / or one or more biological pathways.
[0154] As used herein, a “cell state transition” or “cellular transition” refers to a transition in a cell’s state from a first cell state to a second cell state. In some embodiments, the second cell state is an altered cell state (e.g., a healthy cell state to a diseased cell state). In some embodiments, one of the respective first cell state and second cell state is an unperturbed state and the other of the respective first cell state and second cell state is a perturbed state caused by an exposure of the cell to a condition. The perturbed state can be caused by exposure of the cell to a compound. A cell state transition can be marked by a change in cellular constituent abundance in the cell, and thus by the identity and quantity of cellular constituents (e.g., mRNA, transcription factors) produced by the cell (e.g., a perturbation signature).
[0155] As used herein, the term “dataset” in reference to cellular constituent abundance measurements for a cell or a plurality of cells can refer to a high-dimensional set of data collected from a single cell (e.g., a single-cell cellular constituent abundance dataset) in some contexts. In other contexts, the term “dataset” can refer to a plurality of high-dimensional sets of data collected from single cells (e.g., a plurality of single-cell cellular constituent abundance datasets), each set of data of the plurality collected from one cell of a plurality of cells.
[0156] As used herein, the term “differential abundance” or “differential expression” refers to differences in the quantity and / or the frequency of a cellular constituent present in a first entity (e.g., a first cell, plurality of cells, and / or sample) as compared to a second entity e.g.,a second cell, plurality of cells, and / or sample). In some embodiments, a first entity is a sample characterized by a first cell state (e.g., a diseased phenotype) and a second entity is a sample characterized by a second cell state (e.g., a normal or healthy phenotype). For example, a cellular constituent can be a polynucleotide e.g., an mRNA transcript) which is present at an elevated level or at a decreased level in entities characterized by a first cell state compared to entities characterized by a second cell state. In some embodiments, a cellular constituent can be a polynucleotide which is detected at a higher frequency or at a lower frequency in entities characterized by a first cell state compared to entities characterized by a second cell state. A cellular constituent can be differentially abundant in terms of quantity, frequency or both. In some instances, a cellular constituent is differentially abundant between two entities if the amount of the cellular constituent in one entity is statistically significantly different from the amount of the cellular constituent in the other entity. For example, a cellular constituent is differentially abundant in two entities if it is present at least about 120%, at least about 130%, at least about 150%, at least about 180%, at least about 200%, at least about 300%, at least about 500%, at least about 700%, at least about 900%, or at least about 1000% greater in one entity than it is present in the other entity, or if it is detectable in one entity and not detectable in the other. In some instances, a cellular constituent is differentially expressed in two sets of entities if the frequency of detecting the cellular constituent in a first subset of entities (e.g., cells representing a first subset of annotated cell states) is statistically significantly higher or lower than in a second subset of entities (e.g., cells representing a second subset of annotated cell states). For example, a cellular constituent is differentially expressed in two sets of entities if it is detected at least about 120%, at least about 130%, at least about 150%, at least about 180%, at least about 200%, at least about 300%, at least about 500%, at least about 700%, at least about 900%, or at least about 1000% more frequently or less frequently observed in one set of entities than the other set of entities.
[0157] As used herein, the term “sample,” “biological sample,” or “patient sample,” refers to any sample taken from a subject, which can reflect a biological state associated with the subject. Examples of samples include, but are not limited to, blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid of the subject. A sample can include any tissue or material derived from a living or dead subject. A sample can be a cell-free sample. A sample can comprise one or more cellular constituents. For instance, a sample can comprise a nucleic acid (e.g., DNA or RNA) or a fragment thereof, or a protein. The term “nucleic acid” can refer todeoxyribonucleic acid (DNA), ribonucleic acid (RNA) or any hybrid or fragment thereof. The nucleic acid in the sample can be a cell-free nucleic acid. A sample can be a liquid sample or a solid sample (e.g., a cell or tissue sample). A sample can be a bodily fluid. A sample can be a stool sample. A sample can be treated to physically disrupt tissue or cell structure (e.g., centrifugation and / or cell lysis), thus releasing intracellular components into a solution which can further contain enzymes, buffers, salts, detergents, and the like which can be used to prepare the sample for analysis.
[0158] The terms “sequence reads” or “reads,” used interchangeably herein, refer to nucleotide sequences produced by any sequencing process described herein or known in the art. Reads can be generated from one end of nucleic acid fragments (“single-end reads”), and sometimes are generated from both ends of nucleic acids (e.g., paired-end reads, double-end reads). The length of the sequence read is often associated with the particular sequencing technology. High-throughput methods, for example, provide sequence reads that can vary in size from tens to hundreds of base pairs (bp). In some embodiments, the sequence reads are of a mean, median or average length of about 15 bp to 900 bp long (e.g., about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, or about 500 bp. In some embodiments, the sequence reads are of a mean, median or average length of about 1000 bp or more. Nanopore sequencing, for example, can provide sequence reads that vary in size from tens to hundreds to thousands of base pairs. Illumina parallel sequencing can provide sequence reads vary to a lesser extent (e.g, where most sequence reads are of a length of about 200 bp or less). A sequence read (or sequencing read) can refer to sequence information corresponding to a nucleic acid molecule (e.g, a string of nucleotides). For example, a sequence read can correspond to a string of nucleotides (e.g., about 20 to about 150) from part of a nucleic acid fragment, can correspond to a string of nucleotides at one or both ends of a nucleic acid fragment, or can correspond to nucleotides of the entire nucleic acid fragment. A sequence read can be obtained in a variety of ways, e.g., using sequencing techniques or using probes (e.g., in hybridization arrays or capture probes) or amplification techniques, such as the polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification.
[0159] As disclosed herein, the terms “sequencing,” “sequence determination,” and the like refer generally to any and all biochemical processes that may be used to determine the orderof biological macromolecules such as nucleic acids or proteins. For example, sequencing data can include all or a portion of the nucleotide bases in a nucleic acid molecule such as a DNA fragment.
[0160] As used herein, the term “tissue” corresponds to a group of cells that group together as a functional unit. More than one type of cell can be found in a single tissue. Different types of tissue may consist of different types of cells (e.g., hepatocytes, alveolar cells or blood cells), but also can correspond to tissue from different organisms (mother vs. fetus) or to healthy cells vs. tumor cells. The term “tissue” can generally refer to any group of cells found in the human body (e.g., heart tissue, lung tissue, kidney tissue, nasopharyngeal tissue, oropharyngeal tissue). In some aspects, the term “tissue” or “tissue type” can be used to refer to a tissue from which a cell-free nucleic acid originates. In one example, viral nucleic acid fragments can be derived from blood tissue. In another example, viral nucleic acid fragments can be derived from tumor tissue.
[0161] I. Exemplary System Embodiments
[0162] Now that an overview of some aspects of the present disclosure and some definitions used in the present disclosure have been provided, details of an exemplary system are described in conjunction with Figure 1.
[0163] Figure 1 illustrates a computer system 100 for manufacturing a myelofibrosis-related detection system. In typical embodiments, computer system 100 comprises one or more computers. For purposes of illustration in Figure 1, the computer system 100 is represented as a single computer that includes all of the functionality of the disclosed computer system 100. However, the present disclosure is not so limited. The functionality of the computer system 100 may be spread across any number of networked computers and / or reside on each of several networked computers and / or virtual machines. One of skill in the art will appreciate that a wide array of different computer topologies is possible for the computer system 100 and all such topologies are within the scope of the present disclosure.
[0164] Turning to Figure 1 with the foregoing in mind, the computer system 100 comprises one or more processing units (CPUs) 52, a network or other communications interface 54, a user interface 56 (e.g., including an optional display 58 and optional input 60 (e.g. keyboard or other form of input device)), a memory 92 (e.g, random access memory, persistent memory, or combination thereof), and one or more communication busses 94 for interconnecting the aforementioned components. To the extent that components of memory 92 are not persistent, data in memory 92 can be seamlessly shared with non-volatile memory(not shown) or portions of memory 92 that are non-volatile / persistent using known computing techniques such as caching. Memory 92 can include mass storage that is remotely located with respect to the central processing unit(s) 52. In other words, some data stored in memory 92 may in fact be hosted on computers that are external to computer system 100 but that can be electronically accessed by the computer system 100 over an Internet, intranet, or other form of network or electronic cable using network interface 54. In some embodiments, the computer system 100 makes use of models that are run from the memory associated with one or more graphical processing units in order to improve the speed and performance of the system. In some alternative embodiments, the computer system 100 makes use of models that are run from memory 92 rather than memory associated with a graphical processing unit.
[0165] The memory 92 of the computer system 100 stores:• an optional operating system 102 that includes procedures for handling various basic system services;• an analysis module 103 for manufacturing a myelofibrosis-related detection system;• first information 104 comprising below referenced single-nucleus or single-cell transcriptome data 106 and first metadata 116;• single-nucleus or single-cell transcriptome data 106 for a plurality of genes 112 (112- 1-1, ..., 112-1-Q; 112-P-l, ...., 112-P-Q) for each nucleus or cell 108 (108-1, ..., 108- P) in a first plurality of nuclei or cells, each nucleus or cell including a barcode 110 (110-1, ..., 110-P);• first metadata 116 for each respective subject 118 (118-1, ..., 118-X) in a cohort of subjects indicating their myelofibrosis-related state 120 (e.g., 120-1), their barcode 110 (e.g., 110-1), sex 124 (e.g., 124-1), race 126 (e.g., 126-1), age 128 (e.g., 128-1), and one or more physical conditions 130 (e.g., 130-1, e.g., body mass index);• Clusters 132, each cluster 134 (134-1, ..., 134-Z) representing a subset 136 (136-1, ... 136-Z) of the plurality of nuclei or cells found in the single-nucleus or single-cell transcriptome data 106 that cluster together based on their respective transcriptome data.
[0166] In some implementations, one or more of the above identified data elements or modules of the computer system 100 are stored in one or more of the previously mentioned memory devices, and correspond to a set of instructions for performing a function described above. The above identified data, modules or programs (e.g., sets of instructions) need not be implemented as separate software programs, procedures or modules, and thus various subsets of these modules may be combined or otherwise re-arranged in various implementations. Insome implementations, the memory 92 optionally stores a subset of the modules and data structures identified above. Furthermore, in some embodiments the memory 92 stores additional modules and data structures not described above.
[0167] IL Methods for manufacturing myelofibrosis-related detection system.
[0168] Referring to block 200 of Fig. 2A, a method for manufacturing a myelofibrosis- related detection system is provided at a computer system having one or more processors and memory storing one or more programs for execution by the one or more processors.
[0169] Referring to block 202 first information is obtained that comprises single-nucleus, or single-cell, transcriptome data for a plurality of genes for each nucleus or cell in a first plurality of nuclei or cells. Each nucleus or cell in the first plurality of nuclei or cells is obtained from a sample in a plurality of samples. Each sample in the plurality of samples is from a different subject in a cohort of subjects. The first plurality of nuclei or cells includes a different subset of nuclei or cells from a sample from each subject in the cohort of subjects. In some embodiments, the first plurality of nuclei or cells comprises at least 1000 nuclei or cells, at least 2,000 nuclei or cells, at least 3,000 nuclei or cells, at least 4,000 nuclei or cells, at least 5,000 nuclei or cells, or at least 10,000 nuclei or cells. In some embodiments the first plurality of nuclei or cells includes at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, or 500 nuclei or cells from each sample and each such sample is from a different subject in the cohort of subjects. Throughout the present disclosure it will be understood that while certain advantages may be incurred by using single-nucleus data, single-cell data can be used. As such, each metric or range that is referenced in relation to the plurality of nuclei (such as the number of nuclei) is equally applicable to embodiments that make use of single-cell transcriptome data.
[0170] In some embodiments single-nucleus sequence or single cell sequencing is used to obtain a plurality of sequence reads from each nucleus or cell in a plurality of nuclei or cells in the sample. See Jiang et al., February 23, 2023, “Isolated nuclei from frozen tissue are the superior source for single cell RNA-seq compared with whole cells,” available on the Internet at doi.org / 10.1101 / 2023.02.19.529150, which is hereby incorporated by reference.
[0171] In some embodiments, first information comprises a respective plurality of sequence reads for each respective nuclei or cell in the first plurality of nuclei or cells. In some embodiments each respective plurality of sequence reads comprises at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10,000, at least 50,000, at least 100,000, at least 500,000, at least 1 million, at least 2 million, at least 3 million, at least 4 million, at least 5 million, at least 6 million, at least 7 million, at least 8 million, at least 9 million, or more sequence reads. In some embodiments, each respective plurality of sequence reads comprises at least 1 x 107, at least 2 x 107, at least 3 x 107, at least 4 x 107, at least 5 x 107, at least 6 x 107, at least 7 x 107, at least 8 x 107, at least 9 x 107, at least 1 x 108, at least 2 x 108, at least 3 x 108, at least 4 x 108, at least 5 x 108, at least 6 x 108, at least 7 x 108, at least 8 x 108, at least 9 x 108, at least 1 x 109, or more sequence reads. In some embodiments, each respective plurality of sequence reads consists of no more than 5 x 107, no more than 1 x 107, no more than 5 x 106, no more than 4 x 106, no more than 3 x 106, no more than 2 x 106, no more than 1 x 106, no more than 500,000, no more than 100,000, no more than 50,000, no more than 30,000, no more than 20,000, no more than 10,000, no more than 9000, no more than 8000, no more than 7000, no more than 6000, no more than 5000, no more than 4000, no more than 3000, no more than 2000, no more than 1000, or less sequence reads.
[0172] In some embodiments, each respective plurality of sequence reads consists of between 1000 to 5000, from 1000 to 10,000, from 2000 to 20,000, from 5000 to 50,000, from 10,000 to 100,000, from 100,000 to 500,000 from 10,000 to 500,000, from 500,000 to 1 million, from 1 million to 30 million, from 30 million to 80 million, or from 10 million to 500 million sequence reads. In some embodiments, the respective plurality of sequence reads falls within another range starting no lower than 1000 sequence reads and ending no higher than 1 x 109sequence reads.
[0173] In some embodiments, the Cell Ranger Single Cell Software Suite (10X Genomics, Inc.) is used for processing each respective plurality of sequence reads. In some such embodiments this includes library demultiplexing, fastq file generation, read alignment and unique molecular identification quantification against a lOx Genomics, Inc. pre-built human genome. This produces a unique molecular identifier (UMI) count for each gene for each nuclei or cell that is used in such embodiments as a basis for determining gene expression in each nuclei or cell. In some embodiments these UMI counts are log normalized.
[0174] In some embodiments the sequence reads are obtained by a whole-genome sequencing. In some embodiments the sequence reads are obtained by targeted DNAsequencing using a plurality of nucleic acid probes. In some embodiments, the plurality of sequence reads is determined by scTag-seq. In some embodiments, the plurality of sequence reads is determined by single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq), CyTOF / SCoP, E-MS / Abseq, miRNA-seq, CITE-seq, or any combination thereof.
[0175] In some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes in the first plurality of nuclei or cells are preprocessed. In some embodiments, the preprocessing includes one or more of filtering, normalization, mapping (e.g., to a reference sequence), quantification, scaling, deconvolution, cleaning, dimension reduction, transformation, statistical analysis, and / or aggregation.
[0176] For example, in some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes in the first plurality of nuclei or cells is filtered based on a desired quality, e.g., size and / or quality of a nucleic acid sequence, or a minimum and / or maximum abundance value for a respective gene in a particular nucleus or cell. In some embodiments, filtering is performed in part or in its entirety by various software tools, such as Skewer. See, Jiang, H. etal., BMC Bioinformatics 15(182): 1-12 (2014). In some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes in the first plurality of nuclei or cells is filtered for quality control, for example, using a sequencing data QC software such as AfterQC, Kraken, RNA-SeQC, FastQC, or another similar software program. In some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes in the first plurality of nuclei or cells is normalized, e.g., to account for pull-down, amplification, and / or sequencing bias (e.g., mappability, GC bias etc ). See, for example, Schwartz et al., PLoS ONE 6(l):el6685 (2011) and Benjamini and Speed, Nucleic Acids Research 40(10):e72 (2012), the contents of which are hereby incorporated by reference, in their entireties, for all purposes. For instance, in some embodiments the single-nucleus or single-cell transcriptome data for the plurality of genes in the first plurality of nuclei or cells is scaled to have a mean value of zero and a standard of deviation of 1. In some embodiments, the preprocessing the single-nucleus or single-cell transcriptome data for the plurality of genes in the first plurality of nuclei or cells improves (e.g, lowers) a high signal -to-noise ratio.
[0177] Thus, in some embodiments, the single-nucleus or single-cell transcriptome data for a plurality of genes for each nucleus or cell in a first plurality of nuclei or cells comprises any one of a variety of forms, including, without limitation, raw abundance values, absolute abundance values (e.g, transcript number), relative abundance values (e.g., relativefluorescent units, transcriptome analysis, and / or gene set expression analysis (GSEA)) for each gene in each nucleus or cell, compound or aggregated abundance value for each gene in each nucleus or cell, transformed abundance value (e.g., Iog2 and / or log transformed) for each gene in each nucleus or cell, a change (e.g., fold- or log-change) relative to a reference (e.g., a normal sample, matched sample, reference dataset, housekeeping gene, and / or reference standard) for each gene in each nucleus or cell, a standardized abundance value for each gene in each nucleus or cell, a measure of central tendency (e.g., mean, median, mode, weighted mean, weighted median, and / or weighted mode) for each gene in each nucleus or cell, a measure of dispersion (e.g, variance, standard deviation, and / or standard error) for each gene in each nucleus or cell, or an adjusted abundance value (e.g, normalized, scaled, and / or error- corrected) for each gene in each nucleus or cell.
[0178] Any one of a number of counting techniques may be used to obtain the abundance value for each gene in each nucleus or cell in the first plurality of nuclei or cells.
[0179] In some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in a first plurality of nuclei or cells is determined using one or more methods including microarray analysis via fluorescence, chemiluminescence, electric signal detection, polymerase chain reaction (PCR), reverse transcriptase polymerase chain reaction (RT-PCR), digital droplet PCR (ddPCR), solid-state nanopore detection, RNA switch activation, a Northern blot, and / or a serial analysis of gene expression (SAGE).
[0180] In some embodiments, single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is determined by a colorimetric measurement, a fluorescence measurement, a luminescence measurement, or a resonance energy transfer (FRET) measurement.
[0181] In some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is gene expression data. In some embodiments, gene expression in a respective nucleus or cell in the first plurality of nuclei or cells is measured by sequencing the RNA transcripts of the nuclei or cells and then counting the quantity of each gene transcript mapping to a gene in the first plurality of genes, identified during the sequencing. In some embodiments, the gene transcripts sequenced and quantified include RNA, such as mRNA. In some embodiments, the gene transcripts sequenced and quantified include a downstream product of mRNA, such as a protein e.g., a transcription factor). In general, as used herein, the term “gene transcript”may be used to denote any downstream product of gene transcription or translation, including post-translational modification, and “gene expression” and “cellular constituent abundance” may interchangeably be used to refer generally to any measure of gene transcripts.
[0182] In some embodiments, the abundance of a respective gene in the first plurality of cellular constituents is RNA abundance (e.g., gene expression), and the abundance of the respective gene in a nucleus or cell in the plurality of nuclei or cells is determined by measuring polynucleotide levels of one or more nucleic acid molecules corresponding to the respective gene in a nucleus or cell. The transcript levels of the respective gene can be determined from the amount of mRNA, or polynucleotides derived therefrom, present in the nucleus or cell. Polynucleotides can be detected and quantitated by a variety of methods including, but not limited to, microarray analysis, polymerase chain reaction (PCR), reverse transcriptase polymerase chain reaction (RT-PCR), Northern blot, serial analysis of gene expression (SAGE), RNA switches, RNA fingerprinting, ligase chain reaction, Qbeta replicase, isothermal amplification method, strand displacement amplification, transcription based amplification systems, nuclease protection assays (Si nuclease or RNAse protection assays), and / or solid-state nanopore detection. See, e.g., Draghi ci, Data Analysis Tools for DNA Microarrays, Chapman and Hall / CRC, 2003; Simon et al., Design and Analysis of DNA Microarray Investigations, Springer, 2004; Real-Time PCR: Current Technology and Applications, Logan, Edwards, and Saunders eds., Caister Academic Press, 2009; Bustin A-Z of Quantitative PCR (IUL Biotechnology, No. 5), International University Line, 2004;Velculescu etal., (1995) Science 270: 484-487; Matsumura et al, (2005) Cell. Microbiol. 7: 11-18; Serial Analysis of Gene Expression (SAGE): Methods and Protocols (Methods in Molecular Biology), Humana Press, 2008; each of which is hereby incorporated herein by reference in its entirety.
[0183] In some embodiments, single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is obtained from expressed RNA or a nucleic acid derived therefrom (e.g., cDNA or amplified RNA derived from cDNA that incorporates an RNA polymerase promoter) from the first plurality of nuclei or cells, including naturally occurring nucleic acid molecules, as well as synthetic nucleic acid molecules. Thus, in some embodiments, the abundance of each gene in the first plurality of genes of the single-nucleus or single-cell transcriptome data for each nucleus or cell in the first plurality of nuclei or cells is obtained from such non-limiting sources as total cellular RNA, poly(A)+ messenger RNA (mRNA) or a fraction thereof, cytoplasmic mRNA, or RNA transcribed from cDNA (e.g., cRNA). Methods for preparing total and poly(A)+RNA are well known in the art, and are described generally, e.g., in Sambrook, et al. , Molecular Cloning: A Laboratory Manual (3rd Edition, 2001). RNA can be extracted from a cell or nucleus of interest using guanidinium thiocyanate lysis followed by CsCl centrifugation (see, e.g., Chirgwin et al., 1979, Biochemistry 18:5294-5299), a silica gelbased column (e.g., RNeasy (Qiagen, Valencia, Calif.) or StrataPrep (Stratagene, La Jolla, Calif.)), or using phenol and chloroform, as described in Ausubel et al., eds., 1989, Current Protocols In Molecular Biology, Vol. Ill, Green Publishing Associates, Inc., John Wiley & Sons, Inc., New York, at pp. 13.12.1-13.12.5). Poly(A)+ RNA can be selected, e.g., by selection with oligo-dT cellulose or, alternatively, by oligo-dT primed reverse transcription of total cellular RNA. RNA can be fragmented by methods known in the art, e.g., by incubation with ZnCh, to generate fragments of RNA.
[0184] In some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is determined by sequencing transcripts from the plurality of cells. In some embodiments, the singlenucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is determined by single-cell ribonucleic acid (RNA) sequencing (scRNA-seq), scTag-seq, single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq), CyTOF / SCoP, E-MS / Abseq, miRNA-seq, CITE-seq, or any combination thereof.
[0185] In some embodiments, scRNA-seq, scTag-seq, and miRNA-seq are used to measure RNA expression in order to obtain the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells. scRNA-seq measures expression of RNA transcripts, scTag-seq allows detection of rare mRNA species, and miRNA-seq measures expression of micro-RNAs. CyTOF / SCoP and E-MS / Abseq can be used to measure protein expression in the cell. CITE-seq simultaneously measures both gene expression and protein expression in the cell, and scATAC-seq measures chromatin conformation in the cell. Table 1 below provides example protocols for performing each of the measurement techniques described above. In some embodiments, any of the protocols described in Table 1 of Shen et al., 2022, “Recent advances in high-throughput single-cell transcriptomics and spatial transcriptomics,” Lab Chip 22, p. 4774, is used to measure the abundance of cellular constituents, such as genes, in order to obtain the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells.
[0186] Table 1 - Example Measurement Protocols
[0187] Referring to block 204, in some embodiments, the plurality of genes comprises 100 or more genes, 200 or more genes, 500 or more genes, 1000 or more genes, 2000 or more genes, 5000 or more genes, 10000 or more genes, or 15,000 or more genes. In some embodiments, the plurality of genes consists of between 50 and 15,000 genes. In some embodiments, the plurality of genes consists of between 100 and 15,000 genes. In some embodiments, the plurality of genes consists of between 200 and 12,000 genes. In some embodiments, the plurality of genes consists of between 250 and 10,000 genes. In some embodiments, the plurality of genes consists of between 300 and 5,000 genes. In embodiments, the plurality of genes comprises at least 2, at least 5, at least 10, at least 15, atleast 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 genes. In embodiments, the plurality of genes comprises at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1000 genes. In embodiments, the plurality of genes comprises at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 10,000, at least 30,000, at least 50,000, or more than 50,000 genes. In embodiments, the plurality of genes comprises between 2 and 20, between 20 and 50, between 50 and 100, between 100 and 200, between 200 and 500, between 500 and 1000, between 1000 and 5000, between 5000 and 10,000 genes, or between 10,000 and 50,000 genes.
[0188] In some embodiments, the first plurality of genes is limited to those genes that are highly variable in their expression across the first plurality of nuclei or cells. In some embodiments these highly variable genes are identified using the Seurat function FindVariableFeatures. See, Stuart el al., 2019, “Comprehensive integration of single-cell data,” Cell 177, 1888-1902, which is hereby incorporated by reference. In some embodiments, the most highly variable genes selected in this manner was limited to between 50 genes and 15,000 genes. In some embodiments. In some embodiments, the most highly variable genes selected in this manner was limited to between 100 genes and 15,000 genes. In some embodiments, the most highly variable genes selected in this manner was limited to between 200 genes and 12,000 genes. In some embodiments, the most highly variable genes selected in this manner was limited to between 250 genes and 10,000 genes. In some embodiments, the most highly variable genes selected in this manner was limited to between 300 genes and 5,000 genes.
[0189] In some embodiments, only genes expressed in at least 10, at least 20, at least 30, at least 40, or at least 50 nuclei or cells are kept in the first plurality of genes.
[0190] Referring to block 206, in some embodiments, the first plurality of nuclei or cells comprises 10,000 or more nuclei or cells, 50,000 or more nuclei or cells, 100,000 or more nuclei or cells, 250,000 or more nuclei or cells, 500,000 or more nuclei or cells, 600,000 or more nuclei or cells, or 1 x 106or more nuclei or cells. In some embodiments, the first plurality of nuclei or cells consists of between 1000 nuclei or cells and 100,000 nuclei or cells. In some embodiments, the first plurality of nuclei or cells consists of between 2000 nuclei or cells and 200,000 nuclei or cells. In some embodiments, the first plurality of nuclei or cells consists of between 3000 nuclei or cells and 250,000 nuclei or cells. In some embodiments, the first plurality of nuclei or cells consists of between 4000 nuclei or cells and500,000 nuclei or cells. In some embodiments, the first plurality of nuclei or cells consists of between 5000 nuclei or cells and 600,000 nuclei or cells. In some embodiments, the first plurality of nuclei or cells consists of between 6000 nuclei or cells and 1 x 106nuclei or cells. In some embodiments, the first plurality of nuclei or cells consists of between 7000 nuclei or cells and 2 x 106nuclei or cells.
[0191] In some embodiments, the first plurality of nuclei or cells comprises at least 5, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 1000, at least at least 2000, at least 3000, at least 4000, at least 5000, at least 10,000, at least 20,000, at least 30,000, at least 50,000, at least 80,000, at least 100,000, at least 500,000, or at least 1 million nuclei or cells. In some embodiments, the first plurality of nuclei comprises no more than 5 million, no more than 1 million, no more than 500,000, no more than 100,000, no more than 50,000, no more than 10,000, no more than 5000, no more than 1000, no more than 500, no more than 200, no more than 100, or no more than 50 nuclei or cells. In some embodiments, the first plurality of nuclei or cells comprises from 5 to 100, from 10 to 50, from 20 to 500, from 200 to 10,000, from 1000 to 100,000, from 50,000 to 500,000, or from 10,000 to 1 million nuclei or cells. In some embodiments, the first plurality of nuclei or cells falls within another range starting no lower than 5 nuclei or cells and ending no higher than 10 million nuclei or cells.
[0192] In some embodiments, outlier nuclei or cells are removed from the first plurality of nuclei or cells using the Scater R isOutlier function. See, McCarthy et al., 2017, “Scater: preprocessing, quality control, normalization and visualization of single-cell RNA-seq data in R,” Bioinformatics, 33, 1179-1186. which is hereby incorporated by reference.
[0193] In some embodiments, only those nuclei or cells that express at least 20, at least 40, at least 50, at least 100, or at least 200 genes are kept in the first plurality of nuclei or cells.
[0194] Referring to block 208, in some embodiments, the cohort of subjects comprises 25 or more subjects, 50 or more subjects, 75 or more subjects, or 100 or more subjects. In some embodiments, the plurality of subjects comprises 3 or more subjects, 5 or more subjects, 10 or more subjects, 15 or more subjects, 20 or more subjects, 25 or more subjects, 30 or more subjects, 35 or more subjects, 40 or more subjects, 45 or more subjects, 50 or more subjects, 55 or more subjects, 60 or more subjects, 65 or more subjects, 70 or more subjects, 75 or more subjects, 80 or more subjects, 85 or more subjects, 90 or more subjects, 95 or more subjects, or 100 or more subjects. In some embodiments, the plurality of subjects consists of between 3 subjects and 500 subjects. In some embodiments, the plurality of subjects consistsof between 5 subjects and 1000 subjects. In some embodiments, the plurality of subjects consists of between 10 subjects and 2000 subjects. In some embodiments, the plurality of subjects consists of between 15 subjects and 2500 subjects. In some embodiments, the plurality of subjects consists of between 20 subjects and 3000 subjects.
[0195] Referring to block 210, in some embodiments, each sample in the plurality of samples is from a different subject in the cohort of subjects and the plurality of samples comprises 5 samples, 20 samples, 50 samples, or 100 or more samples. In some embodiments, each sample in the plurality of samples is from a different subject in the cohort of subjects and the plurality of samples comprises 3 or more samples, 5 or more samples, 10 or more samples, 15 or more samples, 20 or more samples, 25 or more samples, 30 or more samples, 35 or more samples, 40 or more samples, 45 or more samples, 50 or more samples, 55 or more samples, 60 or more samples, 65 or more samples, 70 or more samples, 75 or more samples, 80 or more samples, 85 or more samples, 90 or more samples, 95 or more samples, or 100 or more samples. In some embodiments, each sample in the plurality of samples is from a different subject in the cohort of subjects and the plurality of samples consists of between 3 samples and 500 samples. In some embodiments, each sample in the plurality of samples is from a different subject in the cohort of subjects and the plurality of samples consists of between 5 samples and 1000 samples. In some embodiments, each sample in the plurality of samples is from a different subject in the cohort of subjects and the plurality of samples consists of between 10 samples and 2000 samples. In some embodiments, each sample in the plurality of samples is from a different subject in the cohort of subjects and the plurality of samples consists of between 15 samples and 2500 samples. In some embodiments, each sample in the plurality of samples is from a different subject in the cohort of subjects and the plurality of samples consists of between 20 samples and 3000 samples. In some embodiments, each of the samples is flash frozen after being biopsied or removed from the corresponding subject.
[0196] Referring to block 212, the first information obtained further comprises first metadata for each respective subject in the cohort of subjects. The first information indicates at least for each respective subject in the cohort of subjects whether the respective subject has a first myelofibrosis-related state or a second myelofibrosis-related state. At least a first subset of subjects in the cohort of subjects have the first myelofibrosis-related state and a second subset of subjects in the cohort of subjects have the second myelofibrosis-related state.
[0197] Referring to block 214, in some embodiments, the first myelofibrosis-related state or the second myelofibrosis-related state is a myeloproliferative neoplasm. In some embodiments the first myelofibrosis-related state or the second myelofibrosis-related state comprises a primary myelofibrosis, a secondary myelofibrosis, essential thrombocythemia, or polycythemia vera. In some embodiments the first myelofibrosis-related state or the second myelofibrosis-related state is associated with osteosclerosis, extramedullary hematopoiesis (EMH), inefficient hematopoiesis, inflammation, splenomegaly, cytopenia, portal hypertension, thromboembolism, infection, or acute myeloid leukemia (AML).
[0198] Referring to block 216, in some embodiments, the first subset of subjects comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects and the second subset of subjects is other than the first subset of subjects and comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects. In some embodiments, the first subset of subjects comprises 3 or more subjects, 5 or more subjects, 10 or more subjects, 15 or more subjects, 20 or more subjects, 25 or more subjects, 30 or more subjects, 35 or more subjects, 40 or more subjects, 45 or more subjects, 50 or more subjects, 55 or more subjects, 60 or more subjects, 65 or more subjects, 70 or more subjects, 75 or more subjects, 80 or more subjects, 85 or more subjects, 90 or more subjects, 95 or more subjects, or 100 or more subjects. In some embodiments, the first subset of subjects consists of between 3 subjects and 500 subjects. In some embodiments, the plurality of subjects consists of between 5 subjects and 1000 subjects. In some embodiments, the first subset of subjects consists of between 10 subjects and 2000 subjects. In some embodiments, the first subset of subjects consists of between 15 subjects and 2500 subjects. In some embodiments, the first subset of subjects consists of between 20 subjects and 3000 subjects. In some embodiments, the second subset of subjects comprises 3 or more subjects, 5 or more subjects, 10 or more subjects, 15 or more subjects, 20 or more subjects, 25 or more subjects, 30 or more subjects, 35 or more subjects, 40 or more subjects, 45 or more subjects, 50 or more subjects, 55 or more subjects, 60 or more subjects, 65 or more subjects, 70 or more subjects, 75 or more subjects, 80 or more subjects, 85 or more subjects, 90 or more subjects, 95 or more subjects, or 100 or more subjects. In some embodiments, the second subset of subjects consists of between 3 subjects and 500 subjects. In some embodiments, the plurality of subjects consists of between 5 subjects and 1000 subjects. In some embodiments, the second subset of subjects consists of between 10 subjects and 2000 subjects. In some embodiments, the second subset of subjects consists of between 15 subjects and 2500 subjects. In some embodiments, the second subset of subjects consists of between 20 subjects and 3000 subjects.
[0199] Referring to block 218, in some embodiments, the first metadata for each respective subject in the cohort of subjects further indicates one or more features of the subject selected from the group consisting of sex, race, age, and physical condition. In embodiments, the physical condition comprises one or more features selected from sex, body mass index (BMI), age, ethnicity, cause of death, serology (CMV, EBV status), final lab profile, alcohol consumption, tobacco use, illicit drug use, medical history, family medical history, disease diagnostic information, medication profile, etc.
[0200] In some embodiments, the first metadata is any combination of sex, BMI, age, ethnicity, cause of death, serology (CMV, EBV status), final lab profile, alcohol use, tobacco use, illicit drug use, International Prognostic Scoring System (IPSS) score, Dynamic IPSS (DIPSS) score, DIPSS Plus score, Mutation-Enhanced International Prognostic Score System (MIPSS70) score, MIPSS70+ score, Genetically Inspired Prognostic Scoring System (GIPSS) score, Myelofibrosis Secondary to PV and ET-Prognostic Model (MYSEC-PM) score, Myelofibrosis Transplant Scoring System (MTSS) score, Response to Ruxolitinib after 6 Months (RR6) score, Artificial Intelligence Prognostic Scoring System for Myelofibrosis (AIPSS-MF) score, medical history, and / or medications.
[0201] In embodiments, the disease diagnostic information comprises one or more features selected from pathology notes and key features.
[0202] Referring to block 220, in some embodiments, the first metadata for each respective subject in the cohort of subjects indicting at least for each respective subject in the cohort of subjects whether the respective subject has a first myelofibrosis-related state or a second myelofibrosis-related state comprises a histologically graded disease status for the respective subject.
[0203] Referring to block 222, in some embodiments, the first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has a first myelofibrosis-related state or a second myelofibrosis-related state comprises a histologically graded disease status for the respective subject determined in accordance with a consistent, verified handling of each biological sample in the plurality of biological samples.
[0204] Referring to block 224, in some embodiments, the first metadata includes one or more first biomarker annotations for each nucleus or cell in the plurality of nuclei or cells drawn from a first standardized set of biomarkers (e.g., genes, proteins, etc.) that each definea unique cell type. Examples of such first biomarker annotations are given below in block 256 in conjunction with Table 2 below.
[0205] Referring to block 226, in some embodiments, sex, race, age, and / or physical condition are each represented in a balanced manner in the cohort of subjects. For instance, in some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is obtained in batches due to possible throughput limitations in obtaining such data. In such embodiments, each batch of the single-nucleus or single-cell transcriptome data is single-nucleus transcriptome or single-cell data for the plurality of genes for each nucleus or cell in a subset of the plurality of nuclei or cells.
[0206] In some embodiments the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is obtained in 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, between 21 and 100, or between 100 and 1000 batches where each respective batch is single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in a corresponding subset of the plurality of nuclei or cells from a corresponding subset of the subjects in the cohort of subjects. In some embodiments, each respective batch is examined relative to all other batches to ensure that the subjects represented in the respective batch are balanced for sex, race, age, and / or physical condition. For instance, in some such embodiments, this is done using one-way analysis of variance (ANOVA). That is, the ANOVA is used to determine whether there are any statistically significant differences between the means in sex, race, age, and / or physical condition (e.g., BMI) of the plurality of batches. The one-way ANOVA compares the means between the k batches and determines whether any of those means are statistically significantly different from each other by testing the null hypothesis:HQ - Pi=P2=P3= =Pk where p is the respective group mean for sex, race, age, and / or physical condition and k is the number of batches. If, however, the one-way ANOVA returns a statistically significant result, the alternative hypothesis HA), which is that there are at least two batches whose means are statistically significantly different from each other. In some embodiments, this is calculated using the F-statistic: n(variance of batch means) (mean of batch variances)In this equation, n is number of subjects represented in each of the batches. Each respective batch mean is the mean value for the property under study (sex, race, age, and / or physical condition) within the respective batch. Each respective batch variance is the variance in the value for the property under study (sex, race, age, and / or physical condition) within the respective batch. This F value can then be used to compute a p-value for the null hypothesis. See, Smith, 1991, Statistical Reasoning, Third Edition, Allyn and Bacon, Boston, Chapter 16, which is hereby incorporated by reference. In some embodiments, a p-value of less than 0.05 is used to accept the null hypothesis. That is, in some embodiments, ANOVA of sex, race, age, or physical condition across the batches yields a p-value of 0.05 or less. In some such embodiments, a p-value of less than 0.05 is required to accept the null hypothesis (that the batches are all balanced for the sex, race, age, and / or physical condition).
[0207] In some embodiments, ANOVA of sex, race, age, and / or physical condition across the batches yields a p-value of 0.10 or less. In some such embodiments, a p-value of less than 0.10 is required to accept the null hypothesis (that the batches are all balanced for the sex, race, age, and / or physical condition).
[0208] In some embodiments, ANOVA of sex, race, age, and / or physical condition across the batches yields a p-value of 0.01 or less. In some such embodiments, a p-value of less than 0.01 is required to accept the null hypothesis (that the batches are all balanced for the sex, race, age, and / or physical condition).
[0209] In some embodiments, an alternative statistical test is used to ensure sex, race, age, and / or physical condition is balanced across the batches. For instance, in some embodiments a t-test is used. See Smith, 1991, Statistical Reasoning, Third Edition, Allyn and Bacon, Boston, Chapter 9, which is hereby incorporated by reference. In some embodiments, a calculation of a p-value of 0.15 or less for the null hypothesis that each batch is balanced for sex, race, age, and / or physical condition using the alternative statistical test is used to ensure that the batches are balanced. In some embodiments, a calculation of a p-value of 0.10 or less for the null hypothesis that each batch is balanced for sex, race, age, and / or physical condition using the alternative statistical test is used to ensure that the batches are balanced. In some embodiments, a calculation of a p-value of 0.05 or less for the null hypothesis that each batch is balanced for sex, race, age, and / or physical condition using the alternative statistical test is used to ensure that the batches are balanced. In some embodiments, a calculation of a p-value of 0.01 or less for the null hypothesis that each batch is balanced for sex, race, age, and / or physical condition using the alternative statistical test is used to ensure that the batches are balanced.
[0210] In some embodiments a nonparametric test such as a sign test for median, Wilcoxon- Mann-Whitney Rank Sum test, Rank Correlation Test, or Runs test is used to ensure that the batches are balanced for sex, race, age, and / or physical condition. See Smith, 1991, Statistical Reasoning, Third Edition, Allyn and Bacon, Boston, Chapter 17, which is hereby incorporated by reference.
[0211] In some embodiments, each respective batch is examined relative to the singlenucleus or single-cell transcriptome data for each nucleus or cell in the first plurality of nuclei or cell for the entire cohort of subjects to ensure that the subjects represented in the respective batch are balanced for sex, race, age, and / or physical condition relative to the entire cohort.
[0212] In some embodiments the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is obtained in a single batch and no balancing is performed for sex, race, age, and / or physical condition.
[0213] In some embodiments, sex is balanced additionally or alternatively in the cohort of subjects by ensuring that between 40 percent and 60 percent of the subjects are female. In some embodiments, sex is balanced additionally or alternatively in the cohort of subjects by ensuring that between 45 percent and 55 percent of the subjects are female. In some embodiments, sex is balanced in the cohort of subjects additionally or alternatively by ensuring that between 47.5 percent and 52.5 percent of the subjects are female. In some embodiments, the cohort of subjects is not balanced for sex.
[0214] In some embodiments, sex is balanced in the cohort of subjects additionally or alternatively by ensuring that the participation to prevalence ratio (PPR) for woman (across the entire cohort) is between 0.80 and 1.20, where PPR is defined as:Percentage of woman in the cohort of subjects Percentage of woman among disease population (e. g. , subjects with first myelofibrosis — related state)
[0215] In some embodiments, sex is balanced in the cohort of subjects additionally or alternatively by ensuring that the PPR for woman (across the entire cohort) is between 0.90 and 1.10.
[0216] In some embodiments, race is represented in a balanced manner additionally or alternatively by ensuring that the respective PPR of white subjects, black subjects, and Asian subjects (across the entire cohort) is between 0.80 and 1.20. The PPR of a given race in such embodiments is calculated as:Percentage of subjects in the cohort of subjects that are of race X Percentage of subjects among disease population (e. g. , have the first first myelofibrosis — related state) that are of race X
[0217] In some embodiments, race is represented in a balanced manner additionally or alternatively in the cohort by ensuring that the respective PPR of white subjects, black subjects, and Asian subjects (across the entire cohort) is between 0.90 and 1.10.
[0218] In some embodiments, age is represented in a balanced manner additionally or alternatively by ensuring that the respective PPR of particular age groups (across the entire cohort) is between 0.80 and 1.20. The PPR of a given age group in such embodiments is calculated as:Percentage of subjects in the cohort in age group X Percentage of subjects among disease population (e. g. , have the first myelofibrosis — related state) in age group X
[0219] In some embodiments, age is represented in a balanced manner additionally or alternative by ensuring that the respective PPR of each respective age group in a particular set of age groups (across the entire cohort) is between 0.90 and 1.10. In some embodiments one of the age groups that is balanced in the cohort is the age group defined as being over 65 years in age. In some embodiments one of the age groups that is balanced in the cohort is the age group defined as being aged 45 to 64 years. In some embodiments one of the age groups that is balanced in the cohort is the age group defined as being aged 18-44 years.
[0220] In some embodiments, physical condition is balanced additionally or alternatively in the cohort of subjects by ensuring that the participation to prevalence ratio (PPR) for subjects with the physical condition is between 0.80 and 1.20, where PPR is defined as:Percentage of subjects in the cohort having the physical condition Percentage of subject among disease population (e. g., have the first myelofibrosis — related state) with the physical condition
[0221] In some embodiments, physical condition is balanced in the cohort of subjects additionally or alternatively by ensuring that the participation to prevalence ratio (PPR) for subjects with the physical condition (across the entire cohort) is between 0.90 and 1.10. In some embodiments the physical condition is body mass index. In some embodiments the physical condition is smoking status.
[0222] For further discussion of the use of PPR to analyze whether a cohort is balanced, see Varma et cd.. 2021, “Reporting of Study Participant Demographic Characteristics and Demographic Representation in Premarketing and Postmarketing Studies of Novel Cancer Therapeutics,” JAMA Netw Open Apr; 4(4): e217063, which is hereby incorporated by reference.
[0223] Referring to block 228, in some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is filtered to remove counts of ambient RNA molecules, doublets and / or empty droplets. Ambient RNA is the pool of mRNA molecules that have been released in a cell suspension, likely from cells that are stressed or have undergone apoptosis during singlenucleus or single-cell sequencing. Cross-contamination occurs when the ambient RNA gets incorporated into droplets representing single nuclei or single cells, which did not originate the ambient RNA, and is barcoded and amplified along with the nucleus’ or cell’s native mRNA. Liver biopsies often comprise 90% hepatocytes and the ambient RNA from the hepatocytes tends to drown out other less prevalent cell types in the biopsies through such ambient RNA contamination. Accordingly, in some embodiments the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells are filtered to remove counts of ambient RNA molecules.Contamination from ambient RNA is evident when highly expressed cell type-specific genes are observed at low levels in other cell populations. Different proportions of contamination can be found in different droplets depending on the amount of ambient and native mRNA present.
[0224] In some embodiments, ambient RNA is filtered to remove counts of ambient RNA molecules by assuming that the respective single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells in fact represents a mixture of counts from two multinomial distributions: (1) a distribution of native transcript counts from the nuclei’s or cell’s actual population and (2) a distribution of contaminating transcript counts from all other nuclei or cell populations captured in the assay. In some such embodiments, a program such as DecontX, Yang etal., 2020, “Decontamination of ambient RNA in single-cell RNA-seq with DecontX”, Genome Biology 21 :57, is used to deconvolute a gene-by-nuclei or gene-by-cell count matrix and a vector of nuclei or cell population labels into a matrix of contamination counts and a matrix of native counts that can be used in downstream analyses.
[0225] In some embodiments, ambient RNA is filtered to remove counts of ambient RNA molecules using Cellbender. See Fleming et aL, 2019, “CellBender remove-background: a deep generative model for unsupervised removal of background noise from scRNA-seq datasets,” bioRxiv 791699, doi: 10.1101 / 791699, which is hereby incorporated by reference.
[0226] In some embodiments, ambient RNA is filtered to remove counts of ambient RNA molecules using SoupX. See, Young and Behjati, 2020, “SoupX removes ambient RNA contamination from droplet-based single-cell RNA sequencing data,” Gigascience 9, doi: 10.1093 / gigascience / giaal51, which is hereby incorporated by reference.
[0227] In some embodiments, ambient RNA is filtered to remove counts of ambient RNA molecules using any known method. Additional examples of such methods are disclosed in Caglayan et al., 2022, “Ambient RNA analysis reveals misinterpreted and masked cell types in brain single-nuclei datasets,” Neuron doi: 10.1016 / j. neuron.2022.09.010, which is hereby incorporated by reference.
[0228] In some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells are filtered to remove doublets. This arises when more than one cell or nucleus is captured in a droplet, also known as a “doublet” or “multiplet.” In microfluidic systems, the occurrence of doublets is proportional to the concentration of cells or nuclei in the suspension and capture rate of the device. In some embodiments Scrublet, Wolock et al., 2019, “Scrublet: computational identification of cell doublets in single-cell transcriptomic data,” Cell Syst. 8(4):281-91, or DoubletFinder, McGinnis et al., 2019, “Doubletfinder: doublet detection in single-cell rna sequencing data using artificial nearest neighbors,” Cell Syst. 8(4):329-37, is used to remove doublets or multiplets in the single-nucleus or single-cell transcriptome data. These programs simulate artificial doublets from the original data coordinates in a reduced-dimensional representation, then create doublet score for each barcode by calculating the similarity of its representation with artificial doublets. In some embodiments demuxlet, Kang et al., 2018, “Multiplexed droplet single-cell rna-sequencing using natural genetic variation,” Nat Biotechnol. 36(1):89, or scds, Bais and Kostka, 2019, “scds: computational annotation of doublets in single cell RNA sequencing data,” bioRxiv. 2019564021 available on the Internet at biorxiv.org / content / 10.1101 / 564021vl, is used to remove doublets or multiplets in the single-nucleus or single-cell transcriptome data. These programs model gene expression from the original data, then assign doublet to barcodes that have observed expression from genes that are likely to not occur simultaneously. In some embodiments any known methodis used to remove doublets or multiplets in the single-nucleus or single-cell transcriptome data.
[0229] In some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cell are filtered to remove empty droplets or droplets representing damaged nuclei or cells. See Lun et al., 2019, “EmptyDrops: distinguishing cells from empty droplets in droplet-based single-cell RNA sequencing data,” Genome Biol 20, 63; and Heiser etal., 2021, “Automated quality control and cell identification of droplet-based single-cell data using dropkick,” Genome Res 31, 1742-1752. In some embodiments, the single-cell or single-nucleus RNA sequencing utilizes abundantly more cell barcodes than the number of targeted cells to ensure single-cell capture per droplet. This results in a relatively small number of droplets that capture real cells or nuclei as most droplets are “empty droplets” and do not contain real cells in such embodiments. However, ambient material created during sample processing can also contain transcripts called “ambient RNA” that can be captured by empty droplets, making them appear non-empty in data analysis. As a result, separation of empty droplets from real cells becomes a significant problem in single-cell RNA-seq analysis. Since real cells or nuclei contain more transcripts (and more unique molecular identifiers -UMIs- that denote unique reads) than ambient RNAs captured in empty droplets, it is possible to apply a cutoff based on number of UMIs to retain real cells or nuclei. However, this can be inaccurate as UMI distribution of cell barcodes is not completely discrete which makes a hard cutoff arbitrary. Moreover, certain cell types may be transcriptomically more silent than others which could lead to filtering out by a UMI-based cutoff. In some embodiments this problem of distinguishing real cells or nuclei and empty droplets is addressed by using other metrics such as expression profile and nuclear fraction.2,4, 5 In some such embodiments DropletQC, Muskovic and Powel, 2021, “DropletQC: improved identification of empty droplets and damaged cells in single-cell RNA-seq data.,” Genome Biol. 22, 329, is used to remove empty droplets or droplets representing damaged nuclei or cells in the single-nucleus or single-cell transcriptome data. In some such embodiments and known method is used to remove empty droplets or droplets representing damaged nuclei or cells in the single-nucleus or single-cell transcriptome data.
[0230] Referring to block 230, in some embodiments, the first myelofibrosis-related state is myelofibrosis and the second myelofibrosis-related state is absence of myelofibrosis.
[0231] Referring to block 230, in some embodiments, the first myelofibrosis-related state is absence of myelofibrosis and the second myelofibrosis-related state is myelofibrosis.
[0232] Myelofibrosis (MF) is a chronic, rare, and fatal myeloproliferative neoplasm characterized by a progressive replacement of bone marrow with fibrous scar tissue. Secondary primary malignancies represent a major cause of morbidity and mortality in MF patients.
[0233] Referring to block 234, in some embodiments, the first myelofibrosis-related state is a first stage of myelofibrosis and the second myelofibrosis-related state is a second stage of myelofibrosis.
[0234] Referring to block 236, in some embodiments, the myelofibrosis-related state is low risk myelofibrosis and the second myelofibrosis-related state is intermediate- 1 risk myelofibrosis, in accordance with the International Prognostic Scoring System (IPSS). See, Cervantes et al., 2009, “New prognostic scoring system for primary myelofibrosis based on a study of the International Working Group for Myelofibrosis Research and Treatment,” Blood 113(13), pp. 2895-901, which is hereby incorporated by reference.
[0235] In accordance with IPSS:
[0236] Low Risk: patients in this stage have relatively mild symptoms and a good prognosis. Their survival rate is typically longer, and they may require minimal treatment.
[0237] Intermediate- 1 Risk: This stage includes patients with slightly more advanced disease, but their prognosis is still relatively favorable. They may experience moderate symptoms, and treatment strategies aim to manage these symptoms and maintain quality of life.
[0238] Intermediate-2 Risk: Patients in this stage have more severe symptoms and a less favorable prognosis. They may require more aggressive treatment approaches to manage their symptoms and improve their quality of life.
[0239] High Risk: Patients in this stage have the most advanced form of myelofibrosis, with severe symptoms and a poor prognosis. Treatment options focus on symptom management and potentially more intensive interventions, such as stem cell transplantation, if the patient is eligible.
[0240] Referring to block 238, in some embodiments, the first myelofibrosis-related state is intermediate- 1 risk myelofibrosis and the second myelofibrosis-related state is intermediate-2 risk myelofibrosis.
[0241] Referring to block 240, in some embodiments, the first myelofibrosis-related state is intermediate-2 risk myelofibrosis and the second myelofibrosis-related state is high risk myelofibrosis.
[0242] Referring to block 242, in some embodiments, the first myelofibrosis-related state is low risk or intermediate- 1 risk myelofibrosis and the second myelofibrosis-related state is intermediate-2 risk myelofibrosis or high risk myelofibrosis.
[0243] Referring to block 244, in some embodiments, the first myelofibrosis-related state is low risk myelofibrosis and the second myelofibrosis-related state is absence of myelofibrosis.
[0244] Referring to block 246, in some embodiments, first myelofibrosis-related state is low risk myelofibrosis and the second myelofibrosis-related state is intermediate- 1 risk or intermediate-2 risk myelofibrosis.
[0245] Referring to block 248, in some embodiments, the first information further comprises second metadata for each subject in the cohort of subjects comprising age, body mass index, and / or sex. The second information is used to prune the cohort of subjects based on age, body mass index, and / or sex. This causes the cohort of subjects to be free of confounding for age, body mass index, and / or sex. Referring to block 250, in some embodiments, the first information further comprises second metadata for each subject in the cohort of subjects comprising age, body mass index, and sex. The second information is used to prune the cohort of subjects based on age, body mass index, and sex. This causes the cohort of subjects to be free of confounding for age, body mass index, and sex. The pruning causes a subset of subjects to be removed from the cohort of subjects. In some embodiments the pruning balances the cohort of subjects for age, body mass index, and / or sex as discussed above in conjunction with block 226.
[0246] In some embodiments, the pruning of block 248 or 250 causes a subset of subjects to be removed from the cohort of subjects. In some embodiments the pruning balances the cohort of subjects for age, body mass index (BMI), and / or sex as discussed above in conjunction with block 226.
[0247] Referring to block 252, in some embodiments, the respective single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is barcoded with the subject in the cohort of subject originating the respective single-nucleus or single-cell transcriptome data.
[0248] In some embodiments a unique barcode for a respective single nucleus or cell is provided in the form of an oligonucleotide that comprises a nucleic acid barcode sequence that is attached to the nucleic acid molecules originating from the respective single nucleus or cell. The oligonucleotide is partitioned such that as between the nucleic acid molecules originating from the respective single nucleus or cell, the nucleic acid barcode sequence is the same, but as between nucleic acids originating from other nuclei or cells, the barcode can, and preferably have differing sequences. In preferred embodiments, only one nucleic acid barcode sequence is associated with a given nuclei or cell, although in some embodiments, two or more different barcode sequences are associated with a given nucleus or cell.
[0249] The nucleic acid barcode sequences will typically include from 6 to about 20 or more nucleotides within the sequence of the barcode. In some embodiments, these nucleotides are completely contiguous, e.g., in a single stretch of adjacent nucleotides. In alternative embodiments, they are separated into two or more separate subsequences that are separated by one or more nucleotides. Typically, separated subsequences are separated by about 4 to about 16 intervening nucleotides.
[0250] In some embodiments, each barcode for each nucleus or cell encodes a unique predetermined value selected from the set { 1, ..., 1024}, { 1, ..., 4096}, { 1, ..., 16384}, { 1, ..., 65536}, { 1, ..., 262144}, { 1, ..., 1048576}, { 1, ..., 4194304}, { 1, ..., 16777216}, { 1, ..., 67108864}, or { 1, ..., 1 x 1012}. For instance, consider the case in which the barcode sequence is represented by a set of five nucleotide positions. In this instance, each nucleotide position contributes four possibilities (A, T, C or G), giving rise, when all five positions are considered, to 4 x 4 x 4 x 4 x 4 = 1024 possibilities. As such, the five nucleotide positions form the basis of the set { 1,..., 1024}. In other words, when the barcode sequence is a 5-mer, the barcode encodes a unique predetermined value selected from the set { 1,..., 1024}.Likewise, when the barcode sequence is represented by a set of six nucleotide positions, the six nucleotide positions collectively contribute 4 x 4 x 4 x 4 x 4 x 4 = 4096 possibilities. As such, the six nucleotide positions form the basis of the set { 1, ... , 4096} . In other words, when the barcode sequence is a 6-mer, the barcode encodes a unique predetermined value selected from the set { 1,..., 4096}.
[0251] By contrast, in some embodiments, the barcode of a sequence read in the plurality of sequence reads is localized to a noncontiguous set of oligonucleotides within the sequence read. In one such exemplary embodiment, the predetermined noncontiguous set of nucleotides collectively consists of N nucleotides, where N is an integer in the set {4, . . ., 20}. As an example, in some embodiments, a barcode sequence comprises a first set of contiguousnucleotide positions at a first position in an oligonucleotide tag and a second set of contiguous nucleotide positions at a second position in an oligonucleotide tag, that is displaced from the first set of contiguous nucleotide positions by a spacer. In one specific example, the barcode sequence comprises (Xl)nYz(X2)m, where XI is n contiguous nucleotide positions, Y is a constant predetermined set of z contiguous nucleotide positions, and X2 is m contiguous nucleotide positions. In this example, the barcode in the second portion of a sequence read produced by a schema invoking this exemplary barcode is localized to a noncontiguous set of oligonucleotides, namely (Xl)nand (X2)m. This is just one of many examples of noncontiguous formats for a barcode.
[0252] Further discussion of the use of barcodes in single-nucleus RNA-seq is discussed in Jiang et al., February 23, 2023, “Isolated nuclei from frozen tissue are the superior source for single cell RNA-seq compared with whole cells,” available on the Internet at doi.org / 10.1101 / 2023.02.19.529150, which is hereby incorporated by reference. In some embodiments a 10X Genomics Chromium 3’ gene expression assay is used, and the barcoding is performed as part of this assay. See Jiang et al., Id, which is hereby incorporated by reference.
[0253] Referring to block 254, in some embodiments, the first plurality of nuclei or cells is clustered into a plurality of clusters by (i) computing a plurality of distances using the singlenucleus or single-cell transcriptome data for the plurality of genes for each unique pair of nuclei or cells in the first plurality of nuclei or cells and (ii) evaluating the plurality of distances with a criterion function.
[0254] Referring to block 256, in some embodiments, the plurality of distances includes a separate distance for each unique pair of nuclei or cells in the first plurality of nuclei or cells. Each respective distance in the plurality of distances represents a different pair of nuclei or cells in the first plurality of nuclei or cells and quantifies a distance between (i) a respective first vector formed by the single-nucleus or single-cell transcriptome data for the plurality of genes for a respective first nucleus or cell in the different pair of nuclei or cells and (ii) a respective second vector formed by the single-nucleus transcriptome data for the plurality of genes for a respective second nucleus or cell in the different pair of nuclei or cells, and each respective cluster in the plurality of clusters represents a corresponding subset of nuclei or cells of the first plurality of nuclei or cells that are clustered together based on evaluation of distances in the plurality of distances representing different pairs of nuclei or cells within the corresponding subset of nuclei or cells with the criterion function.
[0255] Clustering is described at pages 211-256 of Duda and Hart, Pattern Classification and Scene Analysis, 1973, John Wiley & Sons, Inc., New York, (hereinafter “Duda 1973”) which is hereby incorporated by reference in its entirety. As described in Section 6.7 of Duda 1973, the clustering problem is described as one of finding natural groupings in a dataset. To identify natural groupings, two issues are addressed. First, a way to measure similarity (or dissimilarity) between two nuclei or cells is determined. This metric (similarity measure) is used to ensure that the nuclei or cells in one cluster are more like one another than they are to nuclei or cells in other clusters based on their gene expression. Second, a mechanism for partitioning the nuclei or cells into clusters using the similarity measure is determined.
[0256] Similarity measures are discussed in Section 6.7 of Duda 1973, where it is stated that one way to begin a clustering investigation is to define a distance function and to compute the matrix of distances between all pairs of nuclei or cells in a training set. If distance is a good measure of similarity, then the distance between reference nuclei or cells in the same cluster will be significantly less than the distance between the reference nuclei or cells in different clusters. However, as stated on page 215 of Duda 1973, clustering does not require the use of a distance metric. For example, a nonmetric similarity function s(x, x') can be used to compare two vectors x and x'. Conventionally, s(x, x') is a symmetric function whose value is large when x and x' are somehow “similar.” An example of a nonmetric similarity function s(x, x') is provided on page 218 of Duda 1973.
[0257] Once a method for measuring “similarity” or “dissimilarity” between vectors in a dataset has been selected, clustering requires a criterion function that measures the clustering quality of any partition of the data. Partitions of the data set that extremize the criterion function are used to cluster the data. See page 217 of Duda 1973. Criterion functions are discussed in Section 6.8 of Duda 1973.
[0258] More recently, Duda el al.. Pattern Classification, 2nd edition, John Wiley & Sons, Inc. New York, has been published. Pages 537-563 describe clustering in detail. More information on clustering techniques can be found in Kaufman and Rousseeuw, 1990, Finding Groups in Data: An Introduction to Cluster Analysis, Wiley, New York, N.Y.; Everitt, 1993, Cluster analysis (3d ed.), Wiley, New York, N.Y.; and Backer, 1995, Computer- Assisted Reasoning in Cluster Analysis, Prentice Hall, Upper Saddle River, New Jersey, each of which is hereby incorporated by reference. Particular exemplary clustering techniques that can be used in the present disclosure include, but are not limited to, hierarchical clustering (agglomerative clustering using nearest-neighbor algorithm, farthest- neighbor algorithm, the average linkage algorithm, the centroid algorithm, or the sum-of-squares algorithm), k-means clustering, fuzzy k-means clustering algorithm, Jarvis-Patrick clustering, Louvain clustering, or Leiden clustering. See, Blondel et al., July 25, 2008, “Fast unfolding of communities in large networks,” arXiv:0803.0476v2 [physical. coc-ph]; and Heumos et al., 2023, “Best practices for single-cell analysis across modalities,” Nature Review Genetics 24, 550-572, each of which is hereby incorporated by reference. Such clustering can be on the vectors formed directly from the single-nucleus or single-cell transcriptome data for the plurality of genes or vectors of principal components derived from such single-nuclei or single-cell transcriptome data. In some embodiments, the clustering comprises unsupervised clustering where no preconceived notion of what clusters should form when the training set is clustered are imposed.
[0259] In some embodiments before each respective first and second vector is generated, the single-nucleus or single-cell transcriptome data is subjected to principal component transformation to produce principal components. In such embodiments, the respective first vector formed by the single-nucleus or single-cell transcriptome data for the plurality of genes for a respective first nucleus or cell in the different pair of nuclei or cells is than the set of principal components derived for the respective first nucleus or cell. Likewise, the respective second vector formed by the single-nucleus or single-cell transcriptome data for the plurality of genes for a respective second nucleus or cell in the different pair of nuclei or cells is the set of principal components derived for the respective second nucleus or cell. Principal component analysis (PCA) algorithms that may be used to transform vectors of gene expression data to vectors of principal components are described in Jolliffe, 1986, Principal Component Analysis, Springer, New York, which is hereby incorporated by reference. PCA is also described in Draghici, 2003, Data Analysis Tools for DNA Microarrays, Chapman & Hall / CRC, which is hereby incorporated by reference. Principal components (PCs) are uncorrelated and are ordered such that the kth PC has the kthlargest variance among PCs. The kthPC can be interpreted as the direction that maximizes the variation of the projections of the data points such that it is orthogonal to the first k-1 PCs. The first few PCs capture most of the variation in dataset. In contrast, the last few PCs are often assumed to capture only the residual 'noise' in the dataset. In some embodiments, each respective vector that is clustered in accordance with block 254 or 256 corresponds to a nucleus or cell in the plurality of nuclei or cells and contains between four and five hundred principal components. In some embodiments, each respective vector that is clustered in accordance with block 254 or 256 corresponds to a nucleus or cell in the plurality of nuclei or cells and contains between five and six hundred principal components. In someembodiments, each respective vector that is clustered in accordance with block 254 or 256 corresponds to a nucleus or cell in the plurality of nuclei or cells and contains between six and second hundred principal components. In some embodiments, each respective vector that is clustered in accordance with block 254 or 256 corresponds to a nucleus or cell in the plurality of nuclei or cells and contains at least 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 principal components. In some embodiments, the principal component analysis is performed on the first plurality of genes using the Seurat function RunPCA. In some embodiments, clusters are identified with Seurat function FindClusters, optionally using the shared nearest neighbor modular optimization based on the first 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or between 20 and 500 principal components identified by principal component analysis. In some such embodiments the clustering resolution is set to a value between 0.2 and 0.6, such as 0.4. See, Stuart et al., 2019, “Comprehensive integration of single-cell data,” Cell 177, 1888-1902, which is hereby incorporated by reference.
[0260] In some embodiments the clustering comprises k-means clustering of the nuclei or cells into a predetermined number of clusters. The goal of k-means clustering is to cluster the nuclei or cells based upon either the original transcriptome data or the principal components derived from the original transcriptome data for the plurality of nuclei or cells into K partitions. In some embodiments, the k-means algorithm computes like clusters of entities from the higher dimensional data (where each dimension is either a different gene or a different principal component) and then after some resolution, the k-means clustering tries to minimize error. In this way, the k-means clustering provides cluster assignments.
[0261] In some embodiments, K is a number between 2 and 50 inclusive. In some embodiments, the number K is set to a predetermined number such as 10. In some embodiments, the number K is optimized. In some embodiments, the number K is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more than 30. In some embodiments, the number K is at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100.
[0262] In some embodiments, the clustering provides a suitable two or three dimensional clustering for visualization. Community based clustering, such as Louvain or Leiden clustering, are examples of such clustering. See, Heumos et al., 2023, “Best practices for single-cell analysis across modalities,” Nature Review Genetics 24, 550-572, which is hereby incorporated by reference. In some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes (where the number of dimensions are either the number of genes in the plurality of genes or, alternatively, in embodiments in which principalcomponents are used, the number of principal components) for the first plurality of nuclei or cells are subjected to dimension reduction, e.g., in accordance with a manifold, into two or three dimensions and the resulting points in the dimension reduced plot, each representing a respective nucleus or cell in the plurality of nuclei or cells are coded by their cluster assignments. Dimension reduction programs such as UMAP, t-SNE, and PHATE can be used for this purpose. See, Heumos et al., 2023, “Best practices for single-cell analysis across modalities,” Nature Review Genetics 24, 550-572. In either approach, the high dimensional nature of the single-nucleus or single-cell transcriptome data is reduced to either a two-dimensional or three-dimensional plot in which each point represents a different nucleus or cell in the plurality of nuclei or cells and the identity of which cluster each nucleus or cell is in can be further emphasized by color coding the nucleus or cell by a unique color associated with each cluster.
[0263] Referring to block 256, in some embodiments, the first metadata is then used to identify a first cluster in the plurality of clusters with the first myelofibrosis-related state by determining that the first cluster includes nuclei or cells from subjects in the cohort of subjects that have the first myelofibrosis-related state.
[0264] In some embodiments gene markers set forth in Table 2 below are used to identify cell types set forth in Table 2 below. In some embodiments, for a respective cell type set forth in Table 2, at least 1, 2, 3, 4, 5, 6, 7, or 8 of the gene markers corresponding to the respective cell type set forth in Table 2 below are used to identify the respective cell type.
[0265] Table 2. Cell Types and associated Gene Markers
[0266] Identification of a cluster associated with a first myelofibrosis-related state allows for advantageous segmentation of myelofibrosis-related abnormalities into meaningful subcategories (myelofibrosis-related states) based on relevant and actual differences in transcriptome expression between the different subcategories, among other practical applications.
[0267] Moreover, the identification of such a cluster can be used to determine whether any of the information in the first metadata is a covariate with respect to the first myelofibrosis- related state. For instance, the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the identified cluster, or for the entire first plurality of nuclei or cells, can be fitted to a varying coefficient model to determine to examine the effect various covariates have on the first myelofibrosis-related state. See, Hastie et al. , 2001, The Elements of Statistical Learning, Data Mining, Inference, and Prediction,, Springer- Verlag, New York, New York, Section 6.4.2 beginning at p. 177, which is hereby incorporated by reference.
[0268] As an example, referring to block 258, in some such embodiments, the first cluster is used to determine the extent to which race affects (is a covariate for) whether or not a subject incurs the first myelofibrosis-related state and / or the degree they incur the first myelofibrosis- related state.
[0269] As another example, referring to block 260, in some such embodiments, the first cluster is used to determine the extent to which sex affects (is a covariate for) whether or not a subject incurs the first myelofibrosis-related state and / or the degree they incur the first myelofibrosis-related state.
[0270] As still another example, referring to block 262, in some such embodiments, the first cluster is used to determine the extent to which age affects (is a covariate for) whether or not a subject incurs the first myelofibrosis-related state and / or the degree they incur the first myelofibrosis-related state.
[0271] As still another example, referring to block 264, in some such embodiments, the first cluster is used to determine the extent to which a particular biomarker, such as the expression of a biomarker in the first standardized set of biomarkers of block 224, affects (is a covariate for) whether or not a subject incurs the first myelofibrosis-related state and / or the degree they incur the first myelofibrosis-related state. For instance, whether an abundance (expression level) of the particular biomarker affects (is a covariate for) whether or not a subject incurs the first myelofibrosis-related state and / or the degree they incur the first myelofibrosis-related state.
[0272] As still another example, referring to block 266, in some embodiments, the first metadata includes one or more second biomarker annotations for each nucleus or cell in the plurality of nuclei or cells drawn from a second standardized set of biomarkers that each define a unique cell type at a unique stage.
[0273] In some such embodiments the first cluster is used to determine an extent to which a second biomarker, in the second standardized set of biomarkers, affects (is a covariate for) whether or not a subject incurs the first myelofibrosis-related state and / or the degree they incur the first myelofibrosis-related state. For instance, whether an abundance (expression level) of the second biomarker affects (is a covariate for) whether or not a subject incurs the first myelofibrosis-related state and / or the degree they incur the first myelofibrosis-related state.
[0274] Block 268 details an additional practical application for the discovery of the first cluster as detailed in block 257. In some such embodiments, genotype data for each subject in the plurality of subjects is obtained and this genotype data is overlay ed in each subject represented in the first cluster with a myelofibrosis-related state of each subject in the first cluster. Alternatively or additionally, such genotype data case is overlay ed with each subject represented in any cluster. Since each point in a cluster represents a nuclei or a cell from aparticular subject on the cohort, and the identity of the subject is known because of the barcoding, each point in the cluster can be coded by a genotypic status of the corresponding subject. For instance, in some embodiments the genotypic status is the absence or presence of a major allele for a single nucleotide polymorphism (SNP) at a particular genetic locus in a genome. In one such embodiment, the nuclei or cell in the cluster that are from subjects that have the major allele for the SNP are indicated with a first color and the nuclei or cells in the cluster that are from subjects that have the minor allele for the SNP are indicated with a second color. In another example, in some embodiments the genotypic status is the absence or presence of a particular restriction fragment length polymorphism (RFLP) at a particular genetic locus in a genome. In one such embodiment, the nuclei or cells in the cluster that are from subjects that have the RFLP (or have a first allele for the RFLP) are indicated with a first color and the nuclei or cells in the cluster that are from subjects that do not have the RFLP (or have a second allele for the RFLP) are indicated with a second color. In alternative embodiments, rather than being a SNP or RFLP, the genotypic data overlayed on the first cluster is the allelic status of the cohort of subjects for particular copy number variations, insertions, or deletions. That is, for each subject having a respective nucleus or cell in the first cluster, the respective nucleus or cell is labeled by the allelic status of a particular copy number variation, a particular genetic insertion, or a particular genetic deletion that arises in the genome of the species of the cohort of subjects. In alternative embodiments, the genotypic data overlayed on the first cluster is the allelic status of the cohort of subjects for a particular haplotype. That is, for each subject having a respective nucleus or cell in the first cluster, the respective nucleus or cell is labeled by the haplotype of a particular haplotype arising in the genome of the species of the cohort of subjects. In alternative embodiments, the genotypic data overlayed on the first cluster is the allelic status of the cohort of subjects for a particular microsatellite or short tandem repeat. That is, for each subject having a respective nucleus or cell in the first cluster, the respective nucleus or cell is labeled by the allelic status of a particular microsatellite or short tandem repeat arising in the genome of the species of the cohort of subjects.
[0275] Referring to block 270, in some such embodiments, genotype data is obtained for each subject in the plurality of subjects, and upon overlay on the clustered transcriptome data, the clustered transcriptome data (e.g., the first cluster, 2 or more of the clusters, etc.) is used to determine an extent to which a genotype affects (is a covariate for) whether or not a subject incurs the first myelofibrosis-related state and / or the degree they incur the first myelofibrosis-related state.
[0276] Referring to block 272, in some embodiments, the myelofibrosis-related detection system is used to associate a test subject with the first myelofibrosis-related state. In some such embodiments, second information is obtained that comprises single-nucleus or singlecell transcriptome data for the plurality of genes for each nucleus or cell in a second plurality of nuclei or cells. Each nucleus or cell in the second plurality of nuclei or cells is obtained from a sample obtained from the test subject. The transcriptome data is co-clustered with the transcriptome data of the first plurality of nuclei or cells into the plurality of clusters. In instances where at least some of the nuclei or cells cluster into a particular cluster that is known from the metadata for the first plurality of nuclei or cells to be associated with a first myelofibrosis-related state, some embodiments of the present disclosure associate the test subject with this first myelofibrosis-related state.
[0277] Referring to block 274, in some embodiments, the method informs a response to a drug compound in a patient or in a plurality of patients. Referring to block 276, in some embodiments, the method informs a response to a dosing amount, duration, and / or frequency of a drug in a patient or in a plurality of patients.
[0278] Referring to block 278, in some embodiments, the first metadata is used to identify a second cluster in the plurality of clusters with the second myelofibrosis-related state by determining that the second cluster includes cells or nuclei from subjects in the cohort that have the second myelofibrosis-related state.
[0279] Referring to block 280, in some embodiments, the first myelofibrosis-related state is a first stage of myelofibrosis and the second myelofibrosis-related state is a second stage of myelofibrosis. In some embodiments, at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the cells or nuclei in the first cluster are cells or nuclei from subjects in the cohort of subjects that have the first stage of myelofibrosis and at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the cells or nuclei in the second cluster are cells or nuclei from subjects in the cohort of subjects that have the second stage of myelofibrosis.
[0280] Referring to block 282, in some embodiments, a determination is made that the first cluster comprises basophil / mast cells or nuclei of basophil / mast cells and the second cluster comprises cells or nuclei of cells other than basophil / mast cells. In some embodiments, at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the nuclei or cells in the first cluster are basophil / mast cellsor nuclei of basophil / mast cells. In some embodiments, at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the nuclei or cells in the second cluster are other than basophil / mast cells or are nuclei from other than basophil / mast cells.
[0281] Referring to block 284, in some embodiments, a determination is made that the first cluster comprises monocytes or nuclei of monocytes and the second cluster comprises other than monocytes or nuclei of other than monocytes. In some embodiments, at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the nuclei or cells in the first cluster are monocytes or nuclei of monocytes. In some embodiments at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the nuclei or cells in the second cluster are other than monocytes or nuclei of other than monocytes.
[0282] Referring to block 286, in some embodiments, a determination is made that the first cluster comprises erythroid lineage cells or nuclei of erythroid lineage cells and the second cluster comprises other than erythroid lineage cells or nuclei of other than erythroid lineage cells. In some embodiments, at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the nuclei or cells in the first cluster are erythroid lineage cells or nuclei of erythroid lineage cells. In some embodiments at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the nuclei or cells in the second cluster are other than erythroid lineage cells or nuclei of other than erythroid lineage cells.
[0283] Referring to block 288, in some embodiments, a determination is made that the first cluster comprises hematopoietic precursor cells or nuclei of hematopoietic precursor cells and the second cluster comprises other than hematopoietic precursor cells or nuclei of other than hematopoietic precursor cells. In some embodiments, at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the cells or nuclei in the first cluster are hematopoietic precursor cells or nuclei of hematopoietic precursor cells. In some embodiments at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the cells or nuclei in the second cluster are other than hematopoietic precursor cells or nuclei of other than hematopoietic precursor cells.
[0284] Referring to block 290, in some embodiments, a determination is made that the first cluster comprises lymphoid lineage cells or nuclei of lymphoid lineage cells and the secondcluster comprises other than lymphoid lineage cells or nuclei of other than lymphoid lineage cells. In some embodiments, at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the cells or nuclei in the first cluster are lymphoid lineage cells or nuclei of lymphoid lineage cells. In some embodiments at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the cells or nuclei in the second cluster are other than lymphoid lineage cells or nuclei of other than lymphoid lineage cells.
[0285] Referring to block 292, in some embodiments, a further determination is made that the first cluster comprises megakaryocyte-erythroid progenitor cells or nuclei of megakaryocyte-erythroid progenitor cells and the second cluster comprises other than megakaryocyte-erythroid progenitor cells or nuclei of other than megakaryocyte-erythroid progenitor cells. In some embodiments, at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the cells or nuclei in the first cluster are megakaryocyte-erythroid progenitor cells or nuclei of megakaryocyte-erythroid progenitor cells. In some embodiments at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the nuclei or cells in the second cluster are other than megakaryocyte-erythroid progenitor cells or nuclei of other than megakaryocyte-erythroid progenitor cells.
[0286] Referring to block 294, in some embodiments, a further determination is made that the first cluster comprises megakaryocyte cells or nuclei of megakaryocyte cells and the second cluster comprises other than megakaryocyte cells or nuclei of other than megakaryocyte cells. In some embodiments, at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the cells or nuclei in the first cluster are megakaryocyte cells or nuclei of megakaryocyte cells. In some embodiments at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the nuclei or cells in the second cluster are other than megakaryocyte cells or nuclei of other than megakaryocyte cells.
[0287] In some embodiments, a further determination is made that the first cluster comprises myeloid lineage cells or nuclei of myeloid lineage cells and the second cluster comprises other than myeloid lineage cells or nuclei of other than myeloid lineage cells. In some embodiments, at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the cells or nuclei in the first cluster are myeloid lineage cells or nuclei of myeloid lineage cells. In some embodiments at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent orat least 90 percent of the nuclei or cells in the second cluster are other than myeloid lineage cells or nuclei of other than myeloid lineage cells.
[0288] Referring to block 296, in some embodiments, the first cluster represents a myelofibrosis disease state and the second cluster represents a healthy state, and the method further comprises identifying a metabolic pathway comprising a set of genes whose activity discriminates between the two states. One or more genes in the set of genes are overexpressed or under-expressed in the first cluster relative to the second cluster. Any art excepted definition of overexpression and under-expression can be used to identify genes that are overexpressed or under-expressed in the first cluster relative to the second cluster. In some embodiments a measure of central tendency of each gene in the plurality of genes is determined across the nuclei or cells in the first cluster as well as in the nuclei or cells in the second cluster. Examples of measure of central tendency include e.g., mean, median, mode, weighted mean, weighted median, and / or weighted mode. Then the fold-change for each respective gene is determined between the measure of central tendency for the respective gene across the first cluster versus across the second cluster. A fold change represents how many times the expression level of a gene has increased in the first cluster compared to the second cluster. In some embodiments a gene is considered overexpressed in the first cluster relative to the second cluster if its fold change (measure of central tendency across the first cluster divided by measure of central tendency across the second cluster) exceeds a predetermined threshold (e.g., 2 times higher, 2.5 times higher, 3 times higher, etc . In some embodiments a gene is considered under-expressed in the first cluster relative to the second cluster if its fold change (measure of central tendency across the first cluster divided by measure of central tendency across the second cluster) is below a predetermined threshold (e.g., below 0.5 times, below .25 times, below 0.10 times).
[0289] In some embodiments a gene is considered overexpressed in the first cluster relative to the second cluster if a t-test or analysis of variance (ANOVA) test determines that differences in expression levels of the gene are statistically significant between the first cluster and the second cluster. Genes whose expression level are higher in the first cluster relative to the second cluster where a t-test or ANOVA indicates this higher expression has a p-value below a first threshold value are considered over-expressed in the first cluster relative to the second cluster in some embodiments. In some embodiments the first threshold is 0.10, 0.05, or 0.01. Genes whose expression level are lower in the first cluster relative to the second cluster where a t-test or ANOVA indicates this lower expression has a p-value below a second threshold value are considered under-expressed in the first cluster relative to thesecond cluster in some embodiments. In some embodiments the second threshold is 0.10, 0.05, or 0.01.
[0290] In some embodiments one or more genes are overexpressed in the first cluster relative to the second cluster while, at the same time one or more genes are under-expressed in the first cluster relative to the second cluster.
[0291] In some embodiments, the first and second clusters are further filtered to select a single cell type in the first cluster and a single cell type in the second cluster before performing differential expression analysis to find one or more genes that are overexpressed in the first cluster relative to the second cluster and / or one or more genes that are underexpressed in the first cluster relative to the second cluster In some embodiments selection of such cell types within the cluster is done by looking for expression of a respective gene signature in such cells that is characteristic of such cell types. The gene signature for many such cell types are known in the art and thus the filtering of the clusters for many specific cell types is done using conventional methods in such embodiments.
[0292] In some embodiments, the differential expression analysis identifies, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more genes that are over-expressed in the first cluster relative to the second cluster. In some embodiments, the differential expression analysis identifies, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more genes that are under-expressed in the first cluster relative to the second cluster. For instance, in some embodiments, using thresholds for differential expression, gene sets are derived that represent the most significant over-expressed and under-expressed genes.
[0293] In some embodiments, differentially expressed genes are identified using the Seurat FindAllMarkers function. See, Stuart et al., 2019, “Comprehensive integration of single-cell data,” Cell 177, 1888-1902, which is hereby incorporated by reference.
[0294] In some embodiments, these over-expressed and under-expressed genes (differential gene set expression signatures) are then characterized using algorithms that measure statistical enrichment for genes in particular pathways, with particular functions or with particular structural characteristics attained from publicly available databases. In some embodiments the statistical significance of enrichment is determined using a hypergeometric distribution or equivalently a one-tailed version of Fisher’s exact test. This and other methods for determining over-expressed and under-expressed genes in the differential expression analysis are disclosed in Plaisier et al., 2010, “Rank-rank hypergeometric overlap:identification of statistically significant overlap between gene-expression signatures,” Nucleic Acids Research 38(17): el69, which is hereby incorporated by reference.
[0295] In some embodiments an examination of publicly curated databases is performed with each of the genes that are over-expressed and / or under-expressed in the first cluster relative to the second cluster to identify a metabolic pathway comprising a set of genes. Examples of such publicly curated databases is provided in Table 3, below.
[0296] Table 3 - publicly curated databases:
[0297] Each of the libraries listed in Table 3 is downloadable from the Internet at maayanlab.cloud / Enrichr / index.jsp#libraries. Another database where such pathways is found is MsigDB (v7.4). See Liberzon et al., 2015, The Molecular Signatures Database (MSigDB) hallmark gene set collection,” Cell Syst 1 : 417-425; and Liberzon et al., 2011, “Molecular signatures database (MSigDB) 3.0,” Bioinformatics 27: 1739-1740, each of which is hereby incorporated by reference.
[0298] Referring to block 296, in some embodiments, the first cluster represents a myelofibrosis disease state and the second cluster represents a healthy state. In some such embodiments a plurality of compound-specific differential transcriptional signatures is accessed. Each respective compound-specific differential transcriptional signature is a difference between (i) a respective compound-treated transcriptional signature in a plurality of compound-treated transcriptional signatures and (ii) a baseline transcriptional signature data set. The baseline transcriptional signature data set is from a control sample of one ormore cells of a cell type (e.g., a cell type whose expression signature serves as a good proxy for the cells found in the first cluster or the second cluster). Each respective compound- treated transcriptional signature is from a corresponding compound-treated sample of one or more cells of the cell type separately treated with a different compound in a plurality of at least 10 compounds. An example of compound-specific differential transcriptional signatures is disclosed in United States Patent Application No. 63 / 580,612, entitled “Cellular Data Library and Methods of Using Same,” filed September 5, 2023, which is hereby incorporated by reference. A test differential transcription signature is generated from a differential comparison of the transcriptional signature of the nuclei or cells of the first cluster and the second cluster. In some embodiments, the differential comparison requires that the baseline transcriptional signature data set and each respective compound-treated transcriptional signature be run two, three, four times or more. In some embodiments, DESeq2, or an analogous algorithm, is used to determine the differential transcriptional signature. See, Love et al, 2014, “Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2,” which is hereby incorporated by reference. In some embodiments, edgeR, or an analogous algorithm, is used to determine the differential transcriptional signature. See, Robinson and Smyth, 2007, “Moderated statistical tests for assessing differences in tag abundance,” Bioinformatics. 2007, 23: 2881-2887; and McCarthy et al., 2012, “Differential expression analysis of multifactor RNA-seq experiments with respect to biological variation,” Nucleic Acids Res. 40: 4288-4297, each of which is hereby incorporated by reference. In some embodiments, BBSeq, or an analogous algorithm, is used to determine the differential transcriptional signature. See, Zhou, 2011, “powerful and flexible approach to the analysis of RNA sequence count data,” Bioinformatics 27, pp. 2672-2678, which is hereby incorporated by reference. In some embodiments, DSS, or an analogous algorithm, is used to determine the differential transcriptional signature. See, Wang, 2013, “A new shrinkage estimator for dispersion improves differential expression detection in RNA-seq data,” Biostatistics 14: 232-243, which is hereby incorporated by reference. In some embodiments, baySeq, or an analogous algorithm, is used to determine the differential transcriptional signature. See, Hardcastle and Kelly, 2010, “baySeq: empirical Bayesian methods for identifying differential expression in sequence count data,” BMC Bioinformatics 11 : 422, which is hereby incorporated by reference. In some embodiments, ShrinkBayes [, or an analogous algorithm, is used to determine the differential transcriptional signature. See, Van De Wiel, 2013, “Bayesian analysis of RNA sequencing data by estimating multiple shrinkage priors,”Biostatistics 14: 113-128, which is hereby incorporated by reference.
[0299] The test differential transcription signature is compared to each respective compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures, thereby identifying a first compound associated with a compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures that matches the test differential transcription signature. In some embodiments, each of the compound-specific differential transcriptional signatures is compared to the test differential transcription signature using a program such as Rank-rank Hypergeometric Overlap (RRHO), or an equivalent algorithm. See, Plaiser, 2010, “Rankrank hypergeometric overlap: identification of statistically significant overlap between geneexpression signatures,” Nucleic Acids Research 38(17), el 69, which is hereby incorporated by reference. In some embodiments, a compound-specific differential transcriptional signature is considered to match the test differential transcription signature when a program such as RRHO identifies the compound-specific differential transcriptional signatures as being upon the top N matches among the plurality of compound-treated transcriptional signatures compared to the test differential transcription signature, where N is a positive integer (e.g., 1, 5, 10, 20, 100, etc.). For instance in embodiments where N is 5, a compoundspecific differential transcriptional signatures is considered to match the test differential transcription signature when a program such as RRHO identifies the compound-specific differential transcriptional signatures as being upon the top 5 matches among the plurality of compound-treated transcriptional signatures compared to the test differential transcription signature. In some embodiments, a statistical test (e.g., t-test, ANOVA) is used to compare the plurality of compound-treated transcriptional signatures to the test differential transcription signature and each compound-treated transcriptional signature that has statistically significant similarity (e.g., P-value less than 0.15, less than 0.10, less than 0.05, less than 0.01) to the test differential transcription signature is considered to match the test differential transcription signature.
[0300] In some embodiments the plurality of compound-treated transcriptional signatures includes at least 5, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 800, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 8000, at least 10,000, at least 20,000, at least 30,000, at least 50,000, at least 80,000, at least 100,000, at least 200,000, at least 500,000, at least 800,000, at least 1 million, or at least 2 million transcriptional signatures, each representing a different compound.
[0301] In some embodiments, the plurality of compound-treated transcriptional signatures includes no more than 10 million, no more than 5 million, no more than 1 million, no more than 500,000, no more than 100,000, no more than 50,000, no more than 10,000, no more than 8000, no more than 5000, no more than 2000, no more than 1000, no more than 800, no more than 500, no more than 200, or no more than 100 compound-treated transcriptional signatures, each representing a different compound. In some embodiments, the plurality of compound-treated transcriptional signatures consists of from 10 to 500, from 100 to 10,000, from 5000 to 200,000, or from 10,000 to 1 million compound-treated transcriptional signatures, each representing a different compound.
[0302] In some embodiments, the plurality of compound-treated transcriptional signatures is between 10 and 1 x 106compound-treated transcriptional signatures, each representing a different compound. In some embodiments, the plurality of compound-treated transcriptional signatures is between 100 and 100,000 compound-treated transcriptional signatures, each representing a different compound. In some embodiments, the plurality of compound-treated transcriptional signatures is between 1000 and 100,000 compound-treated transcriptional signatures, each representing a different compound.
[0303] In some embodiments, the method further comprises formulating the first compound for use alleviating the myelofibrosis disease state. In some embodiments formulating the first compound for use in a therapy comprises manufacturing a composition comprising the first compound and one or more excipients and / or one or more pharmaceutically acceptable carriers and / or one or more diluents.
[0304] Such excipients and / or carriers include all conventional solvents, dispersion media, fillers, solid carriers, coatings, antifungal and antibacterial agents, dermal penetration agents, surfactants, isotonic and absorption agents and the like. It will be understood that the compositions of the present disclosure may also include other supplementary physiologically active agents.
[0305] An exemplary carrier is pharmaceutically “acceptable” in the sense of being compatible with the other ingredients of the composition (e.g., the composition comprising the test chemical compound) and not injurious to a subject. The compositions may conveniently be presented in unit dosage form and may be prepared by any methods well known in the art of pharmacy. Such methods include the step of bringing into association the active ingredient with the carrier that constitutes one or more accessory ingredients. In general, the compositions are prepared by uniformly and intimately bringing into associationthe active ingredient with liquid carriers or finely divided solid carriers or both, and then, if necessary, shaping the product.
[0306] Exemplary compounds, compositions or combinations of the present disclosure (e.g., the first compound) formulated for intravenous, intramuscular or intraperitoneal administration, or a pharmaceutically acceptable salt, solvate or prodrug thereof may be administered by injection or infusion.
[0307] Injectables for such use can be prepared in conventional forms, either as a liquid solution or suspension or in a solid form suitable for preparation as a solution or suspension in a liquid prior to injection, or as an emulsion. Carriers can include, for example, water, saline (e.g., normal saline (NS), phosphate-buffered saline (PBS), balanced saline solution (BSS)), sodium lactate Ringer's solution, dextrose, glycerol, ethanol, and the like; and if desired, minor amounts of auxiliary substances, such as wetting or emulsifying agents, buffers, and the like can be added. Proper fluidity can be maintained, for example, by using a coating such as lecithin, by maintaining the required particle size in the case of dispersion and by using surfactants.
[0308] The compound, composition or combinations of the present disclosure (e.g., the first compound) may also be suitable for oral administration and may be presented as discrete units such as capsules, sachets or tablets each containing a predetermined amount of the active ingredient; as a powder or granules; as a solution or a suspension in an aqueous or nonaqueous liquid; or as an oil-in-water liquid emulsion or a water-in-oil liquid emulsion. The active ingredient may also be presented as a bolus, electuary or paste.
[0309] A tablet may be made by compression or molding, optionally with one or more accessory ingredients. Compressed tablets may be prepared by compressing in a suitable machine the active ingredient (e.g., the first compound) in a free-flowing form such as a powder or granules, optionally mixed with a binder (e.g., inert diluent, preservative disintegrant (e.g. sodium starch glycolate, cross-linked polyvinyl pyrrolidone, cross-linked sodium carboxymethyl cellulose) surface-active or dispersing agent). Molded tablets may be made by molding in a suitable machine a mixture of the powdered compound moistened with an inert liquid diluent. The tablets may optionally be coated or scored and may be formulated so as to provide slow or controlled release of the active ingredient therein using, for example, hydroxypropylmethyl cellulose in varying proportions to provide the desired release profile. Tablets may optionally be provided with an enteric coating, to provide release in parts of the gut other than the stomach.
[0310] The compound, composition or combinations of the present disclosure (e.g., the first compound) may be suitable for topical administration in the mouth including lozenges comprising the active ingredient in a flavored base, usually sucrose and acacia or tragacanth gum; pastilles comprising the active ingredient in an inert basis such as gelatine and glycerin, or sucrose and acacia gum; and mouthwashes comprising the active ingredient in a suitable liquid carrier.
[0311] The compound, composition or combinations of the present disclosure (e.g., the first compound) may be suitable for topical administration to the skin may comprise the compounds dissolved or suspended in any suitable carrier or base and may be in the form of lotions, gel, creams, pastes, ointments and the like. Suitable carriers include mineral oil, propylene glycol, polyoxyethylene, polyoxypropylene, emulsifying wax, sorbitan monostearate, polysorbate 60, cetyl esters wax, cetearyl alcohol, 2-octyldodecanol, benzyl alcohol and water. Transdermal patches may also be used to administer the compounds of the invention.
[0312] The compound, composition or combination of the present disclosure (e.g., the first compound) may be suitable for parenteral administration include aqueous and non-aqueous isotonic sterile injection solutions which may contain anti-oxidants, buffers, bactericides and solutes which render the compound, composition or combination isotonic with the blood of the intended recipient; and aqueous and non-aqueous sterile suspensions which may include suspending agents and thickening agents. The compound, composition or combination may be presented in unit-dose or multi-dose sealed containers, for example, ampoules and vials, and may be stored in a freeze-dried (lyophilized) condition requiring only the addition of the sterile liquid carrier, for example water for injections, immediately prior to use.Extemporaneous injection solutions and suspensions may be prepared from sterile powders, granules and tablets of the kind previously described.
[0313] It should be understood that in addition to the active ingredients particularly mentioned above, the composition or combination of this present disclosure (e.g., the first compound) may include other agents conventional in the art having regard to the type of composition or combination in question, for example, those suitable for oral administration may include such further agents as binders, sweeteners, thickeners, flavoring agents disintegrating agents, coating agents, preservatives, lubricants and / or time delay agents. Suitable sweeteners include sucrose, lactose, glucose, aspartame or saccharine. Suitable disintegrating agents include cornstarch, methylcellulose, polyvinylpyrrolidone, xanthan gum, bentonite, alginic acid or agar. Suitable flavoring agents include peppermint oil, oil ofwintergreen, cherry, orange or raspberry flavoring. Suitable coating agents include polymers or copolymers of acrylic acid and / or methacrylic acid and / or their esters, waxes, fatty alcohols, zein, shellac or gluten. Suitable preservatives include sodium benzoate, vitamin E, alpha-tocopherol, ascorbic acid, methyl paraben, propyl paraben or sodium bisulphite. Suitable lubricants include magnesium stearate, stearic acid, sodium oleate, sodium chloride or talc. Suitable time delay agents include glyceryl monostearate or glyceryl distearate.
[0314] Referring to block 298, in some embodiments, the myelofibrosis disease state is selected from low risk, intermediate- 1 risk, intermediate-2 risk, and high risk myelofibrosis.
[0315] Referring to block 300, in some embodiments, the control sample and each corresponding compound-treated sample is exposed to a solvent. In the case of the control sample, the cells are exposed to the solvent for a particular period of time and the solvent does not include a compound. In the case of the corresponding compound-treated sample, the ells are exposed to the solvent for a particular period of time with the compound dissolved in the solvent. Thus, in typical embodiments the solvent is the same solvent for the control sample and each corresponding compound-treated sample with the exception that the solvent for each corresponding compound-treated sample includes a particular compound. Nonlimiting examples of solvents include water or aqueous-based solvents, alcohol-based solvents (e.g., ethanol or isopropyl alcohol), an ether such as bi s(2-methoxy ethyl) ether, P- cyclodextrin, or a polar aprotic solvent such as dimethylsulfoxide (DMSO), dimethylformamide (DMF), N,N-dimethylacetamide (DMAc), tetrahydrofuran (THF) acetonitrile (CEECN), or mixtures thereof. In some embodiments, the solvent is a polar aprotic solvent. In embodiments, the solvent is dimethylsulfoxide (DMSO). In some embodiments, the solvent comprises DMSO.
[0316] Referring to block 302, in some embodiments, the control sample and each corresponding compound-treated sample is exposed to a polar aprotic solvent. In the case of the control sample, the cells are exposed to the polar aprotic solvent for a particular period of time and the polar aprotic solvent does not include a compound. In the case of the corresponding compound-treated sample, the compound is exposed to the polar aprotic solvent with the corresponding compound dissolved in it. Thus, in typical embodiments, the polar aprotic solvent is the same polar aprotic solvent for the control sample and each corresponding compound-treated sample with the exception that the solvent for each corresponding compound-treated sample includes a particular compound. Optionally, the polar aprotic solvent comprises a mixture of a polar aprotic solvent and water. In some such embodiments, the polar aprotic solvent concentration ranges from about 0.01 pM to about 10pM, about 0.1 pM to about 5 pM, or about 0.5 pM to about 2 pM. In some embodiments, the polar aprotic solvent concentration (in water) is about 0.01 pM, about 0.1 pM, about 0.5 pM, about 1.0 pM, about 5 pM, or about 10 pM. In some embodiments, the polar aprotic solvent (in water) is about 1 pM.
[0317] Referring to block 304, in some embodiments, the control sample and each corresponding compound-treated sample is exposed to a DMSO solvent. Optionally, the DMSO solvent comprises a mixture of DMSO and water. In the case of the control sample, the cells are exposed to the DMSO solvent for a particular period of time and the DMSO solvent does not include a compound. In the case of the corresponding compound-treated sample, the compound is exposed to the DMSO solvent with the corresponding compound dissolved in it. Thus, in typical embodiments, the DMSO solvent is the same DMSO solvent for the control sample and each corresponding compound-treated sample with the exception that the solvent for each corresponding compound-treated sample includes a particular compound. Optionally, the DMSO solvent comprises a mixture of a DMSO and water. In some such embodiments, the DMSO concentration in the DMSO solvent ranges from about 0.01 pM to about 10 pM, about 0.1 pM to about 5 pM, or about 0.5 pM to about 2 pM. In some embodiments, the DMSO concentration in the DMSO solvent is about 0.01 pM, about 0.1 pM, about 0.5 pM, about 1.0 pM, about 5 pM, or about 10 pM. In some embodiments, the DMSO concentration in the DMSO solvent is about 1 pM.
[0318] Referring to block 306, in some embodiments, each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nucleus or single-cell assay and / or single-cell assay data. Optionally, the single-nucleus assay and / or single-cell assay data is selected from single-cell ribonucleic acid (RNA) sequencing (scRNA-seq) data, single-nucleus RNA sequencing (snRNA-seq) data, scTag-seq data, single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data, CyTOF / SCoP data, E-MS / Abseq data, miRNA-seq data, CITE-seq data, or any combinations thereof.
[0319] Referring to block 308, in some embodiments, each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises or consists of single-cell RNA sequencing (scRNA-seq) data.
[0320] Referring to block 310, in some embodiments, each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises scRNA-seq data.
[0321] Referring to block 312, in some embodiments, each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures consists of scRNA-seq data.
[0322] Referring to block 314, in some embodiments, each corresponding compound- treated sample of one or more cells comprises basophils / mast cells, monocytes, CD14+ cells, erythroid lineage cells, hematopoietic precursor cells, lymphoid lineage cells, megakaryocyte-erythroid progenitor cells, megakaryocytes, or myeloid lineage cells.
[0323] Referring to block 316, in some embodiments, each corresponding compound- treated sample of one or more cells comprises basophils / mast cells.
[0324] Referring to block 318, in some embodiments, each corresponding compound- treated sample of one or more cells comprises monocytes.
[0325] Referring to block 320, in some embodiments, each corresponding compound- treated sample of one or more cells consists of basophils / mast cells.
[0326] Referring to block 322, in some embodiments, each corresponding compound- treated sample of one or more cells consists of monocytes.
[0327] Referring to block 324, in some embodiments, each corresponding compound- treated sample of one or more cells consists of erythroid lineage cells.
[0328] Referring to block 326, in some embodiments, each corresponding compound- treated sample of one or more cells comprises or consists of one or more cells from an organ, cells from a tissue, stem cells, human cells, cells from umbilical cord blood, cells from peripheral blood, bone marrow cells, cells from a solid tissue, differentiated cells, and / or cells from a cell line.
[0329] Referring to block 328, in some embodiments, each corresponding compound- treated sample of one or more cells is a frozen sample. In some embodiments, each corresponding compound-treated sample of one or more cells is an unfrozen sample.
[0330] Referring to block 330, in some embodiments, each respective compound in the plurality of compounds has a molecular weight of less than 2000 Daltons. In some embodiments, each respective compound in the plurality of compounds is inorganic or organic. In some embodiments, each respective compound in the plurality of compounds is an organic compound having a molecular weight of less than 2000 Daltons (Da). In some embodiments, each respective compound in the plurality of compounds has a molecular weight of at least 10 Da, at least 20 Da, at least 50 Da, at least 100 Da, at least 200 Da, atleast 500 Da, at least 1 kDa, at least 2 kDa, at least 3 kDa, at least 5 kDa, at least 10 kDa, at least 20 kDa, at least 30 kDa, at least 50 kDa, at least 100 kDa, or at least 500 kDa. In some embodiments, each respective compound in the plurality of compounds has a molecular weight of no more than 1000 kDa, no more than 500 kDa, no more than 100 kDa, no more than 50 kDa, no more than 10 kDa, no more than 5 kDa, no more than 2 kDa, no more than 1 kDa, no more than 500 Da, no more than 300 Da, no more than 100 Da, or no more than 50 Da. In some embodiments, each respective compound in the plurality of compounds has a molecular weight of from 10 Da to 900 Da, from 50 Da to 1000 Da, from 100 Da to 2000 Da, from 1 kDa to 10 kDa, from 5 kDa to 500 kDa, or from 100 kDa to 1000 kDa. In some embodiments, each respective compound in the plurality of compounds has a molecular weight that falls within another range starting no lower than 10 Daltons and ending no higher than 1000 kDa.
[0331] Referring to block 332, in some embodiments, each respective compound in the plurality of compounds satisfies at least three criteria of the Lipinski rule of five criteria, or optionally each of the Lipinski rule of five criteria. In some each respective compound in the plurality of compounds is an organic compound that satisfies two or more rules, three or more rules, or all four rules of the Lipinski's Rule of Five: (i) not more than five hydrogen bond donors (e.g., OH and NH groups), (ii) not more than ten hydrogen bond acceptors (e.g. N and O), (iii) a molecular weight under 500 Daltons, and (iv) a LogP under 5. The “Rule of Five” is so called because three of the four criteria involve the number five. See, Lipinski, 1997, Adv. Drug Del. Rev. 23, 3, which is hereby incorporated herein by reference in its entirety. In some embodiments, a respective compound of the present disclosure satisfies one or more criteria in addition to Lipinski's Rule of Five. For example, in some embodiments, a compound of the present disclosure has five or fewer aromatic rings, four or fewer aromatic rings, three or fewer aromatic rings, or two or fewer aromatic rings.
[0332] REFERENCES CITED AND ALTERNATIVE EMBODIMENTS
[0333] All references cited herein are incorporated herein by reference in their entirety and for all purposes to the same extent as if each individual publication or patent or patent application was specifically and individually indicated to be incorporated by reference in its entirety for all purposes.
[0334] The present invention can be implemented as a computer program product that includes a computer program mechanism embedded in a non-transitory computer readable storage medium. For instance, the computer program product could contain the programmodules shown in any combination of Figures 1-2. These program modules can be stored on a CD-ROM, DVD, magnetic disk storage product, or any other non-transitory computer readable data or program storage product.
[0335] The present description includes example systems, methods, techniques, instruction sequences, and computing machine program products that embody illustrative implementations. For purposes of explanation, numerous specific details are set forth in order to provide an understanding of various implementations of the inventive subject matter. It will be evident, however, to those skilled in the art that implementations of the inventive subject matter may be practiced without these specific details. In general, well-known instruction instances, protocols, structures and techniques have not been shown in detail.
[0336] Many modifications and variations of this invention can be made without departing from its spirit and scope, as will be apparent to those skilled in the art. The specific embodiments described herein are offered by way of example only. The embodiments were chosen and described in order to best explain the principles of the invention and its practical applications, to thereby enable others skilled in the art to best utilize the invention and various embodiments with various modifications as are suited to the particular use contemplated. The invention is to be limited only by the terms of the appended claims, along with the full scope of equivalents to which such claims are entitled.
Claims
WHAT IS CLAIMED IS:
1. A method for manufacturing a myelofibrosis-related detection system, the method comprising: at a computer system having one or more processors, and memory storing one or more programs for execution by the one or more processors: obtaining, in electronic form, first information comprising:(i) single-nucleus or single-cell transcriptome data for a plurality of genes for each nucleus or cell in a first plurality of nuclei or cells, wherein each nucleus or cell in the first plurality of nuclei or cells is obtained from a sample in a plurality of samples, each sample in the plurality of samples is from a different subject in a cohort of subjects, the first plurality of nuclei or cells includes a different subset of nuclei or cells from a sample from each subject in the cohort of subjects, and the first plurality of nuclei or cells comprises at least 1000 nuclei or cells, and(ii) first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has a first myelofibrosis-related state or a second myelofibrosis-related state, wherein at least a first subset of subjects in the cohort of subjects have the first myelofibrosis-related state and a second subset of subjects in the cohort of subjects have the second myelofibrosis- related state, and wherein the respective single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is barcoded with the subject in the cohort of subjects originating the respective single-nucleus or single-cell transcriptome data; clustering the first plurality of nuclei or cells into a plurality of clusters by (i) computing a plurality of distances using the single-nucleus or single-cell transcriptome data for the plurality of genes for each unique pair of nuclei or cells in the first plurality of nuclei or cells and (ii) evaluating the plurality of distances with a criterion function, wherein the plurality of distances includes a separate distance for each unique pair of nuclei or cells in the first plurality of nuclei or cells,each respective distance in the plurality of distances represents a different pair of nuclei or cells in the first plurality of nuclei or cells and quantifies a distance between (i) a respective first vector formed by the single-nucleus or singlecell transcriptome data for the plurality of genes for a respective first nucleus or first cell in the different pair of nuclei and (ii) a respective second vector formed by the single-nucleus or single-cell transcriptome data for the plurality of genes for a respective second nucleus or second cell in the different pair of nuclei or cells, and each respective cluster in the plurality of clusters represents a corresponding subset of nuclei or cells of the first plurality of nuclei or cells that are clustered together based on evaluation of distances in the plurality of distances representing different pairs of nuclei or cells within the corresponding subset of nuclei or cells with the criterion function; and using the first metadata to identify a first cluster in the plurality of clusters with the first myelofibrosis-related state by determining that the first cluster includes nuclei or cells from subjects in the cohort of subjects that have the first myelofibrosis-related state.
2. The method of claim 1, wherein the first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has the first myelofibrosis-related state or the second myelofibrosis-related state comprises a histologically graded disease status for the respective subject.
3. The method of claim 1, wherein the first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has the first myelofibrosis-related state or the second myelofibrosis-related state comprises a histologically graded disease status for the respective subject determined in accordance with a consistent, verified handling of each sample in the plurality of samples.
4. The method of any one of claims 1-3, wherein each sample in the plurality of samples is from a different subject in the cohort of subjects and the plurality of samples comprises 5 samples, 20 samples, 50 samples, or 100 or more samples.
5. The method of any one of claims 1-4, wherein the first metadata for each respective subject in the cohort of subjects further indicates one or more features of the subject selected from the group consisting of sex, race, age, and physical condition.
6. The method of claim 5, wherein the method further comprises using the first cluster to determine an extent to which race is a covariate with respect to the first myelofibrosis-related state.
7. The method of claim 5, wherein the method further comprises using the first cluster to determine an extent to which sex is a covariate with respect to the first myelofibrosis-related state.
8. The method of claim 5, wherein the method further comprises using the first cluster to determine an extent to which age is a covariate with respect to the first myelofibrosis-related state.
9. The method of any one of claims 1-8, wherein sex, race, age, and / or physical condition are each represented in a balanced manner in the cohort of subjects.
10. The method of any one of claims 1-9, the method further comprising filtering the single nucleus-or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells to remove counts of ambient RNA molecules, doublets and / or empty droplets.
11. The method of any one of claims 1-10, wherein the first metadata includes one or more first biomarker annotations for each nucleus or cell in the plurality of nuclei or cells drawn from a first standardized set of biomarkers that each define a unique cell type.
12. The method claim 11, wherein the method further comprises using the first cluster to determine an extent to which a first biomarker in the first standardized set of biomarkers is a covariate with respect to the first myelofibrosis-related state.
13. The method of any one of claims 1-12, wherein the first metadata includes one or more second biomarker annotations for each nucleus or cell in the plurality of nuclei or cells drawn from a second standardized set of biomarkers that each define a unique cell type at a unique stage.
14. The method claim 11, wherein the method further comprises using the first cluster to determine an extent to which a second biomarker in the second standardized set of biomarkers is a covariate with respect to the first myelofibrosis-related state.
15. The method of any one of claims 1-14, the method further comprising: obtaining genotype data for each subject in the plurality of subjects; and overlaying genotype data for each subject represented in the first cluster with a myelofibrosis-related state of each subject in the first cluster.
16. The method of any one of claims 1-15, the method further comprising: obtaining genotype data for each subject in the plurality of subjects, and determining an extent to which a genotype is a covariate for the first myelofibrosis- related state using the genotype data for each subject represented in the first cluster.
17. The method of any one of claims 1-16, the method further comprising using the myelofibrosis abnormality detection system to associate a test subject with the first myelofibrosis-related state by a procedure comprising: obtaining, in electronic form, second information comprising single-nucleus or singlecell transcriptome data for the plurality of genes for each nucleus or cell in a second plurality of nuclei or cells, wherein each nucleus or cell in the second plurality of nuclei or cells is obtained from a sample obtained from the test subject; co-clustering the first plurality of nuclei or cells and the second plurality of nuclei or cells into the plurality of clusters; and identifying the test subject as having the first myelofibrosis-related state when nuclei or cells from the second plurality of nuclei or cells co-cluster into the first cluster.
18. The method of any one of claims 1-17, wherein the first myelofibrosis-related state is myelofibrosis and the second myelofibrosis-related state is absence of myelofibrosis.
19. The method of any one of claims 1-17, wherein the first myelofibrosis-related state is absence of myelofibrosis and the second myelofibrosis-related state is myelofibrosis.
20. The method of any one of claims 1-17, wherein the first myelofibrosis-related state is a first stage of myelofibrosis and the second myelofibrosis-related state is a second stage of myelofibrosis.
21. The method of any one of claims 1-17, wherein the first myelofibrosis-related state is low risk myelofibrosis and the second myelofibrosis-related state is intermediate- 1 risk myelofibrosis.
22. The method of any one of claims 1-17, wherein the first myelofibrosis-related state is intermediate- 1 risk myelofibrosis and the second myelofibrosis-related state is intermediate-2 risk myelofibrosis.
23. The method of any one of claims 1-17, wherein the first myelofibrosis-related state is intermediate-2 risk myelofibrosis and the second myelofibrosis-related state is high risk myelofibrosis.
24. The method of any one of claims 1-17, wherein the first myelofibrosis-related state is low risk or intermediate- 1 risk myelofibrosis and the second myelofibrosis-related state is intermediate-2 risk myelofibrosis or high risk myelofibrosis.
25. The method of any one of claims 1-17, wherein the first myelofibrosis-related state is low risk myelofibrosis and the second myelofibrosis-related state is absence of myelofibrosis.
26. The method of any one of claims 1-17, wherein the first myelofibrosis-related state is low risk myelofibrosis and the second myelofibrosis-related state is intermediate- 1 risk or intermediate-2 risk myelofibrosis.
27. The method of any one of claims 1-26, wherein the first information further comprises second metadata for each subject in the cohort of subjects comprising age, body mass index, or sex, and the method further comprises using the second information to prune the cohort of subjects based on age, body mass index, or sex thereby causing the cohort of subjects to be free of confounding for age, body mass index, or sex, wherein the pruning causes a subset of subjects to be removed from the cohort of subjects prior to the clustering.
28. The method of any one of claims 1-27, wherein the first information further comprises second metadata for each subject in the cohort of subjects comprising age, body mass index, and sex, and the method further comprises using the second information to prune the cohort of subjects based on age, body mass index, and sex thereby causing the cohort of subjects to be free of confounding for age, body mass index, and sex, wherein the pruning causes a subset of subjects to be removed from the cohort of subjects prior to the clustering.
29. The method of any one of claims 1-28, the method further comprising: using the first metadata to identify a second cluster in the plurality of clusters with the second myelofibrosis-related state by determining that the second cluster includes nuclei or cells from subjects in the cohort that have the second myelofibrosis-related state.
30. The method of claim 29, wherein the first myelofibrosis-related state is a first stage of myelofibrosis and the second myelofibrosis-related state is a second stage of myelofibrosis.
31. The method of claim 29, the method further determining that the first cluster comprises basophil / mast cells or nuclei of basophil / mast cells and the second cluster comprises cells or nuclei of cells other than basophil / mast cells.
32. The method of claim 29, the method further determining that the first cluster comprises monocytes or nuclei of monocytes and the second cluster comprises cells or nuclei of cells other than monocytes.
33. The method of claim 29, the method further determining that the first cluster comprises erythroid lineage cells or nuclei of erythroid lineage cells and the second cluster comprises cells or nuclei of cells other than erythroid lineage cells.
34. The method of claim 29, the method further determining that the first cluster comprises hematopoietic precursor cells or nuclei of hematopoietic precursor cells and the second cluster comprises cells or nuclei of cells other than hematopoietic precursor cells.
35. The method of claim 29, the method further determining that the first cluster comprises lymphoid lineage cells or nuclei of lymphoid lineage cells and the second cluster comprises cells or nuclei of cells other than lymphoid lineage cells.
36. The method of claim 29, the method further determining that the first cluster comprises megakaryocyte-erythroid progenitor cells or nuclei of megakaryocyte-erythroid progenitor cells and the second cluster comprises cells or nuclei of cells other than megakaryocyte- erythroid progenitor cells.
37. The method of claim 29, the method further determining that the first cluster comprises megakaryocyte cells or nuclei of megakaryocyte cells and the second cluster comprises cells or nuclei of cells other than megakaryocyte cells.
38. The method of claim 29, the method further determining that the first cluster comprises myeloid lineage cells or nuclei of myeloid lineage cells and the second cluster comprises cells or nuclei of cells other than myeloid lineage cells.
39. The method of any one of claims 1-38, wherein the plurality of genes comprises 100 or more genes, 200 or more genes, 500 or more genes, 1000 or more genes, 2000 or more genes, 5000 or more genes, 10000 or more genes, or 15,000 or more genes.
40. The method of any one of claims 1-39, wherein the first plurality of nuclei or cells comprises 10,000 or more nuclei or cells, 50,000 or more nuclei or cells, 100,000 or more nuclei or cells, 250,000 or more nuclei or cells, 500,000 or more nuclei or cells, 600,000 or more nuclei or cells, or 1 x 106or more nuclei or cells.
41. The method of any one of claims 1-40, wherein the cohort of subjects comprises 25 or more subjects, 50 or more subjects, 75 or more subjects, or 100 or more subjects.
42. The method of any one of claims 1-41, wherein the first subset of subjects comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects and the second subset of subjects is other than the first subset of subjects and comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects.
43. The method of any one of claims 1-42, wherein the first myelofibrosis-related state or the second myelofibrosis-related state is associated with osteosclerosis, extramedullary hematopoiesis (EMH), inefficient hematopoiesis, inflammation, splenomegaly, cytopenia, portal hypertension, thromboembolism, infection, or acute myeloid leukemia (AML).
44. The method of any one of claims 1-43, wherein the method informs a response to a drug compound in a patient or in a plurality of patients.
45. The method of any one of claims 1-44, wherein the method informs a response to a dosing amount, duration, and / or frequency of a drug in a patient or in a plurality of patients.
46. The method of claim 29, wherein the first cluster represents a myelofibrosis disease state and the second cluster represents a healthy state, and the method further comprises identifying a metabolic pathway comprising a set of genes, wherein one or more genes in the set of genes are overexpressed or under-expressed in the first cluster relative to the second cluster.
47. The method of any one of claims 29-46, wherein the first cluster represents a myelofibrosis disease state and the second cluster represents a healthy state, and the method further comprises: accessing, in electronic form, a plurality of compound-specific differential transcriptional signatures, wherein each respective compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures is a difference between (i) a respective compound-treated transcriptional signature in a plurality of compound-treated transcriptional signatures and (ii) a baseline transcriptional signature data set, wherein the baseline transcriptional signature data set is from a control sample of one or more cells of a cell type, and each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures is from a corresponding compound- treated sample of one or more cells of the cell type separately treated with a different compound in a plurality of at least 10 compounds; determining a test differential transcription signature by differential comparison of a transcriptional signature of the nuclei or cells of the first cluster and the second cluster; andcomparing the test differential transcription signature to each respective compoundspecific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures, thereby identifying a first compound associated with a compoundspecific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures that matches the test differential transcription signature.
48. The method of claim 47, wherein the myelofibrosis disease state is selected from low risk, intermediate- 1 risk, intermediate-2 risk, and high risk myelofibrosis.
49. The method of claim 47 or 48, wherein the control sample and each corresponding compound-treated sample is exposed to a solvent, wherein the solvent is the same solvent for the control sample and each corresponding compound-treated sample, optionally wherein the solvent comprises dimethyl sulfoxide (DMSO).
50. The method of any one of claims 47 to 49, wherein the control sample and each corresponding compound-treated sample each comprises a polar aprotic solvent, wherein the polar aprotic solvent is the same polar aprotic solvent for the control sample and each corresponding compound-treated sample.
51. The method of any one of claims 47 to 50, wherein the control sample and each corresponding compound-treated sample comprises DMSO.
52. The method of any one of claims 47 to 51, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nucleus assay and / or single-cell assay data, optionally wherein the singlenucleus assay and / or single-cell assay data is selected from single-cell ribonucleic acid (RNA) sequencing (scRNA-seq) data, single-nucleus RNA sequencing (snRNA-seq) data, scTag-seq data, single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data, CyTOF / SCoP data, E-MS / Abseq data, miRNA-seq data, CITE-seq data, or any combinations thereof.
53. The method of any one of claims 47 to 52, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises or consists of single-cell RNA sequencing (scRNA-seq) data.
54. The method of any one of claims 47 to 52, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nuclei or single-cell RNA sequencing (scRNA-seq) data.
55. The method of any one of claims 47 to 52, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures consists of single-nuclei or single-cell RNA sequencing (scRNA-seq) data.
56. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells comprises basophils / mast cells, monocytes, CD14+ cells, erythroid lineage cells, hematopoietic precursor cells, lymphoid lineage cells, megakaryocyte-erythroid progenitor cells, megakaryocytes, or myeloid lineage cells.
57. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells comprises basophils / mast cells.
58. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells comprises monocytes.
59. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells consists of basophils / mast cells.
60. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells consists of monocytes.
61. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells consists of erythroid lineage cells.
62. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells comprises hematopoietic precursor cells.
63. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells comprises lymphoid lineage cells.
64. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells comprises megakaryocyte-erythroid progenitor cells.
65. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells consists of macrophages.
66. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells consists of megakaryocytes.
67. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells consists of myeloid lineage cells.
68. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells comprises or consists of one or more cells from an organ, cells from a tissue, stem cells, human cells, cells from umbilical cord blood, cells from peripheral blood, bone marrow cells, cells from a solid tissue, differentiated cells and / or cells from a cell line.
69. The method of any one of claims 47 to 68, wherein each corresponding compound- treated sample of one or more cells is a frozen sample.
70. The method of any one of claims 47 to 69, wherein each respective compound in the plurality of compounds has a molecular weight of less than 2000 Daltons.
71. The method of any one of claims 47 to 70, wherein each respective compound in the plurality of compounds satisfies at least three criteria of the Lipinski rule of five criteria, or optionally each of the Lipinski rule of five criteria.
72. A computer system for manufacturing a myelofibrosis-related detection system, the computer system comprising: one or more processors; andmemory addressable by the one or more processors, the memory storing at least one program for execution by the one or more processors, the at least one program comprising instructions for: obtaining, in electronic form, first information comprising:(i) single-nucleus or single-cell transcriptome data for a plurality of genes for each nucleus or cell in a first plurality of nuclei or cells, wherein each nucleus or cell in the first plurality of nuclei or cells is obtained from a sample in a plurality of samples, each sample in the plurality of samples is from a different subject in a cohort of subjects, the first plurality of nuclei or cells includes a different subset of nuclei or cells from a sample from each subject in the cohort of subjects, and the first plurality of nuclei or cells comprises at least 1000 nuclei or cells, and(ii) first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has a first myelofibrosis-related state or a second myelofibrosis-related state, wherein at least a first subset of subjects in the cohort of subjects have the first myelofibrosis-related state and a second subset of subjects in the cohort of subjects have the second myelofibrosis- related state, and wherein the respective single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is barcoded with the subject in the cohort of subjects originating the respective single-nucleus or single-cell transcriptome data; clustering the first plurality of nuclei or cells into a plurality of clusters by (i) computing a plurality of distances using the single-nucleus or single-cell transcriptome data for the plurality of genes for each unique pair of nuclei or cells in the first plurality of nuclei or cells and (ii) evaluating the plurality of distances with a criterion function, wherein the plurality of distances includes a separate distance for each unique pair of nuclei or cells in the first plurality of nuclei or cells, each respective distance in the plurality of distances represents a different pair of nuclei or cells in the first plurality of nuclei or cells and quantifies a distance between (i) a respective first vector formed by the single-nucleus or single-cell transcriptome data for the plurality of genes for a respective first nucleus or first cell in the different pair of nuclei and (ii) a respective second vector formed by the single-nucleus or single-cell transcriptome data for the plurality of genes for a respective second nucleus or second cell in the different pair of nuclei or cells, and each respective cluster in the plurality of clusters represents a corresponding subset of nuclei or cells of the first plurality of nuclei or cells that are clustered together based on evaluation of distances in the plurality of distances representing different pairs of nuclei or cells within the corresponding subset of nuclei or cells with the criterion function; and using the first metadata to identify a first cluster in the plurality of clusters with the first myelofibrosis-related state by determining that the first cluster includes nuclei or cells from subjects in the cohort of subjects that have the first myelofibrosis-related state.
73. The computer system of claim 72, wherein the first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has the first myelofibrosis-related state or the second myelofibrosis-related state comprises a histologically graded disease status for the respective subject.
74. The computer system of claim 72, wherein the first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has the first myelofibrosis-related state or the second myelofibrosis-related state comprises a histologically graded disease status for the respective subject determined in accordance with a consistent, verified handling of each sample the plurality of samples.
75. The computer system of any one of claims 72-74, wherein each sample in the plurality of samples is from a different subject in the cohort of subjects and the plurality of samples comprises 5 samples, 20 samples, 50 samples, or 100 or more samples.
76. The computer system of any one of claims 72-75, wherein the first metadata for each respective subject in the cohort of subjects further indicates one or more features of the subject selected from the group consisting of sex, race, age, and physical condition.
77. The computer system of claim 76, wherein the method further comprises using the first cluster to determine an extent to which race is a covariate with respect to the first myelofibrosis-related state.
78. The computer system of claim 76, wherein the method further comprises using the first cluster to determine an extent to which sex is a covariate with respect to the first myelofibrosis-related state.
79. The computer system of claim 76, wherein the method further comprises using the first cluster to determine an extent to which age is a covariate with respect to the first myelofibrosis-related state.
80. The computer system of any one of claims 72-79, wherein sex, race, age, and / or physical condition are each represented in a balanced manner in the cohort of subjects.
81. The computer system of any one of claims 72-80, the method further comprising filtering the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells to remove counts of ambient RNA molecules, doublets and / or empty droplets.
82. The computer system of any one of claims 72-81, wherein the first metadata includes one or more first biomarker annotations for each nucleus or cell in the plurality of nuclei or cells drawn from a first standardized set of biomarkers that each define a unique cell type.
83. The computer system claim 82, wherein the method further comprises using the first cluster to determine an extent to which a first biomarker in the first standardized set of biomarkers is a covariate with respect to the first myelofibrosis-related state.
84. The computer system of any one of claims 72-83, wherein the first metadata includes one or more second biomarker annotations for each nucleus or cell in the plurality of nuclei or cells drawn from a second standardized set of biomarkers that each define a unique cell type at a unique stage.
85. The computer system claim 84, wherein the method further comprises using the first cluster to determine an extent to which a second biomarker in the second standardized set of biomarkers is a covariate with respect to the first myelofibrosis-related state.
86. The computer system of any one of claims 72-85, the method further comprising: obtaining genotype data for each subject in the plurality of subjects; and overlaying genotype data for each subject represented in the first cluster with a myelofibrosis-related state of each subject in the first cluster.
87. The computer system of any one of claims 72-86, the method further comprising: obtaining genotype data for each subject in the plurality of subjects; and determining an extent to which a genotype is a covariate for the first myelofibrosis- related state using the genotype data for each subject represented in the first cluster.
88. The computer system of any one of claims 72-87, the method further comprising using the myelofibrosis abnormality detection system to associate a test subject with the first myelofibrosis-related state by a procedure comprising: obtaining, in electronic form, second information comprising single-nucleus or singlecell transcriptome data for the plurality of genes for each nucleus or cell in a second plurality of nuclei or cells, wherein each nucleus or cell in the second plurality of nuclei or cells is obtained from a sample obtained from the test subject; co-clustering the first plurality of nuclei or cells and the second plurality of nuclei or cells into the plurality of clusters; and identifying the test subject as having the first myelofibrosis-related state when nuclei or cells from the second plurality of nuclei or cells co-cluster into the first cluster.
89. The computer system of any one of claims 72-88, wherein the first myelofibrosis-related state myelofibrosis and the second myelofibrosis-related state is absence of myelofibrosis.
90. The computer system of any one of claims 72-88, wherein the first myelofibrosis-related state is absence of myelofibrosis and the second myelofibrosis-related state is myelofibrosis.
91. The computer system of any one of claims 72-88, wherein the first myelofibrosis-related state is a first stage of myelofibrosis and the second myelofibrosis-related state is a second stage of myelofibrosis.
92. The computer system of any one of claims 72-88, wherein the first myelofibrosis-related state is low risk myelofibrosis and the second myelofibrosis-related state is intermediate- 1 risk myelofibrosis.
93. The computer system of any one of claims 72-88, wherein the first myelofibrosis-related state is intermediate- 1 risk myelofibrosis and the second myelofibrosis-related state is intermediate-2 risk myelofibrosis.
94. The computer system of any one of claims 72-88, wherein the first myelofibrosis-related state is intermediate-2 risk myelofibrosis and the second myelofibrosis-related state is high risk myelofibrosis.
95. The computer system of any one of claims 72-88, wherein the first myelofibrosis-related state is low risk or intermediate- 1 risk myelofibrosis and the second myelofibrosis-related state is intermediate-2 risk myelofibrosis or high risk myelofibrosis.
96. The computer system of any one of claims 72-88, wherein the first myelofibrosis-related state is low risk myelofibrosis and the second myelofibrosis-related state is absence of myelofibrosis.
97. The computer system of any one of claims 72-88, wherein the first myelofibrosis-related state is low risk myelofibrosis and the second myelofibrosis-related state is intermediate- 1 risk or intermediate-2 risk myelofibrosis.
98. The computer system of any one of claims 72-97, wherein the first information further comprises second metadata for each subject in the cohort of subjects comprising age, body mass index, or sex, and the method further comprises using the second information to prune the cohort of subjects based on age, body mass index, or sex thereby causing the cohort of subjects to befree of confounding for age, body mass index, or sex, wherein the pruning causes a subset of subjects to be removed from the cohort of subjects prior to the clustering.
99. The computer system of any one of claims 72-98, wherein the first information further comprises second metadata for each subject in the cohort of subjects comprising age, body mass index, and sex, and the method further comprises using the second information to prune the cohort of subjects based on age, body mass index, and sex thereby causing the cohort of subjects to be free of confounding for age, body mass index, and sex, wherein the pruning causes a subset of subjects to be removed from the cohort of subjects prior to the clustering.
100. The computer system of any one of claims 72-99, the method further comprising: using the first metadata to identify a second cluster in the plurality of clusters with the second myelofibrosis-related state by determining that the second cluster includes nuclei or cells from subjects in the cohort that have the second myelofibrosis-related state.
101. The computer system of claim 100, wherein the first myelofibrosis-related state is a first stage of myelofibrosis and the second myelofibrosis-related state is a second stage of myelofibrosis.
102. The computer system of claim 100, the method further determining that the first cluster comprises basophil / mast cells or nuclei of basophil / mast cells and the second cluster comprises cells or nuclei of cells other than basophil / mast cells.
103. The computer system of claim 100, the method further determining that the first cluster comprises monocytes or nuclei of monocytes and the second cluster comprises cells or nuclei of cells other than monocytes.
104. The computer system of claim 100, the method further determining that the first cluster comprises erythroid lineage cells or nuclei of erythroid lineage cells and the second cluster comprises cells or nuclei of cells other than erythroid lineage cells.
105. The computer system of claim 100, the method further determining that the first cluster comprises hematopoietic precursor cells or nuclei of hematopoietic precursor cells and the second cluster comprises cells or nuclei of cells other than hematopoietic precursor cells.
106. The computer system of claim 100, the method further determining that the first cluster comprises lymphoid lineage cells or nuclei of lymphoid lineage cells and the second cluster comprises cells or nuclei of cells other than lymphoid lineage cells.
107. The computer system of claim 100, the method further determining that the first cluster comprises megakaryocyte-erythroid progenitor cells or nuclei of megakaryocyte-erythroid progenitor cells and the second cluster comprises cells or nuclei of cells other than megakaryocyte-erythroid progenitor cells.
108. The computer system of claim 100, wherein the first myelofibrosis-related state is absence of fibrosis or inflammation and the second myelofibrosis-related state is presence of fibrosis or inflammation.
109. The computer system of claim 100, the method further determining that the first cluster comprises megakaryocyte cells or nuclei of megakaryocyte cells and the second cluster comprises cells or nuclei of cells other than megakaryocyte cells.
110. The computer system of claim 100, the method further determining that the first cluster comprises myeloid lineage cells or nuclei of myeloid lineage cells and the second cluster comprises cells or nuclei of cells other than myeloid lineage cells.
111. The computer system of any one of claims 72-110, wherein the plurality of genes comprises 100 or more genes, 200 or more genes, 500 or more genes, 1000 or more genes, 2000 or more genes, 5000 or more genes, 10000 or more genes, or 15,000 or more genes.
112. The computer system of any one of claims 72-111, wherein the first plurality of nuclei or cells comprises 10,000 or more nuclei or cells, 50,000 or more nuclei or cells, 100,000 or more nuclei or cells, 250,000 or more nuclei or cells, 500,000 or more nuclei or cells, 600,000 or more nuclei or cells, or 1 x 106or more nuclei or cells.
113. The computer system of any one of claims 72-112, wherein the cohort of subjects comprises 25 or more subjects, 50 or more subjects, 75 or more subjects, or 100 or more subjects.
114. The computer system of any one of claims 72-113, wherein the first subset of subjects comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects and the second subset of subjects is other than the first subset of subjects and comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects.
115. The computer system of any one of claims 72-114, wherein the first myelofibrosis- related state or the second myelofibrosis-related state is associated with osteosclerosis, extramedullary hematopoiesis (EMH), inefficient hematopoiesis, inflammation, splenomegaly, cytopenia, portal hypertension, thromboembolism, infection, or acute myeloid leukemia (AML).
116. The computer system of any one of claims 72-115, wherein the method informs a response to a drug compound in a patient or in a plurality of patients.
117. The computer system of any one of claims 71-116, wherein the method informs a response to a dosing amount, duration, and / or frequency of a drug in a patient or in a plurality of patients.
118. The computer system of claim 100, wherein the first cluster represents a myelofibrosis disease state and the second cluster represents a healthy state, and the method further comprises identifying a metabolic pathway comprising a set of genes, wherein one or more genes in the set of genes are overexpressed or under-expressed in the first cluster relative to the second cluster.
119. The computer system of any one of claims 72-118, wherein the first cluster represents a myelofibrosis disease state and the second cluster represents a healthy state, and the method further comprises: accessing, in electronic form, a plurality of compound-specific differential transcriptional signatures, wherein each respective compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptionalsignatures is a difference between (i) a respective compound-treated transcriptional signature in a plurality of compound-treated transcriptional signatures and (ii) a baseline transcriptional signature data set, wherein the baseline transcriptional signature data set is from a control sample of one or more cells of a cell type, and each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures is from a corresponding compound- treated sample of one or more cells of the cell type separately treated with a different compound in a plurality of at least 10 compounds; determining a test differential transcription signature by differential comparison of the transcriptional signature of the nuclei or cells of the first cluster and the second cluster; and comparing the test differential transcription signature to each respective compoundspecific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures, thereby identifying a first compound associated with a compoundspecific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures that matches the test differential transcription signature.
120. The method of claim 119, wherein the myelofibrosis disease state is selected from low risk, intermediate- 1 risk, intermediate-2 risk, and high risk myelofibrosis.
121. The computer system of claim 119 or 120, wherein the control sample and each corresponding compound-treated sample each comprises a solvent, wherein the solvent is the same solvent for the control sample and each corresponding compound-treated sample, optionally wherein the solvent is dimethyl sulfoxide (DMSO).
122. The computer system of claim 119 or 120, wherein the control sample and each corresponding compound-treated sample each comprises a polar aprotic solvent, wherein the polar aprotic solvent is the same polar aprotic solvent for the control sample and each corresponding compound-treated sample.
123. The computer system of any one of claims 119 or 120, wherein the control sample and each corresponding compound-treated sample comprises DMSO.
124. The computer system of any one of claims 119 to 123, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nucleus assay and / or single-cell assay data, optionally wherein the single-nucleus assay and / or single-cell assay data is selected from single-cell ribonucleic acid (RNA) sequencing (scRNA-seq) data, single-nucleus RNA sequencing (snRNA-seq) data, scTag-seq data, single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data, CyTOF / SCoP data, E-MS / Abseq data, miRNA-seq data, CITE-seq data, or any combinations thereof.
125. The computer system of any one of claims 119 to 124, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises or consists of single-cell RNA sequencing (scRNA-seq) data.
126. The computer system of any one of claims 119 to 124, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nucleus RNA sequencing (scRNA-seq) data.
127. The computer system of any one of claims 119 to 124, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures consists of single-nucleus RNA sequencing (scRNA-seq) data.
128. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells comprises basophils / mast cells, monocytes, CD14+ cells, erythroid lineage cells, hematopoietic precursor cells, lymphoid lineage cells, megakaryocyte-erythroid progenitor cells, megakaryocytes, or myeloid lineage cells.
129. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells comprises basophils / mast cells.
130. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells comprises monocytes.
131. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells consists of basophils / mast cells.
132. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells consists of monocytes.
133. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells consists of erythroid lineage cells.
134. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells comprises hematopoietic precursor cells.
135. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells comprises lymphoid lineage cells136. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells comprises megakaryocyte-erythroid progenitor cells.
137. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells consists of macrophages.
138. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells consists of megakaryocytes.
139. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells consists of myeloid lineage cells.
140. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells comprises or consists of one or more cells from an organ, cells from a tissue, stem cells, human cells, cells from umbilical cord blood, cells from peripheral blood, bone marrow cells, cells from a solid tissue, differentiated cells and / or cells from a cell line.
141. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells is a frozen sample.
142. The computer system of any one of claims 119 to 141, wherein each respective compound in the plurality of compounds has a molecular weight of less than 2000 Daltons.
143. The computer system of any one of claims 119 to 142, wherein each respective compound in the plurality of compounds satisfies at least three criteria of the Lipinski rule of five criteria, or optionally each of the Lipinski rule of five criteria.
144. A non-transitory computer readable storage medium, wherein the non-transitory computer readable storage medium stores instructions, which when executed by a computer system, cause the computer system to perform a method for manufacturing a myelofibrosis- related abnormality detection system, the method comprising: obtaining, in electronic form, first information comprising:(i) single-nucleus or single-cell transcriptome data for a plurality of genes for each nucleus or cell in a first plurality of nuclei or cells, wherein each nucleus or cell in the first plurality of nuclei or cells is obtained from a sample in a plurality of samples, each sample in the plurality of samples is from a different subject in a cohort of subjects, the first plurality of nuclei or cells includes a different subset of nuclei or cells from a sample from each subject in the cohort of subjects, and the first plurality of nuclei or cells comprises at least 1000 nuclei or cells, and(ii) first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has a first myelofibrosis-related state or a second myelofibrosis-related state, wherein at least a first subset of subjects in the cohort of subjects have the first myelofibrosis-related state and a second subset of subjects in the cohort of subjects have the second myelofibrosis- related state, and wherein the respective single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the firstplurality of nuclei or cells is barcoded with the subject in the cohort of subjects originating the respective single-nucleus or single-cell transcriptome data; clustering the first plurality of nuclei or cells into a plurality of clusters by (i) computing a plurality of distances using the single-nucleus or single-cell transcriptome data for the plurality of genes for each unique pair of nuclei or cells in the first plurality of nuclei or cells and (ii) evaluating the plurality of distances with a criterion function, wherein the plurality of distances includes a separate distance for each unique pair of nuclei or cells in the first plurality of nuclei or cells, each respective distance in the plurality of distances represents a different pair of nuclei or cells in the first plurality of nuclei or cells and quantifies a distance between (i) a respective first vector formed by the single-nucleus or singlecell transcriptome data for the plurality of genes for a respective first nucleus or first cell in the different pair of nuclei and (ii) a respective second vector formed by the single-nucleus or single-cell transcriptome data for the plurality of genes for a respective second nucleus or second cell in the different pair of nuclei or cells, and each respective cluster in the plurality of clusters represents a corresponding subset of nuclei or cells of the first plurality of nuclei or cells that are clustered together based on evaluation of distances in the plurality of distances representing different pairs of nuclei or cells within the corresponding subset of nuclei or cells with the criterion function; and using the first metadata to identify a first cluster in the plurality of clusters with the first myelofibrosis-related state by determining that the first cluster includes nuclei or cells from subjects in the cohort of subjects that have the first myelofibrosis-related state.
145. A method for identifying a compound that transitions a myelofibrosis state to a healthy state, the method comprising: obtaining first information comprising:(i) single-nucleus or single-cell transcriptome data for a plurality of genes for each nucleus or cell in a first plurality of nuclei or cells, wherein each nucleus or cell in the first plurality of nuclei or cells is obtained from a sample in a plurality of samples from a cohort of subjects, and(ii) first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has a myelofibrosis disease state or a healthy state,wherein at least a first subset of subjects in the cohort of subjects have the myelofibrosis disease state and a second subset of subjects in the cohort of subjects have the healthy state, and wherein the respective single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is barcoded with the subject in the cohort of subjects originating the respective single-nucleus or single-cell transcriptome data; clustering the first plurality of nuclei or cells into a plurality of clusters based on the single-nucleus or single-cell transcriptome data for the plurality of genes for each pair of nuclei or cells in the first plurality of nuclei or cells; using the first information to identify a first cluster in the plurality of clusters with the myelofibrosis disease state by determining that the first cluster includes nuclei or cells from subjects in the cohort of subjects that have the myelofibrosis disease state; using the first information to identify a second cluster in the plurality of clusters with the healthy state by determining that the second cluster includes nuclei or cells from subjects in the cohort that have the healthy state; accessing a plurality of compound-specific differential transcriptional signatures, wherein each respective compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures is a difference between (i) a respective compound-treated transcriptional signature in a plurality of compound-treated transcriptional signatures and (ii) a baseline transcriptional signature data set, wherein, the baseline transcriptional signature data set is from a control sample of one or more cells of a cell type, and each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures is from a corresponding compound-treated sample of one or more cells of the cell type separately treated with a different compound in a plurality of compounds; determining a test differential transcription signature by differential comparison of a transcriptional signature of the nuclei or cells of the first cluster and the second cluster; and comparing the test differential transcription signature to each respective compoundspecific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures, thereby identifying a first compound associated with a compoundspecific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures that matches the test differential transcription signature.
146. The method of claim 145, wherein the first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has the myelofibrosis disease state or the healthy state comprises a histologically graded disease status for the respective subject.
147. The method of claim 145, wherein the first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has the myelofibrosis disease state or the healthy state comprises a histologically graded disease status for the respective subject determined in accordance with a consistent, verified handling of each sample in the plurality of samples.
148. The method of any one of claims 145-147, wherein each sample in the plurality of samples is from a different subject in the cohort of subjects and the plurality of samples comprises 5 samples, 20 samples, 50 samples, or 100 or more samples.
149. The method of any one of claims 145-148, wherein the first metadata for each respective subject in the cohort of subjects further indicates one or more features of the subject selected from the group consisting of sex, race, age, and physical condition.
150. The method of any one of claims 145-149, the method further comprising filtering the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells to remove counts of ambient RNA molecules, doublets and / or empty droplets.
151. The method of any one of claims 145-150, wherein the myelofibrosis state is low risk myelofibrosis intermediate- 1 risk myelofibrosis, intermediate-2 risk myelofibrosis, or high risk myelofibrosis.
152. The method of any one of claims 145-151, the method further determining that the first cluster comprises basophil / mast cells or nuclei of basophil / mast cells and the second cluster comprises cells or nuclei of cells other than basophil / mast cells.
153. The method of any one of claims 145-151, the method further determining that the first cluster comprises monocytes or nuclei of monocytes and the second cluster comprises cells or nuclei of cells other than monocytes.
154. The method of any one of claims 145-151, the method further determining that the first cluster comprises erythroid lineage cells or nuclei of erythroid lineage cells and the second cluster comprises cells or nuclei of cells other than erythroid lineage cells.
155. The method of any one of claims 145-151, the method further determining that the first cluster comprises hematopoietic precursor cells or nuclei of hematopoietic precursor cells and the second cluster comprises cells or nuclei of cells other than hematopoietic precursor cells.
156. The method of any one of claims 145-151, the method further determining that the first cluster comprises lymphoid lineage cells or nuclei of lymphoid lineage cells and the second cluster comprises cells or nuclei of cells other than lymphoid lineage cells.
157. The method of any one of claims 145-156, wherein the plurality of genes comprises 100 or more genes, 200 or more genes, 500 or more genes, 1000 or more genes, 2000 or more genes, 5000 or more genes, 10000 or more genes, or 15,000 or more genes.
158. The method of any one of claims 145-157, wherein the first plurality of nuclei or cells comprises 10,000 or more nuclei or cells, 50,000 or more nuclei or cells, 100,000 or more nuclei or cells, 250,000 or more nuclei or cells, 500,000 or more nuclei or cells, 600,000 or more nuclei or cells, or 1 x 106or more nuclei or cells.
159. The method of any one of claims 145-158, wherein the cohort of subjects comprises 25 or more subjects, 50 or more subjects, 75 or more subjects, or 100 or more subjects.
160. The method of any one of claims 145-159, wherein the first subset of subjects comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects and the second subset of subjects is other than the first subset of subjects and comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects.
161. The method of any one of claims 145-160, wherein the myelofibrosis disease state is associated with osteosclerosis, extramedullary hematopoiesis (EMH), inefficient hematopoiesis, inflammation, splenomegaly, cytopenia, portal hypertension, thromboembolism, infection, or acute myeloid leukemia (AML).
162. The method of any one of claims 145-161, wherein the method informs a response to a drug compound in a patient or in a plurality of patients.
163. The method of any one of claims 145-162, wherein the method informs a response to a dosing amount, duration, and / or frequency of a drug in a patient or in a plurality of patients.
164. The method of any one of claims 145-163, wherein the method further comprises identifying a metabolic pathway comprising a set of genes, wherein one or more genes in the set of genes are overexpressed or under-expressed in the first cluster relative to the second cluster.
165. The method of any one of claims 145-164, wherein the control sample and each corresponding compound-treated sample is exposed to a solvent, wherein the solvent is the same solvent for the control sample and each corresponding compound-treated sample, optionally wherein the solvent comprises dimethyl sulfoxide (DMSO).
166. The method of any one of claims 145-165, wherein the control sample and each corresponding compound-treated sample each comprises a polar aprotic solvent, wherein the polar aprotic solvent is the same polar aprotic solvent for the control sample and each corresponding compound-treated sample.
167. The method of any one of claims 145-165, wherein the control sample and each corresponding compound-treated sample comprises DMSO.
168. The method of any one of claims 145-167, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nucleus assay and / or single-cell assay data, optionally wherein the single — nucleus-assay and / or single-cell assay data is selected from single-cell ribonucleic acid (RNA) sequencing (scRNA-seq) data, single-nucleus RNA sequencing (snRNA-seq) data,scTag-seq data, single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data, CyTOF / SCoP data, E-MS / Abseq data, miRNA-seq data, CITE-seq data, or any combinations thereof.
169. The method of any one of claims 145-168, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises or consists of single-cell RNA sequencing (scRNA-seq) data.
170. The method of any one of claims 145-168, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nucleus RNA sequencing (scRNA-seq) data.
171. The method of any one of claims 145-168, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures consists of single-nucleus RNA sequencing (scRNA-seq) data.
172. The method of any one of claims 145-171, wherein each corresponding compound- treated sample of one or more cells comprises basophils / mast cells, monocytes, CD14+ cells, erythroid lineage cells, hematopoietic precursor cells, lymphoid lineage cells, megakaryocyte-erythroid progenitor cells, megakaryocytes, or myeloid lineage cells.
173. The method of any one of claims 145-171, wherein each corresponding compound- treated sample of one or more cells comprises basophils / mast cells.
174. The method of any one of claims 145-171, wherein each corresponding compound- treated sample of one or more cells comprises monocytes.
175. The method of any one of claims 145-171, wherein each corresponding compound- treated sample of one or more cells consists of basophils / mast cells.
176. The method of any one of claims 145-171, wherein each corresponding compound- treated sample of one or more cells consists of monocytes.
177. The method of any one of claims 145-171, wherein each corresponding compound- treated sample of one or more cells consists of erythroid lineage cells.
178. The method of any one of claims 145-171, wherein each corresponding compound- treated sample of one or more cells comprises hematopoietic precursor cells.
179. The method of any one of claims 145-171, wherein each corresponding compound- treated sample of one or more cells comprises lymphoid lineage cells.
180. The method of any one of claims 145-171, wherein each corresponding compound- treated sample of one or more cells comprises megakaryocyte-erythroid progenitor cells.
181. The method of any one of claims 145-171, wherein each corresponding compound- treated sample of one or more cells consists of macrophages.
182. The method of any one of claims 145-171, wherein each corresponding compound- treated sample of one or more cells consists of megakaryocytes.
183. The method of any one of claims 145-171, wherein each corresponding compound- treated sample of one or more cells consists of myeloid lineage cells.
184. The method of any one of claims 145-183, wherein each corresponding compound- treated sample of one or more cells comprises or consists of one or more cells from an organ, cells from a tissue, stem cells, human cells, cells from umbilical cord blood, cells from peripheral blood, bone marrow cells, cells from a solid tissue, differentiated cells and / or cells from a cell line.
185. The method of any one of claims 145-184, wherein each respective compound in the plurality of compounds has a molecular weight of less than 2000 Daltons.
186. The method of any one of claims 145-185, wherein each respective compound in the plurality of compounds satisfies at least three criteria of the Lipinski rule of five criteria, or optionally each of the Lipinski rule of five criteria.
Citation Information
Patent Citations
Method for constructing myelodysplastic syndrome progression gene prediction model
CN113764044A
Single cell transcriptome disease specific cell analysis method based on iteration EM clustering
CN116486920A
Data based cancer research and treatment systems and methods
US20230223121A1