Non-invasive bone marrow diagnostic methods
The method addresses the lack of understanding of HSPC heterogeneity by using scRNA-seq and machine learning to analyze CD34-positive cells from peripheral blood, enabling non-invasive detection and prediction of bone marrow pathologies and risk scores, thus overcoming the limitations of invasive sampling.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- YEDA RES & DEV CO LTD
- Filing Date
- 2024-06-09
- Publication Date
- 2026-07-24
AI Technical Summary
Current methods lack a comprehensive understanding of inter-individual heterogeneity in hematopoietic stem cell (HSPC) transcription programs across healthy individuals, particularly in relation to age-related blood phenomena like clonal hematopoiesis and CBC deviations, and there is a need for non-invasive methods to assess bone marrow pathology.
A non-invasive method using single-cell RNA sequencing (scRNA-seq) of CD34-positive cells from peripheral blood to analyze target cell datasets against control datasets, applying machine learning models to predict bone marrow pathology and calculate IPSS-M risk scores, without requiring invasive bone marrow sampling.
Enables accurate detection of bone marrow pathologies such as MDS, CMML, AML, and other conditions, and predicts blast cell percentages and IPSS-M risk scores, providing a non-invasive, efficient diagnostic and prognostic tool for bone marrow health assessment.
Smart Images

Figure 2026524765000001_ABST
Abstract
Description
Technical Field
[0001] Cross - reference to related applications This application claims the benefit of priority of Israeli Patent Application No. 303582, filed on June 8, 2023, the content of which is hereby incorporated by reference in its entirety into this specification.
[0002] The present invention belongs to the field of bone marrow diagnostic methods.
Background Art
[0003] The basis for understanding and defining the pathophysiological state of humans is a detailed description of the heterogeneity among healthy individuals. The variability among healthy humans is multifactorial and is determined by the interaction between germline / somatic mutations and the environment. The identification of inter - individual variations in the complete blood count (CBC) in a large cohort of healthy individuals revealed different age - related deviations from reference values. Such studies have revealed age - related macrocytic anemia associated with an increase in RDW and a decrease in absolute lymphocyte count. The mechanisms involved in both phenomena remain unclear. Another aspect of heterogeneity in the blood is the emergence of somatic mutations in hematopoietic stem cells and hematopoietic progenitor cells (HSPCs). All HSPCs acquire somatic mutations, but specific mutations in leukemia - related genes, namely pre - leukemia mutations - pLMs, can lead to clonal expansion of HSPCs, a phenomenon called clonal hematopoiesis (CH). CH is very common among the elderly, but little is understood about why pLMs cause clonal expansion and how CH is related to other age - related blood phenomena.
[0004] One of the major gaps in understanding these age-related phenomena in the blood is our insufficient knowledge about HSPC variability across healthy individuals of diverse ages. While various HSPC subpopulations and their functions have been widely studied, how they differ between individuals is not well understood. Heterogeneity in the frequency of CD34+ peripheral blood (PB) HSPCs has been reported in the past and has been associated with age, smoking, sex, and genetic factors, as well as different pathological conditions. Several studies have analyzed HSPC heterogeneity at higher resolution, but their sample sizes were limited. No studies have specifically determined inter-individual heterogeneity in the HSPC transcription program in large cohorts of healthy individuals, and how these correlate with CBC, CH, and age.
[0005] Such reference maps have not yet been described, despite the recent development of tools to characterize transcriptional programs in HSPCs at single-cell resolution with minimal bias. Furthermore, since most HSPCs reside in the bone marrow (BM), access to these cells, especially from healthy donors, is problematic. However, previous studies have demonstrated that most HSPC populations can be identified in PB, including those based on scRNAseq analysis, and that functional stem cells have been identified in mouse and human PB. PB can enrich intrinsic stem cell populations because it connects BM to other extramedullary stem cell sites. All of this suggests that PB HSPCs could be a good substitute for studying inter-individual HSPC transcriptional heterogeneity. Therefore, there is a great need for novel, accurate, non-invasive tests to assess bone marrow MSPCs by examining HSPCs in PB. [Overview of the project] [Means for solving the problem]
[0006] The present invention provides a non-invasive method for detecting bone marrow pathology, comprising: receiving a target cell dataset based on single-cell RNA sequencing (scRNA-seq) of CD34-positive cells from the peripheral blood of a subject; and analyzing the received target cell dataset in relation to a control dataset comprising multiple cell datasets, wherein each cell dataset of the multiple cell datasets is analyzed based on scRNA-seq of CD34-positive cells from the peripheral blood of a healthy subject. A non-invasive method for predicting the percentage of blast cells in the bone marrow and calculating an IPSS-M risk score is also provided, as well as a system for carrying out the method of the present invention.
[0007] According to the first aspect, a non-invasive method for detecting a pathological condition of the bone marrow in a subject requiring such detection, wherein the method is a. To receive a target cell dataset based on single-cell RNA sequencing (scRNA-seq) of CD34-positive cells from the target peripheral blood, b. Analyzing a received target cell dataset in relation to a control dataset containing multiple cell datasets, wherein each cell dataset in the multiple cell datasets is based on scRNA-seq of CD34-positive cells from the peripheral blood of a healthy subject, and the deviation of the target cell dataset from the control dataset indicates a pathological condition in the bone marrow. This provides a method for detecting pathological conditions in the bone marrow.
[0008] According to some embodiments, the cell dataset includes statistical data on the total number of CD34-positive cells in a peripheral blood sample.
[0009] According to some embodiments, the analysis involves generating feature vectors that represent the deviation of the target cell data from the control cell data.
[0010] According to some embodiments, the analysis involves applying a trained machine learning model to a received dataset, the machine learning model being trained on a training set containing multiple cell datasets, and the machine learning model classifying the bone marrow in question as healthy or not.
[0011] According to some embodiments, the training set further includes a cell dataset based on scRNA-seq of CD34-positive cells derived from the peripheral blood of subjects suffering from a bone marrow condition, and labels indicating that the cell dataset originates from a healthy subject or a subject with a bone marrow condition, and the machine learning model classifies the subjects as either healthy or suffering from a bone marrow condition.
[0012] According to some embodiments, the analysis involves applying a trained machine learning model to feature vectors, the machine learning model being trained on a training set including feature vectors from healthy subjects and subjects suffering from bone marrow disease, and labels indicating that the feature vectors are from healthy subjects or subjects suffering from bone marrow disease, and the machine learning model classifying subjects as either healthy or suffering from bone marrow disease.
[0013] According to some embodiments, the analysis involves applying a trained machine learning model to parameters extracted from a cell dataset, the machine learning model being trained on a training set containing parameters extracted from cell datasets of healthy subjects and, optionally, subjects suffering from myelopathy, and the machine learning model classifying subjects as healthy or not.
[0014] According to some embodiments, the cell dataset is selected from a metacell model of all CD34-positive cells in a peripheral blood sample, the transcriptome of each CD34-positive cell in the peripheral blood sample, and an annotated cell atlas of CD34-positive cell types present in the peripheral blood sample.
[0015] According to some embodiments, the bone marrow pathology is selected from myelodysplastic syndrome (MDS), chronic myelomonocytic leukemia (CMML), acute myeloid leukemia (AML), polycythemia vera (PV), essential thrombocythemia (ET), mastocytosis, chronic eosinophilic leukemia, myelofibrosis (MF), acute lymphoblastic leukemia (ALL), ambiguous lineage acute leukemia, multiple myeloma (MM), myeloproliferative disorders (MPN), and blastic plasmacytoid dendritic cell leukemia.
[0016] According to some embodiments, the method is a method for detecting MDS, in which deviations in the frequency of erythrocyte progenitor cells (ERYP), basophil / eosinophil / mast cell progenitor cells (BEMP), and / or megakaryocyte / erythrocyte / basophil / eosinophil / mast cell progenitor cells (MEBEMP) indicate the presence of MDS.
[0017] According to some embodiments, the method is a method for detecting CMML, in which deviations in the frequency of early granulocyte-monocyte progenitor cells (GMP-E) indicate the presence of CMML.
[0018] According to some embodiments, the method is a method for detecting AML, in which deviations in the frequency of common lymphocyte progenitor cells (CLPs) and / or natural killer / T / dendritic cell progenitor cells (NKTDPs) indicate the presence of AML.
[0019] According to some embodiments, the deviations are higher or lower levels of cell type than those present in healthy subjects.
[0020] According to some embodiments, deviations in CLP frequency also indicate MDS, where the deviation is a lower level of CLP than is present in healthy subjects.
[0021] According to some embodiments, deviations in CLP frequency also indicate CMML, MF, or MPN, where the deviation is a lower level of CLP than is present in healthy subjects.
[0022] According to some embodiments, the method is a method for detecting MDS, and a decrease in the frequency of CLP, NKTDP, or both compared to a healthy subject indicates MDS.
[0023] According to some embodiments, a decrease in the frequency of both CLP and NKTDP compared to a healthy subject indicates MDS.
[0024] According to some embodiments, the bone marrow pathology includes an increased percentage of blasts, there is an increased deviation, and a deviation in the frequency of early common lymphoid progenitor cells (CLP-E) indicates the presence of an increased percentage of blasts.
[0025] According to some embodiments, it further includes administering at least one therapeutic agent to a subject determined to have a bone marrow pathology.
[0026] According to another aspect, a non-invasive method for predicting the percentage of blasts in the bone marrow of a subject in need thereof, the method comprising receiving a measurement value of CLP-E cells in the peripheral blood of the subject, the measurement value being proportional to the percentage of blasts in the subject's bone marrow, thereby predicting the percentage of blasts in the subject's bone marrow.
[0027] According to some embodiments, the method further includes analyzing the received measurement value in relation to a control data set including multiple measurement values of CLP-E cells in the peripheral blood of healthy subjects and subjects suffering from bone marrow pathology, and the percentage of blasts in the bone marrow is known for each subject in the control data set.
[0028] According to another aspect, a non-invasive method for predicting the percentage of blasts in the bone marrow of a subject in need thereof, the method a. receiving a subject cell data set based on single-cell RNA sequencing (scRNA-seq) of CD34-positive cells from the peripheral blood of the subject, and b. Applying a trained machine learning model to a received dataset, wherein the machine learning model is trained on a training set comprising multiple cell datasets, and each cell dataset is based on scRNA-seq of CD34-positive cells from the peripheral blood of a control subject and labels indicating the percentage of blast cells in the bone marrow of the control subject that provided each cell dataset. The machine learning model includes outputting the predicted percentage of blast cells in the target bone marrow. This provides a method for predicting the percentage of blast cells in the bone marrow of a target.
[0029] According to some embodiments, the subject is suffering from leukemia.
[0030] According to some embodiments, the control subjects include subjects with leukemia and non-leukemia subjects.
[0031] According to some embodiments, the cell dataset is selected from a metacell model of all CD34-positive cells in a peripheral blood sample, the transcriptome of each CD34-positive cell in the peripheral blood sample, and an annotated cell atlas of CD34-positive cell types present in the peripheral blood sample.
[0032] According to some embodiments, the cell dataset is a metacell model, a. Obtaining peripheral blood samples from the subject, b. Isolating CD34-positive hematopoietic stem cells and progenitor cells (HSPCs) from peripheral blood samples, c. Perform scRNA-seq on isolated HSPCs to generate a transcriptome for each isolated HSPC, d. To generate a metacell model of HSPC based on those transcriptomes, It is produced by a method that includes [a specific method].
[0033] According to some embodiments, metacells are clusters of cells having similar transcriptomes.
[0034] According to some embodiments, the cell dataset includes grouping cells into cell types that share common differentiation within the HSPC differentiation spectrum.
[0035] According to some embodiments, the cell type is selected from BEMP, ERYP, MEBEMP-L, MEBEMP-E, GMP-E, pluripotent progenitor cells (MPP), hematopoietic stem cells (HSC), CLP-E, CLP-M, CLP-L, and NKTDP.
[0036] According to some embodiments, the method is a method for detecting MDS and / or leukemia, wherein a percentage of blast cells exceeding a predetermined threshold indicates that the subject has MDS and / or leukemia.
[0037] According to some embodiments, the method further includes administering at least one anti-cancer treatment to a subject suffering from MDS and / or leukemia.
[0038] In another embodiment, a non-invasive method for calculating a Molecular International Prognostic Scoring System (IPSS-M) risk score for a subject suffering from a myeloma, wherein the method is a. Predicting the percentage of blast cells in the target bone marrow by the method of the present invention, b. Detecting the presence of bone marrow mutations and karyotype abnormalities based on scRNA-seq reads from CD34-positive cells derived from the target peripheral blood, c. Obtaining hemoglobin levels and platelet counts from the subject's peripheral blood, d. Calculate the IPSS-M risk score based on the predicted blast percentage, detected mutations and karyotype, as well as the received hemoglobin level and platelet count. Includes, This provides a method for calculating the IPSS-M risk score.
[0039] According to some embodiments, the method further comprises administering a treatment regimen based on an IPSS-M risk score, where a more potent treatment regimen is administered to subjects with higher scores and a reduced treatment regimen is administered to subjects with lower scores.
[0040] In another embodiment, a system for evaluating the health of a subject's bone marrow, wherein the system is scRNA sequencing device, A non-temporary memory device in which instruction code modules are stored, A memory device is associated with at least one processor configured to execute a module of instruction code, and when the module of instruction code is executed, the at least one processor Using an scRNA sequencing instrument, we obtained single-cell transcriptomes from target peripheral blood-derived CD34-positive cells. A cell dataset is generated based on the obtained single-cell transcriptome. Cell datasets generated in association with a control dataset containing multiple cell datasets were analyzed, and each cell dataset in the multiple cell datasets was based on scRNA-seq of CD34-positive cells from the peripheral blood of healthy subjects, and Based on the deviation of the target cell dataset from the control dataset, the results of healthy bone marrow or bone marrow pathology in the subjects are output. We provide a system that is configured in this way.
[0041] According to some embodiments, the cell dataset is a metacell model having similar transcriptomes from acquired single-cell transcriptomes clustered into metacells.
[0042] Further embodiments and the full scope of applicability of the present invention will become apparent from the detailed description given below. However, it should be understood that the detailed description and specific examples, while illustrating preferred embodiments of the present invention, are given merely as examples, as various changes and modifications within the spirit and scope of the invention will become apparent to those skilled in the art from this detailed description. [Brief explanation of the drawing]
[0043] [Figure 1A-1H] (1A) Experimental design. (1B) Annotated 2D UMAP projection of the metacell manifold after filtering of metacells with low CD34 expression. (1C-D) Symmetric and (1D) asymmetric regulation of specific HSC transcription factors at branching into (1C) CLP (right) and MEBEMP (left) lines. Each panel shows the expression of one gene (Y axis). Metacells in all panels are ordered by increasing AVP expression in the MEBEMP line and decreasing AVP expression in the CLP line (left to right). The unit of gene expression in all figure panels is log2 of the fractional expression of each gene. (1E) Metacell population of interest linking BEMP to their MEBEMP-L precursor (dotted line). (1F) Positively and negatively regulated TFs involved in early BEMP differentiation. (1G) Gene-gene plots of IRF8 against TCF7 expression as prominent markers of DC and T cell differentiation, respectively. A population of interest consisting of high ACY3 NKTDP metacells is shown (dotted line). (1H) This population forms a gradient consisting of NK / T cell-like precursors showing high expression of both T cell and dendritic cell regulators, along with high expression of other T cell features such as CD7, MAF, IL7R, and TRBC2, and DC-like precursors showing low TCF7 / IRF8 expression ratios along with high expression of other DC features such as myeloid TF PU.1 and the MHC class II gene CD74. [Figure 2A-2H](2A) Characterization of inter-individual HSPC composition variability (scheme). (2B) Box plots (logarithmic scale) of inter-individual cell state frequency distributions. Percentages were calculated from the CD34+ population. The box plot center, hinge, and whiskers represent the median, first and third quartiles, and 1.5 × interquartile range, respectively. The numbers represent the mean + / -SD for each distribution. (2C) Correlation of cell state frequencies between 20 biological replicates and their original samples for the CLP (CLP-E, CLP-M, CLP-L, NKTDP - upper row) and MEBEMP (MEBEMP-E, MEBEMP-L, ERYP, BEMP - lower row) populations. All biological replicates were sampled one year after the original blood collection. (2D) (Top) - Individual cell state frequency profiles (colored lines) across HSC-MEBEMP and HSC-CLP differentiation gradients for 6 subjects, each representing one of the six archetypes (classes) of HSPC composition in healthy individuals. The dashed lines represent the median (black) and the 5th and 95th percentiles (gray) of the tested population (bottom panel). Cell state enrichment maps across 15 differentiation bins (rows) for all tested individuals (columns) clustered into 6 classes. Classes I and II represent individuals with relatively abundant lymphoid precursors, while classes V and VI represent individuals with relatively depleted lymphoid precursors. Individuals are sorted by stem cell nature within each class. Age and sex are shown for each individual. (2E) CBC correlation to cell type frequency: lymphocyte % (calculated from WBC for the entire cohort, left), HCT (males only, center), RDW (males only, right). Missing individuals lacked sufficient cells for analysis. The p-value for the sort-based test is shown for each correlation. See Methods for details on sort-based tests. (2F) Box plots of CLP frequency distribution in individuals with (right) and without (left) clonal hematopoiesis. (2G) Relative cell state frequencies in mutant cells (right) and non-mutant cells (left) after GoT in sample N122 (DNMT3A R882 mutation, VAF=0.07). (2H) CH frequencies (by gene) in age- and sex-matched high (red) and low (black) RDW individuals in a cohort of 18,147 individuals. [Figure 3A-3K](3A) Analysis of age-related compositional differences in MEBEMP (MEBEMP-E, MEBEMP-L, ERYP, BEMP, left) and CLP (CLP-E, CLP-M, CLP-L, right) populations, comparing the frequencies of specific cellular states (within the total CD34+ population) in young (<50 years) vs. older (>60 years) individuals without clonal hematopoiesis, performed separately for males (blue) and females (red). The Kruskal-Wallis p-values of group differences are shown in the upper panel. (3B) Analysis of age-related compositional differences within the MEBEMP differentiation trajectory, comparing the abundance of more differentiated (MEBEMP-L) states and more undifferentiated (MPP) states in young (<50 years) vs. older (>60 years) individuals, performed separately for males (blue) and females (red). The Kruskal-Wallis p-values of these group differences are shown in the upper panel. (3C) For the HSC population, see 3A. (3D) Analysis of age-related differences in (CD34+)cHSPC frequency from total PBMCs in a recent cohort of 1000 healthy individuals who underwent PBMC scRNA-seq28. (3E) True age (x) and predicted age (y) based on composition-controlled MEBEMP expression. (3F) Gene-gene correlation heatmap calculated across controlled individual levels of HSC-MEBEMP gene expression for HSC-MEBEMP composition. (3G) Intra-individual correlations of LMNA signatures in CLP and MEBEMP for both males (blue) and females (red). (3H) Analysis of age-related differences in LMNA signature expression for CLP (right) and MEBEMP (left) populations in young (<50 years) vs. older (>60 years) individuals. The Y-axis shows log2 (observed / predicted expression) normalized for composition. The box plot center, hinge, and whiskers represent the median, first and third quartiles, and 1.5 × interquartile range, respectively. (3I) Individual heatmaps of single cell counts across 20 bins for stem cell origin (AVP signature, y-axis) and MEBEMP differentiation (GATA1 signature, x-axis). Individual identifiers, as well as their RBC and MCV, are shown in the upper panel. (3J) Comparison of individual synchronization scores and clinical parameters (RBC / MCV) across males.High and low synchronization scores (shown as red and black dots, respectively) define clinically distinct populations. (3K) Correlation between individual synchronization scores and cellular state composition. Sorting test p-values are shown in the upper panel. [Figure 4A-4G] (4A) Age-related variation in composition bias score. (4B) Cell type-specific comparison of S phase signatures in circulating (left) HSPC vs. BM (right) HSPC. (4C) Age-related variation in S phase signature in late MEBEMP trajectory. (4D) Corresponding individual S phase signatures (X axis) and composition bias scores (Y) for individuals under 65 years (left) and 65 years and older (right). (4E-4G) are similar to 4D, but instead of S phase, (4E) LMNA signature, (4F) synchronization score, and (4G) RDW are shown. [Figures 5A-5F](5A) Diagnostic approaches to leukemia analysis using the cHSPC reference atlas (scheme): 1. Construction of scRNA-seq and patient-specific metacell models on CD34-enriched PB, 2. Projection of patient-derived metacells onto the healthy reference atlas, 3. Compositional (relative cell state frequency) analysis, and 4. Composition-controlled differential gene expression analysis, 5. Mutation and CNV analysis using targeted DNA sequencing and RNA-based karyotype analysis. (5B) Projection of metacells from two new healthy individuals (not included in the reference model), three CMML patients, two MDS patients (one of whom is pre- and post-treatment), one MDS / MPN duplication patient, one MF patient, and two AML patients onto a healthy HSPC reference metacell model. Gene expression correlations between patients (projections) and reference metacells are color-coded according to the legend on the right. (5C~5D) Individual cell state frequency profiles across HSC-MEBEMP and HSC-CLP differentiation gradients for two healthy samples and seven patient samples (one of which was both pre- and post-treatment). The dashed line in 5D represents the median (black) and 5th and 95th percentiles (gray) of the healthy population. (5E) Quantification of compositional regulatory differential gene expression against a normal reference for two healthy samples and seven patient samples (one of which was both pre- and post-treatment), quantification of the number of differentially expressed genes, and identification of specific genes that are repeatedly induced or repressed in the disease. (5F) scRNA-seq karyotype analysis of two AML patients and one MDS patient before and after treatment. Metacell models were created for each MDS / AML patient and projected onto our health reference map. The combined reference metacells and projected (patient) metacells were then used to calculate expression ratios across all expressed genes on all chromosomes. This shows the Log2 ratio variation in expression (patient / healthy) for all expressed genes across all chromosomes. The red line represents the median of the ratio variation distribution for each chromosome. [Figure 6] This is a block diagram showing a computing device that may be included in a system for determining the state of hematopoietic stem cells (HSCs) in a subject, according to some embodiments of the present invention. [Figures 7A-7B] This is a block diagram showing a system (7A) and display or (7B) for determining the IPSS-M score of a target according to some embodiments of the present invention. [Figure 8] This flowchart illustrates a method for determining the HSC state in a subject according to several embodiments of the present invention. [Figure 9] Factors involved in the differentiation of BEMP and NKTDP. These factors are positively regulated (CNRIP1, HPGDS, TET2, TNFSF10) and negatively regulated (CD34, HBD, CD74, and BLVRB) in the early stages of BEMP specification. [Figure 10] True age (x) versus predicted age based on composition-controlled CLP expression (y). [Figure 11] Gene-gene correlation heatmaps calculated across controlled individual levels of CLP gene expression for CLP compositions. [Figure 12A-12C] (12A) Heatmap of individual LMNA signature expression across the MEBEMP trajectory. Individual age and sex are color-coded in the upper row. (12B) LMNA signature expression correlation between 39 technical replicates and 20 biological replicates and their original samples. (12C) Synchronization score correlation between 39 technical replicates and 20 biological replicates and their original samples. All biological replicates were sampled one year after the original blood collection. [Figure 13A-13C](13A) Each of the four panels points to a different cell state gene signature shown on the x-axis. Top panel - Box plot of gene module expression distribution for different cell states in our reference atlas. Bottom panel - Gene signature expression density plot for each AML subclone. Using the reference gene signature distribution (top panel), we identified subpopulations of AML cells with CLP, MEBEMP, HSC, and NKTDP features (bottom panel). The dashed line represents the threshold for gene signature expression, listing the percentage of cells expressing the signature per AML clone. (13B) Left - Correlation heatmap of differentially expressed gene signatures for AML-1 (N186, top panel) and AML-2 (N205, bottom panel). Malignant states are characterized by multiple novel gene signatures in addition to the abnormal expression of differentiation-related modules from "healthy" cells. Right: UMAP projections of metacell models of AML-1 (top) and AML-2 (bottom), color-coded by the relative expression of differentially expressed genes. Overexpression of BCL2 can be observed in AML-1-2 compared to AML-1-1. AML-1 gene signatures: BCL2, VPREB1, RUNX3, GATA2, SELL, LMNA, ID2. AML-2 gene signatures: CCL4, MPO, LMNA, MME, JCHAIN, ACY3, DNTT, GATA2. (13C) Expression heatmaps of several genes across reference cell states and AML subclones. Malignant states differ significantly from healthy states in both reference gene expression and multiple additional gene expression signatures. [Modes for carrying out the invention]
[0044] For the sake of simplicity and clarity, it should be understood that the elements shown in the diagrams are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to others for clarity. Furthermore, reference numbers may be repeated between drawings to indicate corresponding or similar elements where deemed appropriate.
[0045] In some embodiments, the present invention provides a non-invasive method for detecting bone marrow pathology, comprising receiving a target cell dataset based on single-cell RNA sequencing (scRNA-seq) of CD34-positive cells from the peripheral blood of a subject, and analyzing the received target cell dataset in relation to a control dataset. Also provided is a non-invasive method for predicting the percentage of blast cells in the bone marrow, comprising applying a trained machine learning model to the received target cell dataset based on single-cell RNA sequencing (scRNA-seq) of CD34-positive cells from peripheral blood. A non-invasive method for calculating an IPSS-M risk score is also provided. A system for carrying out the methods of the present invention is also provided.
[0046] This invention is based, at least in part, on the remarkable finding that single-cell RNA sequencing (scRNA-Seq) of HSPCs in blood can be used to reconstruct the state of HSPCs in the bone marrow, thereby detecting bone marrow pathology, detecting the presence and percentage of myeloblasts, and predicting clinical outcomes and treatments based on differences from those observed in healthy controls. This study characterizes inter-individual heterogeneity in cHSPCs across 148 healthy individuals and analyzes 627K PB CD34+ cells via scRNA-seq. The size of the cohort, along with the efficacy and resolution of state-of-the-art single-cell technology and the computational methods used herein, allowed the inventors to refine and enhance previous findings from much smaller cohorts to characterize in detail the transcriptional programs of diverse, and sometimes rare (NKTDP, BEMP) HSPC subpopulations (Figure 1). This disclosure defines a normal reference range for cHSPC subpopulation frequencies within a large, age- and sex-diverse healthy population, demonstrating that while cHSPC subtype compositions are highly diverse among individuals, the cellular state itself is remarkably universal (Figure 2). These compositions remained stable over a one-year follow-up period, making them strong individual characteristics. At the population level, a significant correlation was found between low CLP frequency, CH, and increased RDW (Figure 2). This disclosure further demonstrates that the known age-related myeloid bias in HSPC is significantly male-dominant (Figure 3). Analysis of compositional regulatory transcriptional dispersion identified an age-correlated RNA expression clock (Figure 3) and a novel gene module (LMNA) whose expression increased with age in CLP. Healthy individuals were found to regulate the transition from stem cell to myeloid differentiation differently, and early downregulation of stem cell genes was found to correlate with high MCV anemia (Figure 3). Finally, we demonstrated how this resource can be used to effectively identify and diagnose pathological cases based on abnormal cHSPC subpopulation frequencies, differential gene signatures, and chromosomal abnormalities from all peripheral blood samples (Figure 5). A detailed pipeline for the identification and characterization of such conditions based on conventional reference is provided.
[0047] A method for analyzing a target bone marrow according to the first aspect, wherein the method is a. Obtain a dataset based on CD34-positive cells from the target blood sample, b. Including analyzing the target dataset received in relation to the control dataset, This provides a method for analyzing the target bone marrow.
[0048] In some embodiments, the method is a diagnostic method. In some embodiments, the method is a prognostic diagnostic method. In some embodiments, the method is a non-invasive method. In some embodiments, the method is an in vitro method. In some embodiments, the method is an ex vivo method. In some embodiments, the method is a treatment method. In some embodiments, the method is a computerized method. In some embodiments, the method is executed by at least one processor. In some embodiments, the method requires the analysis of data beyond the capabilities of human cognitive abilities.
[0049] As used herein, the term “non-invasive” refers to a method that does not require the extraction of a sample from the bone marrow. Bone marrow biopsy and aspiration are invasive, painful, and expensive procedures that provide a diagnostician with a sample of cells in the bone marrow. The method avoids the drawbacks of invasive bone marrow sampling by analyzing the bone marrow via circulating CD34-positive cells found in the blood. Thus, the method is highly beneficial because it is non-invasive. In some embodiments, the blood is peripheral blood. In some embodiments, the blood is venous blood. In some embodiments, the blood is circulating blood. In some embodiments, the blood is not of organ origin. In some embodiments, the blood is not of tissue origin. In some embodiments, the blood is not of bone marrow origin. In some embodiments, the blood is a blood sample.
[0050] In some embodiments, CD34-positive cells are hematopoietic stem progenitor cells (HSPCs). CD34 is a transmembrane cell surface protein that marks hematopoietic stem cells (HSCs) and early progenitor cells differentiated from HSCs. CD34-positive cells range from complete stem cells (HSCs) to cells that have begun to differentiate into one of two lineage programs: the common lymphocyte precursor (CLP) lineage or the megakaryocyte / erythrocyte / basophil / eosinophil / mast cell precursor (MEBEM-P) lineage. The human CD34 protein sequence can be found in Uniprot entry P28906, while the Entrez gene ID is #947. Drugs that bind to and / or identify CD34-expressing cells, as well as kits for isolating CD34-positive cells, are well known in the art. Examples include, but are not limited to, the Dynabead CD34 Positive Isolation Kit (ThermoFisher), IO Human CD34+Cell Isolation Kit (Creative Biolabs), EasySep Human CD34 Positive Selection Kit (Stemcell Technologies), and CD34 MicroBead Kit, as well as human (Miltenyi Biotec).
[0051] In some embodiments, the dataset is based on CD34-positive cells derived from a blood sample from a subject. In some embodiments, the dataset is based on all CD34-positive cells in the sample. In some embodiments, the dataset is a cell dataset. In some embodiments, the dataset is a collection of CD34-positive cells in blood. In some embodiments, the dataset is a cell-by-cell dataset. In some embodiments, the dataset includes entries for each CD34-positive cell. In some embodiments, the data is data about all CD34-positive cells in blood. In some embodiments, the dataset is statistical data. In some embodiments, the statistical data is statistical data that is a data transformation of cell data. In some embodiments, the dataset is based on single-cell data. In some embodiments, the dataset includes single-cell data. In some embodiments, the dataset consists of single-cell data. In some embodiments, the single-cell data is single-cell RNA data. In some embodiments, the single-cell RNA data is single-cell RNA sequencing (scRNA-seq) data. In some embodiments, the data is reads. In some embodiments, the reads are sequencing reads. In some embodiments, the data is transcriptome data. In some embodiments, the single-cell data is protein data. In some embodiments, single-cell data are proteome data. In some embodiments, the dataset includes the transcriptome of each CD34-positive cell. In some embodiments, the dataset includes the proteome of each CD34-positive cell. In some embodiments, the dataset is a cell atlas. In some embodiments, the cell atlas is annotated. In some embodiments, the annotations are cell types.
[0052] In some embodiments, the method further includes receiving a blood sample from a subject. In some embodiments, the method further includes extracting a blood sample from a subject. In some embodiments, the blood sample is a peripheral blood sample. In some embodiments, the method further includes generating a dataset from the sample. In some embodiments, the method further includes isolating CD34-positive cells from the sample. In some embodiments, isolation includes extraction. In some embodiments, isolation is positive selection. In some embodiments, isolation is negative selection.
[0053] In some embodiments, the method includes sequencing CD34-positive cells. In some embodiments, the sequencing is single-cell sequencing. In some embodiments, the sequencing is next-generation sequencing. In some embodiments, the sequencing is high-throughput sequencing. In some embodiments, the sequencing is large-scale parallel sequencing. In some embodiments, the dataset is a dataset of sequences. In some embodiments, the dataset is a dataset of expression. In some embodiments, expression is gene expression.
[0054] In some embodiments, CD34 cells are clustered into cell types. In some embodiments, cell types are defined by their transcriptional profiles. In some embodiments, cell types are defined by their transcriptomes. In some embodiments, cell types are defined by their proteomes. In some embodiments, cell types are defined by their differentiation levels. In some embodiments, cell types are defined by their differentiation states. In some embodiments, cell types are defined by how similarly they have differentiated.
[0055] In some embodiments, the dataset is a metacell model of CD34-positive cells. In some embodiments, the model is a model of all CD34-positive cells. Metacell modeling calculates cell partitions by similarity to generate substantially homogeneous groups (e.g., cell types) defined as metacells. In some embodiments, a cell type includes multiple metacells. In some embodiments, a cell type includes metacells with similar differentiation. Methods for generating metacells from single-cell data are well known and are described herein below, as well as, for example, Baran et al., "MetaCell: analysis of single-cell RNA-seq data using K-nn graph partitions," Genome Biol. 2019 Oct 11;20(1):206 and Ben-Kiki et al., "Metacell-s: a divide and conquer metacell algorithm for scalable scRNA-seq analysis," Genome Biol. 2022 Apr 19;23(1):100, the contents of which are incorporated herein by reference in their entirety. Furthermore, the metacell program is freely available at github.com / tanaylab / metacells. In some embodiments, the method includes generating metacells from scRNA-seq data.
[0056] In some embodiments, the control dataset contains data of the same type as the control dataset. In some embodiments, the control dataset contains multiple control datasets. In some embodiments, the control dataset contains multiple datasets. In some embodiments, each of the multiple datasets in the control dataset comes from a different control subject. In some embodiments, the control dataset contains multiple control subject datasets. In some embodiments, each of the multiple datasets is based on scRNA-seq of CD34-positive cells. In some embodiments, the CD34-positive cells are blood-derived. In some embodiments, the CD34-positive cells come from a control subject. In some embodiments, the control subject is a healthy subject. In some embodiments, the control subject is a subject with a bone marrow disorder. In some embodiments, the control subject is both a healthy subject and a subject with a bone marrow disorder. In some embodiments, the control dataset is an atlas of control cells. In some embodiments, the control dataset is an atlas of metacells from the control subject. In some embodiments, the atlas is an atlas of datasets.
[0057] In some embodiments, the dataset includes grouping cells into cell types. In some embodiments, metacells are classified into cell types. In some embodiments, cell types share a common transcriptional profile. In some embodiments, cell types share a common differentiation state. In some embodiments, the differentiation state is within the HSPC spectrum of differentiation. In some embodiments, the control dataset includes the amount of a control cell type. In some embodiments, the amount is a range. In some embodiments, the cell type is a type of metacell. In some embodiments, the cell type is a differentiation state. In some embodiments, the amount is a relative amount. In some embodiments, the amount is the amount of all cell types in the control. In some embodiments, the range is the range of all cell types in the control.
[0058] In some embodiments, the cell type is selected from different differentiation states of CD34-positive cells. In some embodiments, the cell type is selected from hematopoietic stem cells (HSCs), common lymphocyte progenitor cells (CLPs), natural killer / T / dendritic cell progenitor cells (NKTDPs), multipotent progenitor cells (MPPs), early granulocyte-monocyte progenitor cells (GMP-E), megakaryocyte / erythrocyte / basophil / eosinophil / mast cell progenitor cells (MEBEMPs), erythrocyte progenitor cells (ERYPs), and basophil / eosinophil / mast cell progenitor cells (BEMPs). In some embodiments, CLPs include early CLPs (CLP-E), intermediate CLPs (CLP-M), and late CLPs (CLP-L). In some embodiments, MEBEMPs include early MEBEMPs (MEBEMP-E) and late MEBEMPs (MEBEMP-L). In some embodiments, the cell type is selected from BEMP, ERYP, MEBEMP-L, MEBEMP-E, GMP-E, MPP, HSC, CLP-E, CLP-M, CLP-L, and NKTDP. In some embodiments, CLP includes NKTDP. In some embodiments, CLP includes CLP-E, CLP-M, CLP-L, and NKTDP. In some embodiments, MEBEMP includes BEMP. In some embodiments, MEBEMP includes ERYP. In some embodiments, MEBEMP includes BEMP, ERYP, and MEBEMP-L.
[0059] In some embodiments, the control dataset includes a control range for each cell type. In some embodiments, the control range is a relative range. In some embodiments, the relative range is a relative abundance. In some embodiments, the control relative range is the relative percentage of all CD34-positive cells. In some embodiments, the percentage is the percentage of CD34-positive cells in the sample. In some embodiments, the control range is shown in Figure 2B. In some embodiments, the control range for BEMP is about 4.4% of CD34-positive cells. In some embodiments, about 4.4% is 4.4+ / -4.1%. In some embodiments, the control range for ERYP is about 1.4% of CD34-positive cells. In some embodiments, about 1.4% is 1.4+ / -0.7%. In some embodiments, the control range for MEMBEMP-L is about 8.2% of CD34-positive cells. In some embodiments, about 8.2% is 8.2+ / -2.2%. In some embodiments, the control range for MEMBEMP-E is approximately 38.0% of CD34-positive cells. In some embodiments, approximately 38.0% is 38.0 + / - 6.5%. In some embodiments, the control range for GMP-E is approximately 3.0% of CD34-positive cells. In some embodiments, approximately 3.0% is 3.0 + / - 0.9%. In some embodiments, the control range for MPP is approximately 21.6% of CD34-positive cells. In some embodiments, approximately 31.6% is 31.6 + / - 4.7%. In some embodiments, the control range for HSC is approximately 1.8% of CD34-positive cells. In some embodiments, approximately 1.8% is 1.8 + / - 1.1%. In some embodiments, the control range for CLP-E is approximately 2.5% of CD34-positive cells. In some embodiments, approximately 2.5% is 2.5 + / - 0.8%. In some embodiments, the control range for CLP-M is approximately 7.9% of CD34-positive cells. In some embodiments, approximately 7.9% is 7.9+ / -5.2%. In some embodiments, the control range for CLP-L is approximately 5.7% of CD34-positive cells. In some embodiments, approximately 5.7% is 5.7+ / -3.6%.In some embodiments, the control range for NKTDP is approximately 5.1% of CD34-positive cells. In some embodiments, it is approximately 45.1%, or 5.1% + / - 3.0%.
[0060] In some embodiments, analysis is comparison. In some embodiments, analysis includes projecting a dataset onto a control dataset. In some embodiments, analysis is determining the difference in cell types between the target dataset and the control dataset. In some embodiments, the change is the loss of cells in a cell type. In some embodiments, the change is the gain of cells in a cell type. In some embodiments, cells are metacells. In some embodiments, analysis is analyzing the entire target dataset. In some embodiments, analysis is analyzing the target dataset in relation to all of the multiple datasets within the control dataset.
[0061] In some embodiments, analyzing the bone marrow includes detecting a pathological condition of the bone marrow. In some embodiments, detection includes determining a pathological condition of the bone marrow. In some embodiments, analysis includes diagnosing a pathological condition of the bone marrow. In some embodiments, analysis includes prognostic determination of a pathological condition of the bone marrow. In some embodiments, analysis includes determining appropriate treatment for a pathological condition of the bone marrow. In some embodiments, analysis includes determining the amount of blast cells in the bone marrow. In some embodiments, determination is prediction. In some embodiments, determination is estimation. In some embodiments, determination is approximation. In some embodiments, determination does not involve actually counting blast cells in the bone marrow.
[0062] In some embodiments, deviations of the target dataset from the control dataset indicate a myeloid pathology. In some embodiments, deviations of the target dataset from the control dataset indicate a specific myeloid pathology. In some embodiments, deviations of the target dataset from the control dataset indicate a bone marrow disorder. In some embodiments, the deviation includes the difference when the target dataset is projected onto the control dataset. In some embodiments, the deviation is a higher level / amount of cell types present in the target than in the healthy control. In some embodiments, the deviation is a higher frequency of cell types in the target than in the healthy control. In some embodiments, the deviation is a lower level / amount of cell types present in the target than in the healthy control. In some embodiments, the deviation is a lower frequency of cell types in the target than in the healthy control. In some embodiments, a smaller amount means the absence of a cell type. In some embodiments, a larger amount means the presence of a new cell type.
[0063] As used herein, the term “bone marrow pathology” refers to any disease or condition affecting the bone marrow of a human being. In some embodiments, the pathology is a disease. In some embodiments, the pathology is a bone marrow abnormality. Examples of bone marrow pathologies include, but are not limited to, myelodysplastic syndrome (MDS), chronic myelomonocytic leukemia (CMML), chronic myeloid leukemia (CML), acute myeloid leukemia (AML), polycythemia vera (PV), essential thrombocythemia (ET), mastocytosis, chronic eosinophilic leukemia, primary myelofibrosis (MF), post-ET myelofibrosis, post-PV myelofibrosis, acute lymphoblastic leukemia (ALL), acute leukemia of an ambiguous lineage, multiple myeloma (MM), myeloproliferative disorders (MPN), and blastic plasmacytoid dendritic cell leukemia. In some embodiments, the pathology is cancer. In some embodiments, cancer is hematopoietic carcinoma. In some embodiments, cancer is leukemia. In some embodiments, the pathophysiology is MDS. MDS is a well-known group of cancers in which immature blood cells (HSPCs) in the bone marrow do not mature into healthy blood cells. In some embodiments, the pathophysiology is CMML. In some embodiments, the pathophysiology is MF. In some embodiments, MF is selected from primary MF, post-ET MF, and post-PV MF. In some embodiments, the pathophysiology is MPN. In some embodiments, the pathophysiology is MDS / MPN. In some embodiments, the pathophysiology is AML. In some embodiments, the pathophysiology is not AML. In some embodiments, the pathophysiology is selected from the group consisting of MDS, CMML, MF, and MPN. In some embodiments, the pathophysiology is selected from the group consisting of MDS, CMML, MF, and MDS / MPN. In some embodiments, the pathophysiology is selected from the group consisting of MDS, CMML, MF, MPN, and AML. In some embodiments, the pathophysiology is selected from the group consisting of MDS, CMML, MF, MDS / MPN, and AML. In some embodiments, MDS is MDS with a del5q mutation.
[0064] In some embodiments, the method is a method for detecting MDS. In some embodiments, a deviation in the quantity or frequency of ERYP cells indicates the presence of MDS. In some embodiments, a deviation in the quantity or frequency of BEMP cells indicates the presence of MDS. In some embodiments, a deviation in the quantity or frequency of MEBEMP cells indicates the presence of MDS. In some embodiments, a deviation in the quantity or frequency of any one of ERYP, BEMP, and MEBEMP cells indicates the presence of MDS. In some embodiments, a deviation in the quantity or frequency of all of ERYP, BEMP, and MEBEMP cells indicates the presence of MDS. In some embodiments, MEBEMP is MEBEMP-L or MEBEMP-E. In some embodiments, MEBEMP is MEBEMP-L and MEBEMP-E. In some embodiments, the deviation is an increase. In some embodiments, a deviation in the quantity or frequency of CLP cells indicates the presence of MDS. In some embodiments, the deviation is a decrease. In some embodiments, CLP is CLP-L, CLP-M, or CLP-E. In some embodiments, CLP is any two of CLP-L, CLP-M, and CLP-E. In some embodiments, CLP is CLP-L, CLP-M, and CLP-E. In some embodiments, a decrease in CLP quantity or frequency indicates the presence of MDS. In some embodiments, MDS is MDS / MPN. In some embodiments, a decrease in CLP quantity or frequency indicates the presence of MDS or MPN. In some embodiments, a deviation in the quantity or frequency of NKTDP cells indicates the presence of MDS. In some embodiments, a decrease in NKTDP quantity or frequency indicates the presence of MDS. In some embodiments, MDS is MDS / MPN. In some embodiments, a decrease in NKTDP quantity or frequency indicates the presence of MDS or MPN.
[0065] In some embodiments, the method is a method for detecting CMML. In some embodiments, a deviation in the quantity or frequency of GMP cells indicates the presence of CMML. In some embodiments, GMP is GMP-E. In some embodiments, the deviation is an increase. In some embodiments, a deviation in the quantity or frequency of CLP cells indicates the presence of CMML. In some embodiments, the deviation is a decrease. In some embodiments, CLP is CLP-L, CLP-M, or CLP-E. In some embodiments, CLP is any two of CLP-L, CLP-M, and CLP-E. In some embodiments, CLP is CLP-L, CLP-M, and CLP-E.
[0066] In some embodiments, the above method is a method for detecting MF. In some embodiments, a deviation in the quantity or frequency of CLP cells indicates the presence of MF. In some embodiments, the deviation is a decrease. In some embodiments, a deviation in the quantity or frequency of NKTDP cells indicates the presence of MF. In some embodiments, a decrease in the quantity or frequency of NKTDP cells indicates the presence of MF.
[0067] In some embodiments, the method is a method for detecting MPN. In some embodiments, a deviation in the quantity or frequency of CLP cells indicates the presence of MPN. In some embodiments, the deviation is a decrease. In some embodiments, CLP is CLP-L, CLP-M, or CLP-E. In some embodiments, CLP is any two of CLP-L, CLP-M, and CLP-E. In some embodiments, CLP is CLP-L, CLP-M, and CLP-E. In some embodiments, a decrease in the quantity or frequency of CLP indicates the presence of MPN. In some embodiments, MPN is MDS / MPN. In some embodiments, a decrease in the quantity or frequency of CLP indicates the presence of MDS or MPN. In some embodiments, a deviation in the quantity or frequency of NKTDP cells indicates the presence of MPN. In some embodiments, a decrease in the quantity or frequency of NKTDP indicates the presence of MPN. In some embodiments, MPN is MDS / MPN. In some embodiments, a decrease in the quantity or frequency of NKTDP indicates the presence of MDS or MPN.
[0068] In some embodiments, the method is a method for detecting AML. In some embodiments, a deviation in the quantity or frequency of CLP cells indicates the presence of AML. In some embodiments, CLP is CLP-L, CLP-M, or CLP-E. In some embodiments, CLP is any two of CLP-L, CLP-M, and CLP-E. In some embodiments, CLP is CLP-L, CLP-M, and CLP-E. In some embodiments, a deviation in the quantity or frequency of NKTDP cells indicates the presence of AML. In some embodiments, the deviation is an increase. In some embodiments, an increase in the quantity or frequency of NKTDP cells indicates the presence of AML.
[0069] In some embodiments, the bone marrow pathology includes an increased percentage of blasts. In some embodiments, the bone marrow pathology is characterized by an increased percentage of blasts. In some embodiments, the bone marrow pathology is selected from AML and MDS. In some embodiments, MDS is MDS / MPN. In some embodiments, the bone marrow pathology is selected from AML, MPN, and MDS. In some embodiments, AML and MDS are characterized by an increased percentage of blasts. In some embodiments, a deviation in the frequency of CLP-E indicates the presence of an increase in the amount of blasts. In some embodiments, the deviation is an increase. In some embodiments, the increase in CLP-E is a deviation. In some embodiments, the magnitude of the increase is proportional to the increase in the amount of blasts. In some embodiments, the increase in blasts is compared to the amount of blasts in healthy controls. In some embodiments, healthy controls are a healthy cohort. In some embodiments, the healthy cohort is the subjects that constitute the control dataset. In some embodiments, linear regression predicts the amount of blasts from the amount of CLP-E.
[0070] In some embodiments, the analysis involves generating a feature vector representing the deviation of the target cell data from the control cell data. In some embodiments, the feature vector contains multiple entries. In some embodiments, each entry corresponds to a specific cell type. In some embodiments, each entry corresponds to a quantity of each cell type. In some embodiments, the quantity is the number. In some embodiments, the quantity is the frequency. In some embodiments, the frequency is the percentage of all CD34-positive cells. In some embodiments, each entry represents or corresponds to a deviation from a reference value. In some embodiments, the deviation is the magnitude of the deviation. In some embodiments, the reference value is the value from the control dataset. In some embodiments, the reference value is the range of quantities of a cell type. In some embodiments, the cell type is a cell population. In some embodiments, the range is the control range. In some embodiments, the range is the healthy range.
[0071] In some embodiments, the analysis involves applying a trained machine learning model to a received dataset. In some embodiments, the machine learning model is trained on a training set. In some embodiments, the training set includes a control dataset. In some embodiments, the training set includes multiple cell datasets. In some embodiments, the machine learning model outputs a classification of the bone marrow of interest. In some embodiments, the machine learning model outputs a classification of the subject. In some embodiments, the machine learning model outputs an analysis of the bone marrow of interest. In some embodiments, the classification is whether or not it is healthy.
[0072] In some embodiments, the training set includes a dataset from healthy subjects. In some embodiments, the training set includes a dataset from subjects suffering from bone marrow conditions. In some embodiments, the training set includes a dataset from subjects suffering from multiple bone marrow conditions. In some embodiments, the training set further includes labels. In some embodiments, labels label the datasets. In some embodiments, labels indicate whether the datasets originate from healthy subjects or subjects with bone marrow conditions. In some embodiments, labels indicate bone marrow conditions. In some embodiments, labels indicate the type of condition. In some embodiments, the classification is either healthy or suffering from multiple bone marrow conditions. In some embodiments, the classification includes classifying what the conditions are. In some embodiments, the classification includes classifying the types of bone marrow conditions.
[0073] In some embodiments, the analysis includes applying a trained machine learning model to parameters extracted from a dataset. In some embodiments, the analysis includes applying a trained machine learning model to feature vectors. In some embodiments, the feature vectors are vectors of cell types. In some embodiments, the cell types are all cell types of CD34-positive cells in the sample. In some embodiments, the cell types are the complete set of CD34-positive cells in the sample. In some embodiments, the machine learning model is trained on a training set. In some embodiments, the training set includes feature vectors from healthy subjects. In some embodiments, the training set includes parameters extracted from a dataset of healthy subjects. In some embodiments, the training set includes feature vectors from subjects suffering from myelopathy. In some embodiments, the training set includes parameters extracted from a dataset of subjects suffering from myelopathy. In some embodiments, the training set includes labels. In some embodiments, the labels indicate that the feature vectors originate from healthy subjects or subjects with myelopathy. In some embodiments, the labels indicate that the extracted parameters originate from healthy subjects or subjects with myelopathy.
[0074] In some embodiments, the analysis further includes applying a trained machine learning model to at least one clinical parameter. In some embodiments, the clinical parameter is the clinical parameter of interest. In some embodiments, the clinical parameter is age. In some embodiments, the clinical parameter is sex. In some embodiments, the clinical parameters are sex and age. In some embodiments, the machine learning model is trained on a training set that includes at least one clinical parameter.
[0075] In another aspect, the present invention provides a method for predicting the amount of blast cells in the bone marrow of a subject, comprising receiving a measurement of CLP-E cells in the peripheral blood of the subject, thereby predicting the amount of blast cells in the bone marrow of the subject.
[0076] In some embodiments, the CLP-E cell measurement is proportional to the amount of blast cells in the bone marrow of the subject. In some embodiments, the proportionality is linear. In some embodiments, linear regression indicates the amount of blast cells from the CLP-E measurement. In some embodiments, indicating is predicting. In some embodiments, a measurement above a predetermined threshold indicates blast cells above a predetermined threshold. In some embodiments, the CLP-E cell measurement is the amount of CLP-E cells. In some embodiments, the CLP-E cell measurement is the number of CLP-E cells. In some embodiments, the CLP-E cell measurement is the percentage of CLP-E cells among CD34-positive cells in peripheral blood. CLP-E cells can be measured by any method known in the art, including flow cytometry, immunostaining, sequencing, and metacell generation from scRNA-seq. Methods for identifying these cells in samples, including blood samples, are known in the art, and any such method may be used. For example, a method for identifying CLP-E cells is provided below, as well as in Ding and Morrison, "Haematopoietic stem cells and early lymphoid progenitors occupy distinct bone marrow niches," Nature. 2013, Mar 14;495(7440):231-235, the contents of which are incorporated herein by reference in their entirety.
[0077] In some embodiments, the method further includes receiving a peripheral blood sample. In some embodiments, the method further includes measuring CLP-E cells in the sample. In some embodiments, the measurement is counting. In some embodiments, the method further includes receiving scRNA-seq data from CD34-positive cells in the blood and calculating the number / quantity / percentage of CLP-E cells in the blood. In some embodiments, the blood is present in the sample. In some embodiments, the method further includes analyzing the received measurements in relation to a control dataset.
[0078] In another aspect, a method for predicting the amount of blast cells in a target bone marrow, wherein the method is a. Obtain a dataset based on CD34-positive cells from the target blood sample, b. Applying a trained machine learning model to a received dataset, wherein the machine learning model outputs a predicted quantity of blast cells in the bone marrow of interest, This provides a method that includes predicting the amount of blast cells in the bone marrow of a subject.
[0079] In some embodiments, the subject is a mammal. In some embodiments, the mammal is a human. In some embodiments, the subject requires the method of the present invention. In some embodiments, the subject is male. In some embodiments, the subject is female. In some embodiments, the subject suffers from a bone marrow pathology. In some embodiments, the bone marrow pathology is a myeloma. In some embodiments, the subject suffers from leukemia. In some embodiments, the leukemia is selected from AML, CMML, CML, mastocytosis, chronic eosinophilic leukemia, ambiguous lineage acute leukemia, and blastic plasmacytoid dendritic cell leukemia.
[0080] In some embodiments, the blast quantity is the number of blasts. In some embodiments, the blast quantity is the frequency of blasts. In some embodiments, the blast quantity is the percentage of blasts in the bone marrow. In some embodiments, the percentage is relative to all cells in the bone marrow. In some embodiments, all cells are CD34-positive cells.
[0081] In some embodiments, the training set includes subjects with MDS. In some embodiments, the training set includes non-MDS subjects. In some embodiments, the training set includes leukemia subjects. In some embodiments, the training set includes leukemia and non-leukemia subjects. In some embodiments, the training set further includes labels. In some embodiments, the labels label the dataset. In some embodiments, the labels indicate the amount of blasts in the subjects that provided the dataset. In some embodiments, the percentage of blasts in the bone marrow is known for each subject in the control dataset. In some embodiments, the subjects in the control dataset are the subjects that provided the data for the control dataset. In some embodiments, the dataset is a dataset of multiple datasets. In some embodiments, the dataset is a control dataset. In some embodiments, the machine learning model outputs the amount of blasts in the subjects.
[0082] In some embodiments, the method is a method for detecting MDS, where a number of blast cells exceeding a predetermined threshold indicates that the subject has MDS. In some embodiments, the method is a method for detecting leukemia, where a number of blast cells exceeding a predetermined threshold indicates that the subject has leukemia. In some embodiments, the threshold is 0%. In some embodiments, the threshold is 5%. In some embodiments, the threshold is 9%. In some embodiments, the threshold is 10%. In some embodiments, the threshold is 15%.
[0083] In some embodiments, the method further includes the step of not administering the therapeutic agent to subjects having a quantity of blast cells below a predetermined threshold. In some embodiments, the method further includes the step of administering the therapeutic agent to subjects determined to be suffering from a bone marrow pathology. In some embodiments, the method further includes the step of administering the therapeutic agent to subjects having a quantity of blast cells exceeding a predetermined threshold. In some embodiments, the drug is an anticancer agent, and the subjects have cancer. In some embodiments, the cancer is MDS. In some embodiments, the drug is an anti-MDS agent. In some embodiments, the anti-MDS agent is lenalidomide. In some embodiments, the drug is an anti-leukemia agent. Anticancer agents are well known in the art, and any such drug may be used, including but not limited to chemotherapy, radiotherapy, immunotherapy, and targeted therapy. In some embodiments, the drug is chemotherapy. In some embodiments, the drug is radiotherapy. In some embodiments, the drug is immunotherapy. In some embodiments, the immunotherapy is immune checkpoint inhibition. In some embodiments, the checkpoint is PD-1 / PD-L1. In some embodiments, the immunotherapy is CAR-T therapy or CAR-NK therapy. In some embodiments, the anticancer agent is a hypomethylating agent. In some embodiments, the hypomethylating agent is azacitidine. In some embodiments, the hypomethylating agent is decitabine. In some embodiments, the anticancer agent is azacitidine in combination with venetoclax. In some embodiments, the subject has leukemia and the anticancer agent is venetoclax. In some embodiments, the subject has MDS and the drug is azacitidine. In some embodiments, the subject has MDS and the drug is azacitidine in combination with venetoclax. In some embodiments, the leukemia is chronic lymphocytic leukemia, small lymphocytic lymphoma, or acute myeloid leukemia. In some embodiments, the method further comprises performing a bone marrow transplant on a subject determined to have a bone marrow pathology. In some embodiments, the method further comprises performing a bone marrow transplant on a subject having a quantity of blast cells exceeding a predetermined threshold.In some embodiments, the subject is suffering from MPN, and the drug is interferon. In some embodiments, the subject is suffering from MPN, and the method further comprises administering interferon therapy. In some embodiments, the interferon is interferon alpha. In some embodiments, the interferon is type I interferon. In some embodiments, the interferon is interferon beta. In some embodiments, interferon beta is interferon beta 1 (IFNB1). In some embodiments, the interferon is interferon alpha. In some embodiments, interferon alpha is selected from interferon alpha 1, 2, 4, 5, 6, 7, 8, 10, 13, 14, 16, 17, and 21. In some embodiments, the interferon is interferon alpha-2b. In some embodiments, the drug is LOPEG interferon alpha-2b (Besremi).
[0084] A method for calculating the Molecular International Prognostic Scoring System (IPSS-M) risk score in another aspect, wherein the method is a. Predicting the percentage of blast cells in the target bone marrow by the method of the present invention, b. To receive data regarding the presence of bone marrow mutations and / or karyotype abnormalities in the subjects, c. Obtaining hemoglobin levels and / or platelet counts in the peripheral blood from the subject, d. Calculating an IPSS-M risk score based on the predicted blast percentage, received mutation and / or karyotype analysis data, and received hemoglobin levels and / or platelet count, This provides a method for calculating the IPSS-M risk score.
[0085] In some embodiments, the method further includes detecting the presence of bone marrow mutations. In some embodiments, the method further includes detecting karyotype abnormalities. In some embodiments, detection is in scRNA data. In some embodiments, detection is non-invasive detection. In some embodiments, detection does not involve detection within the bone marrow. It will be understood that all steps of the method can be performed non-invasively, and one of the main advantages of the method of the present invention is that it does not require bone marrow samples to learn important information about the bone marrow (e.g., IPSS score). A method for performing karyotype and mutation analysis from scRNA-seq data is described below. Furthermore, these are disclosed in the relevant art, such as Weissbein et al., "Analysis of chromosomal aberrations and recombination by allelic bias in RNA-Seq," Nature Communications volume 7, Article number: 12144 (2016), and Petti et al., "A general approach for detecting expressed mutations in AML cells using single cell RNA sequencing," Nature Communications volume 10, Article number: 3660 (2019), and are incorporated herein by reference in their entirety.
[0086] In some embodiments, the mutation or karyotype abnormality is del(5q). In some embodiments, the mutation or karyotype abnormality is -7 / del(7q). In some embodiments, the mutation or karyotype abnormality is -17 / del(17p). In some embodiments, the mutation or karyotype abnormality is a complex karyotype. In some embodiments, the mutation or karyotype abnormality is del(11q). In some embodiments, the mutation or karyotype abnormality is del(5q). In some embodiments, the mutation or karyotype abnormality is del(12p). In some embodiments, the mutation or karyotype abnormality is del(20q). In some embodiments, the mutation or karyotype abnormality is del(7q). In some embodiments, the mutation or karyotype abnormality is +8. In some embodiments, the mutation or karyotype abnormality is +19. In some embodiments, the mutation or karyotype abnormality is i(17q). In some embodiments, the mutation or karyotype abnormality is -Y. In some embodiments, the mutation or karyotype abnormality is -7. In some embodiments, the mutation or karyotype abnormality is (inv)3 / t(3q) / del(3q).
[0087] In some embodiments, the mutation is a mutated allele. In some embodiments, the mutation is a mutation within the tumor protein p53 (TP53). In some embodiments, the mutation is the number of mutations. In some embodiments, the mutation or karyotype abnormality is a loss of heterozygosity at the TP53 locus. In some embodiments, the mutation is a MLL (lysine methyltransferase 2A (KMT2A)) mutation. In some embodiments, the mutation is an fms-associated receptor tyrosine kinase 3 (FLT3) mutation. In some embodiments, the mutation is an ASXL transcription regulator 1 (ASXL1) mutation. In some embodiments, the mutation or karyotype abnormality is a Cbl proto-oncogene (CBL) mutation. In some embodiments, the mutation is a DNA methyltransferase 3 alpha (DNMT3A) mutation. In some embodiments, the mutation is an ETS variant transcription factor 6 (ETV6) mutation. In some embodiments, the mutation is an Enhancer of Zeste 2 Polycomb Repressive Complex 2 Subunit (EZH2) mutation. In some embodiments, the mutation is an isocitrate dehydrogenase (NADP(+))2 (IDH2) mutation. In some embodiments, the mutation is a KRAS proto-oncogene GTPase (KRAS) mutation. In some embodiments, the mutation is a nucleophosmin 1 (NPM1) mutation. In some embodiments, the mutation is a NRAS proto-oncogene GTPase (NRAS) mutation. In some embodiments, the mutation is a RUNX family transcription factor 1 (RUNX1) mutation. In some embodiments, the mutation is a splicing factor 3b subunit 1 (SF3B1) mutation. In some embodiments, the mutation is a serine and arginine-rich splicing factor 2 (SRSF2) mutation. In some embodiments, the mutation is a U2 micronuclear RNA cofactor 1 (U2AF1) mutation. In some embodiments, the mutation is a BCL6 corepressor (BCOR) mutation. In some embodiments, the mutation is a BCL6 corepressor-like 1 (BCORL1) mutation. In some embodiments, the mutation is a CCAAT enhancer-binding protein alpha (CEBPA) mutation.In some embodiments, the mutation is in ethanolamine kinase 1 (ETNK1). In some embodiments, the mutation is in GATA-binding protein 2 (GATA2). In some embodiments, the mutation is in G protein subunit β1 (GNB1). In some embodiments, the mutation is in isocitrate dehydrogenase (NADP(+)) 1 (IDH1). In some embodiments, the mutation is in neurofibromin 1 (NF1). In some embodiments, the mutation is in PHD finger protein 6 (PHF6). In some embodiments, the mutation is in protein phosphatase, Mg2+ / Mn2+-dependent 1D (PPM1D). In some embodiments, the mutation is in pre-mRNA processing factor 8 (PRPF8). In some embodiments, the mutation is in protein tyrosine phosphatase nonreceptor type 11 (PTPN11). In some embodiments, the mutation is in SET-binding protein 1 (SETBP1). In some embodiments, the mutation is a mutation in the STAG2 cohesin complex component (STAG2). In some embodiments, the mutation is a mutation in the WT1 transcription factor (WT1).
[0088] In some embodiments, hemoglobin levels are received. In some embodiments, the method further includes measuring hemoglobin levels. In some embodiments, the method includes receiving a blood sample from a subject. In some embodiments, hemoglobin levels are calculated in the blood sample. In some embodiments, platelet counts are received. In some embodiments, the method further includes counting platelets. In some embodiments, platelets are present in the blood sample. In some embodiments, the method further includes receiving neutrophil counts. In some embodiments, the method further includes counting neutrophils. In some embodiments, neutrophils in the sample are counted. In some embodiments, the age of the subject is also received. In some embodiments, the sex / gender of the subject is also received.
[0089] In some embodiments, the IPSS-M risk score is calculated based on any combination of received data. In some embodiments, the IPSS-M risk score is calculated based on the predicted blast percentage. In some embodiments, the IPSS-M risk score is calculated based on the predicted blast percentage and the received mutation and karyotype analysis. In some embodiments, the IPSS-M risk score is calculated based on the predicted blast percentage and the received hemoglobin level and platelet count. In some embodiments, the IPSS-M risk score is calculated based on the predicted blast percentage, the received mutation and karyotype analysis and the received hemoglobin level and platelet count. In some embodiments, the IPSS-M risk score is further calculated based on neutrophil count and / or patient age.
[0090] The IPSS-M score is well known in the art. It ranges from 0 to 16. The score is divided into six risk categories: very low (VL) risk, low (L) risk, moderate-low (ML) risk, moderate-high (MH) risk, high risk (H), and very high (VH) risk. Low-risk subjects may receive no treatment or treatment to manage symptoms, such as erythrocyte stimulants (ESAs) to treat anemia. Patients with thrombocytopenia may receive romiplostim or eltrombopag. Similarly, ruspatercept may be administered if ESA is ineffective (and / or if there is a mutation in SF3B1 or a cyclic sideroblast is present). High-risk subjects may receive hypomethylating agents or other anticancer drug treatments. High-risk subjects may receive bone marrow transplantation.
[0091] In some embodiments, the method further includes administering a treatment regimen to subjects based on a calculated IPSS-M score. In some embodiments, subjects with higher scores are administered a more potent treatment regimen. In some embodiments, subjects with lower scores are administered a reduced treatment regimen. In some embodiments, a stronger intensity is increased. In some embodiments, a weaker intensity is lower.
[0092] In another embodiment, a method for detecting AML in a subject is provided, which includes detecting the presence of the R353K mutation in GATA3 in a sample from the subject, thereby detecting AML in the subject.
[0093] In some embodiments, the sample includes cells. In some embodiments, the cells are hematopoietic cells. In some embodiments, the cells are blast cells. In some embodiments, the cells are CD34-positive cells. In some embodiments, the mutation is a mutation at arginine 353 in GATA3. In some embodiments, the arginine is mutated to lysine. In some embodiments, the mutation indicates AML.
[0094] Next, refer to Figure 6, a block diagram showing a computing device that may be included in one embodiment of a system for analyzing bone marrow or calculating an IPSS-M risk score, according to several embodiments.
[0095] Computing device 1 may include, for example, a processor or controller 2 which can be a central processing unit (CPU) processor, chip, or any suitable computing or calculating device, an operating system 3, memory 4, executable code 5, a storage system 6, an input device 7, and an output device 8. The processor 2 (or one or more controllers or processors that may span multiple units or devices) may be configured to perform the methods described herein and / or to perform or function as various modules, units, etc. Two or more computing devices 1 may be included in a system according to embodiments of the present invention, and one or more computing devices 1 may function as components thereof.
[0096] Operating System 3 may be or include any code segment (for example, similar to executable code 5 described herein) designed and / or configured to coordinate, schedule, arbitrate, supervise, control, or otherwise manage the operation of computing device 1, for example, scheduling the execution of software programs or tasks, or enabling software programs or other modules or units to communicate. Operating System 3 may be a commercial operating system. It should be noted that Operating System 3 may be an optional component, and for example, in some embodiments, the system may include computing devices that do not require or include Operating System 3.
[0097] Memory 4 may be, or may include, random access memory (RAM), read-only memory (ROM), dynamic RAM (DRAM), synchronous DRAM (SD-RAM), double data-rate (DDR) memory chips, flash memory, volatile memory, non-volatile memory, cache memory, buffer, short-term memory unit, long-term memory unit, or other suitable memory unit or storage unit. Memory 4 may, in some cases, be multiple different memory units, or may include them. Memory 4 may be a non-temporary readable medium of a computer or processor, or a non-temporary storage medium of a computer, such as RAM. In one embodiment, a non-temporary storage medium such as memory 4, a hard disk drive, or another storage device may store instructions or code that, when executed by the processor, cause the processor to execute the methods described herein.
[0098] The executable code 5 may be any executable code, such as an application, program, process, task, or script. The executable code 5 may optionally be executed by a processor or controller 2 under the control of the operating system 3. For example, the executable code 5 may be an application capable of calculating a target IPSS-M score, as further described herein. For clarity, a single executable code 5 is shown in Figure 6, but systems according to some embodiments of the present invention may include multiple executable code segments similar to the executable code 5 that can be loaded into memory 4. Processor 2 is made to perform the method described herein.
[0099] The storage system 6 may be, for example, a flash memory as known in the art, a memory located inside or embedded in a microcontroller or chip as known in the art, a hard disk drive, a CD-Recordable (CD-R) drive, a Blu-ray disc (BD), a Universal Serial Bus (USB) device, or other suitable removable and / or fixed storage units, or may include them. Data associated with single-cell RNA sequencing (scRNA-seq) reads may be stored in the storage system 6, loaded from the storage system 6 into memory 4, where it may be processed by a processor or controller 2. In some embodiments, some of the components shown in Figure 1 may be omitted. For example, memory 4 may be a non-volatile memory having the storage capacity of the storage system 6. Thus, although shown as a separate component, the storage system 6 may be embedded in or included in memory 4.
[0100] Input device 7 may be or may include any suitable input device, component, or system, such as a detachable keyboard or keypad, a mouse, etc. Output device 8 may include one or more (potentially detachable) displays or monitors, speakers, and / or any other suitable output devices. Any applicable input / output (I / O) devices may be connected to computing device 1 as shown by blocks 7 and 8. For example, a wired or wireless network interface card (NIC), a Universal Serial Bus (USB) device, or an external hard drive may be included in input device 7 and / or output device 8. It will be recognized that any suitable number of input devices 7 and output devices 8 may be operably connected to computing device 1 as shown by blocks 7 and 8.
[0101] Systems according to some embodiments of the present invention may include, but are not limited to, a plurality of central processing units (CPUs) or any other suitable multipurpose or specific processor or controller (for example, similar to element 2), a plurality of input units, a plurality of output units, a plurality of memory units, and a plurality of storage units.
[0102] The terms neural networks (NN) or artificial neural networks (ANN), for example, neural networks that implement machine learning (ML) or artificial intelligence (AI) functions, may be used herein to refer to an information processing paradigm that may include nodes called neurons, organized into layers with links between neurons. Links may transmit signals between neurons and may be associated with weights. An NN may be configured or trained for a specific task, e.g., pattern recognition or classification. Training an NN for a specific task may involve adjusting these weights based on examples. Each neuron in a hidden or final layer may receive an input signal, e.g., a weighted sum of output signals from other neurons, and may process the input signal using a linear or nonlinear function (e.g., an activation function). The results of the input and hidden layers may be forwarded to other neurons, and the results of the output layer may be provided as the output of the NN. Typically, neurons and links in an NN are represented by activation functions and mathematical constructs such as matrices consisting of data elements and weights. The related calculations may be performed by at least one processor, such as one or more CPUs or graphics processing units (GPUs) (e.g., Processor 2 in Figure 6), or by a dedicated hardware device.
[0103] Next, we refer to Figure 7, which shows a system 10 for analyzing a target bone marrow according to several embodiments.
[0104] According to some embodiments of the present invention, the system 10 may be implemented as a software module, a hardware module, or any combination thereof. For example, the system may be a computing device such as element 1 in Figure 6, or may include one, which may be adapted to execute one or more modules of executable code (e.g., element 5 in Figure 6) to analyze the bone marrow of interest, as further described herein.
[0105] As shown in Figure 7, the arrows may represent the flow of one or more data elements to and from system 10, and / or between modules or elements of system 10. For clarity, some arrows have been omitted in Figure 7.
[0106] In some embodiments, the analysis includes generating feature vectors that represent the deviation of the target cell data from the control cell data.
[0107] As shown in Figure 7A, the system 10 may include, or be communicably connected to, a single-cell RNA sequencing (scRNA-seq) module or device 20 which may be configured to generate scRNA-seq data 20S (or for short, “data 20S”), as detailed herein.
[0108] The analysis module 100 of system 10 may be configured to analyze data 20S in order to extract feature vectors 150F. As detailed herein, feature vectors 150F may include one or more values that indicate a CD34-positive population in peripheral blood samples of the subject of interest (e.g., a patient).
[0109] For example, feature vector 150F may contain multiple entries, each corresponding to a specific cell type. The value of each entry in feature vector 150F may represent a relationship to a reference value, a deviation from the reference value, or a range of cell populations.
[0110] Referring to the example in Figure 2D (upper panel), reference values for the frequency range and mean of each type in the stem cell population may be shown by gray and dashed lines. In such embodiments, an entry in feature vector 150F for a particular subject (e.g., #115, green line) may include a value for that subject's stem cell population. Additionally or alternatively, entries in feature vector 150F may include statistical values representing deviations from a reference. Such a reference may include the mean (black dashed line) and / or normal range (gray line) of the stem cell population in the subject's cohort.
[0111] In some embodiments, the analysis includes applying a trained machine learning (ML)-based module 200, also referred to herein as classifier 200, to a received dataset 20S. Additionally or alternatively, the analysis may include applying the ML 200 to feature vectors 150F. In some embodiments, the ML module is trained on a training set. In some embodiments, the training set includes a control dataset. In some embodiments, the training set includes multiple cell datasets.
[0112] In some embodiments, the ML200 may output a display 30 (for example, via the output device 8 in Figure 6). The display 30 may be, for example, a classification of the bone marrow of the subject. In some embodiments, the display 30 may include the classification of the subject and the analysis of the bone marrow of the subject. Additionally or alternatively, the display 30 may include notifications regarding the health status of the subject (e.g., whether it is healthy or not), the diagnosis of the subject (e.g., suspected bone marrow condition), the prognosis of the subject's condition, etc.
[0113] In some embodiments, the analysis may include applying ML200 to parameters extracted from the dataset. In some embodiments, the analysis may include applying a trained machine learning model to the feature vector 150F.
[0114] A block diagram showing a non-limiting example for implementing system 10 according to some embodiments of the present invention is shown in Figure 7B. System 10 in Figure 7B may be the same as system 10 in Figure 7A.
[0115] As shown in Figure 7B, the analysis module 100 may include a feature extraction module 110. As detailed herein, the feature extraction module 110 may be configured to extract from the data 20 a plurality of features 110F, or parameters relating to or revealing the characteristics of a plurality of specific cells in a peripheral blood sample. These features may be gene expression profiles of informational value or other transcriptional data extracted from scRNA-seq data. The features may be a valuable portion of the whole transcriptome or the cellular transcriptome.
[0116] The analysis module 100 can use feature 110F to bin or cluster feature 110F to form a high-level representation of the cell population in the peripheral blood sample.
[0117] For example, the target module 130 of the analysis module 100 may be configured to generate at least one target-specific model 130M. The target-specific model 130M may be associated with a specific peripheral blood test taken from a particular target. In some embodiments, the target-specific model 130M may include multiple metacell entities, each representing an abstraction of cell population data belonging to its target, as detailed herein.
[0118] Additionally or alternatively, the cohort reference generation module 120 of the analysis module 100 may be configured to generate reference data 120M or cohort data model 120M, also referred to herein as HSPC atlas 120M. In some embodiments, the reference data 120M may include multiple metacell entities, each representing an abstraction of cell population data belonging to the cohort in question, as detailed herein.
[0119] As shown in Figure 7B, the analysis module 100 may include a projection module 150 configured to project or compare specific object features 110F, as revealed by a target-specific model 130M, onto features of a cohort of objects, as revealed by reference data (e.g., HSPC atlas) 120M.
[0120] According to some embodiments, based on this comparison or projection, the projection module 150 may generate a feature vector 150F, also referred to herein as a “normal vector” 150F. The normal vector 150F may represent the state of a particular object.
[0121] According to some embodiments, the system 10 may infer a classifier 200 on feature vectors (e.g., normal vectors) 150F to generate the representation 30 of Figure 7A. In such embodiments, the classifier 200 may be, or include, an ML-based classification model that can be trained on a training dataset containing a plurality of labeled or annotated normal vectors 150F. The annotations on the normal vectors 150F may include, for example, expert representations 30 (e.g., diagnoses) of the corresponding peripheral blood sample. The ML-based classification model 200 may be trained by a supervised training scheme to generate representations 30 of the input normal vectors 150F using the annotations as training data.
[0122] Additionally or alternatively, system 10 may infer a classifier 200 on the target-specific model 130M data to generate the representation 30. In such embodiments, the classifier 200 may be, or include, an ML-based classification model that can be trained on a training dataset containing multiple labeled or annotated target-specific model 130M data entities. The annotations on the target-specific model 130M in the dataset may include, for example, expert representations 30 (e.g., diagnoses) of the corresponding peripheral blood sample. Thus, the ML-based classification model 200 can be trained to generate the representation 30 by a supervised training scheme using the annotations as training data.
[0123] Additionally or alternatively, system 10 may infer a classifier 200 on feature vectors (e.g., normal vectors) 150F to generate predictions of blast cell levels 210B in the bone marrow. In such embodiments, the classifier 200 may be, or include, an ML-based classification model 210 that can be trained on a training dataset containing multiple labeled or annotated normal vectors 150F. The annotations on the normal vectors 150F may include the levels of blast cells 210B in the bone marrow corresponding to each patient peripheral blood sample. The ML-based classification model 210 can be trained to predict bone marrow blast cell levels 210B by a supervised training scheme using the annotations as training data.
[0124] Additionally or alternatively, the system 10 may include an auxiliary data extraction module 140 (or abbreviated as “auxiliary module 140”). For example, the auxiliary module 140 may be configured to generate auxiliary information 140A, such as karyotype data 140A or mutation data, from the data 20, as is known in the art. In such embodiments, the classifier module 200 may include an IPSS-M risk score calculation module 220, configured to calculate an IPSS-M risk score 220S based on predicted myeloblast levels 210B, calculated karyotype data 140A, mutation data, and other clinical blood measurements, as is known in the art.
[0125] Here, we refer to Figure 8, a flowchart illustrating a method for analyzing target bone marrow using at least one processor, according to several embodiments.
[0126] As shown in step S1005, at least one processor (e.g., processor 2 in Figure 6) may receive a target cell dataset 20S based on single-cell RNA sequencing (scRNA-seq) 20 of CD34-positive cells from the target peripheral blood.
[0127] As shown in step S1010, at least one processor may analyze the received target cell dataset (e.g., 130M) in relation to a control dataset (e.g., 120M) comprising multiple cell datasets, using the analysis module 100 (e.g., as detailed herein in relation to Figures 7A, 7B). Each of the multiple cell datasets may be based on scRNA-seq 20S of CD34-positive cells from the peripheral blood of a healthy subject, and deviations of the target cell dataset from the control dataset (e.g., feature vectors, or normal vectors 150F) may indicate bone marrow pathology. Thereafter, embodiments of the present invention may generate a display 30 that represents or indicates the detection of bone marrow pathology in the subject.
[0128] In another aspect, a system for carrying out the method of the present invention is provided.
[0129] In some embodiments, the system is used to assess bone marrow health. In some embodiments, the system is used to measure the number of blast cells in the bone marrow. In some embodiments, the system is a non-invasive system.
[0130] In some embodiments, the system includes an scRNA sequencing device. In some embodiments, the sequencing device is an scRNA sequencer. In some embodiments, the system includes a non-temporary memory device in which modules of instruction code are stored. In some embodiments, the system includes at least one processor. In some embodiments, the processor is associated with the memory device. In some embodiments, the processor is configured to perform the method of the present invention. In some embodiments, the processor is configured to execute modules of instruction code, and when modules of instruction code are executed, at least one processor is configured to perform the method of the present invention.
[0131] In some embodiments, the method includes obtaining a single-cell transcriptome from CD34-positive cells in peripheral blood using an scRNA sequencing device. In some embodiments, the peripheral blood is derived from a subject. In some embodiments, the method includes generating a cell dataset from the obtained single-cell transcriptome. In some embodiments, the method includes generating a cell dataset based on the obtained single-cell transcriptome. In some embodiments, the method includes generating a cell dataset derived from the obtained single-cell transcriptome. In some embodiments, the method includes analyzing the generated dataset. In some embodiments, the analysis relates to a control dataset. In some embodiments, the method includes accessing the control dataset. In some embodiments, the control dataset is a control database. In some embodiments, the control dataset is multiple datasets. In some embodiments, the method includes outputting findings. In some embodiments, the findings are the health status of the subject. In some embodiments, the findings are the health status of the bone marrow. In some embodiments, the findings are healthy. In some embodiments, the findings are the presence of a bone marrow pathology. In some embodiments, the findings are what the bone marrow pathology is. In some embodiments, the findings are based on whether or not there is a deviation of the subject dataset from the control dataset.
[0132] As used herein, the term "approximately," when combined with a value, refers to a range of ±10% of the reference value. For example, a length of approximately 1000 nanometers (nm) refers to a length of 1000 nm ± 100 nm.
[0133] It should be noted that, as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include multiple referents unless otherwise explicitly indicated in the context. Thus, for example, a reference to “polynucleotide” includes multiple such polynucleotides, and a reference to “polypeptide” includes a reference to one or more polypeptides and their equivalents known to those skilled in the art, and so on. It should be further noted that the claims may be drafted to exclude any element. Therefore, this statement is intended to serve as an antecedent for the use of exclusive terms such as “alone,” “only,” etc., in connection with the enumeration of elements in the claims or the use of “negative” limitations.
[0134] Where conventions similar to “at least one of A, B, and C, etc.” are used, such configurations are generally intended to be understood by those skilled in the art (for example, “a system having at least one of A, B, and C” includes, but is not limited to, systems having only A, only B, only C, A and B together, A and C together, B and C together, and / or systems having A, B and C together). It will further be understood by those skilled in the art that substantially any disjunctive word and / or phrase presenting two or more alternative terms should be understood as construing the possibility of including one of the terms, either of the terms, or both of the terms, whether in the specification, claims, or drawings. For example, the phrase “A or B” is understood to include the possibilities of “A” or “B” or “A and B”.
[0135] For clarity, certain features of the Invention described in the context of separate embodiments may be provided in combination in a single embodiment. Conversely, various features of the Invention described in the context of a single embodiment for brevity may be provided separately or in any suitable partial combination. All combinations of embodiments relating to the Invention are specifically encompassed by the Invention and are disclosed herein as if each and all combinations were individually and expressly disclosed. Furthermore, all partial combinations of various embodiments and their elements are also specifically encompassed by the Invention and are disclosed herein as if each and all such partial combinations were individually and expressly disclosed herein.
[0136] As used herein and in the appended claims, the singular forms “a,” “an,” and “the” refer to multiple subjects unless otherwise explicitly indicated by the context. The terms “a” (or “an”), as well as “one or more” and “at least one,” are interchangeable.
[0137] Furthermore, “and / or” should be interpreted as a specific disclosure of each of the two specified features or components, with or without the other. Thus, the term “and / or” as used in phrases such as “A and / or B” is intended to include A and B, A or B, A (alone), and B (alone). Similarly, the term “and / or” as used in phrases such as “A, B, and / or C” is intended to include A, B, and C; A, B, or C; A or B; A or C; B or C; A and B; A and C; B and C; A (alone); B (alone); and C (alone).
[0138] Whenever an embodiment is described using the word “comprising,” it includes other similar embodiments described using the terms “consisting of” and / or “consisting essentially of.”
[0139] Further objectives, advantages, and novel features of the present invention will become apparent to those skilled in the art by considering the following examples, which are not intended to be limiting. In addition, each of the various embodiments and aspects of the present invention described above herein and claimed in the following claims section will find experimental support in the following examples.
[0140] The various embodiments and aspects of the present invention described above and claimed in the following claims section will be experimentally supported in the following examples. [Examples]
[0141] In general, the nomenclature used herein and the laboratory procedures utilized in the present invention include molecular, biochemical, microbiological, and recombinant DNA techniques. Such techniques are well described in the literature. For example, "Molecular Cloning: A Laboratory Manual," Sambrook et al., (1989); "Current Protocols in Molecular Biology," Volumes I-III, Ausubel, RM, ed. (1994); Ausubel et al., "Current Protocols in Molecular Biology," John Wiley and Sons, Baltimore, Maryland (1989); Perbal, "A Practical Guide to Molecular Cloning," John Wiley & Sons, New York (1988); Watson et al., "Recombinant DNA," Scientific American Books, New York; Birren et al. (eds), "Genome Analysis: A Laboratory Manual Series," Vols. 1-4, Cold Spring Harbor Laboratory Press, New York York (1998); Methods described in U.S. Patent Nos. 4,666,828, 4,683,202, 4,801,531, 5,192,659, and 5,272,057; "Cell Biology: A Laboratory Handbook," Volumes I-III, Cellis, JE, ed. (1994); "Culture of Animal Cells - A Manual of Basic Technique," by Freshney, Wiley-Liss, NY (1994), Third Edition; "Current Protocols in Immunology," Volumes I-III, Coligan, JE, ed.(1994); See Stites et al. (eds), "Basic and Clinical Immunology" (8th Edition), Appleton & Lange, Norwalk, CT (1994); Mishell and Shiigi (eds), "Strategies for Protein Purification and Characterization - A Laboratory Course Manual," CSHL Press (1996). All of these are incorporated by reference. Other general references are provided throughout this document.
[0142] material and method Sample Procurement and Handling: Fresh peripheral blood samples were collected from 148 healthy individuals (79 males and 69 females) aged 23–91 years. All sample donors were considered healthy, their CBCs were within the normal range, and they were not known to have CH-defining mutations prior to sequencing. In accordance with the Declaration of Helsinki, written informed consent was obtained from all participants, allowing access to longitudinal CBC and sequencing data (CH and genotyping panel). All protocols were approved by the Weizmann Institute of Science Ethics Committee (under IRB Protocol 283-1), in accordance with all applicable ethical rules.
[0143] Recruitment was intended to enable the characterization of normal fluctuations in cHSPC status. Since such profiling had not been performed previously, little could be assumed beforehand regarding the population variance. Therefore, the objective was to profile volunteers with normal blood cell counts, balanced by sex, and to determine a varianced age distribution biased towards older individuals. This strategy was re-evaluated after initial sampling, and a remarkable homogeneity of transcriptional status was observed across individuals sampled directly from the community, as well as from HMO outpatient clinics, emergency medical centers, hospital wards, etc. This was crucial as it confirmed the universality of the model across individuals.
[0144] 50 ml of peripheral blood (PB) was collected from each individual in lithium heparin tubes. 1 ml of blood was used for DNA preparation, and the remaining volume was used for PBMC isolation by Ficoll using Lymphoprep-filled Sepmate tubes (StemCell Technologies), followed by enrichment based on CD34 magnetic beads using the EasySep Human CD34 Positive Selection Kit II (StemCell Technologies). This enrichment strategy has been found to be simple and reproducible and was chosen for several reasons: 1) RNA-seq data were most reproducible when cells were not sorted and were enriched using beads (lower mitochondrial gene fraction). 2) CD34 purity can be highly controlled by this method to achieve enrichment of 50-95% of CD34-positive cells, which can later be easily distinguished based on their single-cell expression data. Regarding cell count, 50 ml of blood was expected to generate 50 to 100 million PBMCs after Ficoll, with 1 / 1000 of them being CD34+. Consequently, the presentation of this population increased from 0.1% in the periphery to at least 50% of the cells loaded for analysis.
[0145] CD34+PBMC scRNA-seq: Single-cell RNA libraries were generated using the 10× genomic scRNA-seq platform (Chromium Next Gem single-cell 3' reagent kit V3.1). Flow cytometry was performed before tip loading to confirm successful enrichment and sufficient CD34+CD45 int We confirmed that viable cells were collected. All blood samples were freshly collected at the Weizmann Institute of Science on the morning of each experimental day, and the time from blood collection to 10x loading was limited to 5 hours. The motivation for working with fresh samples was based on previous experience that PB CD34+ cells are fragile to repeated freezing and thawing and prolonged handling.
[0146] All 10x libraries were sequenced using two alternative platforms (Illumina / Ultima Genomics). Twelve libraries were sequenced simultaneously on both platforms for comparison and to demonstrate the scalability of this approach. The Ultima sequence data were observed to be highly similar to the Illumina sequence data.
[0147] Genotype-based demultiplexing: Genotype-based demultiplexing was used to trace all cells back to their source sample. This method allowed pooling of blood samples immediately after DNA aliquot extraction, resulting in CD34 enrichment across the entire pool of resulting PBMCs. The use of SNP-based multiplexing has several advantages over alternative antibody-based cell hashing methods: 1) It is extremely cost-effective, as the cost of sequencing a single organism with a 2000 SNP Molecular Inversion Probe (MIP) panel at a depth of 1000X per SNP (suitable for demultiplexing purposes) is several times cheaper than antibody staining; 2) Genotyping does not require antibody incubation periods and multiple wash centrifugations, eliminating the need to separate samples before loading, resulting in shorter handling times and less cell manipulation. This was very evident in cell viability before tip loading. Similar to other sample multiplexing methods, genotype-based multiplexing enabled robust doublet detection during data analysis, allowing for the loading of 30–40K cells from 4–6 individuals on each chromium tip lane, resulting in 15–25K cells per library.
[0148] Molecular inversion probe (MIP) panels: Both the CH and genotyping panels are molecular inversion probe-based (MIP) panels, as previously described in detail in Biezuner, T. et al., “An improved molecular inversion probe based targeted sequencing approach for low variant allele frequency,” NAR Genom Bioinform 4, (2022), which is incorporated herein by reference in its entirety. The CH panel contains 705 probes covering preleukemic SNVs and indels in 47 genes, complemented by a two-amplicon sequencing reaction to cover GC-rich regions in SRSF2 and ASXL1. Because MIP sequencing is cost-effective but noisy, an in-house variant calling method was designed to identify low-VAF CH events. This is described in Biezuner et al., “An improved molecular inversion probe based targeted sequencing approach for low variant allele frequency,” NAR Genom Bioinform 4 (2022), which is incorporated herein by reference in its entirety. The genotyping panel enables the simultaneous detection of over 2000 common genetic variants, all of which are extensively covered across all cell types in the data. This includes heterozygous sites with at least 5% minor allele frequencies from the 1K Genome Project, highly covered by RNA molecules in the data (at least 80 UMIs across all cells in the 10× test library), excluding repeat elements and sites within sex chromosomes. Both panels were designed using MIPgen to ensure capture uniformity and specificity.
[0149] CH sequencing of high-RDW samples and controls: To compare trends in CH and high-risk CH mutations in high-RDW cases and normal RDW controls, 602 high-RDW (>15%) individuals (11.5g / dL≦Hg≦15.5g / dL[F], 13g / dL≦Hg≦17g / dL[M], 80fL≦MCV≦96fL, PLT≧100×10⁹ / L, Abs Neut≧1.8×10⁹ / L) who did not show signs of anemia and whose blood cell counts did not meet the MDS criteria were included, as well as 602 normal RDW individuals (11.5g / dL≦Hg≦15.5g / dL[F], 13g / dL≦Hg≦17g / dL[M], 80fL≦MCV≦96fL, PLT≧100×10⁹ / L, Abs Neut≧1.8×10⁹ / L). Deep targeted sequencing was performed on DNA samples of age- and sex-matched controls (Neut ≥ 1.8 × 10⁹ / L). Case-control matching was performed on a total of 18,147 individuals with longitudinal complete blood count data and available DNA, using the R MatchIt package, balanced by age and sex, with method = "nearest" and ratio = 1. All DNA samples were collected in accordance with the Declaration of Helsinki after obtaining written informed consent and received anonymized from the Tel Aviv Sourasky Medical Center (TASMC) Integrative Cancer Prevention Clinic. All protocols were approved by the TASMC Ethics Committee (under IRB Protocol 02-130) in accordance with all applicable ethical rules.
[0150] scRNA-seq processing: FastQ files were processed by running Cellranger with the hg-38 reference genome. Cells were filtered to have at least 20% mitochondrial expression and UMI of less than 500 from unfiltered genes.
[0151] Doublet calling: Several steps were performed to assign cells to their respective organisms and detect doublets. The pipeline consists of the following steps: 1. Demultiplex cells based on SNPs found in scRNAseq data and call doublets. 2. Construct a metacell model using cells from all libraries, including those previously marked as doublets, and identify and remove metacells created as doublets. 3. Identify doublet metacells based on marker gene expression. 4. Construct the final metacell model and mark the metacells as doublets based on the expression markers.
[0152] In the first step, cells were assigned to each individual using Vireo and Souporcell, which identify doublets and cluster cells based on SNPs found in the sequenced RNA molecules. Vireo (preceded by running cellsnp) and Souporcell were run individually for each library. Both methods used SNPs from a genotyping panel covered by at least 20 UMIs in the library (at least 10 from major and minor alleles, respectively, in Souporcell). High agreement was observed in doublet calling between these two methods.
[0153] In the next step, a metacell model was constructed using cells from all libraries. This model included cells that had already been identified as doublets. The model was constructed in metacells (see Lee-Six, H. et al., "Population dynamics of normal human blood inferred from somatic mutations"; Nature 561, 473-478 (2018), the entire article is incorporated herein by reference), with a target metacell size of 200 cells. All metacells in which at least 35% of the cells were already marked as doublets, and all metacells expressing key markers of different cell types, were then marked as doublet metacells. Subsequently, all cells belonging to the doublet metacells were marked as doublets. Then, an additional metacell model (see below) was constructed without cells marked as doublets.
[0154] Correction for Sequencing Platform Bias: Since a small number of the 10× libraries were sequenced with the Ultima Genomics sequencer and most libraries were processed through the standard Illumina pipeline, it was desirable to minimize batch effects associated with these sequencing platform variations. To this end, using libraries sequenced on both platforms, an Illumina-Ultima correction factor per gene was calculated as the mean log2 change in gene expression across the rearranged libraries. Each Ultima-sequencing library was then normalized by downsampling genes with at least 0.28 log2 Ultima overexpression and resampling genes with at least 0.2 Illumina overexpression. Downsampling and resampling were performed independently for each gene across all cells in each Ultima library. Downsampling and resampling thresholds were selected so that the total number of UMIs per cell remained similar. 87 genes with at least a 4-fold change between Ultima and Illumina were excluded from further processing.
[0155] The reference metacell model was computer-processed: The metacell model was constructed using metacell 2 with a target metacell size of 200 cells. Histone, cell cycle, ribosome, sex linkage, and stress response genes (including FOS, JUN) were marked as prohibited genes, as were genes with high technical variability, such as genes with high or inconsistent differences between technical repeats of Illumina sequences and technical repeats of Ultima sequences. These genes were not used to calculate gene-gene similarity but were included in downstream analyses. Metacells were annotated using known markers. Metacells with low CD34 expression, such as mature monocytes, B cells, T cells, NK cells, DCs, and endothelial cells, were excluded from most downstream analyses. UMAP projection of metacell expression vectors for cell type-specific enriched genes was used to visualize the metacell manifold.
[0156] Comparison and Projection of BMs: Three BM datasets were used for comparison purposes. These were a dataset containing CD34-enriched cells from two individual BMs collected by the inventors and processed similarly to PBs (Figure 1A), the Human Cell Atlas (HCA) dataset, and the CD34+ bead-enriched BM dataset from Setty et al., "Characterization of cell fate probabilities in single-cell data with Palantir," Nat Biotechnol 37, 451-460 (2019), the entire contents of which are incorporated here by reference. The HCA dataset was previously processed and annotated into a metacell model. Metacell models for the two collected CD34-enriched BM samples were constructed using metacells in the same manner as previously described, and a third BM metacell model was created from the data by downloading the sequencing data from Setty et al. and processing it by running cellranger. To project both PB and Setty datasets onto the HCA dataset, HCA metacells, as well as Setty and PB metacells, were correlated across genes showing high variance across HCA metacell models. Each Setty metacell was annotated using the mode and gene marker expression of the five most correlated HCA metacells. Metacells from each of these models were projected onto the HCA UMAP using the mean x and y values of its five most correlated HCA metacells. To compare S-phase genes between PB and BM (Figure 4B), S-phase signatures (mean expression of six cell cycle genes: thymidylate synthase (TYMS), H2A.Z variant histone 1 (H2AZ1 / H2AFZ), proliferating cell nuclear antigen (PCNA), minichromosome maintenance complex component 4 (MCM4), helicase, lymphoid-specific (HELLS), and proliferation marker Ki-67 (MKI67)) were calculated for each PB and HCA metacell, and the distribution of these scores across metacells for each cell type was plotted.
[0157] HSC Differentiation Gene Program: To visualize transcriptional dynamics in HSC cells, MEBEMP and CLP metacells were sorted based on their AVP expression. To calculate differential expression (DE) between HSCs and adjacent cell types, the geometric mean of each gene was calculated across HSC, CLP-E, and MPP metacells, and the differences between HSCs and MPPs, and between HSCs and CLP-Es, were selected.
[0158] Differential expression between individuals not explained by the metacell model: Pooled expression profiles of each individual were compared to matched expression profiles based on the distribution of individuals across metacells. The analysis was performed separately for MPP / MEBEMP (BEMP, ERYP, MEBEMP-E / L, GMP-E, and MPP) and CLP (CLP-E / M / L, NKTDP). In each cell type group, each cell was downsampled to have 500 UMIs, the UMIs across all cells in each individual were summed, the sum was normalized to 1, and log2 was calculated to obtain the observed expression. To calculate matched expression, each metacell was downsampled to have 90K UMIs, and the UMIs across all metacells to which each cell belonged were summed for each individual. This matched expression was normalized to a sum of 1, and log2 was calculated. All genes expressed in any individual (log2 expression level > 2^-14.5) with either the observed or matched expression, showing at least a twofold change between the observed and matched expression in at least one individual, were plotted. Genes showing strong batch effects were excluded.
[0159] HSPC Composition Analysis: To investigate the variance of cell type composition between individuals, the cell distribution of each individual across CD34+ cell states was first calculated. Furthermore, cells from the CD34+ state were distributed into finer bins, totaling 15 bins, using one HSC bin, four CLP bins, and ten MEBEMP / MPP bins. Based on the AVP expression gradient, HSC cells were assigned to bin 0, CLP-E cells to CLP bin 1, and CLP-M / L cells to CLP bins 2-4, so that each of these bins contained an equal number of cells. Similarly, MPP and MEBEMP-E / L cells were assigned to equally sized MPP / MEBEMP bins 1-10 based on the decrease in AVP expression.
[0160] The lower panel of Figure 2D shows the individual enrichment across the bins (log2 of the ratio of the cell frequency of each individual in each individual to the central cell frequency in that bin across the individual). Individuals were divided into three groups based on the average enrichment across CLP bins 2–4 – those with an average enrichment > 0.5 were high CLP, those <-0.5 were low, and the rest were intermediate. Next, a stem cell score was defined as the ratio of the number of cells in MPP / MEBEMP bins 1–5 to the total number of MPP / MEBEMP cells (cells in bins 1–10). Individuals with a stem cell score > 0.5 were enriched in stem cell properties. Individuals within each cluster were further sorted based on their stem cell properties scores. The combination of CLP enrichment and stem cell properties defines the six classes shown in the figure.
[0161] Testing the relationship between cell state compositions and numerical labels: Using a sorting test, we tested the relationship between the cell state distribution and labels such as the CBC index or synchronization score. The inventors sorted CD34+ cell states into 11 bins, from late MEBEMP differentiation through HSC to late CLP differentiation (as shown in Figure 2B). We examined triplets of adjacent cell types in this order and calculated the sum of the individual cell state frequencies for each triplet, obtaining nine vectors of length 150. Each of these nine vectors was then correlated with a label vector, and the largest absolute correlation value was taken as the test statistic. After sorting the labels 10,000 times, we repeated this process and derived p-values using the test statistics from the sorting.
[0162] Variableally expressed gene modules: Gene modules with high variability were detected between individuals while controlling for variations in composition. This was done separately for bone marrow and lymphoid states in the following manner:
[0163] A) For each individual, the fifth percentile of the UMI number was calculated across all MPP metacell cells, and all cells were downsampled to this number. Then, all downsampled cells were pooled, normalized to sum to 1, and log2 was calculated. This allowed observation of the expression profile for each individual.
[0164] B) Next, the predicted expression profiles for each individual were created as follows: All MPP metacells were distributed into 30 equally sized bins based on their AVP expression, and the metacells were downsampled to 90K UMI. The average expression of each gene across the downsampled metacells in each bin was calculated. This defined the expression profile for each of the 30 bins. To obtain the predicted expression for each individual, the weighted average expression profile of the bins was calculated. Here, the weight of each bin was proportional to the proportion of individual cells from that bin, normalized to a sum of 1, and log2 was calculated. Then, the difference between the observed expression profile and the predicted expression profile was calculated.
[0165] C) The data showed several batch effects that distinguished samples collected during two calendar periods. Since this effect may introduce inter-gene covariance between individuals, a correction was applied to control it. This was done using a linear model that fitted each gene to the sample collection period. The inferred period factor was then subtracted from the samples collected during the second period. This approach was found to significantly reduce the appearance of gene clusters associated with sample collection date bias.
[0166] D) Genes with high variance that are unlikely to be residually affected by the main manifold differentiation process were screened. Genes with high batch effects (Kruskal-Wallis p-value < 1e-3 when using a 10 × batch of individuals as a covariate), genes with high AVP correlation (absolute Pearson correlation > 0.65), and genes highly correlated with the module of genes differentially expressed between the first and second collection periods (absolute Pearson correlation > 0.5) were removed. The variance of each gene was then calculated as the difference between observed and expected expression between individuals. Since some of this variance can be explained by sampling noise, the variance of each gene was plotted between individuals against the mean expression between individuals. The genes were sorted by this expression value, and the rolling mean of the variances of 100 neighboring genes in its ordering was subtracted from the variance of each gene. Genes with a variance at least 0.08 higher than the rolling mean variance were selected.
[0167] E) For highly variance genes, gene-gene Spearman correlation matrices were calculated, and correlation profiles were clustered using hierarchical clustering. Genes with low mean correlations to clusters (<0.2) were removed, and then gene clusters with low mean correlations between those genes were removed (mean correlation <0.25 for all gene pairs). Gene-gene correlations were further calculated using only samples from the initial library collection period, and gene clusters had to have high mean correlations between those genes (>0.25) when using only these samples. Additional gene modules resulting from this analysis were removed due to batch effects or trace amounts of MEBEMP differentiation not normalized by this approach. This yielded Figure 3F.
[0168] A similar analysis was performed for CLP (Figure 11), but there were few differences. The analysis included cells derived from CLP-M metacells. The cells were distributed into six equally sized bins, and the distribution was based on their mean DNTT and VPREB1 expressions. Genes with a high absolute correlation to mean DNTT and VPREB1 expressions were excluded. This was followed by hierarchical clustering of gene-gene correlation profiles and gene removal as described for MEBEMP.
[0169] Age Regression: Age regression models were developed separately for MEBEMP and CLP expression. To predict age, the difference between observed and predicted gene expression in an individual was used as described above. Genes with minimum expression of ≥2^-14.5 for MEBEMP and ≥2^-15.5 for CLP across the entire individual were used. The LASSO model was trained using nested single-out cross-validation. Sample cross-validation was performed on the remaining samples to select the LASSO λ parameter, the model was trained using the selected λ, and predictions were made for the excluded samples.
[0170] LMNA Signature: The difference between observed and expected gene expression in an individual was used, and this difference correlated separately with ΔLMNA for MEBEMP and CLP. The MEBEMP and CLP correlation values were then summed, and genes with a summed correlation >0.7 were retained. Furthermore, genes with high technical variance were removed, and 17 genes were retained in the LMNA signature. To calculate the individual LMNA signature, the mean values of these 17 genes in the observed-expected matrix for each individual for MEBEMP and CLP were separately selected. To plot Figure 12A, the geometric mean of LMNA signature gene expression was calculated for each individual in each of the 10 MPP / MEBEMP bins described earlier (Figure 2D). GoT Analysis: GoT analysis was performed on sample N122. 20 This allowed us to mark the cells of this individual as either wild-type or mutant. Due to the low VAF of the DNMT3A mutation in N122, and to improve detection power, cells whose DNMT3A mutation status could not be determined by GoT were marked as wild-type cells. Figure 2G shows the distribution of N122 cells across cellular states.
[0171] Synchronization score: The AVP signature was defined as containing genes with a high correlation (>0.6) to AVP across HSC, MPP, and MEBEMP metacells, and the GATA1 signature was defined as containing genes with a high correlation (>0.7) to GATA1. To eliminate the possibility of a few genes dominating the signature, genes with an average relative expression greater than 2^-10 were filtered in these metacells. All HSC, MPP, MEBEMP-E, and MEBEMP-L cells were then scored by their UMI ratio from the AVP and GATA1 signatures, and all cells were allocated to 20 equal-sized bins for AVP signature expression and 20 equal-sized bins for GATA1 signature expression. The synchronization score was then defined as the percentage of cells with an AVP bin of 9 or higher (top 3 quintiles for AVP expression) and a GATA1 bin of 13 or higher (top 2 quintiles for GATA1 expression).
[0172] To visualize the synchronization score (Figure 3I), this 20x20 bin matrix was normalized to a sum of 1, the resulting matrix was smoothed by averaging the cells using a moving window of length 3, and log2 was calculated.
[0173] Differential gene expression for age and CBC: Differential expression was performed separately for MPP / MEBEMP and CLP cells, as well as for males and females. MPP and CLP-M matrices previously used to detect mutant gene modules were also included. Individual gene expressions were correlated with age, maximum VAF of CH mutations, and 20 CBC index using Spearman correlation, and the correlations were tested for significance. p-values were FDR corrected (Benjamini-Hochberg) separately for each label. For maximum VAF, the Mann-Whitney test was additionally performed to compare individuals with and without the detected mutation. Differential expression between males and females was performed using the Mann-Whitney test on the same expression matrix.
[0174] Patient scRNA-seq initial processing: 10× libraries containing all patients were multiplexed with additional healthy samples. These were processed using cellranger as previously described. Doublets were detected using Vireo and Souporcell, and cells were assigned to individuals as described above. All patient data were sequenced on the Ultima platform and corrected by downsampling and resampling of UMI as described above. Metacell models were then created separately for each of the 12 samples: 2 healthy individuals, 2 MDS patients (one of whom was a del5q patient sampled twice before and after treatment initiation), 3 CMML patients, 1 MDS / MPN duplication patient, 1 myelofibrosis patient, and 2 AML patients. As previously described, cells with <500 UMI, >20% mitochondrial gene expression, or high megakaryocyte gene expression were excluded from these models. The same set of ignored genes previously used in the healthy model were employed, and the target number of cells per metacell was set so that each metacell had approximately 300K UMI.
[0175] Projection of disease data onto the HSPC model: To project patient metacells onto a healthy reference, patient (query) metacells were correlated with reference metacells. Due to sequencing depth variability, query metacells were initially downsampled to 150K UMI per metacell. Correlation was performed on a log2 scale using variable genes from the reference. Next, query metacells were annotated using the five most correlated modes (most common cellular states) of reference metacells. Query metacells mapped to CD34-negative reference metacells were discarded from downstream analysis. Figure 4B shows the mean correlation between each query metacell and its five most correlated reference metacells, placing each query metacell at the mean x and y coordinates of these metacells on the reference UMAP. Figures 4C–4D are based on single-cell projections. Patient query cells were correlated with reference metacells using raw UMI vectors. Figure 4C then shows the distribution of projected metacell annotations (determined by the five most correlated modes of reference metacells for each query cell). To create Figure 4D, each query cell was assigned to the most common bin (as in Figure 2D) within the metacell it was mapped to. To create Figure 4E, the observed expression for each patient was calculated as the geometric mean expression across all metacells annotated as either MPP / MEBEMP-E / MEBEMP-L. To calculate the expected expression, the geometric mean expressions of the five most correlated reference metacells for each query metacell were selected. Due to periodic systematic differences between query and control samples, genes were sorted based on their mean expression in observed and expected values, and expected values were normalized by subtracting the rolling average (moving average) of the difference between expected and observed values across 500 consecutive genes from the expected values. Differential expression (DE) genes were defined as those with a difference of more than 2x between observed and expected values.
[0176] Karyotype Analysis: To perform karyotype analysis, the geometric mean of the five most correlated reference metacells was subtracted from the expression of each query metacell (normalized to sum to 1, with log2 calculated) (expression difference). Each chromosome was divided into equal-sized bins, each containing 40 or more genes, and the median expression difference was calculated across all genes in each combination. For this analysis, only genes with a mean expression of at least 2^-15.5 in either the query or matching reference metacell were considered. This analysis provides karyotype at metacell resolution, as shown in Figure 4F. Signature profiling in disease cases: To create Figure 13, AML-1 (patient N186) metacells were separated into AML-1-1 and AML-1-2 based on their karyotype. Lists of differentially (over)expressed genes were created in healthy HSC, MEBEMP, NKTDP, and CLP metacells. Each AML metacell was scored by the geometric mean of all genes in each of these cell state-specific gene lists, as well as the LMNA signature genes. State-specific expression thresholds were then established by observing the expression of each state-specific gene program across all reference metacells belonging to the relevant cell state (e.g., NKTDP metacell for the NKTDP gene program, see dashed line in Figure 13A). To discover de novo gene programs in AML samples (Figure 13B), highly variable genes were selected from each AML metacell model, their correlations were calculated across metacells, their correlation profiles were clustered, and clusters with low mean correlations (<0.4) and genes with low mean correlations to those clusters (<0.4) were filtered. Several target genes were manually added to the displayed correlation matrix (Figure 13B). For the heatmap in Figure 13C, state-specific genes were selected from the gene list described above, as well as genes that were higher in AML-1 / AML-2 compared to the reference across all cell states, and genes that were higher in AML-1-2 compared to AML-1-1 in MPP, MEBEMP-E, and CLP-E metacells.
[0177] Example 1: Universal stem cell and precursor states observed across humans in CD34+ peripheral blood To assess inter-individual diversity in subtype distribution and regulation of circulating HSPCs (cHSPCs) from healthy individuals, multiplexed scRNA-seq was combined with genotyping and integrated clinical data. Multiplexing was resolved using SNPs identified in the 3'UTR of cHSPC RNA, thereby facilitating accurate cell matching to individuals and improving control against batch effects and doublets (Figure 1A). Overall, HSPCs were collected from 79 males and 69 females aged 23–91 years (median 61.5). Technical replication was performed on 39 individuals, and biological replication was performed on a follow-up cohort of 20 individuals (1 year after the original sampling date). To identify cases of CH, longitudinal CBCs were collected up to 5 years prior to scRNA-seq, and deep targeted somatic mutation analysis was performed on DNA generated from their blood at the time of sampling. Following quality control and filtering, 846,762 single-cell profiles were retained, normalized to a control for sequencing platform batch effects, and combined to construct and annotate a metacell manifold model. 672,000 CD34+ single cells were retained for downstream analysis. These formed a rich repertoire of states related to cHSCs and their differentiation trajectories (Figure 1B). The derived model replicated and deepened previous efforts to characterize HSPC states from bone marrow (BM), and while it did not fully reflect BM dynamics, it fit with previously generated scRNA-seq data, suggesting that cHSPCs may function as a highly accessible substitute for hematopoietic dynamics both intraorganically and interorganically. However, one notable feature specific to cHSPCs was the repression of cell cycle gene expression. Importantly, the cHSPC model was found to be consistent across individuals. The median number of individuals contributing cells to each metacell was 84, and all metacells contained cells from at least 47 individuals. After controlling the cell distribution of each sample across the atlas state, individual-specific differential expression was restricted.
[0178] Example 2: High-resolution circulating HSC map shows HLF, GATA3, HOXB5, and TLE4 as different HSC TFs. One of the notable features of this cHSPC model is the distinct HSC states transcriptionally linked to two major differentiation gradients: the first represents a continuum of the common lymphocyte precursor (CLP) program; the second, more general branch represents the multipotency precursor (MPP) state and their differentiation into granulocyte-monocyte precursor (GMP), erythrocyte precursor (ERYP), and basophil / eosinophil / mast cell precursor (BEMP). Technical limitations of cell dissociation in scRNA-seq hindered accurate megakaryocyte program modeling. Therefore, since megakaryocyte / erythrocyte / basophil / eosinophil / mast cell precursor (MEBEMP) is also presumed to be the cell of origin of megakaryocytes, the states underlying this trajectory were annotated as these.
[0179] Early HSCs are characterized by high AVP and HLF expression and have been previously shown to represent a rare cell population with self-renewal capabilities in BM and umbilical cord blood. This model, which includes data on approximately 14,440 HLF / AVP HSCs that can be fitted with cells from an independent BM atlas, suggests that HSCs with potential self-renewal capabilities exist in peripheral blood under steady state. Along with HLF and AVP, 14 genes were identified that were expressed at least 1.75 times higher in HSCs compared to their two immediate differentiation branches. Several transcription factors (TFs) enriched in HSCs, including the genes HOXB5, TLE4, and GATA3, were specifically identified (Figure 1C). While HSC status is defined by intrinsic markers that are symmetrically downregulated at the end of the CLP and MEBEMP trajectories (Figure 1C), it should be noted that several lineage-specific regulators at intermediate levels, which asymmetrically branch off from the HSC status at the end of the CLP and MEBEMP trajectories, are also expressed (Figure 1D). This may suggest that the multipotency of HSCs is related to the intermediate expression of multiple regulatory factors that dissipate during differentiation.
[0180] Example 3: NK-T-dendritic cell precursors and basophil-eosinophil-mast precursors are enriched in circulating HSPCs. The cHSPC atlas was enriched for basophil-eosinophil-mast cell precursors (BEMPs) and mapped as one possible terminal of HSC differentiation. Classical studies associated these cells with granulocyte / monocyte precursor (GMP) origin, but more recent studies have suggested that they arise at least partially from mouse and human erythrocyte precursors. This analysis allowed us to focus on a small population of metacells that link BEMPs to their MEBEMP-L precursors (Figure 1E). This highlighted that TF (Figure 1F) and other factors (Figure 9) are positively or negatively regulated in this hypothetical early-stage BEMP specificization. Another rare HSPC population that we can now focus on involves a lymphoid state with high ACY3 expression and moderate to low DNTT levels, a combination rarely seen in human BM but present in peripheral blood. Interestingly, covariances of key T cell regulators were observed within this population, but inverse correlations of these factors with several features of the dendritic cell (DC) program were also observed. This can be demonstrated by the comparison of TCF7 and IRF8 expression (Figure 1G), as well as the consistent TCF7-bound dynamics of CD7, MAF, and IL7R, or the IRF8-bound dynamics of myeloid TF SPI1 (PU.1) and multiple MHC-II genes (Figure 1H). Therefore, this subpopulation was referred to as NK / T / DC precursor (NKTDP). In summary, the mapping of circulating HSPCs revealed a rich spectral differentiation trajectory and precursor state that refined previous analyses, as well as remarkable state universality, providing an opportunity to decipher inter-individual hematopoietic variability.
[0181] Example 4: Inter-individual variability of cHSPC stem cell type and lymphoid / myelin differentiation bias To study inter-individual cHSPC variability, we first examined individual-specific cell state compositions. This was done by quantifying the relative frequency of cell states within single-cell aggregates in each individual (Figure 2A). These frequencies varied widely (Figure 2B). For example, HSC and CLP-Ms, representing 2.4% and 12.6% of the CD34+ population on average, showed standard deviations of 1.0% and 6.8%, respectively. Abundant MPP and MEBEMP-E states (average frequencies of 20.7% and 37.6%, respectively) showed smaller relative variability (SD of 4.9% and 5.8%, respectively). To analyze cell state frequencies over time and the stability of sampling cases, 20 individuals were resampled one year after the original sampling date. Both lymphocyte precursor frequencies (CLP-M, CLP-L, NKTDP) and MEBEMP frequencies (MEBEMP-E, MEBEMP-L, ERYP, BEMP) were stable over time within the same individual (Figure 2C).
[0182] To analyze the composition with higher resolution, the enrichment level of each individual was profiled across MEBEMP and CLP trajectories. Clustering of these enrichment profiles resulted in six archetypes (classes I-VI) of cHSPC composition within the healthy population (Figure 2D). These consisted of individuals with relative lymphoid enrichment (classes I-III) or depletion (classes V-VI), and were further subdivided by stem cell gradients, with classes II, IV, and VI being enriched and classes I, III, and V being depleted. Analysis of technical and biological replicates confirmed that this variability is robust and individual-specific. In summary, this analysis provides the first cHSPC subpopulation normal reference range (Figure 2B) characterized by widespread variability among healthy individuals, demonstrating that these compositional differences are true individual characteristics with potential clinical significance.
[0183] Example 5: Circulating HSPC frequency correlates with CBC and CH. Analysis of CBC correlations with this single-cell atlas reinforced previous findings regarding inter-individual variability in cHSPC composition. All CBC correlation analyses were performed using the median values of each complete blood count parameter over a five-year period prior to scRNA-seq. The mean and median number of complete blood counts per individual over this five-year period were 8 and 6, respectively. A significant positive correlation (P<0.01) was observed between PB mature lymphocyte percentage and CLP frequency (Figure 2E, left). Considering the very high variability of female erythrocyte (RBC) count and size during young adulthood (with the effects of menstruation and pregnancy) and the long-term perimenopause period, RBC indices, including RBC, hematocrit (HCT), mean corpuscular volume (MCV), and RDW, were analyzed separately for men and women. A significant negative correlation (P<0.02) was observed between CLP frequency and HCT (men, Figure 2E-center). A similarly significant positive correlation (P<0.01) was observed between increased RDW—cHSPC bone marrow bias—and relative CLP depletion (men, Figure 2E-right).
[0184] Previous studies have correlated increased RDW with a high risk of CH and predisposition to acute myeloid leukemia (AML). A low CLP frequency has been demonstrated to be associated with CH (two-sided Mann-Whitney test, Figure 2F), and we further reinforced this observation by performing transcriptome genotyping in one of DNMT3A R882 cases and identifying a lower fraction of CLP cells in the mutant clone (P<0.005, Fisher's exact test, Figure 2G). To further investigate this association, we studied a cohort of 18,147 healthy individuals for whom both longitudinal CBC and DNA were available. We identified 602 individuals with high RDW (≥15%, not meeting the minimum criteria for myelodysplastic syndrome (MDS) diagnosis) and 602 age- and sex-matched normal RDW controls. Deep targeted sequencing was performed to identify preleukemic mutations (pLMs) in both high-RDW individuals and controls, and a significant enrichment of CH+ cases was found in the high-RDW group (Fisher's exact test p-value < 0.002, Figure 2H). Overall, the data demonstrate a three-way association between reduced CLP frequency, high RDW, and CH.
[0185] Example 6: Age-related bone marrow bias is observed mainly in males. Blood aging is a complex, multifactorial process likely driven by endogenous factors such as pre-leukemic mutations, as well as exogenous effects such as cytokine and hormonal changes. To isolate these factors as much as possible, we studied age-related changes in the cHSPC population in individuals without CH mutations. Analysis of age-related compositional changes in cHSPCs within this group showed a significant increase in the bone marrow (MEBEMP) to lymphocyte (CLP) ratio in men (comparing individuals <50 to >60 years old, Figure 3A). This effect was not significant in women. In this regard, it is important to note that while a decrease in lymphocyte count can be observed in both older men and women, it appears later in women. Interestingly, women show a surge in lymphocyte count immediately after menopause, contributing to this delay in lymphopenia. Within the MEBEMP differentiation trajectory, aging correlated with over-presentation of more differentiated states, again in men (Figure 3B). Notably, the frequency of cHSCs did not change significantly with age (Figure 3C). Previous studies have suggested that aging is associated with an increase in HSC frequency; however, such an increase was not observed when determined using the limited definition used herein, as well as from CD34+PB HSPC frequency in a recent cohort of 1000 healthy individuals who underwent PBMC scRNA-seq (Figure 3D). The sex-specific correlation between age and cHSPC myeloid bias highlights the role of such non-endogenous effects in this classic feature of hematological aging.
[0186] Example 7: Composition-controlled HSPC expression correlates with age. As described above, the individual cHSPC composition provides an initial blueprint of hematopoietic dynamics along the stem cell axis and the CLP / MEBEMP axis. Here, further analysis of transcriptional variations can be performed while controlling for the dominant effect of the cHSPC composition in order to characterize further gene expression signatures that can distinguish individuals. Individual expression profiles controlled by composition showed high informativeness when correlated with age, enabling age prediction based only on normalized expression (Figure 3E, 10). Next, we searched for gene sets (signatures) that co-varied between individuals, excluding sex-linked signatures and those showing strong batch effects. The most prominent of these signatures included, among others, lamin-A (LMNA), as well as annexin A1 (ANXA1), AHNAK nucleoprotein (AHNAK), myeloid-associated differentiation marker (MYADM), tetraspanin 2 (TSPAN2), and vimentin (VIM) (Figure 3F, 11). Individual LMNA signature expression varied over a range of more than twofold (Figure 12A), showing high variability in HSCs and early myeloid and lymphoid cell states, as well as uniformly low expression in late MEBEMP and CLP. Individual LMNA signature expression was consistent across myeloid and lymphoid cell states (Figure 3G) and stable in the follow-up cohort (Figure 12B). Interestingly, an age-related increase in LMNA signature expression was observed in lymphoid cHSPCs but not in myeloid cHSPCs (Figure 3H). In summary, this suggests that, in addition to pLM accumulation in HSPCs, aging is strongly associated with changes in the distribution of progenitor cell states in PBs, and that there are significant differences in the expression of specific gene signatures.
[0187] Example 8: Rapid suppression of stem cell signatures in MEBEMP is associated with lower red blood cell counts and higher red blood cell volume. The differentiation of HSPCs into MEBEMP and CLP fate involves the coordinated activation of specific transcriptional programs, which were generally universal across individuals. However, screening of individual-specific gene signatures suggested that individuals differ in how they synchronize the opposing effects of these stem cell and differentiation programs. To quantify this variability, AVP (stem cell) signatures and GATA1 (MEBEMP differentiation) signatures were compared in 20×20 bin expression matrices (Figure 3I). Most individuals exhibited near-diagonal dynamics (e.g., individuals N16 and N86), following the typical transition from stem cell to differentiation, but some individuals deviated from the diagonal, showing biased synchronization between the AVP and GATA1 signatures. This deviation (i.e., frequency of off-diagonal) was quantified using a novel synchronization score. This facilitated the identification of individuals with low synchronization scores of 0.12 (e.g., N122 and N172, Figure 3I, top panel), demonstrating delayed GATA1 activation compared to AVP suppression. Specifically, these individuals exhibited a rapid decrease in AVP expression, but a delayed increase in GATA1 and GATA1-related genes. In contrast, individuals with high synchronization scores (e.g., N98 and N121, e.g., Figure 3I, bottom panel) showed early activation of GATA1 expression preceding AVP suppression. The stability of synchronization scores was detected in a follow-up cohort (Figure 12C). Inter-individual synchronization score variability was positively correlated with RBC levels and consistently negatively correlated with MCV in males (P<0.01 for both RBC and MCV (Spearman); Figure 3J). Analysis of correlations between individual synchronization scores and cHSPC composition in males demonstrated negative correlations with ERYP and BEMP (Figure 3K). In summary, variations in the regulation of stem cell characteristics and MEBEMP differentiation programs correlated with red blood cell count and volume were demonstrated.
[0188] Example 9: Aging perturbations of HSPC compositions and transcription signatures Blood aging represents a complex, multifactorial process likely caused by endogenous hematopoietic effects (e.g., pre-malignant mutations) and exogenous physiological effects (e.g., hormonal changes). Therefore, we anticipated several characteristics defining multi-layered aging HSPC correlations. We first examined the association between HSPC composition and age and observed no clear directional increase or decrease in HSPC subtypes with aging. We demonstrated an increase in variance of cellular state frequencies with significantly higher variance after age 65 (p<0.01). To quantify the deviation of each individual from expected cellular state frequencies, we calculated an HSPC composition bias score, which significantly increased with age (Figure 4A, p<0.02, Spearman's rho test). This supported the concept of multiple aging processes disrupting the highly homogeneous and robust HSPC landscape observed in younger adults.
[0189] The inventors further investigated inter-individual variability in age-related hematopoiesis using several HSPC signatures, including the LMNA signature and synchronization signature described above, as well as an S-phase signature that quantifies the expression of S-phase-related cell cycle genes, which have previously been shown to have high inter-individual compositional normalization gene expression correlations (Figure 3B). The S-phase signature was robust in the follow-up cohort, supporting its role in characterizing individual quality rather than being a transient effect. Circulating HSPCs generally did not express S-phase transcriptional signatures, in contrast to their bone marrow counterparts (Figure 4B). However, in some individuals, weak but significant expression of DNA replication genes was observed in the late MEBEMP trajectory, strongly positively associated with age (Figure 4C, p<0.04, Spearman's rho test). Comparison of S-phase signatures to HSPC compositional bias scores suggested that the two increased independently of each other with age (Figure 4D). In contrast, increased HSPC bias scores could be associated with lower LMNA signatures (Figure 4E), reinforcing the association between CH and low LMNA expression. Despite the association with RBCs and MCVs as described above, the synchronization score did not directly correlate with age (Figure 4F).
[0190] Case studies of individuals with highly abnormal HSPC distributions, and their integration with clinical markers and mutation profiling, illustrate the multi-mode nature of hematopoietic aging. Individual #151, an 80-year-old male diagnosed with MDS, defined by the TET2 / DNMT3A / CBL clone, exhibited an extreme HSPC bias, a low LMNA signature, and a high S-phase signature, with a high mutation allele frequency (VAF; TET2 VAF=70%) and high RDW anemia (Figure 4G). Individual #98, a 69-year-old male, showed another distinct behavior with polycythemia, a high synchronous signature, and high RDW. In summary, analysis of HSPC composition and transcriptional signatures provided insights into various mechanisms driving hematopoietic aging. In particular, our analysis separates the spectrum of CH-related effects from those related to changes in HSPC regulation and differentiation. High-resolution characterization of these effects enables analysis of hematological malignancies at high molecular depth.
[0191] Example 10: Use of the cHSPC Atlas for Mapping, Anatomy, and Annotation of Myeloid Malignancies The diagnosis of myeloid malignancies requires the identification of clonal markers (mutations or structural variants), as well as the detection and quantification of blast cells by microscopy and flow cytometry. Figure 5A illustrates a novel stepwise approach to the analysis of myeloid diseases based on the sampling of cHSPCs and their composition, normalized expression, and copy number variation (karyotyping) compared to a new normal reference. As proof of concept, cells sampled from two healthy individuals, three patients with chronic myelomonocytic leukemia (CMML), two patients with MDS, one patient with MDS / myeloproliferative disorder (MPN), and one patient with myelofibrosis (MF) were analyzed. Furthermore, two AML cases were sampled and analyzed to demonstrate how acute disease manifests when projected onto the cHSPC model. The projection of patient metacells showed high gene expression correlations between metacell pairs for all pathological cases except AML (Figure 5B, color-coded dots). However, the composition of cellular status was biased in all pathological cases (Figures 5C-5D), but not in healthy controls. All patient samples showed a significant decrease in the CLP population (Figures 5C-5D). Two of the CMML patients (N192, N235) showed highly abnormal enrichment of specific (basophil and bone marrow) cell states, while the rest showed a relatively balanced distribution across the MEBEMP differentiation spectrum (with slight enrichment of stem cell states). Compositional regulatory gene expression comparisons between patient samples and normal models identified specific genes that are repeatedly induced or repressed in the disease (Figure 5E). Both healthy individuals and treated MDS cases showed minimal DE genes, but all leukemia cases showed a substantial increase in the number of DE genes compared to the healthy population model (Figure 5E).
[0192] The detection of karyotype abnormalities based on gene expression levels, previously suggested and implemented in several tools, can be readily applied to cHSPCs, as shown in Figure 5F, and clear chromosomal abnormalities are observed in two AML cases analyzed. Furthermore, a deletion of the long arm of chromosome 5, del(5q), was detected in one of the MDS samples that showed complete remission of cytoplasticity after lenalidomide treatment (N211A & N211B, Figure 5F). In conclusion, the atlas of normal cHSPC states presented herein enables high-resolution characterization of myeloid disorders based on their cellular state, compositionally controlled transcriptional dispersion, and abnormal karyotype.
[0193] While the present invention has been described in conjunction with its specific embodiments, it will be apparent that many alternative, modified, and variant forms are obvious to those skilled in the art. Therefore, it is intended to encompass all such alternative, modified, and variant forms that fall within the spirit and broad scope of the appended claims.
Claims
1. A non-invasive method for detecting bone marrow pathology in a subject requiring such detection, wherein the method is a. To receive a target cell dataset based on single-cell RNA sequencing (scRNA-seq) of CD34-positive cells from the peripheral blood of the subject, b. Analyzing the received target cell dataset in relation to a control dataset containing multiple cell datasets, wherein each cell dataset in the multiple cell datasets is based on the scRNA-seq of CD34-positive cells from the peripheral blood of a healthy subject, and the deviation of the target cell dataset from the control dataset indicates a pathological condition of the bone marrow. A method for detecting pathological conditions in the bone marrow.
2. The method according to claim 1, wherein the cell dataset includes statistical data on all CD34-positive cells in a peripheral blood sample.
3. The method according to claim 1 or 2, wherein the analysis includes generating a feature vector representing the deviation of the target cell data from the control cell data.
4. The method according to claim 1 or 2, wherein the analysis comprises applying a trained machine learning model to the received dataset, the machine learning model being trained on a training set comprising the plurality of cell datasets, and the machine learning model classifying the bone marrow in question as healthy or not.
5. The method according to claim 4, wherein the training set further comprises a cell dataset based on scRNA-seq of CD34-positive cells derived from the peripheral blood of a subject suffering from a bone marrow condition, and labels indicating that the cell dataset originates from a healthy subject or a subject suffering from a bone marrow condition, and the machine learning model classifies the subject as healthy or suffering from the bone marrow condition.
6. The method according to claim 3, wherein the analysis comprises applying a trained machine learning model to the feature vectors, the machine learning model being trained on a training set comprising feature vectors from healthy subjects and subjects suffering from the bone marrow disease, and labels indicating that the feature vectors are from healthy subjects or subjects suffering from the bone marrow disease, and the machine learning model classifying the subjects as healthy or suffering from the bone marrow disease.
7. The method according to claim 1 or 2, wherein the analysis comprises applying a trained machine learning model to parameters extracted from the cell dataset, the machine learning model being trained on a training set comprising parameters extracted from cell datasets of healthy subjects and, optionally, subjects suffering from myelopathy, and the machine learning model classifying the subjects as healthy or not.
8. The method according to any one of claims 1 to 7, wherein the cell dataset is selected from a metacell model of all CD34-positive cells in a peripheral blood sample, the transcriptome of each of the CD34-positive cells in the peripheral blood sample, and an annotated cell atlas of CD34-positive cell types present in the peripheral blood sample.
9. The method according to any one of claims 1 to 8, wherein the aforementioned pathological condition of the bone marrow is selected from myelodysplastic syndrome (MDS), chronic myelomonocytic leukemia (CMML), acute myeloid leukemia (AML), polycythemia vera (PV), essential thrombocythemia (ET), mastocytosis, chronic eosinophilic leukemia, myelofibrosis (MF), acute lymphoblastic leukemia (ALL), ambiguous lineage acute leukemia, multiple myeloma (MM), myeloproliferative disorder (MPN), and blastic plasmacytoid dendritic cell leukemia.
10. The method according to claim 9, wherein the method is a method for detecting MDS, and deviations in the frequency of erythrocyte progenitor cells (ERYP), basophil / eosinophil / mast cell progenitor cells (BEMP) and / or megakaryocyte / erythrocyte / basophil / eosinophil / mast cell progenitor cells (MEBEMP) indicate the presence of MDS.
11. The method according to claim 9, wherein the method is a method for detecting CMML, and deviations in the frequency of early granulocyte-monocyte progenitor cells (GMP-E) indicate the presence of CMML.
12. The method according to claim 9, wherein the method is a method for detecting AML, and deviations in the frequency of common lymphocyte progenitor cells (CLPs) and / or natural killer / T / dendritic cell progenitor cells (NKTDPs) indicate the presence of AML.
13. The method according to any one of claims 10 to 12, wherein the deviation is at a higher or lower level of cell type than that present in the healthy subject.
14. The method according to claim 10 or 13, wherein a deviation in the frequency of CLP also indicates MDS, and the deviation is at a level of CLP lower than that present in the healthy subject.
15. The method according to claim 10 or 14, wherein the deviation in CLP frequency also indicates CMML, MF, or MPN, and the deviation is at a lower level of CLP than is present in the healthy subject.
16. The method according to claim 9, wherein the method is a method for detecting MDS, and a decrease in the frequency of CLP, NKTDP, or both, compared to a healthy subject, indicates MDS.
17. The method according to claim 16, wherein a reduction in the frequency of both CLP and NKTDP compared to a healthy subject indicates MDS.
18. The method according to any one of claims 1 to 8, wherein the pathological condition of the bone marrow includes an increased percentage of blasts, the deviation is increased, and the deviation in the frequency of early common lymphocyte progenitor cells (CLP-E) indicates the presence of an increased percentage of blasts.
19. The method according to any one of claims 1 to 18, further comprising administering at least one therapeutic agent to a subject determined to be suffering from a bone marrow disorder.
20. A non-invasive method for predicting the percentage of blast cells in the bone marrow of a subject requiring such a method, the method comprising receiving a measurement of the CLP-E cells in the peripheral blood of the subject, wherein the measurement is proportional to the percentage of blast cells in the bone marrow of the subject, thereby predicting the percentage of blast cells in the bone marrow of the subject.
21. The method according to claim 20, further comprising analyzing the received measurements in relation to a control dataset containing multiple measurements of CLP-E cells in the peripheral blood of healthy subjects and subjects suffering from bone marrow pathologies, wherein the percentage of blast cells in the bone marrow is known for each subject in the control dataset.
22. A non-invasive method for predicting the percentage of blast cells in the bone marrow of a subject requiring such a method, wherein the method is a. To receive a target cell dataset based on single-cell RNA sequencing (scRNA-seq) of CD34-positive cells from the peripheral blood of the subject, b. Applying a trained machine learning model to the received dataset, wherein the machine learning model is trained on a training set comprising multiple cell datasets, and each of the multiple cell datasets is based on the scRNA-seq of CD34-positive cells from the peripheral blood of a control subject and a label indicating the percentage of blast cells in the bone marrow of the control subject that provided each of the multiple cell datasets. The machine learning model includes outputting a predicted percentage of blast cells in the bone marrow of the subject, A method for predicting the percentage of blast cells in the bone marrow of a subject.
23. The method according to any one of claims 20 to 22, wherein the subject is suffering from leukemia.
24. The method according to any one of claims 19 to 23, wherein the control subjects include subjects suffering from leukemia and subjects without leukemia.
25. The method according to any one of claims 20 to 24, wherein the cell dataset is selected from a metacell model of all CD34-positive cells in a peripheral blood sample, the transcriptome of each CD34-positive cell in a peripheral blood sample, and an annotated cell atlas of CD34-positive cell types present in a peripheral blood sample.
26. The aforementioned cell dataset is a metacell model, a. Obtaining peripheral blood samples from the subject, b. Isolating CD34-positive hematopoietic stem cells and progenitor cells (HSPCs) from the peripheral blood sample, c. Perform scRNA-seq on the isolated HSPCs to generate transcriptomes for each isolated HSPC, d. To generate a metacell model of the HSPC based on those transcriptomes, The method according to any one of claims 8 to 19 and 22 to 25, which is produced by a method including the method.
27. The method according to claim 26, wherein the metacells are clusters of cells having similar transcriptomes.
28. The method according to any one of claims 1 to 19 and 22 to 27, wherein the cell dataset comprises grouping cells into cell types that share common differentiation within the HSPC differentiation spectrum.
29. The method according to claim 28, wherein the cell type is selected from BEMP, ERYP, MEBEMP-L, MEBEMP-E, GMP-E, pluripotent progenitor cells (MPP), hematopoietic stem cells (HSC), CLP-E, CLP-M, CLP-L, and NKTDP.
30. The method according to any one of claims 20 to 29, wherein the method is a method for detecting MDS and / or leukemia, and the percentage of blast cells exceeding a predetermined threshold indicates that the subject is suffering from MDS and / or leukemia.
31. The method according to claim 30, further comprising administering at least one anticancer treatment to a subject suffering from MDS and / or leukemia.
32. A non-invasive method for calculating a Molecular International Prognostic Scoring System (IPSS-M) risk score for a subject suffering from a myeloma, wherein the method is: a. Predicting the percentage of blast cells in the bone marrow of the subject by the method described in any one of claims 20 to 31, b. To detect the presence of bone marrow mutations and karyotype abnormalities based on scRNA-seq reads from CD34-positive cells derived from the peripheral blood of the subject, c. Obtaining hemoglobin levels and platelet counts in peripheral blood from the subject, d. Calculating the IPSS-M risk score based on the predicted blast percentage, detected mutations and karyotype, and received hemoglobin levels and platelet counts, A method for calculating the IPSS-M risk score.
33. The method according to claim 32, further comprising administering a treatment regimen based on the IPSS-M risk score to the subject, wherein a more potent treatment regimen is administered to the subject with a higher score, and a reduced treatment regimen is administered to the subject with a lower score.
34. A system for evaluating the health of a target bone marrow, wherein the system scRNA sequencing device, A non-temporary memory device in which instruction code modules are stored, The memory device is associated with at least one processor configured to execute the instruction code module, and when the instruction code module is executed, the at least one processor From the aforementioned scRNA sequencing device, a single-cell transcriptome is obtained from CD34-positive cells in the target peripheral blood. Based on the single-cell transcriptome obtained above, a cell dataset is generated. The generated cell datasets are analyzed in relation to a control dataset containing multiple cell datasets, and each cell dataset of the multiple cell datasets is based on the scRNA-seq of CD34-positive cells from the peripheral blood of healthy subjects, and Based on the deviation of the target cell dataset from the control dataset, findings of healthy bone marrow or bone marrow pathology in the subject are output. A system that is configured in such a way.
35. The system according to claim 35, wherein the cell dataset is a metacell model having similar transcriptomes from the acquired single-cell transcriptomes clustered into metacells.