Transformer based artificial intelligence on single cell clinical data

WO2026207139A1PCT designated stage Publication Date: 2026-10-01THE GENERAL HOSPITAL CORP +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/020800
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-25
Filing Date
2026-03-25
Publication Date
2026-10-01

Smart Images

  • Figure US2026020800_01102026_PF_FP_ABST
    Figure US2026020800_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for receiving data representing multiple single-cell distributions, where each of the multiple single-cell distributions represents a corresponding one or more blood cell measurements for one or more corresponding blood cell types of multiple blood cell types, the one or more blood cell measurements collected using a complete blood count (CBC) test performed on a blood sample of a patient. The method includes generating, using a permutation invariant first neural network and the data representing the multiple single-cell distributions, an encoded distribution for each of the multiple single-cell distributions, and predicting, using the respective encoded distributions, a value indicative of a blood cell attribute of a blood cell population of the patient.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0002] TRANSFORMER BASED ARTIFICIAL INTELLIGENCE ON SINGLE CELL CLINICAL DATA

[0003] CLAIM OF PRIORITY

[0004] This application claims the benefit of U.S. Provisional Application Serial No.

[0005] 63 / 777,075. filed on March 25, 2025. The entire contents of the foregoing are incorporated herein by reference.

[0006] FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0007] This invention was made with government support under grant numbers 1922658 and 2145542 from the National Science Foundation; and grant numbers HL148248, DK123330, and HD104756 from the National Institutes of Health. The government has certain rights in the invention.

[0008] BACKGROUND

[0009] The circulating populations of blood cells in humans are regulated to maintain homeostasis in changing environments and to respond to acute and chronic disease processes. Homeostasis is observed at the level of cell population size and composition, where steady states reflect integrated dynamics of production, maturation, activation, clearance, and other cellular-scale regulatory processes. Insights into cellular-scale mechanisms can be inferred from distributions of characteristics of cells that constitute these populations. Examples of these insights from the clinical setting are provided by a few well-established clinical biomarkers obtained from single-cell data that is collected during routine blood counts: the reticulocyte count measures the rate of red blood cell (RBC) production, the “Left Shift'’ flag identifies increased white blood cell (WBC) proliferation, and the immature platelet fraction reflects platelet (PLT) kinetics. Single-cell data is collected during routine complete blood counts (CBCs), which measure single-cell characteristics like volume, protein content, and nuclear morphology for more than -50,000 individual cells and typically output the sizes of the RBC, WBC, and PLT populations.

[0010] SUMMARY

[0011] Disclosed herein are diagnostic and predictive technologies that leverage transformer-based deep learning models to predict blood cell populations and identify, from the predicted populations, biomarkers from single-cell data of routine complete blood counts (CBC). These technologies can include a first transformer-based deep learning modelAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0012] used to predict the blood cell population sizes of particular blood cell populations, and a second transformer-based deep learning model used to interpret the physiological basis for prediction accuracy of the blood cell populations and, as such, assist in identifying biomarkers for connected single-cell features to cross-lineage dynamics, including a singlecell predictor of sepsis, heart failure, and diabetes.

[0013] In general, in a first aspect, a method includes receiving data representing a plurality of single-cell distributions. Each of one or more single-cell distributions represents a corresponding plurality of blood cell measurements for one or more corresponding blood cell types of a plurality of blood cell types. One or more blood cell measurements collected using a complete blood count (CBC) test performed on a blood sample of a patient. The method includes generating, using a permutation invariant first neural network and the data representing one or more single-cell distributions, an encoded distribution for each of one or more single-cell distributions, and predicting, using the respective encoded distributions, a value indicative of a blood cell attribute of a blood cell population of the patient.

[0014] In general, in a second aspect, combinable with the first aspect, the blood cell attribute includes a size of a blood cell population of the patient.

[0015] In general, in a third aspect, combinable with the first aspect, the blood cell attribute includes a blood cell measurement for the one or more corresponding blood cell types not provided as input.

[0016] In general, in a fourth aspect, combinable with the first aspect, the blood cell attribute includes a corresponding plurality of blood cell measurements for one or more corresponding blood cell types, one or more blood cell measurements collected using a second complete blood count (CBC) test performed on a second blood sample of a patient taken after the first CBC.

[0017] In general, in a fifth aspect, combinable with any of the preceding aspects, one or more single-cell distributions include: (i) a first single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of red blood cells, (ii) a second single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of red blood cells and platelets, (iii) a third single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of nuclei of white blood cells, and (iv) a fourth single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of a peroxidase reaction in white blood cells.Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0018] In general, in a sixth aspect, combinable with the fifth aspect, the blood cell population is: (i) a red blood cell population of the patient, (ii) a white blood cell population of the patient, or (iii) a platelet population of the patient.

[0019] In general, in a seventh aspect, combinable with the sixth aspect, the blood cell population is the red blood cell population of the patient, and the value includes a hematocrit of the red blood cell population or a total hemoglobin mass of the red blood cell population.

[0020] In general, in a eight aspect, combinable with any of the preceding aspects, generating the encoded distribution for each of one or more single-cell distributions includes generating, by applying the permutation invariant neural network, a plurality of encoded vectors for each of one or more single-cell distributions, one or more encoded vectors characterizing a plurality of blood cells of a corresponding single cell-distribution of one or more single-cell distributions.

[0021] In general, in a ninth aspect, combinable with the eighth aspect, generating one or more encoded vectors includes for each of one or more single-cell distributions includes: using a permutation equivariant encoder of the permutation invariant neural network.

[0022] In general, in a tenth aspect, combinable with the ninth aspect, generating one or more encoded vectors for each of one or more single-cell distributions includes normalizing the data representing the corresponding single-cell distribution, projecting the normalized data to an embedding space to generate a plurality of embedding vectors, and processing, using one or more attention blocks, one or more embedding vectors to generate one or more encoded vectors.

[0023] In general, in an eleventh aspect, combinable with the tenth aspect, each of the one or more attention blocks includes one or more multi-head attention blocks.

[0024] In general, in a twelfth aspect, combinable with the eighth through eleventh aspects, generating the encoded distribution for each of one or more single-cell distributions includes applying a permutation invariant aggregation function to one or more encoded vectors for the corresponding single-cell distribution.

[0025] In general, in a thirteenth aspect, combinable with the twelfth aspect, applying the permutation invariant aggregation function to one or more encoded vectors for the corresponding single-cell distribution includes processing, using one or more pooling attention blocks, one or more encoded vectors.Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0026] In general, in a fourteenth aspect, combinable with the thirteenth aspect, the one or more pooling attention blocks include one or more multi -head attention blocks.

[0027] In general, in a fifteenth aspect, combinable with the eighth through fourteenth aspects, one or more encoded vectors includes a covariate vector.

[0028] In general, in a sixteenth aspect, combinable with the fifteenth aspect, the additional covariate vector includes covariates containing at least one of: (i) age, (ii) year of measurement, or (iii) identity of the blood analyzer used for the measurements.

[0029] In general, in a seventeenth aspect, combinable with any of the preceding aspects, predicting the value indicative of the size of the blood cell population includes combining the encoded distribution of each of one or more single-cell distributions to generate a combined distribution vector, and decoding, using a decoder function, the combined distribution vector to generate the predicted value.

[0030] In general, in an eighteenth aspect, combinable with the seventeenth aspect, decoding the combined distribution vector includes processing, using one or more attention blocks, the combined distribution vector.

[0031] In general, in a nineteenth aspect, combinable with the eighteenth aspect, the one or more attention blocks include one or more self-attention blocks.

[0032] In general, in a twentieth aspect, combinable with any of the preceding aspects, the permutation invariant neural network includes a Transformer neural network.

[0033] In general, in a twenty first aspect, combinable with any of the preceding aspects, the permutation invariant neural network is pretrained on training datasets for the value.

[0034] In general, in a twenty second aspect, a method includes receiving data representing a plurality of single-cell distributions. Each of one or more single-cell distributions represents a plurality of blood cell measurements for one or more corresponding blood cell types of a plurality of blood cell types. One or more blood cell measurements collected using a complete blood count (CBC) test performed on a blood sample of a patient. The method includes predicting, using a permutation invariant first neural network and the respective encoded distributions, a value indicative of a size of a blood cell population of the patient, and generating, for each of one or more single-cell distributions and by using a second neural network, a plurality of contribution values indicative of contributions of one or more blood cell measurements to prediction of the value.

[0035] In general, in a twenty third aspect, combinable with the twenty second aspect, one or more single-cell distributions include: (i) a first single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of red blood cells, (ii)Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0036] a second single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of red blood cells and platelets, (iii) a third single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of nuclei of white blood cells, and (iv) a fourth single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of a peroxidase reaction in white blood cells.

[0037] In general, in a twenty fourth aspect, combinable with the twenty second or twenty third aspect, the blood cell population is: (i) a red blood cell population of the patient, (ii) a white blood cell population of the patient, or (iii) a platelet population of the patient.

[0038] In general, in a twenty fifth aspect, combinable with the twenty fourth aspect, the blood cell population is the red blood cell population of the patient, and the value includes a hematocrit of the red blood cell population or a total hemoglobin mass of the red blood cell population.

[0039] In general, in a twenty sixth aspect, combinable with any of the twenty second through twenty fifth aspects, one or more contribution values include predictions of Shapley values.

[0040] In general, in a twenty seventh aspect, combinable with any of the twenty second through twenty sixth aspects, one or more contribution values include at least one positive contribution value indicating that a first blood cell corresponding to a first blood cell measurement of one or more blood cell measurements contributes to an increase in the predicted size, and at least one a negative contribution value of the blood cell indicating that a second blood cell corresponding to a second blood cell measurement of one or more blood cell measurements contributes to a decrease in the predicted size.

[0041] In general, in a twenty eighth aspect, combinable with any of the twenty second through twenty seventh aspects, the second neural network is pretrained, using a machine learning model that generates a second value indicative of a size of blood cell populations from a plurality of training single-cell distributions, to generate one or more contribution values.

[0042] In general, in a twenty ninth aspect, combinable with the twenty eighth aspect, the machine learning model and the first neural network have the same architecture.

[0043] In general, in a thirtieth aspect, combinable with the twenty eighth or twenty ninth aspects, the machine learning model is pretrained on a plurality of randomly-masked training single-cell distributions generated by randomly masking one or more training single-cell distributions.Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0044] In general, in a thirty first aspect, combinable with the thirtieth aspect, one or more randomly -masked single-cell distributions are randomly masked using a plurality of randomly-generated ellipsoids overlaid on one or more training single-cell distributions.

[0045] In general, in a thirty second aspect, combinable with any of the twenty eighth through thirty first aspects, the second neural network is pretrained on one or more training single-cell distributions to generate one or more contribution values, the second neural network being pretrained based on a loss computed from the second predicted size and a sum of output values from the second neural network using one or more training single-cell distributions.

[0046] In general, in a thirty third aspect, a method includes receiving data representing (i) a predicted value indicative of a size of a blood cell population of a patient and (ii) a plurality’ of contribution values indicative of contributions of a plurality' of blood cell measurements to prediction of the predicted value. One or more blood cell measurements collected using a complete blood count (CBC) test performed on a blood sample of the patient. The method includes generating, from the predicted value and one or more contribution values, an interpretability map corresponding to the predicted value, determining, from the interpretability map, a region of the interpretability map including a subset of contribution values of one or more contribution values corresponding to a subset of blood cell measurements of one or more blood cell measurements is associated with an expected change in the predicted value, and in response to determining the region of the interpretability map, selecting the patient for workup or treatment for a pathophysiological state and / or a risk to develop the pathophysiological state.

[0047] In general, in a thirty fourth aspect, combinable with the thirty third aspect, the expected change includes either (i) an expected increase in the predicted value, or (ii) an expected decrease in the predicted value.

[0048] In general, in a thirty fifth aspect, combinable with the thirty third or thirty fourth aspect, one or more blood cell measurements are a plurality of measurements of white blood cells, and the blood cell population is a red blood cell population of the patient.

[0049] In general, in a thirty sixth aspect, combinable with the thirty’ fifth aspect, the value is indicative of: (i) a size of the red blood cell population, (ii) a hematocrit of the red blood cell population, or (iii) a total hemoglobin mass of the red blood cell population, and the region corresponds to a downshift region associated with the expected change.

[0050] In general, in a thirty seventh aspect, combinable with any of the thirty’ third or thirty sixth aspect, the region is at least partially defined by a percentile threshold.Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0051] In general, in a thirty eighth aspect, combinable with the thirty sixth or th i rty seventh aspect, the region indicates an expected decrease in the at least one of: (i) a value indicative of the size of the population of the red blood cells, (ii) a hematocrit of the red blood cell population, or (iii) a total hemoglobin mass of the red blood cell population.

[0052] In general, in a thirty ninth aspect, combinable with the thirty eighth aspect, the pathophysiological state includes at least one of sepsis, respiratory distress, coronary disease, stroke, heart failure, or diabetes.

[0053] In general, in a fortieth aspect, combinable with any of the thirty third through thirty eighth aspects, the interpretability map indicates a plurality of map values computed using a mesh and one or more contribution values.

[0054] In general, in a forty first aspect, combinable with the fortieth aspect, each of one or more map values includes a mean and a standard deviation for a corresponding cell of the mesh.

[0055] In general, in a forty second aspect, combinable with the fortieth aspect, one or more contribution values is normalized for signed analysis or absolute analysis.

[0056] In some implementations, the technologies described in this specification can allow for granular analysis of individual blood cells of a subject to compute a more accurate population-level attribute of a blood cell population of the subject. The technologies described in this specification can allow for the first transformer-based deep learning model to learn separate functions for each blood cell population while also modeling interactions between the encoded representations of the blood cell populations. In particular, the first transformer-based deep learning model can separately normalize and encode each singlecell distribution, and then combine the encoded distributions before decoding a combined distribution to enable both individualized set-normalization and cross-talk between the blood cell populations. As a result, the first transformer-based deep learning model can more accurately compute a population-level attribute of a blood cell population.

[0057] In some implementations, the technologies described in this specification, in determining a population-level attribute of a blood cell population of the subject, identify the contribution of each of the blood cells of a sample of blood cells from the subject used to determine the population-level attribute. In particular, the technologies described in this specification can quantify the contribution of each single cell in a sample of blood cells from the subj ect to the predicted population-level attribute by predicting a Shapley value for each of the blood cells. The second transformer-based deep learning model can efficiently estimate Shapley values with a single forward pass of the neural network.Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0058] In some implementations, the technologies described in this specification can be used to identify a set of blood cells of a sample of blood cells from a subject that had the highest contribution for determining a population-level attribute of the subject, and this identified set of blood cells can be used to determine a workup or treatment plan for the subject. In particular, the technologies described in this specification can generate and leverage interpretability maps to identify regions of a single-cell distribution associated with changes in population-level attributes of the subject. The interpretability maps can be used to find a set of blood cells of the sample of blood cells that had the highest contribution for determining a population-level attribute, and in response, this identified set of blood cells can be used to determine a pathophysiological state and / or risk of a pathophysiological state that can determine the workup or treatment plan for the subject.

[0059] In some implementations, the technologies described in this specification can be used to assist in deriving blood biomarkers of a sample of blood cells from a subject, and this derived biomarker can be used to determine a workup or treatment plan for the subject. In particular, the technologies described in this specification can leverage the interpretability maps to identify regions of single-cell distribution associated with changes in populationlevel attributes of the subject, and from these regions, a blood biomarker can be derived. For example, a region of large negative values for prediction of a blood cell population in a leftside region of a single-cell distribution can be used to identify a known blood biomarker “left-shift. ” As another example, the left-side region of the single-cell distribution and values indicative of hematocrit, hemoglobin, and red blood cell populations with both large positive and negative Shapley values can be used to assist in derivation of a blood biomarker “down-shift'’. More specifically, from the juxtaposition of the regions with positive and negative values, a shift of white blood cells from the top-left of the single-cell distribution to the bottom-left can be determined to be associated with a decrease in the size of hematocrit, hemoglobin, and red blood cell populations. In response, the “down-shift” biomarker can be derived and a workup or treatment plan for the patient can be determined.

[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Methods and materials are described herein for use in the present invention; other, suitable methods and materials known in the art can also be used. The materials, methods, and examples are illustrative only and not intended to be limiting. All publications, patent applications, patents, sequences, database entries, and other referencesAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0061] mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control.

[0062] Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and the claims.

[0063] BRIEF DESCRIPTION OF THE DRAWINGS FIG. 1 is a block diagram of an example system including a first neural network and a second neural network for determining a recommending treatment for a patient based on a more predicted population value of a blood cell population.

[0064] FIG. 2A is a block diagram of an example first neural network for predicting a value indicative of a population attribute of a blood cell population.

[0065] FIG. 2B is a block diagram of an example second neural network for predicting contribution values of blood cells from a blood sample on the predicted population values.

[0066] FIG. 3 is a graph indicative of a decrease in blood cell population based on a predicted blood cell population and a ground-truth blood cell population.

[0067] FIG. 4 is a flowchart of an example process for predicting a population value of a blood cell population.

[0068] FIG. 5 is a flowchart of an example process for predicting contribution values of blood cells from a blood sample on the predicted population values.

[0069] FIG. 6 is a flowchart of an example process for selecting a patient for treatment for a pathophysiological state and / or risk to develop a pathophysiological state based on the predicted contribution values.

[0070] FIG. 7 shows a set of graphs indicative of four input single-cell distributions.

[0071] FIG. 8 shows a training pipeline for the second neural network.

[0072] FIG. 9A shows a set of graphs indicative of an interpretability analysis for different blood cell populations.

[0073] FIG. 9B shows a set of graphs indicative of an interpretability7analysis for two sampling procedures of different single-cell distributions.

[0074] FIG. 10 shows a set of graphs indicative of an interpretability map for prediction of white blood cell population size.

[0075] FIG. 11 shows a set of graphs indicative of an interpretability map for prediction of platelet population size.

[0076] FIG. 12 shows a set of graphs indicative of an interpretability map for prediction of red blood cell volume fraction (hematocrit HCT).Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0077] FIG. 13 shows a set of graphs indicative of an interpretability map for prediction of total red blood cell hemoglobin mass.

[0078] FIG. 14 shows a set of graphs indicative of an interpretability map for prediction of red blood cell population size.

[0079] FIG. 15 A shows a set of graphs indicative of an ablation analysis of the accuracy of the first neural network for predicting each blood cell population size.

[0080] FIG. 15B shows a set of graphs indicative of Shapley value maps for each of the single-cell distributions when predicting each of the blood cell populations.

[0081] FIG. 16 shows a set of graphs indicative of a white blood cell distribution.

[0082] FIG. 17 shows a set of graphs indicative of the results of different prediction methods for prediction of a blood cell population size, including the first neural network.

[0083] FIG. 18A shows a set of graphs indicative of the error of the first neural network based on clinical reference intervals.

[0084] FIG. 18B shows a set of graphs indicative of the error of the first neural network based on sex.

[0085] FIG. 18C shows a set of graphs indicative of the error of the first neural network based on age-range

[0086] FIG. 19 shows a set of graphs indicative of Shapley values of a white blood cell biomarker associated with reduced red blood cell population size.

[0087] FIG. 20A shows a set of graphs indicative of the patients with and without a downshift biomarker and the association of lower hematocrit and white blood population values with the downshift biomarker.

[0088] FIG. 20B shows a graph indicative of an odds ratio for positive test results for standard inflammatory markers for those with and without the downshift biomarker.

[0089] FIG. 20C shows a graph indicative of the presence of downshift biomarkers for 25 diagnoses.

[0090] DETAILED DESCRIPTION

[0091] As described in this disclosure, blood cell population attributes, including size, can be predicted from single-cell measurements. Additionally, contribution values indicative of the contribution of single-cell measurements to the predictive blood cell population size can be predicted. Single-cell measurements can be one or more measurements corresponding to a single blood cell of one or more blood cells of a particular blood cell population from a sample collected using a complete blood count (CBC) test performed on a blood sample of the patient.Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0092] Described herein are one or more transformer-based neural networks trained and configured to receive single-cell measurements as input and predict one of (i) a blood cell population size for a particular blood cell population, or (ii) one or more contribution values indicative of a contribution of the one or more measurements on the predicted blood cell population size. Blood cell population size can correspond to the concentration and total count of a particular type of blood cell, or blood cell population, circulating in a patient’s bloodstream. For example, the blood cell population size can correspond to the concentration and total count of red blood cells, white blood cells, platelets, etc. circulating in a patient’s bloodstream. The contribution values can correspond to one or more predicted Shapley values representing the contribution of blood cell measurements of a blood cell to the predicted blood cell population size, or other blood cell population attribute. Upon generation of the one or more blood cell population sizes for a particular blood cell population or (ii) the contribution values, systems and methods described herein determine a pathophysiological state of the patient or a risk for the patient to develop a pathophysiological state with greater sensitivity and specificity.

[0093] The systems and method of the present disclosure can have one or more of the following advantages.

[0094] This technology can predict blood cell population size from single-cell measurements with high accuracy. In particular, the transformer-based neural network can utilize crosstalk between blood cell populations, leveraging the presence of co-regulatory processes that link the single-cell characteristics of one population (e.g. WBC) to the sizes of the others (e.g. RBC).

[0095] Further the accuracy of the transformer-based model can hint that single-cell data encodes substantial information about the dynamic physiologic processes regulating blood cell populations, including production, maturation, and clearance. Standard CBC parameters are already essential markers of an individual’s hematologic, immunologic, and hemodynamic states, and small differences in steady states or CBC setpoints have recently been shown to be associated with significant differences in risk of major diseases. By explaining more than 50% of the variance in cell population size, single-cell blood count data could have a dramatic effect on clinical inference from routine CBCs.

[0096] Additionally, this technology can validate the clinical utility of existing single-cell markers, generate data-driven hypotheses for the existence of previously unrecognized co-regulatory mechanisms, and enable identification of one or more single-cell markersAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0097] associated with subsequent diagnosis of inflammatory states and maj or diseases including sepsis, heart attack, stroke, and diabetes.

[0098] FIG. 1 is a block diagram of an example system 100 that obtains CBC data 110 from one or more CBCs 112 performed on a subject. The CBC data 110 represents one or more single-cell distributions 120. Each of the one or more single-cell distributions 120 represents one or more blood cell measurements for one or more corresponding blood cell types, or populations. For example, each of the one or more CBCs 112 can measure cell-level characteristics of individual cells derived from a sample collected from the subject. The number of blood cells measured using the CBCs 112 can vary in implementations. In some implementations, the number of blood cells can be at least 100, e.g., at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1,000, at least 10,000, at least 50,000, at least 100,000, or more. Blood cell types can include red blood cells, white blood cells, and platelets. In some implementations, the number of blood cells measured by the CBCs differs from the number of blood cell measurements used for forming the single-cell distributions 120. For example, the blood cell measurements forming the single-cell distributions 120 can correspond to a subset of the blood cells sampled using the CBCs. The number of blood cells whose measurements form the single-cell distributions 120 can be at least 100, e.g., at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1,000, at least 10,000, at least 50,000. at least 100,000, or more.

[0099] The system can predict, using a first neural network 130, a predicted blood cell attribute of a blood cell population of the patient, such as, for example, a blood cell population size 140 of one or more blood cell populations based on the one or more singlecell distributions 120. The system 100 can. using a second neural network 150, analyze the blood cell measurements for the one or more corresponding blood cell types of the one or more single cell distributions 120 and the predicted blood cell population sizes 140 of the one or more blood cell populations and generate the contribution values 155. The contribution values 155 indicate a measure of how much each of the measured blood cells contributed to the prediction of the predicted blood cell population sizes 140. Using the one or more predicted blood cell population sizes 140 and / or the one or more contribution values 155, a medical practitioner can determine a recommended workup of treatment plan 155 for a patient for a pathophysiological state and / or a risk to develop the pathophysiological state.Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0100] The system 100 can include one or more computing devices for performing these analyses. For example, the system 100 can include an input device, a network, and one or more computers (e.g., one or more local or cloud-based processors). The one or more computers can include an input processing engine that receives the data representing the multiple measurements from the one or more CBC tests 112, the first neural network 130 (or the one or more first neural networks as described in more detail below) to predict the blood cell population size 140, and the second neural network 150 to identify a blood biomarker from the data representing the one or more single cell distributions 120 and the predicted blood cell population sizes 140. In some implementations, the computer can be a server.

[0101] In the process illustrated in FIG. 1, the CBC data 110 is generated and received by the system 100. The CBC data 110 is generated through performing the one or more CBCs 112 in which a sample of blood cells is taken from a subject, e.g., a patient, and then analyzed using a CBC machine, e.g., CBC machines that analyze blood based on electrical impedance, flow cytometry, digital imaging, fluorescence, etc. In some implementations, the CBC data 110 can include data collected from only a single CBC performed on a sample of the subject. In some implementations, the CBC data 110 can include data collected from multiple CBCs performed on multiple different samples from the subject. The system 100 can obtain the data from the CBC test 112 in any appropriate manner. For example, the system 100 can include the CBC machine for generating the data from the CBC tests 112. In some implementations, the system 100 can obtain the data from the CBC tests 112 from an electronic data warehouse that can obtain the data, e.g., by accessing medical records of a patient, and transmit the patient data to another device such as a computer across a network. In some implementations, the electronic data warehouse can obtain the data that can be accessed by one or more other input devices such as a computer (e.g., desktop, laptop, tablet, etc.), a smartphone, or a server. In such instances, the one or more other input devices can access the patient data obtained by the electronic data warehouse and transmit the obtained patient data to a computer via a network. The network can include one or more of a wired Ethernet network, a wired optical network, a wireless WiFi network, a LAN, a WAN, a Bluetooth network, a cellular network, the Internet, or other suitable network, or any combination thereof. In some implementations, the electronic data warehouse and the computer are the same.

[0102] The one or more single-cell distributions 120 are determined from the CBC data 110. The one or more single-cell distributions 120 can be any single-cell distributionAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0103] representing one or more blood cell measurements for one or more corresponding blood cell types collected during a CBC, e.g., collected from the one or more CBCs 112 performed on one or more samples from the subject. A single-cell distribution of the one or more singlecell distributions 120 includes or represents the set of the individual measurements of the blood cells, and the single-cell distribution can include measurements for one type of blood cell or multiple types of blood cells. In the example shown in FIG. 1, the one or more single cell distributions 120 include a (1 ) single-cell distribution 122 including blood cell measurements of red blood cells 122, (2) a single-cell distribution 124 including blood cell measurements of red blood cells and platelets 124, (3) a single-cell distribution 126 including blood cell measurements of nuclei of white blood cells 126, and (4) a single-cell distribution including blood cell measurements of a peroxidase reaction in white blood cells 128.

[0104] The blood cells measured for the single-cell distributions 120 may vary in implementations. The blood cell measurements of the single-cell distributions 120 can be measurements of red blood cells, platelets, white blood cells (e.g., untreated white blood cells, nuclei of white blood cells that have treated with surfactant and stained, intensity of optical scatter of white blood cells, intensity of fluorescence of white blood cells), or combination of these types of cells. In examples in which the blood cell measurements are measurements of white blood cells, the measurements can be represented as a 3D vector with different regions corresponding to a white blood cell type, e.g., neutrophil, lymphocyte, monocyte, eosinophil, unclassified and debris.

[0105] The number of the single-cell distributions 120 can vary in implementations. The one or more single-cell distributions can be one single-cell distribution (e.g., any one of the single-cell distributions 122, 124, 126, or other single-cell distributions described in this specification). The one or more single-cell distributions can be two or more single-cell distributions (e.g., two or more of the single-cell distributions 122, 124, 126, or other singlecell distributions described in this specification).

[0106] After the system 100 determines the one or more single-cell distributions 120, the system 100 can use the first neural network 130 to predict a value indicative of a blood cell attribute of a population of one or more populations of blood cells of the subject. The data representing the one or more single-cell distributions 120 represents measurements of a sample of blood cells of the subject, and the first neural network 130 is trained to determine, from the one or more single-cell distributions 120, a population-level attribute of the sample of the blood cells of the subject. The system 100 can, using the first neural network 130 andAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0107] the data representing the one or more single-cell distributions 120 generate an encoded distribution for each of the one or more single-cell distributions. These encoded distributions, as described in greater detail in this specification, can be used to predict values indicative of sizes of populations of blood cells of the subject. The system 100 can then, using the first neural network 130 and the encoded distributions, predict the value indicative of a blood cell attribute of the population of blood cells of the subject. That is, the system 100 can leverage blood cell measurements of the one or more single-cell distributions, e.g., the single-cell distributions 122, 124, 126, and 128, to determine a predicted blood cell population attribute for a particular blood cell population. In other words, for example, the system 100 can leverage blood cell measurements of red blood cells, red blood cells and platelets, nuclei of white blood cells, and a peroxidase reaction in white blood cells, to determine a blood cell population size of red blood cells 142 in the patient’s bloodstream. The architecture of the first neural network 130 can vary in implementations. The first neural network 130 can be a transformer-based neural network. The first neural network 130 can be a permutation invariant transformer-based generative neural network such that an order of the blood cells with each of the single-cell distributions input to the first neural network 130 does not change the predictive blood cell population sizes 140.

[0108] In some implementations, the first neural network 130 can be a trained neural network. For example, the first neural network 130 can be trained on significantly large training datasets of one or more single-cell measurements and ground-truth values. The training dataset can include (i) one or more single-cell distributions from multiple subjects, and (ii) one or more ground-truth blood cell attribute data for a blood cell type for the subjects. In some implementations, the blood cell attribute predicted by the first neural network 130 is a specific blood cell population size. In some implementations, the system 100 can include one or more first neural networks 130, or a first neural network 130 and one or more additional neural networks identical in architecture to the first neural network 130, where each neural network is trained to predict a blood cell population size of a blood cell population of the subject. For example, the predicted blood cell population size can be a red blood cell population size, while an additional neural network that is identical in architecture to the first neural network 130 can be trained to predict a value indicative of a white blood cell size.

[0109] In some implementations, the blood cell attribute predicted by the first neural network 130 is a blood cell measurement for the one or more corresponding blood cell types not provided as input. That is, the first neural network 130 can be trained to predict a bloodAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0110] cell attribute that may be missing during collection of the blood cell measurements. As a result, the first neural network 130 can learn the structure or autocorrelations of the blood cell measurements and in some cases, can be fine-tuned to predict a variety of outcomes or features of interest based on the learned structure of the data. As an example, the training dataset can include (i) one or more single-cell measurements with one or more missing or masked measurements, and (ii) one or more ground-truth measurements corresponding to the missing or masked measurements.

[0111] In some implementations, the blood cell attribute predicted by the first neural network 130 includes corresponding one or more blood cell measurements for one or more corresponding blood cell types, the plurality of blood cell measurements collected using a second complete blood count (CBC) test performed on a second blood sample of a patient taken after the first CBC. That is, the first neural network 130 can be trained to predict one or more blood cell measurements of a subsequent CBC test for a particular patient. As a result, the first neural network 130 can be used to accurately predict a change in the blood cell measurements of the patients, e.g., an increase or decrease to determine a pathophysiological state and / or a risk to develop a pathophysiological state or a response to current treatment of a pathophysiological state. For example, in a patient with an undiagnosed infection, the first neural network 130 can be used to predict an increase in white blood cell measurements. As another example, in a patient with undiagnosed anemia or a risk to develop anemia, the first neural network 130 can be used to predict a decrease in red blood cell measurements. For example, the first neural network 130 can be used to predict a change, e.g., an increase or decrease in blood cell measurements, to predict a response of a patient to current treatment for a pathophysiological state. As a specific example, the first neural network 130 can be used to determine no change in blood cell measurements to determine a patient is not responding to treatment, or an increase / decrease in blood cell measurements to determine a patient is responding either in a negative or positive way to a treatment for a pathophysiological state. From this prediction, the system 100 can be used to enable faster optimization and personalization of treatment plans for patients for a pathophysiological state. For example, the first neural network 130 can be trained to predict whether a blood cell measurement of a subsequent CBC test will increase or decrease. As an example, the training dataset can include (i) one or more single-cell measurements from a first time point, and (ii) one or more ground-truth single-cell measurements from a second time point after the first time point.Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0112] The system 100 can utilize a loss function to train the first neural network 130 on the single cell measurements and the ground truth values. More specifically, the loss function can compare the predicted values indicative of a blood cell attribute by the first neural network 130 with the example ground-truth values indicative of a blood cell attribute in the training dataset. The loss function can be any appropriate loss function, including a mean squared error (MSE) loss function.

[0113] The system 100 can train the first neural network 130 to minimize the loss function by computing a gradient of the loss function with respect to the parameters of the first neural network, e.g., through backpropagation. The training system 100 can then apply an optimizer to the gradients to update the parameters of the first neural network 130. In some implementations, each of the components of the first neural network 130 can be trained jointly. In some implementations, some of the components of the first neural network 130 are trained separately.

[0114] In some implementations, the first neural network 130 can be fine-tuned. For example, the first neural network 130 can be fine-tuned to predict any clinical scenario of interest. As described above, in some implementations, the first neural network 130 can be fine-tuned to diagnose undetected infection, anemia, cancer, metabolic disease, and autoimmune diseases in all settings, as well as to predict response to current treatment.

[0115] In some implementations, the first neural network 130 can include one or more permutation equivariant encoders, one or more permutation invariant aggregation functions, and a decoder. In the example illustrated in FIG. 2A, the first neural network 130 includes multiple permutation equivariant encoders 131a-131d (collectively encoders 131), multiple permutation invariant aggregation functions 132a-132d (collectively aggregation functions 132), and a transformer decoder 133. The encoders 131 receive the single-cell distributions 120 and generate encoded vectors corresponding to cells measured for singlecell distributions 120. Each encoded vector characterizes a corresponding blood cell that was sampled in a CBC and encodes latent features of the blood cell. For each blood cell in each of the single-cell distributions, the encoders 131 generate an encoded vector representing the measurements of that blood cell. Moreover, each of the encoders 131 can be used for a corresponding one of the single-cell distributions 120. For example, the encoder 131 a generates a set of encoded vectors for the single-cell distribution 122, the encoder 131b generates a set of encoded vectors for the single-cell distribution 124, the encoder 131c generates a set of encoded vectors for the single-cell distribution 126. and the encoder 131 d generates a set of encoded vectors for the single-cell distribution 128. TheAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0116] number of blood cells, and hence the number of encoded vectors varies in implementations. For example, as shown in FIG. 2A. the single-cell distributions 120 can include measurements of 1000 blood cells, e.g., 1000 RBCs, 1000 platelets, and 1000 WBCs, although in other implementations, the number of measured blood cells may vary, as described in this specification.

[0117] The encoders 131 can be any appropriate permutation equivariant encoder, each configured to receive a single-cell distribution, encode data representing each blood cell within each single-cell distribution to generate one or more encoded vectors. Each of the encoders 131 can have any appropriate transformer architecture with any appropriate number of neural network layers, such as, for example, one or more transformer layers per single-cell distribution, one or more normalization layers per single-cell distribution, pooling layers, etc. As an example, each of the encoders 131 can have four transformer layer blocks, such as induced set attention blocks, each block including a normalization layer and two multi-head attention blocks.

[0118] To generate the one or more encoded vectors, the encoders 131 can, using a normalization function based on the corresponding single-cell distribution, normalize the data representing the corresponding single-cell distribution. The normalization function can be any appropriate type of normalization, such as batch normalization, layer normalization, group normalization, or set normalization. As a specific example, the encoders 131 can perform set-normalization, or normalization per set, on each of the single-cell distributions. Set normalization can correspond to techniques to normalize data within a specific ‘'set”, or unordered collection of items, to maintain permutation invariance. In some cases, the encoders 131 can leverage set normalization by normalizing the data representing the corresponding single-cell distribution. For example, the encoder 131a can normalize the single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of red blood cells 122, while the encoder 131b can normalize the single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of red blood cells and platelets 124. As such, each of the single-cell distributions 120 can be separately normalized for the data belonging to the single-cell distribution.

[0119] The encoders 131 can project the normalized data to an embedding space to generate one or more embedding vectors representing each of the blood cells of each of the corresponding single-cell distributions and then process, using one or more attention blocks, the one or more embedding vectors to generate the one or more encoded vectors. The one orAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0120] more attention blocks can be any appropriate attention mechanisms, such as multi-head attention mechanisms within induced set attention blocks. The multi-head attention mechanisms can be any appropriate type of attention, such as cross-attention. For example, the one or more attention mechanisms within an induced set attention block can include a set of learnable parameters (“inducing points”) that act as a working memory for the set and a first multi-head attention mechanism can update the learnable parameters based on the data, or features, of the set and a second multi-head attention mechanism can infuse the embedding vectors of the set based on the updated learnable parameters. As such, each embedding vector can be infused with the global context of the single-cell distribution, e.g., the one or more measurements of the one or more blood cells within the single-cell distribution, and a relational context to reflect the positioning of the corresponding blood cell within the single-cell distribution, e.g., is it an outlier, a central point, etc.. Additionally, the induced set attention block can maintain permutation invariance as the permutation equivariant encoder does not utilize positional encoding to know which blood cell is received first in the single cell distribution and the input is treated as a cloud of points, not a sequence of data.

[0121] In some implementations, the single-cell distributions 120 can further include a covariate vector. The covariate vector can correspond to a vector representing one or more input features of a blood cell of the single-cell distribution, the single-cell distribution, or the patient to which the blood sample used to obtain the measurements for the single-cell distribution belongs. For example, the additional covariate vector can include covariates containing at least one of: (i) age of the patient, (ii) year of measurement, or (iii) identity of the blood analyzer used for the measurements.

[0122] The aggregation functions 132 of the first neural network 130 can generate encoded distributions 202a, 202b, 202c, 202d (collectively encoded distributions 202) for the singlecell distributions 120, respectively. The aggregation functions 132 can be any appropriate permutation invariant aggregation functions that are each configured to receive one or more encoded vectors of a single-cell distribution of the single-cell distributions 120 and aggregate the vectors to generate a single encoded vector, or distribution, of the encoded distributions 202. The aggregation functions 132 can include any transformer architecture with any number of transformer layers, including, for example, a pooling attention block that includes one or more pooling layers and one or more attention mechanism. For example, the aggregation functions can include one or more multi-head attention blocks that can perform pooling by including a learnable seed vector as the query vector to determineAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0123] the most important features for attention and limit the number of output vectors to a single vector representing the encoded distribution.

[0124] More specifically, the first neural network 130 can combine, using the aggregation functions 132, the one or more encoded vectors corresponding to a single-cell distribution of the one or more single-cell distributions 120 to generate the encoded distributions 202. To aggregate the one or more encoded vectors, in an example in which the aggregation functions 132 include one or more pooling attention layers, the aggregation functions 132 can process, using the one or more pooling attention blocks, the one or more encoded vectors corresponding to the single-cell distributions 120 to generate the encoded distributions 202. More specifically, each of the aggregation functions 132 can utilize a seed vector to probe the contextualized data of the one or more encoded vectors and extract only the most important features into a single, encoded distribution vector using a multi-head attention mechanism. The first neural network 130 can utilize the aggregation functions 132 to aggregate one or more encoded vectors for each of the single cell distributions 120. For example, the first neural network 130 can utilize the aggregation function 132a to aggregate one or more encoded vectors of the single-cell distribution in which the corresponding one or more blood cell measurements are measurements of one or more red blood cells 122 to generate the encoded distribution 202a for the corresponding single-cell distribution 122.

[0125] Using the one or more encoded distributions 202, the first neural network 130 can predict the value indicative of the blood cell attribute 204. The value indicative of a blood cell attribute 204 can be any value indicative of a blood cell value predictable by the neural network 130 from single-cell distribution data. The values 204, as described in this specification, can include the blood cell population sizes 140 shown in and described with respect to FIG. 1. Examples of the value 204 indicative of a blood cell attribute includes a value indicative of a blood cell population of the patient, a value indicative of a blood cell measurement for the one or more corresponding blood cell types not provided as input, and a value indicative of corresponding one or more blood cell measurements for one or more corresponding blood cell types, the one or more blood cell measurements collected using a second complete blood count (CBC) test performed on a second blood sample of a patient taken after the first CBC. For example, the system 100 can predict a value indicative of a blood cell population of a particular blood cell population type of the patient. Examples of blood cell population types can include a red blood cell population of the patient, a white blood cell population of the patient, and a platelet population of the patient. As an example, the blood cell population can be the red blood cell population of the patient and the valueAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0126] can include a hematocrit of the red blood cell population, and / or a total hemoglobin mass of the red blood cell population. As an example, the system 100 can predict a value indicative of a blood cell measurement for the one or more corresponding blood cell types not provided as input. For example, the system 100 can predict one or more measurements of the one or more single-cell distribution not provided to the system 100, or in some cases, masked to the system 100. As an example, the system 100 can predict a value indicative of corresponding one or more blood cell measurements for one or more corresponding blood cell types, the one or more blood cell measurements collected using a second complete blood count (CBC) test performed on a second blood sample of a patient taken after the first CBC. For example, the system 100 can predict a value indicative of single-cell blood cell measurements for the patient for a next CBC test at a next time point. As described above, in some implementations, the first neural network 130 can be one or more first neural networks identical in architecture to the first neural network 130, where each neural network is trained to predict a blood cell population size of one or more blood cell population types. For example, the system 100 can include five neural networks identical in architecture to the first neural network 130: (i) a neural network for predicting a red blood cell population of the patient, (ii) a neural network for predicting a white blood cell population of the patient, (iii) a neural network for predicting a platelet population of the patient, (iv) a neural network for predicting a hematocrit of the red blood cell population, and (v) a neural network for predicting a total hemoglobin mass of the red blood cell population.

[0127] The transformer decoder 133 of the first neural network 130 is used to predict based on the encoded distributions 202, a value 140 indicative of a blood cell attribute of blood cell populations of the subject. The decoder 133 can be any appropriate decoder configured to receive one or more encoded distributions and generate a predicted value 140 indicative of a blood cell attribute. The decoder can include any appropriate transformer-based architecture, including one or more decoding layers. The one or more decoding layers can be any appropriate neural network layer, such as one or more self-attention layer blocks.

[0128] To predict the value 204 indicative of a blood cell attribute, the decoder 133 can combine the encoded distributions 202 representing each of the single-cell distributions 120 to generate a combined distribution vector and decode, using a decoder function the combined distribution vector to generate the predicted value. That is, the decoder 133 can combine the one or more encoded set-normalized single-cell distributions 202 to generate a combined distribution vector that is then decoded to generate the predicted value 204. For example, the decoder 133 can process the combined distribution vector using one or moreAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0129] self-attention mechanisms in one or more self-attention layers blocks. The self-atention layer blocks analyze the combined distribution vector and dynamically weigh the importance of different parts of the data relative to one another. As a result, each of the encoded distributions 202 can crosstalk, or influence the predicted values of a different blood cell type depending on the value 204. For example, for a predictive value 204 indicative of a blood cell population size, the decoder 133 can utilize the measurements of each of the single-cell distributions 120, e.g., single-cell distributions of 122, 124, 126, and to predict the red blood cell population of the patient 142. As such, the blood cell measurements of platelets, and various WBCs can influence the prediction of the red blood cell population size 142, for example.

[0130] Referring back to FIG. 1, after the system 100 computes the value 204 indicative of the blood cell attribute, the system 100 can (i) predict the direction of change of the blood cell population size of a particular blood cell type (e.g., increasing or decreasing), and in response, predict a pathophysiological state based on the predicted direction of change of the blood cell population or (ii) generate, using a second neural network 130, one or more contributions values 155 and from the one or more predicted contribution values, select the patient for workup or treatment for a pathophysiological state and / or risk to develop a pathophysiological state.

[0131] The system 100 can predict the direction of change of the blood cell population size of a particular blood cell population (e.g., increasing or decreasing), and in response, a pathophysiological state and / or risk for the patient to develop the pathophysiological state can be predicted based on the predicted direction of change of the blood cell population size. Examples of pathophysiological states can include anemia, cancer, liver and kidney failure, latent infection, or treatment failure for any treatment where the hematological trajectory of the patient (e.g., represented by the blood cell measurements collected in a CBC test) is associated with the response of a patient to the treatment. The risk for the patient to develop the pathophysiological state can include determining the risk to develop at least one of: an infection, a malignancy, a disease, anemia, or diabetes.

[0132] In some implementations, the prediction of the pathophysiological state and / or risk for the patient to develop the pathophysiological state can be based on the predicted direction of change of the population size of a single blood cell population. For example, the prediction of the pathophysiological state and / or risk for the patient to develop the pathophysiological state can be predicted based on the predicted direction of change of the white blood cell population of the patient.Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0133] In some implementations, the predicted direction of change can be determined by comparing the predicted blood cell population size of a blood cell population and the ground truth blood cell population size. In particular, the system 100 can compare the predicted blood cell population size of a particular blood cell population and the ground-truth blood cell population size measured by the CBC, and determine from the comparison a predicted direction of change. For example, if the predicted blood cell population size of a particular blood cell population is lower by a margin than the ground-truth blood cell population size, the blood cell population size can be decreasing. As another example, if the predicted blood cell population size of a particular blood cell population is higher by a margin than the ground-truth blood cell population size, the blood cell population size can be decreasing.

[0134] In some implementations, the predicted direction of change can be determined by predicting the blood cell measurements for one or more corresponding blood cell types, the plurality of blood cell measurements collected using a second complete blood count (CBC) test performed on a second blood sample of a patient taken after the first CBC. That is, the predicted direction of change can be determined by predicting subsequent blood cell measurements and comparing current blood cell measurements with predicted blood cell measurements of a subsequent CBC test for a particular patient. As a result, a predicted direction of change can be determined, e.g., an increase or decrease in the blood cell measurements, or in blood cell population size. In response to the predicted direction of change, the pathophysiological state and / or risk for the patient to develop the pathophysiological state can be predicted. For example, for a prediction of increasing white blood cell population size, latent infection can be predicted and further tests, diagnosis, and treatments can still be ordered / implemented, such as enhanced surveillance, follow-up microbiological testing, or utilization of empiric antibiotic treatment As an example,. for a prediction of increasing white blood cell population size, latent anemia, or a ferntin deficient state) can be predicted and further tests, diagnosis, and treatments can still be ordered / implemented, such as iron studies (e.g., ferritin, total iron binding capacity, transferrin saturation).

[0135] After the system 100 generates the blood cell population sizes 140, the second neural network 150 can be used to generate contributions value 155 and then, from the predicted contribution values 155, select the patient for workup or treatment for a pathophysiological state and / or risk to develop a pathophysiological state. The system 100 can leverage the predicted contribution values 155 to assist in identifying a blood biomarker, and based on blood cell data of the patient and whether the biomarker is present, select the patient forAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0136] workup or treatment for a pathophysiological state and / or risk to develop a pathophysiological state. The contribution values 155 can be indicative of contributions of the one or more blood cell measurements to prediction of the value. In some implementations, the contribution values 155 can include a contribution value for each of the one or more blood cells in the single-cell distribution to predict the value. The contribution values 155 can be any appropriate contribution values, such as predictions of Shapley values. Shapley values measure the contribution of a blood cell, or the blood cell measurements of a particular blood in a single-cell distribution 120, based on the contribution of the blood cell in every possible combination of blood cells in the single-cell distribution 120. In some implementations, the contribution values 155 can include at least one positive contribution value indicating that a first blood cell corresponding to a first blood cell measurement of the plurality of blood cell measurements contributes to an increase in the predicted size, and at least one a negative contribution value of the blood cell indicating that a second blood cell corresponding to a second blood cell measurement of the plurality of blood cell measurements contributes to a decrease in the predicted size. As such, the one or more contribution values 155 can give the system 100 insights on which particular blood cells, and blood cell measurements of the particular blood cells, influence the predicted blood size population the most, either reflecting an increase in population size or a decrease in population size.

[0137] The architecture of the second neural network 150 can vary in implementations. The second neural network 150 can be a transformer-based neural network. The second neural network 150 can have any appropriate transformer architecture with any appropriate number of neural network layers, such as, for example, one or more transformer layers per blood cell population, one or more normalization layers, pooling layers, etc. In the example illustrated in FIG. 2A, the second neural network 150 can have multiple transformer blocks 151a, 151b, 151c, 151 d (collectively transformer blocks 151). For example, the transformer blocks 151 can be induced set attention blocks, each block including a normalization layer and two multi-head attention blocks. The multi-head attention mechanisms can be any appropriate type of attention, such as cross-attention. For example, the one or more attention mechanisms within an induced set attention block can include a set of learnable parameters (“inducing points”) that act as a working memory for the singlecell distribution and a first multi-head attention mechanism can update the learnable parameters based on the data, or features, of the single-cell distribution and a second multihead attention mechanism can infuse the embedding vectors of the single-cell distributionAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0138] based on the updated learnable parameters. As such, each embedding vector can be infused with the global context of the single-cell distribution, e.g., the one or more measurements of the one or more blood cells within the single-cell distribution, and a relational context to reflect the positioning of the corresponding blood cell within the single-cell distribution, e.g., is it an outlier, a central point, etc., to determine the one or more blood cells within the single-cell distribution that influence the contribution values 255 the most.

[0139] In some implementations, the second neural network 150 can be pretrained to generate the contribution values 255a-d (collectively the contribution values 255), e.g., the predicted Shapley values for each of the one or more blood cells of each of the one or more single-cell distributions 120. In particular, the second neural network 150 can be pretrained, using a machine learning model that generates a second value indicative of a size of blood cell populations from one or more training single-cell distributions, to generate the contribution values 255. The machine learning model and the first neural network 130 have the same architecture. That is, to pretrain the second neural network 150, the system 100 can (i) train a machine learning model that is a copy of the first neural network 130 to predict a second value indicative of a blood cell population, and (ii) pre-train the second neural network 150 generate the contribution values 255. To train the machine learning model that is a copy of the first neural network 130 to predict a second value indicative of a blood cell population, the system 100 can train the machine learning model on one or more randomly-masked training single-cell distributions generated by randomly masking the one or more training single-cell distributions. The system 100 can generate the randomly -masked training single-cell distributions using any appropriate technique, including, for example, randomly masking the single-cell distributions using one or more randomly -generated ellipsoids overlaid on the one or more training single-cell distributions. To train the second neural network 150 on the one or more training single-cell distributions to generate the one or more contribution values, the second neural network 150 can be being pretrained based on a loss computed from the second predicted size and a sum of output values from the second neural network using the plurality of training single-cell distributions.

[0140] Referring back to FIG. 1, from the generated contribution values 255 and the predicted blood cell population size 140, the system 100 can select the patient for workup or treatment for a pathophysiological state and / or a risk to develop the pathophysiological state. More specifically, the system 100 can generate, from the predicted value 140 and the contribution values 155, an interpretability map corresponding to the predicted value and determine, from the interpretability map, a region of the interpretability map including aAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0141] subset of contribution values of the one or more contribution values corresponding to a subset of blood cell measurements of the plurality of blood cell measurements is associated with an expected change in the predicted value. It is in response to determining the region of the interpretability map, that the system 100 can select the patient for workup or treatment for a pathophysiological state and / or a risk to develop the pathophysiological state.

[0142] More specifically, the system 100 can generate, from the predicted value, e.g., the predicted blood cell population size 140. and the contribution values 155. an interpretability map corresponding to the predicted value. The interpretability map indicates one or more map values computed using a mesh and the contribution values 155. That is, for every cell of the mesh, the one or more map values can include a mean and a standard deviation. Before computing the one or more map values, the system 100 can normalize the contribution values 155 using one or more methods, such as a signed analysis or an absolute analysis.

[0143] The system 100 can determine, from the interpretability map, a region of the interpretability map including a subset of contribution values of the contribution values 155 corresponding to a subset of blood cell measurements of the plurality of blood cell measurements is associated with an expected change in the predicted value. In particular, the system 100 can determine, from the interpretability map, a region of contribution values corresponding to a set of cells that are associated with an expected change in the predicted value indicative of a blood cell attribute. The expected change can be any appropriate change in the predictive value, such as either (i) an expected increase in the predicted value, or (ii) an expected decrease in the predicted value. In some implementations, the region is at least partially defined by a percentile threshold.

[0144] By identifying regions of single-cell distributions associated with changes in population sizes, the second neural network 150 can provide data-driven hypotheses for single-cell markers of cell population kinetics. That is, the identified regions of the single-cell distributions and the corresponding interpretability maps can be leveraged to derive singlecell blood biomarkers.

[0145] For example, the interpretability maps can be leveraged to determine large negative Shapley values for prediction of a population size of white blood cells (NJVBC ) in the leftside region of the single-cell distribution in which the corresponding blood cell measurements are measurements of nuclei of white blood cells PWBC-BASOS. This region can be associated with and define the clinical ‘“Left Shift” (LS) blood biomarker that is alreadyAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0146] known to be associated with altered WBC kinetics, e.g., how white blood cells move, distribute, or disappear in the blood.

[0147] As another example, the interpretability maps can be leveraged to determine a connection between the left-side region of the single-cell distribution in which the corresponding blood cell measurements are measurements of nuclei of white blood cells PWBC-BASOS and the hematocrit of the red blood cell population N CT), the total hemoglobin mass of the red blood cell population (NHGB), and red blood cell population NRBC). In particular large positive and negative Shapley values in close proximity can be identified from the interpretability maps. The juxtaposition of the regions of the interpretability maps with positive and negative Shapley values can suggest that a shift of white blood cells from the top left of the single-cell distribution in which the corresponding blood cell measurements are measurements of nuclei of white blood cells PWBC-BASOS ) to the bottom left may be particularly strongly associated with a decrease in the size of the hematocrit of the red blood cell population (NHCT), the total hemoglobin mass of the red blood cell population NHGB), and the red blood cell population (NRBC). That is, for example, the blood cell population includes a population of white blood cells, and the identified region of the interpretability maps can correspond to a downshift region associated with an expected decrease in the values indicative of the size of the population of the red blood cells. As a result, the determined downshift region can be leveraged in the derivation of “Down Shift,'’ a blood biomarker.

[0148] In some implementations, the identified region is at least partially defined by a percentile threshold. For example, the determined downshift region associated by the " Down Shift’' (DS) binary marker can be defined to be positive when the y-coordinate of the centroid for the leftmost portion of the single-cell distribution in w hich the corresponding blood cell measurements are measurements of nuclei of white blood cells PWBC-BASOS) was below the 25thpercentile.

[0149] As a result of determining the region of the interpretability map that can be used to identify one or more blood biomarkers, the system 100 can select the patient for workup or treatment for a pathophysiological state and / or a risk to develop the pathophysiological state. Examples of pathophysiological states include sepsis, respiratory distress, coronary disease, stroke, heart failure, and diabetes. In particular, for example with the “Down Shift” biomarker, the system 100 can receive blood cell measurements and based on the leftmost portion of the single-cell distribution in which the corresponding blood cell measurements are measurements of nuclei of white blood cells, determine whether to select the patient forAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0150] workup or treatment for a pathophysiological state and / or a risk to develop the pathophysiological state. For example, if the y-coordinate of the centroid for the leftmost portion of the single-cell distribution in which the corresponding blood cell measurements are measurements of nuclei of white blood cells (PWBC-BASOS ) is below the 25thpercentile, the system 100 can select the patient for workup or treatment for sepsis, respiratory distress, coronary disease, stroke, heart failure, diabetes, or other diseases associated with a decrease in the size of the hematocrit of the red blood cell population (NHCT). the total hemoglobin mass of the red blood cell population (NHGB), and the red blood cell population (NRBC FIG. 3 shows a graph indicative of a predicted blood cell population by the first neural network, e g., the first neural network 130 of FIG. 1, lower than the ground truth value of the blood cell population. In particular, the graph is indicative of a predicted white blood cell population 344 by the first neural network that is a lower value than the ground truth value 354 of the white blood cell population, accurately reflecting that the white blood cell population is decreasing.

[0151] The graph showcases ground-truth white blood cell populations taken from multiple CBCs at multiple time points. The graph can display data, from multiple CBC tests taken over a period of time, representing multiple measurements of a white blood cell test, where each of the multiple measurements is collected from a corresponding CBC test. The multiple CBC tests can be taken with any interval of time in between the tests such as, for example, 7 or more days. 14 or more days. 30 or more days. 60 or more days. 90 or more days, 180 or more days, between 7 and 14 days, between 14 and 30 days, between 7 and 60 days, between 7 and 90 days, between 7 and 180 days, etc. In some implementations, the interval of time between each test is at least 90 days apart. For example, the multiple CBC tests can be 2 tests, each taken a year apart, spanning 2 years.

[0152] As depicted in FIG. 2, from the previous white blood cell population 356 to the ground truth white blood cell population 354, the graph showcases that the white blood cell population is decreasing for the patient. The predicted white blood cell population 344 by the first neural netw ork 130 of FIG. 1 can accurately reflect the decreasing nature of the white blood cell population. That is, the system 100 can predict the decreasing nature of the white blood cell population by comparing the predicted white blood cell population with the ground-truth white blood cell population measured by the CBC. As such, the system 100 can predict a direction of change of a white blood cell population of a patient from a single CBC test, and does not require previous CBCs to determine a direction of change and a potential pathophysiological state and / or the risk of developing a pathophysiological state.Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0153] Further, the system 100 can predict directions of change of blood cell populations as compared to a patient-specific hematological setpoint, providing further insights to the hematological state of the patient. The patient-specific hematological setpoint can represent a personalized baseline for a patient within a population wide reference interval for each of one or more hematological parameters measured in a CBC test. Examples of hematological parameters include red blood cell (RBC) volume 122 (also referred to as HCT), RBC hemoglobin mass (also referred to as MCH). RBC size (also referred to as MCV), RBC hemoglobin concentration (also referred to as MCHC), hemoglobin (also referred to as HGB), platelet size (also referred to as MPV), platelet count (also referred to as PLT), RBC count (also referred to as RBC), RBC size variation (also referred to as RDW), and white blood cell (WBC) count (also referred to as WBC). For example, as depicted in FIG. 3, a hematological setpoint for a white blood cell count can be calculated based on a history of white blood cell counts of the patient. That is, the subsequent calculated hematological setpoints can represent homeostatic measurements of a patient over multiple years. For the white blood cell population, or white blood cell count, that is decreasing towards the hematological setpoint of the patient, the first neural network 130 can predict a blood cell population lower than the current white blood cell population. In response, insights into the patient’s hematological setpoint, or homeostatic normal, can be provided to assist in determination of a pathophysiological state and / or risk to develop a pathophysiological state.

[0154] FIG. 4 is a flowchart of an example process for predicting a population value of a blood cell population.

[0155] For convenience, the process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, a system, e.g., the system 100 depicted in FIG 1, appropriately programmed in accordance with this specification, can perform the process 400.

[0156] The system can receive data representing one or more single-cell distributions, where each of the one or more single-cell distributions represents a corresponding plurality of blood cell measurements for one or more corresponding blood cell types of one or more blood cell types, the one or more blood cell measurements collected using a complete blood count (CBC) test performed on a blood sample of a patient (402). As described above with reference to FIG. 1 and 2A, the one or more single-cell distributions include: (i) a first single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of red blood cells, (ii) a second single-cell distribution in which theAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0157] corresponding plurality of blood cell measurements are measurements of red blood cells and platelets, (iii) a third single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of nuclei of white blood cells, and (iv) a fourth singlecell distribution in which the corresponding plurality of blood cell measurements are measurements of a peroxidase reaction in white blood cells.

[0158] The system can generate, using a permutation invariant first neural network and the data representing the one or more single-cell distributions, an encoded distribution for each of the one or more single-cell distributions (404). As described above with reference to FIG.

[0159] 1, and 2 A, generating the encoded distribution for each of the one or more single-cell distributions includes generating, by applying the permutation invariant neural network, one or more encoded vectors for each of the one or more single-cell distributions, the one or more encoded vectors characterizing one or more blood cells of a corresponding single cell distribution of the one or more single-cell distributions.

[0160] As described above with reference to FIG. 1 and 2A, the first neural network can include a permutation equivariant encoder, a permutation invariant aggregation, and a decoding function. The system can generate the one or more encoded vectors for each of the one or more single-cell distributions using a permutation equivariant encoder of the permutation invariant neural network. Generating the one or more encoded vectors for each of the one or more single-cell distributions includes: normalizing the data representing the corresponding single-cell distribution, projecting the normalized data to an embedding space to generate one or more embedding vectors, and processing, using one or more attention blocks, the one or more embedding vectors to generate the one or more encoded vectors. Each of the one or more attention blocks includes one or more multi-head attention blocks. The system can also generate the encoded distribution for each of the one or more single-cell distributions by applying a permutation invariant aggregation function to the one or more encoded vectors for the corresponding single-cell distribution. Applying the permutation invariant aggregation function to the one or more encoded vectors for the corresponding single-cell distribution includes processing, using one or more pooling attention blocks, the one or more encoded vectors. The one or more pooling attention blocks include one or more multi-head attention blocks.

[0161] In some implementations, the one or more encoded vectors includes a covariate vector. The additional covariate vector includes covariates containing at least one of: (i) age, (ii) year of measurement, or (iii) identity of the blood analyzer used for the measurements.Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0162] The permutation invariant neural network can include a Transformer neural network. In some implementations, the permutation invariant neural network is pretrained on training datasets for the value.

[0163] The system can predict, using the respective encoded distributions, a value indicative of a blood cell attribute of a blood cell population of the patient (406). As described above with reference to FIG. 1 and 2A, in some implementations, the blood cell attribute includes a size of a blood cell population of the patient. In some implementations, the blood cell population is: (i) a red blood cell population of the patient, (ii) a white blood cell population of the patient, or (iii) a platelet population of the patient. In some implementations, the blood cell population is the red blood cell population of the patient, and the value includes a hematocrit of the red blood cell population or a total hemoglobin mass of the red blood cell population.

[0164] In some implementations, the blood cell attribute includes a blood cell measurement for the one or more corresponding blood cell types not provided as input.

[0165] In some implementations, the blood cell attribute includes a corresponding plurality of blood cell measurements for one or more corresponding blood cell types, the one or more blood cell measurements collected using a second complete blood count (CBC) test performed on a second blood sample of a patient taken after the first CBC.

[0166] Predicting the value indicative of the size of the blood cell population includes combining the encoded distribution of each of the one or more single-cell distributions to generate a combined distribution vector; and decoding, using a decoder function, the combined distribution vector to generate the predicted value. In some implementations, decoding the combined distribution vector includes: processing, using one or more attention blocks, the combined distribution vector. The one or more attention blocks include one or more self-attention blocks.

[0167] FIG. 5 is a flowchart of an example process for predicting contribution values of blood cells from a blood sample on the predicted population values.

[0168] For convenience, the process 500 will be described as being performed by a system of one or more computers located in one or more locations. For example, a system, e.g., the system 100 depicted in FIG 1, appropriately programmed in accordance with this specification, can perform the process 500.

[0169] The system can receive data representing one or more single-cell distributions, where each of the one or more single-cell distributions represents one or more blood cell measurements for one or more corresponding blood cell types of one or more blood cellAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0170] types, the one or more blood cell measurements collected using a complete blood count (CBC) test performed on a blood sample of a patient (502). As described above with reference to FIG. 1 and 2B, the one or more single-cell distributions include: (i) a first single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of red blood cells, (ii) a second single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of red blood cells and platelets, (iii) a third single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of nuclei of white blood cells, and (iv) a fourth singlecell distribution in which the corresponding plurality of blood cell measurements are measurements of a peroxidase reaction in white blood cells.

[0171] The system can predict, using a permutation invariant first neural network and the respective encoded distributions, a value indicative of a size of a blood cell population of the patient (504). As described above with reference to FIG. 1 and 2A, the permutation invariant first neural network can include a permutation equivariant encoder, a permutation invariant aggregate function and a decoder.

[0172] In some implementations, the blood cell population is: (i) a red blood cell population of the patient, (ii) a white blood cell population of the patient, or (iii) a platelet population of the patient. In some implementations, the blood cell population is the red blood cell population of the patient, and the value includes a hematocrit of the red blood cell population or a total hemoglobin mass of the red blood cell population.

[0173] The system can generate, for each of the one or more single-cell distributions and by using a second neural network, one or more contribution values indicative of contributions of the one or more blood cell measurements to prediction of the value (506). As described above with reference to FIG. 1 and 2B. the system can generate the multiple contribution values indicative of which blood cell measurement of the blood cells of the single-cell distribution are most influential to prediction of the value.

[0174] In some implementations, the one or more contribution values include predictions of Shapley values. In some implementations, the one or more contribution values include at least one positive contribution value indicating that a first blood cell corresponding to a first blood cell measurement of the one or more blood cell measurements contributes to an increase in the predicted size, and at least one a negative contribution value of the blood cell indicating that a second blood cell corresponding to a second blood cell measurement of the one or more blood cell measurements contributes to a decrease in the predicted size.Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0175] In some implementations, the second neural network is pretrained, using a machine learning model that generates a second value indicative of a size of blood cell populations from one or more training single-cell distributions, to generate the one or more contribution values, second neural network is pretrained on the one or more training single-cell distributions to generate the one or more contribution values, the second neural network being pretrained based on a loss computed from the second predicted size and a sum of output values from the second neural network using the one or more training single-cell distributions.

[0176] In some implementations, the machine learning model and the first neural network have the same architecture. In some implementations, the machine learning model is pretrained on one or more randomly -masked training single-cell distributions generated by randomly masking the one or more training single-cell distributions. The one or more randomly -masked single-cell distributions are randomly masked using one or more randomly -generated ellipsoids overlaid on the one or more training single-cell distributions.

[0177] FIG. 6 is a flowchart of an example process for selecting a patient for treatment for a pathophysiological state and / or risk to develop a pathophysiological state based on the predicted contribution values.

[0178] For convenience, the process 600 will be described as being performed by a system of one or more computers located in one or more locations. For example, a system, e.g., the system 100 depicted in FIG 1, appropriately programmed in accordance with this specification, can perform the process 600.

[0179] The system can receive data representing (i) a predicted value indicative of a size of a blood cell population of a patient and (ii) one or more contribution values indicative of contributions of one or more blood cell measurements to prediction of the predicted value, the one or more blood cell measurements collected using a complete blood count (CBC) test performed on a blood sample of the patient (602). As described in more detail above with reference to FIG. 1 and 2B, in some implementations, the system can receive data representing a predictive value indicative of a size of a blood cell population of a patient from a first stage of the system and multiple contribution values from a second stage of the system. The predicted value indicative of a size of a blood cell population of a patient can be, for example, the predicted value indicative of a white blood cell population, a red blood cell population, a platelet population, a hematocrit of the red blood cell population or a total hemoglobin mass of the red blood cell population. The contribution values can be indicative of contributions of blood cell measurements to the prediction of the blood cell populationAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0180] size, such that each blood cell of a single-cell distribution of a blood cell population can be associated with a contribution value to the prediction of the blood cell population size.

[0181] The system can generate, from the predicted value and the one or more contribution values, an interpretability map corresponding to the predicted value (604). As described above with reference to FIG. 1 and FIG. 2B, the system can generate, from the predicted value of the blood cell population and the multiple contribution values, an interpretability map that includes a contribution value for each of the blood cells of the single-cell distribution.

[0182] As described above with reference to FIG. 1, the interpretability map can indicate one or more map values computed using a mesh and the one or more contribution values, where each of the one or more map values includes a mean and a standard deviation for a corresponding cell of the mesh. In some implementations, the one or more contribution values are normalized for signed analysis or absolute analysis.

[0183] The system can determine, from the interpretability map, a region of the interpretability map including a subset of contribution values of the one or more contribution values corresponding to a subset of blood cell measurements of the one or more blood cell measurements is associated with an expected change in the predicted value (606). As described above with reference to FIG. 1 and 2B, the region of the interpretability map can indicate an expected change in the predicted value, such as an expected increase in the predicted value or an expected increase in the predicted value. In some implementations, the region can be at least partially defined by a percentile threshold. For example, the region can be at least partially defined by a percentile threshold, such as a twenty' fifth percentile threshold, a fiftieth percentile threshold, a seventy fifth percentile threshold, etc.

[0184] For example, the blood cell population includes a population of white blood cells, and the region corresponds to a downshift region associated with the expected change. As such an example, the region can indicate an expected decrease in the value indicative of the size of the population of the white blood cells.

[0185] In response to determining the region of the interpretability' map, the system can select the patient for workup or treatment for a pathophysiological state and / or a risk to develop the pathophysiological state (608). As described above with reference to FIG. 1, the pathophysiological state includes at least one of sepsis, respiratory’ distress, coronary disease, stroke, heart failure, or diabetes.

[0186] For example, for the downshift region, the system can select the patient for a workup or treatment for latent infection or anemia, or for a risk to develop an infection or anemia.Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0187] For example, for a prediction of increasing white blood cell population size, latent infection can be predicted and further tests, diagnosis, and treatments can still be ordered / implemented, such as enhanced surveillance, follow-up microbiological testing, or utilization of empiric antibiotic treatment As an example, for a prediction of increasing white blood cell population size, latent anemia, or a ferritin deficient state) can be predicted and further tests, diagnosis, and treatments can still be ordered / implemented, such as iron studies (e.g., ferritin, total iron binding capacity, transferrin saturation).

[0188] This specification uses the term “configured’’ in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions. Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.

[0189] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g.. an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates anAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0190] execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0191] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.

[0192] In this specification, the term “database” is used broadly to refer to any collection of data: the data does not need to be structured in any particular w ay, or structured at all, and it can be stored on storage devices in one or more locations. Thus, for example, the index database can include multiple collections of data, each of which may be organized and accessed differently.

[0193] Similarly, in this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.

[0194] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0195] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices.

[0196] Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.

[0197] Computer readable media suitable for storing computer program instructions and data include all forms of non volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks.

[0198] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0199] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and compute-intensive parts of machine learning training or production, i.e., inference, workloads.

[0200] Machine learning models can be implemented and deployed using a machine learning framework, e.g., aTensorFlow framework, a Microsoft Cognitive Toolkit framework, an Apache Singa framework, or an Apache MXNet framework.

[0201] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middlew are, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.

[0202] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.

[0203] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, oneAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0204] or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.

[0205] Similarly, while operations are corresponded to in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0206] EXAMPLES

[0207] The following examples further describe examples related to using transformerbased deep learning models for (i) predicting blood cell attributes of a blood cell population of a patient and (ii) predicting contribution values of each blood cell of one or more singlecell distributions used to predict the blood cell attributes. The following examples further describe other aspects of the diagnostic processes and systems described in this disclosure and do not limit the scope of the invention described in the claims.

[0208] Methods

[0209] The following methods were used in the Examples below.

[0210] Data collection

[0211] Data was collected at Massachusetts General Hospital between 2006 and 2012 for 402,490 individuals and ~3 million CBCs. Only one laboratory test per individual was considered, selected at random. Clinical information was obtained retrospectively from medical records. The Mass General Brigham Institutional Review Board approved the study and w aived the requirement for informed consent based on a determination of minimal risk. There were no exclusion criteria, and all available individuals were included in the analysis.

[0212] CBCs were measured on Siemens Advia2120i blood analyzers which utilize flow cytometric principles to measure light scatter properties of individual cells. Three single-cell distributions are generated: one that includes RBCs and PLTs, one that includes basophilsAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0213] and nuclei from all other WBCs (PWBC-BASOS), and one that includes all WBCs and a cytoplasmic peroxidase stain (PWBC PEROX). For the RBC and PLT distribution, two versions were considered: the original distribution containing PLT measurements PPLT+RBC), and a version processed to remove PLTs and project remaining measurements using the Mie Transform to produce measurements of single RBC volume and hemoglobin content (PRBC). The PLT distribution was not chosen to be separated because it is difficult to reliably distinguish platelets from debris and RBC sediment. As a result, some information on the N z,r / NflBc ratio was available in PPLT+RBC. The potential impact of this information was minimal because the ratio itself explained only a small amount of the variance in population sizes (~5%) with the exception of Np / j for which it explained about 20%. The MIST ablation analysis also suggests the ratio information contained in PPLT+RBC had a minimal effect on the prediction of NHCT, NHGB, NRBC and NWBCC because the predictions using all distributions other than PPLT+RBC have accuracy comparable to that obtained using all four single-cell distributions. Additional details on CBC data are available in the below. Each of the generated distributions is two-dimensional and contains up to 50,000 cells (visualization in Figure 7C). Each of these distributions was subsampled to 1,000 samples, removing any information about counts and making model training computationally manageable with deep learning on a single 24GB consumer GPU, while retaining cell population statistics with minimal additional variance. Each distribution was then normalized by the mean and standard deviation across the population. The Advia also provided the CBC indices which were used as dependent variables: hematocrit NHCT), hemoglobin (NHGB), RBC count NRBC), PLT count (NPLT, and WBC count (NWBCC). During the period considered, three different Advia instruments were used in the clinical laboratory, and information on analyzer identity was collected from raw data associated with the CBC laboratory test.

[0214] Multi-input Set Transformer ++ (MIST) for single-cell CBC data

[0215] A key property of single-cell data that differentiates it from tabular, image, or sequential data types is the lack of information in the ordering of cellular measurements. This property can be encoded computationally by models that respect permutation invariance, i.e., any reshuffling of the input produces the same output. Each input distribution (PPLT+RBC, PRBC, PWBC-BASOS, and PWBC_PEROX) is represented as matrix

[0216] X ∈ RN × 2where N is the number of cells. For each individual i G [1,..., M], the full data input is represented by

[0217]

[0218] A" = XABC+PLT<XWBCBASOS’XWBCPER0X'XCOV } whereAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0219] Xicov∈ Rdis a vector of covariates containing age, year of measurement, and the identity of the blood analyzer used for the test.

[0220] MIST respects permutation invariance within each distribution of cells per individual, meaning re-ordering of the input cells does not affect the output prediction. Specifically, let H be the dimension of the latent space, φ: RN × 2→ RN × Hbe a permutation equivariant encoder, meaning any reordering of cells for a given input results in the corresponding ordering of the output, and φ: RN × H→ RHbe a permutation -invariant aggregation function that maps each cell ty pe to a latent vector. Then, each input single-cell distribution is encoded as

[0221]

[0222] φt(φt(Xt)) ∈ RHfor t ∈ {RBC, PLT +

[0223] RBC, WBCBASOS, WBCPEROX}. The encoded distributions in the latent space can be used optionally alongside the additional covariate vector Xcov. The concatenated vector is then passed into a decoder function ρ: R(4×H+d)-> that produces the output. Overall, the multi-input single-cell deep architecture of MIST can be defined as:

[0224] /

[0225]

[0226] (X; θ) = ρ([{φt(φt(Xt)) | t = [RBC, PLT + RBC, WBCBASOS, WBCPEROX}; Xcov])

[0227] where θ denotes parameters of the overall architecture. The values φ, φ, and ρ were instantiated to follow the Set Transformer++ architecture.

[0228] Three datasets were created with similar demographics of N=268326 training samples and N=134164 test samples such that the test splits of each dataset are nonoverlapping. The validation set was selected to be a random 10% subset of the training set. For each dependent variable (e.g., NHCT, NWBC), a separate model was trained for each split, for a total of 15 models, or three models per output. The models use a hidden dimension H=128, and 8 transformer layers: four transformer layers in the encoder per distribution via two Induced Set Attention Blocks (ISAB++), each consisting of two Multihead Attention Blocks (MAB++); one Pooling by Multihead Attention block (PMA++) for the aggregation; and three Self Attention Blocks (SAB++) for the decoder. The models were trained for 30 epochs with a batch size of 64, with the final model chosen based on validation loss. Models were trained on an NVIDIA TITAN RTX 24GB or Tesla V100-PCIE-32GB GPU. To understand the role of demographic and machine information on the prediction task, MIST was trained both with and without the additional covariates vector. A gradient-boosted tree (GBT) model was also trained to predict cell population sizes using this covariates vector alone. The MIST model performed similarly with and without the additional covariatesAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0229] vector, suggesting that the additional covariates do not add additional information beyond that contained in the single-cell data. Consistent with this hypothesis, a population sizes were predicted using only this vector and found performance only marginally better than that of a constant baseline prediction and significantly worse than prediction with the singlecell data. Given the lack of additional information provided by the available covariates, the main results in the text are presented for a MIST model without them.

[0230] Prediction with standard clinical CBC parameters

[0231] The CBC laboratory test typically produces ten CBC parameters. In addition to the five populations sizes used as main outputs in this study, the CBC also provides mean corpuscular volume (MCV), mean corpuscular hemoglobin (MCH). red cell distribution width (RDW), mean corpuscular hemoglobin concentration (MCHC), and mean platelet volume (MPV). To evaluate the efficacy of single-cell data compared to the existing ten CBC parameters, a GBT model was trained to predict the five populations using every other CBC parameter. Due to the correlation between NHCT, NHGB, and NRBC, these population sizes were not used as input when predicting any of these indices as output: NPLT, NWBCC, MCV, MCH, RDW, MCHC, and MPV were used as input for prediction of NHCT, NHGB, and NRBC, NHC, NHGB, and NRBC, NWBCC, MCV, MCH, RDW, MCHC, and MPV for the prediction of NPLT; and NHCT, NHGB, and NRBC, NPLT, MCV, MCH. RDW, MCHC, and MPV for the prediction of NWBCC. AGBT model was used with ten estimators and sweep hyperparameters using grid search cross-validation to select learning rate (0.001, 0.01, 0.1) and minimum number of samples allowed for a split (2, 4, 8).

[0232] MIST comparison with gradient-boosted trees and Deep Sets+ +

[0233] To assess the efficacy of a transformer-based pipeline on single cell data, MIST was compared with two baselines. The first was a multi-input variation of Deep Sets++ using the same single-cell data used as input for MIST. The Deep Sets++ model consisted of a 2-layer multi-layer perceptron (MLP) encoder, a sum aggregation, and a 3-layer MLP decoder, all with He-style residual connections and set norm. This model was used to evaluate the utility of a transformer-based architecture. The second comparison was made with a GBT model trained on hand-engineered statistics from the single-cell distributions. This comparison was carried out to determine whether simpler methods on easily interpretable hand-engineered features capture as much information as the complex pipeline. The GBT model included up to the fourth-order moment (mean, standard deviation, skew, kurtosis) and deciles of eachAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0234] distribution along both dimensions (i.e., cell size and hemoglobin content). A grid search cross validation was performed over the number of trees (100 or 200), learning rate (.001, .01, .1), and minimum number of samples allowed for a split (2, 4, 8).

[0235] Comparison with measurement noise

[0236] Measurement noise denoted as a coefficient of variation (CV) was collected from available literature. For comparison with population-wide RMSE the measurement noise at the average value of each population size in the analyzed cohort was calculated. For example, in the case of RCT the CV was 1.8%, and measurement noise was calculated as 1.8 × |NHCT|

[0237] - - - where i denotes a different individual. To assess how many predictions fell within measurement noise, the calculation was performed at the individual level. Therefore, if NHCTis the measured hematocrit level for the individual i and N̂HCTis MIST prediction, the prediction was defined to be within measurement error if \N_HCT*i _i — N_HCT*i | < (1.8 X V_HCTAz) / 100.

[0238] Single-cell Shapley values

[0239] To evaluate the role of each cell in the predictions, their Shapley values were estimated. Shapley values are defined based on a particular choice of value function dictating the value or importance of an input subset, and Shapley values attribute the contribution to the overall value across different members of the input set (in this case, cells) according to the desirable theoretical properties of efficiency, symmetry’, and additivity.

[0240] The Shapley value φc(v) of a cell c depends on the choice of value function v(c) and is calculated as the marginal contribution of the cell c averaged over all possible subsets of cells. The marginal contribution is computed by comparing the value of a randomly sampled subset s with and without cell c. The subsets s will have different sizes to account for interactions and complementary’ effects among the cells. Let n ~ Unif(N') denote that integer n is sampled uniformly from 0 to N — 1, and let

[0241] s ~ Unif(PX\{c}(n)) denote that subset s is sampled from a uniform distribution over cell subsets of size n that do not include cell c. Then, the Shapley value pcis defined formally as

[0242] φ_c(v) = E_(n ~ Unif(N)) E_(s ~ Unif(PX\{c}(n))) [v(s ∪ {c}) − v(s)]Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0243] Our choice of value function was the conditional expectation (CE) of the output given the subset as input: vCE(s) = E[p(y|s)]. Given this value function, a positive Shapley value for a cell means that the inclusion of the cell increases the predicted output on average over all possible input cell subsets. On the other hand, a negative Shapley value means that the cell decreases the predicted output on average over all possible input cell subsets.

[0244] In addition to the advantageous properties of Shapley values, single-cell FastShap favors non-encoding explanations, due to the fact that the model sees only the selected cells without information about the selection mask. This design mitigates the risk that the explanation leaks information about the label beyond the information contained in the cells selected by the explanation, meaning that the predictive accuracy measures the quality of the explanation accurately.

[0245] Single-cell FastShap training

[0246] Since the Shapley value for a given cell is an average over all possible subsets, it is prohibitively expensive to compute directly for each cell for each patient. Instead, a singlecell FastShap model was trained to predict the outcome of this computation directly. At a high level, training proceeds as follows (Figure 8):

[0247] 1. A surrogate model was trained to predict the output y from a random subset of cells.

[0248] This model has the same architecture of MIST.

[0249] 2. A single-cell FastShap model was trained following the training objective of the original FastShap, which utilizes the prediction of the surrogate model to compute the FastShap loss. The model architecture matches that of the MIST encoder.

[0250] Surrogate model objective function

[0251] The surrogate model fEwas trained to predict the output given randomly masked inputs via the following objective, following Jethani et al:

[0252] f_E(y | x; β) = (max)_β / Σ_(t=1) E_(r[_0 =

[0253]

[0254] ~ B) [log (y_(i) | f_mask(r_(j,) x_(i, -(j = 1)m)]]]

[0255] where / ? are the parameters of the surrogate model and B is a distribution over input masks chosen to favor contiguous regions rather than uniformly sampled cells. In practice, for aAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0256] given individual, new sets of input distributions were produced by selecting five ellipsoids across all available data types (see Ellipsoids for data augmentation). Then, the surrogate model was trained to predict the output of interest using as input only the cells that fall \within the random ellipsoid regions. In other words, the training process for the surrogate model matched that of the full MIST model except that the surrogate model saw an augmented input dataset containing both the full 1000-cell distributions as well as arbitrary subsets of the single cells randomly sampled. Performance of the surrogate model on the full input was marginally worse than that of the full model but was generally within a 95% confidence interval, implying that the surrogate was a good approximation.

[0257] Single-cell FastShap objective function

[0258] Training the single-cell FastShap model involved four steps (Figure 8): (1) for an individual i and their input single-cell distributions, the FastShap model predicts single-cell Shapley values, one value for each cell and each population size; (2) a random mask k is sampled to produce a subset of cells to pass to the surrogate model to get a predicted output VSURR; (3) the same random mask is applied to the single-cell Shapley values which are then summed to produce an estimate of the predicted output ŷSHAP,i; (4) the loss is computed as the mean squared error between ŷSURRand ŷSHAP- Formally, given a subset of cells s and a boolean mask msfor the subset s, the FastShap model (pfastshap-

[0259]

[0260] with parameters rj is trained via the following objective:

[0261] φ_fastshap(X; η) = argmin_(v'(s)) E_p(X) E_p(s) [-1 / v_X(s) −

[0262]

[0263] msTη) − v_X(0))]2,

[0264] where p(X) is the distribution of inputs in the training data and p(s) is the Shapley kernel p(s) ∝dN−1. This approach takes advantage of the weighted least squares

[0265]

[0266] characterization of the Shapley values by directly optimizing for the model to predict the value of an input subset, v (s), minus the value of the empty set as input, vx(0). It was ensured that the Shapley values summed to the correct total for each individual by forcing the single-cell FastShap model's predictions to satisfy the efficiency constraint 1Tφ(X; η) = vx(X) − vx(0) via additive efficient normalization. In practice, duringAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0267] inference, for each cell

[0268]

[0269] − vx(0) − 1Tφ(X; η)) was added to the model's Shapley value prediction.

[0270] Two different training approaches were considered for the single-cell model: the first default approach samples a subset s for each example from the Shapley kernel and computes the FastShap objective for each; the second paired-sampling approach samples a subset s for a given cell and then computes the FastShap objective for both s and its complement and averages the result. The latter has been shown to reduce gradient variance which can improve optimization efficacy. The results from the paired sampling approach in the main paper and the results from default nonpaired approach were reported. Compared to simply retraining a separate model for different ablations of entire distributions, this interpretability pipeline offers more fine-grained insights with a comparable amount of total computational effort.

[0271] To test the robustness of the single-cell FastShap interpretability procedure, the inclusion plots were generated and compared the interpretability maps for both sampling schemes (default and paired), noting that results look similar (Figures 9A and 9B). It was verified that the single-cell FastShap model outperformed a baseline model that assumes all cells for all individuals are equally important for prediction, and that it was significantly better at matching the subset-based marker outputs.

[0272] Interpretability maps

[0273] To visualize the single-cell Shapley values at a population level. 3000 individuals for each of the predicted markers were randomly sampled: 1000 whose results were below the reference interval, 1000 falling within, and 1000 above. The Shapley values for their input distributions were computed and considered both their signed and absolute values. A 100 x 100 mesh for each individual was computed and normalized single-cell Shapley values in each interval [-100, 100] for signed analysis and [0, 100] for absolute analysis (Figures 10-15). Normalization was performed across all input distributions, and it was necessary to account for differences in population sizes (NHCT, NHGB, NRBC, NWBCC, or NPLT) across individuals. Within each individual, the Shapley values for the cells falling within each mesh boundary were averaged and computed mean and standard deviation at the mesh level for visualization. The parts of the mesh that had more than 80% missing data (i.e., more than 80% of the individuals considered had no cells falling in that 2D space) were given a value of zero and not shown.Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0274] Down Shift definition

[0275] Down Shift was discovered through visual inspection of the interpretability maps in 15A and 15B for the prediction of NHCT, NHGB, and NRBC. Given the split behavior on the left side of the PHBC-BASOS distribution, it was decided to isolate the left-most subpopulation of cells in that distribution. A mixture of Gaussians with 2 components was fit using full covariance matrices with le-6 regularization added to the diagonals to ensure matrices were positive semi -definite. Expectation-Maximization was then performed on five separate random initializations and chose the best result based on likelihood. The y-coordinate of these left centroids across the population of individuals was collected. Individuals with ay-coordinate below the 25thpercentile across the population were assigned a positive Down Shift. The alternative of simply considering the y-axis mean was evaluated, eliminating the need for clustering, and results were consistent. The centroid-based approach was proceeded with because it is more consistent with the biological hypothesis generated by the Shapley value maps (see Difference between Left Shift and Down Shift below).

[0276] Down Shift association analysis

[0277] The clinical utility of Down Shift was evaluated compared to Left Shift alone, as well as that of the combination of both Left Shift and Down Shift. Left Shift is defined by the Advia as 0 if absent or as 1,2.3 depending on the strength. In the analysis LS was considered to be a binary variable where any non-zero strength was considered a positive Left Shift. The association of these three exposures was analyzed with two markers of inflammation, Erythrocyte Sedimentation Rate (ESR) and C-Reactive Protein (CRP), as well as future diagnoses. The analysis was performed with a logistic regression model adjusted for age, sex, and NWBCC. All the diagnoses (26) for the individuals analyzed from their Electronic Health Records was collected and mapped them to PheCodes, considering only those diagnoses that were made in the 30-day period after the CBC measurement, had not been made in the year prior to the CBC, and had at least 1% prevalence in the examined cohort. After observing multiple diagnoses of signs related to sepsis (respiratory’ failure, fever, pulmonary congestion, pleural effusion, etc.), a diagnosis of sepsis was included in the analysis even though its prevalence was less than 1%. In total 27 diagnoses were tested. Significance was evaluated at an alpha of 0.05 with Bonferroni adjustment for multiple testing (p=0.0056 for ESR and CRP and p=0.0002 for diagnosis).Attorney Docket No. 29539-0868WO1 / MGH 2025-273

[0278] Single-cell complete blood count data

[0279] The single-cell complete blood count data collected by the Siemens Advia 2120i instruments1consists of optical intensity signals measured flow cytometrically under different staining and solution conditions. Forward light scattering provides an estimate of the relative size of individual cells by measuring the angular distribution of the light scattered by the cell, which depends not only on the volume of the particle but also on its shape, orientation, and refractive-index distribution. The forward scatter (FSC) measurement reflects scatter along the path of the laser, and the side scatter (SSC) measures scatter at a ninety-degree or other angle relative to the laser.

[0280] FSC intensity is proportional to the cell's diameter and is primarily determined by light diffraction around the cell. The FSC signal is often used for differentiating cells by size. SSC, on the other hand, is determined by the light refracted or reflected at the interface betw een the laser and intracellular structures, such as granules and the nucleus, and provides information about the internal complexity (i.e. granularity’) of a cell. Depending on the cell staining, incubation, or other treatment conditions prior to optical analysis, SSC will measure different characteristics, for instance as reported by the Siemens Advia BASOS and PEROX methods1detailed below’.

[0281] 1. RBC distribution (PRBC)

[0282] For the red blood cell (RBC) input distribution, a transformed version of the raw’ data generated by the Advia instrument w’as used. The Mie Transform maps the raw FSC and SSC to measurements of volume (fL) and hemoglobin content (pg), which allows us to reason about this distribution in units that are used in clinical practice. We refer to Tycko et al. for a more in-depth discussion on how the raw’ data can be transformed.

[0283] 2. PLT+RBC distribution (PPLT+RBC)

[0284] The PLT+RBC distribution consists of the raw FSC and SSC measurements of red blood cells and platelets. Clean separation of PLTs is not possible because of the presence of debris and RBC sediments.

[0285] 3. BASOS distribution (PWBC BASOS)

[0286] The white blood cell basophil distribution (BASOS) is raw scatter light intensity’ data that measures the size and lobularity of white blood cell nuclei. For this measurement, allAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0287] white blood cells except basophils are stripped of their membrane and cytoplasm and are then categorized as more mononuclear or more polymorphonuclear based on the shape and complexity of their nuclei. The basophils can be distinguished from the smaller cell nuclei based on size.

[0288] 4. PEROX distribution. (PWBC PEROX)

[0289] The white blood cell peroxidase method (PEROX) helps identify different types of white blood cells based on the size and intensity of the peroxidase reaction. Neutrophils, eosinophils, and monocytes are stained while lymphocytes and basophils remain unstained.

[0290] The " Left Shift" and " Down Shift” single-cell clinical biomarkers

[0291] In the main analysis the novel Down Shift biomarker was discussed as a complement to the established Left Shift clinical marker. Here more detail on Left Shift was provided and how it differs from Down Shift. Mature neutrophils can be differentiated into polymorphonuclear (PMN) and mononuclear (MN) groups based on their nuclear morphology and how many lobes the nucleus has. The most immature white blood cells are called blast cells. FIG. 16 below shows a schematic of a PWBC BASOS distribution and labels the areas where the three neutrophil types (blasts, MN, PMN) typically appear: green, pink, and orange areas respectively. During an acute inflammatory response, for instance during the early stages of a bacterial infection, a larger number of immature neutrophils will be released from the bone marrow. This change in the morphology associated with the presence of a larger fraction of immature neutrophils w as first observed visually many decades ago and was called a "shift to the left” or “Left Shift.” Modem automated hematology analyzers like the Siemens Advia 2120i quantify the “Left Shift,” as shown below, based on the fact that the presence of a larger fraction of immature neutrophils will reduce the relative distance between the MN / PMN valley and the leftmost (MN) peak (i.e., the distance d in the right panel of FIG.

[0292] 16).

[0293] In contrast, Down Shift captures a shift of probability density or cells in a largely orthogonal direction along the y-axis and is independent of the unimodality or bimodality of they-axis marginal distribution. In practice, the presence in Down Shift likely implies a higher percentage of blasts or immature cells and might therefore be expected to complement Left Shift.

[0294] Single-cell FastShap model vs. baselineAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0295] To test the ability of single-cell FastShap to capture useful signals reported as singlecell Shapley values, model performance was compared to a baseline model that assigns the same Shapley value to all cells for each individual. The baseline single-cell Shapley values from the training set (size N) were calculated by separately sampling a random mask for each individual and computing the best least-squares estimate for a system of equations with N random input subsets (one for each individual) and with N corresponding surrogate outputs for the subset. Results can be found in Table S9. All FastShap models are systematically better than a constant Shapley baseline, consistent with the conclusion that the single-cell FastShap model is learning a more meaningful interpretation than assuming all cells are equally important.

[0296] Ellipsoids for data augmentation

[0297] The surrogate models used to train single-cell FastShap required data augmentation and generation of random masks. To generate the masks, a random number of ellipsoids was sampled and selected the cells falling within the ellipsoid regions. The blood cell distribution in which the ellipsoid fell was also chosen at random, as were the parameters defining the shape itself: the mean was randomly sampled in the interval between the 2.5thand 97.5thquantiles of each distribution along both the x- and y-axes, while variance was sampled in the interval [0.5,0.8] for the PEROX distribution and [1, 2] for the other distributions, with covariance in the interval [0, 0.5] for the PEROX distribution and [0, 1] for other distributions. Note that there is evidence that the tails of the single-cell distributions, which normally have lower density than the center, carry important biological information. Therefore, ellipsoids were chosen in place of random sampling to increase the probability that lower density regions would be selected in the data augmentation process.

[0298] Example 1: A transformer-based general-purpose deep learning pipeline for interpretable inference from single-cell data

[0299] A pipeline composed of two novel transformer-based Al models was developed. The first, MIST, is a prediction model for sets as input that extends state-of-the-art permutationinvariant neural network architectures to accept any number of single-cell distributions. MIST encodes each single-cell distribution separately and then merges these encoded representations to predict the desired target (Figure 2A). This flexible architecture allows the model to learn separate functions for each cell population while also modeling interactionsAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0300] between their encoded representations. The second model in the pipeline, single-cell FastShap (Figure 2B), quantifies the contribution of each single cell to the predicted population size. Single-cell FastShap combines MIST with recent advances in efficient amortized computation of Shapley values. Shapley values provide a formal grounding for interpretability in terms of game theory but are expensive to compute. Single-cell FastShap efficiently estimates Shapley values with a single forward pass of a neural network (see Figure 8). Single-cell Shapley values quantify each cell’s marginal contribution to the average prediction across all combinations of cells. For the choice of a value function, a positive Shapley value indicates that inclusion of the cell increases the predicted population size on average, while a negative Shapley value indicates a decrease.

[0301] Example 2: CBC measurements of cell population sizes and single-cell data

[0302] CBCs were measured on Siemens Advia 2120i automated hematology analyzers for 402,490 individuals treated at Massachusetts General Hospital between 2006-2012 with no exclusion criteria. CBCs measure the sizes of the RBC, WBC, and PLT populations in a microliter of blood, with population size quantified as simple cell counts (NRBC, NWBC, and NPLT), and the RBC population size also quantified in terms of the fraction of volume it occupies (NHCT) and the total hemoglobin mass it contains NHGB). On the same blood sample, CBCs also measure optical scatter and fluorescence properties for >50,000 individual RBCs, WBCs, and PLTs. The single-cell data consists of four two-dimensional distributions measured on subsets of cells under different conditions that include stains or surfactants: PPLT+RBC contains both RBCs and PLTs, PRBC contains only RBCs, and PWBC_BASOS and PWBC_PEROX contain only WBCs (Figure 7C,). For analysis, 1,000 cells from each distribution were randomly subsampled, eliminating any information on population size from the singlecell data. This approach also eliminates any information on the relative sizes of NWBCCINRBC and NWBC / NPLT. PPLT+RBC contains some information on relative sizes of NRBC and NPLT because it contains both PLTs and RBCs in proportions that are associated with their relative proportions in the blood sample.

[0303] Example 3: MIST explained 70%-82% of the variance in cell population size

[0304] Five different MIST models were trained to predict the cell populations sizes NHCT, NHGB, NRBC, NWBCC, or NPLT) given all available single-cell data PRBC+PLT, PRBC, PWBC_BASOS, and PWBC PEROX). Figure 17 shows that MIST explained at least 70% and up to 82% of the baseline in population sizes among individuals. Standard automated hematology analyzersAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0305] provide precise measurements but do have some measurement noise. MIST predictions were within measurement noise for 40% of NHCT. 30% of NHGB, 63% of NRBC, 64% of NPLT, and 27% of NWBCC MIST’s root mean squared error (RMSE) was 1.9x measurement noise for NHCT, 2.7x for NHGB., 1.1x for NRBC, 1.1x for NPLT, and 3.5x for NWBCC. RMSE was lower for counts within or below standard clinical reference intervals (Figure 18A) and showed slight differences based on reported sex and age (Figure 18A, 18B, 18C). However, the addition of sex and age as covariates did not significantly improve prediction accuracy. MIST’s high accuracy provides evidence that the single-cell data contains rich information on cell population sizes that may yield insights into homeostatic processes.

[0306] Example 4: Single-cell data had 15x more information on WBC population size than standard CBC parameters

[0307] One reasonable null hypothesis for MIST’s accuracy is that the simple statistics of single-cell measurements (e.g., their average) are associated with cell population size. It was tried to predict NWBCC. NHCT, NHGB, NRBC, and NPLT from standard CBC parameters: the mean volumes of RBCs and PLTs. the mean hemoglobin mass and hemoglobin concentration of RBCs, and the other population sizes (e g., NRBC and NPLT for the prediction of NWBCC). Gradient-boosted tree (GBT) models trained on these parameters were able to explain less than 5% of the variance compared to 70% for MIST for the prediction of NWBCC. GBT models explained only 14-19% of the variance in NHCT, NHGB. NRBC, or NPLT compared to 70%-82% for MIST (see Figure 17,). Next, the inputs of the gradient-boosted tree models were expanded to include the first four moments (mean, variance, skewness, kurtosis) of each of the eight single-cell marginal distributions, as well as all 10 distribution deciles along each of the eight single-cell dimensions, providing 112 inputs total. Performance improved but still failed to explain at least 10% of the variance in population sizes explained by MIST. Adding the PLT / RBC ratio to simulate the possible availability of information in the PPLT+RBC distribution on that ratio explained only a small additional amount of variance (~5%) for all population sizes with the exception of NPLT where it explained 20% of the variance. The effect of the transformer-based architecture was investigated by comparing MIST to a nontransformer deep learning model for sets. Deep Sets++, and found that MIST explained significantly more variance except for NWBCC where the methods were comparable. Even when population sizes for the other cell types were added as inputs, along with other standard reported CBC parameters, only an additional ~4% of the variance was explained compared to single-cell input only. These results imply that single-cell data contains a substantial amountAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0308] of information on the dynamic processes regulating cell population size and that transformerbased methods can capture more of this information than other simpler approaches.

[0309] Example 5: Single-cell data contained cross-population information on size

[0310] Prior studies have identified evidence for correlation among NRBC, NWBCC, and NPLT both at steady state and in response to multiple disease processes. It was therefore investigated whether single-cell data from one cell population (e.g., PRBC was being used by MIST to infer the sizes of other cell populations (e.g., NWBCC. Ablation analysis comparing MIST’s prediction accuracy using all permutations of single-cell distributions as input was performed (Figure 15A). Accuracy was generally lowest for predictions that used only one single-cell distribution at a time and typically improved as other single-cell distributions were added. An exception to this pattern was NPLT prediction performance where PPLT+RBC on its own yielded significantly greater accuracy than the other three single-cell distributions combined, possibly reflecting the effect of information on the NPLT / NRBC ratio available in that distribution. This evidence for crosstalk and the greater accuracy of MIST compared to simpler methods together suggest that models limited to one single-cell distribution or to population-level information will fail to leverage significant predictive information available and will underperform. Therefore single-cell FastShap was applied to help interpret the relationships between all single-cell data and each cell population size in physiologic terms (Figure 15B).

[0311] Example 6: WBC population size was most strongly associated with single-WBC and single-PLT data

[0312] Ablation analysis showed that PWBC PEROX contained the most information about NWBCC. Adding PPLT+RBC as input next improved accuracy the most, suggesting that single-PLT data, single-RBC data, or both provided significant complementary information about NWBCC, and since the performance of PWBC PEROX combined with PRBC was not significantly different from that of PWBC PEROX alone, it is most likely that single-PLT data helps predict NWBCC (Figure 15A). Consistent with this ablation analysis, single-cell FastShap found large positive Shapley values in PWBC PEROX, particularly in the top-left boundary region which is dominated by lymphocytes, in contrast to the top-right boundary region which had moderately negative values and is dominated by neutrophils7(Figure 15B, Figure 10). These results suggest that patients with higher NWBCC in this study cohort had more lymphocytes and fewer neutrophils. Negative Shapley values were found in PWBC-BASOS, with stronger values in theAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0313] left side of the distribution. PPLT+RBC had large negative Shapley values in the bottom left region, suggesting a negative correlation between NWBCC and the sub-region in PPLT+RBC specific to platelets. The Shapley values for PRBC showed areas with moderate influence on NWBCC in both directions, with positive Shapley values near the top right of PRBC and the lower left tail. Overall, this analysis finds evidence that changes to the compositions of the single-WBC and PLT distributions are associated with changes to NWBCC and suggests that the composition of the single-RBC distribution is only modestly linked to NWBCC.

[0314] Example 7: PLT population size was associated with single-cell data for PLTs but not WBCs

[0315] Ablation analysis showed that PPLT+RBC on its own enabled accurate prediction of NPLT with an RMSE similar to that achieved with all inputs, in contrast to PRBC, PWBC-BASOS, and PWBC PEROX which had negligible effects on their own and when combined with PPLT+RBC (Figure 15A). Single-cell FastShap analysis was consistent with the ablation analysis, showing high Shapley values concentrated in the extreme bottom left region of PPLT+RBC where PLTs are located and low signal in PWBC-BASOS and negligible signal in PRBC and PWBC PEROX (Figure 15B, Figure 11). These results may in part reflect information on the ratio between NPLT and N RBC contained in the PPLT+RBC distribution. It is therefore difficult to make reliable inferences about biological mechanisms that relate changes in the single-PLT distribution to NPLT regulation, but this single-cell FastShap analysis supports the hypothesis that changes to compositions of single-WBC distributions are not systematically associated with changes in NPLT.

[0316] Example 8: RBC population size was associated with all single-cell distributions Ablation analysis showed generally steady improvement in MIST’s accuracy estimating the RBC population sizes (NHCT, NHGB, NRBC) as additional single-cell distributions were included as inputs (Figure 15A). The PWBC PEROX distribution had the smallest effect on accuracy, and when excluded, the residual accuracy was close to that for all inputs. Singlecell Fastshap interpretability maps were consistent and provide further detail on potential relationships (Figure 15B, Figures 12-14). The PLT-specific sub-region of PPLT+RBC had negative Shapley values for NHCT, NHGB, and NRBC, and the left side of PWBC-BASOS had large positive Shapley values at the top (top left) and moderately negative at the bottom (bottom left). The Shapley values for PRBC were consistently moderately positive in the top right where reticulocytes appear. It is well-established that the presence of more reticulocytes isAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0317] associated with a faster RBC birth rate, and NHCT, NHGB, and NRBC would then be expected to increase if mean RBC lifespan and other cellular characteristics remain stable. Different Shapley value variation in this reticulocyte region for the three measures of RBC population size was observed, suggesting that RBC volume fraction, hemoglobin mass, and count may be regulated differently in some settings. Single-cell FastShap analysis thus finds evidence that all single-cell distributions are systematically associated with changes in NHCT, NHGB, and NRBC, implying that RBC, WBC. and PLT population compositions are altered in the course of regulating RBC population size.

[0318] Example 9: Rational discovery of single-cell markers of cellular kinetics

[0319] By identifying regions of single-cell distributions associated with changes in population sizes, single-cell FastShap provides data-driven hypotheses for single-cell markers of cell population kinetics. Indeed, single-cell FastShap found large negative Shapley values for prediction of NWBCC in the left-side region of PWBC-BASOS which is used to define the clinical "‘Left Shift” (LS) marker that is already known to be associated with altered WBC kinetics. Single-cell FastShap also found a connection between this region of PWBC-BASOS and NHCT, NHGB, and NRBC, with large positive and negative Shapley values in close proximity. The juxtaposition of regions with positive and negative Shapley values suggests that a shift of white blood cells from the top left of PWBC-BASOS distribution to the bottom left may be particularly strongly associated with a decrease in the size of NHCT, NHGB. and NRBC. This hypothesis was tested by deriving “Down Shift” (DS), a binary marker that was defined to be positive when the y-coordinate of the centroid for the leftmost portion of PWBC-BASOS was below the 25thpercentile (Figure 19). Figure 19 shows that in the study cohort a positive DS was associated with a significant decrease in NHCT, -1.98 % (as well a significant decrease in NHGB of -0.67 g / dL and in NRBC of -0.20 103 / μL) and a significantly higher average NWBCC. The presence of LS was associated with a smaller but significant decrease in NHCT, -1.54 %, and an increase in NWBC, 2.54 103 / μL. The existing clinical LS flag is sensitive to changes in WBC dynamics, and because DS appeared to be more sensitive to RBC dynamics despite being a single-WBC marker, it was hypothesized that DS might complement LS in some clinical situations where changes in both WBC and RBC populations occur.

[0320] Example 10: Down Shift complemented Left Shift to improve risk stratification for multiple major diseasesAttorney Docket No. 29539-0868WO1 / MGH 2025-273

[0321] Diagnostic associations for LS, DS, and their combination (LS+DS) were analyzed in the study cohort. Individuals were divided into four overlapping groups based on whether they had neither, LS, DS, or both (LS+DS). DS with and without LS was associated with a decrease in NHCT and an increase in NWBCC (Figure 20A). Their association with standard clinical markers of inflammation were tested: erythrocyte sedimentation rate (ESR) and C-reactive protein (CRP). Figure 20B shows that both LS and DS were significantly associated with elevated ESR, with DS and LS+DS having higher odds ratios. LS and DS were also significantly associated with elevated CRP, with DS having a slightly higher odds ratio. ESR tests available for the cohort were elevated 54% of the time overall, and the positive predictive value (PPV) for elevated ESR increased to 65%, 68% and 72% based on the presence of LS, DS. and LS+DS, with a higher sensitivity for DS of 44% compared to 17% for LS and 11% for LS+DS. CRP tests are often ordered for cardiovascular disease risk screening in healthy patients and therefore had a lower rate of positivity in the study cohort than ESR, 15%. PPVs for elevated CRP were 18%, 18% and 19%, and sensitivity followed a pattern similar to that for ESR, with a 44% sensitivity for DS compared to 19% for LS and 11% for LS+DS. The associations of LS, DS, and LS+DS with future diagnoses was investigated. Figure 20C shows that LS and DS were both significantly associated with the appearance of several new diagnoses during the 30-day period after the CBC measurement, including sepsis, respiratory distress, coronary disease, stroke, heart failure, and diabetes. After adjusting for age, sex. and NWBCC, LS+DS was more strongly associated with many diagnoses than either LS or DS alone, suggesting that DS complements LS in providing an early signal of the presence of inflammatory and other pathologic processes

[0322] OTHER EMBODIMENTS

[0323] It is to be understood that while the invention has been described in conjunction with the detailed description thereof, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.

Claims

Attorney Docket No. 29539-0868WO1 / MGH 2025-273WHAT IS CLAIMED IS:

1. A method comprising:receiving data representing a plurality of single-cell distributions, wherein each of the plurality of single-cell distributions represents a corresponding plurality of blood cell measurements for one or more corresponding blood cell types of a plurality of blood cell types, the plurality of blood cell measurements collected using a complete blood count (CBC) test performed on a blood sample of a patient;generating, using a permutation invariant first neural network and the data representing the plurality of single-cell distributions, an encoded distribution for each of the plurality of single-cell distributions: andpredicting, using the respective encoded distributions, a value indicative of a blood cell attribute of a blood cell population of the patient.

2. The method of claim 1, wherein the blood cell attribute comprises a size of a blood cell population of the patient.

3. The method of claim 1, wherein the blood cell attribute comprises a blood cell measurement for the one or more corresponding blood cell types not provided as input.

4. The method of claim 1, wherein the blood cell attribute comprises a corresponding plurality of blood cell measurements for one or more corresponding blood cell types, the plurality of blood cell measurements collected using a second complete blood count (CBC) test performed on a second blood sample of a patient taken after the first CBC test.

5. The method of any one of the preceding claims, wherein the plurality of single-cell distributions comprise: (i) a first single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of red blood cells, (ii) a second single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of red blood cells and platelets, (iii) a third single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of nuclei of white blood cells, and (iv) a fourth single-cell distribution inAttorney Docket No. 29539-0868WO1 / MGH 2025-273which the corresponding plurality of blood cell measurements are measurements of a peroxidase reaction in white blood cells.

6. The method of claim 5, wherein the blood cell population is: (i) a red blood cell population of the patient, (ii) a white blood cell population of the patient, or (iii) a platelet population of the patient.

7. The method of claim 6, wherein the blood cell population is the red blood cell population of the patient, and the value comprises a hematocrit of the red blood cell population or a total hemoglobin mass of the red blood cell population.

8. The method of any one of the preceding claims, wherein generating the encoded distribution for each of the plurality7of single-cell distributions comprises:generating, by applying the permutation invariant neural network, a plurality7of encoded vectors for each of the plurality of single-cell distributions, the plurality of encoded vectors characterizing a plurality of blood cells of a corresponding single cell-distribution of the plurality of single-cell distributions.

9. The method of claim 8, wherein generating the plurality of encoded vectors comprises for each of the plurality of single-cell distributions comprises:using a permutation equivariant encoder of the permutation invariant neural network.

10. The method of claim 9, wherein generating the plurality of encoded vectors for each of the plurality of single-cell distributions comprises:normalizing the data representing the corresponding single-cell distribution; projecting the normalized data to an embedding space to generate a plurality7of embedding vectors; andprocessing, using one or more attention blocks, the plurality of embedding vectors to generate the plurality of encoded vectors.

11. The method of claim 10, wherein each of the one or more attention blocks comprises one or more multi-head attention blocks.Attorney Docket No. 29539-0868WO1 / MGH 2025-27312. The method of any one of the claims 8-11, wherein generating the encoded distribution for each of the plurality of single-cell distributions comprises:applying a permutation invariant aggregation function to the plurality of encoded vectors for the corresponding single-cell distribution.

13. The method of claim 12, wherein applying the permutation invariant aggregation function to the plurality of encoded vectors for the corresponding single-cell distribution comprises:processing, using one or more pooling attention blocks, the plurality' of encoded vectors.

14. The method of claim 13, wherein the one or more pooling attention blocks comprise one or more multi-head attention blocks.

15. The method of any of claims 8-14, wherein the plurality of encoded vectors comprises a covariate vector.

16. The method of claim 15, wherein the covariate vector comprises covariates containing at least one of: (i) age, (ii) year of measurement, or (iii) identity' of a blood analyzer used for collection of the measurements.

17. The method of any one of the preceding claims, wherein predicting the value indicative of the size of the blood cell population comprises:combining the encoded distribution of each of the plurality of single-cell distributions to generate a combined distribution vector; anddecoding, using a decoder function, the combined distribution vector to generate the predicted value.

18. The method of claim 17, wherein decoding the combined distribution vector comprises:processing, using one or more attention blocks, the combined distribution vector.

19. The method of claim 18, the one or more attention blocks comprise one or more self-attention blocks.Attorney Docket No. 29539-0868WO1 / MGH 2025-27320. The method of any one of the preceding claims, wherein the permutation invariant neural network comprises a Transformer neural network.

21. The method of any one of the preceding claims, wherein the permutation invariant neural network is pretrained on training datasets for the value.

22. A method comprising:receiving data representing a plurality of single-cell distributions, wherein each of the plurality of single-cell distributions represents a plurality of blood cell measurements for one or more corresponding blood cell types of a plurality of blood cell types, the plurality of blood cell measurements collected using a complete blood count (CBC) test performed on a blood sample of a patient;predicting, using a permutation invariant first neural network and respective encoded distributions, a value indicative of a size of a blood cell population of the patient; and generating, for each of the plurality of single-cell distributions and by using a second neural netw ork, a plurality of contribution values indicative of contributions of the plurality' of blood cell measurements to prediction of the value.

23. The method of claim 22, wherein the plurality of single-cell distributions comprise: (i) a first single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of red blood cells, (ii) a second single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of red blood cells and platelets, (iii) a third single-cell distribution in which the corresponding plurality of blood cell measurements are measurements of nuclei of white blood cells, and (iv) a fourth single-cell distribution in which the corresponding plurality’ of blood cell measurements are measurements of a peroxidase reaction in white blood cells.

24. The method of claim 22 or 23, wherein the blood cell population is: (i) a red blood cell population of the patient, (ii) a white blood cell population of the patient, or (iii) a platelet population of the patient.Attorney Docket No. 29539-0868WO1 / MGH 2025-27325. The method of claim 24, wherein the blood cell population is the red blood cell population of the patient, and the value comprises a hematocrit of the red blood cell population or a total hemoglobin mass of the red blood cell population.

26. The method of any one of claims 22-25, wherein the plurality of contribution values comprise predictions of Shapley values.

27. The method of any one of claims 22-26, wherein:the plurality of contribution values comprise:at least one positive contribution value indicating that a first blood cell corresponding to a first blood cell measurement of the plurality of blood cell measurements contributes to an increase in the predicted size, andat least one negative contribution value of the blood cell indicating that a second blood cell corresponding to a second blood cell measurement of the plurality of blood cell measurements contributes to a decrease in the predicted size.

28. The method of any one of claims 22-27, wherein the second neural network is pretrained, using a machine learning model that generates a second value indicative of a size of blood cell populations from a plurality of training single-cell distributions, to generate the plurality of contribution values.

29. The method of claim 28, wherein the machine learning model and the first neural network have the same architecture.

30. The method of claim 28 or 29, wherein the machine learning model is pretrained on a plurality of randomly-masked training single-cell distributions generated by randomly masking the plurality of training single-cell distributions.

31. The method of claim 30, wherein the plurality of randomly -masked singlecell distributions are randomly masked using a plurality of randomly-generated ellipsoids overlaid on the plurality of training single-cell distributions.

32. The method of any one of claims 28-31, wherein the second neural network is pretrained on the plurality of training single-cell distributions to generate the plurality of contribution values, the second neural network being pretrained based on a loss computedAttorney Docket No. 29539-0868WO1 / MGH 2025-273from the second predicted size and a sum of output values from the second neural network using the plurality of training single-cell distributions.

33. A method comprising:receiving data representing (i) a predicted value indicative of a size of a blood cell population of a patient and (ii) a plurality of contribution values indicative of contributions of a plurality of blood cell measurements to prediction of the predicted value, the plurality of blood cell measurements collected using a complete blood count (CBC) test performed on a blood sample of the patient;generating, from the predicted value and the plurality of contribution values, an interpretability map corresponding to the predicted value;determining, from the interpretability map, a region of the interpretability map comprising a subset of contribution values of the plurality of contribution values corresponding to a subset of blood cell measurements of the plurality of blood cell measurements is associated with an expected change in the predicted value; andin response to determining the region of the interpretability map, selecting the patient for workup or treatment for a pathophysiological state and / or a risk to develop the pathophysiological state.

34. The method of claim 33, wherein the expected change comprises either (i) an expected increase in the predicted value, or (ii) an expected decrease in the predicted value.

35. The method of claim 33 or 34, wherein the plurality of blood cell measurements are a plurality of measurements of white blood cells, and the blood cell population is a red blood cell population of the patient.

36. The method of claim 35, wherein the value is indicative of: (i) a size of the red blood cell population, (ii) a hematocrit of the red blood cell population, or (iii) a total hemoglobin mass of the red blood cell population, and the region corresponds to a downshift region associated with the expected change.

37. The method of any one of claims 33-36, wherein the region is at least partially defined by a percentile threshold.Attorney Docket No. 29539-0868WO1 / MGH 2025-27338. The method of claim 36 or 37, wherein the region indicates an expected decrease in the at least one of: (i) a value indicative of the size of the population of the red blood cells, (ii) a hematocrit of the red blood cell population, or (iii) a total hemoglobin mass of the red blood cell population.

39. The method of claim 38, wherein the pathophysiological state comprises at least one of sepsis, respiratory distress, coronary disease, stroke, heart failure, or diabetes.

40. The method of any one of claims 33-39, wherein the interpretability map indicates a plurality of map values computed using a mesh and the plurality of contribution values.

41. The method of claim 40, wherein each of the plurality of map values comprises a mean and a standard deviation for a corresponding cell of the mesh.

42. The method of claim 40, wherein the plurality of contribution values are normalized for signed analysis or absolute analysis.