Automated screening of acid sphingomyelinase deficiency
Through a system implemented on a computer, a machine learning model is used to process the electronic medical record database to generate a classification output of whether the subject has ASMD. This solves the problem of difficulty in effectively screening and diagnosing ASMD in existing technologies, and achieves efficient and accurate diagnosis and treatment.
Patent Information
- Application Number
- CN202480012686.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-15
- Filing Date
- 2024-02-09
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies make it difficult to effectively screen and diagnose acid sphingomyelinase deficiency (ASMD), resulting in many patients with this condition not being diagnosed and treated in a timely manner.
The system, implemented on a computer, processes an electronic medical record database using a machine learning model to generate a classification output indicating whether a subject has ASMD. The system, which includes a feature generation engine and a diagnostic machine learning model, is capable of rapidly screening large populations of subjects.
It has achieved efficient automatic screening of large-scale subject groups, improved the diagnostic accuracy and efficiency of ASMD, and ensured that potential patients can be diagnosed and treated in a timely manner.
Smart Images

Figure CN120641996A_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority to Provisional Application No. 63 / 445,649, filed on February 14, 2023, entitled “AUTOMATED SCREENING FOR ACIDSPHINGOMYELINASE DEFICIENCY,” Provisional Application No. 63 / 519,957, filed on August 16, 2023, entitled “AUTOMATED SCREENING FOR ACIDSPHINGOMYELINASE DEFICIENCY,” and Provisional Application No. 63 / 610,665, filed on December 15, 2023, entitled “AUTOMATED SCREENING FOR ACIDSPHINGOMYELINASE DEFICIENCY,” the contents of which are hereby incorporated by reference. Technical Field
[0002] The present disclosure relates to automated screening of a population of subjects to identify subjects with acid sphingomyelinase deficiency (ASMD). Background Art
[0003] ASMD is a rare genetic disorder caused by a deficiency of the enzyme acid sphingomyelinase (ASM). It is part of a group of disorders known as lysosomal storage disorders. ASMD prevents cells from breaking down sphingomyelin (a type of fat found in cell membranes). As a result, sphingomyelin accumulates in cells and forms lysosomes containing it. These lysosomes eventually rupture and release their contents into the cell, causing damage. ASMD is a multisystem disease that affects various organs and is potentially fatal. Treatment of ASMD may involve enzyme replacement therapy, which involves administering a human recombinant enzyme to replace the deficient ASM enzyme. ASMD can occur in various forms, such as type A ASMD, type A / B ASMD, and type B ASMD. Each type of ASMD can be associated with different rates of progression, age of onset, and severity. Summary of the Invention
[0004] This specification describes a system implemented as a computer program on one or more computers at one or more sites that can screen an electronic medical record database of a population of subjects to identify subjects predicted to have ASMD.
[0005] According to one aspect, a method performed by one or more computers is provided, the method comprising: screening an electronic medical record database of a subject population to identify subjects predicted to have acid sphingomyelinase deficiency, the screening comprising, for each subject in the subject population: obtaining a set of features characterizing the subject from the electronic medical record database; processing a model input comprising the set of features characterizing the subject using a diagnostic machine learning model having a set of diagnostic machine learning model parameters to generate a model output characterizing a likelihood that the subject has acid sphingomyelinase deficiency; and classifying whether the subject has acid sphingomyelinase deficiency based on the model output of the diagnostic machine learning model; and providing an output identifying subjects classified as having acid sphingomyelinase deficiency from the subject population.
[0006] In some embodiments, the diagnostic machine learning model has been trained using a machine learning training technique to determine trained values for the set of diagnostic machine learning model parameters.
[0007] In some embodiments, training the diagnostic machine learning model via the machine learning training technique comprises: obtaining a set of training examples, wherein each training example comprises: (i) a training input comprising a set of features of a training subject, and (ii) a target output based on whether the training subject is designated as having acid sphingomyelinase deficiency; and training the diagnostic machine learning model on the set of training examples.
[0008] In some embodiments, obtaining the set of training examples comprises: obtaining a set of positive training examples, wherein each positive training example corresponds to a subject having acid sphingomyelinase deficiency; and selecting a plurality of negative training examples, each of the negative training examples corresponding to a corresponding subject not having acid sphingomyelinase deficiency, the selecting comprising, for each positive training example: identifying a set of corresponding negative training examples based on matching one or more confounding features in the set of negative training examples with the positive training example; and selecting the set of negative training examples for training the diagnostic machine learning model.
[0009] In some embodiments, the one or more confounding characteristics include demographic characteristics.
[0010] In some embodiments, training the diagnostic machine learning model on the set of training examples includes training the diagnostic machine learning model to, for each training example, process the training input of the training example to generate a model output that matches a target output for the training example.
[0011] In some embodiments, obtaining the set of training examples includes generating, for one or more training examples, a training example from one or more electronic medical records of a training subject corresponding to the training example.
[0012] In some embodiments, obtaining the set of training examples comprises: generating a set of features of the subject using a generative neural network; and generating training examples comprising training inputs based on the set of features generated by the generative neural network.
[0013] In some embodiments, a generative neural network has been trained by operations comprising: generating a set of output features using the generative neural network; processing the set of output features generated by the generative neural network using a discriminator neural network to generate a discriminant score that defines a likelihood that the set of output features was generated by the generative neural network rather than originating from one or more electronic medical records; and training the generative neural network to optimize an objective function that is dependent on the discriminant score generated by the discriminator neural network.
[0014] In some embodiments, generating a set of features for a subject using the generative neural network includes: instantiating initial values of the set of features for the subject as random values sampled from a probability distribution; and denoising the set of features in a series of denoising iterations, the denoising including: processing current values of the set of features using the generative neural network to generate a denoised version of the set of features at each denoising iteration prior to a last denoising iteration in the series of denoising iterations; and providing the denoised version of the set of features for processing at a next denoising iteration; wherein the denoised version of the set of features generated at the last denoising iteration defines the set of features for the subject.
[0015] In some embodiments, the method further includes: generating a confidence score for each subject in the subject population, the confidence score characterizing the confidence of the diagnostic machine learning model in classifying whether the subject has acid sphingomyelinase deficiency; selecting an appropriate subgroup of the subject population for diagnostic testing for acid sphingomyelinase deficiency based on the confidence scores; retraining the diagnostic machine learning model based on the results of the diagnostic testing of the appropriate subgroup of the subject population; and reclassifying whether each subject in the subject population has acid sphingomyelinase deficiency using the retrained diagnostic machine learning model.
[0016] In some embodiments, selecting an appropriate subset of the subject population for diagnostic testing for acid sphingomyelinase deficiency based on the confidence scores comprises selecting subjects from the subject population associated with the lowest confidence scores for diagnostic testing for acid sphingomyelinase deficiency.
[0017] In some embodiments, selecting an appropriate subset of the subject population for diagnostic testing for acid sphingomyelinase deficiency based on the confidence scores includes: selecting multiple subjects associated with the lowest confidence scores from the subjects classified as having acid sphingomyelinase deficiency for diagnostic testing for acid sphingomyelinase deficiency; and selecting multiple subjects associated with the lowest confidence scores from the subjects classified as not having acid sphingomyelinase deficiency for diagnostic testing for acid sphingomyelinase deficiency.
[0018] In some embodiments, generating a confidence score for each subject in the subject population that characterizes the confidence of the diagnostic machine learning model in the classification of whether the subject has acid sphingomyelinase deficiency includes, for each subject in the subject population: generating, using each diagnostic machine learning model in an ensemble of multiple diagnostic machine learning models, a corresponding model output that characterizes the likelihood that the subject has acid sphingomyelinase deficiency; and generating a confidence score based on a variance measure of the model outputs generated by the ensemble of diagnostic machine learning models for the subject.
[0019] In some embodiments, the subject population includes at least 10,000 subjects.
[0020] In some embodiments, screening electronic medical record data of a population of subjects is performed in 1 minute or less.
[0021] In some embodiments, the method further comprises: for one or more subjects classified as having acid sphingomyelinase deficiency, administering to the subject a drug to treat the acid sphingomyelinase deficiency in the subject.
[0022] In some embodiments, the drug comprises oripase alpha or oripase alpha-rpcp.
[0023] In some embodiments, the set of features characterizing the subject includes features of a sample derived from a bodily fluid or bodily tissue of the subject.
[0024] In some embodiments, the diagnostic machine learning model comprises one or more of: a decision tree model, a random forest model, a neural network model, or a support vector machine model.
[0025] In some embodiments, the subject population is pre-screened to include only subjects who have been diagnosed with one or more of: interstitial lung disease (ILD), splenomegaly, or nonalcoholic steatohepatitis (NASH).
[0026] According to another aspect, a system is provided that includes: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the methods described herein.
[0027] According to another aspect, one or more non-transitory computer storage media are provided that store instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the methods described herein.
[0028] Throughout this specification, the term "subject" refers to a human subject, ie, a human.
[0029] Particular embodiments of the subject matter described in this specification can be implemented to realize one or more of the following advantages.
[0030] ASMD is often undiagnosed or misdiagnosed in subjects, for example, due to the rarity of the disease (and thus low disease awareness and knowledge), and because the disease can present with a complex array of symptoms. Consequently, many subjects with ASMD are unaware that they have the disease. To address this issue, the system described herein can automatically screen large patient populations for ASMD based on electronic medical record data to identify patients predicted to have ASMD. ASMD patients identified by this system can have their diagnosis confirmed through diagnostic testing, and the ASMD can be treated, for example, with appropriate medications. Performing diagnostic testing on every patient in a large patient population would be impractical. This system can automatically identify a subset of patients (e.g., less than 1% of the patient population) who are likely to have ASMD, thereby enabling these patients to be precisely targeted for confirmatory diagnostic testing.
[0031] To classify whether a subject has ASMD, the system can extract a set of features characterizing the subject from the subject's electronic medical record and process the set of features using a diagnostic machine learning model to generate a classification output. This classification output can define, for example, a probability that the subject has ASMD, or a hard (e.g., binary) classification predicting whether the subject has ASMD (as described in more detail below). The screening system can train the diagnostic machine learning model on a set of training examples using machine learning training techniques. The diagnostic machine learning model provides a data-driven approach for classifying whether a subject potentially has ASMD based on complex patterns and correlations in electronic medical record data that far exceed what humans, or even the human mind alone, can analyze.
[0032] The system can train the diagnostic machine learning model in a series of iterations to improve its predictive accuracy. At each iteration, the system can generate a corresponding ASMD classification and a corresponding confidence score for each subject in the subject population. The subject's confidence score characterizes the confidence of the diagnostic machine learning model in generating the subject's ASMD classification. The system can then mark appropriate subgroups of the patient population for diagnostic testing based on the confidence score, particularly by selecting subjects with the ASMD classification associated with the lowest confidence score for diagnostic testing. After applying the diagnostic machine learning model and subsequent patient diagnostic testing, the results of the diagnostic test can later be used to refine the system. The system can significantly improve the predictive accuracy of the diagnostic machine learning model by training subjects with ASMD classifications associated with low confidence scores.
[0033] ASMD is a rare disease, predicted to affect only 1 in 250,000 individuals. Consequently, only a very limited amount of training data is available to train diagnostic machine learning models to perform ASMD classification. To address this issue, the system can use generative modeling techniques to generate "synthetic" training examples (e.g., as distinct from "real-world" training examples derived from electronic medical record data of real-world patients) for training diagnostic machine learning models. Augmenting the training data used to train diagnostic machine learning models with synthetic training examples can improve the robustness and predictive accuracy of the diagnostic machine learning models.
[0034] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] FIG1 illustrates an exemplary screening system.
[0036] FIG2 illustrates an exemplary training system.
[0037] FIG3 illustrates an exemplary generative system.
[0038] Figure 4A shows an example of a diagnostic machine learning model.
[0039] FIG4B shows a table of experimental results of applying the diagnostic machine learning model of FIG4A to a subject population.
[0040] FIG4C provides a bar graph showing the relative prevalence of certain features in: (i) ASMD subjects, and (ii) control (non-ASMD) subjects.
[0041] Figure 4D shows an example of a diagnostic machine learning model that has been trained to screen for ASMD in a patient population with unexplained interstitial lung disease (ILD).
[0042] FIG4E shows a table of experimental results of applying the diagnostic machine learning model of FIG4D to a population of subjects with unexplained ILD.
[0043] 5 is a flow chart of an exemplary process for screening an electronic medical record database of a population of subjects to identify subjects predicted to have ASMD.
[0044] FIG6 illustrates an exemplary process for training a diagnostic machine learning model to identify subjects with ASMD and then using the diagnostic machine learning model to screen a population of subjects.
[0045] Like reference numbers and designations throughout the various drawings indicate like elements. DETAILED DESCRIPTION
[0046] 1 illustrates an exemplary screening system 100. Screening system 100 is an example of a system implemented as a computer program on one or more computers at one or more locations in which the systems, components, and techniques described below are implemented.
[0047] Screening system 100 is configured to screen an electronic medical record database 102 (EMR database 102) of a subject population to identify subjects predicted to have ASMD 112. The subject population can include any suitable number of subjects, such as 10,000 subjects, 100,000 subjects, or 1 million subjects.
[0048] The EMR database 102 may include electronic medical records from one or more sources (e.g., a database for each of one or more departments within a medical center), or from one or more medical centers, etc. The EMR database 102 may include one or more electronic medical records for each subject in a population of subjects.
[0049] The subject's electronic medical record may include any appropriate data characterizing the subject. For example, the subject's electronic medical record may characterize clinical laboratory data obtained by analyzing a sample of the subject's blood, urine, or other bodily fluid, such as data characterizing bilirubin levels, cholesterol levels, hemoglobin levels, platelet levels, lymphocyte levels, red blood cell levels, triglyceride levels, alanine aminotransferase levels, aspartate aminotransferase levels, and the like. As another example, the subject's electronic medical record may characterize one or more medical symptoms or medical diagnoses of the subject, such as diabetes, interstitial lung disease, obesity, chronic pain, chest pain, joint pain, limb pain, back pain, abnormal gas levels, weight change, confusion, hyperlipidemia, cardiac arrhythmia, sleep apnea, liver failure, osmotic hyponatremia, renal failure, heart disease, dyspnea, cardiomegaly, menorrhagia, keratosis, macular degeneration, hyperopia, cataplexy, abnormal facial features, bronchiolitis, splenomegaly, lung infection, anemia, abdominal pain, abnormal bleeding, and the like. As another example, a subject's electronic medical record may characterize one or more demographic characteristics of the subject, such as the subject's age, the subject's gender, and the like.
[0050] In some cases, EMR database 102 can be pre-screened to include only medical records of subjects with one or more characteristics (e.g., symptoms or diagnoses) associated with ASMD. For example, EMR database 102 can be pre-screened to include only medical records of subjects diagnosed with interstitial lung disease (ILD), splenomegaly, non-alcoholic steatohepatitis (NASH), or any combination thereof. Pre-screening EMR database 102 to include only subjects with characteristics associated with ASMD can reduce the consumption of computing resources (e.g., memory and computing power) by screening system 100. Furthermore, pre-screening EMR database 102 can allow screening system 100 to perform the screening using a machine learning model that is trained specifically to identify ASMD in subjects with specific characteristics selected through pre-screening.
[0051] The screening system 100 includes a feature generation engine 104 and a diagnostic machine learning model 108, each of which will be described in more detail below.
[0052] The feature generation engine 104 is configured to generate a set of subject features 106 for each subject based on one or more electronic medical records of the subject from the EMR database 102. The set of subject features 106 for the subject may represent, for example, clinical laboratory data of the subject, medical symptoms of the subject, medical diagnosis of the subject, demographic characteristics of the subject, etc. The set of subject features for the subject may be represented in a standardized format, such as a tensor (e.g., a vector or a matrix) of cells, where each cell includes data associated with a predefined semantic category (e.g., numeric data, textual data, or a combination of both). For example, one cell in the set of subject features for a subject may include a categorical (e.g., Boolean) variable defining whether the subject has been diagnosed with a particular medical condition (e.g., diabetes), while another cell may include a continuous-valued feature defining the level of an enzyme in the subject's blood.
[0053] To generate a set of subject features 106 for a subject, the feature generation engine 104 can scan the EMR database 102 to identify one or more electronic medical records for the subject, for example, by identifying an electronic medical record that includes a patient identifier associated with the subject. The feature generation engine 104 can then extract data from the subject's electronic medical records to populate the set of subject features 106 for the subject.
[0054] In some cases, the feature generation engine 104 may format or normalize data extracted from the subject's electronic medical record as part of generating the subject's set of subject features 106. For example, the feature generation engine 104 may normalize a continuous-valued feature characterizing a subject (e.g., an enzyme level in the subject's blood) into predefined intervals, such as the interval As another example, the feature generation engine 104 may format text data from an electronic medical record indicating a subject's diagnosis into a categorical value, eg, a Boolean value indicating that the subject has received the diagnosis.
[0055] The diagnostic machine learning model 108 is configured to process a model input comprising a set of subject features 106 of the subject to generate a model output that characterizes the likelihood that the subject has ASMD. Several examples of possible model outputs of the diagnostic machine learning model 108 are described below.
[0056] In some embodiments, a model output of a diagnostic machine learning model may include a hard classification that identifies a subject as being included in one of a set of diagnostic categories, the set of diagnostic categories including: a non-ASMD category (i.e., indicating that the subject does not have an ASMD) and one or more ASMD categories (i.e., indicating that the subject has an ASMD). In some cases, the set of diagnostic categories may include a single ASMD category, and if the subject has any type of ASMD, e.g., type A, type A / B, or type B, the subject may be classified as being included in the ASMD category. In some cases, the set of diagnostic categories may include a corresponding category for each type of ASMD (e.g., type A, type A / B, and type B), and if the subject is predicted to have the corresponding type of ASMD, the subject may be classified as being included in the ASMD category.
[0057] In some embodiments, the model output of a diagnostic machine learning model can include a soft (probabilistic) classification of a score distribution defined over a set of diagnostic categories. As described above, the set of diagnostic categories can include a non-ASMD category and one or more ASMD categories. The score for each diagnostic category can define the likelihood (probability) of a subject being included in the diagnostic category.
[0058] The diagnostic machine learning model 108 can have any appropriate machine learning model architecture that enables the diagnostic machine learning model to perform its described functions. For example, the diagnostic machine learning model 108 can be implemented, for example, as a neural network model, or a random forest model, or a support vector machine model, or a decision tree model, or a linear regression model. In an embodiment in which the diagnostic machine learning model 108 is implemented as a neural network model, the diagnostic machine learning model 108 can include any appropriate number (e.g., 5 layers, 10 layers, or 50 layers) of any appropriate type of neural network layers (e.g., fully connected layers, convolutional layers, attention layers, etc.) and connected in any appropriate configuration (e.g., as a sequence of linear layers). In an embodiment in which the diagnostic machine learning model 108 is implemented as a decision tree model, the diagnostic machine learning model 108 can include any appropriate number of vertices and can implement any appropriate splitting function at each vertex.
[0059] The diagnostic machine learning model 108 may include a set of machine learning model parameters. For example, for the diagnostic machine learning model 108 implemented as a neural network model, the set of machine learning model parameters may define the weights and biases of the neural network layers of the diagnostic machine learning model. As another example, for the diagnostic machine learning model 108 implemented as a decision tree, the set of machine learning model parameters may define the parameters of the corresponding split function used at each vertex of the decision tree. To generate a model output, the diagnostic machine learning model 108 may process the model input according to the values of the set of machine learning model parameters.
[0060] The screening system 100 can use a training system 200 to train the diagnostic machine learning model 108 on a set of training examples. More specifically, the training system 200 can determine, through machine learning training techniques, trained values for the set of machine learning model parameters of the diagnostic machine learning model 108. An example of a training system 200 for training the diagnostic machine learning model 108 is described in more detail with reference to FIG.
[0061] The screening system 100 can generate an ASMD classification 110 for the subject based on the model output generated by the diagnostic machine learning model 108 by processing a set of subject features 106 that characterize the subject. For example, if the probability that the subject has ASMD (as defined by the diagnostic machine learning model output) exceeds a threshold, the screening system 100 can classify the subject as being predicted to have ASMD. The threshold can be, for example 、 、 or any other appropriate threshold.
[0062] The screening system 100 can use the feature generation engine 104 and the diagnostic machine learning model 108 to generate a corresponding ASMD classification 110 for each subject in the subject population. More specifically, for each subject, the screening system 100 can use the feature generation engine 104 to generate a set of subject features 106 based on one or more medical records of the subject from the EMR database 102. The screening system 100 can then process the set of subject features 106 for the subject using the diagnostic machine learning model 108 to generate an ASMD classification 110 for the subject. The screening system 100 can then identify each subject classified by the diagnostic machine learning model 108 as having an ASMD (ASMD subject 112) and output data identifying the ASMD subject 112. For example, the screening system 100 can output a corresponding patient identifier for each subject in the subject population classified as having an ASMD. The data identifying the ASMD subject 112 can be used in any of a variety of ways, as described in more detail below.
[0063] The operations performed by the feature generation engine 104 and the diagnostic machine learning model 108 are automated and parallelizable, and the screening system 100 can quickly screen large populations of subjects. For example, in some embodiments, the screening system 100 can screen a population of at least 100,000 subjects in less than 1 hour, or less than 30 minutes, or less than 10 minutes, or less than 1 minute.
[0064] In some embodiments, the screening system 100 can include an ensemble of multiple diagnostic machine learning models 108, rather than just a single diagnostic machine learning model 108. For example, the screening system 100 can include 10 diagnostic machine learning models, or 100 diagnostic machine learning models, or 1000 diagnostic machine learning models, or any other suitable number of diagnostic machine learning models 108. Each diagnostic machine learning model in the ensemble can have a different architecture, or different model parameter values, or both, relative to the other diagnostic machine learning models in the ensemble. An exemplary technique for training an ensemble of diagnostic machine learning models is described in more detail with reference to FIG2 .
[0065] The screening system 100 can use the ensemble of diagnostic machine learning models 108 to generate an ASMD classification 110 for a subject in any of a variety of ways. For example, the screening system 100 can process a set of subject features 106 for a subject using each diagnostic machine learning model 108 in the ensemble to generate a corresponding model output from each diagnostic machine learning model 108 in the ensemble. The screening system 100 can aggregate the model outputs from the diagnostic machine learning models 108 in the ensemble to generate an ensemble output, and then determine an ASMD classification 110 for the subject based on the ensemble output. For example, each diagnostic machine learning model 108 can generate a corresponding probability value defining the probability that the subject has an ASMD, and the ensemble output can be a measure of central tendency (e.g., a mean, median, or mode) of the probability values generated by the ensemble of diagnostic machine learning models 108. Generating an ASMD classification using an ensemble of diagnostic machine learning models 108 can result in more accurate and stable ASMD classifications.
[0066] The screening system 100 can use the ensemble of diagnostic machine learning models 108 to generate both: (i) an ASMD classification 110 for a subject, and (ii) a confidence score for the subject's ASMD classification 110. The confidence score for the ASMD classification can define the screening system 100's confidence in the accuracy of the ASMD classification. To generate the confidence score for the subject's ASMD classification 110, the screening system 100 can determine a variance measure of the model outputs generated for the subject by the ensemble of diagnostic machine learning models. The screening system 100 can then generate the confidence score for the ASMD classification based on the variance measure of the model outputs. For example, each diagnostic machine learning model 108 in the ensemble can generate a model output that defines a corresponding probability that the subject has an ASMD, and the variance measure can be the standard deviation of the probability values generated by the ensemble. In this example, the screening system 100 can determine the confidence score based on, for example, the inverse of the variance measure, such that a low variance in the predictions generated by the ensemble of diagnostic machine learning models 108 indicates a high confidence, and vice versa.
[0067] The screening system 100 can predict whether a subject has an ASMD based on: (i) the subject's ASMD classification 110, and (ii) a confidence score for the subject's ASMD classification 110. For example, in some embodiments, the screening system 100 designates a subject as having an ASMD only if: (i) the subject's ASMD classification 110 indicates that the subject is predicted to have an ASMD, and (ii) the confidence score meets (e.g., exceeds) a confidence threshold. Thus, the screening system 100 can reduce the likelihood of false positive classifications of subjects with ASMD by filtering out ASMD classifications associated with low confidence scores (e.g., scores below the confidence threshold).
[0068] In some embodiments, the screening system 100 may include a set of multiple diagnostic machine learning models 108, each specifically configured to generate an ASMD classification for a particular class of subjects. More specifically, before generating the ASMD classification for a subject, the screening system 100 may assign the subject to one of a set of patient categories based on the subject's set of subject characteristics 106. Patient categories may be defined in any suitable manner. For example, a patient category may be associated with a particular symptom (e.g., chronic pain, or pulmonary manifestations such as abnormal blood gas levels, abnormal lungs, abnormal lung tests, asphyxial hypoxemia, abnormal breathing, bronchiolitis, chest pain, symptoms of chronic obstructive pulmonary disease, dyspnea, lung infection, progressive lung function impairment, pulmonary fibrosis, or respiratory failure, or a combination thereof), and the screening system 100 may be configured to assign the subject to the patient category only if the subject has that symptom. As another example, a patient category may be associated with a particular medical diagnosis (e.g., interstitial lung disease), and the screening system 100 may be configured to assign the subject to the patient category only if the subject has that medical diagnosis. As another example, a patient category may be associated with a particular demographic characteristic (eg, age), and the screening system may be configured to assign a subject to a patient category only if the value of the subject's demographic characteristic meets a threshold.
[0069] The set of diagnostic machine learning models 108 can include a respective diagnostic machine learning model 108 corresponding to each patient category in the set of patient categories. A diagnostic machine learning model 108 for a patient category can be specifically used to generate an ASMD classification for patients included in the patient category. For example, the training system 200 can train a diagnostic machine learning model 108 for a patient category only on training data derived from subjects included in the patient category. Exemplary techniques for training a diagnostic machine learning model 108 specifically for a patient category are described in more detail with reference to FIG2 .
[0070] To generate an ASMD classification 110 for a subject, the screening system 100 can assign the subject to a patient class and then process a set of subject features 106 for the subject using a diagnostic machine learning model corresponding to the patient class to generate a model output. The screening system 100 can then generate an ASMD classification 110 for the subject based on the model output generated by the diagnostic machine learning model 108 that is specific to the patient class of the subject. Specificizing the diagnostic machine learning model 108 to a particular patient class can allow the diagnostic machine learning model 108 to learn to infer implicit classification rules and criteria that are optimized for accurately classifying subjects included in the patient class.
[0071] In some embodiments, for each subject in the subject population, the screening system can generate a corresponding importance score for each feature in the set of subject features 106 for the subject. The importance score for a feature can represent the influence of the feature on the ASMD classification 110 generated for the subject using the diagnostic machine learning model 108. The screening system 100 can generate the importance scores for the set of subject features using any suitable technique. For example, the importance scores can be "SHAP" values (i.e., "Shapley Additive Explanations") or "locally interpretable model-agnostic explanations" (i.e., "LIME" values), which can be used to essentially determine the importance scores of the features by evaluating how the predictions generated by the diagnostic machine learning model change as the features are altered. The screening system 100 can provide the importance scores for the set of subject features 106 for the subject as interpretable data, e.g., the interpretable data explains and provides a rationale for the ASMD classification 110 generated for the subject by the diagnostic machine learning model 108. Thus, generating the importance scores can increase user confidence in the validity of the ASMD classification 110 generated by the diagnostic machine learning model 108.
[0072] Optionally, the screening system 100 can aggregate the importance scores of the subject features 106 across the subject population to generate an aggregated importance score for each feature in the set of subject features 106. More particularly, for each subject in the subject population, the screening system 100 can generate an importance score distribution that defines the corresponding importance score for each feature in the set of subject features. The screening system 100 can aggregate the importance score distributions of the subjects in the subject population to generate an aggregated importance score distribution. For example, for each feature in the set of subject features, the aggregated importance score distribution can assign an aggregated importance score to the feature, which is a measure of the central tendency of the importance scores assigned to the feature by the importance score distribution specific to each subject. (The measure of central tendency can be, for example, the mean, median, or mode). The aggregated importance score distribution can be used in any of a variety of ways. For example, the aggregated importance score distribution can be used as interpretable data to explain the rationale used by the diagnostic machine learning model 108 to generate the ASMD classification 110 for subjects across the subject population.
[0073] The ASMD classification 110 generated by the screening system 100 can be used in any of a variety of applications. Several exemplary applications of the ASMD classification 110 are described below.
[0074] In some cases, ASMD classification 110 can provide the basis for diagnostic testing 114 of subjects in the subject population. For example, for each subject classified by the screening system 100 as having ASMD, a diagnostic (clinical) test can be performed to confirm the diagnosis. The subject's diagnostic testing can include a blood test to measure the amount of ASMase activity in the subject. For example, if the test shows decreased ASMase activity, a diagnosis of ASMD can be confirmed.
[0075] In some cases, ASMD classification 110 can provide a basis for treating 116 a subject diagnosed with ASMD, e.g., by classification by screening system 100 and by subsequent confirmatory diagnostic testing. Treating a subject for ASMD can include administering a drug, e.g., the drug orlipase alpha or orlipase alpha-rpcp, e.g., Xenpozyme™, to the subject.
[0076] In some cases, the ASMD classification can be used to facilitate further training 118 of the diagnostic machine learning model 108. More specifically, as described above, the screening system 100 can generate a confidence score for each ASMD classification, which represents the confidence level of the diagnostic machine learning model 108 in the ASMD classification. The screening system 100 can use the ASMD classification and the confidence score to select an appropriate subset of the subject population for diagnostic testing 114 for ASMD. The diagnostic test for ASMD can be performed on the selected subjects, and the training system 200 can use the results of the diagnostic test to generate training examples for further training of the diagnostic machine learning model 108. In particular, for each subject selected for diagnostic testing for ASMD, the training system 200 can generate a corresponding training example, which includes: (i) a model input comprising a set of subject features of the subject, and (ii) a target ASMD classification. The target ASMD classification is based on the diagnostic test for ASMD and defines whether the subject has ASMD. The training system 200 may train the diagnostic machine learning model 108 on training examples using machine learning training techniques, as will be described in more detail with reference to FIG. 2 .
[0077] The screening system 100 can designate patients for diagnostic testing for ASMD based on the ASMD classification and confidence score to maximize the improvement in predictive accuracy achieved by further training the diagnostic machine learning model 108. For example, the screening system 100 can designate multiple subjects with the lowest confidence scores for diagnostic testing for ASMD. As another example, the screening system 100 can designate: (i) a first plurality of subjects with the lowest confidence scores from subjects classified as having ASMD, and (ii) a second plurality of subjects with the lowest confidence scores from subjects classified as not having ASMD for diagnostic testing for ASMD. Training the diagnostic machine learning model 108 on training examples corresponding to subjects with ASMD classifications associated with low confidence scores can result in a significant improvement in the predictive accuracy of the diagnostic machine learning model 108.
[0078] The screening system 100 and the training system 200 can operate in series to train the diagnostic machine learning model 108 in a series of update iterations. At each update iteration, the screening system 100 can screen the subject population to generate a corresponding ASMD classification and confidence score for each subject in the subject population. An appropriate subgroup of the subject population can be selected for diagnostic testing based on the ASMD classification and confidence score. The training system 200 can generate new training examples based on the results of the diagnostic test for ASMD and further train the diagnostic machine learning model 108. The process can then be repeated at the next update iteration. The prediction accuracy of the diagnostic machine learning model 108 can be improved at each update iteration, and thus the screening system 100 can refine the screening of the subject population to reduce the number of false negatives and false positives.
[0079] In some embodiments, the screening system 100 can be used to generate an ASMD classification for a single subject, e.g., as an alternative to or in combination with generating a corresponding ASMD classification for each subject in a population of subjects. That is, the screening system 100 can be used to screen individual subjects for ASMD as well as to screen populations of subjects for ASMD.
[0080] 2 illustrates an exemplary training system 200. Training system 200 is an example of a system implemented as a computer program on one or more computers at one or more locations in which the systems, components, and techniques described below are implemented.
[0081] The training system 200 is configured to train the diagnostic machine learning model 108 included in the screening system 100 described with reference to Figure 1. The diagnostic machine learning model 108 is configured to process a model input comprising a set of subject characteristics of the subject to generate a model output that characterizes the likelihood that the subject has ASMD.
[0082] The training system 200 uses a training engine 206 to train a set of machine learning model parameters 208 for the diagnostic machine learning model 108 on a set of training examples 204. Each training example can correspond to a subject (referred to for convenience as a "training subject") and can include: (i) a model input comprising a set of subject features characterizing the training subject, and (ii) a target ASMD classification for the training subject. For each training example 204, the training engine 206 trains the diagnostic machine learning model 108 to process the model input of the training example to generate a model output that matches the target ASMD classification for the training example. More specifically, the training engine 206 trains the diagnostic machine learning model 108 using a machine learning training technique to optimize an objective function that measures the error between: (i) the model output generated by the diagnostic machine learning model 108 for the training subject, and (ii) the target ASMD classification for the training subject. The objective function can measure the error between the model output and the target ASMD classification in any suitable manner, for example, as a squared error or an absolute error.
[0083] The training engine 206 may train the diagnostic machine learning model 108 using any machine learning training technique appropriate to the architecture of the diagnostic machine learning model 108. For example, if the diagnostic machine learning model 108 is implemented as a neural network model, the training engine 206 may train the diagnostic machine learning model using stochastic gradient descent.
[0084] The training system 200 can obtain training examples for training the diagnostic machine learning model 108 in a variety of possible ways. For example, the training system 200 can generate "positive" training examples (i.e., corresponding to training subjects with ASMD) and "negative" training examples (i.e., corresponding to training subjects without ASMD) by scanning an EMR database. In particular, the training system 200 can scan the EMR database to identify one or more training subjects who have been diagnosed with ASMD (e.g., through a diagnostic test) and generate corresponding positive training examples corresponding to each patient identified as having ASMD. Similarly, the training system 200 can scan the EMR database and generate negative training examples based on patients who have not been diagnosed with ASMD.
[0085] Optionally, the training system 200 can construct a set of training examples by utilizing statistical techniques for matching based on confounding features (e.g., demographic features such as age, gender, geographic region of residence, etc.). More particularly, for each positive training example in one or more positive training examples, the training system 200 can identify one or more corresponding negative training examples that "match" the positive training example based on matching criteria evaluated based on the values of the confounding feature dimensions. The training system 200 can then select the negative training examples that match the positive training examples to include in a set of training examples for training the diagnostic machine learning model. If the confounding feature value of the negative training example is the same as the confounding feature value of the positive training example (or within a tolerance threshold thereof), the matching criteria can define that the negative training example matches the positive training example. Using statistical matching to construct a set of training examples can reduce the impact of confounding variables and enhance the robustness of the trained diagnostic machine learning model.
[0086] ASMD is predicted to affect 1 in 250,000 people and is therefore a rare disease. Generating positive training examples (i.e., corresponding to training subjects with ASMD) by scanning the EMR database may only produce a small number of positive training examples (e.g., a few hundred positive training examples or less than a hundred). However, training the diagnostic machine learning model 108 on only a small number of positive training examples may not be sufficient to allow the diagnostic machine learning model to achieve acceptable predictive accuracy. To address this issue, the training system 200 can optionally use a generative system 300 to generate "synthetic" training examples for training the diagnostic machine learning model 108. An example of a generative system is described in more detail with reference to FIG3 .
[0087] In some embodiments, as described with reference to FIG1 , the screening system 100 includes an ensemble of diagnostic machine learning models 108, each having a different architecture, or different model parameter values, or both, relative to the other diagnostic machine learning models in the ensemble. The training engine 206 can train each diagnostic machine learning model 108 in the ensemble on a set of training examples, e.g., as described above. Optionally, the training engine 206 can train each diagnostic machine learning model 108 in the ensemble on a different appropriate subset of the training examples 204 so as to introduce variation between the trained parameter values of the diagnostic machine learning models 108 in the ensemble. For example, the training engine 206 can train each diagnostic machine learning model 108 in the ensemble on a different randomly selected appropriate subset of a set of training examples.
[0088] In some embodiments, as described with reference to FIG1 , the screening system 100 includes a set of diagnostic machine learning models 108, each of which is specialized for generating an ASMD classification for a particular class of subjects. For each class of subjects, the training engine 206 can train the corresponding diagnostic machine learning model 108 only on training examples based on training patients included in the class of subjects. For example, for a class of patients that includes only patients with a particular medical diagnosis (e.g., pneumonia, anemia, or interstitial lung disease), the training engine 206 can train the corresponding diagnostic machine learning model 108 only on training subjects with that medical diagnosis.
[0089] In some embodiments, as described with reference to FIG1 , the training system 200 can train the diagnostic machine learning model 108 in a series of update iterations. At each update iteration, one or more "test" subjects are selected for diagnostic testing for ASMD, and corresponding new training examples are generated for each test subject. At each update iteration, the training system 200 can add the new training examples to the set of training examples 204 and then train the diagnostic machine learning model 108 on the set of training examples.
[0090] In some embodiments, before training the diagnostic machine learning model 108 to perform the ASMD classification task, the training engine 206 pre-trains the diagnostic machine learning model 108 to perform an auxiliary prediction task. That is, the training engine 206 may pre-train the diagnostic machine learning model 108 to perform the auxiliary prediction task and then fine-tune the diagnostic machine learning model on the ASMD classification task. Pre-training the diagnostic machine learning model 108 to perform the auxiliary prediction task can provide efficient initialization for the parameter values of the diagnostic machine learning model 108, thereby accelerating the training of the diagnostic machine learning model on the ASMD classification task. The auxiliary prediction task may, for example, process a set of subject features to classify whether the subject has a medical condition other than ASMD (e.g., diabetes, heart disease, etc.). The auxiliary prediction task may be a prediction task for which significantly more training data is available than the training data used for the ASMD prediction task.
[0091] 3 illustrates an exemplary generative system 300. Generative system 300 is an example of a system implemented as a computer program on one or more computers at one or more locations, in which the systems, components, and techniques described below are implemented.
[0092] Generative system 300 is configured to generate a set of synthetic training examples 304 based on a set of real-world training examples 302 for use in training a diagnostic machine learning model. Synthetic training examples are training examples that do not directly correspond to real-world subjects but are computationally synthesized by generative system 300. Real-world training examples are training examples that directly correspond to real-world subjects. Generative system 300 can provide synthetic training examples 304 to a training system for use in training a diagnostic machine learning model, as described with reference to FIG2 .
[0093] Generative system 300 can generate synthetic training examples 304 in any of a number of possible ways. Several exemplary implementations of generative system 300 are described next.
[0094] In some implementations, generative system 300 can generate synthetic training examples 304 by selecting real-world training examples 302 and adding random noise to a set of subject features included in the model input for real-world training examples 302. The generative system can define the target ASMD classification for synthetic training examples 304 to be the same as the target ASMD classification for real-world training examples 302. Thus, generative system 300 can generate synthetic training examples 304 as noisy versions of real-world training examples 302. In particular, generative system 300 can generate synthetic positive training examples 304 as noisy versions of real-world positive training examples 302. Generative system 300 can sample the random noise values added to the model input for real-world training examples 302 from an appropriate probability distribution (e.g., a standard normal distribution). (This implementation of generative system 300 is illustrated in FIG3 as “Add Random Noise 306”).
[0095] In some embodiments, to generate synthetic training examples, generative system 300 may condition generative neural network 308 on a target ASMD classification and then generate a set of subject features as output of the generative neural network. Generative system 300 may then instantiate a synthetic training example that includes: (i) a model input comprising a set of subject features generated by generative neural network 308, and (ii) the target ASMD classification for conditioning generative neural network 308. ("Conditioning" the neural network on the target ASMD classification may refer to providing the target ASMD classification as input to the neural network). Generative system 300 may train generative neural network 308 to generate sets of realistic subject features, e.g., drawn from the same distribution as the sets of subject features of the real-world training examples. Next, some exemplary techniques by which generative system 300 may train generative neural network 308 are described.
[0096] In one example, the generative system 300 can train a generative neural network as a generative adversarial neural network 310. More specifically, the generative system 300 can jointly train the generative neural network with another neural network, referred to as a discriminator neural network. The generative neural network is configured to process: (i) a tensor of random values (e.g., a vector or matrix) and (ii) a target ASMD classification to generate a set of synthetic subject features. The discriminator neural network is configured to process: (i) the target ASMD classification and (ii) a set of subject features to generate a discrimination score. The discrimination score defines the likelihood that the set of subject features was generated by the generative neural network rather than a set of subject features representative of a real-world subject. The generative system 300 alternates between training the generative neural network and the discriminator neural network. Specifically, the generative system 300 trains the generative neural network to optimize an objective function based on the discrimination scores generated by the discriminator neural network. The objective function encourages the generative neural network to generate sets of subject features that the discriminator neural network incorrectly classifies as representative of a real-world subject. The generative system 300 trains the discriminator neural network to optimize an objective function that encourages the discriminator neural network to accurately distinguish between several sets of real-world subject features and several sets of synthetic subject features. After training, the discriminator neural network can be discarded and the generative neural network can be used to generate synthetic training examples for training the diagnostic machine learning model 108.
[0097] In another example, the generative system 300 can train the generative neural network 308 as a diffusion model 312. More specifically, the generative system 300 can train the generative neural network to process: (i) a target ASMD classification, and (ii) a noisy version of a set of subject features for a subject with the target ASMD classification to generate a denoised version of the set of subject features. A "noisy version" of a set of subject features refers to a set of subject features that has been combined with random noise sampled from a probability distribution (e.g., a standard normal distribution). A "denoised" version of a set of subject features refers to a set of subject features from which the random noise has been removed. The generative system 300 can train the generative neural network on both noisy and denoised versions of real-world training examples. After training, the generative system 300 can use the generative neural network 308 to generate synthetic training examples.
[0098] To generate synthetic training examples, generative system 300 may select a target ASMD classification and generate an initial set of synthetic subject features by sampling each feature from a probability distribution (e.g., a standard normal distribution). Generative system 300 may then iteratively denoise the combined subject features in a series of denoising iterations. More specifically, during a first denoising iteration, generative system 300 may process the initial set of synthetic subject features using generative neural network 308 to generate a denoised version of the combined subject features. Generative system 300 may then provide the denoised version of the combined subject features for processing by generative neural network 308 during the next denoising iteration. Thus, generative system 300 may iteratively denoise the combined subject features using generative neural network 308 during the series of denoising iterations. Generative system 300 may generate synthetic training examples that include: (i) a model input comprising the set of synthetic subject features generated during the last denoising iteration, and (ii) the target ASMD classification.
[0099] Figure 4A illustrates an example of a diagnostic machine learning model implemented as a decision tree that a screening system can use to screen a population of subjects (or a single subject) for ASMD. Each node in the decision tree (except for the leaf nodes) is associated with a split function that defines the rules for assigning subjects to the node's corresponding child nodes. For example, root node 402 is associated with a split function that assigns the subject to a first child node 404 if the subject's HDL cholesterol level is below 29.16 mg / dl, and to a second child node 406 if the subject's HDL cholesterol level is above 29.16 mg / dl. The decision tree classifies subjects by progressing from the root node of the decision tree to the leaf nodes of the decision tree. At each node before the leaf node, the decision tree applies the node's associated split function to a set of features of the subject to assign the subject to a child node. Each leaf node in the decision tree has an associated ASMD classification; for example, leaf node 410 is associated with a classification of "control" (e.g., identifying the subject as not suffering from ASMD), and leaf node 408 is associated with a classification of ASMD (e.g., identifying the subject as suffering from ASMD). The decision tree can classify the subject based on the ASMD classification of the leaf node to which the subject is assigned. The exemplary diagnostic machine learning model in FIG4A is provided for illustrative purposes only and is non-limiting. Other implementations of the diagnostic machine learning model are also possible, for example, using larger decision trees, using different machine learning model architectures (e.g., random forests or neural networks), using ensembles of machine learning models, etc.
[0100] FIG4B shows a table of experimental results of applying the diagnostic machine learning model of FIG4A to a subject population. Rows 412, 414, and 416 of the table each correspond to a respective leaf node of a decision tree and show: (i) a series of split decisions applied by the decision tree to reach the leaf node; (ii) a number of ASMD subjects classified into the leaf node; and (iii) a number of control (non-ASMD) subjects classified into the leaf node.
[0101] FIG4C provides a bar graph showing the relative prevalence of certain features in (i) ASMD subjects and (ii) control (non-ASMD) subjects. It will be appreciated that certain features are significantly more prevalent in ASMD subjects than in control subjects. Diagnostic machine learning models can learn to identify and utilize combinations of such features to accurately classify whether a subject has ASMD, as described throughout this specification.
[0102] FIG4D shows an example of a diagnostic machine learning model implemented as a decision tree that has been trained specifically to perform ASMD screening on a population of subjects with unexplained ILD.
[0103] FIG4E shows a table of experimental results of applying the diagnostic machine learning model of FIG4D to a population of subjects with unexplained ILD.
[0104] FIG5 is a flow chart of an exemplary process 500 for screening an electronic medical record database of a population of subjects to identify subjects predicted to have ASMD. For convenience, process 500 will be described as being performed by a system comprised of one or more computers located at one or more locations. For example, a screening system (e.g., screening system 100 of FIG1 ) appropriately programmed according to this specification can perform process 500.
[0105] The system obtains, for each subject, a set of features characterizing the subject from the electronic medical record database (502). The set of features characterizing the subject includes features that may include features derived from a sample of a body fluid or body tissue of the subject.
[0106] The system processes, for each subject, a model input comprising the set of features characterizing the subject using a diagnostic machine learning model to generate a model output characterizing a likelihood that the subject has ASMD (504). The diagnostic machine learning model can include one or more of a decision tree model, a random forest model, a neural network model, or a support vector machine model.
[0107] The diagnostic machine learning model may have been trained using a machine learning training technique to determine trained values for the set of diagnostic machine learning model parameters. In particular, the diagnostic machine learning model may have been trained on a set of training examples, each training example comprising: (i) a training input based on a set of features of a training subject, and (ii) a target output based on whether the training subject was designated as having ASMD. Some positive training examples (i.e., for training subjects designated as having ASMD) may be generated by a generative neural network. The generative neural network may be trained as a generative adversarial neural network or implemented as a diffusion model, as described above with reference to FIG3 .
[0108] For each subject and based on the model output of the diagnostic machine learning model for the subject, the system classifies whether the subject has ASMD (506). For example, if the likelihood that the subject has ASMD (as defined by the model output for the subject) meets a threshold, the system may classify the subject as having ASMD.
[0109] The system provides an output that identifies subjects from the population of subjects classified as having ASMD (508). In some cases, such as after confirmatory diagnostic testing, a drug for treating ASMD (e.g., oripase alpha or oripase alpha-rpcp) can be administered to one or more of the subjects classified as having ASMD.
[0110] 6 illustrates an exemplary process for training a diagnostic machine learning model to identify subjects with ASMD and then using the diagnostic machine learning model to screen a study subject population. More specifically, in a training phase 602, the screening system trains a diagnostic machine learning model (in this illustration, a decision tree 610) to perform ASMD classification on a training subject population 606. In an application phase 604, the screening system uses the diagnostic machine learning model to screen a study subject population 608 for ASMD. In this example, the study subject population is pre-screened to include only subjects with unexplained ILD and who are less than or equal to 50 years of age. The results of the screening are, Study subjects (from In a study subject population, subjects are labeled as potentially having ASMD. Subjects labeled as potentially having ASMD can undergo subsequent clinical testing (e.g., enzyme and genetic testing in leukocytes, dried blood spots, or cultured fibroblasts) and, if appropriate, treatment for ASMD, as described throughout this specification.
[0111] Certain aspects of the subject matter of this specification are further described in Appendix A.
[0112] This specification uses the term "configured" when referring to systems and computer program components. For a system consisting of one or more computers to be configured to perform a particular operation or action, it means that the system has installed thereon software, firmware, hardware, or a combination thereof that, when run, causes the system to perform those operations or actions. For one or more computer programs to be configured to perform a particular operation or action, it means that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform those operations or actions.
[0113] The embodiments of the subject matter and functional operations described in this specification can be implemented in digital electronic circuits, tangibly embodied computer software or firmware, computer hardware (including the structures disclosed in this specification and their structural equivalents), or a combination of one or more thereof. The embodiments of the subject matter described in this specification can be implemented as one or more computer programs (i.e., one or more computer program instruction modules encoded on a tangible, non-transitory storage medium) for execution by a data processing device or for controlling the operation of the data processing device. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more thereof. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagation signal (e.g., a machine-generated electrical signal, optical signal, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiver device for execution by the data processing device.
[0114] The term "data processing apparatus" refers to data processing hardware and encompasses all types of devices, equipment, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. The apparatus may also be or further include special-purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). In addition to the hardware, the apparatus may optionally include code that creates an execution environment for computer programs, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these.
[0115] A computer program (which may also be referred to or described as a program, software, software application, application, module, software module, script, or code) may be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions). A computer program may be deployed to execute on a single computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communications network.
[0116] In this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Typically, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a specific engine; in other cases, multiple engines may be installed on or run on the same computer or computers.
[0117] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry (e.g., an FPGA or ASIC), or by a combination of special purpose logic circuitry and one or more programmed computers.
[0118] A computer suitable for executing computer programs can be based on a general-purpose microprocessor, a special-purpose microprocessor, or both, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a central processing unit for executing instructions and one or more memory devices for storing instructions and data. The central processing unit and memory may be supplemented by or incorporated into dedicated logic circuitry. Typically, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or be operatively coupled to one or more mass storage devices for storing data, to receive data from or transfer data to, or both. However, a computer need not include such devices. Furthermore, a computer may be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive).
[0119] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example: semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0120] To provide for user interaction, embodiments of the subject matter described in this specification can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide for user interaction; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including auditory, voice, or tactile input. Furthermore, a computer can interact with a user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser on a user's device in response to a request received from the web browser. Furthermore, a computer can interact with a user by sending text messages or other forms of messages to a personal device (e.g., a smartphone running a messaging application) and receiving response messages from the user.
[0121] The data processing apparatus used to implement the machine learning model may also include, for example, dedicated hardware accelerator units for handling the typical compute-intensive parts of machine learning training or production (i.e., inference) workloads.
[0122] Machine learning models can be implemented and deployed using a machine learning framework (e.g., the TensorFlow framework or the Jax framework).
[0123] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component (e.g., as a data server), a middleware component (e.g., an application server), a front-end component (e.g., a client computer with a graphical user interface, a web browser, or an application program through which a user can interact with an implementation of the subject matter described in this specification), or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (LANs) and wide area networks (WANs) (e.g., the Internet).
[0124] A computing system may include a client and a server. The client and server are typically remote from each other and typically interact via a communication network. The client-server relationship arises from computer programs running on the respective computers and having a client-server relationship with each other. In some embodiments, the server transmits data (e.g., an HTML page) to a user device acting as a client, for example, to display data to a user interacting with the device and to receive user input from the user. Data generated at the user device (e.g., the results of the user interaction) can be received from the device at the server.
[0125] Although this specification contains many specific implementation details, these details should not be interpreted as limitations on the scope of any invention or the scope of what may be claimed, but rather as descriptions of features that may be specific to a particular embodiment of a particular invention. Certain features described in this specification in the context of separate embodiments may also be implemented in a single embodiment in combination. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments, either individually or in any suitable subcombination. Furthermore, although features may be described above as functioning in a certain combination and even initially claimed as such, in some cases one or more features from a claimed combination may be deleted from that combination, and a claimed combination may involve subcombinations or variations of subcombinations.
[0126] Similarly, although operations are depicted in a particular order in the drawings and recited in a particular order in the claims, this should not be construed as requiring that such operations be performed in the particular order shown or in a sequential order, or that all illustrated operations be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the embodiments described above should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0127] Specific embodiments of the present subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve the desired results. As an example, the processes depicted in the accompanying figures do not necessarily require the particular order shown or sequential order to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous.
[0128] Claims:
Claims
1. A method performed by one or more computers, the method comprising: Screening an electronic medical record database of a subject population to identify subjects predicted to have acid sphingomyelinase deficiency comprises, for each subject in the subject population: obtaining a set of characteristics characterizing the subject from the electronic medical record database; processing a model input comprising the set of features characterizing the subject using a diagnostic machine learning model having a set of diagnostic machine learning model parameters to generate a model output characterizing a likelihood that the subject has acid sphingomyelinase deficiency; as well as classifying whether the subject has acid sphingomyelinase deficiency based on the model output of the diagnostic machine learning model; and An output is provided identifying a subject classified as having acid sphingomyelinase deficiency from the population of subjects.
2. The method according to claim 1, wherein The diagnostic machine learning model has been trained using a machine learning training technique to determine trained values for the set of diagnostic machine learning model parameters.
3. The method according to claim 2, wherein: Training the diagnostic machine learning model through the machine learning training technique includes: obtaining a set of training examples, wherein each training example comprises: (i) a training input comprising a set of features of a training subject, and (ii) a target output based on whether the training subject is designated as having acid sphingomyelinase deficiency; and The diagnostic machine learning model is trained on the set of training examples.
4. The method according to claim 3, wherein: The training examples obtained include: obtaining a set of positive training examples, wherein each positive training example corresponds to a subject having acid sphingomyelinase deficiency; and Select a plurality of negative training examples, each corresponding to a respective subject not having acid sphingomyelinase deficiency, the selection comprising, for each positive training example: identifying a corresponding set of negative training examples based on matching one or more confounding features in the set of negative training examples with the positive training examples; and The set of negative training examples is selected for training the diagnostic machine learning model.
5. The method according to claim 4, wherein: The one or more confounding characteristics include demographic characteristics.
6. The method according to any one of claims 3 to 5, wherein Training the diagnostic machine learning model on the set of training examples includes: The diagnostic machine learning model is trained to process, for each training example, a training input for the training example to generate a model output that matches a target output for the training example.
7. The method according to any one of claims 3 to 6, wherein Obtaining the set of training examples includes, for one or more training examples: The training examples are generated from one or more electronic medical records of a training subject corresponding to the training examples.
8. The method according to any one of claims 3 to 7, wherein The training examples obtained from this set include: generating a set of features for the subject using a generative neural network; and Training examples comprising training inputs are generated based on the set of features generated by the generative neural network.
9. The method of claim 8, wherein: The generative neural network has been trained with operations including: generating a set of output features using the generative neural network; processing the set of output features generated by the generative neural network using a discriminator neural network to generate a discriminant score that defines a likelihood that the set of output features was generated by the generative neural network rather than originating from one or more electronic medical records; as well as The generative neural network is trained to optimize an objective function that depends on the discriminant score generated by the discriminator neural network.
10. The method of claim 8, wherein: The set of features generated for the subject using this generative neural network includes: Instantiating initial values of the set of characteristics of the subject as random values sampled from a probability distribution; and Denoising the set of features in a series of denoising iterations, the denoising including at each denoising iteration before the last denoising iteration in the series of denoising iterations: processing the current values of the set of features using a generative neural network to generate a denoised version of the set of features; and providing the denoised version of the set of features for processing in a next denoising iteration; The denoised version of the set of features generated at the last denoising iteration defines the set of features for the subject.
11. The method of any preceding claim, further comprising: generating, for each subject in the population of subjects, a confidence score characterizing a confidence of the diagnostic machine learning model in classifying whether the subject has acid sphingomyelinase deficiency; selecting an appropriate subgroup of the subject population for diagnostic testing for acid sphingomyelinase deficiency based on these confidence scores; retraining the diagnostic machine learning model based on results of a diagnostic test on an appropriate subgroup of the subject population; as well as The retrained diagnostic machine learning model is used to reclassify whether each subject in the subject population has acid sphingomyelinase deficiency.
12. The method of claim 11, wherein: Selection of appropriate subgroups of this subject population for diagnostic testing for acid sphingomyelinase deficiency based on these confidence scores includes: A number of subjects associated with the lowest confidence scores are selected from the subject population for diagnostic testing for acid sphingomyelinase deficiency.
13. The method of claim 11, wherein: Selection of appropriate subgroups of this subject population for diagnostic testing for acid sphingomyelinase deficiency based on these confidence scores includes: selecting a plurality of subjects associated with a lowest confidence score from among the subjects classified as having acid sphingomyelinase deficiency for diagnostic testing for acid sphingomyelinase deficiency; and A number of subjects associated with the lowest confidence score are selected from the subjects classified as not having acid sphingomyelinase deficiency for diagnostic testing for acid sphingomyelinase deficiency.
14. The method according to any one of claims 11 to 13, wherein Generating, for each subject in the subject population, a confidence score characterizing the confidence of the diagnostic machine learning model in the classification of whether the subject has acid sphingomyelinase deficiency comprises, for each subject in the subject population: generating, using each diagnostic machine learning model in the ensemble of the plurality of diagnostic machine learning models, a corresponding model output characterizing a likelihood that the subject has acid sphingomyelinase deficiency; as well as The confidence score is generated based on a variance measure of the model outputs generated by the ensemble of diagnostic machine learning models for the subject.
15. A method as claimed in any preceding claim, wherein The subject population includes at least 10,000 subjects.
16. A method as claimed in any preceding claim, wherein Screening of electronic medical record data for this subject population was performed in 1 minute or less.
17. The method of any preceding claim, further comprising: For one or more subjects classified as having acid sphingomyelinase deficiency, A drug is administered to the subject to treat acid sphingomyelinase deficiency in the subject.
18. The method of claim 17, wherein: This medicine includes orlipase alfa or orlipase alfa-rpcp.
19. A method as claimed in any preceding claim, wherein The set of characteristics characterizing the subject includes characteristics of a sample of bodily fluid or bodily tissue derived from the subject.
20. A method as claimed in any preceding claim, wherein The diagnostic machine learning model includes one or more of the following: a decision tree model, a random forest model, a neural network model, or a support vector machine model.
21. A method as claimed in any preceding claim, wherein The subject population was pre-screened to include only subjects who had been diagnosed with one or more of the following: interstitial lung disease (ILD), splenomegaly, or nonalcoholic steatohepatitis (NASH).
22. A system comprising: one or more computers; as well as One or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the corresponding method of any one of claims 1-21.
23. One or more non-transitory computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the corresponding method of any one of claims 1-21.