Method and system for obtaining adverse outcome information
A machine learning-based method using adversarial networks and model selection addresses missing data challenges in pre-eclampsia predictions, ensuring accurate and uncertainly quantified risk assessments in diverse healthcare environments.
Patent Information
- Application Number
- PCT/GB2025/051580
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-17
- Filing Date
- 2025-07-17
- Publication Date
- 2026-01-22
AI Technical Summary
Current predictive models for pre-eclampsia-related adverse maternal outcomes struggle with missing data during deployment in real-world settings, leading to inaccurate predictions and a lack of guidance on handling incomplete data, especially in resource-constrained environments.
A method utilizing a multi-step hierarchical classification tool that employs machine learning to handle missing data through synthetic data generation using adversarial networks and model selection, allowing for accurate risk prediction even with incomplete data.
The method provides reliable risk predictions for adverse maternal outcomes in pre-eclampsia by minimizing data imputation requirements and quantifying uncertainty, enabling effective decision-making in various healthcare settings.
Smart Images

Figure GB2025051580_22012026_PF_FP_ABST
Abstract
Description
[0001] Method and System for Obtaining Adverse Outcome Information
[0002] Field
[0003] The present invention relates to a method and system for obtaining adverse maternal outcome information, for example, a risk prediction such as a risk level and / or probability associated with one or more adverse maternal outcomes for a subject with suspected or confirmed preeclampsia.
[0004] Background
[0005] Pre-eclampsia is a pregnancy condition affecting 2 to 8% of pregnancies which can cause adverse maternal outcomes including 50,000-100,000 maternal deaths annually and serious morbidities. Current models recommended in clinical practice for prediction of adverse outcomes use logistic regression or survival analysis and are healthcare system specific.
[0006] An example of current models include the logistic regression-based Pre-eclampsia Integrated Estimate of RiSk (PIERS) model described in “Prediction of adverse maternal outcomes in pre-eclampsia: development and validation of the fullPIERS model” by von Dadelszen P et aL, Lancet 2011 ; 377(9761): 219-27. The fullPIERS model was developed for well-resourced settings and using variables such as gestational age, symptoms (including chest pain or dyspnea), signs (oxygen saturation), and laboratory tests (platelet count, and creatinine and aspartate transaminase concentrations).
[0007] “A machine-learning-based algorithm improves prediction of preeclampsia-associated adverse outcomes” Am J Obstet Gynecol 2022; 227(1): 77 e1- e30, by Schmidt et al. describes an algorithm for determining a composite foetal and maternal adverse outcome at any time after a visit, up until 14 days after delivery of the neonate.
[0008] Missing data in a development phase of predictive modelling is a widely discussed topic. However, missing data at the model deployment stage (i.e. during prediction) is less discussed. Many predictive modelling methods, especially commonly used regression-based methods, require complete observations with no missing values for the purposes of prediction. In general, known methods for handling missing data during model development (multiple imputation) does not allow for the imputation method to be trained on the development dataset and then deployed on a single new observation in most cases. The few cases that do allow this require the development dataset to be available for the new observation to be appended to, as the parameters of the imputation are estimated from the development data each time.
[0009] Moreover, as data is often collected in a controlled environment of a study for model development, it may not be appropriate to use the same method of imputation as the development stage when the models are deployed in much less controlled real-world settings, where a range of additional factors may influence data availability. In such, real-world settings, a missing value may be indicative or linked to an outcome even if they were not during the development stage. Missing values may also be indicative of the available resources, or of personal bias of the decision-maker.
[0010] While national and international guidelines are in place that recommend the use of predictive models for pre-eclampsia patients, the guidelines do not offer suggestions on what to do in the absence of the required variables for these models are not available. The models recommended in guidelines are regression-based models, which are not able to make predictions on data with missing values.
[0011] Summary
[0012] According to a first aspect, there is provided a method comprising: receiving or otherwise obtaining input data for a subject with suspected or confirmed pre-eclampsia, wherein the input data comprises data representing a plurality of input variables comprising clinical data variables and / or other input variables associated with the subject and / or healthcare setting, wherein the input variables are grouped into a plurality of variable groups, wherein the method comprises:processing the input data to identify each of the plurality of variable groups as one of: empty, partially complete or complete; performing a data completion process to complete the input data for the one or more partially complete variable group; combining the data of the complete variable group and the completed data of the partially complete variables groups to form combined data; performing a model selection process based on the received or otherwise obtained input data to select at least one model from a plurality of trained models; applying the selected model to at least some of the combined data to obtain a risk level and / or probability associated with one or more adverse maternal outcomes for a subject with suspected or confirmed pre-eclampsia. At least part of the method, optionally all of the method, may be a computer- implemented method.
[0013] The input data may represent a single observation and / or more than one observation for a single subject. The input data represents available data for the subject. The input data may represent available data for the subject. The input data may correspond to a data availability scenario.
[0014] The variable groups may be pre-determined based on one or more data availability criteria and / or a degree of similarity between variables of the data groups.
[0015] The variable group may comprise at least one of, optionally all of: a baseline variable group, a patient data group, medical history group, signs group, symptom group, one or more blood test groups, urine test group
[0016] The method may comprise performing one or more data collection tasks to collect at least partial data for one or more variable groups and generating synthetic data for a variable group containing only partially collected data.
[0017] The data collection tasks may comprise: a) obtaining patient data; b) obtaining medical history data; c) obtaining vital sign data; d) obtaining symptom data; e) obtaining a first set of one or more blood samples; f) obtaining a second set of one or more blood samples; g) obtaining a urine sample.
[0018] The method may comprise determining data that is missing from the input data that can improve the prediction, for example, to reduce uncertainty and / or increase accuracy. The method may further comprise obtaining data to replace the missing data and obtaining a further prediction using at least the obtained data.
[0019] The method of any preceding claim further comprising determining a measure of a loss in accuracy or other performance metric of the selected model due to missing data for one or more input variable. The method may further comprise performing at least one further data collection process to obtain data for said one or more input variables. The method may further comprise updating the combined data with the collected data for said one or more input variables and obtaining a further prediction and / or risk level. The method may further comprise updating the combined data with the collected data and performing one or more steps of the method in response to obtaining the further data.
[0020] The variable groups comprise a first blood test group and a second blood test group.
[0021] The baseline variable group comprises one or more, optionally all of: Gestational age on admission, maternal age at expected delivery date, systolic and diastolic blood pressure, national maternal mortality rate and national per capita GDP
[0022] The patient data group comprises one or more, optionally all, of: singleton / multiple pregnancy, ethnicity, parity
[0023] The medical history group comprises one or more, optionally all, of: renal disease, chronic hypertension, pre-gestational diabetes, gestational diabetes in previous pregnancy, history of smoking
[0024] The signs group comprises one or more, optionally all of: SpO2, height on admission, weight on admission
[0025] The symptoms variable group comprises one or more of: vomiting or nausea, right upper quadrant or epigastric pain, headache or visual disturbances, chest pain or dyspnoea
[0026] The first blood tests group comprises: total leucocyte count, platelet count, mean platelet volume, uric acid, hematocrit, serum creatinine, aspartate transaminase, alanine transaminase, lactase dehydrogenase, serum albumin
[0027] The second blood tests group comprises: International normalised ratio, fibrinogen, activated partial thromboplastin time The urine test group comprises: Dipstick proteinuria
[0028] The data of an identified partially complete variable group may comprise missing data and wherein the data completion process comprises generating synthetic data for the missing data and / or wherein the data of an identified empty group comprises no data and / or only missing data and / or wherein the data of an identified complete group comprises no missing data.
[0029] Missing data may comprises an absence of input data for an input variable. The input data may be characterized by a degree of missingness, and / or comprises at least some missing data.
[0030] The data for each partially complete variable group may comprise at least one missing data value from the input data, wherein generating the synthetic data comprises replacing the at least one missing data value with at least one generated value for that variable group.
[0031] Generating the synthetic data may be independently performed for the data of each partially complete variable group
[0032] Generating the synthetic data may comprise applying an adversarial network, for example, a GAIN, to at least part of the received or otherwise obtained input data, optionally applying an adversarial network, for example, a GAIN to each partially complete group.
[0033] The adversarial network may comprise a general adversarial imputation network (GAIN).
[0034] Generating the synthetic data may comprise applying a network forming at least part of pre-trained adversarial network to at least part of the input data.
[0035] Generating the synthetic data may comprise applying one of a plurality of trained networks, each of the plurality of networks train to generate data for a subset of variable groups, to at least part of the input data corresponding to said sub-set of variable groups. The network may comprise: a first network model configured to receive incomplete data and indicators of missing data as its input and outputs at least synthetic data wherein synthetic data is generated for missing data; a second network model configured to receive a combination of real and synthetic data as the input and classify the input data as real or synthetic, wherein the training of the network comprises applying a penalty to the second model for incorrectly classifying a synthetic variable as real and the first model is penalised if generated synthetic data is classified as real by the second network model.
[0036] Generating the synthetic data may comprise replacing missing data for one or more input variables with a plurality or range of generated values for that variable, wherein applying the selected model comprises obtaining a plurality and / or range of predictions for the missing data using the plurality or range of generated values
[0037] The selection may be based on at least a pre-determined performance metric associated with the model. The pre-determine performance metric may be calculated per risk level of an adverse outcome.
[0038] The selection process may comprises using a performance metric for one or more of a plurality of risk levels thereby to select the model for said one or more plurality of risk levels.
[0039] The performance metric may comprise at least one of a positive or negative likelihood ratio, a number of patients and outcome rate, an area under a precision-recall curve and / or an F1 value
[0040] Selecting the model may comprise filtering a plurality of models to obtain a filtered set of models based on the identified partially complete and complete variable groups and selecting a model from the filtered set of models, for example, based on a performance metric associated with the model.
[0041] The plurality of models may comprise at least one trained model for each possible combination of variable groups and the filtered set comprises at least one trained model for each possible combination of identified partially complete and complete groups
[0042] The selection process may be further based on a ruling in or ruling out criteria, optionally, based on a selection by a user.
[0043] The model is configured to rule out and / or rule in one or more of a plurality of risk levels.
[0044] Selecting based on a selected ruling in or ruling out criteria comprises selecting a model based on ruling in and / or ruling out one of a plurality of risk levels.
[0045] Each model of the plurality of models may comprise at least one of: logistic regression, random forest, LASSO, ridge regression, artificial neural networks and Bayesian Model averaging.
[0046] The selected model may be trained to receive complete data for a corresponding combination of variable groups, wherein applying the selected model to the at least some combined data comprises applying the selected model to a subset of the combined data for that combination of variable groups.
[0047] The selected model may be configured to output a prediction using only part of the combined data. The selected model may be configured to output a prediction using some, optionally all, of the input variables of the combined data.
[0048] The method may comprise determining a measure of a loss in accuracy or other performance metric of the selected model due to missing data for or one or more variables, optionally providing a recommendation for improvement in accuracy or other performance metric based on the determined measure. The recommendation may comprise a recommendation on including one or more input variable corresponding to missing data in the input data.
[0049] The method may comprise displaying an interface for receiving input data and / or receiving the user input, optionally via a user input device and displaying output of the model via the interface. The adverse maternal outcome may comprises at least one of: maternal death; an adverse central nervous system event; a cardiorespiratory event; a hematologic event; a hepatic event; a renal event or one or more of: placental abruption, severe ascites, bell’s palsy.
[0050] The risk level and / or probability may represent the risk level and / or probability of an occurrence of the adverse maternal outcome within one or more predefined time periods, optionally, wherein the predefined time period comprises two days and / or seven days.
[0051] The model may be configured to classify the subject into one of a plurality of risk levels based on a probability of the occurrence of one or more adverse maternal events in a predetermined time period.
[0052] Obtaining input data may comprise receiving user input data representative of the country and / or region of the health care system and retrieving the health care system data representing a value for at least one statistic associated with the country and / or region of the health care system. Obtaining input data may comprise performing a vital sign measurement to obtain the vital sign data. Obtaining input data may comprise receiving user input data representative of the demographic data. Obtaining input data may comprise obtaining blood or other sample test data.
[0053] Obtaining the input data may comprise obtaining a sample from the subject and performing a sample analysis on the sample to obtain at least some of the sample data.
[0054] The sample may comprise a blood sample and wherein the sample analysis comprises a blood sample analysis to obtain at least one of: a) one or more haematological parameters including haematocrit, mean platelet volume, uric acid, platelet count, total leukocyte count, serum creatinine; b) one or more renal parameters including lactate dehydrogenase, aspartate transaminase, alanine transaminase, serum albumin. The sample may comprise a blood sample and the sample analysis comprises a blood sample analysis to obtain data relating to blood coagulation, for example, at least one of: International normalised ratio, fibrinogen, activated partial thromboplastin time.
[0055] The sample may comprise a urine sample and wherein the sample analysis comprises a urine sample analysis to obtain dipstick proteinuria value.
[0056] The method may further comprise receiving user input data representing at least some input data and / or displaying the obtained risk level and / or probability. User input data may be received via one or more user input devices.
[0057] The risk level and / or probability may represent the risk level and / or probability of an occurrence of the adverse maternal outcome within one or more predefined time periods. The predefined time period may comprise two days and / or seven days.
[0058] The model and / or the application of the model to data to obtain an output may form part of a machine learning derive procedure.
[0059] The output of the model may comprise risk level and / or probability associated with one or more adverse maternal outcomes for a subject with suspected or confirmed preeclampsia. One or more further actions may be performed based on the output of the model and / or a score derived from the output of the model and / or one or more model performance metrics. The one or more further actions may be performed based on an evaluation of the output of the model. The evaluation of the model may comprise comparing an output risk level to a pre-determined level or risk and / or an output probability to a threshold probability value.
[0060] The one or more further action may comprise performing one or more further data collection processes and / or medical scans and / or displaying a recommendation.
[0061] The method may further comprise determining a drug administration scheme and / or suitable medical intervention based on the output of the model or an evaluation of the output. The method may comprise displaying a recommendation for drug administration and / or performing the medical intervention based on the output of the model. The method may comprise performing the drug administration and / or performing the medical intervention based on the output of the model. The method may comprise performing a further data collection process based on the output of the model or an evaluation of the output. The method may comprise obtaining a sample, for example, a blood sample and performing an analysis of said sample, based on the determined output. The method may comprise performing a medical scan, for example, an obstetric ultrasound scan, based on the output or an evaluation of the output.
[0062] The method may comprises determining to admit or to not admit or to transfer a patient to a health care facility based on the output of the model. The method may comprise selecting a health care facility based on the output model. The method may comprise evaluating the output of the model and admitting or not admitting a patient based on said evaluation.
[0063] The method may comprise displaying a recommendation based on evaluation of the output. The evaluation may comprise comparing the output to a pre-determined threshold value or risk level. The recommendation may comprise a recommendation to obtain further data that is missing from the initial data. The recommendation may include a recommendation to admit the patient to a healthcare facility and / or to transfer a patient to a further health care facility. The recommendation may include a recommendation to administer or to defer or not administer medication and / or to perform or defer or not perform a medical intervention.
[0064] The output of the model may be stored and processed for subsequent training and / or revising the model.
[0065] According to a second aspect, there is provided a method comprising: receiving or otherwise obtaining input data for a subject with suspected or confirmed pre-eclampsia, wherein the input data comprises data corresponding to one of a plurality of data availability scenarios, wherein the method comprises: processing the input data to determine a corresponding data availability scenario; generating synthetic data to form complete data for the data availability scenario; performing a model selection process based on data availability scenario; and applying the selected model to at least some of the complete data to obtain a risk level and / or probability associated with one or more adverse maternal outcomes for a subject with suspected or confirmed pre-eclampsia. According to a third aspect, there is provided a method of training a plurality of models comprising: obtaining training data associated with a plurality of subjects, wherein the training data represents data for a plurality of input variables comprising clinical data variables and / or other input variables associated with the subject and / or healthcare setting, wherein the input variables are grouped into a plurality of variable groups; generating sets of training data for combinations of the variable groups; performing a training process to train one or more models for each variable group combination; obtaining a performance metric for at least one more for each variable group combination; storing said trained models for selection.
[0066] The method may further comprise performing an imputation for each set of training data.
[0067] According to a fourth aspect there is provided a method of training a synthetic data generator comprising: a first network model configured to receive incomplete data and indicators of missing data as its input and output at least synthetic data wherein synthetic data is generated for the missing data; a second network model configured to receive a combination of real and synthetic data as its input and classify the input data as real or synthetic, wherein the second network model is arranged to receive, as at least part of its input, the synthetic data generated by the first network model; wherein the training of the network comprises applying a penalty to the second model for incorrectly classifying a synthetic variable as real and applying a penalty to the second model for classifying generated synthetic data as real.
[0068] The synthetic data generator may be trained to generate synthetic data per input variable group.
[0069] According to a fifth aspect, there is provided an apparatus comprising a processing resource configured to perform the method of any of the first, second, third or fourth aspect. According to a sixth aspect, there is provided an apparatus comprising a processing resource configured to: receive or otherwise obtaining input data for a subject with suspected or confirmed pre-eclampsia, wherein the input data comprises data representing a plurality of input variables comprising clinical data variables and / or other input variables associated with the subject and / or healthcare setting, wherein the input variables are grouped into a plurality of variable data groups, wherein the method comprises: process the input data to identify each of the plurality of variable data groups as one of: empty, partially complete or complete; perform a data completion process to complete the input data for the one or more partially complete variable group; combine the data of the complete variable group and the completed data of the partially complete variables groups to form combined data; perform a model selection process based on the received or otherwise obtained input data to select at least one model from a plurality of trained models; applying the selected model to at least some of the combined data to obtain a risk level and / or probability associated with one or more adverse maternal outcomes for a subject with suspected or confirmed preeclampsia.
[0070] The apparatus may further comprise a user input device configured to receive user input data representing values for input variables and a display for displaying the risk prediction. The apparatus may further comprise a user input device configured to receive user input data representing values for input variables. The apparatus may comprise a display for displaying the risk level and / or probability.
[0071] According to a seventh aspect, there is provided a computer program product comprising computer-readable instructions that are executable to receive or otherwise obtaining input data for a subject with suspected or confirmed pre-eclampsia, wherein the input data comprises data representing a plurality of input variables comprising clinical data variables and / or other input variables associated with the subject and / or healthcare setting, wherein the input variables are grouped into a plurality of variable data groups, wherein the method comprises: process the input data to identify each of the plurality of variable data groups as one of: empty, partially complete or complete; perform a data completion process to complete the input data for the one or more partially complete variable group; combine the data of the complete variable group and the completed data of the partially complete variables groups to form combined data; perform a model selection process based on the received or otherwise obtained input data to select at least one model from a plurality of trained models; applying the selected model to at least some of the combined data to obtain a risk level and / or probability associated with one or more adverse maternal outcomes for a subject with suspected or confirmed pre-eclampsia.
[0072] According to an eight aspect, there is provided an apparatus comprising a processing resource configured to perform the method of any of the first, second, third or fourth aspect.
[0073] Features in one aspect may be applied as features in any other aspect, in any appropriate combination. For example, features of the first aspect may be provided as features of the second to eight aspect and vice versa. For example, method features may be provided as apparatus features and / or computer program product features or vice versa.
[0074] Brief Description of Drawings
[0075] Various aspects of the invention will now be described by way of example only, and with reference to the accompanying drawings, of which:
[0076] Figure 1 is a schematic diagram of a data processing system in accordance with an embodiment;
[0077] Figure 2 is a flowchart of a method of obtaining a risk prediction in accordance with an embodiment;
[0078] Figure 3 is a schematic diagram showing in overview a method of obtaining a risk prediction, in accordance with an embodiment;
[0079] Figure 4 is a flowchart showing a workflow of training a synthetic data generator;
[0080] Figure 5 is a flowchart showing a workflow of training a procedure for obtaining a risk prediction;
[0081] Figure 6 is a workflow of using a method according to an embodiment,
[0082] Figure 7 is a table outlining detail of data used for training;
[0083] Figure 8 is part of a results table; Figure 9 is a table illustrating combinations of variable groups;
[0084] Figure 10 is a table of adverse maternal outcomes.
[0085] Detailed Description
[0086] Missing data in a development phase of predictive modelling is a widely discussed topic. Multiple imputation, is a method that may minimise bias introduced from imputation very efficiently provided there is a sufficiently large amount of data, strong relationships between variables, and the data is imputed a large number of times (for example, at least the percentage of missingness). When dealing with large datasets with missing data, multiple imputations may be easy to use and can be very effective. However, such an approach is not suitable when using a single observation with one or more missing values to obtain a prediction. In general, missing data at the model deployment stage (i.e. during prediction) is less discussed.
[0087] In healthcare settings where clinical resources may be scarce, one approach to missing data is to use a model trained to operate with as few input variables as possible, without a significant drop in performance. In such circumstances, after a set of variables is selected and the model trained, the trained model requires complete data for the selected set of variables. The following embodiments provide a method that can operate in the presence of incomplete and / or unavailable input data. In some embodiments, even if a user only has access to data variable that are less predictive than an optimum case, a prediction can be made, together with an assessment of uncertainty.
[0088] The following embodiments relate to a multi-step hierarchical classification tool trained using machine learning. The tool may minimise the amount of imputation required in the process. Embodiments make use of the currently available data, handle missing information without heavily relying on multiple imputation to be deployable without the development dataset, make a risk prediction using this subset of data and, in some embodiments, the method quantifies the uncertainty around the prediction.
[0089] Data collected for modelling adverse maternal outcomes of pre-eclampsia may be categorised as follows: baseline patient information, medical history, signs, symptoms, clinical assessment, and laboratory tests. Data from different categories are not all available at the same time, some are available immediately on admission while others become available over time. Generally, baseline patient information and medical history are available immediately on admission through patient records, although access to patient records is not always available in low resource settings. Signs and symptom data may be recorded along with a clinical assessment given that the patient is conscious and a healthcare professional is available to perform the assessment. Therefore, variables in these categories may be available at different times depending on waiting times, and the condition of the patient and also the clinical environment. Variables in the laboratory test groups may become available later than all other groups, as they require a sample to be taken and analysed. Additionally, patients with pre-eclampsia are treated in low-, middle- and high-income countries and are treated both as in- and outpatients in each setting, leading to limited access to certain tests in some cases and different waiting times for results.
[0090] Figure 1 is a schematic diagram depicting a computing apparatus 10 for performing a method of obtaining a risk level and / or a probability associated with one or more adverse maternal outcomes for a subject with suspected or confirmed preeclampsia, in accordance with embodiments.
[0091] Figure 1 depicts a computing apparatus 12. The computing apparatus 12 has a processing resource 14, one or more data storage resources 16, a display screen 18 and an input device 20.
[0092] It will be understood that while Figure 1 depicts a single computing apparatus for the purposes of the following description, the computing apparatus may be a distributed computing apparatus. For example, the computing apparatus 12 may comprise two or more computing apparatuses over a network. Likewise, the one or more data storage resources may be distributed data storage resources, for example, data may be retrieved from databases over a network.
[0093] The computing apparatus 12 comprises a processing resource 14. In the present embodiment, the processing resource 14 comprises a Central Processing Unit (CPU). For the purposes of the following description, the processing resource 14 has training circuitry 22, data completion circuitry 24, model selection circuitry 26 and prediction circuitry 28. The prediction circuitry may also be referred to as classification circuitry. While each of these circuitries is depicted in the processing resource 14, it will be understood that in some embodiments, one or more of these circuitries may be provided as part of a further processing resource, for example, of a network connected computing resource. In particular, it will be understood that training may be performed on a further computing resource. In some embodiments, the training circuitry may be implemented on a dedicated processing unit, for example, a GPU. It will be understood that, in some embodiments, processing steps can be performed on a combination of computer processing units (CPUs) and graphics processing units (GPUs) and other dedicated processing units. As an example, prediction including classification data processing may be performed separately to other data processing steps, such as training, data completion and model selection. As such, a separate computing apparatus, for example, a portable computing device, may be provided to receive user input and perform the prediction steps based on the received user input.
[0094] The circuitries may also be referred to, in some embodiments, as modules such that the apparatus has a training module configured to train a model or procedure, data completion module configured to perform a data completion process, model selection module configured to perform a model selection process and prediction module configured to apply a prediction model.
[0095] In the present embodiment, the various circuitries of the processing resource 14 are each implemented in the processing resource 14 by means of a computer program having computer-readable instructions that are executable to perform the method of the embodiment. However, in other embodiments each circuitry may be implemented in software, hardware or any suitable combination of hardware and software. In some embodiments, the various circuitries may be implemented as one or more ASICs (application specific integrated circuits) or FPGAs (field programmable gate arrays).
[0096] The computing apparatus 12 also includes a hard drive and other components including RAM, ROM, a data bus, an operating system including various device drivers, and hardware devices including a graphics card. Such components are not shown in Figure 1 for clarity.
[0097] Turning to the data storage resources 16, for the purposes of the description, Fig. 1 depicts the data storage resources 16 having a number of different storage modules. The data storage resource 16 may include a hard drive or other suitable memory resource, for example, a network connected database. It will be understood that the data storage resources 16 may be a network distributed storage resource.
[0098] As described in the following, the data storage resources 16 stores training data for use during training, data processing. The data storage resource may also store data used during prediction data processing steps. The data storage resources can also provide a storage resource for trained model data, in some embodiments, for example, trained model parameters and model architecture parameters. In some embodiments, one or more of these types of data has a corresponding data storage resource and the processing resource is configured to perform the required data operations on the corresponding resource as required.
[0099] The display screen 18 may also be referred to as the display, for brevity. The input device 20 is configured to receive user input data representative of a user input. The user input and display may be considered to form a user interface. The display screen 18 may display a graphical user interface for a user to interact. The input device may also be referred to as a user input device. The display screen and input device may together form a user interface.
[0100] It will be understood that while Figure 1 depicts an apparatus for training a procedure and for applying the trained procedure, in some embodiments, the training process will be performed separately to the application of the trained procedure. In such embodiments, learned procedure parameters, such as weights and / or coefficients may be retrieved from local storage and / or over a network or any suitable data communication interface.
[0101] Figure 2 depicts a workflow of a method of obtaining adverse maternal outcome information for a subject.
[0102] At step 202, input data is obtained from a user. The input data represents values for a set of input variables associated with a subject. Generally, data for making predictions may be categorised as baseline patient information, medical history, signs, symptoms, clinical assessment, and laboratory tests. In the following embodiments, the input variable are grouped into the following variable groups of Table 1.
[0103] Table 1
[0104] In the embodiment of Figure 2, the display screen 18 displays prompts for the user input parameters. The display screen displays a graphical user interface, for example, via an internet browser or other suitable interface. In some embodiments, one or more controllable display elements, such as drop down menu elements and / or toggle element are presented on display screen. In some embodiments, the display screen 18 may display a form for providing numerical input. In the present embodiment, user input data is generated in response to user input being received at input device 20. The user input data is representative of values for the input variables. For example, in embodiments, in which a selected set of input variables is used, the user input data may represent values for the set of input variables. In other embodiments, the input may be provided in the form of a suitable data file. The user input may be provided in any suitable format, such as numerical input.
[0105] The arrangement of input variables into groups is based on the cohort (for example the group of symptoms, or the group of medical history) and, for example, clinical input (separating coagulation blood tests from other blood tests). As described in further detail, for example, with reference to Figure 3, the received user data may include one or more missing entries. As such, the received user data may be partially complete. It will be understood that the received input data is representative of data available at the time of receiving the user input, for example, representing clinical observations for a single subject. As such, the user input data may include one or more missing values depending on a number of factors.
[0106] The input data is received via an input device 20. For example, an interface is presented to the user via display screen 18 asking the user to input values for a number of input variables. This set of input variables is pre-determined and input data is received via input device 20. In the present embodiment, the displayed user interface includes an entry box or other display element for each variable listed in Table 1 .
[0107] As step 204, a data completion process is performed to complete the data. In the present embodiments, the data completion process includes an initial step of processing the input data to identify each of the plurality of variable groups listed above as one of: empty, partially complete or complete. For any identified empty groups, the group is discarded. For any identified partially complete (or partially empty) groups the data is subject to a completion process. A data completion process, is described with reference to Figure 3 and 4. The completed groups are combined with the initially complete groups to form combined data. The combined data is then used, at step 208 to obtain a prediction.
[0108] At step 206, a model or machine learning procedure is selected from a library of trained models. The model selection process is described in further detail with reference to Figure 3. In some embodiments, the model selection process is based on at least the input data obtained at step 202. In some embodiments, the model selection process is based on the identification of the variables groups as empty, partially complete or complete.
[0109] At step 208, the selected model is applied to the combined data set to obtain a prediction. The trained machine learning procedure operates on the received input data and calculates an output. In this embodiment, the machine learning procedure is a trained classifier as described with reference to Figure 2 and the output is a classification of risk level. In this embodiment, the output is displayed on display screen 18.
[0110] In some embodiments, in case where the data entered by the user contained missing values, obtaining the risk prediction includes the further step of determining the uncertainty of the prediction.
[0111] At this step, the method returns to the original data with missing values (i.e. the input data received from the user) and the variables that have been used in the previously selected model are isolated for uncertainty quantification. For example, input variables that form the complete and reduced data set are isolated. The input variables that were missing are used to determine an uncertainty. In the present embodiment, a range of values for the missing variables are calculated. In some embodiments, this range corresponds to values with 2.5 and 97.5 quantile values, however, it will be understood that a different ranges may be used.
[0112] For a single missing variable, the range of values are provided to the selected model together with the other values of the original input data to obtain an upper and lower bound for the predicted probability. In detail, at least two values of the range are combined with the original input data to obtain at least two predictions. For example, the lower value of the range for the missing variable is determined and combined with the original data and provided to the selected model to obtain a first prediction and then the higher value lower value of the range for the missing variable is determined and combined with the original data and provided to the selected model to obtain second prediction. The first and second prediction provide a range of predictions. A third and / or further prediction may be obtained using one or more intermediate values in the range.
[0113] If more than one variable is missing a value, a data grid is constructed for all missing input variables. The data grid includes all possible combinations of replacements for the missing data. These combinations are combined with the original non-missing data. Predictions are then generated for each combination by providing the combination to the model to obtain a corresponding grid of predictions. The highest and lowest probabilities are then selected as the output range. By providing an uncertainty range, insights into the range of potential outcomes, accounting for both conservative and optimistic estimates can be made, thereby supporting informed decision-making.
[0114] In addition to obtaining a range of predictions, step 208 may also include an assessment of the impact of missing values have on the output prediction. In some embodiments, recommendations are provided to the user. In some embodiments, an evaluation of alternative models from the library that are not available due to lack of data for one or more input variables is performed to determine if including data for further input variables can improve a performance metric, for example, an accuracy, of the prediction.
[0115] In some embodiments, an evaluation of the impact of a missing input variable is performed. For example, in some embodiments, a range of predictions can be determined by varying values of the missing input variable in a range about the imputed value, which would be decreased if a value for that input variable was obtained. Such an evaluation may be limited to impact of variable from partially complete groups or may be extended, in further embodiments, to assess impact from empty variable groups. A recommendation may be made based on the evaluation and displayed on display. The recommendation may include a recommendation to obtain further data that is missing from the initial data.
[0116] At step 210, a visualisation of the predicted risk is displayed to the user. The predicted risk may include a risk level, optionally an uncertainty. A comparisons with average predicted risks of patients at the same gestational age in weeks may also be created and displayed. Text summaries of the risk, the missing values and their impact and detail regarding the ranking of models and the models selected may also displayed for the user.
[0117] Figure 3 illustrates, in further detail, the data completion process and the machine learning procedure library, in accordance with an embodiment.
[0118] As described with reference to Figure 2, the input variables are arranged into a number of groups. As described above, data may be missing for a number of variables. The input data is illustrated by partially complete data 302. The partially complete data set is formed by a number of data blocks. The data collected from the user input for each variable group is referred to in the following as a data block or block. In this embodiment, the input variables are grouped into eight groups and the data for each group is in a respective data block such that each block is formed of data for a group of variables. The data blocks include a first data block 302a, second data block 302b, third data block, 302c, fourth data block 302d, fifth data block 302e, sixth data block 302f, seventh data block 302g and eight data block 302h. In the present embodiment, the eight blocks correspond to the variable groups of Table 1.
[0119] As described above the input data is data for a single subject and may have one or more missing values. As such the input data may be characterised by a degree of missingness. The input data therefore is associated with one or a number of data availability scenarios. In the present embodiment, each data availability scenario corresponds to a combination of groups. There are 128 data availability scenarios. In embodiments, the input data is matched to one of the plurality of data availability scenarios and the data is completed to obtain a complete data set for that data availability scenario. The combination of groups making up the data scenarios for the present embodiment are depicted in Figure 9.
[0120] In the embodiment of Figure 3, a dashed smaller box, such as 304, indicates missing data. Missing data may be due to lack of resources. A larger dashed box, such as 306, indicate a partially complete block, with at least one missing value. In the embodiment of Figure 3, the first (302a), second (302b) and fifth blocks (302e) are complete blocks. These completed blocks have data for each of their variables and have no missing values. The third (302c), sixth (302f) and seventh (302g) are partially filled or populated blocks and have one or more missing values. The fourth (302d) and eight (302h) blocks are empty blocks and include no data values.
[0121] In the present embodiment, the first group 302a is a baseline variable group. While the method can be performed for missing data, data for each variable of the baseline group is required to be present for the method. The baseline variable group represents variables for which data will be available in almost all clinical environments. As such, a preliminary check that the baseline group is complete may be performed, and if the baseline group is not complete, the method may display an indication that the procedure may not continue to the user and / or that the results may not be reliable.
[0122] As described with reference to Figure 2, a data completion process is performed. In the present embodiment, at a first step, each block is identified as one of empty, partially complete and fully complete.
[0123] For each identified empty group, the group is discarded and / or not used when combining. In the illustrated example of Figure 3, block 302d and 302h are empty and discarded.
[0124] In a second step, for each group that is identified as partially complete, synthetic data is generated to complete and the data blocks. The synthetic data is generated using a data imputation process. A number of data imputation processes may be used, for example, the data imputation process described with reference to Figure 4.
[0125] In the illustrative example of Figure 3, the third (310c), sixth (31 Of) and seventh (310g) data blocks are identified as partially complete and synthetic data is generated to complete these blocks. In particular, any missing values are populated. In the illustration of Figure 3, synthetic data is indicated by a patterned box, such as box 310c.
[0126] Following the discarding and data completion of empty and partially complete blocks, respectively, the initially complete data blocks (first block 302a, 310a and fifth block 302e, 31 Oe) are combined with the completed blocks (second block 310b, third block 310c, sixth block 31 Of and seventh block 310g) to form combined data set 310. The fourth and eight block are discarded and do not form part of the combined data. The combined data is then used as input to a machine learning procedure for obtaining a prediction. The combined data can be understood as a reduced or truncated and complete data set.
[0127] A plurality of pre-trained models are available for making prediction using the data. The model library includes a plurality of pre-trained models. Each pre-trained model is trained on a particular combination of completed data blocks. In this embodiment, as there are eight data blocks including the baseline block, there are 128 possible combinations of blocks. The plurality of pre-trained models are stored for retrieval, for example, in a model library 312. The model library 312 is illustrated in Figure 3 as having a plurality of models for selection, represented by circles. In the present embodiment, each model is stored together with a pre-determined performance metric that indicates its performance.
[0128] It will be understood that a number of different performance metrics may be used to rank the filtered models. For example, the performance metric may comprise at least one of a positive or negative likelihood ratio, a number of patients and outcome rate, an area under a precision-recall curve or an F1 value.
[0129] In some embodiments, a particular combination of variable groups may have more than one model available. For example, for a given combination of variable groups, there may be at least a random forest based and a logistic regression model available for selection.
[0130] The selection of model is based on the variable group combination identified previously. As an example, if only the baseline group is available, and all other groups are empty, then only one model is available for selection. At the other extreme, if none of the groups having missing data so that none are discarded, then all 128 models are available for selection. In that case, the model with the best performance will be selected for the prediction. In the present example represented in Figure 3, 5 blocks including baseline are complete.
[0131] It will be understood that the identification of the variable group combination for the combined data may be determined using a number of different methods. As a first example, the variable group combination may be identified at the same stage as identifying empty blocks so that the empty blocks are discarded and the variable group combination is determined based on non-empty blocks. As second example, the variable group combination may be determined at a later stage by processing the combined data and determining the variable group combination from the combined data.
[0132] At a first step the models are filtered depending on the data availability. In particular, the filtering of the models is based on the variable group combination identified above. In further detail, in the present embodiment, the model is filtered based on the partially complete and complete variables groups. In some embodiments, the model is filtered based on the completed data blocks. The result is a filtered set of models that is compatible with at least part of the combined data. In this sense, compatible can mean that each model of the filtered models is configured to provide a prediction based on that at least part of the combined data. In the present embodiment, the models are filtered to remove one or more models that are trained to output a prediction using data from an identified empty group.
[0133] In further detail, each model is trained to output data based on complete data sets with no missing data. The model library has at least one trained model for each possible combination of variable groups and the filtered set comprises at least one trained model for each possible combination of identified partially complete and complete groups.
[0134] The result of the filtering is a filtered set of models 314 (indicated by the dashed line). The filtered set of models have a corresponding pre-determined performance metric such that the filtered set of models can be ranked.
[0135] As an example, for input data that is completed such that the combined data for groups 2, 5, 6 are completed (either initially or due to date completion) corresponding to data scenario 39, then the filtered list of models includes models configured to be applied to and / or trained on data for the following combinations of variable groups: baseline group (1 combination), baseline group plus any of group 2, 3 or 5 (3 combinations), baseline group plus any two of group 2, 5 and 6 (3 combinations), baseline group plus all of groups 2, 5 and 6 (1 combination) which results in 8 combinations in total. In embodiments in which one model is available for each combination, the filtered set of models includes 8 models.
[0136] Based on the ranking, a model is selected. In the present embodiment, the selected model 316 is the highest ranking model of the filtered models 314.
[0137] In the present embodiment, the selected model is selected based on a ruling-in or ruling out criteria. The criteria may be selected by a user when providing the initial user input, i.e. via the interface. The ruling in criteria corresponds to a focus on accurately predicting patients in a high and very high risk groups, even if it decreases the accuracy of the low and very low risk groups). The ruling out criteria focuses on accurately predicting patients into the low and very low risk groups, even if it decreases the accuracy of the high and very high risk groups) the outcome.
[0138] In further detail, for a ruling in criteria to rule in an outcome, the listed models were ranked based on decreasing positive likelihood ratio and decreasing percentage of patients classed in the very high risk group, and the highest ranking model was selected. This ensures that model were selected that could identify the most patients at the highest risk of an adverse outcome.
[0139] In some embodiments, likelihood ratios are used to determine risk strata were data- defined, based on likelihood ratios, as follows: very-low risk (by a negative likelihood ratio <0.10), low risk (negative likelihood ratio of 0.1 to 0.2), high risk (positive likelihood ratio of 5.0 to 10 10.0), very-high risk (positive likelihood ratio greater than 10.0), and moderate risk otherwise. In some embodiments, positive likelihood ratios for very high and high risk are calculated by splitting the testing data into "very-high risk" and "not very high risk", and "high risk" and "less than high risk" groups, then calculating likelihood ratios for a two group prediction using sensitivity and specificity. Similarly, negative likelihood ratios for very low and low risk are calculated by creating "very-low risk" and "not very low risk", and "low risk" and "higher than low risk" groups.
[0140] In some embodiments, a positive likelihood ratios may be understood as the ratio of the probability that a person with the disease tested positive divided probability that a person without the disease tested positive. A negative likelihood ratio may be understood as the ratio of the probability that a person with the disease tested a person with the disease tested negative to the probability that a person without the disease tested negative.
[0141] For ruling out criteria, the ranking was based on increasing percentage of outcomes in the very low risk group and decreasing percentage of patients in the very low risk group, and the highest ranking model was selected. This method was chosen over the negative likelihood ratio as models could have very low negative likelihood ratios by only classifying 1 patient with no outcome into the very low risk group. Ensuring that a high percentage of patients were classified into the group with a very low outcome rate, while still making sure that the negative likelihood ratio was acceptably low may give better confidence in the system’s utility for ruling out the outcome.
[0142] In some embodiment, the selection of the model is performed via a look-up table. In some embodiment, the selection of the model is performed at an earlier stage of the process, for example, on initial processing of the input data. In particular, as the input data indicated the variables that are missing and hence the data availability scenario and the ruling in or ruling out criteria, the selection of the model can be performed at that stage. For example, in some embodiments, each variable group combination has a stored top ranking model for ruling in and for ruling and the model is retrieved based on the determined variable group combination.
[0143] The selected model forms part of a machine learning procedure 318. Input data 320 formed from the combined data 311 is provided to the machine learning procedure 318 to obtain a risk prediction 322.
[0144] As shown in Figure 3, the selected machine learning derived procedure 318 is for obtaining a risk prediction 322 for one or more adverse maternal outcomes for a subject with suspected or confirmed preeclampsia. In the present embodiments, the machine learning derived procedure can be considered as a set of learned rules and / or instructions for operating on input data. In the present embodiments, the machine learning procedure corresponds to a set of rules and / or a procedure for classifying a set of received input data representative of a plurality of parameters as a risk level, where the risk level of the occurrence of one or more adverse maternal outcomes for a subject with suspected or confirmed preeclampsia. The risk level may be a risk level from one of a group of classes: very low, low, moderate, high, very high. In the present embodiment, the machine learning derived procedure is a machine learning classifier, and may be referred to, for brevity, as a classifier.
[0145] In the above described embodiments, generation of synthetic data is described to complete one or more partially complete data blocks. Figure 4 is a schematic diagram of training of a synthetic data generator, specifically a general adversarial imputation network (GAIN) suitable for completing at least part of an incomplete data set, in accordance with an embodiment. Figure 4 shows a first artificial neural network 402 configured to generate synthetic data. The first network is referred to as a generator or a generator network. Figure 4 shows a second artificial neural network 404. The second artificial neural network is referred to as a discriminator or discriminator network. The discriminator is configured to receive data and classify the data, specifically data for each data variable, as real or synthetic.
[0146] Training is performed by introducing a penalty term for both networks. The first ANN 402 is penalized if generated synthetic data is classified as real by the second ANN model 404. The second ANN model 404 is penalized for incorrectly classifying a synthetic variable as real.
[0147] In further detail, a real data set 406 is obtained. The real data set comprises clinical and / or data for a plurality of observations. In this embodiment, the real data for each observation is a complete and non-reduced data set (i.e. no missing data is present).
[0148] Samples of a number of observations (in this embodiment 200 observations), were obtained. Data for a number of data variables are then removed to turn the batch data into an incomplete data set 407. Missing data indicators 408 are generated to indicate that a particular variable is missing. These can be binary 0 and 1 values indicating missing or present, respectively.
[0149] The incomplete data set 407 is provided to the first generator and synthetic data values are generated for any missing variable, as indicated by the missing data indicator. In the incomplete data, the missing values are replaced by “noise”, this and the missing indicator (a matrix of 0s and 1s, where 1 represents a missing value and 0 represents an observed value) are provided to the generator.
[0150] The generated synthetic data (also referred to as fake data) is combined with the real data sample to form a complete and partially synthetic data sample. The complete and partially synthetic data sample is provided to the discriminator network 404. The discriminator network 404 receives the data sample and generate labels for each variable. The labels represent prediction for whether each data value is real or synthetic. The output is therefore a labelled data sample 410. To train the network, feedback is provided for both the generator network and the discriminator network simultaneously. As set out above, training is performed by penalizing both networks. The first ANN 402 is penalized if generated synthetic data is classified as real by the second ANN model. The second ANN model is penalized for incorrectly classifying a synthetic variable as real.
[0151] As a result of the training, a trained synthetic data generator 402 is provided. To use the trained GAINs for imputation of new data, an R function was created which, moving through each variable group one by one, would first select the rows of data with either all variables observed or all variables missing and return these rows without any changes. A GAIN was then applied to each incomplete data group to impute missing values(s), thus only imputing variable groups if at least one variable of the group was observed otherwise imputation would be based on random noise only. If all values are missing, all values are replaced by random noise, so the generated output is not based on any actual observations.
[0152] It will be understood that a plurality of GAINs are trained in accordance with embodiment. As described with reference to Figure 3, the variables are grouped into variable groups or blocks and a GAIN is created and trained for each variable group to allow synthetic data to be generated per group thereby to complete the data entries for each variable group or block.
[0153] For the base variable group, only the base variables were used for the corresponding GAIN. For each other group, the variables of the group together with the variables of the base variable group are used.
[0154] Further information on how to create and train the GAINS is provided in the following. To create the GAINs, first training datasets for each were created (in this embodiment, using software program R) by selecting variables as specified above and creating a subset of the data by keeping complete observations only. A second dataset was also created for each by simulating missing values for each variable by randomly deleting 20% of values per variable. GAINs were created for each variable group using a combination of R and python software. The dimensions and variables can be varied. In one embodiment, each network is a 28 dimensional network with two hidden layers of 28 nodes.
[0155] During training of the GAINs, samples of 200 observations (a batch) were taken from the dataset with generated missingness, which was the batch size of the original GAIN code, then the generator was used to replace the missing values. The completed dataset was then supplied to the discriminator which labelled the data points as real or fake data.
[0156] Without limitation, an examples of the loss function for the generator 402 is a follows:
[0157] In the above equations, X is the original, complete dataset of size NrowxNcohM is a matrix the same size of X with values of 0 and 1 to indicate missing values, Icontis a vector of 0s and 1s indicating continuous variables, X is the generated dataset and is the matrix of labels generated by the discriminator for X based on the hint matrix H. LM1is the sum of MSE of missing values in numeric columns and LM2is the sum of cross entropy of missing values in the categorical variables. The generator and the discriminator were then both backpropagated using the calculated losses, the generator is trained to minimise Gioss, while the discriminator is trained to maximise D0SS, and losses were recorded to to track performance over training instances. The process was then repeated with a new sample 3000 times to allow the loss functions to converge.
[0158] To use the trained GAINs for imputation of new data, an R function was created, which, moving through each variable group one by one, would first select the rows of data with either all variables observed or all variables missing and return these rows without any changes, then take the remaining data and apply the corresponding GAIN to impute missing values(s), thus only imputing variable groups if at least one variable of the group was observed, otherwise imputation would be based on random noise only.
[0159] In embodiments, the input for each GAIN is a dataset of 1 row and N columns (where N is the number of baseline variables + number of variables in the variable group), with at least 1 missing value, and the output is a dataset of 1 row and N columns, with no missing values, and variables that had an observed value in the input have the same value in the output.
[0160] In the embodiments described in the following, input data is received from a user, for example, a clinician, via a user interface and represent clinical and / or other data for a subject with suspected or confirmed pre-eclampsia. In embodiments, the data provided to the machine learning procedure does not include all the data received from the user. In particular, the data is a complete data set.
[0161] A notable distinction between a GAIN and a typical GAN structure, is the input to the generator. While a GAN generator is provided with noise as an input, the generator of a GAIN has access to further information in the form of 0 / 1 indicator values of missingness, and the observed values in the dataset with missing data. By considering this extra information, GAINs may better model and determine the underlying patterns and dependencies in the dataset, enabling imputed values more appropriate to each individual sample, rather than values which could belong to any sample in the dataset. Once trained, the generator (G) observes some components of a real data vector, imputes the missing components conditioned on what is actually observed, and outputs a completed vector Furthermore, the model of the machine learning procedure 318 or the procedure itself is selected from a plurality of models or procedures. The selection of model is dependent on the data availability of the input data received from a user and / or the combined data formed by proceeding the user data. Such embodiments offer advantages for example, these embodiments may allow performance to be improved in the case where not all data is available, for example, in healthcare settings where obtaining certain types of subject data may be more problematic.
[0162] The testing and validation datasets were used for model assessment. For each variable group combination, GAINs per variable group were used on the testing and validation datasets to impute missing data. These GAINs were only used for their corresponding variable groups if at least one variable in the group was observed, thus not all missing values were imputed. Moreover, using the same GAIN on the same data will always create the same prediction, therefore multiple imputation of the training and validation datasets was not needed, and only one imputed dataset was created per variable group combination. Models corresponding to the variable group combination were used on the complete cases in the training and validation datasets. As models were created on multiple imputed datasets, in most cases of variable group combinations, more than 1 model was created, so to make predictions on a new, unseen dataset, each of the corresponding models were used to make predictions, the mean of which was returned as the final predicted probability.
[0163] Thresholds for the risk classification groups were determined on the testing dataset, and patients in the validation dataset were classified into risk groups accordingly. On the validation dataset the positive likelihood ratio, and the number and percentage of outcomes for the very high risk group was recorded and later used to assess the ability of the models to rule in outcome. Similarly, the negative likelihood ratio as well as the number and percentage of outcomes was recorded for the very low risk group for assessment for the ability to rule out the outcome. Other rule in and rule out criteria based on other risk categories may be used.
[0164] As described in further detail in the following, the machine learning derived models or procedures available for selection from the model library are obtained by performing a training process on one or more data sets. The machine learning derived procedures may refer to or be dependent on a trained machine learning model and the machine learning derived procedure may comprise applying the trained model to input data. The machine learning derived procedure may be dependent on a set of trained model weights or model coefficients and the procedure may comprise a set of rules or instructions for combining the trained model weights or model coefficients to input data.
[0165] Figure 5 depicts a flow-chart of a method of training the plurality of models in the model library.
[0166] At step 502, training data was obtained. In the present embodiment, data was used from 11472 patients combined from 9 studies, made up from 7 studies, data of 2901 patients from the Fetal Medicine Foundation, and a small dataset of 208 patients from Brazil. The Fetal Medicine Foundation (FMF) data was collected from electronic health records as part of a prospective observational cohort UK women with singleton pregnancies, diagnosed with pre-eclampsia. Data was collected between December 2013 and December 2021 at King’s College Hospital, London, and Medway Maritime Hospital, Gillingham. A breakdown of the adverse maternal outcomes can be found at Figure 10.
[0167] The Brazil dataset came from a validation study of the fullPIERS model, a cross- sectional study of women diagnosed with pre-eclampsia admitted to the Women’s Hospital at the University of Campinas, Brazil between January 2017 and February 2018 (159). The aim of the study was to validate the fullPIERS model in a new, unseen cohort, the study included all women admitted for childbirth who were diagnosed with pre-eclampsia. Only 27 patients had an adverse maternal outcome in this dataset, all within 2 days of admission. The most severe outcome recorded was one case of maternal death, while the most common adverse outcomes were eclamptic seizure, blood transfusion and acute renal failure (8 occurrences each), followed by 5 placental abruptions, 2 cases of hepatic dysfunction.
[0168] Data from the studies were combined into a single dataset. Only data from day of or day after admission were used. For each patient, the worst value of each variable with repeated measurements was taken. Data was split into development (2 / 3 data), testing (1 / 6 data) and validation (1 / 6 data) data. As described above, the 35 predictor variables were grouped into baseline always available information plus 7 variable groups: patient data, medical history, signs, symptoms, blood tests part 1 , blood test part 2 (the measures of coagulation) and urine. These variables were grouped together as they were expected to be mostly likely to be available at the same time - i.e. if one symptom variable was available, it is likely all symptom variables would be available, and collected simultaneously. The coagulation variables were separated from the other blood test variables as these variables are less likely to be collected routinely in low resource settings. The variables are shown in Table 1 above.
[0169] At step 504, the training data is divided in to a plurality of training data sets, a development, testing and validation data set was created for each possible combination of the 7 variable groups (128 in total).
[0170] At step 506, for each combination, the variables belonging to the selected variable group(s), the baseline variables and the outcome variable were selected, and all other variables were treated as missing, thereby artificially creating every possible data availability scenario.
[0171] After the selection of variables, for each combination in the development dataset, M, = proportion of missing in the dataset corresponding to the ith variable group combination, rounded to the nearest percent was calculated. At step 506, each development dataset was imputed M, times using multiple chained random forests. Multiple imputation was not performed on the testing and validation datasets.
[0172] At step 508, models were trained for each variable group combination. In the present embodiment, logistic regression, random forest and LASSO methods were used for modelling. Models were fitted using tidymodels in R. Tidymodels are specified using a “model recipe”, which describes the steps of pre-processing and model formula; model specification, including the computational engine to be used for modelling (a package or software), the mode of modelling (regression or classification) and any additional model options or parameters; and workflow consisting of adding the model specification to the recipe and fitting the model. While the present embodiments describes using logistic regression, random forest and LASSO based models, it will be understood that alternative models may be used, for example, ridge regression, artificial neural networks and Bayesian Model averaging.
[0173] The following modelling specifications and steps were completed for each variable group combination. The model recipe was the same for logistic regression and LASSO methods, consisting of the steps of updating the ID variable to have an ID role and thus not be used as a predictor variable, creating dummy variables with 0 and 1 values of categorical variables, normalising numerical variables and finally modelling the outcome variable by all variables except ID. While other functions used in R to create logistic regression models, such as Im or glm, do not require this step of creating dummy variables as part of the data pre-processing as the functions identify any categorical variables and create the dummy variables as part of the model fitting process, tidymodels includes this step in the pre=processing stage to allow for greater flexibility in modelling.
[0174] An extra step was also included to remove any variables with zero variance (in other words, variables that have the same value for all patients), however, during training, no variables met this condition.
[0175] The model recipe for random forest was much simpler since the method could easily handle categorical variables and variables with very different variances. The model recipe for random forest consisted of updating the ID variable to have an ID role and thus not be used as a predictor variable in modelling, followed by modelling the outcome variable by all other variables.
[0176] All model specifications included “classification” mode, which indicates that the outcome is a categorical variable rather than continuous. For logistic regression the glm engine was used, random forest was fitted with randomForest engine and the glmnet engine was used for LASSO.
[0177] At step 510, an additional step of simplifying models by selecting variables was perfomed. For each method, steps were taken to simplify models and create a final model per method that only included variables that were significant predictors of the outcome to minimise the number of variables required to make predictions. The exact steps of model simplification is dependent on the type of model.
[0178] First, logistic regression models were fitted using all predictor variables with the recipe and specifications above. A model was fitted on each imputed dataset. For all resulting models, the p-value for each predictor variable was obtained, associated with the observed T-statistic used to test the hypothesis that the corresponding regression term is non-zero. Variables that had a p-value<0.05 in at least one of the imputed datasets were kept as predictor variables for the final model, while the rest were removed. A new model was then fitted on each imputed dataset using only the remaining predictor variables, with the same specifications and recipe as before. For each variable group combination / , if M, was greater than 1 , and thus multiple imputed datasets were created for the combination, the resulting final model was a list of the final model from each imputed dataset. If M, was not greater than 1 , and thus the combination was imputed only once, the resulting final model was the single final model created on the imputed dataset.
[0179] Variable selection for random forest was carried out similarly. First, random forest models were fitted on each imputed dataset using all predictor variables, with the recipe and specifications above. For all resulting models, variable importance was calculated for each variable using mean decrease in Gini index. The average importance was also calculated per model. Variables that had an importance greater than the corresponding average importance in at least one of the imputed datasets were kept as predictor variables for the final model, while the rest were removed. A new model was then fitted on each imputed dataset using only the remaining predictor variables, with the same specifications and recipe as before. If the variable groups combination had more than one imputed dataset, the resulting final models were combined into a list.
[0180] LASSO automatically carries out variable selection as part of its fitting process as the coefficients of less significant predictors are shrunk to or close to zero, hence the extra step of variable selection was not necessary for this method. However, this method had to be tuned to find the optimal penalty value. Here extra information was provided in the model specification: within the glmnet engine, mixture was set the value of 1 as that specifies LASSO method (rather than ridge regression or elastic net), and at first penalty was set as 0. Bootstrapping was used to tune the penalty on a tuning grid with 50 levels. Best penalty was chosen based on highest AUC. Once penalty was tuned, model specification was changed to include the new penalty (using the same penalty on all imputed datasets) and models were refitted. If the variable groups combination had more than one imputed dataset, the resulting final models were combined into a list.
[0181] At step 512, for variable group combination for which more than one model was trained, a selection process is performed to select one of the models to be stored in the model library.
[0182] The result is a plurality of trained models forming the model library, for example the model library 312. There is at least one trained model provided per variable group combination.
[0183] In the above described embodiments, a table of input variables was provided, arranged in to variable groups. Further detail on the variables is provided in the table of Figure 7. The baseline variables include: national per capita gross domestic product (USD), national maternal mortality ratio (maternal deaths per 100,000 live births), maternal age at expected date of delivery (years), gestational age at eligibility (weeks), and systolic and diastolic blood pressure (mm Hg). The patient information variables include: ethnicity, parity and multiple or singleton pregnancy indicator. The medical history variables include: pre-gestational renal disease, pre-gestational diabetes, chronic hypertension, gestational diabetes on admission and any history of cigarette smoking. The symptoms variables include: nausea or vomiting, headache or visual disturbances, right upper quadrant or epigastric pain and chest pain or dyspnoea. The signs group variables include: height (cm), weight (kg) and oxygen saturation (SpC>2 - Oxygen saturation by pulse oximetry). The first blood test group includes: hematocrit (%), total leucocyte count (x109per litre), platelet count (x109per litre), mean platelet volume (fl_), serum creatinine (pmol / L), uric acid (pmol / L), aspartate transaminase (U / L), alanine transaminase (U / L), serum albumin (g / L) and lactase dehydrogenase (U / L). The second group of blood tests includes: international normalised ratio, fibrinogen (g / L) and activated partial thromboplastin time (seconds). The second group of blood tests includes: international normalised ratio, fibrinogen (g / L) and activated partial thromboplastin time (seconds). GDP per capita may be defined according to https: / / data.worldbank.org / indicator / NY.GDP.PCAP.CD. NMR per capita may be defined at https: / / databank.worldbank.org / reports.aspx?source=world-development- indicators#. In some embodiments, the most informative maternal mortality ratio (e.g. one of national or regional) may be used.
[0184] In the embodiments described above, the procedures include a classifier that classifies received input data in terms of a risk level is described. In the present embodiment, the output of the classifier is a risk level associated with a Delphi-derived composite outcome of maternal mortality or severe morbidity within two days of first assessment with pre-eclampsia. In the present embodiment, the outputted risk level is one of five risk levels: “very low”; “low”; “moderate”, “high”, “very high”. However, it will be understood that in other embodiments, the procedure is a machine learning derived procedure that can output alternative output. In other embodiments, the output from the procedure may be a predicted probability for a specific maternal outcome and / or a score or percentage representing the risk level.
[0185] In the present embodiments, the outcome is defined as the first occurrence of one or more of: maternal mortality or severe maternal morbidity within two days of first assessment for pre-eclampsia.
[0186] In some embodiments, the risk categories were determined based on likelihood ratios (positive or negative) from the observed risk in the testing dataset,
[0187] In the above-described embodiments, since many variable group combinations are subsets of others, it is possible to use not only the corresponding model on a variable group combination but also models corresponding to combinations that are subsets of it. For example, in the combination consisting of the signs and symptoms groups with the baseline group, naturally it is possible to use the model created on the signs and symptoms combination, the model created on the signs group, the model created on the symptoms group, and the model made for the baseline group. However, it is also possible that after variable selection, a model created on a completely different variable group combination, for example symptoms and medical history, would not use variables from all groups in the combination. In case of a model created on symptoms and medical history, a model could be made only using symptom variables, and thus would be useable on the combination of signs and symptoms. Because of the use of variable selection, the models may not use every variable in the variable group combination. The model could not use any variables from a variable group, even if it is present. So instead of checking that the variable group combinations on which the model was trained on are available, instead it may be checked if the model variables themselves are available.
[0188] Figure 6 depicts a method, in accordance with a further embodiment. The process begins with data entry into the application by a user, at step 702, where missing values may be present in the dataset.
[0189] At step 704, dataset with 1 row and columns for each variable are split into variable groups. At step 706, each group is checked to determine one of: a) if all values are missing (in which case the group is discarded), b) all values observed, then the data is used as is and c) partially complete groups are imputed and then used. At step 708, the imputation is performed using GAIN for each variable group. The GAINs are applied on a per variable group basis, where if at least one variable in a group has observed values, the corresponding GAIN is used to impute missing values for that group. The imputed values are recorded, and information about the available variables is stored.
[0190] Next, at step 710, the models are filtered to retain only those in which all variables are available, ensuring that no missing values exist. From the filtered models, the highest ranked model is selected at step 712 based on the ruling in or ruling out criteria, described above, as selected by the user. This model is then utilized to predict outcomes on the GAIN-imputed dataset, incorporating the imputed values.
[0191] At step 709, the model are ranked by performance. As described above, this may be performed prior to the method.
[0192] At step 710, in cases where the data entered by the user contained missing values, the next step returns to the original data with missing values, and the variables used in the previously selected model are isolated for uncertainty quantification. If any of the variables used in the model contain missing values, a lower and upper bound for the predicted probability by replacing the missing values with the 2.5 and 97.5 quantile values calculated from the observed training data. These ranges are provided as an example and may be modified in further embodiments. If more than one variable is missing a value, a data grid is constructed, considering all possible combinations of replacements for the missing data.
[0193] Predictions are generated for each combination, and the lowest and highest predicted probabilities are selected and reported, at step 712 and 716. This reporting strategy provides insights into the range of potential outcomes, accounting for both conservative and optimistic estimates, thereby supporting informed decision-making.
[0194] Visualisations of the predicted risk, as well as comparisons with average predicted risks of patients at the same gestational age in weeks are created and returned on the user interface of the application, at step 718. Text summaries of the risk, missing values and ranking of models are also displayed for the user.
[0195] At step 720, a patient consultation is performed to support cinical decision-making and patient awareness. At step 722 and 724, the workflow of Figure 6 moves to an inpatient / induction stage and a decision may be made.
[0196] Figure 7 is a table showing the characteristics of data for subjects with and without outcome at any point in the total combined dataset with no imputation. Of the observed values, there was a significant difference between the group of patients in all baseline variables, patient information and symptom variables, variables in the second part of blood tests, and urine dipstick. All but one blood test variables were significantly different between patients with and without an outcome at any point.
[0197] In the combined dataset of 11472 patients, all variables in the patient information variable group were missing for 3057 (26.7%) patients, symptoms were all missing for 1930 (8.6%), at least one variable for medical history was recorded for all patients, blood tests part 1 was all missing for 1265 (11 .0%), blood tests part 2 for 6364 (55.5%), and urine for 3486 (30.4%).
[0198] For 994 (8.7%) patients some, but not all variables in the patient information group were missing, similarly some but not all variables were missing in the symptoms, medical history, signs, and bloods part 1 and part 2 were missing for 1930 (16.8%), 3444 (30.0%), 4611 (40.2%), 8258 (72.0%), and 1414 (12.3%) patients, respectively.
[0199] It will be understood that each data scenario has a corresponding set of useable models. The selected model of the useable models (i.e. that with the best performance metric) is selected depending on the ruling on or ruling out scenario.
[0200] For completeness, a table showing all combinations of variable groups and their correspondence with data scenarios is shown in Figure 9. It will be understood that a check mark indicate presence of a group in a particular data scenario. It will be understood that baseline variable group is present in all 128 scenarios in accordance with embodiments.
[0201] A skilled person will appreciate that variations of the enclosed arrangement are possible without departing from the invention. Accordingly, the above description of the specific embodiments are made by way of example only and not for the purposes of limitations.
[0202] Without being bound by theory, further comments on the modelling process are provided, for completeness.
[0203] Figure 8 depicts part of a table showing model performance and ranking of each model in the testing dataset corresponding. Figure 8 shows performance metric and ranking for the first 10 models. As an example, the performance of model 13, which was built on the development dataset corresponding to scenario 13 (baseline data, patient information and the first part of blood test results), when applied to the testing dataset corresponding to scenario 13. Some models may have a missing ruling in or ruling out performance indicating that no threshold could be determined for that model to classify patients into a very high or a very low risk stratum, respectively. Models reaching the desirable performance on ruling in and / or ruling out are also highlighted in the table, ruling in performance of a model is highlighted if the positive likelihood ratio was at least 10, and highlighted in yellow if the positive likelihood ratio was less than 10 but no less than 9. Similarly, the ruling out performance of a model is highlighted in green if the negative likelihood ratio was no more than 0.1 , and highlighted in yellow for values greater 0.1 but no more than 0.2. The table depicts a data scenario corresponding to the combined group of complete / complete data (i.e. after discarding any empty variable groups). The table shows the selected method (i.e. model type of one of logistic regression, random forest and LASSO) and the selected model number. The third column depicts the likelihood ratio. For brevity only the first 10 of 128 scenarios are shown.
[0204] Alternative tables were also constructed for each modelling method. These show the performance of models selected based on the hierarchy for ruling in and ruling out. These tables lists the models that would be used for each scenario, determined by availability of all model variables in the variable group combination and the existence of ranking for ruling in / out of the model. The table also shows the selected model from the useable model for ruling in / out, and its performance on the testing dataset corresponding to the given scenario. For example, if models 1 , 2, 5 and 13 could be used for scenario 13, but model 5 was selected for ruling in, then the table would show the performance of model 5, which was built on the development dataset corresponding to scenario 5, when applied to the testing dataset corresponding to scenario 13. The scenarios where the positive likelihood ratio on ruling in is at least 10, and / or the negative likelihood ratio on ruling out is at most 0.1 are highlighted in green. If no models could be found for ruling in and / or ruling out, the corresponding ruling in and / or ruling out performance is highlighted in red.
[0205] For each model-type (logistic regression, random forest and LASSO) performance of that model type for rule in and rule out the composite maternal outcome for each variable group combination was calculated. The following comments on logistic regression modelling are provided.
[0206] Risk strata thresholds were selected on a testing dataset to achieve desired likelihood ratios, as previously described. If a threshold for the very high risk group could be found such that the positive likelihood ratio on the testing dataset was at least 10, this threshold was used to determine the ruling in performance on the validation dataset, shown in the table. If a threshold could not be found, the ruling in performance for that model is left blank in the table. The ruling out performance was determined similarly. It was possible to classify patients into a very high risk group with 46.1% (59 of 128) models. In the validation datasets, 33.9% (20 out of the 59) models achieved a positive likelihood ratio (LR+) greater than 10, while 10.2% (6 of 59) had LR+ between 9 and 10. Variable group combinations with larger number of variable groups had less models which could rule in the outcome, however a bigger proportion of the models had LR+>10 It should be noted that the top 5 ranked models each classified only a small proportion of the population into the very high risk groups, at or under 1% of the population each.
[0207] It was possible to classify patients into a very low risk group with only 61.7% (79 of 128) of the models. In the validation datasets, 10.1 % (8 out of the 79) models achieved a negative likelihood ratio (LR-) less than 0.1 , while 2.5% (2 of 79) had LR- between 0.1 and 0.2. All models classified a larger proportion of the population into the very low risk group than into the very high risk group, and there were less than 5 outcomes in the very low risk group in the validation dataset for most models, even if the LR- was greater than 0.2.
[0208] Based on the ruling in criterion for choosing any model for which all variables were available in each of the variable group combinations, 13 models were selected (models 1 , 8, 14, 15, 16, 18, 19, 23, 29, 32, 45, 51 , and 103). Of the chosen models, only one model had LR+<9 (model 1), four models showed LR+ values between 9 and 10, and the remaining 7 demonstrated LR+ values exceeding 10, indicating their effectiveness in ruling in the composite maternal outcome.
[0209] One of the selected 13 models was chosen in all 128 variable group combination, meaning that a prediction to rule in outcome could be made regardless of what information was provided, with the confidence of LR+>10 48.4% (62 of 128) of the time.
[0210] Seven models (1 , 18, 29, 44, 48, 60, and 90) were needed to cover all variable group combinations when selecting models based on the ruling out criterion for choosing any model for which all variables were available. Importantly, all of these models achieved LR-<0.1 on their corresponding validation dataset at model ranking, indicating that they can be used effectively to rule out the composite maternal outcome with LR-<0.1 .
[0211] The following comments on random forest are provided. It was possible to classify patients into a very high risk group with 50% (64 of 128) of the models. In the validation datasets, 12.5% (8 out of the 64) models achieved LR+>10, while 3.1% (2 of 64) had LR+ between 9 and 10. It should be noted that the top 5 ranked models each classified only a small proportion of the population into the very high risk groups, under 1% of the population each.
[0212] It was possible to classify patients into a very low risk group with only 8.6% (11 of 128) of the models. In the validation datasets, 10 of the 11 models achieved LR-<0.1 , while the last model had LR->0.2. It was found that there were not enough models able to classify into the very low risk group to look for any relationship between the number of models with desirable LR- and the number of variable groups in the combination.
[0213] Based on the ruling in criterion for choosing any model for which all variables were available in each of the variable group combinations, only four models were selected (models 34, 45, 73, and 68). Of these models, only three (45, 73 and 68) were also among the top 5 ranked model, which were also the only models of the chosen four with LR+>10 on their corresponding validation dataset. When tested based on the ruling in criterion, choosing the best model for each variable group combination, I could rule in an outcome for all combinations, however 9.4% (12 of 128) of combinations did not reach the LR+ of 10. As one of the selected 3 models was chosen in all 128 variable group combination, a prediction to rule in outcome could be made regardless of what information was provided, with the confidence of LR+>10 in most cases.
[0214] Only two models (models 45 and 105 were used in 75% (96 of 128) of variable group combinations when selecting models based on the ruling out criterion for choosing any model for which all variables were available. The outcome could not be ruled out in the remaining 32 variable group combinations. The models used were ranked number 1 (model 45) and number 12 (model 105), and only model 45 achieved LR-<0.1 on its corresponding validation dataset at model ranking. When tested using the ruling out criterion, the desired LR- for 26% (25 of 96) of the variable group combinations that had a selected model could be achieved, meaning that while the outcome for the majority of variable group combinations could not be ruled out, there was not the desired level of confidence.
[0215] The following comments on LASSO models are provided. It was possible to classify patients into a very high risk group with 39.1 % (50 of 128) models. In the validation datasets, 24% (12 out of the 50) models achieved a positive likelihood ratio (LR+) greater than 10, while 2% (1 of 50) had LR+ between 9 and 10. Variable group combinations with larger number of variable groups had less models which could rule in the outcome. Similarly to before, the top 5 ranked models each classified only a small proportion of the population into the very high risk groups, under 1 % of the population each.
[0216] It was possible to classify into a very low risk group with the most models, with 85.9% (110 of 128). In the validation datasets, 10.1% (50 out of the 110) models achieved a negative likelihood ratio (LR-) less than 0.1 , while 2.5% (3 of 110) had LR- between 0.1 and 0.2. Unlike with the other two methods, most but not all models classified a larger proportion of the population into the very low risk group than into the very high risk group, and all but 4 models had less than 5 outcomes in the very low risk group, even if the LR- was greater than 0.2.
[0217] Based on the ruling in criterion for choosing any model for which all variables were available in each of the variable group combinations, 11 models were selected (models 1 , 3, 7, 17, 29, 34, 54, 71 , 74, 90, and 105). Of these chosen models, only five had LR+<9 (models 17, 34, 71 , 74, and 105), one model (model 90) showed LR+ values between 9 and 10, and the remaining five demonstrated LR+ values exceeding 10, indicating their effectiveness in ruling in the composite maternal outcome. One of the selected 11 models was chosen in all 128 variable group combination, meaning that a prediction to rule in outcome could be made regardless of what information was provided, with the confidence of LR+>10 31.3% (40 of 128) of the time.
[0218] Ten models (models 1 , 3, 17, 40, 41 , 51 , 69, 73, 100, and 122) were needed to cover all variable group combinations when selecting models based on the ruling out criterion for choosing any model for which all variables were available. One of the models were chosen in all 128 variable groups and achieved LR-<0.1 , meaning that a prediction to rule out outcome could be made regardless of what information was provided.
[0219] A selection of models from all available model-types including was also combined. As the best model for each scenario was previously selected for each modelling method, only the performance of the best model per method was compared for each scenario. Altogether 12 models were needed to rule in the outcome in all 128 scenarios, eight logistic regression models (model 1 , 14, 15, 188, 23, 26, 51 , and 103), three random forest models (models 45, 68, and 73) and one LASSO model (model 3). As a model was selected in all 128 scenarios, the system was able to rule in an outcome within two days, regardless of what information was available. Moreover, nearly all scenarios, 87.5% (112 of 128), showed a positive likelihood ratio of at least 10, of the remaining 16 scenarios, in the six scenarios where logistic regression models 15 or 18 were used, the positive likelihood ratio was between 9 and 10, and only the 10 scenarios using logistic regression model 1 , or LASSO model 3, which were the simplest models, were positive likelihood ratios under 9 observed. Generally, better performance of models in scenarios with more variable groups available, with no scenarios with more than four available variable groups having a positive likelihood ratio under 10.
[0220] Without being bound by theory, the following comments regarding methods are provided. The 2017 NICE Evidence Review identified risk scoring as key to determining place of HDP care, particularly for preeclampsia. Their review of externally-validated prognostic models identified: (i) our fullPIERS (Pre-eclampsia Integrated Estimate of RiSk) and miniPIERS models, and (ii) PREP (Prediction of Risks in early-onset Preeclampsia) specifically at <34 weeks. In the development of these models, data with missing observations was used by imputing the missing values to create multiple complete datasets. Whilst a suitable approach for model development, the resulting tools all require complete data to be able to make predictions - that is, it is unusable for patients with incomplete information. To create a truly useful tool, in both high and low and middle income countries, an artificial intelligence (Al) multistep hierarchical classifier was created that pragmatically uses the data available to it at any point in time in a woman’s care to make the best estimate of maternal risk of an adverse outcome (with quantification of uncertainty around that estimate).
[0221] The method referred to as PanPIERS has synthesised data from sources, which either developed or validated the PIERS or mini PIERS tools and data from the Fetal Medicine Foundation which has been used to validate PIERS-AL
[0222] The resulting observations were combined into a single dataset. For each patient, the worst value of each variable with repeated measurements on day 1 or 2 was taken. Data was split randomly into sets for development (2 / 3 data), testing (1 / 6 data) and validation (1 / 6 data). The potential predictor variables were grouped into 7 blocks (also referred to as variable groups) sets according to their likely availability - (i) baseline “always available” information (Gestational age on admission, maternal age at expected delivery date, systolic and diastolic blood pressure, national maternal mortality rate and national per capita GDP (both provided by the model given country of use)) (ii) patient data (singleton / multiple pregnancy, ethnicity, parity), (iii) medical history (renal disease, chronic hypertension, pre-gestational diabetes, gestational diabetes in previous pregnancy, history of smoking) (iv) signs (SpC>2, height on admission, weight on admission), (v) symptoms (vomiting or nausea, right upper quadrant or epigastric pain, headache or visual disturbances, chest pain or dyspnoea), (vi) blood tests part 1 (total leucocyte count, platelet count, mean platelet volume, uric acid, hematocrit, serum creatinine, aspartate aminotransferase - AST, alanine transaminase - ALT, lactate dehydrogenase - LDH, serum albumin) and (vii) blood tests part 2 (international normalized ratio - INR, fibrinogen, activated partial thromboplastin time APTT), and urine (dipstick proteinuria).
[0223] For each of the 7 variable groups, at the time of using the tool, the data for all of the fields may be either missing, partially complete or complete. To build the hierarchical multistep classifier, the best prediction model given the 128 realistic combination of variables. In this way a hierarchy of prediction models is created, the best of which are then chosen to make up the classifier. The following steps are performed.
[0224] (1 ) for each possible combination of the 7 variable groups (block), datasets were created for development, testing and validation of the hierarchical classifier
[0225] (2) For each combination each combination in the development dataset, M, = proportion of missing in the dataset corresponding to the ith variable group combination was calculated. Each development dataset was imputed M, times using, for example, multiple chained random forests. Multiple imputation was not used on the testing and validation datasets.
[0226] (3) A prediction model was then built for each combination. Methods considered were logistic regression, random forest and LASSO. Model fit was assessed, and a final best fitting model selected for that combination.
[0227] (4) Given the best fitting models, predictions were then made using the testing data. For the test data, the best model from step (4) given the blocks of variables which data was available for, is selected and risk prediction is made. If a block had no data, it was not used. For blocks with at least one data entry recorded but other data was missing, pre-trained Generative Adversarial Imputation Nets (GAINS) were used to impute the data and then make the prediction.
[0228] (5) These model predictions were then split into risk groups - very low, low, high and very high. The very low threshold was set first as the highest predicted probability to obtain a negative likelihood ratio<0.1. The low threshold was set to obtain a negative likelihood ratio<0.2. The very high risk group was selected as the lowest probability threshold with positive likelihood ratio>10. The high risk group set as the lowest probability threshold that was less than the very high risk threshold and had a positive likelihood ratio>5. These thresholds may be varied depending on embodiments and models.
[0229] (6) The ability of each of the combination models was then assessed in its ability to accurately classify individuals according to their outcomes into each of the risk groupings using the validation data based on rule in or rule out criteria.
[0230] The data from all studies were compared and mapped across. As the studies were collected either as part of the primary PIERS / minPIERS studies or validation studies of these, they adhered to the same study protocol and definitions which allowed mapping of variables. Data were not missing equally between studies. Race was more likely to be missing in certain data (the fullPIERS cohort). Symptoms were rarely recorded within a day of first admission in the PETRA dataset and, except for right upper quadrant pain, in the PREP dataset. Blood pressure and oxygen saturation were missing for >70% on day of admission in the PREP data; however, blood pressure was often recorded after day of admission, with 90% of women in the PREP data having blood pressure recorded on or the day after of admission. The highest rate of missingness was reported for the laboratory tests, some tests missing entirely from some datasets, due to the study design. Fibrinogen, activated partial thromboplastin time, aspartate transaminase and albumin were rarely measured in the FINNPEC data. Mean platelet volume and albumin were rarely measured in the miniPIERS data. Total leucocyte count, mean platelet volume, and albumin were not measured, and fibrinogen and activated partial thromboplastin time were rarely measured in the PETRA data. Haematocrit and mean platelet volume were not measured in the PREP data. The fullPIERS data had the least amount of missing data.
[0231] While rates of missingness were different between datasets and not all datasets recorded all variables, it was assumed that patients within all cohorts were similar enough in their presentation for the data to be combined. The data is likely illustrative of “real-world” data missingness in clinical practice, emphasising the need for a tool that can make risk predictions with data missing.
[0232] The multipstep classifier was successfully coded in R studio and thresholds created according to the steps outlined in the methods. For each prediction method, logistic regression, random Forest and LASSO, 128 models were built and then ranked in terms of their ability to accurately classify risk groups in the validation data. The positive likelihood ratio, and the number and percentage of outcomes for the very high risk group was recorded and later used to assess the ability of the models to rule in outcome. Similarly, the negative likelihood ratio as well as the number and percentage of outcomes was recorded for the very low risk group for assessment of ability to rule out the outcome. When aiming to rule in an outcome, the models were ranked based on decreasing positive likelihood ratio and decreasing percentage of patients classed in the very high risk group, and the highest ranking model was selected. For ruling out, ranking was based on increasing percentage of outcomes in the very low risk group and decreasing percentage of patients in the very low risk group, and the highest ranking model was selected.
[0233] The performance of the classifier using the differing prediction methods to rule out / rule in outcomes, using the validation data is summarised here.
[0234] Logistic regression models
[0235] Looking at the rule in performance, it was observed that approximately 32% of the models achieved a positive likelihood ratio (LR+) greater than 10 on the validation dataset, indicating their effectiveness in ruling in the composite maternal outcome. It was noted that the selected logistic regression models had a relatively small number of patients classified as being in the very high-risk group, accounting for less than 1 % of the total patient population of the corresponding variable group combination’s validation dataset. However, each of these models demonstrated a significant proportion, equal to or greater than 50%, of patients who experienced the composite maternal outcome.
[0236] In terms of ruling out the composite maternal outcome, approximately 9% of the logistic regression models achieved a negative likelihood ratio (LR-) less than 0.1 during internal validation. Unlike ruling in the outcome, many of the selected models classed a much larger proportion of patients classified as being in the very low-risk group, some even over 10% of the total patient population of the corresponding variable group combination’s validation dataset. The outcome rates, however, were higher than expected, and while many were under 1 %, the highest observed outcome rate among these models was 2.2%.
[0237] Random Forest models
[0238] Looking at the rule in performance, it was observed that approximately 34% of the models achieved a positive likelihood ratio (LR+) greater than 10. Similarly to the performance of selected logistic regression models, all of the chosen models had a relatively small number of patients classified as being in the very high-risk group, accounting for approximately 1 % of the corresponding total patient population.
[0239] In terms of ruling out the composite maternal outcome, the performance of random forest models was found to be inadequate. Only three variable groups had models capable of effectively excluding the outcome, with two out of the three models achieving a negative likelihood ratio (LR-) less than 0.1 during internal validation.
[0240] LASSO models
[0241] During internal validation of ruling in the outcome, approximately 34% of the models achieved a positive likelihood ratio (LR+) greater than 10. Similarly to the ruling in performance of logistic regression models, the chosen LASSO models covered all possible combinations of variable groups. However, it was observed that all of these models had a relatively small number of patients classified as being in the very high- risk group, accounting for approximately 1% of the total patient population.
[0242] LASSO models performed best at ruling out outcome out of all methods, with approximately 44% of the models achieving a negative likelihood ratio (LR-) less than 0.1 during internal validation.
[0243] The chosen LASSO models covered all possible combinations of variable groups, and for all of the selected models, between 5% to 10% of the patients were classified as being in the very low-risk group. Most importantly, none of the patients in this group experienced the composite maternal outcome for any of the selected models. Selection of models for the classifier
[0244] Finally, for each of the variable group combinations, models were selected to rule, and rule out the outcome from the list of models including all models created with all modelling methods. LASSO was chosen most often both for ruling in and for ruling out, however random forest and logistic regression models were also selected in on some occasions. Only LASSO and logistic regression models were selected to rule out outcome, none of the three random forest models made it into the selection showing that this modelling method was not as appropriate to rule out an outcome based on this data than the regression-based methods.
[0245] Comments regarding the user interface are provided. The tool imbeds the multistep classifier tool in its backend and provides an interface for the user (maternity care provider) to input data for the patient. As in the multistep classifier, complete data is not required, the tool will make the best possible prediction, using the optimal model based on the data available. There are options available to prioritise either ruling in or ruling out disease. Upon entering the data, users are provided with a risk estimate for an outcome within 48 hours. The estimate is also stratified into risk groups and the uncertainty in the estimate, driven by the level of data missingness, is reflected in the output.
[0246] A multistep hierarchical classifier is described to estimate the risk of an adverse outcome in women diagnosed with pre-eclampsia within 48 hours of admission. By collating data from across 11 studies, the degree of missingness in these data for all possibly predictive covariates, demonstrated the need to create a tool which can flexibly use the data available to it at that point in time and in that care setting. Within variable groupings GAINS may be applied - a technique generally used in image analysis - to generate likely values for missing data and allow predictions to be generated.
[0247] Exploring the potential predictive modelling algorithmic approaches demonstrated that no one technique is fit for purpose - the algorithmic choice may vary dependent on the data available. The classifier may seamlessly uses the best algorithmic choice and model for the set of variables available by employing a ranking system based on achieving accurate classification to risk groups, which allow to vary dependent on a preference to rule in or rule out the potential adverse outcome. Estimates of risk are created, alongside uncertainty intervals, which may reflect how the risk estimate may change if more information were available.
[0248] In accordance with embodiments, one or more further steps may be performed using outputs of the model and / or a score obtained using output of the model. In accordance with embodiments, the output and / or further data derived from the output may be used to inform decisions and / or initiate further investigations.
[0249] In some embodiments, the method may include performing or deferring additional investigative steps based on the output of the model or further data obtained from the output. In embodiments, a recommendation to perform said steps may be displayed based on the output of the model.
[0250] For example, the further steps may include performing a blood test or medical scan, for example, an ultrasound obstetric scan. In embodiments, one or more samples or further scans are performed based on the output of the model. Performing said steps may be performed in response to the output of the model corresponding to or representing a higher degree of risk. As a non-limiting example, said steps may be performed based on the output representing a high risk or very high risk. As described above, in embodiments, those risk levels corresponding to a positive likelihood ratio above 5.0.
[0251] In some embodiments the recommendation may be to defer said further steps. Such recommendations may be displayed in response to the output corresponding to low or very-low risk. In embodiments, these risk levels correspond to negative likelihood ratios below 0.2.
[0252] In some embodiments, the output of the model or further score derived from said output is used to determine a recommendation. That recommendation may be displayed to the user on the display.
[0253] In some embodiments, a further decision or recommendation on admitting or not admitting a patient may be made based on the output. For example, a decision or recommendation to not admit a patient at a given point in time as the negative likelihood ratio (-LR) <0.2 is sufficient to rule out maternal risk over the subsequent 48 hours. Conversely, a decision to admit the woman at a given point in time as the positive likelihood ratio (+LR) >5.0 is insufficient to rule out maternal risk over the subsequent 48 hours.
[0254] In some embodiments, a decision or recommendation to initiate birth may be made based on the output. For example, a decision or recommendation to defer initiating birth may be taken as the negative likelihood ratio (-LR) <0.2 may be sufficient to rule out maternal and fetal risk over the subsequent 48 hours.
[0255] Furthermore, a decision or recommendation to transfer a patient to a different facility or level of care. For example, in certain countries this may be from a remote health facility to a mainland facility or from a primary health centre to a referral hospital.
[0256] In embodiments, a decision or recommendation to administer treatment and / or medication may be taken based on the output of the model. For example, a decision or recommendation to administer antenatal corticosteroids to a woman at high or very- high risk at a gestation <34 weeks and 0 days of pregnancy may be taken to reduce the risks of prematurity. As a further example, a decision or recommendation to initiate magnesium sulphate in a woman at high / very-high risk to prevent grand mal seizures of eclampsia may be taken, for example, based on positive likelihood ratio >5 for an adverse maternal event is associated with increased seizure risk. Conversely, the decision or recommendation to withhold magnesium sulphate in a woman at low / very- low risk, as her risks for eclampsia are minimal.
[0257] A decision or recommendation for further medical intervention may be taken based on the output of the model. For example, a decision or recommendation to induce or initiate birth in a woman at high / very-high risk for herself (for example, based on positive likelihood ratio >5 for an adverse maternal event). The decision or recommendation to initiate birth in a woman may be based on output representing a high / very-high risk for her fetus (based on positive likelihood ratio greater than 5 for an adverse maternal event associated with increased stillbirth risk). In embodiments, performance of maternity units may be compared using the output of the model, for example, to compare performance of any given maternity unit with its peers, by standardising and adjusting for baseline risk between sites.
Claims
CLAIMS:
1. A method comprising: receiving or otherwise obtaining input data for a subject with suspected or confirmed pre-eclampsia, wherein the input data comprises data representing a plurality of input variables comprising clinical data variables and / or other input variables associated with the subject and / or healthcare setting, wherein the input variables are grouped into a plurality of variable data groups, wherein the method comprises: processing the input data to identify each of the plurality of variable data groups as one of: empty, partially complete or complete; performing a data completion process to complete the input data for the one or more partially complete variable group; combining the data of the complete variable group and the completed data of the partially complete variables groups to form combined data; performing a model selection process based on the received or otherwise obtained input data to select at least one model from a plurality of trained models; applying the selected model to at least some of the combined data to obtain at least a risk level and / or probability associated with one or more adverse maternal outcomes for a subject with suspected or confirmed pre-eclampsia.
2. The method of claim 1 , wherein the variable groups are pre-determined based on one or more data availability criteria and / or a degree of similarity between variables of the data groups.
3. The method of any preceding claim, wherein the variable groups comprise at least one of: a baseline variable group, a patient data group, medical history group, signs group, symptom group, one or more blood test groups, urine test group4. The method of any preceding claim further comprising determining a measure of a loss in accuracy or other performance metric of the selected model due to missing data for one or more input variable, wherein the method further comprise performing at least one further data collection process to obtain data for said one or more input variables.
5. The method of any preceding claim wherein the data of an identified partially complete variable group comprises missing data and wherein the data completion process comprises generating synthetic data for the missing data and / or wherein the data of an identified empty group comprises no data and / or only missing data and / or wherein the data of an identified complete group comprises no missing data6. The method of any preceding claim, wherein the data for each partially complete variable group comprises at least one missing data value from the input data, wherein generating the synthetic data comprises replacing the at least one missing data value with at least one generated value for that variable group.
7. The method of any preceding claim wherein generating the synthetic data is independently performed for the data of each partially complete variable group8. The method of any preceding claim wherein generating the synthetic data comprises applying a general adversarial network to at least part of the received or otherwise obtained input data, optionally applying a general adversarial network to the data of each partially complete group.
9. The method of any preceding claim wherein generating the synthetic data comprise replacing missing data for one or more input variables with a plurality or range of generated values for that variable, wherein applying the selected model comprises obtaining a plurality and / or range of predictions for the missing data using the plurality or range of generated values10. The method of any preceding claim wherein the selection is based on at least a pre-determined performance metric associated with the model11 . The method of any preceding claim wherein selecting the model comprises filtering a plurality of models to obtain a filtered set of models based on the identified partially complete and complete variable groups and selecting a model from the filtered set of models, for example, based on a performance metric associated with the model.
12. The method of any preceding claim wherein the plurality of models comprises at least one trained model for each possible combination of variable groups and the filtered set comprises at least one trained model for each possible combination of identified partially complete and complete groups13. The method of any preceding claim wherein the selection process is further based on a ruling in or ruling out criteria, optionally, based on a selection by a user.
14. The method of any preceding claim, wherein each model of the plurality of models comprises at least one of: logistic regression, random forest, LASSO, ridge regression, artificial neural networks and Bayesian Model averaging.
15. The method of any preceding claim, wherein the selected model is trained to receive complete data for a corresponding combination of variable groups, wherein applying the selected model to the at least some combined data comprises applying the selected model to a subset of the combined data for that combination of variable groups16. The method of any preceding claim, wherein the method comprises determining a measure of a loss in accuracy or other performance metric of the selected model due to missing data for or one or more variables, optionally providing a recommendation for improvement in accuracy or other performance metric based on the determined measure17. The method of any preceding claim, further comprising displaying an interface for receiving input data and / or receiving the user input via an user input device and displaying output of the model via the interface18. The method as claimed in any preceding claim, wherein the adverse maternal outcome comprises at least one of: maternal death; an adverse central nervous system event; a cardiorespiratory event; a hematologic event; a hepatic event; a renal event or one or more of: placental abruption, severe ascites, bell’s palsy.
19. The method as claimed in any preceding claim, wherein the model is configured to classify the subject into one of a plurality of risk levels based on a probability of the occurrence of one or more adverse maternal events in a predetermined time period.
20. The method as claimed in any preceding claim, wherein obtaining input data may comprise at least one of: a) receiving user input data representative of the country and / or region of the health care system and retrieving the health care system data representing a value for at least one statistic associated with the country and / or region of the health care system; b) performing a vital sign measurement to obtain the vital sign data; c) receiving user input data representative of the demographic data; d) obtaining blood or other sample test data.21 . An apparatus comprising a processing resource configured to: receive or otherwise obtaining input data for a subject with suspected or confirmed pre-eclampsia, wherein the input data comprises data representing a plurality of input variables comprising clinical data variables and / or other input variables associated with the subject and / or healthcare setting, wherein the input variables are grouped into a plurality of variable data groups, wherein the method comprises: process the input data to identify each of the plurality of variable data groups as one of: empty, partially complete or complete; perform a data completion process to complete the input data for the one or more partially complete variable group; combine the data of the complete variable group and the completed data of the partially complete variables groups to form combined data; perform a model selection process based on the received or otherwise obtained input data to select at least one model from a plurality of trained models; applying the selected model to at least some of the combined data to obtain a risk level and / or probability associated with one or more adverse maternal outcomes for a subject with suspected or confirmed pre-eclampsia22. The apparatus of claim 21 , further comprising a user input device configured to receive user input data representing values for input variables and a display for displaying the risk prediction.
23. A computer program product comprising computer-readable instructions that are executable to: receive or otherwise obtaining input data for a subject with suspected or confirmed pre-eclampsia, wherein the input data comprises data representing a plurality of input variables comprising clinical data variables and / or other input variables associated with the subject and / or healthcare setting, wherein the input variables are grouped into a plurality of variable data groups, wherein the method comprises: process the input data to identify each of the plurality of variable data groups as one of: empty, partially complete or complete; perform a data completion process to complete the input data for the one or more partially complete variable group; combine the data of the complete variable group and the completed data of the partially complete variables groups to form combined data; perform a model selection process based on the received or otherwise obtained input data to select at least one model from a plurality of trained models; applying the selected model to at least some of the combined data to obtain a risk level and / or probability associated with one or more adverse maternal outcomes for a subject with suspected or confirmed pre-eclampsia.
Citation Information
Patent Citations
A risk prediction method for preeclampsia based on MLP multi-platform calibration
CN113724873B
Preeclampsia poor pregnancy outcome prediction method based on COX proportional risk model
CN116705314A
Maternal and infant health insights & cognitive intelligence (MIHIC) system and score to predict the risk of maternal, fetal, and infant morbidity and mortality
US20240038402A1