Spatiotemporal Attention for Clinical Outcome Prediction

The disease prognosis model using a recurrent neural network with spatiotemporal attention and feedforward components addresses the challenge of capturing both short-term and long-term dependencies in health data, enabling precise clinical outcome predictions for diseases like COVID-19.

JP2025532100APending Publication Date: 2025-09-29GENENTECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025517224
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-23
Filing Date
2023-09-22
Publication Date
2025-09-29

AI Technical Summary

Technical Problem

Existing approaches to predicting long-term clinical outcomes from longitudinal health data, such as electronic health records, struggle to accurately capture both short-term and long-term dependencies and feature importance, which is crucial for identifying high-risk cohorts for severe complications like chronic COVID-19 symptoms.

Method used

A disease prognosis model incorporating a recurrent neural network, spatiotemporal attention mechanism, and feedforward neural network to jointly weight feature importance across time and feature space, effectively capturing short-term and long-term dependencies in longitudinal health data.

Benefits of technology

The model accurately predicts clinical outcomes like cure, persistence, relapse, worsening, or death by identifying important time points and feature patterns, enhancing treatment planning and healthcare resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025532100000001_ABST
    Figure 2025532100000001_ABST
Patent Text Reader

Abstract

A disease prognosis model can be trained to determine a clinical outcome of a disease based on longitudinal data including health records for each time point in a sequence of time points. The disease prognosis model can include a recurrent neural network trained to extract, from each health record, a feature set representing local dependencies present in the health record. The disease prognosis model can include spatiotemporal attention trained to determine the importance of each feature in the feature set at each time point in the sequence of time points. The disease prognosis model can include a feedforward neural network trained to determine a clinical outcome of the disease based on the importance of each feature in the feature set at each time point in the sequence of time points. The trained disease prognosis model can be applied to determine a clinical outcome of the disease for one or more patients.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application No. 63 / 376,957, entitled "SPACETIME ATTENTION FOR CLINICAL OUTCOME PREDICTION," filed September 23, 2022, the disclosure of which is incorporated by reference in its entirety into this specification.

[0002]

[0002] The subject matter described herein relates generally to machine learning, and more specifically to machine learning-based prognostic models for predicting clinical outcomes of diseases. [Background technology]

[0003] Many diseases can lead to serious long-term complications. For example, infection with the novel coronavirus SARS-CoV-2 can cause coronavirus disease 2019 (COVID-19), characterized by clinical symptoms such as cough, headache, and fever. A significant number of patients infected with SARS-CoV-2 develop persistent sequelae after infection. Patients with COVID-19 may experience persistent COVID-19 symptoms, such as fatigue, shortness of breath, and memory impairment, for at least two months after the initial acute infection. In some cases, COVID-19-related symptoms may appear, recur, and persist for months or even years after the initial acute infection. Chronic COVID-19 symptoms can be life-threatening in the most severe cases. Therefore, identifying cohorts at high risk for severe long-term disease complications can aid in treatment planning and healthcare resource allocation. Summary of the Invention

[0004]

[0004] Systems, methods, and articles of manufacture, including computer program products, are provided for predicting clinical outcomes using machine learning. In one aspect, a system for predicting clinical outcomes using machine learning is provided. The system may include at least one processor and at least one memory. The at least one memory may include program code that, when executed by the at least one processor, provides operations. The operations may include training a disease prognosis model to determine a clinical outcome of the disease based on at least longitudinal data, the longitudinal data including health records for each time point in a sequence of time points, training the disease prognosis model including training a recurrent neural network, a spatio-temporal attention network, and a feed-forward neural network, wherein the recurrent neural network is trained to extract from each health record a feature set representing one or more local dependencies present in the health record, the spatio-temporal attention network is trained to determine an importance of each feature in the feature set at each time point in the sequence of time points, and the feed-forward neural network is trained to determine the clinical outcome of the disease based at least on the importance of each feature in the feature set at each time point in the sequence of time points; and applying the trained disease prognosis model to determine a clinical outcome of the disease for a patient associated with the first health record and the second health record based at least on the first health record from the first time point and the second health record from the second time point.

[0005]

[0005] In another aspect, a method for predicting clinical outcomes using machine learning is provided. The method may include training a disease prognosis model to determine a clinical outcome of the disease based on at least longitudinal data, the longitudinal data including health records for each time point in a sequence of time points, wherein training the disease prognosis model includes training a recurrent neural network, a spatio-temporal attention, and a feed-forward neural network, wherein the recurrent neural network is trained to extract from each health record a feature set representing one or more local dependencies present in the health record, the spatio-temporal attention is trained to determine an importance of each feature in the feature set at each time point in the sequence of time points, and the feed-forward neural network is trained to determine the clinical outcome of the disease based at least on the importance of each feature in the feature set at each time point in the sequence of time points; and applying the trained disease prognosis model to determine a clinical outcome of the disease for a patient associated with the first health record and the second health record based at least on the first health record from the first time point and the second health record from the second time point.

[0006] In another aspect, a computer program product for predicting clinical outcomes using machine learning is provided. The computer program product may include a non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operations to occur. The operations may include training a disease prognosis model to determine a clinical outcome of the disease based on at least longitudinal data, the longitudinal data including health records for each time point in a sequence of time points, training the disease prognosis model including training a recurrent neural network, a spatio-temporal attention, and a feed-forward neural network, wherein the recurrent neural network is trained to extract from each health record a feature set representing one or more local dependencies present in the health record, the spatio-temporal attention is trained to determine an importance of each feature in the feature set at each time point in the sequence of time points, and the feed-forward neural network is trained to determine the clinical outcome of the disease based at least on the importance of each feature in the feature set at each time point in the sequence of time points; and applying the trained disease prognosis model to determine a clinical outcome of the disease for a patient associated with the first health record and the second health record based at least on the first health record from the first time point and the second health record from the second time point.

[0007]

[0007] Some variations of the methods, systems, and non-transitory computer-readable media may optionally include one or more of the following features in any possible combination.

[0008]

[0008] In some variations, the recurrent neural network may be a bidirectional recurrent neural network (RNN), a long short-term memory (LSTM) network, a local long short-term memory (LSTM) network with a given window size for the time point, or a gated recurrent unit (GRU) network.

[0009]

[0009] In some variations, the feedforward neural network may be a multi-layer perceptron model.

[0010]

[0010] In some variations, the trained disease prognosis model can determine the clinical outcome of a disease for a patient associated with a first health record and a second health record by applying at least a trained recurrent neural network to extract from a first health record a first set of feature values ​​for a hidden feature set representing a first set of local dependencies present in the first health record, and applying a trained recurrent neural network to extract from a second health record a second set of feature values ​​for a hidden feature set representing a second set of local dependencies present in the second health record.

[0011]

[0011] In some variations, the trained recurrent neural network can output a feature map including a first set of feature values ​​from a first time point and a second set of feature values ​​from a second time point for incorporation by the trained spatiotemporal attention.

[0012]

[0012] In some variations, the trained disease prognosis model can further determine a clinical outcome of a disease for a patient associated with the first health record and the second health record by applying at least the trained spatiotemporal attention to determine the importance of each feature in the hidden feature set at each of the first and second time points based at least on a feature map including a first set of feature values ​​and a second set of feature values.

[0013]

[0013] In some variations, the trained spatiotemporal attention may include one or more two-dimensional convolutional layers trained to determine the importance of each feature in the hidden feature set across the time dimension and the feature dimension.

[0014]

[0014] In some variations, one or more two-dimensional convolutional layers may include a 1x1 convolutional filter configured to perform joint weighting of the importance of each feature in the hidden feature set across the time dimension and the feature dimension.

[0015]

[0015] In some variations, the trained spatiotemporal attention can determine, for a first feature from a hidden feature set, a first importance of the first feature at a first time point and a second importance of the first feature at a second time point.

[0016]

[0016] In some variations, the trained spatiotemporal attention can further determine, for a second feature from the hidden feature set, a third importance of the second feature at the first time point and a fourth importance of the second feature at the second time point.

[0017]

[0017] In some variations, the trained disease prognosis model may further determine a clinical outcome of the patient's disease associated with the first health record and the second health record by applying at least the trained feedforward neural network to determine the clinical outcome of the patient's disease based at least on the importance of each feature in the hidden feature set at each of the first and second time points.

[0018] In some variations, the disease prognosis model may be further trained to determine a clinical outcome of the disease based on non-longitudinal data. The non-longitudinal data may be static across a sequence of time points. The non-longitudinal data may be linked to health records associated with each time point in the sequence of time points.

[0019] In some variations, one or more missing values ​​of the non-longitudinal variables that make up the non-longitudinal data may be identified, and the one or more missing values ​​may be replaced with the mean value of the non-longitudinal variable observed in the available dataset.

[0020]

[0020] In some variations, the non-longitudinal data may include medical image data and / or electrogram data corresponding to a metric that quantifies the severity of a disease depicted in one or more medical images and / or electrograms.

[0021]

[0021] In some variations, the health record associated with each time point in the sequence of time points may include a value for each of a plurality of vital sign statistics.

[0022]

[0022] In some variations, the health record associated with each time point in the sequence of time points may include values ​​for each of a plurality of specimen test variables.

[0023] In some variations, the health record associated with each time point in the sequence of time points includes medical image data and / or electrogram data, the medical image data including metrics that quantify the severity of a disease depicted in one or more medical images and / or electrograms.

[0024] In some variations, the longitudinal data may be determined to include a missing value for a longitudinal variable at a first time point. The missing value can be replaced with (i) a first value of the longitudinal variable from a second time point that precedes the first time point, (ii) a second value of the longitudinal variable from a third time point that follows the first time point, or (iii) a third value determined based on the first and second values.

[0025]

[0025] In some variations, the disease may be coronavirus disease (COVID-19), Alzheimer's disease, or age-related macular degeneration.

[0026]

[0026] In some variations, the clinical outcome of the disease may include a probability associated with one or more of cure, deterioration, and death.

[0027]

[0027] Implementations of the present subject matter may include, but are not limited to, methods according to the description provided herein, as well as articles comprising tangibly embodied machine-readable media operable to cause one or more machines (e.g., computers, etc.) to perform operations that implement one or more of the described features. Similarly, computer systems are described, which may include one or more processors and one or more memories coupled to the one or more processors. The memory may include a non-transitory computer-readable or machine-readable storage medium and may include, encode, store, etc., one or more programs that cause the one or more processors to perform one or more operations described herein. Computer-implemented methods according to one or more implementations of the present subject matter may be executed by one or more data processors present in a single computing system or in multiple computing systems. Such multiple computing systems may be connected via one or more connections, including, for example, a connection over a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.), via a direct connection between one or more of the multiple computing systems, etc., and may exchange data and / or commands or other instructions, etc.

[0028]

[0028] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the following description. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. While certain features of the subject matter of the present disclosure have been described for illustrative purposes in connection with predicting clinical outcomes in the context of post-acute sequelae of COVID-19, it should be readily understood that such features are not intended to be limiting. The claims following this disclosure are intended to define the scope of the protected subject matter.

[0029]

[0029] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate certain aspects of the subject matter disclosed in this specification and, together with the description, serve to explain some of the principles associated with the disclosed embodiments. [Brief explanation of the drawings]

[0030] [Figure 1] FIG. 1 is a system diagram illustrating an example of a prognosis system, according to some exemplary embodiments. [Figure 2] FIG. 1 is a schematic diagram illustrating an example of spatio-temporal attention computing feature importance across time and feature dimensions, according to some exemplary embodiments. [Figure 3A] FIG. 1 is a schematic diagram illustrating an example of a disease prognosis model, according to some exemplary embodiments. [Figure 3B] FIG. 1 is a schematic diagram illustrating an example of a disease prognosis model, according to some exemplary embodiments. [Figure 4A] 1 is a flowchart illustrating an example of a process for training a disease prognosis model to perform clinical outcome prediction, according to some exemplary embodiments. [Figure 4B] 1 is a flowchart illustrating an example of a process for clinical outcome prediction using machine learning, according to some exemplary embodiments. [Figure 4C] 1 is a flowchart illustrating an example of a process for clinical outcome prediction using machine learning, according to some exemplary embodiments. [Figure 5A] 1 is a graph illustrating an example of classification of a patient's disease severity based on Acute Physiology and Chronic Health Evaluation II (APACHE II) scores over time, according to some exemplary embodiments. [Figure 5B] 1 is a graph illustrating an example of a patient's Acute Physiology and Chronic Health Evaluation II (APACHE II) score being broken down by various physiological variables, according to some exemplary embodiments. [Figure 5C] 10 is a table illustrating example outputs of spatio-temporal attention, according to some exemplary embodiments. [Figure 6] FIG. 1 is a block diagram illustrating an example of a computing system, in accordance with some illustrative embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0031]

[0041] Wherever practical, like reference numerals refer to like structures, features, or elements.

[0032]

[0042] Identifying cohorts at high risk for severe long-term morbidity can aid in treatment planning and healthcare resource allocation. In the case of COVID-19, identifying patients at high risk for severe long-term complications may be essential for timely medical intervention. Accurate prediction of long-term disease outcomes, such as long-term morbidity and mortality, requires the integrated evaluation of multimodal data, including static, non-longitudinal data, such as patients' initial condition at disease onset, and longitudinal data tracking disease progression. However, because of the diversity of patient phenotypes and chronic conditions, predicting patient outcomes from longitudinal data, such as a series of electronic health records, acquired at different time points can be challenging. In particular, longitudinal data exhibit a combination of short-term and long-term dependencies that evade traditional approaches to disease outcome prediction. For example, in patients hospitalized with COVID-19 pneumonia, fibrotic abnormalities commonly observed in the first 3 months largely disappear by 1 year, whereas consolidation typically resolves within 6 months. Therefore, in COVID-19, where symptoms may persist or recur, the feature importance of fibrosis should not only change over time but also differ from that of consolidation. However, previous approaches to disease outcome prediction can either recognize time-dependent feature importance in the absence of feature diversity, or provide spatial feature importance that is static in time.

[0033]

[0043] Thus, in some exemplary embodiments, a disease prognosis model for determining the clinical outcome of a disease can include a spatiotemporal attention mechanism that can jointly weight feature importance from the time dimension and feature space of longitudinal data. An example of longitudinal data is medical data in the form of health records (e.g., electronic health records (EHRs)) for each time point in a sequence of consecutive time points. As used herein, the term "health record" can refer to an assemblage of data across various modalities. For example, each health record can include values ​​for each of multiple vital statistics, such as systolic blood pressure, diastolic blood pressure, pulse rate, respiratory rate, etc. Alternatively and / or additionally, each health record can include values ​​for each of multiple specimen test variables. Examples of analyte test variables include fibrinogen, C-reactive protein, prothrombin international normalized ratio, prothrombin time, lactate dehydrogenase, D-dimer, albumin, ferritin, alanine aminotransferase, aspartate aminotransferase, chloride, protein, alkaline phosphatase, bilirubin, calcium, creatinine, glucose, hematocrit, hemoglobin, potassium, platelets, red blood cells, sodium, white blood cells, etc. In some cases, the health record associated with each time point in the sequence of time points may include medical imaging data associated with an x-ray, magnetic resonance imaging (MRI) scan, computed tomography (CT) scan, positron emission tomography (PET) scan, optical coherence tomography (OCT) scan, etc. Additionally, in some cases, each health record may also include one or more electroencephalogram (EEG), electrocorticogram (ECoG or iEEG), electrooculogram (EOG), electroretinogram (ERG), electrocardiogram (ECG), electromyogram (EMG), etc.

[0034]

[0044] In some exemplary embodiments, the disease prognosis model may include a recurrent neural network and a feedforward neural network coupled with a spatiotemporal attention mechanism. For example, the recurrent neural network may be trained to extract, from each health record included in the longitudinal data, a feature set representing one or more local dependencies present in the health record. The spatiotemporal attention may be trained to determine the importance of each feature in the feature set at each time point in a sequence of time points. Furthermore, the feedforward neural network may be trained to determine a clinical outcome of the disease based at least on the importance of each feature in the feature set at each time point in the sequence of time points. By combining the recurrent neural network and spatiotemporal attention, the disease prognosis model can effectively capture short-term and long-term dependencies present in longitudinal data, such as a series of health records (e.g., electronic health records) from different time points. In particular, the trained disease prognosis model may be applied to identify important time points and feature patterns for determining a clinical outcome of the disease, such as the probability of cure, persistence, relapse, worsening, and / or death from COVID-19.

[0035]

[0045] FIG. 1 is a system diagram illustrating an example of a prognosis system 100 according to some exemplary embodiments. Referring to FIG. 1, the prognosis system 100 may include a prognosis engine 110, a client device 120, and a data store 130. As shown in FIG. 1, the prognosis engine 110, the client device 120, and the data store 130 may be communicatively coupled via a network 140. The client device 120 may be a processor-based device, including, for example, a workstation, a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable device, etc. The data store 130 may be a relational database, an unstructured query language (NoSQL) database, an in-memory database, a graph database, a key-value store, a document store, etc. The network 140 may be a wired and / or wireless network, including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, etc.

[0036]

[0046] In some exemplary embodiments, the prognosis engine 110 may sequentially apply disease prognosis models 115 to longitudinal data 133 from the data store 130 to determine a clinical outcome of a disease for a patient associated with the longitudinal data 133. For example, in some cases, the longitudinal data 133 may be represented by a sequence {x i |x1,x2,…,x T}, where T denotes the length of the sequence (or the number of time points), and x i denotes the data at the i-th time point. In some cases, x i may be health records (e.g., electronic health records (EHRs)) represented as vectors, matrices, or tensors. Thus, the longitudinal data 133 may include, for example, a first health record x1 of a patient from a first time point t1 and a second health record x2 of the patient from a second time point t2. The disease prognosis model 115 may be based on the sequence {xi |x1,x2,…,x T}, the disease prognosis model 115 may be trained to determine a clinical outcome of a disease, including, for example, the probability of cure, persistence, recurrence, worsening, and / or death of a patient associated with the longitudinal data 133. In some cases, the clinical outcome prediction performed by the disease prognosis model 115 may be formulated as a sequence-to-one problem. For example, the disease prognosis model 115 may extract, from the longitudinal data 133, a feature set (e.g., hidden features) that represents short-term dependencies present in the longitudinal data 133. Furthermore, the disease prognosis model 115 may determine the importance of each feature in the feature set for each time point in the longitudinal data 133. The clinical outcome of the disease may be determined based on the importance of each feature in the feature set at each time point in the longitudinal data 133.

[0037]

[0047] The aforementioned sequence-to-one problem for disease outcome prediction can be challenging due to the presence of various features (e.g., disease symptoms) at different time points. In some cases, features (e.g., disease symptoms) present at earlier time points may be correlated with features present at later time points. Accurate prognosis prediction requires a comprehensive analysis of feature and temporal information, but existing approaches to outcome prediction can consider either time-dependent feature importance or spatial feature importance, but not both. In contrast, various embodiments of the disease prognosis model 115 disclosed herein include spatiotemporal attention 200 trained to jointly weight feature importance across time and feature space.

[0038]

[0048] 2 is a schematic diagram illustrating an example of spatio-temporal attention 200 that computes feature importance across time and feature dimensions, according to some example embodiments. As shown in FIG. 2, the feature importance of a feature f at any particular time point is calculated by multiplying the value of each feature f at all time points. ijIn some cases, spatio-temporal attention 200 may be implemented as one or more two-dimensional convolutional layers trained to calculate the key, query, and value of each feature extracted from longitudinal data 133 using a convolutional filter (e.g., a 1×1 convolutional filter) to calculate an alignment score as feature importance. The advantage of using a 1×1 convolutional filter for attention calculation is the joint weighting of feature importance from two dimensions, including the time dimension and the feature dimension. Equation (1) below shows that features (e.g., hidden features) extracted from longitudinal data 133 may be adjusted by spatio-temporal attention 200 with a weighting factor γ. f'=f+γ·a(f) (1) where f denotes the hidden features extracted from the longitudinal data 133 (as described in detail below), and f′ denotes the adjusted value of each feature determined by spatiotemporal attention a(·).

[0039]

[0049] The following equation (2) shows the calculation of the spatiotemporal attention a(·). TIFF2025532100000002.tif8170 where H denotes the size of the feature space extracted from the longitudinal data, and T denotes the number of time points in the longitudinal data 133. As shown in equation (2), the feature importance of a given feature f is calculated by the relative importance of other features {f ij |i=1,2,...,H;j=1,2,...,T}. In some cases, feature importance may be measured by calculating an alignment score via a key k( ), value v( ), and query q( ) operation performed by a convolution filter (e.g., a 1x1 convolution filter).

[0040]

[0050] As described above, the disease prognosis model 115 can extract a set of features (e.g., hidden features) from the longitudinal data 133 that represent short-term dependencies present in the longitudinal data 133. In some cases, the disease prognosis model 115 can include a recurrent neural network (RNN) trained to extract the set of features from the longitudinal data 133 before spatiotemporal attention 200 is applied to determine the importance of each feature at each time point in the longitudinal data 133. FIGS. 3A-B are schematic diagrams illustrating an example of a disease prognosis model 115 in which spatiotemporal attention 200 is integrated into a recurrent neural network 300. In some cases, the recurrent neural network 300 can be a bidirectional recurrent neural network (RNN), a long short-term memory (LSTM) network, a local long short-term memory (LSTM) network with a given window size for the time points, a gated recurrent unit (GRU) network, etc.

[0041]

[0051] Referring to FIG. 3A , in some cases, the recurrent neural network 310 may be a long-short-term memory (LSTM) network trained to extract short-term and long-term dependencies from the longitudinal data 133 to form a feature map 325. The size of the feature map 325 may be N×H, where N corresponds to the number of stacked recurrent layers and H corresponds to the number of hidden features in the LSTM network. A spatiotemporal attention 200 may operate on the feature map 325 to learn correlations among the H features across the T time points. As shown in FIG. 3A , the spatiotemporal attention 200 may output an adjusted feature map indicating the importance of each of the H features at each of the T time points for ingestion by the feedforward neural network 350. For example, in some cases, the adjusted feature map may indicate a first importance of a first feature at a first time point and a second importance of the first feature at a second time point. Additionally, in some cases, the adjusted feature map may indicate a third importance of a second feature at the first time point and a fourth importance of the second feature at the second time point. A feedforward neural network 350, which may in some cases be a multi-layer perceptron model, may operate on the adjusted feature map to determine a clinical outcome of the disease.

[0042]

[0052] Spatiotemporal attention200 may be adept at learning long-term dependencies but may lack the ability to order and model short-term (or local) dependencies. Some attention-based models, such as Transformers, rely on positional embeddings to encode the ordering of short-term dependencies. However, unlike images or text, where the order of elements has contextual meaning, health records (e.g., electronic health records (EHRs)) contained in longitudinal data133 are not strictly ordered. For example, a patient may undergo a lab test before undergoing a medical imaging test, and vice versa. Thus, positional embeddings are not well suited to encoding the ordering of short-term dependencies present in longitudinal data133. However, long-short-term memory (LSTM) networks do not impose a strict ordering, at least because signals stored in memory cells can propagate even when the local order changes. To decouple learning of short-term and long-term dependencies, learning of short-term dependencies may be restricted to a set of local long-short-term memory (LSTM) networks, and learning of long-term dependencies may be restricted to spatiotemporal attention 200. An example of a disease prognosis model 115 in which a recurrent neural network 115 is implemented as a set of local long-short-term memory (LSTM) networks is shown in Figure 3B. The local long-short-term memory (LSTM) networks are restricted to learning continuous patterns that exist within a certain window size and extracting the corresponding local patterns as hidden features.

[0043]

[0053] Referring again to Figure 3B, after concatenating the hidden states from each local long-short-term memory (LSTM) network, the stacked hidden states can become a matrix containing T vectors of length H. Spatiotemporal attention 200 then refines the hidden states by mining feature and long-term dependencies, outputting a set of refined hidden states that are then fed into a feedforward neural network 350 (e.g., a multilayer perceptron model) to determine clinical outcomes. In some cases, the set of local long-short-term memory (LSTM) networks and spatiotemporal attention 200 can form an R-Transformer. This differs from the Transformer models used in computer vision and natural language processing, where location information is encoded in location embeddings.

[0044]

[0054] In some exemplary embodiments, the prognosis engine 110 can apply the disease prognosis model 115 to a combination of the ordered longitudinal data 133 and the non-longitudinal data 135 to determine a clinical outcome of a patient's disease associated with the longitudinal data 133 and the non-longitudinal data 135. In this context, the non-longitudinal data 135 may include data whose values ​​remain fixed across different time points. Examples of the non-longitudinal data 135 include demographic information, medical history (e.g., pre-existing conditions such as hypertension, obesity, hyperlipidemia, diabetes, etc.), medical imaging data (e.g., grades of disease severity depicted on one or more medical images), and / or electrogram data (e.g., grades of disease severity indicated by one or more electrograms). The prognosis engine 110 can combine the longitudinal data 133 with the non-longitudinal data 135 by concatenating at least the non-longitudinal data 135 with the longitudinal data 133 from each time point. For example, non-longitudinal data 135 associated with a patient may be concatenated with a first health record x1 from a first time point t1 and a second health record x2 from a second time point t2.

[0045]

[0055] In some exemplary embodiments, the prognosis engine 110 may preprocess the longitudinal data 133 and / or non-longitudinal data 135 before applying the disease prognosis model 115. For example, preprocessing the longitudinal data 133 and / or non-longitudinal data 135 may include normalizing one or more of the values ​​present therein. For at least some medical images, the preprocessing performed by the prognosis engine 110 may include determining a metric that quantifies the severity of the disease depicted in the medical image. For example, in the case of chest x-rays, the prognosis engine 110 may calculate a Radiological Assessment of Pulmonary Edema (RALE) score to characterize the disease severity of acute respiratory distress syndrome (ARDS) in COVID-19-positive patients. In the case of numeric variables, the prognosis engine 110 may apply min-max normalization to render the values ​​of each variable on the same scale (e.g., from 0 to 1). The values ​​of binary variables present in the longitudinal data 133 and / or non-longitudinal data 135 may be represented as one of two values ​​(e.g., 0 and 1).

[0046]

[0056] In some cases, longitudinal data 133 and non-longitudinal data 135 may exhibit sparsity, where values ​​for one or more variables are missing. Thus, in some exemplary embodiments, preprocessing of longitudinal data 133 and / or non-longitudinal data 135 may include excluding from further analysis one or more variables that are available for a number of patients below a threshold (e.g., 95% of the patients). Additionally, if prognosis engine 110 detects that a patient's non-longitudinal data 135 contains one or more missing values ​​for a non-longitudinal variable (e.g., a variable with static values ​​across multiple time points), prognosis engine 110 may perform mean imputation, replacing the one or more missing values ​​with the mean value of the variable observed in the available dataset (e.g., the mean value of the variable associated with other patients). If the prognosis engine 110 detects that a patient's longitudinal data 133 includes a missing value for a longitudinal variable (e.g., a variable whose value varies between time points) at a first time point, the prognosis engine 110 can perform forward imputation, replacing the missing value with a first value of the longitudinal variable from a second time point preceding the first time point. Alternatively, the prognosis engine 110 may perform backward imputation, replacing the missing value with a second value of the longitudinal variable from a third time point following the first time point. In some cases, the missing value of the longitudinal variable at the first time point may be interpolated based on the first value of the longitudinal variable from a second time point preceding the first time point and the second value of the longitudinal variable from a third time point following the first time point.

[0047]

[0057] 4A is a flowchart illustrating an example of a process 400 for training a disease prognosis model 115 to perform clinical outcome prediction, according to some exemplary embodiments. With reference to FIGS. 1-2, 3A-3B, and 4A, the process 400 may be performed, for example, by the prognosis engine 110 to train the disease prognosis model 115.

[0048]

[0058] At 402, the prognosis engine 110 may train the disease prognosis model 115 by training at least the recurrent neural network 300, the spatio-temporal attention 200, and the feedforward neural network 350 included in the disease prognosis model 115. For example, in some exemplary embodiments, the disease prognosis model 115 may be trained to determine a clinical outcome of a patient's disease based on at least the patient's longitudinal data 133. As shown in FIGS. 3A-3B , the disease prognosis model 115 may include the recurrent neural network 330, the spatio-temporal attention 200, and the feedforward neural network 350. Furthermore, the longitudinal data 133 may include, for each time point in the sequence of time points, a corresponding health record. Thus, training the disease prognosis model 115 may include training the recurrent neural network 330, the spatio-temporal attention 200, and the feedforward neural network 350. For example, the recurrent neural network 330 may be trained to extract a feature set from each health record included in the longitudinal data 133 that represents one or more local dependencies present in the health record. The spatio-temporal attention 200 may be trained to determine the importance of each feature in the feature set at each time point in the sequence of time points. The feedforward neural network 350 may be trained to determine a clinical outcome of a disease based at least on the importance of each feature in the feature set at each time point in the sequence of time points.

[0049]

[0059] At 404, the prognosis engine 110 may apply the trained disease prognosis model 115 to determine a clinical outcome of a disease for one or more patients. In some exemplary embodiments, the trained disease prognosis model 115 may be applied to determine clinical outcomes of various diseases, including, for example, coronavirus disease (COVID-19), Alzheimer's disease, age-related macular degeneration, etc. The output of the trained disease prognosis model 115 may include a probability associated with one or more of cure, persistence, recurrence, worsening, and death as the clinical outcome of the disease.

[0050]

[0060] 4B is a flowchart illustrating an example of a process 430 for clinical outcome prediction using machine learning, according to some exemplary embodiments. With reference to FIGS. 1-2, 3A-3B, and 4A-4B, process 430 may be executed by prognosis engine 110 to implement operation 404 of process 400 shown in FIG. 4A.

[0051]

[0061] At 432, the prognosis engine 110 may receive the patient's longitudinal data 133. For example, in some exemplary embodiments, the prognosis engine 110 may receive the patient's first health record at a first time point and the patient's second health record at a second time point as part of the patient's longitudinal data 133. In some cases, the patient's longitudinal data 133 may be combined with the patient's non-longitudinal data 135. The patient's non-longitudinal data 135 may include one or more non-longitudinal variables whose values ​​remain static across successive time points, while the patient's longitudinal data 133 may include one or more longitudinal variables whose values ​​fluctuate across successive time points. Thus, combining the patient's longitudinal data 133 and non-longitudinal data 135 may include concatenating values ​​of the longitudinal variables (e.g., included in the health record associated with each time point) with values ​​of the non-longitudinal variables at each time point.

[0052]

[0062] At 434, the prognosis engine 110 may apply the trained disease prognosis model 115 to determine a clinical outcome of the patient's disease based on at least the patient's longitudinal data 133. For example, in some exemplary embodiments, the trained disease prognosis model 115 may be applied to determine a clinical outcome of the patient's disease based on at least a first health record of the patient at a first time point and a second health record of the patient at a second time point. The disease prognosis model 115, including spatio-temporal attention 200 integrated with a recurrent neural network 300, may be able to identify short-term and long-term dependencies present in the longitudinal data 133. In particular, the recurrent neural network 300 may be trained to extract various features representing short-term dependencies present in the longitudinal data 133, while the spatio-temporal attention 200 may be trained to recognize changes in feature importance of each feature across different time points. In doing so, the disease prognosis model 115 may generate an accurate prognosis of the patient's clinical outcome of the disease.

[0053]

[0063] 4C is a flowchart illustrating an example of a process 450 for clinical outcome prediction using machine learning, according to some exemplary embodiments. With reference to FIGS. 1-2, 3A-3B, and 4A-4C, process 450 may be performed by disease prognosis model 115 to perform operation 434 of operation 430 shown in FIG. 4B.

[0054]

[0064] At 452, the disease prognosis model 115 may apply the recurrent neural network 300 to generate a feature map including a first set of feature values ​​for the hidden feature set representing a first set of local dependencies present in the patient's first health record from a first time point and a second set of feature values ​​for the hidden feature set representing a second set of local dependencies present in the patient's second health record from a second time point. In some exemplary embodiments, the recurrent neural network 300 may generate a feature map including a first set of feature values ​​for the hidden feature set representing a first set of local dependencies present in the patient's second health record from a second time point. i|x1,x2,…,x T}, for example, between a first health record x1 from a first time point t1 and a second health record x2 from a second time point t2. As shown in FIGS. 3A-3B, the recurrent neural network 300 can extract H features (e.g., hidden features) from each of the first health record x1 and the second health record x2. Accordingly, FIGS. 3A-3B further illustrate that the output of the recurrent neural network 300 can be a feature map 325 having a size of N×H. For example, the feature map 325 can include a corresponding vector of length H for health records from each of the T time points. If the recurrent neural network 300 is implemented as a local long-short-term memory (LSTM) network, the recurrent neural network 300 can be limited to learning continuous patterns that exist within a specific window size and extracting the corresponding local patterns as hidden features.

[0055]

[0065] At 454, the disease prognosis model 115 can apply spatio-temporal attention 200 to determine the importance of each feature at each of the first and second time points based on at least the feature map. In some exemplary embodiments, the spatio-temporal attention 200 can operate on the feature map 325 to determine correlations among the H features across the T time points and output a corresponding adjusted feature map indicating the importance of each of the H features at each of the T time points. For example, the spatio-temporal attention 200 can determine a first importance of a first feature at the first time point t1 and a second importance of the first feature at the second time point t2. Furthermore, the spatio-temporal attention 200 can determine a third importance of a second feature at the first time point t1 and a fourth importance of the second feature at the second time point t2.

[0056]

[0066] As mentioned above, the importance of different features may vary over time, consistent with clinical observations in COVID-19 patients, such as those in which fibrotic abnormalities commonly seen in the first three months largely disappear after one year, while consolidation typically disappears within six months. Therefore, the adjusted feature map may include adjusted feature values ​​for each of the first and second features. For example, the value of the first feature at the first time point t1 and the second time point t2 may be adjusted based on the first importance of the first feature at the first time point t1 and the second importance of the first feature at the second time point t2, respectively. Similarly, the value of the second feature at the first time point t1 and the second time point t2 may be adjusted based on the third importance of the second feature at the first time point t1 and the fourth importance of the second feature at the second time point t2, respectively.

[0057]

[0067] At 456, the disease prognosis model 115 can apply a feedforward neural network 350 to determine a clinical outcome of the patient's disease based on the importance of each feature at each of at least the first and second time points. In some exemplary embodiments, the disease prognosis model 115 can apply a feedforward neural network 350 that can operate on the adjusted feature map to determine a clinical outcome of the patient's disease. For example, the feedforward neural network 350 can be implemented as a multilayer perceptron model. Furthermore, the feedforward neural network 350 operating on the adjusted feature map can determine a clinical outcome of the patient's disease based at least on the importance of each feature at different time points in the longitudinal data 133.

[0058]

[0068] In some exemplary embodiments, the performance of the disease prognosis model 115 in predicting COVID-19 clinical outcomes was evaluated based on data associated with a cohort of 365 patients hospitalized with severe COVID-19 pneumonia. Non-longitudinal data 135 for each patient collected at initial hospitalization includes demographic information, medical history, and medical imaging data (e.g., radiological assessment of pulmonary edema (RALE), which quantifies disease severity reflected on chest x-rays). Longitudinal data 133 associated with each patient (including laboratory test results and vital signs) is collected at follow-up visits. The mean number of time points for each patient is 10, with a standard deviation of 6. Clinical outcomes, such as survival status, for each patient are collected 60 days after initial hospitalization. Table 1 below shows patient characteristics at initial hospitalization.

[0059]

[0069] Table 1 TIFF2025532100000003.tif186170

[0060]

[0070] Table 2 below shows the laboratory variables for the example patient on Day 1. As shown in Table 2, the laboratory variables may include fibrinogen, C-reactive protein, prothrombin international normalized ratio, prothrombin time, lactate dehydrogenase, D-dimer, albumin, ferritin, alanine aminotransferase, aspartate aminotransferase, chloride, protein, alkaline phosphatase, bilirubin, calcium, creatinine, glucose, hematocrit, hemoglobin, potassium, platelets, red blood cells, sodium, and white blood cells.

[0061]

[0071] Table 2 TIFF2025532100000004.tif245170TIFF2025532100000005.tif94170

[0062]

[0072] In some exemplary embodiments, the performance of the disease prognosis model 115 may be evaluated by training, validating, and testing the disease prognosis model 115 using a training set, a validation set, and a test set generated by splitting the aforementioned dataset in a 7:1:2 ratio. In some cases, training the disease prognosis model 115 includes stochastic gradient optimization, and validating the disease prognosis model 115 includes tuning hyperparameters of the disease prognosis model 115 on the validation set. For example, if the disease prognosis model 115 includes an R-Transformer formed by a set of local long-short-term memory (LSTM) networks and spatiotemporal attention 200, a window size of 6 may be used to restrict the local long-short-term memory (LSTM) networks to learning short-term dependencies. Furthermore, the size of the hidden feature set (e.g., the value of H) may be set to 32 in some cases. Using the tuned hyperparameters, the disease prognosis model 115 is trained for 50 epochs with a batch size of 2. The training of the disease prognosis model 115 may follow a learning rate schedule, such as an annealed learning rate that gradually decreases from 1e-3 to 1e-5. The disease prognosis model 115 that performs best on the validation set is evaluated on the test set.

[0063]

[0073] In some exemplary embodiments, evaluating the performance of the disease prognosis model 115 may include evaluating the prognostic value of different data modalities by incrementally incorporating them into the disease prognosis model 115. For example, spatiotemporal attention 200 for clinical outcome prediction may be interpreted by first visualizing a spatiotemporal feature map (e.g., varying feature importance across successive time points). The critical time points identified by spatiotemporal attention 200 (e.g., time points with the highest feature importance) may be compared with those identified by the Acute Physiology and Chronic Health Evaluation II (APACHE II) system (e.g., time points with the greatest increase in Apache II score). The latter is a clinical nomogram for quantifying disease severity. The Apache II measures physiological variables, age, and previous health conditions, assigning a score from 0 to 71, with higher scores indicating a higher risk of death. A Mann-Whitney U test is performed to compare the critical time points identified by spatiotemporal attention 200 with those identified by the Apache II system.

[0064]

[0074] Table 3 shows the area under the curve (AUC) when using a disease prognosis model 115 (e.g., implemented using a long-short-term memory (LSTM) network) to determine the clinical outcome of COVID-19 patients. As shown in Table 3, when the disease prognosis model 115 operates only on laboratory test data (as one type of longitudinal data), the disease prognosis model 115 can achieve an AUC of 0.63 on the test set. By incorporating vital signs (as another type of longitudinal data), the performance of the disease prognosis model 115 improved to an AUC of 0.70. These two results demonstrate the effectiveness of the disease prognosis model 115 in modeling longitudinal data. Furthermore, by incorporating different types of non-longitudinal data135 (e.g., static data including demographic data, medical history data, and medical data), the performance of the disease prognosis model115 further improved to area under the curve (AUC) of 0.73, 0.75, and 0.76, respectively.

[0065]

[0075] Table 3 TIFF2025532100000006.tif68170

[0066]

[0076] In some exemplary embodiments, the performance of the disease prognosis model 115 is further evaluated for different network architectures (e.g., a long short-term memory (LSTM) network with spatio-temporal attention 200 and an R-transformer with spatio-temporal attention 200) shown in FIGS. 3A-3B. The performance of the disease prognosis model 115 is also compared with that of two conventional models, including a long short-term memory (LSTM) network with temporal attention and a transformer with spatial attention. Table 4 shows the performance of the various models.

[0067]

[0077] Table 4 TIFF2025532100000007.tif52170

[0068]

[0078] The clinical model in Table 4, which only considered non-longitudinal data collected at the time of initial hospitalization, had an area under the curve (AUC) of 0.61. The poor performance of the clinical model is likely due to its inability to consider nonlinear interactions between variables and its exclusion of longitudinal data. A traditional long-short-term memory (LSTM) network alone can achieve an AUC of 0.76, while temporal attention can achieve an AUC of 0.77. With the addition of spatiotemporal attention, the disease prognosis model 115 achieves an AUC of 0.80 when implemented using a long-short-term memory (LSTM) network and an AUC of 0.94 when implemented as an R-Transformer. These results demonstrate that separating the learning of short-term and long-term dependencies, as in the case of the disease prognosis model 115 implemented using an R-Transformer, can improve the accuracy of clinical outcome prediction.

[0069]

[0079] The adjusted feature map output by spatiotemporal attention 200 can include various nonlinear interactions between features. Explaining spatiotemporal attention 200 with hidden features may be less straightforward and intuitive than explaining variables with physical meaning, such as heart rate or body temperature. Therefore, in some exemplary embodiments, the Apache II system can be used as a bridge to interpret where spatiotemporal attention 200 is looking. Figure 5A is a graph showing an example of a patient's Apache II score over time. Figure 5B is a graph showing the growth of the Apache II score broken down by the patient's various physiological variables. Figure 5C is a table visualizing an example output of spatiotemporal attention 200. As shown in Figure 5C, the importance of individual features not only changes between successive time points, but may also differ from the importance of other features.

[0070]

[0080] Based on statistical testing of the critical time points identified by Spatiotemporal Attention 200 and the Apache II system, feature values ​​present at the first time point (e.g., corresponding to disease onset) and the last time point (e.g., corresponding to the most recent condition) tend to be more important than those at other time points in determining a patient's COVID-19 clinical outcome, including the likelihood of symptom persistence and / or recurrence. The importance of the first time point is consistent with findings that experiencing five or more symptoms in the first week after illness onset is associated with COVID-19. Significant time points significantly correlated (p<0.05) between Spatiotemporal Attention 200 and the Apache II system with respect to COVID-19 clinical outcomes include: i) abnormal respiratory rate at the first time point and at 60 days later; and ii) both heart rate and creatinine levels in the abnormal range at any time point.

[0071]

[0081] In view of the above-described embodiments of the subject matter, the present application discloses the following list of examples, wherein one feature of an example alone or a combination of features of said example, optionally combined with one or more features of one or more additional examples, are also further examples within the scope of the present application's disclosure.

[0072]

[0082] Item 1: A computer-implemented method including: training a disease prognosis model to determine a clinical outcome of a disease based on at least longitudinal data, the longitudinal data including health records for each time point in a sequence of time points; training the disease prognosis model including training a recurrent neural network, a spatio-temporal attention, and a feed-forward neural network, wherein the recurrent neural network is trained to extract from each health record a feature set representing one or more local dependencies present in the health record, the spatio-temporal attention is trained to determine an importance of each feature in the feature set at each time point in the sequence of time points, and the feed-forward neural network is trained to determine the clinical outcome of the disease based at least on the importance of each feature in the feature set at each time point in the sequence of time points; and applying the trained disease prognosis model to determine a clinical outcome of the disease for a patient associated with the first health record and the second health record, based at least on the first health record from the first time point and the second health record from the second time point.

[0073]

[0083] Item 2: The method of item 1, wherein the recurrent neural network is a bidirectional recurrent neural network (RNN), a long short-term memory (LSTM) network, a local long short-term memory (LSTM) network with a given window size for a time point, or a gated recurrent unit (GRU) network.

[0074]

[0084] Item 3: The method according to any one of items 1 to 2, wherein the feedforward neural network is a multilayer perceptron model.

[0075]

[0085] Item 4: A method according to any one of items 1 to 3, wherein the trained disease prognosis model determines a clinical outcome of a disease for a patient associated with the first health record and the second health record by applying at least a trained recurrent neural network to extract from the first health record a first set of feature values ​​for a hidden feature set representing a first set of local dependencies present in the first health record, and applying a trained recurrent neural network to extract from the second health record a second set of feature values ​​for the hidden feature set representing a second set of local dependencies present in the second health record.

[0076]

[0086] Item 5: The method of item 4, wherein the trained recurrent neural network outputs a feature map including a first set of feature values ​​from a first time point and a second set of feature values ​​from a second time point for incorporation by the trained spatiotemporal attention.

[0077]

[0087] Item 6: The method of any of Items 4 to 5, wherein the trained disease prognosis model further determines a clinical outcome of a disease of a patient associated with the first health record and the second health record by applying at least the trained spatiotemporal attention to determine the importance of each feature in the hidden feature set at each of the first and second time points based at least on a feature map including a first set of feature values ​​and a second set of feature values.

[0078]

[0088] Item 7: The method of item 6, wherein the trained spatiotemporal attention includes one or more two-dimensional convolutional layers trained to determine the importance of each feature in the hidden feature set across the time dimension and the feature dimension.

[0079]

[0089] Item 8: The method of item 7, wherein the one or more two-dimensional convolutional layers include a 1x1 convolutional filter configured to perform joint weighting of the importance of each feature in the hidden feature set across the time dimension and the feature dimension.

[0080]

[0090] Item 9: The method of any of Items 6 to 8, wherein the trained spatiotemporal attention determines, for a first feature from the hidden feature set, a first importance of the first feature at a first time point and a second importance of the first feature at a second time point.

[0081]

[0091] Item 10: The method of item 9, wherein the trained spatiotemporal attention further determines, for a second feature from the hidden feature set, a third importance of the second feature at the first time point and a fourth importance of the second feature at the second time point.

[0082]

[0092] Item 11: The method of any of items 6 to 10, wherein the trained disease prognosis model further determines a clinical outcome of the patient's disease associated with the first health record and the second health record by applying at least the trained feedforward neural network to determine the clinical outcome of the patient's disease based at least on the importance of each feature in the hidden feature set at each of the first time point and the second time point.

[0083]

[0093] Item 12: The method of any of items 1 to 11, wherein the disease prognosis model is further trained to determine a clinical outcome of the disease based on non-longitudinal data, wherein the non-longitudinal data is static across the sequence of time points, and wherein the non-longitudinal data is linked with health records associated with each time point in the sequence of time points.

[0084]

[0094] Item 13: The method of item 12, further comprising identifying one or more missing values ​​for the non-longitudinal variables that constitute the non-longitudinal data, and replacing the one or more missing values ​​with the mean values ​​of the non-longitudinal variables observed in the available dataset.

[0085]

[0095] Item 14: The method of any of items 12 to 13, wherein the non-longitudinal data includes at least one of demographic information and medical history.

[0086]

[0096] Item 15: The method of any of items 12 to 14, wherein the non-longitudinal data includes medical image data and / or electrogram data corresponding to a metric that quantifies the severity of disease depicted in one or more medical images and / or electrograms.

[0087]

[0097] Item 16: The method of any of items 1 to 15, wherein the health record associated with each time point in the sequence of time points includes a value for each of a plurality of vital sign statistics.

[0088]

[0098] Item 17: The method of item 16, wherein the plurality of vital sign statistics includes one or more of systolic blood pressure, diastolic blood pressure, pulse rate, and respiratory rate.

[0089]

[0099] Item 18: The method of any of items 1 to 17, wherein the health record associated with each time point in the sequence of time points includes a value for each of a plurality of specimen testing variables.

[0090]

[0100] Item 19: The method of Item 18, wherein the plurality of specimen test variables includes one or more of fibrinogen, C-reactive protein, prothrombin international normalized ratio, prothrombin time, lactate dehydrogenase, D-dimer, albumin, ferritin, alanine aminotransferase, aspartate aminotransferase, chloride, protein, alkaline phosphatase, bilirubin, calcium, creatinine, glucose, hematocrit, hemoglobin, potassium, platelets, red blood cells, sodium, and white blood cells.

[0091]

[0101] Item 20: The method of any of items 1 to 19, wherein the health record associated with each time point in the sequence of time points includes medical imaging data and / or electrogram data.

[0092]

[0102] Item 21: The method of item 20, wherein the medical image data includes a metric that quantifies the severity of a disease depicted in one or more medical images and / or electrograms.

[0093]

[0103] Item 22: The method of any of items 1 to 21, further comprising determining that the longitudinal data includes a missing value for the longitudinal variable at the first time point; and replacing the missing value with (i) a first value for the longitudinal variable from a second time point that precedes the first time point, (ii) a second value for the longitudinal variable from a third time point that follows the first time point, or (iii) a third value determined based on the first value and the second value.

[0094]

[0104] Item 23: The method of any one of items 1 to 22, wherein the disease is coronavirus disease (COVID-19), Alzheimer's disease, or age-related macular degeneration.

[0095]

[0105] Item 24 : The method of any of items 1 to 23, wherein the clinical outcome of the disease includes a probability associated with one or more of cure, deterioration, and death.

[0096]

[0106] Item 25: The method of any of items 1 to 24, wherein the health record associated with each time point in the sequence of time points is an electronic health record (EHR).

[0097]

[0107] Item 26: A system comprising at least one data processor and at least one memory storing instructions, the instructions, when executed by the at least one data processor, causing operations including a method described in any of items 1 to 25.

[0098]

[0108] Item 27: A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, produce operations including the method described in any of items 1 to 25.

[0099]

[0109] 6 is a block diagram illustrating an example of a computing system 600, in accordance with an embodiment of the present subject matter. Referring to FIGS. 1-6, computing system 600 may be used to implement database management system 110 and / or any components therein.

[0100]

[0110] 6, computing system 600 may include a processor 610, a memory 620, a storage device 630, and an input / output device 640. The processor 610, the memory 620, the storage device 630, and the input / output device 640 may be interconnected via a system bus 650. The processor 610 may process instructions for execution within the computing system 600. The instructions thus executed may implement, for example, one or more components of the database management system 110. In some exemplary embodiments, the processor 610 may be a single-threaded processor. Alternatively, the processor 610 may be a multi-threaded processor. The processor 610 may process instructions stored in the memory 620 and / or the storage device 630 to display graphical information for a user interface provided via the input / output device 640.

[0101]

[0111] Memory 620 is a computer-readable medium, such as a volatile or non-volatile medium, that stores information within computing system 600. Memory 620 may store, for example, data structures representing a configuration object database. Storage device 630 may provide persistent storage for computing system 600. Storage device 630 may be a solid-state drive, a floppy disk device, a hard disk device, an optical disk device, a tape device, or other suitable persistent storage means. Input / output device 640 provides input / output operations for computing system 600. In some exemplary embodiments, input / output device 640 includes a keyboard and / or a pointing device. In various implementations, input / output device 640 includes a display device for displaying a graphical user interface.

[0102]

[0112] According to some demonstrative embodiments, I / O device(s) 640 may provide input / output operations for network devices. For example, I / O device(s) 640 may include an Ethernet port or other networking port for communicating with one or more wired and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).

[0103]

[0113] In some exemplary embodiments, computing system 600 may be used to execute various interactive computer software applications that can be used for organizing, analyzing, and / or storing various types of data. Alternatively, computing system 600 may be used to execute any type of software application. These applications may be used to perform various functions, such as planning functions (e.g., creating, managing, and editing spreadsheet documents, word processing documents, other objects, etc.), computing functions, communication functions, etc. Applications may include various add-in functions or may be standalone computing products or functions. When active within an application, these functions may be used to generate a user interface that is provided via input / output devices 640. The user interface may be generated by computing system 600 and presented to a user (e.g., on a computer screen monitor, etc.).

[0104]

[0114] One or more aspects or features of the subject matter described herein may be implemented in digital electronic circuitry, integrated circuits, specially designed ASICs, field programmable gate array (FPGA) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features may include implementation in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be special-purpose or general-purpose, coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communications network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0105]

[0115] These computer programs, which may also be referred to as programs, software, software applications, applications, components, or code, contain machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language and / or assembly / machine language. As used herein, the term “machine-readable medium” refers to any computer program product, apparatus, and / or device used to provide machine instructions and / or data to a programmable processor, such as, for example, magnetic disks, optical disks, memory, and programmable logic devices (PLDs), including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. A machine-readable medium may non-transitory store such machine instructions, such as, for example, a non-transitory solid-state memory, a magnetic hard drive, or any equivalent storage medium. Alternatively or additionally, a machine-readable medium may temporarily store such machine instructions, such as, for example, a processor cache or other random query memory associated with one or more physical processor cores.

[0106]

[0116] To provide for user interaction, one or more aspects or features of the subject matter described herein may be implemented on a computer having a display device, such as a cathode ray tube (CRT) or liquid crystal display (LCD) or light-emitting diode (LED) monitor, for displaying information to a user, and a keyboard and pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices may also be used to provide for user interaction. For example, the recurrent provided to the user may be any form of sensory recurrent, such as visual recurrent, auditory recurrent, or tactile recurrent, and input from the user may be received in any form, including acoustic input, voice input, or tactile input. Other possible input devices include touchscreens or other touch-sensitive devices such as single-point or multi-point resistive or capacitive trackpads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, etc.

[0107]

[0117] In the above description and in the claims, phrases such as "at least one of" or "one or more of" may be preceded by a conjunctive list of elements or features. The term "and / or" may also be used with lists of two or more elements or features. Unless otherwise implicitly or explicitly stated by the context of use, such phrases are intended to refer to any of the listed elements or features individually, or any of the listed elements or features in combination with any of the other listed elements or features. For example, the phrases "at least one of A and B," "one or more of A and B," and "A and / or B" are intended to mean "A only, B only, or A and B together," respectively. A similar interpretation is intended for lists containing more than two items. For example, the phrases "at least one of A, B, C," "one or more of A, B, C," and "A, B, and / or C" are intended to mean "A only, B only, C only, A and B together, A and C together, B and C together, or A, B and C together," respectively. Use of the term "based on" above and in the claims means "based at least in part on," and implies that unrecited features or elements are also permitted.

[0108]

[0118] The subject matter described herein may be embodied in systems, devices, methods, and / or articles, depending on the desired configuration. The embodiments set forth in the foregoing description do not represent all embodiments consistent with the subject matter described herein. Rather, they are merely some examples consistent with aspects related to the described subject matter. While several variations have been detailed above, other modifications and additions are possible. In particular, additional features and / or variations may be provided in addition to those described herein. For example, the embodiments described above may be directed to various combinations and subcombinations of the disclosed features and / or combinations and subcombinations of certain additional features disclosed above. In addition, the logic flow depicted in the accompanying figures and / or described herein does not necessarily require the particular order shown or sequential order to achieve desirable results. Other embodiments may be within the scope of the following claims.

Claims

1. training a disease prognosis model to determine a clinical outcome of a disease based on at least longitudinal data, the longitudinal data including health records for each time point in a sequence of time points, training the disease prognosis model including training a recurrent neural network, a spatio-temporal attention, and a feed-forward neural network, wherein the recurrent neural network is trained to extract from each health record a feature set that represents one or more local dependencies present in the health record, the spatio-temporal attention is trained to determine an importance of each feature in the feature set at each time point in the sequence of time points, and the feed-forward neural network is trained to determine the clinical outcome of the disease based at least on the importance of each feature in the feature set at each time point in the sequence of time points; applying the trained disease prognosis model to determine, based at least on a first health record from a first time point and a second health record from a second time point, the clinical outcome of the disease for the patient associated with the first health record and the second health record; 10. A computer-implemented method comprising:

2. 2. The method of claim 1, wherein the recurrent neural network is a bidirectional recurrent neural network (RNN), a long short-term memory (LSTM) network, a localized long short-term memory (LSTM) network with a given window size for time points, or a gated recurrent unit (GRU) network.

3. The method of claim 1 or 2, wherein the feedforward neural network is a multi-layer perceptron model.

4. The trained disease prognosis model comprises at least: applying the trained recurrent neural network to extract from the first health record a first set of feature values ​​for a set of hidden features representing a first set of local dependencies present in the first health record; applying the trained recurrent neural network to extract from the second health record a second set of feature values ​​for the set of hidden features representing a second set of local dependencies present in the second health record; and The method of claim 1 , further comprising determining the clinical outcome of the disease of the patient associated with the first health record and the second health record by:

5. 5. The method of claim 4, wherein the trained recurrent neural network outputs a feature map including the first set of feature values ​​from the first time point and the second set of feature values ​​from the second time point for incorporation by the trained spatiotemporal attention.

6. 6. The method of claim 4, wherein the trained disease prognosis model further determines the clinical outcome of the disease for the patient associated with the first health record and the second health record by applying at least the trained spatio-temporal attention to determine the importance of each feature in the hidden feature set at each of the first time point and the second time point based at least on a feature map including the first set of feature values ​​and the second set of feature values.

7. 7. The method of claim 6, wherein the trained spatio-temporal attention comprises one or more two-dimensional convolutional layers trained to determine the importance of each feature in the set of hidden features across a time dimension and a feature dimension.

8. 8. The method of claim 7, wherein the one or more two-dimensional convolutional layers comprise a 1x1 convolutional filter configured to perform joint weighting of the importance of each feature in the set of hidden features across the time dimension and the feature dimension.

9. 9. The method of claim 6, wherein the trained spatio-temporal attention determines, for a first feature from the set of hidden features, a first importance of the first feature at the first time point and a second importance of the first feature at the second time point.

10. 10. The method of claim 9 , wherein the trained spatio-temporal attention further determines, for a second feature from the hidden feature set, a third importance of the second feature at the first time point and a fourth importance of the second feature at the second time point.

11. 11. The method of claim 6, wherein the trained disease prognosis model further determines the clinical outcome of the disease for the patient associated with the first health record and the second health record by applying at least the trained feedforward neural network to determine the clinical outcome of the disease for the patient based at least on the importance of each feature in the set of hidden features at each of the first time point and the second time point.

12. 12. The method of claim 1, wherein the disease prognosis model is further trained to determine the clinical outcome of the disease based on non-longitudinal data, wherein the non-longitudinal data is static across the sequence of time points, and wherein the non-longitudinal data is linked with the health records associated with each time point in the sequence of time points.

13. identifying one or more missing values ​​for non-longitudinal variables comprising the non-longitudinal data; replacing the one or more missing values ​​with the mean value of the non-longitudinal variable observed in the available dataset; The method of claim 12 further comprising:

14. 14. The method of claim 12 or 13, wherein the non-longitudinal data comprises medical image data and / or electrogram data corresponding to a metric quantifying the severity of the disease depicted in one or more medical images and / or electrograms.

15. 15. The method of any one of claims 1 to 14, wherein the health record associated with each time point in the sequence of time points includes a value for each of a plurality of vital sign statistics.

16. 16. The method of any one of claims 1 to 15, wherein the health record associated with each time point in the sequence of time points includes a value for each of a plurality of analyte test variables.

17. 17. The method of any one of claims 1 to 16, wherein the health record associated with each time point in the sequence of time points comprises medical image data and / or electrogram data, the medical image data comprising a metric quantifying the severity of the disease depicted in one or more medical images and / or electrograms.

18. determining that the longitudinal data includes missing values ​​for longitudinal variables at a first time point; replacing the missing value with (i) a first value of the longitudinal variable from a second time point that precedes the first time point, (ii) a second value of the longitudinal variable from a third time point that follows the first time point, or (iii) a third value determined based on the first value and the second value; 18. The method of any one of claims 1 to 17, further comprising:

19. 19. The method of any one of claims 1 to 18, wherein the disease is coronavirus disease (COVID-19), Alzheimer's disease, or age-related macular degeneration.

20. 20. The method of any one of claims 1 to 19, wherein the clinical outcome of the disease comprises a probability associated with one or more of cure, deterioration, and death.

21. at least one data processor; at least one memory storing instructions; 22. A system comprising: instructions that, when executed by the at least one data processor, result in operations comprising the method of any one of claims 1 to 21.

22. A non-transitory computer readable medium storing instructions that, when executed by at least one data processor, result in operations comprising the method of any one of claims 1 to 21.