Predicting risk of diseases with artificial intellegince foundation model on electronic health records

The Transformer-based predictive model, TRADE, enhances EHR-based AD/ADRD/MCI prediction by integrating diverse medical information, achieving superior accuracy and addressing healthcare inequities through self-supervised learning and finetuning, thereby improving early diagnosis and intervention.

WO2025226739A1PCT designated stage Publication Date: 2025-10-30NEW YORK UNIV +5
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/025859
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-25
Filing Date
2025-04-22
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing EHR-based models for predicting Alzheimer's disease (AD) and related dementias (ADRD) lack precision due to the exclusion of important factors like specific medications and lab values, leading to reduced sensitivity in marginalized communities and inequities in care.

Method used

A Transformer-based predictive model, TRADE, is developed using a self-supervised learning approach to pretrain a foundation model with a large-scale EHR dataset, incorporating medications and discretized lab values, and is finetuned for predicting AD/ADRD/MCI, leveraging diverse medical information.

Benefits of technology

The model achieves improved accuracy in identifying AD/ADRD/MCI risks, with AUROC of 0.772 in 1 year and 0.735 in 5 years, and positive predictive values of 39.2% and 27.8% for the top 1% and 5% highest estimated risks, outperforming current clinical screening programs and addressing disparities in healthcare access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025025859_30102025_PF_FP_ABST
    Figure US2025025859_30102025_PF_FP_ABST
Patent Text Reader

Abstract

Exemplary systems, methods, and computer-accessible medium according to the exemplary embodiments of the present disclosure are provided for generating at least one medical condition prediction. Thus, the exemplary systems, methods, and computer-accessible medium can obtain, by at least one computer processor, structured data from a pool of electronic health records (EHRs), the structured data comprising a plurality of variables from one or more tables, create a training data set from the structured data, encode for each patient in the pool of EHRs an age, an index of days, and one or more encounters via a positional embedding mechanism, train a machine learning model using the training data and encoding, receive patient data, and generate the at least one medical condition prediction on the received patient data with the trained machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION SYSTEM, METHOD, AND COMPUTER ACCESSIBLE MEDIUM FOR PREDICTING RISK OF DISEASES WITH ARTIFICIAL INTELLEGINCE FOUNDATION MODEL ON ELECTRONIC HEALTH RECORDS CROSS REFERENCE TO RELATED APPLICATION(S)

[0001] This application relates to and claims the benefit of priority from U.S. Provisional Patent Application Nos.63 / 637,169, filed on April 22, 2024, and 63 / 638,866, filed on April 25, 2024, the entire disclosures of which are incorporated herein by reference. STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH

[0002] This invention was made with government support under R01 AG085617, R01 AG079175, and P30 AG066512 awarded by the National Institutes of Health, and 1922658 awarded by the National Science Foundation. The government has certain rights in the invention. FIELD OF THE DISCLOSURE

[0003] The present disclosure relates generally to systems, methods and computer-accessible medium for predicting a risk of disease with an artificial intelligence (AI) foundation model on electronic health records. BACKGROUND INFORMATION

[0004] Alzheimer’s disease (AD) and AD related dementias (ADRD) are irreversible conditions affecting over 6 million people in the United States (see, e.g., Refs.1, 2). While the pathogenesis of AD / ADRD is complex, several modifiable cardiovascular risk factors (such as hypertension, smoking, obesity, and diabetes) are recognized to contribute to its pathogenesis and progression. (see, e.g., Refs. 3-6). Early diagnosis of AD / ADRD or mild cognitive impairment (MCI) are important for a number of reasons. First, the diagnosis may serve as a motivating factor for addressing modifiable risk factors amongst both clinicians and persons living with dementia (PLWD) or mild cognitive impairment (MCI) as it is one of the only currently known ways to slow cognitive decline. Second, when diagnosed earlier, it is easier for advance care planning to occur with input from the PLWD or MCI, and allows for greater planning for the eventual decline, including caregiving situation, housing or moving closer to family, financial planning. While controversial in nature, early diagnosis is also important as the sole disease-modifying FDA- approved therapy currently on the market for AD / ADRD / MCI is only available and shown some efficacy in PLWD or MCI at the early stage of impairment. (see, e.g., Ref. 7). Accordingly, anAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION early identification of AD / ADRD at the prodromal stage can be a critical component in optimal care delivery. However, at the same time, there are significant inequities in care, including lower rates of early diagnosis in racial and ethnic minoritized groups and worse management of modifiable cardiovascular disease factors. These inequities are multifactorial, and new methods are needed for systematically improving the healthcare system.

[0005] Algorithms and risk models for detecting probable undiagnosed MCI or AD / ADRD are one way to facilitate such early diagnosis and intervention. Approaches with wearable signals (see, e.g., Ref. 8), neuroimaging (see, e.g., Ref. 9) and blood biomarkers (see, e.g., Ref. 10) are promising, but cannot be readily applied to the general population due to the cost and effort , and limited use, especially in underserved and historically marginalized communities. Assessing risks with Electronic Health Records (EHR)-based models has the potential for direct integration into EHR systems and clinical workflow at the point of care for every patient. A few retrospectively validated EHR-based models exist to date. eRADAR, a statistical model (see, e.g., Ref. 11), which was developed and validated using participants of Adult Changes in Thought study, is the most robust model tested to date in scientific studies and is undergoing prospective validation and randomized trials. (see, e.g., Refs.12 and 13). One limitation of the eRADAR model however is its focus on using known risk factors including aging, vascular diseases and diabetes, while other potential factors that can improve precision, such as specific medications, lab values, and additional diagnoses are not included. This is particularly important as many of the eRADAR risk factors have correlations with socioeconomic status and racial and ethnic minoritized groups which can reduce sensitivity in these populations.

[0006] While existing models can be beneficial, there remains a need for an EHR-based risk assessment model leveraging utilizing the additional information available in the EHR. Such assessment model can exhibit improved performances on subgroups on subgroups with different characteristics, which can address and / or overcome at least some of the deficiencies described herein above. SUMMARY OF EXEMPLARY EMBODIMENTS

[0007] The following is intended to be a brief summary of the exemplary embodiments of the present disclosure and is not intended to limit the scope of the exemplary embodiments.

[0008] In some exemplary embodiments of the present disclosure, exemplary systems, methods, and computer accessible medium can be provided which can generate at least one medicalAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION condition prediction. The exemplary systems, methods, and computer accessible medium can ingest structured data from a pool of electronic health records (EHRs) where the structured data can comprise a plurality of variables from one or more tables, create a training data set from the structured data, encode for each patient in the pool of EHRs, an age, an index of days, and one or more encounters via a positional embedding mechanism, train a machine learning model using the training data and encoding, receive patient data, and generate the at least one medical condition prediction on the received patient data with the trained machine learning model.

[0009] The exemplary systems, methods, and computer accessible medium can further tokenize the plurality of variables, wherein numeric variables can be tokenized into a plurality of number ranges based on standard deviations from accepted normal values and wherein International Classification of Diseases codes, medications, and demographics variables are tokenized as indicator tokens. In some exemplary embodiments, the training data set can be created from the tokenized plurality of variables.

[0010] In some exemplary embodiments of the present disclosure, exemplary systems, methods, and computer accessible medium can further finetune the machine learning model based the structured data. The at least one medical condition prediction can be generated on the received patient data with the trained and finetuned machine learning model. Moreover, the at least one medical condition prediction can comprise a dementia prediction, a pancreatic cancer prediction, or any other medical condition prediction.

[0011] These and other objects, features and advantages of the exemplary embodiments of the present disclosure will become apparent upon reading the following detailed description of the exemplary embodiments of the present disclosure, when taken in conjunction with the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Further objects, features and advantages of the present disclosure will become apparent from the following detailed description taken in conjunction with the accompanying Figures showing illustrative embodiments of the present disclosure, in which:

[0013] Figure 1 is an illustration of an exemplary design of ML architecture according to exemplary embodiments of the present disclosure;

[0014] Figure 2 is a set of exemplary graphs illustrating prediction performance of a machine learning model according to an exemplary embodiment of the present disclosure;Attorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION

[0015] Figure 3 is a set of exemplary bar charts illustrating an exemplary machine learning model performance across different subcohorts according to an exemplary embodiment of the present disclosure;

[0016] Figure 4 is a set of exemplary graphs illustrating prediction performance of a machine learning model in different outcome windows according to an exemplary embodiment of the present disclosure;

[0017] Figure 5 is a set of exemplary graphs illustrating an exemplary prediction performance of a machine learning model with and without excluding new-onset patients diagnosed in a 1-year gap window according to an exemplary embodiment of the present disclosure;

[0018] Figure 6 is an exemplary set of graphs illustrating an exemplary analysis of MMSE scores for partial patients with cognitive exam records within 1-year of the index date according to an exemplary embodiment of the present disclosure;

[0019] Figure 7 is a block diagram of an exemplary embodiment of a system according to the present disclosure;

[0020] Figure 8 is an exemplary flow chart for creating, training, and using an exemplary machine learning algorithm according to an exemplary embodiment of the present disclosure;

[0021] Figure 9 is a set of exemplary bar charts illustrating distributions of EHR data for a pretraining cohort (top row) and a 1-year feature window in the AD / ADRD / MCI finetuning cohort (bottom row) according to an exemplary embodiment of the present disclosure;

[0022] Figure 10A is a set of exemplary graphs illustrating comparison of performance on predicting AD / ADRD / MCI onset in a 1-year outcome window according to an exemplary embodiment of the present disclosure; and

[0023] Figure 10B is a set of exemplary graphs illustrating comparison of performance on predicting AD / ADRD / MCI onset in a 2-year outcome window according to an exemplary embodiment of the present disclosure.

[0024] Throughout the drawings, the same reference numerals and characters, unless otherwise stated, are used to denote like features, elements, components or portions of the illustrated embodiments. Moreover, while the present disclosure will now be described in detail with reference to the figures, it is done so in connection with the illustrative embodiments and is not limited by the particular embodiments illustrated in the figures and the appended claims.Attorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION DETAILED DESCRIPTION OF EXAMPLARY EMBODIMENTS

[0025] The following description of exemplary embodiments provides non-limiting representative examples referencing numerals to particularly describe features and teachings of different exemplary aspects and exemplary embodiments of the present disclosure. The exemplary embodiments described should be recognized as capable of implementation separately, or in combination, with other exemplary embodiments from the description of the exemplary embodiments. A person of ordinary skill in the art reviewing the description of the exemplary embodiments should be able to learn and understand the different described aspects of the present disclosure. The description of the exemplary embodiments should facilitate understanding of the exemplary embodiments of the present disclosure to such an extent that other implementations, not specifically covered but within the knowledge of a person of skill in the art having read the description of embodiments, would be understood to be consistent with an application of the exemplary embodiments of the present disclosure.

[0026] The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can train an artificial intelligence foundation model to understand the electronic health records (EHR) with a vast cohort of e.g., 1 million patients. Building upon this foundation EHR model, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can develop a predictive model capable of identifying risks for a medical condition such as AD / ADRD and mild cognitive impairment (MCI), by analyzing the past sequential visit records. This exemplary model can generate risk predictions for various future timeframes. For example, on the held-out validation set, the model can achieve an area under the receiver operating characteristic (AUROC) at 0.769 for identifying the AD / ADRD / MCI risks in 1 year, and AUROC at about 0.734 in 5 years, specifically among patients aged 65 and above. The positive predictive values (PPV) in 5 years among patients with top 1% and 5% highest estimated risks can be about 39.2% and about 27.8%. These exemplary results demonstrate improvements upon the current clinical AD / ADRD / MCI screening program with statistical risk scores, which can potentially benefit the prognosis of dementia by preemptive intervention.

[0027] Recent advancements in deep learning can offer neural networks with a stronger capacity to understand the EHR. (see, e.g., Refs.14-17). A graph neural network, taking the connections among various EHR features into account, can outperform traditional statistical models andAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION multiple layer perceptron networks in predicting potential AD. (see, e.g., Ref.18) However, such study only used cross-sectional EHR without temporal information. Transformer (see, e.g., Ref. 19), a powerful network architecture used to represent high-dimensional data like images and languages, is now widely used as the state-of-the-art method to represent sequential relationship across visits and interconnection of information within a single visit (see, e.g., Refs.20-22). The increment in the scale of the Transformer model usually leads to better performance, while it also requires larger training data to achieve generalizable performance (see, e.g., Ref.23). To address the insufficiency of labeled data, a self-supervised learning approach, masked token modeling, can be used to pretrain these large models on a broader domain with unlabeled data (see, e.g., Refs. 24, 25). Then the large pretrained model, termed foundation Model, can be partially or fully finetuned for some downstream prediction tasks with smaller labeled datasets (see, e.g., Refs.26, 27). The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can employ this framework to develop a Tranformer-based predictive model, namely TRADE, where the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can pretrain a foundation model with a large-scale inclusive EHR dataset, and finetune the model with a selective cohort for predicting AD / ADRD / MCI, and can outperform a vanilla Transformer model trained from scratch. Compared to Transformers for EHR in previous literature (e.g. BEHRT (see, e.g., Ref.20) and MED-BERT (see, e.g., Ref.21)), TRADE and its base foundation model can further incorporate the medications and discretized lab values, in addition to diagnosis codes, to enrich the EHR representation. Also, the exemplary model can be trained with a curated cohort with more medical concepts, visits and longer span of presence in the healthcare system.

[0028] The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can achieve an area under the receiver operating characteristic (AUROC) at 0.772 (95% CI: 0.770, 0.773) in identifying AD / ADRD / MCI patients in 1 year, and 0.735 (95% CI: 0.734, 0.736) in about 5 years, specifically among patients aged 65 and above. The positive predictive values (PPV) in 5 years among patients with the top 1% and 5% highest estimated risks can be 39.2% and 27.8%. The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can build a comparison model on the same cohort with eRADAR that can be deployed in clinical practices (see, e.g., Ref.11). The exemplary systems, methods, and computer accessible mediumAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION according to the exemplary embodiments of the present disclosure can demonstrate that the proposed deep-learning model of exemplary embodiments of the present disclosure is more accurate on the overall heldout validation set and subcohorts separated by different characteristics. These results suggest that incorporating the AI foundation model according to exemplary systems, methods, and computer accessible medium according to exemplary embodiments of the present disclosure, in the EHR system is promising to advance the current identification and screening of patients at high risk for AD / ADRD / MCI.

[0029] Figure 1 shows an exemplary design of machine learning architecture according to the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure. For instance, at 130, a patient’s EHR data can be collected / input. This can include data from one or more visits across time, and can also include demographics, diagnoses, lab values, medication etc. At 120, a larger pool of patient data can be partitioned – e.g., into a training, validation, and held-out validation set. A pretraining cohort, formed by patients in the training set, can be used to pretrain an exemplary model of the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure. An finetuning cohort (e.g., for AD / ADRD / MCI or any other condition), can be filtered by age, and can be used to train and validate a predictive model according to the exemplary embodiments of the present disclosure. The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can pretrain the EHR model at 130. The exemplary model can process variables in EHR as tokens, and can encode indices of visits, days, and ages for e.g., 3 distinct positional encodings, which can be added to the token embeddings. The exemplary model can be pretrained by masked token modeling, which can teach the exemplary model to reconstruct the randomly masked information in the input EHR sequence. The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can train the predictive model at 140. The exemplary predictive model can be constructed by stacking a linear network upon the output representation of <CLS> token, estimating the risks based on previous EHR. At 150, patient sample timelines can be created and at 160, disease samples can be created. The samples for the predictive model of the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can be generated by, e.g., sliding windows on various index visits over patient’s longitudinal EHR. For example, records inAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION the 1-year feature window prior to the index visit can be used as inputs and the disease status can be used in various outcome windows following the index visit as labels. Samples with no follow- ups by the end of the outcome window can be considered censored and excluded. A minimum gap period between the end of the feature window and the beginning of the outcome window can be specified, and samples with disease onset in the gap window can be excluded to avoid data leakage due to lengthy disease diagnosis processes.

[0030] Figure 8 illustrates an exemplary process according to at least some exemplary embodiments of the present disclosure. For example, at step 810, structured EHR data may be ingested. This data can be structured as part of the ingestion step. The structure of the EHR data can be compliant with FHIR specifications. At step 820, the ingested data may be stored in a database such as MongoDB. At step 830, relevant data can be moved to a computation-efficient sparse-array shelf database. A bespoke tokenizer can be built and trained at step 840. This tokenizer can incorporate EHR variables from the ingested EHR data. At step 850, a machine learning model (e.g., a Foundation Model) can be pre-trained for EHR data with a pretraining cohort. The tokenized EHR variables can then be fed into the learning model. At step 860, a disease prediction can be provided by the learning model, based on the tokenized EHR variables. Finally, at step 870, results can be provided, and specific treatments can be directed based on the specific predicted disease and the timing and / or severity of the predicted disease.

[0031] For example, while much of the disclosure that follows describes the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure through the lens of AD / ADRD / MCI, the applicability of the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure are not so limited. In fact, the exemplary system described in Figure 1 and the exemplary method illustrated by Figure 8 can be applied to generate a prediction as to any medical condition or disease. As described in detail below, EHR records of patients can be distilled into a number of variables that can be tokenized. Once the EHR data is tokenized, the exemplary model may be finetuned to generate predictions for a desired medical condition based on the tokenized EHR data. For example, Table 1a provides exemplary definition criteria for AD / ADRD / MCI onset while Table 1b provides definition criteria for pancreatic cancer. Criteria for other medical conditions can be considered and used as desired. Exemplary ResultsAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION Exemplary Study Participants

[0032] The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can use EHR data from a large health system, covering a large number of patients (e.g., 1,288,333 patients) with 587 million medical concept tokens from Jan 2013 to Jan 2023 with at least 5 visits. The EHR of each patient can include a sequence of encounters consisting of International Classification of Diseases (ICD-10) diagnostic codes (all diagnoses types, including encounter, problem list, billing and medical history), lab values coded under Logical Observation Identifiers Names and Codes (LOINC) and medications at the therapeutic level (see e.g., 110, Figure 1). The onset of AD / ADRD and MCI can be defined as the visit with the first occurrences of AD / ADRD / MCI-related diagnosis codes or medication (see, e.g., Table 1a). Condition Criteria (ICD-10 codes and Medications) F01.*: Any Vascular Dementia F02.*: Dementia in other diseases classified elsewhere with or without behavioral disturbance F03.*: Unspecified dementia with / without behavioral disturbance F04.*: Amnestic disorder due to known physiological condition AD / ADRD G23.1: progressive supranuclear palsy G30.*: Any Alzheimer’s disease G31.01: Pick’s disease G31.09: Other frontotemporal dementia G31.83: Dementia with Lewy bodies G31.9: Degenerative disease of nervous system, unspecified G31.1: Senile degeneration of brain, not elsewhere classified Mild Cognitive G31.84: Mild cognitive impairment of uncertain or unknown etiology Impairment G31.85: Corticobasal degeneration DONEPEZIL GALANTAMINE Dementia Medications MEMANTINE RIVASTIGMINE TACRINE Table 1a. Definition criteria for AD / ADRD / MCI onsetAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION Table 1b. Definition criteria for pancreatic cancer

[0033] According to exemplary systems, methods, and computer accessible medium according to exemplary embodiments of the present disclosure, Element 120 of Figure 1 provides an exemplary illustration indicating that patients can be randomly partitioned into the training, validation, and a held-out validation set, with a ratio of 80% (n=1,030,438), 10% (n=129,127), and 10% (n=128,768). The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can construct two cohorts to train the exemplary model: (1) a pretraining cohort for self-supervised learning for the foundation Transformer model; (2) an AD / ADRD / MCI finetuning cohort for training and evaluating the predictive model for AD / ADRD / MCI.

[0034] The exemplary pretraining cohort can include, e.g., all the patients in the training set (n=1,030,340). The AD / ADRD / MCI finetuning cohort may only contain patients over the age of 65 (by the end of records) (n=142,702). To simulate the predictive performance of the screening model in a retrospective manner, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can utilize a sliding window for each patient to generate pairs of features and labels (see, e.g., elements 150 and 160 of Figure 1), where, given an index visit, the features can be extracted from the EHR data prior to that index visit (feature window), and the labels can be based on the EHR data following the index window (outcome window). Samples with no follow-ups by the end of the outcome window can be considered censored and excluded. Patients who already had AD / ADRD / MCI by the end of theAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION feature window can also be excluded. In the analysis of AD / ADRD / MCI risk predictions in 5 years, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can use a 1-year feature window and 5-year outcome window, which can result in 445, 142 records among 142,702 unique patients with 7.20% of those developing AD / ADRD / MCI in 5 years. These samples belonging to the training, validation, and held-out validation set of the full dataset can form the corresponding sets of AD / ADRD / MCI finetuning sets. Table 2 shows demographic characteristics of both patient cohorts. The distributions of each patient’s length of presence, number of encounters, and medical codes are illustrated in Figure 9. Exemplary Identifying Risk of AD / ADRD / MCI with EHR

[0035] The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can build or otherwise generate a classification model on top of a Transformer network (see, e.g., Ref. 19), namely TRADE, to Predict Risk of AD / ADRD / MCI based on the previous EHR (see, e.g., 140 of Figure 1). The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can assess the predictive performance of exemplary predictive models using two key metrics: the area under the receiver operating characteristic curve (AUC) and positive predictive values (PPV) at varying sensitivities. For example, Figure 2 shows the performance on predicting diseases (e.g., AD / ADRD / MCI) in 5 years with different methods: “Scratch” is training the Transformer from random initial weights; “Linear” is training a linear classifier on top of the foundation model representations; “Finetuning” is training the classifier layer together with the foundation model layers. The performance can be evaluated with ROC curves on the left graph, and precision-recall curves on the right graph. The PPVs and sensitivities of the patients with top 1, 5 and 10% risks are highlighted in the right graph by ⋆, ▲ and ■, respectively. Similarly, Figure 4 shows a comparison of an exemplary predictive model performance on predicting disease (e.g., AD / ADRD / MCI) onset in different outcome windows. Prediction in the 1-year timeframe can obtain the highest AUROC due to the least uncertainty. Prediction in the 5-year timeframe can top the PPVs as the corresponding prevalence may be the highest.

[0036] For predicting the onset of AD / ADRD / MCI within 5 years from the reference index visit, the fine-tuned transformer according to the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can achieve anAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION AUROC of 0.735 (95% CI: 0.734-0.736) and average PPV (AP) of 0.199 (95% CI: 0.198-0.200). Since this predictive model aims to suggest cognitive screening for high-risk patients to providers, exemplary embodiments also present PPVs of the high-risk patient groups can be predicted by the model according to the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure (see, e.g., Figure 2). These high-risk groups can be delineated by thresholds of the predicted risks from the top k% of patients with the highest predicted risks. For the 5-year prediction horizon, the PPVs for the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure are 39.2% and 27.8% and 22.4% for thresholds corresponding to the top 1, 5, and 10 percentiles, respectively. These values demonstrate the precision of the exemplary model at different screening operation points. Table 2. Exemplary characteristics of the large-scale pretraining cohort and AD / ADRD / MCI finetuning cohort. The proportion or standard deviation is marked in parentheses.

[0037] Prior studies have also developed and validated EHR-based methods for early ADRD detection and prediction, yielding promising results. eRADAR (see, e.g., Ref. 11) is a statistical model estimating ADRD risks based on EHR variables, widely employed and validated in practice (see, e.g., Ref. 13). Comparing the TRADE model of the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure with eRADAR for predicting AD / ADRD / MCI over five years (see, e.g., Figure 2), the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments ofAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION the present disclosure can observe that the foundation model can significantly outperform the statistical model which achieves AUROC of 0.688 (95% CI: 0.687-0.689) and AP of 0.143 (95% CI: 0.142, 0.143). Even without pre-training, the model, according to exemplary systems, methods, and computer accessible medium according to exemplary embodiments of the present disclosure, can achieve better performance, and linear probing (i.e only fine-tuning a linear classifier over the foundation model’s pre-learned features) bypasses the performance of the eRADAR model. Exemplary Subcohort Analysis

[0038] The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can compare the model performance of TRADE and eRADAR across various sub-cohorts according to e.g., gender, age, race, and other health conditions, like vascular risk (see, e.g., Figure 3). Age and vascular risks are two major dementia risk factors significantly impacting prediction performance. Both models can obtain higher PPVs among older patients and patients with vascular risks since AD / ADRD / MCIs are more prevalent among these subcohorts. However, the AUROCs can be lower in the higher-risk subcohorts since it can be harder to identify the negative patients among them. Notably, the AUROCs of eRADAR in age and vascular risk subgroups can drop more than those of TRADE. This phenomenon indicates that the predictions of eRADAR are mostly attributed to age and vascular diseases, while the Transformer model, according to the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can leverage diverse medical information to further distinguish risks of patients at similar ages or health conditions.

[0039] Among subcohorts separated by demographic characteristics, the exemplary TRADE exhibited superior performance among females compared to males. Regarding the different races, while TRADE of the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can demonstrate lower performance among the underrepresented Black population when compared to the White population is still can consistently enhance performance across all racial groups compared to eRADAR, with a most notable improvement observable among the Asian population. This finding suggests that TRADE has the potential to benefit all gender and race groups. with a more accurate assessment of AD / ADRD / MCI risks.Attorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION Exemplary Improving Performance with a Pretrained Foundation EHR Model

[0040] An exemplary large-scale EHR dataset contains rich information on patients. To enable a model to understand EHR best, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can design a prediction framework with two stages (see, e.g., 110 of Figure 1): (1) pertaining, where the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure pre-train a Foundation Model for EHR with Transformer architecture with a pretraining cohort. The model according to the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can be trained without labels and merely by reconstructing randomly masked information from the EHR; (2) fine-tuning, where the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can fine-tuned the exemplary model with medical history and AD / ADRD / MCI outcomes in a fine-tuning cohort to identify the high-risk patients (see the Exemplary Method section below for more details). The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can conduct pretraining only on the pretraining cohort. The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can fine-tune the model with the training set in the AD / ADRD / MCI finetuning cohort and use the validation set to examine the performance for different hyperparameter settings, which can guide model selection. The performance of the selected models can be evaluated on the fully held-out validation set of patients and can be reported as an estimate of performance in new patient cohorts.

[0041] The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can compare the performance of the fine-tuned AD / ADRD / MCI predictive model with and without self-supervised pre- training. Figure 2 shows that for the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure under the same Transformer architecture, finetuning from pretrained foundation model can obtain both higher AUROC and PPVs as compared to training from scratch. In addition, a simplified linear classifier, trained over the representations extracted from the pretrained foundation model, and comparable in complexity toAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION eRADAR, can still achieve superior performance to eRADAR. This can demonstrate the value of a self-supervised pre-trained foundation AI model in building EHR-based predictive models. Exemplary Prediction for Different Time Intervals

[0042] The exemplary purpose of identifying patients at risk of dementia at different timelines varies. By predicting dementia across different time intervals, healthcare providers can intervene effectively at each stage. The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can develop models for predicting AD / ADRD / MCI onset within 1, 2, and 5 years from the visit being assessed. The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can construct samples of feature and label pair with different lengths of intervals for the outcome window (see, e.g., 130 of Figure 1). The corresponding characteristics of samples for various timelines are reported in Table 5. Exemplary figure 4 illustrates the prediction performance of the fine-tuned Transformer according to the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure for different time intervals. Comparing the AUCs among them, prediction for the outcome window with 1 year is highest with AUC at 0.772 (95% CI: 0.770, 0.773). This can indicate the predictive model of the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure is more discriminative on the outcomes in the near future, while longer time interval can introduce more inherent uncertainty of the potential outcomes. Similarly, Figures 10A and 10B provide exemplary illustration of an exemplary comparison of performance on predicting disease (e.g., AD / ADRD / MCI) onset in different 1-year (Figure 10A) and 2-year (Figure 10B) outcome windows with different baselines. These graphs illustrate that the finetuned foundation model can consistently outperform the other baselines.

[0043] The precision-recall curves provide additional insights into the differences among outcome windows. For example, Figure 4 indicates the prediction on the longer outcome window has a high PPV compared to shorter ones at the same operation points. For patients at the top 1% risk level, the PPV of AD / ADRD / MCI onset in 5 years reached 39.2%, and decreased to 20.7% in the 1-year window. The false positive patients from the prediction in 1 year might develop dementia in the future and thus become a true positive in the 5-year window. On the other hand, the longer outcomeAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION window can lead to lower sensitivity as the early-stage AD / ADRD / MCI patients might not exhibit any risk factors at the current or past visits. Exemplary Prediction with Gap Windows

[0044] Since the development and diagnosis of AD / ADRD / MCI is a prolonged process that can last, e.g., take several months, the recorded diagnosis on the EHR may be delayed. Considering the situation where a patient has onset of AD / ADRD / MCI, but has not yet been diagnosed, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can conduct another analysis by excluding recently diagnosed patients within a gap window of about 1 year after the end of the feature window, as shown in exemplary elements 150 and 160 of Figure 1. Figure 5 shows exemplary performance of exemplary systems, methods, and computer accessible medium according to exemplary embodiments of the present disclosure, after the exclusion of these patients, where Gap =1 indicates these samples were excluded and Gap = 0 indicates no exclusion. As expected, the AUROC within 2 years dropped from 0.760 (95% CI: 0.760, 0.761) to 0.731 (95% CI: 0.730, 0.733), and the AUROC within 5 years dropped from 0.735 (95% CI: 0.734, 0.736) to 0.721 (95% CI: 0.720, 0.722). Less impact on the prediction in the 5-year time frame appears to confirm the ability of models according to the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure to identify long-term disease (e.g., AD / ADRD / MCI) risks after eliminating the potential leakage due to the prolonged diagnosis process. Exemplary Model Predictions Correlating With Cognitive Impairment Levels

[0045] The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can hypothesize that patients with higher predicted risk according to the model, if undergoing cognitive screening, would also show a higher degree of impairment. From the technical standpoint, the correlation between the model score and cognitive test can indicate the calibration of the model predictions and is a desired model behavior. (see, e.g., Ref.28). Exemplary figure 6 demonstrates the relationship between the risk estimated by the model and the Mini-Mental Status Exam (MMSE) cognitive scores. (see, e.g., Ref.29). For patients in the held-out validation set who had undergone MMSE examination in the 1-year time interval after the feature window, the exemplary systems, methods, and computer accessibleAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION medium according to the exemplary embodiments of the present disclosure can extract the MMSE scores mentioned in their clinical notes via ChatGPT with high accuracy. (see, e.g., Ref.30). The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can then compare these scores with the predicted risk from the 1-year AD / ADRD / MCI prediction Transformer model. As seen in in the box plot of Figure 6, the higher the risk model estimates, the more likely the patient would have severe cognitive impairment (i.e. MMSE < 20). Exemplary Model Interpretation

[0046] To understand the relationship between the model’s predictions and patients’ EHR tokens contributing to each predicted score, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can compute the Integrated Gradients (see, e.g., Refs.31 and 32) of the prediction, with respect to the embedding layers of the EHR tokens. This can provide a personalized explanation per patient. A higher magnitude of the integrated gradient for a variable indicates greater influence from that variable in the model output. These gradients for each patient can be aggregated to explain overall variables associated with a higher AD / ADRD / MCI risk. The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can establish that this explanation method does not draw direct causal relationships. To aggregate the gradients, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can use the mean of gradient norms across all positively labeled AD / ADRD / MCI patients in the heldout validation set as the metric. Tables 3 and 7 list the variables with the greatest magnitudes of gradients among the positively labeled AD / ADRD / MCI patients in the held-out validation set, unveiling the features most associated with high AD / ADRD / MCI risks. Only tokens with more than 10 occurrences may be included to avoid outliers.

[0047] Several established and previously postulated risk factors of AD / ADRD / MCI emerge from these explanations, including vascular risk factors (such as heart disease, vascular disease, type 2 diabetes, hyperlipidemia, overweight) (see, e.g., Ref.33), weight loss (see, e.g., Ref.34), chronic kidney disease (see, e.g., Ref. 35), gait abnormality (see, e.g., Ref. 36), cataracts (see, e.g., Ref. 37), sleep apnea (see, e.g., Ref.38), convulsions / seizures (see, e.g., Ref.39), and mood disorderAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION (including anxiety and depression) (see, e.g., Refs.40, 41). In addition to those factors, the model according to the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can recognize symptoms or diseases known to be precursors to AD / ADRD syndromes, including urinary incontinence, falls, and parkinsonism / Parkinson’s disease. Furthermore, certain medications can also be highlighted, which may represent diagnoses that are AD / ADRD / MCI risk factors in the absence of formal diagnosis codes (e.g. escitalopram oxalate for depression or anxiety, atorvastatin for hyperlipidemia, metoprolol for hypertension), or may be linked via alternative relationships. This underscores the importance of utilizing interconnections among all EHR data to reveal a holistic picture of an individual’s health status. Exemplary Discussion

[0048] Recent studies have consistently highlighted the delayed diagnosis of AD / ADRD in primary care (see, e.g., Ref. 42), often attributed to inadequate training and resource constraints. Notably, later diagnosis of AD / ADRD varies by socioeconomic, geolocation, racial and ethnic groups due to the disparities in access to timely diagnosis and specialists. (see, e.g., Refs.43-45). Previous research has revealed lower documentation rates of cognitive assessments 1-5 years prior to AD / ADRD diagnosis among Black / African American patients, older individuals, those with non-commercial health insurance, or residing in neighborhoods with lower mean income levels. (see, e.g., Ref. 46). In this context, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure showcase the potential of TRADE, an EHR-based predictive model, to address this care gap by serving as a valuable tool for identifying high-risk patients who may benefit from targeted screening. Notably, the model according to exemplary systems, methods, and computer accessible medium according to exemplary embodiments of the present disclosure, exhibits improved performance compared to the best currently available tools, particularly among minority and high-risk groups, as illustrated in Figures 3 and 4.

[0049] While improvements can be consistently observed across various demographic groups, disparities in accuracy can persist notably among racial subgroups, particularly between White and Black populations. While the difference in prevalence might partially explain the lower performance within the Black population, the higher precision among the less common AsianAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION group contradicts such a hypothesis. This highlights the critical need not only for refining model development but also for recognizing the influence of social determinants of health (see, e.g., Refs. 47, 48) in deploying these models effectively. The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can facilitate analysis that delves into subcohort disparities in Mini-Mental Status Exam (MMSE) scores of patients who underwent screening and were diagnosed with AD / ADRD / MCI within the first year following the index date. In Figure 6a, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can illustrate a comparison of MMSE score distributions between White and Black subgroups, revealing a higher density in the Black group within the 20-24 MMSE range. While MMSE scores can be influenced by factors such as education level and education quality (see, e.g., Ref.49), this finding may also suggest that a significant portion of Black patients may receive diagnoses at later disease stages. Indeed, the lower Positive Predictive Value (PPV) in the Black population may stem from delayed or underdiagnosis, further emphasizing the value of deploying decision support tools such as TRADE in the EHR systems in achieving a more inclusive screening program for all demographics.

[0050] In addition to promising performance for identifying patients at risk for AD / ADRD / MCI, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can include an interpretability analysis, demonstrating that TRADE effectively extracts pertinent features from high-dimensional EHR data associated with an elevated AD / ADRD / MCI risk. The feature importance analysis underscores several established risk factors (e.g. vascular risks), and precursors strongly associated with AD / ADRD / MCI (e.g. monoclonal gammopathy). It is crucial to approach these associations with caution, recognizing them as correlations with potential AD / ADRD / MCI rather than direct causes. While further validation will help establish causality, these associations serve as valuable insights for driving future research into risk factors.

[0051] The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can demonstrate an improved performance of TRADE based on the self-supervised pre-trained foundation EHR model over the standard Transformer without pre-training. This suggests that our foundation EHR model has broader applicability in enhancing other EHR-based clinical tasks through fine-tuning with diverseAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION datasets. Furthermore, while current EHR data predominantly includes generic diagnoses, lab values, and medications, integration of additional patient data such as medical images and clinical notes, if available, could offer richer contexts. The architecture of our foundation model allows seamless integration of multi-modal data. Building prediction frameworks capable of leveraging the extra information can be an important direction for future research.

[0052] The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure represent a risk assessment framework for AD / ADRD / MCI using structured EHR. The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure demonstrate the efficacy of this model on the vast EHR of a healthcare system and its ability to significantly enhance the current clinical practices for the identification of AD / ADRD / MCI risks and screening programs. This improvement is particularly impactful for underrepresented and high-risk populations, where the model, according to the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure consistently outperforms existing methods. Meanwhile, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can show that the elevated predicted model scores are based on several recognizable risk factors or precursors of AD / ADRD / MCI. Finally, despite never directly training the model against cognitive screening scores such as MMSE, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure reveal that the higher risks predicted by the model do correlate with the degree of impairment, adding validity to the role of such models in the identification of undiagnosed AD / ADRD / MCI and tackling late diagnosis. The approach, according to exemplary systems, methods, and computer accessible medium according to exemplary embodiments of the present disclosure, to building predictive models based on the pre-trained foundation model of the EHR can benefit other diseases and similar risk assessment tasks, particularly in scenarios involving small and imbalanced patient cohorts. Exemplary Methods Exemplary Data

[0053] The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can use EHR data from, e.g., NYU LangoneAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION Health including any visits (inpatient, outpatient and ED) between Jan 2013 and Jan 2023. This study has been approved by NYU Langone Institutional Review Board (IRB), as protocol s20- 01095 Understanding and predicting Alzheimer’s Disease. Data can be acquired from NYU Langone DataCore in a de-identified manner. Only patients with more than 5 visits may be used. This provided the EHR of cohorts mentioned in the Exemplary Study Participants section.

[0054] In the records of each patient, each visit is represented as a combination of variables, which include, demographic information (age, gender, race, ethnicity, 30 variables), alongside observed lab results (represented as LOINC codes, 86,529 variables total), medication orders (represented at Pharmaceutical class, 162,761 variables total) and diagnosis codes (represented as ICD-10 codes and SNOMED, 196,983 variables total). To convert these variables into tokens for Transformer inputs, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can build a vocabulary list as follows: Each demographic category, diagnosis and medication code can be treated as a unique token. The continuous lab values can be binned into ranges of -10, -3, -1, -0.5, 0.5, 1, 3, 10 standard deviations from the population mean and then tokenized. Among all the variables, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can select 57,735 variables (6,482 lab values, 42,815 ICD-10 diagnosis codes, 8,387 medications, 50 demographics) with more than 100 occurrences. Any other variables may not be included in the model. Moreover, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can exclude ICD-10 codes related to amnesia, R41.1 / 2 / 3, to avoid potential leakage, which may not indicate AD / ADRD / MCI directly but could be used in preliminary diagnosis stages for these patients. Exemplary Model Development Exemplary ML Architecture

[0055] As the state-of-the-art model for learning representations of the structured EHR data (see, e.g., Refs. 21 and 22), a Transformer network (see, e.g., Ref. 19) can be employed by the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure to encode the EHR at the patient level. The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can organize the longitudinal EHR data for each patient sequentially based on visit dates resulting in a token sequence structured as follows (see, e.g., 130 of Figure 1):Attorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION “<CLS>”, {variables in 1st visit}, “<SEP>”, {variables in 2nd visit} , ..., “<SEP>”. Here, “<CLS> and <SEP>” serve as the placeholder tokens representing the outputs used for classification and separating consecutive visits, respectively. Each medical observation within a visit was represented by a token, with each token encoded using a learnable embedding vector within the Transformer architecture. To capture the sequential nature and continuous variables across visits effectively, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can introduce three learnable positional embeddings for each medical token , accounting for the patient’s age, the visit index, and the number of days since the beginning of the patient’s medical history. This refined approach can offer a more nuanced representation of age and the temporal gap between encounters, compared to BERHT and Med-BERT (see, e.g., Refs. 20, 21). Given that tokens within the same visit are permutation invariant, identical positional embeddings can be applied to all tokens within a given visit. Then these embeddings can be encoded by a Transformer neural network with 12 self- attention layers with 12 heads, 3072 dimensions for the feed-forward layers, and 786 dimensions for the encoder and pooling layers. Exemplary Model Pretraining

[0056] Although the large-scale Transformer network can have a strong capacity for encoding complicated high-dimensional data like EHR, it can also be vulnerable to overfitting when trained on limited data for a single binary classification task, especially when the low prevalence in disease prediction tasks leads to limited positive samples. (see, e.g., Ref. 22). To address this issue, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can conduct self-supervised learning to pretrain the model with a large amount of unlabeled data. The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can employ the masked token modeling as proposed in Bidirectional Encoder Representations from Transformers BERT (see, e.g., Refs.24, 25) for pretraining (see, e.g., 130 of Figure 1) where the network can be trained to predict randomly masked tokens (with a masking rate of 20%) from the model input. During pretaining, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can use a sliding window framework to process each patient’s data. The first visit of the sliding window can be randomly sampled and can includeAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION the EHR data from that visit up to a future visit such that the number of tokens in the EHR segment would be fewer than 512. In each epoch, for each patient in the training set (n=1,030,438), one randomly sampled window can be processed. The foundation model pre-training can converge in 500 epochs. An AdamW optimizer can be used at batch size 256 and a learning rate 2×10^-5 with a cosine learning rate decay. The computation can be executed on 4 Nvidia A100 GPUs. Exemplary Predictive Modeling for AD / ADRD / MCI

[0057] The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can train the AD / ADRD / MCI predictive model, TRADE, based on the pretrained foundation Transformer. To conduct a retrospective analysis, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can use a sliding window to create multiple samples from the medical trajectory of each patient (see, e.g., 150, 160 of Figure 1). The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can select the index visits with at least 180 days stride and assume them as the date for calculating the AD / ADRD / MCI prediction. The model can take the EHR within 1 year before the index date as the model input. Multiple binary labels can be created based on whether the patient was diagnosed with AD / ADRD / MCI within 1, 2, and 5 years after the index date. To avoid unrecorded diagnosis, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure may only keep samples that had follow-up encounters after the end of the outcome window. The max length of the tokens can also be limited to 512 latest tokens in the input. The Transformer model can be trained with the cross-entropy loss between the label and output of a linear classification layer at the Transformer encoding of the “<CLS>” token (see, e.g., 140 of Figure 1).

[0058] The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can use multiple ways to train the Transformer: (1) scratch, train the Transformer with randomly initialized weights; (2) linear probing only update the last linear layer of the Transformer while keeping the other layers frozen; (3) Finetuning using Low-Rank Adaptation (LoRA) method (see, e.g., Ref. 26) which injects and trains the rank decomposition matrices (with rank r = 128 and scale α = 256) into each linear layer of the Transformer architecture while keeping the original pretrained weights frozen. LoRA preventedAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION the immediate overfitting and outperformed the vanilla finetuning method that updated all the parameters. The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can train the model for the downstream task for 10 epochs with an AdamW optimizer at batch size 64 and a learning rate 10−5 with a 10−6 weight decay. The computation can be executed on 1 or 2 Nvidia V100 GPUs. The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can all be early stopped by the optimal average PPVs on the validation set.year models, with (gap=1), and without (gap=0) excluding patients whose AD / ADRD / MCI diagnosis occurred within 1 year of the end of the feature window. These ranked variables were extracted based on the magnitude of their gradients and indicate an association rather than causation relationship between the variable and AD / ADRD / MCI risk. The ranking of variables is determined by the mean of gradient norms across all patients in the heldout validation set with aAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION positive AD / ADRD / MCI label. The color of the cells indicates the category of variables based on either high-level ICD-10 categories or medication. Exemplary eRADAR baselines

[0059] The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can establish the benchmark for eRADAR prediction by evaluating it in the heldout validation set within AD / ADRD / MCI finetuning cohorts, using the same patients and index dates used for the Transformer model finetuning. The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can gather the retrospective encounter data of variables used in the eRADAR model at the index dates for each individual, including age, sex, diagnoses in the past 2 years (e.g. congestive heart failure, cerebrovascular disease, diabetes (complex or any), etc.) based on ICD- 10 codes, most recent vital signs for underweight (BMI <18.5), obese (BMI ≥30), high blood pressure (≥140 SBP or ≥90 DBP), healthcare utilization in the past 2 years (e.g. ≥1 outpatient visit, ≥1 emergency department visit, ≥1 language and learning visit, etc), and utilization of medications in the past 2 years for non-tricyclic antidepressant and sedative-hypnotic. The eRADAR risks were computed with the scoring function and the published weights on each variable. (see, e.g., Ref. 11). Exemplary Extracting cognitive impairment level from clinical notes

[0060] Mini-Mental State Examination (MMSE) (see, e.g., Ref.29) is an 11-question measure that tests the cognitive function of patients, where the scores range from 0-30 (28-30:Cognitively Normal, 25-27:MCI, 0-24:ADRD). The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can retrieve available MMSE scores of patients to study the association between estimated risks and cognitive impairment level. Among 44,222 samples from 14,131 patients identified as having AD / ADRD / MCI in the held-out validation set, 873 patients underwent MMSE cognitive test and were recorded by clinical notes at the institute of this study. From the clinical notes, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can extract the MMSE scores of the patients with OpenAI GPT-4 (see, e.g., Refs. 50,51). A HIPAA-compliant private instance of ChatGPT can be utilized to ensure data privacy.Attorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION Exemplary Performance Metrics

[0061] The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can compute the area under the receiver operating characteristic (AUROC) and the positive predicted values (PPV), at different sensitivity levels, which are widely used for measuring the predictive accuracy of binary classification tasks. To report the statistical significance of descriptive statistics, the exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can compute 95% confidence intervals, and the bootstrapping method with 100 bootstrap iterations.Each column reports variables across patients with varied time-to-dementia from the index date. The color of cells indicates the categorization of code types (e.g., medications) and ICD-10 code categories based on the initial character.Attorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION

[0062] Figure 7 shows a block diagram of an exemplary embodiment of a system according to the present disclosure. For example, exemplary procedures in accordance with the present disclosure described herein can be performed by a processing arrangement and / or a computing arrangement (e.g., computer hardware arrangement) 705. Such processing / computing arrangement 705 can be, for example entirely or a part of, or include, but not limited to, a computer / processor 710 that can include, for example one or more microprocessors, and use instructions stored on a computer- accessible medium (e.g., RAM, ROM, hard drive, or other storage device).

[0063] As illustrated in Figure 7, for example a computer-accessible medium 715 (e.g., as described herein above, a storage device such as a hard disk, floppy disk, memory stick, CD-ROM, RAM, ROM, etc., or a collection thereof) can be provided (e.g., in communication with the processing arrangement 705). The computer-accessible medium 715 can contain executable instructions 720 thereon. In addition, or alternatively, a storage arrangement 725 can be provided separately from the computer-accessible medium 715, which can provide the instructions to the processing arrangement 705 so as to configure the processing arrangement to execute certain exemplary procedures, processes, and methods, as described herein above, for example. Further, the exemplary processing arrangement 705 can be provided with or include an input / output ports 735, which can include, for example a wired network, a wireless network, the internet, an intranet, a data collection probe, a sensor, etc. As shown in Figure 7, the exemplary processing arrangement 705 can be in communication with an exemplary display arrangement 730, which, according to certain exemplary embodiments of the present disclosure, can be a touch-screen configured for inputting information to the processing arrangement in addition to outputting information from the processing arrangement, for example. Further, the exemplary display arrangement 730 and / or a storage arrangement 725 can be used to display and / or store data in a user-accessible format and / or user-readable format.Attorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION Exemplary Features

[0064] The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can utilize structured data which can include a large number of variables (e.g., 470k) from tables that are part of FHIR exchangeable variables (medications, diagnosis, labs, demographics). FHIR specifications make it possible for multiple systems (EPIC, Any patient-approved App) to use exemplary systems, methods, and computer accessible medium according to exemplary embodiments of the present disclosure.

[0065] The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can use comprehensive EHR information, including ICD codes, dedications, lab values and demographics (e.g., age, gender, races) of patients. Exemplary embodiments can build and train a bespoke tokenizer, which can incorporate EHR variables. The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can tokenize continuous lab values by bucketing them based on e.g., -10, -3, -1, -0.5, 0.5, 1, 3, and 10 standard deviations from normal values. ICD codes and medications and demographics can be treated as indicator tokens.

[0066] The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can encode the age, index of days, and encounters via a positional embedding mechanism, which allows exemplary embodiments to handle continuous variables. The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can fine-tune the foundation model for prediction tasks with the low-rank adaptation method which is normally used for large language models before but not on structured EHR datasets.

[0067] Exemplary models according to embodiments of the present disclosure can also provide a gradient-based framework for explaining a feature importance accounting for a model prediction. Previous approaches only visualize connections among different variables via attention maps.

[0068] A dementia risk assessment model currently applied in clinical practices (eRADAR) only uses <30 variables, such as gender, vascular risks, diabetes and etc., and uses a simple logistic regression on these variables to estimate the risk. The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can leverage the entire trajectory of patient EHR in the past 1 year, in a Transformer model, to assess the riskAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION of dementia, and can yield twice the positive predictive values (PPV) under the same sensitivity level, compared to eRADAR.

[0069] The exemplary systems, methods, and computer accessible medium according to the exemplary embodiments of the present disclosure can include clinical intervention. These interventions can include identifying patients with upcoming appointments in a given clinic within a set timeframe, assessing their risk by the AI model of exemplary systems, methods, and computer accessible medium according to exemplary embodiments of the present disclosure, and putting the risk score into EPIC. An alert can be shown to the doctors during the appointment, suggesting cognitive screening, and addressing cardiovascular problems for the patients (for example, addressing blood pressure has shown to delay onset of dementia in a recent large scale clinical trial SPRINT MIND.).

[0070] According to exemplary embodiments of the present disclosure, numerous specific details have been set forth. It is to be understood, however, that implementations of the disclosed technology can be practiced without these specific details. In other instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this description. References to “some examples,” “other examples,” “one example,” “an example,” “various examples,” “one embodiment,” “an embodiment,” “some embodiments,” “example embodiment,” “various embodiments,” “one implementation,” “an implementation,” “example implementation,” “various implementations,” “some implementations,” etc., indicate that the implementation(s) of the disclosed technology so described may include a particular feature, structure, or characteristic, but not every implementation necessarily includes the particular feature, structure, or characteristic. Further, repeated use of the phrases “in one example,” “in one exemplary embodiment,” or “in one implementation” does not necessarily refer to the same example, exemplary embodiment, or implementation, although it may.

[0071] As used herein, unless otherwise specified the use of the ordinal adjectives “first,” “second,” “third,” etc., to describe a common object, merely indicate that different instances of like objects are being referred to, and are not intended to imply that the objects so described must be in a given sequence, either temporally, spatially, in ranking, or in any other manner.

[0072] While certain implementations of the disclosed technology have been described in connection with what is presently considered to be the most practical and various implementations,Attorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION it is to be understood that the disclosed technology is not to be limited to the disclosed implementations, but on the contrary, is intended to cover various modifications and equivalent arrangements included within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.

[0073] The foregoing merely illustrates the principles of the disclosure. Various modifications and alterations to the described embodiments will be apparent to those skilled in the art in view of the teachings herein. It will thus be appreciated that those skilled in the art will be able to devise numerous systems, arrangements, and procedures which, although not explicitly shown or described herein, embody the principles of the disclosure and can be thus within the spirit and scope of the disclosure. Various different exemplary embodiments can be used together with one another, as well as interchangeably therewith, as should be understood by those having ordinary skill in the art. In addition, certain terms used in the present disclosure, including the specification and drawings, can be used synonymously in certain instances, including, but not limited to, for example, data and information. It should be understood that, while these words, and / or other words that can be synonymous to one another, can be used synonymously herein, that there can be instances when such words can be intended to not be used synonymously. Further, to the extent that the prior art knowledge has not been explicitly incorporated by reference herein above, it is explicitly incorporated herein in its entirety. All publications referenced are incorporated herein by reference in their entireties.

[0074] Throughout the disclosure, the following terms take at least the meanings explicitly associated herein, unless the context clearly dictates otherwise. The term “or” is intended to mean an inclusive “or.” Further, the terms “a,” “an,” and “the” are intended to mean one or more unless specified otherwise or clear from the context to be directed to a singular form.

[0075] This written description uses examples to disclose certain implementations of the disclosed technology, including the best mode, and also to enable any person skilled in the art to practice certain implementations of the disclosed technology, including making and using any devices or systems and performing any incorporated methods.Attorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATIONAttorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION Exemplary References: 1. 2023 alzheimer’s disease facts and figures. Alzheimers. Dement.19, 1598–1695 (2023). 2. Livingston, G. et al. Dementia prevention, intervention, and care: 2020 report of the lancet commission. Lancet 396, 413–446 (2020). 3. McGrath, E. R. et al. Blood pressure from mid- to late life and risk of incident dementia. Neurology 89, 2447, DOI: 10.1212 / WNL.0000000000004741 (2017). 4. The SPRINT MIND Investigators for the SPRINT Research Group. Effect of Intensive vs Standard Blood Pressure Control on Probable Dementia: A Randomized Clinical Trial. JAMA 321, 553–561, DOI: 10.1001 / jama.2018.21442 (2019). 5. Jm, M., Bs, M., Db, H. & Aa, L. Impact of pharmacological treatment of diabetes mellitus on dementia risk: systematic review and meta-analysis. BMJ open diabetes research & care 6, DOI: 10.1136 / bmjdrc-2018-000563 (2018). 6. Sabia, S. et al. Association of ideal cardiovascular health at age 50 with incidence of dementia: 25 year follow-up of Whitehall II cohort study. BMJ 366, l4414, DOI: 10.1136 / bmj.l4414 (2019). 7. van Dyck, C. H. et al. Lecanemab in Early Alzheimer’s Disease. The New Engl. J. Medicine 388, 9–21, DOI: 10.1056 / NEJMoa2212948 (2023). 8. Chen, R. et al. Developing Measures of Cognitive Impairment in the Real World from Consumer-Grade Multimodal Sensor Streams. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD ’19, 2145–2155, DOI: 10.1145 / 3292500.3330690 (Association for Computing Machinery, New York, NY, USA, 2019). 9. Liu, S. et al. Generalizable deep learning model for early Alzheimer’s disease detection from structural MRIs. Sci. Reports 12, 17106, DOI: 10.1038 / s41598-022-20674-x (2022). 10. Hansson, O., Blennow, K., Zetterberg, H. & Dage, J. Blood biomarkers for Alzheimer’s disease in clinical practice and trials. Nat. Aging 3, 506–519, DOI: 10.1038 / s43587-023- 00403-3 (2023). 11. Barnes, D. E. et al. Development and validation of eradar: A tool using ehr data to detect unrecognized dementia. J. Am. Geriatr. Soc. 68, 103–111, DOI: https: / / doi.org / 10.1111 / jgs.16182 (2020). https: / / agsjournals.onlinelibrary.wiley.com / doi / pdf / 10.1111 / jgs.16182. 12. Dublin, S. et al. The electronic health record risk of alzheimer’s and dementia assessment rule (eradar) brain health trial: Protocol for an embedded, pragmatic clinical trial of a low- cost dementia detection algorithm. Contemp. Clin. Trials 135, 107356 (2023). 13. Coley, R. Y. et al. External Validation of the eRADAR Risk Score for Detecting Undiagnosed Dementia in Two Real-World Healthcare Systems. J. Gen. Intern. Medicine 38, 351–360, DOI: 10.1007 / s11606-022-07736-6 (2023).Attorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION Che, Z., Purushotham, S., Cho, K., Sontag, D. A. & Liu, Y. Recurrent neural networks for multivariate time series with missing values. CoRR abs / 1606.01865 (2016). 1606.01865. Choi, Y., Chiu, C. & Sontag, D. Learning low-dimensional representations of medical concepts. AMIA Jt. Summits on Transl. Sci. proceedings. AMIA Summit on Transl. Sci.2016, 41–50 (2016). Shickel, B., Tighe, P., Bihorac, A. & Rashidi, P. Deep EHR: A survey of recent advances on deep learning techniques for electronic health record (EHR) analysis. CoRR abs / 1706.03446 (2017). 1706.03446. Steinberg, E. et al. Language models are an effective representation learning technique for electronic health record data. J. Biomed. Informatics 113, 103637, DOI: 10.1016 / j.jbi.2020.103637 (2021). Zhu, W. & Razavian, N. Variationally regularized graph-based representation learning for electronic health records. In Proceedings of the Conference on Health, Inference, and Learning, CHIL ’21, 1–13, DOI: 10.1145 / 3450439.3451855 (Association for Computing Machinery, New York, NY, USA, 2021). Vaswani, A. et al. Attention is all you need. Adv. Neural Inf. Process. Syst.30 (2017). Li, Y. et al. BEHRT: Transformer for Electronic Health Records. Sci. Reports 10, 7155, DOI: 10.1038 / s41598-020-62922-y (2020). Rasmy, L., Xiang, Y., Xie, Z., Tao, C. & Zhi, D. Med-BERT: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction. npj Digit. Medicine 4, 1–13, DOI: 10.1038 / s41746-021-00455-y (2021). Placido, D. et al. A deep learning algorithm to predict risk of pancreatic cancer from disease trajectories. Nat. Medicine 29, 1113–1122, DOI: 10.1038 / s41591-023-02332-5 (2023). Kaplan, J. et al. Scaling Laws for Neural Language Models, DOI: 10.48550 / arXiv.2001.08361 (2020). Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Burstein, J., Doran, C. & Solorio, T. (eds.) Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 4171–4186, DOI: 10.18653 / v1 / N19-1423 (Association for Computational Linguistics, Minneapolis, Minnesota, 2019). iu, Y. et al. RoBERTa: A Robustly Optimized BERT Pretraining Approach, DOI: 0.48550 / arXiv.1907.11692 (2019). ArXiv:1907.11692 [cs]. Hu, E. J. et al. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations (2022). Wornow, M., Thapa, R., Steinberg, E., Fries, J. & Shah, N. EHRSHOT: An EHR Benchmark for Few-Shot Evaluation of Foundation Models. In Oh, A. et al. (eds.) Advances in Neural Information Processing Systems, vol.36, 67125–67137 (Curran Associates, Inc., 2023). Liu, S. et al. Deep probability estimation. In Chaudhuri, K. et al. (eds.) Proceedings of the 39th International Conference on Machine Learning, vol.162 of Proceedings of Machine Learning Research, 13746–13781 (PMLR, 2022).Attorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION Folstein, M., Folstein, S. & McHugh, P. Mini-mental state examination (mms, mmse)[database record]. PsycTESTS Dataset. doi 10 (1975). Zhang, H. et al. Evaluating large language models in extracting cognitive exam dates and scores. medRxiv 2023–07 (2023). Sundararajan, M., Taly, A. & Yan, Q. Axiomatic Attribution for Deep Networks. In Proceedings of the 34th International Conference on Machine Learning, 3319–3328 (PMLR, 2017). ISSN: 2640-3498. Kokhlikyan, N. et al. Captum: A unified and generic model interpretability library for PyTorch, DOI: 10.48550 / arXiv.2009.07896 (2020). ArXiv:2009.07896 [cs, stat]. Viswanathan, A., Rocca, W. A. & Tzourio, C. Vascular risk factors and dementia: how to move forward? Neurology 72, 368–374, DOI: 10.1212 / 01.wnl.0000341271.90478.8e (2009). Wang, C. et al. Weight Loss and the Risk of Dementia: A Meta-analysis of Cohort Studies. Curr. Alzheimer Res.18, 125–135, DOI: 10.2174 / 1567205018666210414112723 (2021). Cheng, K.-C. et al. Patients with chronic kidney disease are at an elevated risk of dementia: a population-based cohort study in Taiwan. BMC nephrology 13, 129, DOI: 10.1186 / 1471- 2369-13-129 (2012). Wisniewski, T. & Masurkar, A. V. Gait dysfunction in Alzheimer disease. Handb. Clin. Neurol.196, 267–274, DOI: 10.1016 / B978-0-323-98817-9.00013-2 (2023). Wang, L., Sang, B. & Zheng, Z. The risk of dementia or cognitive impairment in patients with cataracts: a systematic review and meta-analysis. Aging & Mental Heal.28, 11–22, DOI: 10.1080 / 13607863.2023.2226616 (2024). Bubu, O. M. et al. Obstructive sleep apnea, cognition and Alzheimer’s disease: A systematic review integrating three decades of multidisciplinary research. Sleep Medicine Rev.50, 101250, DOI: 10.1016 / j.smrv.2019.101250 (2020). Stefanidou, M. et al. Bi-directional association between epilepsy and dementia: The Framingham Heart Study. Neurology 95, e3241–e3247, DOI: 10.1212 / WNL.0000000000011077 (2020). Fernandez Fernandez, R., Martin, J. I. & Anton, M. A. M. Depression as a Risk Factor for Dementia: A Meta-Analysis. The J. Neuropsychiatry Clin. Neurosci.36, 101–109, DOI: 10.1176 / appi.neuropsych.20230043 (2024). Patel, P. & Masurkar, A. V. The Relationship of Anxiety with Alzheimer’s Disease: A Narrative Review. Curr. Alzheimer Res.18, 359–371, DOI: 10.2174 / 1567205018666210823095603 (2021). 2021 alzheimer’s disease facts and figures. Alzheimer’s & Dementia 17, 327–406, DOI: https: / / doi.org / 10.1002 / alz.12328 (2021). https: / / alz- journals.onlinelibrary.wiley.com / doi / pdf / 10.1002 / alz.12328. Lin, P.-J. et al. Dementia diagnosis disparities by race and ethnicity. Alzheimer’s & Dementia 16, e043183, DOI: 10.1002 / alz.043183 (2020). Kim, N. Racial disparities in neurological care in the united states: An internal mechanism. HPHR 32, DOI: 10.54111 / 0001 / FF11 (2021).Attorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION Tsoy, E. et al. Assessment of Racial / Ethnic Disparities in Timeliness and Comprehensiveness of Dementia Diagnosis in California. JAMA Neurol.78, 657–665, DOI: 10.1001 / jamaneurol.2021.0399 (2021). Maserejian, N., Krzywy, H., Eaton, S. & Galvin, J. E. Cognitive measures lacking in EHR prior to dementia or Alzheimer’s disease diagnosis. Alzheimer’s & Dementia 17, 1231– 1243, DOI: 10.1002 / alz.12280 (2021). Majoka, M. A. & Schimming, C. Effect of Social Determinants of Health on Cognition and Risk of Alzheimer Disease and Related Dementias. Clin. Ther.43, 922–929, DOI: 10.1016 / j.clinthera.2021.05.005 (2021). Wu, W., Holkeboer, K. J., Kolawole, T. O., Carbone, L. & Mahmoudi, E. Natural language processing to identify social determinants of health in Alzheimer’s disease and related dementia from electronic health records. Heal. Serv. Res.58, 1292–1302, DOI: 10.1111 / 1475-6773.14210 (2023). Sisco, S. et al. The role of early-life educational quality and literacy in explaining racial disparities in cognition in late life. The Journals Gerontol. Ser. B, Psychol. Sci. Soc. Sci.70, 557–567, DOI: 10.1093 / geronb / gbt133 (2015). OpenAI. GPT-4 (2023). Zhang, H. et al. Evaluating large language models in extracting cognitive exam dates and scores. medRxiv: preprint server for health sciences 2023–07 (2024).

Claims

Attorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION WHAT IS CLAIMED IS 1. A method for generating at least one medical condition prediction, comprising: obtaining, by at least one computer processor, structured data from a pool of electronic health records (EHRs), the structured data comprising a plurality of variables from one or more tables; generating, by the at least one computer processor, a training data set from the structured data; encoding for each patient in the pool of EHRs, by the at least one computer processor, an age, an index of days, and one or more encounters via a positional embedding mechanism; training, by the at least one computer processor, a machine learning model using the training data and encoding; receiving patient data; and generating, by the at least one computer processor, the at least one medical condition prediction on the received patient data with the trained machine learning model.

2. The method of claim 1, further comprising, tokenizing, by the at least one computer processor, the plurality of numeric variables into a plurality of number ranges based on standard deviations from accepted normal values and wherein International Classification of Diseases codes, medications, and demographics variables are tokenized as indicator tokens.

3. The method of claim 2, wherein the training data set is generated from the tokenized plurality of variables.Attorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION 4. The method of claim 1, further comprising, fine-tuning, by the at least one computer processor, the machine learning model based the structured data.

5. The method of claim 4, wherein the at least one medical condition prediction is generated on the received patient data with the trained and finetuned machine learning model.

6. The method of claim 1, wherein the at least one medical condition prediction comprises a dementia prediction.

7. The method of claim 1, wherein the at least one medical condition prediction comprises a pancreatic cancer prediction.

8. A system for generating at least one dementia prediction, comprising: at least one computer processor which is configured to: obtain structured data from a pool of electronic health records (EHRs), the structured data comprising a plurality of variables from one or more tables; generate a training data set from the tokenized plurality of variables; encode for each patient in the pool of EHRs, an age, an index of days, and one or more encounters via a positional embedding mechanism; train a machine learning model using the training data and encoding; receive patient data; and generate the at least one dementia prediction on the received patient data with the trained machine learning model.Attorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION 9. The system of claim 8, further comprising, tokenizing, by the at least one computer processor, the plurality of numeric variables into a plurality of number ranges based on standard deviations from accepted normal values and wherein International Classification of Diseases codes, medications, and demographics variables are tokenized as indicator tokens.

10. The system of claim 9, wherein the training data set is generated from the tokenized plurality of variables.

11. The system of claim 8, further comprising, fine-tuning, by the at least one computer processor, the machine learning model based the structured data.

12. The system of claim 11, wherein the at least one dementia prediction is generated on the received patient data with the trained and finetuned machine learning model.

13. The system of claim 8, wherein the at least one medical condition prediction comprises a dementia prediction.

14. The system of claim 8, wherein the at least one medical condition prediction comprises a pancreatic cancer prediction.

15. A non-transitory computer accessible medium which includes software thereon for generating at least one dementia prediction, wherein, when at least one computer processor execute the software, the computer processor is configured to perform the procedures, comprising::Attorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION obtaining structured data from a pool of electronic health records (EHRs), the structured data comprising a plurality of variables from one or more tables; generating a training data set from the tokenized plurality of variables; encoding for each patient in the pool of EHRs, an age, an index of days, and one or more encounters via a positional embedding mechanism; training a machine learning model using the training data and encoding; receiving patient data; and generating the at least one dementia prediction on the received patient data with the trained machine learning model.

16. The non-transitory computer accessible medium of claim 15, further comprising, tokenizing, by the at least one computer processor, the plurality of numeric variables into a plurality of number ranges based on standard deviations from accepted normal values and wherein International Classification of Diseases codes, medications, and demographics variables are tokenized as indicator tokens.

17. The non-transitory computer accessible medium of claim 16, wherein the training data set is generated from the tokenized plurality of variables.

18. The non-transitory computer accessible medium of claim 15, further comprising, fine-tuning, by the at least one computer processor, the machine learning model based the structured data.Attorney Docket No.: 300712.WO.01-109197-0000173 PATENT APPLICATION 19. The non-transitory computer accessible medium of claim 18, wherein the at least one dementia prediction is generated on the received patient data with the trained and finetuned machine learning model.

20. The non-transitory computer accessible medium of claim 15, wherein the at least one medical condition prediction comprises a dementia prediction.

21. The non-transitory computer accessible medium of claim 15, wherein the at least one medical condition prediction comprises a pancreatic cancer prediction.

Citation Information

Patent Citations

  • Medical report coding with acronym / abbreviation disambiguation

    US20170199963A1

  • Machine-learning-based forecasting of the progression of alzheimer's disease

    US20190272922A1

  • Method for detection and diagnosis of lung and pancreatic cancers from imaging scans

    US20200160997A1

  • Monitoring System for Assessing Control of a Disease State

    US20200253547A1

  • Systems and methods for continuous cancer treatment and prognostics

    WO2022232850A1