Application of digital twins in healthcare and medicine

The healthcare platform leverages AI and cloud-based data repositories to process multi-modal healthcare data for personalized and predictive healthcare decisions, addressing integration challenges and enhancing patient care through real-time data analysis.

WO2026026949A1PCT designated stage Publication Date: 2026-02-05ANTINOUS TECHNOLOGY CO LTD

Patent Information

Application Number
PCT/CN2025/112123
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-04-11
Filing Date
2025-08-01
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing healthcare technologies struggle to provide personalized and predictive healthcare solutions due to challenges in integrating and analyzing multi-modal healthcare data, particularly in real-time, while addressing biological heterogeneity and ethical considerations.

Method used

A computer-implemented healthcare platform utilizing a cloud-based data repository and AI agents, including Large Language Models and continuous learning models, processes multi-modal healthcare data to generate personalized healthcare decisions and predictions, integrating data from wearable sensors and patient-specific avatars.

Benefits of technology

Enables real-time, personalized healthcare decisions and predictions, such as disease likelihood, treatment responses, and medical device design, by effectively managing and analyzing diverse patient data, thereby enhancing patient care and treatment efficacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025112123_05022026_PF_FP_ABST
    Figure CN2025112123_05022026_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein is a method of using digital twins for medical diagnosis, prediction, and / or prognosis of a disease or condition, such as diagnosis, prediction, and / or prognosis of one or more cancers, organ-specific age prediction, predicting of female aging, predicting hospital-acquired infection, predicting dialysis time series, predicting diabetes and complications, predicting pregnancy and child outcome, predicting heart failure, predicting myopia, predicting vision in anti-VEGF treatment, or any combination thereof. A system and non-transitory computer-readable storage medium configured for performing the method is also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

APPLICATION OF DIGITAL TWINS IN HEALTHCARE AND MEDICINECROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The present application claims priority to International Patent Application No. PCT / CN2024 / 109378 filed August 2, 2024, entitled “Digital Twins in Healthcare and Medicine, ” International Patent Application No. PCT / CN2024 / 142316 filed December 25, 2024, entitled “Application of Digital Twins in Healthcare and Medicine, ” and International Patent Application No. PCT / CN2025 / 088491 filed April 11, 2025, entitled “Application of Digital Twins in Healthcare and Medicine, ” all of which applications are incorporated herein by reference herein as if set forth in full. BRIEF SUMMARY

[0002] A reference herein to a patent document or any other matter identified as prior art, is not to be taken as an admission that the document or other matter was known or that the information it contains was part of the common general knowledge as at the priority date of any of the claims.

[0003] In some embodiments, a computer-implemented healthcare platform is provided herein. The healthcare platform comprises a) a data repository, wherein the data repository is configured to: i) store multi-modal healthcare data for a plurality of patients in the form of patient-specific avatars, ii) provide secure access to the multi-modal healthcare data for a plurality of patient-specific avatars corresponding to the plurality of patients by at least one Artificial Intelligence (AI) agent; and b) at least one Artificial Intelligence (AI) agent, wherein the at least one AI agent is configured to: iii) process the multi-modal healthcare data for at least one patient-specific avatar as input, and iv) output at least one healthcare-related prediction based on the multi-modal healthcare data for the at least one patient-specific avatar. In some embodiments, the data repository is cloud-based. In some embodiments, the multi-modal healthcare data is real-time data.

[0004] In some embodiments, the computer-implemented healthcare platform further comprises at least one wearable sensor configured to provide real-time healthcare data for at least one patient of the plurality of patients. In some embodiments, the at least one wearable sensor comprises a wearable sensor configured to monitor body temperature, heart rate, blood pressure, blood glucose level, blood oxygen saturation, level of physical activity, or sleep pattern.

[0005] In some embodiments, the healthcare platform is configured to: (i) upload the real-time healthcare data provided by the at least one wearable sensor to the cloud-based data repository, and (ii) update the real-time, multi-modal healthcare data associated with at least one patient-specific avatar corresponding to the at least one patient. In some embodiments, the real-time, multimodal healthcare data for each patient of the plurality comprises patient identification data, patient age data, patient gender data, patient clinical data, patient genomic data, patient proteomic data, patient metabolomic data, patient physiology data, patient lifestyle data, patient-related environmental data, healthcare data for a patient’s family, or any combination thereof. In some embodiments, the cloud-based data repository is configured to allow the patients of the plurality to control access to the real-time, multi-modal healthcare data associated with their corresponding patient-specific avatars. In some embodiments, the at least one AI agent comprises a Large Language Model (LLM) -powered AI model. In some embodiments, the at least one AI agent comprises a continuous learning AI model. In some embodiments, the at least one AI agent comprises a lifelong learning AI model. In some embodiments, the at least one AI agent is stored in the cloud-based repository.

[0006] In some embodiments, the at least one healthcare-related prediction comprises a likelihood that a given patient will develop a disease, a diagnosis of disease for a given patient, a likelihood that a given patient will respond to a given disease treatment, a treatment decision for a given patient diagnosed with a disease, or any combination thereof. In some embodiments, the disease is cancer, a cardiovascular disease, an autoimmune disorder, or an infectious disease. In some embodiments, the at least one healthcare-related prediction comprises a prediction of an optimized medical device design for a given patient. In some embodiments, the at least one healthcare-related prediction comprises a prediction for a surgical outcome for a given patient.

[0007] In some embodiments, a method for making a personalized healthcare decision for a patient is also provided. The method comprises a) accessing a cloud-based data repository configured to: i) store real-time, multi-modal healthcare data for a plurality of patients in the form of patient-specific avatars, ii) provide secure access to the real-time, multi-modal healthcare data for a plurality of patient-specific avatars corresponding to the plurality of patients by at least one Artificial Intelligence (AI) agent; b) accessing and providing the real-time, multi-modal healthcare data for at least one patient-specific avatar corresponding to a given patient as input the at least one AI agent, and c) processing the real-time, multi-modal healthcare data for the at least one patient-specific avatar using the AI agent to output at least one personalized healthcare decision for the patient. In some embodiments, the method further comprises generating the cloud-based data repository. In some embodiments, the method further comprises training the at least one AI agent.

[0008] In some embodiments, the real-time, multi-modal healthcare data for a plurality of patients comprises data from at least one wearable sensor configured to provide real-time healthcare data for at least one patient of the plurality of patients. In some embodiments, the at least one wearable sensor comprises a wearable sensor configured to monitor body temperature, heart rate, blood pressure, blood glucose level, blood oxygen saturation, level of physical activity, or sleep pattern. In some embodiments, the real-time, multimodal healthcare data for each patient of the plurality comprises patient identification data, patient age data, patient gender data, patient clinical data, patient genomic data, patient proteomic data, patient metabolomic data, patient physiology data, patient lifestyle data, patient-related environmental data, healthcare data for a patient’s family, or any combination thereof. In some embodiments, the at least one AI agent comprises a Large Language Model (LLM) -powered AI model. In some embodiments, wherein the at least one AI agent comprises a continuous learning AI model. In some embodiments, the at least one AI agent comprises a lifelong learning AI model.

[0009] In some embodiments, the at least one personalized healthcare decision comprises a lifestyle recommendation decision based on a predicted likelihood that a given patient will develop a disease. In some embodiments, the at least one personalized healthcare decision comprises a patient monitoring decision based on a predicted likelihood that a given patient will develop a disease. In some embodiments, the at least one personalized healthcare decision comprises a preventative treatment decision based on a predicted likelihood that a given patient will develop a disease.

[0010] In some embodiments, the at least one personalized healthcare decision comprises a treatment decision based on a predicted likelihood that a given patient diagnosed with a disease will respond to a given disease treatment. In some embodiments, the disease is cancer, a cardiovascular disease, an autoimmune disorder, or an infectious disease. In some embodiments, the at least one personalized healthcare decision comprises a decision of whether or not to recommend use of a specified medical device to a given patient diagnosed with a disease. In some embodiments, the at least one personalized healthcare decision comprises a decision of whether or not to perform surgery for a given patient.

[0011] In some embodiments, provided herein is the application of digital twins, for instance, using a transformer-based model, for organ-specific age prediction, predicting of female aging, predicting hospital-acquired infection, predicting dialysis time series, predicting diabetes and complications, predicting pregnancy and child outcome, predicting heart failure, predicting myopia, and / or predicting vision in anti-VEGF treatment.

[0012] In some embodiments, provided herein is a method of generating a trained large language model (LLM) , comprising (i) training an LLM using longitudinal multimodal data of a plurality of subjects (e.g., patients) , wherein the LLM comprises an examination encoder, a temporal embedding, and task-specific decoder heads, wherein for each subject, the longitudinal multimodal data comprise longitudinal electronic health record (EHR) data from a chronological sequence of clinical visits of the subject, and (ii) adapting the subject-level longitudinal representations to distinct tasks using the task-specific decoder heads, wherein the distinct tasks comprise first occurrence disease diagnosis and future disease prediction, thereby generating the trained LLM. In some embodiments, the trained LLM is configured to process the multi-modal healthcare data for at least one patient-specific avatar as input, and output at least one healthcare-related prediction based on the multi-modal healthcare data for the at least one patient-specific avatar.

[0013] In some embodiments, provided herein is a method of generating a trained large language model (LLM) , comprising (i) training an LLM using longitudinal multimodal data of a plurality of subjects, wherein for each subject, the longitudinal multimodal data comprise data from a chronological sequence of clinical visits of the subject, wherein the LLM comprises modality-specific alignment encoders, a decoder for cross-modal latent space prediction, and an autoregressive module, and wherein the pretraining comprises: (i) adversarially aligning the longitudinal multimodal data using the modality-specific alignment encoders, wherein each modality-specific alignment encoder embeds data of the corresponding modality into a unified, low-dimensional latent space, and wherein a modality discriminator identifies embedding sources in the latent space, (ii) using the decoder to predict a representation of a target modality in the latent space, wherein data from one or more clinical visits of a subject are missing in the target modality, and wherein the prediction imputes the missing data based on the subject’s historical data of the target modality and the subject’s historical data of one or more source modalities, thereby generating modality-invariant representations of the subjects in the latent space; and (ii) feeding the modality-invariant representations into the autoregressive module to model each subject's historical trajectory, thereby generating the trained LLM. In some embodiments, the trained LLM is configured to process the multi-modal healthcare data for at least one patient-specific avatar as input, and output at least one healthcare-related prediction based on the multi-modal healthcare data for the at least one patient-specific avatar.

[0014] In some embodiments, disclosed herein is a method of generating a trained large language model (LLM) , comprising: (a) using an LLM to categorize longitudinal electronic health record (EHR) data of a plurality of subjects into structured data and unstructured data, wherein for each subject, the longitudinal EHR data are from a chronological sequence of clinical visits of the subject; (b) using the LLM to extract biomedical concepts from both the structured data and the unstructured data as medically relevant tokens; (c) using the LLM to process the medically relevant tokens in chronological order to maintain the temporal coherence of subject health data, thereby predicting the next token in a patient’s timeline using temporal patterns of preceding tokens; and (d) using the LLM to measure and maximize the likelihood of predicting the next token correctly, thereby generating the trained LLM. In some embodiments, the trained LLM is configured to process the multi-modal healthcare data for at least one patient-specific avatar as input, and output at least one healthcare-related prediction based on the multi-modal healthcare data for the at least one patient-specific avatar.

[0015] In some embodiments, provided herein is a method of generating a trained large language model (LLM) , comprising: (a) using longitudinal electronic health record (EHR) data of a plurality of subjects to train an LLM, wherein for each subject, the longitudinal EHR data are from a chronological sequence of clinical visits of the subject and comprise complete blood count (CBC) data; (b) extracting biomedical concepts from the longitudinal EHR data as medically relevant tokens; (c) processing the medically relevant tokens in chronological order to maintain the temporal coherence of subject health data, thereby predicting the next token in a patient’s timeline using temporal patterns of preceding tokens; and (d) measuring and maximizing the likelihood of predicting the next token correctly, thereby generating the trained LLM. In some embodiments, the trained LLM is configured to process the multi-modal healthcare data for at least one patient-specific avatar as input, and output at least one healthcare-related prediction based on the multi-modal healthcare data for the at least one patient-specific avatar.BRIEF DESCRIPTION OF DRAWINGS

[0016] The drawings illustrate certain features and advantages of this disclosure. These embodiments are not intended to limit the scope of the appended claims in any manner.

[0017] FIG. 1. Basics of Digital Twins: A virtual entity, a physical entity, and an input data flow for real-time collection and monitoring of the physical entity's state or physiological functions, along with an output data flow for real-time interaction and communication, such as transmitting diagnosis and treatment solutions.

[0018] FIGS. 2A-2C. Building with artificial intelligence and metaverse. FIG. 2A, Building digital twins with large language models. FIG. 2B, Combining embodied AI with large language model-powered digital twins to construct AI agents. FIG. 2C, Metaverse provides a shared space for physical and virtual entities to communicate regarding patient care.

[0019] FIG. 3. Hallmarks of the Digital Twin platform. Any healthcare Digital Twin should include basic physical-virtual two-way communication, a metaverse of representative data, embodied AI agents based on LLM interfaces, reliable learning and prediction of multi-modal data, real-time patient monitoring, secure data storage, access to patient data, and adherence to ethical standards.

[0020] FIG. 4. Representative data repository as a metaverse. 1, The Digital Twin platform integrates extensive multi-modal biological omics and medical data from patients, generating algorithms for individualized guidance in prevention, risk assessment, and therapies. 2, The platform uses comprehensive patient input to match their virtual counterpart in the deeply phenotyped Digital Twin database, providing personalized treatment and prevention recommendations.

[0021] FIG. 5. Development of the Digital Twin platform. Static twins: These serve as the starting point, where physical entities are digitized, enabling periodic updates to their virtual counterparts. Progressive twins: By integrating temporal or progressive information, these twins can reflect the evolution of the physical entity and reliably forecast future state transitions. Operational twins: With the development of a closed-loop, iterative improvement framework, operational twins enable real-time interaction between physical and virtual entities. This facilitates both a deeper understanding of biological phenomena and the achievement of specific design objectives in biology and healthcare. Autonomous Twins: In the final stage, the digitized physical and virtual worlds merge, representing the highest level of physical-virtual co-existence. Autonomous virtual entities continuously generate information and knowledge for their associated physical entities within the Digital Twin platform.

[0022] FIG. 6 shows a pipeline for cancer detection / prediction using digital twin pretraining based on EHRs, described in Example 1.

[0023] FIG. 7 shows correlation among features.

[0024] FIG. 8 shows prediction of various tumors using digital twin trained using longitudinal EHRs.

[0025] FIG. 9 shows performance of cancer prediction models for different time periods.

[0026] FIG. 10 shows the impact of time series length on model performance.

[0027] FIG. 11 shows that longer time series improve prediction performance.

[0028] FIG. 12 shows SHAP values for features in the trained time-series cancer prediction model.

[0029] FIG. 13 shows the predictive performance of time series prediction models across different cancer categories.

[0030] FIG. 14 shows predicting organ specific biological age using individualized time serial patient EHR data with a transformer based digital twin pretraining paradigm (EHRFormer) . Stage 1: digital twin model pretraining based on time serial patient data; Stage 2: predicting organ specific biological age and age gap and their applications in aging and disease risk stratifications.

[0031] FIGS. 15-22 show data and results for organ-specific age prediction, described in Example 2.

[0032] FIGS. 23A-23D show predicted age generated by EHRFormer demonstrated a strong correlation with chronological age for each gender and age group, and gender differences in biological aging patterns. (A, B) Predicted biological age based on biomarker data from CHAI cohort reveals a strong correlation with chronological age for both females (A) and males (B) . However, a notable divergence occurs in older females, with biological age accelerating sharply between ages 40 to 50 (A) , a trend not observed in males (B) . (C, D) Validation of these findings using UKBiobank data confirms a similar trend, with biological age strongly correlating to chronological age for both genders. Accelerated aging is exclusively observed in women aged 40-50 (C, D) . These findings suggest that the onset of menopause may contribute to the rapid acceleration of biological age in females.

[0033] FIGS. 24A-24C show a reproductive system-specific aging model in females, and results of using female-specific ovarian aging clock to divide 40-70 year old Chinese female population into two categories (top 50%vs bottom 50%) . Women aging is divided into two categories (top 50%vs bottom 50%) according to female-specific age clock in CHN population. (A-C) Reproductive system-specific aging model, developed using biomarkers tied to menopause, demonstrates accelerated ovarian aging in females aged 40-50 in CHAI (A) . After stratification of women based on whether their predicted age is higher or lower than the mean predicted age for their chronological age group, violin plots show women in the higher predicted age to also demonstrated higher reproductive and endocrine (B) , as well as all-organ age (C) .

[0034] FIGS. 25A-25B show validation of the ovarian aging model in an independent UK cohort, and results of using female-specific reproductive aging clock to divide 40-70 UKB female populations are divided into two categories (top 50%vs bottom 50%) . Age predictions made on UKB cohort in women aged 40-70 y. o. showed accelerated aging during midlife as well (A) . After stratification of women based on whether their predicted age is higher or lower than the mean predicted age for their chronological age group, violin plot show the accelerated group (green) to have higher predicted organ age than decelerated group (blue) . The analysis of predicted organ ages in this independent UK cohort confirms the validity of the model, demonstrating that accelerated aging is characterized by significantly older predicted organ ages compared to the decelerated aging group.

[0035] FIGS. 26A-26K show comparative analysis of predicted organ ages in accelerated and decelerated aging groups, and female reproductive aging acceleration (divide into top 50%vs bottom 50%group by ovarian aging clock) is associated with widespread other organ specific aging judged by each organ specific clock in UKB cohort, and results of using female-specific reproductive aging clock to divide 40-70 years old UK female populations (UKB) into two categories (top 50%vs bottom 50%) and finding significant widespread organ specific aging. Female reproductive aging (divide into top 50%vs bottom 50%group by female reproductive aging clock) increases other organ specific aging judged by each organ specific clock in UKB. Comparative analysis of predicted organ ages across various organ systems reveals that females in the accelerated ovarian aging group exhibit significantly older predicted organ ages, while those in the decelerated aging group show significantly younger predicted ages. The analysis shows the least overlap between the two groups in organs such as the heart, liver, and kidneys, highlighting the strong discriminatory power of organ-specific predicted ages in distinguishing between accelerated and decelerated ovarian aging.

[0036] FIGS. 27A-27C show Estradiol (E2) levels in women with accelerated ovarian aging. (A) Estradiol (E2) levels decline significantly after the age of 50, with a more pronounced reduction observed in women exhibiting accelerated ovarian aging. (B, C) In women experiencing accelerated ovarian aging during their menopausal years (40-60 years old) , E2 levels are significantly lower compared to those with typical ovarian aging trajectories. These findings suggest a strong correlation between reduced E2 levels and accelerated ovarian aging, implicating hormonal dysfunctions in the progression of ovarian aging and related multi-organ dysfunctions. FIG. 27A shows that E2 levels are linked with female ovarian aging, left Y axis: AI predicted female age using reproductive aging clock; and Right Y axis: E2. FIG. 27B shows E2 levels in the top 50%vs bottom 50%of predicted aging individuals using AI in 40-50 year old Chinses population. FIG. 27C shows E2 levels in the top 50%vs bottom 50%of predicted aging individuals using AI in 50-60 year old Chinese population.

[0037]

[0038] FIGS. 28-30 show prediction of hospital-acquired infection based on time sequence data and digital twin technology.

[0039] FIGS. 31-35 show the analysis and results using digital twin paradigm to predict diabetes and complications.

[0040] FIGS. 36-37 show the analysis and results using digital twin paradigm to predict pregnancy and child outcome.

[0041] FIGS. 38-39 show the analysis and results using digital twin paradigm to predict heart failure.

[0042] FIGS. 40-48 show the analysis and results using digital twin paradigm to predict myopia.

[0043] FIGS. 49-50 show the analysis and results using digital twin paradigm to predict vision in anti-VEGF therapy using time series data.

[0044] FIG. 51 shows comparison of biological age prediction performance using different subsets of laboratory test data for male and female individuals. The panels illustrate the prediction results based on three feature sets: (1) CBC: complete blood count (CBC) data only, (2) non-CBC: laboratory data excluding CBC, and (3) all available laboratory data (combining CBC and non-CBC features) . The top row presents results for males, while the bottom row corresponds to females.

[0045] FIG. 52 shows comparison of pregnancy gestational age prediction performance using clinical laboratory test data. The three panels depict results based on different subsets of lab data: (1) CBC: complete blood count (CBC) data only, (2) non-CBC: laboratory data excluding CBC and (3) all available laboratory data (combining CBC and non-CBC features) . Each scatter plot shows the relationship between actual gestational age (x-axis) and AI-assessed predicted gestational age (y-axis) , with the diagonal line indicating perfect prediction.

[0046] FIG. 53 shows predicting organ specific biological age using individualized time serial patient EHR data with a transformer based digital twin pretraining paradigm. Using a Transformer model to predict physiological age based on a patient’s past laboratory test records, the physiological age of a specific organ could be obtained by based on the laboratory test indicators related to that organ. By comparing the patient's organ-specific physiological age to the population mean, the aging status of that organ can be assessed. If the participant’s organ physiological age is above the average level, it is considered to be over-aging, on the other hand, if it is below the average level, it is considered to be under-aged. The association between the age gap of different organs with diseases can be used as a biomarker for diagnosis and early warning of diseases in various organs. Predicting organ specific biological age using individualized time serial patient EHR data with a transformer based digital twin pretraining paradigm (EHRFormer) . Stage 1: digital twin model pretraining based on time serial patient data; Stage 2: predicting organ specific biological age and age gap and their applications in aging and disease risk stratifications. EHRFormer is a deep learning framework that conducts pre-training and downstream task fine-tuning based on laboratory tests and Vital Signs indicators from large-scale EHRs to construct biological clocks for various parts of the human body. a. The construction of organ-specific aging clocks mainly involves three steps: the collection and organization of EHR data, pre-training and fine-tuning based on EHR data to predict the aging condition of various organs, and correlation analysis between the aging conditions of the organs and diseases. b. In the pre-training stage, laboratory tests and Vital Signs indicators from time-series EHR records are used to predict the indicators at given future time points based on the previous n test results. c. During downstream task fine-tuning, only the indicators associated with specific organs are used for age prediction and age gap evaluation. d. The structure of EHRFormer.

[0047] FIG. 54 shows (a) : Predicted biological age based on biomarker data from the CHAI cohort reveals a strong correlation with chronological age in females (left) and males (right) ; A notable divergence occurs in older female but not males with biological age accelerating sharply between ages 40 to 55; (b) : UMAP visualizations for the CHAI cohort. The color of a plot indicates predicted age of the individual. In the female (left panel) , a sharp acceleration in predicted age could be seen in individuals between ages 40 and 55. In contrast, the male plots (right panel) do not show such a distinct age-related divergence or acceleration in aging within the same age range. (c-d) : Similar trends of (a) and (b) was also observed in the UK Biobank cohort. (g) - (h) individuals with 2 more standard deviations than the average predicted organ or system age of that chronological age group exhibited higher hazard ratios (g) and cumulative disease rates (h) for that organ or system.

[0048] FIG. 55 shows a female-specific ovarian aging clock. a. Using a female-specific reproductive aging clock, 40-70 year old female populations in CHAI and UKB cohorts were divided into two categories (top 50%vs bottom 50%) with accelerated aging from a 45-55 year old period. b-l. different organ or system ages of ovarian under-and over-aged female in CHAI cohort. Female reproductive aging is associated with widespread organ specific aging measured by each organ specific clock in both CHAI and UKB cohorts.

[0049] FIG. 56 shows diet-related metabolic disorder in POI patients. a. The serum concentration of protein in the POI cohort; b. The skeletal muscle mass of the POI cohort; c. The relative abundance of branched chain amino acid in the serum of POI cohort; d. The relative abundance of ceramide with various length of acyl chains in the serum of POI cohort; e. The level of E2 in the POI cohort; f-g. The level of Estradiol (E2) in women before or after the age of 50 in CHAI (f) or UKB (g) ; h. The predicted biological age of the participants in the POI cohort. *, p<0.05; **, p<0.01; ***, p<0.001; ****, p<0.0001.

[0050] FIG. 57 shows Ramelteon treatment alleviates female reproductive aging. a. The workflow of the phenotypic screening on human granulosa cell KGN cells. N=4; b. The treatment of ramelteon rescued E2 secretion in KGN cells. N=4; c. The treatment of ramelteon upregulated the serum level of E2 in the 18-month-old female mice. N=9-11; d. The treatment of ramelteon improved the fertility in the 18-month-old female mice. N=9-11; (e-g) The serum level of E2 (e) , pregnant rate (f) and changes of ovarian follicles (g) in the young female mice upon dietary BCAA restriction induced POI, with or without ramelteon treatment. N=5-7. *, p<0.05; **, p<0.01; ***, p<0.001; ****, p<0.0001. Error bars stand for SEM of biological repeats. S1, Primordial; S2, Primary; S3, Secondary; S4, Antral; S5, Atretic. Error bars stand for SEM of biological repeats. Chi-square test was used for the analysis of pregnant rate.

[0051] FIG. 58 and FIG. 59 are described in Example 12.

[0052] FIG. 60 shows E2 levels are linked with female ovarian aging. (A) Left Y axis: AI predicted female age using reproductive aging clock; Right Y axis: E2 levels, X axis, chronological age. (B) E2 levels in the top 50%vs bottom 50%of AI-predicted biological age individuals in 40-50 yr old CHAI population. (C) E2 levels in the top 50%vs bottom 50%of predicted aging individuals using AI in 50-60 yr old CHAI population.

[0053] FIG. 61 shows ramelteon treatment improved ovarian functions. (A-B) Gene Set Enrichment Analysis (GSEA) revealed that ramelteon treatment leads to the downregulation of genes related to steroid, cholesterol, thioester or protein synthesis (A) and upregulation of genes related ovarian development or function (B) . N=5; (C-D) The expression of CYP19A1 gene in KGN cells upon the treatment of ramelteon (C) or other 3 melatonin receptor agonists (D) . N=5. ****, p<0.0001. Error bars stand for SEM of biological repeats.

[0054] FIG. 62 shows a framework of biological aging clocks in a full lifecycle using an EHRFormer model leveraging longitudinal EHRs. a. The architecture for generating a comprehensive digital representation from longitudinal EHR data. The EHRFormer processes sequential EHR visits iteratively. The digital representation is optimized via tasks such as reconstruction, cohort discrimination, missing data discrimination, and next-visit prediction, which feed into downstream clinical tasks. b. The input-output mask reconstruction mechanism. An attention-based encoder processes masked input data (where "M" is for masked values and "NA" for missing values) to reconstruct the original data matrix, capturing feature interactions and imputing undetected indicators. c. The mitigation strategies for missing bias and cohort bias. A missing discriminator and a cohort discriminator are used to address biases related to missing data (labeled 0 or 1) and cohort differences (labeled A, B, C, D) , respectively. d. Biological aging clock in the full lifecycle. Depicts two biological aging clocks: a Development Clock for individuals under 18 years old and an aging clock for those over 18 years old, highlighting different life-stage health considerations. e. Diseases risk and development patterns. Shows how a cohort of individuals can be clustered based on disease risk. A graph illustrates the risk over time for different risk groups (high, medium, low) , depicting disease development patterns. f. Disease diagnosis and future predictions. EHRFormer takes longitudinal EHR data from multiple patient visits as input. It is used to generate current disease diagnoses and predict future disease occurrences for patients.

[0055] FIG. 63 shows overall biological age prediction model and difference in prediction for two age phases (CHAI) . a. Strong correction between the chronological age (CA) and biological age (BA) predicted by the EHRFormer-based age model, each dot represented one EHR data point; b. Heatmap showing the average values of multiple clinical parameters (z-score normalized across the population) clustered by their age-related trajectories, with data aggregated at 1-year intervals across the entire study population c. Correlation between CA and BA in the pediatric clock predicted by EHRFormer based age model, each dot represented one EHR data point under 18 years old; d. The SHAP values of the top 20 contributors in the BA prediction for EHR data under 18 years old; e. Correlation between CA and BA in the adult clock predicted by EHRFormer based age model, each dot represents one EHR data point over 18 years old; f. The SHAP values of the top 20 contributors in the BA prediction for EHR data over 18 years old.

[0056] FIG. 64 shows clusters generated based on EHRFormer-representations are informative of current and future health status. a. UMAP projection of patient visits colored by age (yellow to purple: young to old) . Each point represented one EHR data (one patient visit) ; b. Leiden cluster distribution of the UMAP in panel a; c-e. The prevalence proportion of a disease in the average average-aged group (left) ; the prevalence proportion of a disease in the over-aged (middle) , and the incidence proportion of a disease in the over-aged group (right) . f. The heatmap illustrates the adjusted risk ratios (ARRs) of 45 diseases in each of the 14 clusters of <12 years old pediatric EHR data. g. The heatmap illustrated the ARRs of 61 common diseases in each of the 50 clusters using>18 years old adult EHR data. Color from white to red indicates the ARRs in a specific cluster truncated at a maximum of 4.

[0057] FIG. 65 shows the performance of the EHRFormer-based disease predicting model. a. ROC curves of EHRFormer-based diseases predicting model in diagnosing different diseases; b. ROC curves of EHRFormer-based diseases predicting model in predicting future risk of different diseases; c. The AUC changes along with visit number in diagnosing different diseases; d. The AUC changes along with visit numbers in predicting different diseases in the future. e-j. Accumulated risk for various diseases based on a predictive model that categorized individuals into high, middle, and low-risk groups. The model demonstrates good predictive capability for these diseases, with a clear separation of risk groups occurring around the age of 10.As time progresses, the gap in accumulated risk between the high, middle, and low-risk groups becomes more pronounced. Representative diseases include hypertension, diabetes, malnutrition, digestive disorders, systemic lupus erythematosus, and asthma, showcasing the model's ability to predict disease risk and progression. k-p. Accumulated risk for various diseases based on a predictive model that categorized individuals into high, middle, and low-risk groups. The model demonstrates good predictive capability for all diseases, with a distinct separation of risk groups occurring around the age of 40. As time progresses, the gap in accumulated risk between the high, middle, and low-risk groups becomes more pronounced, showcasing the model's ability to predict disease risks and progressions over time. Representative diseases include obesity, meningitis, epilepsy, systemic lupus erythematosus, asthma, and arthritis, showcasing the model's ability to predict disease risks and / or onset and progressions.

[0058] FIG. 66 shows biological age prediction model for two genders. a. Correlation between chronological age (CA) and biological age (BA) predicted by EHRFormer based age model, each plot represents one male EHR data under 18 years old; b. The SHAP values of the top 20 contributors in the BA prediction for male EHR data under 18 years old; c. Correlation between CA and BA predicted by EHRFormer based age model, each plot represents one female EHR data under 18 years old; d. The SHAP values of the top 20 contributors in the BA prediction for female EHR data under 18 years old; e. Correlation between CA and BA predicted by EHRFormer based age model, each plot represents one male EHR data over 18 years old; f. The SHAP values of the top 20 contributors in the BA prediction for male EHR data over 18 years old; g. Correlation between CA and BA predicted by EHRFormer based age model, each plot represents one female EHR data over 18 years old; h. The SHAP values of the top 20 contributors in the BA prediction for female EHR data over 18 years old.

[0059] FIG. 67 shows UMAP visualizations of EHR data before and after batch effect elimination. Data distribution before batch effect elimination; a. Distribution of complete dataset colored by hospital source (Hosp #1-4) ; b. Distribution of individuals under 18 years; c. Distribution of individuals over 18 years; d-f. Corresponding age distribution heatmaps with color scale indicating individual age in years. g-l. Data distribution following batch effect elimination. g. Homogeneous distribution of the complete dataset; h. Distribution of individuals under 18 years; i. Distribution of individuals over 18 years; j-l. Corresponding age distribution heatmaps after batch effect elimination.

[0060] FIG. 68 shows validation of EHRFormer-based diseases predicting model in diagnosing in diagnosing diseases and predicting future diseases. a, c, e. ROC curves of EHRFormer-based diseases predicting model in diagnosing different diseases based on EHR of external validation cohort from 3 different hospitals; b, d, f. ROC curves of EHRFormer-based diseases predicting model in predicting future diseases based on EHR of external validation cohort from 3 different hospitals.

[0061] FIG. 69 shows the performance of EHRFormer-based disease predicting model in diagnosing and predicting common pediatric and adult diseases based on <12 and >18 EHR.

[0062] FIGS. 70A-70D show a large-scale, multimodal dataset that enables the development and validation of a predictive model.

[0063] FIGS. 71A-71D show the model accurately diagnoses prevalent cancers and predicts long-term incidence across major subtypes.

[0064] FIGS. 72A-72E show the model precisely delineates tumor stage, including fine-grained TNM classification.

[0065] FIGS. 73A-73F show the model predicts treatment-specific response to guide personalized therapy.

[0066] FIGS. 74A-74C show the model enables comprehensive prognostic surveillance by modeling tumor dynamics and forecasting clinical outcomes.

[0067] FIG. 75 shows the Universal Twin Framework (MIDAS) for integrating longitudinal multimodal data. a. a schematic diagram illustrating longitudinally collected multimodal data from multiple patients (P1 to Pn) , including data types such as lab tests, fundus images, X-rays, proteomics, and metabolomics. The data exhibits pervasive missingness across both modalities and visits; b. conceptual diagram mapping disparate data modalities into a unified latent space; c. the reconstructed latent space provides an integrated representation of the patient's physiological state, synthesizing imputed data across modalities; d. the detailed architecture of the MIDAS model processes multimodal data at visit n by first encoding available modalities (A, C) into a shared latent space. A generator then synthesizes missing modalities (B) conditioned on the historical state embedding Hn-1, with adversarial training ensuring fidelity through dual discriminators-one distinguishing real versus generated data, and another identifying originally missing modalities. The aggregated current visit embedding is subsequently fused with Hn-1via an autoregressive module to produce the final visit-state embedding Hn, which simultaneously serves as input for downstream tasks and propagates contextual information to the next visit.

[0068] FIG. 76 shows alignment of multimodal data in a unified latent space. a and c. UMAP visualizations of the integrated latent space embeddings for female (a) and male (c) subjects in the UK Biobank testing cohort (n=100,000) demonstrate significant overlap between proteomics-derived (blue) and clinical laboratory-data-derived (yellow) embeddings, confirming their successful alignment into a cohesive physiological manifold; b and d, Quantitative evaluation of cross-modal alignment quality for female (b) and male (d) cohorts. Box plots depict the distributions of pairwise cosine similarity between embedding pairs: Proteomics-Clinical, Proteomics-Metabolomics, and Clinical-Metabolomics. MIDAS achieves significantly higher alignment consistency versus benchmark models (SigLIP, CLIP) and the ablation variant lacking historical context (MIDAS w / o context) , with all box plots visualizing the median (center line) , interquartile range (IQR; box bounds) , and 1.5×IQR extent (whiskers) ; e and f. UMAP visualizations of single-modality embeddings-Proteomics, Clinical, and Metabolomics-for female (e) and male (f) cohorts depict a gradient from blue (younger) to red (older) based on chronological age. These reveal highly congruent biological aging trajectories across all three modalities, thereby affirming the retention of fundamental biological signals in the latent representation.

[0069] FIG. 77 shows identification of distinct biological aging trajectories from the unified latent space. a and m. UMAP visualization of integrated embeddings from female (a) and male (c) subjects in the UK Biobank test set, color-coded by chronological age (blue: younger, red: older) , reveals a strong association between the primary UMAP axis (UMAP1) and aging progression. This demonstrates the embedding space effectively captures fundamental biological aging signatures; b and l. within the identical UMAP space shown in panels (a) and (c) , subjects are stratified by UMAP2 coordinates into three distinct biological aging trajectory groups: Group 1 (red, accelerated aging) , Group 2 (yellow, normal aging) , and Group 3 (blue, delayed aging) , revealing aging patterns independent of chronological age. c-k and n-t. Kaplan-Meier analysis of major age-related disease incidence in the UK Biobank cohort stratifies subjects into biological aging trajectories: accelerated (red) , normal (yellow) , and delayed (blue) . Panels (c-k) and (n-t) display cumulative risk curves for females and males, respectively. The accelerated aging group shows significantly higher cumulative risk for chronic conditions-including chronic obstructive pulmonary disease (COPD) , heart failure, and chronic kidney disease-than the delayed group.

[0070] FIG. 78 shows cross-modal generation of synthetic proteomics improves disease risk prediction. a. schematic illustration of the MIDAS framework for conditional proteomics generation from clinical inputs. Leveraging the cross-modally aligned latent space, the model synthesizes biologically consistent proteomic profiles conditioned solely on clinical data streams; b. quantitative evaluation of synthetic proteomics fidelity. Box plots demonstrate statistically significant cosine similarity between MIDAS-generated proteomics embeddings and actual proteomics embeddings in a paired test cohort (n=10,000) , confirming biologically plausible cross-modal generation; c. workflow comparing the predictive performance of synthetic versus actual proteomics data within a comparative modeling framework. d-i. ROC curves for predicting six different diseases. Models using Basic Risk Factors (BRF) + MIDAS-generated proteomics (orange) achieve performance nearly equivalent to models using BRF+actual proteomics (blue) , validating the clinical utility of the synthetic data; j. Workflow of applying the generative model to a large cohort (n=440,000) lacking actual proteomics data to compare three prediction models; k-p. ROC curves for disease prediction in the cohort without actual proteomics. The model combining BRF with MIDAS-generated proteomics (blue) significantly outperforms models using BRF with clinical data (orange) or BRF alone (green) , demonstrating that synthesized proteomics capture valuable predictive signals beyond standard clinical measurements.

[0071] FIG. 79 shows MIDAS improves myopia progression and systemic disease prediction by integrating fundus images and clinical laboratory tests. a-c. performance comparison of axial length progression predictors over a 5-year risk horizon assessed through three complementary metrics: determination coefficient (R2; a) , mean absolute error (MAE; b) , and Pearson correlation coefficient (PCC; c) . The multimodal MIDAS model (laboratory tests + retinal fundus images, fine-tuned) demonstrates statistically significant superiority to all single-modality baselines-including conventional algorithms and MIDAS variants using only lab tests or fundus alone (p<0.01 all comparisons) . Shaded bands indicate 95%confidence intervals throughout. d-f. AUC performance for 5-year risk prediction of ischemic heart disease (d) , dementia (e) , and heart failure (f) demonstrates the multimodal MIDAS model (laboratory tests + retinal fundus; pretrained-finetuned) outperforms unimodal variants and standard baselines across all disease endpoints.DETAILED DESCRIPTION

[0072] The disclosures of all publications, patents, patent applications and published patent applications referred to herein are hereby incorporated herein by reference in their entirety. All publications, comprising patent documents, scientific articles and databases, referred to in this application are incorporated by reference in their entirety for all purposes to the same extent as if each individual publication were individually incorporated by reference. If a definition set forth herein is contrary to or otherwise inconsistent with a definition set forth in the patents, applications, published applications and other publications that are herein incorporated by reference, the definition set forth herein prevails over the definition that is incorporated herein by reference.

[0073] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0074] The digital twin is a concept used to create digital replicas of physical objects or systems. The dynamic, bi-directional link between the physical entity and its digital counterpart enables a real-time update of the digital entity. It can predict perturbations related to the physical object’s function. The obvious applications of digital twins in healthcare and medicine are extremely attractive prospects that have the potential to revolutionize patient diagnosis and treatment. However, challenges including technical obstacles, biological heterogeneity, and ethical considerations, make it difficult to achieve the desired goal. Advances in multi-modal deep learning methods, embodied AI agents, and the metaverse may mitigate some difficulties.

[0075] At the core of Digital Twins is a mathematical model that utilizes data gathered from the physical entity to update the digital counterpart. This iterative approach then allows data to be generated from the digital entity indistinguishable from the physical entity. Examples include jet engine performance evaluation and the development of smart cities. In the context of medicine, the physical entity can refer to the patient being studied in their “real-world” existence, incorporating all molecular, physiological, lifestyle and environmental information across time. FIG. 1 shows a virtual entity, a physical entity, and an input data flow for real-time collection and monitoring of the physical entity's state or physiological functions, along with an output data flow for real-time interaction and communication, such as transmitting diagnosis and treatment solutions. The virtual entity is therefore a digital replica of the patient, or even a virtual space of many digital patients. These digital replicas have characteristics similar to those of the patient, which enable predictions and simulations of the biological processes or disease states using data collected from the patient. The physical and virtual entities communicate via a physical-to-virtual connection to allow continuous update of the parameters that reflect the state of the physical entity. In essence, a medical Digital Twin represents a virtual representation where clinical and medical decisions can be tested before application in the actual patient. Therefore, this enables the real-time dynamic modeling of biochemical pathways, cells, tissues, diseases, and, ultimately, the entire human body, making personalized medicine a tantalizing reality.

[0076] The unique opportunities offered by Digital Twins address the human population as individuals to provide improved and personalized therapies and preventions. The construction of a Digital Twin in medicine can comprise the integration of diverse information such as clinical data, real-time physiological changes, and the omics of an individual. For example, Digital Twins may enable the provision of a personalized, on-demand risk profile for chronic diseases, offer lifestyle suggestions to mitigate these risks, deliver warnings about immediate health risks, and provide alerts for pre-emptive diagnostic tests. Evaluating individual patient responses to a particular drug and forecasting its efficacy and potential adverse effects will also be possible. Digital Twins therefore represent a faithful implementation of personalized medicine.

[0077] While the application of Digital Twins in healthcare will represent an essential step towards truly personalized medicine, the significant heterogeneity inherent in human populations in genomics, physiology, lifestyle and environment represents a real and significant hurdle. Significant developments in artificial intelligence (AI) , large language models (LLM) , and wearable devices may provide solutions to some of these hurdles.

[0078] One of the major challenges of a medical Digital Twin is the acquisition of sufficient data to make meaningful predictions about the physical entity (the patient) . The “All of Us” research program launched by the U.S. National Institutes of Health (NIH) in 2018 seeks to gather data from at least one million individuals to create one of the largest and most diverse datasets on health and genomics. Furthermore, next-generation sequencing, high-throughput multi-omics profiling and mass spectrometry can evaluate the transcriptome, methylome, proteome, histone post-translational modifications and the microbiome at unprecedented speed and scale. Nevertheless, the All of Us program focuses on the United States population which may limit any predictive outputs. Therefore, the blueprint of the All of Us program should be established in multiple countries to achieve a truly global and representative cohort.

[0079] In addition, to handle large amounts of data, the medical Digital Twin should be able to ensure real-time data collection, integration and interoperability among different platforms and systems. The maintenance of data fidelity is also very important. For instance, constructing a high-fidelity virtual patient is challenging due to the typically sparse communication rate between the physical and virtual entities compared to mechanistic processes. Therefore, continuous monitoring of static multi-modal health data, including clinical phenotypes and multi-omics such as genomics, metabolism, physiology and lifestyle parameters is required. There is also a need for health data standardization to enable data integration and interoperability among different Digital Twin providers. Furthermore, advances in biosensor technology have enabled real-time data capture using small, implanted biosensors. Small broadband acoustic and mechanical sensing devices can accurately and continuously measure respiratory airflow, intestinal motility and other physiological events such as the cardiac cycle. Self-sustaining wireless charging utilizing metamaterial surfaces has been explored to enable battery-less pacemakers. Soft wearable sensors with wireless communication capability will be the next frontier in population health data acquisition. Such wearable digital health technologies are developing rapidly to make previously unavailable data outside of the clinic, such as behavioral and physiological data, available to clinicians so that additional considerations can be provided to clinical decisions and diagnoses. In addition, facial, fundus and tongue images can be used to predict underlying pathologies such as cardiovascular disease and diabetes. These images can be obtained using standard imaging tools. The ultimate application of a medical Digital Twin is therefore an integration of these multiscale data to observe and predict deviations from the normal state of each individual. For example, diabetic patients can be provided with customized recommendations on how to improve their health, by tracking food consumption, physical activities and daily life routines. The medical Digital Twin platform can also search the virtual world for similar patients to glean peer insights on improving quality of life.

[0080] To construct a medical Digital Twin with high performance in making efficient and inclusive decisions, integrating large-scale AI models into healthcare is necessary. In addition, high-quality datasets are required to train these integrated AI modules. Data acquired from super cohorts, such as the All of Us program described above, are ideal for this purpose. Such integrated AI modules can take multi-modal data to make clinical diagnoses, predict treatment outcomes, and interpret radiographical images. Digital Twins created using large multi-modal AI models are more likely to mimic their real-life counterparts. Such approaches have been proposed in oncology, cardiovascular health, and neurodegenerative disorders. Platforms for collecting and providing large amounts of multi-modal data can also be established.

[0081] Recent advances in large language models (LLMs) , embodied AI and the metaverse provide exceptional opportunities to make medical Digital Twins a reality. LLMs refer to deep neural networks trained on vast amounts of text with billions of parameters that can understand and generate human-like text. Multi-modal LLMs, which encompass diverse input modes in addition to textual inputs within a unified framework, set the stage for a comprehensive approach to healthcare. Addressing these challenges will require a foundation AI system that is capable of integration and interpretation of multi-modal data and will output with both holistic and specialist modes. This offers transformative potential for various healthcare scenarios, from answering health questions to clinical diagnostics and mortality prediction. Leveraging LLMs to improve medical Digital Twin models will also lead to better predictions of disease progression and optimized treatment plans (FIG. 2A) . Through enhanced linguistic competencies, LLMs can convincingly replicate human-like thought patterns and emotional responses. These models can emulate behaviors, decision-making patterns, and personality traits previously perceived as human by tailoring their interactions based on natural language inputs. These unique characteristics will allow for customization and personalization, which align well with the need to construct medical Digital Twin models. LLMs can also collaborate with medical Digital Twin models to provide dynamic and real-time simulation of physiological processes and patient-specific health conditions. Moreover, the integration of LLMs with Digital Twins can enhance the ability to make informed decisions in scenarios where the underlying mechanisms are not clear.

[0082] Embodied AI learns from interactions with environments instead of static datasets (FIG. 2B) . In medicine, embodied AI has been playing important roles in mental healthcare and medical robotics. One successful story is the development and use of embodied AI robots as Digital Twins of psychotherapists that provide therapy interventions to children suffering from autism spectrum disorders. Incorporation of LLMs can also seamlessly integrate various models and modalities, as well as orchestrate complex tasks such as planning, scheduling and collaboration. This will pave the way for the development of versatile, general-purpose embodied AI systems. These LLM-powered embodied AI models are often referred to as “AI agents” . Such AI agents can perceive and interpret their surroundings, tackle and resolve problems, as well as exhibit intelligent behaviors (FIG. 2C) . This can be achieved autonomously in conjunction with other AI agents or in synergy with human beings. Combining Digital Twins with AI agents may redefine patient care, diagnosis and treatment. By leveraging insights gained from human behavior and decision-making processes, AI agents can be employed to build virtual patient models, autonomous medical robots or medical assistants. For example, a Digital Twin of a brain tumor patient undergoing surgery can be built, upon which AI agents can simulate the entire surgical procedure. This enables surgeons to plan the operation by assessing different entry points, angles and depths, which can reduce the risk of damaging healthy brain tissue. Moreover, the capability of AI agents to orchestrate multiple Digital Twin models is important, as complex clinical tasks often require the development of more than one Digital Twin.

[0083] The metaverse represents a collective virtual shared space created by converging virtually enhanced physical reality, augmented reality and the internet. It is a space where digital and physical worlds intersect and open new possibilities in various domains, including medicine. This raises the possibility that the metaverse could represent a virtual space where healthcare providers and patients can interact with Digital Twins in real-time, allowing for more collaborative and patient-centered care. For example, doctors could use the metaverse to remotely monitor patients' Digital Twins and adjust their treatment plans based on real-time data. In addition, the metaverse could provide a platform for multidisciplinary teams of healthcare professionals to collaborate and share insights, leading to more informed and effective treatment decisions. Implementing Digital Twins within a metaverse could also facilitate patient engagement and education, empowering patients to take a more active role in their healthcare. To build consistent digital models for physical objects in the metaverse, generative algorithms showed attractive potential. Deep generative models, including the GPT-3 (Generative Pre-trained Transformer) , DALL-E (released by OpenAI, as well as the booming diffusion model, have been used to create dynamic metaverse environments. Deep generative models have also been extensively used for de novo molecular designs, compound optimization and hit identification.

[0084] Though deep learning models have played a key role in solving important problems in computational biology, they are faced with challenges such as interpretability and generalization. The interpretability of AI is one of the key hurdles to building human trust in a medical Digital Twin model, because it requires AI to provide diagnostic or treatment evidence with high transparency and interpretability. Many approaches have been tested in this area. Saliency mapping has been used to demonstrate that networks learn patterns which agree with accepted pathologic features for Alzheimer’s disease. The visualization of convoluted neural network ensembles that classify estrogen receptors has also been used to provide interpretability to breast MRI predictions. Generative discriminative machines can handle confounding variables to increase confidence in predictions. Other approaches include interactive learning, causal reasoning, counterfactual reasoning and mental theory to construct interpretable AI models.

[0085] Integrating expert knowledge and clinical evidence to guide AI development is still challenging, resulting in some difficulty in revealing the underlying AI explanatory structures. In addition, the AI model in a medical Digital Twin platform requires high robustness and generalization when dealing with massive and multisource data. The robustness refers to the tolerance of the model to perturbations in the input data. Models with poor robustness are easily misled by tiny and simple perturbations in the input. Many methods for improving model robustness have been applied in the biomedical field, such as an adversarial attack algorithm in gastric cancer subtype analysis models. Platforms for evaluating the robustness of an AI model have also been proposed. The degradations in the performance of a model when evaluated on previously unseen data compared to data it has already seen is known as generalization. Data augmentation using generative adversarial networks can generate a large amount of training data to solve the problems of insufficient data and uneven distribution. Many generative adversarial networks have been proposed for data augmentation to improve the generalizability of AI models, including CycleGAN, pix2pix GAN, and Self-Attention GAN.

[0086] The performance of AI systems is known to deteriorate on older tasks during training, which is called catastrophic forgetting. This is a particular issue in implementing medical Digital Twins because continuous learning is necessary. Lifelong learning is a paradigm that allows continuous learning, and to retain prior experience with old tasks while learning new tasks. Such approaches should be a part of medical Digital Twin platform so they can adapt to the real world. Lifelong learning has been explored in the processing and interpretation of medical images, and elastic weight consolidation has been applied to learning normal brain structure and white matter lesion segmentation. Reduction of catastrophic forgetting has also been successful in cardiac ultrasound view classification and pneumothorax detection. It has also been used in dealing with modality and task transitions caused by changes in protocols, parameter settings, or different scanners in a clinical setting. Research on improving the lifelong learning ability of models will focus on the following aspects: task transfer and adaptation, overcoming catastrophic forgetting, exploiting task similarity, task-agnostic learning, noise tolerance, and resource efficiency and sustainability.

[0087] Digital Twin platforms require tremendous computing power. Quantum computing is well-suited for large-scale data processing, information modeling process and real-world and virtual world communication process. In addition, quantum imaging techniques, combined with quantum sensors and quantum dots, are likely to usher in a new era of medical imaging, and it is expected that quantum MRI machines will produce extremely precise imaging, with the potential to visualize individual molecules. Combined with AI, quantum computing can also be applied to interpret diagnostic images, identifying anomalies with greater precision than the human eye. Quantum sensors can also be applied to acquire multi-modal data, particularly in wearable devices, to allow for highly sensitive and accurate monitoring of the physical entity. Quantum dots can be used in conjunction with quantum computing to personalize drug design, potentially enabling tailor-made drugs for each patient, maximizing efficacy and minimizing adverse reactions. Similar approaches can also be used to develop radiation plans to kill cancer cells without harming healthy cells. In neurology, quantum computing can simulate complex neural networks, aiding in the understanding and treatment of neurological disorders. This can be applied with Digital Twins of the brain to accurately model the behavior of neurons and synapses, leading to more informed treatment strategies and a better understanding of disorders like Alzheimer's or Parkinson's . Quantum computing and Digital Twins can also be used to optimize hospital operations. Quantum algorithms can analyze patient flow, resource utilization, and staff scheduling datasets. This enables administrators to optimize the allocation of resources, reduce waiting times and enhance overall operational efficiency in healthcare facilities. Nevertheless, significant developments are needed before quantum methods can be scaled up for these approaches described above. Current challenges include the need for error correction, the stability of qubits as well as the development of scalable quantum hardware.

[0088] While acquiring data from a large population makes it possible to realize the application of medical Digital Twins, the challenge of data security and privacy / confidentiality are critical considerations. New rules and regulations prohibit institutions from exchanging medical data without patients’a pproval, resulting in the occurrence of data silos. Therefore, creative methods are required to coordinate data retrieval while protecting privacy. To address these challenges, federated learning (FL) is a promising technology to boost data collaboration across multiple centers rather than sharing raw data. FL sidesteps privacy barriers by allowing clients to update models locally and upload model parameters to the server until the global model gains stable results. Federated multi-modal learning has been implemented in predicting future oxygen requirements of symptomatic patients with COVID-19. Cross-silo FL is also an increasingly attractive solution for predicting heart disease hospitalization through electronic health records, while cross-device FL has been used to handle continuous health data from wearable devices to deliver personalized health insights. Swarm learning is another approach that builds a model independently on private data using blockchain technology. This can track and mediate access to health and genomic records. Additional challenges to accessibility and privacy lie in data heterogeneity, safety and model communication efficiency. Data heterogeneity can result in client shifts and degrade the convergence of predictive outputs. In addition, inversion attacks can reconstruct images from model weights or gradient updates with impressive visual details. Poisoning attacks damage the training of global models by deliberately uploading malicious local models, requiring additional privacy-enhancing techniques. Furthermore, convergence times for FL are limited to communication bandwidths that affect communication delay times, necessitating the development of communication-efficient FL.

[0089] Additionally, the "Privacy by Design" approach can enhance data security and privacy at the infrastructure level by implementing robust authentication and access control measures. This includes encrypting data to prevent unauthorized access, using secure protocols like HTTPS or VPNs to protect sensitive data during transmission, anonymizing or pseudonymizing sensitive information, keeping personally identifiable data locally, and establishing a robust system for regular anomaly detection and prevention. Adapting blockchain technology in a medical Digital Twin platform can also mitigate the problem of data tampering. The decentralized nature of blockchain technology provides transparency in consent management and allows patients to see who has access and for what purposes. This will facilitate data audit and allow data changes to be traced. Moreover, self-executing agreements based on predefined rules and conditions called "Smart Contracts" can convert physical data governance and regulatory requirements into digital processes. Additionally, tokenization capabilities of blockchain technology can facilitate individual data ownership. In summary, addressing data security, privacy protection, and data ownership is crucial in designing and implementing medical Digital Twin technology. This protects sensitive healthcare information, prevents unauthorized access or data manipulation, and fosters trust among stakeholders.

[0090] Several ethical issues related to the extensive collection of sensitive health information arise in any discussions of a medical Digital Twin. Therefore, the protection and governance of such collected data are among the top priorities. For example, a determined adversary can hack into a Digital Twin repository to potentially harm entire populations. The issue of multiple-use of the collected data also needs to be properly addressed and governed, to alleviate the inevitable concern that accumulated sensitive data could be used for purposes other than informing healthcare decisions, such as research, commercialization or surveillance.

[0091] The provision of informed consent will also be critical, particularly with regard to transparency about how the data will be used and who will have access to the data. While the benefits and risks of participating in medical Digital Twin projects can be clearly explained to patients in detail, informed consent for collecting individualized information from wearables is more difficult. Furthermore, the extent of multi-modal data involved in the meaningful implementation of medical Digital Twins raises privacy issues and patient confidentiality. One major ethical hurdle is re-identification from anonymized data. This is a particular problem when highly parametrized models such as neural networks are used, as a significant fraction of the training data can be reconstructed from the trained neural network model. Indeed, a recent systematic review found re-identification rates are high.

[0092] Data acquired from medical Digital Twin platforms will be used to inform clinical decisions, many of which may be life-altering. Where the burden of accountability lies is critically important regarding liability in case of errors or adverse outcomes. In addition, the accumulated data may contain unequal representations of certain demographic groups. This creates a void in the data for machine learning and affects the performance of the medical Digital Twin platform with respect to minority groups. The clinical decisions for patients in such groups, or those with lower socioeconomic status, may therefore contain a certain degree of bias and inequality. Furthermore, it can be envisaged, at least initially, that medical Digital Twin technologies will be implemented in settings where the more affluent will benefit. This may exacerbate the inequality gap and bias. The result may be a disproportionate health improvement for high versus low socioeconomic status patients. Unless addressed, these concerns will severely dampen the enthusiasm for the widespread implementation of a medical Digital Twin platform.

[0093] One additional concern is data ownership. Currently, consent for data provision is received from the data producers, which has led to a digital economy built on centralized data owned by large tech corporations. This system has resulted in a scenario where the data creators (the patients or participants in research) have limited control and oversight of downstream processes. Indeed, tech companies routinely trade and sell personal data for profit. Therefore, in addition to the privacy, ownership and security issues discussed above, there is also an economic and profit issue that complicates the ethics around medical Digital Twin initiatives. Therefore, a fundamentally different approach to data management philosophy may be needed for medical Digital Twin applications. One bold approach may be to empower individuals with full data ownership. This means medical Digital Twin applications will provide the data creator the right to keep, sell, donate or trade their personal data for research or drug discovery. This can be implemented using smart contracts with built-in economic compensation logic. Patients can be compensated when they choose to sell or trade their anonymized data based on their nuanced preference for privacy. Mechanisms to compensate individuals for their health data can also incentivize targeted health data collection for medical research and drug discovery. Such approaches can improve individual digital rights when AI and Big Data become indispensable components of modern medicine.

[0094] Individualized homeostasis monitoring -The classification of experimentally or clinically defined normal or healthy states differs slightly in each individual and cannot be extrapolated to large populations. Currently, treatment personalization relies on low-resolution data and a limited picture of the clinical history of a particular person. For example, there is not yet a clear understanding of “normal” blood pressure. The reasons may be due to the relatively sparse blood pressure measurements and the lack of assessment of the impact of physiological and behavioral patterns in any individual. Without a personalized definition of normal, it is difficult to detect deviations from normal which ultimately constitute the disease state. A medical Digital Twin can define normal in each individual through continuous feedback of information between the patient and their Digital Twin. Deviations from this normal state define disease, and treatments can be leveraged to predict intervention outcomes.

[0095] In some embodiments, disclosed herein are methods of using data relating to the homeostasis of an individual for generating a medical diagnosis, prediction, and / or prognosis for the individual, using a trained model based on the digital twin technology disclosed herein. In some embodiments, the data relating to the homeostasis of the individual comprise routine laboratory results, vital signs, imaging data, or any combination thereof. In some embodiments, the model has been trained using data relating to the homeostasis of a plurality of subjects, e.g., patients. In some embodiments, the data relating to the homeostasis of the plurality of subjects (e.g., patients having particular diseases such as cancers) comprise routine laboratory results, vital signs, imaging data, or any combination thereof, which are not collected for medical diagnosis, prediction, and / or prognosis of the particular diseases such as one or more cancers.

[0096] In some embodiments, the data relating to the homeostasis of the individual (for generating a medical diagnosis, prediction, and / or prognosis for the individual) and / or the data relating to the homeostasis of the plurality of subjects (for training an AI model, such as an LLM) comprise data of setpoints that reflect a deep physiologic phenotype of the individual or each subject. In some embodiments, the setpoints are hematologic setpoints, such as complete blood count (CBC) . In some embodiments, the data of the setpoints comprise hematocrit (HCT) , hemoglobin (HGB) , mean corpuscular hemoglobin (MCH) , mean corpuscular hemoglobin concentration (MCHC) , mean platelet volume (MPV) , platelet count (PLT) , red cell count (RBC) , red cell distribution width (RDW) , white cell count (WBC) , or any combination thereof. In some embodiments, the data of the setpoints comprise any one or more of those listed in Table 1 and / or Table 2. In some embodiments, the medical diagnosis, prediction, and / or prognosis is for all-cause mortality or a disease such as heart attack, stroke, diabetes, kidney disease, osteoporosis, thyroid dysfunction, iron deficiency, myeloproliferative neoplasm, a cancer, or any combination thereof, using data of setpoints and trained AI models disclosed herein. In some embodiments, training of the models comprises generating subject-specific representations (e.g., digital twins) using longitudinal data of the plurality of subjects, such as multimodal longitudinal data.

[0097] Cancer management –Medical Digital Twins can also realize the promise of precision oncology by integrating individual proteome and clinical data with population data. Such models continuously learn from new data, as well as individual patient care decisions from physicians, and can be used for real-time adjustment of treatments. This is particularly important in cancer recurrence or drug resistance, and patients may require different surgical, chemotherapy, or radiation regimens depending on the innate resistance of their particular tumor. Medical Digital Twin platforms can predict the onset of resistance and offer alternative treatment regimens based on the genome of individual patient tumors. Chemotherapy regimens also can be personalized depending on the patient’s metabolism to mitigate toxic side effects. Such models have shown promise in predicting treatment responses in triple-negative breast cancer. A Medical Digital Twin platform can also be used to predict metastatic disease through structured, consecutive radiology reports.

[0098] Cardiovascular disease –Improved survival and quality of life in cardiovascular disease are achieved by effective acute care and guideline-based risk factor management strategies. Digital Twins can be created from traditional simulation models and precursor models at different scales to create real-time, cyber-physical systems to provide tailored therapies. For example, an inverse analytic Digital Twin system can detect abdominal aortic aneurysm and its severity scores using neural networks. The Siemens Digital Heart is used to evaluate the success of cardiac resynchronization therapy by implanting virtual electrodes.

[0099] Immune responses -A medical Digital Twin platform can also play an essential role in autoimmune disorders and infectious diseases. This will require multi-modal, granular and integrated information at the molecular, cellular, tissue, organ and body levels. Such platforms can be used to predict the rejection of transplanted organs and the potential responses to immunosuppressive agents. It will also be useful in infectious diseases, particularly during a pandemic, to identify individuals who are susceptible to certain infections or at risk of fatal cytokine storms. It can also be used to predict protective immune responses and immune memory as a result of vaccination.

[0100] The design, manufacture, and implementation of medical devices -The design of customized medical devices is a great challenge. Creating Digital Twins of different anatomical structures in the body can simplify the design and implementation of customized medical devices. Dassault Systèmes, based in France, developed a model of the structure and function of the human heart using magnetic resonance imaging and electrocardiograms. This Living Heart Project evaluates the use of the model in the insertion, placement and assessment of pacemaker leads, and other cardiac medical devices. Further work from this collaboration will utilize Digital Twin technologies to improve the efficiency of medical devices in clinical trials and leverage simulation data as a source of evidence.

[0101] Surgery -Surgical interventions offer curative potential for many diseases without effective pharmacological treatment alternatives. However, the surgical procedure and some of the interventions given during the perioperative period may adversely affect patient outcomes. Digital Twins can be very useful in the perioperative period for the planning and simulation of the surgery itself, as well as for predicting surgical outcomes. A Digital Twin platform will also be beneficial for assessing the tolerance of surgery outcomes with respect to small human-induced variations during the surgical procedure, such as during complex heart surgeries like transcatheter aortic valve replacements. Digital Orthopedics has generated a Digital Twin of the foot and ankle, allowing surgeons to simulate surgery results and optimize surgical planning.

[0102] Hospital and nursing administration: Outside of the biological and clinical setting, Digital Twin technologies can impact the administration of large healthcare institutions. Digital Twin platforms created from electronic medical records and live physiological data from wearable devices can be leveraged to provide optimized and personalized medical and nursing services. The Verto Flow platform integrates patient data from various sources using AI algorithms, which healthcare professionals can use to optimize patient care. The ThoughtWire platform can simulate the health status of a patient, alerting doctors when a patient is likely to have a life-threatening complication, and offer suggestions based on predictions to mitigate the risks. Digital Twins of entire hospital workflows can be explored to optimize surgery schedules and simplify staffing requirements. This can improve overall hospital efficiency and shorten patient waiting times. Project BreathEasy, developed by OnScale, is a Digital Twin inspired lung model to assist clinicians in the prediction of the ventilation requirements of COVID-19 patients. This is particularly important in low-resource regions where ventilators are in short supply.

[0103] Synthetic biology circuits: The use of nature in engineering and industry holds immense potential, with synthetic biology emerging as a rapidly growing field with significant economic prospects. Advances in DNA synthesis and sequencing technology have greatly reduced the costs of constructing synthetic DNA genomes. Incorporating advanced microfluidics allows for the creation of cell-free biological components. Indeed, synthetic biology closely aligns with the original concept of Digital Twins. For example, the development of an artificial human heart required substantial Digital Twin input. Incorporating Digital Twin technology in synthetic biology could create autonomic biological modules for applications like "smart organs" , drug production or renewable energy. These biological modules can communicate with virtual entities on the medical Digital Twin platform using synthetic genetic circuits and optogenetic tools. The potential for biological computers suggests a future where control of the virtual entity is integrated within the physical entity. Advances in cellular state estimation and control methods can lead to intelligent designs for biological components, furthering the discovery of new technologies in the control and integration of cellular processes.

[0104] In an era marked by groundbreaking technological advancements, healthcare is on the brink of a revolutionary transformation with the integration of Digital Twin technologies. This paradigm shift towards personalized, data-driven healthcare is encapsulated in a medical Digital Twin platform characterized by five key hallmarks (FIG. 3) .

[0105] Representative Data Repository as a Metaverse: At the core of a medical Digital Twin platform lies the concept of a metaverse, a virtual space of large-scale and high-fidelity digital models and entities that can be used to adjust treatment, monitor response and track lifestyle modifications (FIG. 4) . This metaverse should also be where patient-specific Digital Twins and their multi-modal biological omics and medical data coexist, and enables the seamless sharing and interaction of healthcare data. It will also be a dynamic ecosystem that integrates a patient’s comprehensive and multi-modal input to find a matching virtual counterpart and offer individualized treatment and prevention recommendations. The recommendation can be further personalized to the patient based on virtual profile latent-space prototyping if necessary.

[0106] Real-Time Monitoring of Physical Entity: The second hallmark of a medical Digital Twin platform is the real-time monitoring of physical entities. Patient-specific avatars within the metaverse closely mirror patients' real-world health conditions. They can monitor vital signs, physiological parameters and treatment progress. This real-time monitoring ensures timely interventions and immediate response to anomalies and empowers patients with continuous access to their health data.

[0107] Reliable Predictions by Embodied AI Agents: Embodied AI agents are the third hallmark of a medical Digital Twin platform. They are the digital brains of the platform, constantly analyzing vast amounts of data to offer reliable predictions and insights. These AI agents draw from comprehensive datasets within the metaverse, including historical health data, treatment outcomes, and patient-specific parameters. By simulating different scenarios, they can predict how a patient's health will evolve, the effectiveness of potential treatments, and the probability of developing specific conditions. This proactive approach transforms healthcare from reactive to predictive, enabling early intervention and optimized treatment strategies.

[0108] Secure Data Access: The fourth hallmark of a medical Digital Twin platform is a robust emphasis on secure data access, where patient data are protected with the highest security and encryption standards. Patients have control over who can access their data, and healthcare providers are granted secure, role-based access to relevant patient information. This secure data access safeguards patient privacy and complies with stringent data protection regulations.

[0109] Ethical: Ethics is the fifth and perhaps the most vital hallmark of a medical Digital Twin platform. It upholds the principle that every healthcare advancement must align with patients' best interests. This platform adheres to the highest ethical data usage, research and care delivery standards. It ensures that patient consent is always obtained and patient rights are respected. The ethical commitment extends to using healthcare data to improve patient care and the broader healthcare community while maintaining transparency and trust.

[0110] The medical Digital Twin platform can be realized using a four-stage development roadmap based on increasing functionality and complexity (FIG. 5) .

[0111] Stage 1: Static twins

[0112] The simplest Digital Twin model starts with a patient model template based on retrospective data and a continuous learning process. The static twin is a traditional simulation and modeling exercise, where analysis is primarily performed offline and characterized by hypothesis-driven mathematical modeling. The static twin is obtained by modeling the state of a physical system through data collected from sensors. Static twins can therefore be considered as data-driven mathematical models of patients. One example of static twins is the HeartNavigator developed by Philips.

[0113] Stage 2: Progressive twins

[0114] The next step in the evolution of the Digital Twin platform should incorporate observational data to represent the patient’s current state and reliably forecast future state transitions. It will need existing techniques for simulation, model inference, data assimilation, and high-performance computing to build and test real-time, dynamic models on relatively large scales. Progressive twins integrate temporal or progressive information to construct a dynamic statistical machine learning model, which reflects the evolution of the physical entity and reliably forecasts future state transitions. Progressive twins are therefore in silico representations that dynamically reflect molecular, physiological and disease static across time (e.g., aging) . An example of progressive twins is the development of three-dimensional brain organoid cell culture models to recapitulate various aspects of human brain physiology in vitro and replicate basic disease processes of Alzheimer’s disease, amyotrophic lateral sclerosis and microcephaly.

[0115] Stage 3: Operational twins

[0116] One of the critical features of the Digital Twin concept is the physical-to-virtual connection. Operational twins are real-time, cyber-physical systems, which utilize a continuous connection to monitor state changes in the physical environment. Therefore, operational twins represent a real-time interaction between the physical and virtual entities in a closed-loop. For example, an automated insulin injector can be built where changes in the Digital Twin of the patient’s blood (data from glucose monitor) can be continuously monitored to determine accurate insulin dose injections throughout the day instead of relying on a fixed injection schedule.

[0117] Stage 4: Autonomous twins

[0118] In the ultimate stage of the Digital Twin platform evolution, known as the autonomous twins, the digital and physical worlds are merged, representing the pinnacle of physical-virtual co-existence. The self-sustaining virtual worlds operate independently while interacting seamlessly with the physical world. This integration can create a metaverse populated with countless autonomous, high-resolution virtual entities. For example, an autonomic Digital Twin brain could be developed, building on in silico representations and evolving autonomously. This can dynamically reflect the biophysical information of an actual brain over time, enabling effective enhancement interventions. Autonomous twins, combined with advanced virtual reality platforms, can also revolutionize surgical practice by providing realistic performance feedback on simulated procedures tailored to each patient. Autonomous twins can offer valuable insights and guidance for real-world decision-making. The ultimate form of autonomous twins could enable the realization of precision medicine by accelerating the discovery of medical phenomena and disease processes, shortening the timeline for drug discovery, improving surgical outcomes via virtual operations and simulating disease progression statistics.

[0119] Creation of personalized treatment plans: A medical Digital Twin platform can use a cancer patient’s medical history, family history, genetic information and lifestyle factors (diet, exercise, and exposure to environmental toxins) to create an AI agent. This AI agent, utilizing a LLM, captures the patient’s unique physiological responses and medical conditions, leveraging external tools for diagnosis and treatment. It can perform self-diagnosis by accessing the latest medical research, clinical trials and treatments. It can consider various treatment options, such as chemotherapy, radiotherapy, immunotherapy, and targeted therapy and assess their potential effectiveness. The AI agent can also analyze the patient’s genetic data to identify mutations or biomarkers that could be indications or contraindications for specific therapies. The AI agent interacts with the patient and the healthcare provider to address concerns promptly. As new data and feedback are received, the platform uses reinforcement learning to refine its models, improving the accuracy and effectiveness of personalized treatment plans over time. Once developed, this agent can be easily customized for other patients based on their personal data.

[0120] Remote patient monitoring: Patients with chronic illnesses, such as hypertension or diabetes, can be equipped with a wearable device that continuously streams their vital signs and health data to their virtual counterparts within the metaverse. The embodied AI agent can analyze the incoming data to detect any irregularities or signs of deterioration. If a critical situation arises, the system can alert the patient and / or healthcare provider. It can even initiate a predefined emergency response. This proactive approach to remote monitoring allows patients to maintain their health from the comfort of their homes or anywhere while receiving immediate interventions when necessary.

[0121] Virtual clinical trials: Clinical trials are essential for testing new medications and treatments. A medical Digital Twin can revolutionize clinical trials by simulating patient responses in the metaverse. Researchers can use the Digital Twin platform to represent virtual patients with specific conditions and characteristics. These virtual patients are subjected to various treatment regimens, minimizing the risks and ethical concerns associated with actual patients. The embodied AI agents within the Digital Twin platform can analyze the treatment outcomes, providing valuable insights into the potential efficacy and safety of the treatments. This approach expedites the drug development process, reduces costs and accelerates the availability of new therapies to actual patients.

[0122] Hospital administration: Integrating Digital Twin platforms in hospital administration can substantially improve management of healthcare facilities. These technologies create virtual representations of hospital infrastructure, allowing administrators to monitor and manage operations in real-time. For instance, a virtual operating room can provide insights into equipment usage, maintenance needs and staff workflows, optimizing space layout and resource utilization. Clinical workflows and administrative processes can be analyzed and optimized within the virtual environment. Nursing administrators can simulate scenarios to identify bottlenecks and streamline workflows. Virtual entities of nursing staff offer real-time insights into availability, skills and workload. The Digital Twin platform can also simulate patient flows to optimize bed management and predict congestion points. Additionally, the metaverse serves as a training ground for medical staff, allowing them to practice critical care, patient interactions and new technologies in a risk-free environment.

[0123] The Digital Twin concept has proven invaluable in industrial applications, from manufacturing to the safe operation of complex systems. Its potential in developing in vitro and in vivo research models is also evident in biomedical research. However, its most transformative application lies in clinical medicine, where Digital Twin technologies could realize personalized medicine. By combining high-throughput genetic and molecular approaches, single-cell and whole-genome sequencing, Big Data, cloud-based electronic medical records, and AI, Digital Twins can deliver modern healthcare.

[0124] Beyond offering personalized treatment regimens, Digital Twin technologies can monitor and predict adverse drug reactions or interactions. Their greatest impact, however, will be in the day-to-day health monitoring of individuals. The ability to precisely predict health perturbations and provide mitigation suggestions will advance the detection and diagnosis of chronic, non-communicable diseases. I. Early Disease Detection Using Longitudinal Patient Data

[0125] Early detection of diseases such as cancers is one of the most critical factors in improving treatment outcomes and patient survival rates. Identifying cancer at its initial stages allows for more effective treatment and significantly reduces the burden and cost associated with advanced disease progression. Traditional methods of cancer detection, such as imaging, blood tests, and tissue biopsies, are effective in certain cases but may have limited sensitivity and specificity in the early stages of cancer. Moreover, accessibility issues further exacerbate the problem, with advanced screening technologies often unavailable in low-resource settings and high costs limiting access for underprivileged populations. The complex and heterogeneous nature of cancer, characterized by diverse subtypes and genetic mutations, further complicates early detection efforts. Addressing these challenges is crucial for improving early cancer detection and ultimately enhancing patient outcomes.

[0126] Electronic health records (EHRs) have revolutionized the way patient information is collected, stored, and utilized in healthcare settings. By digitizing patient data, EHRs provide a comprehensive and accessible repository of health information, which holds immense potential for clinical research. The analysis of EHRs has become a cornerstone in modern medical research, offering invaluable insights into patient care, disease progression, and healthcare outcomes.

[0127] EHR data is categorized into structured and unstructured formats, and the integration of both in research allows for a more holistic understanding of patient health and disease progression. Structured data includes standardized information (e.g. demographics, laboratory results, medication lists, and diagnosis codes) that have a predefined format, making it easily searchable and analyzable. In contrast, unstructured data consists of narrative text and other free-form entries (e.g. doctor’s notes, imaging reports, and clinical correspondence) that is more challenging to analyze due to its lack of standardization. However, current models, such as BEHRT, G-BERT, and Med-BERT, all primarily rely on structured data that are more readily accessible for computational analysis and integration into predictive models (Johnson et al., 2016) . Models like BEHRT (Bidirectional Encoder Representations from Transformers for EHRs) and G-BERT (Graph Neural Networks and BERT) predominantly utilize structured data, limiting their ability to capture the full scope of patient information (Li et al., 2020; Shang et al., 2019) . BEHRT, for instance, focuses on forecasting disorders within a predefined timeframe using structured data, but struggles with the increasing complexity of forecasting numerous concepts simultaneously. Similarly, G-BERT relies on single-visit samples, which are insufficient for capturing long-term contextual information necessary for comprehensive patient analysis (Shang et al., 2019) .

[0128] The reliance on structured data alone restricts the potential of these models to provide a holistic view of patient health. Recent advancements in natural language processing (NLP) and machine learning offer promising solutions to this challenge, enabling the extraction and analysis of meaningful information from unstructured text (Huang et al., 2020) .

[0129] Electronic health records (EHRs) play a crucial role in early cancer detection. By comprehensively analyzing longitudinal patient health data, including symptoms, laboratory results, imaging reports, and physician notes, potential early signs of cancer can be identified. This integrated approach leverages the wealth of information contained in EHRs to enhance predictive modeling and improve early detection strategies, ultimately contributing to better clinical outcomes for patients.

[0130] In some embodiments, provided herein is a model designed for early cancer detection by leveraging both structured data (such as age, ethnicity, and sex) and unstructured data from Electronic Health Records (EHRs) . In some embodiments, the model excels in performing a variety of tasks without the need for fine-tuning, including forecasting cancer risks, providing differential diagnoses, and suggesting relevant treatments. By analyzing comprehensive datasets from multiple hospitals, the model can identify early signs of cancer across diverse populations and health conditions. The model has demonstrated its efficacy in detecting early-stage cancers through rigorous testing on datasets encompassing both physical and mental health events. Accessible to healthcare professionals and researchers via a web application, the model represents a significant advancement in utilizing AI for personalized and timely cancer care, ultimately aiming to improve patient outcomes through early diagnosis and intervention.

[0131] In some embodiments, the efficacy of the generative transformer model is evaluated in temporal modeling of cancer patient data. In some embodiments, the model integrates both structured and unstructured formats from electronic health records (EHRs) , which contain comprehensive longitudinal information on a patient’s health status and clinical history. Unlike existing approaches that predominantly utilize structured data and focus on single-domain outcomes, the model aims to predict a wide range of future medical outcomes such as disorders, substances (including medications, allergies, and poisonings) , procedures, and findings (related to observations, judgments, or assessments) . By leveraging both free text and structured data within EHRs, the model seeks to enhance the predictive capability across diverse medical scenarios.

[0132] In some embodiments, the model is a transformer-based pipeline developed to convert electronic health record (EHR) text into structured, coded concepts and to forecast future medical events, including disorders, substances, procedures, and findings. In some embodiments, the pipeline comprises four main components: (1) CogStack: handles data retrieval and preprocessing; (2) Medical Concept Annotation Toolkit: organizes free-text data from EHRs; (3) MetaGP Core: employs deep learning for modeling biomedical concepts; and (4) MetaGP Web Application: enables user interaction with the system. In some embodiments, the system is applied to datasets (e.g., hospital datasets) which include both EHRs and laboratory test records. In some embodiments, model performance is assessed using custom metrics based on precision and recall.

[0133] In some embodiments, disclosed herein is a method of generating a trained large language model (LLM) , comprising: (a) using an LLM to categorize longitudinal electronic health record (EHR) data of a plurality of subjects into structured data and unstructured data, wherein for each subject, the longitudinal EHR data are from a chronological sequence of clinical visits of the subject; (b) using the LLM to extract biomedical concepts from both the structured data and the unstructured data as medically relevant tokens; (c) using the LLM to process the medically relevant tokens in chronological order to maintain the temporal coherence of subject health data, thereby predicting the next token in a patient’s timeline using temporal patterns of preceding tokens; and (d) using the LLM to measure and maximize the likelihood of predicting the next token correctly, thereby generating the trained LLM.

[0134] In some embodiments, the method comprises converting free-text information from the longitudinal EHR data into structured data. In some embodiments, the structured data comprise demographics, laboratory results, medication lists, and / or diagnosis codes, and the unstructured data comprise doctor’s notes, imaging reports, and / or clinical correspondence.

[0135] In some embodiments, the longitudinal EHR data comprise routine laboratory results, vital signs, imaging data, or any combination thereof, and the trained LLM is used for early cancer detection, wherein the routine laboratory results, vital signs, and / or imaging data are not collected in association with generating a medical diagnosis, prediction, and / or prognosis of an oncologic indication. In some embodiments, the routine laboratory results comprise results of any one or more of the biomarkers listed in Table 1 or Table 2. In some embodiments, the routine laboratory results comprise complete blood count (CBC) , blood chemistry, coagulation, or any combination thereof; wherein the vital signs comprise heart rate, blood pressure, body mass index (BMI) , or any combination thereof; and / or wherein the imaging data comprise chest X-rays (CXRs) .

[0136] In some embodiments, the method further comprises receiving a natural-language prompt for obtaining the medical diagnosis, prediction, and / or prognosis of an oncologic indication and a set of data related to an individual, and generating the medical diagnosis, prediction, and / or prognosis by inputting the prompt and the set of data in the trained LLM, wherein the set of data comprise routine laboratory results, vital signs, imaging data, or any combination thereof, which are not collected in association with cancer diagnosis, prediction, and / or prognosis.

[0137] In some embodiments, the method further comprises receiving a natural-language prompt for obtaining the medical diagnosis, prediction, and / or prognosis of hospital acquired infection, community acquired infection, and / or non-infection, and a set of data related to an individual, and generating the medical diagnosis, prediction, and / or prognosis by inputting the prompt and the set of data in the trained LLM.

[0138] In some embodiments, the method further comprises receiving a natural-language prompt for obtaining the medical diagnosis, prediction, and / or prognosis of organ specific aging, and a set of data related to an individual, and generating the medical diagnosis, prediction, and / or prognosis by inputting the prompt and the set of data in the trained LLM. In some embodiments, acceleration of the organ specific aging is associated with a disease or condition. In some embodiments, the organ specific aging is ovarian aging.

[0139] In some embodiments, the method further comprises receiving a natural-language prompt for obtaining the medical diagnosis, prediction, and / or prognosis of nosocomial infection, and a set of data related to an individual, and generating the medical diagnosis, prediction, and / or prognosis by inputting the prompt and the set of data in the trained LLM.

[0140] In some embodiments, the method further comprises receiving a natural-language prompt for obtaining dialysis time series prediction, and a set of data related to an individual, and generating the dialysis time series prediction by inputting the prompt and the set of data in the trained LLM.

[0141] In some embodiments, the method further comprises receiving a natural-language prompt for obtaining the medical diagnosis, prediction, and / or prognosis of diabetes, and a set of data related to an individual, and generating the medical diagnosis, prediction, and / or prognosis by inputting the prompt and the set of data in the trained LLM.

[0142] In some embodiments, the method further comprises receiving a natural-language prompt for obtaining pregnancy and / or child outcome prediction, and a set of data related to an individual, and generating the pregnancy and / or child outcome prediction by inputting the prompt and the set of data in the trained LLM.

[0143] In some embodiments, the method further comprises receiving a natural-language prompt for obtaining heart failure prediction, and a set of data related to an individual, and generating the heart failure prediction by inputting the prompt and the set of data in the trained LLM.

[0144] In some embodiments, the method further comprises receiving a natural-language prompt for obtaining myopia prediction, and a set of data related to an individual, and generating the myopia prediction by inputting the prompt and the set of data in the trained LLM.

[0145] In some embodiments, the method further comprises receiving a natural-language prompt for obtaining vision prediction in patients undergoing or to undergo anti-VEGF therapy, and a set of data related to an individual, and generating the vision prediction by inputting the prompt and the set of data in the trained LLM.

[0146] In some embodiments, the method further comprises receiving a natural-language prompt for obtaining biological age prediction, and a set of data related to an individual, and generating the biological age prediction by inputting the prompt and the set of data in the trained LLM.

[0147] In some embodiments, the method further comprises receiving a natural-language prompt for obtaining female reproductive age prediction, and a set of data related to an individual, and generating the female reproductive age prediction by inputting the prompt and the set of data in the trained LLM. In some embodiments,

[0148] In some embodiments, the method further comprises using multimodal longitudinal data of the plurality of subjects to train the LLM, for instance, as disclosed in Section IV, e.g., by using methods described in Example 14 or Example 15.

[0149] In some embodiments, provided herein is a method of generating a trained large language model (LLM) , comprising: (a) using longitudinal electronic health record (EHR) data of a plurality of subjects to train an LLM, wherein for each subject, the longitudinal EHR data are from a chronological sequence of clinical visits of the subject and comprise complete blood count (CBC) data; (b) extracting biomedical concepts from the longitudinal EHR data as medically relevant tokens; (c) processing the medically relevant tokens in chronological order to maintain the temporal coherence of subject health data, thereby predicting the next token in a patient’s timeline using temporal patterns of preceding tokens; and (d) measuring and maximizing the likelihood of predicting the next token correctly, thereby generating the trained LLM. In some embodiments, the CBC data comprise hematocrit (HCT) , hemoglobin (HGB) , mean corpuscular hemoglobin (MCH) , mean corpuscular hemoglobin concentration (MCHC) , mean platelet volume (MPV) , platelet count (PLT) , red cell count (RBC) , red cell distribution width (RDW) , white cell count (WBC) , or any combination thereof. In some embodiments, the method comprises receiving a natural-language prompt for obtaining the medical diagnosis, prediction, and / or prognosis of a disease or condition, and a set of data related to an individual, and generating the medical diagnosis, prediction, and / or prognosis by inputting the prompt and the set of data in the trained LLM. In some embodiments, the data related to the individual comprise longitudinal complete blood count (CBC) data from a chronological sequence of clinical visits of the individual, and the disease or condition is all-cause mortality or a disease such as heart attack, stroke, diabetes, kidney disease, osteoporosis, thyroid dysfunction, iron deficiency, myeloproliferative neoplasm, a cancer, or any combination thereof. II. Aging Clocks

[0150] Aging is traditionally understood as the progressive decline in physiological function that is intrinsic to all living organisms and crucial for survival and reproductive success. The healthy aging process is distinct from age-associated diseases such as cancer or cardiovascular disease, which primarily affect individuals in later stages of life but are not universal.

[0151] Most recent aspects consider aging as a genetically regulated and modifiable pathway through certain genetic and environmental interventions. Research has demonstrated that lifespan can be extended by altering specific genes or dietary factors, implicating molecular pathways in the control of senescence. Counting the years since birth, known as chronological age, is a common quantification of aging, but other hallmarks are able to more accurately quantify the functional decline that characterizes aging.

[0152] This growing understanding of aging mechanisms has driven interest in aging clocks-molecular markers that predict biological age (BA) more precisely than chronological age that measures the passage of time. Unlike chronological age, which is static, BA reflects the efficiency of biological functions using genomic, epigenetic, and clinical markers. Genomic markers are fixed at birth, while epigenetic markers, such as DNA methylation and histone modifications, changes with age.

[0153] Theoretically, individuals of the same chronological age should present similar rates of functional decline, but genetic and environmental factors can manipulate cellular aging rates, making some seemingly older or younger than their age. In some aspects, the difference is quantified by finding the difference between predicted biological age and chronological age, defined as the age gap.

[0154] Studies have shown that increased age gap suggests accelerated aging and is often linked to increased disease risk, as shown in patients with increased brain age gap who exhibit increased signs of broader systemic aging, including sensory-motor decline and an older appearance. This accelerated aging is particularly pronounced in individuals with chronic diseases, indicating that disease burdens can further drive biological aging. By creating reliable measures of biological age, aging clocks hold promise for the development of therapies to extend healthspan and improve life quality during aging, an increasingly pressing goal as human life expectancy rises worldwide.

[0155] In some embodiments, disclosed herein is a biological aging clock, EHRFormer, which leverages electronic health records (EHR) data to predict organ-specific biological age (BA) and its association with disease risk and survival outcomes. In some embodiments, developing the model comprises using unsupervised learning techniques to extract relevant features from patient data, with the goal of providing a more accurate assessment of biological aging compared to chronological age. In some embodiments, EHRFormer integrates data across multiple organ systems, including the blood, immune system, liver, kidney, and reproductive systems, and accounts for gender differences in aging patterns. In some embodiments, the model’s performance is evaluated by comparing predicted biological age to chronological age using metrics such as mean absolute error (MAE) , coefficient of determination (R2) , and Pearson correlation coefficient (PCC) , showing high accuracy for younger participants, particularly in organs with more uniform developmental trajectories. In some embodiments, by focusing on organ-specific aging, age gaps-where biological age diverges significantly from chronological age-are identified as critical biomarkers for predicting disease risk and patient stratification. In some embodiments, age gaps are critical biomarkers in cases of accelerated organ aging which is associated with a higher risk of diseases in both younger and older populations. In some embodiments, through survival analysis, specific organ age differences, particularly in the kidneys, are linked to poor survival outcomes, emphasizing the importance of integrating multi-organ age differences for more precise disease prognosis and patient management.

[0156] FIG. 14 shows an example of digital twin pretraining based on EHRs, followed by organ-specific aging analysis. The figure shows predicting organ specific biological age using individualized time serial patient EHR data with a transformer based digital twin pretraining paradigm -Stage 1: digital twin model pretraining (EHRFormer) based on time serial patient data; and Stage 2: predicting organ specific biological age and age gap and their applications in aging and disease risk stratifications. Example 2 provides an example of using EHRFormer as an advanced tool for predicting biological age across multiple organ systems. III. Female-Specific Aging

[0157] Aging is a complex, multi-scale process that manifests at molecular, cellular, organ, and whole-organism levels, leading to functional decline, age-related diseases, and ultimately mortality. In some embodiments, provided herein is an ovarian-specific aging model trained on clinical laboratory test data from over 1 million human adult subjects across multiple cohorts, including the China Health Aging Investigation (CHAI) and UK Biobank (UKB) . Through an analysis of each individual patient’s longitudinal time-series data using transformer-based digital twin technology, an ovarian organ-specific model is developed based on chronological age and longitudinal lab test data. In some embodiments, ovarian failure, including post-menopause and premature ovarian insufficiency (POI) , significantly accelerates ovarian and other organ-specific aging in women, and the ovarian organ-specific model can be used to address POI and postmenopausal syndromes, which are major female health burdens responsible for infertility and aging-associated disorders in women in modern society.

[0158] Basic and clinical studies have revealed the potential correlation between metabolic disorders and POI. In Example 3, a multi-omics analysis on a population of ~200 patients with POI and ~100 matched healthy volunteers were performed, followed by detailed mechanistic and therapeutic studies. It was identified that the insufficiency of branched-chain amino acids (BCAAs) in weight-loss diets induces a long-term detrimental effect on the ovaries via a signaling cascade involving DNA damage, inflammation, and resolution / repair. Mechanistic studies revealed that the abnormal activation of the IFN-γ-IFNGR1-ceramide axis contributes to the development of POI under conditions of BCAA insufficiency. A small molecule-based phenotypic screen for elevated E2 production in an ovarian granulosa cell line identified Ramelteon, an FDA-approved drug for insomnia, which significantly increased E2 levels and improved ovarian function in both mice and humans. The findings reveal an ovarian aging clock that is accelerated in the postmenopausal period and in POI. Ovarian aging significantly impacts aging in other organs, and BCAA insufficiency in many weight-loss diets may accelerate ovarian aging. These findings offer new insights into ovarian aging and potential new therapies aimed at mitigating age-related reproductive and other vital organ declines. IV. Multimodal Data Integration

[0159] In some embodiments, described herein is a multimodal data integration method using latent space for cross-scale biomedical applications. In some embodiments, high dimensional multimodal data, such as medical images (e.g., X-rays, CT, PET, PET-CT, MRI, OCT, FDG-PET, perfusion imaging, Radiomics, and / or fundus images) , omics data on the cellular / molecular level (e.g., Genomic Alterations: Mutations, copy number variations (CNVs) , chromosomal rearrangements, scDNA-seq; Epigenetics: DNA methylation, histone modifications, scATAC-seq; Transcriptomics: RNA-seq, gene expression profiles, scRNA-seq; Proteomics: Protein expression, post-translational modifications; Metabolomics: Targeted Metabolomics, Untargeted Metabolomics, Lipidomics, Fluxomics; Spatial Multi-omics Integration: Spatial Transcriptomics, Spatial Metabolomics, Spatial Proteomics, Spatial Epigenomics) , Pathological Data (e.g., Histopathology, Immunohistochemistry (IHC) , Cytopathology) , Laboratory Data (e.g., Blood Tests, Tumor Markers, Liquid Biopsies) , and EHRs (e.g., including clinical notes, Demographics, Medical History, Symptoms &Signs, Treatment Records) , are used to build a shared latent space comprising low dimensional data. Techniques such as variational autoencoder (VAE) and / or generative adversarial network (GAN) can be used in the building of the latent space. For instance, VAE is a type of neural network that learns to reproduce its input and map data to a latent space. In some embodiments, the latent space comprises an embedding of a set of items within a manifold in which items resembling each other are positioned closer to one another, based on that many high-dimensional data sets can lie along low-dimensional latent manifolds within that space. In some embodiments, the latent space is a geometric latent space. Using the latent space, the multimodal data can be integrated and used to impute data within a modality and / or synthesize data across modalities. In some embodiments, the multimodal data can be integrated using the latent space to predict aging, disease trajectory, and / or development. In some embodiments, the multimodal data can be integrated using the latent space for cross-scale mapping (e.g., from the molecular level to the cellular level, tissue level, and / or organismal level, or vice versa) , and / or for making mechanism connections.

[0160] In some embodiments, multi-omics data are processed using unsupervised pre-training learning and / or deep learning architectures, such as convolutional neural networks (CNNs) , generative adversarial networks (GANs) , transformers and sequence models, and / or diffusion models. In some embodiments, the processed data are further processed using multi-model diffusion, for instance, using early integration at the data level, intermediate integration at the representation level, and / or late integration at the model level.

[0161] In some embodiments, provided herein is a method of generating a trained large language model (LLM) , comprising (i) training an LLM using longitudinal multimodal data of a plurality of subjects, wherein the LLM comprises an examination encoder, a temporal embedding, and task-specific decoder heads, wherein for each subject, the longitudinal multimodal data comprise longitudinal EHR data from a chronological sequence of clinical visits of the subject, and (ii) adapting the subject-level longitudinal representations to distinct tasks using the task-specific decoder heads, wherein the distinct tasks comprise first occurrence disease diagnosis and future disease prediction, thereby generating the trained LLM.

[0162] In some embodiments, provided herein is a method of generating a trained large language model (LLM) , comprising: pretraining an LLM using longitudinal multimodal data of a plurality of subjects, wherein the LLM comprises an examination encoder, a temporal embedding, and task-specific decoder heads, wherein for each subject, the longitudinal multimodal data comprise EHR data from a chronological sequence of clinical visits of the subject. In some aspects, the examination encoder generates a contextualized representation of subject data from each clinical visit. In some aspects, the temporal embedding captures temporal relationships between clinical visits from the output of the examination encoder to generate a subject-level longitudinal representation for each subject. In some embodiments, the method further comprises finetuning the pretrained LLM by adapting the subject-level longitudinal representations to distinct tasks using the task-specific decoder heads, wherein the distinct tasks comprise first occurrence disease diagnosis and future disease prediction, thereby generating the trained LLM.

[0163] In some embodiments, the plurality of subjects comprise cancer subjects and the longitudinal multimodal data comprise routine laboratory results, vital signs, routine imaging data, or any combination thereof. In some embodiments, the routine laboratory results comprise results of any one or more of the biomarkers listed in Table 1 or Table 2. In some embodiments, the routine laboratory results comprise complete blood count (CBC) , blood chemistry, coagulation, or any combination thereof. In some embodiments, the vital signs comprise heart rate, blood pressure, body mass index (BMI) , or any combination thereof. In some embodiments, the routine imaging data comprise chest X-rays (CXRs) . In some embodiments, the longitudinal multimodal data comprise alterations to the pulmonary vasculature, shifts in bone density, changes in soft tissue composition, systemic inflammation, hormonal dysregulation, hemodynamic shifts, early cachexia, or any combination thereof.

[0164] In some embodiments, the longitudinal multimodal data comprise categorical variables and / or continuous variables. In some embodiments, the continuous variables are discretized to preserve their distributional characteristics. In some embodiments, for each subject, the longitudinal multimodal data comprise tabular clinical variables and image data. In some embodiments, for each modality of subject data, the examination encoder simultaneously captures both a value distribution and a semantic meaning of the modality of subject data.

[0165] In some embodiments, the temporal embedding comprises a linear positional embedding to learn time-dependent patterns in the longitudinal EHR data. In some embodiments, the temporal embedding is autoregressive and comprises causal masking to ensure unidirectional information flow in the autoregressive process.

[0166] In some embodiments, each of the task-specific decoder heads comprises a separate pathway applying a projection layer followed by Rectified Linear Unit (ReLU) activation. In some embodiments, each of the task-specific decoder heads comprises causal masking to prevent information from future clinical visits from influencing prediction at a given clinical visit.

[0167] In some embodiments, the LLM further comprises a missingness discriminator that determines whether a value of a particular feature is missing or not. In some embodiments, the LLM further comprises a gradient reversal layer (GRL) between the missingness discriminator and the examination encoder, wherein the GRL inverts the gradient during backpropagation, compelling the examination encoder to produce a representation that is independent of the missingness status of the particular feature. In some embodiments, the LLM further comprises a cohort discriminator that identifies the cohort label of each subject and forces the examination encoder to suppress cohort-specific information. In some embodiments, the pretraining is a self-supervised pretraining.

[0168] In some embodiments, provided herein is a method of generating a medical diagnosis, prediction, and / or prognosis for a subject, the method comprising receiving a natural-language prompt for obtaining the medical diagnosis, prediction, and / or prognosis and a set of data related to the subject, and generating the medical diagnosis, prediction, and / or prognosis by inputting the prompt and the set of data in the trained LLM disclosed herein.

[0169] In some embodiments, the medical diagnosis, prediction, and / or prognosis comprises stratifying disease risk, predicting disease incidence, classifying disease types and / or stages, assessing prognosis, or any combination thereof. In some embodiments, the medical diagnosis, prediction, and / or prognosis is for a pan-cancer diagnosis, prediction, and / or prognosis. In some embodiments, the medical diagnosis, prediction, and / or prognosis is for one or more specific cancers. In some embodiments, the medical diagnosis, prediction, and / or prognosis is for: prediction of an organ-specific age, predicting of female aging, prediction of a hospital-acquired infection, prediction of a dialysis time series, prediction of diabetes and complications, prediction of pregnancy and child outcome, prediction of heart failure, prediction of myopia, prediction of vision in anti-VEGF treatment, or any combination thereof.

[0170] In some embodiments, the set of data related to the subject comprises longitudinal EHR data of the subject. In some embodiments, at least part of the set of data related to the subject is not collected in association with generating a medical diagnosis, prediction, and / or prognosis of an oncologic indication. In some embodiments, the set of data related to the subject comprises routine laboratory results, vital signs, and routine imaging data. In some embodiments, the set of data related to the subject comprises: results of any one or more of the biomarkers listed in Table 1 or Table 2, optionally complete blood count (CBC) , blood chemistry, coagulation, or any combination thereof; heart rate, blood pressure, body mass index (BMI) , or any combination thereof; and chest X-rays (CXRs) .

[0171] In some embodiments, provided herein is a method of generating a trained large language model (LLM) , comprising (i) training an LLM using longitudinal multimodal data of a plurality of subjects, wherein for each subject, the longitudinal multimodal data comprise data from a chronological sequence of clinical visits of the subject, wherein the LLM comprises modality-specific alignment encoders, a decoder for cross-modal latent space prediction, and an autoregressive module, and wherein the pretraining comprises: (i) adversarially aligning the longitudinal multimodal data using the modality-specific alignment encoders, wherein each modality-specific alignment encoder embeds data of the corresponding modality into a unified, low-dimensional latent space, and wherein a modality discriminator identifies embedding sources in the latent space, (ii) using the decoder to predict a representation of a target modality in the latent space, wherein data from one or more clinical visits of a subject are missing in the target modality, and wherein the prediction imputes the missing data based on the subject’s historical data of the target modality and the subject’s historical data of one or more source modalities, thereby generating modality-invariant representations of the subjects in the latent space; and (ii) feeding the modality-invariant representations into the autoregressive module to model each subject's historical trajectory, thereby generating the trained LLM.

[0172] In some embodiments, provided herein is a method of generating a trained large language model (LLM) , comprising: pretraining an LLM using longitudinal multimodal data of a plurality of subjects, wherein for each subject, the longitudinal multimodal data comprise data from a chronological sequence of clinical visits of the subject, wherein the LLM comprises modality-specific alignment encoders, a decoder for cross-modal latent space prediction, and an autoregressive module, and wherein the pretraining comprises: (i) adversarially aligning the longitudinal multimodal data using the modality-specific alignment encoders, wherein each modality-specific alignment encoder embeds data of the corresponding modality into a unified, low-dimensional latent space, and wherein a modality discriminator identifies embedding sources in the latent space, (ii) using the decoder to predict a representation of a target modality in the latent space, wherein data from one or more clinical visits of a subject are missing in the target modality, and wherein the prediction imputes the missing data based on the subject’s historical data of the target modality and the subject’s historical data of one or more source modalities, thereby generating modality-invariant representations of the subjects in the latent space. In some embodiments, the method further comprises finetuning the pretrained LLM by feeding the modality-invariant representations into the autoregressive module to model each subject's historical trajectory, thereby generating the trained LLM.

[0173] In some embodiments, the multimodal data comprise laboratory tests, imaging scans, and / or omics profiles. In some embodiments, the method further comprises feeding the output of the autoregressive module into a task-specific prediction head for a target application.

[0174] In some embodiments, the target application is prediction of a systemic disease, prediction of a chronic disease, prediction of an ocular disease such as myopia, and / or identification of aging trajectories of the subjects. In some embodiments, the target application is prediction of a disease selected from the group consisting of atrial fibrillation, coronary artery disease, diabetes, hypertension, ischemic stroke, multiple sclerosis, osteoporosis, Parkinson’s disease, and rheumatoid arthritis.

[0175] In some embodiments, the longitudinal multimodal data comprise fundus images, routine laboratory test results. In some embodiments, the longitudinal multimodal data comprise electronic health record (EHR) data, vital signs, imaging data, or any combination thereof. In some embodiments, the routine laboratory results comprise results of any one or more of the biomarkers listed in Table 1 or Table 2. In some embodiments, the routine laboratory results comprise complete blood count (CBC) , blood chemistry, coagulation, or any combination thereof. In some embodiments, the vital signs comprise heart rate, blood pressure, body mass index (BMI) , or any combination thereof. In some embodiments, the imaging data comprise chest X-rays (CXRs) .

[0176] In some embodiments, provided herein is a method of generating a medical diagnosis, prediction, and / or prognosis for a subject, the method comprising receiving a natural-language prompt for obtaining the medical diagnosis, prediction, and / or prognosis and a set of data related to the subject, and generating the medical diagnosis, prediction, and / or prognosis by inputting the prompt and the set of data in the trained LLM disclosed herein. In some embodiments, the medical diagnosis, prediction, and / or prognosis comprises stratifying disease risk, predicting disease incidence, classifying disease types and / or stages, assessing prognosis, or any combination thereof. In some embodiments, the disease comprises one or more cancers.

[0177] In some embodiments, provided herein is a system comprising: at least one hardware processor; and one or more software modules configured to, when executed by the at least one hardware processor, perform the method disclosed herein. In some embodiments, provided herein is a non-transitory computer-readable medium having instructions stored thereon, wherein the instructions, when executed by a processor, cause the processor to perform the method disclosed herein. In some embodiments, provided herein is a system comprising: at least one hardware processor; non-transitory computer-readable medium coupled to at least one hardware processor, optionally wherein the coupling is over a network; and instructions stored in the non-transitory computer-readable medium, wherein the instructions when implemented by the processor, configure the system to perform the method disclosed herein. EXAMPLES

[0178] The following Examples are included for illustrative purposes only and are not intended to limit the scope of the disclosure. Example 1: Early Cancer Detection Using Longitudinal Patient EHRs

[0179] FIG. 6 shows a pipeline for cancer detection / prediction using digital twin pretraining based on EHRs. All patient data were first cleaned and aligned. Based on the feature distribution in different hospitals, 480 shared features were identified, including demographic features, laboratory test features, surgical and procedural features, and medication features. Then, the records of all hospitals were merged and aligned with these 480 features. A threshold for NA values at 0.8 was set to filter out the sparse features. In the end, 341 features remained. In addition to feature selection, the follow-up data were also filtered. All records with missing age or gender information were removed, as well as any anomalies in demographic features due to manual entry errors. Furthermore, data from patients with chronic diseases who had an unusually high number of medical visits and follow-up records with outliers in the test data were excluded. All discrete variables were mapped to integers. For all surgical, procedural, and medication data, the counts for each patient were accumulated. Regarding missing value handling, all demographic and laboratory test data were filled with the mean value of the respective feature, while for all surgical, procedural, and medication data, 0 was used for imputation. Lasso Regression was used to further narrow down the selection from the remaining 341 features. Lasso Regression employs an L1 penalty term to promote sparsity among the coefficients. This approach effectively shrinks some coefficients to zero, thereby eliminating the corresponding features from the model. After applying Lasso Regression, a subset comprising 165 features was successfully identified from the original 341 that are most pertinent and predictive for the outcome variable under consideration. For time-series dataset construction, multiple visit data of patients according to the visit time were sorted. A patient may have one or more visit records. These data were obtained using a sliding window with a length of 5, ensuring that the 5 visits are consecutive. These five visit records were then concatenated in chronological order, making sure that each time series data includes features from the "-1" visit. For patients with less than 5 visits, their time series sequences were padded with zeros. In one example, after processing, the complete time series dataset contains 1, 929, 608 patients with 2, 241, 506 time series records, and time series dataset was constructed based on these data.

[0180] In one aspect, time series datasets were constructed. The cohort contained 2, 495, 119 patients with 6, 187, 026 visits. Each patient HER included vital signs, laboratory test, drug and treatment prescriptions, and operation notes. Features were extracted to build the model. The training dataset contained over 2 million patients with 5 or more visits, containing 2.3 million time series records. Each time series record had 615 features. FIG. 7 shows correlation among features, including Pearson correlation coefficient between features with less than 70%NA rate. Red areas indicate higher correlation coefficients between features, while blue areas indicate lower correlation coefficients between features. FIG. 8 shows tumor prediction in various tumors.

[0181] Cancer occurrence prediction was performed with each time series record using a CatBoost model. One year, 3 year, 5 year and 10 year were used as the window to predict cancer incidence. FIG. 9 shows performance of cancer prediction models for different time periods. As the prediction period extends, the model’s Recall and F1 scores show significant improvement. FIG. 10 shows the impact of time series length on model performance. Input sequences of varying temporal lengths were constructed from the test set. Results demonstrate that increasing the length of time series data leads to improved classification performance by the model. The results showed the model performance improves with an increased time window in recall, f1, auprc, with 10 year recall of 0.5129. This model is significantly better than non-time series models. The time-series analysis demonstrated superior performance in early cancer screening. Different length time series datasets were constructed. As shown in FIG. 11, as patient visits increased from one to five in time series records, the model performance increased accordingly, suggesting that longer time series improve prediction performance.

[0182] For interpreting the effects and relative contributions of the features and clinical parameters on cancer prediction, an explainer SHAP (Shapley Additive exPlanation) was implemented and results are shown in FIG. 12. SHAP values for features in the trained time-series cancer prediction model. The cancer biomarker CEA shows the highest contribution to the model's predictions. As expected, the age features were among the most significant contributors in the cancer occurrence estimation. In addition, clinical parameters, prescriptions, and general clinical characteristics also contributed to the prediction of cancer. Interestingly, additional prognostic markers were also identified, including liver biochemistry markers, electrolyte and acid-base balance, and markers of inflammation, suggesting the overall health and systemic homeostasis played an important role in determining the clinical prognosis of cancer in these patients in terms of their evolving into a cancer status. Furthermore, the features of the visit right before cancer occurrence also ranked high on the list. FIG. 13 shows the predictive performance of time series prediction models across different cancer categories, with the model demonstrating the best performance in lymphohematopoietic cancers.

[0183] Biomedical concept tokens were extracted from EHR using an integration of MedCATintegrates Named Entity Recognition and Linking [NER+L] , MedCAT, which was trained and tested on a dataset comprising 17 282 manual annotations, and the train–test split used was 80%for training and 20%for testing. The patient-level MedCAT achieved a precision of 0·9549 (95%CI 0·9519–0·9579) , recall of 0·8077 (0·8017–0·8137) , and F1 score of 0·8752 (0·8702–0·8802) , whereas the models without precision bias achieved a precision of 0·9314 (0·9274–0·9354) , recall of 0·8959 (0·8909–0·9009) , and F1 score of 0·9133 (0·9093–0·9173) . The model’s performance in identifying the context of these tokens were then evaluated in terms of relationship to a patient or person (experiencer contextualization) and whether the token is affirmed or negated (negation contextualization) , and F1 scores were 0·9280 (0·9240–0·8320) for experiencer and 0·9490 (0·9460–0·952) for negation.

[0184] The model was validated using three datasets, KCH, SLaM, and MIMIC-III, by comparing the prediction for the next biomedical concept with the ground truth. To measure the model’s generalizability to unseen data, train and validation loss was measured to be 3·01 and 3·14 for KCH, 2·95 and 3·23 for SLaM, and 3·77 and 3·93 for MIMIC-III. The model's average precision and recall for identifying all positive instances in KCH improved from 0.55 (SD 0.0018) and 0.47 (0.0013) to 0.64 (0.0018) and 0.54 (0.0015) , respectively, when extending the time. Increasing the number of top N forecasts was considered from @1 to @10, an average precision of 0.84 (0.0015) and recall of 0.76 (0.0005) across new and recurring concepts were achieved. For predicting new disorders, the model achieved precision@10 of 0.68 (SD 0.0027) for KCH, 0.76 (0.0032) for SLaM, and 0.88 (0.0018) for MIMIC-III. For forecasting the next new biomedical concepts precision@10 was 0.80 (0.0013) for KCH, 0.81 (0.0026) for SLaM, and 0.91 (0.0011) for MIMIC-III. Recurring concepts were forecasted more accurately than new ones. Standard deviation using bootstrapping was 0.00631 for KCH, 0.0154 for SLaM, and 0.0066 for MIMIC-III.

[0185] Through experimentation with network size, adding more layers or increasing the product of attention heads and layers up to 32 × 32 did not improve performance but rather led to significant overfitting. A smaller bucket size, which is the size of the smallest timespan discernable, of 1 day outperformed models trained on larger spans, such as 3, 7, 14, 30, and 365 days. Comparative performance testing against an LSTM-based approach the task of next-concept prediction (for the category of all concept types) demonstrated it to be 40%inferior in performance to the GPT-based model on the KCH dataset.

[0186] In a qualitative analysis involving five clinicians, 34 synthetic timelines resembling clinical scenarios were generated and evaluated using the trained model. Emphasizing relevance over strict accuracy, clinicians assessed the significance of five forecasted disorder concepts. The findings showed that the proportions of relevant forecasted concepts were high: 97% (95%CI 91–100) for one concept, 96% (89–100) for two concepts, 90% (80–100) for three concepts, 89%(87–100) for four concepts, and 88% (88–99) for five concepts.

[0187] Inter-annotator agreement among clinicians stood at 86%, with scenarios of high agreement demonstrating 93% (95%CI 88–98) relevance for the top five concepts. Illustrating clinician reasoning, a specific case exemplified how diagnoses were contextually linked to preceding symptoms and procedures. Contextual errors included forecasting high-probability, low-impact events like systemic arterial hypertensive disorder, which clinicians considered irrelevant. Despite these errors, the majority of concepts forecasted by trained model were clinically relevant, showcasing the model's effective contextual attentional transformer mechanism.

[0188] Furthermore, trained model demonstrated its capability to generate detailed patient timelines with minimal input, as evidenced by a case study involving a 43-year-old Black female. Employing top-k sampling (k=100) , the model generated timelines encompassing a total of 21 concepts, comprising 6 base concepts and introducing 15 new concepts.

[0189] The model in this example is a deep-learning model designed to generate patient timelines across mental and physical health domains using EHR data, with a specific focus on early cancer detection. The model integrates natural language processing to extract biomedical concepts such as disorders, procedures, substances, and findings from both structured and unstructured clinical text. It is scalable and adaptable across diverse patient populations and medical conditions, benefiting from increased data volume to enhance predictive accuracy over time. This generative model not only predicts the next steps in patient trajectories but also simulates long-term timelines, facilitating simulations ranging from brief inpatient episodes to extensive chronic condition management scenarios. This capability supports digital health twins, enabling researchers to explore hypothetical health scenarios and conduct counterfactual modeling for causal inference.

[0190] In the context of early cancer detection, the model's ability to analyze comprehensive patient data allows for the identification of subtle patterns and risk factors indicative of early-stage cancers. By integrating and analyzing vast amounts of longitudinal data, the model can detect anomalies and trends that may signify the onset of cancer before it manifests clinically. This early identification capability aids in timely diagnosis and intervention, which are crucial for improving patient outcomes. By simulating long-term health trajectories, the model can predict the likelihood of cancer development, thus facilitating proactive healthcare strategies. Additionally, in medical education, the model serves as a tool for simulating complex case studies, enabling interactive learning, clinical reasoning practice, and discussions on ethical considerations in healthcare within a safe environment. This application is particularly valuable in training healthcare professionals to recognize early signs of cancer and understand the importance of early intervention.

[0191] In some examples, the model disclosed herein uses historical data. In some examples, the model disclosed herein aligns with current clinical guidelines and has the ability to incorporate emerging treatments. In some examples, the model disclosed herein uses likelihood of a concept occurring rather than its urgency and impact. In some examples, the model disclosed herein prioritizes high-impact, high-urgency events. In some examples, by focusing on the likelihood of a concept occurring rather than its urgency and impact, the model can lead to forecasts of common but contextually irrelevant concepts, such as predicting cataracts in an elderly patient presenting with chest pain. To address this, one or more relevancy filters are introduced, for example through prompt engineering to limit predictions to certain disease types or organs, or by providing a separate relevancy signal. Additionally, the dataset can be expanded, for instance, by including rare diseases in the training set.

[0192] In some examples, the model disclosed herein addresses another issue is the potential for "hallucinations" in transformer-based generative models, where the model generates plausible but incorrect outputs. This is particularly concerning for long-term simulations, emphasizing the need for robust relevancy and mitigation systems before the model can be used for clinical decision support. Ensuring accurate predictions in the context of early cancer detection is critical, as false positives or negatives could have significant consequences for patient care.

[0193] The modular architecture of the model allows for continuous improvement and extension. Enhancements could include fine-tuning natural language processing capabilities, incorporating quantitative data, expanding the dataset to better cover rare diseases, improving temporal quantification, adding primary care data, and integrating external knowledge from clinical guidelines and academic literature. The model represents a novel deep-learning generative model utilizing EHRs, combining natural language processing with longitudinal forecasting. Its broad applicability spans digital health twins, synthetic dataset generation, real-world risk forecasting, longitudinal research, virtual trial emulation, and medical education. As all components are improvable, ongoing refinements can enhance its utility across various healthcare domains, particularly in the early detection of cancer, where timely and accurate predictions are crucial for improving patient outcomes.

[0194] The model pipeline comprises four primary components. First, CogStack is employed for data retrieval and initial data preprocessing. Next, the Medical Concept Annotation Toolkit (MedCAT) is utilized to convert free-text information from EHRs into structured data. The core of the system, MetaGP Core, is a deep learning model designed for biomedical concept modeling. Finally, the MetaGP web application facilitates user interaction with the trained model.

[0195] To train and evaluate the model, three distinct datasets were utilized. For named entity recognition and linking, biomedical concepts were extracted from free text using MedCAT and linked to the SNOMED Clinical Terms database. Extracted concepts included diseases, symptoms, medications, and findings. To maintain privacy and ensure adequate data for analysis, concepts occurring fewer than 100 times were excluded. The remaining data were organized into patient timelines and randomly divided into a training set (95%) and a test set (5%) . The model for biomedical concept forecasting was built on the GPT-2 architecture and is a transformer-based pipeline that processes medically relevant tokens, which are processable unit of biomedical concepts, extracted from clinical narrative. Clinical narratives and tokens are processed in chronological order to maintain the temporal coherence of patient health data, enabling the model to predict the next biomedical concept in a patient’s timeline using temporal patterns. In the model, a group of patients U= {u1, u2, u3, …} , are analyzed where each patient ui is a sequence of tokens in chronological order, ui= {w1, w2, w3, …} . The logarithm function is used to measure the likelihood of predicting the next token correctly, which is to be maximized. Optimal hyperparameters were determined using population-based training on a validation set, resulting in a configuration with 16 layers, 16 attention heads, an embedding dimension of 512, a weight decay rate of 0.01, a learning rate of 0.000314, a batch size of 32, and a warm-up ratio of 0.01. Training on the largest dataset (KCH) was completed in 1–2 days using eight V100 GPUs. A web application was developed to facilitate user interaction with the trained model. This application allows users to evaluate the model by creating or loading patient timelines and includes a gradient-based saliency method to visualize the importance of each concept in forecasting.

[0196] To evaluate the performance of the model, precision and recall metrics were adapted to the nature of EHR data. The model forecasts the next concept in a patient's timeline, and precision and recall are evaluated by comparing the prediction to actual records. Forecasts are considered correct if prediction presents within specified time ranges (30 days, 1 year, and ndefinitely) . Precision@N and recall@N metrics assessed the likelihood that one of the top N forecasts was correct, with N set to 1, 5, or 10. The forecasts had to match the type of the ground truth concept, and a distinction was made between new and recurring concepts in a patient’s timeline.

[0197] To avoid class imbalance, the model is not only forecasting the overrepresented groups of concepts by matching the ground truth (EHR) indications to model’s predictions at that position in the timeline. This encourages it to learn and predict diverse types of medical conditions and events that are relevant to the patient's health journey.

[0198] For each concept, whether each predicted concept is new or recurring in the patient’s timeline were also monitored. New concepts make first time appearances, while recurring concept have prior documentation in patient's medical history. The model's predictions are filtered and evaluated based on whether they align with these distinctions as compared to the ground truth to enable evaluation of model's ability to identify new developments or ongoing conditions in the patient.

[0199] For forecasting the next new disorder in a patient timeline, the model achieved a precision@10 of 0.68 (SD 0.0027) for the KCH dataset, 0.76 (SD 0.0032) for the SLaM dataset, and 0.88 (SD 0.0018) for the MIMIC-III dataset. When predicting the next new biomedical concept, the model achieved a precision@10 of 0.80 (SD 0.0013) for the KCH dataset, 0.81 (SD 0.0026) for the SLaM dataset, and 0.91 (SD 0.0011) for the MIMIC-III dataset. Additionally, the model's performance was validated on 34 synthetic patient timelines by five clinicians, showing a relevancy of 33 (97% [95%CI 91–100] ) out of 34 for the top forecasted candidate disorder. As a generative model, the model can continue to predict subsequent biomedical concepts for as many steps as needed. Example 2: Organ-specific Age Prediction

[0200] Using unsupervised learning to extract features from patient data, patients were mapped onto a low-dimensionality t-SNE graph, revealing that the model was able to differentiate patients into distinct clusters by gender and age, confirming the are meaningfulness and relevance of these extracted features (FIG. 15) . Specifically, after unsupervised learning was used to extract key features, patients were mapped onto a t-SNE plot for dimensionality reduction and the plot reveals distinct groupings based on gender (female = green, male = blue) and age (darker = older) , indicating that the extracted features successfully capture meaningful patterns in the data.

[0201] In the WYFE cohort, the model demonstrated overall decent accuracy (low MAE) and explainability of variability (high R2 / PCC) , indicating that predictions based on lab test data can accurately predict age (FIG. 16) . However, the model performed better in younger participants (<18 years old) compared to older ones (>18 years old) , likely due to the larger gap between predicted biological age and chronological age in older individuals. FIG. 16 shows scatter plots comparing AI-predicted age and actual chronological age for different age groups (0–18 years and >18 years) and gender. Each plot shows the mean absolute error (MAE) , coefficient of determination (R2) , and Pearson correlation coefficient (PCC) as measures of prediction accuracy. For younger individuals (0–18 years) , the AI model demonstrates high accuracy (lower MAE and higher R2 / PCC) for both females (top-left) and males. However, for adults (>18 years) , the prediction accuracy decreases, as seen in higher MAE and lower R2 / PCC for both genders.

[0202] Visual explanations of relative feature contributions are created using SHapley Additive exPlanations (SHAP) , highlighting low alkaline phosphatase (ALP) levels, high urea nitrogen, and high mean corpuscular volume (MCV) to contribute to higher predicted biological age. This aligns with existing literature, where low ALP levels and high urea nitrogen have been associated with aging and age-related diseases, and high MCV values have been linked to age-related changes in red blood cell morphology. Thus, these features appear to be consistent markers of biological aging. FIG. 17 shows a SHAP summary plot for feature contribution to model predictions. Features are ranked in descending order of importance based on the color gradient, from blue to red, indicates feature values, with blue representing low values and red representing high values.

[0203] The organ-specific age prediction models demonstrated varying levels of performance across different organ systems, age groups, and genders, reflecting distinct age-related patterns in biological systems. Performance was evaluated using mean absolute error (MAE) , coefficient of determination (R2) , and Pearson correlation coefficient (PCC) , providing insights into prediction accuracy and organ-specific variability.

[0204] Among younger participants, (<18 years old) , the model demonstrated high prediction accuracy for ages of the blood, immune system, liver, and kidney among both genders (FIG. 18) . Among males, the model achieved the highest performance with an MAE of 1.79 years, R2 of 0.73, and PCC of 0.86, with comparable results for the liver. In young females (<18 years old) , the liver demonstrated the most robust performance, with an MAE of 1.79 years, R2 of 0.71, and PCC of 0.85. These results suggest that age-related biological changes in these systems are more predictable, potentially due to their more uniform developmental trajectories during childhood and adolescence. In contrast, predictions for the reproductive system demonstrated greater variability, potentially reflecting increased individual heterogeneity in maturation timing. These findings highlight the utility of organ-specific modeling approaches in capturing age-related biological changes and identifying age gaps for understanding organ-specific aging trajectories.

[0205] In older participants (>18 years old) , MAE and R2 of predictions were significantly higher compared to that of younger participants (FIG. 19) . In the older male cohort, while the model still performed the best for age prediction of the blood system, MAE has increased to 10.58 years and R2 has reduced to 0.34. While the model performed the worst in age prediction of the immune and reproductive system for both genders, the R2 was much smaller in predictions among older males. This could be due to sex-specific differences in aging. Unlike women, whose reproductive aging is marked by the relatively defined onset of menopause, male reproductive aging is more variable, with a gradual decline in testosterone8. Additionally, sex hormones like testosterone significantly impact immune function9, thus the more variable decline in testosterone levels in males may contribute to the reduced predictability of both immune system’s function and age.

[0206] Focusing on individuals with accelerated organ aging, defined as a deviation of more than 3 standard deviations from the average predicted organ age of that chronological age group, it was found that organ-specific aging significantly increases the risk of developing specific diseases, with age-specific differences. The observed difference between among different age cohorts suggest different mechanisms between the association between accelerated organ aging and disease. In younger individuals (<18 years old) , accelerated aging in the blood system is associated with a higher risk of diseases such as diabetes mellitus, hypertension, and congenital abnormalities, suggesting that accelerated organ aging predisposes them to certain conditions (data not shown) . In contrast, in older individuals, the presence of diseases like protozoal infections, developmental delays, and HIV may accelerate biological aging, creating a feedback loop where the disease exacerbates organ dysfunction, leading to accelerated organ aging. Thus, while in younger individuals, accelerated aging appears to be a causal factor for disease, in older individuals, pre-existing diseases seem to accelerate aging. This highlights the age-dependent interplay between organ dysfunction and disease progression.

[0207] Using survival analysis based on the age gap (predicted biological age minus chronological age) across various organ systems, the risk of disease outcomes were assessed and high-and low-risk populations were identified (data not shown) . Notably, a greater age differences in kidney correlated with a over 50%increased risk of poor survival outcomes in renal failure patients, underscoring the critical role of kidney aging in disease progression. In male genitourinary conditions, a larger gap in prostate age was linked to a nearly 50%increased risk of adverse outcomes. Additionally, for arterial diseases, increased age differences in the lung, prostate, skin, and cardiovascular systems were all associated with over 40%increased risk of poor survival outcomes, indicating that accelerated aging in these organs contributes substantially to the prognosis of arterial diseases. These findings emphasize the potential of organ-specific age gaps as biomarkers for identifying individuals at higher risk of poor health outcomes, particularly in those with chronic diseases. However, only the age difference of the kidneys showed a significantly increased risk of poor survival outcomes (HR > 1.5) in renal failure patients, while age differences in other organs were only moderately associated with risk for any disease. Integrating age differences across multiple organs could provide a more accurate prognosis and improve patient stratification, as most diseases affect multiple organs.

[0208] When this age difference is reduced to 2 standard deviations from the average predicted organ age of that chronological age group, it was found accelerated aging of specific organs to be associated with the presence of specific disease (FIG. 20) . In all participants, respiratory, musculoskeletal, and metabolic disorders exhibit stronger associations with accelerated aging in organs such as the immune system, lungs, and blood.

[0209] By examining the hazard ratio of older participants (>18 years old) that exhibited accelerated organ aging, it was found that accelerated aging of almost any organ increased the risk of developing circulation, blood, respiratory, urinary, digestive, endocrine, and nervous system conditions. Notably, accelerated aging of the gastrointestinal system seems to have minimal impact on the risk of developing any conditions, except an increased risk for genitourinary conditions in males. Data for the hazard ratio between ageotypes (Age Diff > 2SD) and diseases is not shown.

[0210] Applying this model to the CHAI dataset, it was found the model to generalize well to other datasets, showing strong correlation (Pearson’s r = 0.8568) and explanation of variance (R2 = 0.7341) . Patients were stratified depending on their deviation from the average predicted age of that chronological age, where patients with predicted age greater than the average considered accelerated aging, and those with lower predicted age considered decelerated aging (FIG. 21A) . When the predicted organ ages of these two groups were investigated, it was found the accelerated aging group to also exhibited significantly older organ age compared to the decelerated aging group (FIG. 21B) . Further analysis comparing the organ-specific aging between the two groups found that the accelerated aging group exhibited older aging of almost all organs, with the largest difference presented in multi-organ, conventional organ, liver, and adipose age (FIG. 22) .

[0211] This example shows the potential of EHRFormer as an advanced tool for predicting biological age (BA) across multiple organ systems, offering insights into organ-specific aging processes and their association with disease outcomes. By leveraging a large cohort of electronic health records (EHR) data, the results underscore the importance of tissue-and organ-specific aging, particularly in how deviations from chronological age-reflected in biological age gaps-are linked to disease susceptibility and survival prognosis. A key factor in this model is the recognition that different organs and tissues undergo aging processes that can be uniquely characterized by specific biomarkers, including DNA methylation patterns, offering an opportunity to enhance the understanding of aging beyond a one-size-fits-all chronological measure.

[0212] Tissue-specific methylation has emerged as a critical component in the development of aging clocks, as it can reflect both the intrinsic and extrinsic factors that contribute to the aging process. Previous studies have identified distinct epigenetic differences across different tissue types that is correlated to the tissue function. For instance, increased promoter methylation variability in sperm correlates with lower live birth rates following intrauterine insemination (IUI) , revealing a direct link between altered epigenetic variability in diseased tissues and their healthy counterparts. In this example, the biological age predictions based on tissue-specific biomarkers, such as those derived from the blood, liver, kidney, and reproductive systems, revealed distinct patterns of aging that are reflective of each organ’s unique response to aging and environmental influences.

[0213] The results demonstrate that the model’s ability to differentiate biological age across organs is not only dependent on age-related changes in cellular function but also influenced by sex-specific factors. As demonstrated by the significant differences in organ-specific performance between genders, particularly in the immune and reproductive systems, sex hormones and their decline over time likely play a major role in shaping organ-specific aging patterns. The incorporation of a more comprehensive panel of inflammatory markers into the current example could significantly enhance the understanding of organ-specific aging and its associated risks. Chronic inflammation has been identified as a central driver of aging and age-related diseases, influencing various conditions such as dementia, type 2 diabetes mellitus (T2DM) , osteoarthritis, and chronic obstructive pulmonary disease (COPD) . As the current model has not demonstrated outstanding performance in prediction of immune systems, addition of other known inflammatory markers involved in disease, such as interleukin-6 (IL-6) , C-reactive protein (CRP) , or PGC1α, can enhance prediction of immune age as well as disease prognosis.

[0214] The findings from this example suggest that organ-specific aging clocks, such as EHRFormer, could be used to predict disease risk with greater accuracy than chronological age. In clinical practice, this could lead to better patient stratification, particularly in chronic diseases that affect multiple organs. For example, in patients with renal failure, the greater the age gap between predicted biological age and chronological age, the higher the risk of poor survival outcomes. This demonstrates that organ-specific age gaps could serve as biomarkers for more accurate prognosis and personalized treatment strategies.

[0215] In conclusion, the integration of tissue-specific methylation patterns into biological aging models has the potential to revolutionize the understanding of aging and its relationship to disease. These insights can guide the development of more accurate biomarkers for aging, enable earlier disease detection, and pave the way for personalized medicine approaches that account for the unique aging trajectories of different organ systems. As the field continues to evolve, the ability to refine aging clocks based on organ-specific methylation and other molecular markers will be crucial for improving healthspan and longevity outcomes across diverse populations. Example 3: Female Aging Clock

[0216] Aging is a complex, multifaceted process involving changes at the molecular, cellular, and organ levels, ultimately impacting whole-organism health and survival. Deciphering how these changes contribute to increased disease susceptibility and mortality risk is essential for advancing interventions that extend healthspan. Biological age, a measure of accumulated biological damage relative to an average individual of the same chronological age, has emerged as a key metric for assessing age-related disease risk. Unlike chronological age, biological age can diverge, providing a valuable metric for age-related disease risk.

[0217] Initially, biological age relied on DNA methylation patterns, but recent advancements have expanded aging clocks to include additional "omics" layers to provide more accurate predictions of biological age. These innovations highlight the variability in aging across organs and their differential response variably to external influences, such as lifestyle or medications, offering new avenues for tailored anti-aging strategies.

[0218] Recent studies have demonstrated organ-specific aging patterns, with proteomic profiling revealing accelerated aging in certain tissues. For example, Oh et al. showcased that plasma proteome profiles across populations could reveal organ-specific aging markers. The UK Biobank, with its extensive phenotypic data, has become a valuable resource for exploring proteome-based aging signatures and evaluating lifestyle and therapeutic interventions. Using high-throughput technologies such as the Olink Explore 3072 platform, protein abundance in plasma from over 53,000 participants was assessed. These findings underscore the link between organ-specific aging and age-associated diseases, emphasizing the potential of lifestyle modifications and therapeutic interventions to mitigate aging and extend healthspan.

[0219] To further advance organ-specific aging research, digital twin technology was integrated into the approach. Digital twins, which are virtual representations of individual patients, were generated using longitudinal electronic health record (EHR) data through a transformer-based model, EHRFormer. Digital twin pretraining was use to model and analyze aging processes with high granularity, linking ovarian aging to multi-organ aging and improving the understanding of the interplay between biological aging and disease risk.

[0220] The ovary, as a critical reproductive and endocrine organ, plays a vital role in maintaining female physiological homeostasis. Ovarian aging not only impairs fertility but also predisposes individuals to various age-related conditions, including cardiovascular diseases, osteoporosis, and mental disorders. While the significance of ovarian aging is well-established, it is crucial to understand its unique aspects in humans, particularly given the prolonged life after reproduction (PRLS) seen in humans compared to most vertebrates. This extended lifespan after ovarian dysfunction in humans and certain animals, including chimpanzee, cetaceans and insects, highlights the distinct nature of ovarian aging and the need to study it in human physiological and pathological contexts.

[0221] Ovarian aging can occur as part of the natural aging process or be accelerated by conditions such as premature ovarian insufficiency (POI) , a disorder characterized by ovarian dysfunction before the age of 40. POI presents with elevated serum follicle-stimulating hormone (FSH) levels (FSH >25 U / L) and provides a more defined model for studying ovarian aging compared to the natural aging process in healthy women. Although many studies have aimed to identify genetic factors associated with ovarian aging, only a small fraction of the pathogenesis can be explained by genetic factors alone. Increasing evidence suggests that metabolic dysfunction plays a critical role in ovarian aging, with dietary factors influencing ovarian health.

[0222] Branched-chain amino acids (BCAAs) , including leucine, isoleucine, and valine, are essential amino acids with pivotal roles in metabolic homeostasis. Recent research has linked BCAA insufficiency to POI, suggesting a metabolic underpinning for ovarian aging. In this example, an ovarian-specific aging clock was developed and transient BCAA insufficiency was identified as a key pathogenic factor in POI. It was found that BCAA deficiency induced ovarian dysfunction through a lasting metabolic memory involving the upregulation of IFN-γ. This immune mediator elevated ceramide levels in granulosa cells via enhanced phosphorylation of the ceramide transporter CERTL. Phenotypic screening revealed that targeting ceramide accumulation with the approved drug Ramelteon could protect the ovaries from aging and restore ovarian function. These findings highlight a metabolic memory-driven mechanism in POI and offer new insights into potential targeted therapies for ovarian aging.

[0223] Ovarian-specific Aging Clock

[0224] To investigate the dynamics of ovarian aging, an aging clock was developed that leverages electronic health records (EHR) , including longitudinal laboratory data from a cohort of 2 million healthy male and female individuals sourced from CHAI. By employing a transformer-based digital twin approach, the EHRFormer model was developed to predict biological age based on EHR (FIG. 14) . Using unsupervised learning, the model identified patterns and extracted age-and sex-specific features that could potentially serve as biomarkers for female biological aging. The EHRFormer was then applied to individual patients' longitudinal EHR data to generate personalized digital representations, known as digital twins. These digital twins form the foundation for this aging clock, as they simulate each complex detail of a patient’s health data and provide personalized predictions of health states. The model was trained separately for each gender and age group (age >18 years) (amore detailed description of the model construction procedure follows) . The predicted age generated by EHRFormer demonstrated a strong correlation with chronological age for each gender and age group (FIGS. 23A-23D) , providing a robust framework for predicting biological age based on routine clinical laboratory data.

[0225] Given that males and females exhibit distinct aging patterns, a biomarker-based prediction model was constructed for each gender, using lab test biomarker data from one million healthy individuals collected from CHAI. The analysis revealed that the predicted biological age demonstrated a strong correlation with chronological age for both genders (FIG. 23A, FIG. 23B) , but diverged significantly from chronological age in older women (FIG 23A) . Notably, from ages 40 to 50, the rate of biological aging in females increased sharply (FIG. 23A) , a phenomenon not observed in males of the same age group (FIG. 23B) . This finding was further validated by applying a biomarker-based prediction model to data from the UK Biobank (UKB) . Consistent with earlier observations, the predicted age strongly correlated with chronological age for both genders, with an accelerated aging trend exclusively observed in women aged 40-55 (FIG. 23C, FIG. 23D) .

[0226] The sharp acceleration of biological age observed in women around midlife coincides with the onset of menopause, suggesting that ovarian aging may play a pivotal role in driving the accelerated aging process. To investigate this, a reproductive system-specific aging model was developed based on biomarkers closely tied to menopause. Indeed, the aging model exhibited accelerated aging in females between ages 40-50 in both the CHAI (FIG. 24A, FIG. 24B, FIG. 24C) and UK Biobank cohort FIG. 25A, FIG. 25B) . In this model, participants were stratified into two groups based on their predicted reproductive and endocrine age (FIG 24A) . Patients were categorized into two groups based on their predicted age relative to the average predicted age for their chronological age group. Those with a predicted age higher than the average for their chronological age were placed in the accelerated aging group, indicating faster biological aging. Conversely, those with a predicted age lower than the average were placed in the decelerated aging group, reflecting slower biological aging. By comparing these two groups, it was observed that the accelerated aging group consistently exhibited a higher predicted age across other organ systems, indicating a strong correlation between reproductive system aging and overall aging (FIG. 24B, FIG. 24C) . Further analysis of organ-specific aging revealed similar age acceleration across multiple organ types corresponding to the accelerated reproductive age, despite predictions by different models. The results show a correlation between reproductive system aging and multi-organ aging. Using female-specific reproductive aging clock to divide 40-70 years old Chinese female populations (CHAI) into two categories (top 50%vs bottom 50%) , it was found that there is significant widespread organ specific aging. Further analysis reveals a strong correlation between accelerated reproductive system aging and aging across other organ systems. Females with higher predicted reproductive age also exhibit significantly higher predicted age in organs such as the heart, liver, and kidneys, indicating that ovarian aging may have systemic implications. These findings raise questions regarding the direction of causality between ovarian aging and the aging of other organs, suggesting potential interdependencies between these processes. This raises the question of causality: whether accelerated aging in the reproductive system drives aging in other organs, or vice versa.

[0227] Ovarian Failure Is Associated with Many Other Organ Failures

[0228] These findings were subsequently validated in an independent UK cohort, ensuring the robustness and generalizability of the model (FIG. 25A) . Participants were categorized into accelerated aging and decelerated ovarian aging groups, following the same criteria as in the CHAI dataset. The accelerated aging group exhibited higher predicted organ ages above the mean. In both the CHAI and UK cohorts, the accelerated aging group showed higher predicted ages for all organs compared to the decelerated aging group (FIG. 25B) . Comparative analysis of predicted organ ages in both cohorts consistently demonstrated that females in the accelerated ovarian aging group had significantly older predicted organ ages, while those in the decelerated cohort exhibited significantly younger predicted ages (FIGS. 26A-26K) . Additionally, the least overlap between the two groups was observed in the predicted age of multiple organs, followed by the artery, intestine, and pituitary age, highlighting the strong discriminatory power of these organs' predicted ages in distinguishing accelerated from decelerated ovarian aging.

[0229] E2 Levels Are Significantly Lower in Accelerated Aging Populations

[0230] The analysis revealed a significant decline in estradiol (E2) levels in women, particularly after the age of 50 (FIG. 27A) , with a more pronounced reduction in women exhibiting accelerated ovarian aging. In cohorts in their menopausal years and experiencing accelerated reproductive age, E2 levels were significantly lower compared to those with typical ovarian aging trajectories (FIG. 27B, FIG. 27C) . This reduction in E2 levels was consistent with the clinical hallmark of premature ovarian insufficiency (POI) , a condition characterized by an earlier onset of menopause and associated with higher risks for cardiovascular disease, osteoporosis, and other age-related conditions.

[0231] These findings suggest a strong correlation between reduced E2 levels and accelerated ovarian aging across multiple organ systems. This suggests that POI is not merely an isolated reproductive issue but rather a systemic condition involving multi-organ dysfunction. Moreover, the observed E2 reductions in individuals with accelerated aging point to potential hormonal dysfunctions contributing to the premature aging process in the reproductive system. Understanding the underlying mechanisms behind these hormonal changes could lead to more precise interventions aimed at slowing the progression of ovarian aging and mitigating associated health risks.

[0232] Ramelteon Treatment Alleviates Ovarian Aging

[0233] Given the observed hormonal changes and accelerated ovarian aging in individuals with POI, potential therapies targeting E2 levels were explored. Through phenotypic screening , Ramelteon, a melatonin receptor agonist, was identified as a promising candidate. Ramelteon has primarily been used to treat sleep disorders, but it has also shown neuroprotective and anti-inflammatory properties in various contexts. It was hypothesized that it may treat ovarian aging by influencing cellular metabolism and hormone regulation.

[0234] In previous work, it was confirmed that the perturbation of ceramide synthesis can prevent the development of POI, while ceramide accumulation exacerbates ovarian dysfunction. To assess the impact of Ramelteon on ovarian health, human granulosa (KGN) cells were used, which serve as a reliable model for studying ovarian function through E2 secretion and follicle count. To induce ovarian dysfunction, KGN cells were challenged with ceramide, which significantly reduced E2 secretion and follicle count, hallmarks of ovarian failure. Treatment with Ramelteon, however, rescued E2 secretion. To further investigate the therapeutic potential of Ramelteon, ovarian damage in mice were induced by subjecting them to a branched-chain amino acid (BCAA) restriction diet, a known model for POI. In these mice, Ramelteon treatment not only restored E2 levels but also improved fertility and ovarian function . Additionally, Ramelteon reduced ovarian damage and infertility in mice with established POI and also improved E2 levels and fertility in naturally aged female mice.

[0235] Gene Set Enrichment Analysis revealed that Ramelteon modulated gene expression in KGN cells, downregulating genes related to steroid, cholesterol, thioester and protein synthesis, while upregulating genes related to ovarian development and function. Notably, Ramelteon upregulated CYP19A1, a key factor for estrogen synthesis, whereas it has not been observed in treatment using other melatonin receptor agonists. This observation, along with the low expression of melatonin receptors in KGN cells, led us to speculate that Ramelteon exerts its effects through via a non-canonical mechanism, independent of the traditional melatonin receptor pathways. Structural modifications of Ramelteon abolished its effects on CYP19A1 expression and E2 secretion in KGN cells.

[0236] In clinical trials with POI patients and menopausal volunteers, it was found that Ramelteon treatment for one month to significantly decreased biological age, increased serum E2 levels, and improved menopausal symptoms. Remarkably, some volunteers even experienced restoration of their menstrual cycle after Ramelteon treatment. These results highlight the potential of Ramelteon treatment in enhancing ovarian functions and delaying ovarian aging in mammals.

[0237] Discussion

[0238] This example introduces a digital twin-based transformer model for predicting ovarian aging and demonstrates that accelerated ovarian aging, marked by a significant decline in estradiol (E2) levels, is a pivotal driver of broader biological aging in women, particularly around the perimenopausal period. Using a cohort of over 2 million individuals from CHAI and the UK Biobank, it was observed that women between the ages of 40 and 50 experience a sharp increase in biological age, coinciding with the onset of menopause. This accelerated aging is tightly linked to diminished ovarian function, as indicated by reduced E2 secretion and follicular health. The sharp decline in E2 levels in individuals with accelerated ovarian aging underscores the central role of ovarian dysfunction in systemic aging, highlighting the need for targeted therapies aimed at restoring E2 levels and promoting ovarian health.

[0239] The aging process of ovaries is a critical aspect of aging research, particularly in understanding conditions like POI, which affects young women. POI can be caused by many factors, and studying its mechanisms can provide insights into ovarian aging and inform targeted therapies. While most existing POI research focuses on genetic factors, recent findings reveal that only 23.5%of POI cases can be explained by coding variants. Notably, mutations in mitochondrial leucyl-tRNA synthetase, a key factor in leucine synthesis, have been linked to POI. The findings here are consistent with the clinical hallmark of premature ovarian insufficiency (POI) , a condition characterized by early-onset menopause and reduced E2 levels. In individuals with accelerated ovarian aging, E2 levels were significantly lower than those observed in women with typical ovarian aging trajectories. This reduction in E2 is not only a reproductive issue but also a systemic concern, as it correlates with aging across multiple organ systems. These results suggest that POI may represent a broader, multi-organ dysfunction, where early ovarian aging accelerates the aging process in other tissues, potentially contributing to an increased risk for cardiovascular diseases, osteoporosis, and other age-related conditions.

[0240] A striking feature of accelerated ovarian aging in women is its overlap with inflammatory and fibrosis pathways, which are known to exacerbate cellular dysfunction. Thus, targeting these pathways may slow ovarian aging and its systemic effects on health. Ceramide, a lipid that accumulates during ovarian aging, plays a central role in promoting inflammation and fibrosis. Previous work has shown that disrupting ceramide synthesis can prevent POI, while ceramide buildup accelerates ovarian dysfunction. This accumulation triggers inflammatory and fibrosis cascades, further reducing ovarian function and E2 secretion. CERTL, a key regulator of ceramide metabolism, facilitates ceramide transfer between cellular compartments, contributing to its accumulation in tissues. In ovarian aging, elevated ceramide levels driven by CERTL activation may intensify inflammation and fibrosis, impairing ovarian function. Modulating CERTL expression or activity could help reduce ceramide accumulation, mitigate inflammation, and preserve E2 secretion, potentially slowing ovarian aging.

[0241] In both in vitro and in vivo models, Ramelteon was able to significantly restore E2 secretion and improve ovarian function. In human granulosa cells (KGN cells) treated with ceramide, Ramelteon reversed the ceramide-induced decline in E2 secretion and follicular health, suggesting its potential to modulate ceramide-driven inflammatory and fibrosis pathways. In a mouse model of POI, Ramelteon not only restored E2 levels but also improved fertility and reduced ovarian damage, further validating its therapeutic potential. These results are promising, as they suggest that Ramelteon may act through mechanisms that both restore hormonal balance and reduce the underlying inflammation and fibrosis that accelerate ovarian aging.

[0242] Currently, the approved therapies for ovarian aging and related complications are mostly hormone replacement therapy (HRT) . Though HRT can certainly relieve the symptoms and complications of ovarian aging, precision medicine is still desired for a cure. There are a couple of clinical trials based on repurposing approved drugs for vasomotor symptoms nowadays, but the targeted therapies to directly enhance ovarian functions and fertility are still missing. Ramelteon is the first approved drug that has shown beneficial effects on ovarian aging in both mice and humans. Although the effects of Ramelteon for sleep disorders have been well-established since its clinical approval in 2005, there is limited data on its use in large epidemiology dataset like the UK Biobank. It is encouraging to identify the therapeutic effects of Ramelteon in menopausal women through this investigator-initiated trial, though a larger, more rigorous clinical trial is needed for full validation.

[0243] Despite certain limitations, such as the observational nature of the datasets and potential biases in the longitudinal cohort, this example highlights the effectiveness of organ-specific aging models in capturing critical information that conventional models may overlook. The strong predictive accuracy demonstrated by the digital twin-based aging clock underscores its potential for future applications and research. While aging clocks primarily estimate biological age using specific biomarkers, digital twins extend this capability by integrating diverse data sources and providing a more dynamic, holistic approach to studying the aging process. By integrating diverse data sources, digital twins enhance traditional aging clocks, transforming them from static tools into dynamic, continuously updated systems52. In the future, incorporating wearable devices, cloud medical records, and environmental sensors can enable aging clocks to use the most current data, improving their adaptability and accuracy53. This synergy between digital twins and aging clocks creates a powerful framework for advancing personalized healthcare strategies aimed at promoting healthy aging and facilitating timely interventions aimed at mitigating aging-related decline.

[0244] The findings suggest that anti-aging interventions should focus on specific organ systems and advocate for a more personalized approach, and that chronic diseases could reflect accelerated aging processes, while emphasizing the intricate dynamics of organismal aging. Methods Study Populations

[0245] The UK Biobank initiated a comprehensive recruitment campaign between 2012 and 2023, enrolling approximately 500,000 individuals aged 40 to 69 years from various regions in the United Kingdom. International Digital twin Consortium China Aging study enrolled approximately 5,000,000 individuals aged 0 to 95 years from various regions in China with EHR. This extensive dataset encompasses a wealth of personal details and medical histories. Recently, Eldjarn et al. (2023) employed the high-throughput Olink Explore 3072 platform to quantify 3,072 plasma proteins in a cohort of 53, 015 UK Biobank participants. Additionally, they assessed 1,463 proteins in plasma samples from 1,132 participants during the imaging visit in 2014 and from 1,006 participants during a follow-up imaging visit in 2019. For validation, datasets from US and unrelated EHR datasets in China were utilized. Data preprocessing

[0246] In the CHAI dataset, health records mainly consist of two types of features: continuous features and categorical features. Continuous features were also converted into categorical features to facilitate subsequent model computations. Specifically, for a given continuous feature, the maximum value fmax and the minimum value fmin within the population were first calculate. Then, all values between fmin and fmax were discretize into N bins of equal width. In this way, each continuous feature is transformed into a categorical feature with N distinct values. Input structure

[0247] A chronological sequence of examination results was assembled for each patient. Each sequence has a form S= {s0, s1, s2, …, sN} , where si is the i-th examination event. Each examination event contains the values of various lab tests for the patient during that examination. Due to the data preprocessing steps, all continuous features have been converted into categorical features. Model architecture

[0248] The model consists of three components: an embedding layer, encoders, and task-specific decoders. The encoder is a transformer-based model, and the decoders are fully connected neural networks. Inputs and embedding layer

[0249] The embedding layer transforms the raw life-sequence into the format that life2vec can process. Given a sequence Sp, representations of tokens were looked up in the embedding matrix where each row of corresponds to a token in the vocabulary (d is the number of hidden dimensions) . Additionally, the segment embedding was looked up in the  matrix. Both and matrices are optimized during the model training. To improve the representation of rare concept tokens and the overall isotropy of the concept embedding space, the global mean from each row of the matrix was removed. That is, each time the token embedding was looked up, the mean was subtracted.

[0250] For each token v in s, sum the associated token embedding in and the temporal embedding of the sentence,  The input to the life2vec model is a concatenated sequence of these token representations (that is, a multidimensional tensor) . Encoder component

[0251] For the multiple encoder blocks, each block processes input representations and passes the results to the next encoder (or decoder) . The architecture of each block is identical and comprises multi-head attention, a position-wise layer, and two residual connections.

[0252] The multi-head attention module comprises several attention heads, which separately process the input representations. The original BERT uses softmax self-attention heads. Devlin, J., Chang, M.-W., Lee, K. &Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. in Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) 4171-4186 (2019) . Each head takes input representations and transforms these with several dense layers-query, key, and value. These layers output linearly transformed representations Q, K, V∈RL×d, where L is the length of the sequence and d is the dimensionality of embeddings. The contextualized representations are computed as:  where D=diag (A1L) . Softmax attention is suboptimal for sequences of length more than 512 tokens. Therefore, softmax attention heads were used only to model local interactions; that is, the span of these heads was limited to 38 neighboring tokens.

[0253] To capture global interactions, performer-style attention heads were used, as they can handle longer sequences. Instead of calculating the precise attention matrix A∈RL×L, performer-heads approximate it via matrix factorization. Entries of the approximated attention matrix are computed using kernels:   (indexes stand for the rows of matrices) . The kernel function is defined as: K (x, y) =E [φ (x) Tφ (y) ] , where φ (u) is a random feature map that projects input into the r-dimensional space. Random mapping φ is constrained to contain features that are positive and exactly orthogonal. If φ is applied to Q, K, Q′, K′∈RL×r is obtained, where r <<L. The attention is now defined as:  where Each multi-head attention module of life2vec has four performer-style attention heads and four softmax attention heads. The output of these heads is concatenated and transformed with one more dense layer.

[0254] The encoder blocks also have a position-wise feed-forward module (PFF) . This consists of two fully connected feed-forward layers that apply additional nonlinear transformations to each representation: fPFF (x) =swish (xW1+b1) W2+b2, where swish (x) =x·sigmoid (x) . Typically, the output representations of each module add up to the input representations via so-called residual connections: y=x+f (x) where f is a multi-head attention module or a position-wise feed-forward module. In this work, ReZero connections were used, which consist of a single scalar, α. This scalar controls the fraction of information that each layer contributes to the contextualized representations: y=x+α·f (x) . At the start of training, each α is initialized to zero (meaning none of the encoder layers contribute at the beginning) .

[0255] Several modifications to the BERT architecture were introduced, such as ReZero, ScaleNorm, Swish, and Weight Tying to speed up the convergence and reduce the size of the model. Protein Annotation

[0256] To identify proteins enriched in specific organs, organ specific clinical lab test and tissue transcription data from the Genotype-Tissue Expression (GTEx) project (Lonsdale et al., 2013) were utilized. Organs were categorized based on criteria set forth in Supplementary Table 2 from Oh et al. (2023) . A protein was deemed organ-enriched if its average GTEx gene counts for that organ were at least four times greater than in other organs. This transcriptional approach was chosen over proteomics data to better indicate the organ of origin for each protein.

[0257] Organ-specific aging models were trained using only the proteins and electrolytes enriched in respective organs, mirroring the methodology employed for traditional aging models. Models that maintained a correlation coefficient (r) exceeding 10%for both chronological and mortality-based aging models were retained for analysis across training and test datasets.

[0258] For analyses depicted, the human proteome was sourced from Uniprot on February 27, 2024. Secreted proteins were those annotated as part of the extracellular space in Gene Ontology cellular components (GOcc) or marked as secreted in Uniprot. Proteins categorized as extracellular membrane were defined as non-secreted proteins annotated as part of the plasma membrane, while all remaining proteins were classified as intracellular. Longitudinal Analysis and External Validation

[0259] Longitudinal analyses were focused on the 100 secreted proteins and electrolytes measured in the EHR. After removing proteins and elctrolytes with over 50%missing values, mortality-based models were retrained using the remaining proteins. Statistical Analysis

[0260] To evaluate the impact of individual plasma proteins on chronological age, ordinary linear regression was performed. Example 4: Hospital-acquired Infection Based on Time Sequence Data and  Digital Twin Technology

[0261] In this example, only pediatric data were selected, including 13, 108 hospital-acquired infections, 40, 608 community-acquired infections, and 101, 467 non-infections. For each patient, there are five types of features: vital signs, diagnostic records, medication records, surgical records, and laboratory tests. Positive samples are patients with hospital-acquired infections. The data includes the date of confirmed hospital-acquired infection and the data from 1 to 10 days before the infection. Negative samples are patients without hospital-acquired infections. Random dates were selected, along with data from 1 to 10 days before these dates. The number of samples is different multiples of positive samples, with zeros padding if fewer than 10 days. The data was randomly split by patient ID into a 9: 1 ratio for training and validation sets. The model was trained on the training set, and performance was reported on the validation set. Label Definition

[0262] Hospital-acquired infection: Confirmed by doctors according to the hospital-acquired infection diagnostic criteria issued by the Chinese Ministry of Health. Community-acquired infection: Not hospital-acquired infection and diagnosed with an infection with any type with the diagnosis date less than or equal to 2 days from the admission date. Non-infection: No infection-related diagnosis during the hospital stay.

[0263] FIG. 28 shows distribution of the time from hospital admission to the diagnosis of nosocomial infection. As the number of days of examination used increases, the performance of predicting whether a nosocomial infection will occur at the next checkup gradually improves (data not shown) . The performance of predicting nosocomial (hospital-acquired) infections can be influenced by both the department and the number of examination days included (data not shown) . The impact of different numbers of examination days on the performance of predicting nosocomial infections, community-acquired infections, and non-infections, was observed separately for adults and children (data not shown) . The impact of different numbers of examination days on the performance of predicting nosocomial infections, community-acquired infections, and non-infections for pediatric patients was also observed (FIG. 29) . FIG. 30 shows visualization of the feature contributions for nosocomial infections, where the color of the points represents the value of the sample for that feature, and the left-right position represents the direction of the contribution to the prediction result. Example 5: Dialysis Time Series Prediction

[0264] In this example, for dialysis time series prediction, data included follow-up records of patients with detailed records from the first dialysis, aligned by a maximum of 24 quarters, filling missing follow-ups with null values.

[0265] The tasks included: Mortality Prediction within 1-5 years -whether the patient will die within N years after the current follow-up; Indicator Prediction -predict the values of each indicator at the next follow-up; and Intermediate Outcome Prediction -Select CA, P, iPTH, DBP, SBP, ALB, and Hb as intermediate outcome values. Under intermediate outcome prediction: Improvement Prediction -positive samples: Abnormal segment followed by normal segment in follow-up records; Negative samples: Always abnormal; Deterioration Prediction -positive samples: Normal segment followed by abnormal segment in follow-up records; negative samples: Always normal. The model was based on Transformer, modify the feature embedding method and multi-task prediction. Example 6: Diabetes and Complications Prediction

[0266] In this example, the trained model was used in time series analysis of diabetic complications. FIG. 31 to FIG. 35 show the analysis and results using digital twin paradigm to predict diabetes and complications. Example 7: Pregnancy and Child Outcome Prediction

[0267] Maternal and infant data input features included diagnoses, examinations, surgeries, medications, and vital signs information. The target diseases included seven types: miscarriage, gestational hypertension, gestational diabetes, preterm birth, premature rupture of membranes, placental abruption, and postpartum hemorrhage. The total number of positive and negative samples was 135, 822. Data and labels within one year after the first record were selected for input and outputand results.

[0268] FIG. 36 shows predicting the gestational age of a pregnant woman during pregnancy. To predict the duration of pregnancy for pregnant women using their lab test features, the dataset was divided into training and testing sets with a ratio of 4: 1. FIG. 37 shows using the lab test features of already pregnant women to predict whether they would develop any of the aforementioned 7 diseases. The dataset was divided into training and testing sets with a ratio of 9:1. The first N visit records were selected to predict whether the corresponding disease would occur at the N+1th visit, with a maximum of 4 visit records. Example 8: Heart Failure Prediction

[0269] FIG. 38 to FIG. 39 show the analysis and results for using a model trained as disclosed herein to predict heart failure. Example 9: Myopia Prediction

[0270] FIG. 40 to FIG. 48 show the analysis and results for using a model trained as disclosed herein to predict myopia. Example 10: Vision Prediction in Anti-VEGF Therapy Using Time Series Data

[0271] FIG. 49 to FIGS. 50A-50B show the analysis and results for using a model trained as disclosed herein to predict vision in patients receiving anti-VEGF therapy. Example 11: Age Prediction Performance Using Lab Test Data

[0272] The classification of experimentally or clinically defined normal or healthy states differs in each individual and cannot be extrapolated to large populations. Currently, treatment personalization relies on low-resolution data and a limited picture of the clinical history of a particular person. For example, there is not yet a clear understanding of “normal” blood pressure or other phenotypes such as complete blood count (CBC) . The reasons may be due to the relatively sparse measurements of these phenotypes and the lack of assessment of the impact of physiological and behavioral patterns in any individual.

[0273] Without a personalized definition of normal, it is difficult to detect deviations from normal which ultimately constitute the disease state. A medical Digital Twin can define normal in each individual through feedback of information between the patient and their Digital Twin and the feedback can be continuous. Deviations from this normal state define disease, and treatments can be leveraged to predict intervention outcomes.

[0274] This example shows age prediction performance comparison in lab test data among CBC only, lab data without CBC, or full lab data including CBC. The results here demonstrate superior performance with full lab data, then lab data without CBC, and the least accurate is using CBC only. Names of the 200 full lab tests are listed in Table 1 below. The biomarkers are divided into three categories: CBC (highlighted in bold) , clinical chemistry and tumor biomarker tests. Each biomarker is listed with its name, the organ it targets, the organ where it is produced and its substance type. Clinical chemistry tests use chemical processes to measure levels of chemical components in body fluids and tissues. The most common specimens used in clinical chemistry are blood and urine. Many different tests exist to detect and measure almost any type of chemical component in blood or urine. Components may include blood glucose, electrolytes, enzymes, hormones, lipids (fats) , other metabolic substances, and proteins.

[0275] Table 1. Comprehensive List of Biomarkers in Clinical Laboratory Tests.

[0276] FIG. 51 shows comparison of biological age prediction performance using different subsets of laboratory test data for male and female individuals. The best performance in biological age prediction was achieved using full lab data, with the lowest MAE (males: 8.14, females: 6.83) , highest R2 (males: 0.60, females: 0.74) , and strongest PCC (males: 0.78, females: 0.86) . Lab data without CBC performed slightly worse but remained accurate (males: MAE 8.35, females: MAE 6.94) . Predictions based only on CBC data were the least accurate, with higher MAE (males: 12.79, females: 13.08) , lower R2, and weaker PCC, highlighting the value of integrating CBC and non-CBC features for optimal performance.

[0277] FIG. 52 shows comparison of pregnancy gestational age prediction performance using clinical laboratory test data. Predictions using full lab data achieved the best performance (MAE = 14.40, R2 = 0.82, PCC = 0.90) , followed closely by non-CBC lab data (MAE = 14.81, R2 =0.80, PCC = 0.90) . Predictions based only on CBC data were the least accurate (MAE = 38.72, R2 = 0.19, PCC = 0.43) .

[0278] These data demonstrate that hematological phenotypes such as CBC laboratory test data, non-CBC laboratory test data, and particularly combining CBC and non-CBC laboratory test data, can be patient-specific values that are useful in accurately predicting health-related conditions or disease states, such as aging. The stable phenotypes can be used for evaluating common conditions, such as those described in Examples 1-9 above, in healthy individuals and patients. Example 12: Reproductive Aging Promotes Whole Body Organ Specific Aging in  Women

[0279] Aging is a complex, multi-scale process that manifests at molecular, cellular, organ, and whole-organism levels, leading to functional decline, age-related diseases, and ultimately mortality. This example presents a whole body as well as female reproductive system-specific AI model for aging prediction trained on clinical laboratory tests data from over 2 million human adult subjects across multiple cohorts, including the China Health Aging Investigation (CHAI) and UK Biobank (UKB) . Each patient’s longitudinal time serial laboratory data were analyzed using a transformer-based digital twin technology to predict the biological age. The results showed that adult women in perimenopause or menopause experienced accelerating reproductive aging, as well as aging in many other organs. A dramatic acceleration of aging in reproductive system was also observed in a separated cohort of patients with premature ovarian insufficiency (POI) , a disease featured by ovarian dysfunction and failure before 40. The results showed unhealthy dietary habits led to branch-chain amino acids nutritional intake deficiencies and POI. A phenotypic screen identified ramelteon, an FDA approved drug for sleeping disorder, significantly rejuvenated the aged ovaries. The results offer new insights into the aging process of human and potentially new therapies aimed at mitigating age-related reproductive and other vital organ declines.

[0280] Aging is a complex, multifaceted process involving changes at the molecular, cellular, and organ level, ultimately impacting whole-organism health and survival. Deciphering how these changes contribute to increased disease susceptibility and mortality risk is essential for advancing interventions that extend health span. Biological age, a measure of accumulated biological damage relative to an average individual of the same chronological age, has emerged as a key metric for assessing age-related disease risk. Unlike chronological age, biological age can diverge, providing a valuable metric for age-related disease risk.

[0281] Initially, biological age relied on the measurement of DNA methylation patterns, but recent advancements have expanded aging clocks to include additional "omics" layers to provide more accurate predictions of biological age. For example, the applications of mass spec or antibody-based proteomics analysis on the biological samples of large populations have generated many useful resources for aging study. These innovations highlight the variability in aging across organs and their differential response variably to external influences, such as lifestyle or medications, offering new avenues for tailored anti-aging strategies.

[0282] The reproductive system is important for women’s fertility, endocrine functions and many related physiological processes. The aging of female reproductive system happens in the perimenopause (45-55 age) , which is much earlier than the aging of male productive system and many other organs. In addition to the naturally aging process, dysfunction of the female reproductive system may be accelerated in women before 40 due to a disease called premature ovarian insufficiency (POI) . Besides the loss of fertility, the aging of reproductive system leads to an increased risk for several severe diseases including cardiovascular diseases in these patients. Though hormone replacement therapy (HRT) can mitigate some of the symptoms related to natural aging or POI, it also increases the risk of neoplasm in tissues such as endometrium or breast. Therefore, there remains a significant need to systemically understand the aging process of female reproductive system and to develop specific therapies.

[0283] A dataset, the China Health Aging Investigation (CHAI) dataset, was established to contain the longitudinal medical records of two million healthy Chinese patients. Based on CHAI, a transformer-based model, EHRFormer, was established to create the digital twins of two million Chinese and predict their biological age. With the whole-body and reproductive system-specific age prediction model, it was discovered that reproductive aging is a major driver of whole-body accelerated aging during perimenopause and menopause in women, which was further validated in a cohort of patients with POI. The unhealthy dietary pattern induced metabolic disorder of branched chain amino acid (BCAA) contributed to the development of female reproductive aging. Phenotypic screen identified that the already FDA-approved drug ramelteon enhanced the production of E2 and prevented the female reproductive aging in mice models. The results in this example link reproductive aging to multi-organ aging and improve the understanding of the interplay between biological aging and disease risk.

[0284] A blood test-based aging clock identified the accelerated aging in women during peri-menopause.

[0285] To investigate the dynamics of ovarian aging, an aging clock that leveraged electronic health records (EHR) was built, including longitudinal laboratory data from a cohort of 2 million healthy male and female individuals sourced from CHAI. By employing a transformer-based digital twin approach, a model called EHRFormer was developed for predicting biological age based on the EHR (FIG. 53) . Using unsupervised learning, the model identified patterns and extracted age-and gender-specific features that potentially serve as biomarkers for gender-specific biological aging. Then the EHRFormer was applied to individual patient’s longitudinal EHR data to generate personalized digital representations of each patient, known as digital twins. These digital twins served as the foundation for this aging clock as they simulate each complex detail of a patient’s health data in order to make personalized predictions of health states in the future. The model was trained separately for each gender and age group (age >20 years) . The predicted age made by EHRFormer demonstrated a strong correlation with chronological age for each gender and age group (FIG. 54A) , providing a robust framework for predicting biological age based on routine clinical laboratory data. The results were replicated in the UKB cohort (FIG. 54B) .

[0286] It has been speculated that males and females exhibit distinct aging patterns, though this hypothesis has not been tested and exact patterns have not been delineated in a systemic manner with large populations. The analysis in this example revealed that the predicted biological age demonstrated a strong correlation with chronological age for both genders but diverges markedly from chronological age in older women (FIG. 54B, FIG. 54C, FIG. 58) . Notably, from ages of perimenopause for women (45-55 years old) , the rate of accelerated biological aging in females increases markedly, a phenomenon not observed in males of the same age group. This finding was further validated by applying the biomarker-based prediction model to data from the UK Biobank (UKB) . Consistent with the observations on CHAI, the predicted age strongly correlated to chronological age for both genders, with accelerated aging trend exclusively observed in women experiencing perimenopause or menopause in UKB (FIG. 54C, FIG. 54D) . The predicted age for the patients with POI, a disease featured by premature dysfunction of ovaries before the age of 40, were then calculated. Comparing to the matched healthy volunteers, patients with POI presented a dramatically accelerated aging independent of HRT, which confirmed the effects of menopause on the aging process of women (FIG. 56H, FIG. 60) .

[0287] Reproductive aging is a major driver of aging in multiple organs in women.

[0288] The sharp acceleration of biological age observed in women experiencing perimenopause or menopause suggested that ovarian aging may play a pivotal role in driving the accelerated aging process. To investigate this aspect, organ-and system-specific aging models were developed. Specifically, laboratory tests relevant to different organs or systems were selected for fine-tuning based on the pre-trained EHRFormer, hence 13 organ-and system-specific aging models were obtained. To evaluate these models, the correlation between their predicted ages and chronological age were first examined. As expected, the predicted ages from all models exhibited strong correlations with chronological age. Moreover, these models were applied to calculate the organ-and system-specific ages of patients with various diseases in the CHAI cohort and their predicted ages were compared with those of healthy individuals. The results reveal that patients with specific organ or system related diseases exhibited older age of that organ or system compared to healthy individuals and patients with unrelated disease. Furthermore, if individuals showed more than 2 standard deviations than the average predicted organ or system age of that chronological age group, they would have higher hazard ratios and cumulative disease rates for that organ or system (FIG. 54E, FIG. 54H) .

[0289] By using the reproductive system-specific aging model, an accelerated aging was found in females between ages 45-55 in CHAI cohort (FIG. 55) . In this model, participants were stratified into the two groups based on their predicted reproductive system specific age (the top 50%and bottom 50%, FIG. 55A) . Those with a predicted age higher than the average for their chronological age were placed in the accelerated aging group, indicating faster biological aging. Conversely, those with a predicted age lower than the average are placed in the decelerated aging group, reflecting slower biological aging. By comparing these two groups, it was observed the accelerated aging group consistently exhibited higher predicted age in not only the reproductive system (FIG. 55A) but almost all organ systems (FIG. 55B-55K, FIG. 59) in CHAI, indicating a strong correlation between reproductive aging and whole-body aging. Furthermore, accelerated aging of reproductive system and whole-body was also identified in the patients with POI independent of HRT, compared to otherwise matched control (FIG. 56H) .

[0290] These findings were subsequently validated in an independent UKB cohort with the same reproductive-system specific aging model. Participants were categorized into accelerated aging and decelerated ovarian aging groups, following the same criteria as in the CHAI dataset. In UKB cohorts, the accelerated reproductive aging group showed higher predicted ageing for the whole-body and other organ systems in comparison to the decelerated aging group. The widespread accelerated organ aging in CHAI and UKB highlighted the significant impacts of reproductive aging on the status of overall aging process in women.

[0291] Unhealthy dietary pattern is a key factor in female reproductive aging in POI patients.

[0292] Based on clinical observation as well as reports from the literatures, the patients with POI are in general leaner than the normal population, indicating a distinct metabolic feature of female reproductive aging. Therefore, the metabolic features in POI patients were investigated. Downregulation of both serum protein level and skeletal muscle mass in the patients with POI (FIG. 56A, FIG. 56B) were observed suggesting a disorder of metabolic homeostasis. The body composition is largely affected by diet. the frequency and amount of food intake in POI patients were measured and matched healthy volunteers. For the frequency-based analysis, it observed that the intake of coarse grain (rich in fiber) , ferment food (rich in probiotics) , vegetables (rich in fiber) , river / sea food (rich in high-quality protein) , and nut (rich in unsaturated fatty acid) was downregulated, while the intake of red meat and processed meat was upregulated for the patients with POI upon adjustment to age and BMI. The patients with POI followed an unhealthy dietary style as their dietary habits exhibited a negative association to the score of Alternate Mediterranean Diet (aMED) or Chinese Healthy Eating Index (CHEI) (PMID: 28872591) . For the amount-based analysis, the deceased intake of protein and carbohydrate were observed but not fat for the patients with POI comparing to the healthy volunteers. Upon adjustment to energy intake, age and BMI, it was identified that the intake of tubers, corn, coarse grain, ferment foods, vegetables and river / sea food was significantly reduced, while the intake of processed meat, tea, soft drink and beer was increased for the patients with POI. Similar to the results of frequency-based study, these patients presented an unhealthy dietary style as their negative association to the score of CHEI.

[0293] It has been demonstrated that the abundance of metabolites in the serum was affected by dietary patterns. Thus, the abundance of metabolites in the serum of POI patients and healthy volunteers were further investigated. Untargeted metabolomics and lipidomic analysis were performed on the serum of POI cohort. Consistent to the results of a pilot study with a small cohort, a remarkable decrease of BCAA including valine, and leucine-isoleucine (FIG. 56C) and an increase in ceramide (FIG. 56D) were identified in the serum of patients with POI compared to healthy controls.

[0294] Ramelteon enhances E2 production and alleviated the female reproductive aging.

[0295] The serum level of estradiol (E2) , mostly driven by the synthesis of E2 in the ovaries, has been recognized as a key hormone for homeostasis in women’s health. In both the CHAI and UKB cohorts, a significant decline in serum estradiol (E2) levels was observed in women experiencing menopause (after the age of 50) . Based on the reproductive system-specific prediction model, serum E2 levels were significantly lower in women with accelerated ovarian aging relative to those with a typical aging trajectory in CHAI (FIG. 56F) and UKB (FIG. 56G, FIG. 60) . Serum E2 levels were also significantly reduced in the serum of patients with POI comparing to the matched healthy volunteers (FIG. 56G, FIG. 60) . These findings suggest a strong correlation between reduced E2 levels and accelerated reproductive system aging in women.

[0296] In previous work, it was confirmed that the abnormal elevation of ceramide in patients with POI, which was validated in a much larger cohort in current study, disrupted the production of E2 from the ovarian granulosa cells. Through ceramide-based phenotypic screening (FIG. 57A) , ramelteon, a melatonin receptor agonist approved for sleeping disorder, was identified to increase the E2 production from human ovarian granulosa cell line KGN cells (FIG. 57B, FIG. 57C) . In mice with dietary BCAA restriction induced POI, ramelteon treatment not only prevented the downregulation of E2 levels (FIG. 57D) but also protected the fertility (FIG. 57E) , increased the number of primordial follicles and decreased the number of atretic follicles in the ovaries (FIG. 57F) . Importantly, the treatment of ramelteon affected the naturally aging process of female reproductive system by increasing the serum E2 levels and improving fertility in 18-month-old mice (FIG. 57G) .

[0297] The mechanism of ramelteon was investigated by transcriptome analysis of KGN cells. Gene Set Enrichment Analysis on the transcriptomics data revealed that ramelteon modulated gene expression in KGN cells, downregulating genes related to steroid, cholesterol, thioester and protein synthesis, while upregulating genes related to ovarian development and function (FIG. 61A, FIG. 61B) . Notably, ramelteon upregulated the expression of CYP19A1 gene, which encodes a key protein for estrogen synthesis (FIG. 61C) . Similar upregulation was not observed when KGN cells were treated by other melatonin receptor agonists (FIG. 61D) , indicating the specific effects of ramelteon. These data suggested that ramelteon can prevent the female reproductive aging by directly enhancing the function of granulosa cells.

[0298] Aging is a very complicated process. On one hand, all the organs may contribute to aging in different manner; on the other hand, the interactions among different organs may have significant impacts. Therefore, to study the whole organism systemically is essential for aging studies. With a cohort of over 2 million individuals from CHAI and a digital twin-based transformer model, it was observed that women experiencing perimenopause presented a sharp increase in the biological age of their female reproductive system, which may induce the accelerated aging of the whole-body. The sharp decline in E2 levels in individuals with accelerated reproductive aging underscores the central role of ovarian dysfunction in systemic aging, highlighting the need for targeted therapies aimed at restoring E2 levels and promoting ovarian health.

[0299] Currently, the approved therapies for ovarian aging and related complications are mostly HRT. Though HRT can certainly relieve the symptoms and complications of ovarian aging, precision medicine is still desired for cure. There are a couple of clinical trials based on re-purposing approved drug for vasomotor symptoms nowadays (PMID: 37678251) , but the targeted therapies to directly enhance ovarian functions and fertility are still missing. Ramelteon is the first approved drug which presented beneficial effects on ovarian aging in mice and humans. Though the effects of ramelteon for sleeping disordered has been validated for a long time (PMID: 16173650; PMID: 16309958) and it has been approved for clinical application since 2005, there is a lack of relevant information for it in large epidemiology dataset such as UKBB. It is encouraging to identify the therapeutic effects of ramelteon on the menopausal with mouse models in the current study.

[0300] This example highlights the effectiveness of organ-specific aging models in capturing critical information that conventional models may overlook. The strong predictive accuracy demonstrated by the digital twin-based aging clock underscores its potential for future applications and research. While aging clocks primarily estimate biological age using specific biomarkers, digital twins extend this capability by integrating diverse data sources and providing a more dynamic, holistic approach to studying the aging process. By integrating diverse data sources, digital twins enhance traditional aging clocks, transforming them from static tools into dynamic, continuously updated systems. Incorporating wearable devices, cloud medical records, and environmental sensors, enables aging clocks to use the most current data, improving their adaptability and accuracy. This synergy between digital twins and aging clocks creates a powerful framework for advancing personalized healthcare strategies aimed at promoting healthy aging and facilitating timely interventions aimed at mitigating aging-related decline.

[0301] The findings suggest anti-aging interventions to focus on a specific organ system and advocate for a more personalized approach, proposing that chronic diseases could reflect accelerated aging processes, while emphasizing the intricate dynamics of organismal aging.

[0302] Data preprocessing. In the CHAI dataset, health records mainly consist of two types of features: continuous features and categorical features. Continuous features were converted into categorical features to facilitate subsequent model computations. Specifically, for a given continuous feature, the maximum value fmax and the minimum value fmin within the population were first calculated. All values between fmin and fmax were discretized into N bins of equal width. In this way, each continuous feature was transformed into a categorical feature with N distinct values.

[0303] Input structure. A chronological sequence of examination results for each patient was assembled. Each sequence has a form S= {s0, s1, s2, …, sN} , where si is the i-th examination event. Each examination event contains the values of various lab tests for the patient during that examination. Due to the data preprocessing steps, all continuous features have been converted into categorical features.

[0304] Model architecture. The model consists of three components: an embedding layer, encoders, and task-specific decoders. The encoder is a transformer-based model, and the decoders are fully connected neural networks.

[0305] Inputs and embedding layer. The embedding layer transforms the raw life-sequence into the format that life2vec can process. Given a sequence Sp, representations of tokens were looked up in the embedding matrix  where each row of corresponds to a token in the vocabulary (d is the number of hidden dimensions) . Additionally, the segment embedding was looked up in the  matrix. Both and matrices were optimized during the model training. To improve the representation of rare concept tokens and the overall isotropy of the concept embedding space, the global mean was removed from each row of the matrix. That is, each time the token embedding was looked up, the mean was subtracted. For each token v in s, the associated token embedding in and the temporal embedding of the sentence,  were summed. The input to the life2vec model is a concatenated sequence of these token representations (that is, a multidimensional tensor) .

[0306] Encoder component. Like the original BERT, life2vec consists of multiple encoder blocks. Each block processes input representations and passes the results to the next encoder (or decoder) . The architecture of each block is identical and consists of multi-head attention, a position-wise layer, and two residual connections. The multi-head attention module consists of several attention heads, which separately process the input representations. The original BERT uses softmax self-attention heads. Each head takes input representations and transforms these with several dense layers-query, key, and value. These layers output linearly transformed representations Q, K, V∈RL×d, where L is the length of the sequence and d is the dimensionality of embeddings. The contextualized representations are computed as: Att (Q, K, V) =softmax where D=diag (A1L) . Softmax attention is suboptimal for sequences of length more than 512 tokens. Therefore, softmax attention heads were use only to model local interactions; that is, the span of these heads was limited to 38 neighboring tokens.

[0307] To capture global interactions, performer-style attention heads were use, as they can handle longer sequences. Instead of calculating the precise attention matrix A∈RL×L, performer-heads approximate it via matrix factorization. Entries of the approximated attention matrix are computed using kernels:   (indexes stand for the rows of matrices) . The kernel function is defined as: K (x, y) =E [φ (x) Tφ (y) ] , where φ (u) is a random feature map that projects input into the r-dimensional space. Random mapping φ is constrained to contain features that are positive and exactly orthogonal. If φ is apply to Q, then K, Q′, K′∈RL×r is obtained, where r <<L. The attention is now defined as:  where Each multi-head attention module of life2vec has four performer-style attention heads and four softmax attention heads. The output of these heads is concatenated and transformed with one more dense layer.

[0308] The encoder blocks also have a position-wise feed-forward module (PFF) . This consists of two fully connected feed-forward layers that apply additional nonlinear transformations to each representation: fPFF (x) =swish (xW1+b1) W2+b2, where swish (x) =x·sigmoid (x) . Typically, the output representations of each module add up to the input representations via so-called residual connections: y=x+f (x) where f is a multi-head attention module or a position-wise feed-forward module. Here, ReZero connections were used, which consist of a single scalar, α. This scalar controls the fraction of information that each layer contributes to the contextualized representations: y=x+α·f (x) . At the start of training, each α is initialized to zero (meaning none of the encoder layers contribute at the beginning) .

[0309] Several modifications were introduced to the BERT architecture, such as ReZero, ScaleNorm, Swish, and Weight Tying to speed up the convergence and reduce the size of the model.

[0310] Protein Annotation. To identify proteins enriched in specific organs, organ specific clinical lab test and tissue transcription data from the Genotype-Tissue Expression (GTEx) project (Lonsdale et al., 2013) were utilized. Organs were categorized based on criteria set forth in Supplementary Table 2 from Oh et al. (2023) . A protein was deemed organ-enriched if its average GTEx gene counts for that organ were at least four times greater than in other organs. This transcriptional approach was chosen over proteomics data to better indicate the organ of origin for each protein. Organ-specific aging models were trained using only the proteins and electrolytes enriched in respective organs, mirroring the methodology employed for traditional aging models. Models that maintained a correlation coefficient (r) exceeding 10%for both chronological and mortality-based aging models were retained for analysis across training and test datasets. For analyses depicted, the human proteome was sourced from Uniprot on February 27, 2024. Secreted proteins were those annotated as part of the extracellular space in Gene Ontology cellular components (GOcc) or marked as secreted in Uniprot. Proteins categorized as extracellular membrane were defined as non-secreted proteins annotated as part of the plasma membrane, while all remaining proteins were classified as intracellular.

[0311] Longitudinal Analysis and External Validation. For longitudinal analyses, the focus was on the 100 secreted proteins and electrolytes measured in the EHR. After removing proteins and electrolytes with over 50%missing values, mortality-based models were retrained using the remaining proteins.

[0312] Dietary Analysis. Information on dietary intake was assessed using a 32-item food frequency and quantity questionnaire. Using the frequency data, the daily frequency for food groups (i.e., rice, leafy vegetables, poultry) and dietary pattern scores (i.e., Alternate Mediterranean Diet [aMED] , Dietary Approaches to Stop Hypertension [DASH] , Plant-based Diet Index [PDI] ) were calculated. Using the frequency and quantity data, the daily quantity for macronutrients (i.e., protein, carbohydrate) and micronutrients (i.e., vitamins, minerals) , daily quantity for food groups, and dietary pattern scores was calculated. The baseline characteristics of study participants were presented as mean and standard deviation for continuous variables, and number and percentage for categorical variables. For group comparisons, student’s t test was used for continuous variables, and the chi-square test was used for categorical variables. When comparing the dietary intake frequency between controls from TAP participants and cases from POI cohort, cases and controls were matched based on age (±2 years) and BMI (±1 kg / m2) at a ratio of 1: 4. Conditional logistic regression was used to investigate the relationships between daily frequency and dietary pattern scores and odds ratio (OR) for POI. When comparing the dietary intake quantity between controls from PUNCH participants and cases from POI cohort, all individuals were included. Logistic regression was used to investigate the relationships between daily quantity for nutrients, daily quantity for food groups, and dietary pattern scores and OR for POI.

[0313] Untargeted Metabolomics. To extract the metabolites, 240μL MeOH: ACN (v: v, 1: 1) was added to 60μl of homogenized sample, vortexed for 30s and sonicated for 10 min. Proteins were prone to precipitate by being incubated 1 h at 20℃. Then the sample was centrifuged at 13000 rpm for 15min at 4℃, and the supernatant was collected and evaporated to dryness in a vacuum concentrator. The dry extracts were then resuspended in 50μL of 1: 1 ACN: H2O and sonicated 10min. The mixture was centrifuged at 13,000 rpm for 15min at 4℃, and the supernatant was collected and stored at -80℃. The supernatant was analyzed by HPLC–MS / MS on Orbitrap Exploris TM 480 mass spectrometer coupled to an Vanquish liquid chromatography system (Thermo Fisher Scientific, USA) . For Hilic system separation, the ACQUITYUPLC BEH C18 column (100mm × 2.1mm i.d., 1.7μm; Waters) was used. 5μL sample was injected and separated with a 12min gradient. The column flow rate was maintained at 500μL / min with the column temperature of 40℃. The electrospray ionization mass spectra were acquired in positive ion mode and negative ion mode respectively. Data dependent acquisition (DDA) was used to collect full scan MS and MSMS information simultaneously. The ion spray voltage was set to 3,500V for positive mode and 2,800V for negative mode. The survey of full scan MS spectra (m / z 70-1200) was acquired in the Orbitrap with 60,000-resolution. The normalized automatic gain control (AGC) target at 100%and the maximum injection time was 100ms. Then the precursor ions were selected into collision cell for fragmentation by higher-energy collision dissociation (HCD) , the collection energy was 20, 30, 40. The MS / MS resolution was set at 30,000, the normalized automatic gain control (AGC) target at 100%, the maximum injection time was 60ms, isolation window was 1m / z, and dynamic exclusion was 4 seconds. ProteoWizard (version 3.0.6150) was used to convert raw MS data files to mzXML format and MS2datafilesto mgf format. All of MS files (mzXML format) were processed using R package “XCMS” (version 1.46.0) for peak detection and alignment. Metabolites identification was achieved by MetDNA (metdna. zhulab. cn / ) withtheMS1 peak table and MS2 data files (mgf format) (PMID: 30944337) . The data was statistically analyzed with MetaboAnalyst according to the guideline (PMID: 38693118) .

[0314] Untargeted Lipidomics. The untargeted lipidomics method was modified from a published method (PMID: 28496395) . Lipid samples were resuspended in 50 μL of isopropanol: acetonitrile: water (vol / vol / vol, 30: 65: 5) , and 10 μL was injected into Orbitrap Exploris 480 LC-MS / MS (Thermo, USA) coupled to HPLC system (Shimadzu, Kyoto, Japan) . Lipids were eluted via C30 by using a 3 μm, 2.1 mm × 150 mm column (Waters) with a flow rate of 0.26 mL / min using buffer A (10 mM ammonium formate at a 60: 40 ratio with acetonitrile: water) and buffer B (10 mM ammonium formate at a 90: 10 ratio with isopropanol: acetonitrile) . Gradients were held in 32%buffer B for 0.5 min and run from 32%buffer B to 45%buffer B at 0.5–4 min; from 45%buffer B to 52%buffer B at 4–5 min; from 52%buffer B to 58%buffer B at 5–8 min; from 58%buffer B to 66%buffer B at 8–11 min; from 66%buffer B to 70%buffer B at11–14 min; from 70%buffer B to 75%buffer B at 14–18 min; from 75%buffer B to 97%buffer B at 18–21 min; 97%buffer B was held from 21–25 min; from 97%buffer B to 32%buffer B at 25–25.01 min; and 32%buffer B was held for 8 min. All the ions were acquired by non-targeted MRM transitions associated with their predicted retention time in a positive and negative mode switching fashion. ESI voltage was +5500 and -4500V in positive or negative mode, respectively. All the lipidomics . RAW files were processed on LipidSearch 4.0 (Thermo Fisher Scientific) for the lipid identification. The data were statistically analyzed with LINT-web according to the guideline (PMID: 34928054) . Example 13: Full Lifecycle Biological Clock and Its Impact in Health and Diseases

[0315] Aging research has primarily focused on adult aging clocks, leaving a critical gap in understanding biological aging across the full lifecycle, particularly during infancy and childhood. This example introduces LifeClock, a biological aging model that predicts biological age (BA) across all life stages using routine electronic health records (EHR) and laboratory test data. To enhance individualized predictions, digital twin technology was integrated, generating virtual patient representations from 27, 763, 357 longitudinal clinical visits across 11, 733, 272 individuals. The approach leverages EHRFormer, a transformer-based model, to analyze aging dynamics with high precision and develop accurate BA clocks spanning infancy to old age. The findings reveal distinct biological aging patterns across different life stages and their strong associations with disease risks and health outcomes. This work establishes a novel framework for studying aging and age-related diseases across the full lifecycle, utilizing widely available longitudinal EHR data to advance precision medicine and early disease detection.

[0316] The growing interest in aging mechanisms and interventions has driven interest in aging clocks-molecular markers that predict BA more precisely than CA which measures the passage of time. Unlike CA, which is static, BA reflects the efficiency of biological functions using genomic, epigenetic, and clinical markers. Genomic markers are fixed at birth, while epigenetic markers, such as DNA methylation and histone modifications, change with age. In theory, individuals of the same CA should exhibit similar rates of functional decline. However, genetic and environmental factors influence cellular aging, making some individuals appear biologically older or younger than their CA. This discrepancy, quantified as the difference between predicted BA and CA, is known as the age gap. Studies have shown that an increased age gap is associated with accelerated aging and heightened disease risk. For example, patients with increased brain age gaps often exhibit systemic aging features, such as sensory-motor decline and an older appearance. Accelerated aging is particularly evident in individuals with chronic diseases, suggesting that disease burden further drives biological aging. By developing reliable BA measures, aging clocks hold promise for extending healthspan and improving quality of life, a crucial goal as human life expectancy continues to rise.

[0317] Despite significant progress in adult aging clocks, the understanding of a full life-cycle clock, particularly during infancy and childhood, and its impact on health and disease remains limited. LifeClock is a full life-cycle BA clock leveraging 30 million electronic health records (EHR) , including laboratory test data, to predict BA across all life stages and assess its association with disease risk and survival outcomes. Physicians traditionally focus on EHR indicators that exceed reference ranges, yet normal values also contain valuable insights. Integrating longitudinal data-regardless of whether values are normal or abnormal-can help identify individual-specific setpoints and fluctuations, improving disease risk assessment and detection of critical aging transitions. While deep learning models have the capacity to extract such information, prior studies have largely focused on specific diseases within narrow age ranges. To further advance precision aging health research and translation, digital twin technology was implemented. Digital twins, which are virtual representations of individual patients, were generated using massive 27, 763, 357 longitudinal EHR data through EHRFormer, a transformer-based model. This approach enabled high-granularity modeling of aging processes, enhancing the understanding of the interplay between biological aging and disease risk.

[0318] Using EHR data, it was shown that there is a strong correlation between BA and CA. In addition, age-correlated changes in 184 clinical laboratory test results, vital sign indicators, and basic metadata were discovered and subjected to a cluster analysis using a standard single-cell cluster analysis method (FIG. 64) . The resultant 64 clusters displayed distinct age characteristics, and disease features which were subjected to aging assessment and disease risk predictions.

[0319] Using unsupervised learning, EHRFormer was trained to extract features from vast patient data spanning birth and childhood to adulthood and geriatric phase, with the goal of providing more accurate BA estimates than CA. The model integrates data across multiple organ systems-including blood, immune, liver, and kidney-while also accounting for gender differences in aging patterns. Model performance was evaluated by comparing predicted BA to CA using R2, Pearson correlation coefficient (PCC) , and mean absolute error (MAE) , demonstrating high accuracy, particularly in younger individuals with more uniform developmental trajectories. By focusing on biological aging, age gaps-significant divergences between BA and CA-were also identified as critical biomarkers for disease risk prediction and patient stratification. This is especially relevant in cases of accelerated aging, which correlates with increased disease risk across both younger and older populations. This example presents a novel framework for studying aging and age-related diseases across the full lifecycle, leveraging widely available and cost-effective EHR data to advance precision medicine in aging research.

[0320] Digital Twin Representation of Full Lifecycle Longitudinal EHRs from Multiple Cohorts.

[0321] Rich longitudinal real-world clinical data from electronic health records (EHRs) of millions of individuals across multiple cohorts serve as the cornerstone for constructing healthcare digital twins. However, real-world EHRs are inherently heterogeneous and prone to biases, necessitating careful processing and curation before they can be effectively applied to downstream clinical analyses. Disclosed herein is a framework for longitudinal electronic health record analysis using EHRFormer. FIG. 62A shows temporal trajectories of various physiological and laboratory measurements across longitudinal patient visits. Each line represents an individual patient, illustrating the heterogeneity and temporal dynamics of clinical parameters within the study population; FIG. 62B shows a computational architecture of EHRFormer. For each patient visit, the model extracts latent representations from both current and historical visits. An attention-based iterator integrates features across multiple visits, capturing temporal relationships and clinical progression patterns; FIG. 62C shows three primary applications of EHRFormer: (top) age prediction using latent representations; (middle) population clustering to identify high and low-risk groups for current and future disease manifestations; and (bottom) predictions of current and future disease states from latent representations along longitudinal visit EHR data.

[0322] A critical first step in constructing a digital twin is modeling clinical indicators and their interactions from longitudinal EHRs (FIG. 62A) . To achieve this, an input-output dual stochastic masking strategy was employed for digital twin encoding and reconstruction (FIG. 62B) . This approach captures complex feature interactions while imputing missing clinical indicators. Additionally, an adversarial module was introduced to remove undetected pattern information, which inherently contains subjective biases related to an individual's physiological state. The foundation model incorporated 184 carefully selected clinical indicators (Table 2) , including widely used, low-cost laboratory test markers and vital signs essential for understanding health and disease across the full lifecycle. EHR data from different cohorts (hospitals) often exhibit batch effects that can introduce biases. To address this, a cohort-agnostic adversarial training model was implemented to eliminate batch effects, ensuring that the learned representations capture universal patient phenotypes while minimizing data-collection biases and artifacts (FIG. 62C) . Table 2: Comprehensive clinical indicators

[0323] To ensure that each visit’s representation encapsulates not only an individual's current physiological state but also the temporal dynamics of past states and future trends, an autoregressive training approach was adopted. This method iteratively learns from longitudinal historical and current EHRs, creating a latent representation optimized to predict future EHR records. Over time, this process refines an individual's digital twin, capturing their evolving health trajectory. The resulting pretrained foundation model integrates multi-cohort longitudinal EHRs, enabling the generation of comprehensive and robust individual representations that can be used for missing clinical indicator imputation and cross-cohort batch effect correction.

[0324] The pretrained foundation model, trained on multi-cohort longitudinal EHRs spanning the full lifecycle, has multiple important applications. Leveraging its ability to analyze full lifecycle populations, the model was first applied to predict BA across all life stages. Additionally, it facilitates investigations into child development, adult aging, and age-related geriatric diseases.

[0325] A Blood Test-based Biological Aging Clock In A Full Lifecycle.

[0326] To investigate a full life-cycle aging clock, a BA clock was built, leveraging extensive EHRs, including longitudinal laboratory test data from 3, 037, 153 longitudinal clinical visits involving 1, 999, 897 healthy individuals. These data were sourced from the China Healthy Aging Investigation (CHAI) , a cohort of 4 large hospitals, and spanned a full lifecycle from birth to geriatrics periods (0-92 years old) . Using the EHRFormer, digital representations from each visit of healthy individuals were generated and a task-specific model was then developed to predict BA based on these digital representations (FIG. 62D) . The model demonstrated strong overall accuracy (low mean absolute error [MAE] ) and good explainability (high R2 and Pearson correlation coefficient [PCC] ) in the validation cohort (n=~1 million individuals, FIG. 63A) , indicating that predictions based on laboratory test data alone can reliably estimate BA.

[0327] However, two distinct aging patterns were identified: one from birth to 18 years (pediatric phase) and another from 18 years onward (adult phase) (FIG. 63A) . As anticipated, the profiles of laboratory test markers differed significantly between these two phases (FIG. 63B) . Consequently, separate prediction models were trained for each phase, which resulted in significantly improved accuracy in each respective phase (FIG. 63C, FIG. 63E) . Visual explanations of the feature contributions, generated using SHapley Additive exPlanations (SHAP) , revealed that key contributors to the pediatric clock included low aspartate aminotransferase (AST) levels, high creatinine (crea) levels, and high total protein (TP) levels (FIG. 63D) . For the adult clock, the most influential features included high urea levels, low albumin (ALB) levels, and high red cell distribution width (RDW) (FIG. 63F) . Notably, the top 20 markers for each clock were completely different (compare FIG. 63D to FIG. 63F) . Additionally, the performance of EHRFormer in predicting age separately for each gender was assessed and similar MAE and R2 values for both males and females were found (FIG. 66A, FIG. 66C, FIG. 66E, FIG. 66G) . However, the contribution of features varied slightly between genders within each phase of the aging clock (FIG. 66B, FIG. 66D, FIG. 66F, FIG. 66H) .

[0328] LifeClock Indicates Current And Future Disease Risks In Both Children And Adults.

[0329] EHRFormer-derived representations were applied for dimensionality reduction using PCA and UMAP, followed by Leiden clustering analysis. The results revealed that among healthy individuals, different CA groups could be clearly clustered, indicating that EHR data contain age-related information (FIG. 67) . Furthermore, data from different hospitals or cohorts were evenly distributed across clusters, particularly when separating those under 18 and over 18 of age, indicating the successful elimination of batch effects (FIG. 67) .

[0330] Given the link between BA and disease risks, dimensionality reduction was performed, followed by Leiden clustering analysis on the entire CHAI dataset, to examine whether individuals with higher age differences (over-aged) were more likely to develop diseases and potential associations between the clusters and diseases (FIG. 64) . The aging model, built on the EHRFormer framework and trained on healthy individuals, computes an age difference for each individual in CHAI dataset by quantifying deviations from the individual’s BA relative to same-CA peers through analysis of EHR profiles and adjusted for age, sex, and hospital. A total of 64 Leiden clusters were obtained from all EHR representations (FIG. 64B) . Adult EHRs were classified into two categories, average-aged (age difference within ±1 standard deviation (SD) ) and over-aged (age difference>3 SD) , and then the prevalence and incidence proportions of different diseases within each cluster were calculated. For most diseases, a significantly higher disease prevalence proportion was present in over-aged individuals when compared to average-aged individuals within the same cluster, which would be further increased in the future (FIG. 64C, FIG. 64D) . In addition, some diseases within certain clusters, such as hypoglycemia, may exhibit a higher incidence proportion in the future in the over-aged individuals, even these over-aged individuals may have not demonstrated a higher prevalence proportion (FIG. 64E) . In summary, these results suggest that the EHRFormer-based aging model not only demonstrates the present health status but also indicates future disease risk based on current EHR profiles.

[0331] Since clusters can serve as indicators of future disease risks, and the EHR representations for children (<18 years old, clusters 1-14) were well separated from those for adults (>18 years old, clusters 15-64) , the adjusted risk ratios (ARRs) were analyzed for incidence of pediatric and adult diseases separately within the respective child and adult clusters. For each identified cluster, ARRs were computed for incidence in reference to the remainder of the study population using log-binomial regression, and applying multivariate adjustment for age, sex, and hospital to minimize potential confounding from demographic factors and institutional variations. As a result, in clusters 1-14, by calculating disease ARRs for incidence using EHR data from individuals <12 years old, individuals within different clusters exhibited distinct tendencies to develop specific disease conditions. For instance, individuals in cluster 14 had 11.8 times, 6.5 times, 6.3 times, and 4.6 times higher risk of developing pituitary hyperfunction, obesity, malnutrition, and vitamin deficiency, respectively; individuals in cluster 12 had an 18 times higher risk of developing hernia; and individuals in cluster 13 had 12.5 times higher risk of developing viral meningitis (FIG. 64F) .

[0332] In parallel, in clusters 15-64, individuals in cluster 55 had a more than 20 times increased risk of vascular-related disorders, including hypotension (22.9 times) and renal failure (11.5 times) (FIG. 64E) . Similarly, diabetes showed increased ARRs for incidence by 5, 4.8, and 5.4 times in clusters 33, 41, and 49, respectively (FIG. 64E) . These findings demonstrate that the model can effectively identify individuals at high risk of developing diseases based on their longitudinal EHR data.

[0333] Fine-tuning EHRFormer For Individual Disease Risk Predictions.

[0334] Since the success of the EHRFormer-based BA prediction model in indicating current disease diagnosis and future disease predictions suggests that EHR data contain information beyond aging, encompassing overall health status and disease progression, the EHRFormer model could be fine-tuned with the introduction of disease labels to develop a model to hospitals. This approach enhances the model’s ability to diagnose current disease status and predict future disease, enabling a quantitative assessment of its predictive capabilities (FIG. 62F) . For each predicted disease, the population was stratified into high, middle, and low-risk cohorts based on model-generated probability scores, then quantified cumulative risk profiles across age groups. This stratification approach enables age-specific risk assessment and potentially facilitates the identification of critical intervention windows within the disease trajectory (FIG. 62E) .

[0335] The EHRFormer-based disease prediction model demonstrated strong current diagnostic performance across multiple diseases. Specifically, it achieved a high prediction accuracy in cardiovascular diseases (Atrial fibrillation AUC=0.95, CAD AUC=0.98, Hypertension AUC=0.95, Stroke AUC=0.96) , neurological disorders (Multiple Sclerosis AUC=0.96, Parkinson's AUC=0.94) , and systemic conditions (Osteoporosis AUC=0.96, Rheumatoid Arthritis AUC=0.96, Diabetes AUC=0.98) (FIG. 65A) . Additionally, the model effectively predicted future risks of these diseases (AUC≥ 0.8) (FIG. 65B) . For comparison, EHRFormer was also evaluated against other models such as RNN and XGBoost. The RNN, similar to the EHRFormer, accepts sequential data and follows the autoregressive paradigm but lacks an attention mechanism. XGBoost can also handle sequential data, yet it does not operate under the autoregressive framework. As shown in the table below, EHRFormer demonstrated superior performance compared to XGBoost and RNN in current disease diagnosis tasks across nine diseases. For example, in atrial fibrillation diagnosis, EHRFormer achieved an AUROC of 0.954, while XGBoost got 0.777 and RNN got 0.840; in coronary artery disease future prediction, EHRFormer's AUROC was 0.889, versus 0.665 for XGBoost and 0.717 for RNN. See table below for comparison of AUROC performance across methods for current disease diagnosis and future prediction tasks. This table presents a performance comparison of three methods (XGBoost, RNN, and EHRFormer) across nine diseases for both current diagnosis and future prediction tasks. Values represent the Area Under the Receiver Operating Characteristic curve (AUROC) for each model-disease combination.

[0336] To further validate its predictive abilities, the model was tested on an external validation cohort consisting of 3, 130, 332 longitudinal clinical visits from 2, 052, 508 individuals collected across 3 hospitals. See table below for demographic characteristics of the study cohorts. This table displays the distribution of patient records by age group and sex across ten different hospital cohorts used for model development and validation. Cohorts Hosp #1, Hosp #2, Hosp #3 and Hosp #4 were utilized for model development and internal validation. Cohorts from 3 hospitals served as external validation sets, primarily comprising normal individuals (healthy controls) . For model development and internal validation, data were aggregated from four hospital cohorts (Hosp #1, Hosp #2, Hosp #3 and Hosp #4) . This combined training and internal validation dataset comprised a total of 9, 680, 764 unique patients, accounting for 24,633, 025 records (visits) . The external validation set consisted of data from three additional cohorts, primarily representing normal individuals (healthy controls) . Based on the available summary data presented in the table, this external validation set included 2, 052, 508 unique patients and 3, 130, 332 visits.

[0337] The table below shows disease distribution in the study cohorts. The table presents the frequency counts of various diseases diagnosed within the internal development dataset and the external validation dataset. Diseases are categorized based on major physiological systems according to ICD10 code system. The counts represent the total individuals in the EHRdata for each system category.

[0338] The model was fine-tuned using each individual hospital’s EHR data and observed consistent good predictive performance across all cohorts (FIG. 68) . Therefore, by benchmarking against baseline models and analyzing the correlation between the number of visits and predicting accuracy, EHRFormer markedly improved current and future disease prediction performance by integrating the whole lifecycle of aging and disease information (FIG. 65C-FIG. 65D) .

[0339] The prediction model was also applied to predict the future disease risks in both pediatric populations (using EHR data before <12 years old) and adult populations (using EHR data >18 years old) . Using EHR data collected before the age of 12, the future common pediatric disease risks were predicted, achieving AUC ranging from 0.70 to 0.96. Similarly, using EHR data from individuals over 18 years old, future adult diseases were predicted with comparable accuracy (FIG. 69) . Furthermore, individuals under 10 years old were stratified into three risk-level groups based on their predicted probabilities: the highest 1 / 3 as the high-risk group, the middle 1 / 3 as the medium-risk group, and the bottom 1 / 3 as the low-risk group. Cumulative incidence plots provide a useful visual tool for comparing disease incidence over time among these groups, revealing large differences in future disease risk for various conditions, including obesity, meningitis, epilepsy, systemic lupus erythematosus (SLE) , asthma, and juvenile arthritis (FIG. 65E-FIG. 65J) . Similarly, the same stratification approach was applied to individuals under 40 years old, dividing them into three risk-level groups based on predicted probabilities. The cumulative incidence curves for these groups demonstrated substantial differences in future disease risk for hypertension, diabetes, anemia, hernia, appendicitis, and miscarriage after age 40 (FIG. 65K-FIG. 65P) . These results suggest that risk stratification based on early-life pediatric EHR data and early-adulthood EHR data can effectively reveal differential long-term disease risks.

[0340] This example highlights the potential of EHRFormer as a powerful tool for predicting BA across the full lifecycle, providing novel insights into aging processes and their association with disease risks. By leveraging a large longitudinal cohort of EHR data, the results reveal distinct biological aging clocks in the pediatric and adult phases and demonstrate how deviations from CA-captured as differences between CA and predicted BA-are linked to disease susceptibility. These insights offer a unique opportunity to enhance understanding of aging across the lifespan.

[0341] The results underscore the effectiveness of EHRFormer as a digital twin technology capable of capturing critical health information and providing a novel framework for leveraging widely available EHR data. The strong predictive performance of the digital twin-based aging clock highlights its potential for future applications in aging research. While traditional aging clocks estimate BA based on specific biomarkers, digital twins extend this capability by integrating diverse data sources, offering a dynamic and holistic approach to aging analysis. By continuously updating with new information, digital twins transform aging clocks from static estimators into adaptive, real-time systems. Incorporating wearable devices, cloud medical records, and environmental sensors can enable aging clocks to use the most current data, improving their adaptability and accuracy. The convergence of digital twins and aging clocks establishes a robust framework for advancing personalized healthcare strategies, promoting healthy aging, facilitating timely interventions, and mitigating aging-related decline. In this example, full lifespan aging clock, EHRFormer, offers greater accuracy in predicting disease risk compared to CA alone. The integration of longitudinal EHR data into biological aging models can drive the development of more precise aging biomarkers, enable prompt disease detection, and guide personalized treatments tailored to unique aging trajectories in diverse populations.

[0342] Data representation. EHR data were structured as chronological sequences of clinical visits for each patient. Each visit sequence is represented as S= {s0, s1, s2, …, sN} , where sidenotes the i-th clinical visit containing various laboratory test results and clinical measurements. For continuous variables, a uniform quantization approach was implemented where each feature x was discretized according to the formula resulting in integer values between 0 and dcont, with values exceeding the defined range truncated to the maximum boundary and missing values encoded as -1. This discretization strategy preserved the distributional characteristics of the original variables while transforming the infinite continuous value space into a finite discrete representation suitable for downstream modeling like categorical features. Temporal information was normalized by converting examination dates to days elapsed since the initial examination, with the first visit designated as day 0, establishing a standardized temporal reference frame across all patient records.

[0343] The EHR data was structured as chronological sequences of clinical visits for each patient. Each visit sequence was represented as S= {s0, s1, s2, …, sN} , where si denotes the i-th clinical visit containing laboratory test results and clinical measurements. Continuous variables were discretized using a uniform quantization approach:  This resulted in integer values between 0 and dcont, with values exceeding the defined range capped at the maximum and missing values encoded as -1. This strategy preserved the distribution of the original variables while converting continuous data into discrete representations suitable for categorical modeling. Temporal information was standardized by converting examination dates into days since the initial visit (day 0) , ensuring a consistent temporal reference across all patients.

[0344] EHRFormer architecture. EHRFormer, a hierarchical Transformer-based architecture, is specifically designed to process longitudinal EHR data. The model comprises three main components: an examination encoder, a temporal encoder, and task-specific decoders.

[0345] Examination encoder: the examination encoder transforms heterogeneous clinical measurements into a unified representation space suitable for downstream analysis. Each patient encounter consists of categorical variables (e.g., qualitative test results) and discretized continuous measurements (e.g., laboratory values, vital signs) represented as and respectively, where L denotes the sequence length of clinical visits. To embed these features, distinct mapping functions εcat:  and εcont:  were used, where each row corresponds to a token in the vocabulary and d represents the hidden dimension size. Both embedding matrices were jointly optimized during training, with a shared special vector reserved to represent missing examinations, enabling the model to distinguish between absent tests and actual clinical observations. To capture complex interdependencies between clinical variables, a bidirectional transformer-based architecture that processes the embedded features through multiple self-attention layers was implemented. This encoding process can be expressed as Fexam (Xcat, Xcont) = Encoder(Embed ( [Xcat; Xcont] ) ) , whereEncoder is the encoder of a Transformer model, yielding a rich contextual representation for each clinical variable.

[0346] Temporal encoder: to model disease progression and capture the longitudinal nature of patient trajectories, a temporal encoder that processes the encoded sequence of clinical visits was implemented. From the examination encoder output, the special token embeddings Evisitwere extracted as visit-level summaries for each encounter. The temporal relationships between visits were incorporated by using days elapsed since the initial encounter as position embeddings within an auto-regressive transformer architecture. This approach captures both short-term fluctuations and long-term trends in EHR records. The attention mechanism enables the model to weigh the relevance of historical encounters when assessing current patient state, a critical capability when analyzing disease trajectories with varied progression rates. To capture the temporal dynamics between clinical visits, the number of days elapsed since the patient's first visit was incorporated as position embeddings within the transformer architecture, enabling the model to learn time-dependent patterns in the longitudinal EHR data. The temporal encoding process follows Ftemp (Evisit, T) =Decoder(Evisit+TimeEmbed (T) ) , where Decoder is the decoder of a Transformer model.

[0347] Task-specific decoders: following established clinical prediction frameworks, a task-specific decoder with separate pathways for discrete outcomes (e.g., diagnosis prediction) and continuous measurements (e.g., biomarker estimation and BA prediction) was designed. Each pathway begins with dimensionality reduction via a projection layer, followed by ReLU activation to capture non-linear relationships, and concludes with task-specific output layers. This bifurcated architecture enables the model to simultaneously address diverse clinical prediction tasks while maintaining specialization for each prediction type. The classification pathway generates probabilities for binary clinical outcomes, while the regression pathway produces continuous estimates with clinical relevance. This multi-task framework facilitates knowledge transfer between related clinical objectives while respecting their distinct statistical properties by optimizing learnable parameters jointly across tasks.

[0348] Training procedure. The training procedure was split into two stages: pretraining the EHRFormer to learn the underlying relationships between multiple examination results and performing task-specific fine-tuning. The pretraining stage employs self-supervised learning on unlabeled EHR data to develop robust clinical representations, while the fine-tuning stage adapts these representations for specific prediction tasks. This two-stage approach enables the model to first develop a comprehensive understanding of clinical measurements and their relationships before specializing in downstream applications. By separating these stages, the model learns generalizable patterns from large-scale unlabeled data that may not be apparent in smaller labeled datasets used for specific clinical prediction tasks. Throughout both stages, specialized loss functions tailored to each learning objective were employed and strategies to mitigate dataset-specific biases were incorporated.

[0349] Mitigating missing bias and cohort bias through an adversarial method. Missing values in EHRs lead to incomplete or biased digital representations, as models may inadvertently learn to rely on the missing-state biases rather than the true clinical meaning of the feature expression values. Drawing inspiration examples from the concept of adversarial learning in other domains, a missingness discriminator was designed. A feature encoder is tasked with generating digital representations that encapsulate the clinical essence of the data. Concurrently, the missingness discriminator is designed to determine whether a specific feature value is missing or not.

[0350] A gradient reversal layer (GRL) between the feature encoder and the missingness discriminator was implemented. During backpropagation, the GRL inverts the gradient, compelling the feature encoder to produce representations that are independent of the missingness status. This forces the encoder to focus on encoding the inherent clinical significance of the features rather than being influenced by whether a value is present or absent. The loss function for the missingness discrimination task is defined as:  where zj, m is a binary indicator denoting whether the feature value of sample j has a missing status m (0 for present, 1 for missing) ,  is the predicted probability, M is the number of samples with potential missing values. By minimizing this loss, the model is encouraged to learn missing-invariant representations. This approach enables the creation of more robust digital representations that can better generalize across datasets with different missing value patterns, ultimately improving the accuracy and reliability of clinical outcome predictions.

[0351] A cohort bias, also known as a batch effect, is a significant challenge in multi-center data studies. Clinical data collected from different hospitals often exhibit systematic variations due to differences in patient demographics, practice patterns, and measurement protocols, potentially leading to biased models. Similar to the missingness discrimination, a cohort discriminator that aims to identify the cohort label of each sample was designed, while the encoder is forced to suppress cohort-specific information. The cohort classification loss is formulated as:  where yi, d is a binary indicator of whether sample i belongs to domain d,  is the predicted probability, N is the number of samples, and D is the number of domains (clinical cohorts) . This approach encourages the model to learn cohort-invariant representations that generalize across healthcare settings while maintaining predictive performance for clinical outcomes.

[0352] Pretraining step. A self-supervised pretraining approach with multiple complementary objectives was employed to enable the model to learn comprehensive representations of EHR data. 50%of the valid test results in the current examination was randomly masked as input and the model was trained to predict 50%masked values in the current examination and next examination. The masked language modeling (MLM) loss function is defined as:  where is the set of masked indices in examination event si and next examination after si,  is the total number of masked tokens across all examination events, vi, j is the true value of the j-th test in examination si and next examination after si, and is the predicted value. To quantify uncertainty in clinical data, a variational framework was incorporated with evidence lower bound (ELBO) maximization as balancing reconstruction fidelity against latent space regularization. Additionally, the domain adversarial loss and Lmissing were incorporated to promote cohort-invariant and missing-invariant representations. For the age regression task, the model was trained to predict patients' ages at each examination event using only clinical measurements (with all age-related information explicitly removed from inputs) to assess biological aging patterns, using the mean squared error loss function defined as where is the predicted age at examination event si and ai is the true age. Therefore, the final pretraining objective combined these components with appropriate weighting coefficients:  where the negative sign reflects the gradient reversal mechanism.

[0353] Fine-tuning step for disease state prediction tasks. Two distinct disease prediction tasks that reflect different clinical scenarios were implemented: current disease diagnosis and future disease prediction. For current disease diagnosis, the model was trained to identify the presence of specific diseases at the time of each examination event. The loss function is:  where is the predicted probability of disease d being present at the examination event si, ci, d is the true disease state (0 or 1) , D is the total number of diseases considered, and is the binary cross-entropy loss. For the future disease prediction task, a more nuanced labeling strategy was developed to identify patients at risk before disease manifestation. For a patient who never developed a specific disease throughout their recorded history, all examination events are labeled as negative (0) . For a patient who develops disease d at an examination event si, all preceding events {s0, s1, …, si-1} are labeled as positive (1) , indicating that the disease will develop in the future, while s_i and subsequent events are excluded from the loss calculation as they represent post-diagnosis data. The future state prediction loss is:  where is the predicted probability of future development of disease d based on the examination event si, fi, dis the true label (0 or 1) ,  is the set of valid (examination event, disease) pairs for future prediction, and is an indicator function that equals 1 if the pair (i, d) is valid and 0 otherwise. The total loss for disease prediction tasks is balancing the model's ability to recognize both existing conditions and identify patients at risk for future disease development.

[0354] Implementation details. The EHRFormer architecture was implemented using a combination of transformer models. Specifically, a 24-layer transformer encoder with a hidden dimension of 1024 was utilized as the examination encoder to process individual examination events, and a 12-layer autoregressive transformer decoder with a hidden dimension of 768 as the temporal encoder was used to capture longitudinal patterns across the sequence of examinations. This design leverages the bidirectional attention capabilities of the multi-headed self-attention mechanism for understanding relationships between clinical measurements within each examination while employing the causal masked attention mechanism to model the temporal progression of patient health. The model was implemented using PyTorch and trained using a two-stage approach. For the pretraining phase, the model for 200 epochs was trained using the Adam optimizer with a learning rate of 10-3 and a weight decay of 10-6. The subsequent fine-tuning phase for the downstream tasks was conducted for 50 epochs using the Adam optimizer with a reduced learning rate of 10-4 while maintaining the same weight decay of 10-6. Before training and evaluation, the dataset was randomly split into ten equal-sized subsets, with each fold using an 8: 1: 1 ratio for training, validation, and testing, respectively.

[0355] Age difference calculation. To quantify biological aging deviations, standardized age differences for each individual were calculated using the aging model. First, the non-linear relationship between predicted BA Ab and CA Ac was modeled using locally weighted scatterplot smoothing (LOWESS) with a bandwidth parameter of 2 / 3 via the statsmodels Python package (version 0.14.4) using EHR data from healthy individuals. The resulting function f (Ac) represents the expected BA for a given CA based on population trends. For each individual i, the raw age difference as Δi = Ab, i -f (Ac, i) was calculated, representing a deviation from the expected aging trajectory. These calculations were performed separately within each cohort to account for cohort-specific variations. Finally, standardized age differences were computed as zi = Δi / σ where σ represents the standard deviation of raw age differences within the model. Adjusted age differences were derived from EHRFormer models rigorously controlled for CA, sex, and hospital as covariates, ensuring robustness against confounding demographic and hospital-related biases.

[0356] Visualization of latent space and disease risk analysis. Visualization and clustering of EHRFormer-derived latent vectors were performed by first extracting laboratory and vital sign features, followed by principal component analysis (PCA) with 50 components. The resulting embeddings were processed using a neighbor graph approach (15 neighbors, euclidean metric) and visualized with uniform manifold approximation and projection (UMAP parameters: min_dist=0.3, spread=1.0, 2 components, spectral initialization) . Cluster identification was performed using the Leiden community detection algorithm, revealing distinct patient groups that correspond predominantly to pediatric and adult populations. For disease visualization, prevalence and incidence proportions were calculated per cluster, and each point was colored according to its corresponding cluster-specific disease prevalence or incidence proportion. Disease-cluster associations were quantified using adjusted risk ratios (ARRs) , calculated for each cluster in reference to the remainder of the study population using log-binomial regression models. These models incorporated multivariate adjustment for patient demographics (age and sex) and hospital to minimize potential confounding. These associations were visualized using a heatmap with ARR values truncated at a maximum of 4 to enhance interpretability while preserving meaningful signal contrast. The ARR was calculated using the statsmodels Python package (version 0.14.4) .

[0357] Statistical analysis. The performance of regression models for continuous value predictions were evaluated using Mean Absolute Error (MAE) , R-squared (R2) , and Pearson Correlation Coefficient (PCC) . Binary classification models were evaluated using Receiver Operating Characteristic (ROC) curves showing sensitivity versus 1-specificity, with the Area Under the Curve (AUC) reported along with 95%Confidence Intervals (CIs) . AUCs were calculated using the scikit-learn package (version 1.6.1) . Example 14: Full Spectrum Cancer Outcome Predictions Using Longitudinal  Multi-modal Health Records

[0358] Cancer care is often reactive, with diagnoses made only after symptoms or clinical suspicion arises, resulting in late-stage diagnosis, limited treatment options, and poorer prognosis. Post-diagnosis monitoring typically follows standardized protocols that overlook individual risk. This example provides a unified, multimodal transformer architecture designed to predict, stage, and monitor cancer across the continuum of care. Trained on EHRs and chest X-ray data from over 3 million patients in a multi-center discovery cohort, the model performs multimodal learning at an unprecedented scale. Validated on both external and internal cohorts, the model accurately predicts the onset of cancer up to three years before diagnosis, infers TNM (T (tumor size and extent) , N (lymph node involvement) , and M (metastasis or spread to distant sites) ) staging from data collected prior to diagnosis, and simulates tumor biomarkers and growth dynamics over time. In post-treatment settings, the model stratifies patients by recurrence-free survival (p < 0.001) , advancing beyond task-specific models toward a data-driven, multi-stage care system. Moreover, unsupervised analysis of model’s learned embeddings reveals a novel phenotypic landscape of cancer, identifying previously unrecognized subtype associations and mapping patient trajectories through health and disease. The example establishes a new paradigm for proactive, personalized, and data-driven cancer care.

[0359] Cancer remains one of the most formidable public health challenges, imposing a profound burden on global health systems and economies (Siegel, R. L., Giaquinto, A. N. &Jemal, A. Cancer statistics, 2024. CA Cancer J Clin 74, 12-49 (2024) ) . Clinically, cancer does not emerge abruptly but unfolds along a complex continuum, from early risk accumulation through diagnosis, treatment, and potential recurrence (Ellrott, K., et al. Classification of non-TCGA cancer samples to TCGA molecular subtypes using compact feature sets. Cancer Cell 43, 195-212. e111 (2025) ) . Despite significant advances in diagnostics and therapeutics, cancer care today remains largely reactive. Diagnoses are frequently made at advanced, less treatable stages; risk assessments rely on static, population-level metrics that fail to reflect individual health dynamics; and precise tumor staging depends on invasive procedures and high-resolution imaging, both unsuitable for routine monitoring (Orcutt, X., Chen, K., Mamtani, R., Long, Q. &Parikh, R. B. Evaluating generalizability of oncology trial results to real-world patients using machine learning-based trial emulations. Nat Med 31, 457-465 (2025) ) . Evaluating tumor burden through biomarkers or anatomical measurements often requires targeted, non-routine tests, and post-treatment surveillance typically follows standardized protocols, lacking the ability to adapt to individual risks. This fragmented, one-size-fits-all approach underscores a pressing need for scalable, dynamic, and non-invasive tools that can map the entire trajectory of cancer in a personalized and longitudinal manner.

[0360] The global digitization of healthcare has created vast repositories of real-world data, particularly through electronic health records (EHRs) , offering an unprecedented opportunity to transform cancer surveillance. Yet, current analytical approaches often treat these records as static snapshots, failing to capture the temporal physiological fluctuations that often precede clinical diagnosis. This example demonstrates that predictive signals lie not in isolated data points, but in the longitudinal sequence of routine clinical encounters. Moreover, this example demonstrates that this temporal physiological signature can be meaningfully enhanced with integration of structural information from routine imaging, such as chest X-rays (CXRs) . While often collected for non-oncologic indications, these images encode latent anatomical information that, when combined with temporal EHR data, can yield a high-resolution representation of the patient's evolving health.

[0361] This example shows that a multimodal transformer architecture trained on EHR data can map the full tumor lifecycle by integrating EHR and sequential radiographic images to: (i) detect current tumors and predict future incidence across major cancer subtypes; (ii) infer detailed clinical and TNM staging; (iii) estimate tumor biomarker levels and tumor size from routine inputs; and (iv) forecast recurrence risk after treatment with high fidelity. In one aspect, the approach integrates longitudinal EHRs with sequential imaging. In another aspect, the approach provides an end-to-end design that spans the entire cancer care continuum. In another aspect, the approach involves rigorous validation on large, real-world, multi-center cohort. Together, these advances establish a new paradigm for data-driven, opportunistic cancer surveillance, advancing the future of personalized, predictive, and proactive oncology. Results:

[0362] A large-scale, multimodal dataset enables the development and validation of the COMPASS predictive framework

[0363] COMPASS (China Oncology Multimodal Prediction and Surveillance Study) , a large-scale, predictive framework designed to model the full cancer trajectory using multi-center, multimodal, longitudinal patient data, was developed. The study incorporated over 4 million individuals across three independent cohorts: a discovery cohort (COMPASS-Main; n=2, 810, 742) , an internal validation cohort from the same healthcare system (COMPASS-Replication Cohort; n=1, 230, 468) , and an external validation cohort (UK Biobank (UKB) ; n=44, 275) (FIG. 70A) .

[0364] To capture the evolving physiological and structural signatures of cancer, longitudinal EHRs and CXRs were integrated, spanning pre-diagnosis to post-treatment follow-up (FIG. 70B) . The COMPASS dataset was organized into three primary modalities for model input (FIG. 70C) : (1) a comprehensive panel of 150 laboratory features (e.g., complete blood counts, blood chemistry, coagulation profiles) ; (2) a set of 9 core vital signs, (e.g. heart rate, blood pressure, BMI) ; and (3) quantitative features derived from CXRs using image encoding techniques to assess latent tissue characteristics. Model performance improved with the inclusion of more historical visits (as the prediction period extends, the model’s AUC and F1 scores show significant improvement; data not shown for plot showing performance of cancer prediction models for different time periods 1 year, 3 years, 5 years, and 10 years) and greater temporal depth (input sequences of varying temporal lengths were constructed from the test set, results demonstrating that increasing the length of time series data leads to improved classification performance by the model; data not shown for the impact of time series length on model performance) , underscoring the value of longitudinal input. The integrated data were analyzed using the model, which demonstrated robust predictive performance across key oncology tasks, including cancer screening, tumor staging, patient risk stratification, and prognostic monitoring (FIG. 70D) .

[0365] The model accurately diagnoses prevalent cancers and predicts long-term incidence across major subtypes

[0366] The COMPASS dataset provided a large-scale foundation for model development, encompassing longitudinal data from millions of patient visits across major malignancies, including lung, breast, and thyroid cancer (FIG. 71A) . The data captures a comprehensive timeline extending several years prior to the formal diagnosis, enabling the model to learn predictive temporal patterns. The model successfully transforms this high-dimensional data into clinically meaningful representations, as visualized by the UMAP plots in FIG. 71B. These plots show that the model learns to map patients onto an embedding space where individuals with high cancer risk are clearly separated. Furthermore, these embeddings contain distinct signatures that correspond to specific cancer subtypes.

[0367] The model demonstrated excellent discriminatory performance in both detecting current (prevalent) cancer and predicting future incidence. For cancer diagnosis, the model achieved a high Area Under the Receiver Operating Characteristic Curve (AUC) of 0.918 in the COMPASS-Main cohort, with its strong performance robustly validated in the COMPASS-Replication cohort (AUC = 0.904) and the external UK Biobank cohort (AUC = 0.909) (FIG. 71C, top row) . The model also showed strong capability in forecasting long-term cancer risk, accurately predicting future incidence at 1, 3, and 5 years. In the main cohort, the AUC for 3-year prediction was 0.886, highlighting its reliability for prospective surveillance (FIG. 71C, bottom row) .

[0368] This high level of accuracy extended across a wide range of major cancer subtypes. As shown in FIG. 71D, the model maintained strong diagnostic and predictive performance for diverse cancers such as prostate, lung, colon, and breast cancer. For most subtypes, diagnostic AUCs surpassed 0.90, while predictive AUCs consistently exceeded 0.80. Together, these results highlight The model’s capacity to detect both general oncologic signals and resolve them into subtype-specific predictions, enabling both broad and targeted cancer surveillance.

[0369] The model precisely delineates tumor stage, including fine-grained TNM classification

[0370] Beyond predicting cancer incidence, the model demonstrated strong performance in tumor staging using EHR data before the predicted diagnosis. It accurately distinguished early-stage (I / II) from late-stage (III / IV) cancers, achieving high AUCs across all cancers (FIG. 72A) .

[0371] The model’s performance in more granular TNM classification was assessed and high pan-cancer accuracy in predicting each individual component of TNM staging was observed (FIG. 72B) . The model achieved high AUCs for predicting tumor size and extent (T-stage) : 0.935 (T1) , 0.906 (T2) , 0.927 (T3) , and 0.930 (T4) . Comparable accuracy was observed for predicting nodal involvement (N-stage, pan-cancer AUC) and distant metastasis (M-stage, pan-cancer AUC) , the latter being a key prognostic factor typically requiring advanced imaging (FIG. 72C) .

[0372] Importantly, the model’s performance remained robust across cancer types, including digestive (esophageal, colon) , respiratory, and endocrine cancers. Unlike conventional staging methods relying on invasive tests and imaging, the model infers stages prospectively using routine, pre-diagnostic clinical data, to not just evaluate cancer risk but also its potential anatomical extent and aggressiveness.

[0373] The Model Predicts Treatment-Specific Response to Guide Personalized Therapy

[0374] A critical challenge in oncology is predicting how an individual patient will respond to a specific treatment. This example demonstrates that the model can prospectively distinguish between patients who are likely to be sensitive or resistant to various targeted and endocrine therapies. By analyzing pre-treatment longitudinal EHR data, the model effectively stratifies patients, showing a clear separation between actual resistant and sensitive individuals in a t-SNE projection (FIG. 73A) .

[0375] To further validate this capability, an AI-based causal inference model was developed to forecast treatment outcomes. This model accurately predicts tumor size dynamics following therapeutic intervention across multiple cancer types. For instance, in non-small cell lung cancer (NSCLC) patients treated with EGFR inhibitors, the model’s predictions of tumor response closely mirrored actual clinical trajectories, achieving a high correlation (R2=0.858, p<0.001) (FIG. 73B) . Similar high-fidelity predictions were observed for breast cancer (BRC) patients undergoing endocrine therapy (R2=0.782, p<0.001) , thyroid cancer (TC) with BRAF inhibitors (R2=0.721, p<0.001) , prostate cancer (PCa) with androgen deprivation therapy (ADT) based on PSA changes (R2=0.683, p<0.001) , and renal cancer (RC) with anti-VEGF therapy (R2=0.734, p<0.001) (FIG. 73C-73F) .

[0376] These findings highlight the model's potential as a powerful decision-support tool, enabling clinicians to perform in silico evaluation of treatment efficacy. By forecasting patient-specific drug sensitivity, it paves the way for personalizing therapeutic strategies, thereby maximizing the likelihood of positive outcomes while avoiding ineffective treatments.

[0377] The model enables comprehensive prognostic surveillance by modeling tumor dynamics and forecasting clinical outcomes.

[0378] The model effectively transformed routine health data into actionable insights for non-invasive cancer monitoring and prognosis. Leveraging routine laboratory tests, vitals, and CXRs, the model accurately estimated tumor biomarker levels, demonstrating strong correlations with ground-truth values across diverse malignancies (FIG. 74A) . Strong correlations across multiple biomarkers, including CA199 (R2=0.14) , NSE (R2=0.32) , CA242 (R2=0.21) , CA153 (R2=0.20) , ProGRP (R2=0.19) , and CYFRA211 (R2=0.24) , confirm its ability to capture systemic tumor signatures. Moreover, the model effectively tracked tumor dynamics, classifying tumor growth, stability, or shrinkage over a 24-week period with high accuracy (0.878) , precision (0.89) , and recall (0.88) . This offers a practical alternative to invasive and costly assessment treatment response by minimizing the amount of testing needed (FIG. 74B) .

[0379] Additionally, the model generated individualized risk scores to stratify patients by likelihood of cancer progression post-treatment, showing robust predictive performance for progression-free survival (log-rank p < 0.001; FIG. 74C) . This capability generalized across a wide spectrum of cancer types, including esophageal (p=1.59e-14) , liver (p=6.30e-27) , colon (p=1.81e-68) and breast cancer (p=3.54e-88) .

[0380] This example provides a unified framework for clinical predictive modelling across the full spectrum of cancer care. Unlike traditional models, which are typically fragmented and task-specific, the approach provided in this example integrates multimodal data and performs diverse predictive tasks within a single, longitudinally informed architecture to address various clinical questions such as screening, staging, and recurrence prediction. In one aspect, the approach leverages a multimodal transformer to synergistically integrate longitudinal, irregularly sampled laboratory data with sequential, opportunistic radiographic imaging, enabling a holistic patient representation that serves as a foundational substrate for all downstream tasks. This unified approach not only streamlines clinical workflows but also fosters transfer learning across tasks, allowing the model to learn more generalizable and robust features. The success of the model signals a paradigm shift from narrow, specialized tools to comprehensive, end-to-end systems capable of capturing the systemic complexity of disease.

[0381] Beyond prediction, the model can be repurposed as a tool for scientific discovery through unsupervised clustering of patient embeddings, revealing a data-driven "phenotypic landscape" in which patients are organized by intrinsic biological similarity, independent of pre-defined labels. Specific cancer subtypes showed non-random, organ-specific distributions across this landscape, confirming the model’s sensitivity to disease biology. More importantly, it uncovered unexpected relationships, such as the co-localization of breast, ovarian, and cervical cancers, suggesting a shared, systemic signature that transcends anatomical origin. The longitudinal structure of the COMPASS dataset further enabled dynamic tracking of patient trajectories across this landscape. Trajectories that map the transition from health to disease onset, followed by staging evolution, treatment response, and recurrence, were observed. These dynamic transitions provide a mechanistic framework to study cancer as a temporally evolving process and promote identification of novel disease subtypes through disease evolution.

[0382] The model’s strong performance using only routine clinical data and images demonstrates that standard, everyday tests contain critical but often overlooked information about cancer. Acting as a sophisticated decoder, the model translates sparse, irregular, and noisy data streams into coherent representations of patient health states. Its ability to prospectively predict fine-grained TNM staging using only pre-diagnostic data implies that early tumor growth leaves detectible, systemic signatures in routine clinical measurements long before clinical detection. Similarly, the model’s accurate estimation of tumor biomarker levels and growth dynamics from non-invasive data reflects a capacity to infer the underlying biological state of malignancy, what can be referred to as the tumor's “systemic echo" . These findings within a systems biology framework reinforce the concept that cancer is a systemic disease with trajectories that can be non-invasively monitored through longitudinal, multimodal patient data.

[0383] A particularly striking finding is the unanticipated power of CXRs, a modality traditionally limited to thoracic pathology, in pan-cancer predictions. This example shows that chest radiographs functions can capture subtle systemic health signals that escape human visual inspection. These can include alterations to the pulmonary vasculature, shifts in bone density, changes in soft tissue composition, which can indicate systemic inflammation, hormonal dysregulation, hemodynamic shifts, or early cachexia, all of which are commonly associated with malignancy. The deep learning model quantifies these latent patterns as high-dimensional features, reframing CXRs as a low-cost, accessible tool for systemic health profiling.

[0384] The versatility of the framework provided herein underscores the generalizability of the patient embeddings learned by the model. The same pretrained model successfully addressed diverse downstream tasks, spanning screening, staging, prognostic forecasting and unsupervised clustering, demonstrating that it encodes a foundational representation of patient health status rather than task-specific shortcuts. This inherent scalability positions the framework as a platform technology, with potential applications to extend beyond oncology with minimal modification. Incorporating additional data modalities, such as genomics, digital pathology, or wearable sensor data, may further enrich patient representations, for instance, creating personalized “digital twins” for precision medicine.

[0385] In conclusion, this example describes and validates a unified, multimodal framework capable of comprehensively mapping the full continuum of cancer care using only routine clinical data. By integrating longitudinal laboratory results and opportunistic CXRs, the model in this example decodes the complex, dynamic patterns of patient health, not only forecasting future cancer incidence years in advance but also prospectively inferring tumor aggressiveness and monitoring disease progression non-invasively over time. Beyond prediction, the approach uncovers novel, data-driven patient subtypes and elucidates temporal trajectories of cancer evolution. Validated at scale through the COMPASS framework, the model represents a paradigm shift toward proactive, predictive, and personalized oncological care. Methods

[0386] Study Populations.

[0387] The China Oncology Multimodal Prediction and Surveillance Study (COMPASS) , as a project of the International Consortium of Digital Twin in Medicine, is an ongoing study to develop and validate artificial intelligence models for comprehensive tumor prediction. The study was registered at clinicaltrial. gov (NCT06791473) . The project leverages longitudinal electronic health records (EHRs) and medical imaging to stratify cancer risk, predict incidence, classify tumor types and stages, and assess prognosis.

[0388] Multimodal Data representation.

[0389] The patient's clinical history was structured as a chronological sequence of...

Claims

1.A method of generating a trained large language model (LLM) , comprising:(a) pretraining an LLM using longitudinal multimodal data of a plurality of subjects, wherein for each subject, the longitudinal multimodal data comprise data from a chronological sequence of clinical visits of the subject, wherein the LLM comprises modality-specific alignment encoders, a decoder for cross-modal latent space prediction, and an autoregressive module, and wherein the pretraining comprises:(i) adversarially aligning the longitudinal multimodal data using the modality-specific alignment encoders, wherein each modality-specific alignment encoder embeds data of the corresponding modality into a unified, low-dimensional latent space, and wherein a modality discriminator identifies embedding sources in the latent space,(ii) using the decoder to predict a representation of a target modality in the latent space, wherein data from one or more clinical visits of a subject are missing in the target modality, and wherein the prediction imputes the missing data based on the subject’s historical data of the target modality and the subject’s historical data of one or more source modalities,thereby generating modality-invariant representations of the subjects in the latent space; and(b) finetuning the pretrained LLM by feeding the modality-invariant representations into the autoregressive module to model each subject's historical trajectory, thereby generating the trained LLM.2.The method of claim 1, wherein the multimodal data comprise laboratory tests, imaging scans, and / or omics profiles.3.The method of claim 1 or claim 2, further comprising feeding the output of the autoregressive module into a task-specific prediction head for a target application.4.The method of claim 3, wherein the target application is prediction of a systemic disease, prediction of a chronic disease, prediction of an ocular disease such as myopia, and / or identification of aging trajectories of the subjects.5.The method of claim 3 or claim 4, wherein the target application is prediction of a disease selected from the group consisting of atrial fibrillation, coronary artery disease, diabetes, hypertension, ischemic stroke, multiple sclerosis, osteoporosis, Parkinson’s disease, and rheumatoid arthritis.6.The method of any one of claims 1-5, wherein the longitudinal multimodal data comprise fundus images, routine laboratory test results.7.The method of any one of claims 1-6, wherein the longitudinal multimodal data comprise electronic health record (EHR) data, vital signs, imaging data, or any combination thereof.8.The method of claim 7, wherein the routine laboratory results comprise results of any one or more of the biomarkers listed in Table 1 or Table 2, optionally wherein the routine laboratory results comprise complete blood count (CBC) , blood chemistry, coagulation, or any combination thereof.9.The method of claim 7 or claim 8, wherein the vital signs comprise heart rate, blood pressure, body mass index (BMI) , or any combination thereof.10.The method of any one of claims 7-9, wherein the imaging data comprise chest X-rays (CXRs) .11.A method of generating a medical diagnosis, prediction, and / or prognosis for a subject, the method comprising receiving a natural-language prompt for obtaining the medical diagnosis, prediction, and / or prognosis and a set of data related to the subject, and generating the medical diagnosis, prediction, and / or prognosis by inputting the prompt and the set of data in the trained LLM of any one of claims 1-10.12.The method of claim 11, wherein the medical diagnosis, prediction, and / or prognosis comprises stratifying disease risk, predicting disease incidence, classifying disease types and / or stages, assessing prognosis, or any combination thereof.13.A method of generating a trained large language model (LLM) , comprising:(a) pretraining an LLM using longitudinal multimodal data of a plurality of subjects,wherein the LLM comprises an examination encoder, a temporal embedding, and task-specific decoder heads,wherein for each subject, the longitudinal multimodal data comprise longitudinal electronic health record (EHR) data from a chronological sequence of clinical visits of the subject,wherein the examination encoder generates a contextualized representation of subject data from each clinical visit, and the temporal embedding captures temporal relationships between clinical visits from the output of the examination encoder to generate a subject-level longitudinal representation for each subject; and(b) finetuning the pretrained LLM by adapting the subject-level longitudinal representations to distinct tasks using the task-specific decoder heads, wherein the distinct tasks comprise first occurrence disease diagnosis and future disease prediction,thereby generating the trained LLM.14.The method of claim 13, wherein the plurality of subjects comprise cancer subjects and the longitudinal multimodal data comprise routine laboratory results, vital signs, routine imaging data, or any combination thereof.15.The method of claim 14, wherein the routine laboratory results comprise results of any one or more of the biomarkers listed in Table 1 or Table 2, optionally wherein the routine laboratory results comprise complete blood count (CBC) , blood chemistry, coagulation, or any combination thereof.16.The method of claim 14 or claim 15, wherein the vital signs comprise heart rate, blood pressure, body mass index (BMI) , or any combination thereof.17.The method of any one of claims 13-16, wherein the routine imaging data comprise chest X-rays (CXRs) .18.The method of any one of claims 13-17, wherein the longitudinal multimodal data comprise alterations to the pulmonary vasculature, shifts in bone density, changes in soft tissue composition, systemic inflammation, hormonal dysregulation, hemodynamic shifts, early cachexia, or any combination thereof.19.The method of any one of claims 13-18, wherein the longitudinal multimodal data comprise categorical variables and / or continuous variables.20.The method of claim 19, wherein the continuous variables are discretized to preserve their distributional characteristics.21.The method of any one of claims 13-20, wherein for each subject, the longitudinal multimodal data comprise tabular clinical variables and image data.22.The method of any one of claims 13-21, wherein for each modality of subject data, the examination encoder simultaneously captures both a value distribution and a semantic meaning of the modality of subject data.23.The method of any one of claims 13-22, wherein the temporal embedding comprises a linear positional embedding to learn time-dependent patterns in the longitudinal EHR data.24.The method of any one of claims 13-23, wherein the temporal embedding is autoregressive and comprises causal masking to ensure unidirectional information flow in the autoregressive process.25.The method of any one of claims 13-24, wherein each of the task-specific decoder heads comprises a separate pathway applying a projection layer followed by Rectified Linear Unit (ReLU) activation.26.The method of any one of claims 13-25, wherein each of the task-specific decoder heads comprises causal masking to prevent information from future clinical visits from influencing prediction at a given clinical visit.27.The method of any one of claims 13-26, wherein the LLM further comprises a missingness discriminator that determines whether a value of a particular feature is missing or not.28.The method of claim 27, wherein the LLM further comprises a gradient reversal layer (GRL) between the missingness discriminator and the examination encoder, wherein the GRL inverts the gradient during backpropagation, compelling the examination encoder to produce a representation that is independent of the missingness status of the particular feature.29.The method of any one of claims 13-28, wherein the LLM further comprises a cohort discriminator that identifies the cohort label of each subject and forces the examination encoder to suppress cohort-specific information.30.The method of any one of claims 13-29, wherein the pretraining is a self-supervised pretraining.31.A method of generating a medical diagnosis, prediction, and / or prognosis for a subject, the method comprising receiving a natural-language prompt for obtaining the medical diagnosis, prediction, and / or prognosis and a set of data related to the subject, and generating the medical diagnosis, prediction, and / or prognosis by inputting the prompt and the set of data in the trained LLM of any one of claims 13-30.32.The method of claim 31, wherein the medical diagnosis, prediction, and / or prognosis comprises stratifying disease risk, predicting disease incidence, classifying disease types and / or stages, assessing prognosis, or any combination thereof.33.The method of claim 32, wherein the medical diagnosis, prediction, and / or prognosis is for a pan-cancer diagnosis, prediction, and / or prognosis.34.The method of claim 32, wherein the medical diagnosis, prediction, and / or prognosis is for one or more specific cancers.35.The method of any one of claims 31-34, wherein the medical diagnosis, prediction, and / or prognosis is for: prediction of an organ-specific age, predicting of female aging, prediction of a hospital-acquired infection, prediction of a dialysis time series, prediction of diabetes and complications, prediction of pregnancy and child outcome, prediction of heart failure, prediction of myopia, prediction of vision in anti-VEGF treatment, or any combination thereof.36.The method of any one of claims 31-35, wherein the set of data related to the subject comprises longitudinal EHR data of the subject.37.The method of any one of claims 31-36, wherein at least part of the set of data related to the subject is not collected in association with generating a medical diagnosis, prediction, and / or prognosis of an oncologic indication.38.The method of any one of claims 31-37, wherein the set of data related to the subject comprises routine laboratory results, vital signs, and routine imaging data.39.The method of any one of claims 31-38, wherein the set of data related to the subject comprises: results of any one or more of the biomarkers listed in Table 1 or Table 2, optionally complete blood count (CBC) , blood chemistry, coagulation, or any combination thereof; heart rate, blood pressure, body mass index (BMI) , or any combination thereof; and chest X-rays (CXRs) .40.A method of generating a trained large language model (LLM) , comprising:(a) using an LLM to categorize longitudinal electronic health record (EHR) data of a plurality of subjects into structured data and unstructured data, wherein for each subject, the longitudinal EHR data are from a chronological sequence of clinical visits of the subject;(b) using the LLM to extract biomedical concepts from both the structured data and the unstructured data as medically relevant tokens;(c) using the LLM to process the medically relevant tokens in chronological order to maintain the temporal coherence of subject health data, thereby predicting the next token in a patient’s timeline using temporal patterns of preceding tokens; and(d) using the LLM to measure and maximize the likelihood of predicting the next token correctly, thereby generating the trained LLM.41.The method of claim 40, wherein the method comprises converting free-text information from the longitudinal EHR data into structured data.42.The method of claim 40 or claim 41, wherein the structured data comprise demographics, laboratory results, medication lists, and / or diagnosis codes, and the unstructured data comprise doctor’s notes, imaging reports, and / or clinical correspondence.43.The method of any one of claims 40-42, wherein the longitudinal EHR data comprise routine laboratory results, vital signs, imaging data, or any combination thereof, and the trained LLM is used for early cancer detection, wherein the routine laboratory results, vital signs, and / or imaging data are not collected in association with generating a medical diagnosis, prediction, and / or prognosis of an oncologic indication.44.The method of claim 43, wherein the routine laboratory results comprise results of any one or more of the biomarkers listed in Table 1 or Table 2, optionally wherein the routine laboratory results comprise complete blood count (CBC) , blood chemistry, coagulation, or any combination thereof; wherein the vital signs comprise heart rate, blood pressure, body mass index (BMI) , or any combination thereof; and / or wherein the imaging data comprise chest X-rays (CXRs) .45.The method of any one of claims 40-44, comprising receiving a natural-language prompt for obtaining the medical diagnosis, prediction, and / or prognosis of an oncologic indication and a set of data related to an individual, and generating the medical diagnosis, prediction, and / or prognosis by inputting the prompt and the set of data in the trained LLM, wherein the set of data comprise routine laboratory results, vital signs, imaging data, or any combination thereof, which are not collected in association with cancer diagnosis, prediction, and / or prognosis.46.The method of any one of claims 40-44, comprising receiving a natural-language prompt for obtaining the medical diagnosis, prediction, and / or prognosis of hospital acquired infection, community acquired infection, and / or non-infection, and a set of data related to an individual, and generating the medical diagnosis, prediction, and / or prognosis by inputting the prompt and the set of data in the trained LLM.47.The method of any one of claims 40-44, comprising receiving a natural-language prompt for obtaining the medical diagnosis, prediction, and / or prognosis of organ specific aging, and a set of data related to an individual, and generating the medical diagnosis, prediction, and / or prognosis by inputting the prompt and the set of data in the trained LLM.48.The method of claim 47, wherein accelerated organ specific aging is associated with a disease or condition.49.The method of claim 47 or claim 48, wherein the organ specific aging is ovarian aging.50.The method of any one of claims 40-44, comprising receiving a natural-language prompt for obtaining the medical diagnosis, prediction, and / or prognosis of nosocomial infection, and a set of data related to an individual, and generating the medical diagnosis, prediction, and / or prognosis by inputting the prompt and the set of data in the trained LLM.51.The method of any one of claims 40-44, comprising receiving a natural-language prompt for obtaining dialysis time series prediction, and a set of data related to an individual, and generating the dialysis time series prediction by inputting the prompt and the set of data in the trained LLM.52.The method of any one of claims 40-44, comprising receiving a natural-language prompt for obtaining the medical diagnosis, prediction, and / or prognosis of diabetes, and a set of data related to an individual, and generating the medical diagnosis, prediction, and / or prognosis by inputting the prompt and the set of data in the trained LLM.53.The method of any one of claims 40-44, comprising receiving a natural-language prompt for obtaining pregnancy and / or child outcome prediction, and a set of data related to an individual, and generating the pregnancy and / or child outcome prediction by inputting the prompt and the set of data in the trained LLM.54.The method of any one of claims 40-44, comprising receiving a natural-language prompt for obtaining heart failure prediction, and a set of data related to an individual, and generating the heart failure prediction by inputting the prompt and the set of data in the trained LLM.55.The method of any one of claims 40-44, comprising receiving a natural-language prompt for obtaining myopia prediction, and a set of data related to an individual, and generating the myopia prediction by inputting the prompt and the set of data in the trained LLM.56.The method of any one of claims 40-44, comprising receiving a natural-language prompt for obtaining vision prediction in patients undergoing or to undergo anti-VEGF therapy, and a set of data related to an individual, and generating the vision prediction by inputting the prompt and the set of data in the trained LLM.57.The method of any one of claims 40-44, comprising receiving a natural-language prompt for obtaining biological age prediction, and a set of data related to an individual, and generating the biological age prediction by inputting the prompt and the set of data in the trained LLM.58.The method of any one of claims 40-44, comprising receiving a natural-language prompt for obtaining female reproductive age prediction, and a set of data related to an individual, and generating the female reproductive age prediction by inputting the prompt and the set of data in the trained LLM.59.The method of any one of claims 40-58, comprising using multimodal longitudinal data of the plurality of subjects to train the LLM.60.A method of generating a trained large language model (LLM) , comprising:(a) using longitudinal electronic health record (EHR) data of a plurality of subjects to train an LLM, wherein for each subject, the longitudinal EHR data are from a chronological sequence of clinical visits of the subject and comprise complete blood count (CBC) data;(b) extracting biomedical concepts from the longitudinal EHR data as medically relevant tokens;(c) processing the medically relevant tokens in chronological order to maintain the temporal coherence of subject health data, thereby predicting the next token in a patient’s timeline using temporal patterns of preceding tokens; and(d) measuring and maximizing the likelihood of predicting the next token correctly, thereby generating the trained LLM.61.The method of claim 60, wherein the CBC data comprise hematocrit (HCT) , hemoglobin (HGB) , mean corpuscular hemoglobin (MCH) , mean corpuscular hemoglobin concentration (MCHC) , mean platelet volume (MPV) , platelet count (PLT) , red cell count (RBC) , red cell distribution width (RDW) , white cell count (WBC) , or any combination thereof.62.The method of claim 60 or claim 61, comprising receiving a natural-language prompt for obtaining the medical diagnosis, prediction, and / or prognosis of a disease or condition, and a set of data related to an individual, and generating the medical diagnosis, prediction, and / or prognosis by inputting the prompt and the set of data in the trained LLM.63.The method of claim 62, wherein the data related to the individual comprise longitudinal complete blood count (CBC) data from a chronological sequence of clinical visits of the individual, and the disease or condition is all-cause mortality or a disease such as heart attack, stroke, diabetes, kidney disease, osteoporosis, thyroid dysfunction, iron deficiency, myeloproliferative neoplasm, a cancer, or any combination thereof.64.A system comprising:at least one hardware processor; andone or more software modules configured to, when executed by the at least one hardware processor, perform the method of any one of claims 1-63.65.A non-transitory computer-readable medium having instructions stored thereon, wherein the instructions, when executed by a processor, cause the processor to perform the method of any one of claims 1-63.66.A system comprising:at least one hardware processor;non-transitory computer-readable medium coupled to at least one hardware processor, optionally wherein the coupling is over a network; andinstructions stored in the non-transitory computer-readable medium, wherein the instructions when implemented by the processor, configure the system to perform the method of any one of claims 1-63.

Citation Information

Patent Citations

  • Multi-modal medical missing data completion method and device based on data relevance mining

    CN116795826A

  • Method and system for predicting death rate of patient based on multi-modal missing data

    CN117954090A

  • Multi-modal information fusion method based on large language model

    CN117992908A

  • Medical data fusion method and system based on multi-modal feature filling

    CN118364423A

  • System and method for medical data governance using large language models

    US12001464B1

Cited By

  • Breast focus progress trend prediction method based on multimode data fusion

    CN121885161A