Machine learning for multi-state disease models
The machine learning method using neural networks for multi-state disease models addresses the limitations of linear assumptions in Cox PH models by providing accurate, individualized disease progression predictions through transition-specific probability distributions, enhancing the precision of disease prognosis.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- DASSAULT SYSTEMES SA
- Filing Date
- 2021-12-21
- Publication Date
- 2026-05-20
AI Technical Summary
Existing methods for predicting disease progression in multi-state models, such as the Cox PH model, rely on strong statistical assumptions and are limited by linear relationships between covariates and event risks, leading to inaccuracies, especially in big data applications.
A machine learning method using neural networks to output transition-specific probabilities for each interval in a multi-state disease model, incorporating covariate-shared and transition-specific subnetworks, with a softmax layer and a loss function that penalizes weight differences across time intervals, allowing for nonlinear processing of covariates.
This approach provides accurate, individualized predictions of disease progression by modeling true transitions between states, reducing approximation errors and improving accuracy compared to conventional discrete-time methods.
Smart Images

Figure 0007862948000143 
Figure 0007862948000144 
Figure 0007862948000145
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of biometrics, and more specifically, to learning a machine learning function configured to output a distribution of transition-specific probabilities based on input covariates representing a patient's medical characteristics with respect to a multi-state model of a disease having states and transitions between states. The present disclosure also relates to methods, data structures, and apparatuses related thereto.
Background Art
[0002] The prognosis of a disease is very important when a doctor makes a medical decision and includes a special algorithm for estimating the risk of a patient. Event history analysis, also known as survival analysis, aims to predict the time until one or more future events of interest occur and is used in a plurality of fields including healthcare in the context of disease prognosis. In particular, survival analysis is very frequently used in healthcare to model the survival outcome of a patient and evaluate the treatment effect. In clinical practice, a clinician may be interested in not only unique or composite events but also the complete progression of a disease. Therefore, a multi-state approach has been developed to generalize survival analysis in which multiple events may occur sequentially over time (see Reference 1).
[0003] A disease-death model is a specific multi-state model consisting of three states: "healthy", "recurrence" or "affected", and "death". This is the most frequently used structure for tracking the progression of diseases in cancer patients who experience a moderate non-fatal recurrence state and a death state, such as ovarian cancer (see Reference 2) and chronic myeloid leukemia (see Reference 3). Other applications of the disease-death model include Alzheimer's disease (see Reference 4) and cardiovascular disease (see Reference 5).
[0004] Regarding event history analysis in this context, there are two main streams of literature.
[0005] The first approach is based on conventional statistical theory and includes three methods: (i) Nonparametric methods, particularly the Kaplan-Meier estimator (see reference 6) and the Nelson-Erlang estimator (see reference 7). These have traditionally been used to model the risk of events without assuming a distribution of event times and cannot perform individual modeling. (ii) Parametric methods enable individual modeling. Event times are associated with individual covariates via linear regression functions and are distributed according to the underlying probability distribution. (iii) Semiparametric methods allow for a compromise between nonparametric and parametric models. These introduce covariate effects via linear regression functions but make no assumptions about the distribution of event times. The Cox proportional hazards (PH) model is the most widely used semiparametric model in multivariate survival analysis (see reference 8). In multistate analysis, most of the existing literature describes the risk of transitioning between states using transition-specific Cox PH models as (semi)Markov processes (see references 9 and 10). However, these conventional methods rely on strong statistical assumptions about the distribution of event times, or the relationship between covariates and event times. In particular, the Cox PH model makes a linear assumption about the relationship between covariates and event risk. This assumption imposes constraints in many real-world applications, as the effects of covariates can change non-linearly in response to variability in event risk. The Cox model does not consider interactions between covariates by default. This limits the applicability of this model to big data, because most variables are not directly related to the modeled outcome and interact with covariate effects. For example, genetic variations in metabolic pathways may not directly affect the risk of cancer recurrence, but they may reduce or enhance the effectiveness of antitumor therapy. These constraints of the Cox model are widely known in clinical practice.
[0006] To address these challenges, a second type of literature employing new machine learning algorithms has been developed. In particular, neural networks have been developed to extend the Cox PH model in a framework without statistical assumptions. Conventional artificial neural networks, in particular, have been successfully implemented in survival analysis by Faraggi and Simon (see reference 11). More recently, deep neural networks have been extended, notably by Luck et al. (see reference 12), Katzman et al. (see reference 13), Fotso (see reference 14), and Kvamme et al. (see reference 15) (see also references 16, 17). By employing prior art deep learning methods and large clinical datasets, significant improvements in predicting patient survival are seen compared to the Cox PH model. Nevertheless, their methods are still limited to the case of intrinsic clinical events. Also, most recent methods directly predict the discrete-time distribution of event times as the output of the neural network. As an approximation of continuous-time survival data, they all divide the continuous-time scale into discrete-time intervals. This introduces relatively significant approximation errors. [Overview of the project] [Problems that the invention aims to solve]
[0007] Therefore, in this context, with respect to multi-state disease-mortality models, improved solutions are needed to determine accurate patient data for multi-state disease models, for example, in order to provide accurate individual predictions of disease progression.
[0008] 1. Webster AJ. Multi-stage models for the failure of complex systems, cascading disasters, and the onset of disease. PloS one 2019;14(5):e0216422. 2.Eulenburg C,Mahner S,Woelber L,Wegscheider K.A systematic model specification procedure for an illness-death model without recovery.PloS one 2015;10(4):e0123489. 3.Iacobelli S,Carstensen B.Multiple time scales in multi-state models.Statistics in medicine 2013;32(30):5315-5327. 4.Commenges D,Joly P,Letenneur L,Dartigues JF.Incidence and mortality of Alzheimer’s disease or dementia using an illness-death model.Statistics in medicine 2004;23(2):199-210. 5.Ramezankhani A,Blaha MJ,Mirbolouk hM,Azizi F,Hadaegh F.Multi-state analysis of hypertension and mortality:application of semi-Markov model in a longitudinal cohort study.BMC Cardiovascular Disorders 2020;20(1):1-13. 6.Kaplan EL,Meier P.Nonparametric estimation from incomplete observations.Journal of the American statistical association 1958;53(282):457-481. 7.Aalen O.Nonparametric inference for a family of counting processes.The Annals of Statistics 1978:701-726. 8.Cox DR.Regression models and life-tables.Journal of the Royal Statistical Society:Series B (Methodological) 1972;34(2):187-202. 9.De Wreede LC,Fiocco M,Putter H.The mstate package for estimation and prediction in non-and semi-parametric multistate and competing risks models.Computer methods and programs in biomedicine 2010;99(3):261-274. 10.Andersen PK,Borgan O,Gill RD,Keiding N.Statistical models based on counting processes.Springer Science&Business Media .2012. 11.Faraggi D,Simon R.A neural network model for survival data.Statistics in medicine 1995;14(1):73-82. 12.Luck M,Sylvain T,Cardinal H,Lodi A,Bengio Y.Deep learning for patient-specific kidney graft survival analysis.arXiv preprint arXiv:1705.10245 2017. 13.Katzman JL,Shaham U,Cloninger A,Bates J,Jiang T,Kluger Y.DeepSurv:personalized treatment recommender system using a Cox proportional hazards deep neural network.BMC medical research methodology 2018;18(1):24. 14.Fotso S.Deep neural networks for survival analysis based on a multi-task framework.arXiv preprint arXiv:1801.05512 2018. 15.Kvamme H,Borgan O,Scheel I.Time-to-event prediction with neural networks and Cox regression.Journal of Machine Learning Research 2019;20(129):1-30. 16.Giunchiglia E,Nemchenko A,Schaar v.dM.Rnn-surv:A deep recurrent model for survival analysis.In:Springer.;2018:23-32. 17.Ren,K.,Qin,J.,Zheng,L.,Yang,Z.,Zhang,W.,Qiu,L.,&Yu,Y.(2019).Deep Recurrent Survival Analysis.Proceedings of the AAAI Conference on Artificial Intelligence,33(01),4798-4805. 18.Gensheimer MF,Narasimhan B.A scalable discrete-time survival model for neural networks.PeerJ 2019;7:e6257. 19.Lee C,Zame WR,Yoon J,Schaar v.dM.Deephit:A deep learning approach to survival analysis with competing risks.In:;2018. 20.Kalbfleisch JD,Prentice RL.The statistical analysis of failure time data.360.John Wiley&Sons .2011. 21.Friedman M,others .Piecewise exponential models for survival data with covariates.The Annals of Statistics 1982;10(1):101-113. 22.Kvamme H,Borgan O.Continuous and discrete-time survival prediction with neural networks.arXiv preprint arXiv:1910.06724 2019. 23.Andersen PK,Borgan O.Counting Process Models for Life History Data:A Review.Preprint series.Statistical Research Report http: / / urn.nb.no / URN:NBN:no-23420 1984. 24.Meira-Machado L,Una-Alvarez dJ,Cadarso-Suarez C,Andersen PK.Multi-state models for the analysis of time-to-event data.Statistical methods in medical research 2009;18(2):195-222. 25.Meira-Machado L,Sestelo M.Estimation in the progressive illness-death model:A nonexhaustive review.Biometrical Journal 2019;61(2):245-263. 26.Triebel H.Theory of Function Spaces,vol.78 of.Monographs in mathematics 1983. 27.Putter H,Fiocco M,Geskus RB.Tutorial in biostatistics:competing risks and multi-state models.Statistics in medicine 2007;26(11):2389-2430. 28.Ruder S.An overview of multi-task learning in deep neural networks.arXiv preprint arXiv:1706.05098 2017. 29.Lee C,Yoon J,Van Der Schaar M.Dynamic-deephit:A deep learning approach for dynamic survival analysis with competing risks based on longitudinal data.IEEE Transactions on Biomedical Engineering 2019;67(1):122-133. 30.Most S.Regularization in discrete survival models.PhD thesis.lmu,2014. 31.Tibshirani R,Saunders M,Rosset S,Zhu J,Knight K.Sparsity and smoothness via the fused lasso.Journal of the Royal Statistical Society:Series B (Statistical Methodology) 2005;67(1):91-108. 32.AalenOO,JohansenS.An empirical transition matrix for non-homogeneous Markov chains based on censored observations.Scandinavian Journal of Statistics 1978:141-150. 33.Jacqmin-Gadda H,Blanche P,Chary E,Touraine C,Dartigues JF.Receiver operatingcharacteristic curveestimation for time to event with semicompeting risks and interval censoring.Statistical methods in medical research 2016;25(6):2750-2766. 34.Spitoni C,Lammens V,Putter H.Prediction errors for state occupation and transition probabilities in multi-state models.Biometrical Journal 2018;60(1):34-48. 35.Jackson CH.flexsurv:a platform for parametric survival modeling in R.Journal of statistical software 2016;70. 36.Royston P,Altman DG.External validation of a Cox prognostic model:principles and methods.BMC medical research methodology 2013;13(1):33. 37.Bergstra J,Bengio Y.Random search for hyper-parameter optimization.The Journal of Machine Learning Research 2012;13(1):281-305. 38.Wilcoxon F.Individual comparisons by ranking methods.In:Springer.1992 (pp.196-202). 39.Bender R,Augustin T,Blettner M.Generating survival times to simulate Cox proportional hazards models.Statistics in medicine 2005;24(11):1713-1723. 40.Van Wieringen WN,Kun D,Hampel R,Boulesteix AL.Survival prediction using gene expression data:a review and comparison.Computational statistics&data analysis 2009;53(5):1590-1603. 41.Yousefi S,Amrollahi F,Amgad M,et al.Predicting clinical outcomes from large scale cancer genomic profiles with deep survival models.Scientific reports 2017;7(1):1-11. 42.Ching T,Zhu X,Garmire LX.Cox-nnet:an artificial neural network method for prognosis prediction of high-throughput omics data.PLoS computational biology 2018;14(4):e1006076. 43.Kuleshov MV,Jones MR,Rouillard AD,et al.Enrichr:a comprehensive gene set enrichment analysis web server 2016 update.Nucleic acids research 2016;44(W1):W90-W97. 44.Liberzon A,Birger C,Thorvaldsdottir H,Ghandi M,Mesirov JP,Tamayo P.The molecular signatures database hallmark gene set collection.Cell systems 2015;1(6):417-425. 45.Wu Y,Sarkissyan M,Vadgama JV.Epithelial-mesenchymal transition and breast cancer.Journal of clinical medicine 2016;5(2):13. 46.Lee DS,Ezekowitz JA.Risk stratification in acute heart failure.Canadian Journal of Cardiology 2014;30(3):312-319. 47.Park S,Han W,Kim J,et al.Risk factors associated with distant metastasis and survival outcomes in breast cancer patients with locoregional recurrence.Journal of breast cancer 2015;18(2):160. 48.Lip GY,Nieuwlaat R,Pisters R,Lane DA,Crijns HJ.Refining clinical risk stratification for predicting stroke and thromboembolism in atrial fibrillation using a novel risk factor-based approach:the euro heart survey on atrial fibrillation.Chest 2010;137(2):263-272. 49.Sparano JA, Crager MR, Tang G, Gray RJ, Stemmer SM, Shak S.Development and validation of a tool integrating the 21gene recurrence score and clinical-pathological features to individualize prognosis and prediction of chemotherapy benefit in early breast cancer.Journal of Clinical Oncology 2020:JCO-20. 50.Ancona M, Ceolini E, Oztireli C, Gross M.Towards better understanding of gradient-based attribution methods for deep neural networks.arXiv preprint arXiv:1711.06104 2017. 51.Jackson C.Multi-state modeling with R:the msm package.Cambridge,UK 2007:1-53. [Means for solving the problem]
[0009] Accordingly, a computer-implemented method is provided for machine learning a function. The function is configured to output a distribution of transition-specific probabilities for each interval in a set of intervals, based on input covariates representing the medical characteristics of a patient with respect to a multi-state model of a disease having states and transitions between states. The set of intervals forms subdivisions of the follow-up period. The machine learning method includes providing an input dataset of covariates and time-versus-event data for a set of patients. The machine learning method further includes training the function based on the input dataset.
[0010] This machine learning method may include one or more of the following:
[0011] The function includes, for each transition, a covariate-shared subnetwork and / or a transition-specific subnetwork.
[0012] Each covariate-sharing subnetwork contains a fully connected neural network, and / or at least one transition-specific subnetwork contains a fully connected neural network.
[0013] Each covariate-shared subnetwork contains its respective nonlinear activation function, and / or at least one transition-specific subnetwork contains its respective nonlinear activation function.
[0014] Each transition-specific subnetwork is followed by a softmax layer.
[0015] The multi-state model includes competing transitions, and the transition-specific subnetworks of these competing transitions share a common softmax layer.
[0016] Training involves minimizing a loss function that includes a likelihood term and / or a regularization term that penalizes the first difference of weights associated with two adjacent time intervals in the weight matrix and / or the first difference of biases associated with two adjacent time intervals in the bias vector.
[0017] The multi-state model is a disease-mortem model.
[0018] Diseases include, for example, cancerous diseases such as breast cancer, or diseases characterized by intermediate and final stages.
[0019] A disease is characterized by a state representing the patient's condition related to the disease, namely, clinical symptoms (e.g., bleeding, dyspnea at rest), or radiological findings (e.g., osteopenia, fractures), or biological measurements (e.g., creatinine clearance). And / or, Medical characteristics include features that describe the patient's general condition and / or features that describe the patient's condition in relation to the disease.
[0020] Furthermore, we provide a method for using the function trained by the machine learning method described above. This method includes providing input covariates that represent the medical characteristics of the patient, and applying the function to the input covariates. As a result, for a multi-state model of a disease with states and transitions between states, this method outputs a distribution of transition-specific probabilities for each interval in a set of time intervals. The set of time intervals forms a subdivision of the follow-up period.
[0021] This method of use may further include one or more of the following:
[0022] The cumulative occurrence rate function specific to one or more transitions is calculated, and optionally, the cumulative occurrence rate function specific to one or more transitions is displayed.
[0023] Identify patient-related risks of recurrence and / or death. For example, determining treatment, indications for treatment, follow-up visits, and / or monitoring with diagnostic tests based on identified recurrence risk and / or mortality risk.
[0024] Furthermore, we provide a data structure that includes functions trained using the machine learning methods described above.
[0025] Furthermore, a computer program is provided that includes instructions for executing the machine learning methods described above.
[0026] Furthermore, a computer program is provided that includes instructions for performing the above-described usage.
[0027] Furthermore, a device is provided that includes a data storage medium recording either or both of the above-described data structures and / or the above-described computer programs. The device may form or function as a non-temporary computer-readable medium, for example, on a SaaS (Software as a Service) or other server or cloud-based platform. Alternatively, the device may include a processor connected to the data storage medium. Thus, the device may form a computer system in whole or in part (for example, it may be a subsystem of the entire system). The system may further include a graphical user interface connected to the processor.
[0028] Here, a non-limiting example will be explained with reference to the attached diagram. [Brief explanation of the drawing]
[0029] [Figure 1] Let's take a disease-mortality model as an example. [Figure 2] This shows the discretization of the time axis into K time intervals. [Figure 3] An exemplary neural network function architecture is shown. [Figure 4] The results obtained are shown below. [Figure 5] The results obtained are shown below. [Figure 6] The results obtained are shown below. [Figure 7] An example of this system is shown. [Modes for carrying out the invention]
[0030] A computer-implemented method is provided for machine learning a function (i.e., determining a neural network function, i.e., a function containing or consisting of one or more neural networks, through a machine learning process). This function is configured such that, given an input, it provides a specific output. The input is a set of covariates (i.e., multiple data) representing the medical characteristics of a patient. The output is defined with respect to (i.e., by reference to) a multi-state model of the disease (e.g., a disease-death model). In other words, the multi-state model is predetermined. The (e.g., disease-death) multi-state model has (i.e., is characterized by) (e.g., three) states and (e.g., irreversible) transitions between states. Given this, the output is a distribution of transition-specific (i.e., transition-specific) probabilities for each interval in a set of intervals. The set of intervals is such that it forms a subdivision of the follow-up period. In other words, the output is a distribution of probability constant values, with exactly one probability constant value for each time interval of the follow-up period and for each transition of the multi-state model. The aforementioned probability constant (e.g., a single number, e.g., between 0 and 1) represents the probability that a patient will experience each transition (and corresponding state change) during each of the aforementioned time intervals. The machine learning method includes providing a covariate input dataset and (e.g., disease-death) time-versus-event data for a set of patients. The machine learning method further includes training a function based on the input dataset.
[0031] This leads to an improved solution for determining accurate patient data for multi-state disease models, which, in the case of a multi-state disease-mortality model, provides accurate individual predictions of disease progression.
[0032] In particular, after such offline learning / training, the machine-learned function may be used online / inline by providing patient input covariates, i.e., multiple covariates (i.e., multiple data) representing the medical characteristics of the patient. The patient may be a new patient, i.e., not part of the patients related to the dataset used in the machine learning method. The input covariates for this patient may be the same as those used during the training process. This method of use may then include applying the function to the input covariates. This allows for outputting the distribution of transition-specific probabilities for each interval in a set of time intervals that form subdivisions of the follow-up period, for a multi-state model (e.g., disease-death).
[0033] Such distributions of transition-specific probabilities form medical / healthcare data that can be used in any way, particularly in the context of patient follow-up for prognosis and / or treatment decision-making. Therefore, these methods enable disease prognosis, in particular, within an event history analysis framework. This method forms algorithms for individual predictions in multi-state disease models (such as the three-state disease-mortality model described later). By applying multi-state analysis, the true transitions between real-world disease states are modeled, thus achieving high accuracy.
[0034] By applying neural network techniques to multi-state analysis (e.g., disease-death processes), this method can perform nonlinear processing of covariates. The function (i.e., neural network) can be fitted to nonlinear patterns in the data, for example, by using multiple (e.g., fully connected) layers and / or one or more nonlinear activation functions. This addresses the linear constraints of the Cox PH model within the (e.g., disease-death) modeling framework.
[0035] The output distribution of transition-specific probabilities represents the density of the probability of each transition occurring during the follow-up period. Therefore, this method provides a well-approximated continuous-time framework (instead of a discrete-time framework which is prone to higher approximation errors). In other words, this method achieves high accuracy. However, since this method determines one probability (a constant value) for each interval, it is assumed that the transition-specific density probabilities are piecewise constant. This limits the approximation error compared to, for example, the discrete-time method of conventional techniques.
[0036] Section 1 below provides an overview of this method and its optional features. Sections 2-4 describe specific implementations. Sections 5-8 describe the performance of specific implementations. System 9 describes this system.
[0037] 1 General expressions The multi-state model of this method represents a stochastic process and can be any disease model having at least two states and at least two transitions, and optionally at least three states and at least three transitions. Here, the patient is in exactly one state at a time.
[0038] A multi-state model could be, for example, a disease-death model.
[0039] The disease-mortality model includes an initial state (e.g., "no events") representing a patient who is not diseased, an intermediate state (e.g., "relapsed" or "disease") representing a patient who is diseased, and an absorbed state (e.g., "death") representing a patient who has died. The intermediate state is sometimes called "relapsed," but this is merely a convention. For some patients, the transition to the intermediate state may represent either an actual relapse (e.g., if the disease in question is cancer) or the first event caused by the disease (e.g., if the disease in question is an infection). Similarly, the state representing a patient who is not diseased is called the "initial" state, but this is merely a convention. However, in the example, all patients in the dataset may start in the initial state. Similarly, this usage may be applied to patients in the initial state. In another example, the dataset may include patients who start in an intermediate state depending on the time-vs-event data. Similarly, in such a case, this usage may optionally be applied to patients who are diseased and therefore may have a different state (when this usage is performed) than the "initial" state.
[0040] In the example, patients in this use, and / or most (e.g., all or substantially all) of the patient set (included in the dataset) may be either initially affected by the disease (i.e., patients showing symptoms) or previously affected by the disease (i.e., asymptomatic, i.e., in remission). For example, patients in this use, and / or most (e.g., all or substantially all) of the patient set (included in the dataset) may have previously been affected by the disease but are now unaffected and in an initial state (i.e., asymptomatic, i.e., in remission). In other words, the initial state is remission, and all patients covered by this method are in remission. Therefore, an intermediate state entry may represent an actual relapse. Thus, this use may be useful in determining healthcare data for patients already tracked for the disease. In other examples, patients in this use, and / or most (e.g., all or substantially all) of the patient set may initially be unaffected and have never been affected by the disease. Therefore, an initial entry in an intermediate state may represent the first impairment due to the disease. This method of use may be helpful in determining healthcare data for patients who are not tracked for a disease, and for patients who have never been tracked for a disease. In other examples, the patient set may include a significant number of patients of both types and / or conditions, and this method of use may be applied indiscriminately to patients of both types and / or conditions. In such cases, machine learning can be considered a generalist.
[0041] In certain cases, this method relies on specific multi-state models, which are disease-death models, and specific disease-death models that do not allow a return to the initial state. This method may apply disease-death models to patients with a fatal disease (such as cancer) in remission. This method may allow monitoring for the occurrence of relapse or death in these patients. Therefore, in these specific cases, all patients may start in the initial state.
[0042] In other specific examples, such as when applied to infectious diseases, the multi-state model may alternatively include the possibility of returning to the initial state. The multi-state model may also be an infectious disease model.
[0043] The disease-mortality model may optionally consist of exactly these three states. Alternatively, the states may include one or more other states, such as at least one additional intermediate state. For example, in the case of multiple myeloma, these states may be the same, except that there are two intermediate states: "biological relapse" and "clinical relapse."
[0044] The death-disease model includes two competing transitions: a first transition from the initial state to an intermediate state, and a second transition from the initial state to an absorbed state. The death-disease model further includes a third transition from the intermediate state to an absorbed state. The transition to the absorbed state is irreversible. The transition from the intermediate state to the absorbed state is continuous. The transition from the initial state to the intermediate state may also be irreversible (i.e., the model does not include any other transitions), or the disease-death model may further include a fourth transition from the intermediate state to the initial state. In such cases, the transitions from the intermediate state are competing.
[0045] The disease may be any disease that can be characterized by an intermediate state and an attenuation state. In the example where the multi-state model is a disease-mortality model, the disease may be any fatal / lethal disease, i.e., a disease known to cause death in a statistically significant proportion (e.g., more than 1% or 5%) of the whole population. The disease may be, for example, a cancerous disease, e.g., breast cancer. In other examples, the disease may be Alzheimer's disease or a cardiovascular disease. Alternatively, the attenuation state may be non-lethal. For example, age-related macular degeneration may progress toward blindness, which is the final attenuation state.
[0046] Covariates are the data points that form the arguments of the machine learning function. Each such data point represents a medical characteristic of a patient, which is related to the disease. Each data point may be a single value, a vector with multiple values / coordinates, or other types of data structures that convey information about the patient, such as images (e.g., scan images or radio images) or a list of biological analysis results. The neural network approach of this method provides high flexibility in terms of the number and heterogeneity of data in the input covariates.
[0047] Optionally, the input covariates may be provided to the function in the form of unique vectors, or the function may process the input covariates to form such vectors. For example, the function may concatenate values or vectors into a single vector. In the case of image covariates, the function may include a convolutional neural network capable of converting images into vectors and concatenating them with other values and / or vectors. The function may optionally be trained to handle missing data, i.e., null values. Additionally or alternatively, missing data may be completed upstream.
[0048] Medical characteristics may include characteristics that represent the patient's general condition (i.e., so-called "core characteristics") and / or characteristics that represent the patient's condition with respect to the disease (i.e., "condition characteristics").
[0049] Core patient characteristics include, for example, age, body mass index (BMI), comorbidities, sex, and / or other medically relevant general characteristics. Conditional characteristics may be more specific to the disease under study. For example, in the case of breast cancer, patient conditional characteristics may include any one or any combination of the following: menopausal status, Nottingham Prognostic Index (NPI), immunohistochemical estrogen receptor (ER) status, number of positive lymph nodes, cancer grade, tumor size, tumor histological type, cellularity, Her2 copy number, Her2 expression, ER expression, progesterone (PR) expression, type of breast surgery, cancer molecular subtype (pam50 subgroup, integrated cluster), chemotherapy regimen, hormone regimen, and / or radiotherapy regimen.
[0050] The number of covariates may be 10, 20, 100, 500, or even more than 1000. A larger number of covariates generally leads to higher accuracy in the results.
[0051] The dataset provided by this machine learning method includes training samples (i.e., data that form a training pattern) corresponding to each (real) patient, each including, for each patient, the (actual) values of the covariates, as well as time-versus-event data, i.e., date information related to the disease state. The number of training samples may exceed 500, 1000, 5000, 10000, or 20000. A larger number of training samples leads to higher accuracy of results. The time-versus-event data represents a time series of the patient's actual transitions from one state of the model to another state of the model over a specific time interval (i.e., a follow-up period), associated with time information (i.e., date and time, and / or duration). Thus, the time-versus-event data forms the ground truth value that the function adheres to in its output. The dataset consists of data related to patients with no transitions, data related to patients with only one transition, and data related to patients with multiple transitions. The time-versus-event data may provide time information in any way, for example, using the actual date of each event, or a counter (e.g., the number of days and / or starting from 0).
[0052] The input covariates may represent baseline medical characteristics for each patient (in the machine learning dataset and / or the data provided in this usage). “Baseline” means the beginning of the follow-up period. In other words, the function is trained to take patient characteristics taken (e.g., measured) all on the same date (i.e., the start date of the patient's follow-up) as input, and based on that, output a distribution of transition-specific probabilities for the subsequent follow-up period (i.e., starting from that date). Specifically, in the dataset, the covariate data may be limited to baseline data, i.e., data dated to the start point of the time series for each patient.
[0053] This function may be configured to output a distribution of transition-specific probabilities for a given follow-up period or time interval. The follow-up period or time interval may be predetermined and defined by the time-vs-event data provided by this machine learning method. In other words, the follow-up period may be a function of the time-vs-event data. A given follow-up period may be any period exceeding, for example, one month, six months, one year, two years, three years, five years, seven years, or ten years. In such cases, the function may be configured to output a distribution of transition-specific probabilities for a fixed follow-up period or a truncated follow-up period shorter than the follow-up period of the dataset.
[0054] In the example, substantially all training samples in the dataset exhibit the same follow-up period (i.e., the time-to-event data duration is substantially the same for all patients), and the follow-up period of the trained function may be predetermined to be equal to or shorter than that same follow-up period.
[0055] Here, we will describe an example of a neural network architecture for functions.
[0056] The function may include a covariate-shared subnetwork and / or a transition-specific subnetwork (e.g., a total of three transition-specific subnetworks) for each transition. In other words, the covariates (e.g., in a unique vector) may be input to and processed in a single common subnetwork of the function. The output of such a shared processing can then be fed in parallel to each subnetwork for each transition in a multi-state model. This can improve the accuracy of the function.
[0057] In particular, a covariate-shared subnetwork may include its own fully connected neural network (for example, it may consist only of neural networks). Additionally or alternatively, at least one (e.g., each) transition-specific subnetwork may include its own fully connected neural network (for example, it may consist only of neural networks). A fully connected neural network includes one or more sets of fully connected layers that connect all neurons in one layer to all neurons in another layer. Optionally, a covariate-shared fully connected neural network may include several fully connected layers. Additionally or alternatively, at least one (e.g., each) transition-specific fully connected neural network may include several fully connected layers. Further optionally, a covariate-shared subnetwork may include its own nonlinear activation function. Additionally or alternatively, at least one (e.g., each) transition-specific subnetwork may include its own nonlinear activation function. By using fully connected layers (especially multiple layers) and / or nonlinear activation functions, the nonlinear dependence of covariates can be captured, resulting in higher accuracy compared to conventional methods based on linearity assumptions.
[0058] The number of fully connected layers and / or the number of neurons in each such fully connected neural network may depend on the number of training samples and / or the number of covariates. The number of fully connected layers and / or the number of neurons may be determined by random search or predetermined. The number of fully connected layers may be, for example, 1, 2, 3, or greater than 3, and / or the number of neurons may be greater than 10 or 15 and / or less than 100, 50, or 30.
[0059] Each transition-specific subnetwork may be followed by a softmax layer or any other stochastic layer. This allows for the output of values between 0 and 1, which can be interpreted as probability values or a constant probability density. Specifically, if a multi-state model includes competing transitions, the transition-specific subnetworks of the competing transitions may share a common softmax layer or other stochastic layer (for example, the function may contain exactly two softmax layers). This allows for joint learning of probabilities. Such softmax layers ensure that the probability of leaving the initial state converges to 1 over time (since death is certain, whatever the cause), and / or that the probability of leaving an intermediate state converges to 1 over time (since death is certain, whatever the cause, this is also assumed to be some other recurrence).
[0060] Training may include, for example, minimizing the loss using stochastic gradient descent.
[0061] The loss function may include a likelihood term. The likelihood term may penalize the model for not fitting the dataset. In other words, the likelihood term rewards each training sample to the extent that the output distribution of transition-specific probabilities increases the likelihood of ground truth. The likelihood term can impose such a penalty in any way, for example, by being equal to a negative value of the log-likelihood.
[0062] In the disease-mortality model, the log-likelihood may include a first subterm representing the log contribution from the initial state and a second subterm representing the log contribution from the intermediate state. The log-likelihood may be equal to the sum of the two terms. The first subterm may optionally be derived under a time-inhomogeneous Markov assumption for the transition from the initial state, and / or the second subterm may optionally be derived under a time-uniform semi-Markov assumption for the transition from the intermediate state.
[0063] This function outputs the probability distribution specific to each transition for a set of intervals. In other words, for each transition in a multi-state model, and for each interval in the set of intervals, this function outputs a single probability value (i.e., a real number between 0 and 1). The set of intervals is a division of a (predetermined) follow-up period into time intervals. In other words, the time intervals in the set are continuous and non-overlapping, completely covering the follow-up period.
[0064] The number of time intervals and the cut points between them may be predetermined, for example, each interval lasting from one week to six months, optionally from two weeks to three months, or, for example, about one month. The choice of the number of time intervals may affect the model's performance. A regularization term may smooth (e.g., mitigate) the effects of suboptimal time intervals, thereby limiting the risks of overfitting and underfitting. For example, the loss function may further include a regularization term that penalizes the function's weight matrix (e.g., the weight matrix of a fully connected layer), a first-order difference of the weights associated with two adjacent time intervals, and / or a first-order difference of the biases associated with two adjacent time intervals in the function's bias vector (e.g., the bias vector of an output layer such as a softmax layer and / or activation function). In other words, the regularization term penalizes the difference in the neural network's weights and biases that results in a high probability of transitions in consecutive time intervals.
[0065] For a predetermined fixed follow-up period, the cut points (i.e., boundaries) of the intervals may be selected to be equidistant or based on the distribution of transition times.
[0066] The loss function may include the sum of the likelihood term (e.g., negative log-likelihood) and the regularization term. For example, it may consist only of the sum of the likelihood term (e.g., negative log-likelihood) and the regularization term.
[0067] Therefore, this method may extend deep neural networks for disease-death processes. This method may propose a deep learning architecture that models the probability of transitions between states without linear assumptions. The multitask neural network may include one subnetwork shared across all transitions and three transition-specific subnetworks. A novel form of the log-likelihood for a piecewise constant disease-death process may be derived. Experiments conducted on simulated nonlinear datasets and real-world datasets of breast cancer patients are provided later to illustrate the performance of this method.
[0068] Here, we will explain how to use this product.
[0069] The outputted transition-specific probability distribution allows for the performance of various types of medical analyses on the patient regarding the disease, such as prognosis. Examples of such medical analyses are provided in references 9 and 27, which are incorporated herein by reference. This method of use may include any of the medical analyses disclosed in those references. Such types of analyses are known in the field of stochastic processes and are applied here to healthcare.
[0070] This method of use may include calculating the probability of being in one or more states of the model at specific times during the follow-up period. Additionally or alternatively, this method of use may include identifying one or more transition risks, i.e., the risk of the patient following each transition of the multi-state model (changing states accordingly), based on the distribution of transition-specific probabilities. In particular, this method of use may include identifying patient-related relapse risks and / or mortality risks based on the distribution of transition-specific probabilities for each time interval. Relapse and / or mortality risks may be of any type of data, such as percentage values. The output of the function is a small probability of experiencing each transition. The probability of experiencing a transition from an intermediate state to a final state is defined by the time of entering the intermediate state conditionally. The identification of relapse and / or mortality risks may be automated or performed manually by a healthcare professional.
[0071] Additionally, or alternatively, this method of use may include calculating one or more transition-specific cumulative incidence functions (CIFs), each corresponding to a transition in a multi-state model. Examples of methods for calculating such functions are provided in Section 5.1. The CIFs may depend on any type of medical analysis, such as the identification of the recurrence risk and / or mortality risk described above.
[0072] In the example, the method may further include determining treatment (patients who have not yet received treatment for the disease) or treatment indication (patients who are already receiving treatment for the disease) based on the outputted distribution of transition-specific probabilities. Additionally or alternatively, the method may determine follow-up visits (e.g., the method may include specifying the time and / or nature of such follow-up visits) and / or monitoring with diagnostic tests (e.g., the method may include specifying the time and / or nature of such diagnostic tests). This determination may be based on any of the medical analyses described above. For example, the risk of cancer recurrence may be calculated and used to stratify populations with a high / moderate / low risk of recurrence. This classification may determine the nature and intensity of perioperative or adjuvant therapy. Similarly, in atrial fibrillation, the individual cumulative risk of embolic events may be calculated and anticoagulation therapy may be specified to exceed guideline-defined thresholds.
[0073] This usage may include, for example, displaying at least one (e.g., each) CIF for transition directions in a 2D cross plot (e.g., with time on the x-axis), and / or, for example, displaying multiple (e.g., all) CIFs simultaneously (e.g., on the same screen). Additionally or alternatively, this usage may include displaying any graphic representation of the outputted distribution of transition-specific probabilities, and / or any graphic representation of automatically identified (i.e., calculated) recurrence and / or mortality risks.
[0074] Here, we will describe a specific implementation of this method.
[0075] Multi-state models can capture various patterns of disease progression. In particular, disease-death models can be used to track disease progression from a healthy state to an intermediate disease state and then to a final state associated with death. Specific implementations aim to use these models to adapt treatment decisions in accordance with disease progression. In prior art methods, the risk of transitions between states is modeled via (semi)Markov processes and transition-specific Cox proportional hazards (PH) models. However, Cox PH models make linear assumptions about the relationship between covariates and transition risk, which have shown constraints in many real-world applications. To address this challenge, specific implementations propose a neural network architecture (hereinafter referred to as "function example") that relaxes the linear Cox PH assumption within the disease-death process. This function example employs a multi-task architecture to learn the probabilities of transitions between states, using a set of fully connected subnetworks to capture nonlinearity in the data. The following sections investigate various configurations of the architecture through simulations and demonstrate the added value of the model. Compared to prior art methods, function example significantly improves predictive performance on simulated datasets and real-world breast cancer datasets.
[0076] 2. Disease-Mortem Model This section introduces the conventional disease-mortality model. For a complete representation, see the work of Andersen et al. in reference 10.
[0077] 2.1 Notation A multi-state model is a generalization of the survival model when multiple events of interest may occur over time.
number
[0078] In a concrete implementation, we consider a disease-death multi-state process with three states. That is, state 0 is the initial state "no event," state 1 is the intermediate state "relapse," and state 2 is the absorbed state "death." The disease-death process is characterized by three irreversible transitions: from 0 to 1 (0→1), from 0 to 2 (0→2), and from 1 to 2 (1→2), where the transitions from state 0 are competitive, and the transitions 0→1 and 1→2 are continuous. It is assumed that all subjects are in state 0 at time t=0 (i.e.,
number
[0079] The progression of the disease-death process is,
number
number
number
number
[0080] In clinical practice, due to right-side censoring, the true transition time is only partially observed. Let C be a non-negative censored random variable independent of (T0, T2) which would interfere with this observation.
number
number
number
number
number
number
[0081] 2.2 Model Definition The disease-death process has traditionally been defined by the theory of counting processes (see, e.g., reference 23). In the disease-death model, the observation of process E is a trivariate process.
number
number
number
[0082] In certain implementations, two more predictable processes Y0 and Y1 are also included.
number
[0083] A three-variable counting process N is traditionally defined as a set of transition strengths, or transition-specific hazard functions.
number
[0084] Depending on the context, the progression of the three time-related processes can be formulated under various assumptions. In general, the processes associated with each of the three transitions may depend on the time t to reach a state (i.e., a Markov process) or the time d spent in the current state (i.e., a semi-Markov process). The choice of time scale (see, e.g., Reference 3) is a crucial step in disease modeling because it guides how the disease progresses over time. This choice is discussed in Section 3.1.
[0085] 2.3 Log-likelihood Log-likelihood related to the observation of a 3D counting process
number
number
number
number
[0086] Here,
number
number
number
number
[0087] 2.4 Transition-specific Cox PH model The Cox PH method is widely used to assess the impact of covariates on patient survival. We extended the transition-specific Cox PH model to multi-state processes to evaluate the covariate effects for each transition (see reference 9). The covariate effects are introduced into the transition intensity of the log-linear risk function as follows:
number
[0088] Here [Number] is the baseline transition intensity associated with the transition k→l (i.e., the potential hazard when all covariates are equal to zero), and [Number] is the vector of estimated regression coefficients. The coefficient [Number] To estimate, conventionally, the log-likelihood of Equation (1) is used.
[0089] 3 Methodology In this section, a method for modeling the disease-death process is introduced.
[0090] 3.1 Function of Interest Here, modeling the disease-death process may include making specific assumptions about the progression of cancer over time and defining three processes (see Reference 24). For transitions 0→1 and 0→2, in a specific implementation, a non-homogeneous Markov process over time may be considered. For transition 1→2, the modeling may include performing a time transformation and considering a homogeneous semi-Markov process over time (the probability of transitioning from state 1 to state 2 at time t depends only on the period t - T0 already spent). Whenever convenient, the modeling may include using the period variable d = t - T0 instead of the time variable t for transition 1→2. Under these assumptions, the modeling may aim to model the transition probability f kl (.). Refer to the diagram of the disease-death process in Figure 1. The modeling may include defining f 01 and f 02 as the infinitesimal probabilities of experiencing transitions 0→1 and 0→2, respectively.
number
[0091] Modeling is F 01 and F 02 This may include defining them as their cumulative counterparts as follows: That is, for l=1,2,
number
number
number
[0092] In the case of transition 1→2, the function of interest is defined by the condition T0,D0=1. To simplify the notation, this condition may be included in the definition. The modeling is f 12 Let this be the small probability of experiencing the transition 1→2.
number
number
[0093] Modeling is F 12 It may also be defined as its cumulative counterpart as follows:
number
number
[0094] F 01 F 02 , and F 12 This is generally referred to as the cumulative incidence function (CIF) in the literature (see, e.g., reference 25) and is the main function targeted for prediction. Modeling may include a formulation of the disease-mortality process using these probabilities. The disease-mortality process may also be defined by transition intensity (see, e.g., reference 23), as described in Section 2. Both formulations are related.
[0095] 3.2 A systematic approach In certain implementations, it may be assumed that the transition-specific density probabilities are piecewise constant. This should be understood as an approximation, but if the true function is smooth, the approximation error can be limited (see, for example, reference 26).
[0096] First, the modeling process is as follows:
number
number
number
number
number
[0097] Assuming a piecewise constant model, we can consider heterogeneous models (see Section 3.1) while maintaining the hypothesis of homogeneity within the same time interval. This means that the density probability is constant within each interval. Therefore, the modeling is f 0l (l=1,2) and f 12 as a step function
number
number
number
[0098] Here
number
[0099] In actual clinical data, the random variables T0 and T2 are
number
number
number
[0100] Here
number
[0101] 3.3 Data Observation In clinical practice, patient characteristics are observed as P-dimensional covariates along with the medical history. Therefore, the data
number
number
number
number
number
number
number
[0102] 3.4 Rewriting the log-likelihood The modeling may involve rewriting the log-likelihood of a piecewise constant disease-mortality process under the assumption of a piecewise constant model. The conventional disease-mortality log-likelihood, defined in equation (1), can be rewritten in terms of the density probability function introduced in Section 3.1.
[0103] The modeling involves splitting the contribution into two different parts to obtain the log-likelihood.
number
number
[0104] Here,
number
number
number
number
number
number
number
number
[0105] The log-likelihood in equation (5) is derived under the time-inhomogeneous Markov assumptions for transitions 0→1 and 0→2, and the time-uniform semi-Markov assumption for transition 1→2. Depending on the application, the time assumptions can be reformulated, which may change the log-likelihood above.
[0106] 4. Explanation of Function Examples This section describes how a specific implementation uses a new deep learning technique to perform intervals
number
[0107] 4.1 Network Architecture Figure 3 shows an example of the architecture of the function example. The architecture of the function example includes three task-specific subnetworks related to three transitions of the disease-death process. Multi-task learning is performed using hard parameter sharing (see reference 28) to extract general and specific patterns from the characteristics of the patients (i.e., the baseline covariates). This consists of a first subnetwork shared between the three transitions of the covariates (coordinate values of the matrix X) and the three transition-specific subnetworks. Two different softmax output layers are used to convert the transition-specific subnetwork outputs into time-dependent probabilities. One softmax layer is related to the termination from state 0 (i.e., transitions 0→1 and 0→2), and the other is related to the termination from state 1 (i.e., transition 1→2). The first subnetwork and the transition-specific subnetworks are fully connected (Fully Connected: FC), and each fully connected subnetwork may be composed of multiple FC layers as shown in the figure.
[0108] Input Layer The input layer may consist of a matrix X (offline) of P baseline covariates of n individuals or a vector X (online) of P baseline covariates of one individual.
[0109] Transition-specific subnetworks Each transition-specific subnetwork may also take the output of the shared network z as its input, l kl Unit and L kl It includes a fully connected hidden layer. Its output is a vector.
number
[0110] Stochastic output layer The network output is the transition-specific result y kl It may consist of two stochastic layers that map to time-dependent probabilities.
[0111] Under the constraint of equation (4), the function example is the supplementary interval.
number
number
[0112] Thus, each output of the network is (i) a linear activation function g linear is used to (y 01 , y 02 ) T (or y 12 ) into a vector
Number
Number
Number
Number
Number
[0113] 4.2 Reduction of the Effect of the Number of Time Intervals by the Loss Function and Regularization To learn an example of the function parameters, the training is the total loss function that sums the negative log-likelihood and the regularization term
Number
[0114] First term - l K+1 This is a modification of the negative log-likelihood -l defined in equation (5) under the constraint of equation (4) (i.e., considering the supplementary interval K+1).
[0115] The second term P λ is l K+1 This is a regularization term related to the number of suboptimal time intervals (i.e., a suboptimal value of K), which can smooth the effects of the number of suboptimal time intervals (i.e., a suboptimal value of K). The choice of K can affect performance: the number of nodes increases with K, which can lead to overfitting (when K is large) or underfitting. One option may be to choose a large K while preventing overfitting by using L1 regularization on the output layer weights. Another option may be to modify K to a small value in order to make the output layer size as small as possible. Yet another option may be to modify the automatic selection of K by applying a time smoothing technique. Training may include applying a temporal smoothness constraint by penaltying the first-order difference (or bias) of weights associated with two adjacent time intervals in the output layer weight matrix (or bias vector), where W=(W1,W 12 ) and B=(B1,B 12 ), the weight and bias parameters associated with the two output layers,
number
number
number
number
number
[0116] The weights and bias differences are determined by the transition k→l, neuron j, and adjacent time compartment ν. k ν k+1 It is related to the following. Next, the penalty term of the loss function in equation (6) takes the following form:
number
[0117] Here
number
number
number
number
[0118] 4.3 Section cut point α k 'S selection Under certain piecewise assumptions (see Section 3.2), the definition of probability density may be such that time takes the following form.
number
[0119] One option for selecting cut points is to select cut points that are equidistant K at [0,τ] (i.e., uniform intervals of length K / τ). Another option is to select cut points based on the distribution of transition times. This method involves selecting different cut points for the two output layers of the function example.
number
number
number
number
number
number
number
number
[0120] Finally, F0 and F 12 The Aalen-Johansen estimator
number
number
[0121] 5. Prediction Tasks and Benchmarks 5.1 Forecast of individual CIFs This subsection defines the target prediction according to the time scale defined below: the network output (i.e., the step function)
number
[0122] Baseline covariate X j For a new patient j having , from equations (2) and (3)
number
number
[0123] The estimated CIF is used to evaluate the predictive performance of the function example.
[0124] In a particular implementation, such usage may include, in the example, calculating and / or displaying such CIF.
[0125] 5.2 Prediction Evaluation Criteria The performance metrics commonly used in event history analysis are time-dependent AUC (see reference 33 for identification) and time-dependent Briar score (BS, see reference 34) (for calibration). The definitions of time-dependent AUC and time-dependent Briar score may be adapted based on the transition-specific time characteristics defined in Section 3.1.
[0126] To evaluate predictive performance related to transitions 0→1 and 0→2, all patients are considered. For transition 1→2, the prediction for the target is conditionally formulated so that the patient reaches state 1 at time T0 and from period t-T0. Therefore, only patients at risk of experiencing a transition are considered (i.e., only patients who have already experienced transition 0→1).
[0127] For two patients i and j, the transition-specific time-dependent AUC measures the probability that patient i, who experienced transition kl before time t, has a higher probability of experiencing the transition than patient j, who survived until the transition. Transitions 0→1 and 0→2 may be defined as follows, and transitions 1→2 as follows.
number
[0128] The transition-specific time-dependent Briar score measures the difference between the predicted probability of a transition occurring at time t and the actual transition situation, as follows:
number
[0129] Using the censored weighted inverse probability (IPCW) method, several estimators have been developed to account for the loss of information due to censoring. Conventional IPCW estimators already developed in the literature may be used to evaluate predictive performance related to transitions 0→1 and 0→2 (see, e.g., references 33 or 34). For transition 1→2, the conventional estimators may be rewritten under the semi-Markov assumption, considering only patients at risk of experiencing the transition.
[0130] Time-dependent AUC and time-dependent BS can be extended to the interval [0,τ] by calculating the combined AUC (iAUC) and combined Briar score (iBS), respectively, as follows.
number
[0131] 5.3 Benchmarks and Verification The prediction performance of the function examples in CIF prediction is compared in terms of discrimination (using iAUC) and calibration (using iBS) using two prior art statistical methods: the multi-state Cox PH model (msCox) defined in section 2.4 from the R library mstate (Reference 9), and the spline-based version of the Cox multi-state model (msSplineCox) from the R library flexsurv (Reference 35). The function examples are also compared with a simplified linear version (linear function example). The simplified linear version does not have covariate-sharing subnetworks or transition-specific subnetworks.
[0132] (1) Two sets of experiments were conducted on a simulated dataset and (2) on one actual clinical dataset. In the two sets of experiments, each dataset
number
number
number
[0133] The example function can be implemented in Python within a Tensorflow environment. The example function may use at least one of the following deep learning techniques: an L2 regularization layer to avoid overfitting, an L1 regularized output layer, a XavierGaussian initialization scheme, an Adam optimizer, mini-batch learning, learning rate weight decay, and early stopping. The example function may be optimized in two steps, including hyperparameter optimization and parameter optimization. The hyperparameters of the example function may include at least one of the number of nodes and hidden layers in each subnetwork, a regularization penalty parameter, an output-specific penalty parameter, and an activation function for each subnetwork. The hyperparameters may be tuned using RandomSearch. For each hyperparameter, the discrete search space may be modified by manual search. For each model, the search set is tuned according to the input dataset.
[0134] The median (± standard deviation (sd)), iAUC (higher is better), and iBS (lower is better) are compared on the validation set. The performance of the function examples using the other three methods is statistically compared using the two-tailed Wilcoxon (reference 38) signed-rank test. In the results,
number
number
number
number
[0135] 6. Experiments on the simulated dataset Through simulations, we first demonstrate the impact of parameterizing the example function on predictive performance, and then compare the performance of the example function with the other methods described above.
[0136] 6.1 Data Simulation In the initial set of experiments, the simulation may include the generation of continuous-time disease-mortality data. The horizontal time window is fixed at τ=100, and three datasets are generated with sizes n=5000, 20000, and 60000, respectively.
number
number
[0137] Here, the entry of the matrix
number
[0138] The simulation aims to generate processes T0 (and D0) and T2.
number
number
number
number
number
number
number
[0139] The fixed-effect coefficients are fixed at arbitrary values. Therefore, Cox's linearity assumption no longer holds in this simulation scheme.
[0140] Simulated T k1 Therefore, the simulation of processes T0 and T2 must adhere to the constraints of the model in order to generate an identifiable Cox model (i.e., T 01 and T 02 They compete, T 01 and T 12(This is repeated). r=30% is fixed so that 30% of patients in state 0 are censored, and 30% of patients at risk of transition 1→2 are censored from state 1 (see Table 1). Table 1 shows descriptive statistics on the number of observations (No.) in the simulated dataset. [Table 1]
[0141] 6.2 Investigation of the simulation 6.2.1 Understanding the effects of K and n To deepen our understanding of the methodologies described in Sections 4.2 and 4.3, we will examine simulations. Here, we will explore various methods for selecting the number of time intervals K used in the piecewise approximation of the dataset size n, and the selection of interval cut points.
[0142] The evaluation uses the internal verification procedures described in Section 5.3. The evaluation uses the estimated CIF.
number
number
number
[0143] To discretize the time scales of the three datasets, the number of time intervals K is varied by K = 25, 50, 100, and 200. The simulation investigation involves applying two methods to select interval cut points, either using equidistant cut points (uniform) or cut points obtained using Aalen-Johansen quantiles (quantilesAJ).
[0144] Figure 4 shows an example plot of transition-specific validated criteria (iAUC, iBS, iMAE) and K and n values for the cut point selection method. In Figure 4, the solid line represents the quantilesAJ method, and the dotted line represents the uniform method. AUC and BS measurements are aggregated at all four equidistant time points in [0,] (i.e., for computational cost reasons, times t=4,8,12,...,96). MAE measurements are aggregated at all 100 discrete time points in [0,] (i.e., times t=1,2,3,...,100). Figure 4 shows that the interval cut point selection method gives similar performance for iAUC and iMAE. However, for iBS, the uniform selection method performs slightly worse than the quantilesAJ selection method when K is small (K=25). For iBS and iMAE of the three transitions, it can be seen that increasing the value of n improves performance for all K values. Furthermore, in the case of transition 1→2, increasing the value of n improves performance from the perspective of iAUC. However, in the cases of transitions 0→1 and 0→2, the performance difference depending on the value of n is very small.
[0145] 6.2.2 P λ Understanding the role The regularization term P in the loss function of equation (6) λ To deepen our understanding of its role, we used a simulated dataset.
number
[0146] Figure 5 shows an example of plotting transition-specific validated criteria (iAUC, iBS, iMAE) and the value of K. Figure 5 shows the regularization term P of the loss function in equation (6). λ The dataset simulated by
number
[0147] For the larger value of K (i.e., K=200), the regularized model outperforms the unregularized model. Therefore, it is clear that as the value of K increases, the performance of the regularized model remains stable while the performance of the unregularized model declines. Thus, the regularization term P λ Using this approach can mitigate the impact of excessively large K values on the performance of example functions.
[0148] Also, the regularization term Pλ The role of is to provide a smooth estimate of the density probability function. In fact, it may be reasonable to assume that large temporal variations in the density function over consecutive time intervals should be smoothed out. If the value of K is too large, a large jump may occur when calculating CIF. This problem is P λ This can be corrected by applying a time smoothing method, as done by [method name]. To illustrate this point, the simulated dataset is shown.
number
number
number
number
[0149] 6.3 Benchmark Nonlinear simulation dataset
number
[0150] Table 2 shows the integrated prediction performance. In particular, Table 2 shows the performance of the nonlinear simulation dataset.
number
[0151] 7. Application to actual clinical datasets The performance of a particular implementation is evaluated based on actual disease-mortality data from a single clinical trial for breast cancer.
[0152] 7.1 Dataset Description The performance of a specific implementation is evaluated based on data from the molecular classification of the Breast Cancer International Consortium (METABRIC) cohort. In the case of the METABRIC dataset, 1903 patients were followed for RFS and OS for 360 months (30 years), including 38% censored from state 0 and 13% censored from state 1 among patients at risk. Descriptive statistics for the dataset are shown in Tables 3 and 4. Table 3 shows descriptive statistics on the number of observations (No.) of the actual clinical dataset. Table 4 shows descriptive statistics on the transition time (in months) of the actual clinical dataset. [Table 3] [Table 4]
[0153] This dataset includes clinical, histopathological, copy number, and gene expression features used to determine breast cancer subgroups. Based on the literature, 17 baseline clinical, histopathological, and copy number features (12 categories, 5 numerical values) are selected. These features include: age, menopausal status, Nottingham Prognostic Index (NPI), immunohistochemical estrogen receptor (ER) status, number of positive lymph nodes, cancer grade, tumor size, tumor histological type, cellularity, Her2 copy number by SNP6, Her2 expression, ER expression, progesterone (PR) expression, type of breast surgery, cancer molecular subtype (pam50 subgroup, integrated cluster), chemotherapy regimen, hormone regimen, and radiotherapy regimen.
[0154] The evaluation may include substituting missing values with the median of numerical features and the mode of categorical features. The evaluation may include applying one-hot encoding to categorical features and standardizing numerical features using Z-scores. The evaluation may include the time interval ν of months j. jThe application of fixing a fixed length in time intervals up to one month in the daily time interval [(j-1)×30.5,j×30.5)] may include the application of fixing K=120 (i.e., τ=120 months = 10 years), dividing the time axis into monthly intervals from 0 to 120 months, and setting the event time 120 months later in the last interval ν121.
[0155] 7.2 Gene selection using independent datasets related to breast cancer Several methods have been reported for integrating gene expression data into survival models. These methods are based on dimensionality reduction, gene or metagene selection (see reference 40), or, more recently, the use of a large number of gene expression values (>1000) with the development of deep learning methods (see references 41 and 42). Evaluation may also use selection methods to extract a ranked list of cancer-related genes based on p-values from independent transition-specific Cox PH models. Genes are selected using independent datasets to control for selection bias: the invasive breast cancer (TCGA) PanCancer dataset is used for evaluation. The invasive breast cancer (TCGA) PanCancer dataset is available at the following URL.
[0156] https: / / www.cbioportal.org / study / clinicalData?id=brca_tcga_pan_can_atlas_2018. The selection process will be explained in the following two steps.
[0157] 1. Regarding the function of each gene expression.
[0158] • Fit an independent transition-specific Cox PH model that includes clinical covariates and the function of each gene expression. For each of the three transitions, calculate the p-value using the Wald test for the estimated coefficient associated with a specific gene.
[0159] 2. Regarding each transition.
[0160] • Adjust the p-values individually using Benjamini-Hochberg multi-test correction for each transition; • Rank the p-values of each transition; • Various p-value thresholds (indicated as α) are used for gene selection. Table 5 shows the number of genes selected for each threshold.
[0161] Table 5 shows the number of genes selected for each transition in the TCGAPanCancer dataset for invasive breast cancer at each threshold. [Table 5] [Table 6]
[0162] In the case of transition 0→1, several genes are selected (up to 110 genes at a threshold of 0.1). In the cases of transitions 0→2 and 1→2, only a few genes are selected. To evaluate the biological significance of the selected genes, the evaluation may include performing gene set enrichment analysis (see reference 43) for each transition-specific gene list against the Hallmark collection (see reference 44) of the Molecular Signatures Database (MSigDB). In the cases of transitions 0→2 and 1→2, there are no significant enriched terms. In the case of transition 0→1, nine significant enriched terms are shown in Table 6. Table 6 shows significant enriched terms for the list of genes specific to transition 0→1. In particular, the most important enriched term, “epithelial-mesenchymal transition,” relates to a core biological element in the metastatic process of breast cancer (see reference 45) and is therefore strongly associated with the progression of primary cancer due to recurrence of the condition. Thus, the genes selected in transition 0→1 exhibit a consistent functional profile associated with cancer recurrence. Therefore, when a selected gene is added to a disease-mortality model, it is expected that the performance of the transition from 0 to 1 will improve compared to other genes.
[0163] 7.3 Benchmark Tables 7 and 8 show the integrated predictive performance. Evaluation may include comparing different models with various lists of genes that vary the p-value threshold α. For each of the four algorithms compared, the models exhibit different performance with respect to iAUC. The algorithms msCox and msSplineCox show similar iAUC for transition 0→1, better iAUC for transition 0→2, and worse iAUC for transition 1→2 as the value of α increases (except for α=0.1). The function example and linear function example algorithms improve their iAUC for transition 0→1 as the value of α increases. As expected, the function example and linear function example show similar iAUC for transitions 0→2 and 1→2 as the value of α increases. Thus, the two deep learning methods have better iAUC for transition 0→1 thanks to the added gene function selected in the modeling. These results are consistent with the gene set enrichment analysis performed in Section 7.2.2, where the list of genes selected for transition 0→1 is particularly relevant to the transition process, while the list of genes selected for transitions 0→2 and 1→2 does not contain any enriched terms. Finally, Model M BH 0.1 However, it is selected as the best model for predicting transitions 0→1 and 1→2. Model M BH 0.1 The example function shows that the iAUC for transition 0→1 is 0.714 and the iBS is 0.142, while the iAUC for transition 1→2 is 0.730 and the iBS is 0.167, significantly outperforming other algorithms. Regarding the performance of transition 0→2, experiencing transition 0→2 means that the patient died from a cause other than cancer, but the clinical and biological features provided to the model are not related to non-cancer causes of death (except age).
[0164] Table 7 shows the AUC (median ± sd) integrated into the validation set (internal validation) of the METABRIC dataset, with a uniformly subdivided time scale of K=120 (months). AUC and BS measurements are integrated at all 30 equidistant time points in [90,τ] (i.e., from 3 months to monthly). Model M0 uses only clinical features, and each model M BH teeth,
number
[0165] Table 8 shows the BS (median ± sd) integrated into the validation set (internal validation) of the METABRIC dataset, with a uniformly subdivided time scale of K=120 (months). AUC and BS measurements are integrated at all 30 equidistant time points in [90,] (i.e., from 3 months to monthly). Model M0 uses only clinical features, and each model M BH teeth,
number
[0166] 8.Result analysis The function example forms a novel deep learning method for modeling the disease-mortality process and predicting the two-stage evolution of disease based on baseline covariates. This is the first deep learning architecture developed in the context of multi-state analysis, in which it outperforms standard methods. In clinical practice, the function example could be useful in personalized medicine by providing predictions of the risk of recurrence and death. This would help physicians adapt treatment guidelines to specific patients.
[0167] Disease prediction is a crucial momentum in the medical decision-making process for various diseases, for example, identifying populations at risk of cardiovascular complications or cancer recurrence. The most commonly used methods are based on nomograms and scores calculated based on several clinical and biological characteristics. For example, the CHAD2-DS2VASC score is commonly used by physicians to identify patients who require anticoagulation therapy after a diagnosis of cardiac atrial arrhythmia (Reference 48). Similarly, the RSClin tool, which combines clinical, pathological, and genetic information, is commonly used in oncology to predict the risk of breast cancer recurrence and to more accurately determine patients who require chemotherapy in addition to surgery (see Reference 49). However, much of the medical information contained in patient records remains underutilized. In contrast to the development of new biomarkers that rely on expensive and highly specialized technologies, another approach focuses on optimizing prognostic models based on large but accessible information. In the field of digital medicine, which aims to collect and combine vast amounts of data, specialized algorithms of function examples support healthcare professionals' practice based on standardized guidelines on the one hand and individualized assessments on the other.
[0168] The example function learns to estimate the density probability of state transitions occurring in a disease-mortality process using a multitasking architecture and transition-specific subnetworks. It provides accurate predictions of the cumulative probability of state transitions using piecewise approximations. The example function models the relationship between covariates and transition risks without assumptions using multiple fully connected layers and a nonlinear activation function. This is trained by minimizing a loss function designed to capture the relationship between covariates and transition risks and provide a smooth piecewise approximation of the density probabilities.
[0169] Experiments on the simulated dataset investigate various configurations of the function example, revealing the advantages of the loss function. Comparison of the method's predictive performance with prior art methods using discrimination and calibration criteria, and evaluation of the function example in predicting the cumulative probability of state transitions in the simulated dataset and actual breast cancer datasets, demonstrate that the function example provides a significant improvement over other methods when nonlinear patterns are found in the data. Note that the dataset used in the test was relatively small, and the number of medical characteristics was relatively limited; therefore, the results may be further improved with larger datasets.
[0170] Furthermore, the combination of heterogeneous individual functions improves medical decision-making. Regarding real-world datasets related to breast cancer, Section 7 demonstrates how to easily adapt function examples to integrate various types of data (clinical, biological, molecular, gene expression, etc.), showing a consistent and significant improvement compared to statistical methods.
[0171] Time-versus-event data collected in clinical settings can present problems with high censoring rates, as is the case with rare diseases. Applying survival models to datasets with fewer events than censored observations can make it difficult to model the relationship between covariates and event risk. Furthermore, deep learning methods improve with larger amounts of training data. Therefore, developing reliable multi-state models with few events and limited training is challenging, potentially leading to inadequate predictive models. This is often seen in clinical trial datasets, which have relatively small patient populations compared to databases commonly used in deep learning. The example function architecture is adapted to such datasets from clinical trials because it uses hyperparameters that simplify the architecture according to the training data available in each of the three transitions. Moreover, given the growing popularity of using real-world data (RWD) collections such as electronic health records (EHRs) and disease registries, certain implementations aim to leverage this data in the future to help people with problems make more efficient and reliable decisions.
[0172] The example function was developed for application to the disease-mortem process and may be applied to many cancers to predict two-stage progression. This method can be generalized to more complex disease processes by adapting states and transitions.
[0173] 9. System All methods described herein are performed on a computer. This means that the steps (or substantially all steps) of the method are performed by at least one computer or any similar system. Thus, the steps of the method may be performed by the computer fully automatically or semi-automatically. In the example, at least some of the steps of the method may be triggered through user interaction. The required level of user-computer interaction may depend on the expected level of automation and may be balanced with the need to implement the user's requests. In the example, this level may be set by the user and / or predefined.
[0174] A typical example of implementing the method by computer is to perform the method using a system suitable for this purpose. The system may include a processor connected to memory storing a computer program containing instructions for performing the method, and a graphical user interface (GUI). The memory may also store a database. The memory is any hardware suitable for such storage and may include several physically distinct parts (e.g., one for the program and possibly one for the database).
[0175] Figure 7 shows an example of this system, which is a client computer system, such as a user's workstation.
[0176] The client computer in this example comprises a central processing unit (CPU) 1010 connected to an internal communication bus 1000, and random access memory (RAM) 1070 also connected to the bus. The client computer further comprises a graphics processing unit (GPU) 1110 associated with video random access memory 1100 connected to the bus. The video RAM 1100 is also known in the art as a frame buffer. A mass storage device controller 1020 manages access to mass storage devices such as a hard drive 1030. Mass memory devices suitable for concretely implementing computer program instructions and data include all forms of non-volatile memory, including, for example, semiconductor memory devices such as EPROMs, EEPROMs, and flash memory devices, magnetic disks such as internal hard disks and removable disks, magneto-optical disks, and CD-ROM disks 1040. Any of the above may be complemented or incorporated by specially designed ASICs (Application-Specific Integrated Circuits). A network adapter 1050 manages access to the network 1060. The client computer may also include a cursor control device, a tactile device 1090 such as a keyboard. A cursor control device is used within a client computer to allow the user to selectively position the cursor at any desired location on the display 1080. Furthermore, the cursor control device allows the user to select various commands and input control signals. The cursor control device includes a number of signal generating devices for inputting control signals to the system. Typically, the cursor control device may be a mouse, and the mouse buttons are used to generate signals. Alternatively, or additionally, the client computer system may include a sensing pad and / or a sensing screen.
[0177] The computer program may include instructions that can be executed by the computer, and the instructions include means for causing the system to execute the Method. The program may be recordable on any data storage medium, including the system's memory. The program may be implemented, for example, in digital electronic circuits, or in computer hardware, firmware, software, or a combination thereof. The program may be implemented as a device, such as a product specifically realized in a machine-readable storage device for execution by a programmable processor. The steps of the Method may be performed by a programmable processor executing a program of instructions and performing the functions of the Method by manipulating input data to produce outputs. Thus, the processor may be programmable to receive data and instructions from and to a data storage system, at least one input device, and at least one output device, and may be connected in such a way as to send data and instructions to them. The application program may be implemented in a high-level procedural or object-oriented programming language, or, if necessary, in assembly language or machine code. In either case, the language may be a compiled language or an interpreted language. The program may be a full installation program or an update program. In either case, applying the program to a system provides instructions for executing the Method.
Claims
1. A computer-based machine learning method for learning a function that, upon receiving input covariates representing the medical characteristics of a patient, outputs a distribution of transition-specific probabilities for each interval in a set of intervals, with respect to a multi-state model of a disease having states and transitions between states, the function is configured to output a transition-specific probability distribution for each interval in a set of intervals. The set of the aforementioned intervals forms a subdivision of the follow-up period. The distribution of the probability specific to the transition is piecewise constant. The aforementioned multi-state model is a disease-death model, The aforementioned disease is a cancerous disease or a disease characterized by an intermediate state and a final state. The medical characteristics include features representing the patient's general condition and / or features representing the patient's condition with respect to the disease. The aforementioned machine learning method is For each set of patients, the present invention provides an input dataset including covariates representing the aforementioned medical characteristics and time-vs-event data, wherein the time-vs-event data represents a time series of the actual transitions of the patient from one state of the model to another state of the model, associated with time information, over a given duration of time. Training the function based on the input dataset, wherein the training is Likelihood term, and A regularization term that penalizes the linear difference of weights associated with two adjacent time intervals in the weight matrix, and / or the linear difference of biases associated with two adjacent time intervals in the bias vector. This includes minimizing a loss function that includes and Machine learning methods, including [specific examples].
2. The machine learning method according to claim 1, wherein the function includes, for each transition, a covariate-shared subnetwork and / or a transition-specific subnetwork.
3. The machine learning method according to claim 2, wherein the covariate-sharing subnetworks each include a fully connected neural network, and / or at least one transition-specific subnetwork includes a fully connected neural network.
4. The covariate sharing subnetwork includes its respective nonlinear activation function, and / or at least one transition eigensubnetwork includes its respective nonlinear activation function. The machine learning method according to claim 2 or 3.
5. A machine learning method according to any one of claims 2 to 4, wherein each transition-specific subnetwork is followed by a softmax layer.
6. The machine learning method according to any one of claims 2 to 5, wherein the multi-state model includes competing transitions, and the transition-specific subnetworks of the competing transitions share a common softmax layer.
7. The aforementioned disease is breast cancer. A machine learning method according to any one of claims 1 to 6.
8. A method for using a machine learning function according to any one of claims 1 to 7, To provide input covariates that represent the medical characteristics of the patient, This includes applying the function to the input covariates, thereby outputting the distribution of transition-specific probabilities for each interval in a set of time intervals, with respect to a multi-state model of a disease having states and transitions between those states. The set of time intervals described above forms a subdivision of the follow-up period.
9. Based on the distribution of the outputted transition-specific probabilities, Calculate one or more transition-specific cumulative occurrence rate functions (CIF), (a) Identify the risk of recurrence and / or death associated with the patient; and (b) Determine at least one of the following: treatment, indications for treatment, follow-up visits, and monitoring with diagnostic tests. The method according to claim 8, further comprising:
10. In the step of calculating the cumulative occurrence rate function (CIF) specific to one or more transitions, the cumulative occurrence rate function specific to one or more transitions is displayed. The method according to claim 9.
11. The determination of at least one of the aforementioned treatments, indications for treatment, follow-up visits, and monitoring by diagnostic tests is further made based on the identified risk of recurrence and / or death. The method according to claim 9.
12. A computer program for causing a computer to perform the method according to any one of claims 1 to 11.
13. An apparatus comprising a data storage medium on which the computer program described in claim 12 is recorded.
14. The apparatus according to claim 13, further comprising a processor connected to the data storage medium.