Method and device for evaluating complex pathological state and comprehensive risk of elderly patient
By obtaining 19 characteristic variables of elderly patients, using the complex pathological state subtype prediction model and the first survival probability prediction model, the shortcomings of the identification of complex pathological states and the prediction of disease severity in elderly patients were solved, and a more accurate comprehensive risk assessment was achieved, supporting precise medical care and personalized treatment for elderly patients.
Patent Information
- Application Number
- CN202510172679.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art lacks practical methods to accurately identify complex pathological states in elderly patients and predict the comprehensive severity of the disease while taking into account their complex pathological states.
By obtaining the values of 19 characteristic variables of elderly patients, using the complex pathological state subtype prediction model to predict the subtype category of their complex pathological state, and combining the first survival probability prediction model, predict the survival probability in different cumulative time intervals, and finally calculate the comprehensive risk value to evaluate the comprehensive risk level.
It realizes accurate identification of complex pathological status in elderly patients and predicts the comprehensive severity of the disease, provides a more accurate and reference value-added comprehensive risk assessment, and supports precise medical care and personalized treatment.
Smart Images

Figure CN120089369A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of medical information processing, and particularly relates to a method and device for evaluating the complex pathological conditions and comprehensive risks of elderly patients. Background Art
[0002] With the accelerating aging of the global population, the health problems of the elderly population have become increasingly prominent. According to the statistics of the World Health Organization (WHO), the elderly population aged 65 and above is expected to increase significantly in the next few decades. Due to physiological and psychological changes, the elderly face more complex and serious health challenges than the young. First of all, the physiological functions of the elderly gradually decline, and the function of the immune system decreases, making them more susceptible to infections and various diseases. In addition, the elderly often suffer from multiple chronic diseases, which usually interact with each other to form complex pathological conditions. Identifying the complex pathological conditions of patients as early as possible and predicting the comprehensive severity of diseases contribute to the precise and personalized treatment of elderly patients.
[0003] The methods for risk assessment of patients in current clinical practice include the following. The SOFA score (Sequential Organ Failure Assessment) was proposed in 1996 and has been widely used globally to date. It calculates and analyzes 9 indicators / interventions including oxygenation index, whether mechanical ventilation is used, platelet count, serum bilirubin concentration, mean arterial pressure, whether vasopressors such as dopamine are used and the dosage, Glasgow Coma Scale (GCS), serum creatinine concentration, and urine output to achieve a quantitative assessment of the degree of failure of 6 organ systems, including the respiratory, circulatory, renal, coagulation, hepatic, and nervous systems. The disease risk level weights assigned to each organ system are the same, all ranging from 0 to 4 points, and the higher the score, the greater the risk. Another scoring system used in the ICU but with more complex calculations is the Simplified Acute Physiology Score II (SAPSII). The SAPSII score was proposed in 1993 and calculates and analyzes 14 indicators including age, heart rate, systolic blood pressure, body temperature, oxygenation index, urine output, blood urea nitrogen, white blood cell count, potassium ion concentration, bicarbonate, bilirubin concentration, Glasgow Coma Scale, whether there are chronic diseases (metastatic cancer, hematological malignancies, AIDS), and the type of ICU admission to achieve an assessment of the in-hospital death risk of adult patients. There is also a more complex and time-consuming score that has been widely used in tertiary or teaching hospitals in recent years, the Acute Physiology and Chronic Health Evaluation IV (APACHE IV). The APACHE IV was proposed in 2006 and calculates and analyzes 149 indicators to obtain the in-hospital and ICU mortality rates and the length of hospital stay of critically ill adult patients. In addition, some machine learning methods, such as logistic regression, have also been applied to the prediction of patient mortality.
[0004] The commonly used scoring systems mentioned above, such as SOFA and SAPSII, were proposed in earlier periods. However, with the changes in patients' baseline characteristics and the continuous improvement of medical intervention measures, these scoring systems are becoming increasingly inapplicable. APACHE IV is relatively new, but its calculation process is complex, and a large number of indicators need to be measured and evaluated, which limits its practicality in clinical applications. In addition, the above-mentioned tools, including machine learning methods, are not specifically designed for elderly patients, resulting in limited effectiveness in evaluating elderly patients. Therefore, there is a lack of practical methods in the prior art to accurately identify the complex pathological conditions of elderly patients and predict the comprehensive severity of diseases considering their complex pathological conditions. Summary of the Invention
[0005] The present application is provided to solve the above-mentioned defects existing in the prior art. There is a need for a method and device for evaluating the complex pathological conditions and comprehensive risks of elderly patients, which can predict the complex pathological conditions of elderly patients, and considering the complex pathological conditions, realize the prediction of the survival probability of elderly patients in different cumulative time intervals, as well as a more accurate and valuable comprehensive risk assessment, providing stronger support for the precision medicine and personalized treatment of elderly patients.
[0006] According to the first aspect of the present application, there is provided a method for evaluating the complex pathological conditions and comprehensive risks of elderly patients, including: by a processor: obtaining the values of 19 feature variables of the elderly patient to be evaluated, the 19 feature variables including gender, age, body mass index, mechanical ventilation, dialysis, lowest hemoglobin, highest lactate, urine volume, lowest oxygen saturation, average heart rate, average systolic blood pressure, average respiratory rate, highest blood glucose, lowest Glasgow coma score, highest blood urea nitrogen, lowest chloride ion concentration, SOFA score, delirium marker, and laboratory frailty index; based on the values of the 19 feature variables, using a complex pathological condition subtype prediction model to predict the probability distribution of each subtype category of the complex pathological conditions of the elderly patient, and taking the subtype category of the complex pathological conditions to which the elderly patient belongs determined based on the probability distribution of each subtype category as the value of the 20th feature variable; based on the values of the 20 feature variables, using a first survival probability prediction model to predict the survival probability of the elderly patient in at least one cumulative time interval; calculating a comprehensive risk value of the elderly patient based on the probability distribution of each subtype category and its corresponding first weight, and the survival probability of each cumulative time interval and its corresponding second weight, so as to give a comprehensive risk level of the elderly patient based on the comprehensive risk value.
[0007] According to a second aspect of the present application, there is provided a device for evaluating the complex pathological conditions and comprehensive risks of elderly patients, including an interface and at least one processor. The interface is configured to receive the values of 19 characteristic variables of the elderly patient to be evaluated, and the 19 characteristic variables include gender, age, body mass index, mechanical ventilation, dialysis, lowest hemoglobin, highest lactate, urine output, lowest blood oxygen saturation, average heart rate, average systolic blood pressure, average respiratory rate, highest blood glucose, lowest Glasgow Coma Scale score, highest blood urea nitrogen, lowest chloride ion concentration, SOFA score, delirium marker, and laboratory frailty index; at least one processor is configured to execute the method for evaluating the complex pathological conditions and comprehensive risks of elderly patients as described in various embodiments of the present application.
[0008] For the method and device for evaluating the complex pathological conditions and comprehensive risks of elderly patients provided in various embodiments of the present application, first, based on the values of 19 characteristic variables of the elderly patient, the subtype category to which the complex pathological condition of the elderly patient belongs is predicted. Then, in the prediction process of the survival probability in subsequent multiple cumulative time intervals, and further in the prediction process of the comprehensive risk of the elderly patient, the predicted subtype category of the complex pathological condition is taken into consideration. In this way, the prediction results of the survival probability in each cumulative time interval and the evaluation results of the comprehensive risk can be more accurate in the scenario of the elderly population. Moreover, by setting corresponding weights for different subtype categories and the survival probabilities in each cumulative time interval, the influence of the complex pathological condition subtypes and the generation probabilities in different cumulative time intervals on the disease severity and comprehensive risk status of elderly patients can be reflected with finer granularity, which can make the method of the present application have a wider applicability in different elderly patient sample sets, and the prediction results are more accurate and reliable.
[0009] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are hereinafter specifically exemplified.
[0010] It should be understood that the foregoing general description and the following detailed description are merely illustrative and explanatory, and are not restrictive of the claimed invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0012] Figure 1The flowchart shows a method for evaluating the complex pathological conditions and comprehensive risks of elderly patients according to an embodiment of the present application.
[0013] Figure 2 The matrix heatmap shows the multicollinearity detection of input feature variables according to an embodiment of the present application.
[0014] Figure 3 The schematic diagram shows the subtype categories of complex pathological conditions obtained by clustering through the k-means algorithm according to an embodiment of the present application.
[0015] Figure 4 The confusion matrix shows the prediction using a complex pathological condition subtype prediction model according to an embodiment of the present application.
[0016] Figure 5(a) shows a comparison of the patient survival probability prediction curves obtained using 20 feature variables and 19 feature variables according to an embodiment of the present application.
[0017] Figure 5(b) shows the confidence curves of the patient survival probabilities predicted using 20 feature variables and 19 feature variables according to an embodiment of the present application.
[0018] Figure 6 The interpretability analysis results of 20 feature variables of the first survival probability prediction model according to an embodiment of the present application are shown.
[0019] Figure 7 The schematic diagram shows the rules and steps for determining the first weight and the second weight according to an embodiment of the present application.
[0020] Figure 8 The partial composition of the device for evaluating the complex pathological conditions and comprehensive risks of elderly patients according to an embodiment of the present application is shown. Detailed implementation
[0021] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.
[0022] Unless otherwise defined, technical terms or scientific terms used in this application shall have the ordinary meanings as understood by those of ordinary skill in the art to which this application pertains. Words such as "including" or "comprising" and the like mean that the elements or items appearing before the word cover the elements or items listed after the word and their equivalents, without excluding other elements or items.
[0023] The "first", "second" and similar words used in this application do not denote any order, quantity or importance, but are only used for distinction. Words such as "including" or "comprising" and the like mean that the elements before the word cover the elements listed after the word, and do not exclude the possibility of also covering other elements. The execution order of each step in the methods described in this application in combination with the drawings is not limited. As long as the logical relationship between each step is not affected, several steps can be integrated into a single step, a single step can be decomposed into multiple steps, or the execution order of each step can be adjusted according to specific requirements.
[0024] It should also be understood that the term "and / or" in this application is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this application generally represents an "or" relationship between the associated objects before and after.
[0025] To keep the following description of the embodiments of this application clear and concise, detailed descriptions of known functions and known components are omitted in this application.
[0026] According to an embodiment of this application, a method for evaluating the complex pathological state and comprehensive risk of elderly patients is provided. Figure 1 A flowchart showing the method for evaluating the complex pathological state and comprehensive risk of elderly patients according to an embodiment of this application is shown.
[0027] As Figure 1 shown, first, in step 101, a processor may obtain the values of 19 characteristic variables of the elderly patient to be evaluated. The 19 characteristic variables include gender, age, body mass index, mechanical ventilation, dialysis, lowest hemoglobin, highest lactate, urine output, lowest oxygen saturation, average heart rate, average systolic blood pressure, average respiratory rate, highest blood glucose, lowest Glasgow Coma Scale score, highest blood urea nitrogen, lowest chloride ion concentration, SOFA score, delirium marker, and laboratory frailty index.
[0028] The above 19 feature variables were selected from a total of 148 feature variables of common demographic characteristics, laboratory indicators, and comprehensive scoring indicators by using a large amount of sample data of elderly patients from multiple hospitals, combined with empirical surveys of clinicians, the latest scientific research results, and a large number of experimental verification results. The selected features are potentially more important for elderly patients. Then, through multicollinearity detection, 19 variables that are significantly important for the two prediction models (complex pathological state subtype prediction model and first survival probability prediction model) in this application and contribute significantly to the discrimination of the complex pathological state subtype prediction model were finally selected, including gender, age, body mass index, mechanical ventilation, dialysis (i.e., renal replacement therapy), lowest hemoglobin, highest lactate, urine output, lowest oxygen saturation, average heart rate, average systolic blood pressure, average respiratory rate, highest blood glucose, lowest Glasgow coma score, highest blood urea nitrogen, lowest chloride ion concentration, SOFA score, delirium marker, and laboratory frailty index. Figure 2 Shows a matrix heat map of the multicollinearity detection of the input feature variables according to an embodiment of the present application. As Figure 2 shown, except for the values on the diagonal, none of the other values in the matrix exceed 0.7, which is the threshold for multicollinearity detection in the art. The maximum value is approximately 0.48. Therefore, each feature variable in this application has passed the multicollinearity detection.
[0029] Then, in step 102, based on the values of the above 19 feature variables, the complex pathological state subtype prediction model is used to predict the probability distribution of each subtype category of the complex pathological state of the elderly patient, and the subtype category of the complex pathological state to which the elderly patient belongs determined based on the probability distribution of each subtype category is used as the value of the 20th feature variable of the elderly patient. That is, when the values of the 19 feature variables of the elderly patient sample are input into the complex pathological state subtype prediction model, the output of the model is an output vector composed of several probability values, and each probability value in the output vector corresponds to the probability that the sample belongs to each subtype category of the complex pathological state. In some embodiments, for example, the subtype category of the complex pathological state corresponding to the maximum value among the probability values can be determined as the subtype category of the complex pathological state to which the elderly patient belongs.
[0030] Next, in step 103, further, based on the values of the 20 feature variables, the first survival probability prediction model is used to predict the survival probability of the elderly patient in at least one cumulative time interval.
[0031] In some embodiments, the number of cumulative time intervals and the specific settings of each cumulative time interval can be set according to the experience and concerns of medical staff users. For example, clinically, time intervals such as 1 - 30 days, 1 - 90 days, and 1 - 360 days are usually of concern. This application does not make specific limitations in this regard. The resolution of the cumulative time intervals on the time axis is consistent with the first survival probability prediction model, such as 1 day or 1 hour. This application does not make limitations in this regard.
[0032] Finally, in step 104, based on the probability distributions of each subtype category and their corresponding first weights, as well as the survival probabilities of each cumulative time interval and their corresponding second weights, the comprehensive risk value of the elderly patient is calculated, so as to give the comprehensive risk level of the elderly patient based on the comprehensive risk value.
[0033] Although only the determined subtype category is used when the subtype category participates in the prediction of the survival probability of the elderly patient by the first survival probability prediction model as the 20th feature variable, when calculating the comprehensive risk value, the original probability distributions of each subtype category output by the complex pathological state subtype prediction model can more comprehensively and accurately reflect the impact of the subtype category on the comprehensive risk of the elderly patient, and can, to a certain extent, make up for some possible errors in the aforementioned subtype clustering.
[0034] In some embodiments, for example, thresholds for different comprehensive risk levels can be set according to the experience of medical staff, and based on the calculated comprehensive risk value of the elderly patient, the elderly patient can be classified into different risk levels. Specific risk levels can include, for example, good, mild, moderate, severe, critical, etc. This application does not make limitations in this regard. Thus, as a user, medical staff can have a clear and comprehensive understanding of the physical condition of the elderly patient, the subtype category of the complex pathological state to which they belong, the multi - period death risk, the severity of the disease or risk, etc., which helps doctors grasp the condition of the elderly patient, intervene in a timely manner, and reasonably allocate medical resources, including ordinary resources such as hospital beds and scarce resources such as dialysis machines.
[0035] Merely as an example, the user can, for example, select a cumulative time interval with special significance that they are concerned about, such as the expected time for arranging the patient's discharge. By predicting the survival probability of the patient in the concerned time interval and the evaluation result of the comprehensive risk to assist in decision - making. For example, when the evaluation result shows that the comprehensive risk is relatively high within this cumulative time interval, a suggestion to postpone the discharge can be given, and so on. Thus, the method according to the embodiments of this application can provide a convenient and beneficial auxiliary tool for doctor users to assess the risks of elderly patients.
[0036] The method for evaluating the complex pathological state and comprehensive risk of elderly patients according to the embodiments of the present application fully considers the health characteristics of the elderly patient group in terms of physiology and psychology, and thus realizes that due to the decline of physiological functions and the decrease of immune system function, they are prone to various diseases, and various diseases may interact with each other, resulting in a complex pathological state in elderly patients. Therefore, in the embodiments of the present application, first, the values of 19 characteristic variables of the elderly patient are used to predict the classification of their complex pathological state. And the subtype categories of the predicted complex pathological state and the 19 characteristic variables of the elderly patient are used together to predict the survival probability in multiple cumulative time intervals. Experimental data show that compared with predicting only using 19 characteristic variables, this way of enhancing the subtype of the complex pathological state will make the prediction results of the survival probability in each cumulative time interval more accurate. Further, in the process of evaluating the comprehensive risk of the elderly patient, the probability distribution of the patient in each subtype category is considered again. And through the proper setting of the first weight of the probability distribution of each subtype category and the second weight of the survival probability in each cumulative time interval, not only can it reflect the influence of the complex pathological state subtype and the generation probability in different cumulative time intervals on the disease severity and comprehensive risk status of the elderly patient with finer granularity and more accuracy, thus making the evaluation result more accurate, but also it can make the method of the present application have a wider applicability in different elderly patient sample sets, and the evaluation result has a higher credibility. It can be seen that compared with various general scoring systems in the prior art, or systems that simply stitch together multiple models, the method in the embodiments of the present application fully and deeply applies the unique complex pathological state characteristics of the elderly patient group multiple times and from different angles in multiple stages. Therefore, it more accurately adapts to the application scenarios of the disease state and comprehensive risk assessment of elderly patients, the evaluation result has a higher accuracy, and it has good generalization ability on elderly patient sample sets with different data characteristics.
[0037] In some embodiments, when it is monitored that the values of the 19 characteristic variables of the elderly patient are updated, the probability distribution of each subtype category is re-predicted, and the values of the updated 20 characteristic variables are used to re-predict the survival probability of the elderly patient in at least one cumulative time interval. And based on the updated probability distribution of each subtype category and its corresponding first weight, and the updated survival probability of each cumulative time interval and its corresponding second weight, the updated comprehensive risk value and comprehensive risk level of the elderly patient are calculated. In this way, the health status and comprehensive risk of the elderly patient can be monitored more real-time, so as to take corresponding diagnosis and treatment measures in time.
[0038] In some embodiments, each subtype category of the complex pathological state is obtained by unsupervised clustering using the k-means algorithm (k-means clustering algorithm, an iterative clustering analysis algorithm) based on the values of the 19 characteristic variables of each elderly patient sample in a training sample set including multiple elderly patient samples.
[0039] Figure 3 FIG. shows a schematic diagram of the subtype categories of the complex pathological state obtained by k-means clustering according to an embodiment of the present application. In Figure 3 , according to the above 19 characteristic variables, based on elderly ICU inpatients over 65 years old in the MIMIC database (patients who died within the first day after entering the ICU were excluded), a data set containing 40,928 elderly patient samples was selected, and the k-means algorithm was used for unsupervised clustering. Among them, the most suitable number of clusters (also known as clusters) can be obtained, for example, by the elbow method. In this embodiment, the most suitable number of clusters is 4, corresponding to 4 typical classifications. After determining the number of clusters, each elderly patient sample can be included in these 4 clusters respectively, that is, a corresponding subtype category of the complex pathological state is assigned to each elderly patient sample. As Figure 3 shown, the sample points of different colors and different markers respectively belong to the clusters corresponding to different typical classifications after clustering, and each typical classification corresponds to a subtype category of the complex pathological state. Among them, the number of samples included in subtype category 1 - subtype category 4 are: 10,393, 9,862, 13,942, and 6,731 respectively.
[0040] It should be noted that when dividing each elderly patient sample into the corresponding cluster, clustering can be considered in a 1D - 19D space. Furthermore, the within-cluster distance metric and / or between-cluster distance metric of the 4 clusters in different dimensional spaces can be used to determine in which dimensional space the final sample clustering is performed. Only as an example, corresponding weights can be assigned to the within-cluster distance metric and the between-cluster distance metric, so as to more reasonably determine the final clustering dimensional space under the comprehensive consideration of the within-cluster aggregation degree and the between-cluster discrimination degree. Among them, the above within-cluster distance metric can usually be characterized by the average distance from the sample points within the cluster to its clustering center, and the between-cluster distance metric can be characterized, for example, by the average distance between the centroids of each cluster, etc. The specific method of the distance metric can adopt, for example, Euclidean distance, Manhattan distance, Manhattan square distance, Chebyshev distance, etc., and the present application does not make specific limitations on this. As Figure 3As shown, in this embodiment, the clustering in the two-dimensional space has better distance metric performance overall. Therefore, it is selected to cluster the samples of each elderly patient in the two-dimensional space. When visualizing, the PCA (Principal Component Analysis) method is used to select the visualization dimensions. Among them, dim1 and dim2 represent the weighted sum of the two orthogonal maximum variabilities in the data, and this weighted sum is obtained by the linear combination of the original variables.
[0041] In addition to determining the various subtype categories of the complex pathological state, it is also necessary to construct a complex pathological state subtype prediction model for predicting the complex pathological state subtype of a specific elderly patient based on the values of 19 characteristic variables of the elderly patient. Only as an example, the complex pathological state subtype prediction model in this application can be constructed based on the SketchBoost model, for example. Among them, SketchBoost is a fast gradient boosting decision tree model for solving multi-output problems. This model aims to solve the problem that the existing GBDT (Gradient Boosting Decision Tree) has poor scalability in the case of multi-dimensional output, and can accelerate the training process of GBDT in the multi-output scenario. Through a large number of experimental verifications in this application, it is found that compared with the more widely used XGBoost model currently, the SketchBoost model not only has a faster operation speed but also has higher accuracy when dealing with the complex pathological state subtype prediction task of elderly patients.
[0042] In some embodiments, the complex pathological state subtype prediction model can be trained in the following manner: First, as described above, after determining the various subtype categories of the complex pathological state, the subtype categories of the complex pathological state can be truthfully labeled for each elderly patient sample in the training sample set. Then, the constructed complex pathological state subtype prediction model is trained using each elderly patient sample with subtype category labels, so as to obtain a trained complex pathological state subtype prediction model. The specific training method can, for example, divide the data samples of elderly patients according to the ratio of training set: test set of 7:3, that is, 28,842 cases are used for model training and 12,086 cases are used for test evaluation, and the model that passes the test is externally verified in an external validation data set that is screened from the eICU database and completely contains 37,166 patients. Figure 4 The schematic diagram of the confusion matrix for prediction using the complex pathological state subtype prediction model according to the embodiment of the present application is shown. Figure 4 Adopt Figure 3The clustering method shown is used, and approximately 30% of the samples are randomly selected from the external validation dataset for prediction. From the values on the main diagonal of the confusion matrix of the obtained prediction results, it can be seen that the complex pathological state subtype prediction model of the embodiment of the present application can classify most samples into the subtype category to which the sample truly belongs, with ideal performance and meeting the requirements of practical applications.
[0043] In some other embodiments, the complex pathological state subtype prediction model can also be used to screen specific clustering methods. For example, using the sample set labeled with the subtype categories of the complex pathological states of the samples according to the clustering in different dimensional spaces, the complex pathological state subtype prediction model is trained respectively, and the model with the best performance is selected as the trained complex pathological state subtype prediction model. The present application does not specifically limit the method for selecting the clustering method.
[0044] After using the complex pathological state subtype prediction model to predict the complex pathological state subtype of the elderly patient to be evaluated, it is also necessary to further use the first survival probability prediction model to predict the survival probability of the elderly patient in at least one cumulative time interval. Among them, the cumulative time interval can usually include, for example, 0 - 30 days, 0 - 60 days, 0 - 360 days, etc., and can be set by the user as needed. The present application does not limit this.
[0045] In some embodiments, the first survival probability prediction model can be constructed based on, for example, the Random Survival Forest (RSF) model. In some other embodiments, the first survival probability prediction model can also be constructed based on other machine learning models such as the Cox model of neural networks (such as Deepsurv, Cox - CC, Cox - Time, etc.). The present application does not make specific limitations, but in the embodiments of the present application, through the comparison of experimental results, the random survival forest model is preferably used.
[0046] Next, it is necessary to use the training sample set to train the constructed first survival probability prediction model. In the training sample set, each elderly patient sample should be labeled with the true death time. In addition, since each elderly patient sample in the training sample set has been clustered and the clustering results are used for the labeling of the subtype categories of the complex pathological states of the samples in the sample set, the subtype category of the complex pathological state labeled for the sample can be used as the newly added 20th feature variable, and together with the 19 feature variables, it is used as the input of the first survival probability prediction model, while the true death time labeling is used as the reference object for the model output to complete the training of the first survival probability prediction model.
[0047] In the embodiments of the present application, the loss functions of the complex pathological state subtype prediction model and the first survival probability prediction model can be constructed based on multi-classification losses such as multi-classification cross-entropy loss. The specific optimization method is an ensemble learning method based on decision trees, which uses the gradient boosting algorithm. The prediction error of the model is reduced by gradually adding new trees. The update method of the first survival probability prediction model is constructed by evaluating the model performance using methods such as the log-rank test, maximization of log-likelihood, and mean squared error in survival analysis. The specific optimization methods can be, for example, gradient descent, stochastic gradient descent, Adam, RMSprop, etc. commonly used in machine learning algorithms. In addition, for the random survival forest model, the split points of the trees can be optimized by the log-rank test, and the model performance can be improved by using bootstrap sampling and random feature selection. The present application does not make specific limitations on this.
[0048] Figure 5(a) shows a comparison schematic diagram of the patient survival probability prediction curves obtained using 20 feature variables and 19 feature variables according to an embodiment of the present application. The first survival probability prediction model adopted in Figure 5(a) is constructed based on the random survival forest model. The dataset of the elderly patient samples is consistent with that in the Figure 3 embodiment, and each elderly patient sample in the dataset is labeled with the true value of the death time and has completed the true value labeling of the subtype categories of the complex pathological state.
[0049] As can be seen from Figure 5(a), the survival probability prediction curves obtained by predicting each sample using 20 feature variables and 19 feature variables do not completely overlap. Moreover, for some samples (such as sample 1), the survival probability is higher overall after considering the subtype categories of the complex pathological state, while for some samples (such as sample 3), the survival probability is lower after considering the subtype categories of the complex pathological state. Thus, it can be seen that different subtype categories of the complex pathological state do indeed exert an impact in the corresponding direction on the patient survival probability prediction result. In other words, the present application incorporates the subtype categories of the complex pathological state into the prediction of the elderly patient survival probability, emphasizing the internal connection between the subtype categories of the complex pathological state to which the elderly patients belong and their survival probability.
[0050] Figure 5(b) shows a schematic diagram of the confidence curves of the patient survival probabilities predicted using 20 feature variables and 19 feature variables according to an embodiment of the present application. In Figure 5(b), the C-index (concordance index) is used as the metric for confidence. As can be seen from the figure, the confidence of the prediction result of the first survival probability prediction model for the elderly patient survival probability using 20 feature variables is slightly higher than that when using 19 feature variables throughout the time axis, further verifying the improvement in the accuracy of the prediction of the elderly patient survival probability by incorporating the subtype categories of the complex pathological state into consideration in the present application.
[0051] After the first survival probability prediction model is trained, next, based on the parameters after the first survival probability prediction model is trained and the survival probability results of each cumulative time interval output by the first survival probability prediction model, an interpretability analysis model can be used to perform an interpretability analysis on the relationship between 20 feature variables and the confidence level of the output result of the first survival probability prediction model, and sort the contribution degrees of each feature variable to the confidence level of the output result of the first survival probability prediction model from high to low.
[0052] Among them, the interpretability analysis model can adopt any applicable model framework. Figure 6 The permutation importance calculation method in the Scikit-learn (also simply referred to as sklearn) machine learning toolkit of Python is used as the interpretability analysis tool, and a schematic diagram of the interpretability analysis results of 20 feature variables of the first survival probability prediction model according to the embodiments of the present application is drawn. It can be seen from Figure 6 that the subtype category of the complex pathological state ranks 6 in the importance ranking of each feature variable, indicating that the subtype category of the complex pathological state contributes greatly to the accurate prediction of the output result of the first survival probability prediction model. This not only further explains the internal reason why including the subtype category of the complex pathological state in the survival probability prediction can improve the model performance, but also proves the typicality and effectiveness of the present application in identifying and classifying the complex pathological state as a potential feature of the elderly patient group. In this case, the first survival probability prediction model can be retained to predict the survival probability of each cumulative time interval, and in the next step, the comprehensive risk value of the elderly patients can be further calculated.
[0053] In the embodiments of the present application, the comprehensive risk value of the elderly patients is obtained by performing a weighted operation on the probability distribution of each predicted subtype category and the survival probability of each cumulative time interval. R Specifically, it can be expressed as the following formula (1):
[0054] Among them, m is the number of subtype categories of the complex pathological state obtained by clustering, and n is the number of predicted cumulative time intervals; P pi is the predicted probability of the i-th subtype category, w pi is the corresponding first weight; P sj is the predicted survival probability of the j-th cumulative time interval, w sj is the corresponding second weight.
[0055] In some embodiments, each of the first weights and the second weights in formula (1) may be set according to Figure 7 the rules and steps for determining the first weights and the second weights as shown.
[0056] As Figure 7 shown, first, in step 701, after sorting the contribution degrees of 20 feature variables by using the interpretability analysis model, in combination with the sorting position of the 20th feature variable (i.e., the subtype category of the complex pathological state) in the sorting of the contribution degrees of each feature variable, set the weight ratio between the sum of each first weight and the sum of each second weight, wherein the higher the sorting position of the 20th feature variable, the larger the weight ratio. Only as an example, when the contribution degree sorting of the 20th feature variable is greater than 10, for example, it is 6, the weight ratio can be set to 1, that is, the sum of the first weights and the sum of the second weights are both 0.5. The specific contribution degree sorting threshold and the corresponding weight ratio correspondence can be set as needed. For example, it can be set and adjusted according to the experience of medical staff for the elderly patient group and / or according to the data characteristics (including data reliability, age distribution, etc.) in the elderly patient dataset. The present application does not make specific limitations on this.
[0057] Then, in step 702, specifically set each first weight. For example, after completing the subtype clustering of the complex pathological state in the training sample set, statistically calculate the average mortality rate of the elderly patient samples corresponding to each subtype category, and assign the first weights from high to low to the probability distributions of each subtype category according to the average mortality rate from high to low. It can be understood that the higher the average mortality rate of a subtype category, the greater the impact on the comprehensive risk of elderly patients. Setting its corresponding first weight higher can effectively reflect the high-risk state that elderly patients may have due to the potential and unpredictable impact of the complex pathological state.
[0058] Next, in step 703, specifically set each second weight. For example, the second weights corresponding to the survival probabilities of each cumulative time interval can be set in association with the time periods of each cumulative time interval, wherein the second weight corresponding to a longer cumulative time interval is not higher than the second weight corresponding to a shorter cumulative time interval. It can be understood that since the 19 feature variables of the elderly patients adopted in the embodiments of the present application are mainly measured and recorded within the first 24 hours after the patient is admitted to the hospital, the predicted short-term survival probability is more accurate and reliable than the long-term survival probability. Therefore, the second weight can be set higher. In some other embodiments, each second weight can also be set to be equal. The present application does not make specific limitations on this.
[0059] Figure 7The setting methods of the first weight and the second weight shown have been verified for their effectiveness in the sample dataset of elderly patients in the embodiments of this application. In the embodiments according to this application, it is allowed to adjust the first weight for each subtype category and the ratio of the first weight to the second weight. Therefore, even if the classification methods and interpretability analysis results of subtypes are completely different on different datasets, the evaluation of the comprehensive risk can still be maintained with high accuracy by adjusting the weights and the ratio of the weights. When specifically applied in other datasets, after setting the weights according to the above rules, the weight ratio of the sum of the first weights to the sum of the second weights and the specific values of the ratios of each first weight / second weight can be adaptively adjusted according to data such as the annotation of the death time of the elderly patient samples in the dataset. For example, the specific adjustment method can first use the data in a shorter cumulative time interval to predict the comprehensive risk of the patient, and compare and verify the overall trend of the predicted comprehensive risk of the patient with the survival / death data corresponding to a longer cumulative time interval, and adjust the weights accordingly, which will not be elaborated here.
[0060] In some other embodiments, if in the interpretability analysis result of the 20 feature variables by the interpretability analysis model, the ranking position of the 20th feature variable in the contribution degree of each feature variable is lower than the first threshold, for example, lower than the 10th, it can be considered that the 20th feature variable, that is, the subtype category of the complex pathological state, does not contribute much to the confidence of the output result of the first survival probability prediction model. In this case, it can be considered to construct a survival probability prediction model that only uses 19 feature variables. For example, a second survival probability prediction model can be constructed based on the random survival forest model, and 19 feature variables in each elderly patient sample with the true value annotation of the death time are used for training. And when the second survival probability prediction model is applied, only the values of the 19 feature variables that do not include the 20th feature variable are used as the model input to predict the survival probability of the elderly patient in at least one cumulative time interval. Then, the confidence levels of the output results of the first survival probability prediction model and the second survival probability prediction model are compared, and according to the comparison result, the one with the higher confidence level in the first survival probability prediction model and the second survival probability prediction model is selected as the final model for predicting the survival probability of the elderly patient in at least one cumulative time interval. The subsequent calculation process of the comprehensive risk value does not change due to the selection of the survival probability prediction model, that is, the calculation method of the comprehensive risk value R shown in formula (1) is still adopted.
[0061] According to the embodiments of this application, there is also provided a device for evaluating the complex pathological state and comprehensive risk of elderly patients. Figure 8 The partial composition schematic diagram of the device for evaluating the complex pathological state and comprehensive risk of elderly patients according to the embodiments of this application is shown.
[0062] AsFigure 8 As shown, device 800 includes at least interface 801 and at least one processor 802. Among them, interface 801 can be configured to receive the values of 19 characteristic variables of the elderly patient to be evaluated. The 19 characteristic variables include gender, age, body mass index, mechanical ventilation, dialysis, lowest hemoglobin, highest lactic acid, urine output, lowest blood oxygen saturation, average heart rate, average systolic blood pressure, average respiratory rate, highest blood glucose, lowest Glasgow Coma Scale score, highest blood urea nitrogen, lowest chloride ion concentration, SOFA score, delirium marker, and laboratory frailty index. As previously mentioned, the values of the above characteristic variables are usually measured and recorded within the first 24 hours after the elderly patient is admitted to the hospital. In the case of hospitalization, they will also be entered into a database such as a hospital information management system at any time. Therefore, in some embodiments, interface 801 can directly receive the values of the above 19 characteristic variables of the elderly patient from a unified hospital information system. Interface 801 can include a network adapter, cable connector, serial connector, USB connector, parallel connector, high-speed data transfer adapter (such as optical fiber, USB 3.0, Thunderbolt interface, etc.), wireless network adapter (such as WiFi adapter), telecommunications (3G, 4G / LTE, etc.) adapter, etc. The present application does not limit this. In some other embodiments, interface 801 can also obtain the values of the required 19 characteristic variables from multiple sources through different interface types, including receiving immediate input executed by the user through interactive operations such as typing and voice. In this case, device 800 can also include input devices (not shown) such as a keyboard, mouse, trackball, and / or a microphone device that can convert voice into an electrical signal, so that device 800 can support the above various operations of the user. In some other embodiments, these input devices can also be integrally provided with the display medium (not shown) in device 800, for example, performing various operations on the touch screen surface of the display medium, and the present application does not limit this.
[0063] In some embodiments, at least one processor 802 can be configured to execute the steps of the method for evaluating the complex pathological state and comprehensive risk of elderly patients as described in various embodiments of the present application.
[0064] In some embodiments, the processor 802 may be a processing device including more than one general-purpose processing device, such as a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), etc. More specifically, the processor may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor running other instruction sets, or a processor running a combination of instruction sets. The processor may also be more than one dedicated processing device, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a system on a chip (SoC), etc.
[0065] According to an embodiment of the present application, there is also provided a non-transitory computer-readable storage medium having computer-executable instructions stored thereon, and when the computer-executable instructions are executed by a processor, the steps of the method for evaluating the complex pathological state and comprehensive risk of elderly patients described in various embodiments of the present application are implemented.
[0066] In some embodiments, the above-mentioned non-transitory computer-readable medium may be, for example, a read-only memory (ROM), a random access memory (RAM), a phase change random access memory (PRAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), an electrically erasable programmable read-only memory (EEPROM), other types of random access memory (RAM), a flash drive or other forms of flash memory, a cache, a register, a static memory, a compact disc read-only memory (CD-ROM), a digital versatile disc (DVD) or other optical memory, a cassette tape or other magnetic storage device, or any other possible non-transitory medium used to store information or instructions accessible by a computer device, etc.
[0067] Moreover, although exemplary embodiments have been described herein, the scope includes any and all embodiments based on the present application having equivalent elements, modifications, omissions, combinations (e.g., schemes that cross various embodiments), adaptations, or alterations. The elements in the claims will be broadly interpreted based on the language employed in the claims and are not limited to the examples described in this specification or during the implementation of the present application, and the examples will be interpreted as non-exclusive. Therefore, this specification and the examples are intended to be considered only as examples, and the true scope and spirit are indicated by the full scope of the claims and their equivalents.
[0068] The above description is intended to be illustrative and not restrictive. For example, the above examples (or one or more aspects thereof) may be used in combination with each other. For example, those of ordinary skill in the art may use other embodiments when reading the above description. Additionally, in the above detailed description, various features may be grouped together to simplify the present application. This should not be construed as an intention that the disclosed features not claimed are necessary for any claim. On the contrary, the subject matter of the present application may be less than all of the features of a particular disclosed embodiment. Thus, the claims are hereby incorporated by way of example or embodiment into the detailed description, where each claim stands on its own as a separate embodiment, and it is contemplated that these embodiments may be combined with each other in various combinations or permutations. The scope of the present application should be determined with reference to the claims and the full scope of equivalents to which these claims are entitled.
[0069] The above embodiments are only exemplary embodiments of the present application and are not used to limit the present application. The protection scope of the present application is defined by the claims. Those skilled in the art may make various modifications or equivalent substitutions within the spirit and protection scope of the present application, and such modifications or equivalent substitutions should also be regarded as falling within the protection scope of the present application.
Claims
1. A method for assessing complex pathological conditions and comprehensive risks in elderly patients, characterized in that: Included by the processor: Obtaining values of 19 characteristic variables of the elderly patient to be evaluated, the 19 characteristic variables including gender, age, body mass index, mechanical ventilation, dialysis, lowest hemoglobin, highest lactate, urine volume, lowest blood oxygen saturation, average heart rate, average systolic blood pressure, average respiratory rate, highest blood glucose, lowest Glasgow Coma Scale score, highest urea nitrogen, lowest chloride ion concentration, SOFA score, delirium markers, and laboratory frailty index; Based on the values of the 19 characteristic variables, the probability distribution of each subtype category of the complex pathological state of the elderly patient is predicted using the complex pathological state subtype prediction model, and the subtype category of the complex pathological state to which the elderly patient belongs determined based on the probability distribution of each subtype category is used as the value of the 20th characteristic variable; Based on the values of the 20 characteristic variables, the first survival probability prediction model is used to predict the survival probability of the elderly patient in at least one cumulative time interval; Based on the probability distribution of each subtype category and its corresponding first weight, as well as the survival probability of each cumulative time interval and its corresponding second weight, the comprehensive risk value of the elderly patient is calculated so that the comprehensive risk level of the elderly patient is given based on the comprehensive risk value.
2. The method according to claim 1, characterized in that Each subtype category of the complex pathological state is obtained by performing unsupervised clustering using the k-means algorithm based on the values of the 19 characteristic variables of each elderly patient sample in a training sample set including multiple elderly patient samples.
3. The method according to claim 2, characterized in that The complex pathological state subtype prediction model is constructed based on the SketchBoost model, and After labeling the subtype categories of complex pathological conditions for each elderly patient sample in the training sample set, the training is performed using each elderly patient sample with subtype category labeling.
4. The method according to claim 2, characterized in that: The first weight corresponding to the probability distribution of each subtype category is determined according to the following steps: In the training sample set, the average mortality rate of elderly patient samples corresponding to each subtype category is counted, and the probability distribution of each subtype category is assigned a first weight from high to low according to the average mortality rate from high to low.
5. The method according to claim 1, characterized in that The second weight corresponding to the survival probability of each cumulative time interval is set in association with the time period of each cumulative time interval, wherein the second weight corresponding to a longer cumulative time interval is not higher than the second weight corresponding to a shorter cumulative time interval.
6. The method according to claim 2, characterized in that Each elderly patient sample in the training sample set is labeled with a true value of death time; The first survival probability prediction model is constructed based on a random survival forest model and is obtained by training using samples of elderly patients with true value annotations of death time and subtype category annotations.
7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: Based on the parameters of the first survival probability prediction model after training, and the survival probability results of each cumulative time interval output by the first survival probability prediction model, an interpretable analysis model is used to perform an interpretable analysis on the relationship between the 20 characteristic variables and the confidence of the output results of the first survival probability prediction model, and the contribution of each characteristic variable to the confidence of the output results of the first survival probability prediction model is ranked from high to low.
8. The method according to claim 7, characterized in that The method further comprises: In combination with the ranking position of the 20th characteristic variable in the contribution degree of each characteristic variable, a weight ratio between the sum of each first weight and the sum of each second weight is set, wherein the higher the ranking position of the 20th characteristic variable, the greater the weight ratio.
9. The method according to claim 7, characterized in that: The method further comprises: When the ranking position of the contribution of the 20th feature variable in each feature variable is lower than the first threshold, a second survival probability prediction model is constructed based on the random survival forest model, and each elderly patient sample with the true value of the death time is used for training, and, Based only on the values of the 19 characteristic variables excluding the 20th characteristic variable, the second survival probability prediction model is used to predict the survival probability of the elderly patient in at least one cumulative time interval; The confidence levels of the output results of the first survival probability prediction model and the second survival probability prediction model are compared, and a model with higher confidence level is selected to predict the survival probability of elderly patients in at least one cumulative time interval.
10. A device for evaluating complex pathological conditions and comprehensive risks of elderly patients, characterized in that: include: The interface is configured to: receive values of 19 characteristic variables of an elderly patient to be evaluated, wherein the 19 characteristic variables include gender, age, body mass index, mechanical ventilation, dialysis, lowest hemoglobin, highest lactate, urine output, lowest blood oxygen saturation, average heart rate, average systolic blood pressure, average respiratory rate, highest blood glucose, lowest Glasgow Coma Scale, highest urea nitrogen, lowest chloride ion concentration, SOFA score, delirium markers, and laboratory frailty index; At least one processor configured to execute the method for assessing complex pathological conditions and comprehensive risks of elderly patients as described in any one of claims 1-9.
Citation Information
Cited By
Elderly hospitalized patient weakness risk prediction model construction method and prediction system
CN120565093A