Cascade evaluation method for prognostic risk of mental disease
Through the cascading evaluation method, first- and second-level risk assessment models are constructed, and multi-level risk assessment is used to perform multi-level risk assessment, which solves the problems of single-dimensional and multi-category risk assessment in the existing technology, and improves the accuracy and comprehensiveness of prognostic risk assessment of mental illness.
Patent Information
- Application Number
- CN202411501802.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-25
- Publication Date
- 2025-06-06
AI Technical Summary
The existing prognostic risk assessment method for mental illness is a single dimension, and it is impossible to fully consider the various possible risks of patients. It lacks multi-level risk assessment and data-driven methods, and there are learning difficulties in multi-category risk assessment.
The cascade evaluation method is used to obtain the patient's clinical data for preprocessing and feature screening, and a primary and secondary risk assessment model is constructed, and a neural network model is used for multi-level risk assessment.
It improves the accuracy of various prognostic risks, solves the problem of learning difficulties in multi-classification models, provides a more comprehensive risk assessment, and supports clinical decision-making and reasonable allocation of medical resources.
Smart Images

Figure CN120108712A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical artificial intelligence technology, and in particular to a cascade assessment method for the prognosis risk of mental illness. Background Art
[0002] Current methods for assessing the prognosis risk of mental illness are usually single-dimensional and only focus on some risk factors. The limitation of this method is that they cannot fully consider the various possible risks of patients, nor can they provide multi-level risk assessment, and may ignore important risk factors that some patients may face. Current methods for assessing the prognosis risk of mental illness usually do not make full use of modern computer science and data analysis technology, and lack comprehensive, data-driven methods for assessing the prognosis risk of mental illness. In addition, there are many types of prognosis risks for mental illness, and there are certain limitations in using a single model for multi-category risk assessment, such as the difficulty of the model in learning when dealing with multi-classification problems. The use of a multi-model, hierarchical strategy can effectively solve the above problems and help improve the accuracy of prognosis risk assessment for mental illness. To this end, we propose a cascaded method for assessing the prognosis risk of mental illness. Summary of the invention
[0003] In order to address the deficiencies in the above-mentioned prior art, the purpose of the present invention is to provide a cascade assessment method for the prognosis risk of mental illness, which solves the problem of difficulty in learning multi-classification models and improves the accuracy of various prognosis risks.
[0004] The technical solution adopted by the present invention to solve the technical problem is: a cascade assessment method for the prognosis risk of mental illness, comprising the following steps:
[0005] Obtain clinical data of the patient to be predicted;
[0006] Preprocess and feature screen the acquired clinical data to obtain feature data;
[0007] Constructing primary risk assessment model and secondary risk assessment model;
[0008] The characteristic data is input into a primary risk assessment model, and the primary risk assessment model outputs a risk value for drug side effects, a risk value for complications, a risk value for deterioration or recurrence, and a risk value for death;
[0009] When the drug side effect risk value and the complication risk value are greater than the set threshold, the drug side effect risk value and the complication risk value are input into the secondary risk assessment model to obtain the prognostic risk value, and the prognostic risk analysis result is output;
[0010] When the drug side effect risk value and complication risk value are less than the set threshold, the primary risk assessment model outputs the prognostic risk analysis results.
[0011] Optionally, the clinical data includes demographic data, biochemical test data, clinical diagnosis data, medication data, genetic disease history and prognosis information within three years.
[0012] Optionally, the preprocessing of the clinical data comprises the steps of:
[0013] Completely duplicate clinical data were eliminated;
[0014] Clinical data with missing data rate greater than the set threshold were deleted.
[0015] Optionally, the method for feature screening of the clinical data is:
[0016] Use the XGBoost model to rank the feature importance of the preprocessed clinical data;
[0017] Screening was performed based on sorted clinical data.
[0018] Optionally, the drug side effect risk value includes an insomnia risk value, an abdominal pain risk value, a headache risk value, a dizziness risk value and a drug poisoning risk value, and the complication risk value includes a hypertension risk value, a hyperglycemia risk value, a liver damage risk value, a cardiovascular disease risk value and a metabolic system risk value.
[0019] Optionally, both the primary risk assessment model and the secondary risk assessment model adopt a neural network model, and the steps of constructing the neural network model are:
[0020] The prognostic data are converted into a multi-hot format as sample labels, where the prognostic data include the risk value of drug side effects, the risk value of complications, the risk value of deterioration or recurrence, and the risk value of death;
[0021] The prognostic data were divided into training set, validation set and test set in a ratio of 6:2:2. The training set and validation set were used for model training and parameter adjustment, and the test set was used for model prediction effect evaluation.
[0022] Determine the hidden layer parameters and number of neurons of the neural network through a grid search strategy;
[0023] The model is trained using the back propagation algorithm, and the loss function is the binary cross entropy function BCELoss, with the formula:
[0024] ;
[0025] Among them, y i is the sample binary label 0 or 1, for y i The corresponding predicted probability value, n is the number of predicted categories.
[0026] The present invention provides a cascaded method for assessing the prognosis risk of mental illness. In view of the wide variety of prognosis risks, the present invention solves the problems of difficulty in learning multiple classification models through a multi-model and hierarchical strategy, thereby improving the accuracy of various prognosis risks. At the same time, the present invention can provide clinical decision support, rationally allocate medical resources, reduce adverse reactions, and lower various prognosis risks. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a schematic diagram of the process of the present invention;
[0028] Figure 2 It is a schematic diagram of the prediction process of the primary risk assessment model and the secondary risk assessment model of the present invention. DETAILED DESCRIPTION
[0029] The present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It is also necessary to explain that, for ease of description, only the parts related to the invention are shown in the accompanying drawings.
[0030] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0031] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0032] Figure 1 A cascade assessment method for the prognosis risk of mental illness is disclosed, comprising the following steps:
[0033] S1. Obtain clinical data of the patient to be predicted.
[0034] The patients to be predicted are patients with mental illness who are treated with psychotropic drugs. Their clinical data include demographic data, biochemical test data, diagnostic data, medication data, genetic history, and prognosis information within three years.
[0035] S2. Preprocess and feature screen the acquired clinical data to obtain feature data.
[0036] In the present invention, the preprocessing method of clinical data is:
[0037] S21. Eliminate completely duplicate clinical data;
[0038] S22. Delete clinical data with a missing rate greater than a set threshold. At the same time, fill in clinical data with a missing rate less than the set threshold with the median data of the data.
[0039] Since the original clinical data obtained has the characteristics of missing data, duplication and high dimensionality, the missing values of the data will cause a lot of information loss, and the repeated values may cause model bias, thereby reducing the generalization ability of the model. The original high-dimensional data contains a lot of redundant or useless information, and has a great risk of overfitting and computational overhead. Therefore, clinical data needs to be preprocessed and feature screened before modeling.
[0040] In one embodiment of the present invention, the feature screening of clinical data uses an XGBoost model to sort the feature importance of the preprocessed clinical data, and then screens the sorted clinical data after the data is sorted.
[0041] The present invention sets a threshold for the missing data rate. For example, in one embodiment, the threshold is set to 20%, that is, clinical data with a missing rate greater than 20% are deleted so that they do not participate in model training, and clinical data with a missing rate less than 20% are filled with the median of the clinical data. After completing the missing data processing, completely identical clinical data are eliminated, that is, duplicate clinical data are removed.
[0042] After completing the above data processing, the obtained data is the feature data required for the subsequent step.
[0043] S3. Construct the first-level risk assessment model and the second-level risk assessment model.
[0044] In the present invention, the primary risk assessment model is used to analyze the characteristic data and output the drug side effect risk value, complication risk value, deterioration or recurrence risk value and death risk value. The secondary risk assessment model is used to evaluate the drug side effect risk value and complication risk value, and finally output various prognosis risk values.
[0045] The risk values of drug side effects include insomnia risk value, abdominal pain risk value, headache risk value, dizziness risk value and drug poisoning risk value; the risk values of complications include hypertension risk value, hyperglycemia risk value, liver damage risk value, cardiovascular disease risk value and metabolic system risk value.
[0046] In the present invention, both the primary risk assessment model and the secondary risk assessment model adopt a neural network model, and the construction steps of the two are the same, which are:
[0047] 1) Convert the prognostic data into a multi-hot format as sample labels, where the prognostic data information includes the risk value of drug side effects, the risk value of complications, the risk value of deterioration or recurrence, and the risk value of death;
[0048] 2) The prognostic data were divided into training set, validation set and test set in a ratio of 6:2:2. The training set and validation set were used for model training and parameter adjustment, and the test set was used for evaluating the prediction effect of the model.
[0049] 3) Determine the hidden layer parameters and the number of neurons of the neural network through a grid search strategy. In this embodiment, the final network structures of the primary model and the secondary model are [20, 60, 20, 4] and [20, 40, 20, 5] respectively;
[0050] 4) The back propagation algorithm is used for model training, and the loss function is the binary cross entropy function BCELoss, the formula is as follows:
[0051]
[0052] in is the sample binary label 0 or 1, for The corresponding predicted probability value is is the number of predicted categories.
[0053] In the present invention, prognostic risk labels are extracted from three years of clinical data information, involving three labels of the primary risk assessment model, which are illustrated below with examples:
[0054] 1) The first-level risk assessment model label is [1, 0, 1, 0], indicating that the patient experienced drug side effects and complications after treatment, but did not experience disease progression, recurrence, or death;
[0055] 2) The first-level risk assessment model label is [1, 1, 0, 0, 0], indicating that there are drug side effects. The patient experienced insomnia and abdominal pain after treatment, but did not experience headache, dizziness, or drug poisoning symptoms;
[0056] 3) The first-level risk assessment model label is [1, 0, 0, 0, 0], indicating that there are complications. After treatment, the patient only develops hypertension, without hyperglycemia, liver damage, cardiovascular disease, and metabolic system disease.
[0057] In the above risk label extraction process, the selected clinical data are all common symptoms that have a relatively large impact on treatment. After the clinical data preprocessing is completed, the screened clinical data and the corresponding prognostic risk labels form a data set, and the XGBoost model is trained using this data set. After the training is completed, the feature importance ranking of each variable is output, and the top 20 features are retained as sample data for the final modeling.
[0058] In the present invention, after the characteristic data is input into the primary risk assessment model, the primary risk assessment model outputs the drug side effect risk value, complication risk value, deterioration or recurrence risk value and death risk value. At the same time, the drug side effect risk value and complication risk value are input into the secondary risk assessment model, and the secondary risk assessment model outputs various prognostic risk values of the patient, and comprehensively obtains the prognostic risk analysis results, such as Figure 2 As shown, the specific steps are as follows:
[0059] S31. Input the characteristic data into the primary risk assessment model to obtain the risk value of each category, including the risk value of drug side effects , deterioration or recurrence risk , Complications risk value and mortality risk .
[0060] S32, when and When it is greater than the set threshold, the above corresponding characteristic data will continue to be input into the secondary risk assessment model to obtain further subdivided prognostic risk values.
[0061] For example, in one embodiment of the present invention, the threshold is set to 0.5. , it indicates that there is a risk value for drug side effects. At this time, the specific type of side effects needs to be determined through the secondary risk assessment model. The output of this model is the risk value for insomnia. , abdominal pain risk value Headache risk value , dizziness risk value and drug poisoning risk .
[0062] when , it indicates that there is a complication risk value. At this time, the specific complication type needs to be determined by the complication risk value. The output of the model is the hypertension risk value. , high blood sugar risk value , liver injury risk value , cardiovascular disease risk value and metabolic risk .
[0063] S33, when and When it is less than the set threshold, the corresponding prognostic risk value is obtained.
[0064] S34. The predicted values of the comprehensive first-level risk assessment model and the second-level risk assessment model are the results of prognostic risk analysis.
[0065] The predicted values of the integrated primary risk assessment model and the secondary risk assessment model are the final prognostic risk analysis results. The comprehensive prognostic risk values of various types are [P 11 (P 111 , P 112 , P 113 , P 114 , P 115 ), P 12 , P 13 (P 131 , P 132 , P 133 , P 134 , P 135 ), P 14 ].
[0066] The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the above features are replaced with (but not limited to) technical features with similar functions disclosed in the present application.
[0067] Except for the technical features described in the specification, the remaining technical features are known technologies to those skilled in the art. In order to highlight the innovative features of the present invention, the remaining technical features will not be described here in detail.
Claims
1. A cascade assessment method for the prognosis risk of mental illness, characterized in that: The following steps are involved: Obtain clinical data of the patient to be predicted; Preprocess and feature screen the acquired clinical data to obtain feature data; Constructing primary risk assessment model and secondary risk assessment model; The characteristic data is input into a primary risk assessment model, and the primary risk assessment model outputs a risk value for drug side effects, a risk value for complications, a risk value for deterioration or recurrence, and a risk value for death; When the drug side effect risk value and the complication risk value are greater than the set threshold, the drug side effect risk value and the complication risk value are input into the secondary risk assessment model to obtain the prognostic risk value, and the prognostic risk analysis result is output; When the drug side effect risk value and complication risk value are less than the set threshold, the primary risk assessment model outputs the prognostic risk analysis results.
2. A cascade assessment method for prognosis risk of mental illness according to claim 1, characterized in that: The clinical data include demographic data, biochemical test data, clinical diagnosis data, medication data, genetic disease history and prognosis information within three years.
3. A cascade assessment method for prognosis risk of mental illness according to claim 2, characterized in that: The preprocessing of the clinical data comprises the steps of: Completely duplicate clinical data were eliminated; Clinical data with missing data rate greater than the set threshold were deleted.
4. A cascade assessment method for prognosis risk of mental illness according to claim 3, characterized in that: The method for feature screening of the clinical data is: Use the XGBoost model to rank the feature importance of the preprocessed clinical data; Screening was performed based on sorted clinical data.
5. A cascade assessment method for prognosis risk of mental illness according to claim 4, characterized in that: The drug side effect risk values include insomnia risk value, abdominal pain risk value, headache risk value, dizziness risk value and drug poisoning risk value, and the complication risk values include hypertension risk value, hyperglycemia risk value, liver damage risk value, cardiovascular disease risk value and metabolic system risk value.
6. A cascade assessment method for prognosis risk of mental illness according to claim 5, characterized in that: The first-level risk assessment model and the second-level risk assessment model both adopt a neural network model, and the steps of constructing the neural network model are as follows: The prognostic data are converted into a multi-hot format as sample labels, where the prognostic data include the risk value of drug side effects, the risk value of complications, the risk value of deterioration or recurrence, and the risk value of death; The prognostic data were divided into training set, validation set and test set in a ratio of 6:2:
2. The training set and validation set were used for model training and parameter adjustment, and the test set was used for model prediction effect evaluation. Determine the hidden layer parameters and number of neurons of the neural network through a grid search strategy; The model is trained using the back propagation algorithm, and the loss function is the binary cross entropy function BCELoss, with the formula: ; Among them, y i is the sample binary label 0 or 1, for y i The corresponding predicted probability value, n is the number of predicted categories.