Cerebral stroke prognosis prediction model training method based on quantum computing and artificial intelligence

Through a prediction model combining quantum computing and artificial intelligence, indicators related to stroke prognosis are screened out and indicators related to stroke prognosis are constructed, which solves the limitations of traditional models in variable correlation and high-dimensional data calculation efficiency, and achieves higher prediction accuracy and model generalization capabilities.

CN120356685APending Publication Date: 2025-07-22THE AFFILIATED HOSPITAL OF XUZHOU MEDICAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510463892.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing stroke prognosis prediction model has limitations in correlation analysis between variables and dynamic interaction characteristics, which cannot meet the needs of causal inference, and the calculation efficiency of high-dimensional data is low. Traditional machine learning models need to improve their prediction accuracy and generalization capabilities.

Method used

Using prediction model training methods based on quantum computing and artificial intelligence, the white blood cell community parameters were detected through flow cytometry technology, combined with Lasso regression, recursive feature elimination and quantum computing, indicators closely related to stroke prognosis were screened, integrated models such as LightGBM were constructed, and high-dimensional data feature screening and prediction were carried out.

Benefits of technology

The prediction accuracy and model generalization ability of stroke prognosis were significantly improved. The LightGBM model performed best in tenfold cross-validation and external verification. The introduction of leukocyte community parameters improved the sensitivity and objectivity of predicting neurological damage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356685A_ABST
    Figure CN120356685A_ABST
Patent Text Reader

Abstract

The invention relates to a stroke prognosis prediction model training method based on quantum computing and artificial intelligence, and relates to the technical field of prognosis effect prediction. According to the method, leukocyte community parameters are introduced into an acute ischemic stroke prognosis model for the first time, morphological function characteristics of leukocytes are detected through a flow cytometry, and compared with traditional inflammation indexes, the leukocyte prognosis model has higher sensitivity and objectivity and becomes an important index for predicting neurological function impairment; the indexes closely related to cerebral apoplexy prognosis are screened out through quantum calculation for the first time, and a new thought is provided for high-dimensional data feature screening; meanwhile, by comparing 14 traditional machine learning, integrated learning and deep learning models, it is judged that the LightGBM model is an acute ischemic stroke prognosis prediction model, and compared with the prior art, the LightGBM model is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of prognosis effect prediction, and particularly to a training method for a stroke prognosis prediction model based on quantum computing and artificial intelligence. Background Art

[0002] Stroke, commonly known as a stroke, is divided into a vascular occlusion type (ischemic stroke) and a vascular rupture type (hemorrhagic stroke), with the former accounting for about 60%-70% of clinical cases, mainly due to thrombus formation or embolism events in cerebral blood vessels. Stroke has the characteristics of high incidence, disability rate, recurrence rate and mortality rate. With the aging of the population structure and the evolution of lifestyle, the incidence of cerebrovascular diseases has shown a continuous upward trend and has now become the second leading cause of death and disability globally. More seriously, three-quarters of survivors have permanent neurological damage, and the direct medical expenditure exceeds 40 billion yuan.

[0003] Among them, ischemic stroke is a complex disease involving multiple pathophysiological pathways. Common poor post-disease prognoses mainly include death, disability and stroke recurrence, etc. For such diseases, a variety of treatment methods have been developed, including intravenous thrombolysis and endovascular thrombectomy, etc., which can effectively improve the prognosis. Timely and accurate prognosis prediction is of great significance for guiding doctors' treatment decisions and patient selection. Prognosis prediction also helps to select the best rehabilitation strategy for patients in the revascularization stage.

[0004] The current mainstream prediction models mainly include three types of statistical models: Poisson distribution model, Cox proportional hazards regression model and Logistic regression model. Among them, the Logistic model constructs a prediction equation through the classic path of single-factor screening combined with multi-factor modeling, while the Cox model shows statistical efficacy in functional outcome prediction through risk ratio calculation. However, traditional models have two limitations: the correlation analysis between variables cannot meet the needs of causal inference, and the fixed coefficient hypothesis conflicts with the dynamic interaction characteristics of clinical variables. With the development of medical informatics, non-parametric models based on machine learning can integrate high-dimensional heterogeneous data and break through the variable capacity limitation of traditional models. The machine learning prediction system constructed by the Brugnara team is of milestone significance (Brugnara et al., Stroke, 2020, 51(12): 3541-3551). By integrating pre-vascular treatment clinical characteristics, CT parameters (ASPECTS score, ischemic lesion volume) and multi-modal imaging characteristics (penumbra / infarct core volume), it realizes the accurate prediction of the 90-day mRS score (AUC = 0.856), but it relies on radiomics and does not solve the problem of high-dimensional data calculation efficiency.

[0005] As an emerging information processing paradigm, the technological development of quantum computing is driving the innovation of computational science. Through the superposition and entanglement effects of quantum states, the optical quantum system provides an exponentially accelerated solution for combinatorial optimization problems. The optical quantum architecture based on the Coherent Ising Model (CIM) has achieved quantum state control at room temperature, and its scalability to 100,000 qubits provides a new idea for high-dimensional clinical data feature screening. Existing research shows that the parallel state space search mechanism of quantum computing can break through the efficiency bottleneck of traditional algorithms and has potential advantages in biomarker identification.

[0006] In view of this, the present invention is hereby proposed. Summary of the Invention

[0007] The purpose of the present invention is to provide a method for screening clinical indicators related to the neurological function prognosis of stroke patients three months after discharge, and based on these indicators, use machine learning, deep learning, and quantum computing to construct a prediction model, deeply explore the key factors leading to poor prognosis of stroke patients, and thus provide reference for clinicians when choosing treatment plans.

[0008] To achieve the above object, the present invention provides the following technical solutions: A method for training a stroke prognosis prediction model based on quantum computing and artificial intelligence, characterized by comprising the following steps: Step S1: Establish a data case library; collect the electronic medical records of patients, and collect the electronic medical records of patients with follow-up records and a primary diagnosis of acute ischemic stroke. Step S2: Extract the medical feature data of the patients: extract the medical features of acute ischemic stroke from the electronic medical record data obtained in Step S1 to obtain medical features and medical feature values; the medical features of acute ischemic stroke include demographic features, vital signs, severity of illness, blood routine, white blood cell colony parameters, and blood biochemistry, which are used as prediction indicators; the medical feature values include the specific values of the demographic features, vital signs, severity of illness, blood routine, white blood cell colony parameters, and blood biochemistry medical features. Step S3: Extract the target result feature data; extract the post-stroke score three months after the patient is discharged; the binary classification results of the post-stroke score are: good prognosis for a post-stroke score of 0-2, and poor prognosis for a post-stroke score of 3-6; after scoring all the patients' post-stroke scores, obtain the big data of the clinical manifestations of acute ischemic stroke. Among them, the poor prognosis includes moderate disability, severe disability, and death. Step S4: Feature data standardization and data cleaning: Standardize the feature data of the big data on the clinical manifestations of acute ischemic stroke obtained in Step S3. Adopt a missing data strategy to exclude patients with more than 20% missing feature variables. Fill the remaining missing feature data in the mode of the existing data of the same feature. Use the Expectation-Maximization algorithm for filling, and use the Wilcoxon rank-sum test to judge the filling effect of the missing values. If the P-value of the difference in the distribution of the feature data before and after filling the missing values is greater than 0.05, the filling effect is good; Step S5: New feature construction and data balancing: Construct new inflammation indices and blood glucose and lipid metabolism indicators from the medical feature data obtained in Step S2. Implement undersampling and oversampling of the data through the RandomUnderSampler and SMOTENC algorithms in Python; Step S6: Feature selection: Select a feature selection method for feature screening. The feature selection methods include Lasso regression, recursive feature elimination, and quantum computing. Take the intersection of the features selected by the three as the key features; Step S7: Establish 14 machine learning prediction models: Based on the key features, establish 14 machine learning prediction models, including logistic regression, quadratic discriminant analysis, naive Bayes, Gaussian process regression, K-nearest neighbor algorithm, decision tree, AdaBoost, Bagging, gradient boosting, extremely randomized tree, XGBoost, CatBoost, LightGBM, and multi-layer perceptron; Step S8: Compare the results obtained in Step S7 and select a model.

[0009] Further, in Step S2, the demographic characteristics include: age, gender, height, weight, body mass index; The vital signs and disease severity include systolic blood pressure, diastolic blood pressure, mean arterial pressure, and admission NIHSS score; The blood routine includes white blood cell count, neutrophil count, lymphocyte count, monocyte count, platelet count, mean platelet volume, C-reactive protein, hemoglobin concentration, hematocrit, mean corpuscular volume, mean corpuscular hemoglobin content, mean corpuscular hemoglobin concentration; The white blood cell community parameters include the lateral scattered light intensity of the Neu region on the DIFF channel, the lateral fluorescence intensity of the Neu region on the DIFF channel, the forward scattered light intensity of the Neu region on the DIFF channel, the lateral scattered light intensity of the Lym region on the DIFF channel, the lateral fluorescence intensity of the Lym region on the DIFF channel, the forward scattered light intensity of the Lym region on the DIFF channel, the lateral scattered light intensity of the Mon region on the DIFF channel, the lateral fluorescence intensity of the Mon region on the DIFF channel, the forward scattered light intensity of the Mon region on the DIFF channel, and the standard deviation of platelet distribution width; The blood biochemistry includes blood glucose, triglyceride, total cholesterol, high-density lipoprotein cholesterol, serum creatinine, alanine aminotransferase, albumin, and uric acid.

[0010] Further, in step S5, the inflammatory indices and construction methods include NLR = Neu / Lym, MLR = Mon / Lym, PLR = PLT / Lym, SII = PLT × Neu / Lym, SIRI = Mon × Neu / Lym, AISI = Neu × PLT × Mon / Lym, PIV = Mon × PLT × Neu, HALP = HGB × Alb × Lym / PLT, IBI = FR-CRP × Neu / Lym, NPHR = Neu / WBC × 100 / HGB, and ALI = BMI × Alb / NLR; The blood glucose and lipid metabolism indices and construction methods include TyG = ln[FBG (mg / dL) × TG (mg / dL)] / 2, NHHR = (TC - HDL) / HDL, NPAR = Neu / WBC × 100 / Alb, PHR = PLT / HDL, MHR = Mon / HDL, UHR = SUA / HDL, and TyG-BMI = BMI × TyG.

[0011] Further, in step S6: Lasso regression: By calculating the conditional probability that a sample is assigned to the poor prognosis group and matching it with the good prognosis group with similar characteristics, a good prognosis group with characteristics similar to those of the poor prognosis group is obtained, so that the covariates tend to be balanced after the two groups of patients are matched, in order to better study the relationship between the characteristics and the target outcome; Recursive feature elimination: Train an extremely randomized tree model on the candidate feature set with all the indicators, representing all the candidate clinical indicators; calculate the importance measure of each feature in the model and sort the feature set according to the size of this measure; eliminate the feature corresponding to the smallest measure to obtain a new feature subset; repeat the "training - calculation - elimination" process until the optimal feature set is obtained; Quantum computing: Establish a QUBO model, that is, a quadratic unconstrained binary optimization model.

[0012] Furthermore, in step S7: The logistic regression uses the sigmoid function. Through this function, function values greater than 0.5 are assigned to the label 1 indicating poor prognosis, and function values less than or equal to 0.5 are assigned to the label 0 indicating good prognosis; this algorithm can see the impact of different indicators on the prognosis of stroke patients from the weights of features, and the output result is a probability value.

[0013] The quadratic discriminant analysis: Calculate the mean and covariance matrix of each category. For a new observation, calculate its posterior probability of belonging to each category, and classify the observation into the category with the highest posterior probability; The naive Bayes simplifies the calculation of the aforementioned conditional probability through the independence assumption of features; The Gaussian process regression: Define the similarity between input points through a covariance function, and output the predicted mean and variance corresponding to the prediction points; The K-nearest neighbor algorithm: After inputting data without labels, compare each feature of this data without labels with the corresponding features of the data in the sample set, and then extract the classification labels of the data with the closest features in the sample; The decision tree: Each node of the decision tree represents a judgment condition of a feature, each branch represents a judgment result, and each leaf node corresponds to the predicted category or value. The construction process of the decision tree is based on the features of the training data, and the optimal feature is selected for data division. By continuously dividing features until all samples are correctly classified or a preset stop condition is reached; The AdaBoost: Set the weight of each sample in the training set to equal weights initially, construct a classifier, calculate the weighted error rate of this classifier, calculate the weight of this classifier, update the weight of each sample, and according to the new weights, sample with replacement to obtain new samples, repeat the above process of training the classifier until a classifier is trained, and output the final classifier; The Bagging: Generate multiple different sub-training sets by sampling the training data set with replacement multiple times, train an independent model on each subset, and then perform majority voting or weighted voting on the prediction results of all models to solve the classification problem; The gradient boosting: Utilize a weak learning model and construct a model to fit the residuals through gradient descent, thereby improving the prediction performance of the model; The extremely randomized trees: In the training process, enhance the generalization ability of the model by randomizing the construction process of the decision tree. The extremely randomized trees randomly select features and randomly select split points at each node; The XGBoost: optimizes the objective function by introducing a regularization term into the loss function, and uses the method of finding the extreme point of the objective function and second-order Taylor expansion to approximate to obtain the structure of the tree; The CatBoost: processes categorical features during the training process, avoids overfitting when calculating leaf nodes, converts floating-point features, statistical data, and one-hot encoded features into binary form, and incorporates these features into a vector for scoring the model output; The LightGBM: uses the negative gradient of the loss function as the residual approximation of the current decision tree to fit a new decision tree, that is, in each iteration, the original model remains unchanged, and a new function is added to the model to make the predicted value continuously approach the true value; The multi-layer perceptron: uses the backpropagation algorithm to gradually optimize the weights through the gradient descent method. The backpropagation algorithm transmits the error layer by layer from the output layer, calculates the gradient of the weights of each layer and updates them.

[0014] The beneficial effects of the present invention are as follows: 1. By comparing 14 traditional machine learning, ensemble learning, and deep learning models, it is found that ensemble models (such as LightGBM, XGBoost, Catboost) are significantly superior to traditional models (such as logistic regression, K-nearest neighbor algorithm) in terms of prediction performance. Among them, LightGBM performs best in both ten-fold cross-validation and external validation (external validation AUC = 0.845), with both high accuracy and generalization ability, significantly superior to the traditional NIHSS score (AUC = 0.748).

[0015] 2. For the first time, quantum computing is used to screen out indicators closely related to stroke prognosis, providing a new idea for high-dimensional data feature screening.

[0016] 3. For the first time, leukocyte community parameters are introduced into the acute ischemic stroke prognosis model, which detects the morphological and functional characteristics of leukocytes through flow cytometry technology, has higher sensitivity and objectivity compared with traditional inflammatory indicators, and becomes an important indicator for predicting nerve function damage.

[0017] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly and implement it in accordance with the content of the specification, the following describes in detail with reference to the preferred embodiments of the present invention and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a flowchart of a method for training a stroke prognosis prediction model based on quantum computing and artificial intelligence shown in an embodiment of the present invention; Figure 2 It is a QUBO value convergence curve graph shown in an embodiment of the present invention; Figure 3 The F1 value chart of 14 models shown in an embodiment of the present invention; Figure 4 The confusion matrix of the LightGBM model and NIHSS shown in an embodiment of the present invention; Figure 5 The local SHAP value chart shown in an embodiment of the present invention. Detailed implementation manners

[0019] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0020] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0021] Please refer to Figure 1 , a training method for a stroke prognosis prediction model based on quantum computing and artificial intelligence shown in a preferred embodiment of the present application, includes the following steps: Step S1: Establish a data case library; Prepare medical record data, collect patients' electronic medical records from the hospital electronic medical record platform, and collect the electronic medical records of patients with acute ischemic stroke with follow-up records; the medical record data of the cases with the first diagnosis of acute ischemic stroke are regarded as qualified electronic medical record data; Step S2: Extract patients' medical feature data; Extract the medical features of acute ischemic stroke from the qualified electronic medical record data obtained in step S1 to obtain medical features and medical feature values; the acute ischemic stroke features include demographic features, vital signs and severity of illness, blood routine, white blood cell colony parameters Cellular Population Data (CPD), and blood biochemistry; which are used as prediction indicators; The medical feature values are the specific values of each medical feature in demographic features, vital signs and severity of illness, blood routine, white blood cell colony parameters, and blood biochemistry; The demographic features include: age (Age), gender (Sex), height (Height), weight (Weight), BMI; The vital signs and severity of illness include systolic blood pressure (SBP), diastolic blood pressure (DBP), mean arterial pressure (MAP), and the admission NIHSS score; The blood routine examination includes white blood cell count (WBC), neutrophil count (Neu), lymphocyte count (Lym), monocyte count (Mon), platelet count (PLT), mean platelet volume (MPV), C-reactive protein (FR-CRP), hemoglobin concentration (HGB), hematocrit (HCT), mean corpuscular volume (MCV), mean corpuscular hemoglobin (MCH), and mean corpuscular hemoglobin concentration (MCHC); The white blood cell community parameters are detected through the DIFF channel of a fully automatic hematology analyzer (such as Mindray BC-7500CS), including the side scatter light intensity (Neu-X) of the Neu region on the DIFF channel for neutrophils, the side fluorescence intensity (Neu-Y) of the Neu region on the DIFF channel, the forward scatter light intensity (Neu-Z) of the Neu region on the DIFF channel, the side scatter light intensity (Lym-X) of the Lym region on the DIFF channel, the side fluorescence intensity (Lym-Y) of the Lym region on the DIFF channel, the forward scatter light intensity (Lym-Z) of the Lym region on the DIFF channel, the side scatter light intensity (Mon-X) of the Mon region on the DIFF channel, the side fluorescence intensity (Mon-Y) of the Mon region on the DIFF channel, the forward scatter light intensity (Mon-Z) of the Mon region on the DIFF channel, and the standard deviation of platelet distribution width (PDW-SD); The blood biochemistry includes fasting blood glucose (FBG), triglyceride (TG), total cholesterol (TC), high-density lipoprotein cholesterol (HDL-C), serum creatinine (Scr), alanine aminotransferase (ALT), albumin (Alb), and serum uric acid (SUA); Step S3: Extract target result characteristic data; Extract the post-stroke score three months after discharge; The binary classification results of the post-stroke score are: good prognosis with a post-stroke score of 0-2, and poor prognosis with a post-stroke score of 3-6. The situation may be moderate or severe disability, or death; Through steps S1 to S3, big data on the clinical manifestations of acute ischemic stroke can be obtained; Step S4: Standardize characteristic data and clean the data; Standardize the characteristic data of the big data on the clinical manifestations of acute ischemic stroke obtained in step S3. Adopt a missing data strategy, exclude patients with more than 20% missing characteristic variables, and fill in the remaining missing characteristic data in the mode of the existing data of the same characteristic, using the expectation-maximization algorithm for filling; Use the Wilcoxon rank-sum test to judge the filling effect of the missing values. If the P-value of the difference in the distribution of characteristic data before and after filling the missing values is greater than 0.05, the filling effect is good; Step S5: Construct new features and balance the data Construct new inflammation indices and blood glucose and lipid metabolism indicators from the medical feature data obtained in step S2; implement undersampling and oversampling of the data through "RandomUnderSampler" and "SMOTENC" in Python, and use an undersampling to oversampling ratio of 2:1 to balance the class distribution and ensure the robustness of the model; The inflammation indices and their construction methods include NLR = Neu / Lym, MLR = Mon / Lym, PLR = PLT / Lym, SII = PLT × Neu / Lym, SIRI = Mon × Neu / Lym, AISI = Neu × PLT × Mon / Lym, PIV = Mon × PLT × Neu, HALP = HGB × Alb × Lym / PLT, IBI = FR-CRP × Neu / Lym, NPHR = Neu / WBC × 100 / HGB, ALI = BMI × Alb / NLR; The blood glucose and lipid metabolism indicators and their construction methods include TyG = ln[FBG (mg / dL) × TG (mg / dL)] / 2, NHHR = (TC - HDL) / HDL, NPAR = Neu / WBC × 100 / Alb, PHR = PLT / HDL, MHR = Mon / HDL, UHR = SUA / HDL, TyG-BMI = BMI × TyG; Step S6: Feature selection: Lasso regression, recursive feature elimination, quantum computing; Select three feature selection methods for feature screening, and take the intersection of the screening results of Lasso regression, recursive feature elimination, and quantum computing as the key features. The key indicators of the features include Age, NIHSS, Neu-Z, Lym-X, Lym-Z, Mon-Y, Mon-Z, ALT, SUA, PHR, UHR, TyG-BMI; The Lasso regression: By calculating the conditional probability that a sample is assigned to the poor prognosis group and matching it with the good prognosis group of similar features, a good prognosis group with features similar to those of the poor prognosis group is obtained, so that the covariates tend to be balanced after the two groups of patients are matched, in order to better study the relationship between features and the target outcome; The recursive feature elimination: Train an extremely randomized tree model on the candidate feature set with all indicators, representing all candidate clinical indicators; calculate the importance measure of each feature in the model, and sort the feature set according to the size of this measure; eliminate the feature corresponding to the smallest measure to obtain a new feature subset; repeat this "training - calculation - elimination" process until the optimal feature set is obtained; The quantum computing: Establish a QUBO model, i.e., a quadratic unconstrained binary optimization model. Quantum computing realizes rapid screening of high-dimensional features through the QUBO model, with a running time of 5.044 milliseconds, showing a breakthrough advantage compared with the classical computing paradigm. Its core mechanism is as follows: A parallel computing architecture based on quantum superposition states and entanglement states significantly improves the operation dimension (quantum parallelism) while maintaining the algorithm accuracy. To a certain extent, this provides a reference for the modeling of larger-scale clinical examination and follow-up datasets of acute ischemic stroke.

[0022] It should be noted that the running times of Lasso regression and recursive feature elimination are 7.77 seconds and 146.8 seconds respectively, and the operation efficiency has been greatly improved compared with the prior art.

[0023] Step S7: Establish 14 machine learning prediction models; Establish 14 machine learning prediction models based on key features, including logistic regression, quadratic discriminant analysis, naive Bayes, Gaussian process regression, K-nearest neighbor algorithm, decision tree, AdaBoost, Bagging (Bootstrap aggregating), gradient boosting, extremely randomized trees, XGBoost (Extreme Gradient Boosting‌), CatBoost (CategoricalBoosting), LightGBM (Light Gradient Boosting Machine), and multi-layer perceptron; The logistic regression: The sigmoid function is used in logistic regression to solve binary classification problems. Through this function, function values greater than 0.5 are assigned to label 1 (poor prognosis), and function values less than or equal to 0.5 are assigned to label 0 (good prognosis). This algorithm can see the influence of different indicators on the prognosis of stroke patients from the weights of features, and the output result is a probability value.

[0024] The quadratic discriminant analysis: Calculate the mean and covariance matrix of each category; for a new observation value, calculate its posterior probability of belonging to each category; classify the observation value into the category with the highest posterior probability; The naive Bayes: Naive Bayes simplifies the calculation of the aforementioned conditional probability through the independence assumption of features; The Gaussian process regression: Define the similarity between input points through a covariance function (kernel function), and output the predicted mean and variance corresponding to the prediction points; The K-nearest neighbor algorithm: After inputting data without labels, compare each feature of this data without labels with the corresponding features in the sample set, and then extract the classification label of the data with the closest features (nearest neighbor) in the sample; The decision tree: Each node of the decision tree represents a judgment condition of a feature, each branch represents a judgment result, and each leaf node corresponds to a predicted class or value. The construction process of the decision tree is based on the features of the training data, and the data is partitioned by selecting the optimal feature. By continuously partitioning the features until all samples are correctly classified or a preset stopping condition is reached; The AdaBoost: Set the weight of each sample in the training set to be equal weights for each sample initially; construct a classifier, calculate the weighted error rate of this classifier; calculate the weight of this classifier, update the weight of each sample; according to the new weights, sample with replacement to obtain new samples, and repeat the above process of training the classifier until a classifier is trained; output the final classifier; The Bagging: Generate multiple different sub-training sets by sampling the training data set with replacement multiple times; train an independent model on each subset, and then perform majority voting or weighted voting on the prediction results of all models to solve the classification problem; The gradient boosting: Utilize a weak learning model and construct a model to fit the residuals through gradient descent (differentiable loss function), thereby improving the prediction performance of the model; The extremely randomized trees: During the training process, enhance the generalization ability of the model by randomizing the construction process of the decision tree; different from the traditional decision tree, the extremely randomized trees randomly select features and randomly select split points at each node; The XGBoost: Optimize the objective function by introducing a regularization term into the loss function; adopt the method of finding the extreme point of the objective function and second-order Taylor expansion to approximate to obtain the structure of the tree; The CatBoost: Process categorical features during the training process and avoid overfitting when calculating leaf nodes, which helps to select a suitable tree structure; convert floating-point features, statistical data, and one-hot encoded features into binary forms and incorporate these features into a vector for scoring the model output; The LightGBM: Use the negative gradient of the loss function as an approximation of the residual of the current decision tree to fit a new decision tree, that is, keep the original model unchanged in each iteration, and then add a new function to the model to make the predicted value continuously approach the true value; The multi-layer perceptron: Use the backpropagation algorithm to gradually optimize the weights through the gradient descent method; the backpropagation algorithm transmits the error layer by layer from the output layer, calculates the gradient of the weights of each layer and updates them; Step S8: Result comparison and model selection; By comparing 14 traditional machine learning, ensemble learning, and deep learning models, it was found that ensemble models (such as LightGBM, XGBoost, Catboost) were significantly superior to traditional models (such as logistic regression, K-nearest neighbor algorithm) in terms of prediction performance.

[0025] Please refer to Figure 3 , the LightGBM model performed best in both ten-fold cross-validation (the internal ten-fold cross-validation F1 value was 0.787 (0.783 - 0.789), and the external validation F1 value was 0.742) and external validation (the external validation AUC = 0.85), with both high accuracy and generalization ability, significantly superior to the traditional NIHSS score (AUC = 0.748); Furthermore, please refer to Figure 4 , where the left figure is the confusion matrix of LightGBM, and the right figure is the confusion matrix of NIHSS. Based on this, the accuracy and recall rate of LightGBM were calculated to be 93.3% and 78% respectively, and the accuracy and recall rate of NIHSS were 100% and 72.7% respectively, determining it as a prediction model for the prognosis of acute ischemic stroke.

[0026] Step S9: Model interpretability analysis: Quantify the contribution of each feature to the prognosis through SHAP (Shapley Additive Explanations) values to identify key risk factors (such as NIHSS, Mon-Y, PHR).

[0027] Please refer to Figure 5 , where the upper figure shows patients with good prognosis, and the lower figure shows patients with poor prognosis.

[0028] It should be noted that the application of quantum computing in feature screening in step S6 includes the following steps: (I) Quantum computing input data preprocessing Convert the acute ischemic stroke prognosis dataset obtained through feature construction before into a row, column matrix , where each column represents a clinical indicator, and each row represents the corresponding data value of the clinical sample.

[0029] ; The neurological prognosis of the patient is represented as a element vector : ; Among them, the data value representing the patient's prognosis is a 0-1 variable, where 0 indicates good prognosis and 1 indicates poor prognosis.

[0030] To establish the QUBO model, it is necessary to calculate the correlation between indicators and the correlation between each indicator and the patient prognosis. Since clinical indicators usually do not follow a normal distribution, the Spearman correlation coefficient is selected when calculating the autocorrelation matrix between indicators. In addition, since the patient prognosis is a binary variable, the point-biserial correlation coefficient is selected when calculating the correlation between the indicator and the prognosis. The calculation formula is as follows: When calculating, the point-biserial correlation coefficient is selected, and the calculation formula is as follows: ; Where and represent the mean values of the continuous indicator when the prognosis classification takes 1 and 0, respectively; represents the standard deviation of the continuous indicator; is the total sample size; and represent the sample sizes of the prognosis classification taking 1 and 0, respectively.

[0031] (II) Quantum Feature Screening The goal of feature selection is to find the columns of the matrix that are relevant to the vector but not correlated with each other. Let represent the correlation between the th column and the th column of the matrix , and represent the correlation between the th column of and the vector .

[0032] To find the best subset, binary variables are introduced, and their mathematical meaning is: ; They are combined into a vector , that is, obtaining the best indicator subset is converted into finding the value of corresponding to the minimum value of the objective function. The objective function consists of two parts: the first part represents the role of the indicator in the stroke prognosis: ; The second part represents the mutual independence between indicators: ; Introduce the parameter β to represent the relative weight of independence and role, and the final objective function is obtained as: ; In addition, the mathematical expression of the QUBO model is: ; Among them, is the binary variable to be solved, taking values of 0 or 1; is the quadratic coefficient, is a known quantity.

[0033] Convert into linear algebra: ; Solve it using a 550-qubit coherent optical quantum computer (CPQC-550). The convergence process of the QUBO value is as Figure 2 , and the screened index subset is obtained as: ; When the QUBO value is -10952, 20 clinical indicators are screened, namely age, gender, height, systolic blood pressure, MCV, Neu-X, Neu-Z, Lym-X, Lym-Z, Mon-Z, PDW-SD, Scr, ALT, mean arterial pressure, NLR, SII, PIV, ALI, NPAR, and TyG-BMI.

[0034] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0035] The above-described embodiments only represent several implementation manners of the present invention. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent should be subject to the appended claims.

Claims

1. A training method for a stroke prognosis prediction model based on quantum computing and artificial intelligence, characterized in that, It includes the following steps: Step S1: Establish a data case library; collect electronic medical records of patients, and collect the electronic medical records of patients with follow-up records and the first diagnosis of acute ischemic stroke; Step S2: Extract medical feature data of patients: extract the medical features of acute ischemic stroke from the electronic medical record data obtained in Step S1 to obtain medical features and medical feature values; the medical features of acute ischemic stroke include demographic features, vital signs and severity of illness, blood routine, white blood cell community parameters, and blood biochemistry, which are used as indicators for prediction; the medical feature values include the specific values of the medical features of demographic features, vital signs and severity of illness, blood routine, white blood cell community parameters, and blood biochemistry; Step S3: Extract target result feature data; extract the post-stroke score three months after the patient is discharged; the binary classification results of the post-stroke score are: good prognosis with a post-stroke score of 0-2, and poor prognosis with a post-stroke score of 3-6; after scoring all patients' post-stroke scores, obtain the big data of the clinical manifestations of acute ischemic stroke; Among them, the poor prognosis includes moderate disability, severe disability, and death; Step S4: Feature data standardization and data cleaning: Standardize the feature data of the big data of the clinical manifestations of acute ischemic stroke obtained in Step S3, adopt a missing data strategy, exclude patients with more than 20% missing feature variables, and fill the remaining missing feature data in the mode of the existing data of the same feature, fill it with the expectation-maximization algorithm, and use the Wilcoxon rank-sum test to judge the filling effect of the missing values. If the P-value of the difference in the feature data distribution before and after filling the missing values is greater than 0.05, the filling effect is better; Step S5: New feature construction and data balancing: Construct new inflammation indices and blood glucose and lipid metabolism indices through the medical feature data obtained in Step S2, and implement undersampling and oversampling of the data through the RandomUnderSampler and SMOTENC algorithms in Python, and use an undersampling to oversampling ratio of 2:1 to balance the class distribution; Step S7: Feature selection: Select 3 feature selection methods for feature screening, and take the intersection of the Lasso regression, recursive feature elimination, and quantum computing screening results as the key features; Step S8: Establish 14 machine learning prediction models: Based on the key features, establish 14 machine learning prediction models, including logistic regression, quadratic discriminant analysis, naive Bayes, Gaussian process regression, K-nearest neighbor algorithm, decision tree, AdaBoost, Bagging, gradient boosting, extremely randomized tree, XGBoost, CatBoost, LightGBM, and multi-layer perceptron; Step S9: Compare the results obtained in Step S7 and select a model; Step S10: Model interpretability analysis: Quantify the contribution of each feature to the prognosis through SHAP values, and identify key risk factors, such as NIHSS, Mon-Y, PHR.

2. The training method of the stroke prognosis prediction model based on quantum computing and artificial intelligence according to claim 1, wherein In Step S2, the demographic features include: age, gender, height, weight, body mass index; The vital signs and severity of the condition include systolic blood pressure, diastolic blood pressure, mean arterial pressure, and the NIHSS score at admission; The blood routine includes white blood cell count, neutrophil count, lymphocyte count, monocyte count, platelet count, mean platelet volume, C-reactive protein, hemoglobin concentration, hematocrit, mean corpuscular volume, mean corpuscular hemoglobin content, and mean corpuscular hemoglobin concentration; The white blood cell population parameters are detected through the DIFF channel of an automatic hematology analyzer, including the lateral scatter light intensity of neutrophils, lymphocytes, and monocytes in the Neu region of the DIFF channel, the lateral fluorescence intensity of the Neu region of the DIFF channel, the forward scatter light intensity of the Neu region of the DIFF channel, the lateral scatter light intensity of the Lym region of the DIFF channel, the lateral fluorescence intensity of the Lym region of the DIFF channel, the forward scatter light intensity of the Lym region of the DIFF channel, the lateral scatter light intensity of the Mon region of the DIFF channel, the lateral fluorescence intensity of the Mon region of the DIFF channel, the forward scatter light intensity of the Mon region of the DIFF channel, and the standard deviation of platelet distribution width; The blood biochemistry includes blood glucose, triglyceride, total cholesterol, high-density lipoprotein cholesterol, serum creatinine, alanine aminotransferase, albumin, and uric acid.

3. The training method of the stroke prognosis prediction model based on quantum computing and artificial intelligence according to claim 1, characterized in that, In step S5, the inflammation indices and construction methods include NLR = Neu / Lym, MLR = Mon / Lym, PLR = PLT / Lym, SII = PLT × Neu / Lym, SIRI = Mon × Neu / Lym, AISI = Neu × PLT × Mon / Lym, PIV = Mon × PLT × Neu, HALP = HGB × Alb × Lym / PLT, IBI = FR-CRP × Neu / Lym, NPHR = Neu / WBC × 100 / HGB, and ALI = BMI × Alb / NLR; The blood glucose and lipid metabolism indices and construction methods include TyG = ln[FBG (mg / dL) × TG (mg / dL)] / 2, NHHR = (TC - HDL) / HDL, NPAR = Neu / WBC × 100 / Alb, PHR = PLT / HDL, MHR = Mon / HDL, UHR = SUA / HDL, and TyG-BMI = BMI × TyG.

4. The training method of the stroke prognosis prediction model based on quantum computing and artificial intelligence according to claim 1, characterized in that In step S6: Lasso regression: By calculating the conditional probability that a sample is assigned to the poor prognosis group and matching it with the good prognosis group with similar characteristics, a good prognosis group with characteristics similar to those of the poor prognosis group is obtained, so that the covariates tend to be balanced after the two groups of patients are matched, in order to better study the relationship between the characteristics and the target outcome; Recursive Feature Elimination: Train an Extra Trees model on the candidate feature set with all metrics, representing all candidate clinical metrics; calculate the importance measure of each feature in the model, and sort the feature set according to the size of this measure; eliminate the feature corresponding to the smallest measure to obtain a new feature subset; repeat this "training - calculation - elimination" process until the optimal feature set is obtained; Quantum Computing: Establish a QUBO model, that is, a Quadratic Unconstrained Binary Optimization model.

5. The training method of the stroke prognosis prediction model based on quantum computing and artificial intelligence according to claim 1, wherein In step S7: The logistic regression uses the sigmoid function. Through this function, the function values greater than 0.5 are assigned to the label 1 indicating poor prognosis, and the function values less than or equal to 0.5 are assigned to the label 0 indicating good prognosis; this algorithm can see the impact of different metrics on the prognosis of stroke patients from the weights of the features, and the output result is a probability value; The quadratic discriminant analysis: Calculate the mean and covariance matrix of each category. For a new observation, calculate its posterior probability of belonging to each category, and classify the observation into the category with the highest posterior probability; The Naive Bayes simplifies the calculation of the aforementioned conditional probability through the independence assumption of features; The Gaussian process regression: Defines the similarity between input points through a covariance function, and outputs the predicted mean and variance corresponding to the prediction points; The K-nearest neighbor algorithm: After inputting data without labels, compare each feature of this data without labels with the corresponding features in the sample set, and then extract the classification label of the data with the most similar features in the sample; The decision tree: Each node of the decision tree represents a judgment condition of a feature, each branch represents the judgment result, and each leaf node corresponds to the predicted category or value. The construction process of the decision tree is based on the features of the training data, and the optimal feature is selected for data division. By continuously dividing the features until all samples are correctly classified or the preset stop condition is reached; The AdaBoost: Set the weight of each sample in the training set to equal weights initially, construct a classifier, calculate the weighted error rate of this classifier, calculate the weight of this classifier, update the weight of each sample, and perform sampling with replacement according to the new weights to obtain new samples. Repeat the above process of training the classifier until a classifier is trained, and output the final classifier; The Bagging: Generate multiple different sub-training sets by sampling the training data set with replacement multiple times, train an independent model on each subset, and then perform majority voting or weighted voting on the prediction results of all models to solve the classification problem; The Gradient Boosting: Utilize a weak learning model and construct a model to fit the residuals through gradient descent, thereby improving the prediction performance of the model; The Extra Trees: During the training process, enhance the generalization ability of the model by randomizing the construction process of the decision tree. The Extra Trees randomly selects features and randomly selects split points at each node; The XGBoost: Optimize the objective function by introducing a regularization term into the loss function, and use the method of finding the extreme point of the objective function and second-order Taylor expansion to approximate to obtain the structure of the tree; The CatBoost: Processes categorical features during the training process and avoids overfitting when calculating leaf nodes, converts floating-point features, statistical data, and one-hot encoded features into binary form, and incorporates these features into a vector for scoring in the model output; The LightGBM: Uses the negative gradient of the loss function as an approximation of the residual of the current decision tree to fit a new decision tree, that is, in each iteration, the original model remains unchanged, and a new function is added to the model to continuously approximate the predicted value to the true value; The multi-layer perceptron: Uses the backpropagation algorithm to gradually optimize the weights through the gradient descent method. The backpropagation algorithm transmits the error layer by layer from the output layer, calculates the gradient of the weights of each layer, and updates them.

Citation Information

Cited By

  • Inflammation signal expression profile fused nerve injury prognosis prediction system and method

    CN121565473A