Prognosis prediction method based on combination of senescence marker and cerebral apoplexy clinical data

By adopting the bilinear interpolation and forgetting weight mechanism of the Eulera Grangian equation in the prognosis prediction of stroke, combined with the feedforward neural network model of the drift compensation module, the noise, redundancy and dynamic drift problems of existing methods when processing high-dimensional medical data is solved, significantly improving the accuracy and robustness of the prediction.

CN120199491APending Publication Date: 2025-06-24LIAOCHENG PEOPLES HOSPITAL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510299103.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing stroke prognosis prediction methods have problems such as data noise, redundant information, model overfitting, poor generalization ability, and dynamic drift when processing high-dimensional and complex medical data, resulting in insufficient prediction accuracy and robustness.

Method used

A bilinear interpolation data augmentation method based on Eulera Grangian equation is adopted, and the sample importance and data offset are automatically processed in combination with the forgetting weight mechanism, and the neural network weight is dynamically adjusted through the drift compensation module to build a prognostic prediction model based on feedforward neural network.

Benefits of technology

Generate high-quality training data to improve the generalization ability and adaptability of the model, significantly improve the accuracy, robustness and adaptability of stroke prognosis prediction, and solve the problems of high-dimensional features, noise and data scarcity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199491A_ABST
    Figure CN120199491A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical prognosis, in particular to a prognosis prediction method based on combination of senescence markers and cerebral apoplexy clinical data, which specifically comprises the following steps: collecting training data, and uniformly storing the collected data in a data warehouse; carrying out data cleaning and multi-modal data fusion on the collected data, and carrying out normalization processing on the fused data; manually marking the collected data by an expert; the method comprises the following steps: expanding acquired senescence marker data and cerebral apoplexy clinical data by adopting a bilinear interpolation method based on an Euler-Lagrange equation to obtain a training set; constructing a machine learning model based on a feedforward neural network of drift compensation, and inputting the training set into the machine learning model for training; and selecting an optimal model, and inputting new data to obtain a final prediction result. The accuracy, the robustness and the adaptability of cerebral apoplexy prognosis prediction can be improved, and the defects in medical data analysis in the prior art are overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical prognosis, and particularly to a prognosis prediction method based on the combination of aging markers and stroke clinical data. Background Art

[0002] Most of the existing stroke prognosis prediction methods rely on traditional clinical data, such as blood pressure, blood glucose levels, imaging data, etc. However, these data often have limitations, cannot comprehensively reflect the overall health status of patients, and are limited by individual differences, making it difficult to provide accurate personalized predictions. In recent years, with the progress of research on aging markers, more and more biomarkers have been found to be closely related to the aging process and the occurrence of various chronic diseases. As a potential disease prediction indicator, aging markers have gradually attracted attention in medical research. However, in-depth analysis of the relationship between aging markers and stroke, especially in clinical applications, still faces many technical challenges.

[0003] There have been many methods for disease prediction in the prior art, for example: A stroke prognosis prediction method and system with the publication number of CN119417747A, which segments and extracts features from CTP images, and then processes the extracted features based on a stroke prognosis prediction model constructed by machine learning. Furthermore, it quantifies and extracts vascular blood flow and perfusion features more comprehensively and precisely, and realizes the entire prediction process automatically.

[0004] A method for establishing a hyperuricemia prediction model using machine learning algorithms with the publication number of CN118486471B, which is trained based on the UKBB cohort of a large number of people, and establishes and validates a reliable and practical stacked multi-modal machine learning model trained based on genetic and clinical data, used to timely identify the HUA phenotype and early predict the risks of gout and metabolism-related outcomes. The involved ISHUA markers can quantify the risks of gout and metabolism-related results at an early stage, and can provide positive guidance for daily health checks and risk factor management.

[0005] A machine learning-based bladder tumor prognosis prediction method with the publication number of CN118888159B, which can solve the problem that the reference values of different samples are different, reduce the misclassification rate of samples in the prediction model, and enhance the accuracy of the results.

[0006] However, the above technical solutions still have the following problems that need to be further solved: (1) When facing medical data, existing data augmentation methods often struggle to handle high-dimensional and complex features. Especially when there is a lot of data noise and redundant information, traditional augmentation methods (such as SMOTE, GAN, etc.) are difficult to generate high-quality training samples, resulting in the model being prone to overfitting or having poor generalization ability. (2) Traditional machine learning models usually cannot cope with the dynamic drift phenomenon in medical data that occurs over time, treatment methods, or disease progression. As a result, when the model is applied to new patient data, the accuracy drops significantly. (3) Redundant and irrelevant features in high-dimensional data will have a negative impact on the learning process of traditional models. Existing technologies often rely on simple feature selection methods and are difficult to handle complex medical data relationships, resulting in limited prediction ability of the model. (4) Traditional neural networks are prone to gradient vanishing or explosion when facing complex medical data. Especially when training deep networks, it is easy to lead to unstable training, which in turn affects the performance and usability of the final model.

[0007] Therefore, the present invention proposes a prognosis prediction method based on the combination of aging markers and stroke clinical data to solve the above problems. Summary of the Invention

[0008] In view of the deficiencies of the prior art, the present invention develops a prognosis prediction method based on the combination of aging markers and stroke clinical data, aiming to improve the generalization ability of the model, enhance the accuracy, robustness, and adaptability of stroke prognosis prediction, and make up for the deficiencies of the prior art in medical data analysis.

[0009] The technical solution for the present invention to solve the technical problem is a prognosis prediction method based on the combination of aging markers and stroke clinical data, including the following steps: S1. Data collection: Collect training data, which includes aging marker data and stroke clinical data. Then, after standardizing the data formats of all the collected training data, store them uniformly in a data warehouse. S2. Data preprocessing: Clean the aging marker data and stroke clinical data, then perform multi-modal data fusion on the two types of data, and finally perform normalization processing on the fused data. S3. Data annotation: Manually annotate the aging marker data and stroke clinical data by experts, and the annotation category is the recovery situation of the patients corresponding to the data. S4. Data augmentation: Use the bilinear interpolation method based on the Euler-Lagrange equation to augment the collected aging marker data and stroke clinical data to obtain an augmented training set. S5. Construct a machine learning model and train it: Construct a machine learning model based on a feedforward neural network with drift compensation, input the data in the training set into the machine learning model for training, and perform multiple iterations; S6. Prognosis prediction: Select the optimal model from the machine learning models after multiple iterations, input new senescence marker data and stroke clinical data, and obtain the final prediction result.

[0010] S1 is specifically as follows: Senescence marker data: Collect senescence marker data through professional biochemical detection equipment and methods. The senescence marker data comes from medical clinical trials, long-term biomedical research data sets, and the genomes and biomarker detection results of senescence markers. The data types include, but are not limited to, biomarker data such as specific molecular levels in blood, gene expression data, and hormone levels; Stroke clinical data: The stroke clinical data comes from the clinical research system. The data includes the basic information, medical history, physical examination results, imaging examination results, treatment methods, and clinical observation data of the disease process of the patients.

[0011] S2 is specifically as follows: S2.1. Data cleaning: For outliers, use the method based on locally weighted regression to detect and remove abnormal data points; For missing values, use the K-nearest neighbor algorithm for imputation; S2.2. Multimodal data fusion: By separately encoding each data type into a high-dimensional vector and adopting a weighted strategy, perform weighted fusion on the senescence marker data and the stroke clinical data to obtain the fused features; S2.3. Normalization processing: Perform normalization processing on the fused features through the maximum-minimum normalization method to make all feature scales consistent.

[0012] The recovery status of the patient includes: complete recovery, partial recovery, and no recovery.

[0013] S4 is specifically as follows: The data set composed of the senescence marker data and the stroke clinical data after data collection, preprocessing, and annotation is denoted as , where represents the features of the th medical data sample, represents the number of medical data samples, represents the index of ; S4.1. Through an adaptive mechanism for the data set A forgetting weight is matched to each medical data sample, and the calculation formula is as follows: , wherein, represents the forgetting weight of the -th medical data sample feature, represents the first hyperparameter for adjusting the forgetting intensity, represents the second hyperparameter for adjusting the forgetting intensity, represents the -th importance evaluation function of the medical data sample feature. The similarity between medical data sample features is calculated through similarity measurement, and then the importance of each medical data sample feature is evaluated. The calculation formula is as follows: , wherein, represents another index of, , , represents the -th medical data sample feature, represents the sensitivity parameter for controlling similarity, represents the L2 norm; S4.2. Bilinearly interpolate the medical data sample features in the data set to generate new medical data sample features , represents the new medical data sample features generated after bilinear interpolation; S4.3. Adopt an optimization mechanism based on the Euler-Lagrange equation to adjust the interpolation and forgetting weights by minimizing the energy function so as to optimize the distribution of the new medical data sample features; Then, by solving the Euler-Lagrange equation , the optimized position distribution of the new medical data sample features is obtained; S4.4. Combine the generated new medical data sample features with the medical data sample features in the data set to obtain an expanded training set , , , .

[0014] S5 is as follows: S5.1. Input the medical data sample features in the training set into the machine learning model, and calculate the drift compensation matrix according to the input medical data sample features. The calculation formula is as follows: , Among them, represents the drift compensation matrix, represents the number of medical data sample features input into the machine learning model, represents the index of represents another index of represents the importance coefficient of the th medical data sample feature, represents the th and the th correlation coefficient between medical data sample features, represents the th medical data sample feature, represents the th medical data sample feature, represents and the covariance of; S5.2. The medical data sample features in the training set are input into the feedforward neural network of the machine learning model for training. During the forward propagation of the feedforward neural network, the feedforward neural network performs forward propagation calculations layer by layer. The output of each layer is fine-tuned in combination with the drift compensation matrix. When the distribution of a certain sample feature drifts, the corresponding network layer automatically adjusts its weights to compensate for the impact of the change in the distribution of medical data sample features. The drift compensation is achieved by weighting the drift compensation matrix on the output of each layer. The output of the feedforward neural network after adjustment is as follows: , , Among them, represents the output of the feedforward neural network after compensation adjustment, represents the operation of the Sigmoid activation function, represents the weight matrix of the feedforward neural network, represents the medical data sample features in the training set input into the feedforward neural network, represents the bias term of the feedforward neural network, represents the drift compensation coefficient, represents the offset term based on the change rate of the input medical data sample features, represents the th offset adjustment coefficient of the medical data sample feature. The th offset adjustment coefficient of the medical data sample feature takes the value of the th normalized medical data sample feature value, Indicates The change rate of the number of iterations , Indicates the operation of integer variables in computer languages; S5.3. During the backpropagation process of training a feedforward neural network, perform gradient path backtracking. Backtrack through the historical gradient correction terms, record the historical gradient values of gradient updates, and adjust the current gradient direction to optimize the gradient flow process. The calculation formula is as follows: , Where Indicates the gradient correction term of the current layer of the feedforward neural network, Indicates the gradient of the loss function of the feedforward neural network, Indicates the gradient correction term of the previous layer of the current layer of the feedforward neural network, Indicates the decay coefficient of the historical gradient. For the first layer of the feedforward neural network, its gradient correction term is set to 0.1; Calculate the decay coefficient of the historical gradient through the norm of the gradient of the loss function of the feedforward neural network. The calculation formula is as follows: , Where Indicates the learning rate control factor, Indicates the norm of the gradient of the loss function of the feedforward neural network; S5.4. After each iteration of the feedforward neural network, perform enhancement processing on the intermediate process features. Aggregate the features of the output of the current iteration of the feedforward neural network and filter out irrelevant features. The calculation formula for feature aggregation is as follows: , Where Indicates the output after feature aggregation, Indicates the number of output features during the iteration process of the feedforward neural network, Indicates The index of Indicates Another index of Indicates the th feature during the iteration process of the feedforward neural network, Indicates the th feature during the iteration process of the feedforward neural network, Indicates the th feature's importance coefficient during the iteration process of the feedforward neural network; S5.5. After each iteration of the feedforward neural network, perform a learning rate adjustment of the feedforward neural network. Dynamically adjust the learning strategy according to the output of the current layer and the error of the previous layer. The learning rate adjustment formula of the feedforward neural network is as follows: , Among them, represents the learning rate of the feedforward neural network, represents the number of layers of the feedforward neural network, represents the error of the th layer of the feedforward neural network, represents the error of the th layer of the feedforward neural network, represents the first error adjustment factor, represents the second error adjustment factor, represents the error history correction factor; Combined with the changing trend of the historical gradient, dynamically adjust the current learning rate to obtain the error history correction factor. The calculation formula is as follows: , Among them, represents the adjustment coefficient of the error history correction factor; S5.6. Calculate the class probability through the output of the last layer of the feedforward neural network by the preset Softmax function, and take the class with the largest class probability as the prediction result. The predicted classes include complete recovery, partial recovery, and no recovery; S5.6. Repeat and iterate steps S5.1 to S5.5 until the preset stop iteration condition is met, and the machine learning model training is completed. The preset stop iteration condition is the number of iterations.

[0015] S6 is specifically as follows: S6.1. Model loading: Determine the optimal model according to the manual annotation and prediction results, save the optimal model parameters during the training process of the machine learning model, and load the trained feedforward neural network model from them; S6.2. Input data preprocessing: Input new senescence marker data and stroke clinical data, and perform data cleaning, missing value filling, and feature normalization on the new data; S6.3. Prediction classification: Input the preprocessed new data into the trained feedforward neural network model for prediction to obtain the final classification prediction result.

[0016] The effects provided in the invention content are only the effects of the embodiments, rather than all the effects of the invention. The above technical solutions have the following advantages or beneficial effects: (1) The present invention adopts a bilinear interpolation data augmentation method based on the Euler-Lagrange equation, and combines a forgetting weight mechanism to automatically process sample importance and data drift, which can generate higher-quality training data, avoid the problem that traditional amplification methods are difficult to handle high-dimensional data, and solve the problems of high-dimensional features, noise, and data scarcity in medical data; (2) The present invention adopts a drift compensation module to dynamically adjust the weights of the neural network to cope with the non-linear changes in medical data caused by factors such as time and disease progression, effectively alleviating the impact of data drift on the model performance and enhancing the adaptability of the model to dynamic medical data. (3) By using random forest to evaluate feature importance and combining Pearson correlation coefficient to handle the correlation between features, the present invention can effectively reduce the interference of redundant features in a high-dimensional data environment and enhance the focusing ability of the model on complex features such as gene expression data. (4) During the training process of the feedforward neural network, the present invention combines error history correction, dynamic learning rate adjustment and gradient path backtracking mechanism to ensure the training stability and efficiency of the network when facing complex medical data, avoiding the problems of gradient disappearance or explosion easily occurring in traditional neural networks.

[0017] In summary, in view of the limitations of the existing stroke prognosis prediction methods, the present invention proposes a comprehensive prediction method combining aging markers and stroke clinical data, overcoming the deficiencies of traditional methods in terms of data quality, feature selection, model stability and adaptability. Through bilinear interpolation based on Euler-Lagrange equation and forgetting weight mechanism, the present invention effectively solves the problems of complex high-dimensional features, data scarcity and noise interference in medical data, making the enhanced data more in line with the real distribution and improving the generalization ability of the model. At the same time, the drift compensation mechanism is adopted to enable the model to dynamically adjust the weights and adapt to the feature distribution changes caused by disease progression, significantly enhancing the accuracy, robustness and adaptability of stroke prognosis prediction and making up for the deficiencies in medical data analysis of the existing technology. Description of the Drawings

[0018] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation to the present invention.

[0019] Figure 1 It is a schematic flowchart of the method of the present invention.

[0020] Figure 2 It is a schematic diagram for comparing the performance of different data augmentation methods.

[0021] Figure 3 It is a schematic diagram for comparing the distribution of generated samples and real data.

[0022] Figure 4 It is a schematic diagram for comparing the robustness of different models to data drift.

[0023] Figure 5 It is a schematic diagram for showing the influence of dimensional features on the model performance.

[0024] Figure 6Schematic diagram for comparing training convergence speed. Detailed implementation manners

[0025] In order to clearly illustrate the technical features of this solution, the present invention will be described in detail below through specific implementation manners and in conjunction with its accompanying drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. To simplify the disclosure of the present invention, the components and settings of specific examples are described below.

[0026] Example 1 A prognosis prediction method based on the combination of aging markers and stroke clinical data, comprising the following steps: S1. Data collection: Collect training data, where the training data includes aging marker data and stroke clinical data, and then uniformly store all the collected training data in a data warehouse after standardizing the data format; S2. Data preprocessing: Perform data cleaning on the aging marker data and stroke clinical data, then perform multi-modal data fusion on the two types of data, and finally perform normalization processing on the fused data; S3. Data annotation: Manually annotate the aging marker data and stroke clinical data by experts, and the annotation category is the recovery situation of the patients corresponding to the data; S4. Data augmentation: Use the bilinear interpolation method based on the Euler-Lagrange equation to augment the collected aging marker data and stroke clinical data to obtain an augmented training set; S5. Construct and train a machine learning model: Construct a machine learning model based on a feedforward neural network with drift compensation, input the data in the training set into the machine learning model for training, and perform multiple iterations; S6. Prognosis prediction: Select the optimal model from the machine learning models after multiple iterations, input new aging marker data and stroke clinical data, and obtain the final prediction result.

[0027] S1 is specifically as follows: Aging marker data: Collect the aging marker data through professional biochemical detection equipment and methods, such as enzyme-linked immunosorbent assay or mass spectrometry analysis, etc., to perform high-precision measurement on the aging markers. The aging marker data comes from medical clinical trials, long-term biomedical research data sets, and genomic and biomarker detection results of aging markers. The data types include but are not limited to biomarker data such as specific molecular levels in blood, gene expression data, hormone levels, etc.; Clinical data of stroke: The clinical data of stroke are sourced from the clinical research system. The data include the patient's basic information (such as age, gender), medical history, physical examination results, imaging examination results (such as CT, MRI data), treatment methods (drug treatment, surgery, etc.), and clinical observation data on the disease process (such as stroke type, onset time, etc.).

[0028] Data storage: The standardized data structure can be in the form of CSV, JSON, or a table in a database, which is used to ensure the compatibility and query efficiency of different data sources.

[0029] S2 is as follows: S2.1. Data cleaning: For outliers, a locally weighted regression-based method is used to detect and remove abnormal data points. Specifically, by performing weighted averaging of each data with its neighboring data, locally weighted regression can effectively reduce the error caused by a single outlier; For missing values, the K-nearest neighbor algorithm is used for imputation to ensure that the filled values are more in line with the data distribution pattern; S2.2. Multimodal data fusion: By separately encoding each data type into a high-dimensional vector and adopting a weighted strategy, the aging biomarker data and the clinical data of stroke are weighted and fused to obtain the fused features. The fused feature vector can more effectively represent the multi-dimensional information of the data; S2.3. Normalization processing: The fused features are normalized by the maximum-minimum normalization method to make all feature scales consistent, so as to avoid feature bias during the model training process.

[0030] In the specific implementation manner, the patient's recovery situation includes: complete recovery, partial recovery, and no recovery.

[0031] S4 is as follows: The dataset composed of the aging biomarker data and the clinical data of stroke after data collection, preprocessing, and annotation is denoted as , where represents the features of the th medical data sample, represents the number of medical data samples, represents the index of , ; S4.1. Match a forgetting weight for each medical data sample in the dataset through an adaptive mechanism. The calculation formula is as follows: , where represents the forgetting weight of the features of the th medical data sample, represents the first hyperparameter for adjusting the forgetting intensity, represents the second hyperparameter for adjusting the forgetting intensity, represents the importance evaluation function of the th medical data sample feature. By calculating the similarity between medical data sample features through similarity measurement, the importance of each medical data sample feature is evaluated. The calculation formula is as follows: where, represents another index of , , represents the th medical data sample feature, represents the sensitivity parameter for controlling similarity, set , represents the L2 norm; S4.2. Bilinearly interpolate the medical data sample features in the dataset to generate new medical data sample features. The calculation formula is as follows: , where, represents the interpolation coefficient, represents the new medical data sample features generated after bilinear interpolation; Adjust the interpolation coefficient through the forgetting weight . The calculation formula is as follows: , where, represents the forgetting weight of the th medical sample data feature; S4.3. Adopt an optimization mechanism based on the Euler-Lagrange equation to adjust the interpolation and forgetting weights by minimizing the energy function, and then optimize the distribution of the new medical data sample features. The calculation formula is as follows: , where, represents the energy function, represents the interpolation regularization parameter for controlling the distance between the new sample features and the original medical data sample features, set ; By solving the Euler-Lagrange equation , obtain the optimized position distribution of the new medical data sample features; S4.4. Combine the generated new medical data sample features with the dataset Combine the characteristics of the medical data samples in it to obtain an expanded training set , , , .

[0032] S5 is specifically as follows: S5.1. Input the characteristics of the medical data samples in the training set into the machine learning model, and calculate the drift compensation matrix according to the input characteristics of the medical data samples. The calculation formula is as follows: , where represents the drift compensation matrix, represents the number of characteristics of the medical data samples input into the machine learning model, represents index of represents another index of represents the th importance coefficient of the medical data sample characteristics, represents the th and the th correlation coefficient between the medical data sample characteristics, represents the th medical data sample characteristic, represents the th medical data sample characteristic, represents and covariance of The importance coefficient of the medical data sample characteristics is specifically obtained by presetting a random forest. Train a preset random forest with the training set. The random forest constructs multiple decision trees. Each decision tree uses the medical data sample characteristics for classification. Each time of splitting, the random forest algorithm calculates how to reduce the impurity of the training set according to the sample characteristics (for example, Gini index or information gain). The greater the reduction of the impurity, the more important the sample characteristic is. Therefore, for each decision tree, calculate the contribution of each sample characteristic to the reduction of the impurity of all splits, and then sum up the contributions of all trees to obtain the overall importance coefficient of the sample characteristic; The correlation coefficient between two sample characteristics is specifically obtained by the Pearson correlation coefficient calculation method or can also be obtained by the Spearman rank correlation coefficient calculation method. The role of the correlation coefficient between two sample characteristics is: to make the drift compensation matrix not only based on the individual values of the sample characteristics, but also consider the correlation between the sample characteristics, so as to more accurately compensate for the drift phenomenon in the medical data; S5.2. Training set The medical data sample features in are input into the feedforward neural network of the machine learning model for training. During the forward propagation of the feedforward neural network, the feedforward neural network performs forward propagation calculations layer by layer. The output of each layer is fine-tuned in combination with the drift compensation matrix. When the distribution of a certain sample feature drifts, the corresponding network layer automatically adjusts its weights to compensate for the impact of the change in the distribution of medical data sample features. Drift compensation is achieved by weighting the drift compensation matrix on the output of each layer. The output of the feedforward neural network after adjustment is as follows: , , where, represents the output of the feedforward neural network after compensation adjustment, represents the operation of the Sigmoid activation function, represents the weight matrix of the feedforward neural network, represents the training set input into the feedforward neural network in the medical data sample features, represents the bias term of the feedforward neural network, represents the drift compensation coefficient, set , represents the offset term based on the change rate of the input medical data sample features, represents the th offset adjustment coefficient of the medical data sample feature. The offset adjustment coefficient of the th medical data sample feature takes the value of the th medical data sample feature after normalization, represents the change rate of the iteration number , represents the integer variable operation in the computer language; S5.3. During the backpropagation process of the feedforward neural network training, gradient path backtracking is performed. Backtracking is carried out through the historical gradient correction term. Record the historical gradient values of the gradient update and adjust the current gradient direction to optimize the gradient flow process, ensuring that the gradient is evenly propagated in the multi-layer network and avoiding the phenomenon of gradient disappearance or explosion during the training of the feedforward neural network. The calculation formula is as follows: , where, represents the gradient correction term of the current layer of the feedforward neural network, represents the gradient of the loss function of the feedforward neural network, represents the gradient correction term of the previous layer of the current layer of the feedforward neural network, Represents the decay coefficient of the historical gradient. For the first layer of the feedforward neural network, its gradient correction term is set to 0.1; The role of the gradient correction term of the current layer of the feedforward neural network is as follows: When using the gradient descent method to optimize the parameters of the feedforward neural network, after multiplying with the learning rate of the feedforward neural network, it is added to the update increment of the feedforward neural network parameters to update the parameters of the feedforward neural network in this iteration; Calculate the decay coefficient of the historical gradient through the norm of the gradient of the feedforward neural network loss function. The calculation formula is as follows: , where, Represents the learning rate control factor, set , Represents the norm of the gradient of the feedforward neural network loss function; S5.4. After each iteration of the feedforward neural network, perform enhancement processing on the intermediate process features, aggregate the output of the current iteration of the feedforward neural network, and aggregate the output of the current iteration of the feedforward neural network to ensure that the feedforward neural network can focus on important biomarkers or clinical features, filter out irrelevant features, and improve the selectivity of features. The calculation formula for feature aggregation is as follows: , where, Represents the output after feature aggregation, Represents the number of output features during the iteration process of the feedforward neural network, Represents the index of, Represents another index of, Represents the th feature during the iteration process of the feedforward neural network, Represents the th feature during the iteration process of the feedforward neural network, Represents the th feature importance coefficient during the iteration process of the feedforward neural network; Similar to the calculation method of the feature importance coefficient in the drift compensation matrix, the feature importance coefficient during the iteration process of the feedforward neural network is also calculated by a preset random forest; S5.5. After each iteration of the feedforward neural network, perform an adjustment of the learning rate of the feedforward neural network. The complexity in medical data requires the network to be able to flexibly adjust the learning rate during training, and dynamically adjust the learning strategy according to the output of the current layer and the error of the previous layer. The learning rate adjustment formula of the feedforward neural network is as follows: , where, represents the learning rate of the feedforward neural network, represents the number of layers of the feedforward neural network, represents the layer error of the feedforward neural network, which is calculated by applying a preset Softmax function to the output of the layer of the feedforward neural network, represents the layer error of the feedforward neural network, represents the first error adjustment factor, set , represents the second error adjustment factor, set , represents the error history correction factor; Combined with the changing trend of the historical gradient, dynamically adjust the current learning rate to obtain the error history correction factor. The calculation formula is as follows: , where represents the adjustment coefficient of the error history correction factor, set ; S5.6. Calculate the class probabilities through the output of the last layer of the feedforward neural network by applying the preset Softmax function, and take the class with the highest class probability as the prediction result. The predicted classes include complete recovery, partial recovery, and no recovery; S5.6. Repeat the iterative steps S5.1 to S5.5 until the preset stop iteration condition is met, and the machine learning model training is completed. The preset stop iteration condition is the number of iterations, and the preset maximum number of iterations is set to 1000 times.

[0033] S6 is specifically as follows: S6.1. Model loading: Determine the optimal model based on the manual annotation and prediction results, save the optimal model parameters during the machine learning model training process, and load the trained feedforward neural network model therefrom; S6.2. Input data preprocessing: Input new biomarker data for aging and clinical data for stroke, and perform data cleaning, missing value filling, and feature normalization on the new data; S6.3. Prediction classification: Input the preprocessed new data into the trained feedforward neural network model for prediction to obtain the final classification prediction result. During the prediction process, considering the change in data distribution, the trained feedforward neural network model will automatically adjust its internal weights to adapt to the new input data.

[0034] Embodiment 2 As Figure 2As shown, to verify the advantages of this method over traditional data augmentation methods and compare it with other commonly used data augmentation methods, the accuracy of this method in the test reached 85%, significantly better than SMOTE (76%), GAN (78%) and random interpolation (74%). This indicates that the bilinear interpolation and forgetting weight mechanism effectively improve the data quality. This method can generate more representative medical data and improve the generalization ability of the model.

[0035] In addition, to verify the reasonable distribution of the generated data, the kernel density distribution is visualized in two feature dimensions, as Figure 3 shown. The experimental results show that the kernel density distribution of the samples generated by this method highly coincides with the real data, indicating that the optimization mechanism based on the Euler-Lagrange equation effectively maintains the data distribution characteristics.

[0036] Example 3 By simulating performance tests under different degrees of data distribution drift, the adaptability of the feedforward neural network model to dynamic data changes is verified. In the experiment, this method is compared with the standard feedforward neural network (FFN), batch normalization FFN and L2-regularized FFN, as Figure 4 shown. The results show that as the drift degree increases from 0 to 1, the accuracy of the traditional methods drops by more than 30%, while this method only drops by 4.2%. Its drift compensation matrix significantly alleviates the negative impact brought by the feature distribution shift by dynamically adjusting the weights. Especially in the scenario of non-linear fluctuations in hormone levels caused by disease progression, the accuracy always remains stable above 85%, indicating the specific optimization for the time-series drift of medical data.

[0037] Regarding the challenge of medical high-dimensional feature processing, the model performance when the number of features increases from 10 dimensions to 40 dimensions is analyzed in the experiment, as Figure 5 shown. Compared with the traditional FFN, the AUC of this method only drops by 5.9% under 40-dimensional features, while the traditional method drops by up to 20.7%. This indicates that the feature correlation compensation mechanism calculates the correlation between features through the Pearson correlation coefficient and combines the feature importance extracted by the random forest, enabling the model to effectively focus on key biomarkers in high-dimensional scenarios such as gene expression data and reducing the interference of redundant features, demonstrating the advantage of multi-dimensional feature collaborative compensation.

[0038] By injecting Gaussian noise to simulate clinical data acquisition errors, the noise robustness of the model is tested, as Figure 6 shown. When the noise level reaches 0.3, the F1 score of this method only drops by 9.8%, while the traditional method drops by 21.3%. This indicates that the dynamic offset compensation mechanism is effective. It real-time corrects the network output based on the feature change rate and combines the weight decay of the irrelevant features by the feature aggregation layer, enabling the model to accurately identify key indicators in noisy data.

[0039] Although the specific embodiments of the invention have been described above in conjunction with the accompanying drawings, they are not intended to limit the scope of protection of the present invention. Based on the technical solutions of the present invention, various modifications or variations that can be made by those skilled in the art without creative efforts are still within the scope of protection of the present invention.

Claims

1. A prognosis prediction method based on the combination of aging markers and stroke clinical data, characterized in that: The following steps are involved: S1. Data collection: Collect training data, including aging marker data and stroke clinical data, and then standardize the data format of all the collected training data and store them in the data warehouse; S2. Data preprocessing: clean the aging marker data and stroke clinical data, then perform multimodal data fusion on the two data, and finally normalize the fused data; S3. Data annotation: Experts manually annotate the aging marker data and stroke clinical data, and the annotated categories are the recovery status of the patients corresponding to the data; S4. Data enhancement: The collected aging marker data and stroke clinical data were expanded using the bilinear interpolation method based on the Euler-Lagrange equation to obtain the expanded training set; S5. Build a machine learning model and perform training: Build a machine learning model based on a drift-compensated feedforward neural network, input the training set data into the machine learning model for training, and perform multiple iterations; S6. Prognosis prediction: Select the optimal model from multiple iterations of machine learning models, input new aging marker data and stroke clinical data, and obtain the final prediction results.

2. The prognosis prediction method based on the combination of aging markers and stroke clinical data according to claim 1, characterized in that: S1 is as follows: Aging marker data: Aging marker data is collected through professional biochemical testing equipment and methods. Aging marker data comes from medical clinical trials, long-term biomedical research data sets, and genome and biomarker test results of aging markers. Data types include but are not limited to specific molecular levels in the blood, gene expression data, hormone levels and other biomarker data; Stroke clinical data: Stroke clinical data comes from the clinical research system. The data include the patient's basic information, medical history, physical examination results, imaging examination results, treatment methods, and clinical observation data of the disease process.

3. The prognosis prediction method based on the combination of aging markers and stroke clinical data according to claim 2, characterized in that S2 The details are as follows: S2.

1. Data cleaning: For outliers, a local weighted regression method is used to detect and remove abnormal data points; For missing values, K nearest neighbor algorithm was used for interpolation; S2.2, Multimodal data fusion: By encoding each data type separately into a high-dimensional vector, a weighted strategy is adopted to perform weighted fusion of aging marker data and stroke clinical data to obtain fused features; S2.3, Normalization: The fused features are normalized by the maximum-minimum normalization method to make all feature scales consistent.

4. The prognosis prediction method based on the combination of aging markers and stroke clinical data according to claim 3, characterized in that: The patients' recovery status included: complete recovery, partial recovery, and no recovery.

5. The prognosis prediction method based on the combination of aging markers and stroke clinical data according to claim 4, characterized in that: S4 is as follows: The dataset consisting of aging marker data and stroke clinical data after data collection, preprocessing and annotation is represented as ,in, Indicates The characteristics of a medical data sample, represents the number of medical data samples, express The index of ; S4.

1. Adaptive mechanism for datasets Each medical data sample in is matched with a forgetting weight, and the calculation formula is as follows: , in, Indicates The forgetting weight of the features of medical data samples, represents the first hyperparameter used to adjust the forgetting strength, represents the second hyperparameter used to adjust the forgetting strength, Indicates The importance evaluation function of each medical data sample feature calculates the similarity between the medical data sample features through similarity measurement, and then evaluates the importance of each medical data sample feature. The calculation formula is as follows: , in, express Another index of , , Indicates Features of medical data samples, represents the sensitivity parameter controlling the similarity, represents the L2 norm; S4.

2. Dataset Perform bilinear interpolation on the medical data sample features in to generate new medical data sample features , Represents the new medical data sample features generated after bilinear interpolation; S4.3, using the optimization mechanism based on the Euler-Lagrange equation, by minimizing the energy function To adjust the interpolation and forgetting weights, and thus optimize the distribution of new medical data sample features; Then by solving the Euler-Lagrange equation , obtain the optimized new medical data sample feature position distribution; S4.

4. New medical data sample features will be generated With the dataset The medical data sample features in the dataset are combined to obtain the expanded training set. , , , .

6. The prognosis prediction method based on the combination of aging markers and stroke clinical data according to claim 5, characterized in that: S5 is as follows: S5.

1. The training set The medical data sample features in are input into the machine learning model, and the drift compensation matrix is ​​calculated according to the input medical data sample features. The calculation formula is as follows: , in, represents the drift compensation matrix, represents the number of features of medical data samples input into the machine learning model, express The index of express Another index of Indicates The importance coefficient of the sample features of medical data, Indicates and The correlation coefficient between the characteristics of medical data samples, Indicates Features of medical data samples, Indicates Features of medical data samples, express and The covariance of S5.

2. Training set The sample features of medical data in the feedforward neural network are input into the machine learning model for training. During the forward propagation process of the feedforward neural network, the feedforward neural network performs forward propagation calculation layer by layer, and the output of each layer is fine-tuned in combination with the drift compensation matrix. When the distribution of a sample feature drifts, the corresponding network layer automatically adjusts its weight to compensate for the influence of the change in the distribution of the sample features of the medical data. Drift compensation is achieved by adding a weighted drift compensation matrix to the output of each layer. The output of the feedforward neural network after adjustment is as follows: , , in, represents the output of the feedforward neural network after compensation adjustment, Represents the operation of the Sigmoid activation function, represents the weight matrix of the feedforward neural network, Represents the training set input to the feedforward neural network The characteristics of medical data samples in represents the bias term of the feedforward neural network, represents the drift compensation coefficient, represents the offset term based on the rate of change of the input medical data sample characteristics, Indicates The offset adjustment coefficient of the medical data sample characteristics, The offset adjustment coefficient of the medical data sample feature is taken as the normalized The value of the medical data sample feature, express For the number of iterations The rate of change, Represents integer variable operations in computer language; S5.

3. During the back propagation process of feedforward neural network training, the gradient path is backtracked, and the historical gradient correction term is used to backtrack, record the historical gradient value of the gradient update, and adjust the current gradient direction to optimize the gradient flow process. The calculation formula is as follows: , in, represents the gradient correction term of the current layer of the feedforward neural network, represents the gradient of the loss function of the feedforward neural network, represents the gradient correction term of the previous layer of the current layer of the feedforward neural network, Represents the attenuation coefficient of the historical gradient. For the first layer of the feedforward neural network, its gradient correction term is set to 0.1; The attenuation coefficient of the historical gradient is calculated by the modulus of the gradient of the feedforward neural network loss function. The calculation formula is as follows: , in, represents the learning rate control factor, Represents the magnitude of the gradient of the loss function of the feedforward neural network; S5.

4. After each iteration of the feedforward neural network, the intermediate process features are enhanced, and the output of the current iterative feedforward neural network is aggregated to filter out irrelevant features. The calculation formula for feature aggregation is as follows: , in, represents the output after feature aggregation, represents the number of output features during the iterative process of the feedforward neural network, express The index of express Another index of represents the first Features, represents the first Features, represents the first The importance coefficient of each feature; S5.

5. After each iteration of the feedforward neural network, the learning rate of the feedforward neural network is adjusted. The learning strategy is dynamically adjusted according to the output of the current layer and the error of the previous layer. The learning rate adjustment formula of the feedforward neural network is as follows: , in, represents the learning rate of the feedforward neural network, represents the number of layers of the feedforward neural network, Represents the feedforward neural network The error of the layer, Represents the feedforward neural network The error of the layer, represents the first error adjustment factor, represents the second error adjustment factor, represents the error history correction factor; Combined with the changing trend of historical gradients, the current learning rate is dynamically adjusted to obtain the error history correction factor. The calculation formula is as follows: , in, represents the adjustment coefficient of the error history correction factor; S5.

6. Calculate the class probability through the output of the last layer of the preset Softmax function feedforward neural network, and take the class with the largest class probability as the prediction result. The predicted categories include complete recovery, partial recovery and no recovery. S5.

6. Repeat iterative steps S5.1 to S5.5 until the preset stop iteration condition is met and the machine learning model training is completed. The preset stop iteration condition is the number of iterations.

7. The prognosis prediction method based on the combination of aging markers and stroke clinical data according to claim 6, characterized in that S6 The details are as follows: S6.1, Model loading: Determine the optimal model based on manual annotation and prediction results, save the optimal model parameters in the machine learning model training process, and load the trained feedforward neural network model from it; S6.

2. Input data preprocessing: Input new aging marker data and stroke clinical data, and perform data cleaning, missing value filling, and feature normalization on the new data; S6.

3. Prediction and classification: Input the preprocessed new data into the trained feedforward neural network model for prediction to obtain the final classification prediction result.

Citation Information

Patent Citations

  • A method for establishing a hyperuricemia prediction model using machine learning algorithms

    CN118486471B

  • A bladder tumor prognosis prediction method based on machine learning

    CN118888159B

  • Cerebral stroke prognosis prediction method and system

    CN119417747A