Brain age prediction system and method based on improved local linear embedding and sparse coding

Through improved local linear embedding and sparse coding technology, combined with the sparse constraint terms of L1 norm and L0 norm and the adaptive step size adjustment mechanism, the feature extraction and model training process are optimized, and the problems of low brain age prediction accuracy and high model complexity in the existing technology are solved, achieving a more efficient and robust brain age prediction system.

CN120163791APending Publication Date: 2025-06-17ANHUI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510261929.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prediction accuracy of brain age prediction methods in the prior art is low, and complex features of brain images cannot be captured. The model is complex and has poor generalization ability, making it difficult to generalize to clinical applications.

Method used

A brain age prediction system based on improved local linear embedding and sparse encoding is adopted. Through MRI image preprocessing, feature extraction, sparse processing, feature selection and regression prediction, combined with L1 norm and L0 norm as sparse constraint terms, an adaptive step adjustment mechanism and feature selection module are introduced to optimize model parameters and feature selection methods.

Benefits of technology

It significantly improves the accuracy and robustness of brain age prediction, reduces the complexity of the model, improves the generalization ability and operation efficiency of the model, simplifies the promotion of clinical applications, and can be applicable to various types of MRI images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163791A_ABST
    Figure CN120163791A_ABST
Patent Text Reader

Abstract

The invention discloses a brain age prediction system and method based on improved local linear embedding and sparse coding, and relates to the technical field of brain age prediction. The system comprises an MRI image preprocessing module, the output end of the MRI image preprocessing module is electrically connected with a feature extraction module, and the output end of the feature extraction module is electrically connected with a sparse processing module. According to the method, improved local linear embedding and sparse coding technologies are combined, the accuracy and robustness of brain age prediction are realized, and the precision of feature extraction and sparse processing is improved by introducing an L1 norm and an L0 norm as sparse constraint terms; meanwhile, through an adaptive step length adjustment mechanism and an iterative optimization strategy, the precision and robustness of brain age prediction are further improved, in addition, a feature selection module is further introduced to further screen features subjected to sparse processing so as to remove redundant and noise features, and the explanatory and generalization ability of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of numerical control machining, and particularly to a brain age prediction system and method based on improved locally linear embedding and sparse coding. Background Art

[0002] The existing technologies mainly rely on traditional brain age prediction methods, such as methods based on statistical analysis and methods based on deep learning. Statistical analysis methods usually predict brain age by calculating statistical features (such as gray values, texture features, etc.) of MRI images, but the prediction accuracy is low and complex features of brain images cannot be captured; while deep learning methods, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), etc., require a large amount of training data, have high requirements for preprocessing of input data, have high model complexity, are difficult to interpret, and are prone to overfitting.

[0003] Traditional statistical analysis methods have high requirements for data quality and limited prediction accuracy, and cannot capture complex features of brain images, resulting in inaccurate prediction results; deep learning methods require a large amount of data, have a complex training process, are difficult to interpret, are not conducive to the popularization of clinical applications, and are prone to overfitting, resulting in poor generalization ability of the model.

[0004] Therefore, we provide a brain age prediction system and method based on improved locally linear embedding and sparse coding to solve the above problems. Summary of the Invention

[0005] The purpose of the present invention is to provide a brain age prediction system and method based on improved locally linear embedding and sparse coding. Through the cooperation of improved locally linear embedding, an adaptive step size adjustment mechanism and an iterative optimization strategy, sparse coding technology, introducing L1 norm and L0 norm as sparse constraint terms, and introducing a feature selection module, the problems in the prior art of inaccurate prediction results, high model complexity, poor generalization ability, and inconvenience for the popularization of clinical applications are solved.

[0006] To solve the above technical problems, the present invention is realized through the following technical solutions:

[0007] The present invention is a brain age prediction system based on improved local linear embedding and sparse coding, including an MRI image preprocessing module. The output end of the MRI image preprocessing module is electrically connected to a feature extraction module. The output end of the feature extraction module is electrically connected to a sparsification processing module. The output end of the sparsification processing module is electrically connected to a feature selection module. The output end of the feature selection module is electrically connected to a regression prediction module. The output end of the regression prediction module is electrically connected to a result evaluation module. The output end of the result evaluation module is electrically connected to a model optimization module. The output end of the model optimization module is electrically connected to the input end of the MRI image preprocessing module.

[0008] The present invention is further configured such that the MRI image preprocessing module includes an image reading sub-module, an image normalization sub-module, an image denoising sub-module, an image registration sub-module, and an image segmentation sub-module. The image reading sub-module, the image normalization sub-module, the image denoising sub-module, the image registration sub-module, and the image segmentation sub-module are connected in series. The feature extraction module includes a feature extraction sub-module based on local linear embedding and other feature extraction sub-modules. The feature extraction sub-module based on local linear embedding and other feature extraction sub-modules are connected in parallel.

[0009] The present invention is further configured such that the sparsification processing module includes a dictionary learning sub-module and a sparse coding sub-module. The dictionary learning sub-module and the sparse coding sub-module are connected in series. The feature selection module includes a feature evaluation sub-module and a feature screening sub-module. The feature evaluation sub-module and the feature screening sub-module are connected in series.

[0010] The present invention is further configured such that the regression prediction module includes a model selection sub-module, a model training sub-module, and a brain age prediction sub-module. The model selection sub-module, the model training sub-module, and the brain age prediction sub-module are connected in series. The result evaluation module includes an index calculation sub-module and a result visualization sub-module. The index calculation sub-module and the result visualization sub-module are connected in series.

[0011] The present invention is further configured such that the model optimization module includes a parameter adjustment sub-module and an algorithm improvement sub-module. The parameter adjustment sub-module and the algorithm improvement sub-module are connected in series.

[0012] Method of a brain age prediction system based on improved local linear embedding and sparse coding, the method comprising the following steps: Step a: MRI image preprocessing: The MRI image preprocessing module sequentially performs denoising, normalization, registration, and segmentation operations on the input MRI image. Denoising reduces noise interference, normalization unifies the image gray scale range, registration aligns the anatomical structures of different images, and segmentation extracts the brain regions. The preprocessed image provides an accurate data basis for subsequent feature extraction; Step b: Feature extraction: The feature extraction module processes the preprocessed MRI image using an improved local linear embedding algorithm. The algorithm introduces the L1 norm as a sparse constraint term and captures the internal structural relationship of the features with the help of the Laplacian matrix, thereby improving the accuracy and robustness of feature extraction. By adjusting the algorithm parameters such as the neighborhood size and the embedding dimension, the optimal feature representation is obtained; Step c: Sparsification processing: The sparsification processing module sparsifies the extracted features using an improved sparse coding algorithm. The algorithm uses the L0 norm as a sparse constraint term and introduces an adaptive step size adjustment mechanism (ASSA) to reduce the feature dimension and improve the feature sparsity, thereby improving the brain age prediction accuracy. Adjust the sparsification parameters such as the sparsity and the number of iterations to achieve the optimal sparsification effect; Step d: Feature selection: The feature selection module screens the sparsified features, removes redundant and noisy features, and retains the features that are most valuable for brain age prediction; Step e: Regression prediction: The regression prediction module inputs the screened features into a linear regression model for brain age prediction. By adjusting the model parameters such as the regularization parameter and the learning rate, the optimal prediction performance is achieved; Step f: Result evaluation: The result evaluation module evaluates the prediction result based on indicators such as accuracy, recall rate, mean square error, and F1 score. According to the evaluation result, adjust the model parameters and the feature selection method to improve the prediction accuracy and robustness; Step g: Model optimization: The model optimization module inputs the adjusted model parameters and the feature selection method back into the feature extraction module, the sparsification processing module, and the regression prediction module for iterative optimization until the prediction result meets the actual requirements; Step h: Result output and visualization: Output the final brain age prediction result and provide a visualization interface. The interface displays information such as the prediction result and the feature importance to help users understand the model prediction process and result.

[0013] The present invention is further configured such that the local linear embedding algorithm in the step b includes the following steps:

[0014] ① Data preparation: The preprocessed MRI image dataset, denoted as X = x1, x2,..., x n

[0015] where x i is a high-dimensional vector representing the features of the i-th MRI image (which can be the extracted pixel values, texture features, etc.);

[0016] ② Local neighborhood selection: For each data point xi , use the k-nearest neighbor algorithm to find its k nearest neighbor points; for example, if n = 1000 (i.e., 1000 MRI images) and k = 10 (i.e., find 10 neighbors for each point), then for each x i , the 10 nearest neighbors will be found;

[0017] ③ Calculation of the weight matrix W:

[0018] In the local linear embedding (LLE) of this scheme, the weight w ij is determined by minimizing the reconstruction error;

[0019] The reconstruction error is the error of approximating the current point x i using a linear combination of neighbor points; specifically, the objective function is:

[0020]

[0021] where neighbors(i) represents the set of neighbors of x i ;

[0022] At the same time, the weight w ij needs to satisfy:

[0023] w ij = 0 if x j is not a neighbor of x i , and

[0024] ④ Introduce the L1-norm sparse constraint:

[0025] To increase sparsity, modify the objective function for weight calculation and introduce the L1-norm as a sparse constraint term:

[0026]

[0027] where λ is the regularization parameter used to balance the reconstruction error and sparsity; for example, if λ = 0.1, it means that while minimizing the reconstruction error, the weight w ij is also made as sparse as possible (i.e., many weights are close to 0);

[0028] ⑤ Construct the Laplacian matrix L:

[0029] Use the weight matrix W to construct the Laplacian matrix L, which is defined as:

[0030] L = D - W

[0031] where D is a diagonal matrix whose diagonal elements

[0032] ⑥ Feature mapping:

[0033] Finally, solve the following optimization problem to find the low-dimensional embedding \(Y = \{y_1, y_2, \cdots, y\}\) n

[0034]

[0035] subject to \(Y\) T \(DY = I\)

[0036] where \(tr(Y^TLY)\) represents the trace of the matrix \(Y^TLY\) (i.e., the sum of the diagonal elements), and \(I\) is the identity matrix; T \(Y^TLY\), and \(I\) is the identity matrix; T the trace of the matrix \(Y^TLY\) (i.e., the sum of the diagonal elements), and \(I\) is the identity matrix;

[0037] The optimization problem can be achieved by solving the generalized eigenvalue problem \(LY=\lambda DY\);

[0038] Specifically, find the eigenvectors corresponding to the \(d\) smallest non-zero eigenvalues of \(L\) (where \(d\) is the embedding dimension), and these eigenvectors form the low-dimensional embedding \(Y\).

[0039] The present invention is further configured such that the improved sparse coding algorithm in step c is as follows: there is an input feature vector \(x\in\mathbb{R}\) n , a dictionary matrix \(D\in\mathbb{R}\) n×m , (where \(m > n\)), and a sparse weight vector \(s\in\mathbb{R}\) m , and the goal of sparse coding is to find \(s\) such that \(Ds\) is as close to \(x\) as possible while \(s\) is sparse;

[0040] The objective function is expressed as:

[0041]

[0042] where \(\|\cdot\|_2\) is the \(L_2\) norm (Euclidean norm), \(\|s\|_0\) is the \(L_0\) norm, and \(\lambda\) is a regularization parameter used to balance the reconstruction error and sparsity;

[0043] To further improve sparsity, an adaptive step size adjustment mechanism (ASSA) can also be introduced in this solution. ASSA accelerates the convergence of the sparse weight vector \(s\) by dynamically adjusting the step size during the optimization process. Specifically, the step size \(\lambda\) is adjusted according to the sparsity of the current weight vector in each iteration.

[0044] The present invention is further configured such that the feature selection method in step d uses a statistics-based feature selection method, and the chi-square test is used to evaluate the strength of the relationship between the features and the target variable (brain age);

[0045] The chi-square test includes the following steps:

[0046] For each feature, construct a 2x2 contingency table, where the rows represent the different categories of the feature and the columns represent the classification of brain age;

[0047] Calculate the chi-square statistic for each contingency table using the following formula:

[0048]

[0049] where is the observed frequency and is the expected frequency;

[0050] Based on the degrees of freedom and the significance level, look up the chi-square distribution table to determine the critical value;

[0051] If the calculated chi-square statistic is greater than the critical value, reject the null hypothesis and conclude that there is a significant association between the feature and the brain age classification;

[0052] Sort all the features according to the magnitude of the chi-square statistic and select the feature with the strongest association with the brain age classification.

[0053] The present invention is further configured such that in step e, the regression model has the form y = wx + b, where y is the predicted brain age, x is the selected feature vector, w is the weight vector, and b is the bias term;

[0054] S1: Model parameter initialization: Assume there is a set of data containing 10 features. The model parameters are initialized as follows: w = [0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0], b = 0, regularization parameter λ = 0.01, learning rate η = 0.01;

[0055] S2: Construct the loss function: Use the mean squared error. The loss function formula is:

[0056] L(w,b) = 1 / 2N * ∑(yi - wxi - b) 2 + λ / 2 * ||w|| 2

[0057] where N is the number of samples, yi is the true value, and xi is the feature vector;

[0058] S3: Model training: Use the gradient descent algorithm for training. The iteration formula is:

[0059]

[0060] S4: Model evaluation and optimization: Use the validation dataset to evaluate the model performance and adjust the regularization parameter λ and the learning rate η according to the mean squared error metric;

[0061] S5: Repeat steps S3 to S4 until a model with optimal prediction performance is obtained;

[0062] S6: Use the optimized linear regression model for brain age prediction. Input the new feature vector into the model and calculate the predicted brain age value y = wx + b.

[0063] The present invention has the following beneficial effects:

[0064] 1. By integrating the improved locally linear embedding algorithm and sparse coding technology, the present invention significantly improves the accuracy and robustness of brain age prediction. This combination not only optimizes the feature extraction process but also enhances the feature expression ability and model generalization ability by introducing the L1 norm and L0 norm as sparse constraint terms. The improved LLE algorithm captures the internal structural relationship between features by minimizing the reconstruction error and introducing the Laplacian matrix, while the sparse coding technology reduces the feature dimension and improves sparsity through L0 norm optimization and the adaptive step size adjustment mechanism (ASSA). The combination of the two enables the model to more accurately capture the complex features related to brain age when processing MRI images.

[0065] 2. By reducing the model complexity, the present invention reduces the dependence on computing resources, which not only improves the running efficiency of the model but also reduces the deployment and maintenance costs. In terms of model interpretability, the feature selection module further screens out the most valuable features for brain age prediction, which not only improves the transparency of the model but also facilitates clinicians to understand and trust the prediction results of the model, thus promoting the popularization of the model in actual clinical applications.

[0066] 3. The present invention can be applied to various types of MRI images. Regardless of the source and quality of the MRI images, the model can provide accurate brain age prediction. This wide applicability makes it not only have important application value in the research field but also have potential practical value in clinical diagnosis and treatment. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments.

[0068] Figure 1 It is a system diagram of a brain age prediction system and method based on improved locally linear embedding and sparse coding;

[0069] Figure 2 It is a flowchart of image feature extraction in a brain age prediction system and method based on improved locally linear embedding and sparse coding;

[0070] Figure 3 It is a flowchart of model training in a brain age prediction system and method based on improved locally linear embedding and sparse coding;

[0071] Figure 4It is the overall flowchart of the brain age prediction system and its method based on improved local linear embedding and sparse coding. Detailed implementation manners

[0072] Next, the technical solutions in the embodiments of the present invention will be described with reference to the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0073] Embodiment 1

[0074] Please refer to Figures 1-4 , a brain age prediction system based on improved local linear embedding and sparse coding, including an MRI image preprocessing module. The output end of the MRI image preprocessing module is electrically connected to a feature extraction module. The output end of the feature extraction module is electrically connected to a sparsification processing module. The output end of the sparsification processing module is electrically connected to a feature selection module. The output end of the feature selection module is electrically connected to a regression prediction module. The output end of the regression prediction module is electrically connected to a result evaluation module. The output end of the result evaluation module is electrically connected to a model optimization module. The output end of the model optimization module is electrically connected to the input end of the MRI image preprocessing module. The MRI image preprocessing module includes an image reading sub-module, an image normalization sub-module, an image denoising sub-module, an image registration sub-module, and an image segmentation sub-module. The image reading sub-module, the image normalization sub-module, the image denoising sub-module, the image registration sub-module, and the image segmentation sub-module are connected in series. The feature extraction module includes a feature extraction sub-module based on local linear embedding and other feature extraction sub-modules. The feature extraction sub-module based on local linear embedding and other feature extraction sub-modules are connected in parallel. The sparsification processing module includes a dictionary learning sub-module and a sparse coding sub-module. The dictionary learning sub-module and the sparse coding sub-module are connected in series. The feature selection module includes a feature evaluation sub-module and a feature screening sub-module. The feature evaluation sub-module and the feature screening sub-module are connected in series. The regression prediction module includes a model selection sub-module, a model training sub-module, and a brain age prediction sub-module. The model selection sub-module, the model training sub-module, and the brain age prediction sub-module are connected in series. The result evaluation module includes an index calculation sub-module and a result visualization sub-module. The index calculation sub-module and the result visualization sub-module are connected in series. The model optimization module includes a parameter adjustment sub-module and an algorithm improvement sub-module. The parameter adjustment sub-module and the algorithm improvement sub-module are connected in series.

[0075] Method of a brain age prediction system based on improved local linear embedding and sparse coding, the method comprising the following steps: Step a: MRI image preprocessing: The MRI image preprocessing module sequentially performs denoising, normalization, registration, and segmentation operations on the input MRI image. Denoising reduces noise interference, normalization unifies the image gray scale range, registration aligns the anatomical structures of different images, and segmentation extracts the brain regions. The preprocessed image provides an accurate data basis for subsequent feature extraction; Step b: Feature extraction: The feature extraction module processes the preprocessed MRI image using an improved local linear embedding algorithm. The algorithm introduces the L1 norm as a sparse constraint term and captures the intrinsic structural relationship of the features with the help of the Laplacian matrix, thereby improving the accuracy and robustness of feature extraction. By adjusting the algorithm parameters such as the neighborhood size and the embedding dimension, the optimal feature representation is obtained. The local linear embedding algorithm comprises the following steps:

[0076] ① Data preparation: The preprocessed MRI image dataset, denoted as X = x1, x2,..., x n

[0077] where x i is a high-dimensional vector representing the features of the i-th MRI image (which can be the extracted pixel values, texture features, etc.);

[0078] ② Local neighborhood selection: For each data point x i , the k-nearest neighbor algorithm is used to find its k nearest neighbor points; for example, if n = 1000 (i.e., 1000 MRI images) and k = 10 (i.e., 10 neighbors are found for each point), then for each x i , 10 neighbors closest to it will be found;

[0079] ③ Calculation of the weight matrix W:

[0080] In the local linear embedding (LLE) of this solution, the weight w ij is determined by minimizing the reconstruction error;

[0081] The reconstruction error is the error of approximating the current point x i using a linear combination of neighbor points; specifically, the objective function is:

[0082]

[0083] where neighbors(i) represents the set of neighbors of x i ;

[0084] At the same time, the weight w ij needs to satisfy:

[0085] w ij = 0, if x j is not a neighbor of xi neighbors, and

[0086] ④ Introduce the L1-norm sparsity constraint:

[0087] To increase sparsity, modify the objective function for weight calculation and introduce the L1-norm as the sparsity constraint term:

[0088]

[0089] where λ is the regularization parameter used to balance the reconstruction error and sparsity; for example, if λ = 0.1, it means that we hope the weights w ij while minimizing the reconstruction error, are also as sparse as possible (i.e., many weights are close to 0);

[0090] ⑤ Construct the Laplacian matrix L:

[0091] Use the weight matrix W to construct the Laplacian matrix L, which is defined as:

[0092] L = D - W

[0093] where D is a diagonal matrix whose diagonal elements

[0094] ⑥ Feature mapping:

[0095] Finally, solve the following optimization problem to find the low-dimensional embedding Y = y1, y2,..., y n

[0096]

[0097] subject to Y T DY = I

[0098] where tr(Y T LY) represents the trace of the matrix Y T LY (i.e., the sum of the diagonal elements), and I is the identity matrix;

[0099] The optimization problem can be achieved by solving the generalized eigenvalue problem LY = λDY;

[0100] Specifically, find the eigenvectors corresponding to the d smallest non-zero eigenvalues of L (where d is the embedding dimension), and these eigenvectors form the low-dimensional embedding Y; Step c: Sparsification processing: The sparsification processing module sparsifies the extracted features using an improved sparse coding algorithm. The algorithm uses the L0 norm as the sparse constraint term and introduces an adaptive step size adjustment mechanism (ASSA) to reduce the feature dimension, improve feature sparsity, and thus enhance the brain age prediction accuracy. Adjust the sparsity and iteration number of the sparsification parameters to achieve the optimal sparsification effect. The improved sparse coding algorithm is as follows: Given an input feature vector x ∈ R n , a dictionary matrix D ∈ R n×m , (where m > n), and a sparse weight vector s ∈ R m , the goal of sparse coding is to find s such that Ds is as close to x as possible while s is sparse;

[0101] The objective function is expressed as:

[0102]

[0103] where, ||·||2 is the L2 norm (Euclidean norm), ||s||0 is the L0 norm, and λ is a regularization parameter used to balance the reconstruction error and sparsity;

[0104] To further improve sparsity, an adaptive step size adjustment mechanism (ASSA) can also be introduced in this solution. ASSA accelerates the convergence of the sparse weight vector s by dynamically adjusting the step size during the optimization process. Specifically, the step size λ is adjusted according to the sparsity of the current weight vector in each iteration; Step d: Feature selection: The feature selection module screens the features after sparsification processing, removes redundant and noisy features, and retains the features most valuable for brain age prediction. The feature selection method uses a statistics-based feature selection method and uses the chi-square test to evaluate the relationship strength between the features and the target variable (brain age);

[0105] The chi-square test includes the following steps:

[0106] For each feature, construct a 2x2 contingency table, where the rows represent different categories of the feature and the columns represent the classification of brain age;

[0107] Calculate the chi-square statistic for each contingency table, and the formula is as follows:

[0108]

[0109] where, is the observed frequency, and is the expected frequency;

[0110] According to the degrees of freedom and significance level, look up the chi-square distribution table to determine the critical value;

[0111] If the calculated chi-square statistic is greater than the critical value, the null hypothesis is rejected, and it is considered that this feature has a significant association with brain age classification;

[0112] According to the magnitude of the chi-square statistic, all features are sorted, and the feature with the strongest association with brain age classification is selected; Step e: Regression prediction: The regression prediction module inputs the filtered features into a linear regression model for brain age prediction. By adjusting the model parameters such as the regularization parameter and the learning rate, optimal prediction performance is achieved. The form of the regression model is y = wx + b, where y is the predicted brain age, x is the filtered feature vector, w is the weight vector, and b is the bias term;

[0113] S1: Model parameter initialization: Assume there is a set of data containing 10 features. The model parameters are initialized as follows: w = [0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0], b = 0, regularization parameter λ = 0.01, learning rate η = 0.01;

[0114] S2: Construct the loss function: Use the mean squared error. The formula for the loss function is:

[0115] L(w,b) = 1 / 2N * ∑(yi - wxi - b) 2 + λ / 2 * ||w|| 2

[0116] where N is the number of samples, yi is the true value, and xi is the feature vector;

[0117] S3: Model training: Use the gradient descent algorithm for training. The iterative formula is:

[0118]

[0119] S4: Model evaluation and optimization: Use the validation dataset to evaluate the model performance, and adjust the regularization parameter λ and the learning rate η according to the mean squared error metric;

[0120] S5: Repeat steps S3 to S4 until a model with optimal prediction performance is obtained;

[0121] S6: Use the optimized linear regression model to predict brain age. Input the new feature vector into the model and calculate the predicted brain age value y = wx + b; Step f: Result evaluation: The result evaluation module evaluates the prediction results based on the indicators of accuracy, recall rate, mean square error, and F1 score. According to the evaluation results, adjust the model parameters and feature selection methods to improve the prediction accuracy and robustness; Step g: Model optimization: The model optimization module re-enters the adjusted model parameters and feature selection methods into the feature extraction module, sparsification processing module, and regression prediction module for iterative optimization until the prediction results meet the actual requirements; Step h: Result output and visualization: Output the final brain age prediction results and provide a visualization interface. The interface displays the prediction results and information on feature importance to help users understand the model prediction process and results.

[0122] Example 2

[0123] For example, it is necessary to develop a brain age prediction system to assist in the early diagnosis of neurodegenerative diseases. If there is an MRI brain image dataset of 1000 volunteers, these data include individuals of different ages, genders, and health conditions. The dataset contains the following features:

[0124] Age: The actual age of the volunteer, ranging from 20 to 80 years old.

[0125] Gender: Male and female.

[0126] Health condition: Including healthy individuals and those with mild cognitive impairment.

[0127] MRI image: The brain MRI image of each volunteer, with a resolution of 256x256 pixels and a gray value range from 0 to 255.

[0128] MRI Image Preprocessing

[0129] Denoise each MRI image using a Gaussian filter with a standard deviation of 1.5.

[0130] Normalize the image by scaling the pixel values to the range from 0 to 1.

[0131] Register the images to ensure that all images are in the same spatial position.

[0132] Use a segmentation algorithm to extract the brain region, and the extracted brain region accounts for 60% of the image area.

[0133] Feature Extraction

[0134] Use the improved locally linear embedding (LLE) algorithm to extract features.

[0135] Set the neighborhood size k = 10 and the embedding dimension d = 50.

[0136] The L1 norm is introduced as a sparse constraint term, and the regularization parameter λ = 0.1.

[0137] Construct the Laplacian matrix L and solve the generalized eigenvalue problem to find the low-dimensional embedding Y.

[0138] Sparsification

[0139] An improved sparse coding algorithm is adopted, using the L0 norm as the sparse constraint term. An adaptive step-size adjustment mechanism (ASSA) is introduced, with the initial step size set to 0.01, and the step size is dynamically adjusted according to the sparsity of the weight vector in each iteration. Adjust the sparsification parameter, such as the sparsity α = 0.05, and the number of iterations T = 200 times.

[0140] Feature selection

[0141] Use the chi-square test to evaluate the independence between features and the brain age classification label. Construct a 2x2 contingency table, calculate the chi-square statistic, and compare it with the critical value. Select the features with a p-value less than 0.05. According to the magnitude of the chi-square statistic, select the top 20 features with the strongest correlation with the brain age classification.

[0142] Regression prediction

[0143] Input the selected features into the linear regression model for brain age prediction. Adjust the model parameters, such as the regularization parameter C = 1.0 and the learning rate η = 0.01.

[0144] Result evaluation

[0145] Use metrics such as accuracy, recall, mean squared error (MSE), and F1 score to evaluate the prediction results. Adjust the model parameters and feature selection methods according to the evaluation results to improve the prediction accuracy and robustness.

[0146] Model optimization

[0147] Re-input the adjusted model parameters and feature selection methods into the feature extraction module, sparsification processing module, and regression prediction module for iterative optimization. The optimization process continues until the MSE of the prediction result is lower than the preset threshold, such as MSE < 10.

[0148] Output result

[0149] Output the final brain age prediction results, showing the prediction results and feature importance. The interface displays the predicted brain age, actual age, prediction error, and feature contribution degree of each volunteer.

[0150] Example 3

[0151] Data preparation

[0152] This embodiment collected 1000 brain MRI images from a large hospital, covering healthy people and patients with neurological diseases of different age groups (10 - 80 years old). The data storage format is DICOM, stored in the hospital's PACS system, and transmitted to the MRI image preprocessing module of the brain age prediction system through a data interface.

[0153] MRI Image Preprocessing

[0154] Image Reading Sub-module: Read DICOM format MRI image data from the PACS system, convert it into a format that the system can process, and at the same time read relevant metadata of the image, such as patient age, gender, etc.

[0155] Image Normalization Sub-module: Adopt the Z-score normalization method to unify the image gray values to the range with a mean of 0 and a standard deviation of 1, ensuring the comparability of gray features between different images.

[0156] Image Denoising Sub-module: Use a three-dimensional Gaussian filtering algorithm to remove image noise. According to the resolution and noise level of the image, set the standard deviation of the Gaussian kernel to 2, effectively reducing the random noise in the image and improving the image quality.

[0157] Image Registration Sub-module: Select a rigid registration algorithm based on mutual information to register all MRI images to the MNI standard space, aligning the brain anatomical structures of different individuals spatially. After registration, the consistency of brain region positions between different images is significantly improved, providing a good basis for subsequent feature extraction.

[0158] Image Segmentation Sub-module: Use the U-Net model based on deep learning to segment the registered images, accurately extracting brain tissue regions such as gray matter, white matter, and cerebrospinal fluid. The Dice similarity coefficient of the segmentation result reaches above 0.9, ensuring the accuracy of brain region extraction.

[0159] Feature Extraction

[0160] Feature Extraction Sub-module Based on Locally Linear Embedding: For the preprocessed MRI images, set the neighborhood size to 10 and the embedding dimension to 50, and use an improved locally linear embedding algorithm to extract features. By introducing the L1 norm as a sparse constraint term, feature redundancy is effectively reduced; the Laplacian matrix is used to capture the internal structural relationship between features, making the extracted features more representative. For example, when calculating the local reconstruction weights, considering the geometric relationship between features improves the accuracy and robustness of feature extraction.

[0161] Other feature extraction sub-module: The gray-level co-occurrence matrix is used to extract the texture features of the image, and the texture feature parameters such as contrast, correlation, energy, and entropy are calculated in the directions of 0°, 45°, 90°, and 135° of the image. The features extracted based on locally linear embedding are fused with the texture features to obtain a richer feature representation.

[0162] Sparsification processing

[0163] Dictionary learning sub-module: The K-SVD algorithm is used to learn the dictionary, and the size of the dictionary is set to 1000×500. Through the learning of a large amount of feature data, a dictionary that can effectively represent the features is obtained, providing a basis for sparse coding.

[0164] Sparse coding sub-module: An improved sparse coding algorithm is used, with the L0 norm as the sparse constraint term, and an adaptive step size adjustment mechanism (ASSA) is introduced. The sparsity is set to 0.1 and the number of iterations is 20, and the fused features are sparsely coded. After sparsification processing, the feature dimension is reduced from the original 550 dimensions to 100 dimensions, effectively improving the feature sparsity and at the same time enhancing the brain age prediction accuracy.

[0165] Feature selection

[0166] Feature evaluation sub-module: The information gain algorithm is used to evaluate the features after sparsification processing, and the information gain value between each feature and the brain age is calculated. The larger the information gain value, the greater the contribution of the feature to brain age prediction.

[0167] Feature screening sub-module: According to the information gain value, the top 50 features are selected as the final features for brain age prediction, removing redundant and noisy features and retaining the most valuable features for brain age prediction.

[0168] Regression prediction

[0169] Model selection sub-module: The support vector regression (SVR) model is selected as the brain age prediction model because of its good performance in small sample and non-linear regression problems.

[0170] Model training sub-module: The selected 50-dimensional features and the corresponding true brain ages are used as training data to train the SVR model. The kernel function is set to the radial basis function (RBF), the regularization parameter C is adjusted to 10, and the learning rate is 0.01. The model parameters are optimized through cross-validation to improve the generalization ability of the model.

[0171] Brain age prediction sub-module: The trained SVR model is used to predict the brain age of the new MRI image features to obtain the predicted brain age value.

[0172] Result evaluation

[0173] Index calculation sub-module: Calculate metrics such as the accuracy, recall, mean squared error (MSE), and F1 score of the prediction results. For example, on the test set, the accuracy reaches 85%, the recall is 80%, the mean squared error is 3.5, and the F1 score is 0.82.

[0174] Result visualization sub-module: Plot the prediction results and the true brain ages on a scatter plot to visually show the relationship between the two; at the same time, display the feature importance through a bar chart to help users understand the contribution degree of each feature to brain age prediction.

[0175] Model optimization

[0176] Parameter adjustment sub-module: According to the result evaluation metrics, if it is found that the mean squared error is large, adjust the regularization parameter C and the kernel function parameter γ of the SVR model, and re-train and predict the model until the mean squared error reaches an acceptable range.

[0177] Algorithm improvement sub-module: If the model performance still does not meet the expectations after multiple parameter adjustments, consider improving the feature extraction algorithm or the sparse coding algorithm. For example, try to improve the neighborhood search strategy in the locally linear embedding algorithm to improve the accuracy of feature extraction, and then improve the overall performance of the model. After multiple iterative optimizations, the mean squared error of the final prediction result is reduced to 2.5, meeting the actual application requirements.

[0178] Result output and visualization

[0179] Output the final brain age prediction result to the hospital's electronic medical record system, and at the same time display the comparison chart of the predicted brain age and the true brain age and the feature importance ranking on the visualization interface of the system. Doctors can quickly understand the brain age prediction situation of patients and the key features affecting brain age prediction through the visualization interface, providing strong support for clinical diagnosis and treatment.

[0180] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor limit the invention to the specific embodiments described. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the present invention, so that those skilled in the art in the relevant technical field can understand and utilize the present invention well.

Claims

1. A brain age prediction system based on improved local linear embedding and sparse coding, including an MRI image preprocessing module, characterized in that: The output end of the MRI image preprocessing module is electrically connected to a feature extraction module, the output end of the feature extraction module is electrically connected to a sparse processing module, the output end of the sparse processing module is electrically connected to a feature selection module, the output end of the feature selection module is electrically connected to a regression prediction module, the output end of the regression prediction module is electrically connected to a result evaluation module, the output end of the result evaluation module is electrically connected to a model optimization module, and the output end of the model optimization module is electrically connected to the input end of the MRI image preprocessing module.

2. The brain age prediction system based on improved local linear embedding and sparse coding according to claim 1, characterized in that: The MRI image preprocessing module includes an image reading submodule, an image normalization submodule, an image denoising submodule, an image registration submodule, and an image segmentation submodule. The image reading submodule, the image normalization submodule, the image denoising submodule, the image registration submodule, and the image segmentation submodule are connected in series. The feature extraction module includes a feature extraction submodule based on local linear embedding and other feature extraction submodules. The feature extraction submodule based on local linear embedding and other feature extraction submodules are connected in parallel.

3. The brain age prediction system based on improved local linear embedding and sparse coding according to claim 1, characterized in that: The sparse processing module includes a dictionary learning submodule and a sparse coding submodule, and the dictionary learning submodule and the sparse coding submodule are connected in series. The feature selection module includes a feature evaluation submodule and a feature screening submodule, and the feature evaluation submodule and the feature screening submodule are connected in series.

4. The brain age prediction system based on improved local linear embedding and sparse coding according to claim 1, characterized in that: The regression prediction module includes a model selection submodule, a model training submodule and a brain age prediction submodule, and the model selection submodule, the model training submodule and the brain age prediction submodule are connected in series. The result evaluation module includes an indicator calculation submodule and a result visualization submodule, and the indicator calculation submodule and the result visualization submodule are connected in series.

5. The brain age prediction system based on improved local linear embedding and sparse coding according to claim 1, characterized in that: The model optimization module includes a parameter adjustment submodule and an algorithm improvement submodule, and the parameter adjustment submodule and the algorithm improvement submodule are connected in series.

6. The method of the brain age prediction system based on improved local linear embedding and sparse coding according to any one of claims 1 to 5, characterized in that: The method comprises the following steps: Step a: MRI image preprocessing: The MRI image preprocessing module performs denoising, normalization, registration and segmentation operations on the input MRI image in sequence. Denoising reduces noise interference, normalization unifies the image grayscale range, registration aligns the anatomical structures of different images, and segmentation extracts the brain area. The preprocessed image provides an accurate data basis for subsequent feature extraction. Step b: Feature extraction: The feature extraction module uses an improved local linear embedding algorithm to process the preprocessed MRI images. The algorithm introduces the L1 norm as a sparse constraint term and uses the Laplace matrix to capture the intrinsic structural relationship of the features to improve the accuracy and robustness of feature extraction. The optimal feature representation is obtained by adjusting the algorithm parameters of the neighborhood size and embedding dimension. Step c: Sparse processing: The sparse processing module uses an improved sparse coding algorithm to sparse the extracted features. The algorithm uses the L0 norm as the sparse constraint term and introduces an adaptive step size adjustment mechanism (ASSA) to reduce the feature dimension and improve the feature sparsity, thereby improving the accuracy of brain age prediction. The sparsity parameters of sparsity and number of iterations are adjusted to achieve the optimal sparse effect. Step d: Feature selection: The feature selection module screens the features after sparse processing, removes redundant and noisy features, and retains the most valuable features for brain age prediction; Step e: Regression prediction: The regression prediction module inputs the filtered features into the linear regression model for brain age prediction, and achieves the optimal prediction performance by adjusting the model parameters of the regularization parameter and the learning rate; Step f: Result evaluation: The result evaluation module evaluates the prediction results based on the indicators of accuracy, recall, mean square error, and F1 score. Based on the evaluation results, the model parameters and feature selection methods are adjusted to improve the prediction accuracy and robustness. Step g: Model optimization: The model optimization module re-inputs the adjusted model parameters and feature selection methods into the feature extraction module, the sparse processing module and the regression prediction module for iterative optimization until the prediction results meet the actual needs; Step h: Result output and visualization: Output the final brain age prediction results and provide a visualization interface that displays the prediction results and feature importance information to help users understand the model prediction process and results.

7. The method of the brain age prediction system based on improved local linear embedding and sparse coding according to claim 6, characterized in that: The local linear embedding algorithm in step b comprises the following steps: ① Data preparation: preprocessed MRI image dataset, denoted as X = x1, x2, ..., x n where x i is a high-dimensional vector representing the features of the i-th MRI image (which can be extracted pixel values, texture features, etc.); ② Local neighborhood selection: For each data point x i , use the k nearest neighbor algorithm to find its k nearest neighbor points; for example, if n = 1000 (i.e. 1000 MRI images), k = 10 (i.e. find 10 neighbors for each point), then for each x i , will find the 10 nearest neighbors; ③Calculation of weight matrix W: In the local linear embedding (LLE) of this scheme, the weight w ij Determined by minimizing the reconstruction error; The reconstruction error is to approximate the current point x using a linear combination of neighboring points. i The error; specifically, the objective function is: Among them, neighbors(i) represents x i The set of neighbors of ; At the same time, the weight w ij Need to meet: w ij =0, if x j Not x i neighbors, and ④Introduce L1 norm sparse constraint: In order to increase sparsity, the objective function of weight calculation is modified, and the L1 norm is introduced as a sparse constraint term: Among them, λ is a regularization parameter used to balance the reconstruction error and sparsity; for example, if λ = 0.1, it means that the weight w is expected to ij While minimizing the reconstruction error, it is also as sparse as possible (i.e. many weights are close to 0); ⑤Construct the Laplace matrix L: The weight matrix W is used to construct the Laplacian matrix L, which is defined as: L=DW Where D is a diagonal matrix whose diagonal elements ⑥Feature Mapping: Finally, solve the following optimization problem to find the low-dimensional embedding Y = y1, y2, ..., y n subject toY T DY=I where tr(Y T LY) represents the matrix Y T The trace of LY (i.e., the sum of the diagonal elements), I is the identity matrix; The optimization problem can be achieved by solving the generalized eigenvalue problem LY = λDY; Specifically, find the eigenvectors corresponding to the smallest d non-zero eigenvalues ​​of L (where d is the embedding dimension), and these eigenvectors constitute the low-dimensional embedding Y.

8. The method of the brain age prediction system based on improved local linear embedding and sparse coding according to claim 6, characterized in that: The improved sparse coding algorithm in step c is: there is an input feature vector x∈R n , a dictionary matrix D∈R n×m , (where m>n), and a sparse weight vector s∈R m ,The goal of sparse coding is to find s so that Ds is as close to x as possible, while s is sparse; The objective function is expressed as: Among them, ||·||2 is the L2 norm (Euclidean norm), ||s||0 is the L0 norm, and λ is a regularization parameter used to balance the reconstruction error and sparsity; In order to further improve the sparsity, an adaptive step size adjustment mechanism (ASSA) can also be introduced in this scheme. ASSA accelerates the convergence of the sparse weight vector s by dynamically adjusting the step size in the optimization process. Specifically, the step size λ is adjusted according to the sparsity of the current weight vector in each iteration.

9. The method of the brain age prediction system based on improved local linear embedding and sparse coding according to claim 6, characterized in that: The feature selection method in step d uses a statistical feature selection method, and uses a chi-square test to evaluate the strength of the relationship between the feature and the target variable (brain age); The chi-square test consists of the following steps: For each feature, a 2x2 contingency table is constructed, where the rows represent the different categories of the feature and the columns represent the classification of brain age; Calculate the chi-square statistic for each contingency table using the following formula: Among them, is the observed frequency, is the expected frequency; According to the degrees of freedom and significance level, find the chi-square distribution table and determine the critical value; If the calculated chi-square statistic is greater than the critical value, the null hypothesis is rejected and it is considered that the feature is significantly associated with the brain age classification; All features were ranked according to the size of the chi-square statistic, and the features with the strongest correlation with brain age classification were selected.

10. The method of the brain age prediction system based on improved local linear embedding and sparse coding according to claim 6, characterized in that: The regression model in step e is in the form of y=wx+b, where y is the predicted brain age, x is the filtered feature vector, w is the weight vector, and b is the bias term; S1: Model parameter initialization: Assume there is a set of data containing 10 features, and the model parameters are initialized as follows: w = [0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0], b = 0, regularization parameter λ = 0.01, learning rate η = 0.01; S2: Construct loss function: Using mean square error, the loss function formula is: <h2 style=";text-align:left;direction:ltr">L(w,b) = 1 / 2N*∑(yi-wxi-b)<h2 style=";text-align:left;direction:ltr"> 2 <h2 style=";text-align:left;direction:ltr"> +λ / 2*||w||<h2 style=";text-align:left;direction:ltr"> 2 Where N is the number of samples, yi is the true value, and xi is the feature vector; S3: Model training: Use the gradient descent algorithm for training, and the iteration formula is: S4: Model evaluation and optimization: Use the validation dataset to evaluate the model performance and adjust the regularization parameter λ and learning rate η according to the mean square error indicator; S5: Repeat steps S3 to S4 until a model with optimal prediction performance is obtained; S6: Use the optimized linear regression model to predict brain age, input the new feature vector into the model, and calculate the predicted brain age value y=wx+b.