Brain disease classification system based on deformable convolution and Vision transfer
By adopting deformable convolution and Vision Transformer in the brain disease classification system, combining imagingomics and global attention mechanisms, the problems of large data demand, high risk of overfitting and high computing resource consumption in the existing technology are solved, and more efficient and accurate classification of brain disease is achieved.
Patent Information
- Application Number
- CN202510045953.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-09
AI Technical Summary
The existing medical image processing technology has problems such as large data demand, high risk of overfitting and high computing resource consumption in the classification of brain diseases.
The brain disease classification system based on deformable convolution and Vision Transformer is adopted to extract brain region features through imaging omics, perform correlation analysis, and combine deformable convolution and global attention mechanisms to dynamically adjust the shape of the convolution kernel and capture the global image information.
It significantly improves model performance, adds available features, goes beyond the limitations of changes in individual brain regions, fully understands the synergistic relationships between brain regions, and improves the accuracy and overall performance of feature extraction.
Smart Images

Figure CN119963910A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and in particular to a brain disease classification system based on deformable convolution and Vision transformer. Background Art
[0002] In the field of medical image processing, computer-assisted analysis uses structural magnetic resonance imaging (sMRI) to screen the population, providing medical staff with important disease prediction and classification references, especially in the classification of early mild cognitive impairment (EMCI), late mild cognitive impairment (LMCI) and Alzheimer's disease (AD). First, sMRI provides high-resolution brain images that can clearly show structural changes associated with brain diseases, such as hippocampal atrophy, and the lack of ionizing radiation and non-invasiveness make it safer in elderly patients. In addition, sMRI supports early detection and multimodal analysis, improving the accuracy and efficiency of diagnosis. At present, researchers usually use machine learning to analyze patients' sMRI images to identify the characteristics of different stages of Alzheimer's disease.
[0003] Compared with traditional methods, convolutional neural networks (CNNs) have significant advantages, such as automatic extraction of deep features, high accuracy, and strong adaptability, making them excellent in the detection of neurodegenerative diseases such as Alzheimer's disease. CNNs can process large-scale data and reduce the burden of clinical work, but they also have some disadvantages, such as large data requirements, overfitting risks, and high consumption of computing resources. Summary of the invention
[0004] In view of the shortcomings of the prior art, the present invention provides a brain disease classification system based on deformable convolution and Vision transformer. The sMRI data of the brain is subjected to imaging omics, the characteristics of each brain region are extracted, correlation analysis is performed, and the classification results are obtained through the network.
[0005] The present invention adopts the following scheme:
[0006] A brain disease classification system based on deformable convolution and Vision transformer, including:
[0007] Data processing module, including image data preprocessing, feature extraction and feature fusion;
[0008] Network modules, including deformable convolution and Vision transformer networks;
[0009] Classification module, including feature selection and classifier.
[0010] Furthermore, the processing steps of the data processing module are as follows:
[0011] Both T1w data and T2w data were corrected for anterior commissure-posterior commissure. T1w preprocessing included skull removal and tissue segmentation to obtain gray matter images; T1w / T2w preprocessing included radiographic registration and intensity calibration to obtain T1w / T2w images;
[0012] The hippocampal ontology template was selected, and the features of 112 brain regions were extracted using radiomics. A 112×1288 matrix was obtained for each subject, and a correlation analysis was performed on the matrix to obtain a 112×112 matrix. The two correlation matrices of T1w and T1w / T2w were weighted and added to obtain a new 112×112 matrix as the input of the network module.
[0013] Furthermore, the processing steps of the network module are as follows:
[0014] The network module divides the 112×112 matrix into 49 16×16 matrices, performs deformable convolution on each of them, and obtains a new feature map. After passing through a linear projection layer, it is used as the input of the Vision transformer. Finally, the extracted learnable vector output by the Vision transformer is selected as the basis for classification.
[0015] Furthermore, the processing steps of the classification module are as follows: the classification module uses the variance method as the basis for feature selection, and the classifiers used are support vector machine, random forest and naive Bayes.
[0016] Compared with the prior art, the present invention has the following beneficial effects:
[0017] (1) Based on the traditional method of using only T1w data, the present invention integrates T1w and T2w data for classification, effectively increasing the available features of the network, thereby significantly improving the performance of the model.
[0018] (2)2 The present invention focuses on exploring the interconnections between brain regions through in-depth correlation analysis of various brain regions, aiming to reveal the complex interactions between brain regions. This method goes beyond the limitations of studying changes in a single brain region, allowing the study to fully understand the synergistic relationship between brain regions.
[0019] (3) The present invention adopts deformable convolution, which successfully overcomes the defect of fixed convolution kernel size in traditional convolution. By learning the offset, the method can dynamically adjust the position of convolution calculation, so that the shape of the convolution kernel is no longer fixed, thereby significantly improving the model's ability to extract image features.
[0020] (4) The present invention adopts Vision Transformer and can effectively capture the global information in the image by introducing the global attention mechanism. This mechanism enables the model to focus on the relationship between the various parts of the image, thereby more comprehensively understanding the image content and improving the accuracy of feature extraction and overall performance. Through global attention, the model can not only identify local features, but also integrate global information. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a schematic diagram of the overall framework of the present invention;
[0022] Figure 2 It is a structural schematic diagram of a data processing module of the present invention;
[0023] Figure 3 It is a structural schematic diagram of the network module of the present invention;
[0024] Figure 4 It is a structural schematic diagram of the classification module of the present invention. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0026] The present invention is further described in detail below in conjunction with the accompanying drawings:
[0027] A brain disease classification system based on deformable convolution and vision transformer, reference Figure 1 ,The present invention is divided into three modules, namely data processing module, network module and classification module.
[0028] refer to Figure 2 ,The structure of the data processing module includes: data preprocessing, ,radiological omics feature extraction, ,correlation analysis and feature fusion.
[0029] T1w data and T2w data were preprocessed to obtain gray matter images and T1w / T2w images, respectively.
[0030] The images were registered to the HOA template, and features were extracted through radiomics. For each subject, 1,288 radiomics features of 112 brain regions in gray matter images and T1w / T2w images were obtained. All radiomics features included three categories, namely shape, intensity, and texture.
[0031] Correlation analysis was performed on the two 112×1288 matrices of each subject to obtain two 112×112 correlation matrices.
[0032] The convolution kernel size of convolution layer 1 is 1×1, and the step size is set to 1. The goal is to fuse the features of the two correlation matrices according to the corresponding positions instead of adding them according to a single weight.
[0033] Normalization layer 1 is batch normalization (BN), which helps improve the stability of the model and the training effect.
[0034] The activation function layer is a Sigmoid function, which aims to calculate the weights to achieve feature fusion based on the values of each corresponding position in the two correlation matrices of each subject. Through this mechanism, the model can effectively integrate the information of different correlation matrices, thereby improving the accuracy and reliability of feature representation.
[0035] After the data processing module, each subject received a matrix of size 112×112 as the input of the network module.
[0036] refer to Figure 3 , the network modules include deformable convolution and Vision transformer networks.
[0037] The convolution kernel size of convolution layer 2 is 16×16, and the stride is set to 16. The purpose is to split the 112×112 matrix output by the data processing module into 49 16×16 sub-matrices. This configuration can effectively extract local features while achieving the required matrix partitioning.
[0038] Convolutional layer 3 is a deformable convolutional layer, which is designed to extract features from 49 16×16 matrices. The deformable convolutional layer introduces a learnable offset so that the convolution kernel can flexibly adapt to the geometry of the input feature map, thereby capturing richer local features and complex spatial relationships.
[0039] The linear projection layer converts 49 16×16 matrices into 49 one-dimensional vectors of length 256. The main purpose of linear projection is to reduce the dimension of features while retaining key information for subsequent processing and analysis.
[0040] The position encoding layer position encodes 49 one-dimensional vectors of length 256, and the feature dimension after embedding is [49, 256]. At the same time, a vector of the same length as the previous one-dimensional vector is introduced, named class token, and spliced into the previous feature. In the process of extracting classification information, the class token will interact with other features to learn the feature information of different one-dimensional vectors. The class token is spliced with the previous vector, and the feature dimension after splicing is [50, 256].
[0041] Normalization (Norm) helps speed up the training process and improve the stability of the model by adjusting the data to a similar scale.
[0042] The multi-head attention layer uses multiple attention heads in parallel to learn different features and relationships of the input data, thereby enhancing feature representation. Each head can focus on different parts of the input, capturing diverse information, making the model more comprehensive in extracting features.
[0043] Multi-layer perceptron layers are able to learn complex nonlinear relationships and extract deep features of the data, which helps to significantly improve the overall classification accuracy of the model.
[0044] The output of the network module is a one-dimensional vector of length 256, which serves as the input of the classification module.
[0045] Reference Figure 4 ,The classification module is composed of the feature selection layer and the ,classification module, and finally obtains the classification probability.
[0046] The feature selection layer is a variance method, which is mainly used to select the most useful features for model prediction from high-dimensional data. The basic idea of this method is to determine the importance of each feature in the data by calculating its variance. The features left after screening are combined into a new feature set for subsequent model training. This can not only reduce the computational complexity, but also improve the performance and generalization ability of the model.
[0047] In terms of classifiers, the present invention selects multiple classifiers, including support vector machines, random forests, and naive Bayes. In the classification tasks of AD vs NC, LMCI vs NC, EMCI vs NC, and LMCI vs EMCI, the three classifiers of support vector machines, random forests, and naive Bayes are respectively applied. By evaluating the classification results of each task, the model with the best classification effect is selected as the final classification result of the present invention.
[0048] The present invention adopts ten-fold cross validation, and the results are shown in Table 1.
[0049] Table 1
[0050]
[0051] Table 1 lists in detail the accuracy (ACC), area under the curve (AUC), sensitivity (SEN), specificity (SPE), precision (PRE) and F1 score (F1 Score, F1) of the results of the present invention. The results show that the present invention has a significant improvement in accuracy.
[0052] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A brain disease classification system based on deformable convolution and vision transformer, characterized in that: include: Data processing module, including image data preprocessing, feature extraction and feature fusion; Network modules, including deformable convolution and Vision transformer networks; Classification module, including feature selection and classifier.
2. A brain disease classification system based on deformable convolution and vision transformer as claimed in claim 1, characterized in that: The processing steps of the data processing module are as follows: Both T1w data and T2w data were corrected for anterior commissure-posterior commissure. T1w preprocessing included skull removal and tissue segmentation to obtain gray matter images. T1w / T2w preprocessing included radiographic registration and intensity calibration to obtain T1w / T2w images. The hippocampal ontology template was selected, and the features of 112 brain regions were extracted using radiomics. A 112×1288 matrix was obtained for each subject, and a correlation analysis was performed on the matrix to obtain a 112×112 matrix. The two correlation matrices of T1w and T1w / T2w are weightedly added to obtain a new 112×112 matrix as the input of the network module.
3. A brain disease classification system based on deformable convolution and vision transformer as claimed in claim 1, characterized in that: The processing steps of the network module are as follows: The network module divides the 112×112 matrix into 49 16×16 matrices, performs deformable convolution on each of them, and obtains a new feature map. After passing through a linear projection layer, it is used as the input of the Vision transformer. Finally, the extracted learnable vector output by the Vision transformer is selected as the basis for classification.
4. A brain disease classification system based on deformable convolution and vision transformer as claimed in claim 1, characterized in that: The processing steps of the classification module are as follows: the classification module uses the variance method as the basis for feature selection, and the classifiers used are support vector machine, random forest and naive Bayes.