Alzheimer's disease MRI diagnosis method based on width neural network learning
By employing a lightweight U-NET and a width neural network method with dynamic contrast-compensated binarization, morphological optimization, and K-means++ feature enhancement, the problems of hippocampal segmentation and subtle variations in early-stage Alzheimer's disease MRI diagnosis were solved, thus improving the accuracy and precision of diagnosis.
Patent Information
- Application Number
- CN202510961760.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-11-04
AI Technical Summary
Existing techniques for MRI diagnosis of Alzheimer's disease are susceptible to interference from cerebrospinal fluid grayscale fluctuations, leading to oversegmentation or undersegmentation of the hippocampal boundaries and difficulty in capturing subtle changes in gray matter density in early-stage patients, thus affecting measurement accuracy and diagnostic precision.
Lightweight U-NET was used to process brain MRI images, combined with dynamic contrast-compensated binarization, morphological optimization, and density peak-guided K-means++ feature extraction. Diagnosis was performed through a wide neural network, including adaptive ROI cropping, local contrast-compensated threshold segmentation, watershed algorithm to fill holes, gradient-guided grayscale backpropagation, and multi-scale feature enhancement.
It effectively eliminates differences in grayscale distribution and the influence of noise, improves the segmentation accuracy of the hippocampus, enhances the ability to capture subtle changes in gray matter density in early Alzheimer's patients, and improves the accuracy and reliability of diagnosis.
Smart Images

Figure CN120895207A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of medical image recognition, and particularly relates to an Alzheimer's disease MRI diagnosis method based on a width neural network learning. BACKGROUND
[0002] Alzheimer's disease (AD) is one of the high-incidence diseases in the elderly, which is a neurodegenerative disease with insidious onset, and is clinically manifested as memory impairment, aphasia, apraxia, agnosia, visual-spatial skill impairment and personality behavior change; magnetic resonance imaging (MRI, Nuclear Magnetic Resonance Imaging) uses the principle of nuclear magnetic resonance to detect electromagnetic waves to draw the internal structure image of an object; it embodies the development of medical image recognition and provides convenience for doctors to diagnose diseases, but there are still many problems in the field of Alzheimer's disease in medical image recognition.
[0003] The prior art usually adopts traditional global threshold binarization or a model based on manually designed features for diagnosis and recognition, but has corresponding defects: in brain MRI, it is easy to be disturbed by cerebrospinal fluid gray level fluctuation, and the traditional global threshold binarization is sensitive to the gray level overlapping area of cerebrospinal fluid and atrophic hippocampus in the MRI image, which is easy to cause hippocampus boundary over-segmentation or under-segmentation, affecting the measurement accuracy, and when a morphological closing operation is used to fill the hole, a false image is easy to be introduced, which destroys the original texture feature. The traditional K-means clustering is sensitive to the initial center, and is easy to fall into a local optimal condition in the AD image, in addition, the model based on the area or the length of the manually designed features is difficult to capture the subtle changes of the gray matter density of the early AD patients. SUMMARY
[0004] In order to solve the technical problems that the existing mainstream technology has limitations in hippocampus segmentation and the subtle changes of early AD patients cannot be captured, the purpose of the application is to provide an Alzheimer's disease MRI diagnosis method based on a width neural network learning, and the technical scheme is as follows:
[0005] Step 1, adaptively ROI cropping the brain MRI image;
[0006] The light-weight U-NET is used to process the brain MRI image, the pyramid pooling is used to enhance the small target detection capability, and the maximum hippocampus volume layer in the brain transverse section is extracted, the center region is cropped, and the ROI image is obtained;
[0007] Step 2, dynamic contrast compensation binarization;
[0008] Based on the local contrast compensation threshold segmentation algorithm, the gray mean value and standard deviation in the 5*5 window of the extracted ROI image are calculated, the binary decision is made, the ROI image is converted to black and white, and the binary mask is dynamically generated;
[0009] Step 3, morphological optimization;
[0010] The binary mask holes are filled by using the watershed algorithm, the area threshold of the binary mask is greater than 50 pixels², and the optimized binary mask is output;
[0011] Step 4, gray scale back;
[0012] The original gray scale information of the ROI image is combined with the optimized binary mask by using the gradient-guided weighted fusion formula, and the optimized image with fused original gray scale information is output;
[0013] Step 5, K-means++ feature extraction guided by density peak value;
[0014] The input image is defined as the optimized image with fused original gray scale information, the initial clustering center is calculated by density peak initialization of the input image, the data points of the input image are classified according to the initial clustering center, and the clustering label is obtained;
[0015] The multi-scale feature construction is performed on the input image and the initial clustering center, and the multi-dimensional feature vector is obtained;
[0016] The dynamic feature weighting is performed on the multi-dimensional feature vector to obtain the weighted feature vector, and the feature mapping enhancement is performed to obtain the final enhanced feature map;
[0017] Step 6, obtain the width neural network, input the final enhanced feature map into the width neural network for final classification prediction, and realize the diagnosis of Alzheimer's disease MRI.
[0018] Preferably, in step 5, the initial clustering center is calculated by density peak initialization of the input image, the data points of the input image are classified according to the initial clustering center, and the clustering label is obtained, specifically:
[0019] The mask mark of the hippocampus region is 1, the gray value of the hippocampus region is extracted to establish a sample set, denoted as , The total number of samples in the sample set is represented by , the neighborhood density of any sample is calculated, and the corresponding calculation formula is:
[0020]
[0021] Among them, the neighborhood density of sample is represented. , denote any sample in the sample set; denote the Euclidean distance of gray value; denote the truncated distance, i.e. the median of the distance between samples;
[0022] Based on the neighborhood density, the minimum distance between the sample set is calculated, and the corresponding calculation formula is:
[0023]
[0024] wherein, denote the minimum distance; denote the neighborhood density of the sample ;
[0025] Determine the initial clustering center in combination with the neighborhood density and the minimum distance, generate a temporary label set, denoted as , denote the total number of samples in the temporary label set, and the corresponding calculation formula is:
[0026]
[0027]
[0028] wherein, denote the th temporary label in the temporary label set; denote the cluster number of the initial clustering center; denote the Euclidean distance of the sample to the initial clustering center; denote the th cluster of the initial clustering center;
[0029] For any cluster, calculate the average gray value of the cluster, update the initial clustering center, and the corresponding calculation formula is:
[0030]
[0031] wherein, denote the th cluster of the updated initial clustering center;
[0032] Based on the hippocampus region weighting, the corresponding calculation formula is:
[0033]
[0034]
[0035] wherein, denote the cluster 0 of the updated initial clustering center; indicates a weight;
[0036] a temporary label set is generated According to the pixel coordinates of the ROI image, a 64*64 matrix is restored, wherein the pixels in the binary mask are filled with label values according to the corresponding positions; the pixels outside the binary mask are assigned with labels; and a label matrix is obtained.
[0037] Preferably, the initial clustering center is determined in combination with the neighborhood density and the minimum distance, and specifically, the initial clustering center is determined as follows:
[0038] The result value obtained by multiplying the neighborhood density and the minimum distance is calculated, denoted as The result value is sorted in descending order according to the result value, and a result value ranking is established.
[0039] The samples with the top several result values are screened as the initial clustering center.
[0040] Preferably, a temporary label set is generated, and specifically, the category division is performed in sequence according to the temporary label set, wherein indicates a high gray region; indicates a medium gray region; indicates a low gray region.
[0041] The present application has the following beneficial effects:
[0042] The binary segmentation method with dynamic contrast compensation is designed, the local adaptive threshold, the morphological topology optimization, and the gradient-guided gray backtracking are used to eliminate the problems of the fixed threshold being unable to adapt to different scanning devices and imaging parameters, leading to the difference in gray distribution, and being sensitive to noise; the multi-scale feature enhancement method driven by K-means++ is designed to capture the local texture fragmentation, edge blur and other subtle abnormalities of the gray matter density of the early AD patients, thereby providing effective guarantee for the accuracy of the early AD patient detection; that is, the present application combines the binary segmentation method with dynamic contrast compensation and the multi-scale feature enhancement method driven by K-means++, and is dedicated to the hippocampus segmentation problem in the image recognition of Alzheimer's disease and the problem that the subtle abnormalities of the early AD patients cannot be captured, thereby providing convenience for doctors. BRIEF DESCRIPTION OF DRAWINGS
[0043] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.
[0044] Figure 1An implementation flowchart of an Alzheimer's MRI diagnosis method based on a width neural network learning is provided for an embodiment of the present application. DETAILED DESCRIPTION
[0045] In order to further clarify the technical means and effects taken by the present application to achieve the predetermined object of the application, the specific implementation, structure, features and effects of an Alzheimer's MRI diagnosis method based on a width neural network learning according to the present application are described in detail as follows in combination with the drawings and preferred embodiments. Different "one embodiment" or "another embodiment" in the following description do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs.
[0047] The specific scheme of the Alzheimer's MRI diagnosis method based on a width neural network learning provided by the present application is specifically described below in combination with the drawings.
[0048] Please refer to Figure 1 which shows an implementation flowchart of an Alzheimer's MRI diagnosis method based on a width neural network learning provided by an embodiment of the present application, the method comprising:
[0049] Step 1, adaptive ROI cropping of brain MRI images;
[0050] The original brain MRI image is processed by a lightweight U-NET, the small target detection capability is enhanced by pyramid pooling, and the maximum hippocampus volume layer in the brain cross-section slice is extracted, the central region is cropped, and the ROI image is obtained.
[0051] Preferably, before processing the original brain MRI image, a gray scale normalization operation and a down-sampling operation are performed in advance to adjust the gray scale value range of the image to a unified standard range and eliminate the gray scale difference; the down-sampling is to reduce the resolution of the image, reduce the calculation amount and storage demand by reducing the number of pixels, and also reduce the influence of noise to a certain extent.
[0052] Specifically, the original brain MRI image is denoted as The corresponding calculation formula is:
[0053]
[0054]
[0055] wherein, represents a grayscale processing of the original brain MRI image ; represents a pixel coordinate of the original brain MRI image ; represents a down-sampling processing of the original brain MRI image , ; represents a pixel coordinate of the original brain MRI image after the down-sampling processing represents a bilinear weight matrix.
[0056] According to the output , a light-weight U-NET feature extraction is adopted, and a small target detection capability is enhanced through pyramid pooling (SPP, Spatial Pyramid Pooling), and the corresponding logical expression is:
[0057]
[0058]
[0059]
[0060]
[0061]
[0062] wherein, represents an output feature map; represents a feature after pyramid pooling; , , all represent feature maps obtained by light-weight U-NET processing in turn; , , represent weight matrices corresponding to the obtained feature maps in turn; represents an activation function; represents a max-pooling processing; represents an average-pooling processing; represents a splicing operation; represents a convolution operation.
[0063] Based on the output , a maximum hippocampus volume layer positioning is performed, including mask prediction and volume calculation, and the corresponding calculation formula is:
[0064]
[0065]
[0066] wherein, denotes a binary mask, ; denotes a key layer index; denotes a weight matrix of the maximum hippocampus volume layer.
[0067] The center region is cropped according to the key layer index, that is, the ROI (Region of Interest) image is cropped after the centroid calculation, and the corresponding calculation formula is:
[0068]
[0069]
[0070]
[0071] wherein, denotes the output ROI image; denotes the pixel coordinates of the original brain MRI image; denotes the centroid.
[0072] Step 2, dynamic contrast compensation binarization
[0073] Based on the threshold segmentation algorithm of local contrast compensation, the average gray value and standard deviation in the 5x5 window of the extracted ROI image are calculated, the binarization decision is made, and the ROI image is converted to pure black and white, and a binary mask is dynamically generated.
[0074] Specifically, each pixel point in the ROI image is traversed, the local mean and standard deviation are calculated with a 5x5 sliding window, and then a dynamic binarization decision is made to obtain a binary mask, denoted as B, and the corresponding calculation formula is:
[0075]
[0076]
[0077]
[0078]
[0079] wherein, denotes the standard deviation; denotes the local mean; denotes dynamic contrast compensation binarization.
[0080] Step 3, morphological optimization
[0081] The watershed algorithm is used to fill the holes in the binary mask, and the area threshold of the binary mask is greater than 50 pixels², and the optimized binary mask is output.
[0082] Specifically, the binary mask B is morphologically optimized by a watershed algorithm, and is sequentially filled with holes by dilation operation and erosion operation, with a kernel size of , and then the largest connected region is reserved, where the largest connected region refers to a region in the binary mask with an area threshold greater than 50 pixels²; that is, the connected regions are labeled, and the unique integer label of each connected region is output, and the area is screened to retain regions with an area greater than 50 pixels², to obtain an optimized binary mask, and the corresponding calculation formula is:
[0083]
[0084]
[0085]
[0086] wherein, represents the morphologically optimized binary mask; represents the closing operation; represents the dilation operation; represents the erosion operation; represents the connected branch processing; represents the optimized binary mask; represents an indicator function, which is 1 when the condition is met, and 0 otherwise.
[0087] Step 4, gray scale back
[0088] The original gray scale information of the ROI image is combined with the optimized binary mask by a gradient-guided weighted fusion formula, and an optimized image with fused original gray scale information is output.
[0089] Specifically, the optimized binary mask is normalized and weighted to obtain a fused image, and the corresponding calculation formula is:
[0090]
[0091]
[0092] wherein, represents the output of the normalized optimized binary mask, and the value range of each number is ; represents the optimized image with fused original gray scale information; represents a weight coefficient; represents a gradient amplitude.
[0093] Step 5, density peak guided K-means++ feature extraction
[0094] The optimized image, which integrates the original grayscale information, is defined as the input image. Density peak initialization calculation is performed on the input image to obtain the initial cluster centers. The data points of the input image are classified according to the initial cluster centers to obtain the cluster labels.
[0095] Furthermore, the optimized image, incorporating the original grayscale information, is defined as the input image. Density peak initialization calculations are performed on the input image to obtain initial cluster centers. Based on these initial cluster centers, the data points of the input image are categorized to obtain cluster labels. Specifically:
[0096] The mask marking defining the hippocampal region is set to 1. The grayscale values of the hippocampal region are extracted to create a sample set, denoted as . , Represents the total number of samples in the sample set, and calculates the total number of samples for any sample. The neighborhood density is calculated using the following formula:
[0097]
[0098] in, Indicates sample The neighborhood density; , Each represents any sample in the sample set; Indicates the Euclidean distance between grayscale values; This represents the cutoff distance, which is the median of the distances between samples;
[0099] Based on neighborhood density, the minimum distance between sample sets is calculated using the following formula:
[0100]
[0101] in, Indicates the minimum distance; Indicates sample The neighborhood density;
[0102] Initial cluster centers are determined by combining neighborhood density and minimum distance, generating a temporary label set, denoted as . , The total number of samples in the temporary label set is represented by the following formula:
[0103]
[0104]
[0105] in, Represents the first in the temporary tag set A temporary tag; This indicates the number of clusters at the initial cluster centers; representing samples Euclidean distance to initial cluster center; representing the first cluster of the initial cluster center;
[0106] For any cluster, calculate the average gray value of the cluster, update the initial cluster center, and the corresponding calculation formula is:
[0107]
[0108] wherein, representing the first cluster of the updated initial cluster center;
[0109] Based on the hippocampus region weighting, the corresponding calculation formula is:
[0110]
[0111]
[0112] wherein, representing the cluster 0 of the updated initial cluster center; representing the weight;
[0113] The temporary label set According to the pixel coordinates of the ROI image, it is restored to a 64x64 matrix, wherein the pixels in the binary mask are filled with label values according to the corresponding positions; the pixels outside the binary mask are assigned a mark; and a label matrix is obtained.
[0114] Further, the initial cluster center is determined in combination with the neighborhood density and the minimum distance, specifically:
[0115] The result value of multiplying the neighborhood density and the minimum distance is calculated, denoted as , and the result value is ranked in descending order according to the result value, and a result value ranking is established.
[0116] The top several samples in the result value ranking are screened as the initial cluster center.
[0117] It is explained that in the embodiment, the several is 3, that is, it respectively corresponds to the hippocampus, the edge and the background.
[0118] Further, a temporary label set is generated, specifically: according to the temporary label set, the category division is performed in turn, wherein, representing a high gray region; representing a medium gray region; representing a low gray region; It can be explained that the high gray region corresponds to the hippocampus gray matter; the medium gray region represents the edge or part of atrophy; and the low gray region corresponds to the cerebrospinal fluid or noise.
[0119] Specifically, the hippocampus center is weighted, and a higher weight is given to cluster 0, i.e., low gray pixels, and the weight increases when the gray value is less than 150; wherein the pixels outside the binary mask are marked, preferably in this embodiment, the mark is 2, i.e., the background, to obtain a label matrix, denoted as .
[0120] Multi-scale feature construction is performed on the input image and the initial clustering center to obtain a multi-dimensional feature vector;
[0121] Dynamic feature weighting is performed on the multi-dimensional feature vector to obtain a weighted feature vector, and feature mapping enhancement is performed to obtain a final enhanced feature map.
[0122] Specifically, the hippocampus details are captured from different scales, divided into local, regional and global, and a 21-dimensional feature vector is output for each pixel, and the multi-dimensional feature vector is enhanced, i.e., dynamic feature weighting is performed, feature mapping enhancement is performed to enhance the discriminative features sensitive to AD, and a final enhanced feature map is output, denoted as .
[0123] Step 6, obtain a width neural network, input the final enhanced feature map into the width neural network for final classification prediction, and realize the diagnosis of Alzheimer's disease MRI.
[0124] As an optional implementation, the width neural network refers to a fully connected network with a wider layer, or a convolutional neural network (CNN) with a wider hidden layer; it is a deep learning model, which mainly has a large number of neurons and a shallower network structure, and increases the expression ability of the model by increasing the number of neurons, rather than increasing the number of network layers, which can more effectively capture and utilize feature information to improve the performance and accuracy of the model, and to a certain extent, avoid the occurrence of overfitting phenomenon, and better handle high-dimensional data.
[0125] Specifically, the final enhanced feature map is expanded into a feature vector and input into the width neural network, and the corresponding logical expression is:
[0126]
[0127] Then, a linear transformation is performed to generate a feature node matrix , and a nonlinear transformation is performed to generate an enhanced node matrix , connecting the feature nodes and the enhanced nodes, calculating the output weight, obtaining the output matrix, analyzing the output matrix, performing final classification prediction through softmax to obtain the classification probability, and performing classification loss calculation with the label, and the corresponding calculation formula is:
[0128]
[0129]
[0130]
[0131]
[0132]
[0133]
[0134] wherein, , denotes the weight matrix; , denotes the bias term; denotes the output weight; denotes the output matrix; denotes the unit matrix; denotes the matrix processed by the softmax function; denotes the classification probability; denotes the total number of categories; denotes the label.
[0135] It can be understood that the dynamic contrast compensation binarization segmentation method is designed to eliminate the problems of fixed threshold value that cannot adapt to different scanning devices, imaging parameters, resulting in gray scale distribution difference, and sensitivity to noise through local adaptive threshold value, morphological topology optimization, and gradient guided gray scale feedback. The K-means++ driven multi-scale feature enhancement method captures the subtle abnormalities of the local texture fragmentation and edge blur of the gray matter density of early AD patients, and provides effective guarantee for the accuracy of early AD patient detection. That is, the present application combines the dynamic contrast compensation binarization segmentation method and the K-means++ driven multi-scale feature enhancement, and is committed to the hippocampus segmentation problem in image recognition of Alzheimer's disease and the problem that subtle abnormalities of early AD patients cannot be captured, thereby providing convenience for doctors.
[0136] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or can be advantageous.
[0137] The various embodiments described in this specification are presented by way of example, and each embodiment is not necessarily composed of all features described with respect to other embodiments.
Claims
1. A method for MRI diagnosis of Alzheimer's disease based on wide neural network learning, characterized in that, The method includes: Step 1: Adaptive ROI cropping of brain MRI images; Lightweight U-NET was used to process brain MRI images. Pyramid pooling was used to enhance the detection of small targets. The largest volume layer of the hippocampus in the brain cross-sectional slices was extracted, and the central region was cropped to obtain the ROI image. Step 2: Dynamic contrast compensation binarization; The threshold segmentation algorithm based on local contrast compensation calculates the mean and standard deviation of grayscale within a 5×5 window of the extracted ROI image, performs binarization decision, converts the ROI image into pure black and white, and dynamically generates a binary mask. Step 3: Morphological optimization; The watershed algorithm is used to fill the holes in the binary mask, and the area in the binary mask with an area threshold greater than 50 pixels² is retained. The optimized binary mask is then output. Step 4: Grayscale return; The original grayscale information of the ROI image is combined with the optimized binary mask using a gradient-guided weighted fusion formula, and the resulting optimized image with fused original grayscale information is output. Step 5: K-means++ feature extraction guided by density peaks; The optimized image, which integrates the original grayscale information, is defined as the input image. Density peak initialization calculation is performed on the input image to obtain the initial cluster centers. The data points of the input image are classified according to the initial cluster centers to obtain the cluster labels. Multi-scale feature construction is performed on the input image and the initial cluster centers to obtain a multi-dimensional feature vector; Dynamic feature weighting is applied to the multidimensional feature vector to obtain a weighted feature vector, and feature mapping enhancement is performed to obtain the final enhanced feature map; Step 6: Obtain the width neural network, input the final enhanced feature map into the width neural network for final classification prediction, and realize the diagnosis of Alzheimer's disease by MRI.
2. The MRI diagnostic method for Alzheimer's disease based on a wide neural network according to claim 1, characterized in that, In step 5, density peak initialization calculation is performed on the input image to obtain initial cluster centers. Based on the initial cluster centers, the data points of the input image are classified into categories to obtain cluster labels, specifically: The mask marking defining the hippocampal region is set to 1. The grayscale values of the hippocampal region are extracted to create a sample set, denoted as . , Represents the total number of samples in the sample set, and calculates the total number of samples for any sample. The neighborhood density is calculated using the following formula: ; in, Indicates sample The neighborhood density; , Each represents any sample in the sample set; Indicates the Euclidean distance between grayscale values; This represents the cutoff distance, which is the median of the distances between samples; Based on neighborhood density, the minimum distance between sample sets is calculated using the following formula: ; in, Indicates the minimum distance; Indicates sample The neighborhood density; Initial cluster centers are determined by combining neighborhood density and minimum distance, generating a temporary label set, denoted as . , The total number of samples in the temporary label set is represented by the following formula: ; ; in, Represents the first in the temporary tag set A temporary tag; This indicates the number of clusters at the initial cluster centers; Indicates sample Euclidean distance to the initial cluster center; The first cluster center represents the initial cluster center. A cluster; For any cluster, calculate the average gray value of the cluster and update the initial cluster centers. The corresponding calculation formula is: ; in, Represents the first cluster center after the update. A cluster; Based on hippocampal region weighting, the corresponding calculation formula is: ; ; in, Cluster 0 represents the updated initial cluster centers; Indicates weight; temporary tag set The ROI image is reconstructed into a 64×64 matrix based on pixel coordinates. Pixels within the binary mask are filled with label values according to their corresponding positions, while pixels outside the binary mask are assigned labels, resulting in a label matrix.
3. The MRI diagnostic method for Alzheimer's disease based on a wide neural network according to claim 2, characterized in that, The initial cluster centers are determined by combining neighborhood density and minimum distance, specifically as follows: The result of multiplying the neighborhood density and the minimum distance is denoted as . Sort the results in descending order to create a ranking of the results; The top-ranked samples in the screening results are used as the initial cluster centers.
4. The MRI diagnostic method for Alzheimer's disease based on a wide neural network according to claim 2, characterized in that, Generate a temporary tag set, specifically by classifying the temporary tag set into categories sequentially, where... Indicates a high grayscale area; Indicates a medium grayscale area; This indicates a low-grayscale area.