Medical image analysis method based on multi-modal craniocerebral MR structural feature recognition

Through multimodal cranial MR image analysis, combined with multiple MRI sequences and feature fusion algorithms, the limitations of single-modality image analysis are overcome, high-precision feature extraction and classification of complex brain structures are achieved, and diagnostic accuracy is improved.

CN120707869APending Publication Date: 2025-09-26INST OF NEUROLOGY AFFILIATED HOSPITAL OF ANHUI MEDICAL UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510941808.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

When dealing with complex cranial structures, existing technologies cannot fully reflect the characteristics of lesions with single-modality imaging features, and lack an effective feature fusion mechanism, resulting in insufficient diagnostic accuracy and limited analytical capabilities.

Method used

A multimodal brain MR structural feature recognition method was used to acquire T1-weighted imaging, T2-weighted imaging, fluid-attenuated inversion recovery sequence, and diffusion-weighted imaging data. The data were decomposed into different resolution levels using an image pyramid algorithm. Features were extracted using a combination of Gaussian filtering and convolutional neural networks, and feature fusion and classification were performed using LASSO regression and recursive feature elimination using random forests.

Benefits of technology

It achieves comprehensive feature extraction and accurate classification of complex brain structures, improves the ability to distinguish image areas and diagnostic accuracy, and provides a reliable analysis tool for clinical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707869A_ABST
    Figure CN120707869A_ABST
Patent Text Reader

Abstract

The invention discloses a medical image analysis method based on multi-modal craniocerebral MR structural feature recognition, and belongs to the technical field of medical image analysis, and the method comprises the steps: obtaining multi-modal craniocerebral MRI image data, and taking the multi-modal craniocerebral MRI image data as an original image; the method comprises the following steps: preprocessing an original image to obtain image hierarchies with different resolutions; constructing a feature extraction module, and performing feature extraction on the image hierarchies with different resolutions to obtain feature maps extracted under different scales; constructing a feature fusion module, and fusing the feature maps extracted under different scales in a channel splicing and element-by-element adding mode to obtain fused features; and inputting the fused features into a classifier constructed by screening key features based on LASSO regression and a recursive feature elimination method of a random forest to obtain a classification result of the image region. According to the method, the analysis precision of brain complex structure image features can be improved, and the distinguishing capability of similar image areas is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image analysis, and in particular relates to a medical image analysis method based on multimodal brain MR structural feature recognition. Background Art

[0002] In the field of medical image analysis, traditional image analysis methods primarily rely on the extraction and analysis of image features from a single modality. For example, they process only a single sequence of magnetic resonance imaging (MRI) data, such as T1-weighted imaging (T1WI) or T2-weighted imaging (T2WI). These methods extract key features from images through specific image processing algorithms, such as edge detection and texture analysis, and perform tasks such as lesion detection and tissue classification based on these features. In addition, some advanced technologies incorporate machine learning algorithms, such as support vector machines (SVM) or random forests, to improve diagnostic accuracy. These methods can provide clinicians with useful diagnostic information to a certain extent, but their analysis scope and depth are limited by the use of single-modality data.

[0003] However, existing technologies have obvious limitations when dealing with complex cranial structures. Brain MRI images contain a variety of tissue structures and complex pathological information, and the image features of a single modality are often difficult to fully reflect the characteristics of the lesions. For example, some lesions may appear as slight signal changes on T1WI, but have more obvious abnormal signals on T2WI or FLAIR sequences. In addition, existing methods lack an effective feature fusion mechanism when processing multi-dimensional images (such as two-dimensional slices and three-dimensional volume data), resulting in insufficient recognition of complex structural features and an inability to effectively distinguish similar image areas. This not only affects the accuracy of diagnosis, but also limits the further development of medical image analysis technology in clinical applications. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention proposes a medical image analysis method based on multimodal cranial MR structural feature recognition to solve the problems existing in the above-mentioned prior art.

[0005] To achieve the above objectives, in a first aspect, the present invention provides a medical image analysis method based on multimodal brain MR structural feature recognition, comprising:

[0006] Acquire multimodal cranial MRI image data including T1-weighted imaging, T2-weighted imaging, fluid-attenuated inversion recovery sequence, and diffusion-weighted imaging as original images;

[0007] Preprocessing the original image to obtain image layers of different resolutions;

[0008] Construct a feature extraction module to extract features from image layers of different resolutions and obtain feature maps extracted at different scales;

[0009] Construct a feature fusion module to fuse the feature maps extracted at different scales through channel splicing and element-by-element addition to obtain the fused features;

[0010] The fused features are input into a classifier constructed by screening key features based on LASSO regression and random forest recursive feature elimination method to obtain the classification results of the image area.

[0011] Preferably, the step of preprocessing the original image includes:

[0012] Using the image pyramid algorithm, the original image is decomposed into N image levels with different resolutions. The original image is used as the base layer, and the subsequent levels are downsampled in sequence to build an image pyramid from high resolution to low resolution.

[0013] The image data of each layer is normalized to obtain normalized image data.

[0014] Preferably, the steps of constructing a feature extraction module and extracting features from image layers of different resolutions include:

[0015] Before inputting image data of different layers into the convolutional neural network model, Gaussian filtering is applied to the images of each layer. For layers with high resolution and more noise, a larger Gaussian kernel is selected;

[0016] For large-scale image levels, a shallow convolutional neural network with a larger convolution kernel is used to extract macro features;

[0017] For small-scale image levels, a deep convolutional neural network with small convolution kernels is used to extract microscopic features.

[0018] Preferably, the step of constructing the feature extraction module further includes:

[0019] According to the dimension of the image data, the input and convolution operations of the convolutional neural network model are adjusted. For two-dimensional slice data, two-dimensional convolution is used; for three-dimensional volume data, three-dimensional convolution is used.

[0020] Preferably, the step of constructing a feature fusion module to fuse feature maps extracted at different scales by channel splicing and element-by-element addition includes:

[0021] The feature maps extracted at different scales are spliced ​​according to the channel dimension;

[0022] For the feature maps after channel splicing, perform element-by-element addition operation at the same spatial position.

[0023] Preferably, the fused features are input into a classifier constructed by screening key features based on LASSO regression and random forest recursive feature elimination method, and the step of obtaining the classification result of the image area includes:

[0024] Perform dimension conversion on the fused feature map to turn it into a one-dimensional vector;

[0025] Inputting the one-dimensional vector into a fully connected neural network comprising multiple hidden layers, each hidden layer being followed by a rectified linear unit activation function;

[0026] The number of nodes in the output layer is determined according to the specific analysis task, and the Softmax activation function is used in the output layer to output the probability value of each category, and the classification result of the image area is obtained through the probability value.

[0027] Preferably, the classifier is a decision tree classifier.

[0028] In a second aspect, the present invention further discloses a computer device comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method described in the first aspect.

[0029] In a third aspect, the present invention further discloses a computer-readable storage medium having a computer program stored thereon, which implements the steps of the method described in the first aspect when executed by a processor.

[0030] In a fourth aspect, the present invention further discloses a computer program product, comprising a computer program, which implements the steps of the method described in the first aspect when executed by a processor.

[0031] Compared with the prior art, the present invention has the following advantages and technical effects:

[0032] The present invention provides a medical image analysis method based on multimodal cranial MR structural feature recognition, comprising: first, obtaining multimodal cranial MRI image data including T1-weighted imaging, T2-weighted imaging, fluid-attenuated inversion recovery sequence, and diffusion-weighted imaging as original images; second, preprocessing the original images to obtain image levels of different resolutions; then, constructing a feature extraction module to extract features from image levels of different resolutions respectively to obtain feature maps extracted at different scales; further, constructing a feature fusion module to fuse the feature maps extracted at different scales by channel splicing and element-by-element addition to obtain fused features; finally, inputting the fused features into a classifier constructed by screening key features based on LASSO regression and random forest recursive feature elimination method to obtain classification results of the image area.

[0033] The present invention automatically segments areas of interest and extracts brain structure texture features by analyzing conventional brain magnetic resonance imaging sequences (T1-weighted imaging, T2-weighted imaging, fluid-attenuated inversion recovery sequence, and diffusion-weighted imaging) under multimodality. Multi-dimensional image feature data is fused through a Gaussian algorithm and a convolutional neural network (CNN). Key features are screened using the recursive feature elimination method combining LASSO regression and random forests to construct a classifier model, thereby improving the ability to recognize texture features of complex structural images.

[0034] Through feature fusion analysis of multimodal brain MR, the present invention can extract medical image texture features more comprehensively and accurately, improve the analysis accuracy of complex brain structure image features, and effectively enhance the ability to distinguish similar image areas. Compared with traditional single-dimensional analysis methods, it improves the accuracy in image classification and recognition tasks, and provides clinicians with a reliable structural image analysis tool. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:

[0036] Figure 1 Flowchart of a method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0037] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0038] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0039] Example 1

[0040] like Figure 1 As shown, this embodiment provides a medical image analysis method based on multimodal cranial MR structural feature recognition, including:

[0041] S1. Acquire multimodal cranial MRI image data including T1-weighted imaging, T2-weighted imaging, fluid-attenuated inversion recovery sequence, and diffusion-weighted imaging as original images;

[0042] Specifically, this embodiment collects conventional brain magnetic resonance sequences (T1WI, T2WI, FLAIR, DWI) under multimodality, ensures the integrity and accuracy of the patient's MRI T1WI sequence and clinical data, and imports them into the system in a standard medical imaging format (such as DICOM).

[0043] S2. Preprocessing the original image to obtain image layers of different resolutions;

[0044] Furthermore, the step of preprocessing the original image includes:

[0045] S201, using an image pyramid algorithm, decomposing the original image into N image layers of different resolutions, taking the original image as the base layer, and sequentially downsampling the subsequent layers to construct an image pyramid from high resolution to low resolution;

[0046] Specifically, the image pyramid algorithm is used to decompose the original image into N image layers of different resolutions. Taking the original image as the base layer, the subsequent layers are sequentially downsampled by a factor of 2 to construct an image pyramid from high resolution (e.g., 512×512 pixels) to low resolution (e.g., 16×16 pixels). This fully captures multi-dimensional image texture data from macro to micro, providing rich information for subsequent feature extraction.

[0047] S202 : performing normalization processing on the image data of each layer to obtain normalized image data.

[0048] Specifically, normalization is performed on the image data at each level, and the pixel value range is uniformly mapped to the [0,1] interval to eliminate image differences caused by different devices and scanning parameters, and improve the stability of subsequent processing. The specific formula used is: norm =xx min / x max -x min , where x is the original pixel value, x min and x max are the minimum and maximum pixel values ​​in the image data, respectively, norm is the normalized pixel value.

[0049] S3. Construct a feature extraction module to extract features from image layers of different resolutions and obtain feature maps extracted at different scales.

[0050] Furthermore, a feature extraction module is constructed to extract features from image layers of different resolutions. The steps include:

[0051] S301, before inputting image data of different layers into the convolutional neural network model, applying Gaussian filtering to the image of each layer, and selecting a larger Gaussian kernel for layers with high resolution and more noise;

[0052] Specifically, before inputting image data from different layers into the CNN module, a Gaussian filter is applied to each layer. For layers with high resolution and high noise, a larger Gaussian kernel (e.g., 5×5) is used to effectively suppress noise interference and highlight the key structural features of the image, laying the foundation for the CNN module to accurately extract features.

[0053] S302. For large-scale image levels, a shallow convolutional neural network with a larger convolution kernel is used to extract macro features.

[0054] Specifically, a shallow CNN network was constructed for large-scale (low-resolution) image layers. The network consists of three convolutional layers with a kernel size of 7×7, a stride of 2, and a padding of 3 to rapidly extract macroscopic features such as the overall brain structure. Each convolutional layer is followed by a batch normalization layer and a rectified linear unit (ReLU) activation function to accelerate network training and alleviate the vanishing gradient problem.

[0055] S303. For small-scale image levels, a deep convolutional neural network with small convolution kernels is used to extract microscopic features.

[0056] Specifically, for the small-scale (high-resolution) image level, a deep CNN network was designed. The network consists of eight convolutional layers with a 3×3 kernel size, a stride of 1, and a padding of 1, capturing subtle structural features such as sulci and gyri. Similarly, a batch normalization layer and a rectified linear unit (ReLU) activation function are added after each convolutional layer, and residual connections are introduced between some convolutional layers to avoid training degradation caused by excessive network depth.

[0057] Furthermore, the step of constructing the feature extraction module further includes:

[0058] According to the dimension of the image data, the input and convolution operations of the convolutional neural network model are adjusted. For two-dimensional slice data, two-dimensional convolution is used; for three-dimensional volume data, three-dimensional convolution is used.

[0059] Specifically, this embodiment provides adaptation to different dimensions. Based on the dimensions of the image data (e.g., 2D slices or 3D volume data), the CNN module's input and convolution operations are flexibly adjusted. For 2D slice data, 2D convolution is used; for 3D volume data, 3D convolution is used. This allows the model to better adapt to the extraction requirements of image features in different dimensions and fully exploit the effective information in the data.

[0060] S4. Construct a feature fusion module to fuse the feature maps extracted at different scales through channel splicing and element-by-element addition to obtain the fused features;

[0061] Furthermore, a feature fusion module is constructed to fuse the feature maps extracted at different scales by channel concatenation and element-by-element addition, including the following steps:

[0062] S401, splicing the feature maps extracted at different scales according to the channel dimension;

[0063] Specifically, the feature maps extracted at different scales are spliced ​​together according to the channel dimension. If the number of channels in the small-scale feature map is 64 and the number of channels in the large-scale feature map is 32, the resulting spliced ​​feature map has 96 channels, so that the fused features contain rich information at different scales. Based on the dimension of the image data (such as two-dimensional slices or three-dimensional volume data), the input and convolution operations of the CNN module are flexibly adjusted. For two-dimensional slice data, two-dimensional convolution is used; for three-dimensional volume data, three-dimensional convolution is used. This allows the model to better adapt to the extraction requirements of image features at different dimensions and fully tap into the effective information in the data.

[0064] S402: Perform element-by-element addition on the feature maps after channel splicing at the same spatial position.

[0065] Specifically, the feature maps after channel splicing are element-by-element added at the same spatial location to further fuse macro and micro feature information and enhance the expressiveness of the features. To avoid excessively large eigenvalues ​​after addition, the addition result can be appropriately scaled, such as by multiplying it by a scaling factor (e.g., 0.5), to ensure the effectiveness and stability of the feature fusion effect.

[0066] S5. The fused features are input into a classifier constructed by screening key features based on LASSO regression and random forest recursive feature elimination method to obtain the classification results of the image area.

[0067] Furthermore, the fused features are input into a classifier constructed by screening key features based on LASSO regression and random forest recursive feature elimination method. The steps of obtaining the classification results of the image area include:

[0068] S501, performing dimension conversion processing on the fused feature map to convert it into a one-dimensional vector;

[0069] S502, inputting the one-dimensional vector into a fully connected neural network comprising multiple hidden layers, each hidden layer being connected to a rectified linear unit activation function;

[0070] S503 : Determine the number of nodes in the output layer according to the specific analysis task, and use the Softmax activation function in the output layer to output the probability value of each category, and obtain the classification result of the image area through the probability value.

[0071] Specifically, the fused feature map is flattened into a one-dimensional vector and input into a fully connected neural network. This fully connected neural network consists of three hidden layers, with 256, 128, and 64 nodes, respectively. Each hidden layer is followed by a ReLU activation function. The number of nodes in the output layer depends on the specific analysis task. For example, for binary classification of image regions (normal / abnormal), the number of output layer nodes is 2. For multi-classification (such as classification of different lesion types), the number of output layer nodes is the corresponding number of categories, and a Softmax activation function is used to output the probability value of each category.

[0072] Furthermore, the classifier is a decision tree classifier.

[0073] In addition, this embodiment also includes:

[0074] The model was trained using a dataset of labeled brain MRI images, using a cross-entropy loss function to measure the difference between the predicted results and the true labels. The Adam optimizer was selected, with a learning rate of 0.001, β1 = 0.9, and β2 = 0.999. The model parameters were updated using a backpropagation algorithm to continuously reduce the loss function value and improve the model's classification accuracy.

[0075] After training is completed, the multimodal cranial MRI image data to be analyzed is input into the model according to the above process. The model outputs the corresponding classification results or other analysis results, such as the location of the lesion area, the assessment of the degree of lesion, etc., to realize medical image analysis based on multimodal cranial MRI structural feature recognition.

[0076] Use the test dataset to evaluate the trained model, calculate indicators such as accuracy, recall, F1 value, AUC (area under the curve), and comprehensively measure the performance of the model in medical image analysis tasks.

[0077] Based on the evaluation results, optimize and adjust the model. If the model is found to be overfitting, you can add data augmentation strategies (such as rotation, translation, flipping, etc.) or add a dropout layer to the fully connected neural network. If the model accuracy is low, you can adjust the structure and hyperparameters of the CNN module (such as the number of convolution kernels and network depth), or try to replace it with a more advanced classifier to further improve model performance.

[0078] Beneficial effects of this embodiment:

[0079] This embodiment innovatively combines multidimensional decomposition with CNN feature extraction modules of different structures, achieving an analysis method that moves from the effective extraction of multidimensional features to the re-integration of unique complex structural features, combining the advantages of features from different dimensions and improving the expressiveness of features. This embodiment is based on conventional MRI data and does not require additional invasive examinations. In addition, the combination of radiomics features and clinical data improves the accuracy of predictions.

[0080] Example 2

[0081] This embodiment further discloses a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method described in the first embodiment.

[0082] Example 3

[0083] This embodiment further discloses a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in the first embodiment are implemented.

[0084] Example 4

[0085] This embodiment further discloses a computer program product, including a computer program, which implements the steps of the method described in the first embodiment when executed by a processor.

[0086] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A medical image analysis method based on multimodal brain MR structural feature recognition, characterized in that: The following steps are involved: Acquire multimodal cranial MRI image data including T1-weighted imaging, T2-weighted imaging, fluid-attenuated inversion recovery sequence, and diffusion-weighted imaging as original images; Preprocessing the original image to obtain image layers of different resolutions; Construct a feature extraction module to extract features from image layers of different resolutions and obtain feature maps extracted at different scales; Construct a feature fusion module to fuse the feature maps extracted at different scales through channel splicing and element-by-element addition to obtain the fused features; The fused features are input into a classifier constructed by screening key features based on LASSO regression and random forest recursive feature elimination method to obtain the classification results of the image area.

2. The method according to claim 1, characterized in that The step of preprocessing the original image includes: Using the image pyramid algorithm, the original image is decomposed into N image levels with different resolutions. The original image is used as the base layer, and the subsequent levels are downsampled in sequence to build an image pyramid from high resolution to low resolution. The image data of each layer is normalized to obtain normalized image data.

3. The method according to claim 1, characterized in that The steps of constructing a feature extraction module and extracting features from image layers of different resolutions include: Before inputting image data of different layers into the convolutional neural network model, Gaussian filtering is applied to the images of each layer. For layers with high resolution and more noise, a larger Gaussian kernel is selected; For large-scale image levels, a shallow convolutional neural network with a larger convolution kernel is used to extract macro features; For small-scale image levels, a deep convolutional neural network with small convolution kernels is used to extract microscopic features.

4. The method according to claim 1, wherein The steps of constructing the feature extraction module also include: According to the dimension of the image data, the input and convolution operations of the convolutional neural network model are adjusted. For two-dimensional slice data, two-dimensional convolution is used; for three-dimensional volume data, three-dimensional convolution is used.

5. The method according to claim 1, wherein The steps of constructing a feature fusion module and fusing feature maps extracted at different scales through channel splicing and element-by-element addition include: The feature maps extracted at different scales are spliced ​​according to the channel dimension; For the feature maps after channel splicing, perform element-by-element addition operation at the same spatial position.

6. The method according to claim 1, wherein The fused features are input into a classifier constructed by filtering key features using LASSO regression and random forest recursive feature elimination. The steps to obtain the classification results of the image area include: Perform dimension conversion on the fused feature map to turn it into a one-dimensional vector; Inputting the one-dimensional vector into a fully connected neural network comprising multiple hidden layers, each hidden layer being followed by a rectified linear unit activation function; The number of nodes in the output layer is determined according to the specific analysis task, and the Softmax activation function is used in the output layer to output the probability value of each category, and the classification result of the image area is obtained through the probability value.

7. The method according to claim 1, characterized in that The classifier is a decision tree classifier.

8. A computer device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • MRI multi-modal image segmentation method and system based on ACU-Net

    CN112215844A

  • Processing method of brain magnetic resonance diffusion weighted image and related product

    CN113538348A

  • Brain age prediction method, system and device based on multi-modal information fusion and medium

    CN119360124A

  • Registration method for multimodality image-guided radiotherapy

    WO2024169341A1