Brain tumor image recognition system and method

By designing a brain tumor image recognition system, using multimodal imaging and three-dimensional deep learning networks, the limitations of single mode and two-dimensional features in the existing technology are solved, and high-precision brain tumor recognition and classification are achieved.

CN120147743APending Publication Date: 2025-06-13THE SECOND HOSPITAL OF HEBEI MEDICAL UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510294518.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art has problems such as single mode limitations, poor adaptability of complex scenarios and insufficient utilization of three-dimensional features in brain tumor recognition, resulting in low recognition efficiency and insufficient accuracy.

Method used

A brain tumor image recognition system was designed to achieve high-precision tumor detection and classification by collecting multimodal images, preprocessing, multimodal feature extraction and fusion, brain tumor segmentation and classification, and interpretability results generation. Adaptive feature fusion, three-dimensional deep learning network and visualization technology are used to achieve high-precision tumor detection and classification.

Benefits of technology

Through multimodal feature fusion and three-dimensional deep learning network, the high-precision recognition and classification capabilities of brain tumors are improved, the limitations of single mode and two-dimensional features are overcome, and more robust tumor detection is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147743A_ABST
    Figure CN120147743A_ABST
Patent Text Reader

Abstract

The invention provides a brain tumor image recognition system and method, and the method comprises the steps: collecting a brain tumor multi-modal image, carrying out the preprocessing operation of the collected multi-modal image, carrying out the multi-modal feature extraction and fusion of the multi-modal image after the preprocessing operation, and carrying out the segmentation and classification of a brain tumor based on the fused multi-modal features. And generating an interpretability result based on the segmentation and classification results. Through adaptive feature fusion, a three-dimensional deep learning network and a visualization technology, the limitation of the prior art is solved, and high-precision and robust tumor detection and classification are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of brain tumor recognition, and particularly to a brain tumor image recognition system and method. Background Art

[0002] Early diagnosis of brain tumors is crucial for clinical treatment. At present, traditional methods mainly rely on doctors' manual interpretation of images such as MRI and CT, which have problems such as low efficiency, strong subjectivity, and easy missed diagnosis of small tumors. In the prior art, image recognition algorithms based on deep learning have been applied to brain tumor detection, but there are still the following deficiencies:

[0003] 1. Single-modal limitation: Most methods are only based on a single modality (such as T1-weighted MRI) and cannot make full use of the complementary information of multi-modal images;

[0004] 2. Poor adaptability to complex scenarios: The recognition accuracy of tumors with low contrast, small tumors, or tumors with blurred boundaries is insufficient;

[0005] 3. Insufficient utilization of three-dimensional features: Most algorithms only process two-dimensional slices and ignore the spatial structure information of tumors.

[0006] Therefore, it is very necessary to design a brain tumor image recognition system and method. Summary of the Invention

[0007] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a brain tumor image recognition system and method.

[0008] To achieve the above purpose, the present invention provides the following solutions:

[0009] The present invention provides a brain tumor image recognition system, including:

[0010] A brain tumor multi-modal image acquisition module, configured to acquire brain tumor multi-modal images;

[0011] A preprocessing module, configured to perform preprocessing operations on the acquired multi-modal images;

[0012] Multi-modal feature extraction and fusion, configured to perform multi-modal feature extraction and fusion on the preprocessed multi-modal images;

[0013] Brain tumor segmentation and classification, configured to perform brain tumor segmentation and classification based on the fused multi-modal features;

[0014] An interpretability result generation module, configured to generate interpretability results based on the segmentation and classification results.

[0015] The present invention also provides a brain tumor image recognition method, including:

[0016] Step 100: Collect multi-modal brain tumor images;

[0017] Step 200: Perform preprocessing operations on the collected multi-modal images;

[0018] Step 300: Extract and fuse multi-modal features from the preprocessed multi-modal images;

[0019] Step 400: Segment and classify brain tumors based on the fused multi-modal features;

[0020] Step 500: Generate interpretable results based on the segmentation and classification results.

[0021] Preferably, in Step 200, the preprocessing operations on the collected multi-modal images are specifically as follows:

[0022] Step 201: Perform registration processing on the collected multi-modal images;

[0023] Step 202: Perform denoising processing on the registered multi-modal images;

[0024] Step 203: Perform standardization and normalization on the denoised multi-modal images;

[0025] Step 204: Perform three-dimensional reconstruction on the standardized and normalized multi-modal images to obtain three-dimensional volume data.

[0026] Preferably, in Step 201, the registration processing on the collected multi-modal images is specifically as follows:

[0027] Process the reference image and the image to be registered in the multi-modal images into grayscale images, initially extract accurately positioned feature points and corresponding descriptors through a feature extraction network, and use the quadtree algorithm to screen the initially extracted feature points to achieve uniform distribution;

[0028] Input the uniformly screened feature points and corresponding feature descriptors into a feature matching network, and perform feature matching and outlier rejection according to a given matching threshold to obtain homologous feature points;

[0029] After obtaining a large number of homologous feature points, use a piecewise linear model to iteratively estimate the transformation model parameters between the reference image and the image to be registered, and finally achieve the registration processing of the multi-modal images.

[0030] Preferably, in Step 202, the denoising processing on the registered multi-modal images is specifically as follows:

[0031] Construct a partial differential equation denoising model based on a hybrid YK&PM model, and perform denoising processing on the registered multi-modal images based on the denoising model.

[0032] Preferably, in step 300, multi-modal feature extraction and fusion are performed on the pre-processed multi-modal images, specifically as follows:

[0033] Step 301: Construct a feature extraction network structure;

[0034] Step 302: Input the three-dimensional volume data into the feature extraction network to obtain high-dimensional feature maps of each modality;

[0035] Step 303: Based on the channel attention mechanism, fuse the high-dimensional feature maps to obtain a fused feature map.

[0036] Preferably, the feature extraction network structure uses a lightweight 3D ResNet-18 network structure as the backbone network, and dilated convolutional layers are embedded in its residual blocks to expand the receptive field.

[0037] Preferably, in step 400, brain tumor segmentation and classification are performed based on the fused multi-modal features, specifically as follows:

[0038] Step 401: Construct a brain tumor segmentation model based on the improved U-Net network and a brain tumor classification model based on the improved SE-Net network;

[0039] Step 402: Train the brain tumor segmentation model and the brain tumor classification model based on a preset data set;

[0040] Step 403: Input the fused feature map into the trained brain tumor segmentation model to obtain the brain tumor segmentation result;

[0041] Step 404: Input the fused feature map and the brain tumor segmentation result into the trained brain tumor classification model to obtain the brain tumor classification result.

[0042] Preferably, the U-Net network is improved as follows:

[0043] Replace the convolutional blocks in the traditional U-Net structure with deep residual blocks, add the CBAM attention mechanism to the skip connections in the traditional U-Net structure, and replace the ReLU activation function in the deep residual blocks with the Dy-ReLU activation function.

[0044] Preferably, the SE-Net network is improved as follows:

[0045] After the BN in the traditional SE-Net network, replace the ReLU activation function with the Swish activation function, add an ECA attention module after the first convolution after the global max pooling in the traditional SE-Net network, add the BAM attention mechanism after the second convolutional layer, and finally, use the improved SE attention mechanism to obtain more channel information.

[0046] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:

[0047] The present invention provides a brain tumor image recognition system and method. The system includes a brain tumor multi-modal image acquisition module, a preprocessing module, multi-modal feature extraction and fusion, brain tumor segmentation and classification, and an interpretable result generation module. The method includes collecting brain tumor multi-modal images, performing preprocessing operations on the collected multi-modal images, performing multi-modal feature extraction and fusion on the preprocessed multi-modal images, performing brain tumor segmentation and classification based on the fused multi-modal features, and generating interpretable results based on the segmentation and classification results. The present invention solves the limitations of the prior art through adaptive feature fusion, three-dimensional deep learning network and visualization technology, and realizes high-precision and robust tumor detection and classification. Description of the Drawings

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0049] Figure 1 It is a flowchart of the method provided by the embodiment of the present invention;

[0050] Figure 2 It is a schematic diagram of the feature extraction network framework;

[0051] Figure 3 It is a schematic diagram of the feature matching network framework;

[0052] Figure 4 It is a schematic diagram of the brain tumor segmentation model structure of the improved U-Net network;

[0053] Figure 5 It is a schematic diagram of Dy-ReLU;

[0054] Figure 6 It is a schematic diagram of the improved residual block structure;

[0055] Figure 7 It is a schematic diagram of the brain tumor classification model structure;

[0056] Figure 8 It is a schematic diagram of the improved SE attention module structure. Detailed Embodiments

[0057] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0058] The object of the present invention is to provide a brain tumor image recognition system and method, which solve the limitations of the prior art through adaptive feature fusion, three-dimensional deep learning network and visualization technology, and realize high-precision and robust tumor detection and classification.

[0059] To make the above objects, features and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0060] The present invention provides a brain tumor image recognition system, including:

[0061] A brain tumor multimodal image acquisition module for acquiring brain tumor multimodal images;

[0062] A preprocessing module for performing preprocessing operations on the acquired multimodal images;

[0063] Multimodal feature extraction and fusion for extracting and fusing multimodal features from the preprocessed multimodal images;

[0064] Brain tumor segmentation and classification for segmenting and classifying brain tumors based on the fused multimodal features;

[0065] An interpretability result generation module for generating interpretability results based on the segmentation and classification results.

[0066] As Figure 1 shown, the present invention also provides a brain tumor image recognition method, including:

[0067] Step 100: Acquire brain tumor multimodal images;

[0068] Step 200: Perform preprocessing operations on the acquired multimodal images;

[0069] Step 300: Extract and fuse multimodal features from the preprocessed multimodal images;

[0070] Step 400: Segment and classify brain tumors based on the fused multimodal features;

[0071] Step 500: Generate interpretability results based on the segmentation and classification results.

[0072] In step 100, multi-modal brain tumor images are acquired, specifically:

[0073] The acquisition of multi-modal images needs to combine anatomical imaging (T1 / T2 MRI, CT), functional imaging (DWI, PWI), and molecular imaging (PET, MRS) to comprehensively characterize the biological behavior of tumors through complementary information. In practical applications, the modal combination needs to be selected according to the clinical scenario (such as emergency, preoperative planning, and efficacy evaluation), and data registration and standardization need to be ensured to provide high-quality input for subsequent AI algorithms;

[0074] An introduction to common multi-modal images is as follows:

[0075] 1. MRI (Magnetic Resonance Imaging) sequences

[0076] MRI is the core modality for brain tumor diagnosis. Different sequences highlight different characteristics of tumors by adjusting magnetic field and radiofrequency pulse parameters, specifically including:

[0077] T1-weighted Imaging (T1WI), T2-weighted Imaging (T2WI), FLAIR (Fluid Attenuated Inversion Recovery), DWI (Diffusion Weighted Imaging) and ADC map, PWI (Perfusion Weighted Imaging), MRS (Magnetic Resonance Spectroscopy);

[0078] 2. CT (Computed Tomography)

[0079] Specifically include: plain CT, enhanced CT;

[0080] 3. PET (Positron Emission Tomography)

[0081] Specifically include: 1 8 F-FDG PET, amino acid-based PET (such as 11 C-MET, 1 8 F-FET);

[0082] Regarding its specific acquisition method, since it is prior art, it will not be described in detail.

[0083] In step 200, preprocessing operations are performed on the acquired multi-modal images, specifically:

[0084] Step 201: Perform registration processing on the acquired multi-modal images;

[0085] Step 202: Denoise the registered multi-modal images;

[0086] Step 203: Standardize and normalize the denoised multi-modal images;

[0087] Step 204: Perform 3D reconstruction based on the standardized and normalized multi-modal images to obtain 3D volume data.

[0088] In step 201, the acquired multi-modal images are registered. Specifically:

[0089] The present invention uses a convolutional neural network to implement the registration process. Among them, the principle of the convolutional neural network will not be introduced in detail here. The specific steps for the registration process of the present invention are as follows:

[0090] Process the reference image and the image to be registered in the multi-modal images into grayscale images, initially extract accurately located feature points and corresponding descriptors through a feature extraction network, and use a quadtree algorithm to screen the initially extracted feature points to achieve uniform distribution;

[0091] Input the uniformly screened feature points and corresponding feature descriptors into a feature matching network, and perform feature matching and outlier rejection according to a given matching threshold to obtain homologous feature points;

[0092] After obtaining a large number of homologous feature points, use a piecewise linear model to iteratively estimate the transformation model parameters between the reference image and the image to be registered, and finally achieve the registration process of the multi-modal images;

[0093] The above steps are described in detail as follows:

[0094] 1. Feature extraction

[0095] The high repeatability and uniform distribution of feature points are essential for image matching, jointly determining the accuracy of subsequent image registration. Existing feature extraction methods can usually obtain a large number of feature points on optical remote sensing image pairs, but these feature points often gather in a certain local area of the image, and their repeatability is also low. Therefore, these feature extraction methods may be ineffective for multi-modal remote sensing images with non-linear radiation and scale differences;

[0096] SuperPoint is a fully convolutional neural network framework for self-supervised multi-scale feature extraction, which can simultaneously obtain key points and corresponding descriptors. The network consists of three parts: a shared encoding layer, a feature point detection layer, and a descriptor decoding layer. However, it has a large number of network layers and a high computational complexity in model training, which not only affects the computational efficiency but also generates a large amount of redundant information. In view of this, the present invention uses a lightweight GhostNet structure to replace the VGG structure in the shared encoding layer of the SuperPoint network to reduce the computational amount, and adds pyramid convolution to the original feature point detection layer and descriptor decoding layer to obtain multi-scale feature maps of the image. The feature extraction network framework proposed by the present invention is as Figure 2 shown;

[0097] GhostNet is a residual structure with few model parameters and rich feature encoding information. It reduces the computational amount by replacing some convolutions with a series of linear operations. For remote sensing images with a large amount of data and complex texture information, this network structure helps to improve the computational efficiency of feature extraction. Therefore, the present invention uses the first 7 layers of the GhostNet structure as the shared encoding layer of the feature extraction network, and at the same time retains the stacked shadow bottleneck layer (Ghostbottleneck, G-bneck) in the 2nd - 7th layers of the GhostNet structure. G-bneck is a residual block structure that increases the receptive field range by first expanding and then compressing the number of channels, so as to aggregate more detailed features;

[0098] In order to adapt to the scale differences between multi-modal remote sensing images and robustly extract a large number of multi-scale features, the present invention uses pyramid convolution composed of different scale convolution kernels to replace the single scale convolution kernel in the above GhostNet structure, that is, in the first convolution layer of the GhostNet structure, 3×3, 5×5, 7×7, 9×9 convolution kernels of different sizes are used to replace the original 3×3 convolution kernel, and the total number of its channels remains unchanged to generate four groups of GhostNet structures with different scales. Then, spatial downsampling operations are performed in each group of GhostNet structures, and the shared feature maps at this scale are generated after six consecutive G-bneck operations;

[0099] For the obtained shared feature maps with different scales The feature point detection layer first uses a 1×1 convolution kernel to combine multi-scale feature information to generate a feature map Then, two convolutional operations are performed on the feature map F to make the size of the feature map become Finally, the values of the feature map are mapped to (0, 1) through a Softmax operation and normalized, so as to output an interest point score map. The descriptor decoding layer then upsamples the obtained semi-dense descriptors using bicubic interpolation to generate complete descriptors, and at the same time performs L2 norm normalization to finally obtain dense descriptors

[0100] 2. Feature Matching

[0101] Traditional feature matching methods usually use manually designed feature descriptors and Euclidean distance to identify pairs of homologous feature points. When there are obvious non - linear radiation differences between two images, the extracted feature descriptors may have large differences, and the matching results are prone to falling into local extrema, resulting in the loss of a large number of correct matching point pairs. To solve the above problems, the present invention uses SuperGlue as a feature matching network to match the multi - scale feature points extracted;

[0102] SuperGlue feature matching is a simple and robust feature matching network. It transforms feature matching into an optimal assignment problem between two sets of feature points, and can simultaneously achieve feature matching and outlier rejection. Different from the classical NN matching method that only uses feature descriptors for feature matching, the SuperGlue feature matching network takes the positions and descriptors of feature points as network inputs together to construct a more representative feature representation. The feature matching network mainly includes two modules: the attention graph neural network layer and the optimization matching layer, and its network framework is as Figure 3 shown;

[0103] The attention graph neural network simulates the matching process of repeated browsing of the human visual system by designing a single complete graph with nodes as each feature point in the image. This graph structure includes two different undirected edges, which are used to connect the internal feature points in the image and all the feature points of the two images respectively. The key - point encoder maps the feature point coordinates and feature descriptors into one - dimensional feature vectors, and alternately updates through Self and Cross operations to generate more representative feature vectors, which are defined as shown in the following formula:

[0104]

[0105] In the formula, W and b represent weights and biases respectively;

[0106] The optimization matching layer then solves the optimal assignment matrix by calculating the similarity between the feature vectors generated by the attention graph neural network, so as to obtain correct matches and reject incorrect matches. Suppose image A and image B have M and N feature points respectively. First, establish an assignment matrix P ∈ [0,1] M×N , and calculate the inner product of the corresponding descriptors and of all possible matching points to obtain the similarity S i,j to generate a score matrix s, which is calculated using the following formula. Then, add a dusbin channel to the last row or last column of the score matrix s to reject outliers, and perform the operation of maximizing the total score, so as to solve and obtain the optimal assignment matrix P;

[0107]

[0108] In step 202, the registered multi-modal images are denoised. Specifically:

[0109] A partial differential equation denoising model based on a hybrid YK&PM model is constructed, and the registered multi-modal images are denoised based on the denoising model;

[0110] Introduce the model. The traditional PM and YK denoising models have deficiencies that can be improved in the process of image denoising. Especially for some internal texture structures and edge information of edge pixels and corner pixels in the image, only using the gradient operator to diffuse cannot meet the established denoising requirements. A new hybrid model is proposed. This model introduces the patch similarity module, constructs a new diffusion function, makes up for the defect of the weak edge detection ability of the YK model, and then introduces a coupling coefficient to combine the improved YK model with the PM model. This model can effectively filter out noise, retain texture and details while protecting edges;

[0111] The present invention proposes a model for denoising by hybrid PM&YK. This model can take the respective advantages of the two independent models and well compensate for each other's defects. By introducing a coupling coefficient θ, this hybrid model can achieve the optimal effect when denoising;

[0112] The new denoising model is:

[0113]

[0114] In the formula, PSYK(u) represents the result of the noise image denoised by the YK model improved by the patch similarity module and the level set curvature, and PM(u) represents the result of the noise denoised by the PM model.

[0115] The new diffusion equation is:

[0116]

[0117] In the formula, is the gradient operator, is the Laplace operator, c(·) represents the diffusion function, and the diffusion function is as follows:

[0118]

[0119] θ is the coupling coefficient, and its value is as follows:

[0120]

[0121] Wherein, S represents the patch similarity modulus. When the pixel point is in the flat area, the patch similarity modulus value is low, the coupling coefficient θ tends to be one, and the hybrid model is more inclined to use the YK model improved based on patch similarity for denoising, which can better filter out noise. When the pixel point is in the edge area, the patch similarity modulus value is high, the coupling coefficient θ tends to be zero, and the hybrid model is more inclined to use the PM model for denoising, which can better preserve the edge structure of the image;

[0122] The calculation process of the YK model improved based on the patch similarity modulus is as follows:

[0123] When the image undergoes the nth iteration, the Laplacian operator value at the point (i, j) of the image is:

[0124]

[0125] Take the intermediate variable f:

[0126]

[0127] In the formula, represents the patch similarity modulus value at the position of the pixel point (i, j) in the nth iteration, and the calculation formula is as shown below;

[0128] The Laplacian operator value corresponding to the intermediate variable f is:

[0129]

[0130] The iteration formula of the YK model improved based on the patch similarity modulus is as follows:

[0131]

[0132] The calculation process of the PM model is as follows:

[0133] Calculate the differences in each direction at the point (i, j) of the image:

[0134]

[0135] In the formula respectively represent the difference operators in the north, south, east, and west directions;

[0136] The iteration formula of the PM model is as follows:

[0137]

[0138] In the formula, represents the difference operator in a certain direction, and then the coupling coefficient of the model is calculated according to the following formula to obtain the iteration formula of the hybrid model:

[0139]

[0140] The edge conditions of the image are as follows:

[0141]

[0142] Where M and N are the length and width of the image.

[0143] In step 300, multi-modal feature extraction and fusion are performed on the preprocessed multi-modal image, specifically:

[0144] Step 301: Construct a feature extraction network structure and introduce its structure in detail.

[0145] The feature extraction network structure uses a lightweight 3D ResNet-18 network structure as the backbone network, reduces the computational cost by reducing the number of parameters (such as reducing the initial number of convolutional kernels from 64 to 32), and adapts to the high-resolution requirements of medical images;

[0146] Improve the backbone network as follows: multi-scale dilated convolution block: embed dilated convolution layers with dilation rates of [1, 2, 4] in the residual block to expand the receptive field and capture the local details and global spatial distribution of the tumor;

[0147] Step 302: Input the three-dimensional volume data into the feature extraction network to obtain high-dimensional feature maps of each modality;

[0148] Step 303: Based on the channel attention mechanism, fuse the high-dimensional feature maps to obtain a fused feature map, which is a prior art and the specific process will not be introduced in detail.

[0149] In step 400, brain tumor segmentation and classification are performed based on the fused multi-modal features, specifically:

[0150] Step 401: Construct a brain tumor segmentation model based on the improved U-Net network and a brain tumor classification model based on the improved SE-Net network;

[0151] Step 402: Train the brain tumor segmentation model and the brain tumor classification model based on a preset dataset;

[0152] Step 403: Input the fused feature map into the trained brain tumor segmentation model to obtain the brain tumor segmentation result;

[0153] Step 404: Input the fused feature map and the brain tumor segmentation result into the trained brain tumor classification model to obtain the brain tumor classification result.

[0154] It should be noted that the preset dataset is set according to specific requirements and will not be elaborated here. This mainly focuses on the detailed introduction of the brain tumor segmentation model based on the improved U-Net network and the brain tumor classification model based on the improved SE-Net network. Among them, the brain tumor segmentation model based on the improved U-Net network will be introduced first:

[0155] 1. Network Structure of the Brain Tumor Segmentation Model Based on the Improved U-Net Network

[0156] Since the residual block can solve the problem of gradient disappearance in the deep learning network, the present invention replaces the convolutional block in the U-Net basic network with the improved residual block and adds the CBAM attention mechanism to the skip connection to focus on the key information in space and channels, so as to improve the effect of brain tumor segmentation;

[0157] As Figure 4 shown, it is the network structure of the present invention, which consists of five parts, including encoding, bridging, skip connection, decoding and classifier. Among them, the encoding and decoding regions are composed of four improved residual blocks, the bridging region is composed of one improved residual block, and one improved residual block is composed of BN, Dy-ReLU activation function, 3×3 convolutional layer and identity mapping;

[0158] The encoding region is mainly composed of the improved residual block and downsampling. The downsampling mainly uses global max pooling. A total of four operations are performed in the figure. Each time a residual operation is passed, global max pooling will be executed once. The bridge plays an essential role in the network model and is used to connect the encoding and decoding regions. The decoding region is mainly composed of the improved residual block and upsampling. A total of four operations are performed in the figure. Each upsampling operation will reduce the number of channels by half and double the image size. Finally, a graph with the same size as the input feature map is obtained. The classifier is composed of a combination of 1×1 convolution and Sigmoid activation function. The 1×1 convolution is used to reduce the number of channels, and the Sigmoid activation function is used to calculate the category of each pixel in the feature map to obtain the feature segmentation map;

[0159] The skip connection is used to fuse the features of the encoding region and the decoding region, which is achieved by cascading shallow features and deep features. However, since the feature information extracted from the encoding region has poor effect and brings a large amount of redundant feature information, to solve this problem, the present invention introduces the CBAM attention mechanism before feature fusion to extract feature information from both the channel and spatial directions and suppress the redundant regions to improve the accuracy of the model;

[0160] The activation function is used to introduce nonlinear factors so that the model has the ability of nonlinear mapping. The ReLU activation function is frequently used in segmentation, but it does not change according to the changes in experimental data. It treats all input samples equally. The Dy-ReLU activation function is different. It can dynamically adjust the activation function using input information. Therefore, the present invention improves the nonlinear expression ability of the network model by introducing the Dy-ReLU activation function.

[0161] The schematic diagram of Dy-ReLU is as follows Figure 5 As shown, x flows to θ(x) and f respectively. θ (x) (x) ,θ(x) is a learnable parameter and is used to encode the context information of the input x and calculate the parameters of the activation function. is the output value of the activation function, which is:

[0162]

[0163] In the formula, k (1≤k≤K) refers to the number of functions, c (1≤c≤C) refers to the number of channels, and x represents the input vector of the cth channel. is the parameter calculated by the auxiliary function θ(x), which is:

[0164]

[0165] θ(x) is implemented by the SE attention module, which is used to obtain the importance between each feature channel, and then extract useful and useless information in a targeted manner according to the importance, and finally weight the value to each feature channel;

[0166] x passes through the global average pooling layer and the fully connected layer successively, and the fully connected layer passes twice. Between the fully connected layers, the ReLU activation function is used to introduce more nonlinear factors, so that it can better fit the complex correlation between channels. Finally, the Sigmoid activation function is used to standardize the output range. When the calculation in the SE module is completed, the final output is shown in the following formula:

[0167]

[0168] Among them, α k , β k Respectively represent and The initial value of λ a and λ b is a scalar that controls the range, i.e., the weight added to each feature channel;

[0169] The present invention adopts an improved residual block, and its structural diagram is shown in Figure 6As shown in the figure, the ReLU activation function in the figure is replaced by the DyReLU activation function, and the context information is dynamically encoded to collect more favorable feature information. Assuming the input is x, the expected output is H(x). By adding an identity mapping y = x in the shallow network, the data can bypass the intermediate layer and be transmitted to the subsequent network layers. Subsequently, the input and output are superimposed using a shortcut connection, which can not only simplify the training of the network but also enable the rapid transmission of information. At this time, the function to be learned becomes F(x) = H(x) - x, that is, the difference between the input x and the output H(x), which is called a residual unit;

[0170] The residual formula is as follows:

[0171] y = w s x + F(x, {w i )

[0172] In the formula, x is the input, y is the output, W s is the convolution operation, F(x, {w i ) is the residual mapping, which is learned by the network. If the intermediate layer consists of two layers of networks, then F(x, {w i ) = w 2 σ(w 1 x). When there is a problem with the dimension of the feature map, it is very effective to use zero-padding to increase the dimension. Of course, 1×1 convolution can also be used to achieve dimension consistency;

[0173] Regarding the dual attention mechanism, the present invention will not introduce it in detail;

[0174] Since the hybrid loss function has good effects in medical images, therefore, the present invention adopts a hybrid loss function composed of Dice loss and cross-entropy loss. The cross-entropy formula is as follows:

[0175]

[0176] In the formula, N represents the number of pixel points, i is the pixel point serial number, g i refers to the true label value of the i-th pixel point, p i refers to the model prediction value of the i-th pixel point;

[0177] If there is a class imbalance problem, the cross-entropy often has a bias, which will affect the performance of the network. Therefore, the present invention solves such problems by adding a Dice loss function. The Dice loss function is as follows:

[0178]

[0179] In the formula, the smoothing operator is represented by ε, and the Dice loss function can guide the learning of network parameters to make the predicted value closer to the true value. The formula of the hybrid loss function is as follows:

[0180] Loss = Loss CE + Loss Dice

[0181] 2. Brain Tumor Classification Model Based on Improved SE-Net Network

[0182] The brain tumor classification model proposed by the present invention is as Figure 7 shown. First, after BN, the Swish activation function is used to replace the ReLU activation function. Compared with the ReLU activation function, Swish does not mask too many features, which can enable the model to better learn effective features. Secondly, after feature fusion, the ReLU activation function is replaced with the Swish activation function. Since Swish is differentiable everywhere and continuously smooth, compared with ReLU, it can significantly improve the expression ability of the model. In addition, after the first convolutional layer (including convolution, batch normalization, and Swish activation function) after global max pooling, an ECA attention module is added to focus on it in the channel features to maximize the extraction of channel information; then, a BAM attention mechanism is added after the second convolutional layer to enable it to perform feature extraction concurrently in both the spatial and channel directions, making the target features play their utmost role. Finally, an improved SE attention mechanism is adopted to obtain more channel information and improve the expression ability of the model;

[0183] After the input feature map, the model will perform three steps. First, the feature map is preprocessed through Step 1, including convolution, batch normalization, Swish activation function, and global max pooling; then, the processed image is operated in Step 2, which includes 4 sub-steps, namely convolution, batch normalization, and Swish activation function, ECA attention module, BAM attention model, and SE attention module, which are processed 3, 4, 6, and 3 times respectively in Step 2. If the input and output sizes are different, a convolution operation will be added; finally, after global average pooling and fully connected layer in Step 3, the output is obtained;

[0184] An introduction to the Swish activation function:

[0185] The Swish activation function, also known as the self-gating activation function, was proposed by Google in 2017. It has been verified that, under the same conditions, Swish can improve the accuracy of the model more than the ReLU activation function;

[0186] The expression of the Swish activation function is shown as follows, where β can be a constant or a parameter obtained through training. When β → ∞, Swish is the ReLU activation function. When β = 0, Swish becomes a linear function. Therefore, Swish can be regarded as a smooth activation function between the two. When X > 0, there is no case of gradient disappearance. When X < 0, the neuron will not die like ReLU. At the same time, compared with ReLU, the derivative of Swish is not constant, and Swish is differentiable everywhere, continuous and smooth;

[0187] Swish(X) = x · sigmoid(βX)

[0188] The improved SE attention module is as Figure 8 shown, and the main steps are as follows: First is the Squeeze stage. First, through the dual-channel pooling layer, the feature map with the size of W × H × C is transformed into the global information of 1 × 1 × C. Then, through the Excitation stage, this stage uses the fully connected layer to predict the importance between channels. A total of two fully connected layers are experienced. The first is scaling, which scales the input with the size of 1 × 1 × C to the output with the size of 1 × 1 × C / r, where r represents the scaling parameter; the second is restoration, and the input and output sizes are just the opposite of the first time, which restores the input with the size of 1 × 1 × C / r to the output with the size of 1 × 1 × C. Then, use Sigmoid to output the vector of the weights of each layer in the feature map. Finally, it is the Scale operation, which multiplies the output weight vector by the feature map to obtain the feature map with weight information. This step focuses on the effective features, avoids the invalid features, and can better extract the target features, which can improve the accuracy of the model to a certain extent;

[0189] Regarding the ECA attention module and the BAM attention mechanism, no improvement is made to them, so no detailed introduction will be given.

[0190] In step 500, generate an interpretable result based on the segmentation and classification results, specifically:

[0191] Based on the segmentation result, classification result and three-dimensional volume data, generate a heat map (showing the areas concerned by the model, which can be generated based on Grad-CAM), a three-dimensional rendering (tumor volume, location and relationship with surrounding tissues), a quantitative report (tumor volume, malignancy probability, distance to adjacent structures, etc.), and so on.

[0192] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same and similar parts between the various embodiments, reference can be made to each other.

[0193] In this text, specific examples are used to illustrate the principles and implementation modes of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation modes and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. A brain tumor image recognition system, characterized in that: include: Brain tumor multimodal image acquisition module, used to acquire brain tumor multimodal images; A preprocessing module, used for preprocessing the acquired multimodal images; Multimodal feature extraction and fusion, used to extract and fuse multimodal features of multimodal images after preprocessing operations; Brain tumor segmentation and classification, used to segment and classify brain tumors based on fused multimodal features; The interpretable result generation module is used to generate interpretable results based on the segmentation and classification results.

2. A method for identifying a brain tumor image, characterized in that: include: Step 100: Acquire multimodal images of brain tumors; Step 200: performing preprocessing operations on the collected multimodal images; Step 300: extracting and fusing multimodal features from the multimodal images after preprocessing; Step 400: performing brain tumor segmentation and classification based on the fused multimodal features; Step 500: Generate interpretable results based on the segmentation and classification results.

3. The method according to claim 2, characterized in that In step 200, the collected multimodal images are preprocessed, specifically: Step 201: performing registration processing on the collected multi-modal images; Step 202: performing denoising on the multimodal image after the registration process; Step 203: standardizing and normalizing the denoised multimodal image; Step 204: Perform three-dimensional reconstruction based on the standardized and normalized multimodal images to obtain three-dimensional volume data.

4. The method according to claim 3, characterized in that: In step 201, the acquired multimodal images are registered, specifically: The reference image and the image to be registered in the multimodal image are processed into grayscale images, and the feature points with accurate positions and corresponding descriptors are preliminarily extracted through the feature extraction network. The preliminarily extracted feature points are screened using the quadtree algorithm to achieve uniform distribution. The feature points and corresponding feature descriptors after uniform screening are input into the feature matching network, and feature matching and outlier removal are performed according to a given matching threshold to obtain feature points with the same name; After obtaining a large number of feature points with the same name, the piecewise linear model is used to iteratively estimate the transformation model parameters between the reference image and the image to be registered, and finally the registration processing of multimodal images is realized.

5. The method according to claim 4, characterized in that In step 202, the multimodal image after the registration process is subjected to denoising, specifically: A partial differential equation denoising model based on the hybrid YK&PM model is constructed, and the multimodal images after registration are denoised based on the denoising model.

6. The method according to claim 5, characterized in that In step 300, multimodal feature extraction and fusion are performed on the multimodal image after the preprocessing operation, specifically: Step 301: construct a feature extraction network structure; Step 302: input the three-dimensional volume data into a feature extraction network to obtain a high-dimensional feature map of each modality; Step 303: Fuse the high-dimensional feature map based on the channel attention mechanism to obtain a fused feature map.

7. The method according to claim 6, characterized in that The feature extraction network structure adopts a lightweight 3DResNet-18 network structure as the backbone network, and embeds a hole convolution layer in its residual block to expand the receptive field.

8. The method according to claim 7, characterized in that In step 400, brain tumor segmentation and classification are performed based on the fused multimodal features, specifically: Step 401: constructing a brain tumor segmentation model based on an improved U-Net network and a brain tumor classification model based on an improved SE-Net network; Step 402: training a brain tumor segmentation model and a brain tumor classification model based on a preset data set; Step 403: inputting the fused feature map into the trained brain tumor segmentation model to obtain a brain tumor segmentation result; Step 404: Input the fused feature map and the brain tumor segmentation result into the trained brain tumor classification model to obtain a brain tumor classification result.

9. The method according to claim 8, characterized in that Improve the U-Net network, specifically: The convolution blocks in the traditional U-Net structure are replaced with deep residual blocks, the CBAM attention mechanism is added to the jump connections of the traditional U-Net structure, and the ReLU activation function in the deep residual block is replaced with the Dy-ReLU activation function.

10. The method according to claim 9, characterized in that Improve the SE-Net network, specifically: After the BN of the traditional SE-Net network, the ReLU activation function is replaced by the Swish activation function, the ECA attention module is added after the first convolution after the global maximum pooling of the traditional SE-Net network, and the BAM attention mechanism is added after the second convolution layer. Finally, the improved SE attention mechanism is used to obtain more channel information.

Citation Information

Patent Citations

  • Image denoising method based on PM model and four-order YK model

    CN110060211A

  • Improved U-Net brain tumor segmentation method based on attention mechanism and multi-scale feature fusion

    CN115424103A

  • Image segmentation method, system and equipment

    CN116129124A

  • Brain tumor MR image segmentation method based on TWM-UNet

    CN117274275A

  • Rectal tumor magnetic resonance image automatic segmentation method based on improved UNet model

    CN117710971A