Brain glioma segmentation method and system based on region of interest in combination with artificial intelligence

By adopting multi-scale convolutional neural network and attention mechanism in the brain glioma segmentation technology, combining spatial pyramid pooling and feature fusion network, the problem of low segmentation accuracy for different sizes and morphology of brain gliomas in the existing technology is solved, and higher positioning accuracy and segmentation efficiency are achieved.

CN120198451AActive Publication Date: 2025-06-24THE AFFILIATED HOSPITAL OF SHANDONG UNIV OF TCM

Patent Information

Application Number
CN202510668305.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The existing deep learning-based glioma segmentation technology lacks adaptability when dealing with brain gliomas of different sizes and morphology, and the segmentation accuracy is not high, making it difficult to accurately identify the fine structure and boundary information of the tumor.

Method used

The brain glioma segmentation method based on the region of interest combined with artificial intelligence is adopted. Image features are extracted through multi-scale convolutional neural network and hierarchical feature processing mechanism, thermal maps are generated and candidate areas are screened according to the attention mechanism, convolution kernel parameters are adaptively adjusted, multi-scale features are extracted using a spatial pyramid pooling structure, and fusion is carried out through the feature fusion network, and refined segmentation is finally achieved through the semantic segmentation module.

Benefits of technology

It improves the accuracy of positioning of brain glioma areas, reduces background interference, improves segmentation efficiency, enhances the algorithm's adaptability to complex lesion areas, and significantly improves the segmentation accuracy of brain glioma boundaries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198451A_ABST
    Figure CN120198451A_ABST
Patent Text Reader

Abstract

The invention provides a brain glioma segmentation method and system based on a region of interest combined with artificial intelligence, and relates to the technical field of image segmentation, and the method comprises the steps: carrying out the preprocessing of a brain magnetic resonance image; extracting image features by using a multi-scale convolutional neural network; calculating an attention weight to generate a thermodynamic diagram, and binarizing to obtain a candidate region; adaptively adjusting convolution kernel parameters according to the size of the candidate region, extracting multi-scale features by adopting spatial pyramid pooling, and fusing the multi-scale features; semantic segmentation is carried out through a segmentation module; and marking a brain glioma area. According to the method, the accuracy and efficiency of brain glioma segmentation are improved, and the consumption of computing resources is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image segmentation, and in particular to a glioma segmentation method and system based on region of interest combined with artificial intelligence. Background Art

[0002] Traditional glioma segmentation methods mainly rely on manual segmentation or semi-automatic segmentation techniques. These methods usually require a large amount of time and effort from professional physicians, and there are problems such as strong subjectivity and poor repeatability. With the development of computer vision and artificial intelligence technologies, automatic segmentation methods based on deep learning have gradually become a research hotspot. In particular, convolutional neural networks (CNNs) have shown excellent performance in the field of medical image segmentation.

[0003] However, the existing deep learning-based glioma segmentation technologies still have the following defects and deficiencies: Existing methods lack adaptability in dealing with gliomas of different sizes and shapes, and have low segmentation accuracy for tumor regions with blurred boundaries and different sizes, making it difficult to accurately identify the fine structure and boundary information of tumors.

[0004] Traditional convolutional neural networks have high computational complexity when processing whole-brain MRI images, making it difficult to effectively focus on tumor regions, resulting in insufficient feature extraction of key regions by the network and affecting the accuracy and efficiency of segmentation. Summary of the Invention

[0005] Embodiments of the present invention provide a glioma segmentation method and system based on region of interest combined with artificial intelligence, which can solve the problems in the prior art.

[0006] In the first aspect of the embodiments of the present invention, a glioma segmentation method based on region of interest combined with artificial intelligence is provided, including: Receiving a brain magnetic resonance image collected by a medical imaging device, and preprocessing the brain magnetic resonance image to obtain a preprocessed brain magnetic resonance image; Based on a multi-scale convolutional neural network, using a hierarchical feature processing mechanism to extract features from the preprocessed brain magnetic resonance image to obtain image features; Calculating the attention weight of each pixel point in the image features, generating a heat map according to the attention weight, and performing binarization processing on the heat map based on a preset threshold to obtain a candidate region containing the region of interest; Adapting and adjusting the convolution kernel parameters according to the size of the candidate region, using a spatial pyramid pooling structure of different scales to extract multi-scale features of the candidate region, and fusing the multi-scale features through a feature fusion network to obtain enhanced candidate region features; Input the enhanced candidate region features into the segmentation module, and perform semantic segmentation on the region of interest through the segmentation module to obtain the segmentation result of glioma. Annotate the glioma region in the preprocessed brain magnetic resonance image based on the segmentation result to obtain a glioma segmentation image.

[0007] Based on a multi-scale convolutional neural network, use a hierarchical feature processing mechanism to extract features from the preprocessed brain magnetic resonance image to obtain image features, including: Construct a multi-scale convolutional neural network model, input the preprocessed brain magnetic resonance image into the feature extraction layer of the multi-scale convolutional neural network model, and extract the initial features of the preprocessed brain magnetic resonance image through the feature extraction layer. Input the initial features into the feature enhancement layer of the multi-scale convolutional neural network model, and perform feature decomposition on the initial features through convolutional kernels of different scales to obtain feature maps of multiple scales. Calculate the spatial attention coefficient and channel attention coefficient of each scale feature map based on the attention mechanism, and perform weighted fusion on the feature maps of multiple scales according to the spatial attention coefficient and the channel attention coefficient to obtain an enhanced feature map. Input the enhanced feature map into the feature optimization layer of the multi-scale convolutional neural network model, construct a residual connection module, input the enhanced feature map into multiple parallel residual branches respectively, and set different numbers of convolutional layers for each residual branch. Calculate the feature importance score for the output features of each residual branch, perform dynamic fusion on the output features of each residual branch based on the feature importance score to obtain an optimized feature map, and output image features based on the optimized feature map.

[0008] Calculate the attention weight of each pixel point in the image features, generate a heat map according to the attention weight, and perform binarization processing on the heat map based on a preset threshold to obtain a candidate region including the region of interest, including: Construct a spatial attention module and a channel attention module, and input the image features into the spatial attention module and the channel attention module respectively. Construct a multi-scale receptive field unit through the spatial attention module to extract the multi-level spatial context information of the image features, and use a non-local correlation calculation unit to obtain the spatial dependence relationship of pixel points in the multi-level spatial context information to obtain the spatial attention weight. Construct a channel interaction learning unit through the channel attention module to establish the correlation mapping between channels of the image features, and use a channel recalibration unit to adaptively calibrate the channel features to obtain channel attention weights; Fuse the spatial attention weights and the channel attention weights to obtain the hybrid attention weights for each pixel point in the image features; Construct an attention weight enhancement module based on a depthwise separable convolutional network, and use the attention weight enhancement module to perform layer-by-layer feature enhancement and progressive optimization on the hybrid attention weights to generate a high-resolution heat map; Calculate the adaptive segmentation threshold of the high-resolution heat map based on the region growing algorithm, and perform binary segmentation on the high-resolution heat map to obtain candidate regions containing the region of interest.

[0009] Construct an attention weight enhancement module based on a depthwise separable convolutional network, and use the attention weight enhancement module to perform layer-by-layer feature enhancement and progressive optimization on the hybrid attention weights to generate a high-resolution heat map, including: Input the hybrid attention weights into the feature extraction unit of the depthwise separable convolutional network. Extract the channel features of the hybrid attention weights through the pointwise convolutional layer in the feature extraction unit, and extract the spatial features of the channel features through the depthwise convolutional layer in the feature extraction unit to obtain multi-dimensional basic features; Input the multi-dimensional basic features into the feature optimization unit of the depthwise separable convolutional network. Construct a multi-scale residual connection module through the feature optimization unit to decompose the multi-dimensional basic features, decompose the multi-dimensional basic features into feature subgraphs of different scales, and perform residual connection optimization on the feature subgraphs to obtain hierarchically enhanced features; Input the hierarchically enhanced features into the first feature enhancement unit of the depthwise separable convolutional network. Construct a hierarchical feature enhancement network through the first feature enhancement unit, use a dense connection structure to perform progressive feature extraction on the hierarchically enhanced features, and introduce an attention feedback mechanism in each progressive layer to adaptively calibrate the features to obtain calibrated features at multiple levels; Construct a pyramid feature fusion network through the first feature enhancement unit, perform cross-level feature alignment and dynamic weight allocation on the calibrated features at multiple levels, and adaptively fuse the aligned features to generate a high-resolution heat map.

[0010] Adaptive adjust the convolution kernel parameters according to the size of the candidate region, extract the multi-scale features of the candidate region using spatial pyramid pooling structures of different scales, and fuse the multi-scale features through a feature fusion network to obtain enhanced candidate region features, including: Perform multi-dimensional feature analysis on the candidate regions through a deep attention network to generate a spatial dimension feature vector and a channel dimension feature vector, and perform dynamic feature fusion on the spatial dimension feature vector and the channel dimension feature vector to generate geometric feature parameters; Input the geometric feature parameters into a feature mapping network, construct an adaptive parameter generation module through the feature mapping network, perform non-linear mapping and progressive optimization on the geometric feature parameters to generate adaptive convolution parameters matching the size of the candidate regions; Construct a multi-level spatial pyramid pooling network. The multi-level spatial pyramid pooling network is set with multiple pooling branches of variable scales in a cascaded structure, and input the candidate regions and the adaptive convolution parameters into the multi-level spatial pyramid pooling network; Perform hierarchical feature extraction on the candidate regions through the pooling branches of the multi-level spatial pyramid pooling network. In each pooling branch, construct a second feature enhancement unit in combination with the adaptive convolution parameters, and perform adaptive recalibration on the extracted features through the second feature enhancement unit to generate multi-scale feature representations; Construct a cross-scale feature fusion network, and perform feature alignment and hierarchical adaptive fusion on the multi-scale feature representations through the cross-scale feature fusion network to obtain enhanced candidate region features.

[0011] Perform hierarchical feature extraction on the candidate regions through the pooling branches of the multi-level spatial pyramid pooling network. In each pooling branch, construct a second feature enhancement unit in combination with the adaptive convolution parameters, and perform adaptive recalibration on the extracted features through the second feature enhancement unit to generate multi-scale feature representations, including: Perform multi-scale feature analysis on the candidate regions through a deep feature extraction network. The deep feature extraction network uses a residual connection structure to decompose the features of the candidate regions to generate multi-level feature vectors; Input the multi-level feature vectors into an attention parameter network, perform channel feature response calculation and spatial dependence calculation on the multi-level feature vectors through the attention parameter network to generate a channel enhancement factor and a spatial enhancement factor; Construct a dynamic parameter generation module based on the channel enhancement factor and the spatial enhancement factor, and generate adaptive convolution parameters through the dynamic parameter generation module; Perform hierarchical feature extraction on the multi-level feature vectors through the pooling branches of the multi-level spatial pyramid pooling network. The pooling branches extract features layer by layer from fine-grained to coarse-grained to generate multi-level initial features; Construct a second feature enhancement unit in each pooling branch of the multi-level spatial pyramid pooling network. The second feature enhancement unit uses the adaptive convolution parameters to perform a non-linear transformation on the multi-level initial features to obtain transformed features; Use the channel enhancement factor and the spatial enhancement factor to perform adaptive recalibration on the transformed features to generate a multi-scale feature representation.

[0012] Input the enhanced candidate region features into the segmentation module, and perform semantic segmentation on the region of interest through the segmentation module to obtain the segmentation result of glioma, including: Input the enhanced candidate region features into the segmentation module, and perform multi-scale feature decomposition on the enhanced candidate region features through the feature extraction unit in the segmentation module to generate hierarchical semantic features; Input the hierarchical semantic features into the region correlation calculation module, and calculate the cross-region feature dependence of the hierarchical semantic features through the region correlation calculation module to generate a spatial enhancement factor; Perform region of interest detection on the hierarchical semantic features through the region localization unit in the segmentation module to generate a region feature map; Construct a boundary awareness module based on the region feature map, and extract the boundary information of the region of interest through the boundary awareness module to generate boundary constraint features; Perform initial semantic segmentation on the region of interest through the segmentation processing unit in the segmentation module to generate an initial segmentation boundary; Establish a boundary mapping relationship based on the initial segmentation boundary, and perform feature remapping and boundary reconstruction on the initial segmentation boundary by combining the boundary constraint features and the spatial enhancement factor; Re-perform feature remapping and boundary reconstruction on the reconstructed boundary, and generate the segmentation result of glioma through multi-round iterative boundary optimization.

[0013] In the second aspect of the embodiments of the present invention, a glioma segmentation system based on a region of interest combined with artificial intelligence is provided, including: A first unit, configured to receive a brain magnetic resonance image collected by a medical imaging device, and perform preprocessing on the brain magnetic resonance image to obtain a preprocessed brain magnetic resonance image; A second unit, configured to perform feature extraction on the preprocessed brain magnetic resonance image by using a hierarchical feature processing mechanism based on a multi-scale convolutional neural network to obtain image features; A third unit, configured to calculate the attention weight of each pixel point in the image features, generate a heat map according to the attention weight, and perform binarization processing on the heat map based on a preset threshold to obtain a candidate region including the region of interest; The fourth unit is used to adaptively adjust the convolution kernel parameters according to the size of the candidate region, extract multi-scale features of the candidate region by using a spatial pyramid pooling structure with different scales, and fuse the multi-scale features through a feature fusion network to obtain enhanced candidate region features; The fifth unit is used to input the enhanced candidate region features into a segmentation module, and perform semantic segmentation on the region of interest through the segmentation module to obtain a segmentation result of glioma; The sixth unit is used to label the glioma region in the preprocessed brain magnetic resonance image based on the segmentation result to obtain a glioma segmentation image.

[0014] In a third aspect of the embodiments of the present invention, an electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0015] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0016] The beneficial effects of this application are as follows: The present invention uses a multi-scale convolutional neural network and a hierarchical feature processing mechanism to extract image features, combines an attention mechanism to generate a heat map and perform candidate region screening, effectively improves the positioning accuracy of the glioma region, reduces background interference, and improves the segmentation efficiency.

[0017] The present invention adaptively adjusts the convolution kernel parameters according to the candidate region size, and extracts multi-scale features through a spatial pyramid pooling structure, solves the problem that traditional methods are difficult to process gliomas of different sizes and shapes, and enhances the adaptability of the algorithm to complex lesion regions.

[0018] The present invention effectively integrates multi-scale features through a feature fusion network, and uses a semantic segmentation module to achieve refined segmentation, significantly improving the segmentation accuracy of the glioma boundary, and providing a more reliable imaging basis for clinical diagnosis and treatment planning. Description of the Drawings

[0019] Figure 1 It is a flowchart of the glioma segmentation method based on the region of interest combined with artificial intelligence according to the embodiments of the present invention; Figure 2 It is the process of generating a heat map and extracting candidate regions by the hybrid attention mechanism of this application; Figure 3Schematic diagram of a glioma semantic segmentation framework based on multi-layer feature optimization and boundary reconstruction. Detailed implementation manners

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0022] Figure 1 Schematic flow chart of a glioma segmentation method based on regions of interest combined with artificial intelligence in an embodiment of the present invention, as Figure 1 shown, the method includes: Receiving a brain magnetic resonance image collected by a medical imaging device, and preprocessing the brain magnetic resonance image to obtain a preprocessed brain magnetic resonance image; Based on a multi-scale convolutional neural network, using a hierarchical feature processing mechanism to extract features from the preprocessed brain magnetic resonance image to obtain image features; Calculating the attention weight of each pixel point in the image features, generating a heat map according to the attention weight, and performing binary processing on the heat map based on a preset threshold to obtain a candidate region including the region of interest; Adapting and adjusting the convolution kernel parameters according to the size of the candidate region, extracting multi-scale features of the candidate region using a spatial pyramid pooling structure of different scales, and fusing the multi-scale features through a feature fusion network to obtain enhanced candidate region features; Inputting the enhanced candidate region features into a segmentation module, and performing semantic segmentation on the region of interest through the segmentation module to obtain a segmentation result of glioma; Based on the segmentation result, annotating the glioma region in the preprocessed brain magnetic resonance image to obtain a glioma segmentation image.

[0023] In an optional implementation manner, based on a multi-scale convolutional neural network, using a hierarchical feature processing mechanism to extract features from the preprocessed brain magnetic resonance image to obtain image features, includes: Construct a multi-scale convolutional neural network model, input the preprocessed brain magnetic resonance image into the feature extraction layer of the multi-scale convolutional neural network model, and extract the initial features of the preprocessed brain magnetic resonance image through the feature extraction layer; Input the initial features into the feature enhancement layer of the multi-scale convolutional neural network model, perform feature decomposition on the initial features through convolutional kernels of different scales, and obtain feature maps of multiple scales; Calculate the spatial attention coefficient and channel attention coefficient of each scale feature map based on the attention mechanism, and perform weighted fusion on the feature maps of multiple scales according to the spatial attention coefficient and the channel attention coefficient to obtain an enhanced feature map; Input the enhanced feature map into the feature optimization layer of the multi-scale convolutional neural network model, construct a residual connection module, input the enhanced feature map into multiple parallel residual branches respectively, and set different numbers of convolutional layers for each residual branch; Calculate the feature importance score for the output features of each residual branch, perform dynamic fusion on the output features of each residual branch based on the feature importance score to obtain an optimized feature map, and output image features based on the optimized feature map.

[0024] In this embodiment, the brain magnetic resonance image is first preprocessed, including operations such as image normalization, denoising, and enhancement. The pixel values of the preprocessed image are scaled to the range [0, 1], the noise is removed by Gaussian filtering, and the image contrast is enhanced by histogram equalization to make the image more suitable for subsequent feature extraction.

[0025] Construct a multi-scale convolutional neural network model, which includes a feature extraction layer, a feature enhancement layer, and a feature optimization layer. The feature extraction layer consists of 3 convolutional blocks, each convolutional block contains two 3×3 convolutional layers and one max pooling layer. The number of convolutional kernels in the first convolutional block is 32, the second is 64, and the third is 128. The ReLU function is used as the activation function. Input the preprocessed brain magnetic resonance image (size 256×256×1) into the feature extraction layer, and extract the initial features of the image through layer-by-layer convolution and pooling operations. The output feature map size is 32×32×128.

[0026] Input the initial features into the feature enhancement layer. The feature enhancement layer adopts a multi-scale convolutional structure and performs feature decomposition on the initial features through convolutional kernels of different scales. Specifically, use four different sizes of convolutional kernels: 1×1, 3×3, 5×5, and 7×7. The number of each convolutional kernel is 64. Perform convolutional operations on the initial features to obtain four feature maps of different scales, and the size of each feature map is 32×32×64.

[0027] Calculate the spatial attention coefficient and channel attention coefficient of each scale feature map based on the attention mechanism. For spatial attention, perform average pooling and max pooling on each feature map in the channel dimension to obtain two feature maps of 32×32×1. Concatenate these two feature maps and pass them through a 7×7 convolutional layer to output a spatial attention map of 32×32×1, and obtain the spatial attention coefficient after normalization by the Sigmoid function. For channel attention, perform average pooling and max pooling on each feature map in the spatial dimension to obtain two feature vectors of 1×1×64. Concatenate these two feature vectors and pass them through two fully connected layers (the first fully connected layer reduces the number of channels to 1 / 16 of the original, and the second fully connected layer restores the original number of channels) to output a channel attention vector of 1×1×64, and obtain the channel attention coefficient after normalization by the Sigmoid function.

[0028] Weightedly fuse the feature maps of multiple scales according to the spatial attention coefficient and channel attention coefficient. Specifically, multiply each scale feature map by the corresponding spatial attention coefficient and channel attention coefficient to obtain the weighted feature map, and then perform element-wise addition on the four weighted feature maps to obtain the enhanced feature map with a size of 32×32×64.

[0029] Input the enhanced feature map into the feature optimization layer to construct a residual connection module. The feature optimization layer contains three parallel residual branches, and each residual branch is set with a different number of convolutional layers. The first residual branch contains 1 3×3 convolutional layer, the second residual branch contains 2 3×3 convolutional layers, and the third residual branch contains 3 3×3 convolutional layers. The number of convolutional kernels in each convolutional layer is 64, and the activation function is ReLU. Each residual branch also contains a skip connection that adds the input feature to the output feature of the convolutional layer to form a residual learning structure.

[0030] Calculate the feature importance score for the output feature of each residual branch. Convert the output feature of each residual branch (with a size of 32×32×64) to a feature vector of 1×1×64 through global average pooling, and then map the feature vector to a scalar value through a fully connected layer. Obtain the feature importance score of each residual branch after normalization by the Softmax function.

[0031] Dynamically fuse the output features of each residual branch based on the feature importance score. Multiply the output feature of each residual branch by the corresponding feature importance score, and then perform element-wise addition on the weighted features to obtain the optimized feature map with a size of 32×32×64.

[0032] Finally, the optimized feature map is converted into a 1×1×64 feature vector through a global average pooling layer, and then the feature vector is mapped to the final image features through a fully connected layer, with the feature dimension being 256. These features can be used for subsequent brain disease classification or other tasks.

[0033] In practical applications, this method has achieved good results in the task of Alzheimer's disease diagnosis. An experiment was conducted using a dataset containing 1000 brain magnetic resonance images, among which 500 were Alzheimer's disease patients and 500 were healthy control groups. Using five-fold cross-validation, this method achieved a classification accuracy of 93.5% on the test set, which is 5.2 percentage points higher than that of traditional single-scale convolutional neural networks, demonstrating the effectiveness of the multi-scale feature extraction and hierarchical feature processing mechanism.

[0034] In an alternative implementation, calculating the attention weights of each pixel point in the image features, generating a heat map based on the attention weights, and performing binary processing on the heat map based on a preset threshold to obtain a candidate region containing the region of interest, includes: Construct a spatial attention module and a channel attention module, and input the image features into the spatial attention module and the channel attention module respectively; Construct a multi-scale receptive field unit through the spatial attention module to extract the multi-level spatial context information of the image features, and use a non-local correlation calculation unit to obtain the spatial dependence relationship of the pixel points in the multi-level spatial context information to obtain spatial attention weights; Construct a channel interaction learning unit through the channel attention module to establish an association mapping between the channels of the image features, and use a channel recalibration unit to adaptively calibrate the channel features to obtain channel attention weights; Fuse the spatial attention weights and the channel attention weights to obtain the mixed attention weights of each pixel point in the image features; Construct an attention weight enhancement module based on a depthwise separable convolutional network, and perform layer-by-layer feature enhancement and progressive optimization on the mixed attention weights through the attention weight enhancement module to generate a high-resolution heat map; Calculate the adaptive segmentation threshold of the high-resolution heat map based on the region growing algorithm, and perform binary segmentation on the high-resolution heat map to obtain a candidate region containing the region of interest.

[0035] Figure 2 This is the process of generating a heat map and extracting candidate regions for the hybrid attention mechanism of this application, as shown in Figure 2As shown, in this embodiment, a method for calculating the attention weight of each pixel point in image features and generating a heat map is provided. This method first constructs a spatial attention module and a channel attention module, and inputs the image features into these two modules for processing respectively.

[0036] The spatial attention module extracts multi-level spatial context information of the image features by constructing multi-scale receptive field units. Specifically, this module uses three convolutional kernels of different sizes, namely 3×3, 5×5, and 7×7, to perform convolutional operations on the input image features. The stride of each convolutional kernel is set to 1, and the padding method is "same" to keep the feature map size unchanged. Through these three convolutional kernels of different sizes, spatial context information within different receptive field ranges can be obtained. Subsequently, these three convolutional results are subjected to element-wise weighted summation, with weight coefficients of 0.3, 0.3, and 0.4 respectively, to obtain the fused multi-scale spatial features.

[0037] The non-local correlation calculation unit is used to obtain the spatial dependence relationship of pixel points in the multi-level spatial context information. This unit first reduces the dimension of the fused multi-scale spatial features to 1 / 8 of the original number of channels through a 1×1 convolution to obtain the query feature Q and the key feature K. Then, the similarity matrix between Q and K is calculated, and the similarity calculation uses the dot product operation and is processed by softmax normalization. At the same time, the original multi-scale spatial features are obtained through another 1×1 convolution to obtain the value feature V. Finally, the normalized similarity matrix is multiplied by V, and added to the original features through a residual connection to obtain the spatial attention weight containing global spatial dependence.

[0038] The channel attention module first constructs a channel interaction learning unit to establish the correlation mapping between channels of the image features. This unit uses global average pooling to compress the spatial dimension to obtain the global descriptor of each channel. Then, feature transformation is performed through a bottleneck structure composed of two fully connected layers. The first fully connected layer reduces the number of channels to 1 / 16 of the original, and uses the ReLU activation function; the second fully connected layer restores the number of channels to the original number of channels. This design can effectively capture the non-linear relationship between channels while reducing the number of parameters.

[0039] The channel recalibration unit adaptively calibrates the channel features. This unit first performs sigmoid activation on the features output by the channel interaction learning unit to map the values to between 0 and 1 as the importance weights of each channel. Then, these weights are multiplied element-wise with the original input features in the channel dimension to achieve adaptive adjustment of each channel and obtain the channel attention weight.

[0040] Fuse the spatial attention weight and the channel attention weight to obtain the hybrid attention weight. The specific fusion method is to perform an element-wise multiplication operation on the spatial attention weight and the channel attention weight, and then perform feature integration through a 1×1 convolution to obtain the final hybrid attention weight.

[0041] Construct an attention weight enhancement module based on the depthwise separable convolutional network to perform layer-by-layer feature enhancement and progressive optimization on the hybrid attention weight. This module contains three depthwise separable convolutional blocks, each block consists of a depthwise convolution (kernel_size = 3×3, stride = 1, padding = "same") and a pointwise convolution (kernel_size = 1×1). The number of output channels of the first block is 64, the second block is 32, and the third block is 16. After each depthwise separable convolutional block, there is a BatchNormalization layer and a ReLU activation function. Finally, a 1×1 convolution is used to reduce the number of channels to 1, and bilinear interpolation is used to upsample the feature map to the resolution of the original input image to generate a high-resolution heat map.

[0042] Calculate the adaptive segmentation threshold of the high-resolution heat map based on the region growing algorithm. First, calculate the histogram of the heat map to count the pixel value distribution. Divide the histogram into 256 bins, and the width of each bin is 1 / 256. Then calculate the mean μ and standard deviation σ of the heat map, and the initial threshold T0 is set to μ + 0.5σ. Next, perform iterative optimization: divide the heat map into the foreground region F and the background region B according to the current threshold Tk, calculate the means μF and μB of the foreground region and the background region respectively, and update the threshold Tk+1 = (μF + μB) / 2. Stop the iteration when |Tk+1 - Tk| < 0.001 or the number of iterations reaches 50 times to obtain the final adaptive segmentation threshold T.

[0043] Use the obtained adaptive segmentation threshold T to perform binary processing on the high-resolution heat map. For each pixel point in the heat map, if its value is greater than or equal to the threshold T, set it to 1 (foreground), otherwise set it to 0 (background). The binary image obtained in this way contains the candidate regions of the region of interest.

[0044] In practical applications, taking a 512×512 pixel medical image as an example, after processing by the above method, the obtained heat map clearly highlights the lesion area. Using the adaptive threshold T = 0.427 for binary processing, a lesion candidate region with an area of 1243 pixels is successfully segmented, and the coincidence degree with the region manually marked by professional doctors reaches 91.5%, which proves the effectiveness of this method in the detection of regions of interest in medical images.

[0045] Table 1: Table of candidate region generation results for different types of gliomas

[0046] As shown in Table 1, five test cases of gliomas of different types are recorded. Case 1 is low-grade astrocytoma, with a tumor size of 876 pixels, a calculated threshold of 0.382, a coincidence degree of 89.3%, and a false positive rate of 3.5%. Case 2 is glioblastoma multiforme, with the largest tumor area of 1243 pixels, a calculated threshold of 0.427, the highest coincidence degree of 91.5%, and a false positive rate of only 2.8%. Case 3 is oligodendroglioma, with an area of 652 pixels, a calculated threshold of 0.415, and a coincidence degree of 90.7%. Case 4 is ependymoma, with an area of 945 pixels and a lower coincidence degree of 88.6%. Case 5 is anaplastic astrocytoma, with an area of 1087 pixels, a calculated threshold of 0.431, a coincidence degree of 92.3%, and the lowest false positive rate of only 2.5%. The average tumor size is 960.6 pixels, the average calculated threshold is 0.411, the average coincidence degree reaches 90.5%, and the average false positive rate is 3.2%. These data indicate that the method of this application has good adaptability and stability for gliomas of different types and sizes.

[0047] In an alternative embodiment, an attention weight enhancement module is constructed based on a depthwise separable convolutional network, and the hybrid attention weight is subjected to layer-by-layer feature enhancement and progressive optimization through the attention weight enhancement module to generate a high-resolution heat map, including: Input the hybrid attention weight into the feature extraction unit of the depthwise separable convolutional network, extract the channel features of the hybrid attention weight through the pointwise convolutional layer in the feature extraction unit, and extract the spatial features of the channel features through the depthwise convolutional layer in the feature extraction unit to obtain multi-dimensional basic features; Input the multi-dimensional basic features into the feature optimization unit of the depthwise separable convolutional network, construct a multi-scale residual connection module through the feature optimization unit to decompose the multi-dimensional basic features, decompose the multi-dimensional basic features into feature sub-maps of different scales, and perform residual connection optimization on the feature sub-maps to obtain hierarchical enhanced features; Input the hierarchical enhanced features into the first feature enhancement unit of the depthwise separable convolutional network, construct a hierarchical feature enhancement network through the first feature enhancement unit, adopt a dense connection structure to perform progressive feature extraction on the hierarchical enhanced features, and introduce an attention feedback mechanism in each progressive layer to adaptively calibrate the features to obtain calibrated features of multiple levels; Construct a pyramid feature fusion network through the first feature enhancement unit, perform cross-level feature alignment and dynamic weight allocation on the calibrated features of multiple levels, and adaptively fuse the aligned features to generate a high-resolution heat map.

[0048] The mixed attention weights are input into the feature extraction unit of the depthwise separable convolutional network for processing. The feature extraction unit consists of multiple depthwise separable convolutional blocks, and each block contains a pointwise convolutional layer and a depthwise convolutional layer. The pointwise convolutional layer uses a 1×1 convolutional kernel to perform feature extraction on the mixed attention weights in the channel dimension. The number of convolutional kernels is set to 64, the stride is 1, and there is no padding. After pointwise convolution, a channel feature map with a size of H×W×64 is obtained, where H and W are the height and width of the input feature map, respectively. The depthwise convolutional layer uses a 3×3 convolutional kernel to perform feature extraction on the channel features in the spatial dimension. Each channel uses an independent convolutional kernel, the stride is 1, and the padding is 1 to keep the size of the feature map unchanged. After depthwise convolution processing, a spatial feature map with a size of H×W×64 is obtained. Three consecutive depthwise separable convolutional blocks are set in the feature extraction unit, and a batch normalization layer and a ReLU activation function are added between each block. Finally, multi-dimensional basic features with a size of H×W×128 are output.

[0049] The multi-dimensional basic features are input into the feature optimization unit of the depthwise separable convolutional network for processing. The feature optimization unit constructs a multi-scale residual connection module to perform feature decomposition on the multi-dimensional basic features. First, different-scale feature extraction is performed on the input features through three parallel depthwise separable convolutional branches. The first branch uses a 3×3 convolutional kernel, the stride is 1, and the padding is 1; the second branch uses a 5×5 convolutional kernel, the stride is 1, and the padding is 2; the third branch uses a 7×7 convolutional kernel, the stride is 1, and the padding is 3. Each branch generates a feature sub-map with a size of H×W×64. For each feature sub-map, a residual connection structure is constructed, and the original input features are reduced in dimension through a 1×1 convolution and then fused with the corresponding feature sub-map through element-wise addition. The three feature sub-maps after residual connection optimization are then merged through channel concatenation to obtain hierarchical enhanced features with a size of H×W×192.

[0050] The hierarchical enhanced features are input into the first feature enhancement unit of the depthwise separable convolutional network for processing. The first feature enhancement unit constructs a hierarchical feature enhancement network, and uses a dense connection structure to perform progressive feature extraction on the hierarchical enhanced features. This network consists of 4 progressive layers, and each layer is composed of a depthwise separable convolutional block and an attention feedback module. In each progressive layer, the depthwise separable convolutional block uses a 3×3 convolutional kernel to extract features, with a stride of 1, padding of 1, and the number of output channels being 64. The attention feedback mechanism is implemented by combining spatial attention and channel attention. Spatial attention generates two spatial descriptors through global average pooling and max pooling, and after concatenating these two descriptors, a 7×7 convolutional layer and a Sigmoid activation function are used to generate spatial attention weights. Channel attention generates two channel descriptors through global average pooling and max pooling, and after processing these two descriptors separately through a shared multi-layer perceptron and adding them together, a Sigmoid activation function is used to generate channel attention weights. The spatial attention weights and channel attention weights are multiplied to obtain comprehensive attention weights, which are used to adaptively calibrate the features. After being processed by 4 progressive layers, 4 calibrated features at different levels are obtained, with sizes of H×W×64, H×W×64, H×W×64, and H×W×64 respectively.

[0051] The first feature enhancement unit also constructs a pyramid feature fusion network to perform cross-level feature alignment and dynamic weight allocation on the calibrated features at multiple levels. First, upsampling or downsampling operations are performed on the calibrated features at 4 levels to make their sizes consistent, all being H×W×64. For features with lower resolutions, bilinear interpolation is used for upsampling; for features with higher resolutions, 2×2 max pooling with a stride of 2 is used for downsampling. After feature alignment, a channel attention module is constructed to assign dynamic weights to each level of features. The channel attention module compresses the features into a 1×1×64 vector through global average pooling, and then generates a channel weight vector through two fully connected layers and a Sigmoid activation function. Each level of features is multiplied by the corresponding channel weight vector to achieve adaptive weighting of the features. Finally, the weighted 4-level features are fused by element-wise addition to obtain a fused feature with a size of H×W×64. The number of channels of the fused feature is adjusted to 1 through a 1×1 convolutional layer, and then normalized to the 0-1 range through a Sigmoid activation function to generate the final high-resolution heat map with a size of H×W×1.

[0052] In practical applications, the input image size is 224×224×3. After the above processing, the size of the finally generated high-resolution heat map is 56×56×1. It can be upsampled to the same size as the original image, 224×224×1, by bilinear interpolation method for visualizing the importance distribution of the target area. Experimental results show that this method can improve the average precision (AP) by 2.7 percentage points in the object detection task and the accuracy by 1.8 percentage points in the image classification task.

[0053] In an alternative embodiment, the convolution kernel parameters are adaptively adjusted according to the size of the candidate region, multi-scale features of the candidate region are extracted by using a spatial pyramid pooling structure with different scales, and the multi-scale features are fused through a feature fusion network to obtain enhanced candidate region features, including: Performing multi-dimensional feature analysis on the candidate region through a depth attention network to generate a spatial dimension feature vector and a channel dimension feature vector, and performing dynamic feature fusion on the spatial dimension feature vector and the channel dimension feature vector to generate geometric feature parameters; Inputting the geometric feature parameters into a feature mapping network, constructing an adaptive parameter generation module through the feature mapping network, performing non-linear mapping and progressive optimization on the geometric feature parameters to generate adaptive convolution parameters matching the size of the candidate region; Constructing a multi-level spatial pyramid pooling network, the multi-level spatial pyramid pooling network is provided with a plurality of pooling branches with variable scales in a cascaded structure, and inputting the candidate region and the adaptive convolution parameters into the multi-level spatial pyramid pooling network; Performing hierarchical feature extraction on the candidate region through the pooling branches of the multi-level spatial pyramid pooling network, constructing a second feature enhancement unit in each pooling branch by combining the adaptive convolution parameters, and performing adaptive recalibration on the extracted features through the second feature enhancement unit to generate multi-scale feature representations; Constructing a cross-scale feature fusion network, and performing feature alignment and hierarchical adaptive fusion on the multi-scale feature representations through the cross-scale feature fusion network to obtain enhanced candidate region features.

[0054] In this embodiment, the specific implementation process of adaptively adjusting the convolution kernel parameters according to the size of the candidate region, extracting multi-scale features of the candidate region by using a spatial pyramid pooling structure with different scales, and fusing the multi-scale features through a feature fusion network to obtain enhanced candidate region features is as follows: The system first conducts multi-dimensional feature analysis on the candidate regions through a deep attention network. This deep attention network includes a spatial attention module and a channel attention module. The spatial attention module uses a 3×3 convolutional layer to process the input candidate region feature map, generating a spatial weight map that reflects the importance of different positions in the feature map. Specifically, max pooling and average pooling operations are performed on the input feature map F to obtain two feature maps F_max and F_avg. These two feature maps are concatenated in the channel dimension and then passed through a 3×3 convolutional layer and a Sigmoid activation function to generate a spatial dimension feature vector S. The channel attention module generates two channel descriptors through global average pooling and global max pooling respectively. After being processed by a shared multi-layer perceptron, these two descriptors are fused through an element-wise addition operation and finally a channel dimension feature vector C is generated through the Sigmoid function. The size of the spatial dimension feature vector S is H×W×1, and the size of the channel dimension feature vector C is 1×1×C, where H and W represent the height and width of the feature map respectively, and C represents the number of channels.

[0055] Dynamic feature fusion is performed on the spatial dimension feature vector S and the channel dimension feature vector C to generate geometric feature parameters G. Specifically, global average pooling is performed on S to obtain a spatial global descriptor S_g, and spatial expansion is performed on C to obtain C_e so that its size is the same as the input feature. Figure 1 match. Tensor multiplication is performed on S_g and C to obtain an initial fusion feature F_i. Then, a residual connection is made between F_i and the input feature map, and feature compression is performed through a 1×1 convolutional layer to obtain geometric feature parameters G. In practical applications, for an input feature map with a size of 224×224×64, the dimension of the generated geometric feature parameters G is 128.

[0056] The geometric feature parameters G are input into a feature mapping network. An adaptive parameter generation module is constructed through the feature mapping network to perform non-linear mapping and progressive optimization on the geometric feature parameters, generating adaptive convolution parameters that match the size of the candidate region. The feature mapping network adopts a multi-layer perceptron structure, including three fully connected layers, with a ReLU activation function and a Dropout layer (dropout rate of 0.2) added between each layer. The first layer maps the 128-dimensional geometric feature parameters G to 256-dimensional intermediate features, the second layer maps the intermediate features to 512-dimensional high-level features, and the third layer maps the high-level features to the final adaptive convolution parameters. The parameter dimension is dynamically adjusted according to the size of the target convolution kernel. For example, for a 3×3 convolution kernel, the output dimension is 3×3×C_in×C_out, where C_in and C_out are the number of input and output channels respectively. To enhance the generalization ability of the model, L2 regularization is added to the output of the second layer during training, and the regularization coefficient is set to 0.001.

[0057] Construct a multi-level spatial pyramid pooling network, which uses a cascaded structure to set multiple pooling branches with variable scales, and inputs the candidate regions and adaptive convolution parameters into the multi-level spatial pyramid pooling network. The network contains four pooling branches, corresponding to four different pooling scales: 1×1, 2×2, 3×3, and 6×6. Each pooling branch first performs an adaptive average pooling operation on the input feature map at the corresponding scale, and then uses a 1×1 convolutional layer to uniformly adjust the number of feature channels to 256.

[0058] Perform hierarchical feature extraction on the candidate regions through the pooling branches of the multi-level spatial pyramid pooling network. In each pooling branch, a second feature enhancement unit is constructed by combining the adaptive convolution parameters. The second feature enhancement unit adaptively recalibrates the extracted features to generate multi-scale feature representations. The second feature enhancement unit uses an adaptive convolution module, which uses the previously generated adaptive convolution parameters to perform convolution operations on the pooled features. Specifically, for the features Fi extracted by the i-th pooling branch, convolution operations are performed using the corresponding adaptive convolution parameters Wi to obtain the enhanced features Ei. Then, the enhanced features Ei are recalibrated through a channel attention mechanism to generate the final multi-scale feature representation Mi. In practical applications, for the pooled features with a size of 7×7×256, after adaptive convolution and channel attention recalibration, the feature size remains unchanged, but the discriminability of the features is enhanced.

[0059] Construct a cross-scale feature fusion network, and perform feature alignment and hierarchical adaptive fusion on the multi-scale feature representations through the cross-scale feature fusion network to obtain enhanced candidate region features. The cross-scale feature fusion network adopts a top-down and bottom-up bidirectional feature fusion strategy. The top-down path transfers high-level semantic features to low-level features through upsampling operations, and the bottom-up path fuses low-level detailed features into high-level features through skip connections. Specifically, the multi-scale feature representations M_1, M_2, M_3, and M_4 of the four pooling branches are respectively upsampled to the same spatial size (usually 1 / 4 the size of the input feature map), and then channel alignment is performed through a 1×1 convolutional layer. The aligned features are fused by weighted summation, and the weight coefficients are adaptively determined through learnable parameters. Finally, through a 3×3 convolutional layer and a global average pooling layer, the fused features are compressed into a vector representation with a fixed dimension (such as 2048 dimensions) as the enhanced candidate region features. On the test dataset, compared with the traditional ROI pooling method, the average precision of object detection by this method has increased by 3.2 percentage points, and the detection effect on small objects and objects with complex shapes has been improved more significantly.

[0060] In an alternative embodiment, hierarchical feature extraction is performed on the candidate regions through the pooling branches of the multi-level spatial pyramid pooling network. In each pooling branch, a second feature enhancement unit is constructed by combining the adaptive convolution parameters, and the extracted features are adaptively recalibrated by the second feature enhancement unit to generate multi-scale feature representations, including: The candidate regions are subjected to multi-scale feature analysis through a deep feature extraction network. The deep feature extraction network uses a residual connection structure to decompose the features of the candidate regions and generate multi-level feature vectors. The multi-level feature vectors are input into an attention parameter network. Through the attention parameter network, channel feature response calculation and spatial dependence calculation are performed on the multi-level feature vectors to generate a channel enhancement factor and a spatial enhancement factor. A dynamic parameter generation module is constructed based on the channel enhancement factor and the spatial enhancement factor, and adaptive convolution parameters are generated through the dynamic parameter generation module. The multi-level feature vectors are subjected to hierarchical feature extraction through the pooling branches of the multi-level spatial pyramid pooling network. The pooling branches extract features layer by layer from fine-grained to coarse-grained to generate multi-level initial features. A second feature enhancement unit is constructed in each pooling branch of the multi-level spatial pyramid pooling network. The second feature enhancement unit uses the adaptive convolution parameters to perform a non-linear transformation on the multi-level initial features to obtain the transformed features. The transformed features are adaptively recalibrated using the channel enhancement factor and the spatial enhancement factor to generate multi-scale feature representations.

[0061] In this embodiment, the specific implementation process of performing hierarchical feature extraction on candidate regions through the pooling branches of the multi-level spatial pyramid pooling network, constructing a second feature enhancement unit by combining adaptive convolution parameters, and adaptively recalibrating the extracted features to generate multi-scale feature representations is as follows: Candidate regions are first subjected to multi-scale feature analysis through a deep feature extraction network. This deep feature extraction network adopts a residual connection structure to decompose the features of candidate regions and generate multi-level feature vectors. Specifically, the deep feature extraction network consists of 5 convolutional blocks, each convolutional block contains 3 convolutional layers, the convolutional kernel size is 3×3, the stride is 1, and the padding is 1. The number of output channels of the first convolutional block is 64, the number of output channels of the second convolutional block is 128, the number of output channels of the third convolutional block is 256, the number of output channels of the fourth convolutional block is 512, and the number of output channels of the fifth convolutional block is 512. Each convolutional block is connected by a max-pooling layer, the pooling kernel size is 2×2, and the stride is 2. Inside each convolutional block, a residual connection structure is adopted, that is, the input feature is element-wise added to the convolved feature to alleviate the vanishing gradient problem and improve the feature extraction ability. For example, for a candidate region image with an input size of 224×224×3, after passing through the first convolutional block, the feature map size is 112×112×64, after passing through the second convolutional block, the feature map size is 56×56×128, and so on, finally obtaining a multi-level feature vector.

[0062] The multi-level feature vectors are input into the attention parameter network, and through the attention parameter network, channel feature response calculation and spatial dependence calculation are performed on the multi-level feature vectors to generate a channel enhancement factor and a spatial enhancement factor. Specifically, the attention parameter network includes a channel attention branch and a spatial attention branch. The channel attention branch first performs global average pooling and global max pooling on the input features to obtain two channel descriptors, namely the average pooling channel descriptor and the max pooling channel descriptor. These two descriptors are processed by a multi-layer perceptron with shared weights. The multi-layer perceptron contains two fully connected layers, and the ReLU activation function is used in the middle. The first fully connected layer reduces the number of channels to 1 / 16 of the original, and the second fully connected layer restores the number of channels to the original size. The two processed channel descriptors are subjected to element-wise addition operation and the channel enhancement factor is obtained through the Sigmoid function. For example, for an input feature size of 14×14×512, after global pooling, a 1×1×512 channel descriptor is obtained, and after being processed by the multi-layer perceptron, a channel enhancement factor with a size of 1×1×512 is obtained.

[0063] The spatial attention branch performs max-pooling and average-pooling on the input features respectively in the channel dimension to obtain two spatial feature maps. Then, these two feature maps are concatenated in the channel dimension and processed through a 7×7 convolutional layer and the Sigmoid function to obtain the spatial enhancement factor. For example, for an input feature size of 14×14×512, two 14×14×1 spatial feature maps are obtained after channel pooling. After concatenation, a 14×14×2 feature map is obtained. After convolution and Sigmoid function processing, the spatial enhancement factor with a size of 14×14×1 is obtained.

[0064] Based on the channel enhancement factor and the spatial enhancement factor, a dynamic parameter generation module is constructed to generate adaptive convolution parameters through the dynamic parameter generation module. Specifically, the dynamic parameter generation module first fuses the channel enhancement factor and the spatial enhancement factor, combines them using element-wise multiplication operation to obtain the comprehensive enhancement feature. Then, through a 1×1 convolutional layer, the comprehensive enhancement feature is mapped to the parameter space to generate the adaptive convolution parameters. The adaptive convolution parameters include the convolutional kernel weights and bias terms, which are used in the subsequent feature enhancement unit. For example, for a channel enhancement factor with a size of 1×1×512 and a spatial enhancement factor with a size of 14×14×1, a comprehensive enhancement feature with a size of 14×14×512 is obtained after feature fusion, and the adaptive convolution parameters are generated through the 1×1 convolutional layer.

[0065] The multi-level feature vectors are subjected to hierarchical feature extraction through the pooling branches of the multi-level spatial pyramid pooling network. The pooling branches extract features layer by layer from fine-grained to coarse-grained to generate multi-level initial features. Specifically, the multi-level spatial pyramid pooling network contains 4 pooling branches, corresponding to different pooling scales: 1×1, 2×2, 3×3, and 6×6 respectively. Each pooling branch first performs adaptive average pooling on the input features at the corresponding scale, and then unifies the number of channels to 256 through a 1×1 convolutional layer. For example, for an input feature size of 14×14×512, a 1×1×256 feature is obtained after passing through the 1×1 pooling branch, a 2×2×256 feature is obtained after passing through the 2×2 pooling branch, and so on.

[0066] Build a second feature enhancement unit in each pooling branch of the multi-level spatial pyramid pooling network. The second feature enhancement unit performs a non-linear transformation on the multi-level initial features using adaptive convolution parameters to obtain the transformed features. Specifically, the second feature enhancement unit first processes the multi-level initial features through an adaptive convolution layer, and the adaptive convolution layer uses the previously generated adaptive convolution parameters. Then, non-linear transformation is performed through a batch normalization layer and a ReLU activation function. For example, for the 2×2×256 features obtained from the 2×2 pooling branch, after being processed by the adaptive convolution layer, the batch normalization layer, and the ReLU activation function, the transformed features are obtained, and the size remains 2×2×256.

[0067] Adopt a channel enhancement factor and a spatial enhancement factor to adaptively recalibrate the transformed features to generate a multi-scale feature representation. Specifically, first perform an element-wise multiplication operation on the transformed features and the channel enhancement factor to achieve feature recalibration in the channel dimension. Then, perform an element-wise multiplication operation on the result and the spatial enhancement factor to achieve feature recalibration in the spatial dimension. Finally, upsample the recalibrated features of each pooling branch to make their size the same as the original input features, and concatenate them in the channel dimension to obtain the final multi-scale feature representation. For example, for the transformed features with a size of 2×2×256, after channel and spatial recalibration, they are upsampled to 14×14×256, and after being concatenated with the features of other pooling branches, a multi-scale feature representation with a size of 14×14×1024 is obtained.

[0068] In an alternative embodiment, input the enhanced candidate region features into a segmentation module, and perform semantic segmentation on the region of interest through the segmentation module to obtain the segmentation result of glioma, including: Input the enhanced candidate region features into a segmentation module, and perform multi-scale feature decomposition on the enhanced candidate region features through a feature extraction unit in the segmentation module to generate hierarchical semantic features; Input the hierarchical semantic features into a region correlation calculation module, and calculate the cross-region feature dependence of the hierarchical semantic features through the region correlation calculation module to generate a spatial enhancement factor; Perform region of interest detection on the hierarchical semantic features through a region localization unit in the segmentation module to generate a region feature map; Construct a boundary awareness module based on the region feature map, and extract the boundary information of the region of interest through the boundary awareness module to generate boundary constraint features; Perform initial semantic segmentation on the region of interest through a segmentation processing unit in the segmentation module to generate an initial segmentation boundary; Based on the initial segmentation boundary, establish a boundary mapping relationship, and combine the boundary constraint features and the spatial enhancement factor to perform feature remapping and boundary reconstruction on the initial segmentation boundary; remap the features and reconstruct the boundary of the reconstructed boundary again, and generate the segmentation result of glioma through multiple rounds of iterative boundary optimization.

[0069] Figure 3 Schematic diagram of a glioma semantic segmentation framework based on multi-layer feature optimization and boundary reconstruction, as Figure 3 shown, input the enhanced candidate region features into the feature extraction unit in the segmentation module. This feature extraction unit adopts a multi-layer convolutional neural network structure, including 5 convolutional layers. The size of the convolutional kernel in each layer is 3×3, the stride is 1, and the padding is 1. The number of input channels of the first convolutional layer is the same as the number of channels of the enhanced candidate region features, and the number of output channels is 64; the number of input channels of the second convolutional layer is 64, and the number of output channels is 128; the number of input channels of the third convolutional layer is 128, and the number of output channels is 256; the number of input channels of the fourth convolutional layer is 256, and the number of output channels is 512; the number of input channels of the fifth convolutional layer is 512, and the number of output channels is 512. After each convolutional layer, a batch normalization layer and a ReLU activation function are connected, and max-pooling operation is used for downsampling. The size of the pooling kernel is 2×2, and the stride is 2. Through this multi-scale feature decomposition method, hierarchical semantic features containing different resolution information are generated, denoted as F1, F2, F3, F4, and F5 respectively, where F1 represents the lowest-level feature and F5 represents the highest-level feature.

[0070] Input the hierarchical semantic features F1 to F5 into the region correlation calculation module, which uses an attention mechanism to calculate cross-region feature dependencies. Specifically, for each feature layer Fi, first reduce the number of channels to 1 / 8 of the original through a 1×1 convolution to obtain the query feature Qi, the key feature Ki, and the value feature Vi. For the query feature Qi and the key feature Ki, reshape them into a matrix with the shape of (H×W)×C, where H and W are the height and width of the feature map respectively, and C is the number of channels. Calculate the matrix product of Qi and the transpose of Ki to obtain a correlation matrix Ai with the size of (H×W)×(H×W). Perform softmax normalization on Ai, and then perform matrix multiplication with the value feature Vi reshaped into the shape of C×(H×W) to obtain the weighted feature. Reshape the weighted feature back to the original shape H×W×C, restore the original number of channels through a 1×1 convolution, and finally perform a residual connection with the original feature Fi to obtain the spatial enhancement factor Si. For a sample in the glioma dataset, the spatial enhancement factor Si can effectively capture the correlation between the tumor region and the surrounding tissues and enhance the feature representation of the tumor boundary.

[0071] The region localization unit in the segmentation module performs region of interest (ROI) detection on the hierarchical semantic features. This unit adopts the region proposal network structure and processes the feature map F5. First, a 3×3 convolutional layer is used to extract features from F5, and then two 1×1 convolutional layers are used to predict the region scores and bounding box regression parameters respectively. For glioma data, 9 anchor boxes with different scales and ratios are set, including square anchor boxes with sizes of 8×8, 16×16, and 32×32, and rectangular anchor boxes with aspect ratios of 1:2 and 2:1. The top 100 candidate regions are selected according to the predicted scores, and the non-maximum suppression algorithm is applied with an intersection over union (IoU) threshold of 0.7, and finally the top 10 regions with the highest scores are retained. These regions are mapped back to the original image space to generate the region feature map R.

[0072] A boundary-aware module is constructed based on the region feature map R. This module uses an edge detection network to extract the boundary information of the region of interest. Specifically, the Sobel operator is used to calculate the horizontal and vertical gradients of the region feature map R to obtain the gradient magnitude map and the gradient direction map. Then, a 3×3 convolutional layer is used to extract features from the gradient information, with 64 convolutional kernels, a stride of 1, and a padding of 1. Next, bilinear interpolation is used to upsample the extracted boundary features to the same resolution as the original input and fuse them with the low-level feature F1 to generate the boundary constraint feature B. For glioma data, the boundary constraint feature B can effectively represent the fine structure of the tumor edge, which helps to improve the segmentation accuracy.

[0073] The initial semantic segmentation of the region of interest is performed by the segmentation processing unit in the segmentation module. This unit adopts the U-shaped network structure, which includes an encoder and a decoder. The encoder uses the previously generated hierarchical semantic features F1 to F5; the decoder performs upsampling through transposed convolution and makes skip connections with the features of the corresponding layers of the encoder. Specifically, the decoder contains 4 transposed convolutional layers, each with a convolutional kernel size of 4×4, a stride of 2, and a padding of 1. The input channel number of the first transposed convolutional layer is 512, and the output channel number is 512; the input channel number of the second transposed convolutional layer is 1024 (512 + 512), and the output channel number is 256; the input channel number of the third transposed convolutional layer is 512 (256 + 256), and the output channel number is 128; the input channel number of the fourth transposed convolutional layer is 256 (128 + 128), and the output channel number is 64. Finally, a 1×1 convolutional layer is used to map the channel number to the number of classes (for the glioma segmentation task, the number of classes is 4, representing the background, edema area, enhanced tumor area, and necrosis area respectively) to generate the initial segmentation boundary M.

[0074] Based on the initial segmentation boundary M, a boundary mapping relationship is established, and the initial segmentation boundary is subjected to feature remapping and boundary reconstruction by combining the boundary constraint feature B and the spatial enhancement factor S. When specifically implemented, first, the initial segmentation boundary M and the boundary constraint feature B are multiplied element-wise to obtain the boundary enhancement feature EB. Then, EB and the spatial enhancement factor S are weighted and fused with weight coefficients of 0.7 and 0.3 to obtain the remapped feature RM. Next, the RM is processed using a 3×3 convolutional layer with 64 convolutional kernels, a stride of 1, and a padding of 1 to generate the reconstructed boundary RC.

[0075] The reconstructed boundary RC is subjected to feature remapping and boundary reconstruction again, and the final glioma segmentation result is generated through multiple rounds of iterative boundary optimization. In this embodiment, the number of iterations is set to 3. In each iteration, the reconstructed boundary RC of the previous round is used as the input, and the aforementioned feature remapping and boundary reconstruction process is repeated. After 3 rounds of iteration, the softmax function is used to normalize the final feature map to obtain the probability distribution of each pixel belonging to each category, and the category with the highest probability is taken as the segmentation label of the pixel, thereby generating the final segmentation result of glioma. In practical applications, the test results of this method on the BraTS2020 dataset show that the Dice similarity coefficient reaches 0.91, the sensitivity is 0.89, and the specificity is 0.94, improving the segmentation accuracy by about 15% compared with traditional methods.

[0076] The glioma segmentation system based on the region of interest and combined with artificial intelligence in the embodiments of the present invention includes: A first unit, configured to receive the brain magnetic resonance image collected by a medical imaging device, and preprocess the brain magnetic resonance image to obtain a preprocessed brain magnetic resonance image; A second unit, configured to perform feature extraction on the preprocessed brain magnetic resonance image by using a hierarchical feature processing mechanism based on a multi-scale convolutional neural network to obtain image features; A third unit, configured to calculate the attention weight of each pixel point in the image features, generate a heat map according to the attention weight, and perform binarization processing on the heat map based on a preset threshold to obtain a candidate region including the region of interest; A fourth unit, configured to adaptively adjust the convolutional kernel parameters according to the size of the candidate region, extract multi-scale features of the candidate region by using a spatial pyramid pooling structure of different scales, and fuse the multi-scale features through a feature fusion network to obtain enhanced candidate region features; A fifth unit, configured to input the enhanced candidate region features into a segmentation module, and perform semantic segmentation on the region of interest through the segmentation module to obtain a segmentation result of glioma; The sixth unit is configured to label the glioma region in the preprocessed brain magnetic resonance image based on the segmentation result, so as to obtain a glioma segmentation image.

[0077] In a third aspect of the embodiments of the present invention, an electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0078] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0079] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.

[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for glioma segmentation based on region of interest combined with artificial intelligence, characterized in that Including: Receiving a brain magnetic resonance image collected by a medical imaging device, and preprocessing the brain magnetic resonance image to obtain a preprocessed brain magnetic resonance image; Based on a multi-scale convolutional neural network, using a hierarchical feature processing mechanism to extract features from the preprocessed brain magnetic resonance image to obtain image features; Calculating the attention weight of each pixel point in the image features, generating a heat map according to the attention weight, and performing binarization processing on the heat map based on a preset threshold to obtain a candidate region containing the region of interest; Adapting and adjusting the convolution kernel parameters according to the size of the candidate region, using a spatial pyramid pooling structure of different scales to extract multi-scale features of the candidate region, and fusing the multi-scale features through a feature fusion network to obtain enhanced candidate region features; Inputting the enhanced candidate region features into a segmentation module, and performing semantic segmentation on the region of interest through the segmentation module to obtain a segmentation result of glioma; Based on the segmentation result, annotating the glioma region in the preprocessed brain magnetic resonance image to obtain a glioma segmentation image.

2. The method according to claim 1, wherein Based on a multi-scale convolutional neural network, using a hierarchical feature processing mechanism to extract features from the preprocessed brain magnetic resonance image to obtain image features, including: Constructing a multi-scale convolutional neural network model, inputting the preprocessed brain magnetic resonance image into the feature extraction layer of the multi-scale convolutional neural network model, and extracting the initial features of the preprocessed brain magnetic resonance image through the feature extraction layer; Inputting the initial features into the feature enhancement layer of the multi-scale convolutional neural network model, and performing feature decomposition on the initial features through convolution kernels of different scales to obtain feature maps of multiple scales; Calculating the spatial attention coefficient and channel attention coefficient of each scale feature map based on the attention mechanism, and performing weighted fusion on the feature maps of multiple scales according to the spatial attention coefficient and the channel attention coefficient to obtain an enhanced feature map; Inputting the enhanced feature map into the feature optimization layer of the multi-scale convolutional neural network model, constructing a residual connection module, inputting the enhanced feature map into multiple parallel residual branches respectively, and setting different numbers of convolutional layers for each residual branch; Calculating the feature importance score for the output features of each residual branch, dynamically fusing the output features of each residual branch based on the feature importance score to obtain an optimized feature map, and outputting image features based on the optimized feature map.

3. The method according to claim 1, wherein Calculating the attention weight of each pixel point in the image features, generating a heat map according to the attention weight, and performing binarization processing on the heat map based on a preset threshold to obtain a candidate region containing the region of interest, including: Constructing a spatial attention module and a channel attention module, and inputting the image features into the spatial attention module and the channel attention module respectively; Construct a multi-scale receptive field unit through the spatial attention module to extract the multi-level spatial context information of the image features, and use a non-local correlation calculation unit to obtain the spatial dependence relationship of pixel points in the multi-level spatial context information to obtain spatial attention weights; Construct a channel interaction learning unit through the channel attention module to establish an association mapping between channels of the image features, and use a channel recalibration unit to adaptively calibrate the channel features to obtain channel attention weights; Fuse the spatial attention weights and the channel attention weights to obtain the hybrid attention weights of each pixel point in the image features; Construct an attention weight enhancement module based on a depthwise separable convolutional network, and use the attention weight enhancement module to perform layer-by-layer feature enhancement and progressive optimization on the hybrid attention weights to generate a high-resolution heat map; Calculate the adaptive segmentation threshold of the high-resolution heat map based on the region growing algorithm, and perform binary segmentation on the high-resolution heat map to obtain candidate regions containing the region of interest.

4. The method according to claim 3, wherein Construct an attention weight enhancement module based on a depthwise separable convolutional network, and use the attention weight enhancement module to perform layer-by-layer feature enhancement and progressive optimization on the hybrid attention weights to generate a high-resolution heat map, including: Input the hybrid attention weights into the feature extraction unit of the depthwise separable convolutional network. Extract the channel features of the hybrid attention weights through the pointwise convolutional layer in the feature extraction unit, and extract the spatial features of the channel features through the depthwise convolutional layer in the feature extraction unit to obtain multi-dimensional basic features; Input the multi-dimensional basic features into the feature optimization unit of the depthwise separable convolutional network. Construct a multi-scale residual connection module through the feature optimization unit to decompose the multi-dimensional basic features, decompose the multi-dimensional basic features into feature sub-maps of different scales, and perform residual connection optimization on the feature sub-maps to obtain hierarchically enhanced features; Input the hierarchically enhanced features into the first feature enhancement unit of the depthwise separable convolutional network. Construct a hierarchical feature enhancement network through the first feature enhancement unit, use a dense connection structure to perform progressive feature extraction on the hierarchically enhanced features, and introduce an attention feedback mechanism in each progressive layer to adaptively calibrate the features to obtain calibrated features at multiple levels; Construct a pyramid feature fusion network through the first feature enhancement unit, perform cross-level feature alignment and dynamic weight allocation on the calibrated features at multiple levels, and adaptively fuse the aligned features to generate a high-resolution heat map.

5. The method according to claim 1, wherein Adaptive adjust the convolution kernel parameters according to the size of the candidate region, use spatial pyramid pooling structures of different scales to extract multi-scale features of the candidate region, and fuse the multi-scale features through a feature fusion network to obtain enhanced candidate region features, including: Perform multi-dimensional feature analysis on the candidate regions through a deep attention network to generate a spatial-dimensional feature vector and a channel-dimensional feature vector, and perform dynamic feature fusion on the spatial-dimensional feature vector and the channel-dimensional feature vector to generate geometric feature parameters; Input the geometric feature parameters into a feature mapping network, construct an adaptive parameter generation module through the feature mapping network, perform non-linear mapping and progressive optimization on the geometric feature parameters to generate adaptive convolution parameters matching the size of the candidate regions; Construct a multi-level spatial pyramid pooling network, which adopts a cascaded structure to set multiple pooling branches with variable scales, and input the candidate regions and the adaptive convolution parameters into the multi-level spatial pyramid pooling network; Perform hierarchical feature extraction on the candidate regions through the pooling branches of the multi-level spatial pyramid pooling network, construct a second feature enhancement unit in each pooling branch by combining the adaptive convolution parameters, and perform adaptive recalibration on the extracted features through the second feature enhancement unit to generate multi-scale feature representations; Construct a cross-scale feature fusion network, and perform feature alignment and hierarchical adaptive fusion on the multi-scale feature representations through the cross-scale feature fusion network to obtain enhanced candidate region features.

6. The method according to claim 5, characterized in that, Perform hierarchical feature extraction on the candidate regions through the pooling branches of the multi-level spatial pyramid pooling network, construct a second feature enhancement unit in each pooling branch by combining the adaptive convolution parameters, and perform adaptive recalibration on the extracted features through the second feature enhancement unit to generate multi-scale feature representations, including: Perform multi-scale feature analysis on the candidate regions through a deep feature extraction network, and the deep feature extraction network uses a residual connection structure to decompose the features of the candidate regions to generate multi-level feature vectors; Input the multi-level feature vectors into an attention parameter network, and perform channel feature response calculation and spatial dependence calculation on the multi-level feature vectors through the attention parameter network to generate a channel enhancement factor and a spatial enhancement factor; Construct a dynamic parameter generation module based on the channel enhancement factor and the spatial enhancement factor, and generate adaptive convolution parameters through the dynamic parameter generation module; Perform hierarchical feature extraction on the multi-level feature vectors through the pooling branches of the multi-level spatial pyramid pooling network, and the pooling branches extract features layer by layer from fine-grained to coarse-grained to generate multi-level initial features; Construct a second feature enhancement unit in each pooling branch of the multi-level spatial pyramid pooling network, and the second feature enhancement unit performs non-linear transformation on the multi-level initial features by using the adaptive convolution parameters to obtain transformed features; Perform adaptive recalibration on the transformed features by using the channel enhancement factor and the spatial enhancement factor to generate multi-scale feature representations.

7. The method according to claim 1, characterized in that, Input the enhanced candidate region features into a segmentation module, and perform semantic segmentation on the region of interest through the segmentation module to obtain the segmentation result of glioma, including: Input the enhanced candidate region features into the segmentation module, and perform multi-scale feature decomposition on the enhanced candidate region features through the feature extraction unit in the segmentation module to generate hierarchical semantic features; Input the hierarchical semantic features into the region correlation calculation module, and calculate the cross-region feature dependence of the hierarchical semantic features through the region correlation calculation module to generate a spatial enhancement factor; Perform region of interest detection on the hierarchical semantic features through the region localization unit in the segmentation module to generate a region feature map; Construct a boundary awareness module based on the region feature map, and extract the boundary information of the region of interest through the boundary awareness module to generate boundary constraint features; Perform initial semantic segmentation on the region of interest through the segmentation processing unit in the segmentation module to generate an initial segmentation boundary; Establish a boundary mapping relationship based on the initial segmentation boundary, and perform feature remapping and boundary reconstruction on the initial segmentation boundary by combining the boundary constraint features and the spatial enhancement factor; Re-perform feature remapping and boundary reconstruction on the reconstructed boundary, and generate the segmentation result of glioma through multi-round iterative boundary optimization.

8. A glioma segmentation system based on regions of interest combined with artificial intelligence for implementing the method according to any one of claims 1-7, characterized in that, Comprising: A first unit, configured to receive a brain magnetic resonance image collected by a medical imaging device, and preprocess the brain magnetic resonance image to obtain a preprocessed brain magnetic resonance image; A second unit, configured to perform feature extraction on the preprocessed brain magnetic resonance image by using a hierarchical feature processing mechanism based on a multi-scale convolutional neural network to obtain image features; A third unit, configured to calculate the attention weight of each pixel point in the image features, generate a heat map according to the attention weight, and perform binarization processing on the heat map based on a preset threshold to obtain a candidate region including the region of interest; A fourth unit, configured to adaptively adjust the convolution kernel parameters according to the size of the candidate region, extract multi-scale features of the candidate region by using a spatial pyramid pooling structure of different scales, and fuse the multi-scale features through a feature fusion network to obtain enhanced candidate region features; A fifth unit, configured to input the enhanced candidate region features into the segmentation module, and perform semantic segmentation on the region of interest through the segmentation module to obtain the segmentation result of glioma; A sixth unit, configured to label the glioma region in the preprocessed brain magnetic resonance image based on the segmentation result to obtain a glioma segmentation image.

9. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Medical image segmentation method, electronic equipment and storage medium

    CN115861149A

  • MRI (Magnetic Resonance Imaging) brain tumor segmentation method based on multi-scale feature fusion of improved U-Net

    CN117876399A

  • Thyroid nodule prediction method based on subnet attention and scale attention

    CN118379307A

  • Computer-implemented method for providing a positioning score regarding a positioning of an examining region in an x-ray image

    EP4345742A1

  • Apparatus to analyse diffusion magnetic resonance imaging data

    US20230056838A1

Cited By

  • Transfer learning driven intracranial tumor image data full-automatic segmentation method

    CN121305088A

  • Convection newborn intelligent identification method and device based on multi-scale neural network

    CN121392639A

  • Convection initial generation intelligent identification method and device based on multi-scale neural network

    CN121392639B