Human brain extraction method and system for enhancing anatomical structure and surface information perception, electronic equipment and storage medium

By introducing the anatomical structure perception module AA, the surface information perception module SA and the local-global feature enhancement module LGFE, the problem of insufficient modeling of global anatomical structure and low-level features in the existing methods is solved, and more accurate human brain extraction is achieved.

CN120471819AActive Publication Date: 2025-08-12YUNNAN UNITED VISION TECH CO LTD

Patent Information

Application Number
CN202510593322.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-12
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

The existing methods lack effective modeling of global anatomical structures and low-level features of the brain surface in human brain extraction tasks, resulting in insufficient segmentation accuracy, especially in unstable segmentation effects in complex anatomical areas and low-contrast areas.

Method used

The anatomical structure perception module AA, the surface information perception module SA and the local-global feature enhancement module LGFE are used to extract global and local anatomical structure information through multi-layer perceptron MLP branches and convolutional neural network CNN branches, and feature fusion is carried out through jump connection and attention mechanism to enhance the perception of brain surface features.

Benefits of technology

It improves the accuracy and robustness of human brain extraction, especially the segmentation accuracy in complex anatomical structure areas, and optimizes the ability to portray brain tissue boundaries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471819A_ABST
    Figure CN120471819A_ABST
Patent Text Reader

Abstract

The invention relates to a human brain extraction method and system for enhancing anatomical structure and surface information perception, electronic equipment and a storage medium, and belongs to the technical field of image processing. The anatomical structure sensing module aims to enhance understanding of global brain tissues, and the surface information sensing module emphasizes a local mode of a brain surface so as to solve challenges brought by low-level inter-class similarity of voxel intensity between brain tissues and non-brain tissues; furthermore, a novel local-global feature enhancement module is employed to enhance the association between local surface patterns and global brain tissue, providing a robust local-global representation that facilitates precise brain extraction. According to the method, the anatomical structure and surface information perception of the human brain is enhanced, so that accurate brain extraction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a human brain extraction method, system, electronic equipment, and storage medium for enhancing anatomical structure and surface information perception, and belongs to the technical field of image processing. Background Art

[0002] Brain extraction aims to segment the human brain from 3D MRI images to meet clinical needs such as brain volume measurement and analysis of brain morphological changes. This task was originally defined as an image processing problem. Researchers use pattern recognition techniques to distinguish brain from non-brain structures. Common methods include histogram analysis, edge detection, and threshold segmentation. However, these methods often rely heavily on pre-set hyperparameters, such as edge constants and thresholds, making them sensitive to variations in voxel intensity caused by different medical devices. Traditional machine learning methods attempt to learn discriminative feature representations from data, but these methods still rely on handcrafted features. Deep learning has become a leading solution for many medical image segmentation tasks, providing the ability to learn effective deep representations. In particular, convolutional neural networks based on the U-net architecture have achieved excellent performance in medical image segmentation tasks and have been used for brain extraction. Existing work has improved the U-net architecture in various ways, including feature fusion, adversarial mechanisms, and attention mechanisms, to enhance its representation of local features and global semantic information. With the emergence of large visual models, researchers have begun to improve these large visual models for natural images to adapt them to medical image segmentation tasks.

[0003] Limited by the receptive field and pooling operations, CNNs are unable to capture global feature information. Traditional U-Net and its variants rely mainly on convolution operations to extract local features, resulting in insufficient fusion of global semantic information and local surface features, especially limited segmentation capabilities for complex anatomical regions. The surface structure of the human brain is complex, and its low-level features depict the detailed features of the brain's edges, which is crucial for improving the accuracy of human brain segmentation. Existing methods often do not pay attention to this part of the features. In the decoding stage, there is a lack of a module specifically for extracting low-level features of the brain surface, resulting in insufficient boundary segmentation accuracy, especially in low-contrast or inter-class similarity areas. The segmentation effect is unstable. In addition, large-model-based medical image segmentation methods such as SAM-Med3D and SegVol have demonstrated strong robustness and generalization capabilities. However, their effectiveness in extracting human brains remains unclear.

[0004] Figure 1 The inter-class similarity of voxel intensity and the challenges of complex anatomical structure of the head to the segmentation task are intuitively demonstrated. These two factors significantly affect the accuracy of brain tissue extraction. Figure 1 In (a), the voxel intensities of some non-brain tissue regions are similar to those of brain tissue. This inter-class similarity often interferes with existing methods and leads to inaccurate brain segmentation masks. Figure 1As shown in (b), the human head has a complex anatomical structure, especially in the brain surface area, which makes it particularly difficult to accurately distinguish brain tissue from surrounding non-brain tissue.

[0005] Current human brain extraction methods based on U-Net and its variants mainly rely on local convolution operations and lack the ability to model the global anatomical structure of the brain in the long term, resulting in decreased segmentation accuracy for complex anatomical regions. In addition, existing methods do not effectively extract low-level physical features of the brain surface (such as edges and textures) during the decoding stage, making it difficult to accurately distinguish between different anatomical regions with similar voxel intensities, resulting in blurred or erroneous segmentation boundaries. At the same time, in recent years, although the basic medical image segmentation models based on large-scale pre-training have strong generalization capabilities, they have not designed dedicated modules for brain extraction tasks, are insufficiently sensitive to the details of the human brain boundaries, and are difficult to adapt to the needs of fine segmentation of complex brain anatomical structures. Summary of the Invention

[0006] In response to the above-mentioned problems existing in the current technology, the present invention provides a human brain extraction method, system, electronic device, and storage medium for enhancing the perception of anatomical structure and surface information. The present invention solves the shortcomings of existing methods in global anatomical structure modeling and human brain boundary detail processing, and improves the accuracy and robustness of human brain extraction.

[0007] The technical solution of the present invention is: a human brain extraction method for enhancing anatomical structure and surface information perception, the method comprising:

[0008] Step 1: Construct the encoder's anatomical structure perception module AA. Feed the input 3D head MRI image into the encoder. By setting up multiple anatomical structure perception modules AA, anatomical structure perception feature maps of different resolutions are extracted layer by layer. The anatomical structure perception feature maps are then fed into two recursive multi-layer perceptron (RMLP) modules to capture long-range dependencies.

[0009] Step 2: Build the decoder's surface information perception module SA. Two additional RMLP modules are used in the decoder to further enhance the decoding of long-range dependencies. Through skip connections, the low-level features of different encoder layers in Step 1 are used to obtain their human brain surface features using a 3D Sobel surface feature detector. These features are then soft-fused with the high-level semantic information upsampled by the decoder to ensure alignment of surface cues with high-level semantics, enhancing surface information perception.

[0010] Step 3: Construct the local-global feature enhancement module LGFE of the decoder. Use the channel attention branch and spatial attention branch to extract global anatomical features and local surface features from the soft-fused features obtained in Step 2, and fuse the extracted global anatomical features and local surface features to generate enhanced feature expressions.

[0011] Step 4: The output of each layer of the decoder will be upsampled and restored to the same size as the original input. Then a 3D convolution is used to further fuse these upsampled features and finally generate the predicted brain tissue mask.

[0012] Furthermore, in Step 1, the encoder uses three stacked anatomical structure perception modules AA, all of which have the same structure. Each anatomical structure perception module AA has a dual-branch structure, including a multi-layer perceptron (MLP) branch and a convolutional neural network (CNN) branch.

[0013] The multi-layer perceptron (MLP) branch is used to capture the global anatomical structure information of the human brain and enhance the understanding of the overall anatomical structure of the human brain;

[0014] The convolutional neural network (CNN) branch is responsible for extracting local anatomical structure information of the human brain;

[0015] The multi-layer perceptron (MLP) branch includes a spatial shift module that shifts relevant input features along three axes: height (H), width (W), and depth (D). After each shift, the linear layer captures long-range dependencies and preserves the global anatomical structure of the human brain.

[0016] The convolutional neural network (CNN) branch includes a 3D convolution block, a batch normalization layer, a maximum pooling layer, and a ReLU activation function. These modules work together to extract local anatomical structure information of the human brain.

[0017] Furthermore, the Step 1 includes:

[0018] Step 1.1, the convolutional neural network (CNN) branch and the multi-layer perceptron (MLP) branch of the lth layer generate features respectively and

[0019]

[0020] Where I is the given input head voxel block, I∈R W×H×D×1 , where W, H and D represent width, height and depth, and the number of channels is 1; is the anatomical structure perception feature map output by the l-1 layer of the anatomical structure perception module AA, The local anatomical structure information of the human brain output by the convolutional neural network CNN branch, The global anatomical structure information of the human brain output by the multi-layer perceptron MLP branch;

[0021] Step 1.2: Use the output from the MLP branch As the output of the convolutional neural network CNN branch The residual connection is then applied to Used to enable the anatomical structure perception module AA to fully extract the anatomical structure perception feature map Anatomical structure perception feature map The extraction process is expressed as:

[0022]

[0023] in, Represents element-wise multiplication.

[0024] Furthermore, in Step 2, the surface information perception module SA is integrated into the decoder of ASNet, which includes three submodules: a surface extraction submodule SE, a semantic perception submodule SP, and a soft fusion submodule SF;

[0025] The surface extraction submodule SE is used to extract the surface features of the human brain. The surface extraction submodule uses a 3D Sobel surface feature detector as a core component;

[0026] The semantic perception submodule SP is used to extract high-level semantic information;

[0027] The soft fusion submodule SF is used to softly fuse human brain surface features with high-level semantic information using cross attention.

[0028] Furthermore, the Step 2 includes:

[0029] Step 2.1: Surface extraction submodule SE processes low-level features Low-level features It is obtained from the (l-1)th encoder through skip connection; the 3D Sobel surface feature detector uses different kernel weights along three axes, namely axial, coronal and sagittal axes to classify low-level features. Perform convolution; then combine the convolved features and apply them as attention weights to The calculation process is shown as follows;

[0030]

[0031] Among them, ReLU(.) represents ReLU activation function processing, BN(.) represents batch normalization processing, Conv(.) represents convolution processing, Sobel(.) represents 3D Sobel surface feature detector processing, Represents the human brain surface features output by the surface extraction submodule SE;

[0032] Low-level features It is obtained from the (l-1)th encoder layer through a skip connection. This sentence means that the output features of the corresponding layer in the encoder are needed in the decoder. Here, the corresponding features are directly obtained through the skip connection to the decoder. The reason why it is l-1 is that the decoder needs to upsample. The features after upsampling of the l-th layer decoder are consistent with the resolution of the l-1 layer features in the encoder.

[0033] Step 2.2: Use the semantic perception submodule SP to obtain the enhanced feature expression from the local-global feature enhancement module LGFE Extract high-level semantic information from the decoder, which is generated by the lth decoder layer;

[0034]

[0035] in, Represents the high-level semantic information output by the semantic perception submodule SP, Represents the enhanced feature expression obtained by the local-global feature enhancement module LGFE;

[0036] Step 2.3, use cross attention and soft fusion submodule SF to With F l sp Soft Fusion:

[0037]

[0038] Where σ(·) represents the Sigmoid function, The surface information perception module SA from the l-th decoder layer represents the soft-fused features.

[0039] Furthermore, in Step 3, the local-global feature enhancement module LGFE includes a channel attention branch CHA, a spatial attention branch SPA, and a skip connection branch:

[0040] The channel attention branch CHA is used to extract global anatomical features and enhance the global anatomical structure information of the human brain;

[0041] The spatial attention branch SPA is used to capture local surface features and refine boundary information;

[0042] The skip connection branch is used to perform cross-scale feature fusion, so that local and global information complement each other.

[0043] Furthermore, the Step 3 includes:

[0044] Step 3.1: Features after soft fusion in Step 2 Channel attention branch CHA uses channel attention to capture Implicit remote dependency information:

[0045]

[0046] Among them, CHA(·) represents the channel attention branch, Represents global anatomical features;

[0047] Step 3.2, spatial attention branch SPA uses spatial attention to capture The local dependency information contained in:

[0048]

[0049] Among them, SPA(·) represents the spatial attention branch, F l spa Represents local surface features;

[0050] Step 3.3: Use element-wise multiplication to fuse global anatomical features and local surface features F l spa ; Finally, through the residual connection, the fused features are combined with Add to generate enhanced feature expression

[0051]

[0052] The present invention also provides a human brain extraction system for enhancing anatomical structure and surface information perception, the system comprising: a module for executing the human brain extraction method for enhancing anatomical structure and surface information perception.

[0053] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for human brain extraction that enhances perception of anatomical structure and surface information is implemented.

[0054] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the human brain extraction method for enhancing the perception of anatomical structure and surface information.

[0055] CNN in the present invention stands for Convolutional Neural Network, which is a specially designed deep learning model.

[0056] The MLP of the present invention stands for Multilayer Perceptron, which is a feedforward artificial neural network model.

[0057] The present invention proposes spatial attention representation: a mechanism in deep learning that focuses on specific areas or content in the input data.

[0058] Channel Attention Representation: A mechanism in deep learning that emphasizes the importance of different channels in the input feature map.

[0059] This paper proposes a human brain extraction system (ASNet) that enhances the perception of anatomical structure and surface information. The system comprises an anatomical perception (AA) module, a surface perception (SA) module, and a local-global feature enhancement (LGFE) module. These modules capture the anatomical structure and surface information of the human brain and effectively represent local and global features. The AA module has a two-branch structure: a multi-layer perceptron (MLP) branch for characterizing anatomical structure and a convolutional neural network (CNN) branch for extracting low-level features. The MLP branch includes a spatial translation module that shifts input features along three axes. After each shift, a linear layer captures long-range correlations and preserves global anatomical structure information. The convolutional branch comprises a 3D convolution block, a batch normalization layer, a max pooling layer, and a ReLU activation function. The surface perception (SA) module is integrated into the decoder of ASNet and consists of three submodules: a surface extraction submodule (SE), a semantic perception submodule (SP), and a soft fusion submodule (SF). The surface extraction submodule (SE) is responsible for extracting surface features of the human brain. To this end, the present invention introduces a 3D Sobel surface detector as a core component. The local-global feature enhancement (LGFE) module consists of three branches: a channel attention branch (CHA) for global feature extraction, a spatial attention branch (SPA) for local feature extraction, and a skip connection branch for feature fusion. The global feature extraction branch uses channel attention to capture long-range dependencies, while the local feature extraction branch uses spatial attention to capture local dependencies.

[0060] The beneficial effects of the present invention are:

[0061] 1. Based on the U-shaped network structure, this paper introduces the anatomical awareness (AA) module, the surface awareness (SA) module, and the local-global feature enhancement (LGFE) module to improve the accuracy and robustness of human brain extraction. The AA module adopts a dual-branch structure, capturing global anatomical information through the MLP branch, enhancing the model's representation of the overall brain structure. At the same time, the CNN branch extracts local anatomical details, retaining detailed anatomical structure information and improving the segmentation ability of complex areas.

[0062] 2. The SA module of the present invention consists of a surface information extraction layer, a semantic perception layer, and a soft fusion layer. The surface extraction layer is used to extract surface features of the human brain; the semantic perception layer is used to extract high-level semantic information; and the soft fusion layer integrates low-level edge information with high-level semantic features through an attention mechanism, thereby enhancing the model's ability to perceive boundary areas. The use of this module makes up for the shortcomings of existing methods in extracting low-level physical features of the brain surface. In order to further improve the model's feature expression capabilities;

[0063] 3. The LGFE module of the present invention adopts a three-branch structure to enhance the network's fusion of local and global features. Specifically, the channel attention branch is responsible for extracting global anatomical features, the spatial attention branch focuses on capturing local surface details, and the skip connection branch is used for cross-scale feature fusion. This allows local and global information to complement each other, improving the model's ability to accurately depict brain tissue boundaries and thus optimizing segmentation results.

[0064] 4. The method of the present invention enhances the perception of the human brain's anatomical structure and surface information to achieve accurate brain extraction. Specifically, the anatomical structure perception module aims to enhance the understanding of global brain organization, while the surface information perception module emphasizes local patterns on the brain surface to address the challenges posed by the low-level inter-class similarity of voxel intensities between brain and non-brain tissue. In addition, a novel local-global feature enhancement module is used to enhance the association between local surface patterns and global brain organization, providing a robust local-global representation that facilitates accurate brain extraction.

[0065] 5. The present invention has conducted extensive experiments on four public datasets, covering different imaging modes, imaging device parameters, and patient groups. The experimental results show that the ASNet proposed in the present invention outperforms the current state-of-the-art methods in terms of accuracy and robustness. Compared with existing methods, ASNet enhances the global anatomical structure modeling capability through the AA module, refines the human brain boundary segmentation through the SA module, and balances the fusion of local and global features through the LGFE module, achieving more accurate human brain extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 Schematic diagram of two challenges faced in inventing the brain extraction process; Figure 1 (a) is a statistical diagram of the intensity distribution of brain tissue and non-brain tissue in a local area of the present invention; Figure 1 (b) Schematic diagram of brain structures in different axes;

[0067] Figure 2 It is the structural diagram of the ASNet network in the present invention;

[0068] Figure 3 : is a structural diagram of the anatomical structure perception module (AA) in the present invention;

[0069] Figure 4 It is a structural diagram of the surface information perception module (SA) in the present invention; wherein, Figure 4 (a) is a structural diagram of the surface extraction submodule SE, semantic perception submodule SP and soft fusion submodule SF in the surface information perception module SA of the present invention; Figure 4 (b) is a structural diagram of the 3D Sobel surface feature detector in the surface information perception module of the present invention;

[0070] Figure 5 This is a structural diagram of the local-global feature enhancement module (LGFE module) in the present invention. DETAILED DESCRIPTION

[0071] Example 1: Figure 1-Figure 5 A method for human brain extraction that enhances perception of anatomical structure and surface information is shown, the method comprising:

[0072] Step 1: Construct the encoder's anatomical structure perception module AA. Feed the input 3D head MRI image into the encoder. By setting up multiple anatomical structure perception modules AA, anatomical structure perception feature maps of different resolutions are extracted layer by layer. The anatomical structure perception feature maps are then fed into two recursive multi-layer perceptron (RMLP) modules to capture long-range dependencies.

[0073] Furthermore, in Step 1, the anatomical awareness (AA) module is integrated into the shallow encoding stage to enhance the network's ability to represent human brain structure information, such as Figure 3 As shown; Figure 2 In the embodiment, the encoder adopts three stacked anatomical structure perception modules AA, which have the same structure. Each anatomical structure perception module AA has a dual-branch structure, including a multi-layer perceptron MLP branch and a convolutional neural network CNN branch.

[0074] The multi-layer perceptron (MLP) branch is used to capture the global anatomical structure information of the human brain and enhance the understanding of the overall anatomical structure of the human brain;

[0075] The convolutional neural network (CNN) branch is responsible for extracting local anatomical structure information of the human brain;

[0076] The multi-layer perceptron (MLP) branch includes a spatial shift module that shifts relevant input features along three axes: height (H), width (W), and depth (D). After each shift, the linear layer captures long-range dependencies and preserves the global anatomical structure of the human brain.

[0077] The convolutional neural network (CNN) branch includes a 3D convolution block, a batch normalization layer, a maximum pooling layer, and a ReLU activation function. These modules work together to extract local anatomical structure information of the human brain.

[0078] Furthermore, the Step 1 includes:

[0079] Step 1.1, the convolutional neural network (CNN) branch and the multi-layer perceptron (MLP) branch of the lth layer generate features respectively and

[0080]

[0081] Where I is the given input head voxel block, I∈R W×H×D×1 , where W, H and D represent width, height and depth, and the number of channels is 1; is the anatomical structure perception feature map output by the l-1 layer of the anatomical structure perception module AA, The local anatomical structure information of the human brain output by the convolutional neural network CNN branch, The global anatomical structure information of the human brain output by the multi-layer perceptron MLP branch;

[0082] Step 1.2: To help the AA module recognize anatomical structures and extract low-level features, the present invention uses the output from the multi-layer perceptron MLP branch. As the output of the convolutional neural network CNN branch The residual connection is then applied to Used to enable the anatomical structure perception module AA to fully extract the anatomical structure perception feature map Anatomical structure perception feature map The extraction process is expressed as:

[0083]

[0084] in, Represents element-wise multiplication.

[0085] Step 2: Build the decoder's surface information perception module SA. Two additional RMLP modules are used in the decoder to further enhance the decoding of long-range dependencies. Through skip connections, the low-level features of different encoder layers in Step 1 are used to obtain their human brain surface features using a 3D Sobel surface feature detector. These features are then soft-fused with the high-level semantic information upsampled by the decoder to ensure alignment of surface cues with high-level semantics, enhancing surface information perception.

[0086] Furthermore, in Step 2, in order to enable ASNet to effectively represent the surface information of the human brain, the present invention proposes a surface information perception (SA) module. The surface information perception module SA is integrated into the decoder of ASNet, such as Figure 2As shown in Figure 2, integrating the SA module into the decoder of ASNet has two main purposes: (i) enhancing the representation of surface cues, and (ii) ensuring that surface cues are consistent with high-level semantics. The structure of the SA module is shown in Figure 2. Figure 4 Detailed description in Figure 4 As shown in (a), the surface information perception module SA includes three submodules: surface extraction submodule SE, semantic perception submodule SP and soft fusion submodule SF;

[0087] The surface extraction submodule SE is used to extract the surface features of the human brain. The surface extraction submodule uses a 3D Sobel surface feature detector as a core component;

[0088] The semantic perception submodule SP is used to extract high-level semantic information;

[0089] The soft fusion submodule SF is used to softly fuse human brain surface features with high-level semantic information using cross attention.

[0090] Furthermore, the Step 2 includes:

[0091] Step 2.1: Since the surface structure is a low-level physical feature, the surface extraction submodule SE processes the low-level features. Low-level features is obtained from the (l-1)th encoder through skip connection; Figure 4 The left side of (b) shows the structure of the 3D Sobel surface feature detector. Figure 4 (b) As shown on the right, the 3D Sobel surface feature detector uses different kernel weights to classify low-level features along three axes: axial, coronal, and sagittal. Perform convolution; then combine the convolved features and apply them as attention weights to To enhance its ability to represent brain surface features, the computational process is shown as follows;

[0092]

[0093] Among them, ReLU(.) represents ReLU activation function processing, BN(.) represents batch normalization processing, Conv(.) represents convolution processing, Sobel(.) represents 3D Sobel surface feature detector processing, Represents the human brain surface features output by the surface extraction submodule SE;

[0094] Low-level features It is obtained from the (l-1)th encoder layer through a skip connection. This sentence means that the output features of the corresponding layer in the encoder are needed in the decoder. Here, the corresponding features are directly obtained through the skip connection to the decoder. The reason why it is l-1 is that the decoder needs to upsample. The features after upsampling of the l-th layer decoder are consistent with the resolution of the l-1 layer features in the encoder.

[0095] Step 2.2: Use the semantic perception submodule SP to obtain the enhanced feature expression from the local-global feature enhancement module LGFE Extract high-level semantic information from the decoder, which is generated by the lth decoder layer;

[0096]

[0097] in, Represents the high-level semantic information output by the semantic perception submodule SP, Represents the enhanced feature expression obtained by the local-global feature enhancement module LGFE;

[0098] Step 2.3, due to From low-level features In order to reduce this noise and combine important high-level semantic information, the present invention uses cross attention and combines it with the soft fusion submodule SF. With F l sp Soft Fusion:

[0099]

[0100] Where σ(·) represents the Sigmoid function, The surface information perception module SA from the l-th decoder layer represents the soft-fused features.

[0101] Step 3: Construct the local-global feature enhancement module LGFE of the decoder. Use the channel attention branch and spatial attention branch to extract global anatomical features and local surface features from the soft-fused features obtained in Step 2, and fuse the extracted global anatomical features and local surface features to generate enhanced feature expressions.

[0102] Furthermore, in the decoding stage, the full integration of local surface features and global anatomical features is very important for generating high-precision brain segmentation masks. To this end, the present invention designs a local-global feature enhancement module (LGFE) and integrates it into each decoder layer;

[0103] In Step 3, Figure 5As shown, the local-global feature enhancement module LGFE includes a channel attention branch CHA, a spatial attention branch SPA and a skip connection branch:

[0104] The channel attention branch CHA is used to extract global anatomical features and enhance the global anatomical structure information of the human brain;

[0105] The spatial attention branch SPA is used to capture local surface features and refine boundary information;

[0106] The skip connection branch is used to perform cross-scale feature fusion, so that local and global information complement each other.

[0107] Furthermore, the Step 3 includes:

[0108] Step 3.1: Features after soft fusion in Step 2 Channel attention branch CHA uses channel attention to capture Implicit remote dependency information:

[0109]

[0110] Among them, CHA(·) represents the channel attention branch, Represents global anatomical features;

[0111] Step 3.2, spatial attention branch SPA uses spatial attention to capture The local dependency information contained in:

[0112]

[0113] Among them, SPA(·) represents the spatial attention branch, F l spa Represents local surface features;

[0114] Step 3.3: Use element-wise multiplication to fuse global anatomical features and local surface features F l spa ; Finally, through the residual connection, the fused features are combined with Add to generate enhanced feature expression

[0115]

[0116] Step 4: The output of each layer of the decoder will be upsampled and restored to the same size as the original input. Then a 3D convolution is used to further fuse these upsampled features and finally generate the predicted brain tissue mask.

[0117] The present invention also provides a human brain extraction system for enhancing anatomical structure and surface information perception, the system comprising: a module for executing the human brain extraction method for enhancing anatomical structure and surface information perception.

[0118] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for human brain extraction that enhances perception of anatomical structure and surface information is implemented.

[0119] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the human brain extraction method for enhancing the perception of anatomical structure and surface information.

[0120] Figure 2 This paper demonstrates the ASNet network structure proposed in this paper. The ASNet network includes an anatomical structure awareness (AA) module, a surface information awareness (SA) module, and a local-global feature enhancement (LGFE) module. This paper enhances the perception of global anatomical structures and local surface patterns, thereby improving segmentation accuracy and robustness.

[0121] In the encoder part, the ASNet network uses three stacked anatomical structure perception (AA) modules. The anatomical structure perception (AA) module structure is as follows: Figure 3 As shown in the figure, the multi-layer perceptron (MLP) branch is responsible for modeling long-range dependencies, capturing global anatomical information, and enhancing the model's understanding of the brain's overall structure. The convolution branch is responsible for extracting local anatomical details. By combining the MLP and convolution branches, the AA module can effectively improve the model's perception of the brain's anatomical structure. In the subsequent part of the encoder, two recursive MLP (RMLP) modules are introduced to further enhance the network's ability to model long-range dependencies. This enables the model to fully integrate local and global information at different scales, providing more discriminative feature representations for the subsequent decoding stage.

[0122] In the decoder part, ASNet adopts a five-layer decoding structure symmetrical to the encoder, gradually recovering the features extracted during the encoding process and enhancing the boundary perception ability. In order to solve the problems of blurred brain surface boundaries and insufficient extraction of local anatomical patterns, the decoder introduces a surface information perception module (SA). Figure 4 As shown in the figure, the module uses a 3D Sobe filter to enhance the perception of brain surface patterns, enabling it to more accurately distinguish different anatomical tissues with similar voxel intensities, thereby optimizing boundary segmentation accuracy.

[0123] like Figure 5As shown in the figure, the local-global feature enhancement (LGFE) module adopts a three-branch structure: the channel attention branch (CHA) is responsible for extracting global features and enhancing anatomical structure information; the spatial attention branch (SPA) is responsible for capturing local surface features and refining boundary information; and the skip connection branch is responsible for fusing features of different scales.

[0124] Finally, the output of each layer of the decoder is upsampled and restored to the same size as the original input (W, H, D, 4), where 4 is the color channel; then a 3D convolution is used to further fuse these features (referring to the features of each layer output of the decoder after being upsampled, the features of different layers are spliced along the channel dimension, and then a 3D convolution is used to fuse them), and finally generate the predicted brain tissue mask Two of the channels correspond to the predicted probability values of brain tissue and non-brain tissue respectively. Here, two channels mean that the network output has two predicted values for each pixel. The first channel represents the probability of background prediction, that is, the probability of non-brain tissue prediction, and the second channel represents the probability of foreground prediction, that is, the probability of brain tissue prediction. The probabilities of non-brain tissue and brain tissue correspond to these two channels. By comparing the probability values of these two channels, it is determined whether the pixel belongs to brain tissue or non-brain tissue, thereby generating the brain tissue segmentation map. Usually, for a two-classification problem, the network output corresponds to 2 channels. If it is a three-classification problem, the network output corresponds to 3 channels.

[0125] exist Figure 2 The flowchart drawn in the figure is the overall process of the model in the inference stage. Due to the large resolution of 3D brain tissue, it is impossible to directly send the high-resolution image to the network for training. Therefore, in the training stage, a block is cropped for training. In the inference stage, predictions cannot be made according to the cropping mode. It is necessary to process a high-resolution image into multiple voxel blocks of fixed size in the form of a sliding window, and send them to the network for prediction respectively. After all the voxel blocks are predicted, they are spliced together to obtain a complete high-resolution image. Splicing is not required in the model training stage. The resulting image is the cropped brain tissue mask. In order to obtain a complete brain mask without cropping during the testing phase, a sliding window is used to crop a complete brain into multiple voxel blocks of size (W, H, D), and then these brain masks are spliced together to obtain a complete brain mask.

[0126] Specifically, given an input head voxel block I∈R W×H×D×1, where W, H, and D represent width, height, and depth, and the number of channels is 1. During inference, the entire head 3D image is evenly divided into multiple 128×128×128 image blocks, with an overlap step of 48 between these blocks. In the encoding stage, the output feature dimensions of each layer of the encoder are and Following the standard U-Net design, while the spatial size is halved to reduce the amount of computation, the number of channels is doubled (for example, from 16 to 32 after the first layer), ensuring that the model can extract richer hierarchical information.

[0127] In the decoding stage, the output size of each decoder layer matches the corresponding encoder layer. The output of all layers is upsampled to (W,H,D,4). Finally, a 3D convolutional layer is used to merge all upsampled results to generate the human brain segmentation mask.

[0128] Through three modules—anatomy awareness (AA), surface awareness (SA), and local-global feature enhancement (LGFE)—ASNet emphasizes the representation of anatomical and surface information, thereby generating more accurate brain segmentation masks. To validate the effectiveness of our approach, we conducted experiments on four challenging datasets, comparing our ASNet approach with other mainstream brain extraction methods. These datasets are NFGS, SynthStrip, OASIS-1, and CC-359. MSD, RMSD, HD95, and HD99 were used as performance metrics. MSD represents mean surface distance, RMSD represents root mean square surface distance, HD95 represents 95% Hausdorff distance, and HD99 represents 99% Hausdorff distance. The experimental results are shown in Tables 1, 2, 3, and 4. ASNet outperforms other baseline models on all four datasets. The results demonstrate that across all datasets, ASNet generates superior surface contours and overall segmentation masks to existing baseline models.

[0129] Table 1 shows the quantitative comparison results on the NFBS dataset.

[0130]

[0131] As can be seen from Table 1, the method of the present invention has the best performance in terms of MSD, RMSD, HD95 and HD99 indicators on the NFBS dataset; the baseline model MedSAM has the second best performance in terms of MSD, RMSD, HD95 and HD99 indicators on the NFBS dataset.

[0132] Table 2 shows the quantitative comparison results on the SynthStrip dataset

[0133]

[0134]

[0135] As can be seen from Table 2, the performance of the MSD, RMSD, HD95 and HD99 indicators of the proposed method on the SynthStrip dataset is the best; the baseline model UNETR++ has the second best performance in the MSD indicator on the SynthStrip dataset; and the baseline model MA-SAM has the second best performance in the RMSD, HD95 and HD99 indicators on the SynthStrip dataset.

[0136] Table 3 shows the quantitative comparison results on the OASIS-1 dataset

[0137]

[0138] As can be seen from Table 3, the proposed method ASNet has the best performance in terms of MSD, RMSD, HD95, and HD99 indicators on the OASIS-1 dataset; the baseline model SAM-Med3D has the second best performance in terms of MSD and RMSD indicators on the OASIS-1 dataset; the baseline model MSMHA-CNN has the second best performance in terms of HD95 indicators on the OASIS-1 dataset, and the baseline model 3D UNeXt has the second best performance in terms of HD99 indicators on the OASIS-1 dataset.

[0139] Table 4 shows the quantitative comparison results on the CC-359 dataset.

[0140]

[0141]

[0142] As can be seen from Table 4, the proposed method ASNet has the best performance in terms of MSD, RMSD, HD95, and HD99 indicators on the CC-359 dataset; the baseline model MedSAM has the second best performance in terms of MSD, RMSD, and HD99 indicators on the CC-359 dataset; and the baseline model 3D UNeXt has the second best performance in terms of HD95 indicators on the CC-359 dataset.

[0143] Table 5 shows the comparison of ASNet and other baseline models in terms of computational complexity.

[0144]

[0145] Table 5 compares the computational complexity of ASNet with other baseline models. TM, IM, and IT represent training memory, inference memory, and inference time, respectively. The proposed method performs best compared to other baseline models in terms of training memory, inference memory, and inference time.

[0146] Table 5 compares the performance of the proposed method ASNet with the baseline model in terms of computing resource consumption and average inference speed. The results show that the proposed method ASNet significantly outperforms existing methods in terms of computational efficiency. Specifically, the proposed method ASNet has the lowest number of parameters, only 7.73M, and achieves the lowest computational complexity, only 28.02GFLOPs. This result is significantly better than other models: for example, SegVol (716.66GFLOPs) and MedSAM (371.99GFLOPs). In addition, the proposed method ASNet requires less training memory, only 2,404M, and inference memory is only 1,106M, which is much lower than other baseline models, such as MedSAM (10,230M and 3,574M, respectively) and SegVol (11,146M and 3,700M, respectively). It is worth noting that the inference time of the proposed method ASNet for each 3D head image of size 128×128×128 is only 0.21 seconds, which is better than SAM-Med3D (0.28 seconds) and far exceeds MedSAM (10.17 seconds) and SegVol (17.92 seconds).

[0147] The specific embodiments of the present invention are described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Various changes can be made within the knowledge of ordinary technicians in this field without departing from the scope of the present invention.

Claims

1. A method for human brain extraction to enhance the perception of anatomical structure and surface information, characterized by: The method comprises: Step 1: Construct the encoder's anatomical structure perception module AA. Feed the input 3D head MRI image into the encoder. By setting up multiple anatomical structure perception modules AA, anatomical structure perception feature maps of different resolutions are extracted layer by layer. The anatomical structure perception feature maps are then fed into two recursive multi-layer perceptron (RMLP) modules to capture long-range dependencies. Step 2: Build the decoder's surface information perception module SA. Use two additional RMLP modules in the decoder to further enhance the decoding of long-range dependencies. Through skip connections, use the 3DSobel surface feature detector to obtain the human brain surface features of the low-level features of different layers of the encoder in Step 1. Then, soft-fuse them with the high-level semantic information after upsampling in the decoder. This ensures the alignment of surface cues with high-level semantics and enhances surface information perception capabilities. Step 3: Construct the local-global feature enhancement module LGFE of the decoder. Use the channel attention branch and spatial attention branch to extract global anatomical features and local surface features from the soft-fused features obtained in Step 2, and fuse the extracted global anatomical features and local surface features to generate enhanced feature expressions. Step 4: The output of each layer of the decoder will be upsampled and restored to the same size as the original input. Then a 3D convolution is used to further fuse these upsampled features and finally generate the predicted brain tissue mask.

2. The method for human brain extraction for enhancing anatomical structure and surface information perception according to claim 1, characterized in that: In Step 1, the encoder uses three stacked anatomical structure perception modules AA, all of which have the same structure. Each anatomical structure perception module AA has a dual-branch structure, including a multi-layer perceptron (MLP) branch and a convolutional neural network (CNN) branch. The multi-layer perceptron (MLP) branch is used to capture the global anatomical structure information of the human brain and enhance the understanding of the overall anatomical structure of the human brain; The convolutional neural network (CNN) branch is responsible for extracting local anatomical structure information of the human brain; The multi-layer perceptron MLP branch includes a spatial movement module, which moves the relevant input features along three axes, namely, height H axis, width W axis and depth D axis; After each shift, the linear layer captures long-range correlations and preserves the global anatomical structure of the human brain; The convolutional neural network (CNN) branch includes a 3D convolution block, a batch normalization layer, a maximum pooling layer, and a ReLU activation function. These modules work together to extract local anatomical structure information of the human brain.

3. The method for human brain extraction for enhancing anatomical structure and surface information perception according to claim 2, characterized in that: Step 1 includes: Step 1.1, the convolutional neural network (CNN) branch and the multi-layer perceptron (MLP) branch of the lth layer generate features respectively and Where I is the given input head voxel block, I∈R W×H×D×1 , where W, H and D represent width, height and depth, and the number of channels is 1; is the anatomical structure perception feature map output by the l-1 layer of the anatomical structure perception module AA, The local anatomical structure information of the human brain output by the convolutional neural network CNN branch, The global anatomical structure information of the human brain output by the multi-layer perceptron MLP branch; Step 1.2: Use the output from the MLP branch As the output of the convolutional neural network CNN branch The residual connection is then applied to Used to enable the anatomical structure perception module AA to fully extract the anatomical structure perception feature map Anatomical structure perception feature map The extraction process is expressed as: in, Represents element-wise multiplication.

4. The method for human brain extraction for enhancing anatomical structure and surface information perception according to claim 1, characterized in that: In Step 2, the surface information perception module SA is integrated into the decoder of ASNet, which includes three submodules: surface extraction submodule SE, semantic perception submodule SP and soft fusion submodule SF; The surface extraction submodule SE is used to extract the surface features of the human brain. The surface extraction submodule uses a 3D Sobel surface feature detector as a core component; The semantic perception submodule SP is used to extract high-level semantic information; The soft fusion submodule SF is used to softly fuse human brain surface features with high-level semantic information using cross attention.

5. The method for human brain extraction for enhancing anatomical structure and surface information perception according to claim 4, characterized in that: Step 2 includes: Step 2.1: Surface extraction submodule SE processes low-level features Low-level features It is obtained from the (l-1)th encoder through skip connection; the 3D Sobel surface feature detector uses different kernel weights along three axes, namely axial, coronal and sagittal axes to classify low-level features. Perform convolution; then combine the convolved features and apply them as attention weights to The calculation process is shown as follows; Among them, ReLU(.) represents ReLU activation function processing, BN(.) represents batch normalization processing, Conv(.) represents convolution processing, Sobel(.) represents 3D Sobel surface feature detector processing, Represents the human brain surface features output by the surface extraction submodule SE; Low-level features It is obtained from the (l-1)th encoder layer through a skip connection. This sentence means that the output features of the corresponding layer in the encoder are needed in the decoder. Here, the corresponding features are directly obtained through the skip connection to the decoder. The reason why it is l-1 is that the decoder needs to upsample. The features after upsampling of the l-th layer decoder are consistent with the resolution of the l-1 layer features in the encoder. Step 2.2: Use the semantic perception submodule SP to obtain the enhanced feature expression from the local-global feature enhancement module LGFE Extract high-level semantic information from the decoder, which is generated by the lth decoder layer; in, Represents the high-level semantic information output by the semantic perception submodule SP, Represents the enhanced feature expression obtained by the local-global feature enhancement module LGFE; Step 2.3, use cross attention and soft fusion submodule SF to With F l sp Soft Fusion: Among them, σ(·) represents the Sigmoid function, The surface information perception module SA from the l-th decoder layer represents the soft-fused features.

6. The method for human brain extraction for enhancing anatomical structure and surface information perception according to claim 1, characterized in that: In Step 3, the local-global feature enhancement module LGFE includes a channel attention branch CHA, a spatial attention branch SPA, and a skip connection branch: The channel attention branch CHA is used to extract global anatomical features and enhance the global anatomical structure information of the human brain; The spatial attention branch SPA is used to capture local surface features and refine boundary information; The skip connection branch is used to perform cross-scale feature fusion, so that local and global information complement each other.

7. The method for human brain extraction for enhancing anatomical structure and surface information perception according to claim 1, characterized in that: Step 3 includes: Step 3.1: Features after soft fusion in Step 2 Channel attention branch CHA uses channel attention to capture Implicit remote dependency information: Among them, CHA(·) represents the channel attention branch, Represents global anatomical features; Step 3.2, spatial attention branch SPA uses spatial attention to capture The local dependency information contained in: Among them, SPA(·) represents the spatial attention branch, F l spa Represents local surface features; Step 3.3: Use element-wise multiplication to fuse global anatomical features and local surface features F l spa ; Finally, through the residual connection, the fused features are combined with Add to generate enhanced feature expression 8. A human brain extraction system for enhancing anatomical structure and surface information perception, characterized in that: The system comprises: a module for executing a human brain extraction method for enhancing anatomical structure and surface information perception according to any one of claims 1 to 7.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for human brain extraction that enhances perception of anatomical structure and surface information as described in any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for human brain extraction that enhances perception of anatomical structure and surface information is implemented as claimed in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Medical image segmentation method based on CNN and Transform fusion network

    CN117173412A

  • Three-dimensional brain tumor segmentation model based on deformable feature aggregation

    CN118967712A

  • Intelligent pulmonary nodule grading method and system based on multi-modality feature fusion

    WO2025020719A1

Cited By

  • Medical imaging analysis methods, devices, and computer equipment for brain organoids

    CN122573948A