Prostate MRI image segmentation method based on dynamic multi-branch convolutional network

Through the dynamic multi-branch convolutional network method, the problems of insufficient accuracy and high computational complexity in fuzzy areas in prostate MRI image segmentation are solved, efficient multimodal image information fusion and anatomical constraint collaboration are achieved, the accuracy and real-time performance of prostate segmentation are improved, and it is suitable for complex cases, providing a highly robust solution for clinical diagnosis and treatment.

CN120635111APending Publication Date: 2025-09-12CHONGQING UNIV OF POSTS & TELECOMM
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510716813.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing prostate MRI image segmentation methods have problems such as insufficient fuzzy area segmentation accuracy, high computational complexity leading to poor real-time performance, insufficient fusion of multimodal image information, and insufficient synergy between anatomical constraints and deep learning models.

Method used

A method based on a dynamic multi-branch convolutional network is adopted to enhance the multi-scale feature reuse capability by embedding ResNeXt Block and cross-level residual connections in the encoder; spatial attention and channel attention modules are introduced in the decoder to dynamically fuse multimodal data features; independent prostate apex, middle, and base segmentation sub-networks are constructed to optimize the anatomical characteristics of different regions; and a cross-modal interaction module is designed to align the feature maps of T2WI and DWI sequences to address resolution differences.

Benefits of technology

It significantly improves the segmentation accuracy and computational efficiency of prostate multimodal images, enhances the adaptability and robustness of the model, and provides high-precision clinical diagnosis and treatment support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635111A_ABST
    Figure CN120635111A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of medical image processing, and particularly relates to a prostate MRI image segmentation method based on a dynamic multi-branch convolutional network, a segmentation model is constructed based on a 3D U-Net network structure, and the segmentation model comprises an encoder module, a bottleneck layer, a decoder module and a segmentation module; training the segmentation model, and adopting the trained segmentation model to realize the segmentation of the prostate MRI image; according to the method, through an innovative dynamically optimized hybrid network architecture and an adaptive fusion strategy, the segmentation precision and the calculation efficiency of the prostate multi-modal image are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image processing, and in particular relates to a prostate MRI image segmentation method based on a dynamic multi-branch convolutional network. Background Art

[0002] Currently, prostate segmentation models based on 3D UNets generally suffer from insufficient reuse of shallow features, resulting in limited segmentation accuracy in regions with blurred boundaries, such as the prostate apex. While 3D extensions of dense computational architectures (such as DenseNet) enhance feature transfer, they struggle to meet intraoperative real-time requirements due to the dramatic increase in parameter count. Existing multimodal fusion methods (such as T2WI / DWI joint segmentation) often employ simple channel concatenation or early fusion strategies, lack cross-modal attention mechanisms to exploit inter-sequence complementarity, and lack dynamic channel compression modules to adapt to resolution differences across different devices. Some studies have introduced general dynamic pruning techniques to reduce computational complexity, but their threshold setting relies on manual experience, which can easily lead to loss of key features in medical segmentation tasks. Furthermore, while anatomical constraint methods (such as ellipsoid shape priors) optimize boundary smoothness, they are not designed in conjunction with hybrid network architectures, resulting in insufficiently refined segmentation of localized regions (such as the prostate apex). Summary of the Invention

[0003] In order to solve the problems existing in existing prostate MRI image segmentation methods, such as insufficient fuzzy region segmentation accuracy, poor real-time performance due to high computational complexity, insufficient fusion of multimodal image information, and insufficient synergy between anatomical constraints and deep learning models, the present invention proposes a prostate MRI image segmentation method based on a dynamic multi-branch convolutional network. A segmentation model is constructed based on a 3D U-Net network structure, and the segmentation model includes an encoder module, a bottleneck layer, a decoder module, and a segmentation module. The segmentation model is trained, and the trained segmentation model is used to implement prostate MRI image segmentation. The training process of the segmentation model includes:

[0004] S1. Obtaining a prostate multimodal image dataset and preprocessing it to obtain a training set; the training set includes multiple groups of sample images;

[0005] S2. Input the sample image into the encoder module to obtain an encoded image; the encoder module includes multiple improved encoders of different size levels, and a maximum pooling layer is provided between each two adjacent improved encoders;

[0006] S3. Input the encoded image into the bottleneck layer to obtain an intermediate image; the bottleneck layer includes two convolutional layers; a maximum pooling layer is provided between the bottleneck layer and the last improved encoder, and an upsampling layer is provided between the bottleneck layer and the first improved decoder;

[0007] S4. Inputting the intermediate image into the decoder module to obtain a decoded image; the decoder module includes multiple improved decoders of different size levels, with an upsampling layer between each two adjacent improved decoders; and residual connections exist between improved encoders and improved decoders of the same size level;

[0008] S5. Input the decoded image into the segmentation module to obtain the segmentation result, and calculate the loss training model parameters according to the segmentation result until the model parameters converge.

[0009] Beneficial effects of the present invention:

[0010] ResNeXt Block is embedded in the encoder to enhance the multi-scale feature reuse capability through group convolution and cross-layer residual connection.

[0011] The decoder incorporates a spatial attention module (focusing on the spatial distribution of lesions) and a channel attention module (selecting key feature channels) to dynamically fuse the encoder's multi-branch outputs. Through high-level semantic guidance, background noise interference is suppressed, enhancing the complementary utilization of multimodal data (such as T2WI / DWI).

[0012] Based on the lightweight SE Block (Squeeze-and-Excitation Block), channel weights are generated to evaluate the importance of each channel to the segmentation task. By setting a dynamic threshold, the calculation path of low-weight channels is compressed to reduce the number of parameters.

[0013] Independent prostate apex, middle, and base segmentation sub-networks are constructed, each optimized for the anatomical characteristics of the different regions. Based on the distribution probability of lesions in three-dimensional space (e.g., the apex is prone to small lesions), the output weights of each sub-network are dynamically adjusted to prioritize segmentation accuracy in high-risk areas.

[0014] A cross-modal interaction module is designed at the encoder input to dynamically align the feature maps of T2WI and DWI sequences through attention weights to resolve the fusion bias caused by resolution differences.

[0015] Through an innovative dynamically optimized hybrid network architecture and adaptive fusion strategy, the segmentation accuracy and computational efficiency of prostate multimodal images are significantly improved, while reducing model complexity and enhancing adaptability to complex cases, providing a highly robust solution for clinical diagnosis and treatment. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is a structural diagram of the model of the present invention;

[0017] Figure 2 This is a schematic diagram of the improved encoder structure of the present invention;

[0018] Figure 3This is a schematic diagram of the improved decoder structure of the present invention. DETAILED DESCRIPTION

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0020] The present invention provides a prostate MRI image segmentation method based on a dynamic multi-branch convolutional network, including constructing and training a segmentation model, and using the trained segmentation model to implement prostate MRI image segmentation. The present invention constructs a segmentation model based on a 3D U-Net network structure, such as Figure 1 As shown, the segmentation model includes an encoder module, a bottleneck layer, a decoder module and a segmentation module; the encoder module includes multiple improved encoders of different size levels, and a maximum pooling layer is provided between every two adjacent improved encoders; the decoder module includes multiple improved decoders of different size levels, and an upsampling layer is provided between every two adjacent improved decoders; there is a residual connection between the improved encoders and improved decoders of the same size level; a maximum pooling layer is provided between the bottleneck layer and the last improved encoder, and an upsampling layer is provided between the bottleneck layer and the first improved decoder.

[0021] In one embodiment, the training process of the segmentation model includes:

[0022] S1. Obtain a prostate multimodal image dataset and preprocess it to obtain a training set; the training set includes multiple groups of sample images.

[0023] Specifically, the prostate multimodal image dataset includes multiple groups of image pairs, each group of image pairs includes T2WI images and DWI images of the same patient; step S1 specifically includes:

[0024] S11. For each image pair, use affine transformation to register the T2WI and DWI images to the same spatial coordinate system to ensure their spatial alignment. Then, crop the registered T2WI and DWI images separately to create a prostate region of a fixed size (e.g., 128 × 128 × 64). Pad the edges of the cropped images to ensure spatial consistency, resulting in cropped T2WI and DWI images.

[0025] S12. Perform normalization processing (such as Z-score normalization) on the cropped T2WI image and the cropped DWI image to obtain a normalized T2WI image and a normalized DWI image.

[0026] S13. Dynamically align the feature maps of normalized T2WI and normalized DWI images through a cross-modality interaction method to address the resolution difference between different modalities and obtain an aligned feature image.

[0027] Specifically, the normalized T2WI images are dynamically aligned through a cross-modal interaction method. and normalized DWI images The feature maps are fused and aligned feature images F are obtained. aligned , expressed as:

[0028] F aligned =A·F DWI

[0029]

[0030] in, represents the cross-modal attention matrix, H and W represent the image height and width, C represents the number of channels, D represents the depth, and A ij represents the element in row i and column j of the cross-modal attention matrix A, Represents the normalized T2WI image F T2 The i-th eigenvector of Represents the normalized DWI image F DWI The jth eigenvector of Represents the normalized DWI image F DWI The kth eigenvector.

[0031] S14. Adaptive instance normalization (AdaIN) is performed on the aligned feature images to eliminate the dependency of imaging parameters of different MRI devices, enhance the cross-institutional generalization ability of the model, and obtain sample images.

[0032] Specifically, step S14 is expressed as

[0033]

[0034] Among them, AadIN(F aligned ) represents the sample image, γ and β are device parameters, μ(F aligned ) represents the mean of the aligned feature images, σ(F aligned ) represents the standard deviation of the aligned feature images.

[0035] S15. Divide all sample images into training set and test set in a 9:1 ratio.

[0036] Specifically, the prostate multimodal dataset of this embodiment selects the prostate MRI multimodal dataset provided by the PI-CAI2022 challenge, including T2WI data and DWI data, as well as corresponding lesion annotations.

[0037] S2. Input the sample image into the encoder module to obtain the encoded image.

[0038] Specifically, the embodiment of the present invention uses four improved encoders of different size levels to form an encoder module. The four improved encoders have the same structure but different sizes. The improved encoder of the present invention embeds a lightweight SE Block (Squeeze-and-Excitation Block) and ResNeXt Block on the basis of the original 3D U-Net network encoder. Figure 2 The introduction of ResNeXt Block’s grouped convolution reduces computational complexity while preserving the diversity of local features. Furthermore, cross-level residual connections are used to fuse shallow detail information with deep semantic information, significantly improving segmentation accuracy, especially when processing fuzzy boundaries such as the prostate apex.

[0039] Specifically, the processing of each improved encoder includes:

[0040] S21. Pass the input features through a one-dimensional convolution layer to obtain convolution features, and pass the convolution features through a BN layer and a ReLU activation function layer to obtain extracted features.

[0041] S22. Input the extracted features into the grouped convolution layer to obtain grouped features, and pass the grouped features through the BN layer and the ReLU activation function layer to obtain local features.

[0042] Specifically, the grouped convolution layer includes G convolution subgroups, each of which performs an independent convolution operation using a convolution kernel of size 3×3×3. In step S22, the extracted features are input into the grouped convolution layer to obtain grouped features, including:

[0043] The extracted features are divided into G sub-blocks, different convolution subgroups are used to process different sub-blocks to obtain sub-convolution features, and the G sub-convolution features are added together to obtain group features.

[0044] The above process can be expressed as

[0045]

[0046] Among them, F group represents the grouping feature, X g Represents the g-th sub-block, Conv 3×3×3 () represents three-dimensional convolution.

[0047] S23. Input the local features into the SE block to obtain fused local features.

[0048] Specifically, step S24 includes:

[0049] S241. Perform global average pooling on the local features to obtain global features;

[0050] S242. Input the global feature into the dimensionality reduction fully connected layer and the ReLU activation function layer to obtain the first feature, and input the first feature into the dimensionality increase fully connected layer and the Sigmoid activation function layer to obtain the weight vector;

[0051] S243. Multiply the weight vector and the fused local feature element-by-element in the channel dimension to obtain the fused local feature;

[0052] At the same time, the SE module introduced in the encoder is also used to evaluate the importance of each channel to the segmentation task, assign a weight to each channel, and prune low-weight channels by setting a dynamic pruning threshold, thereby reducing redundant calculations and improving efficiency. During the training process, the dynamic pruning threshold τ is calculated based on the weight vector, and channels below the dynamic pruning threshold τ are pruned; the calculation formula of the dynamic pruning threshold τ is

[0053] τ=μ(W SE )-α·σ(W SE )

[0054] Wherein, μ() is a mean calculation operation, σ() is a standard deviation calculation operation, and α is an adjustment factor. In the embodiment of the present invention, α=0.5.

[0055] S24. Add the input features and the fused local features to obtain the output features.

[0056] S3. Input the encoded image into the bottleneck layer to obtain an intermediate image; the bottleneck layer includes two convolutional layers.

[0057] S4. Input the intermediate image into the decoder module to obtain a decoded image.

[0058] Specifically, the embodiment of the present invention uses four improved decoders of different size levels to form a decoder module. The four improved decoders have the same structure but different sizes. The size levels of the improved decoders correspond one to one with the size levels of the improved encoders. The present invention introduces a spatial attention module and a channel attention module into the improved decoder based on the original 3D U-Net network decoder. Figure 3 shown.

[0059] The processing of each improved decoder includes:

[0060] S31. Obtain the output features of the improved encoder of the same size level as the current improved decoder, and concatenate the output features with the input features of the current improved decoder to obtain concatenated features;

[0061] S32. The concatenated features are sequentially passed through a 3D convolutional layer, a batch normalization layer, and a ReLU activation function layer to obtain fused features;

[0062] S33. Fusion feature F c Input channel attention module to obtain channel weight vector W channel , expressed as

[0063] W channel =sig(MLP(AvgPool(F c )))

[0064] Among them, sig() represents the Sigmoid activation function, MLP() is the multi-layer perceptron, and AvgPool() represents average pooling.

[0065] S34. Fusion feature F c Input the spatial attention module to obtain the spatial weight vector W spatial , expressed as

[0066] W spatial =sig(Conv 7×7×7 (AvgPool(F c )+MaxPool(F c )))

[0067] Among them, Conv 7×7×7 () indicates a convolution layer with a convolution kernel of 7×7×7, and MaxPool() indicates maximum pooling.

[0068] S35. Multiply the fusion feature by the spatial weight vector and then by the channel weight vector to obtain the output feature F out , expressed as

[0069] F out =W channel (W spatial ·F c ).

[0070] S5. Input the decoded image into the segmentation module to obtain the segmentation result, and calculate the loss training model parameters according to the segmentation result until the model parameters converge.

[0071] Specifically, the segmentation module includes three parallel sub-networks, each targeting a different anatomical region of the prostate (apex, middle, and base). Each sub-network is optimized for the characteristics of the region it processes. For example, the apex sub-network introduces an ellipsoid shape prior constraint to improve segmentation accuracy in the apex region, while the middle sub-network strengthens the smoothness loss function of the gland edge to ensure smooth boundaries. The output weights of each sub-network are dynamically adjusted based on the distribution probability of the lesion in three-dimensional space, especially in high-risk areas such as the apex, to prioritize segmentation accuracy.

[0072] The processing of the segmentation module includes:

[0073] The decoded image is input into three sub-networks to obtain the prostate tip segmentation result, the prostate middle segmentation result, and the prostate base segmentation result;

[0074] The weight of the prostate tip, the weight of the prostate middle, and the weight of the prostate bottom are calculated based on the three segmentation results; , which is expressed as

[0075]

[0076] Among them, w tip represents the weight of the prostate apex, w mid represents the weight of the middle part of the prostate, w base represents the weight of the prostate base, P tip represents the segmentation result of the prostate apex, P mid represents the segmentation result of the middle part of the prostate, P base represents the segmentation result of the prostate base, ∑ p Indicates summing the segmentation results.

[0077] The segmentation results of the prostate apex, the middle prostate and the base of the prostate are weighted and fused according to the weight of the prostate apex, the weight of the middle prostate and the weight of the base of the prostate to obtain the segmentation result. In the present invention, unless otherwise clearly specified and limited, the terms "install", "set", "connect", "fix", "rotate" and the like should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integrated connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium; it can be the internal connection of two elements or the interaction relationship between two elements. Unless otherwise clearly defined, ordinary technicians in this field can understand the specific meaning of the above terms in the present invention according to the specific circumstances.

[0078] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A prostate MRI image segmentation method based on a dynamic multi-branch convolutional network, characterized in that: Building a segmentation model based on a 3D U-Net network structure, wherein the segmentation model includes an encoder module, a bottleneck layer, a decoder module, and a segmentation module; The segmentation model is trained and the prostate MRI image segmentation is performed using the trained segmentation model. The training process of the segmentation model includes: S1. Obtaining a prostate multimodal image dataset and preprocessing it to obtain a training set; the training set includes multiple groups of sample images; S2. Input the sample image into the encoder module to obtain an encoded image; the encoder module includes multiple improved encoders of different size levels, and a maximum pooling layer is provided between each two adjacent improved encoders; S3. Input the encoded image into the bottleneck layer to obtain an intermediate image; the bottleneck layer includes two convolutional layers; a maximum pooling layer is provided between the bottleneck layer and the last improved encoder, and an upsampling layer is provided between the bottleneck layer and the first improved decoder; S4. Inputting the intermediate image into the decoder module to obtain a decoded image; the decoder module includes multiple improved decoders of different size levels, with an upsampling layer between each two adjacent improved decoders; and residual connections exist between improved encoders and improved decoders of the same size level; S5. Input the decoded image into the segmentation module to obtain the segmentation result, and calculate the loss training model parameters according to the segmentation result until the model parameters converge.

2. The method for prostate MRI image segmentation based on a dynamic multi-branch convolutional network according to claim 1, characterized in that: The prostate multimodal image dataset includes multiple groups of image pairs, each group of image pairs includes T2WI images and DWI images of the same patient; step S1 specifically includes: S11. For each image pair, register the T2WI image and DWI image to the same spatial coordinate system using affine transformation, then perform cropping and edge filling to obtain a cropped T2WI image and a cropped DWI image. S12. Normalizing the cropped T2WI image and the cropped DWI image to obtain a normalized T2WI image and a normalized DWI image; S13. Dynamically aligning the feature maps of the normalized T2WI image and the normalized DWI image using a cross-modality interaction method to obtain an aligned feature image; S14. performing adaptive instance normalization on the aligned feature image to obtain a sample image; S15. Divide all sample images into training set and test set in a 9:1 ratio.

3. The prostate MRI image segmentation method based on a dynamic multi-branch convolutional network according to claim 1, characterized in that: Step S13 dynamically aligns the normalized T2WI image F through a cross-modality interaction method T2 and normalized DWI image F DWI The feature map of the alignment feature image F is obtained aligned , expressed as: F aligned =A·(F T2 ⊙F DWI ) Among them, A represents the cross-modal attention matrix, A ij represents the element in row i and column j of the cross-modal attention matrix A, Represents the normalized T2WI image F T2 The i-th eigenvector of Represents the normalized DWI image F DWI The jth eigenvector of Represents the normalized DWI image F DWI The kth eigenvector of .

4. The method for prostate MRI image segmentation based on a dynamic multi-branch convolutional network according to claim 1, wherein: The processing of each improved encoder includes: S21. Pass the input features through a one-dimensional convolution layer to obtain convolution features, and pass the convolution features through a BN layer and a ReLU activation function layer to obtain extracted features; S22. Input the extracted features into the group convolution layer to obtain group features, and pass the group features through the BN layer and the ReLU activation function layer to obtain local features; S23. Input the local features into the SE module to obtain fused local features; S24. Add the input features and the fused local features to obtain the output features.

5. The method for prostate MRI image segmentation based on a dynamic multi-branch convolutional network according to claim 4, characterized in that: The grouped convolution layer includes G convolution subgroups, each of which uses a convolution with a kernel size of 3×3×3. In step S22, the extracted features are input into the grouped convolution layer to obtain grouped features, including: The extracted features are divided into G sub-blocks, different convolution subgroups are used to process different sub-blocks to obtain sub-convolution features, and the G sub-convolution features are added together to obtain group features.

6. The method for prostate MRI image segmentation based on a dynamic multi-branch convolutional network according to claim 4, characterized in that: Step S23 includes: S231. Perform global average pooling on the local features to obtain global features; S232. Input the global feature into the dimensionality reduction fully connected layer and the ReLU activation function layer to obtain the first feature, and input the first feature into the dimensionality increase fully connected layer and the Sigmoid activation function layer to obtain the weight vector; S233. Multiply the weight vector and the fused local feature element-by-element in the channel dimension to obtain the fused local feature; During the training process, the dynamic pruning threshold τ is calculated according to the weight vector, and the channels below the dynamic pruning threshold τ are pruned; the calculation formula of the dynamic pruning threshold τ is τ=μ(W SE )-a·s(W SE ) Among them, μ() is the mean calculation operation, σ() is the standard deviation calculation operation, and α is the adjustment factor.

7. The method for prostate MRI image segmentation based on a dynamic multi-branch convolutional network according to claim 1, characterized in that: The processing of each improved decoder includes: S31. Obtain the output features of the improved encoder of the same size level as the current improved decoder, and concatenate the output features with the input features of the current improved decoder to obtain concatenated features; S32. The concatenated features are sequentially passed through a 3D convolutional layer, a batch normalization layer, and a ReLU activation function layer to obtain fused features; S33. Input the fused features into the channel attention module to obtain a channel weight vector; S34. Input the fused features into the spatial attention module to obtain a spatial weight vector; S35. Multiply the fusion feature by the spatial weight vector and then multiply it by the channel weight vector to obtain the output feature.

8. The method for prostate MRI image segmentation based on a dynamic multi-branch convolutional network according to claim 1, wherein: The segmentation module includes three parallel sub-networks, which are used to process the three anatomical regions of the prostate tip, the middle prostate, and the prostate base respectively; The processing of the segmentation module includes: The decoded image is input into three sub-networks to obtain the prostate tip segmentation result, the prostate middle segmentation result, and the prostate base segmentation result; The weight of the prostate tip, the weight of the prostate middle, and the weight of the prostate base were calculated based on the three segmentation results; The segmentation results of the prostate apex, the middle prostate and the base of the prostate are weighted and fused according to the weight of the prostate apex, the weight of the middle prostate and the weight of the base of the prostate to obtain the segmentation result.

Citation Information

Patent Citations

  • Method for automatically segmenting whole prostate gland based on deep learning convolutional neural network

    CN114399501A

  • Federal learning-oriented model pruning method and system

    CN115983366A

  • Image recognition method based on fusion attention mechanism

    CN116229234A

  • Low-illumination image defogging method based on lightweight deep neural network

    CN116309110A

  • Cross-modal 3D medical image registration method based on attention mechanism

    CN116309748A