Prostate ultrasound segmentation method and device based on multi-scale feature fusion
Patent Information
- Application Number
- CN202411174853.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-26
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-08-26
AI Technical Summary
然而,2.5D的分割方式,虽然都可以为2D卷积神经网络带来性能的提升并且可以在低计算成本的情况下获取空间轮廓信息,但是却并不适用于经直肠超声这类厚层数据,该类数据由于其大量信息都存在于切面上且立体信息较少,因此无法像CT影像那样很容易就得到三个具有丰富信息的正交平面
[0048]1、本申请针对U-Net模型中编解码器只能进行相同尺度的信息传递而产生的语义鸿沟,引入了多尺度跳跃连接的方式,模型能够更有效地提取前列腺影像中目标区域的多尺度特征信息。同时,加入了通道与空间注意力机制MSA(Multi-Scale-Attention),通过注意力机制筛选出空间或通道中不重要的特征信息,从而保留更为关键的特征信息,并且将MSA模块加入到每个上采样操作中,使模型能够关注到各个尺度切片中有用的特征信息,从而提升分割精度。
Smart Images

Figure CN119131387B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image segmentation technology, and more specifically, to a method and apparatus for prostate ultrasound segmentation based on multi-scale feature fusion. Background Technology
[0002] In recent years, some deep learning methods have been introduced into the segmentation of prostate ultrasound images. For example, U-Net is a representative 2D convolutional neural network. Although U-Net performs well in image segmentation tasks, its single-skip connection limitation prevents the model from fusing feature information from multiple different levels, thus affecting the segmentation results. Furthermore, 2D segmentation networks can only extract features from a two-dimensional plane and cannot obtain three-dimensional structural information from image data. This is because 2D segmentation networks have limitations in processing relationships between adjacent slices; they often fail to identify and construct the contour representation information of the target region. Compared to 2D models, 3D models can focus more comprehensively on the detailed features and spatial contours in image data. In addition, 3D models have the ability to process a series of adjacent image slices simultaneously, which are often described as "stacks" or "volumes." In this way, 3D models can better capture the spatial relationships between objects and structures. Compared to 2D models, 3D models can more comprehensively capture features such as depth, height, and volume in continuous image slices, thus providing more complete and coherent image information. However, training 3D networks often requires higher hardware configurations and more computational costs, so the limitations of computing power and storage resources need to be fully considered in practical applications.
[0003] To bridge the gap between 2D and 3D convolutional neural networks, 2.5D segmentation methods have emerged. Many existing 2.5D segmentation methods integrate volumetric information into 2D convolutions by improving or designing new architectures to achieve efficient volumetric medical image segmentation. However, while 2.5D segmentation methods can improve the performance of 2D convolutional neural networks and acquire spatial contour information at low computational cost, they are not suitable for thick-slice data such as transrectal ultrasound. This type of data contains a large amount of information on the slice plane and less three-dimensional information, making it difficult to easily obtain three orthogonal planes with rich information like CT images. Furthermore, these methods do not effectively integrate information between slices; they simply input three planes (sagittal, coronal, and axial planes) into the 2D network for training and then fuse them according to their positions to obtain the segmentation result. Summary of the Invention
[0004] To overcome at least one deficiency in the prior art, this application provides a method and apparatus for prostate ultrasound segmentation based on multi-scale feature fusion.
[0005] Firstly, a method for prostate ultrasound segmentation based on multi-scale feature fusion is provided, including:
[0006] A prostate ultrasound segmentation model was constructed. The prostate ultrasound segmentation model adopted a U-Net network, which includes an encoder and a decoder. The encoder includes multiple encoding layers, and the decoder includes multiple decoding layers. The convolutional modules in each encoding layer and each decoding layer were replaced with hybrid convolutional modules. A channel and spatial attention module was also set between the encoder and the decoder. The spatial attention module is used to retain useful feature information in the features output by the encoding layer. The useful feature information is input into the decoder for multi-scale feature fusion and outputs the segmentation result.
[0007] The prostate ultrasound segmentation model was trained to obtain the trained prostate ultrasound segmentation model.
[0008] The prostate ultrasound image to be segmented is input into the trained prostate ultrasound segmentation model to obtain the segmentation result.
[0009] In one embodiment, the hybrid convolution module includes: a first 3D convolution block, a first 2D convolution block, a second 2D convolution block, a second 3D convolution block, and a 2.5D convolution block;
[0010] The input to the hybrid convolution module passes through a first 3D convolution block and a first 2D convolution block before being input into a 2.5D convolution block;
[0011] The input to the hybrid convolution module passes through a second 2D convolution block and a second 3D convolution block before being input into a 2.5D convolution block;
[0012] The 2.5D convolution block performs 2.5D convolution operations on the two inputs respectively, and then merges the convolution results for output.
[0013] In one embodiment, the encoder has five encoding layers, referred to as the first encoding layer, the second encoding layer, the third encoding layer, the fourth encoding layer, and the fifth encoding layer; the decoder has four decoding layers, referred to as the first decoding layer, the second decoding layer, the third decoding layer, and the fourth decoding layer.
[0014] The channel and spatial attention module has five channel and spatial attention mechanisms, which are referred to as the first channel and spatial attention mechanism, the second channel and spatial attention mechanism, the third channel and spatial attention mechanism, the fourth channel and spatial attention mechanism, and the fifth channel and spatial attention mechanism.
[0015] The outputs of the first channel and the spatial attention mechanism, the second channel and the spatial attention mechanism, the third channel and the spatial attention mechanism, the fourth channel and the spatial attention mechanism, and the fifth channel and the spatial attention mechanism are all input into the first decoding layer for multi-scale fusion operation to obtain the first fusion result;
[0016] The outputs of the first channel and the spatial attention mechanism, the second channel and the spatial attention mechanism, and the third channel and the spatial attention mechanism, as well as the first fusion result, are all input into the second decoding layer for multi-scale fusion operation to obtain the second fusion result.
[0017] The outputs of the first channel and the spatial attention mechanism, the second channel and the spatial attention mechanism, and the second fusion result are all input into the third decoding layer for multi-scale fusion operation to obtain the third fusion result.
[0018] The output of the first channel and the spatial attention mechanism, along with the third fusion result, are input into the fourth decoding layer for multi-scale fusion operation to obtain the segmentation result.
[0019] In one embodiment, the output of the first channel and the spatial attention mechanism are first max-pooled after being input to the decoding layer, and then fused with other inputs of the decoding layer;
[0020] The output of the second channel and the spatial attention mechanism is first max-pooled after being input into the decoding layer, and then fused with other inputs of the decoding layer.
[0021] The output of the third channel and the spatial attention mechanism is first subjected to a hybrid convolution operation after being input into the decoding layer, and then fused with other inputs of the decoding layer;
[0022] The output of the fourth channel and the spatial attention mechanism is first subjected to bilinear interpolation after being input to the decoding layer, and then fused with other inputs of the decoding layer.
[0023] The output of the fifth channel and the spatial attention mechanism is first subjected to bilinear interpolation after being input to the decoding layer, and then fused with other inputs of the decoding layer;
[0024] The outputs of the first, second, and third decoding layers are first subjected to bilinear interpolation after being input into the next decoding layer, and then fused with other inputs of the decoding layer.
[0025] Secondly, a prostate ultrasound segmentation device based on multi-scale feature fusion is provided, comprising:
[0026] The model building module is used to construct a prostate ultrasound segmentation model. The prostate ultrasound segmentation model adopts a U-Net network, which includes an encoder and a decoder. The encoder includes multiple encoding layers, and the decoder includes multiple decoding layers. The convolutional modules in each encoding layer and each decoding layer are replaced with hybrid convolutional modules. A channel and spatial attention module is also set between the encoder and the decoder. The spatial attention module is used to retain useful feature information in the features output by the encoding layer. The useful feature information is input into the decoder for multi-scale feature fusion and outputs the segmentation result.
[0027] The model training module is used to train the prostate ultrasound segmentation model to obtain the trained prostate ultrasound segmentation model.
[0028] The segmentation module is used to input the prostate ultrasound image to be segmented into the trained prostate ultrasound segmentation model to obtain the segmentation result.
[0029] In one embodiment, the hybrid convolution module includes: a first 3D convolution block, a first 2D convolution block, a second 2D convolution block, a second 3D convolution block, and a 2.5D convolution block;
[0030] The input to the hybrid convolution module passes through a first 3D convolution block and a first 2D convolution block before being input into a 2.5D convolution block;
[0031] The input to the hybrid convolution module passes through a second 2D convolution block and a second 3D convolution block before being input into a 2.5D convolution block;
[0032] The 2.5D convolution block performs 2.5D convolution operations on the two inputs respectively, and then merges the convolution results for output.
[0033] In one embodiment, the encoder has five encoding layers, referred to as the first encoding layer, the second encoding layer, the third encoding layer, the fourth encoding layer, and the fifth encoding layer; the decoder has four decoding layers, referred to as the first decoding layer, the second decoding layer, the third decoding layer, and the fourth decoding layer.
[0034] The channel and spatial attention module has five channel and spatial attention mechanisms, which are referred to as the first channel and spatial attention mechanism, the second channel and spatial attention mechanism, the third channel and spatial attention mechanism, the fourth channel and spatial attention mechanism, and the fifth channel and spatial attention mechanism.
[0035] The outputs of the first channel and the spatial attention mechanism, the second channel and the spatial attention mechanism, the third channel and the spatial attention mechanism, the fourth channel and the spatial attention mechanism, and the fifth channel and the spatial attention mechanism are all input into the first decoding layer for multi-scale fusion operation to obtain the first fusion result;
[0036] The outputs of the first channel and the spatial attention mechanism, the second channel and the spatial attention mechanism, and the third channel and the spatial attention mechanism, as well as the first fusion result, are all input into the second decoding layer for multi-scale fusion operation to obtain the second fusion result.
[0037] The outputs of the first channel and the spatial attention mechanism, the second channel and the spatial attention mechanism, and the second fusion result are all input into the third decoding layer for multi-scale fusion operation to obtain the third fusion result.
[0038] The output of the first channel and the spatial attention mechanism, along with the third fusion result, are input into the fourth decoding layer for multi-scale fusion operation to obtain the segmentation result.
[0039] In one embodiment, the output of the first channel and the spatial attention mechanism are first max-pooled after being input to the decoding layer, and then fused with other inputs of the decoding layer;
[0040] The output of the second channel and the spatial attention mechanism is first max-pooled after being input into the decoding layer, and then fused with other inputs of the decoding layer.
[0041] The output of the third channel and the spatial attention mechanism is first subjected to a hybrid convolution operation after being input into the decoding layer, and then fused with other inputs of the decoding layer;
[0042] The output of the fourth channel and the spatial attention mechanism is first subjected to bilinear interpolation after being input to the decoding layer, and then fused with other inputs of the decoding layer.
[0043] The output of the fifth channel and the spatial attention mechanism is first subjected to bilinear interpolation after being input to the decoding layer, and then fused with other inputs of the decoding layer;
[0044] The outputs of the first, second, and third decoding layers are first subjected to bilinear interpolation after being input into the next decoding layer, and then fused with other inputs of the decoding layer.
[0045] Thirdly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the aforementioned prostate ultrasound segmentation method based on multi-scale feature fusion.
[0046] Fourthly, a computer program product is provided, characterized in that it includes a computer program / instruction, which, when executed by a processor, implements the above-described prostate ultrasound segmentation method based on multi-scale feature fusion.
[0047] Compared with the prior art, this application has the following beneficial effects:
[0048] 1. This application addresses the semantic gap caused by the U-Net model's encoder and decoder only being able to transmit information at the same scale. It introduces a multi-scale skip connection approach, enabling the model to more effectively extract multi-scale feature information of the target region in prostate images. Simultaneously, a channel and spatial attention mechanism, MSA (Multi-Scale Attention), is incorporated. This attention mechanism filters out unimportant features in the spatial or channel dimensions, thus retaining more crucial features. Furthermore, the MSA module is added to each upsampling operation, allowing the model to focus on useful features in slices at various scales, thereby improving segmentation accuracy.
[0049] 2. This application replaces all convolutions in the U-Net model with hybrid convolutions. These hybrid convolutions incorporate both 2D and 3D CNNs, enabling simultaneous capture of local and spatial information. Furthermore, feature fusion processes are implemented in each dimensional branch to merge these two-dimensional and three-dimensional features. This allows the entire 2.5D model to capture representations across all dimensions without requiring a large number of parameters or computational costs. The aim is to achieve good segmentation accuracy while reducing computational costs, thus realizing precise segmentation of prostate images. Attached Figure Description
[0050] This application can be better understood by referring to the description given below in conjunction with the accompanying drawings, which, together with the detailed description below, are incorporated in and form part of this specification. In the drawings:
[0051] Figure 1 A schematic diagram of a prostate ultrasound segmentation model is shown;
[0052] Figure 2 A schematic diagram of the hybrid convolutional module FM is shown;
[0053] Figure 3 A block diagram of a prostate ultrasound segmentation device based on multi-scale feature fusion is shown. Detailed Implementation
[0054] Exemplary embodiments of the present application will be described below with reference to the accompanying drawings. For clarity and brevity, not all features of the actual embodiments are described in the specification. However, it should be understood that many embodiment-specific decisions can be made in the development of any such actual embodiment to achieve the developer’s specific objectives, and these decisions may vary as the embodiments differ.
[0055] It should also be noted that, in order to avoid obscuring this application with unnecessary details, only the device structure closely related to the solution according to this application is shown in the accompanying drawings, while other details that are not closely related to this application are omitted.
[0056] It should be understood that this application is not limited to the described embodiments by virtue of the following description with reference to the accompanying drawings. In this document, embodiments may be combined with each other, features may be substituted or borrowed between different embodiments, and one or more features may be omitted in one embodiment, where feasible.
[0057] This application provides a method for prostate ultrasound segmentation based on multi-scale feature fusion, including:
[0058] Step S1: Construct a prostate ultrasound segmentation model; Figure 1 A schematic diagram of a prostate ultrasound segmentation model is shown. (See attached image) Figure 1 The prostate ultrasound segmentation model uses a U-Net network, which includes an encoder and a decoder. The encoder includes multiple encoding layers, and the decoder includes multiple decoding layers. The convolutional modules in each encoding and decoding layer are replaced with hybrid convolutional modules. A channel and spatial attention module is also set between the encoder and the decoder. The spatial attention module is used to retain useful feature information in the features output by the encoding layer. The useful feature information is input into the decoder for multi-scale feature fusion and outputs the segmentation result.
[0059] Figure 2 The diagram illustrates the hybrid convolutional module FM, which consists of two paths. The purple squares contain both 2D and 3D convolutional blocks, while the blue squares contain only a single 3D block (within which are two 3×3×3 convolutional layers), responsible for extracting 3D features from prostate ultrasound images. The green squares contain a single 2D convolutional block with two 3×3 convolutional layers. The green squares focus on extracting local detail features from 2D slices of prostate ultrasound images. Each convolutional layer is followed by instance normalization and a ReLU activation function to enhance the model's nonlinear representation capabilities and stability.
[0060] Specifically, the hybrid convolution module includes: a first 3D convolution block, a first 2D convolution block, a second 2D convolution block, a second 3D convolution block, and a 2.5D convolution block;
[0061] The input to the hybrid convolution module passes through a first 3D convolution block and a first 2D convolution block before being input into a 2.5D convolution block;
[0062] The input to the hybrid convolution module passes through a second 2D convolution block and a second 3D convolution block before being input into a 2.5D convolution block;
[0063] The 2.5D convolution block performs 2.5D convolution operations on the two inputs respectively, and then merges the convolution results for output.
[0064] Here, a hybrid convolutional module is employed, allowing CNNs of different dimensions to extract feature information at their respective layers. During this process, feature maps of different dimensions can complement and enhance each other, resulting in a richer and more comprehensive feature representation. This feature representation includes local detail information from 2D CNNs, spatial contour information from 3D CNNs, and inter-layer correlation information. Furthermore, the shallow fusion of multiple CNNs generates numerous new paths, constructing a more complex, multi-layered network structure. To integrate this multi-layered network, this embodiment employs a strategy that combines shallow and deep layers. In each hybrid convolutional block (containing 2D / 2.5D / 3D CNNs), if sparse inter-slice information is over-represented, it indicates a lack of local feature information within the slice; in this case, the model will choose to pass through a 2D convolutional branch to focus on extracting intra-slice information. Additionally, when relevant spatial information is needed, the 3D convolutions in the branch will play a role in effectively extracting features and integrating 3D spatial information. This hybrid convolution approach allows for flexible adjustment of processing strategies based on different task requirements, thereby improving overall performance and enabling a comprehensive understanding of the input image.
[0065] Step S2: Train the prostate ultrasound segmentation model to obtain the trained prostate ultrasound segmentation model.
[0066] First, data collection was conducted in collaboration with medical institutions to obtain raw prostate ultrasound images of several patients in DICOM format. Each patient's data was placed in a separate folder containing all their prostate TRUS images, i.e., continuous ultrasound slices. These data were acquired using a transrectal ultrasound probe at different angles and times, and can be viewed using specific software in DICOM format. Because the data provided by each medical institution varied, the obtained data underwent screening as follows: First, images with clear contours, well-defined prostate outlines, and intact prostate shape were selected, as these provide accurate and detailed information on prostate structure and pathology. Second, images with indistinct prostate TRUS outlines and incomplete glandular structures were included in the dataset, even if they did not provide sufficiently detailed prostate information, as they may still contain valuable diagnostic clues. Finally, images of the prostate taken from different angles were also included. These images do not directly reflect the true morphology and pathological characteristics of the prostate, and clinicians cannot make appropriate judgments based on them; therefore, these data were discarded.
[0067] Next, the data underwent preprocessing. Since the acquired data were DICOM format transrectal ultrasound (TRUS) images of the prostate containing rich information, including instrument settings, image acquisition parameters, and patient details, this study primarily focuses on the image information of the prostate TRUS images themselves; other information was considered redundant. Therefore, an algorithmic batch processing method was used to extract the images, focusing only on the entire glandular portion of the prostate TRUS image and adjusting it to a uniform pixel size. After this processing, the extracted prostate TRUS images were saved as JPG format, containing a total of 1800 consecutive prostate TRUS slice images. It is noteworthy that one patient corresponds to multiple consecutive TRUS slices.
[0068] Then, a prostate ultrasound segmentation model is trained based on the preprocessed data. The specific training process is a standard technique in this field and will not be described in detail here.
[0069] Step S3: Input the prostate ultrasound image to be segmented into the trained prostate ultrasound segmentation model to obtain the segmentation result.
[0070] In one embodiment, see Figure 1 The encoder has 5 encoding layers, which are designated as the first encoding layer, the second encoding layer, the third encoding layer, the fourth encoding layer, and the fifth encoding layer; the decoder has 4 decoding layers, which are designated as the first decoding layer, the second decoding layer, the third decoding layer, and the fourth decoding layer.
[0071] The channel and spatial attention module has five channel and spatial attention mechanisms (MSA), which are referred to as the first channel and spatial attention mechanism, the second channel and spatial attention mechanism, the third channel and spatial attention mechanism, the fourth channel and spatial attention mechanism, and the fifth channel and spatial attention mechanism.
[0072] The outputs of the first channel and the spatial attention mechanism, the second channel and the spatial attention mechanism, the third channel and the spatial attention mechanism, the fourth channel and the spatial attention mechanism, and the fifth channel and the spatial attention mechanism are all input into the first decoding layer for multi-scale fusion operation to obtain the first fusion result;
[0073] The outputs of the first channel and the spatial attention mechanism, the second channel and the spatial attention mechanism, and the third channel and the spatial attention mechanism, as well as the first fusion result, are all input into the second decoding layer for multi-scale fusion operation to obtain the second fusion result.
[0074] The outputs of the first channel and the spatial attention mechanism, the second channel and the spatial attention mechanism, and the second fusion result are all input into the third decoding layer for multi-scale fusion operation to obtain the third fusion result.
[0075] The output of the first channel and the spatial attention mechanism, along with the third fusion result, are input into the fourth decoding layer for multi-scale fusion operation to obtain the segmentation result.
[0076] In this embodiment, to address the semantic gap caused by the U-Net model's encoder and decoder only being able to transmit information at the same scale, a multi-scale skip connection approach is introduced. This allows the model to more effectively extract multi-scale feature information of the target region in prostate images. Simultaneously, a channel and spatial attention mechanism (MSA) is incorporated. This attention mechanism filters out unimportant features in the spatial or channel contexts, thus retaining more critical features. Furthermore, the MSA module is added to each upsampling operation, enabling the model to focus on useful features in slices at various scales, thereby improving segmentation accuracy.
[0077] A channel-spatial cross-attention mechanism is incorporated into each upsampling operation. First, multi-scale features of the input data are extracted through different convolutional or pooling layers. For each scale of feature, MSA computes an attention weight. These weights are typically computed using a lightweight attention network (such as a self-attention or channel attention network). This network evaluates the importance of each feature and outputs a weight vector. Next, MSA uses the computed attention weights to weightedly fuse features from different scales. To preserve information from the original input features and prevent gradient vanishing, MSA adds the weighted fused features to the original input features via residual connections. Finally, MSA outputs a feature representation that fuses multi-scale information and attention weights.
[0078] In one embodiment, the output of the first channel and the spatial attention mechanism are first subjected to max pooling after being input to the decoding layer, and then fused with other inputs of the decoding layer;
[0079] The output of the second channel and the spatial attention mechanism is first subjected to max pooling after being input into the decoding layer, and then fused with other inputs of the decoding layer.
[0080] The output of the third channel and the spatial attention mechanism is first subjected to a hybrid convolution operation after being input into the decoding layer, and then fused with other inputs of the decoding layer;
[0081] The output of the fourth channel and the spatial attention mechanism is first subjected to bilinear interpolation after being input to the decoding layer, and then fused with other inputs of the decoding layer.
[0082] The output of the fifth channel and the spatial attention mechanism is first subjected to bilinear interpolation after being input to the decoding layer, and then fused with other inputs of the decoding layer;
[0083] The outputs of the first, second, and third decoding layers are first subjected to bilinear interpolation after being input into the next decoding layer, and then fused with other inputs of the decoding layer.
[0084] To further verify the effectiveness of the proposed method, this application employs three evaluation metrics: similarity coefficient (Dice), intersection-over-union (IoU), and cross-entropy loss function (Loss) to measure and evaluate the method. The proposed method is experimentally compared with multiple segmentation networks, including traditional segmentation models such as U-Net, UNet++, and nnUNet. The experimental results are shown in Table 1 below.
[0085] Table 1. Performance comparison of multiple segmentation networks on the prostate TRUS dataset.
[0086]
[0087]
[0088] As shown in Table 1, the method in this application achieves the highest Dice score, improving it by 4.92% compared to the traditional U-Net. This is mainly attributed to the effective extraction of feature information at various scales through multi-scale skip connections in this method, and the keen attention paid to important information by MSA. Compared to U-Net++, the network structure in this application adopts a simpler and more effective skip connection structure, resulting in a 2.79% improvement in the Dice score. Compared to the latest segmentation methods, the method in this application also demonstrates advantages.
[0089] Employing the same inventive concept as the prostate ultrasound segmentation method based on multi-scale feature fusion, this embodiment also provides a corresponding prostate ultrasound segmentation device based on multi-scale feature fusion. Figure 3 A structural block diagram of a prostate ultrasound segmentation device based on multi-scale feature fusion is shown, including:
[0090] Model building module 31 is used to build a prostate ultrasound segmentation model. The prostate ultrasound segmentation model adopts a U-Net network, which includes an encoder and a decoder. The encoder includes multiple encoding layers, and the decoder includes multiple decoding layers. The convolutional modules in each encoding layer and each decoding layer are replaced with hybrid convolutional modules. A channel and spatial attention module is also set between the encoder and the decoder. The spatial attention module is used to retain useful feature information in the features output by the encoding layer. The useful feature information is input into the decoder for multi-scale feature fusion and outputs the segmentation result.
[0091] Model training module 32 is used to train the prostate ultrasound segmentation model to obtain the trained prostate ultrasound segmentation model;
[0092] The segmentation module 33 is used to input the prostate ultrasound image to be segmented into the trained prostate ultrasound segmentation model to obtain the segmentation result.
[0093] The prostate ultrasound segmentation device based on multi-scale feature fusion in this embodiment has the same inventive concept as the prostate ultrasound segmentation method based on multi-scale feature fusion described above. Therefore, the specific implementation of this device can be found in the embodiment section of the prostate ultrasound segmentation method based on multi-scale feature fusion described above, and its technical effects correspond to the technical effects of the above method, so it will not be repeated here.
[0094] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the above-described prostate ultrasound segmentation method based on multi-scale feature fusion.
[0095] This application provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the above-described prostate ultrasound segmentation method based on multi-scale feature fusion.
[0096] In summary, this application has the following technical effects:
[0097] 1. This application addresses the semantic gap caused by the U-Net model's encoder and decoder only being able to transmit information at the same scale. It introduces a multi-scale skip connection approach, enabling the model to more effectively extract multi-scale feature information of the target region in prostate images. Simultaneously, a channel and spatial attention mechanism, MSA (Multi-Scale Attention), is incorporated. This attention mechanism filters out unimportant features in the spatial or channel dimensions, thus retaining more crucial features. Furthermore, the MSA module is added to each upsampling operation, allowing the model to focus on useful features in slices at various scales, thereby improving segmentation accuracy.
[0098] 2. This application replaces all convolutions in the U-Net model with hybrid convolutions. These hybrid convolutions incorporate both 2D and 3D CNNs, enabling simultaneous capture of local and spatial information. Furthermore, feature fusion processes are implemented in each dimensional branch to merge these two-dimensional and three-dimensional features. This allows the entire 2.5D model to capture representations across all dimensions without requiring a large number of parameters or computational costs. The aim is to achieve good segmentation accuracy while reducing computational costs, thus realizing precise segmentation of prostate images.
[0099] The above descriptions are merely various embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A prostate ultrasound segmentation method based on multi-scale feature fusion, characterized in that, include: A prostate ultrasound segmentation model is constructed. The prostate ultrasound segmentation model adopts a U-Net network, which includes an encoder and a decoder. The encoder includes multiple encoding layers, and the decoder includes multiple decoding layers. The convolutional modules in each encoding layer and each decoding layer are replaced with hybrid convolutional modules. A channel and spatial attention module is also set between the encoder and the decoder. The spatial attention module is used to retain useful feature information in the features output by the encoding layer. The useful feature information is input into the decoder for multi-scale feature fusion and outputs the segmentation result. The prostate ultrasound segmentation model is trained to obtain the trained prostate ultrasound segmentation model; The prostate ultrasound image to be segmented is input into the trained prostate ultrasound segmentation model to obtain the segmentation result; The encoder has five encoding layers, which are designated as the first encoding layer, the second encoding layer, the third encoding layer, the fourth encoding layer, and the fifth encoding layer; the decoder has four decoding layers, which are designated as the first decoding layer, the second decoding layer, the third decoding layer, and the fourth decoding layer. The channel and spatial attention module is equipped with five channel and spatial attention mechanisms, which are respectively referred to as the first channel and spatial attention mechanism, the second channel and spatial attention mechanism, the third channel and spatial attention mechanism, the fourth channel and spatial attention mechanism, and the fifth channel and spatial attention mechanism. The outputs of the first channel and the spatial attention mechanism, the second channel and the spatial attention mechanism, the third channel and the spatial attention mechanism, the fourth channel and the spatial attention mechanism, and the fifth channel and the spatial attention mechanism are all input to the first decoding layer for multi-scale fusion operation to obtain the first fusion result; The outputs of the first channel and spatial attention mechanism, the second channel and spatial attention mechanism, the third channel and spatial attention mechanism, and the first fusion result are all input into the second decoding layer for multi-scale fusion operation to obtain the second fusion result. The outputs of the first channel and the spatial attention mechanism, the second channel and the spatial attention mechanism, and the second fusion result are all input into the third decoding layer for multi-scale fusion operation to obtain the third fusion result. The output of the first channel and the spatial attention mechanism, along with the third fusion result, are input into the fourth decoding layer for multi-scale fusion operation to obtain the segmentation result.
2. The method as described in claim 1, characterized in that, The hybrid convolution module includes: a first 3D convolution block, a first 2D convolution block, a second 2D convolution block, a second 3D convolution block, and a 2.5D convolution block; The input to the hybrid convolution module passes through the first 3D convolution block and the first 2D convolution block before being input to the 2.5D convolution block; The input to the hybrid convolution module passes through the second 2D convolution block and the second 3D convolution block before being input to the 2.5D convolution block; The 2.5D convolution block performs 2.5D convolution operations on the two inputs respectively, and then merges the convolution results for output.
3. The method as described in claim 1, characterized in that, The output of the first channel and the spatial attention mechanism are first max-pooled after being input into the decoding layer, and then fused with other inputs of the decoding layer. The output of the second channel and the spatial attention mechanism are first max-pooled after being input to the decoding layer, and then fused with other inputs of the decoding layer. The output of the third channel and the spatial attention mechanism is first subjected to a mixed convolution operation after being input into the decoding layer, and then fused with other inputs of the decoding layer. The output of the fourth channel and the spatial attention mechanism is first subjected to bilinear interpolation after being input to the decoding layer, and then fused with other inputs of the decoding layer. The output of the fifth channel and the spatial attention mechanism is first subjected to bilinear interpolation after being input to the decoding layer, and then fused with other inputs of the decoding layer. The outputs of the first decoding layer, the second decoding layer, and the third decoding layer are first subjected to bilinear interpolation after being input to the next decoding layer, and then fused with other inputs of the decoding layer.
4. A prostate ultrasound segmentation device based on multi-scale feature fusion, characterized in that, include: A model building module is used to construct a prostate ultrasound segmentation model. The prostate ultrasound segmentation model adopts a U-Net network, which includes an encoder and a decoder. The encoder includes multiple encoding layers, and the decoder includes multiple decoding layers. The convolutional modules in each encoding layer and each decoding layer are replaced with hybrid convolutional modules. A channel and spatial attention module is also provided between the encoder and the decoder. The spatial attention module is used to retain useful feature information in the features output by the encoding layers. The useful feature information is input into the decoder for multi-scale feature fusion and outputs the segmentation result. The model training module is used to train the prostate ultrasound segmentation model to obtain the trained prostate ultrasound segmentation model. The segmentation module is used to input the prostate ultrasound image to be segmented into the trained prostate ultrasound segmentation model to obtain the segmentation result; The encoder has five encoding layers, which are designated as the first encoding layer, the second encoding layer, the third encoding layer, the fourth encoding layer, and the fifth encoding layer; the decoder has four decoding layers, which are designated as the first decoding layer, the second decoding layer, the third decoding layer, and the fourth decoding layer. The channel and spatial attention module is equipped with five channel and spatial attention mechanisms, which are respectively referred to as the first channel and spatial attention mechanism, the second channel and spatial attention mechanism, the third channel and spatial attention mechanism, the fourth channel and spatial attention mechanism, and the fifth channel and spatial attention mechanism. The outputs of the first channel and the spatial attention mechanism, the second channel and the spatial attention mechanism, the third channel and the spatial attention mechanism, the fourth channel and the spatial attention mechanism, and the fifth channel and the spatial attention mechanism are all input to the first decoding layer for multi-scale fusion operation to obtain the first fusion result; The outputs of the first channel and spatial attention mechanism, the second channel and spatial attention mechanism, the third channel and spatial attention mechanism, and the first fusion result are all input into the second decoding layer for multi-scale fusion operation to obtain the second fusion result. The outputs of the first channel and the spatial attention mechanism, the second channel and the spatial attention mechanism, and the second fusion result are all input into the third decoding layer for multi-scale fusion operation to obtain the third fusion result. The output of the first channel and the spatial attention mechanism, along with the third fusion result, are input into the fourth decoding layer for multi-scale fusion operation to obtain the segmentation result.
5. The apparatus as described in claim 4, characterized in that, The hybrid convolution module includes: a first 3D convolution block, a first 2D convolution block, a second 2D convolution block, a second 3D convolution block, and a 2.5D convolution block; The input to the hybrid convolution module passes through the first 3D convolution block and the first 2D convolution block before being input to the 2.5D convolution block; The input to the hybrid convolution module passes through the second 2D convolution block and the second 3D convolution block before being input to the 2.5D convolution block; The 2.5D convolution block performs 2.5D convolution operations on the two inputs respectively, and then merges the convolution results for output.
6. The apparatus as claimed in claim 4, characterized in that, The output of the first channel and the spatial attention mechanism are first max-pooled after being input into the decoding layer, and then fused with other inputs of the decoding layer. The output of the second channel and the spatial attention mechanism are first max-pooled after being input to the decoding layer, and then fused with other inputs of the decoding layer. The output of the third channel and the spatial attention mechanism is first subjected to a mixed convolution operation after being input into the decoding layer, and then fused with other inputs of the decoding layer. The output of the fourth channel and the spatial attention mechanism is first subjected to bilinear interpolation after being input to the decoding layer, and then fused with other inputs of the decoding layer. The output of the fifth channel and the spatial attention mechanism is first subjected to bilinear interpolation after being input to the decoding layer, and then fused with other inputs of the decoding layer. The outputs of the first decoding layer, the second decoding layer, and the third decoding layer are first subjected to bilinear interpolation after being input to the next decoding layer, and then fused with other inputs of the decoding layer.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the prostate ultrasound segmentation method based on multi-scale feature fusion as described in any one of claims 1-3.
8. A computer program product, characterized in that, Includes a computer program / instruction, which, when executed by a processor, implements the prostate ultrasound segmentation method based on multi-scale feature fusion as described in any one of claims 1-3.
Citation Information
Patent Citations
Brain tumor MR image segmentation method
CN114519719A
Cross-species target detection method based on cross-layer feature fusion and linear attention optimization
CN116246147A