Image segmentation method and device, electronic equipment and storage medium

Through the image segmentation model based on the three-dimensional cascade attention mechanism, the problem of low image segmentation accuracy in traditional Chinese medicine in the prior art is solved, and comprehensive capture and precise segmentation of three-dimensional target area features are achieved.

CN120374637APending Publication Date: 2025-07-25TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311706447.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, medical image segmentation accuracy is low, especially in three-dimensional medical image processing, it is difficult to accurately express the three-dimensional shape of the target area.

Method used

An image segmentation model based on the three-dimensional cascade attention mechanism is adopted to achieve comprehensive capture of target area features by extracting multi-scale features and long-range dependencies in the original three-dimensional medical images.

Benefits of technology

The accuracy of the three-dimensional segmentation mask image of the target area is improved, and the three-dimensional shape of the target area is accurately expressed, which significantly improves the accuracy of image segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374637A_ABST
    Figure CN120374637A_ABST
Patent Text Reader

Abstract

The invention discloses an image segmentation method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring a to-be-segmented three-dimensional medical original image containing a target area; inputting a to-be-segmented three-dimensional medical original image containing the target area into a pre-trained three-dimensional image segmentation model to obtain a target area three-dimensional segmentation mask image corresponding to the to-be-segmented three-dimensional medical original image containing the target area output by the three-dimensional image segmentation model; wherein the three-dimensional image segmentation model obtains a target region three-dimensional segmentation mask image based on three-dimensional target region features, extracted by a three-dimensional cascade attention mechanism, in the to-be-segmented three-dimensional medical original image containing the target region. According to the technical scheme, through the three-dimensional image segmentation model based on the three-dimensional cascading attention mechanism, comprehensive capture of the three-dimensional target area features is realized, so that the three-dimensional shape of the target area is accurately expressed, and the precision of the three-dimensional segmentation mask image of the target area is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an image segmentation method, apparatus, electronic device, and storage medium. Background Art

[0002] With the development of image segmentation technology, more and more image segmentation technologies have been widely applied to the field of medical image processing.

[0003] In the prior art, when image segmentation technology is applied to medical image processing, there is a problem of low image segmentation accuracy. Summary of the Invention

[0004] The present invention provides an image segmentation method, apparatus, electronic device, and storage medium to improve the segmentation accuracy of medical images.

[0005] According to one aspect of the present invention, an image segmentation method is provided, including:

[0006] Obtaining a three-dimensional medical original image containing a target region to be segmented;

[0007] Inputting the three-dimensional medical original image containing the target region to be segmented into a pre-trained three-dimensional image segmentation model to obtain a target region three-dimensional segmentation mask image corresponding to the three-dimensional medical original image containing the target region output by the three-dimensional image segmentation model;

[0008] Wherein, the three-dimensional image segmentation model obtains the target region three-dimensional segmentation mask image based on the three-dimensional target region features extracted from the three-dimensional medical original image containing the target region to be segmented by a three-dimensional cascaded attention mechanism.

[0009] According to another aspect of the present invention, an image segmentation apparatus is provided, including:

[0010] A three-dimensional medical original image acquisition module for obtaining a three-dimensional medical original image containing a target region to be segmented;

[0011] A target region segmentation module for inputting the three-dimensional medical original image containing the target region to be segmented into a pre-trained three-dimensional image segmentation model to obtain a target region three-dimensional segmentation mask image corresponding to the three-dimensional medical original image containing the target region output by the three-dimensional image segmentation model;

[0012] Wherein, the three-dimensional image segmentation model obtains the target region three-dimensional segmentation mask image based on the three-dimensional target region features extracted from the three-dimensional medical original image containing the target region to be segmented by a three-dimensional cascaded attention mechanism.

[0013] According to another aspect of the present invention, there is provided an electronic device, which includes:

[0014] at least one processor;

[0015] and a memory communicatively connected to the at least one processor;

[0016] wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the image segmentation method according to any embodiment of the present invention.

[0017] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for implementing the image segmentation method according to any embodiment of the present invention when executed by a processor.

[0018] The technical solution of the embodiment of the present invention can effectively extract multi-scale features and long-range dependencies in three-dimensional medical original images through a three-dimensional image segmentation model based on a three-dimensional cascaded attention mechanism, achieve comprehensive capture of three-dimensional target region features, thereby accurately expressing the three-dimensional shape of the target region, and improving the accuracy of the three-dimensional segmentation mask image of the target region.

[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 is a flowchart of an image segmentation method provided in Embodiment 1 of the present invention;

[0022] Figure 2 is a flowchart of an image segmentation method provided in Embodiment 2 of the present invention;

[0023] Figure 3 is a flowchart of an image segmentation method provided in Embodiment 3 of the present invention;

[0024] Figure 4A is a network framework diagram of a three-dimensional image segmentation model provided in an embodiment of the present invention;

[0025] Figure 4B It is a network framework diagram of a three-dimensional convolutional encoder provided according to an embodiment of the present invention;

[0026] Figure 4C It is a network framework diagram of a three-dimensional convolutional attention module provided according to an embodiment of the present invention;

[0027] Figure 4D It is a network framework diagram of a channel attention module provided according to an embodiment of the present invention;

[0028] Figure 4E It is a network framework diagram of a three-dimensional upsampling module provided according to an embodiment of the present invention;

[0029] Figure 5A It is an MRI image containing a prostate region provided according to an embodiment of the present invention;

[0030] Figure 5B It is a true mask image of a prostate region provided according to an embodiment of the present invention;

[0031] Figure 5C It is a three-dimensional segmentation mask image of a prostate region provided according to an embodiment of the present invention;

[0032] Figure 6 It is a structural schematic diagram of an image segmentation device provided according to Embodiment 4 of the present invention;

[0033] Figure 7 It is a structural schematic diagram of an electronic device for implementing the image segmentation method of the embodiment of the present invention. Detailed implementation manners

[0034] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0035] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. The acquisition, storage, use, processing, etc. of the data in the technical solution of this application all comply with the relevant regulations of national laws and regulations.

[0036] Embodiment 1

[0037] Figure 1 The figure is a flowchart of an image segmentation method provided for Embodiment 1 of the present invention. This embodiment is applicable to the case of segmenting the target area of three-dimensional medical images. This method can be executed by an image segmentation device, which can be implemented in the form of hardware and / or software, and the image segmentation device can be configured in a terminal and / or a server. As Figure 1 shown, the method includes:

[0038] S110. Obtain a three-dimensional medical original image including a target area to be segmented.

[0039] Among them, the three-dimensional medical original image refers to a three-dimensional medical image including a target area. Exemplarily, the three-dimensional medical original image can be a Computed Tomography (CT) image or a Magnetic Resonance Imaging (MRI) image, etc., which is not specifically limited herein. The target area can be any area of the detected object's body. For example, the target area can be areas such as the brain, liver, prostate, etc., which is not specifically limited herein.

[0040] Specifically, the three-dimensional medical original image including the target area to be segmented can be read from a preset storage path of the local electronic device, or downloaded from other devices or the cloud that are communicatively connected to the electronic device, which is not specifically limited herein.

[0041] S120. Input the three-dimensional medical original image containing the target region to be segmented into the pre-trained three-dimensional image segmentation model, and obtain the target region three-dimensional segmentation mask image corresponding to the three-dimensional medical original image containing the target region output by the three-dimensional image segmentation model.

[0042] Among them, the three-dimensional image segmentation model can extract the three-dimensional target region features in the three-dimensional medical original image containing the target region to be segmented based on the three-dimensional cascaded attention mechanism, and obtain the target region three-dimensional segmentation mask image.

[0043] Specifically, the three-dimensional medical original image containing the target region to be segmented is used as input data and input into the three-dimensional image segmentation model based on the three-dimensional cascaded attention mechanism. The three-dimensional image segmentation model based on the three-dimensional cascaded attention mechanism extracts the three-dimensional target region features in the three-dimensional medical original image containing the target region to be segmented, and based on the three-dimensional target region features, segments and obtains the target region three-dimensional segmentation mask image and outputs it.

[0044] It should be noted that through the three-dimensional image segmentation model based on the three-dimensional cascaded attention mechanism, multi-scale features and long-range dependencies in the three-dimensional medical original image can be effectively extracted, the three-dimensional target region features can be comprehensively captured, so as to accurately express the three-dimensional shape of the target region and improve the accuracy of the target region three-dimensional segmentation mask image.

[0045] In the embodiment of the present disclosure, the three-dimensional image segmentation model can be trained by a large number of three-dimensional medical sample images. Specifically, in the trained neural network model, the three-dimensional medical sample images can be pre-processed in advance, the features of the target region are extracted from the pre-processed three-dimensional medical sample images, and then the model parameters in the neural network model are trained based on the extracted three-dimensional target region features. By continuously adjusting the model parameters, the distance deviation between the output result of the model and the true mask image of the target region gradually decreases and tends to be stable.

[0046] The technical solution of the embodiment of the present invention can effectively extract multi-scale features and long-range dependencies in the three-dimensional medical original image through the three-dimensional image segmentation model based on the three-dimensional cascaded attention mechanism, realizes the comprehensive capture of the three-dimensional target region features, thus accurately expressing the three-dimensional shape of the target region and improving the accuracy of the target region three-dimensional segmentation mask image.

[0047] Embodiment Two

[0048] Figure 2The flowchart of an image segmentation method provided in the second embodiment of the present invention. The method in this embodiment can be combined with each optional solution in the image segmentation method provided in the above embodiment. The image segmentation method provided in this embodiment is further optimized. Optionally, before obtaining the three-dimensional medical original image including the target region to be segmented, the method further includes: obtaining a plurality of three-dimensional medical sample images including the target region and the target region true mask images corresponding to each of the three-dimensional medical sample images including the target region; using each of the three-dimensional medical sample images including the target region and the target region true mask images corresponding to each of the three-dimensional medical sample images including the target region as sample data to train the initially established neural network model, and obtaining a three-dimensional image segmentation model.

[0049] As Figure 2 shown, the method includes:

[0050] S210. Obtain a plurality of three-dimensional medical sample images including the target region and the target region true mask images corresponding to each of the three-dimensional medical sample images including the target region.

[0051] S220. Use each of the three-dimensional medical sample images including the target region and the target region true mask images corresponding to each of the three-dimensional medical sample images including the target region as sample data to train the initially established neural network model, and obtain a three-dimensional image segmentation model.

[0052] Specifically, the model parameters of the neural network model can be trained by using the three-dimensional medical sample images including the target region. By continuously adjusting the parameter values of the model parameters, the distance deviation between the target region mask image predicted by the model and the target region true mask image is gradually reduced until the distance deviation tends to be stable, and a three-dimensional image segmentation model is obtained.

[0053] In an alternative implementation of the embodiments of the present disclosure, the neural network model includes a three-dimensional convolutional encoder and a three-dimensional cascaded attention decoder; correspondingly, each three-dimensional medical sample image containing the target region and the true mask image of the target region corresponding to each three-dimensional medical sample image containing the target region are used as sample data to train the initially established neural network model to obtain a three-dimensional image segmentation model, including: for any three-dimensional medical sample image containing the target region, the three-dimensional medical sample image containing the target region is sequentially input into the three-dimensional convolutional encoder to obtain a convolutional encoded feature image corresponding to the three-dimensional medical sample image containing the target region; the convolutional encoded feature image is input into the three-dimensional cascaded attention decoder to obtain a predicted target region mask image; the model loss is determined based on the predicted target region mask image and the true mask image of the target region corresponding to the three-dimensional medical sample image containing the target region, and the network parameters in the neural network model are updated based on the model loss until the training ends to obtain a three-dimensional image segmentation model.

[0054] Among them, when the neural network model includes multiple levels, the number of three-dimensional convolutional encoders can be multiple, which can encode the three-dimensional medical sample image. Specifically, any level of the three-dimensional convolutional encoder can be composed of multiple three-dimensional convolutional modules. The three-dimensional cascaded attention decoder can be composed of multiple associated level decoders, which are used to decode the encoded three-dimensional medical sample image. Specifically, any level decoder can be composed of a three-dimensional convolutional module, a three-dimensional convolutional attention module, a three-dimensional upsampling module, etc.

[0055] It should be noted that the three-dimensional convolutional module can accurately segment the three-dimensional medical sample image. In addition, the three-dimensional cascaded attention decoder can effectively extract multi-scale features and long-term dependencies in the three-dimensional medical sample image, thereby improving the training effect of the three-dimensional image segmentation model.

[0056] Optionally, the three-dimensional cascaded attention decoder includes multiple levels, and any level includes a first three-dimensional convolution module, a three-dimensional convolutional attention module, and a second three-dimensional convolution module; correspondingly, inputting the convolutionally encoded feature image into the three-dimensional cascaded attention decoder to obtain a predicted target region mask image includes: for any level, inputting the convolutionally encoded feature image into the first three-dimensional convolution module to obtain a three-dimensional convolutional feature image corresponding to the convolutionally encoded feature image; when the convolutionally encoded feature image is the highest-level feature image, inputting the three-dimensional convolutional feature image into the three-dimensional convolutional attention module to obtain an attention feature image corresponding to the three-dimensional convolutional feature image; when the convolutionally encoded feature image is a non-highest-level feature image, concatenating the three-dimensional convolutional feature image with the upsampled feature image of the previous level, and inputting the concatenated image into the three-dimensional convolutional attention module to obtain an attention feature image corresponding to the three-dimensional convolutional feature image, where the upsampled feature image of the previous level is obtained by performing three-dimensional upsampling on the attention feature image of the previous level; inputting the attention feature image into the second three-dimensional convolution module to obtain a predicted mask image of the current level; determining the predicted target region mask image based on the predicted mask images of each level.

[0057] It should be noted that the above three-dimensional convolutional attention module can effectively extract multi-scale features and long-range dependencies in the three-dimensional medical sample image.

[0058] Optionally, the three-dimensional convolutional attention module includes a channel attention module, a third three-dimensional convolution module, a first batch normalization module, a first activation module, a fourth three-dimensional convolution module, a second batch normalization module, and a second activation module; the channel attention module includes an adaptive average pooling module, an adaptive max pooling module, a fourth three-dimensional convolution module, a fifth three-dimensional convolution module, a third activation module, a fourth activation module, a sixth three-dimensional convolution module, a seventh three-dimensional convolution module, and a fifth activation module.

[0059] It should be noted that the network architecture form of the above three-dimensional convolutional attention module is set for three-dimensional medical sample image segmentation and can effectively extract multi-scale features and long-range dependencies in the three-dimensional medical sample image.

[0060] Optionally, determining the model loss based on the predicted target region mask image and the target region ground truth mask image corresponding to the three-dimensional medical sample image including the target region includes: determining the level model loss based on the predicted mask images of each level and the target region ground truth mask image corresponding to the three-dimensional medical sample image including the target region; determining the overall model loss based on the predicted target region mask image and the target region ground truth mask image corresponding to the three-dimensional medical sample image including the target region; determining the model loss based on the level model loss and the overall model loss.

[0061] Exemplarily, the loss function for determining the model loss can be the Dice loss function. The hierarchical model loss can be represented by Dice(p i , M), and the overall model loss can be represented by Dice(output, M). Thus, the model loss is loss = Dice(p i , M) + Dice(output, M); where p i represents the predicted mask image of the i-th layer, output represents the predicted target region mask image, and M represents the true mask image of the target region. The calculation formula of the Dice loss function is:

[0062]

[0063] where x can be p i or output, N represents the total number of pixels in the image, M represents the true mask image of the target region, x′ n represents the pixels in x, and m n represents the pixels in M.

[0064] S230. Obtain a three-dimensional medical original image containing a target region to be segmented.

[0065] S240. Input the three-dimensional medical original image containing the target region to be segmented into a pre-trained three-dimensional image segmentation model, and obtain a target region three-dimensional segmentation mask image corresponding to the three-dimensional medical original image output by the three-dimensional image segmentation model.

[0066] The technical solution of the embodiment of the present invention trains the model parameters of the neural network model through a three-dimensional medical sample image containing a target region. By continuously adjusting the parameter values of the model parameters, the distance deviation between the target region mask image predicted by the model and the true mask image of the target region is gradually reduced until the distance deviation tends to be stable, and a three-dimensional image segmentation model is obtained, laying a foundation for the segmentation of three-dimensional medical images, thereby realizing the accurate segmentation of three-dimensional medical original images.

[0067] Embodiment III

[0068] Figure 3 It is a flowchart of an image segmentation method provided by Embodiment III of the present invention. The method of this embodiment can be combined with each optional solution in the image segmentation method provided in the above embodiment. The image segmentation method provided in this embodiment is further optimized. Optionally, the target region is the prostate region in the three-dimensional medical original image.

[0069] As Figure 3 shown, the method includes:

[0070] S310. Obtain a three-dimensional medical original image containing the prostate region to be segmented.

[0071] S320. Input the three-dimensional medical original image containing the prostate region to be segmented into a pre-trained three-dimensional image segmentation model, and obtain a three-dimensional segmentation mask image of the prostate region corresponding to the three-dimensional medical original image output by the three-dimensional image segmentation model.

[0072] Among them, the three-dimensional image segmentation model obtains a three-dimensional segmentation mask image of the prostate region based on the three-dimensional prostate region features in the three-dimensional medical original image containing the prostate region to be segmented extracted by the three-dimensional cascaded attention mechanism.

[0073] Exemplarily, the three-dimensional medical sample image can be a 3D MRI image containing the prostate region, and the target region true mask image can be a manually marked prostate mask image corresponding to the 3D MRI image. Figure 4A It is a network framework diagram of a three-dimensional image segmentation model provided by an embodiment of the present invention; specifically, the image segmentation method includes the following steps:

[0074] S301. Obtain a data set composed of a 3D MRI image containing the prostate region and a manually marked prostate mask image corresponding to the 3D MRI image.

[0075] S302. Preprocess all the 3D MRI images containing the prostate region in the data set respectively to obtain a preprocessed data set. Among them, the data set preprocessing includes one or more of the following steps:

[0076] S3021. Crop the 3D MRI image containing the prostate region.

[0077] S3022. Smooth and denoise the 3D MRI image containing the prostate region. Specifically, a Gaussian filter can be used to smooth and denoise the 3D MRI image containing the prostate region, so as to reduce the noise of the 3D MRI image.

[0078] S3023. Perform gray normalization on the 3D MRI image containing the prostate region to map the gray value of the 3D MRI image containing the prostate region to the range of 0-1.

[0079] S3024. Perform image enhancement processing on the 3D MRI image containing the prostate region. Specifically, one or more of methods such as rotation, flipping, scaling, adding noise, etc. can be used to perform image enhancement processing on the 3D MRI image containing the prostate region.

[0080] S303. Initialize the parameters in the neural network model, and input the 3D MRI images containing the prostate region in the preprocessed dataset into the initially established neural network model for training, and update the parameters in the neural network model to obtain a trained 3D image segmentation model. Among them, the specific process of training the neural network model includes:

[0081] S3031. Input the 3D MRI images containing the prostate region into the 3D convolutional encoder in sequence. The 3D MRI images pass through the 3D convolutional encoder of four levels in sequence. The 3D convolutional encoder of each level can extract convolutional encoding feature images with different resolutions, which are respectively denoted as X1, X2, X3, and X4. Figure 4B is the network framework diagram of a 3D convolutional encoder provided by an embodiment of the present invention; for the 3D convolutional encoder of the i-th layer, the corresponding X i The calculation process is:

[0082] X i = ConvEncoder i (X i-1 )

[0083] Among them, ConvEncoder i represents the 3D convolutional encoder of the i-th layer, and X i represents the output of the 3D convolutional encoder of the i-th layer. When i = 0, X0 represents the original 3D MRI image containing the prostate region.

[0084] S3033. X4 can be input into the 3D cascaded attention decoder. At the same time, X3, X2, and X1 are input as skip connections.

[0085] S3034. For any level i of the 3D cascaded attention decoder, it includes the following steps:

[0086] S30341. Perform 3D convolutional encoding on the convolutional encoding feature image to generate a 3D convolutional feature image. The formula is:

[0087] Y i = σ(BN(Conv(X i )));

[0088] Among them, σ represents the Sigmoid activation function, BN represents the batch normalization module, Conv represents the 3D convolutional module, X i represents the convolutional encoding feature image of the i-th level, and Y i represents the 3D convolutional feature image of the i-th level.

[0089] S30342. Concatenate the three-dimensional convolutional feature image with the upsampled feature image of the previous layer, and input it into the three-dimensional convolutional attention module to generate an attention feature image. Figure 4C is the network framework diagram of a three-dimensional convolutional attention module provided by an embodiment of the present invention; Figure 4D is the network framework diagram of a channel attention module provided by an embodiment of the present invention; its calculation formula is:

[0090]

[0091] where represents the concatenation operation, X′ i-1 represents the upsampled feature image of the previous level, Z i represents the attention feature image of the i-th level, and CAM represents the three-dimensional convolutional attention module. It can be understood that when i = 1, since there is no upsampled feature image of the previous layer, Y i can be directly input into the three-dimensional convolutional attention module to generate Z i , and its calculation formula is:

[0092] Z i = CAM(Y i );

[0093] S30343. Input the attention feature image into the three-dimensional upsampling module to obtain an upsampled feature image. Figure 4E is the network framework diagram of a three-dimensional upsampling module provided by an embodiment of the present invention; its calculation formula is:

[0094] X′ i = UpConv(Z i );

[0095] where X′ i represents the upsampled feature image of the i-th level, and UpConv represents the three-dimensional upsampling module.

[0096] S30344. Perform three-dimensional convolutional encoding on the attention feature image to obtain the predicted mask image of the current level, and its calculation formula is:

[0097] p i = Up(σ(BN(Conv(Z i ))));

[0098] where p i represents the predicted mask image of the i-th level, σ represents the Sigmoid activation function, BN represents the batch normalization module, Conv represents the three-dimensional convolutional module, Up represents the upsampling layer, and the upsampling multiples of different levels are different. The upsampling multiples of the 1st, 2nd, 3rd, and 4th layers are 32, 16, 8, and 4 respectively.

[0099] S3035. Weightedly superimpose the prediction mask images of the four levels to obtain the final predicted target region mask image, and its formula is:

[0100] output = w1p1 + w2p2 + w3p3 + w4p4;

[0101] where w i represents the weight coefficient of p i and output represents the final predicted target region mask image.

[0102] S3036. Determine the model loss, and then update the parameters of the neural network model based on the model loss through backpropagation training. The calculation formula of the model loss is:

[0103] loss = Dice(p1, M) + Dice(p2, M) + Dice(p3, M) + Dice(p4, M) + Dice(output, M)

[0104] where M represents the manually marked prostate mask image corresponding to the 3D MRI image containing the prostate region, and Dice is the Dice loss function.

[0105] After completing the training of the three-dimensional image segmentation model, preprocess the new 3D MRI image containing the prostate region and input it into the trained three-dimensional image segmentation model. The three-dimensional image segmentation model encodes and decodes the input image to generate a three-dimensional segmentation mask image of the prostate region, and this three-dimensional segmentation mask image of the prostate region can be a three-dimensional binary mask image.

[0106] Figure 5A is an MRI image containing the prostate region provided by an embodiment of the present invention, Figure 5B is a true mask image of the prostate region provided by an embodiment of the present invention, Figure 5C is a three-dimensional segmentation mask image of the prostate region provided by an embodiment of the present invention. From Figure 5A , Figure 5B and Figure 5C , it can be seen that the difference between the three-dimensional segmentation mask image of the prostate region obtained by the image segmentation method of this embodiment and the true mask image of the prostate region is small, achieving high-precision prostate segmentation.

[0107] Compared with the prior art, the embodiments of the present disclosure adopt a three-dimensional cascaded attention decoder structure, which makes full use of three-dimensional information, so as to more accurately represent the three-dimensional shape of the prostate, and significantly improves the three-dimensional modeling ability and accuracy of prostate segmentation. It should be emphasized that through the three-dimensional cascaded attention mechanism, the three-dimensional target region features can be comprehensively captured, the three-dimensional shape variation of the prostate can be accurately expressed, and the error caused by only relying on two-dimensional image information is greatly reduced. In addition, the embodiments of the present disclosure adopt an end-to-end three-dimensional training method, avoiding complex manual three-dimensional feature engineering and simplifying the development.

[0108] Embodiment 4

[0109] Figure 6 The following is a schematic structural diagram of an image segmentation device provided in Embodiment 4 of the present invention. As Figure 6 shown, the device includes:

[0110] A three-dimensional medical original image acquisition module 410, configured to acquire a three-dimensional medical original image including a target region to be segmented;

[0111] A target region segmentation module 420, configured to input the three-dimensional medical original image including the target region to be segmented into a pre-trained three-dimensional image segmentation model, and obtain a target region three-dimensional segmentation mask image corresponding to the three-dimensional medical original image including the target region output by the three-dimensional image segmentation model;

[0112] Wherein, the three-dimensional image segmentation model obtains a target region three-dimensional segmentation mask image based on the three-dimensional target region features in the three-dimensional medical original image including the target region extracted by the three-dimensional cascaded attention mechanism.

[0113] The technical solution of the embodiments of the present invention can effectively extract multi-scale features and long-range dependencies in the three-dimensional medical original image through the three-dimensional image segmentation model based on the three-dimensional cascaded attention mechanism, realizes the comprehensive capture of the three-dimensional target region features, thereby accurately expressing the three-dimensional shape of the target region, and improves the accuracy of the target region three-dimensional segmentation mask image.

[0114] In some optional embodiments, the device further includes:

[0115] A model training sample acquisition module, configured to acquire a plurality of three-dimensional medical sample images including the target region and the target region true mask images corresponding to the three-dimensional medical sample images including the target region;

[0116] A three-dimensional image segmentation model training module is used to train an initially established neural network model with each three-dimensional medical sample image containing a target region and the corresponding ground truth mask image of the target region in each three-dimensional medical sample image as sample data, so as to obtain a three-dimensional image segmentation model.

[0117] In some optional embodiments, the neural network model includes a three-dimensional convolutional encoder and a three-dimensional cascaded attention decoder;

[0118] Correspondingly, the three-dimensional image segmentation model training module includes:

[0119] A three-dimensional convolutional encoding unit is configured to input, for any three-dimensional medical sample image containing a target region, the three-dimensional medical sample image containing the target region into the three-dimensional convolutional encoder to obtain a convolutional encoding feature image corresponding to the three-dimensional medical sample image containing the target region;

[0120] A three-dimensional cascaded attention decoding unit is configured to input the convolutional encoding feature image into the three-dimensional cascaded attention decoder to obtain a predicted target region mask image;

[0121] A network parameter updating unit is configured to determine a model loss based on the predicted target region mask image and the ground truth mask image of the target region corresponding to the three-dimensional medical sample image containing the target region, and update the network parameters in the neural network model based on the model loss until the training ends to obtain a three-dimensional image segmentation model.

[0122] In some optional embodiments, the three-dimensional cascaded attention decoder includes multiple levels, and any level includes a first three-dimensional convolutional module, a three-dimensional convolutional attention module, and a second three-dimensional convolutional module;

[0123] Correspondingly, the three-dimensional cascaded attention decoding unit is further specifically configured to:

[0124] For any level, input the convolutionally encoded feature image into the first three-dimensional convolution module to obtain a three-dimensional convolution feature image corresponding to the convolutionally encoded feature image; in the case where the convolutionally encoded feature image is the highest-level feature image, input the three-dimensional convolution feature image into the three-dimensional convolution attention module to obtain an attention feature image corresponding to the three-dimensional convolution feature image; in the case where the convolutionally encoded feature image is a non-highest-level feature image, splice the three-dimensional convolution feature image with the upsampled feature image of the previous level, and input the spliced image into the three-dimensional convolution attention module to obtain an attention feature image corresponding to the three-dimensional convolution feature image, where the upsampled feature image of the previous level is obtained by performing three-dimensional upsampling on the attention feature image of the previous level; input the attention feature image into the second three-dimensional convolution module to obtain a predicted mask image for the current level;

[0125] Determine a predicted target region mask image based on the predicted mask images of each level.

[0126] In some alternative embodiments, the three-dimensional convolution attention module includes a channel attention module, a third three-dimensional convolution module, a first batch normalization module, a first activation module, a fourth three-dimensional convolution module, a second batch normalization module, and a second activation module;

[0127] The channel attention module includes an adaptive average pooling module, an adaptive max pooling module, a fourth three-dimensional convolution module, a fifth three-dimensional convolution module, a third activation module, a fourth activation module, a sixth three-dimensional convolution module, a seventh three-dimensional convolution module, and a fifth activation module.

[0128] In some alternative embodiments, the network parameter update unit can also be used for:

[0129] Determine a hierarchical model loss based on the predicted mask images of each level and the target region ground truth mask image corresponding to the three-dimensional medical sample image including the target region;

[0130] Determine an overall model loss based on the predicted target region mask image and the target region ground truth mask image corresponding to the three-dimensional medical sample image including the target region;

[0131] Determine a model loss based on the hierarchical model loss and the overall model loss.

[0132] In some alternative embodiments, the target region is the prostate region in the three-dimensional medical original image.

[0133] The image segmentation device provided by the embodiments of the present invention can execute the image segmentation method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0134] Example 5

[0135] Figure 7 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0136] As Figure 7 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The I / O interface 15 is also connected to the bus 14.

[0137] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0138] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as an image segmentation method, which includes:

[0139] Obtain a three-dimensional medical original image containing a target region to be segmented;

[0140] Input the three-dimensional medical original image containing the target region to be segmented into a pre-trained three-dimensional image segmentation model, and obtain a target region three-dimensional segmentation mask image corresponding to the three-dimensional medical original image containing the target region output by the three-dimensional image segmentation model;

[0141] Wherein, the three-dimensional image segmentation model obtains a target region three-dimensional segmentation mask image based on the three-dimensional target region features in the three-dimensional medical original image containing the target region to be segmented extracted by the three-dimensional cascaded attention mechanism.

[0142] In some embodiments, the image segmentation method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by the processor 11, one or more steps of the image segmentation method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the image segmentation method by any other suitable means (e.g., by means of firmware).

[0143] The various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs, the one or more computer programs can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0144] A computer program for implementing the method of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0145] In the context of the present invention, a computer-readable storage medium may be a tangible medium that can contain, or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium may be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0146] In order to provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, speech input, or tactile input).

[0147] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0148] The computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0149] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0150] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An image segmentation method, characterized in that, Comprising: Obtain a three-dimensional medical original image containing a target region to be segmented; Input the three-dimensional medical original image containing the target region to be segmented into a pre-trained three-dimensional image segmentation model, and obtain a target region three-dimensional segmentation mask image corresponding to the three-dimensional medical original image containing the target region output by the three-dimensional image segmentation model; Wherein, the three-dimensional image segmentation model obtains a target region three-dimensional segmentation mask image based on three-dimensional target region features in the three-dimensional medical original image containing the target region extracted by a three-dimensional cascaded attention mechanism.

2. The method according to claim 1, wherein Before the step of obtaining the three-dimensional medical original image containing the target region to be segmented, the method further includes: Obtain a plurality of three-dimensional medical sample images containing the target region and target region ground truth mask images corresponding to the three-dimensional medical sample images containing the target region; Use each of the three-dimensional medical sample images containing the target region and the target region ground truth mask image corresponding to each of the three-dimensional medical sample images containing the target region as sample data to train an initially established neural network model, and obtain a three-dimensional image segmentation model.

3. The method according to claim 2, characterized in that, The neural network model includes a three-dimensional convolutional encoder and a three-dimensional cascaded attention decoder; Correspondingly, the step of using each of the three-dimensional medical sample images containing the target region and the target region ground truth mask image corresponding to each of the three-dimensional medical sample images containing the target region as sample data to train an initially established neural network model, and obtain a three-dimensional image segmentation model, includes: For any three-dimensional medical sample image containing the target region, input the three-dimensional medical sample image containing the target region into the three-dimensional convolutional encoder to obtain a convolutional encoded feature image corresponding to the three-dimensional medical sample image containing the target region; Input the convolutional encoded feature image into the three-dimensional cascaded attention decoder to obtain a predicted target region mask image; Determine a model loss based on the predicted target region mask image and the target region ground truth mask image corresponding to the three-dimensional medical sample image containing the target region, and update network parameters in the neural network model based on the model loss until the training ends to obtain a three-dimensional image segmentation model.

4. The method according to claim 3, characterized in that, The three-dimensional cascaded attention decoder includes multiple levels, and any level includes a first three-dimensional convolutional module, a three-dimensional convolutional attention module, and a second three-dimensional convolutional module; Correspondingly, the step of inputting the convolutional encoded feature image into the three-dimensional cascaded attention decoder to obtain a predicted target region mask image, includes: For any level, input the convolutionally encoded feature image into the first three-dimensional convolution module to obtain a three-dimensional convolution feature image corresponding to the convolutionally encoded feature image; when the convolutionally encoded feature image is the highest-level feature image, input the three-dimensional convolution feature image into the three-dimensional convolution attention module to obtain an attention feature image corresponding to the three-dimensional convolution feature image; when the convolutionally encoded feature image is a non-highest-level feature image, concatenate the three-dimensional convolution feature image with the upsampled feature image of the previous level, and input the concatenated image into the three-dimensional convolution attention module to obtain an attention feature image corresponding to the three-dimensional convolution feature image, where the upsampled feature image of the previous level is obtained by performing three-dimensional upsampling on the attention feature image of the previous level; input the attention feature image into the second three-dimensional convolution module to obtain a predicted mask image for the current level; Determine a predicted target region mask image based on the predicted mask images of each level.

5. The method according to claim 4, wherein The three-dimensional convolution attention module includes a channel attention module, a third three-dimensional convolution module, a first batch normalization module, a first activation module, a fourth three-dimensional convolution module, a second batch normalization module, and a second activation module; The channel attention module includes an adaptive average pooling module, an adaptive max pooling module, a fourth three-dimensional convolution module, a fifth three-dimensional convolution module, a third activation module, a fourth activation module, a sixth three-dimensional convolution module, a seventh three-dimensional convolution module, and a fifth activation module.

6. The method according to claim 4, wherein Determining the model loss based on the predicted target region mask image and the target region ground truth mask image corresponding to the three-dimensional medical sample image containing the target region includes: Determine the level model loss based on the predicted mask images of each level and the target region ground truth mask image corresponding to the three-dimensional medical sample image containing the target region; Determine the overall model loss based on the predicted target region mask image and the target region ground truth mask image corresponding to the three-dimensional medical sample image containing the target region; Determine the model loss based on the level model loss and the overall model loss.

7. According to the method as claimed in any one of claims 1 to 6, characterized in that The target region is the prostate region in the three-dimensional medical original image.

8. An image segmentation device, characterized in that, Includes: A three-dimensional medical original image acquisition module for acquiring a three-dimensional medical original image containing the target region to be segmented; A target region segmentation module for inputting the three-dimensional medical original image containing the target region to be segmented into a pre-trained three-dimensional image segmentation model to obtain a target region three-dimensional segmentation mask image corresponding to the three-dimensional medical original image containing the target region output by the three-dimensional image segmentation model; Wherein, the three-dimensional image segmentation model obtains a target region three-dimensional segmentation mask image based on the three-dimensional target region features in the three-dimensional medical original image containing the target region extracted by the three-dimensional cascaded attention mechanism.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the image segmentation method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for implementing the image segmentation method according to any one of claims 1-7 when the computer instructions are executed by a processor.