Three-dimensional medical image segmentation method and device, equipment and storage medium

By cutting and downsampling three-dimensional medical images, non-local and local features are extracted, and feature fusion and self-attention calculation are performed through the INLF model, the problem of loss of local details and global context information in the image chunking method is solved, and the accuracy of medical image segmentation is improved.

CN120015250APending Publication Date: 2025-05-16NINGBO INST OF NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510093150.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing image chunking methods lead to loss of local details or loss of global context information in medical image segmentation, limiting the accuracy of segmentation.

Method used

A three-dimensional medical image segmentation method is adopted to extract non-local and local features by cutting and downsampling of three-dimensional medical images, and feature fusion and self-attention calculation are performed through the INLF model to generate segmented images.

Benefits of technology

This method can simultaneously extract local detail features and global semantic features in three-dimensional medical images, improve the accuracy of segmentation tasks, and avoid the problem of loss of local details and global context information in image tiling method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015250A_ABST
    Figure CN120015250A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional medical image segmentation method and device, equipment and a storage medium, and relates to the technical field of image segmentation, and the method comprises the following steps: carrying out the primary cutting of a three-dimensional medical image, obtaining a non-local image, carrying out the secondary cutting of the non-local image, and obtaining a local image; and performing down-sampling on the non-local image, and inputting the non-local image and the local image after down-sampling into the INLF model to obtain a segmented image. Therefore, local features with detail information and non-local features with more semantic information in the three-dimensional medical image can be extracted at the same time, and the accuracy of the segmentation task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image segmentation technology, and in particular to a three-dimensional medical image segmentation method, device, equipment and storage medium. Background Art

[0002] At present, with the advancement of medical imaging technology and the large number of imaging devices put into use, many hospitals have accumulated a huge number of medical images. These image data, as well as the corresponding image interpretation, disease diagnosis, treatment plan and prognosis and efficacy evaluation, constitute the medical imaging big data about the occurrence, evolution and treatment of such diseases. Therefore, it is an urgent and meaningful research work to make full use of the existing huge computing power, combine the analysis of medical imaging big data with deep learning technology, conduct in-depth research, and discover new problems, new theories, new methods and new tools.

[0003] In medical image segmentation, image data is often more critical than the model in affecting the results. A model with strong generalization ability will inevitably have to process massive amounts of data, but medical images collected by different methods may bring about these problems:

[0004] (1) The sizes of images collected from different datasets are inconsistent and cannot be trained uniformly. If the images are forcibly scaled, the images may be distorted and detail information may be lost.

[0005] (2) The scale of 3D medical images is very large. When using these data for training, there will be a problem of insufficient video memory. Taking the medical image segmentation dataset LiTS and the real-world semantic segmentation dataset ADE20k as examples, the size of an RGB image is usually no more than 1MB, while a CT image in the LiTS dataset may be as large as hundreds of MB. Such large data is difficult to train directly using a single GPU.

[0006] (3) Even when there is sufficient video memory, the model may have difficulty capturing the details of large-scale medical images. Compared with hundreds of MB of CT images, liver tumors only occupy a small part of them. It is undoubtedly a very difficult task to locate these small targets from such a large background.

[0007] (4) Medical image segmentation data is usually privacy-sensitive, and the data is difficult to collect and annotate, and the available data resources are very limited. Real-world semantic segmentation datasets can reach tens of thousands of data points, while medical image segmentation datasets usually only have a few hundred data points.

[0008] The current solution to these problems is to divide the image into blocks. That is, use a sliding window to cut (Crop) a two-dimensional or three-dimensional image into many small patches of the same size as training data. In this way, the image size is unified, and the small image patch is equivalent to the local enlargement of the image, which reduces the video memory and greatly enriches the data set. However, while solving the above problems, image segmentation will cause the loss of local details of the image or the loss of global context information, thereby limiting the accuracy of subsequent image segmentation. Summary of the invention

[0009] The present invention provides a three-dimensional medical image segmentation method, device, equipment and storage medium, which solves the problem that the existing image segmentation method causes the loss of local details of the image or the loss of global context information, thereby limiting the accuracy of segmentation.

[0010] In a first aspect, the present invention provides a three-dimensional medical image segmentation method, comprising the following steps:

[0011] Perform a first cropping on the three-dimensional medical image to obtain a non-local image, and perform a second cropping on the non-local image to obtain a local image;

[0012] The non-local image is downsampled, and the downsampled non-local image and the local image are input into the INLF model to obtain a segmented image, including:

[0013] Based on the non-local feature encoder, the downsampled non-local image is mapped to obtain the non-local feature; based on the local encoder, the local image is mapped to obtain the local feature;

[0014] The local features and non-local features are expanded based on the reshaping layer, and the last two dimensions of the expanded local features and non-local features are transposed;

[0015] Based on the connection layer, the transposed local features and non-local features are concatenated in the second dimension to obtain concatenated features;

[0016] Based on the multi-head self-attention layer, self-attention is calculated on the spliced ​​features to obtain the attention weighted features;

[0017] Based on the feedforward layer, the attention weighted features and the concatenated features are nonlinearly transformed to obtain the transformed features;

[0018] The expanded transformed features are decoded based on the feature decoder to obtain a segmented image.

[0019] Preferably, the local feature encoder and the non-local encoder have the same structure, both comprising 5 resolution stages, each stage comprising convolution, instance normalization and leaky ReLU functions.

[0020] Preferably, the multi-head self-attention layer maps the concatenated features into query Q, key K and value V through a fully connected layer.

[0021] Preferably, the feed-forward layer is stacked in the form of a first fully connected layer-activation function-second fully connected layer-layer normalization, as shown below:

[0022] FFN(X)=max(0,xW1+b1)W2+b2.

[0023] Where FFN(X) is the output of the feedforward layer, W1 and W2 are the weights of the first fully connected layer and the second fully connected layer, b1 and b2 are the biases of the first fully connected layer and the second fully connected layer, respectively, and x is the concatenated feature and the attention weighted feature.

[0024] Preferably, the feature decoder includes a multi-layer structure, each layer of which is a convolution layer-upsampling layer-convolution layer; in each layer of the feature decoder, features of corresponding layers of the local feature encoder and the non-local feature encoder are transmitted through jump connections.

[0025] Preferably, the non-local feature encoder updates weights by an exponential sliding average method.

[0026] In a second aspect, the present invention further provides a three-dimensional medical image segmentation device, comprising:

[0027] A cropping module, used for cropping the three-dimensional medical image once to obtain a non-local image, and cropping the non-local image twice to obtain a local image;

[0028] An input module is used to downsample the non-local image and input the downsampled non-local image and the local image into the INLF model to obtain a segmented image;

[0029] The input module comprises:

[0030] The encoding module is used to map the downsampled non-local image based on the non-local feature encoder to obtain non-local features; and map the local image based on the local encoder to obtain local features;

[0031] A reshaping module, used for expanding the local features and non-local features based on the reshaping layer, and transposing the last two dimensions of the expanded local features and non-local features;

[0032] A splicing module, used for splicing the transposed local features and non-local features in the second dimension based on the connection layer to obtain splicing features;

[0033] A calculation module is used to perform self-attention calculation on the spliced ​​features based on the multi-head self-attention layer to obtain attention weighted features;

[0034] A transformation module is used to perform nonlinear transformation on the attention weighted features and the concatenated features based on the feedforward layer to obtain transformation features;

[0035] The segmentation module is used to decode the expanded transformation features based on the feature decoder to obtain a segmented image.

[0036] In a third aspect, the present invention further provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned three-dimensional medical image segmentation method when executing the program.

[0037] In a fourth aspect, the present invention further provides a computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned three-dimensional medical image segmentation method is implemented.

[0038] Compared with the prior art, the present invention has the following beneficial effects:

[0039] The present invention first crops the three-dimensional medical image to obtain a non-local image, and downsamples the non-local image. The non-local image is cropped twice to obtain a local image. The non-local image has rich spatial context information, but the resolution of the non-local image is high, and the non-local image needs to be downsampled, and the loss of local detail information caused by downsampling can be supplemented by the local image. Therefore, the present invention can simultaneously extract local features with detail information and non-local features with more semantic information in the three-dimensional medical image, thereby improving the accuracy of the segmentation task. Then the downsampled non-local image and local image are input into the INLF model to obtain a segmented image. The encoder of the INLF model can extract features from the non-local image and the local image, and splice the local features and the non-local features, and then use self-attention to weight the spliced ​​features and introduce non-linear features to enhance the expression ability and further improve the accuracy of the segmentation task. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0041] Figure 1 A schematic diagram of a flow chart of a three-dimensional medical image segmentation method of the present invention;

[0042] Figure 2It is a schematic diagram of the structure of the backbone network of the INLF model of the present invention;

[0043] Figure 3 Schematic diagram of the self-attention module of the INLF model of the present invention;

[0044] Figure 4 Schematic diagram of the data processing process of the multi-head self-attention layer of the INLF model of the present invention. DETAILED DESCRIPTION

[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] The present invention provides a three-dimensional medical image segmentation method, and proposes a non-local feature fusion (INLF) model based on a self-attention mechanism to solve the problem of missing global context information when using image blocks, so as to improve the accuracy of medical image segmentation, comprising the following steps:

[0047] The first step is to perform a first cropping on the three-dimensional medical image to obtain a non-local image, and then perform a second cropping on the non-local image to obtain a local image.

[0048] The non-local image refers to the complete image after a cropping, and the local image refers to a part of the complete image that is directly cropped.

[0049] Step 2: Downsample the non-local image, and input the downsampled non-local image and local image into the INLF model to obtain the segmented image.

[0050] Reference Figure 1 The INLF (Incorporate Non-Local Features) model includes two feature encoders, a reshaping layer, a connection layer, a self-attention module, and a feature decoder. The two feature encoders are the local feature encoder branch E l and the non-local feature encoder branch E n , E n The input is the non-local image X obtained by cropping the original 3D medical image n ∈R B×1×2H×2W×2D , where B is the number of images input in one iteration of training, and H, W, and D are the length, width, and height of the cropped image block.

[0051] E l The input is a non-local image X nThe local image X obtained by secondary cropping l ∈R B×1×H×W×D . Non-local image X n There is richer spatial context information for extracting non-local features, but the non-local image X n The resolution of the non-local image X is high, and directly inputting it into the model for training will consume a lot of computing resources. Therefore, before inputting the model, it is necessary to n Perform downsampling operation and let X n ∈R B×1×H×W×D , and the loss of local detail information caused by downsampling can be solved by the local image X l This method enables the model to extract local features with detailed information and non-local features with more semantic information in 3D medical images at a lower cost. At the same time, the non-local feature encoder E n The Exponential Moving Average (EMA) method is used to update the weights, so that the model does not increase the amount of calculation when acquiring non-local features. Finally, the long-range dependency between local features and non-local features is modeled through the self-attention mechanism, so that local features can use non-local features to enhance their expressiveness and improve the accuracy of segmentation tasks.

[0052] Local Encoder E l X l Mapped to local feature F l ∈R B×C×H×W×D , non-local encoder E n X n Mapped to non-local feature F n ∈R B×C×H×W×D , where C is the number of channels in the last convolutional layer of the encoder, and C = 288 in the INLF model. After feature extraction, X n and X l Mapped by the encoder into non-local features F n and local features F l The reshape layer will be F n and F l Expand it in the last three dimensions (H×W×D) and transpose the last two dimensions of the expanded feature to obtain and The connection layer will and Concatenate in the second dimension to get the concatenated feature F that is input to the self-attention module t ∈R B×2HWD×C .

[0053] The INLF model uses a U-shaped structure as the backbone network. Figure 2As shown in the figure, the network has a total of 5 resolution stages, and only ordinary convolution, instance normalization and leaky ReLU are used in each stage. The operation order of each convolution layer is convolution-instance normalization-leakyReLU, and two convolution layers are used in each resolution stage of the encoder and decoder. In the backbone network, downsampling is done by serial convolution (the step size of the first convolution layer is greater than 1 at each new resolution stage), and upsampling is done by deconvolution. The feature decoder is similar to the feature encoder, and both are stacked by convolution layers. The difference is that each layer in the decoder is a convolution layer-upsampling layer-convolution layer, and upsampling is done by a deconvolution layer with a step size of 2. At each layer of the decoder, the features of the corresponding layers of the two encoders are passed through the skip connection, and are concatenated with the corresponding features on the channel and then input into the subsequent convolution layer. After all upsampling is completed, the last convolution layer of the decoder will map the features to the final segmentation result.

[0054] Reference Figure 3 and Figure 4 The self-attention module used in the INLF model consists of a multi-headed self-attention layer and a feed-forward layer. The multi-headed self-attention layer performs self-attention calculations on the concatenated features to obtain attention-weighted features. The multi-headed self-attention layer connects F t Mapped to query Q, key K and value V, the output Attention(Q,K,V) of the multi-head self-attention layer is calculated as:

[0055]

[0056] Where d is F t The number of channels is 2HWD, Attention(Q,K,V) then enters the feedforward layer, which is stacked in the form of the first fully connected layer-activation function-second fully connected layer-layer normalization. The calculation of the feedforward layer can be expressed as:

[0057] FFN(X)=max(0,xW1+b1)W2+b2

[0058] W1, W2, b1, b2 are the weights and biases of the first and second fully connected layers respectively, and x is the input of the feedforward layer. The multi-head self-attention layer is used to model F n and F l The long-range dependence between l and F n The information between them can be exchanged, and the feedforward layer is used to introduce nonlinear factors into the self-attention module. The multi-head self-attention layer completes F nand F l After the fusion of t The function of this step is to use skip connections to supplement the position information between features. After feature enhancement, the self-attention module outputs the transformed feature F t ′∈R B×HWD×C , and then re-expand it to get F d ∈R B×C×H×W×D And input the feature decoder D to get the segmented image.

[0059] Based on the same concept, the present invention also provides a three-dimensional medical image segmentation device, including a clipping module and an input module.

[0060] The cropping module is used to crop the three-dimensional medical image once to obtain a non-local image, and to crop the non-local image twice to obtain a local image.

[0061] The input module is used to downsample the non-local image, and input the downsampled non-local image and the local image into the INLF model to obtain a segmented image.

[0062] The input module includes an encoding module, a reshaping module, a splicing module, a calculation module, a transformation module and a segmentation module.

[0063] The encoding module is used to map the downsampled non-local image based on the non-local feature encoder to obtain non-local features; and to map the local image based on the local encoder to obtain local features.

[0064] The reshaping module is used to expand the local features and non-local features based on the reshaping layer, and transpose the last two dimensions of the expanded local features and non-local features.

[0065] The splicing module is used to splice the transposed local features and non-local features in the second dimension based on the connection layer to obtain splicing features.

[0066] The calculation module is used to perform self-attention calculation on the splicing features based on the multi-head self-attention layer to obtain the attention weighted features.

[0067] The transformation module is used to perform nonlinear transformation on the attention weighted features and concatenated features based on the feedforward layer to obtain transformed features.

[0068] The segmentation module is used to decode the expanded transformation features based on the feature decoder to obtain a segmented image.

[0069] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned three-dimensional medical image segmentation method when executing the program.

[0070] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned three-dimensional medical image segmentation method is implemented.

[0071] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0072] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A three-dimensional medical image segmentation method, characterized in that: The following steps are involved: Perform a first cropping on the three-dimensional medical image to obtain a non-local image, and perform a second cropping on the non-local image to obtain a local image; The non-local image is downsampled, and the downsampled non-local image and the local image are input into the INLF model to obtain a segmented image, including: Based on the non-local feature encoder, the downsampled non-local image is mapped to obtain the non-local feature; based on the local encoder, the local image is mapped to obtain the local feature; The local features and non-local features are expanded based on the reshaping layer, and the last two dimensions of the expanded local features and non-local features are transposed; Based on the connection layer, the transposed local features and non-local features are concatenated in the second dimension to obtain concatenated features; Based on the multi-head self-attention layer, self-attention is calculated on the spliced ​​features to obtain the attention weighted features; Based on the feedforward layer, the attention weighted features and the concatenated features are nonlinearly transformed to obtain the transformed features; The expanded transformed features are decoded based on the feature decoder to obtain a segmented image.

2. A three-dimensional medical image segmentation method as claimed in claim 1, characterized in that: The local feature encoder has the same structure as the non-local encoder, both of which include 5 resolution stages, each of which includes convolution, instance normalization and leaky ReLU functions.

3. A three-dimensional medical image segmentation method as claimed in claim 1, characterized in that: The multi-head self-attention layer maps the concatenated features into query Q, key K and value V through a fully connected layer.

4. A three-dimensional medical image segmentation method as claimed in claim 1, characterized in that: The feedforward layer is stacked in the form of a first fully connected layer-activation function-second fully connected layer-layer normalization, as shown below: FFN(X)=max(0,xW1+b1)W2+b2; Where FFN(X) is the output of the feedforward layer, W1 and W2 are the weights of the first fully connected layer and the second fully connected layer, b1 and b2 are the biases of the first fully connected layer and the second fully connected layer, respectively, and x is the concatenated feature and the attention weighted feature.

5. A three-dimensional medical image segmentation method as claimed in claim 1, characterized in that: The feature decoder includes a multi-layer structure, each of which is a convolution layer-upsampling layer-convolution layer; in each layer of the feature decoder, features of corresponding layers of the local feature encoder and the non-local feature encoder are transmitted through a jump connection.

6. A three-dimensional medical image segmentation method as claimed in claim 1, characterized in that: The non-local feature encoder updates the weights by an exponential sliding average method.

7. A three-dimensional medical image segmentation device, characterized in that: include: A cropping module, used for cropping the three-dimensional medical image once to obtain a non-local image, and cropping the non-local image twice to obtain a local image; An input module is used to downsample the non-local image and input the downsampled non-local image and the local image into the INLF model to obtain a segmented image; The input module comprises: The encoding module is used to map the downsampled non-local image based on the non-local feature encoder to obtain non-local features; and map the local image based on the local encoder to obtain local features; A reshaping module, used for expanding the local features and non-local features based on the reshaping layer, and transposing the last two dimensions of the expanded local features and non-local features; A splicing module, used for splicing the transposed local features and non-local features in the second dimension based on the connection layer to obtain splicing features; A calculation module is used to perform self-attention calculation on the spliced ​​features based on the multi-head self-attention layer to obtain attention weighted features; A transformation module is used to perform nonlinear transformation on the attention weighted features and the concatenated features based on the feedforward layer to obtain transformation features; The segmentation module is used to decode the expanded transformation features based on the feature decoder to obtain a segmented image.

8. A computer device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the three-dimensional medical image segmentation method described in any one of claims 1 to 7 is implemented.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the three-dimensional medical image segmentation method described in any one of claims 1 to 7 is implemented.