A multi-object segmentation method and device based on a U-shaped network
Through the multi-objective segmentation method based on U-shaped network, the context-transforming network self-attention mechanism is used to perform block partitioning and feature fusion of medical images, solving the problems of blurred image and high noise in medical image segmentation, and achieving high-precision and efficient image segmentation.
Patent Information
- Application Number
- CN202210597579.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-05-27
AI Technical Summary
Medical image segmentation faces problems such as blurred image, high noise and low visual contrast, which leads to inefficient diagnosis and increases the work burden of doctors.
The multi-objective segmentation method based on U-shaped network is adopted, and the context-transforming network self-attention mechanism is used to block partition, local semantic feature extraction and global semantic feature fusion of images, and image segmentation is performed by combining encoder and decoder.
It improves the accuracy and robustness of medical image segmentation, reduces the difficulty of diagnosis for doctors, saves time, and improves diagnostic efficiency.
Smart Images

Figure CN115082381B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a multi-object segmentation method and device based on a U-shaped network. Background Art
[0002] With the continuous improvement of medical standards, hospitals have more and more intelligent devices, and doctors can also use more and more medical devices to assist themselves in diagnosis. At the same time, due to changes in the living environment, the continuous increase in modern life and work pressure, and the irregular diet and work and rest of people, more and more people are suffering from physical discomfort, and even some relatively serious diseases. The cases and number of people with organic lesions are also gradually increasing. This has gradually increased the workload of doctors. Especially in the diagnosis of CT scan images, the huge number of image diagnosis tasks has virtually increased the work cost of doctors. When doctors spend a lot of time diagnosing patients' CT images, the patients' conditions are easily delayed and timely treatment cannot be obtained, which also reduces the diagnosis efficiency of doctors. With the continuous development of AI technology, the technologies related to medical image processing have been continuously improved and matured, greatly improving the diagnosis efficiency of doctors. With the further improvement of the scientific and technological level, AI artificial intelligence technology has developed rapidly in the medical field. Now, many studies are applying the principle of deep learning to the research of medical image segmentation technology, and intelligent treatment machines have been produced according to this principle. At the same time, due to the further development of deep networks, the accuracy of medical image segmentation has been improved, and deep learning has also been applied to the segmentation of medical images. This can not only greatly reduce the diagnosis difficulty of doctors, but also enable doctors to save more time and energy for more practical treatment of patients, providing a time advantage for researching some intractable diseases. Doctors do not have to spend too much time on the initial diagnosis of the condition, which undoubtedly makes a great contribution to medicine.
[0003] Biomedical image segmentation technology aims to make the anatomical or pathological tissue structure in the image more vivid and clear. Due to the great improvement in inspection results and accuracy, it often plays an important role in the fields of computer technology-assisted diagnosis and smart hospital technology. Mainstream medical image segmentation operations include resection of heart and liver malignancies, resection of brain and brain tumors, discotomy, cell resection, lung incision, pulmonary nodules, heart image segmentation, etc. With medical imaging devices, X-ray radiation, computed tomography (CT), magnetic resonance imaging (MRI), and ultrasound, they have become the four main image assistance and technical means to assist clinicians in examining clinical conditions, evaluating prognosis, and planning in-hospital equipment treatment surgeries for patients. In actual use, although each of these imaging methods has its own characteristics, they are all beneficial for medical examinations of various parts of the human body.
[0004] To assist clinicians in making correct diagnoses, it is necessary to segment certain important objects in medical images and obtain features in the segmented regions. Early medical image segmentation methods usually relied on edge detection, template matching techniques, statistical graphical modeling, active contours, and machine learning techniques. Zhao et al. proposed a new mathematical morphology edge detection algorithm for lung CT images. Lalonde et al. applied Hausdorff-based template matching to inspections, and Chen et al. also used template matching to perform ventricular segmentation in brain CT images. Tsai et al. proposed a shape-based method using level sets for 2D segmentation of cardiac MRI images and 3D segmentation of prostate MRI images. Li et al. used an active profile model to segment liver tumors from abdominal CT images, and Li et al. proposed a framework for medical body data segmentation by combining level sets and support vector machines (SVM). Held et al. applied Markov random fields (MRF) to brain MRI image segmentation.
[0005] Although a large number of solutions have been studied and some have made great progress in some cases, due to the difficulty of feature expression, image segmentation has always been one of the most challenging important academic topics in the application of computer image vision. In particular, it is more difficult to obtain recognition features in medical images than in ordinary RGB images because the former often faces problems such as image blurring, high noise, and low visual contrast. Summary of the Invention
[0006] To solve the above problems existing in the prior art, the present invention provides a multi-object segmentation method and device based on a U-shaped network. The technical problems to be solved by the present invention are achieved through the following technical solutions:
[0007] An embodiment of the present invention provides a multi-object segmentation method based on a U-shaped network, including the steps of:
[0008] S1. Perform block partitioning on the image to be segmented to obtain an input image;
[0009] S2. Input the input image into an encoding module based on the self-attention mechanism of the context transformation network to extract unified local semantic feature information and locate the segmentation target, obtaining an encoder output image;
[0010] S3. Fuse the encoder output image with the semantic feature information obtained during the process of extracting local semantic feature information by the encoding module, and unify the global semantic feature information of the image to be segmented, obtaining a decoder output image;
[0011] S4. Map and output different targets of the image output by the decoder to obtain a segmentation result map.
[0012] In one embodiment of the present invention, step S2 includes:
[0013] S21. Linearly embed the input image and then input it into two consecutive context transformation modules for representation learning, keeping the feature dimension and resolution of the input image unchanged to obtain first multi-scale features;
[0014] S22. Input the first multi-scale features into a block merging layer for downsampling to obtain first downsampled features;
[0015] S23. Input the first downsampled features into two consecutive context transformation modules for representation learning, keeping the feature dimension and resolution of the first downsampled features unchanged to obtain second multi-scale features;
[0016] S24. Input the second multi-scale features into a block merging layer for downsampling to obtain second downsampled features;
[0017] S25. Input the second downsampled features into two consecutive context transformation modules for representation learning, keeping the feature dimension and resolution of the second downsampled features unchanged to obtain third multi-scale features;
[0018] S26. Input the third multi-scale features into a block merging layer for downsampling to obtain the encoder output image.
[0019] In one embodiment of the present invention, the execution steps in the block merging layer include:
[0020] Connect the input blocks together so that the resolution of the image is downsampled by 2 times while the feature dimension is increased by 4 times to obtain connected features;
[0021] Use a linear layer to unify the feature dimension of the connected features to 2 times the original feature dimension of the input blocks to obtain the downsampled features output by the block merging layer.
[0022] In one embodiment of the present invention, step S3 includes:
[0023] S31. Input the encoder output image into a block expansion layer for upsampling to obtain first upsampled features;
[0024] S32. Input the first upsampled features and the third multi-scale features into two consecutive context transformation modules for fusion to obtain first fusion features;
[0025] S33. Input the first fusion feature into a block expansion layer for upsampling to obtain a second upsampled feature;
[0026] S34. Input the second upsampled feature and the second multi-scale feature into two consecutive context transformation modules for fusion to obtain a second fusion feature;
[0027] S35. Input the second fusion feature into a block expansion layer for upsampling to obtain a third upsampled feature;
[0028] S36. Input the third upsampled feature and the first multi-scale feature into two consecutive context transformation modules for fusion to obtain a third fusion feature;
[0029] S37. Input the third fusion feature into a block expansion layer for upsampling to obtain the decoder output image.
[0030] In an embodiment of the present invention, the execution steps in the block expansion layer include:
[0031] Use a linear layer to increase the feature dimension of the input feature to twice the original feature dimension;
[0032] Use a rearrangement operation to expand the resolution of the input feature to twice the original resolution and reduce the feature size of the input feature to one-fourth of the original feature size to obtain the upsampled feature of the block expansion layer.
[0033] In an embodiment of the present invention, the execution steps in the context transformation module include:
[0034] Define the key, query, and value in the context transformation module respectively;
[0035] Perform k×k grouped convolution on all neighbor keys in the k×k spatial grid to obtain the static context of the input image;
[0036] Perform two consecutive convolutions on the static context and the query to obtain an attention matrix;
[0037] Aggregate the attention matrix and all the values to obtain a dynamic context;
[0038] Fuse the static context and the dynamic context to obtain the output feature of the context transformation module.
[0039] In an embodiment of the present invention, step S4 includes:
[0040] Input the decoder output image into a linear projection layer for mapping output to obtain the segmentation result map.
[0041] In an embodiment of the present invention, between step S2 and step S3, there is further included a step:
[0042] Input the encoder output image into a bottleneck layer for deep feature learning, keeping the feature dimension and depth of the encoder output image unchanged to obtain the decoder input image.
[0043] In an embodiment of the present invention, the bottleneck layer includes two consecutive context transformation modules.
[0044] Another embodiment of the present invention provides a multi-object segmentation device based on a U-shaped network, including:
[0045] An image block partitioning module, configured to perform block partitioning on the image to be segmented to obtain an input image;
[0046] An encoding module, configured to extract unified local semantic feature information of the input image and locate the segmentation target to obtain an encoder output image, and the encoding module includes an encoding module based on the self-attention mechanism of the context transformation network;
[0047] A decoding module, configured to fuse the encoder output image with the semantic feature information obtained during the extraction of local semantic feature information, and unify the global semantic feature information of the image to be segmented to obtain a decoder output image, and the decoding module includes a decoding module based on the self-attention mechanism of the context transformation network;
[0048] A mapping output module, configured to map and output different targets of the decoder output image to obtain a segmentation result map.
[0049] Compared with the prior art, the beneficial effects of the present invention:
[0050] The multi-object segmentation method of the present invention uses a module based on the self-attention mechanism of the context transformation network for multi-object segmentation of images, which is more convenient for extracting local semantic information in the image. At the same time, the encoder output image is fused with the semantic feature information obtained during the extraction of local semantic feature information by the encoding module, taking into account both the optimization of local semantic extraction and the unified optimization of global semantic information. It can overcome the problems of image blur, large noise, and low visual contrast faced by medical images, and has high segmentation accuracy, strong robustness, and high segmentation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic flowchart of a multi-object segmentation method based on a U-shaped network provided by an embodiment of the present invention;
[0052] Figure 2The overall flowchart of a multi-object segmentation method based on a U-shaped network provided by an embodiment of the present invention;
[0053] Figure 3 The structural schematic diagram of a context conversion module provided by an embodiment of the present invention;
[0054] Figure 4 The result diagram of segmenting multi-organ medical images provided by an embodiment of the present invention. Detailed implementation manners
[0055] The following further describes the present invention in detail with reference to specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0056] Embodiment 1
[0057] With the rapid development of deep learning technology, medical image segmentation will no longer be limited to the characteristics of manual annotation. The convolutional neural network (CNN) has successfully obtained hierarchical feature extraction of images and has thus become the hottest research in image information processing and computer image vision applications. Also, because the CNN system in feature learning is insensitive to image noise, blur, low contrast, etc., it has also obtained better segmentation results for medical images. Medical image segmentation not only plays an important role in the application and development of computer vision but also provides great help for actual medical treatment.
[0058] Please refer to Figure 1 and Figure 2 , Figure 1 which is the flowchart of a multi-object segmentation method based on a U-shaped network provided by an embodiment of the present invention, Figure 2 and which is the overall flowchart of a multi-object segmentation method based on a U-shaped network provided by an embodiment of the present invention.
[0059] This multi-object segmentation method based on a U-shaped network combines a context conversion (Cot Transformer) module and a U-shaped network, embeds the Cot conversion module into the U-shaped network for image segmentation processing. First, the image is subjected to patch partition, and then it is input into an encoding and decoding module established based on Cot Transformer for local feature extraction. Finally, the global semantic information is unified in the U-shaped network, and after patch expanding, the segmentation result map is mapped and output. Specifically, it includes the steps:
[0060] S1. Perform patch partition on the image to be segmented to obtain the input image.
[0061] Specifically, partition the input image into patches to obtain 3D matrix input image.
[0062] S2. Input the input image into an encoding module based on the context transformation network self-attention mechanism to extract unified local semantic feature information and locate the segmentation target, obtaining an encoder output image.
[0063] Specifically, it includes the steps:
[0064] S21. Linearly embed the input image and then input it into two consecutive context transformation modules for representation learning, keeping the feature dimension and resolution of the input image unchanged, obtaining a first multi-scale feature.
[0065] Specifically, the 3D matrix input image is linearly embedded (LinearEmbedding) to obtain a C-dimensional tokenized image with a resolution of Then, the C-dimensional tokenized image with a resolution of is fed into two consecutive context transformation (Cot Transformer) modules for representation learning, keeping the feature dimension and resolution of the image unchanged, obtaining the first multi-scale feature.
[0066] Specifically, in the traditional self-attention mechanism, all key-value pairs are learned on a single key-value pair, and the text information between them is not explored, which severely limits the ability of self-attention to learn 2D feature maps for visual representation learning. However, the context transformation module solves this problem by integrating context information mining and self-attention learning into a unified architecture.
[0067] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a context transformation module provided by an embodiment of the present invention. The execution steps in the context transformation module include:
[0068] 1) Define the keys, queries, and values in the context transformation module respectively.
[0069] Specifically, assuming there is the same 2D feature map the keys, queries, and values are respectively defined as K = X, Q = X, and V = XW v .
[0070] 2) Perform k×k grouped convolution on all neighbor keys in the k×k spatial grid to obtain the static context of the input image.
[0071] Specifically, different from traditional convolutions, the Cot module does not encode each key through a 1×1 convolution. The Cot module first performs a k×k grouped convolution on all neighbor keys in a k×k spatial grid to obtain the context between each key, which is defined as the static context. The key values of the learned static context are represented as which reflects the static context information between local adjacent keys. Further, K 1 is regarded as the static context representation of the input X.
[0072] 3) Perform two consecutive convolutions on the static context and the query to obtain the attention matrix.
[0073] Specifically, to obtain the static context K 1 and the query Q, first concatenate the static context K 1 and the query Q to get a feature of H×W×2C, and then pass them through two consecutive 1×1 convolutions (θ: 1×1 and δ: 1×1) to obtain an attention matrix of H×W×(k×k×Ch) (W θ represents with ReLU activation function, W δ represents without activation function): A = [K 1 , Q]W θ W δ . In other words, for each head, the local attention matrix at each spatial position of the attention matrix A is learned based on the query feature and the context key feature, rather than an isolated query-key pair. This way enhances the self-attention learning with the additional guidance of the mined static context K 1 .
[0074] 4) Aggregate the attention matrix and all the values to obtain the dynamic context.
[0075] Specifically, according to the context-aware attention matrix A, calculate the feature map K 2 of the parameters by aggregating all the values V: Given that the feature map K 2 with self-attention captures the dynamic feature interactions between the inputs, therefore, K 2 is named the dynamic context, and its dimension is H×W×C.
[0076] 5) Fuse the static context and the dynamic context to obtain the output feature of the context transformation module.
[0077] Specifically, the final output of the Cot module is measured by the attention mechanism as the fusion of the static context K 1 and the dynamic context K 2 , that is, the K 1and K of H×W×C 2 are fused to obtain the output features of the context conversion module of H×W×C.
[0078] The multi-object segmentation method based on the Cot network proposed in this embodiment further improves the segmentation accuracy by extracting the local semantic features of medical images through the Cot network, effectively improves the accuracy and robustness of medical image segmentation, and the obtained segmentation results are more reliable, with high practicality and popularization value.
[0079] S22. Input the first multi-scale feature into the block merging layer for downsampling to obtain the first downsampled feature.
[0080] Specifically, input the first multi-scale feature into the Patch Merging layer for 2× downsampling to reduce the number of tokens and increase the feature dimension to 2 times the original dimension, obtaining the first downsampled feature.
[0081] In a specific embodiment, the execution steps in the block merging layer include:
[0082] 1) Connect the input blocks together so that the resolution of the image is downsampled by 2 times, and at the same time the feature dimension is increased by 4 times to obtain the connected block.
[0083] Specifically, the input blocks are divided into 4 parts and connected together by the block merging layer. Through such processing, the feature resolution will be downsampled by 2 times, and at the same time, due to the connection operation, the feature dimension is increased by 4 times, thus obtaining the connected feature.
[0084] 2) Use a linear layer to unify the feature dimension of the connected feature to 2 times the original feature dimension of the input block to obtain the downsampled feature output by the block merging layer.
[0085] Specifically, since the connection operation increases the feature dimension by 4 times, a linear layer is applied to the connected feature to unify the feature dimension to 2 times the original dimension.
[0086] S23. Input the first downsampled feature into two consecutive context conversion modules for representation learning, keeping the feature dimension and resolution of the first downsampled feature unchanged to obtain the second multi-scale feature.
[0087] Specifically, input the first downsampled feature into two consecutive Cot Transformer modules for representation learning, keeping the feature dimension and resolution of the image unchanged to obtain the second multi-scale feature of dimension.
[0088] S24. Input the second multi-scale feature into the patch merging layer for downsampling to obtain the second downsampled feature.
[0089] Specifically, input the second multi-scale feature into the patch merging layer for 2× downsampling to reduce the number of tokens and increase the feature dimension to twice the original dimension, obtaining the second downsampled feature.
[0090] S25. Input the second downsampled feature into two consecutive context transformation modules for representation learning, keeping the feature dimension and resolution of the second downsampled feature unchanged, to obtain the third multi-scale feature.
[0091] Specifically, input the second downsampled feature into two consecutive context transformation (CotTransformer) modules for representation learning, keeping the feature dimension and resolution of the image unchanged, to obtain the third multi-scale feature.
[0092] S26. Input the third multi-scale feature into the patch merging layer for downsampling to obtain the encoder output image.
[0093] Specifically, input the third multi-scale feature into the patch merging layer for 2× downsampling to reduce the number of tokens and increase the feature dimension to twice the original dimension, obtaining the third downsampled feature as the encoder output image.
[0094] For the specific execution steps of the context transformation (Cot Transformer) module and the patch merging layer in steps S23 - S26, please refer to steps S21 and S22.
[0095] S3. Fuse the encoder output image with the semantic feature information obtained during the local semantic feature information extraction process of the encoding module, and unify the global semantic feature information of the image to be segmented to obtain the decoder output image.
[0096] Specifically, use the decoder to fuse the encoder output image with the semantic feature information obtained during the local semantic feature information extraction process of the encoding module, and unify the global semantic feature information of the image to be segmented to obtain the decoder output image.
[0097] Furthermore, the decoder is symmetric to the encoder and is also constructed based on the Cot Transformer module.
[0098] S31. Input the encoder output image into the patch expansion layer for upsampling to obtain the first upsampled feature.
[0099] Specifically, compared with the block merging layer used in the encoder, the PatchExpanding layer in the decoder is used to upsample the extracted depth features. The PatchExpanding layer reshapes the feature maps of adjacent dimensions into higher-resolution feature maps (2× upsampling), and correspondingly reduces the feature dimension to half of the original dimension. Specifically in this step, the encoder output image is input into the PatchExpanding layer for upsampling to obtain the first upsampled feature.
[0100] In a specific embodiment, the execution steps in the PatchExpanding layer include:
[0101] 1) Use a linear layer to increase the feature dimension of the input features to 2 times the original feature dimension.
[0102] Specifically, before upsampling, apply a linear layer to the input features to increase the feature dimension to 2 times the original dimension
[0103] 2) Use a rearrangement operation to expand the resolution of the input features to 2 times the original resolution and reduce the feature size of the input features to one-fourth of the original feature size to obtain the upsampled features of the PatchExpanding layer.
[0104] Then, use a rearrangement operation to expand the resolution of the input features to 2 times the input resolution and reduce the feature size to one-fourth of the input size thereby obtaining the upsampled features of the PatchExpanding layer.
[0105] S32. Input the first upsampled feature and the third multi-scale feature into two consecutive context transformation modules for fusion to obtain a first fused feature.
[0106] Specifically, similar to U-Net, skip connections are used to fuse the multi-scale features from the encoder and the upsampled features. Specifically in this step, the first upsampled feature and the third multi-scale feature are input into two consecutive context transformation modules for fusion using skip connections to obtain the first fused feature.
[0107] In this embodiment, connecting the shallow features and the deep features together reduces the loss of spatial information caused by downsampling.
[0108] S33. Input the first fused feature into the PatchExpanding layer for upsampling to obtain a second upsampled feature.
[0109] Specifically, input the first fusion feature of into the patch expanding layer for upsampling to obtain the second upsampling feature of
[0110] S34. Input the second upsampling feature and the second multi-scale feature into two consecutive context transformation modules for fusion to obtain the second fusion feature.
[0111] Specifically, input the second upsampling feature of and the second multi-scale feature of into two consecutive context transformation modules for fusion using skip connections to obtain the second fusion feature of
[0112] S35. Input the second fusion feature into the patch expanding layer for upsampling to obtain the third upsampling feature.
[0113] Specifically, input the second fusion feature of into the patch expanding layer for upsampling to obtain the third upsampling feature of
[0114] S36. Input the third upsampling feature and the first multi-scale feature into two consecutive context transformation modules for fusion to obtain the third fusion feature.
[0115] Specifically, input the third upsampling feature of and the first multi-scale feature of into two consecutive context transformation modules for fusion using skip connections to obtain the third fusion feature of
[0116] S37. Input the third fusion feature into the patch expanding layer for upsampling to obtain the decoder output image.
[0117] Specifically, input the third fusion feature of into the patch expanding layer for upsampling to obtain the fourth upsampling feature of W×H×C(4x) as the decoder output image.
[0118] For the specific execution steps of the context transformation (Cot Transformer) module and the patch expanding layer in steps S32 - S37, please refer to steps S21 and S31.
[0119] S4. Input the encoder output image into the bottleneck layer for deep feature learning, keeping the feature dimension and depth of the encoder output image unchanged to obtain the decoder input image.
[0120] Specifically, since the Transformer is too deep to converge, a bottleneck layer is used to learn deep feature representations. In the bottleneck, the feature dimension and resolution remain unchanged.
[0121] Specifically, two consecutive Context Transformer (Cot Transformer) modules are used to construct the bottleneck.
[0122] S4. Map and output different targets of the decoder output image to obtain a segmentation result map.
[0123] Specifically, a Linear Projection layer is used to map and output different targets of the decoder output image to obtain a segmentation result map. Please refer to Figure 4 , Figure 4 which is a segmentation result map of a multi-organ medical image provided by an embodiment of the present invention, and is composed of Figure 4 It can be seen that the multi-target segmentation method of this embodiment obtains a segmentation result with higher accuracy.
[0124] In this embodiment, an image multi-target segmentation model combining a Cot network and a U-shaped network is used for image segmentation, which has the following advantages: 1) The latest self-attention module is used for multi-target segmentation of medical images. Compared with traditional networks, it is more convenient to extract local semantic information in the image. Due to the embedded structure, the integration of the network makes the model more concise, thus reducing the computational cost. 2) Only the self-attention mechanism needs to be used to achieve multi-target segmentation of images. Compared with traditional self-attention mechanisms, the model has a lower complexity and is easier to implement, which can effectively improve the segmentation efficiency of the network model. 3) The embedded structure proposed by using the latest Cot network combined with the U-shaped network takes into account the characteristics of both local semantic extraction optimization and global semantic information unified optimization compared with traditional network segmentation. The segmentation result has higher accuracy and stronger robustness.
[0125] In this embodiment, key features such as text and image key-value pairs are no longer learned separately, but are dynamically learned together with local and global context information; the Cot conversion module can generate context semantic information very well, and combined with the representation learning of the global text information by the U-shaped network, the whole method has good stability and accuracy for the segmentation of medical images. Specifically, the multi-object segmentation method uses a module based on the self-attention mechanism of the context conversion network for multi-object segmentation of images, which is more convenient to extract local semantic information in the image. At the same time, the semantic feature information obtained during the local semantic feature information extraction process of the encoder output image and the encoding module is fused, taking into account both the optimization of local semantic extraction and the unified optimization of global semantic information. It can overcome the problems of image blurring, large noise, and low visual contrast faced by medical images, with high segmentation result accuracy, strong robustness, high segmentation efficiency, and strong multi-object adaptability. It can identify and segment multiple objects in medical images, including the aorta, spleen, liver, etc., and has high practicality and promotion value.
[0126] In summary, the method of this embodiment is based on the combination of the Cot conversion module and the U-shaped network for multi-organ segmentation of medical images, realizing the unity of local and global semantic information. Especially, it has obtained a large improvement in the HD (average Hausdorff Distance) accuracy and has more potential advantages in the multi-organ segmentation of medical images.
[0127] Embodiment 2
[0128] Based on Embodiment 1, this embodiment provides a multi-object segmentation device based on a U-shaped network. The device includes an image block partitioning module, an encoding module, a decoding module, and a mapping output module.
[0129] Specifically, the image block partitioning module is used to perform block partitioning on the image to be segmented to obtain an input image. The encoding module is connected to the image block partitioning module and is used to extract unified local semantic feature information of the input image and locate the segmentation target to obtain an encoder output image; the encoding module includes an encoding module based on the self-attention mechanism of the context conversion network. The decoding module is connected to the encoding module and is used to fuse the encoder output image and the semantic feature information obtained during the local semantic feature information extraction process, and unify the global semantic feature information of the image to be segmented to obtain a decoder output image; the decoding module includes a decoding module based on the self-attention mechanism of the context conversion network. The mapping output module is connected to the decoding module and is used to map and output different targets of the decoder output image to obtain a segmentation result map.
[0130] For the specific execution steps and achieved technical effects in each module of this embodiment, please refer to Embodiment 1, which will not be elaborated in this embodiment.
[0131] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A multi-object segmentation method based on a U-shaped network, characterized in that, Including the steps: S1. Perform block partitioning on the image to be segmented to obtain an input image; S2. Input the input image into an encoding module based on the self-attention mechanism of the context transformation network to extract unified local semantic feature information and locate the segmentation target, obtaining an encoder output image. It includes: S21. Linearly embed the input image and then input it into two consecutive context transformation modules for representation learning, keeping the feature dimension and resolution of the input image unchanged to obtain a first multi-scale feature; S22. Input the first multi-scale feature into a block merging layer for downsampling to obtain a first downsampled feature; S23. Input the first downsampled feature into two consecutive context transformation modules for representation learning, keeping the feature dimension and resolution of the first downsampled feature unchanged to obtain a second multi-scale feature; S24. Input the second multi-scale feature into a block merging layer for downsampling to obtain a second downsampled feature; S25. Input the second downsampled feature into two consecutive context transformation modules for representation learning, keeping the feature dimension and resolution of the second downsampled feature unchanged to obtain a third multi-scale feature; S26. Input the third multi-scale feature into a block merging layer for downsampling to obtain the encoder output image; S3. Fuse the encoder output image with the semantic feature information obtained during the process of extracting local semantic feature information by the encoding module, and unify the global semantic feature information of the image to be segmented to obtain a decoder output image; S4. Map and output different targets of the decoder output image to obtain a segmentation result map.
2. The multi-object segmentation method based on the U-shaped network according to claim 1, wherein The execution steps in the block merging layer include: Connect the input blocks together, so that the resolution of the image is downsampled by 2 times while the feature dimension is increased by 4 times to obtain a connected feature; Use a linear layer to unify the feature dimension of the connected feature to 2 times the original feature dimension of the input block to obtain the downsampled feature output by the block merging layer.
3. The multi-object segmentation method based on a U-shaped network according to claim 1, characterized in that Step S3 includes: S31. Input the decoder input image into a block expansion layer for upsampling to obtain a first upsampled feature; S32. Input the first upsampled feature and the third multi-scale feature into two consecutive context transformation modules for fusion to obtain a first fusion feature; S33. Input the first fusion feature into a block expansion layer for upsampling to obtain a second upsampled feature; S34. Input the second upsampled feature and the second multi-scale feature into two consecutive context transformation modules for fusion to obtain a second fusion feature; S35. Input the second fusion feature into a block expansion layer for upsampling to obtain a third upsampled feature; S36. Input the third upsampled feature and the first multi-scale feature into two consecutive context transformation modules for fusion to obtain a third fusion feature; S37. Input the third fusion feature into a block expansion layer for upsampling to obtain the decoder output image.
4. The multi-object segmentation method based on the U-shaped network according to claim 3, wherein The execution steps in the block expansion layer include: Use a linear layer to increase the feature dimension of the input feature to 2 times the original feature dimension; Use a rearrangement operation to expand the resolution of the input feature to twice the original resolution and reduce the feature size of the input feature to one-fourth of the original feature size, obtaining the upsampled feature of the block expansion layer.
5. The multi-object segmentation method based on the U-shaped network according to claim 1 or 3, characterized in that The execution steps in the context transformation module include: Define the key, query, and value in the context transformation module respectively; Perform k×k grouped convolution on all neighbor keys in the k×k spatial grid to obtain the static context of the input image; Perform two consecutive convolutions on the static context and the query to obtain the attention matrix; Aggregate the attention matrix and all the values to obtain the dynamic context; Fuse the static context and the dynamic context to obtain the output feature of the context transformation module.
6. The multi-object segmentation method based on the U-shaped network according to claim 1, characterized in that Step S4 includes: Input the decoder output image into a linear projection layer for mapping output to obtain the segmentation result map.
7. The multi-object segmentation method based on a U-shaped network according to claim 1, characterized in that There is also a step between step S2 and step S3: Input the encoder output image into a bottleneck layer for deep feature learning, keeping the feature dimension and depth of the encoder output image unchanged to obtain the decoder input image.
8. The multi-object segmentation method based on the U-shaped network according to claim 7, characterized in that, There are two consecutive context transformation modules in the bottleneck layer.
9. A multi-object segmentation device based on a U-shaped network, characterized in that, Include: An image block partitioning module for partitioning the image to be segmented into blocks to obtain the input image; An encoding module for extracting unified local semantic feature information of the input image and locating the segmentation target to obtain the encoder output image. The encoding module includes an encoding module based on the self-attention mechanism of the context transformation network; Extracting unified local semantic feature information of the input image and locating the segmentation target to obtain the encoder output image includes: linearly embedding the input image and then inputting it into two consecutive context transformation modules for representation learning, keeping the feature dimension and resolution of the input image unchanged to obtain the first multi-scale feature; inputting the first multi-scale feature into a block merging layer for downsampling to obtain the first downsampled feature; inputting the first downsampled feature into two consecutive context transformation modules for representation learning, keeping the feature dimension and resolution of the first downsampled feature unchanged to obtain the second multi-scale feature; inputting the second multi-scale feature into a block merging layer for downsampling to obtain the second downsampled feature; inputting the second downsampled feature into two consecutive context transformation modules for representation learning, keeping the feature dimension and resolution of the second downsampled feature unchanged to obtain the third multi-scale feature; inputting the third multi-scale feature into a block merging layer for downsampling to obtain the encoder output image; A decoding module for fusing the encoder output image with the semantic feature information obtained during the local semantic feature information extraction process and unifying the global semantic feature information of the image to be segmented to obtain the decoder output image. The decoding module includes a decoding module based on the self-attention mechanism of the context transformation network; A mapping output module for mapping and outputting different targets of the decoder output image to obtain the segmentation result map.
Citation Information
Patent Citations
Image segmentation method and device, diagnosis system and storage medium
CN109598728A