Three-dimensional multi-scale image segmentation model and image segmentation method, application and system based on three-dimensional multi-scale image segmentation model

By combining Swin Transformer and large-core 3D convolutional encoder in the three-dimensional multi-scale image segmentation model, the problem of instability in three-dimensional image segmentation with large individual differences is solved, and an efficient and robust three-dimensional image segmentation effect is achieved.

CN119991687AActive Publication Date: 2025-05-13GUANGDONG UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510117218.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-13
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

The prior art lacks adaptability to three-dimensional image segmentation with large individual differences, resulting in unstable segmentation effect and poor robustness.

Method used

A three-dimensional multi-scale image segmentation model is adopted, which includes a Swin Transformer encoder, a large-core three-dimensional convolutional encoder and a space converter-based decoder. Through the coordinated work of these components, effective extraction and fusion of the context information and local information of the three-dimensional image are achieved.

Benefits of technology

Accurate segmentation of images with large-scale and foreground-background proportions is achieved, and the robustness and stability of three-dimensional image segmentation is improved, and objective and accurate fully automatic segmentation can be achieved without manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991687A_ABST
    Figure CN119991687A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image segmentation, and discloses a three-dimensional multi-scale image segmentation model and an image segmentation method, application and system based on the three-dimensional multi-scale image segmentation model, and the model comprises a Swin Transformer encoder, a large-kernel three-dimensional convolution encoder and two decoders based on a space converter. The Swin Transformer encoder and the large-kernel three-dimensional convolution encoder are respectively and correspondingly connected with the two decoders based on the space converter; the coding layers of the Swin Transformer encoder are connected with the coding layers of the large-kernel three-dimensional convolution encoder in a one-to-one correspondence manner; the output of the decoder corresponding to the Swin Transformer encoder is connected with the input of the large-kernel three-dimensional convolution encoder; and each decoding layer of the decoder based on the space converter is provided with one space converter block. The method solves the problem that the prior art is not suitable for segmentation of three-dimensional images with large individual differences, and has the characteristic of being suitable for large-scale images with unbalanced foreground-background proportions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image segmentation technology, and more specifically, to a three-dimensional multi-scale image segmentation model and an image segmentation method, application, and system based thereon. Background Art

[0002] Coronary computed tomography angiography (CCTA) is a key auxiliary tool that is extremely useful for clinicians to evaluate coronary artery stenosis. It is an efficient non-invasive technique with high sensitivity and specificity. At the same time, CCTA helps to evaluate other risk factors for coronary heart disease, such as measuring epicardial fat tissue. However, traditional coronary artery automatic segmentation techniques, such as region growing, centerline tracking, threshold processing, and Hessian matrix analysis, usually require preliminary guidance from medical staff to optimize the processing of vascular images. These techniques often show low robustness when facing different individual data sets or the same sample at different stages of the disease process, and their segmentation effects may vary from case to case. The diagnosis of these coronary artery lesions usually requires experienced doctors to manually extract coronary artery images, a process that is not only time-consuming but also susceptible to differences between observers. Therefore, it is particularly important to develop objective and fully automatic coronary artery segmentation techniques.

[0003] Fully automatic coronary artery segmentation is a difficult problem. First, interference from surrounding tissues such as ribs and spine poses a challenge to segmentation. Second, other blood vessels that are similar to coronary arteries in appearance and morphology may lead to misidentification. Third, the segmentation method needs to solve the significant imbalance between foreground and background. Finally, due to the diversity of coronary artery tree morphology in different patients and the differences in local contextual information of different coronary arteries, the accurate positioning of coronary arteries becomes complicated.

[0004] Coronary artery segmentation technologies based on deep learning are mostly based on convolutional neural network models and Transformer models. The convolutional network model is limited by the receptive field of the convolution kernel, resulting in insufficient modeling of long-distance dependencies of contextual information. Although the Transformer model can capture global contextual information well, it has a weak ability to obtain local information.

[0005] For example, there is a medical image segmentation method based on SwinTransformer fusion of multi-scale features and multi-attention mechanism. A medical image segmentation method based on SwinTransformer fusion of multi-scale features and multi-attention mechanism is provided. It includes S1: establishing an encoding module with multiple downsampling layers based on the SwinTransformer network, continuously downsampling the sample medical image, and obtaining sample image features of four different scales; S2: multiplying and fusing the sample image features of the four different scales to obtain four multi-scale fused sample image features; S3: establishing a decoder with multiple upsampling layers based on the attention module combining space and channels; performing continuous upsampling encoding based on the above four fused sample image features to obtain the corresponding segmentation result of the sample image.

[0006] In summary, the existing technical methods lack adaptability to three-dimensional image segmentation with large individual differences. Therefore, how to invent a method suitable for individuals is a technical problem that urgently needs to be solved in this technical field. Summary of the invention

[0007] In order to solve the problem that the existing technology is not suitable for the segmentation of three-dimensional images with large individual differences, the present invention provides a three-dimensional multi-scale image segmentation model and an image segmentation method, application, and system based thereon, which have the characteristics of being suitable for large-scale images and images with an unbalanced foreground-background ratio.

[0008] In order to achieve the above-mentioned purpose of the present invention, the technical scheme adopted is as follows: A three-dimensional multi-scale image segmentation model comprises a Swin Transformer encoder, a large-core three-dimensional convolutional encoder, and two decoders based on spatial transformers; the Swin Transformer encoder and the large-core three-dimensional convolutional encoder are respectively connected to the two decoders based on spatial transformers; the coding layer of the Swin Transformer encoder is connected to the coding layer of the large-core three-dimensional convolutional encoder in a one-to-one correspondence; the output of the decoder corresponding to the Swin Transformer encoder is connected to the input of the large-core three-dimensional convolutional encoder; each decoding layer of the decoder based on the spatial transformer is provided with a spatial transformer block.

[0009] Preferably, the Swin Transformer encoder comprises 4 cascaded encoding layers, each of which is composed of 2 Swin Transformer blocks; the large-core three-dimensional convolution encoder has 4 encoding layers, each of which is composed of 2 large-core convolution blocks, each large-core convolution block comprises 7x7 convolution, BN normalization, ReLU activation function, and each 7x7 convolution is reparameterized by a spatial frequency matrix.

[0010] Furthermore, the decoder includes four cascaded decoding layers, each decoding layer is residually connected to the encoding layer of its corresponding encoder; each decoding layer of the decoder includes a cascaded upsampling block, a residual block, and a spatial transformer block.

[0011] A three-dimensional multi-scale image segmentation method based on the three-dimensional multi-scale image segmentation model comprises the following specific steps: Obtain a 3D image dataset for training and perform preprocessing; The preprocessed 3D image dataset is cropped and input into the 3D multi-scale image segmentation model for multiple iterative training; The 3D image data to be segmented is preprocessed and input into the trained 3D multi-scale image segmentation model for pixel-level classification to complete automatic segmentation.

[0012] Preferably, a three-dimensional image data set for training is obtained and preprocessed, and the specific steps are as follows: Obtain a 3D image dataset for training, and divide it into a training set, a validation set, and a test set in a preset ratio; The obtained training set, validation set and test set are preprocessed respectively; during preprocessing, all 3D image volumes are resampled to the preset isotropic voxel spacing, and then 3D images of different resolutions are randomly sampled into uniform resolution images according to random patches of foreground / background.

[0013] Furthermore, the three-dimensional multi-scale image segmentation model uses the sum of soft Dice coefficient loss and cross entropy loss as the loss function:

[0014] Where I represents the number of classes; J represents the number of voxels; and Respectively represent Voxel-like The true labels and output probabilities at .

[0015] Furthermore, the preprocessed 3D image dataset is cropped and input into the 3D multi-scale image segmentation model for multiple iterative training. The specific steps are as follows: The preprocessed three-dimensional image data set is cut into a plurality of three-dimensional square images of a preset size; Input the 3D block image into the Swin Transformer encoder and extract features from the global information of the input image; The feature information generated at each stage in the Swin Transformer encoder is input into its corresponding decoder based on the spatial transformer for decoding; each time the decoding is performed, the spatial information retained in the spatial path is fused through the spatial transformer block; The decoded semantic information output by the decoder is fused with the spatial information generated by the spatial path and a convolution operation is performed to obtain the output output1; output1 is outputted as a weight mask of a three-dimensional image through a Sigmoid activation function; The weighted mask of the acquired 3D image is multiplied with the input image at the same position and input into a large-core 3D convolutional encoder to further extract the local information of the coronary arteries; The feature information output by each layer of the large-core 3D convolutional encoder is added to the features output by the encoding layer of the corresponding SwinTransformer encoder at the same position and then input into the corresponding decoder based on the spatial transformer for decoding; each time the decoding is performed, the spatial information retained in the spatial path is fused through the spatial transformer block; The decoded semantic information is fused with the spatial information generated by the spatial path and a convolution operation is performed to obtain the output output2; Output2 and output1 are fused by adding them in the same place, and the fused feature information is passed through a segmentation head to output the 3D image segmentation result; The 3D multi-scale image segmentation model is iteratively trained based on the obtained 3D image segmentation result until a preset end-iteration condition is reached.

[0016] Furthermore, after the input feature information X is input into the spatial transformer block, X is divided into output x1 and output x2 through a 1x1x1 convolution operation; output x1 is convolved according to the three directions of H, W, and D respectively to generate their own static context information representation x H 、x W 、x D ; x H 、x W 、x D After fusion, two 1x1 convolution operations are performed to dynamically learn the multi-head attention matrix; the output feature x o Multiply x2 in the same position to obtain the dynamic context representation xo2 based on different spatial state bases, and add xo2 to X in the same position to obtain the output y.

[0017] An application of the three-dimensional multi-scale cascade-based image segmentation method, wherein the method is used for segmenting three-dimensional coronary artery CT images.

[0018] A three-dimensional multi-scale image segmentation system for implementing the image segmentation method, comprising an image processing module, a model training module, and an automatic segmentation module; The image processing module is used to obtain the three-dimensional image data set for training and to be segmented and perform preprocessing; The model training module is used to crop the preprocessed three-dimensional image data set and input it into the three-dimensional multi-scale image segmentation model for multiple iterative training; The automatic segmentation module is used to input the preprocessed image into a trained three-dimensional multi-scale image segmentation model for pixel-level classification to complete automatic segmentation.

[0019] The beneficial effects of the present invention are as follows: The present invention proposes a 3D multi-scale image segmentation model, which realizes long-distance modeling of context information of 3D images through Swin Transformer encoder, extracts local information of 3D images through large-core 3D convolution encoder, and obtains spatial information of input data images through decoder based on spatial transformer, which effectively improves the segmentation effect of 3D images. Objective and accurate fully automatic segmentation of 3D images is realized. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a schematic diagram of the three-dimensional multi-scale image segmentation model of the present invention.

[0021] Figure 2 It is a schematic diagram of the Swin Transformer encoder in the three-dimensional multi-scale image segmentation model of the present invention.

[0022] Figure 3 It is a schematic diagram of a large-core three-dimensional convolutional encoder in the three-dimensional multi-scale image segmentation model of the present invention.

[0023] Figure 4 It is a schematic flow chart of the three-dimensional multi-scale image segmentation method of the present invention.

[0024] Figure 5 It is a specific flow chart of the process of the three-dimensional multi-scale image segmentation method of the present invention in Example 3.

[0025] Figure 6 It is a decoder based on a spatial transformer of the three-dimensional multi-scale image segmentation method of the present invention in Example 3.

[0026] Figure 7 It is a flow chart of the space converter in Example 3.

[0027] Figure 8 It is a schematic diagram of coronary artery CT, a true label of coronary artery, and a predicted label of coronary artery in Example 4. DETAILED DESCRIPTION

[0028] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0029] Example 1 like Figure 1 As shown, a three-dimensional multi-scale image segmentation model is characterized in that: it includes a Swin Transformer encoder, a large-core three-dimensional convolutional encoder, and two decoders based on spatial transformers; the Swin Transformer encoder and the large-core three-dimensional convolutional encoder are respectively connected to the two decoders based on spatial transformers; the encoding layer of the Swin Transformer encoder is connected to the encoding layer of the large-core three-dimensional convolutional encoder in a one-to-one correspondence; the output of the decoder corresponding to the Swin Transformer encoder is connected to the input of the large-core three-dimensional convolutional encoder; each decoding layer of the decoder based on the spatial transformer is provided with a spatial transformer block.

[0030] In one specific implementation, the Swin Transformer encoder includes 4 cascaded encoding layers, each of which is composed of 2 Swin Transformer blocks; In this embodiment, Figure 2 As shown, the Swin Transformer block includes the first cascaded normalization layer LN, a window multi-head self-attention layer W-MSA, the output of the W-MSA is connected with the input residual of the LN and then input into the second cascaded LN and a normalization and multi-layer perceptron MLP, and the output of the MLP is connected with the input residual of the second LN and then output.

[0031] In this embodiment, the patch size of the Swin Transformer encoder is 2 × 2 × 2, and the feature dimension is 2 × 2 × 2 × 1 = 8. In the Swin Transformer encoder, the size of the embedding space C is set to 48. In order to maintain the hierarchical structure of the encoder, at the end of each encoding layer, a patch merging layer is used to reduce the resolution of the feature representation by a factor of 2. In addition, the patch merging layer groups patches of resolution 2 × 2 × 2 and concatenates them, resulting in a feature embedding of 4C, and then a linear layer is used to reduce the feature size of the representation to 2C.

[0032] In this embodiment, Figure 3As shown, the large-core three-dimensional convolution encoder has 4 coding layers, each of which is composed of 2 large-core convolution blocks, each of which includes 7x7 convolution, BN normalization, ReLU activation function, and each 7x7 convolution is reparameterized by the spatial frequency matrix. After each coding layer, the resolution of the feature image decreases by half, and the number of channels doubles. The spatial frequency matrix is ​​obtained by Fourier transform or other frequency domain transform based on demand, representing the convolution kernel in the frequency domain, which can effectively capture the different frequency components of the input signal; when reparameterizing, the original 7x7 convolution kernel is converted into a frequency form, and a new convolution kernel is constructed using these frequency components; by using the information in the frequency domain, a new convolution kernel is recombined. The reparameterization process is actually to operate on the information in the frequency domain and then transform it back to the spatial domain to obtain the convolution kernel for convolution; reparameterization can perform convolution calculations in the frequency domain more efficiently than direct calculations in the spatial domain in some cases, especially when processing large convolution kernels. Reparameterization can help reduce the amount of computation and speed up inference. By processing spatial frequencies, it can better capture the periodicity and texture information in the data, allowing the convolution layer to learn richer features. Operating the convolution kernel in the frequency domain allows for more flexible changes. In addition, for different input features, the convolution kernel obtained after reparameterization can be adaptively adjusted to improve the robustness of the model. Using frequency components to reparameterize the convolution kernel can help the model focus on more regular features rather than local noise, which may reduce the risk of overfitting to a certain extent.

[0033] In a specific implementation, Figure 4 As shown, the decoder includes 4 cascaded decoding layers Stage i , each decoding layer is residually connected to the encoding layer of its corresponding encoder; each decoding layer of the decoder includes a cascaded upsampling block, a residual block, and a spatial transformer block.

[0034] Example 2 like Figure 5 As shown, a three-dimensional multi-scale image segmentation method based on the three-dimensional multi-scale image segmentation model is characterized by comprising the following specific steps: Obtain a 3D image dataset for training and perform preprocessing; The preprocessed 3D image dataset is cropped and input into the 3D multi-scale image segmentation model for multiple iterative training; The 3D image data to be segmented is preprocessed and input into the trained 3D multi-scale image segmentation model for pixel-level classification to complete automatic segmentation.

[0035] In a specific implementation, a 3D image dataset for training is obtained and preprocessed, and the specific steps are as follows: Obtain a 3D image dataset for training, and divide it into a training set, a validation set, and a test set in a preset ratio; The obtained training set, validation set and test set are preprocessed respectively; during preprocessing, all 3D image volumes are resampled to the preset isotropic voxel spacing, and then 3D images of different resolutions are randomly sampled into uniform resolution images according to random patches of foreground / background.

[0036] In this embodiment, all 3D image volumes are resampled to an isotropic voxel spacing of 0.5 mm2, and all non-zero voxel intensities of the image are normalized to the range of [0,1]. 3D coronary CT images of different resolutions are randomly sampled into uniform size [128, 128, 128] resolution images according to random patches of foreground / background at a ratio of 1:1; In a specific implementation, the three-dimensional multi-scale image segmentation model uses the sum of soft Dice coefficient loss and cross entropy loss as the loss function:

[0037] Where I represents the number of classes; J represents the number of voxels; and Respectively represent Voxel-like The true labels and output probabilities at .

[0038] In a specific implementation, Figure 6 As shown in the figure, the preprocessed 3D image dataset is cropped and input into the 3D multi-scale image segmentation model for multiple iterative training. The specific steps are as follows: The preprocessed three-dimensional image data set is cut into a plurality of three-dimensional square images of a preset size; Input the 3D block image into the Swin Transformer encoder and extract features from the global information of the input image; The feature information generated at each stage in the Swin Transformer encoder is input into its corresponding decoder based on the spatial transformer for decoding; each time the decoding is performed, the spatial information retained in the spatial path is fused through the spatial transformer block; The decoded semantic information output by the decoder is fused with the spatial information generated by the spatial path and a convolution operation is performed to obtain the output output1; output1 is outputted through a 1x1 convolution, a Sigmoid activation function, and a residual block to output the weight mask of the three-dimensional image; The weighted mask of the acquired 3D image is multiplied with the input image at the same position and input into a large-core 3D convolutional encoder to further extract the local information of the coronary arteries; The feature information output by each layer of the large-core 3D convolutional encoder is added to the features output by the encoding layer of the corresponding SwinTransformer encoder at the same position and then input into the corresponding decoder based on the spatial transformer for decoding; each time the decoding is performed, the spatial information retained in the spatial path is fused through the spatial transformer block; In this embodiment, the decoded semantic information is fused with the spatial information generated by the spatial path and a convolution operation is performed to obtain the output output2; Output2 is fused with output1 by adding them in the same place, and the fused feature information is passed through a residual block and a segmentation head, and then a 1x1 convolution operation is used to output a segmentation result with 2 channels; The 3D multi-scale image segmentation model is iteratively trained based on the obtained 3D image segmentation result until a preset end-iteration condition is reached.

[0039] In this embodiment, the output resolution of each encoding layer of the Swin transformer encoder and the large-core three-dimensional convolutional encoder is shown in Table 1: Table 1

[0040] In a specific implementation, after the input feature information X is input into the spatial transformer block, X is divided into output x1 and output x2 through a 1x1x1 convolution operation; the output x1 is convolved according to the three directions of H, W, and D to generate their own static context information representation x H 、x W 、x D ; x H 、x W 、x D After fusion, two 1x1 convolution operations are performed to dynamically learn the multi-head attention matrix; the output feature x o Multiply x2 in the same position to obtain the dynamic context representation xo2 based on different spatial state bases, and add xo2 to X in the same position to obtain the output y.

[0041] In this embodiment, all three-dimensional image volumes are resampled to an isotropic voxel spacing of 0.5 mm2, and all non-zero voxel intensities of the image are normalized to the range of [0,1]. Three-dimensional images of different resolutions are randomly sampled into images of uniform resolution of [128, 128, 128] according to random patches of foreground / background at a ratio of 1:1; a basic residual block is introduced to retain the spatial information of the initial three-dimensional image; the SwinTransformer in the cascade encoder performs long-distance modeling on the contextual global information of the three-dimensional image, and decodes it through a decoder based on a spatial transformer module, further integrating the spatial information retained in the spatial path to preliminarily generate a weight mask of the three-dimensional image, and multiplying it with the three-dimensional image as the input of the large-core three-dimensional convolution encoder in the cascade encoder, obtaining local information based on long-distance modeling, reparameterizing the large-core convolution operation through the spatial frequency matrix, improving the ability to capture local details, and integrating the feature encoding of the Swin Transformer into the decoder for decoding; integrating the feature representation of the two decodings of the cascade network and the spatial information retained in the spatial path, and predicting the final three-dimensional image segmentation result through the segmentation head. Therefore, by establishing a three-dimensional multi-scale cascade network model, the three-dimensional image can be segmented in a single stage and end-to-end manner, achieving high segmentation accuracy and stable segmentation performance. Example 3 An application of the three-dimensional multi-scale cascade-based image segmentation method, wherein the method is used for segmenting three-dimensional coronary artery CT images.

[0042] In this example, the data for training the model came from real clinical cases of Guangdong Provincial People's Hospital, consisting of 3D CTA images acquired from 1000 patients with coronary artery disease by Siemens 128-slice dual-source scanner. The size of the CTA image is 512×512×(206−275) voxels, the plane resolution is 0.29 mm2–0.43 mm2, and the spacing is 0.25 mm–0.45 mm. The patients collected were 414 women and 586 men, with an average age of 59.98 years and 57.68 years, respectively. Patients aged 18 years or older with a history of ischemic stroke, transient ischemic attack, or peripheral artery disease were eligible for inclusion. Patients were excluded due to low CTA imaging quality (assessed by a radiologist at level 3), which may affect coronary artery function. The left and right coronary arteries in each image were independently labeled by two radiologists, and the results were cross-validated. If there was a difference, a third radiologist performed annotations, and the final annotations were based on consensus.

[0043] In this embodiment, after obtaining the three-dimensional image data set for training, it is divided into training set, validation set, and test set in a ratio of 8:1:1. The epoch of all model training is set to 50 rounds (about 40,000 iterations), the training batch size is 1, the learning rate is set to 0.0001, the Adamw optimizer is used, and the entire training process is trained using linear warm-up and cosine annealing learning rate scheduler, and DiceCELoss is used as the loss function. A sliding window method with an overlap of 0.5 is used for adjacent voxels, and the evaluation indicators of the final prediction results mainly use Dice similarity coefficient evaluation, HD95 Hausdorff distance, recall rate Recall and precision Precision.

[0044] In this embodiment, the average dice coefficient is 81.24, the Hausdorff distance HD95 is 14.28, and the three-dimensional coronary artery segmentation results are as follows: Figure 8 As shown, from left to right are coronary artery CT, coronary artery true label, and coronary artery predicted label.

[0045] In this example, the performance of the model on the test set is compared with other segmentation methods. The compared methods are all common 3D medical image segmentation models, namely 3D Unet, SegResNet, 3DUXNET based on convolutional neural networks and nnFormer, UNETR, SwinUNETR based on Transformer architecture. The comparison results are shown in Table 2: Table 2

[0046] In summary, the present invention obtains and preprocesses a three-dimensional coronary CT three-dimensional image data set; establishes a multi-scale cascade coronary artery segmentation model based on Swin Transformer and large-core three-dimensional convolution; cuts the preprocessed coronary CT three-dimensional image data set into a number of three-dimensional blocks of a preset size, and inputs all three-dimensional blocks into the multi-scale cascade coronary artery segmentation model in turn for multiple iterative training; inputs the coronary CT three-dimensional image to be segmented into the trained multi-scale cascade coronary artery segmentation model for pixel-level classification, and completes the automatic segmentation of the three-dimensional coronary artery; the present invention proposes to extract global features and local detail features in the encoding and decoding by using a multi-scale cascade encoding method based on Swin Transformer and large-core three-dimensional convolution, and introduces a 3D space converter module in the decoder, which can extract the correlation information between different spatial states, effectively improve the decoding ability and visual representation ability of the decoder, and improve the accuracy of three-dimensional coronary artery segmentation. Compared with the prior art, the present invention can better overcome the limitations caused by the dependence on manual intervention, the imbalance of foreground and background of coronary artery images, large individual differences, and complex and small vascular structures.

[0047] Example 4 A three-dimensional multi-scale image segmentation system for implementing the image segmentation method, comprising an image processing module, a model training module, and an automatic segmentation module; The image processing module is used to obtain the three-dimensional image data set for training and to be segmented and perform preprocessing; The model training module is used to crop the preprocessed three-dimensional image data set and input it into the three-dimensional multi-scale image segmentation model for multiple iterative training; The automatic segmentation module is used to input the preprocessed image into a trained three-dimensional multi-scale image segmentation model for pixel-level classification to complete automatic segmentation.

[0048] Obviously, the above embodiments of the present invention are only examples for clearly explaining the present invention, and are not intended to limit the implementation methods of the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A three-dimensional multi-scale image segmentation model, characterized in that: It includes a Swin Transformer encoder, a large-core three-dimensional convolutional encoder, and two decoders based on spatial transformers; the Swin Transformer encoder and the large-core three-dimensional convolutional encoder are respectively connected to the two decoders based on spatial transformers; the encoding layer of the Swin Transformer encoder is connected to the encoding layer of the large-core three-dimensional convolutional encoder in a one-to-one correspondence; the output of the decoder corresponding to the Swin Transformer encoder is connected to the input of the large-core three-dimensional convolutional encoder; each decoding layer of the decoder based on the spatial transformer is provided with a spatial transformer block.

2. The three-dimensional multi-scale image segmentation model according to claim 1, characterized in that: The SwinTransformer encoder includes 4 cascaded encoding layers, each of which is composed of 2 Swin Transformer blocks; the large-core three-dimensional convolution encoder has 4 encoding layers, each of which is composed of 2 large-core convolution blocks, each of which includes 7x7 convolution, BN normalization, ReLU activation function, and each 7x7 convolution is reparameterized by a spatial frequency matrix.

3. The three-dimensional multi-scale image segmentation model according to claim 2, characterized in that: The decoder comprises four cascaded decoding layers, each of which is residually connected to the encoding layer of its corresponding encoder; each decoding layer of the decoder comprises a cascaded upsampling block, a residual block, and a spatial converter block.

4. A three-dimensional multi-scale image segmentation method based on the three-dimensional multi-scale image segmentation model according to any one of claims 1 to 3, characterized in that: The specific steps include: Obtain a 3D image dataset for training and perform preprocessing; The preprocessed 3D image dataset is cropped and input into the 3D multi-scale image segmentation model for multiple iterative training; The 3D image data to be segmented is preprocessed and input into the trained 3D multi-scale image segmentation model for pixel-level classification to complete automatic segmentation.

5. The three-dimensional multi-scale cascade image segmentation method according to claim 4, characterized in that: Obtain the 3D image dataset for training and preprocess it. The specific steps are as follows: Obtain a 3D image dataset for training, and divide it into a training set, a validation set, and a test set in a preset ratio; The obtained training set, validation set and test set are preprocessed respectively; during preprocessing, all 3D image volumes are resampled to the preset isotropic voxel spacing, and then 3D images of different resolutions are randomly sampled into uniform resolution images according to random patches of foreground / background.

6. The three-dimensional multi-scale cascade image segmentation method according to claim 4, characterized in that: The three-dimensional multi-scale image segmentation model uses the sum of soft Dice coefficient loss and cross entropy loss as the loss function: Where I represents the number of classes; J represents the number of voxels; and Respectively represent Voxel The true labels and output probabilities at .

7. The three-dimensional multi-scale cascade image segmentation method according to claim 4, characterized in that: The preprocessed 3D image dataset is cropped and input into the 3D multi-scale image segmentation model for multiple iterative training. The specific steps are as follows: The preprocessed three-dimensional image data set is cut into a plurality of three-dimensional square images of a preset size; Input the 3D block image into the Swin Transformer encoder and extract features from the global information of the input image; The feature information generated at each stage in the Swin Transformer encoder is input into its corresponding decoder based on the spatial transformer for decoding; each time the decoding is performed, the spatial information retained in the spatial path is fused through the spatial transformer block; The decoded semantic information output by the decoder is fused with the spatial information generated by the spatial path and a convolution operation is performed to obtain the output output1; Output1 is used to output the weight mask of the three-dimensional image through the Sigmoid activation function; The weighted mask of the acquired 3D image is multiplied with the input image at the same position and input into a large-core 3D convolutional encoder to further extract the local information of the coronary arteries; The feature information output by each layer of the large-core 3D convolutional encoder is added to the features output by the encoding layer of its corresponding Swin Transformer encoder at the same position and then input into its corresponding decoder based on the spatial transformer for decoding; during each decoding, the spatial information retained in the spatial path is fused through the spatial transformer block; The decoded semantic information is fused with the spatial information generated by the spatial path and a convolution operation is performed to obtain the output output2; Output2 and output1 are fused by adding them in the same place, and the fused feature information is passed through a segmentation head to output the 3D image segmentation result; The 3D multi-scale image segmentation model is iteratively trained based on the obtained 3D image segmentation result until a preset end-iteration condition is reached.

8. The three-dimensional multi-scale cascade image segmentation method according to claim 7, characterized in that: After the input feature information X is input into the spatial transformer block, X is divided into output x1 and output x2 through a 1x1x1 convolution operation; output x1 is convolved according to the three directions of H, W, and D to generate their own static context information representation x H 、x W 、x D ; x H 、x W 、x D After fusion, two 1x1 convolution operations are performed to dynamically learn the multi-head attention matrix; the output feature x o Multiply x2 in the same position to obtain the dynamic context representation xo2 based on different spatial state bases, and add xo2 to X in the same position to obtain the output y.

9. An application of the three-dimensional multi-scale cascade image segmentation method according to any one of claims 4 to 8, characterized in that: The method is used for segmenting three-dimensional coronary artery CT images.

10. A three-dimensional multi-scale image segmentation system for implementing the image segmentation method according to any one of claims 4 to 8, characterized in that: Includes image processing module, model training module, and automatic segmentation module; The image processing module is used to obtain the three-dimensional image data set for training and to be segmented and perform preprocessing; The model training module is used to crop the preprocessed three-dimensional image data set and input it into the three-dimensional multi-scale image segmentation model for multiple iterative training; The automatic segmentation module is used to input the preprocessed image into a trained three-dimensional multi-scale image segmentation model for pixel-level classification to complete automatic segmentation.

Citation Information

Patent Citations

  • Lung CT image segmentation method based on mixed Swin Transform U-Net

    CN117274147A

  • Cloud layer detection method integrating Swin transformer and CNN (Convolutional Neural Network)

    CN117975284A

  • System and method for 3D medical image segmentation

    US20240362788A1