Single-slice semi-supervised 3D medical image segmentation method
Through the single-slice semi-supervised method, the static and dynamic correlation extraction modules and bidirectional multi-grained fusion are used to solve the problems of heavy labeling burden and low accuracy in 3D medical image segmentation, and the efficient segmentation effect is achieved.
Patent Information
- Application Number
- CN202510404156.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
The existing 3D medical image segmentation method relies on a large amount of precise annotation data, which leads to heavy annotation burden. The traditional semi-supervised learning method is difficult to capture complex structure and context information, and the segmentation accuracy is limited.
A single-sliced semi-supervised method is used to capture spatial structure and timing change information through static and dynamic correlation extraction modules, and rich feature representations are generated through bidirectional multi-grained fusion, and model training is performed by combining mixed pseudo-labels and consistency losses.
Significantly reduces the annotation burden, improves segmentation accuracy and stability, and is suitable for a variety of 3D medical image segmentation tasks.
Smart Images

Figure CN120339613A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image processing, and particularly to a single-slice semi-supervised 3D medical image segmentation method. Background Art
[0002] With the rapid development of deep learning in the field of computer vision, learning-based models have been widely used in medical image processing, especially in image segmentation tasks. 3D medical image segmentation plays a crucial role in disease diagnosis and surgical navigation, enabling doctors to more accurately identify lesion areas and formulate treatment plans. However, existing 3D medical image segmentation methods rely on a large amount of precisely annotated data, and obtaining such annotated data requires professional clinicians to spend a large amount of time on pixel-by-pixel annotation. In addition, the annotation of 3D medical images is more complex than that of 2D images because 3D images need to be annotated slice by slice, and subjective judgments of different doctors on the same image may lead to inconsistent annotations. Therefore, constructing a high-quality, large-scale annotated dataset is a huge challenge.
[0003] Although existing semi-supervised learning methods alleviate the annotation requirement to a certain extent, they still need to annotate the entire 3D volume, which is still a heavy task for clinicians. Traditional semi-supervised learning methods usually rely on consistency regularization or pseudo-labeling techniques to utilize unlabeled data to enhance the generalization ability of the model. However, when dealing with 3D medical images, these methods often have difficulty capturing complex structures and context information in the images, resulting in limited segmentation accuracy. Summary of the Invention
[0004] To solve the above technical problems existing in the prior art, the present invention proposes a single-slice semi-supervised 3D medical image segmentation method. By using a small amount of single-slice annotations and a large amount of unlabeled data for training, the annotation burden is significantly reduced. This method captures spatial structure information and temporal change information in 3D medical images through static correlation extraction and dynamic correlation extraction modules respectively, and generates rich feature representations through a bidirectional multi-granularity fusion module, thereby improving segmentation accuracy and stability.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A single-slice semi-supervised 3D medical image segmentation method, comprising the following steps:
[0007] Step 1, construct a backbone network based on the MT framework. The backbone network includes a dynamic correlation extraction DCE module and a static correlation extraction SCE module. Use VNet as the backbone network for medical image transmission and processing, which can be described as E s ·D s and Et ·D t , where E s and D s are the encoder and decoder of the student model, and E t and D t are the encoder and decoder of the teacher model. For SCE, an auxiliary restoration decoder is established and denoted as
[0008] Step 2, Data Acquisition and Preprocessing: Obtain 3D medical image data, which is a 3D medical image dataset containing a small amount of single-slice annotation data and a large amount of unannotated data, and perform preprocessing;
[0009] Step 3, Dynamic Correlation Extraction: Construct a pseudo spatio-temporal sequence through recursive slice shifting and extract dynamic features;
[0010] Step 4, Static Correlation Extraction: Perform an image restoration task through a static correlation extraction module and extract static features;
[0011] Step 5, Bidirectional Multi-Granularity Feature Learning and Decision Fusion: Fuse the dynamic features and the static features through a bidirectional multi-granularity feature fusion module to generate a multi-scale feature representation;
[0012] Step 6, Model Training and Optimization: Adopt a teacher-student model framework and combine hybrid pseudo-label generation and consistency loss for model training and optimization;
[0013] Step 7, Model Evaluation: Evaluate the segmentation performance of the model using multiple metrics.
[0014] Preferably, in Step 3, the following steps are specifically included:
[0015] Step 31, Slice Shifting: Set the acquisition direction of the 3D medical image data as the sliding axis, select m consecutive slices as the first group of sliding slices, and move the sliding slices to the end position of the sample. Then, repeat the above operation to recursively move the original three-dimensional volume by 2m, 3m... sm slices to obtain s new 3D volumes;
[0016] Step 32, Sequence Concatenation: Concatenate the s 3D volumes into a pseudo spatio-temporal sequence X P ;
[0017] Step 33, Dynamic Feature Extraction: Add random Gaussian noise and of the pseudo-sequence as perturbations and input them into the student model and the teacher model respectively to extract the dynamic correlation feature fea d volume;
[0018] Step 34, Dynamic Correlation Loss Calculation: Calculate the difference L between the dynamic correlation features and the true annotation ds : 3
[0019]
[0020] where Pr i is the prediction result of the i-th layer, and Y k is the true annotation.
[0021] Preferably, in the said Step 4, it specifically includes the following steps:
[0022] Step 41, Static Feature Extraction: Input the pseudo spatio-temporal sequence X P into the encoder E S to extract the feature fea e . The encoder adopts a multi-layer convolutional neural network structure, which is used to extract multi-level features from the image. After each layer of convolution operation, the ReLU activation function is used for non-linear transformation, and downsampling is performed through the max pooling layer. The output of the encoder is the high-dimensional feature representation fea e after multiple layers of convolution and pooling operations;
[0023] Step 42, Image Restoration: Input the extracted feature fea e into the restoration decoder to generate the restored image . The restoration decoder adopts a transposed convolution layer and an upsampling operation to gradually restore the low-resolution feature map to the resolution of the original image;
[0024] Step 43, Restoration Loss Calculation: Calculate the mean square error between the restored image and the original image X P . The formula is as follows:
[0025]
[0026] Preferably, in the said Step 5, it specifically includes the following steps:
[0027] Step 51, Input Features: Input the multi-scale features fea i from the encoder and the decoder;
[0028] Step 52, Bidirectional Convolution Operation: Perform 3×3 convolution operations on the feature maps of each scale in both directions, including forward convolution: starting from the smallest feature map, gradually performing convolution operations on the larger-scale feature maps to fuse the feature information from small scales to large scales; backward convolution: starting from the largest feature map, gradually performing convolution operations on the smaller-scale feature maps to fuse the feature information from large scales to small scales. The formula for the bidirectional convolution operation is as follows:
[0029] Fh f = Conv 3×3 (Fea f ) + Attention(Fh f-1 )
[0030] Among them, Fea f is the feature map of the f-th layer, and Fh f is the feature map after bidirectional convolution;
[0031] Step 53, Attention mechanism: Filter the context information through the attention mechanism to generate a high-level feature representation with rich semantic information, and use the Sigmoid function to calculate the attention weight. The formula is as follows:
[0032] Att f = Sigmoid(Conv 3×3 (Fh f-1 ))
[0033] Among them, Fh f-1 is the feature map of the previous layer.
[0034] Step 54, Feature filtering and fusion: Multiply the attention weight by the feature map of the current layer to filter out unimportant context information. Feature fusion includes forward feature propagation, backward feature propagation, and feature concatenation;
[0035] Step 55, Decision fusion: Extract the global feature Fea g through a 3D convolution and a Transformer module, and fuse the global feature Fea g with the multi-scale feature representation Fb i to generate the final segmentation result P s .
[0036] Preferably, in step 54, the following feature fusions are specifically included:
[0037] Forward feature propagation: Starting from the smallest feature map, perform convolution operations step by step towards the large-scale feature map to generate the forward feature map Fh f ;
[0038] Backward feature propagation: Starting from the largest feature map, perform convolution operations step by step towards the small-scale feature map to generate the backward feature map Fg f ;
[0039] Feature concatenation: Concatenate the forward and backward feature maps to generate the final feature representation Fb i .
[0040] Preferably, in step 6, the following steps are specifically included:
[0041] Step 61. Generate a mask in advance: The distance regularization level set evolution is introduced to generate a prior mask for unlabeled slices. The mask is calculated according to the single-slice label hint. The label of the k-th slice is used as the segmentation hint for the current nearest neighboring slice, and this is repeated until the prior masks for all slices are calculated. According to the evolved level set function Generate the pseudo-label Y for unlabeled slices P . The formula for level set evolution is as follows:
[0042]
[0043] where μ, λ, and α are control parameters, d p is the distance regularization term, and g is the edge detection function;
[0044] Step 62. Generate mixed pseudo-labels: By combining the mask generated in advance and the prediction results of the teacher model, more reliable mixed pseudo-labels are generated;
[0045] Step 63. Loss design: Use a loss function to constrain the backpropagation of the network model;
[0046] Step 64. Weight update: Use the weights θ of the student model s to update the weights θ of the teacher model through EM4 t , and the weight synchronization function is as follows:
[0047]
[0048] where η is the decay hyperparameter of EM4, which is used to control the update rate.
[0049] Preferably, in step 62, the following steps are specifically included:
[0050] Step 621. Teacher model prediction: Input the unlabeled 3D medical image into the teacher model, and the teacher model generates prediction results P t ;
[0051] Step 622. Confidence calculation: Calculate the confidence Conf of the teacher model prediction results. The formula is as follows:
[0052]
[0053] where p f (i) and p b (i) are the probabilities of the foreground and background respectively;
[0054] Step 623. Mixed weight calculation:
[0055]
[0056] Among them, is the maximum mixing weight, T is the total number of training steps, and v is the decay coefficient;
[0057] The mixed pseudo-label P fu is generated as follows:
[0058] w fu_model = Conf(i)w ct
[0059] w fu_pseudo = (1 - Conf(i))(1 - w ct )
[0060] w model , w pseudo = softmax[w fu_model , w fu_pseudo
[0061] P fu = w model Y p , w pseudo P t
[0062] Among them, w model and w pseudo are the weights of the pre-generated mask and the prediction result of the teacher model, respectively.
[0063] Preferably, in step 63, it specifically includes several key losses to define the total loss, including:
[0064] Using Dice and cross-entropy losses to optimize the supervised segmentation loss on the single-slice label Y k The formula is as follows: The formula is as follows:
[0065]
[0066] Among them, P s is the prediction result, and Y k is the single-slice annotation;
[0067] The student model is also restricted by the mixed pseudo-label P fu , and the loss between P s and P fu is calculated Add
[0068]
[0069] Among them, i and j are the slice index and voxel index respectively, and θ max is the weight decay coefficient;
[0070] Calculate the consistency loss using unlabeled data to ensure the output consistency of the model under different perturbations. The formula is as follows:
[0071] L con = MES(P s , P t )
[0072] where P s is the prediction result of the student model, and P t is the prediction result of the teacher model;
[0073] Total loss:
[0074]
[0075] where λ is the Gaussian warm-up function, and t and t max represent the current training step and the maximum training step respectively.
[0076] Preferably, in step 7, the multiple metrics specifically include the Dice coefficient, Jaccard coefficient, average surface distance, and 95% Hausdorff distance.
[0077] Compared with the prior art, the present invention has the following beneficial effects:
[0078] 1. Reduce the annotation burden: By using a small amount of single-slice annotations and a large amount of unlabeled data for training, the annotation burden on clinicians is significantly reduced.
[0079] 2. Improve the segmentation accuracy: Through static and dynamic correlation extraction and bidirectional multi-granularity fusion, the model can better capture the context information in 3D medical images, significantly improving the segmentation accuracy.
[0080] 3. Enhance the model stability: Through the construction of pseudo spatio-temporal sequences and bidirectional multi-granularity fusion, the model shows stronger stability when facing complex and non-stationary 3D medical images.
[0081] 4. Wide applicability: This method is applicable to a variety of 3D medical image segmentation tasks, such as the segmentation of organs like the left atrium, liver tumors, pancreas, and prostate, etc., with wide applicability. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0083] Figure 1 is the overall flowchart of the present invention;
[0084] Figure 2Schematic diagram of generating pseudo spatio-temporal sequence of the present invention;
[0085] Figure 3 Schematic diagram of the two-way multi-scale feature learning architecture of the present invention;
[0086] Figure 4 Schematic diagram of the feature fusion architecture of the present invention. Detailed implementation manners
[0087] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents the selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0088] As Figure 1 - Figure 4 shown, this embodiment proposes a single-slice semi-supervised 3D medical image segmentation method, including the following steps:
[0089] Step 1: Construct a backbone network based on the MT framework. The backbone network includes a dynamic correlation extraction (DCE) module and a static correlation extraction (SCE) module. Use VNet as the backbone network for medical image transmission and processing, which can be described as E s ·D s and E t ·D t where E s and D s are the encoder and decoder of the student model, and E t and D t are the encoder and decoder of the teacher model. For SCE, establish an auxiliary reduction decoder denoted as
[0090] Step 2: Data acquisition and preprocessing: Obtain 3D medical image data. The 3D medical image data contains a 3D medical image dataset with a small amount of single-slice annotation data and a large amount of unannotated data. The data source can be hospitals, public datasets (such as left atrium, liver tumor, pancreas, and prostate datasets). Perform preprocessing on the dataset, including operations such as normalization (eliminating intensity differences) and denoising (non-local mean filtering) to improve the training effect of the model;
[0091] Step 3, Dynamic Relevance Extraction: Construct a pseudo spatio-temporal sequence through recursive slice shifting to extract dynamic features, which specifically includes the following steps:
[0092] Step 31, Slice Shifting: Set the acquisition direction of the 3D medical image data as the sliding axis, select consecutive m slices as the first group of sliding slices, and move the sliding slices to the end position of the sample. Then, repeat the above operation to recursively move the original three-dimensional volume by 2m, 3m... sm slices to obtain s new 3D volumes;
[0093] Step 32, Sequence Concatenation: Concatenate the s 3D volumes into a pseudo spatio-temporal sequence X P ;
[0094] Step 33, Dynamic Feature Extraction: Add pseudo sequences with random Gaussian noise and as perturbations and input them into the student model and the teacher model respectively to extract dynamic relevance features fea d volume;
[0095] Step 34, Dynamic Relevance Loss Calculation: Calculate the difference L ds :
[0096]
[0097] where Pr i is the prediction result of the i-th layer, and Y k is the true annotation.
[0098] Add Gaussian noise perturbations to the sequence, input it into the student and teacher models, extract spatio-temporal correlation features, constrain the dynamic feature differences between the student and teacher models through consistency loss, the pseudo spatio-temporal sequence simulates the temporal changes of organ morphology, such as heart contraction and tumor infiltration, captures dynamic context information, the noise perturbation enhances the model's robustness to data noise, the pseudo spatio-temporal sequence perturbation and EMA strategy reduce noise interference, and it performs stably in low-contrast images (such as pancreatic CT).
[0099] Step 4, Static Relevance Extraction: Perform an image restoration task through a static relevance extraction module to extract static features. The reconstruction task forces the model to learn the stable anatomical structures between slices, such as organ boundaries and vascular topologies. The static features retain global spatial information and make up for the locality of the dynamic features, which specifically includes the following steps:
[0100] Step 41, Static Feature Extraction: Pass the pseudo spatio-temporal sequence X P through the encoder E S to extract the feature fea e, the encoder adopts a multi-layer convolutional neural network structure to extract multi-level features from the image. After each layer of convolution operation, the ReLU activation function is used for non-linear transformation, and downsampling is performed through the max-pooling layer. The output of the encoder is a high-dimensional feature representation fea after multiple layers of convolution and pooling operations. e ;
[0101] Step 42, Image restoration: Input the extracted feature fea e into the restoration decoder to generate the restored image The restoration decoder adopts transposed convolution layers and upsampling operations to gradually restore the low-resolution feature map to the resolution of the original image;
[0102] Step 43, Restoration loss calculation: Calculate the mean square error between the restored image and the original image X P , and the formula is as follows:
[0103]
[0104] Step 5, Bidirectional multi-granularity feature learning and decision fusion: Fuse the dynamic features and the static features through a bidirectional multi-granularity feature fusion module to generate a multi-scale feature representation. The bidirectional propagation simulates the "local-global" cognitive mechanism of human vision to enhance semantic consistency, and the attention mechanism dynamically adjusts the feature weights to improve the accuracy of the segmentation boundary. Specifically, it includes the following steps:
[0105] Step 51, Input features: Input the multi-scale features fea from the encoder and the decoder i ;
[0106] Step 52, Bidirectional convolution operation: Perform 3×3 convolution operations on the feature maps of each scale in both directions, including forward convolution: starting from the smallest feature map, gradually performing convolution operations on the larger-scale feature maps to fuse the feature information from small scales to large scales; backward convolution: starting from the largest feature map, gradually performing convolution operations on the smaller-scale feature maps to fuse the feature information from large scales to small scales. The formula for the bidirectional convolution operation is as follows:
[0107] Fh f =Conv 3×3 (Fea f )+Attention(Fh f-1 )
[0108] where Fea f is the feature map of the f-th layer, and Fh f is the feature map after bidirectional convolution;
[0109] Step 53. Attention mechanism: Filter the context information through the attention mechanism to generate a high-level feature representation with rich semantic information. Use the Sigmoid function to calculate the attention weights. The formula is as follows:
[0110] Att f =Sigmoid(Conv 3×3 (Fh f-1 ))
[0111] where Fh f-1 is the feature map of the previous layer.
[0112] Step 54. Feature filtering and fusion: Multiply the attention weights by the feature map of the current layer to filter out unimportant context information. Feature fusion includes forward feature propagation, backward feature propagation, and feature concatenation, which specifically include the following feature fusions:
[0113] Forward feature propagation: Starting from the smallest feature map, perform convolutional operations step by step towards the larger-scale feature maps to generate the forward feature map Fh f ;
[0114] Backward feature propagation: Starting from the largest feature map, perform convolutional operations step by step towards the smaller-scale feature maps to generate the backward feature map Fg f ;
[0115] Feature concatenation: Concatenate the forward and backward feature maps to generate the final feature representation Fb i .
[0116] Step 55. Decision fusion: Extract the global feature Fea g through 3D convolution and the Transformer module, and fuse the global feature Fea g with the multi-scale feature representation Fb i to generate the final segmentation result P s .
[0117] Step 6. Model training and optimization: Adopt the teacher-student model framework and combine the generation of mixed pseudo-labels and consistency loss for model training and optimization, which specifically include the following steps:
[0118] Step 61. Generate masks in advance: To more effectively guide the learning of the segmentation network in the initial stage of training, distance regularization level set evolution is introduced to generate prior masks for unlabeled slices. The masks are calculated based on the single-slice label hints. Given that the differences between adjacent slices are extremely small, the label of the k-th slice is used as the segmentation hint for the current nearest adjacent slice. This process is repeated until the prior masks for all slices are calculated. Generate the pseudo-labels Y for unlabeled slices according to the evolved level set function P The formula for level set evolution is as follows:
[0119]
[0120] where μ, λ, and α are control parameters, d p is the distance regularization term, and g is the edge detection function;
[0121] Step 62, Generation of mixed pseudo-labels: Generation of mixed pseudo-labels is a key step in the present invention, aiming to generate more reliable mixed pseudo-labels by combining the pre-generated mask and the prediction results of the teacher model to guide the learning process of the student model in the middle stage of training. Through the generation of mixed pseudo-labels, the model can better utilize the unlabeled data and improve the segmentation accuracy and stability. The specific steps are as follows:
[0122] Step 621, Prediction by the teacher model: Input the unlabeled 3D medical image into the teacher model, and the teacher model generates the prediction result P t ;
[0123] Step 622, Confidence calculation: Calculate the confidence Conf of the prediction result of the teacher model. The formula is as follows:
[0124]
[0125] where p f (i) and p b (i) are the probabilities of the foreground and background respectively;
[0126] Step 623, Calculation of the mixing weight:
[0127]
[0128] where, is the maximum mixing weight, T is the total number of training steps, and v is the decay coefficient;
[0129] Generation of the mixed pseudo-label P fu :
[0130] w fu_model = Conf(i)w ct
[0131] w fu_pseudo = (1 - Conf(i))(1 - w ct )
[0132] w model , w pseudo = softmax[w fu_model , w fu_pseudo
[0133] Pfu = w model Y p , w pseudo P t
[0134] where w model and w pseudo are the weights of the pre-generated mask and the prediction result of the teacher model, respectively.
[0135] Step 63, Loss design: To ensure the normal training of the network, the present invention uses the following loss function to constrain the backpropagation of the network model, including several key losses to define the total loss:
[0136] Use Dice and cross-entropy losses to optimize the supervised segmentation loss on the single-slice label Y k The formula is as follows: The formula is as follows:
[0137]
[0138] where P s is the prediction result, and Y k is the single-slice annotation;
[0139] The student model is also restricted by the mixed pseudo-label P fu Calculate the loss between P s and P fu To avoid the error accumulation caused by slices far from the labeled slices, the present invention adds a distance weight w s to reduce the error, The formula is as follows:
[0140]
[0141] where i and j are the slice index and voxel index, respectively, and θ max is the weight decay coefficient;
[0142] Calculate the consistency loss using unlabeled data to ensure the output consistency of the model under different perturbations. The formula is as follows:
[0143] L con = MES(P s , P t )
[0144] where P s is the prediction result of the student model, and P t is the prediction result of the teacher model;
[0145] Total loss:
[0146]
[0147] Among them, λ is the Gaussian warm-up function, which is used to prevent large prediction errors at the initial stage of training from having a negative impact on the model learning and ensure that the total loss is mainly dominated by the supervision term. t and t max represent the current training step and the maximum training step respectively.
[0148] Step 64, weight update: Use the weights θ of the student model s to update the weights θ of the teacher model through EM4 t , and the weight synchronization function is as follows:
[0149]
[0150] Among them, η is the decay hyperparameter of EM4, which is used to control the update rate.
[0151] Step 7, model evaluation: Use indicators such as Dice coefficient, Jaccard coefficient, average surface distance, and 95% Hausdorff distance to evaluate the segmentation performance of the model;
[0152] To verify the effectiveness of the present invention, the present invention conducts experiments on four publicly available medical image datasets, including the left atrium, liver tumor, pancreas, and prostate datasets. Please refer to Table 1 for the comparison of the segmentation performance of different methods on the left atrium dataset. The experimental results show that the method of the present invention is superior to the existing semi-supervised learning methods in terms of indicators such as Dice coefficient, Jaccard coefficient, average surface distance, and 95% Hausdorff distance.
[0153] Table 1: Comparison of segmentation performance of different methods on the left atrium (LA) dataset
[0154]
[0155] The overall workflow of the present invention:
[0156] Input data, including a 3D medical image dataset with a small amount of single-slice labeled data and a large amount of unlabeled data, and perform normalization and denoising processing on the data to improve the training effect of the model. Slide a window along the slice direction (such as 5 consecutive slices), generate multiple 3D volume blocks and splice them into a pseudo spatio-temporal sequence. Add Gaussian noise perturbation to the sequence and input it into the student model and the teacher model. Extract the spatio-temporal correlation in the sequence through a convolutional network, such as the dynamic changes of heart contraction. Calculate the dynamic feature consistency loss between the student and teacher models. The encoder extracts multi-scale features to assist the decoder (transposed convolution + upsampling) in reconstructing the original 3D volume. Calculate the mean square error between the reconstructed image and the original image to force the model to learn the global anatomical structure (such as organ boundaries, vascular topologies), gradually integrating from small-scale local features to large-scale global features, simulating the "local → global" visual perception, deconstructing from large-scale features to small-scale features to supplement local details, embedding an attention mechanism to filter out irrelevant background information. The teacher model synchronizes the weights of the student model through exponential moving average, combines the prior mask of level set evolution and the prediction results of the teacher model, dynamically adjusts the confidence weights, generates mixed pseudo-labels to guide the training of unlabeled data, calculates the total loss, and uses a Gaussian warm-up function to control the weights of supervised and unsupervised losses to prevent overfitting in the initial stage of training.
[0157] In summary, through dynamic-static dual feature extraction and semi-supervised collaborative training, this method achieves high-precision segmentation of 3D medical images under the condition of only single-slice labeling, significantly reducing the labeling cost (more than 90%), and at the same time supporting multi-organ (heart, liver, prostate) and multi-modal (MRI, CT) scenarios, providing an efficient tool for clinical diagnosis and surgical navigation.
[0158] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations. In addition, in the description of the present invention, unless otherwise stated, the meaning of "a plurality" is two or more.
[0159] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of the present invention. The scope of the present invention is defined by the claims and their equivalents.
Claims
1. A single-slice semi-supervised 3D medical image segmentation method, characterized in that, It includes the following steps: Step 1. Construct a backbone network based on the MT framework. The backbone network includes a dynamic correlation extraction (DCE) module and a static correlation extraction (SCE) module. Use VNet as the backbone network for medical image transmission and processing, which can be described as E s ·D s and E t ·D t , where E s and D s are the encoder and decoder of the student model, and E t and D t are the encoder and decoder of the teacher model. For SCE, establish an auxiliary reduction decoder denoted as Step 2, Data acquisition and preprocessing: Obtain 3D medical image data, which is a 3D medical image dataset containing a small amount of single-slice annotation data and a large amount of unannotated data, and perform preprocessing; Step 3, Dynamic correlation extraction: Construct a pseudo spatio-temporal sequence by recursive slice shifting and extract dynamic features; Step 4, Static correlation extraction: Execute an image restoration task through a static correlation extraction module and extract static features; Step 5, Bidirectional multi-granularity feature learning and decision fusion: Fuse the dynamic features and the static features through a bidirectional multi-granularity feature fusion module to generate a multi-scale feature representation; Step 6, Model training and optimization: Adopt a teacher-student model framework and combine hybrid pseudo-label generation and consistency loss for model training and optimization; Step 7, Model evaluation: Evaluate the segmentation performance of the model using multiple metrics.
2. The single-slice semi-supervised 3D medical image segmentation method according to claim 1, wherein In step 3, it specifically includes the following steps: Step 31, Slice shifting: Set the acquisition direction of the 3D medical image data as the sliding axis, select continuous m slices as the first group of sliding slices, and move the sliding slices to the end position of the sample. Then, repeat the above operation to recursively move the original three-dimensional volume by 2m, 3m... sm slices to obtain s new 3D volumes; Step 32, sequence splicing: Concatenate s 3D volumes into a pseudo spatio-temporal sequence X P ; Step 33, Dynamic Feature Extraction: Add random Gaussian noise and The pseudo-sequences of are input into the student model and the teacher model respectively as perturbations to extract the dynamic correlation feature fea d Volume; Step 34, calculation of dynamic correlation loss: Calculate the difference L between the dynamic correlation features and the true annotation ds : Among them, Pr i is the prediction result of the i-th layer, and Y k is the true annotation.
3. A single-slice semi-supervised 3D medical image segmentation method according to claim 1, characterized in that, In step 4, it specifically includes the following steps: Step 41. Static feature extraction: The pseudo spatio-temporal sequence X P is input into the encoder E S to extract the feature fea e . The encoder adopts a multi-layer convolutional neural network structure to extract multi-level features from the image. After each convolutional operation, the ReLU activation function is used for non-linear transformation, and downsampling is performed through the max-pooling layer. The output of the encoder is a high-dimensional feature representation fea e ; Step 42, Image Restoration: The extracted feature fea e is input into the restoration decoder to generate the restored image The restoration decoder uses deconvolution layers and upsampling operations to gradually restore the low-resolution feature map to the resolution of the original image; Step 43, recovery loss calculation: Calculate the mean square error between the recovered image and the original image X P The formula is as follows:
4. A single-slice semi-supervised 3D medical image segmentation method according to claim 1, characterized in that, In step 5, it specifically includes the following steps: Step 51, Input Feature: Input multi-scale features fae from the encoder and the decoder i ; Step 52, Bidirectional convolution operation: Perform 3×3 convolution operations on the feature maps of each scale bidirectionally, including forward convolution: starting from the smallest feature map, gradually perform convolution operations towards the large-scale feature maps to fuse the feature information from small scales to large scales; backward convolution: starting from the largest feature map, gradually perform convolution operations towards the small-scale feature maps to fuse the feature information from large scales to small scales. The formula for the bidirectional convolution operation is as follows: Fh f = Conv 3×3 (Fea f ) + Attention(Fh f-1 ) Among them, Fea f is the feature map of the f-th layer, and Fh f is the feature map after bidirectional convolution; Step 53, Attention mechanism: Filter the context information through the attention mechanism to generate a high-level feature representation with rich semantic information, and calculate the attention weights using the Sigmoid function. The formula is as follows: Att f = Sigmoid(Conv 3×3 (Fh f-1 )) Among them, Fh f-1 is the feature map of the previous layer. Step 54, Feature filtering and fusion: Multiply the attention weights by the feature maps of the current layer to filter out unimportant context information. Feature fusion includes forward feature propagation, backward feature propagation, and feature concatenation; Step 55, Decision Fusion: Extract the global feature Fea through 3D convolution and Transformer module g , and the global feature Fea g is fused with the multi-scale feature representation Fb i to generate the final segmentation result P s .
5. A single-slice semi-supervised 3D medical image segmentation method according to claim 4, characterized in that In step 54, the following specific feature fusions are included: Forward feature propagation: Starting from the smallest feature map, convolutional operations are gradually performed on the feature maps of larger scales to generate the forward feature map Fh f ; Backward feature propagation: Starting from the largest feature map, convolutional operations are gradually performed on the feature maps with smaller scales to generate the backward feature map Fg f ; Feature concatenation: Concatenate the forward and backward feature maps to generate the final feature representation Fb i .
6. A single-slice semi-supervised 3D medical image segmentation method according to claim 1, characterized in that, In step 6, it specifically includes the following steps: Step 61, generating a mask in advance: The distance regularized level set evolution is introduced to generate a prior mask for the unlabeled slices, and the mask is calculated according to the single-slice label hint. The label of the k-th slice is used as the segmentation hint for the current nearest neighboring slice, and this is repeated until the prior masks for all slices are calculated. According to the evolved level set function Generate the pseudo-label Y for the unlabeled slices P . The formula for the level set evolution is as follows: where μ, λ, and α are control parameters, d p is the distance regularization term, and g is the edge detection function; Step 62, Hybrid pseudo-label generation: Generate more reliable hybrid pseudo-labels by combining a pre-generated mask and the prediction results of the teacher model; Step 63, Loss design: Use a loss function to constrain the backpropagation of the network model; Step 64, weight update: Use the weights θ of the student model s to update the weights θ of the teacher model through EM4 t , and the weight synchronization function is as follows: Among them, η is the decay hyperparameter of EM4, which is used to control the update rate.
7. A single-slice semi-supervised 3D medical image segmentation method according to claim 6, characterized in that, In step 62, it specifically includes the following steps: Step 621, Teacher Model Prediction: Input the unlabeled 3D medical image into the teacher model, and the teacher model generates a prediction result P t ; Step 622, Confidence calculation: Calculate the confidence Conf of the prediction results of the teacher model. The formula is as follows: where p f (i) and p b (i) are the probabilities of the foreground and background, respectively; Step 623, Hybrid weight calculation: Among them, is the maximum mixing weight, T is the total number of training steps, and v is the decay coefficient; Mixed Pseudo-Label P fu Generation: w fu_model = Conf(i)w ct w fu_pseudo = (1 - Conf(i))(1 - w ct ) w model ,w pseudo = softmax[w fu_model ,w fu_pseudo P fu = w model Y p , w pseudo P t where, w model and w pseudo are the weights of the pre-generated mask and the prediction result of the teacher model respectively.
8. A single-slice semi-supervised 3D medical image segmentation method according to claim 6, characterized in that, In step 63, it specifically includes several key losses to define the total loss, including: Optimize the monolithic label Y using Dice and cross-entropy losses k for the supervised segmentation loss The formula is as follows: Among them, P s is the prediction result, and Y k is the single-slice annotation; The student model is also restricted by the mixed pseudo-label P fu to calculate the loss between P s and P fu by adding a distance weight w s to mitigate the error, and the formula is as follows: where \(i\) and \(j\) are the slice index and voxel index respectively, and \(\theta\) max is the weight decay coefficient; Calculate the consistency loss using unannotated data to ensure the output consistency of the model under different perturbations. The formula is as follows: L con = MES(P s , P t ) Among them, P s is the prediction result of the student model, and P t is the prediction result of the teacher model; Total loss: Among them, λ is the Gaussian warm-up function, and t and t max represent the current training step and the maximum training step respectively.
9. A single-slice semi-supervised 3D medical image segmentation method according to claim 1, wherein In the said step 7, the multiple said metrics specifically include Dice coefficient, Jaccard coefficient, average surface distance, and 95% Hausdorff distance.