Semi-supervised medical image segmentation method based on data augmentation strategy
Patent Information
- Application Number
- CN202410025592.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-08
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-01-08
AI Technical Summary
然而,由于医学影像数据的固有特性,不同解剖结构之间的边界常常模糊不清,这给获取标注数据带来了巨大挑战
[0051]首先,以往的半监督医学影像分割方法在训练时,采用的数据增强的策略是随机仿射变换、随机弹性变换和随机对比度变换。这些方法实现起来非常简单,并且经验上已经证明可以减少过拟合并提高看不见的示例的性能。但这种提升随着新的分割方法的出现,取得的收益甚微。本发明通过对有限标注数据和无标记数据进行增强,具体的,使用生成的互补中心伪影掩模与不同标记图像进行体素点积融合,构建新的标记数据以及对应的标签。这部分标记数据在训练中通过子模型输出概率映射与重构标记之间的均方误差、Dice损失来反向传播,进而调整网络参数,优化模型分类结果。通过在合成图像中引入伪影实现体素融合点积,打破训练数据体素级别的相关性,引入新的不确定性和噪声。标签的存在抑制了噪声,并鼓励模型从这种新引入的不确定性中学习。为了进一步增强无标记数据的多样性并改善网络的泛化能力,本发明利用扩散模型进行数据增强,利用拉普拉斯金字塔融合将原始和重构数据的切片融合作为无标记数据的扩充,并通过无监督损失来优化网络参数。本发明提出了一种新型的网络分割模型,旨在通过整合不确定性估计和相互一致性约束来融合来自不同维度和分辨率细节的语义信息,从而进一步提高分割性能。本发明所提出的模型分割性能在广泛使用的医学图像数据集上表现出极佳的分割准确性,具有极佳的适用性。
Smart Images

Figure CN117710681B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer signal processing, specifically relating to a semi-supervised medical image segmentation method based on the consistency of two data augmentation strategies and an attention mechanism. Background Technology
[0002] Medical images reflect the anatomical structures or functional tissues within the human body. They are discrete image representations generated through sampling or reconstruction, mapping numerical values to different spatial locations. Compared to natural images, medical images typically exhibit low contrast, blurred boundaries, and inaccurate visual recognition. Commonly used medical imaging techniques include computed tomography (CT), ultrasound, magnetic resonance imaging (MRI), and X-rays. The goal of medical image segmentation is to separate different organs or lesion regions within an image. This task involves automatically or semi-automatically classifying pixels in a medical image, thereby segmenting the image into different meaningful regions. In medical imaging, these regions often correspond to different tissue types, pathologies, organs, or other biological structures.
[0003] With the rapid development and widespread adoption of medical imaging equipment, imaging technology has been widely used in clinical practice, becoming an indispensable auxiliary tool for disease diagnosis, surgical planning, prognostic assessment, and follow-up. Medical images have become one of the most important sources of evidence for clinical analysis and medical intervention. Medical image segmentation can extract key information from specific tissue images and is a crucial step in realizing medical image visualization. The segmented images are provided to doctors for various tasks such as quantitative analysis of tissue volume, diagnosis, localization of pathological changes, delineation of anatomical structures, and treatment planning. Medical images contain a vast amount of information, and manually delineating target regions in medical images is a time-consuming and laborious task, significantly increasing the burden on clinicians' daily work.
[0004] Compared to traditional segmentation algorithms based on thresholding, clustering techniques, and deformable models, combining medical image segmentation with deep learning methods has achieved great success. Various approaches have been proposed by researchers, including the development of unique segmentation architectures based on convolutional neural networks. Among these methods, fully supervised learning performs exceptionally well in segmentation tasks due to the availability of abundant labeled data. However, due to the inherent characteristics of medical image data, the boundaries between different anatomical structures are often blurred, posing a significant challenge to obtaining labeled data. With the emergence of semi-supervised learning, there is increasing interest in how to utilize abundant unlabeled data in combination with labeled data to improve the overall segmentation performance of deep learning models. Existing semi-supervised segmentation methods are mainly divided into two categories. The first category is based on consistency models, which relies on the assumption that small perturbations in the input should not lead to large deviations in the corresponding outputs. The second category is based on entropy minimization strategies, aiming to achieve compact, low-entropy clustering for each category. In these two mainstream AI semi-supervised medical image segmentation methods, the former can reduce the uncertainty of segmentation results by constraining the consistency between different outputs of sub-models under limited labeled data training, while the latter directly uses entropy minimization to constrain the model to produce high-confidence predictions. The performance of both methods is constrained by the limited labeled data. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of the existing technology and propose a semi-supervised medical image segmentation method based on data augmentation strategy. This method adopts a labeled data augmentation strategy, which achieves labeled data fusion through voxel dot product using complementary masks, and further mines valuable knowledge from the limited labels. At the same time, it utilizes an unlabeled data augmentation strategy, which reconstructs pseudo-training data using a diffusion model, increases the diversity of unlabeled samples, and significantly improves the generalization ability of the network. Finally, it uses pseudo-labels and data consistency to constrain the output of multiple sub-network models of the convolutional neural network to improve the overall segmentation performance.
[0006] The technical solution adopted in this invention is:
[0007] This invention provides a semi-supervised medical image segmentation method based on data augmentation strategies, comprising the following steps:
[0008] 1) Divide the original medical image data into training and testing sets. Then, further divide the training set into labeled and unlabeled data. Next, perform random cropping as input for training the semi-supervised medical image segmentation model.
[0009] 2) The original medical image data is reconstructed using a diffusion model, and the reconstructed data is then fused with the original medical image data using Laplacian pyramid convolution. This is used as an extension of the unlabeled data for training the semi-supervised medical image segmentation model.
[0010] 3) Perform data augmentation on the batch input of the neural network model in each loop. Specifically, perform binary classification on the labeled data in each batch, and then randomly extract one image from each binary classification as the foreground and background of the fused labeled data. Use complementary masks to achieve voxel dot product fusion (a type of voxel fusion). The corresponding operation is also applied to the labels of the corresponding images to obtain fused labeled data with half the amount of the original batch input labeled data and its fused labels as the expansion of labeled data for training the semi-supervised medical image segmentation model.
[0011] 4) A semi-supervised medical image segmentation model incorporating a convolutional neural network is constructed. The specific process is as follows: The encoder extracts features at different scales from the input image through convolutional blocks and multi-layer downsampling. During the upsampling process of the decoder, skip connections are used to fuse features of the same size and dimension. Through this operation, high-level features are directly passed to the decoder across different stages, allowing for better feature fusion during resolution restoration and helping the decoder better restore the details and boundaries of the segmented target. Furthermore, a hierarchical difference structure is adopted, with four decoders using different skip connections to achieve feature fusion. Specifically, skip connections are removed in the i-th layer of decoder i to encourage diversity in decoder outputs. Channel and spatial attention modules are added after layer i to enhance the difference in feature fusion. For labeled data input, supervised learning training is performed using real labels. For unlabeled data input, four different decoders extract complementary features with dimensional differences from the same encoder input. A sharpening function is used to transform the probability maps output by each decoder into pseudo-labels. Consistency constraints are applied using these pseudo-labels to encourage different decoders to learn from each other and output consistent results. The semi-supervised medical image segmentation model was tested using a test set, and evaluation metrics were calculated. If the evaluation metrics met the requirements, the optimal parameters of the convolutional neural network model and data augmentation in the semi-supervised medical image segmentation model were obtained; otherwise, the parameters of the convolutional neural network model and data augmentation were adjusted, and the semi-supervised medical image segmentation model was retrained and tested. The adjusted parameters of the convolutional neural network model included the weight α for fusing labeled data during labeled data supervised learning training, and the weight coefficient λ of the supervised loss in the sum of supervised and unsupervised losses. s With unsupervised loss weight coefficient λ u The adjusted data augmentation parameter is the size factor β of the zero-center mask corresponding to the complementary mask in step 3).
[0012] Preferably, step 1) is as follows:
[0013] a1) Divide the original medical image data into training and test sets, and then further divide the training set into two different semi-supervised experimental settings: 10% labeled data and 20% labeled data.
[0014] a2) The training set images are read according to their indices and randomly cropped into blocks as input for training the semi-supervised medical image segmentation model. Then, the cropped standard-sized images are randomly rotated and flipped as standard data augmentation methods.
[0015] Preferably, step 2) is as follows:
[0016] b1) Utilizing the concept of iterative reconstruction, a two-dimensional diffusion model is constructed to reconstruct the slices, and ADMM is applied to update the differentiator, thereby promoting crosstalk between slices through measurement information. Specifically, regularized reconstruction derives the reconstructed image u based on the measured value h. * :
[0017]
[0018] Where u is the image to be reconstructed; the penalty function F(u) = ||D z u||1;D z To calculate finite differences along the z-axis; ||||2 denotes the 2-norm, and ||||1 denotes the 1-norm.
[0019] b2) Laplacian pyramid convolutional fusion involves downsampling to generate a Gaussian pyramid, then fusing the Gaussian pyramid features from the diffusion model output of step b1) with the same size as the original medical image data. The fused image is then upsampled to achieve the same size as the original medical image. This process is repeated layer by layer from the bottom up until the top layer is reached. The resulting image is used as an augmentation of unlabeled data for training the semi-supervised medical image segmentation model. The k-th layer of the Laplacian pyramid yields the following features:
[0020]
[0021] In the formula, G k To integrate the features of the k-th layer of the Laplace pyramid, and The corresponding Laplacian pyramid layers for the two input images; weight factor β k This determines the contribution of each image to layer k, where β k Let it be 0.5, k∈[1,6].
[0022] Preferably, step 3) specifically includes the following steps:
[0023] c1) During training, the labeled data is randomly divided into two subsets. From each subset, two labeled images are randomly selected. Generate a zero-center mask H, W, and D represent the height, width, and depth of the image, used to distinguish whether a pixel in the image comes from the foreground or background image. If it comes from the foreground image, it is assigned a value of 1; otherwise, it is assigned a value of 0. The size of the zero-center mask is βH×βW×βD, where the size coefficient β∈(0,1).
[0024] c2) Let the image Foreground image, image Background image; Set image Artifact mask of the central voxel region with a set size (zero-center mask) Perform voxel dot product fusion, then process the image. Complementary size center voxel region artifact mask Voxel dot product fusion is performed, and the two fused images are superimposed to generate a blended image with the same size as the original image. In addition, from the image tags The fused voxel dot product is obtained from the image. tags The fused voxel dot product is obtained, and the two are superimposed to obtain a hybrid label that matches the original label size. The specific formula is as follows:
[0025]
[0026]
[0027] in 1 is used to obtain 1,D when pixels in an image originate from the foreground image. L ⊙ indicates a labeled dataset, and ⊙ indicates element-wise multiplication.
[0028] c3) Based on batch input N L Generate N labeled data. L / 2 blended images and N L / 2 mixed labels are used for training the semi-supervised medical image segmentation model.
[0029] Preferably, step 4) specifically includes the following steps:
[0030] d1) Use a 5-layer convolutional network in the encoder with 4 layers of downsampling to extract different dimensional features of each input image. Then, in the upsampling process of the decoder, the same size and dimensional features are fused through skip connections. Furthermore, feature fusion is achieved by using different skip connections in four decoders. Specifically, skip connections are removed in the i-th layer of decoder i, and channel attention modules and spatial attention modules are added after the i-th layer.
[0031] d2) Define a semi-supervised segmentation task, using p(y pred |x;f θ ) represents the generation probability map of x by the backbone network (one decoder has one backbone network) in a semi-supervised medical image segmentation model, where y pred f represents the segmentation result of the backbone network. θ The parameters of the backbone network, x∈X H×W×D Let X represent the input image set in the batch input; and define y... l ∈Y H×W×D Let Y represent the labeled data, and let Y represent the labeled dataset in the batch input. The sets recording the labeled data and their labels in the batch input, and the unlabeled dataset in the batch input, are respectively represented as follows: and N L D L The number of labeled data, N U D U The number of unlabeled data points, N L <<N U D U The unlabeled data includes the expanded unlabeled data after performing step 2).
[0032] D L Each labeled data When training a semi-supervised medical image segmentation model, labels are used. Supervised loss is used to guide the learning process of the backbone network. Supervised loss is defined as follows:
[0033]
[0034] The loss is calculated using the Dice loss function. This indicates that the j-th backbone network is connected to the i-th input image. The generated probability graph, y j pred This represents the segmentation result of the j-th backbone network.
[0035] When training a semi-supervised medical image segmentation model using fused labeled data as input, the supervised loss is defined as follows:
[0036]
[0037] MSE stands for mean squared error.
[0038] The overall supervised loss for the labeled data is as follows:
[0039]
[0040] The parameter α represents the weight of the fused labeled data;
[0041] d3)D U Unlabeled data When training a semi-supervised medical image segmentation model, four probability maps output by different decoders are obtained. These maps are then processed using the following sharpening function to obtain the pseudo-label generated by the j-th decoder for the i-th unlabeled data.
[0042]
[0043] The hyperparameter T controls the sharpening temperature, thereby enhancing regularization by minimizing the entropy constraint. The unsupervised loss is calculated using the mean squared error (MSE) between the pseudo-label and the probability map, defined as follows:
[0044]
[0045] in This represents the k-th decoder in the segmentation network model for the i-th input image. Generated pseudo-tags.
[0046] d4) The final loss function under semi-supervised training is the sum of the supervised loss and the unsupervised loss, defined as:
[0047] L total =λ s L' sup +λ u L unsup
[0048] λ s λ represents the supervised loss weighting coefficient. u This represents the unsupervised loss weighting coefficient.
[0049] d5) Test the segmentation results of the semi-supervised medical image segmentation model using a test set, calculate the evaluation index. If the evaluation index meets the requirements, obtain the optimal parameters of the convolutional neural network model and data augmentation in the semi-supervised medical image segmentation model. Otherwise, adjust the parameters of the convolutional neural network model and data augmentation, and retrain and test the semi-supervised medical image segmentation model.
[0050] Compared with the prior art, the beneficial effects of this invention are as follows:
[0051] First, previous semi-supervised medical image segmentation methods employed data augmentation strategies such as stochastic affine transformation, stochastic elastic transformation, and stochastic contrast transformation during training. These methods are very simple to implement and have been empirically proven to reduce overfitting and improve performance on unseen examples. However, this improvement has diminished with the emergence of new segmentation methods. This invention augments limited labeled and unlabeled data by using a generated complementary center artifact mask to perform voxel dot product fusion with different labeled images, constructing new labeled data and corresponding labels. This labeled data is backpropagated during training through the mean squared error and Dice loss between the sub-model output probability mapping and the reconstructed labels, thereby adjusting network parameters and optimizing the model's classification results. By introducing artifacts into the synthetic image to achieve voxel fusion dot product, the voxel-level correlation of the training data is broken, introducing new uncertainties and noise. The presence of labels suppresses noise and encourages the model to learn from this newly introduced uncertainty. To further enhance the diversity of unlabeled data and improve the generalization ability of the network, this invention utilizes a diffusion model for data augmentation, employs Laplacian pyramid fusion to fuse slices of the original and reconstructed data as an extension of the unlabeled data, and optimizes network parameters through unsupervised loss. This invention proposes a novel network segmentation model that aims to further improve segmentation performance by integrating semantic information from details of different dimensions and resolutions through uncertainty estimation and mutual consistency constraints. The proposed model demonstrates excellent segmentation accuracy on widely used medical image datasets and exhibits excellent applicability. Attached Figure Description
[0052] Figure 1 This is a network structure diagram of the present invention;
[0053] Figure 2 This is a schematic diagram of the label data augmentation strategy in this invention. Detailed Implementation
[0054] The invention will now be further described with reference to the accompanying drawings.
[0055] like Figure 1 As shown, the semi-supervised medical image segmentation method based on data augmentation strategy is as follows:
[0056] 1) Preprocess the raw medical image data, specifically:
[0057] a1) Select a publicly available medical image dataset (3D CT|MRI data or 2D slices can be used), divide it into training and test sets, and then further divide the training set according to two different settings of semi-supervised experiments: 10% labeled data and 20% labeled data.
[0058] a2) The training set images are read according to their indices and randomly cropped into blocks as input for training the semi-supervised medical image segmentation model. Then, the cropped standard-sized images are randomly rotated and flipped as standard data augmentation methods.
[0059] 2) A diffusion model is used to reconstruct the original medical image data, and the reconstructed data is then fused with the original medical image data using a Laplacian pyramid convolution. This fusion serves as supplementary unlabeled data for training the semi-supervised medical image segmentation model. Specifically:
[0060] b1) Utilizing the concept of iterative reconstruction, a traditional two-dimensional diffusion model is used to reconstruct the slices, and ADMM is applied to update the differentiator, thereby promoting crosstalk between slices through measurement information. Specifically, regularized reconstruction derives the image u based on the measured value h. * :
[0061]
[0062] Where u is the image to be reconstructed; the penalty function F(u) = ||D z u||1,D z To calculate finite differences along the z-axis; ||||2 denotes the 2-norm, and ||||1 denotes the 1-norm.
[0063] b2) If only the diffusion model output is used to expand the unlabeled data, its structural similarity index (SSI) is 0.36 and its peak signal-to-noise ratio (PSNR) is 64.45%, which are 12.2% and 3.6% lower than the original data, respectively, indicating a significant loss of slice correlation in the z-direction. It is evident that relying solely on the diffusion model output reduces segmentation performance. This invention proposes a Laplacian pyramid convolutional fusion method to fuse features of the original image and the diffused reconstructed image at different scales, preserving high-frequency components while suppressing noise to maintain slice correlation. The Laplacian pyramid convolutional fusion involves downsampling to generate a Gaussian pyramid, then fusing the diffusion model output from step b1) with the Gaussian pyramid features of the same size as the original medical image data. The fused image is then upsampled to have the same size as the original medical image. Starting from the bottom layer, fusion and upsampling are performed layer by layer until the top layer is reached, and the final image is used as an expansion of the unlabeled data. For the k-th layer of the Laplacian pyramid, the following features can be obtained:
[0064]
[0065] In the formula, G k To integrate the features of the k-th layer of the Laplace pyramid, and The corresponding Laplacian pyramid layers for the two input images; weight factor β k This determines the contribution of each image to layer k, where β k Set to 0.5, k∈[1,6]. Pyramid fusion reduces artifacts, over-enhancement, and discontinuities in transition regions, preserving detail while suppressing noise. The SSI between fused data slices is 0.40, and the PSNR is 65.68%, representing improvements of 11.1% and 1.91% respectively compared to the output of the simple diffusion model. This effectively addresses the correlation loss problem in the z-axis direction.
[0066] 3) such as Figure 2 As shown, data augmentation operations are performed on the labeled data input of each batch in each iteration of the neural network model. Specifically, the labeled data in each batch is binary classified, and then one image is randomly selected from each binary classification as the foreground and background of the fused labeled data. Voxel dot product fusion (a type of voxel fusion) is achieved using complementary masks. The corresponding operation is also applied to the labels of the corresponding images, resulting in fused labeled data with half the amount of the original batch input labeled data and its fused labels. This is used to expand the labeled data for training the semi-supervised medical image segmentation model. The specific steps include:
[0067] c1) During training, the limited labeled data is randomly divided into two subsets. From each subset, two labeled images are randomly selected. Generate a zero-center mask Used to distinguish whether a pixel comes from the foreground image (1) or the background image (0). The size of the zero-center mask is βH×βW×βD, where the size factor β∈(0,1).
[0068] c2) Image (Foreground image) and a set-size center voxel region artifact mask (zero-center mask) Perform voxel dot product fusion, then process the image. (Background image) and artifact mask of complementary size central voxel region Voxel dot product fusion is performed, and the two fused images are superimposed to generate a blended image with the same size as the original image. Furthermore, similarly, from the label ( The fusion voxel dot product is obtained from the label, and then it is combined with the label. ( The fusion voxel dot products obtained on the original labels are superimposed to obtain a hybrid label that matches the size of the original label. The specific formula is as follows:
[0069]
[0070]
[0071] in 1 is used to obtain 1,D when pixels in an image originate from the foreground image. L represents a labeled dataset, while ⊙ represents element-wise multiplication.
[0072] c3) N based on batch input L N tags are generated from data. L / 2 blended images and N L Two mixed labels were used for training a semi-supervised medical image segmentation model. Figure 1 It can be seen that the mixed labels have a constraining effect on the output probability mapping of the four sub-networks.
[0073] 4) Construct a semi-supervised medical image segmentation model that includes a convolutional neural network model, specifically including the following steps:
[0074] d1) A 5-layer convolutional network with 4-layer downsampling is used in the encoder to extract different dimensional features of each input image. Then, during the upsampling process in the decoder, skip connections are used to fuse features of the same size and dimension. Through this operation, high-level features can be directly passed to the decoder across different stages, allowing for better feature fusion during resolution restoration and helping the decoder to better restore the details and boundaries of the segmented target. Furthermore, a hierarchical difference structure is adopted, using different skip connections in four decoders to achieve feature fusion. Specifically, skip connections are removed in the i-th layer of decoder i to encourage diversity in decoder output. Simultaneously, channel attention and spatial attention modules are added after the i-th layer to enhance the difference in feature fusion.
[0075] d2) Define a semi-supervised segmentation task, using p(y pred |x;f θ ) represents the generation probability map of x by the backbone network (one decoder has one backbone network) in a semi-supervised medical image segmentation model, where y pred f represents the segmentation result of the backbone network. θ The parameters of the backbone network, x∈X H×W×D Let X represent the input image set in the batch input; and define y... l ∈Y H×W×DLet Y represent the labeled data, and let Y represent the labeled dataset in the batch input. The sets recording the labeled data and their labels in the batch input, and the unlabeled dataset in the batch input, are respectively represented as follows: and N L D L The number of labeled data, N U D U The number of unlabeled data points, N L <<N U D U The unlabeled data includes the expanded unlabeled data after performing step 2).
[0076] Tag data When inputting as model training, labels are used Supervised loss is used to guide the learning process of the backbone network. Supervised loss is defined as follows:
[0077]
[0078] The loss is calculated using the Dice loss function. This indicates that the j-th backbone network is connected to the i-th input image. The generated probability graph, y j pred This represents the segmentation result of the j-th backbone network.
[0079] When training a semi-supervised medical image segmentation model using fused labeled data as input, the supervised loss is defined as follows:
[0080]
[0081] MSE stands for mean squared error.
[0082] The overall supervised loss for the labeled data is as follows:
[0083]
[0084] The parameter α represents the weight of the fused labeled data;
[0085] d3) Unlabeled data When the input is used for model training, four different decoder output probability maps are obtained. These maps are then processed by the following sharpening function to obtain the pseudo-label generated by the j-th decoder for the i-th unlabeled data.
[0086]
[0087] The hyperparameter T controls the sharpening temperature, thereby enhancing the model's regularization capability by minimizing the entropy constraint. The unsupervised loss is calculated using the mean squared error (MSE) between the pseudo-label and the probability map, defined as follows:
[0088]
[0089] in This represents the k-th decoder in the segmentation network model for the i-th input image. Generated pseudo-tags.
[0090] d4) The final loss function under semi-supervised training is the sum of the supervised loss and the unsupervised loss, defined as:
[0091] L total =λ s L' sup +λ u L unsup
[0092] d5) The segmentation results of the semi-supervised medical image segmentation model are tested using a test set, and evaluation metrics are calculated. If the evaluation metrics meet the requirements, the optimal parameters of the convolutional neural network model and data augmentation in the semi-supervised medical image segmentation model are obtained; otherwise, the parameters of the convolutional neural network model and data augmentation are adjusted, and the semi-supervised medical image segmentation model is retrained and tested. Among them, the adjusted parameters of the convolutional neural network model are the weight α for fused labeled data and the supervised loss weight coefficient λ. s With unsupervised loss weight coefficient λ u The adjusted data augmentation parameter is the size factor β of the zero-center mask.
[0093] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A semi-supervised medical image segmentation method based on data augmentation strategy, characterized in that: Includes the following steps: 1) Divide the original medical image data into training and testing sets, then further divide the training set into labeled and unlabeled data, and then perform random cropping as input for training the semi-supervised medical image segmentation model; 2) The original medical image data is reconstructed using a diffusion model, and the reconstructed data is then fused with the original medical image data using a Laplacian pyramid convolution as an augmentation for unlabeled data; 3) Perform data augmentation on the batch input of the neural network model in each loop. Specifically, perform binary classification on the labeled data in each batch, and then randomly extract an image from each binary classification as the foreground and background of the fused labeled data. Use complementary masks to achieve voxel dot product fusion. The corresponding operation is also applied to the labels of the corresponding images to obtain fused labeled data and its fused labels, which is half the amount of the original batch input labeled data, as an expansion of the labeled data. 4) Construct a semi-supervised medical image segmentation model containing a convolutional neural network model. The specific process is as follows: The encoder extracts features of different scales of the input image through convolutional blocks and multi-layer downsampling. In the upsampling process of the decoder, the fusion of features of the same size and dimension is achieved through skip connections. Moreover, feature fusion is achieved through four decoders using different skip connections. Specifically, skip connections are removed in the i-th layer of decoder i, and channel and spatial attention modules are added after the i-th layer. For labeled data input, supervised learning training is performed using real labels; for unlabeled data input, four different decoders are used to extract complementary features with dimensional differences from the same encoder input. A sharpening function is used to transform the probability maps output by each decoder into pseudo-labels. Consistency constraints are imposed by using these pseudo-labels to encourage different decoders to learn from each other and output consistent results. The segmentation results of the semi-supervised medical image segmentation model are tested using a test set, and evaluation indicators are calculated. If the evaluation indicators meet the requirements, the optimal parameters of the convolutional neural network model and data augmentation in the semi-supervised medical image segmentation model are obtained; otherwise, the parameters of the convolutional neural network model and data augmentation are adjusted, and the semi-supervised medical image segmentation model is retrained and tested. Among them, the adjusted parameters of the convolutional neural network model are the weights of the fused labeled data in the labeled data supervised learning training, as well as the weight coefficients of the supervised loss and the weight coefficients of the unsupervised loss in the sum of supervised loss and unsupervised loss. The adjusted data augmentation parameters are the size coefficients of the zero-center mask corresponding to the complementary mask in step 3). Step 3) specifically includes the following steps: c1) During training, the labeled data is randomly divided into two subsets; from each subset, two labeled images are randomly selected. Generate a zero-center mask. H, W, and D represent the height, width, and depth of the image, used to distinguish whether a pixel in the image comes from the foreground or background image. A value of 1 is assigned if the pixel comes from the foreground image, and 0 otherwise. The size of the zero-center mask is... The size factor of the zero-center mask ; c2) Let the image Foreground image, image Background image; Set image With zero-center mask Perform voxel dot product fusion, then image Complementary size center voxel region artifact mask Voxel dot product fusion is performed, and the two fused images are superimposed to generate a blended image with the same size as the original image. ; Furthermore, from the image tags The fused voxel dot product is obtained from the image. tags The fused voxel dot product is obtained, and the two are superimposed to obtain a hybrid label that matches the original label size. The specific formula is as follows: in , 1 is obtained when a pixel in the image originates from the foreground image. ⊙ represents a labeled dataset; c3) Based on batch input Each labeled data is generated. / 2 blended images and / 2 mixed labels are used for training the semi-supervised medical image segmentation model.
2. The semi-supervised medical image segmentation method based on data augmentation strategy according to claim 1, characterized in that: Step 1) is as follows: a1) Divide the original medical image data into training and testing sets, and then further divide the training set into two different semi-supervised experimental settings: 10% labeled data and 20% labeled data. a2) Read the training set images according to the index, randomly crop them into blocks as the training input for the semi-supervised medical image segmentation model; then perform random rotation and flipping operations on the cropped standard-sized images.
3. The semi-supervised medical image segmentation method based on data augmentation strategy according to claim 1, characterized in that: Step 2) is as follows: b1) Utilizing the concept of iterative reconstruction, a two-dimensional diffusion model is constructed to reconstruct the slices, and ADMM is applied to update the differentiator, thereby promoting crosstalk between slices through measurement information; wherein, regularized reconstruction yields the reconstructed image based on the measured value h. : u is the image to be reconstructed; penalty function ; To calculate the finite difference along the z-axis; Describes the 2-norm. Represents the 1-norm; b2) Laplacian pyramid convolutional fusion involves downsampling to generate a Gaussian pyramid, then fusing the Gaussian pyramid features from the diffusion model output of step b1) with the same size as the original medical image data. The fused image is then upsampled to ensure it has the same size as the original medical image. This process is repeated layer by layer from the bottom up until the top layer is reached. The resulting image is used as an augmentation of unlabeled data for training the semi-supervised medical image segmentation model. The k-th layer of the Laplacian pyramid yields the following features: In the formula, To integrate the features of the k-th layer of the Laplace pyramid, and For the corresponding Laplacian pyramid layers of the two input images; weight factors This determines the contribution of each image to layer k.
4. The semi-supervised medical image segmentation method based on data augmentation strategy according to claim 1, characterized in that: Step 4) specifically includes the following steps: d1) Use a 5-layer convolutional network in the encoder with 4 layers of downsampling to extract different dimensional features of each input image. Then, in the upsampling process of the decoder, the same size and dimensional features are fused through skip connections. Furthermore, feature fusion is achieved by using different skip connections in four decoders. Specifically, skip connections are removed in the i-th layer of decoder i, and channel attention modules and spatial attention modules are added after the i-th layer. d2) Define a semi-supervised segmentation task, using... This represents the generation probability graph of x by the backbone network in a semi-supervised medical image segmentation model, where... This represents the segmentation result of the backbone network. The parameters representing the backbone network, Let X represent the input image set in the batch input; and define... Labels representing marked data, The labeled dataset in the batch input represents the set of labeled data and their labels in the batch input, and the unlabeled dataset in the batch input is represented as follows: and , for Number of data items marked in the middle for The number of unlabeled data points , The unlabeled data includes the expanded unlabeled data after performing step 2); Each labeled data When training a semi-supervised medical image segmentation model, labels are used. Supervised loss is used to guide the learning process of the backbone network. Supervised loss is defined as follows: The loss is calculated using the Dice loss function. This indicates that the j-th backbone network is connected to the i-th input image. The generated probability map, This represents the segmentation result of the j-th backbone network; When training a semi-supervised medical image segmentation model using fused labeled data as input, the supervised loss is defined as follows: MSE stands for mean squared error; The overall supervised loss for the labeled data is as follows: The parameter α represents the weight of the fused labeled data; d3) Unlabeled data When training a semi-supervised medical image segmentation model, four probability maps output by different decoders are obtained. These maps are then processed using the following sharpening function to obtain the pseudo-label generated by the j-th decoder for the i-th unlabeled data. ; The hyperparameter T controls the sharpening temperature; the unsupervised loss is calculated using the mean square error (MSE) between the pseudo-label and the probability map, defined as follows: in This represents the k-th decoder in the segmentation network model for the i-th input image. Generated pseudo-tags; d4) The final loss function under semi-supervised training is the sum of the supervised loss and the unsupervised loss, defined as: The supervised loss weighting coefficients are... The unsupervised loss weight coefficient; d5) Test the segmentation results of the semi-supervised medical image segmentation model using a test set, calculate the evaluation index. If the evaluation index meets the requirements, obtain the optimal parameters of the convolutional neural network model and data augmentation in the semi-supervised medical image segmentation model. Otherwise, adjust the parameters of the convolutional neural network model and data augmentation, and retrain and test the semi-supervised medical image segmentation model.