Training Method and Segmentation Method of Image Segmentation Model Based on Semi-Supervised Learning
By alternately optimizing the teacher and student models in medical image segmentation, and using labeled and unlabeled data sets, the problem of difficulty in obtaining labeled data is solved, and the segmentation performance and generalization ability of the model are improved.
Patent Information
- Application Number
- CN202111338238.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-11-12
AI Technical Summary
The existing medical image segmentation method based on deep learning requires sufficient segmentation of labeled data. It is difficult to obtain labeled data and pseudo labeling noise will affect the performance of the model. The existing semi-supervised learning method cannot effectively utilize the features of unlabeled data.
By determining the second training sample set based on the first training sample set and the initial teacher model, the parameters of the initial student model are optimized, and the loss function is determined in combination with the labeled and unlabeled data sets, the initial teacher and student models are alternately optimized until the preset conditions are met.
The performance of the image segmentation model is improved, and the rational use of labeled and unlabeled data is improved, which improves the prediction accuracy and generalization ability of the model.
Smart Images

Figure CN114255237B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of medical image processing, and particularly relates to a training method and a segmentation method for an image segmentation model based on semi-supervised learning. Background Art
[0002] In recent years, deep learning models have been widely applied in the field of medical image segmentation and have achieved remarkable performance. However, a common problem with segmentation methods based on deep learning technology is that sufficient segmentation annotations are required for model training. For medical images, obtaining annotated data requires professional knowledge and time.
[0003] Therefore, a semi-supervised deep learning approach has been applied to medical image segmentation, and there are generally two commonly used semi-supervised deep learning methods, namely, a method of generating pseudo-annotations based on self-training or co-training, and a method of regularizing the model based on the smoothness assumption. Among them, the method of generating pseudo-annotations based on self-training or co-training is to train a segmentation model based on a training sample set of annotated data, then use the trained segmentation model to predict unannotated data, and select pseudo-annotations with high credibility to add to the training sample set, and continue to train the segmentation model until the model parameters of the segmentation model meet the cut-off conditions. This method requires selecting pseudo-annotations with high credibility, but when there are noisy pseudo-annotations in the training sample set, the model performance of the segmentation model will deteriorate. The method of regularizing the model based on the smoothness assumption is to obtain multiple input images by performing different data transformations on the input image, and then input the multiple input images into the segmentation model respectively to make the prediction results of the multiple input images tend to be similar. This method means that the same sample is close in the data distribution space after different transformations, and the corresponding prediction results are also similar, but it does not explain that the annotated samples and the unannotated samples are close in the data distribution space, and the features of the unannotated data cannot be fully extracted, thus the model performance of the segmentation model cannot be guaranteed.
[0004] Therefore, the existing technology still needs to be improved. Summary of the Invention
[0005] The technical problem to be solved by this application is to provide a training method and a segmentation method for an image segmentation model based on semi-supervised learning in view of the deficiencies of the prior art.
[0006] To solve the above technical problem, in the first aspect of the embodiments of this application, a training method for an image segmentation model based on semi-supervised learning is provided, and the training method includes:
[0007] Determine a second training sample set based on the first training sample set and the initial teacher model, and optimize the model parameters of the initial student model based on the second training sample set to obtain a candidate student model, where each training image in the first training sample set does not carry a true annotation;
[0008] Determine a loss function based on the second training sample set, the candidate student model, the initial teacher model, and the third training sample set, and optimize the model parameters of the initial teacher model based on the loss function, where each training image in the third training sample set carries a true annotation;
[0009] Take the candidate teacher model as the initial teacher model and the candidate student model as the initial student model, and continue to execute the step of determining the second training sample set based on the first training sample set and the initial teacher model until the model parameters of the initial student model meet the preset conditions.
[0010] The training method of the image segmentation model based on semi-supervised learning, where before determining the second training sample set based on the first training sample set and the initial teacher model, the method includes:
[0011] Train a preset network model with a fourth training sample set to obtain a pre-trained teacher model, where each training image in the fourth training sample set carries a true annotation;
[0012] Train the pre-trained teacher model with a fifth training sample set to obtain an initial teacher model, and train the preset network model with the fifth training sample set to obtain an initial student model, where some training images in the fifth training sample set carry true annotations and some training images do not carry true annotations.
[0013] The training method of the image segmentation model based on semi-supervised learning, where determining the second training sample set based on the first training sample set and the initial teacher model specifically includes:
[0014] For each first training image in the first training sample set, input the first training image into the initial teacher model, and output a first predicted annotation corresponding to the first training image through the initial teacher model;
[0015] Take each first training image and its corresponding first predicted annotation as a training sample, and take the training sample set composed of all the obtained training samples as the second training sample set.
[0016] The training method of the image segmentation model based on semi-supervised learning, where determining the loss function based on the second training sample set, the candidate student model, the initial teacher model, and the third training sample set specifically includes:
[0017] Determine a first loss term based on the third training sample set and the candidate student model;
[0018] Determine a second loss term based on the third training sample set and the initial teacher model;
[0019] Determine a third loss term based on the second training sample set and the initial teacher model;
[0020] Determine a loss function based on the first loss term, the second loss term, and the third loss term.
[0021] The training method of the image segmentation model based on semi-supervised learning, wherein determining the third loss term based on the second training sample set and the initial teacher model specifically includes:
[0022] Perform data transformation on each second training image in the second training sample set and its corresponding second predicted annotation respectively;
[0023] Input each data-transformed second training image into the initial teacher model, and determine the predicted annotation corresponding to each data-transformed second training image through the initial teacher model;
[0024] Determine the third loss term based on each transformed second predicted annotation and each predicted annotation.
[0025] The training method of the image segmentation model based on semi-supervised learning, wherein the data transformation is random data transformation, and the random data transformation corresponding to each of the second training image and its corresponding second predicted annotation is the same.
[0026] The second aspect of the embodiments of the present application provides a training device for an image segmentation model based on semi-supervised learning, and the training device includes:
[0027] A first optimization module, configured to determine a second training sample set based on a first training sample set and an initial teacher model, and optimize the model parameters of the initial student model based on the second training sample set to obtain a candidate student model, wherein each training image in the first training sample set does not carry a true annotation;
[0028] A second optimization module, configured to determine a loss function based on the second training sample set, the candidate student model, the initial teacher model, and a third training sample set, and optimize the model parameters of the initial teacher model based on the loss function, wherein each training image in the third training sample set carries a true annotation;
[0029] An execution module, configured to use the candidate teacher model as the initial teacher model and the candidate student model as the initial student model, and continue to execute the step of determining the second training sample set based on the first training sample set and the initial teacher model until the model parameters of the initial student model meet the preset conditions.
[0030] In a third aspect of the embodiments of the present application, a medical image segmentation method is provided. The method uses an image segmentation model trained by the training method of the image segmentation model based on semi-supervised learning as described above. The method includes:
[0031] Input the medical image to be segmented into the image segmentation model; output the target region corresponding to the medical image through the image segmentation model.
[0032] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in any one of the above-mentioned training methods of the image segmentation model based on semi-supervised learning.
[0033] In a fifth aspect of the embodiments of the present application, a terminal device is provided, which includes: a processor, a memory, and a communication bus; a computer-readable program executable by the processor is stored on the memory;
[0034] The communication bus realizes the connection and communication between the processor and the memory;
[0035] When the processor executes the computer-readable program, the steps in any one of the above-mentioned training methods of the image segmentation model based on semi-supervised learning are implemented.
[0036] Beneficial effects: Compared with the prior art, the present application provides a training method and a segmentation method for an image segmentation model based on semi-supervised learning. The method includes determining a second training sample set based on a first training sample set and an initial teacher model, and optimizing the model parameters of an initial student model based on the second training sample set to obtain a candidate student model; determining a loss function based on the second training sample set, the candidate student model, the initial teacher model, and a third training sample set, and optimizing the model parameters of the initial teacher model based on the loss function to obtain a candidate teacher model; using the candidate teacher model as the initial teacher model and the candidate student model as the initial student model, and continuing to execute the step of determining the second training sample set based on the first training sample set and the initial teacher model until the model parameters of the initial student model meet the preset conditions. In the present application, the initial student model optimizes its parameters according to the predicted labels of the initial teacher model on the unlabeled data set, and the initial teacher model optimizes its parameters according to the performance of the optimized initial student model on the labeled data set to generate better predicted labels for the initial student model to learn. In this way, through the alternating optimization of the initial teacher model and the initial student model, the labeled data and the unlabeled data can be reasonably utilized, thereby improving the model performance of the trained segmentation model. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0038] Figure 1 It is a flowchart of the training method for the image segmentation model based on semi-supervised learning provided by the present application.
[0039] Figure 2 It is a schematic flow diagram of the training method for the image segmentation model based on semi-supervised learning provided by the present application.
[0040] Figure 3 It is a schematic diagram of the model structure of the preset network model in the training method for the image segmentation model based on semi-supervised learning provided by the present application.
[0041] Figure 4 It is a schematic diagram of the structure of the training device for the image segmentation model based on semi-supervised learning provided by the present application.
[0042] Figure 5 It is a schematic diagram of the structure of the terminal device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] This application provides a training method and a segmentation method for an image segmentation model based on semi-supervised learning. To make the objectives, technical solutions, and effects of this application clearer and more explicit, the following further elaborates on this application with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0044] Those skilled in the art of this technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of this application means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more of the associated listed items.
[0045] Those skilled in the art of this technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which this application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.
[0046] It should be understood that the sequence numbers and magnitudes of the steps in this embodiment do not imply the order of execution. The order of execution of each process is determined by its function and internal logic and should not constitute any limitation to the implementation process of the embodiments of this application.
[0047] The inventors have found through research that deep learning models are widely used in the field of medical image segmentation and have achieved remarkable performance. However, a common problem with segmentation methods based on deep learning technology is that sufficient segmentation annotations are required for model training. For medical images, the acquisition of annotated data requires professional knowledge and time.
[0048] Accordingly, the semi-supervised deep learning approach is applied to medical image segmentation, and there are two commonly used semi-supervised deep learning methods, namely, the method of generating pseudo-labels based on self-training or co-training, and the method of regularizing the model based on the smoothness assumption. Among them, the method of generating pseudo-labels based on self-training or co-training is to train a segmentation model based on a training sample set of labeled data, then use the trained segmentation model to predict unlabeled data, and select highly credible pseudo-labels to add to the training sample set, and continue to train the segmentation model until the model parameters of the segmentation model meet the cut-off conditions. This method requires selecting highly credible pseudo-labels, but when there are noisy pseudo-labels in the training sample set, the model performance of the segmentation model will deteriorate. The method of regularizing the model based on the smoothness assumption is to obtain multiple input images by performing different data transformations on the input image, and then input the multiple input images into the segmentation model respectively to make the prediction results of the multiple input images tend to be similar. This method means that the same sample is close in the data space after different transformations, and the corresponding prediction results are also similar, but it does not explain that the labeled samples and unlabeled samples are close in the data space, and the features of the unlabeled data cannot be fully extracted, thus the model performance of the segmentation model cannot be guaranteed.
[0049] To solve the above problems, in the embodiments of the present application, a second training sample set is determined based on a first training sample set and an initial teacher model, and the model parameters of the initial student model are optimized based on the second training sample set to obtain a candidate student model; a loss function is determined based on the second training sample set, the candidate student model, the initial teacher model, and a third training sample set, and the model parameters of the initial teacher model are optimized based on the loss function to obtain a candidate teacher model; the candidate teacher model is used as the initial teacher model, and the candidate student model is used as the initial student model, and the step of determining the second training sample set based on the first training sample set and the initial teacher model is continued until the model parameters of the initial student model meet the preset conditions. In the present application, the initial student model optimizes the parameters according to the predicted labels of the initial teacher model in the unlabeled data set, and the initial teacher model optimizes the parameters according to the performance of the optimized initial student model on the labeled data set to generate better predicted labels for the initial student model to learn. In this way, through the alternating optimization of the initial teacher model and the initial student model, the labeled data and unlabeled data can be reasonably utilized, thereby improving the model performance of the trained segmentation model.
[0050] The following further illustrates the application content by describing the embodiments in conjunction with the accompanying drawings.
[0051] This embodiment provides a training method for an image segmentation model based on semi-supervised learning, as Figure 1 and Figure 2 shown. The method includes:
[0052] S10. Determine a second training sample set based on the first training sample set and the initial teacher model, and optimize the model parameters of the initial student model based on the second training sample set to obtain a candidate student model.
[0053] Specifically, both the initial teacher model and the initial student model are deep learning network models, and the model structures of the initial teacher model and the initial student model are the same. The difference between the two is that the model parameters of the initial teacher model are different from those of the initial student model. It can be understood that the initial teacher model and the initial student model are obtained by pre-training the same deep learning network model using different training sample sets respectively. Among them, the initial teacher model can be pre-trained using a training sample set with annotations, and the initial student model can be pre-trained using a training sample set including some training images with annotations and some training images without annotations.
[0054] In an implementation manner of this embodiment, before determining the second training sample set based on the first training sample set and the initial teacher model, the method includes:
[0055] Train a preset network model using a fourth training sample set to obtain a pre-trained teacher model;
[0056] Train the pre-trained teacher model using a fifth training sample set to obtain an initial teacher model, and train the preset network model using the fifth training sample set to obtain an initial student model.
[0057] Specifically, each training image in the fourth training sample set carries a true annotation, and some training images in the fifth training sample set carry true annotations while some do not. It can be understood that for each training image in the fourth training sample set, the training image carries a true annotation, where the true annotation can be obtained through manual annotation. That some training images in the fifth training sample set carry true annotations while some do not means that when the fifth training sample set is divided according to whether it carries true annotations, two sub-sample sets containing training images can be obtained, and the training images in one of the two sub-sample sets all carry true annotations, while the training images in the other sub-sample set do not carry true annotations. For example, the fifth training sample set includes several training batches, and each of the several training batches includes 2 training medical images carrying true annotations and 4 training medical images not carrying true annotations. In a specific implementation manner of this embodiment, the training images in the fourth training sample set and the training images in the fifth training sample set are all doctor images, such as CT images, ultrasound images, and MRI images, etc. In this embodiment, the fifth training sample set carrying true annotations and not carrying true annotations is used to train the student model, so that the student model does not directly learn the labeled data, avoiding overfitting.
[0058] Illustrative example:
[0059] The fourth training sample set includes 3000 training batches, and each training batch includes 4 training medical images carrying true annotations. When training the preset network model based on the fourth training sample set, the stochastic gradient descent algorithm is used to optimize the model parameters of the preset network model. The initial learning rate is 0.01. After every 1000 batches are trained, the learning rate is reduced to 0.1 of the original until the 3000th batch is trained to obtain the pre-trained teacher model.
[0060] The fifth training sample includes 6000 training batches, and each training batch includes 2 training medical images carrying true annotations and 4 training medical images carrying true annotations. When training the pre-trained teacher model and the preset network model based on the fifth training sample set, the initial learning rate is 0.01. After every 2500 batches are trained, the learning rate is reduced to 0.1 of the original until the 6000th batch is trained to obtain the pre-trained teacher model, to obtain the initial teacher model and the initial student model.
[0061] In an implementation manner of this embodiment, when training the preset network model based on the fourth training sample set to obtain the pre-trained teacher model, the cross-entropy loss function can be used as the loss function, where the calculation formula of the loss function can be:
[0062]
[0063] Among them, represents the loss function, x l represents the training image carrying the true annotation, y l represents the true annotation, T(x l ; θ T ) represents the predicted annotation obtained by passing the training image carrying the true annotation through a preset network model, and θ T represents the model parameters for training the preset network model as the initial teacher model.
[0064] In an implementation manner of this embodiment, the preset network model is a neural network for performing a segmentation task. By training the preset network model, an initial teacher model and an initial student model can be obtained. Among them, as Figure 3 shown, the preset network model may include an encoder and a decoder. The encoder includes 4 encoding modules and 4 max-pooling layers. The 4 encoding modules and the 4 max-pooling layers are arranged alternately in sequence, and a max-pooling layer is arranged after each encoding module. Among them, each encoding module includes two convolutional layers, a normalization layer, and an activation function (Relu) layer cascaded in sequence. After passing through the encoding module, the number of channels of the input image becomes twice the original. The input image output by the encoding module is downsampled by a factor of 1 / 2 in size through the max-pooling layer, and the initial number of convolutional channels is 64. The decoder includes 3 decoding modules and 3 upsampling layers. The 3 decoding modules and the 3 upsampling layers are arranged alternately in sequence, and an upsampling layer is arranged before each decoding module. Among them, the decoding module has the same structure as the encoding module. The feature map output by the encoder is magnified by a factor of 2 through the upsampling layer and is concatenated with the same shallow feature map output by the encoding module in the encoder through a skip connection and the magnified feature map. The concatenated feature map is input into the decoding module. The upsampling layer may use the nearest neighbor interpolation method to upsample the feature map.
[0065] In an implementation manner of this embodiment, determining the second training sample set based on the first training sample set and the initial teacher model specifically includes:
[0066] For each first training image in the first training sample set, input the first training image into the initial teacher model, and output the first predicted annotation corresponding to the first training image through the initial teacher model;
[0067] Take each first training image and its corresponding first predicted annotation as a training sample, and use the training sample set composed of all the obtained training samples as the second training sample set.
[0068] Specifically, none of the first training images in the first training sample set carry real annotations. Each training image in the second training sample set is a first training image in the first training sample set, and each training image in the second training sample set carries a predicted annotation, which is obtained by the initial teacher model learning from the corresponding first training image in the first training sample set. That is to say, after obtaining the first training sample set, the initial teacher model performs image segmentation on each first training image in the first training sample set to obtain the first predicted annotation corresponding to each first training image, and then uses the first predicted annotation as the pseudo-annotation corresponding to the first training image to form training samples. Then, the training sample set composed of all the formed training samples is used as the second training sample set. It can be understood that the number of training images included in the second training sample set is the same as the number of first training images included in the first training sample set, and the image content of the first training images included in the first training sample set is the same as that of the training images included in the second training sample set. The difference between the two is that the first training images do not carry annotations, while the training images in the second training samples carry pseudo-annotations learned by the initial teacher model.
[0069] In an implementation manner of this embodiment, after obtaining the second training sample set, the initial student model is trained using the second training sample set. Among them, the initial student model can use the minimized cross-entropy loss to optimize the model parameters. The expression of the optimized model parameters can be:
[0070]
[0071] Among them, represents the model parameters of the candidate student model, x u represents the second training image in the second training sample set, represents the second predicted annotation corresponding to the second training image, S(x u ; θ S ) represents the predicted annotation obtained by predicting through the initial student model, and θ S represents the model parameters of the initial student model before optimization.
[0072] S20. Determine the loss function based on the second training sample set, the candidate student model, the initial teacher model, and the third training sample set, and optimize the model parameters of the initial teacher model based on the loss function.
[0073] Specifically, each training image in the third training sample set carries a true annotation. The third training sample set may be the same as or different from the fourth training sample set. For example, a data set of training images carrying true annotations is preset, and several training images carrying true annotations are selected from this data set to form the third training sample set, and several training images carrying true annotations are selected from this data set to form the fourth training sample set. Among them, the training images in the third training sample set are different from those in the fourth training sample set, or some of the training images in the third training sample set are different from those in the fourth training sample set.
[0074] In an implementation manner of this embodiment, determining the loss function based on the second training sample set, the candidate student model, the initial teacher model, and the third training sample set specifically includes:
[0075] Determining a first loss term based on the third training sample set and the candidate student model;
[0076] Determining a second loss term based on the third training sample set and the initial teacher model;
[0077] Determining a third loss term based on the second training sample set and the initial teacher model;
[0078] Determining the loss function based on the first loss term, the second loss term, and the third loss term.
[0079] Specifically, the first loss term is obtained from the third training sample set and the candidate student model. It can be understood that in the selected third training sample set, the third training images in the third training sample set are input into the candidate student model, and the third predicted annotation corresponding to the third predicted image is output through the initial student model, and the first loss term is determined based on the true annotation corresponding to the third training image and the third predicted annotation.
[0080] The segmentation performance of the candidate student model in the annotated data will be better than that of the initial student model before optimization to achieve a lower cross-entropy loss, that is, to minimize. In addition, the model parameters of the candidate student model are related to the second predicted annotation and the model parameters θ of the initial teacher model T T are related. Thus, the model parameters of the candidate student model can be expressed as so as to optimize the initial teacher model parameters θ by minimizing the feedback loss : T :
[0081]
[0082] Among them, y l represents the true annotation of the third training image in the third training sample set, and x l represents the third training image in the third training sample set.
[0083] In addition, in one interactive learning of the initial student model and the initial teacher model (including one update of the initial student model and one update of the initial teacher model), since optimizing θ S until the initial student model converges completely is very inefficient, and calculating the gradient requires expanding the entire training process of the initial student model, which will affect the operation efficiency and thus the model training speed. Therefore, in a typical implementation manner of this embodiment, borrowing the meta-learning method, one-step gradient update of θ S is approximately equal to That is: Among them, η S is the learning rate for training the initial student model, which can improve the operation efficiency of the gradient It can be understood that when optimizing the initial student model based on the second training sample set, the model parameters of the initial student model can be optimized using stochastic gradient descent. Among them, the optimized model parameters can be expressed as:
[0084]
[0085] Among them, θ' S represents the model parameters of the candidate student model, x u represents the second training image in the second training sample set, represents the second predicted annotation corresponding to the second training image, S(x u ; θ S ) represents the predicted annotation obtained by predicting through the initial student model, and θ S represents the model parameters of the initial student model before optimization, and η S represents the learning rate for training the initial student model.
[0086] Based on this, the first loss term can be determined using the stochastic gradient method, that is, the gradient can be used as the first loss term. Among them, the determination process of the first loss term can include:
[0087] First: According to the chain rule can be transformed into:
[0088]
[0089] The model parameters θ' of the initial student model after optimization S According to the probability distribution T(xu ; θ T ) Sampling is obtained from the expectation, that is
[0090]
[0091] Let:
[0092]
[0093] Since is only related to the predicted annotation and the model parameter θ of the initial teacher model T are related, so the gradient is calculated using the REINFORCE equation:
[0094]
[0095] Finally, using Monte Carlo estimation, sample from T(x u ; θ T ) to obtain approximate with empirical average So the first loss term is:
[0096]
[0097] In an implementation manner of this embodiment, in order to fully optimize the initial teacher model, in addition to the first loss term determined based on the initial student model, the training images carrying the true annotations in the third training sample set are also used to perform full-supervised training on the initial teacher model. It can be understood that the second loss term is obtained by performing full-supervised training on the initial teacher model using the third training sample set, where the expression of the second loss term can be:
[0098]
[0099] Among them, represents the second loss term, x l represents the training image, y l represents the true annotation corresponding to the training image, T((x l )); θ T ) represents the predicted annotation output by the initial teacher model.
[0100] In an implementation manner of this embodiment, determining the third loss term based on the second training sample set and the initial teacher model specifically includes:
[0101] Perform data transformation on each second training image and its corresponding second predicted annotation in the second training sample set respectively;
[0102] Input the second training images after various data transformations into the initial teacher model, and determine the predicted annotations corresponding to each of the second training images after various data transformations through the initial teacher model;
[0103] Based on each of the second predicted annotations after transformation and each predicted annotation, determine the third loss term.
[0104] Specifically, the data transformation is a random data transformation, and the random data transformations corresponding to the second training image and its corresponding second predicted annotation are the same. The random data transformation may include random flipping and random rotation, etc. It can be understood that the third loss term is a regularization direction determined by adopting data conversion consistency. By regularizing the initial teacher model with the regularization direction, the training images without true annotations can be more fully utilized, thereby improving the generalization ability of the segmentation network. In a specific implementation manner, the third loss term can adopt a minimized cross-entropy loss function, where the third loss term can be expressed as:
[0105]
[0106] where π i represents the random data transformation, including random flipping and random rotation of the input image. The rotation angle of the random rotation can be γ·90°, γ ∈ {0, 1, 2, 3}.
[0107] Finally, after obtaining the first loss term, the second loss term, and the third loss term, the SGD can be used to optimize the initial teacher model. The optimization expression of the model parameters of the initial teacher model can be:
[0108]
[0109] where is a time-varying Gaussian increasing function used to balance the regularization loss function and other loss functions. t represents the current training iteration number, and t max represents the total iteration number.
[0110] S30. Use the candidate teacher model as the initial teacher model and the candidate student model as the initial student model, and continue to execute the step of determining the second training sample set based on the first training sample set and the initial teacher model until the model parameters of the initial student model meet the preset conditions to obtain an image segmentation model.
[0111] Specifically, after completing an initial optimization of the teacher model and an initial optimization of the student model, the steps of determining the second training sample set based on the first training sample set and the initial teacher model and optimizing the model parameters of the initial teacher model are repeatedly executed for iterative training until the model parameters of the initial student model meet the preset conditions, and the initial student model is used as the trained medical image segmentation model, where the preset conditions may include that the loss term determined based on the initial student model is less than the preset threshold, or the number of training times reaches the preset number of times threshold. In addition, after the model parameters of the initial student model meet the preset conditions, the initial student model that meets the preset conditions is used as the image segmentation model.
[0112] In an implementation manner of this embodiment, when training the segmentation model by using the initial teacher model and the initial student model for interactive training, the image size of the training image can be scaled according to the video memory size of the electronic device used to execute the training method of the image segmentation model based on semi-supervised learning, which can avoid the limitation of the video memory of the electronic device on the training process. Among them, when scaling the image size of the training image, the SimpleITK resampling method can be used for image size scaling. In addition, after the image size is scaled, the data distribution of the image can be standardized so that its distribution mean is 0 and the variance is 1. Of course, in practical applications, in order to reduce the calculation cost, a fixed-size patch can be randomly selected from each training image as the training image during training. In addition, the true annotation in this embodiment can be the true segmentation region, and the predicted annotation can be the predicted segmentation region.
[0113] In summary, this embodiment provides a training method for an image segmentation model based on semi-supervised learning. The method includes determining a second training sample set based on a first training sample set and an initial teacher model, and optimizing the model parameters of the initial student model based on the second training sample set to obtain a candidate student model; determining a loss function based on the second training sample set, the candidate student model, the initial teacher model, and a third training sample set, and optimizing the model parameters of the initial teacher model based on the loss function; using the candidate teacher model as the initial teacher model and the candidate student model as the initial student model, and continuing to execute the steps of determining the second training sample set based on the first training sample set and the initial teacher model and optimizing the model parameters of the initial teacher model until the preset conditions are met to obtain an image segmentation model. In this application, the initial student model optimizes the parameters according to the predicted annotation of the initial teacher model in the unlabeled data set, and the initial teacher model optimizes the parameters according to the performance of the optimized initial student model on the labeled data set to generate better predicted annotations for the initial student model to learn. In this way, through the alternating optimization of the initial teacher model and the initial student model, the labeled data and the unlabeled data can be reasonably utilized, thereby improving the model performance of the trained segmentation model.
[0114] Based on the above training method of the semi-supervised learning-based image segmentation model, this embodiment provides a training device for the semi-supervised learning-based image segmentation model, as Figure 4 shown, the training device includes:
[0115] A first optimization module 100, configured to determine a second training sample set based on a first training sample set and an initial teacher model, and optimize model parameters of the initial student model based on the second training sample set to obtain a candidate student model, wherein each training image in the first training sample set does not carry a true annotation;
[0116] A second optimization module 200, configured to determine a loss function based on the second training sample set, the candidate student model, the initial teacher model, and a third training sample set, and optimize model parameters of the initial teacher model based on the loss function, wherein each training image in the third training sample set carries a true annotation;
[0117] An execution module 300, configured to use the candidate teacher model as the initial teacher model and the candidate student model as the initial student model, and continue to execute the step of determining the second training sample set based on the first training sample set and the initial teacher model until the model parameters of the initial student model meet a preset condition, so as to obtain an image segmentation model.
[0118] Based on the above training method of the semi-supervised learning-based image segmentation model, this embodiment provides a medical image segmentation method, and the method applies the image segmentation model obtained by the training method of the semi-supervised learning-based image segmentation model described in the above embodiment; the method includes:
[0119] Input the medical image to be segmented into the image segmentation model; output the target region corresponding to the medical image through the image segmentation model.
[0120] Specifically, after obtaining the medical image to be segmented, the image size can be appropriately scaled by resampling using SimpleITK and then gray-scale normalization processing can be performed, and the medical image after gray-scale normalization processing is input into the image segmentation model, so as to output the target region corresponding to the medical image through the learning model, thereby realizing the segmentation of the medical image. Of course, in practical applications, when there is a video memory limit in the electronic device running the image segmentation model, the medical image after gray-scale normalization processing can be windowed to obtain a number of patches, and then the prediction target region of each patch is determined through the image segmentation model, and then the prediction target regions of the number of patches are reconstructed to obtain the target region of the medical image.
[0121] Based on the above training method of the semi-supervised learning-based image segmentation model, this embodiment provides a computer-readable storage medium storing one or more programs, which can be executed by one or more processors to implement the steps in the training method of the semi-supervised learning-based image segmentation model as described in the above embodiment.
[0122] Based on the above training method of the semi-supervised learning-based image segmentation model, the present application also provides a terminal device, as Figure 5 shown, which includes at least one processor 20; a display screen 21; and a memory 22, and may further include a communication interface 23 and a bus 24. Among them, the processor 20, the display screen 21, the memory 22, and the communication interface 23 can complete mutual communication through the bus 24. The display screen 21 is set to display a user guidance interface preset in the initial setting mode. The communication interface 23 can transmit information. The processor 20 can call the logical instructions in the memory 22 to execute the method in the above embodiment.
[0123] In addition, when the logical instructions in the above-mentioned memory 22 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium.
[0124] The memory 22, as a computer-readable storage medium, can be set to store software programs and computer-executable programs, such as the program instructions or modules corresponding to the method in the embodiment of the present disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions, or modules stored in the memory 22, that is, to implement the method in the above embodiment.
[0125] The memory 22 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal device, etc. In addition, the memory 22 may include a high-speed random access memory and may also include a non-volatile memory. For example, various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes can also be transient storage media.
[0126] In addition, the specific processes of loading and executing multiple instructions by the instruction processor in the above storage medium and the terminal device have been described in detail in the above method and will not be repeated here.
[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A training method for an image segmentation model based on semi-supervised learning, characterized in that, The described training method includes: Determining a second training sample set based on a first training sample set and an initial teacher model, and optimizing the model parameters of the initial student model based on the second training sample set to obtain a candidate student model, where each training image in the first training sample set does not carry a true annotation; Determining a loss function based on the second training sample set, the candidate student model, the initial teacher model, and a third training sample set, and optimizing the model parameters of the initial teacher model based on the loss function to obtain a candidate teacher model, where each training image in the third training sample set carries a true annotation; Taking the candidate teacher model as the initial teacher model and the candidate student model as the initial student model, and continuing to execute the step of determining the second training sample set based on the first training sample set and the initial teacher model until the model parameters of the initial student model meet the preset conditions to obtain an image segmentation model; Among them, the determining the loss function based on the second training sample set, the candidate student model, the initial teacher model, and the third training sample set specifically includes: Determining a first loss term based on the third training sample set and the candidate student model; Determining a second loss term based on the third training sample set and the initial teacher model; Performing data transformation on each second training image in the second training sample set and its corresponding second predicted annotation; Inputting each data-transformed second training image into the initial teacher model, and determining the predicted annotation corresponding to each data-transformed second training image through the initial teacher model; Determining a third loss term based on each transformed second predicted annotation and each predicted annotation; Determining the loss function based on the first loss term, the second loss term, and the third loss term.
2. The training method of the image segmentation model based on semi-supervised learning according to claim 1, characterized in that Before determining the second training sample set based on the first training sample set and the initial teacher model, the method includes: Training a preset network model with a fourth training sample set to obtain a pre-trained teacher model, where each training image in the fourth training sample set carries a true annotation; Training the pre-trained teacher model with a fifth training sample set to obtain an initial teacher model, and training the preset network model with the fifth training sample set to obtain an initial student model, where some training images in the fifth training sample set carry true annotations and some training images do not carry true annotations.
3. The training method of the image segmentation model based on semi-supervised learning according to claim 1, characterized in that, The determining the second training sample set based on the first training sample set and the initial teacher model specifically includes: For each first training image in the first training sample set, inputting the first training image into the initial teacher model, and outputting the first predicted annotation corresponding to the first training image through the initial teacher model; Taking each first training image and its corresponding first predicted annotation as a training sample, and taking the training sample set composed of all the obtained training samples as the second training sample set.
4. The training method of the image segmentation model based on semi-supervised learning according to claim 1, wherein, The data transformation is random data transformation, and the random data transformations corresponding to the second training image and its corresponding second predicted annotation are the same.
5. A training device for an image segmentation model based on semi-supervised learning, characterized in that, The training device includes: The first optimization module is used to determine a second training sample set based on the first training sample set and the initial teacher model, and optimize the model parameters of the initial student model based on the second training sample set to obtain a candidate student model, where each training image in the first training sample set does not carry a true annotation; The second optimization module is used to determine a loss function based on the second training sample set, the candidate student model, the initial teacher model, and the third training sample set, and optimize the model parameters of the initial teacher model based on the loss function to obtain a candidate teacher model, where each training image in the third training sample set carries a true annotation; The execution module is used to use the candidate teacher model as the initial teacher model and the candidate student model as the initial student model, and continue to execute the step of determining the second training sample set based on the first training sample set and the initial teacher model until the model parameters of the initial student model meet the preset conditions to obtain an image segmentation model; Among them, the determining the loss function based on the second training sample set, the candidate student model, the initial teacher model, and the third training sample set specifically includes: Determining a first loss term based on the third training sample set and the candidate student model; Determining a second loss term based on the third training sample set and the initial teacher model; Performing data transformation on each second training image in the second training sample set and its corresponding second predicted annotation respectively; Inputting each data-transformed second training image into the initial teacher model, and determining the predicted annotation corresponding to each data-transformed second training image through the initial teacher model; Determining a third loss term based on each transformed second predicted annotation and each predicted annotation; Determining the loss function based on the first loss term, the second loss term, and the third loss term.
6. A medical image segmentation method, characterized in that, The method is applied to an image segmentation model trained by the training method of the image segmentation model based on semi-supervised learning according to any one of claims 1-4; the method includes: Inputting a medical image to be segmented into the image segmentation model; outputting a target region corresponding to the medical image through the image segmentation model.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the training method of the image segmentation model based on semi-supervised learning according to any one of claims 1-4.