Semi-Supervised Contrastive Learning Medical Image Segmentation Method Based on Similarity Measurement

By adopting a semi-supervised contrast learning method based on similarity metrics in medical image segmentation, screening the similarity between voxel data of medical image and building a network of teachers and students, the problem of false negative caused by the failure of existing methods to effectively utilize similarity is solved, and the segmentation performance of the model is improved.

CN116309403BActive Publication Date: 2025-07-01XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310196976.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-03
Publication Date
2025-07-01
Estimated Expiration
2043-03-03

AI Technical Summary

Technical Problem

Existing medical image segmentation methods fail to effectively utilize the similarity between medical image voxel data, resulting in false negatives in negative pairs, affecting the segmentation performance of the model.

Method used

A semi-supervised contrast learning method based on similarity metrics is adopted, and through virtual blocking and similarity metric mechanisms, incomplete positive sets, negative sets and complete positive sets are screened to build a network of teachers and students, and model training is carried out using comparison losses, unsupervised consistency losses and full supervision losses.

Benefits of technology

Reliance on label data is reduced, segmentation performance of the model is improved, false negative errors are reduced, and similarity characteristics between tissues in medical images are fully utilized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116309403B_ABST
    Figure CN116309403B_ABST
Patent Text Reader

Abstract

The present invention discloses a semi-supervised contrastive learning medical image segmentation method based on similarity measurement. The specific process is as follows: dividing the medical image dataset to be segmented in voxel format into a training set, a validation set and a test set, and preprocessing the dataset; constructing a contrastive learning semi-supervised segmentation model; using the training set and the validation set to iteratively train the constructed contrastive learning semi-supervised segmentation model to obtain a trained contrastive learning semi-supervised segmentation model; taking the test set as the input of the trained contrastive learning semi-supervised segmentation model for forward inference to obtain the segmentation score of each test sample. The method of the present invention solves the problem that the existing segmentation methods do not utilize the similarity between voxel data in medical images, resulting in false negatives in the selected negative pairs and affecting the segmentation performance of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image segmentation, and in particular relates to a medical image segmentation method based on semi-supervised contrast learning based on similarity measurement. Background Art

[0002] In the field of medical image segmentation, due to the complexity and variability of medical images, traditional segmentation algorithms have great deficiencies in the accuracy and stability of segmentation results. In response to these problems, scholars have proposed various improved algorithms, such as segmentation algorithms based on deep learning and segmentation algorithms based on traditional machine learning.

[0003] However, deep learning-based medical image segmentation algorithms usually require a large amount of manually annotated data to train the segmentation model, and these annotated data often need to be manually annotated by professional doctors, which is time-consuming, labor-intensive and costly. Therefore, a semi-supervised learning method using a small amount of labeled data and a large amount of unlabeled data is proposed to solve the above problems to a certain extent. Compared with the fully supervised algorithm, the performance of the semi-supervised algorithm has a higher room for improvement. Contrastive learning, as a self-supervised learning method, improves the feature extraction performance of the model by shortening the distance between positive pairs and increasing the distance between negative pairs. However, unlike natural images, medical images are mostly presented in the form of voxels, and different slice images have a high similarity. However, the existing medical image segmentation methods based on contrastive learning lack the similarity between slices of medical image voxel data, which will lead to the problem of false negatives in the selected negative pairs. Summary of the invention

[0004] The purpose of the present invention is to provide a medical image segmentation method based on semi-supervised contrast learning of similarity measurement, which solves the problem that the existing segmentation method does not utilize the similarity between voxel data in medical images, resulting in false negatives in the selected negative pairs, affecting the segmentation performance of the model.

[0005] The technical solution adopted by the present invention is a medical image segmentation method based on semi-supervised contrast learning of similarity measurement, which is specifically implemented in the following steps:

[0006] Step 1, dividing the medical image dataset to be segmented in voxel format into a training set, a validation set and a test set, and preprocessing the dataset;

[0007] Step 2, construct a contrastive learning semi-supervised segmentation model;

[0008] Step 3, iteratively training the contrastive learning semi-supervised segmentation model constructed in step 2 using the training set and the validation set to obtain a trained contrastive learning semi-supervised segmentation model;

[0009] Step 4: Use the test set as the input of the trained contrastive learning semi-supervised segmentation model for forward inference to obtain the segmentation score of each test sample.

[0010] The features of the present invention also lie in that

[0011] In Step 1, the specific process of preprocessing the data set is as follows: perform virtual chunking on the data set to obtain virtual image chunks, calculate the similarity between the virtual image chunks, and screen the incomplete positive pair set, negative pair set and complete positive pair of each virtual image chunk according to the similarity. Store the incomplete positive pair set and negative pair set corresponding to each virtual image chunk in the incomplete positive pair dictionary and incomplete negative pair dictionary respectively in the form of coordinate mapping.

[0012] Use histogram or Bhattacharyya coefficient to calculate the similarity between virtual image chunks.

[0013] The specific process of dividing the incomplete positive pair set, negative pair set and complete positive pair according to the similarity is as follows: when the similarity between two virtual image chunks is greater than 0.7 and less than 1.0, then these two virtual image chunks belong to the incomplete positive pair set; when the similarity between two virtual image chunks is less than 0.5, then these two virtual image chunks belong to the negative pair set.

[0014] In Step 2, the contrastive learning semi-supervised segmentation model consists of a teacher network and a student network. Both the teacher network and the student network use the encoder and decoder in U-Net to extract features, and both the teacher network and the student network add a feature projection head at the end of the encoder to extract features. The feature projection head is used for contrastive learning.

[0015] The loss function adopted by the contrastive learning semi-supervised segmentation model includes contrastive loss, unsupervised consistency loss, and fully supervised loss;

[0016] Then the contrastive loss is:

[0017]

[0018]

[0019]

[0020] L contr (q,k) = L contr1 (q,k) + w × L contr2 (q,k)

[0021] In the formula, L contr1 represents the complete contrastive loss, L contr2 represents the incomplete contrastive loss, q represents the currently selected virtual feature chunk, k +Denote the output virtual feature block obtained by the feature projection head of the teacher network for the virtual image block corresponding to q. NF represents the negative pair feature set, which is the set of virtual image blocks corresponding in the incomplete negative pair dictionary input into the teacher network, and the set of virtual feature sub-blocks output from the feature projection head. PF represents the incomplete positive pair feature set, which is the set of virtual image blocks corresponding in the incomplete positive pair dictionary input into the teacher network, and the set of virtual feature sub-blocks output from the feature projection head. k_ represents a single virtual feature sub-block in NF. Denote a single virtual feature sub-block in PF. sim(·,·) represents the cosine similarity. τ represents the temperature hyperparameter, which is a scalar. w is the weight coefficient.

[0022] For the data with labels, use the cross-entropy loss L ce and the Dice loss L dice as the fully supervised loss L sup , and the expression is:

[0023]

[0024]

[0025]

[0026] In the formula, P is the probability map obtained by the model output and processed by the softmax function, and p i is the probability value at a single position in the probability map. H represents the length of the input voxel image, W represents the width of the voxel image, D represents the depth of the voxel image, and Y represents the original label data.

[0027] For the unlabeled data, use the consistency loss L cons and the entropy minimization loss L ent as the unsupervised loss, and take their average as the final unsupervised loss L unsup , and the expression is:

[0028] L cons (P,P')=(P - P') 2

[0029]

[0030]

[0031] Then the final overall loss function of the model is:

[0032] L = L sup (Y,P)+λ(t)×(L unsup (P,P')+L contr (q,k))

[0033]

[0034] In the formula, L represents the final loss function of the contrastive learning semi-supervised segmentation model, P' represents the predicted probability map output by the student network, t represents the current training iteration number, and t max represents the total number of training iterations, and λ(t) controls the balance between the fully supervised loss, the unsupervised consistency loss, and the contrastive loss.

[0035] In step 3, during the training process, the loss function L = L contr +L sup +L unsup is used for backpropagation to update the hyperparameters of the model.

[0036] The beneficial effects of the present invention are as follows: The method of the present invention adopts a semi-supervised learning method, reduces the dependence on labeled data, introduces teacher and student networks, and fully learns knowledge from a large amount of unlabeled data. While reducing the dependence on labeled data, it improves the performance upper limit of the contrastive learning semi-supervised segmentation model. At the same time, a new contrastive learning strategy is designed, introducing virtual chunking and similarity measurement mechanisms to screen the incomplete positive pair set, negative pair set, and complete positive pairs corresponding to each virtual image chunk. This strategy makes full use of the similarity characteristics between tissues in medical images, enabling the model to fully utilize the feature differences of the learning data during the training process and improving the segmentation performance of the model for the input medical images. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is a flowchart of the method for semi-supervised contrastive learning medical image segmentation based on similarity measurement of the present invention;

[0038] Figure 2 is the coordinate of the medical image to be segmented in the method of the present invention;

[0039] Figure 3 is a schematic diagram of the virtual image chunk in the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The present invention will be described in detail below with reference to the drawings and specific embodiments.

[0041] The method for semi-supervised contrastive learning medical image segmentation based on similarity measurement of the present invention, as Figure 1 shown, is specifically implemented according to the following steps:

[0042] In step 1, the medical image dataset to be segmented in voxel format is divided into a training set, a validation set, and a test set. Then, the dataset is virtually partitioned to obtain virtual image patches of equal size, and the similarity between the virtual image patches is calculated. According to the similarity, the incomplete positive pair set, negative pair set, and complete positive pair of each virtual image patch are screened, and the incomplete positive pair set and negative pair set corresponding to each virtual image patch are stored in the incomplete positive pair dictionary (PositiveDictionary, PD) and the incomplete negative pair dictionary (Negative Dictionary, ND) in the form of coordinate mapping, respectively;

[0043] The training set contains a small amount of labeled data and a large amount of unlabeled data;

[0044] As Figure 3 shown, virtual partitioning means dividing the medical image to be segmented into blocks of the same size in the form of coordinates without cropping to obtain virtual image patches, that is, obtaining the position information of the virtual image patches through the upper left corner coordinates and the lower right corner coordinates; for example, each medical image to be segmented as Figure 2 shown is partitioned into blocks of size 4*4, as Figure 3 shown, that is, each medical image to be segmented is divided into 16 virtual image patches. Suppose the incomplete positive pair of virtual image patch No. 6 is No. 10, and the negative pairs are No. 1, No. 2, and No. 16. The coordinate information corresponding to No. 6 is [64, 128], [128, 192], then the coordinate information corresponding to No. 10 is [64, 64], [128, 128], which is stored in the incomplete positive pair dictionary PD;

[0045] The similarity between virtual image patches is calculated using histograms or Bhattacharyya coefficients;

[0046] The specific process of dividing the incomplete positive pair set, negative pair set, and complete positive pair according to the similarity is as follows:

[0047] When the similarity between two virtual image patches is greater than 0.7 and less than 1.0, then these two virtual image patches belong to the incomplete positive pair set. When the similarity between two virtual image patches is less than 0.5, then these two virtual image patches belong to the negative pair set. When the similarity is equal to 1.0, then the virtual image patch is a complete positive pair, that is, the current virtual image patch itself;

[0048] Step 2, construct a contrastive learning semi-supervised segmentation model;

[0049] The contrastive learning semi-supervised segmentation model consists of a teacher network and a student network. Both the teacher network and the student network use the encoder and decoder in U-Net to extract features, and both the teacher network and the student network add a feature projection head at the end of the encoder to extract features, and this feature projection head is used for contrastive learning;

[0050] The loss functions adopted by the contrastive learning semi-supervised segmentation model include contrastive loss, unsupervised consistency loss, and fully supervised loss;

[0051] During the training process, the virtual image patches obtained in Step 1 are input into the teacher network and the student network, and then the virtual feature patches are output at the feature projection head. Then, the virtual image patch set corresponding to the incomplete positive pair dictionary is input into the teacher network, and the virtual feature patches output from the feature projection head belong to the incomplete positive pair feature set. The virtual image patch set corresponding to the incomplete negative pair dictionary is input into the teacher network, and the virtual feature patches output from the feature projection head belong to the negative pair feature set;

[0052] Then the contrastive loss is:

[0053]

[0054]

[0055]

[0056] L contr (q, k) = L contr1 (q, k) + w × L contr2 (q, k)

[0057] In the formula, L contr1 represents the complete contrastive loss, L contr2 represents the incomplete contrastive loss, q represents the currently selected virtual feature patch, k + represents the output virtual feature patch obtained by passing the virtual image patch corresponding to q through the feature projection head of the teacher network. NF represents the negative pair feature set, which is the set of virtual feature patches output from the feature projection head when the virtual image patches corresponding to the incomplete negative pair dictionary are input into the teacher network. PF represents the incomplete positive pair feature set, which is the set of virtual feature patches output from the feature projection head when the virtual image patches corresponding to the incomplete positive pair dictionary are input into the teacher network. k_ represents a single virtual feature patch in NF, represents a single virtual feature patch in PF, sim(·, ·) represents the cosine similarity, τ represents the temperature hyperparameter, which is a scalar, and w is the weight coefficient, which is a constant set manually;

[0058] For the data with labels in the training set, the cross-entropy loss L ce and the Dice loss L dice are used as the fully supervised loss L sup , and the expression is:

[0059]

[0060]

[0061]

[0062] Wherein, P is the probability map output by the model and processed by the softmax function, and p i is the probability value at a single position in the probability map, H represents the length of the input voxel image, W represents the width of the voxel image, D represents the depth of the voxel image, H×W×D is defaulted to 256*256*24, and Y represents the original label data;

[0063] For the unlabeled data in the training set, the consistency loss L cons and the entropy minimization loss L ent are used as the unsupervised loss, and their average value is taken as the final unsupervised loss L unsup , and the expression is:

[0064] L cons (P, P') = (P - P') 2

[0065]

[0066]

[0067] Then the overall loss function of the final model is:

[0068] L = L sup (Y, P) + λ(t) × (L unsup (P, P') + L contr (q, k))

[0069]

[0070] Wherein, L represents the final loss function of the contrastive learning semi-supervised segmentation model, P' represents the predicted probability map output by the student network, t represents the current training iteration number, and t max represents the total number of training iterations, and λ(t) controls the balance between the fully supervised loss, the unsupervised consistency loss, and the contrastive loss. As the number of training iterations increases, this value gradually increases;

[0071] Step 3, use the training set and the validation set to iteratively train the contrastive learning semi-supervised segmentation model constructed in Step 2. During the training process, the loss function L = L contr + L sup + L unsup is used for backpropagation to update the hyperparameters of the model, and a trained contrastive learning semi-supervised segmentation model is obtained;

[0072] Step 4, use the test set as the input of the trained contrastive learning semi-supervised segmentation model for forward inference to obtain the segmentation score of each test sample.

[0073] Example

[0074] For the segmentation of intracerebral hemorrhage images, the method of the present invention divides the labeled data and unlabeled data in the training set according to the ratio of 1:32 or 1:16 or 1:8 or 1:4 for experimental validation set comparison.

[0075] Table 1 Comparison of the segmentation results of the method of the present invention divided according to the ratio of 1:32 and the prior art

[0076]

[0077]

[0078] As can be seen from Table 1, the semi-supervised method of the present invention can learn additional knowledge from unlabeled data to improve the segmentation performance of the model, and at the same time, the segmentation result of this method is the best.

[0079] Table 2 Comparison of the segmentation results of the method of the present invention carried out according to different division ratios

[0080] Label ratio Dice↑ HD95↓ ASD↓ 1 / 32 0.832 9.214 2.137 1 / 16 0.861 4.845 1.513 1 / 8 0.898 3.137 1.502 1 / 4 0.903 3.128 0.762

[0081] As can be seen from Table 2, with the increase in the number of labels, the segmentation performance of the model can be further improved.

Claims

1. A semi-supervised contrastive learning medical image segmentation method based on similarity measurement, characterized in that, The implementation is specifically carried out according to the following steps: Step 1: Divide the medical image dataset to be segmented in voxel format into a training set, a validation set, and a test set, and preprocess the dataset; In Step 1, the specific process of preprocessing the dataset is as follows: Virtually divide the dataset to obtain virtual image patches, calculate the similarity between the virtual image patches, and screen the incomplete positive pair set, negative pair set, and complete positive pair of each virtual image patch according to the similarity. Store the incomplete positive pair set and negative pair set corresponding to each virtual image patch in an incomplete positive pair dictionary and an incomplete negative pair dictionary respectively in the form of coordinate mapping; Step 2: Construct a contrastive learning semi-supervised segmentation model; In Step 2, the contrastive learning semi-supervised segmentation model consists of a teacher network and a student network. Both the teacher network and the student network use the encoder and decoder in U-Net to extract features, and both the teacher network and the student network add a feature projection head at the end of the encoder to extract features. The feature projection head is used for contrastive learning; Step 3: Iteratively train the contrastive learning semi-supervised segmentation model constructed in Step 2 using the training set and the validation set to obtain a trained contrastive learning semi-supervised segmentation model; Step 4: Use the test set as the input of the trained contrastive learning semi-supervised segmentation model for forward inference to obtain the segmentation score of each test sample.

2. The semi-supervised contrastive learning medical image segmentation method based on similarity measurement according to claim 1, wherein The similarity between virtual image patches is calculated using a histogram or the Bhattacharyya coefficient.

3. The semi-supervised contrastive learning medical image segmentation method based on similarity measurement according to claim 1, wherein The specific process of dividing the incomplete positive pair set, negative pair set, and complete positive pair according to the similarity is as follows: When the similarity between two virtual image patches is greater than 0.7 and less than 1.0, then these two virtual image patches are assigned to the incomplete positive pair set. When the similarity between two virtual image patches is less than 0.5, then these two virtual image patches are assigned to the negative pair set.

4. The semi-supervised contrastive learning medical image segmentation method based on similarity measurement according to claim 3, wherein The loss function adopted by the shown contrastive learning semi-supervised segmentation model includes a contrastive loss, an unsupervised consistency loss, and a fully supervised loss; Then the contrastive loss is: In the formula, represents the complete contrast loss, represents the incomplete contrast loss, represents the currently selected virtual feature block, represents the output virtual feature block obtained by projecting the features of the corresponding virtual image block through the feature projection head of the teacher network. NF represents the negative pair feature set, which is the set of virtual feature sub-blocks output from the feature projection head when the corresponding virtual image blocks in the incomplete negative pair dictionary are input into the teacher network. PF represents the incomplete positive pair feature set, which is the set of virtual feature sub-blocks output from the feature projection head when the corresponding virtual image block set in the incomplete positive pair dictionary is input into the teacher network. represents a single virtual feature sub-block in NF, represents a single virtual feature sub-block in PF, represents the cosine similarity, represents the temperature hyperparameter, which is a scalar, w is the weight coefficient; For data containing labels, use cross-entropy loss and Dice loss as the fully supervised loss , and the expression is: where P is the probability map output by the model and processed by the softmax function, is the probability value at a single position in the probability map, H represents the length of the input voxel image, W represents the width of the voxel image, D represents the depth of the voxel image, and Y represents the original label data; For unlabeled data, use consistency loss and entropy minimization loss as unsupervised losses, and take their average as the final unsupervised loss , and the expression is: Then the final overall loss function of the model is: In the formula, L represents the final loss function of the contrastive learning semi-supervised segmentation model, P’ represents the predicted probability map output by the student network, t represents the current training iteration number, and tmax represents the total number of training iterations, controls the balance among the fully supervised loss, the unsupervised consistency loss, and the contrastive loss.

5. The semi-supervised contrastive learning medical image segmentation method based on similarity measurement according to claim 1, characterized in that In step 3, during the training process, the loss function is used for backpropagation to update the hyperparameters of the model.

Citation Information

Patent Citations

  • Medical image comparison method and device, electronic equipment and storage medium

    CN113935957A

  • Medical image segmentation method of semi-supervised convolutional neural network based on comparative learning

    CN114266739A