Dual-view interactive learning semi-supervised training method based on information bottleneck and comparative learning

By adopting a semi-supervised training method based on information bottlenecks and contrast learning in the auxiliary diagnosis of digestive tract diseases, the problem of parallel network coupling, fuzzy boundaries and micro lesions in the prior art is solved, and higher segmentation accuracy and model performance are achieved.

CN120182248AActive Publication Date: 2025-06-20UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510624700.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-20
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

The existing semi-supervised training methods based on parallel networks have problems such as coupling effect, blurred boundaries and difficult to detect micro lesions in the auxiliary diagnosis of digestive tract diseases, resulting in low accuracy of model segmentation.

Method used

The semi-supervised training method of dual-view interactive learning based on information bottlenecks and contrast learning is adopted to avoid parallel network coupling through dual-view interactive learning, the information bottleneck theory strengthens perception ability, and category-centered contrast learning improves intra-class compactness and inter-class separability.

Benefits of technology

The performance and accuracy of semi-supervised segmentation are significantly improved, effectively solving the problem of difficult detection of blurred boundaries and tiny lesions in the images of digestive tract disease, and improving the model's segmentation ability of the digestive tract injury areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182248A_ABST
    Figure CN120182248A_ABST
Patent Text Reader

Abstract

The invention provides a double-view interactive learning semi-supervised training method based on information bottleneck and comparative learning, and the method comprises the steps: dividing a data stream into labeled data and unlabeled data in a double-view interactive learning process, and for the labeled data, carrying out the data stream processing; two segmentation networks in double-view interactive learning are optimized by calculating cross entropy loss and Dice loss of a predicted value and a real label, as for unlabeled data, the two segmentation networks generate pseudo labels for mutual supervision, and an information bottleneck theory is used for strengthening consistency of a feature layer and the labels (or the pseudo labels) in the segmentation networks, so that the accuracy of data processing is improved. In addition, a class center comparison learning method is applied in the training process, the center feature of each class is firstly constructed, then a certain pixel feature is compared with the center features of all the classes, and finally the pixel feature is classified into the class with the center feature most similar to the pixel feature. Through the scheme of the invention, the performance and accuracy of semi-supervised segmentation can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and image processing, and in particular to a dual-view interactive learning semi-supervised training method based on information bottleneck and contrast learning. Background Art

[0002] Digestive tract diseases are a very common type of disease, such as gastric cancer and esophageal cancer. These diseases seriously threaten the physical and mental health of patients. Early detection and treatment can significantly reduce the mortality rate of this disease. At present, the clinical diagnosis of digestive tract diseases mainly relies on physicians through endoscopic imaging. This method is highly dependent on the doctor's experience and visual perception. For inexperienced doctors, it is very easy to misdiagnose and miss the diagnosis. Computer-aided diagnosis methods based on deep learning can effectively alleviate this problem. At present, many deep learning models have been produced for automatic segmentation of digestive tract diseases with high accuracy. However, deep learning methods are heavily dependent on labeled data sets. It is time-consuming and labor-intensive for doctors to manually label complex digestive tract disease data, and the increasing amount of medical data cannot be manually labeled. Semi-supervised training technology can effectively use unlabeled images to improve the accuracy of segmentation models, but the existing semi-supervised methods still have the following problems when applied to auxiliary diagnosis of digestive tract diseases:

[0003] When existing parallel network-based methods perform semi-supervised training on digestive tract disease images, parallel networks are prone to coupling, which limits the performance of the model.

[0004] The edges of some digestive tract disease lesions are very similar to normal areas. This fuzzy boundary can easily lead to segmentation errors in the training model.

[0005] The lesion areas of some digestive tract diseases are very small, making it difficult for the model to capture them effectively, which in turn affects the segmentation accuracy. Summary of the invention

[0006] To solve the above problems, the present invention proposes a dual-view interactive learning semi-supervised training method based on information bottleneck and contrastive learning. By introducing dual-view interactive learning, information bottleneck theory and category center contrastive learning, the performance and accuracy of semi-supervised segmentation can be significantly improved.

[0007] The present invention proposes a dual-view interactive learning semi-supervised training method based on information bottleneck and contrastive learning, comprising the following steps:

[0008] Step S1: collect medical image data, and then preprocess the image data, the preprocessing includes data enhancement and view conversion;

[0009] Step S2: Construct a semi-supervised training model based on dual-view interactive learning, information bottleneck theory, and class center contrast learning. The segmentation model in the dual-view interactive learning consists of Network A and Network B. The U-Net is used as the segmentation network for both Network A and Network B. The collected RGB images are input into Network A, and the corresponding HSL images of the RGB images are input into Network B. In the parallel network composed of Network A and Network B, dual-view interactive learning is used to avoid the coupling effect of the parallel network. The information bottleneck theory and class center contrast learning are used to enhance the perception and discrimination ability of the network;

[0010] Step S3: Train the segmentation model and train the preprocessed dataset through the semi-supervised training model. During the training process, parameter updates are performed through forward propagation and backward propagation, and the training is stopped when the performance of the network reaches the optimal level, so as to obtain the weight parameters with the best segmentation effect.

[0011] Further, the specific steps of Step S1 include: Divide the collected image data into labeled data and unlabeled data, and then use the color space conversion algorithm to convert the collected RGB images into the corresponding HSL images. Denote the labeled RGB images as , the labeled HSL images as , the unlabeled RGB images as , the unlabeled HSL images as , and the total dataset as , where , , and are the image feature values input into the network, is the label value, and represent the height and width of the image respectively.

[0012] Further, the data augmentation methods in Step S1 include horizontal flipping, vertical flipping, and random scaling.

[0013] Further, the specific steps of the dual-view interactive learning in Step S2 include:

[0014] Use two segmentation networks A and B with the same structure to form a parallel network framework. Denote the segmentation network A as , and the segmentation network B as , Train the collected RGB images, Train the corresponding HSL images. During the training process, for the labeled images, the two segmentation networks are optimized by calculating the cross-entropy loss and Dice loss between the predicted values and the ground truth values. For the unlabeled images, the two segmentation networks are mutually supervised and optimized through pseudo-labels, thus avoiding the coupling caused by training a single view in the parallel networks and realizing the complementary advantages of lesion features in different views, and improving the segmentation accuracy.

[0015] Furthermore, use to train the labeled RGB images , and use the combination of cross-entropy loss and Dice loss as the supervision loss for training , expressed as:

[0016] ;

[0017] where is the supervision loss for RGB labeled data, is the number of labeled RGB images, is the cross-entropy loss, is the Dice loss, The calculation method is as follows:

[0018] ;

[0019] The calculation method is as follows:

[0020] ;

[0021] where and respectively represent the label value and predicted probability of the input image. The label value is obtained through manual annotation, and the predicted probability is obtained by training, and respectively represent the height and width of the image, and respectively represent the label value and predicted probability of the th pixel in the image.

[0022] Furthermore, use to train the labeled HSL images. Similarly, use the combination of cross-entropy loss and Dice loss as the supervision loss for training the labeled HSL images, expressed as:

[0023] ;

[0024] where is the supervision loss for HSL labeled data, is the number of labeled HSL images.

[0025] Furthermore, use to train the unlabeled RGB image data to obtain the corresponding predicted probability values , use to train the unlabeled HSL image data to obtain the corresponding predicted probability values , then, through the network generate the pseudo-label supervised network , and the calculation method of the loss function is as follows:

[0026] ;

[0027] Among them, is the number of unlabeled RGB images, is the pseudo-label generated by , similarly, through the network generate the pseudo-label supervised network , and the calculation method of the loss function is as follows:

[0028] ;

[0029] Among them, is the number of unlabeled HSL images, is the pseudo-label generated by .

[0030] Adding the above losses gives the total loss function of the dual-view interactive learning, denoted as:

[0031] ;

[0032] By using the dual-view interactive learning strategy, the two segmentation networks supervise each other, avoid coupling, and thus improve the segmentation ability of the semi-supervised model for the digestive tract injury area.

[0033] Furthermore, the information bottleneck theory in step S2 specifically includes:

[0034] Adopt the Hilbert-Schmidt Independence Criterion (HSIC) to define the information bottleneck theory, and HSIC is:

[0035] ;

[0036] Among them, , , is the identity matrix, is a column vector with all elements being 1, Represents a variable and the number of rows, represents the trace of a matrix, calculated it is necessary to ensure and have the same number of rows, is the kernel matrix, and the calculation method is as follows:

[0037] ;

[0038] where, and are two different variables in is an adjustable parameter, represents the 2-norm of a vector, also called the Euclidean norm, the definition of is similar to

[0039] ;

[0040] where, and are two different variables in

[0041] Based on the above definition of the information bottleneck, the information bottleneck loss of the RGB image labeled in

[0042] ;

[0043] where, is the labeled RGB image, is the corresponding label value, is the matrix of the feature layer (including convolutional layer and deconvolutional layer) during the training is the number of feature layers, the information bottleneck loss of the unlabeled RGB image in

[0044] ;

[0045] where, is the unlabeled RGB image, is the pseudo-label value generated by is the matrix of the feature layer during the training

[0046] Similarly, it is also possible to obtain The information bottleneck loss of the HSL image annotated in is:

[0047] ;

[0048] in, is the annotated HSL image, is the corresponding label value, for train The feature layer matrix in the process.

[0049] The information bottleneck loss of the unlabeled HSL image in is:

[0050] ;

[0051] in, is an unlabeled HSL image, for The pseudo labels generated are for train The feature layer matrix in the process.

[0052] The total information bottleneck loss is obtained as follows:

[0053] .

[0054] Furthermore, the category center contrast learning in step S2 specifically includes:

[0055] A category center contrast learning method is constructed to calculate the central feature of each category, compare the pixel feature with the central features of all categories, and classify the pixel feature into the category whose central feature is most similar to it according to the cosine similarity. The positive sample pair of the category center contrast learning proposed in the present invention is the comparison between the pixel feature and its corresponding category center, and the negative sample pair is the comparison between the pixel feature and the category center of other categories.

[0056] The category center feature is calculated as:

[0057] The predicted probability value for the labeled RGB data is ,Will middle The eigenvector of the position is recorded as , label middle The eigenvector of the position is recorded as ,Will The pseudo labels generated middle The eigenvector of the position is recorded as ,in, is the feature dimension, and the central feature of each category calculated based on RGB annotation data is:

[0058] ;

[0059] where is the category central feature of category in the RGB annotation image, is the category central feature of category in the RGB annotation image. The category includes two categories: the lesion area and the normal area. According to this formula, the positive sample category central feature of is calculated (denoted as ), and the negative sample category central feature (denoted as ). The cosine similarity is used to measure the similarity between the pixel feature and :

[0060] ;

[0061] || represents the norm of the vector. Similarly, the cosine similarity is used to measure the similarity between the pixel feature and :

[0062] ;

[0063] Then, based on the InfoNCE loss, the category center contrast learning loss for the annotated RGB image during training is constructed:

[0064] ;

[0065] where is the number of pixels in, is the negative sample set, is the temperature parameter, which is used to adjust the scale of the similarity in contrast learning. In the present invention, is uniformly set to 0.1.

[0066] Similarly, the central feature of each category calculated based on the RGB unannotated image is:

[0067] ;

[0068] where is the category central feature of category in the RGB unannotated image, is the number of pixels in;

[0069] Training The contrastive learning loss of the unlabeled RGB image without category center annotation is as follows:

[0070] ;

[0071] where, is the number of pixels in is the predicted probability distribution generated by training the unlabeled data at position in the feature vector, is the corresponding positive sample, is the corresponding negative sample, is the negative sample set. Similarly, the category center contrast loss of the network training labeled HSL data and the category center contrast loss of the network training unlabeled HSL data are obtained. The total category center contrast learning loss is as follows:

[0072] .

[0073] Adding the dual-view interaction learning loss, the information bottleneck loss, and the category center contrast learning loss together, the total loss function of the dual-view interaction learning semi-supervised training method (denoted as DSDNet) proposed in the present invention based on information bottleneck and contrast learning is obtained:

[0074] ;

[0075] where, and are the trade-off coefficients in the total loss function. Through the trade-off coefficients, the model performance can be further optimized.

[0076] ​The present invention proposes a semi-supervised training method for dual-view interactive learning based on information bottleneck and contrast learning. The method mainly includes three parts: dual-view interactive learning, information bottleneck theory, and class center contrast learning. The collected RGB images of digestive tract diseases are converted into HSL views, and then the collected RGB view images and the converted HSL views are respectively used as the inputs of two parallel networks in the dual-view interactive learning. Moreover, the data flow in the dual-view interactive learning process can be divided into a labeled path and an unlabeled path. For the labeled data, the two segmentation networks are optimized by calculating the cross-entropy loss and Dice loss between the predicted value and the true label. For the unlabeled data, the two segmentation networks generate pseudo-labels to supervise each other for optimization, which can effectively alleviate the coupling problem caused by training a single view in the parallel network, improve the learning ability of the parallel network, enable the model to learn richer semantic information, and improve the semi-supervised segmentation accuracy. The information bottleneck theory is used for shape perception. The information bottleneck theory is used to strengthen the consistency between the feature layers (convolutional layer and deconvolutional layer) and the labels (or pseudo-labels) in the segmentation network, and weaken the consistency between the feature layers and the input values. It is like a filter that filters out the redundant information in the feature layers and only retains the key information. By introducing the information bottleneck theory, the neural network is guided to perceive the shape of the lesion, accurately distinguish the lesion area from the normal area, and effectively solve the problem of incorrect segmentation caused by the blurred lesion boundary in the digestive tract disease images. In addition, the present invention applies the class center contrast learning method during the training process. First, the center features of each class are constructed. Second, a certain pixel feature is compared with the center features of all classes. Finally, the pixel feature is classified into the class with which the center feature is most similar. By comparing the pixel feature with the class center feature, the intra-class compactness and inter-class separability are improved, the problem of difficult detection of tiny lesions is effectively solved, and the segmentation ability of the model for tiny lesions is enhanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0078] Figure 1 It is a schematic flowchart of a semi-supervised training method for dual-view interactive learning based on information bottleneck and contrast learning provided by an embodiment of the present invention;

[0079] Figure 2 It is a schematic diagram of the principle of dual-view interactive learning provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0080] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0081] The present invention proposes a semi-supervised training method for dual-view interactive learning based on information bottleneck and contrast learning, as Figure 1 shown. The method includes the following steps:

[0082] Step S1: Collect medical image data, and then preprocess the image data. The preprocessing includes data augmentation and view transformation;

[0083] Step S2: Construct a semi-supervised training model based on dual-view interactive learning, information bottleneck theory, and class center contrast learning. The segmentation model in the dual-view interactive learning consists of Network A and Network B. The U-Net is used as the segmentation network for Network A and Network B. The RGB image of the digestive tract disease collected by the endoscope is input into Network A, and the corresponding HSL image of the RGB image is input into Network B. In the parallel network composed of Network A and Network B, the dual-view interactive learning is used to avoid the coupling effect of the parallel network. The information bottleneck theory and class center contrast learning are used to enhance the perception and discrimination ability of the network;

[0084] Step S3: Train the segmentation model and train the preprocessed data set through the semi-supervised training model. During the training process, parameter updates are performed through forward propagation and backward propagation, and the training is stopped when the performance of the network reaches the optimal, so as to obtain the weight parameters with the best segmentation effect.

[0085] The specific content of step S1 includes: dividing the collected image data into labeled data and unlabeled data, and then using the color space conversion algorithm to convert the collected RGB image into the corresponding HSL image. The labeled RGB image is denoted as , the labeled HSL image is denoted as , the unlabeled RGB image is denoted as , the unlabeled HSL image is denoted as , the total data set is , where , , and are the image feature values input to the network, is the label value, and respectively represent the height and width of the image.

[0086] The data augmentation method in step S1 includes horizontal flipping, vertical flipping, and random scaling.

[0087] As Figure 2 shown, the dual-view interactive learning in step S2 specifically includes:

[0088] A parallel network framework is composed of two segmentation networks A and B with the same structure. The segmentation network A is denoted as , and the segmentation network B is denoted as , Train the endoscopic RGB images collected clinically, Train the corresponding HSL images. During the training process, for the labeled images, the two segmentation networks are optimized by calculating the cross-entropy loss and Dice loss between the predicted value and the true value. For the unlabeled images, the two segmentation networks are mutually supervised and optimized through pseudo-labels, thereby avoiding the coupling caused by training a single view in the parallel network and realizing the complementary advantages of lesion features in different views, and improving the segmentation accuracy.

[0089] Furthermore, use to train the labeled RGB images of digestive tract diseases , and use the combination of cross-entropy loss and Dice loss as the supervision loss for training , expressed as:

[0090] ;

[0091] where is the supervision loss for RGB labeled data, is the number of labeled RGB images, is the cross-entropy loss, is the Dice loss, The calculation method is as follows:

[0092] ;

[0093] The calculation method is as follows:

[0094] ;

[0095] where and respectively represent the label value and predicted probability of the input image. The label value is obtained through manual annotation, and the predicted probability is obtained by training, and respectively represent the height and width of the image, and respectively represent the label value and predicted probability of the th pixel in the image.

[0096] Furthermore, use the labeled HSL images for training. Similarly, use the combination of cross-entropy loss and Dice loss as the supervision loss for training the labeled HSL images, expressed as:

[0097] ;

[0098] where is the supervision loss for the HSL annotation data, is the number of labeled HSL images.

[0099] Furthermore, use to train the unlabeled RGB image data to obtain the corresponding predicted probability value , use to train the unlabeled HSL image data to obtain the corresponding predicted probability value , then, through the network generate the pseudo-label supervision network , and the calculation method of the loss function is as follows:

[0100] ;

[0101] where is the number of unlabeled RGB images, is the pseudo-label generated by , similarly, through the network generate the pseudo-label supervision network , and the calculation method of the loss function is as follows:

[0102] ;

[0103] where is the number of unlabeled HSL images, is the pseudo-label generated by .

[0104] Adding the above losses gives the total loss function of the dual-view interactive learning, expressed as:

[0105] ;

[0106] By using the dual-view interactive learning strategy, the two segmentation networks supervise each other, avoid coupling, and thus improve the segmentation ability of the semi-supervised model for the digestive tract injury area.

[0107] Furthermore, the information bottleneck theory in step S2 specifically includes:

[0108] The Information Bottleneck (IB) theory is a concept in information theory. The core idea of this theory is that the primary task of most learning models is to extract label-related information from raw data and remove label-unrelated information. The IB principle is as follows:

[0109] ;

[0110] Among them, represents the mutual information between two variables, represents the input value of the network, represents the label-related information extracted from In the present invention, represents and the convolutional layer and deconvolutional layer of represents the label value, is the adjustable parameter in

[0111] Considering that the calculation of is very complex, the present invention selects a more convenient metric, namely the Hilbert-Schmidt Independence Criterion (HSIC). The definition of HSIC is as follows:

[0112] ;

[0113] Among them, , , is the identity matrix, is the column vector with all elements being 1, represents the variable and the number of rows of represents the trace of the matrix. Calculating requires ensuring that and have the same number of rows, is the kernel matrix, and its calculation method is as follows:

[0114] ;

[0115] Among them, and are two different variables in is the adjustable parameter, The 2-norm of a representative vector, also known as the Euclidean norm, is defined similarly to :

[0116] ;

[0117] where and are two different variables in

[0118] Based on the above definition of the information bottleneck, the information bottleneck loss of the RGB image labeled in

[0119] ;

[0120] where is the labeled RGB image, is the corresponding label value, is the matrix of the feature layers (convolutional layer and deconvolutional layer) during training is the number of feature layers, the information bottleneck loss of the unlabeled RGB image in

[0121] ;

[0122] where is the unlabeled RGB image, is the pseudo-label value generated by is the matrix of the feature layers during training

[0123] Similarly, the information bottleneck loss of the labeled HSL image in can also be obtained as:

[0124] ;

[0125] where is the labeled HSL image, is the corresponding label value, is the matrix of the feature layers during training

[0126] the information bottleneck loss of the unlabeled HSL image in

[0127] ;

[0128] Among them, is the unlabeled HSL image, is the generated pseudo-label, is the feature layer matrix during training.

[0129] Thus, the total information bottleneck loss is:

[0130] .

[0131] Furthermore, the class center contrast learning in step S2 specifically includes:

[0132] Construct a class center contrast learning method, calculate the center feature of each class, compare the pixel feature with the center features of all classes, and classify the pixel feature into the class whose center feature is the most similar to it according to the cosine similarity. In this way, the intra-class compactness and inter-class separability of the damage classes are improved, and the model's ability to recognize minor lesions is enhanced. The positive sample pair of the class center contrast learning proposed by the present invention is the contrast between the pixel feature and its corresponding class center, and the negative sample pair is the contrast between the pixel feature and the class centers of other classes.

[0133] The calculation method of the class center feature is:

[0134] The predicted probability value of the labeled RGB data is , and the feature vector at the position is denoted as , and the feature vector at the position in the label is denoted as , and the feature vector at the position in the generated pseudo-label

[0135] ;

[0136] Among them, is the class center feature of class in the RGB labeled image, is the number of pixels of class in the RGB labeled image. The class includes two classes: the lesion area and the normal area. According to this formula, the positive sample class center feature of is calculated (denoted as ), and the negative sample class center feature (denoted as ), use cosine similarity to measure the similarity between pixel feature and :

[0137] ;

[0138] || represents the norm of the vector. Similarly, use cosine similarity to measure the similarity between pixel feature and :

[0139] ;

[0140] Then, based on the InfoNCE loss, construct the contrastive learning loss of the labeled RGB image class center during training :

[0141] ;

[0142] Among them, is the number of pixels in is the negative sample set, is the temperature parameter, used to adjust the scale of similarity in contrastive learning. In the present invention, is uniformly set to 0.1.

[0143] Similarly, the class center feature calculated for each class based on the RGB unlabeled image is:

[0144] ;

[0145] Among them, is the class center feature of class in the RGB unlabeled image, is the number of pixels in

[0146] During training the contrastive learning loss of the unlabeled RGB image class center is:

[0147] ;

[0148] Among them, is the number of pixels in is the predicted probability distribution generated by training the unlabeled data at the position in the feature vector, is The corresponding positive samples are the corresponding negative samples is the negative sample set. Similarly, the category center contrast loss of the network training labeled HSL data and the category center contrast loss of the training unlabeled HSL data are obtained. The total category center contrast learning loss is:

[0149] .

[0150] Adding the dual-view interaction learning loss, the information bottleneck loss, and the category center contrast learning loss, the total loss function of the dual-view interaction learning semi-supervised training method (denoted as DSDNet) proposed by the present invention based on information bottleneck and contrast learning is obtained:

[0151] ;

[0152] wherein and are the trade-off coefficients in the total loss function. Through the trade-off coefficients, the model performance can be further optimized.

[0153] The method of the present invention is experimentally verified through specific data as follows:

[0154] Step 1: Construction of the digestive tract disease dataset. Use the private early esophageal cancer and the public colon polyp datasets as the data required for training the model. Label 10% of the data and leave the rest unlabeled.

[0155] Step 2: Preprocessing of the dataset. To improve the robustness of the model and avoid overfitting from affecting the model performance, the present invention uses data augmentation methods such as horizontal flipping, vertical flipping, and random scaling in the training of the model.

[0156] Step 3: Construction of the dual-view interaction learning semi-supervised training method based on information bottleneck and contrast learning. Use U-Net as the segmentation network for Network A and Network B. Introduce the information bottleneck theory and category center contrast learning in the parallel network composed of Network A and B to enhance the perception and discrimination ability of the network, and then avoid the coupling effect of the parallel network through dual-view interaction learning.

[0157] Step 4: Train the segmentation model. Train the preprocessed dataset through the dual-view interaction learning semi-supervised training method based on information bottleneck and contrast learning. During the training process, update the parameters through forward propagation and backward propagation, and stop training when the performance of the network reaches the optimal. In this way, the weight parameters with the best segmentation effect are obtained.

[0158] Use the private early esophageal cancer and the publicly available colon polyp datasets as the data required for training the model. The early esophageal cancer (EEC) dataset consists of early esophageal cancer images collected from a large well-known hospital, with a total of 2,689 images. In this embodiment, 2,189 images are randomly selected as the training set, 500 images are used as the test set, and 10% of the data is labeled, while the rest is unlabeled. The Kvasir-SEG (KS) dataset is a publicly available endoscopic dataset for pixel-level segmentation of colon polyps, including 1,000 lesion images and their corresponding labeled images. In this embodiment, 800 images are randomly selected as the training set, and 200 images are used as the test set.

[0159] The experiments in this embodiment are implemented using an NVIDIA A100 GPU based on the Pytorch deep learning framework. The input data is uniformly cropped to a size of 256×256 pixels. During the training process, stochastic gradient descent (SGD) with a momentum of 0.9 and a weight decay of 0.0001 is used as the optimizer, and the initial learning rate is 0.03. The batch size of the training data in the experiment is set to 16, including 8 labeled data and 8 unlabeled data. The parameters in the present invention take the values of = 1.5, = 2.5, = 0.3, = 0.7. On the KS dataset, the values are = 2.4, = 3.6, = 0.5, = 0.3. In the test phase, the predicted values of Network A are used as the overall predicted values of the model. For fair comparison, all experiments are conducted with the same experimental settings.

[0160] The performance of the model is evaluated using metrics such as the Dice coefficient, intersection over union (IoU), accuracy (Acc), mean absolute error (MAE), and 95% Hausdorff distance (95HD). Among them, the higher the values of the Dice, IoU, and Acc metrics, the better the model performance; the lower the values of the MAE and 95HD metrics, the better the model performance.

[0161] The experimental results are shown in Table 1 and Table 2. Table 1 shows the test results of the method of the present invention (DSDNet) and other semi-supervised segmentation algorithms in the EEC dataset. It can be seen from the table that the method of the present invention leads other methods in terms of Dice, IoU, Acc, MAE, and 95HD metrics. The test results of other methods are not ideal, mainly because in some EEC images, the lesion areas are very similar to the normal areas, resulting in difficult identification of lesion boundaries and easy segmentation errors by the model. In contrast, the method of the present invention has achieved excellent segmentation results, and each index is significantly better than other methods, which reflects that the present invention can effectively solve the problem of incorrect segmentation of the model caused by fuzzy boundaries.

[0162] Table 2 shows the test results of the method of the present invention (DSDNet) and other semi-supervised segmentation algorithms in the KS dataset. Similarly, the method of the present invention is superior to other methods in terms of Dice, IoU, Acc, MAE, and 95HD metrics. It should be noted that there are many tiny polyps in the KS dataset. Due to the inconspicuous features of these tiny lesions, it is easy to result in poor segmentation effects of the model. Therefore, the existing methods often have poor segmentation effects on tiny polyps. However, the method of the present invention can effectively alleviate this problem. It can be seen from Table 2 that DSDNet is significantly superior to other methods in all metrics. Therefore, the method of the present invention can effectively solve the problem of difficult segmentation of tiny polyps. It can be seen that the test metrics of the method of the present invention in the two digestive disease segmentation datasets have reached the optimal results and are far ahead of the existing semi-supervised segmentation methods, indicating that the method of the present invention has excellent performance in the digestive disease segmentation task.

[0163] Table 1 Comparison of evaluation metrics between the DSDNet method of the present invention and other methods in the EEC dataset

[0164]

[0165] Table 2 Comparison of evaluation metrics between the DSDNet method of the present invention and other methods in the KS dataset

[0166]

[0167] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A dual-view interactive learning semi-supervised training method based on information bottleneck and contrastive learning, characterized in that: The method comprises: Step S1: collect medical image data, and then preprocess the image data, the preprocessing includes data enhancement and view conversion; Step S2: A semi-supervised training model is constructed based on dual-view interactive learning, information bottleneck theory and category center contrast learning. The segmentation model in the dual-view interactive learning consists of network A and network B. U-Net is used as the segmentation network of network A and network B. The collected RGB image is input into network A, and the HSL image corresponding to the RGB image is input into network B. In the parallel network composed of network A and network B, dual-view interactive learning is used to avoid coupling between parallel networks. Information bottleneck theory and category center contrast learning are used to enhance the perception and discrimination ability of the network. Step S3: Train the segmentation model and train the preprocessed data set through a semi-supervised training model. During the training process, parameters are updated through forward propagation and back propagation, and the training is stopped when the network performance reaches the optimal level, so as to obtain the weight parameters with the best segmentation effect.

2. The method according to claim 1, characterized in that: The step S1 further comprises: The collected image data is divided into Annotated data and Unlabeled data, and then use the color space conversion algorithm to convert the collected RGB image into the corresponding HSL image, and record the labeled RGB image as , the annotated HSL image is recorded as , the unlabeled RGB image is denoted as , the unlabeled HSL image is denoted as The total data set is ,in, , , and is the image feature value of the network input, is the label value, and the data enhancement methods include horizontal flipping, vertical flipping and random scaling.

3. The method according to claim 1, characterized in that The dual-view interactive learning in step S2 specifically includes: using two segmentation networks A and B with the same structure to form a parallel network framework, and the segmentation network A is recorded as , the segmentation network B is recorded as , RGB images collected for training, The corresponding HSL images are trained. During the training process, for labeled images, the two segmentation networks are optimized by calculating the cross entropy loss and Dice loss between the predicted values ​​and the true values. For unlabeled images, the two segmentation networks are optimized by mutual supervision through pseudo labels.

4. The method according to claim 3, characterized in that: use Training labeled RGB images , through the combination of cross entropy loss and Dice loss as training The supervision loss is expressed as: ; in, yes Supervision loss for RGB annotated data, is the number of annotated RGB images, is the cross entropy loss, is the Dice loss, expressed as: , ,in, and Represent the label value and prediction probability of the input image respectively. The label value is obtained by manual annotation, and the prediction probability is obtained by Trained, and Represents the height and width of the image respectively. and Respectively represent the The label value and predicted probability of each pixel.

5. The method according to claim 4, characterized in that use The HSL images labeled for training also use a combination of cross entropy loss and Dice loss as the supervision loss for training labeled HSL images, expressed as: ; in, yes Supervision loss for HSL labeled data, is the number of annotated HSL images; use Train the unlabeled RGB image data to obtain the corresponding prediction probability value ,use Train the unlabeled HSL image data to obtain the corresponding predicted probability value , then, through the network Generated pseudo-label supervision network , the loss function is calculated as follows: ; in, is the number of unlabeled RGB images, Is Generated pseudo labels; Via the network Generated pseudo-label supervision network , the loss function is calculated as follows: ; in, is the number of unlabeled HSL images, Is Generated pseudo labels; Add the above losses to get the total loss function of dual-view interactive learning , expressed as: 。 6. The method according to claim 1, characterized in that In step S2 The information bottleneck loss of the annotated RGB image is expressed as: ; Among them, HSIC is the Hilbert-Schmidt independence criterion, is the annotated RGB image, is the corresponding label value, for train The feature layer matrix in the process, is the number of feature layers; The information bottleneck loss of the unlabeled RGB image in is expressed as: ; in, is an unlabeled RGB image, for The pseudo label value generated, for train The feature layer matrix in the process; The information bottleneck loss of the HSL image annotated in is: ; in, is the annotated HSL image, is the corresponding label value, for train The feature layer matrix in the process; The information bottleneck loss of the unlabeled HSL image in is: ; in, is an unlabeled HSL image, for The pseudo labels generated are for train The feature layer matrix in the process; The total information bottleneck loss is obtained as follows: 。 7. The method according to claim 6, characterized in that HSIC is defined as: ; in, , , is the identity matrix, is a column vector with all elements equal to 1, Representation variables and the number of rows, represents the trace of the matrix, is the kernel matrix, expressed as: ,in, and yes Two different variables in is an adjustable parameter, represents the 2-norm of a vector; It is expressed as: ,in, and yes Two different variables in .

8. The method according to claim 1, characterized in that: The category center contrast learning in step S2 specifically includes: constructing a category center contrast learning method, calculating the central feature of each category, comparing the pixel feature with the central features of all categories, and classifying the pixel feature into the category whose central feature is most similar to it according to cosine similarity. The positive sample pair of category center contrast learning is the comparison between the pixel feature and its corresponding category center, and the negative sample pair is the comparison between the pixel feature and the category center of other categories.

9. The method according to claim 8, characterized in that The category center feature is calculated as: The predicted probability value for the labeled RGB data is ,Will middle The eigenvector of the position is recorded as , label middle The eigenvector of the position is recorded as ,Will The pseudo labels generated middle The eigenvector of the position is recorded as ,in, is the feature dimension. The central feature of each category calculated based on the RGB labeled data is: ; in, is the category in the RGB labeled image The category center feature of is the category in the RGB labeled image The number of pixels in the image is divided into two categories: lesion area and normal area. The formula is used to calculate The positive sample category center feature and negative sample category center features , using cosine similarity to measure pixel features and The similarity between: ; | | represents the norm of the vector. Similarly, cosine similarity is used to measure pixel features. and The similarity between: ; Then, we build a training set based on the InfoNCE loss. The category-centered contrastive learning loss for labeled RGB images is: ; in, yes The number of pixels in is the set of negative samples, is the temperature parameter, which is used to adjust the scale of similarity in contrastive learning; Similarly, the central feature of each category calculated based on the RGB unlabeled image is: ; in, is the category in the RGB unlabeled image The category center feature of yes The number of pixels in ; train The category-centered contrastive learning loss of unlabeled RGB images is: ; in, yes The number of pixels in for Training on unlabeled data The resulting predicted probability distribution Middle position The characteristic vector of yes The corresponding positive sample is yes The corresponding negative samples are is a set of negative samples; Similarly, we can conclude Network training class center contrast loss of labeled HSL data Compared with the category center loss of training unlabeled HSL data , the total category-centered contrastive learning loss is: 。 10. The method according to claim 5, 6 or 9, characterized in that: Add the dual-view interaction learning loss, information bottleneck loss, and category center contrast learning loss to get the total loss function: ; in, and is the trade-off coefficient in the total loss function.

Citation Information

Patent Citations

  • Weak supervision image semantic segmentation method, system and device and storage medium

    CN116309653A

  • Semi-supervised reference-free image quality evaluation method based on uncertainty estimation

    CN117541562A