A method for image semantic segmentation based on dual network collaborative segmentation model
By introducing a dual network collaborative segmentation model in the semi-supervised semantic segmentation model, the cross-network consistency learning and contrast learning methods are adopted to solve the problems of category imbalance, difficulty in coupling the model, low pseudomark quality and insufficient feature distinction, and more efficient and stable semantic segmentation performance is achieved.
Patent Information
- Application Number
- CN202310309481.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-28
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2043-03-28
AI Technical Summary
The existing semi-supervised semantic segmentation model has problems of inefficiency and performance instability when dealing with problems such as natural image category imbalance, difficulty in model coupling, low pseudo-marking quality and insufficient feature distinction.
A semantic segmentation method based on dual network collaborative segmentation model is proposed. By constructing two segmentation networks with different parameters but the same structure, weak to strong data augmentation and network interference are used to perform cross-network consistency learning and contrast learning, and improve the feature generalization ability and performance of the model.
This method alleviates the pseudo-label overfitting and model degradation problems by improving the consistency training difficulty of the model, improves the separability between classes in the feature space, reduces the dependence on labeled data, and improves the efficiency in practical applications.
Smart Images

Figure CN116468888B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and particularly relates to an image semantic segmentation method based on a dual-network collaborative segmentation model. Background Art
[0002] The purpose of image semantic segmentation is to classify the semantic categories of each pixel in an image. As a basic computer vision task, it plays a very important role in the field of image understanding, such as autonomous driving, medical imaging and other applications. In recent years, driven by the rapid development of deep neural network technology, the availability of large-scale datasets and high computing resources, fully supervised learning methods have made remarkable progress. However, existing fully supervised methods rely on large-scale annotated data, especially for fully supervised semantic segmentation methods that require expensive pixel annotations. To address this limitation, semi-supervised semantic segmentation (requiring only a small amount of labeled data) has recently attracted increasing attention. Semantic segmentation based on semi-supervised learning is a research method that solves problems by adopting semi-supervised learning methods while based on classical semantic segmentation models. The segmentation models used therein are mainly deep learning-based segmentation networks. Different from traditional supervised learning semantic segmentation, semi-supervised methods can utilize both unlabeled data and labeled data that are valueless in supervised learning, making the field of semantic segmentation closer to the way of human learning.
[0003] The current state-of-the-art semi-supervised semantic segmentation models have demonstrated that designing pseudo-label consistency regularization and pixel-level contrast regularization for different views of unlabeled images is beneficial to improving model performance. The former type of work usually uses different data augmentations, feature perturbations and network perturbations to obtain different views. For example, different views are obtained through two networks with different parameters, and the outputs of these two networks are aligned in the label space through cross-pseudo supervision. The latter type of work benefits from designing effective positive and negative sample pairs for the anchor features in the feature space. Usually, a filtering mechanism for additional positive / negative samples also needs to be carefully designed. For example, by forcing the student model and the teacher model that can produce more stable and less noisy results to maintain pixel-level contrast regularization in the projected feature space.
[0004] Although excellent results have been achieved and great potential has been demonstrated in natural image semantic segmentation based on semi-supervised learning, the current work still has the following several research problems to be solved:
[0005] (1) The problem of natural image class imbalance. The problem of class imbalance is naturally widespread in practical applications. When the training data is highly imbalanced, most learning frameworks will show a bias towards the majority class. In some extreme cases, the minority class may be completely ignored, seriously affecting the efficiency of the prediction model. However, to handle the semi-supervised problem, it is usually assumed that the training dataset is evenly distributed across all class labels.
[0006] (2) The problem of model coupling caused by insufficient mutual learning in consistency learning. To enable the model to learn effective knowledge from unlabeled data by producing stable predictions for different views of unlabeled images, it is particularly important to design a method for obtaining different views. If the process of producing stable predictions is too simple, it will lead to the model being unable to learn knowledge due to degradation, such as traditional self-training methods.
[0007] (3) The problem of the quality of pseudo-labels. The quality of pseudo-labels can directly determine the performance of semi-supervised semantic segmentation models. Currently, many studies have tried to solve this problem, but there is not yet a very good solution. The reason for this problem is that the current semi-supervised learning method is the same as the fully supervised learning method. The former learns from pseudo-labels, and the latter learns under manually annotated labels. Currently, most methods alleviate the quality problem of pseudo-labels by using teacher-student models or filtering out unreliable pseudo-labels using confidence.
[0008] (4) In deep learning methods, the distinctiveness of the features learned by the model is closely related to the performance of the model. Recently, in the field of image semantic segmentation, this problem has attracted a large number of researchers due to the success of contrastive learning. Using contrastive learning can not only enable the model to better distinguish between different classes while fitting the labeled data. Moreover, since this problem is still in the initial stage of research, it is a future research trend in semi-supervised semantic segmentation. Summary of the Invention
[0009] To solve the above technical problems, the present invention proposes an image semantic segmentation method based on a dual-network collaborative segmentation model, including:
[0010] Construct a dual-network collaborative segmentation model and train it, and input the image to be segmented into the trained dual-network collaborative segmentation model to obtain the segmentation result;
[0011] The dual-network collaborative segmentation model includes: two segmentation networks f θ1 and f θ2 with different parameters but the same network structure. The f θ1 and f θ2 include: an encoder, a classifier, and a projection head;
[0012] The training process of the dual-network collaborative segmentation model includes the following steps:
[0013] S1: Input the labeled data into the segmentation network f θ1 and f θ2 , and use the traditional fully supervised semantic segmentation method based on pixel-level cross-entropy to train the model. The fully supervised loss function for training is represented by ;
[0014] S2: Perform weak data augmentation A u processing on the unlabeled data x w to obtain the weakly data-augmented version x w = A w (x u );
[0015] S3: Perform strong data augmentation A w twice on the weakly data-augmented version x s to obtain two different strongly data-augmented versions x s1 = A s (x w ) and x s2 = A s (x w );
[0016] S4: Input the weakly data-augmented version x w into the segmentation network f θ1 and f θ2 at the same time. First, encode x w through the encoder, and then use the classifier for classification to obtain the segmentation predictions and
[0017] S5: Input the two strongly data-augmented versions x s1 and x s2 into the segmentation network f θ1 and f θ2 respectively. The two networks encode x s1 and x s2 through the encoder respectively, and then use the classifier for classification to obtain the segmentation predictions and
[0018] S6: Perform hardening operations on the predictions of the weakly data-augmented version respectively to obtain the pseudo-labels y 1 and y 2 ;
[0019] S7: According to the pseudo-labels y 1 and y 2Using a loss function Perform consistency learning across networks in the label space;
[0020] S8: Input the weakly data-augmented version x of the unlabeled data w into the segmentation network f θ1 and f θ2 , and respectively encode x w through the encoders of the two segmentation networks, and project the encoded x w into the contrastive learning feature space through the projection head to obtain the feature sets z 1 and z 2 ;
[0021] S9: Define the features with the same position projected into the same feature space from different networks in the feature sets z 1 and z 2 as a pair of positive samples, and define the features with different pseudo-labels as a pair of negative samples;
[0022] S10: Obtain the loss function of the model's cross-network contrastive learning in the feature space through the positive and negative sample pairs Through optimization perform contrastive learning to make the anchor feature vector attract to the positive samples and repel from the negative samples;
[0023] S11: Obtain the overall loss function of the dual-network collaborative segmentation model according to the loss functions of the fully supervised training and the semi-supervised training. When the loss value of the overall loss function is the smallest, complete the training of the dual-network collaborative segmentation model to obtain the trained dual-network collaborative segmentation model.
[0024] Advantages of the present invention:
[0025] The present invention designs a network model with two different parameters but the same structure. Through weak-to-strong data augmentation and network interference, the difficulty of the model's consistency training is increased, alleviating the overfitting of the model to incorrect pseudo-labels and the problem of model degradation caused by insufficient consistency learning difficulty in consistency training; at the same time, contrastive learning is added to the feature spaces of the two networks, enabling the two networks to learn from each other in the feature space and improving the generalization ability of the features extracted by the model; at the same time, the model of the present invention only requires less labeled data, reducing the data labeling process that consumes a large amount of funds in actual engineering applications. Description of the drawings
[0026] Figure 1 is the semi-supervised semantic segmentation method of cross-network contrastive learning of the present invention;
[0027] Figure 2 is the segmentation result diagram of the embodiment of the present invention. Detailed implementation manners
[0028] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0029] An image semantic segmentation method based on a dual-network collaborative segmentation model includes:
[0030] Construct a dual-network collaborative segmentation model and train it, and input the image to be segmented into the trained dual-network collaborative segmentation model to obtain a segmentation result;
[0031] The dual-network collaborative segmentation model includes: two segmentation networks f θ1 and f θ2 with different parameters but the same network structure. The f θ1 and f θ2 include: an encoder, a classifier, and a projection head;
[0032] The training process of the dual-network collaborative segmentation model includes the following steps:
[0033] Step 1: Input the labeled data into the segmentation networks f θ1 and f θ2 , and use the traditional fully supervised semantic segmentation method based on pixel-level cross-entropy to train the model. The fully supervised loss function for training is represented by ;
[0034] As shown in Figure 1 , Step 2: For the unlabeled data, use the image semi-supervised semantic segmentation method based on cross-network contrast learning to perform semi-supervised training on the model;
[0035] S1: Perform weak data augmentation A u processing on the unlabeled data x w to obtain a weakly data-augmented version x w = A w (x u );
[0036] S2: Perform strong data augmentation A s processing on the weakly data-augmented version x w twice to obtain two different strongly data-augmented versions x s1 = A s (x w ) and x s2 = A s(x w );
[0037] S3: Input the weakly data - augmented version x w into the segmentation networks f θ1 and f θ2 simultaneously. First, encode x w through the encoder, and then classify it using the classifier to obtain the segmentation predictions and
[0038] S4: Input the two strongly data - augmented versions x s1 and x s2 into the segmentation networks f θ1 and f θ2 respectively. The two networks encode x s1 and x s2 through the encoder respectively, and then classify them using the classifier to obtain the segmentation predictions and
[0039] S5: Perform hardening operations on the predictions of the weakly data - augmented version respectively to obtain the pseudo - labels y 1 and y 2 ;
[0040] S6: Use the loss function 1 and y 2 to perform consistency learning across networks in the label space, and further improve the performance of the entire model through complementary learning in the label space;
[0041] S7: Input the weakly data - augmented version xof the unlabeled data w into the segmentation networks f θ1 and f θ2 respectively. Encode x w through the encoders of the two segmentation networks, and project the encoded x w into the contrastive learning feature space through the projection head to obtain the feature sets z 1 and z 2 ;
[0042] S8: Define the features with the same position projected into the same feature space from different networks in the feature sets z 1 and z 2 as a pair of positive samples, and define the features with different pseudo - labels as a pair of negative samples;
[0043] S9: Obtain the loss function of the model's cross - network contrastive learning in the feature space through positive and negative sample pairs
[0044] By optimizing the anchor feature vector can be made to attract positive samples and repel negative samples, thereby improving the inter-class separability in the feature space of the model during the mutual learning process of the two networks;
[0045] Step 3: Obtain the overall loss function of the dual-network collaborative segmentation model according to the loss functions of full-supervised training and semi-supervised training. When the loss value of the overall loss function is minimized, the training of the dual-network collaborative segmentation model is completed, and the trained dual-network collaborative segmentation model is obtained.
[0046] The model is trained using a traditional full-supervised semantic segmentation method based on pixel-level cross-entropy. The loss function for full-supervised training of the labeled data includes:
[0047]
[0048] where l ce () represents the cross-entropy loss function, x l represents the labeled data, y represents the label corresponding to x l , W and H are the width and height of the image x to be segmented respectively, f l represents the segmentation network of the model, T represents the transpose operation, y θ represents the label with index i, i and represents the labeled data with index i.
[0049] The cross-network semi-supervised learning includes cross-network consistency learning and cross-network contrast learning. The former effectively utilizes the knowledge of unlabeled data through cross-supervised and asymmetric data augmentation methods, and uses the consistency based on pseudo-labels of different views for localization. The latter enables two models with the same semantic space in the overall to learn more diverse and discriminative features through the two networks with different parameters in the feature space.
[0050] The strong and weak data augmentation includes:
[0051] The weak data augmentation method includes: random cropping, random horizontal flipping, and random scaling;
[0052] The strong data augmentation method includes: random Gaussian blur, random color distortion, and random gray-scale scaling.
[0053] Obtaining pseudo-labels includes:
[0054] Taking and Perform hardening, and take the maximum value in the category channel as the pseudo-label of this pixel. The pseudo-label is represented by y 1 and y 2 That is,
[0055] After obtaining the pseudo-labels, filtering operations on the pseudo-labels can also be performed. To obtain high-quality pseudo-labels, filtering is performed through a confidence-awareness-based method: If networks with different parameters generate pseudo-labels of different categories at the same location, the pixel is regarded as a difficult pixel. The pseudo-labels generated at this pixel position are considered untrustworthy. On the contrary, the positions of trustworthy pseudo-labels can be found, as well as the positions of the trustworthy pseudo-labels.
[0056] The loss function for the cross-network consistency learning of the model in the label space includes:
[0057]
[0058] Among them, represents the loss function for the cross-network consistency learning of the model in the label space, l ce () represents the cross-entropy loss function, represents the segmentation network f θ1 generates the segmentation prediction on the strongly data-augmented version x s1 , y 1 represents the segmentation network f θ1 generates the pseudo-label on the weakly data-augmented version x w . represents the segmentation network f θ2 generates the segmentation prediction on the strongly data-augmented version x s2 , y 2 represents the segmentation network f θ2 generates the pseudo-label on the weakly data-augmented version x w .
[0059] The weakly data-augmented version x w is encoded through the encoder of the network, and the encoded weakly data-augmented version x w is projected into the common feature space through the projection head to obtain the features z 1 and z 2 , including:
[0060] z 1 =projector 1 (encoder 1 (x w ))
[0061] z 2 =projector 2 (encoder2 (x w ))
[0062] where z 1 and z 2 represent the weakly data-augmented versions of the unlabeled data x w generated by the segmentation networks f θ1 and f θ2 respectively. projector 1 () and projector 2 () represent the projection operations by the projection heads of the segmentation networks f θ1 and f θ2 respectively. encoder 1 () and encoder 2 () represent the encoding operations by the encoders of the segmentation networks f θ1 and f θ2 respectively.
[0063] Define positive and negative sample features, including:
[0064] Downsample the pseudo-labels y 1 and y 2 to make y 1 and y 2 have the same resolution as the feature sets z 1 and z 2 respectively. For a given anchor feature vector from network k, k ∈ {1, 2} with index i and its pseudo-label its positive sample is defined as where is the feature generated by another network and is the same as ; for a given anchor feature vector from network k, k ∈ {1, 2} with index j and its pseudo-label the negative sample is defined as all pixel-level features whose pseudo-labels are different from the pseudo-labels of the corresponding anchor features, represented by a set of size N as
[0065] Construct the pixel-level contrastive learning loss functions of the segmentation networks f θ1 to f θ2 and the pixel-level contrastive learning loss functions of the segmentation networks f θ2 to f θ1 respectively, to obtain the overall loss function for cross-network contrastive learning of the model.
[0066] The segmentation networks f θ1 to f θ2The pixel-level contrastive learning loss function includes:
[0067]
[0068] Among them, r() represents the similarity metric function, and z i 1 represents the anchor feature vector from the segmentation network f θ1 with index i, The positive sample of the anchor feature vector z i 1 is z n 1 The negative sample of the anchor feature vector z i 1 is z
[0069] The pixel-level contrastive learning loss function of the segmentation network f θ2 to f θ1 includes:
[0070]
[0071] Among them, r() represents the similarity metric function, and z i 2 represents the anchor feature vector from the segmentation network f θ2 with index i, The positive sample of the anchor feature vector z i 2 is z n 2 The negative sample of the anchor feature vector z i 2 is z
[0072] The overall loss function for cross-network contrastive learning of the model includes:
[0073]
[0074] Among them, represents the overall loss function for cross-network contrastive learning of the model, represents the pixel-level contrastive learning loss function of the segmentation network f θ1 to f θ2 and represents the pixel-level contrastive learning loss function of the segmentation network f θ2 to f θ1 is
[0075] The overall loss function of the dual-network collaborative segmentation model includes:
[0076]
[0077] Among them, represents the loss function for full-supervised training represents the loss function for the consistency learning of the model across networks in the label space The overall loss function for the cross-network contrast learning of the model. λ represents the loss weight for the consistency learning of the model across networks in the label space, and μ represents the loss weight for the cross-network contrast learning of the model, which are respectively used to control the contribution of each item to the overall loss
[0078] During the specific implementation of network training, the SGD optimizer with momentum is used to iteratively optimize the learning parameters of all networks. The momentum is set to 0.9, and the initial learning rate is set to 0.01 until the model training is completed
[0079] In this implementation, as Figure 2 shown, three pictures are respectively input into the trained dual-network collaborative segmentation model, and the images to be segmented are respectively input into the segmentation networks f θ1 and f θ2 . The images to be segmented are encoded through the encoder, and then classified through the classifier to obtain the pixel-level semantic segmentation results under the segmentation networks f θ1 and f θ2 . Different colors in the segmentation results represent different categories. By comparing and analyzing the pixel-level semantic segmentation results under the segmentation networks f θ1 and f θ2 , the best pixel-level semantic segmentation results under the segmentation networks f θ1 and f θ2 are selected as the final segmentation results
[0080] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents
Claims
1. An image semantic segmentation method based on a dual network collaborative segmentation model, characterized in that: include: Construct a dual-network collaborative segmentation model and train it, input the image to be segmented into the trained dual-network collaborative segmentation model to obtain the segmentation result; The dual network collaborative segmentation model includes: two segmentation networks with different parameters and the same network structure. θ1 and f θ2 , the f θ1 and f θ2 Includes: encoder, classifier, and projection head; The training process of the dual network collaborative segmentation model includes the following steps: S1: Input the labeled data into the segmentation network f θ1 and f θ2 , the traditional pixel-level cross entropy-based fully supervised semantic segmentation method is used to train the model, and the fully supervised loss function of the training is express; S2: for unlabeled data x u Perform weak data enhancement A w Processing, get the weak data enhancement version x of the unlabeled data w =A w (x u ); S3: Enhance the weak data version x w Perform two strong data enhancements A s Processing, get two different strong data enhancement versions x s1 =A s (x w ) and x s2 =A s (x w ); S4: Enhance the weak data version x w At the same time, input the segmentation network f θ1 and f θ2 , first x w Encode through the encoder, and then use the classifier to classify and obtain segmentation predictions respectively and S5: Two strong data enhancement versions x s1 、x s2 Input the segmentation network f respectively θ1 and f θ2 , the two networks are respectively s1 、x s2 Encode through the encoder, and then use the classifier to classify and obtain segmentation predictions respectively and S6: Predictions of the enhanced version of weak data Perform hardening operations respectively to obtain pseudo labels y 1 and 2 ; S7: Use loss function based on pseudo labels y1 and y2 Conduct cross-network consistency learning in label space; S8: The weak data enhancement version x of the unlabeled data w Input segmentation network f θ1 and f θ2 , respectively, through the encoder of the two segmentation networks x w Encode and project the encoded x w Projected into the contrastive learning feature space, we get the feature set z 1 、z 2 ; S9: Set the feature set z 1 、z 2 The features with the same position from different networks projected into the same feature space are defined as a pair of positive samples, and the features with different pseudo labels are defined as a pair of negative samples; S10: Obtain the loss function of the model cross-network contrast learning in the feature space through positive and negative sample pairs By optimizing Perform contrastive learning to make the anchor feature vector attract positive samples and repel negative samples; S11: The overall loss function of the dual-network collaborative segmentation model is obtained according to the loss functions of the fully supervised training and the semi-supervised training. When the loss value of the overall loss function is the smallest, the training of the dual-network collaborative segmentation model is completed to obtain a trained dual-network collaborative segmentation model.
2. According to claim 1, the image semantic segmentation method based on the dual network collaborative segmentation model is characterized in that: The loss function for fully supervised training with labeled data include: Among them, l ce () represents the cross entropy loss function, x l represents label data, y represents x l The corresponding labels, W and H are the images to be segmented x l The width and height, f θ represents the segmentation network of the model, T represents the transposition operation, and y i represents the label with index i, Represents the label data with index i.
3. The image semantic segmentation method based on the dual network collaborative segmentation model according to claim 1, characterized in that: The strong and weak data enhancements include: The weak data enhancement method includes: random cropping, random horizontal flipping and random scaling; The strong data enhancement method includes: random Gaussian blur, random color distortion and random grayscale scaling.
4. The image semantic segmentation method based on the dual network collaborative segmentation model according to claim 1, characterized in that: Get pseudo labels, including: Will and Harden and take the maximum value in the category channel as the pseudo label of this pixel. The pseudo label is y 1 and 2 Indicates that 5. The image semantic segmentation method based on the dual network collaborative segmentation model according to claim 1, characterized in that: The loss function of the model's consistent learning in label space across networks includes: in, The loss function that represents the consistency learning of the model across networks in the label space, l ce () represents the cross entropy loss function, represents the segmentation network f θ1 In the strong data augmentation version x s1 The segmentation prediction generated on 1 Represents the segmentation network f θ1 In the weak data enhancement version x w The pseudo labels generated on Represents the segmentation network f θ2 In the strong data augmentation version x s2 The segmentation prediction generated on 2 Represents the segmentation network f θ2 In the weak data enhancement version x w The pseudo labels generated on .
6. The image semantic segmentation method based on the dual network collaborative segmentation model according to claim 1, characterized in that: A weak data augmentation version x is performed through the network's encoder w Encoding, through the projection head to encode the weak data enhanced version x w Projected into the common feature space, we get the features z of the two networks 1 and z 2 ,include: z 1 =projector 1 (encoder 1 (x w )) z 2 =projector 2 (encoder 2 (x w )) Among them, z 1 、z 2 They represent the weak data enhancement version x of the unlabeled data respectively. w By segmenting the network f θ1 and f θ2 Generated features, projector 1 (), projector 2 () represent the segmentation network f θ1 and f θ2 The projection head performs projection operation, encoder 1 (), encoder 2 () represent the segmentation network f θ1 and f θ2 The encoder performs encoding operation.
7. The image semantic segmentation method based on the dual network collaborative segmentation model according to claim 1, characterized in that: Define the positive and negative sample features, including: For the pseudo label y 1 and 2 Downsampling is performed to make y1 and y2 consistent with the resolution of feature sets z1 and z2. For a given anchor feature vector from network k, k∈{1,2}, index i And its pseudo-label Its positive sample is defined as in is a feature generated from another network, and and Same; for a given anchor feature vector from network k, k∈{1,2}, indexed as j And its pseudo-label Negative samples are defined as all pixel-level features whose pseudo labels are different from the corresponding anchor features, and are represented by a set of size N as 8. The image semantic segmentation method based on the dual network collaborative segmentation model according to claim 1, characterized in that: The overall loss function for cross-network contrastive learning includes: in, Represents the overall loss function of the model for cross-network contrastive learning, Represents the segmentation network f θ1 to f θ2 The pixel-level contrast learning loss function, r() represents the similarity measurement function, z i 1 Represents the segmentation network f θ1 , index is the anchor feature vector of i, represents the anchor feature vector z i 1 The positive sample, z n 1 represents the anchor feature vector z i 1 Negative samples of Represents the segmentation network f θ2 to f θ1 The pixel-level contrastive learning loss function is z i 2 Represents the segmentation network f θ2 , index is the anchor feature vector of i, Anchor feature vector z i 2 The positive sample, z n 2 represents the anchor feature vector z i 2 Negative samples.
9. The image semantic segmentation method based on the dual network collaborative segmentation model according to claim 1, characterized in that: The overall loss function of the dual network collaborative segmentation model includes: in, represents the loss function for fully supervised training, represents the loss function for consistent learning of the model across networks in the label space, The overall loss function of the model for cross-network contrastive learning, λ represents the loss weight of the model's cross-network consistency learning in the label space, and μ represents the loss weight of the model for cross-network contrastive learning.
Citation Information
Patent Citations
Remote sensing image semantic segmentation model training method and device for contrast consistency learning
CN114299380A
Image semi-supervised semantic segmentation method based on conservative aggressive collaborative learning
CN114821053A