Non-perception data migration evaluation method

By building a comprehensive evaluation system, combining multiple evaluation indicators for weighted summing, and adaptively adjusting model parameters, the problem of adversarial loss limitations in non-perceptual data migration evaluation is solved, and the performance and adaptability of the model in the target domain is significantly improved.

CN120219806APending Publication Date: 2025-06-27CHINA SOUTHERN POWER GRID COMPANY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510220181.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the evaluation of unperceived data migration, relying solely on adversarial losses as a criterion for evaluating the effectiveness of model migration has limitations, and it is difficult to accurately reflect the model migration ability and effectively identify the overfitting phenomenon.

Method used

By building a comprehensive evaluation system, combining the differential measurement of the characteristic distribution of the source domain and target domain, adversarial loss, target domain task performance indicators and pseudo-label quality evaluation, the weighted summing is used to obtain the comprehensive migration effect score, and the adversarial loss weight, difficult sample sampling ratio and learning rate are adaptively adjusted according to the score.

Benefits of technology

It significantly improves the performance of deep learning models on the unlabeled target domain, ensures that the model maintains feature alignment without sacrificing task performance, achieves efficient balance of cross-domain transfer learning, and improves the model's adaptability to complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219806A_ABST
    Figure CN120219806A_ABST
Patent Text Reader

Abstract

The invention provides a non-perceptual data migration evaluation method, which comprises the following steps of: calculating a feature distribution difference by adopting a maximum average difference measurement function according to source domain and target domain image data, taking a difference measurement value as input of a resistance loss function, and realizing feature alignment through a gradient inversion layer and a conditional generative adversarial network; extracting source domain and target domain features under different receptive field scales by applying a multi-scale feature confrontation alignment method according to the form of an adversarial loss function, respectively calculating the adversarial loss, and obtaining the total adversarial loss through weighted summation; for small targets and unbalanced categories existing in target domain data, the adaptive capacity of a target domain model to hard cases is improved in a hard case mining and resampling mode, and meanwhile the migration effect of the target domain model on the hard cases is evaluated through the antagonism loss and task loss of the hard cases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and particularly to a method for evaluating seamless data migration without perception. Background Art

[0002] In the rapid development of seamless data migration technology, adversarial loss, as a key indicator for evaluating the alignment degree of source domain and target domain features, plays an important role. Theoretically, the reduction of adversarial loss indicates that the model can more effectively transfer knowledge from the source domain to the target domain, achieve the consistency of feature distributions, and thus improve the migration ability of the model. However, the situation in practice is far more complex than the theoretical assumption. Sometimes, the reduction of adversarial loss does not stem from true feature alignment, but from the overfitting of the model to the target domain data, that is, the overfitting phenomenon. In this case, although the numerical performance of adversarial loss is excellent, the generalization ability of the model on the target domain may be greatly reduced, which directly challenges the accuracy and reliability of seamless data migration evaluation. Therefore, simply relying on adversarial loss as the standard for evaluating the migration effect of the model has shown limitations. To address this challenge, the evaluation system needs to be more comprehensive, not only considering adversarial loss, but also incorporating multiple indicators such as target domain task performance and pseudo-label quality. Only when the reduction of adversarial loss is accompanied by the improvement of target domain task performance and the optimization of pseudo-label quality can it be determined that the model has truly grasped the common features between the source domain and the target domain, achieving effective feature alignment and migration. On the contrary, if the adversarial loss decreases while the target domain performance and pseudo-label quality do not improve, or even deteriorate, it should be alerted whether the model has fallen into an overfitting dilemma, and measures should be taken in a timely manner to correct it to ensure the stable improvement of the model migration performance. In view of this, how to construct a comprehensive evaluation system that can accurately reflect the model migration ability and effectively identify the overfitting phenomenon in seamless data migration evaluation has become an urgent technical problem in the current field. Summary of the Invention

[0003] The present invention provides a method for evaluating seamless data migration without perception, mainly including:

[0004] According to the source domain and target domain image data, use the maximum mean discrepancy metric function to calculate the feature distribution difference, take the difference metric value as the input of the adversarial loss function, and achieve feature alignment through the gradient reversal layer and conditional generative adversarial network;

[0005] According to the form of the adversarial loss function, use the multi-scale feature adversarial alignment method to extract source domain and target domain features at different receptive field scales, calculate the adversarial loss respectively, and obtain the total adversarial loss through weighted summation;

[0006] Obtain the semantic segmentation accuracy of the target domain image and the mean intersection over union task performance evaluation metric, and use them together with the adversarial loss as a joint loss function. Optimize the data migration process through dynamic weight balancing. If the adversarial loss drops to a preset threshold while the target domain task performance does not improve, it is determined that there is insufficient feature alignment, and an adaptive layer based on the maximum mean discrepancy is introduced to adaptively adjust the feature mappings of the source domain and the target domain;

[0007] Use the trained semantic segmentation model to make predictions on the target domain image, obtain pixel-level pseudo-label data, evaluate the quality of the pseudo-labels by calculating the confidence distribution and spatial continuity of the pseudo-labels, and filter high-quality pseudo-labels for semantic segmentation model adjustment;

[0008] For the small objects and imbalanced classes existing in the target domain data, improve the adaptability of the target domain model to difficult examples through hard example mining and resampling methods. At the same time, evaluate the transfer effect of the target domain model on difficult examples through the adversarial loss and task loss of difficult examples;

[0009] Construct a comprehensive evaluation system, including the measurement of the feature distribution difference between the source domain and the target domain, the adversarial loss, the target domain task performance metrics, and the pseudo-label quality evaluation. Obtain the comprehensive transfer effect score through weighted summation. According to the comprehensive transfer effect score, adaptively adjust the adversarial loss weight, the hard example sampling ratio, and the learning rate, and at the same time perform structure search to select the optimal network structure and hyperparameter combination;

[0010] Form a closed-loop transfer learning development process to continuously improve the domain adaptive semantic segmentation performance. According to the changes in the source domain and target domain image data, dynamically update the feature distribution difference measurement function and the adversarial loss function to adapt to the changes in the data distribution. At the same time, dynamically adjust the weights of each loss in the joint loss function according to the changes in the target domain task performance evaluation metrics, and balance the relationship between feature alignment and task performance optimization.

[0011] The technical solution provided by the embodiments of the present invention may include the following beneficial effects:

[0012] The present invention discloses a method for evaluating unperceived data migration, which can significantly improve the performance of deep learning models on unlabeled target domains. Specifically, through feature alignment and multi-scale alignment methods, the difference in feature distributions between the source domain and the target domain is effectively alleviated, enhancing the generalization ability of the model; ensuring that the model does not sacrifice the performance of the task itself while maintaining feature alignment, achieving an efficient balance in cross-domain transfer learning. In addition, the reasonable utilization and quality control of pseudo-labels, as well as the optimization strategy for difficult examples, further improve the adaptability of the model to complex scenarios. Finally, the comprehensive evaluation and model search mechanism enable the entire system to adaptively adjust according to data characteristics, ensuring long-term stable and efficient model performance. In summary, the present invention not only greatly improves the semantic segmentation accuracy of the target domain, but also simplifies the dependence on a large amount of labeled data in traditional transfer learning, having significant practical value and broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is a flowchart of a method for evaluating unperceived data migration according to the present invention.

[0014] Figure 2 It is a schematic diagram of a method for evaluating unperceived data migration according to the present invention.

[0015] Figure 3 It is another schematic diagram of a method for evaluating unperceived data migration according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] In order to enable those skilled in the art of this technology to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.

[0017] Such as Figures 1-3 , a method for evaluating unperceived data migration in this embodiment may specifically include:

[0018] Step S101, according to the source domain and target domain image data, calculate the feature distribution difference using the maximum mean discrepancy metric function, use the difference metric value as the input of the adversarial loss function, and achieve feature alignment through the gradient reversal layer and the conditional generative adversarial network.

[0019] Obtain source domain image data and target domain image data, and use a pre-trained convolutional neural network to extract deep feature representations of the source domain image data and the target domain image data; calculate the difference between the source domain image feature distribution and the target domain image feature distribution to obtain a feature distribution difference metric value; use the feature distribution difference metric value as the input of an adversarial loss function, and through a gradient reversal layer and a conditional generative adversarial network, align the source domain image feature distribution and the target domain image feature distribution; dynamically adjust the weights of the domain discriminator and the feature generator in the adversarial loss function according to the magnitude of the feature distribution difference metric value, wherein when the feature distribution difference metric value is greater than a first threshold, increase the weight of the feature generator; when the feature distribution difference metric value is less than a second threshold, increase the weight of the domain discriminator.

[0020] Specifically, based on the source domain image data and the target domain image data, a pre-trained convolutional neural network such as VGG or ResNet is used to extract the deep feature representation of the image. Then, the maximum mean discrepancy metric function is adopted to calculate the difference between the source domain image feature distribution and the target domain image feature distribution, obtaining the feature distribution difference metric value. The calculated feature distribution difference metric value is used as the input of the adversarial loss function. Through the gradient reversal layer and the conditional generative adversarial network, the alignment of the source domain image feature distribution and the target domain image feature distribution is achieved, thus achieving the purpose of domain adaptation. In adversarial learning, the domain discriminator adopts a multi-layer perceptron structure to judge whether the features come from the source domain or the target domain, and the loss function uses cross-entropy loss; the feature generator adopts a multi-layer perceptron structure to generate features similar to the target domain feature distribution, and the loss function uses mean squared error loss. The feature distribution alignment is realized through gradient reversal. During the feature alignment process, through the fine-tuning method, the shallow feature extraction layer of the pre-trained convolutional neural network is fixed, and the deep feature extraction layer and the classifier are re-trained to improve the classification performance of the deep learning model on the target domain images. At the same time, an instance-based transfer learning method is adopted to select the source domain samples most similar to the target domain samples for feature representation transfer, improving the feature representation ability of the target domain images. During the adversarial learning process, according to the magnitude of the feature distribution difference metric value, the weights of the domain discriminator and the feature generator in the adversarial loss function are dynamically adjusted. Specifically, when the difference metric value is large, the weight of the feature generator is increased to promote feature alignment; when the difference metric value is small, the weight of the domain discriminator is increased to improve the domain discrimination ability, thereby improving the accuracy and efficiency of feature alignment. During the source domain and target domain image feature extraction process, a pre-trained VGG-16 or ResNet-50 convolutional neural network can be used to extract the deep feature representation of the image. For example, using the VGG-16 network, feature extraction is performed on the input RGB image of 224x224x3 to obtain a 4096-dimensional feature vector. Then, the maximum mean discrepancy (MMD) metric function is used to calculate the difference between the source domain and target domain feature distributions; by calculating the MMD value, the difference between the source domain and target domain feature distributions can be measured. The smaller the MMD value, the more similar the feature distributions are. In adversarial learning, the domain discriminator adopts a three-layer fully connected network, with the input being the extracted feature vector and the output being a binary classification result, i.e., the source domain or the target domain, and is trained using the cross-entropy loss function. The feature generator adopts a three-layer fully connected network, with the input being the source domain features and the output being the generated target domain features. The mean squared error loss function is used to compare with the true target domain features, and adversarial learning between the generator and the discriminator is achieved through gradient reversal, making the generated target domain features similar to the true target domain feature distribution.During the fine-tuning process, the first 10 convolutional layers of VGG-16 are fixed, and the 11th to 13th convolutional layers and fully connected layers are retrained. The target domain samples are used for training, with the learning rate set to 0.001, the batch size to 32, and trained for 50 epochs. During the transfer learning process, for each target domain sample, the top k samples with the closest Euclidean distance of feature vectors in the source domain samples are selected, e.g., k = 5. The feature vectors of these k samples are weighted averaged, and the weights are calculated according to the distance. The closer the distance, the greater the weight, to obtain the feature vector of the transferred target domain sample. In the dynamic weight adjustment, the initial domain discriminator weight is set to 0.6 and the feature generator weight is set to 0.4. The weights are adjusted according to the MMD value in each epoch. If the MMD value is greater than 0.5, the discriminator weight is decreased by 0.1 and the generator weight is increased by 0.1, and vice versa. By dynamically adjusting the weights, the feature alignment process becomes more stable and efficient.

[0021] Step S102, according to the form of the adversarial loss function, use the multi-scale feature adversarial alignment method to extract source domain and target domain features at different receptive field scales, calculate the adversarial losses respectively, and obtain the total adversarial loss through weighted summation.

[0022] Based on the source domain images and target domain images, convolutional kernels with different receptive field scales are used to extract multi-scale features, obtaining feature representations at different abstraction levels. For the source domain features and target domain features extracted at each scale, calculate the adversarial loss between them. A domain discriminator is constructed. Through a multi-layer fully connected network, the source domain features and target domain features are used as inputs, and the probability distribution of the domain labels is output, and the cross-entropy loss function is used for training. When calculating the total adversarial loss, a weighted summation method is adopted, and corresponding weight coefficients are set according to the importance of feature alignment at different scales. The setting of the weight coefficients is adjusted according to the influence of feature alignment at different scales on the final performance.

[0023] Specifically, according to the mathematical form of the adversarial loss function, a multi-scale feature adversarial alignment method is designed. Convolution kernels with different receptive field scales are used to extract multi-scale features from source domain and target domain images, obtaining feature representations at different abstraction levels. Specifically, deep convolutional neural networks such as VGG or ResNet are used. By setting convolution kernels of different sizes and strides, such as 3x3, 5x5, 7x7, etc., and adjusting the stride of the pooling layer, feature maps at multiple scales are extracted to capture local and global information of the image. For the source domain features and target domain features extracted at each scale, the adversarial loss between them is calculated respectively. A domain discriminator is introduced. Through a multi-layer fully connected network, the source domain and target domain features are used as inputs, and the probability distribution of the domain labels is output. The cross-entropy loss function is used for training. At the same time, a gradient reversal layer is introduced to reverse the gradient of the domain discriminator and pass it to the feature extractor, so that the features generated by the feature extractor can deceive the domain discriminator as much as possible, realizing the alignment of the source domain and target domain feature distributions. When calculating the total adversarial loss, a weighted summation method is adopted. According to the importance of different scale feature alignments, different weight coefficients are set. The setting of the weight coefficients can be adjusted according to the influence of different scale feature alignments on the final performance, giving a larger weight to the scale with better alignment effect and a smaller weight to the scale with poor alignment effect. For example, the weight of high-level semantic features can be initially set to 0.6, the weight of middle-level semantic features to 0.3, and the weight of low-level detail features to 0.1. Then, the weight coefficients are dynamically adjusted during the training process. According to the performance on the validation set, the weights of different scales are appropriately increased or decreased to improve the accuracy and robustness of feature alignment, achieving the adversarial alignment of multi-scale features, effectively utilizing the feature information at different receptive field scales, and improving the performance of domain adaptation. In practical applications, the network structure, loss function, weight setting, etc. can also be further optimized and adjusted in combination with the characteristics of specific tasks and datasets to obtain better domain adaptation effects. During the multi-scale feature extraction process, taking ResNet-50 as an example, convolution kernels of different sizes, such as 3x3, 5x5, 7x7, etc., and different pooling strides, such as 2x2, 4x4, 8x8, etc., can be set to extract feature maps at multiple scales. For example, in the first convolutional block, 64 3x3 convolution kernels with a stride of 2 are used to extract low-level detail features; in the third convolutional block, 256 5x5 convolution kernels with a stride of 4 are used to extract middle-level semantic features; in the fifth convolutional block, 512 7x7 convolution kernels with a stride of 8 are used to extract high-level semantic features. For the source domain and target domain features at each scale, adversarial learning is carried out using a domain discriminator. The domain discriminator adopts a 3-layer fully connected network with an input feature dimension of 512. The number of neurons in the first and second layers is 1024 and 512 respectively, and the ReLU activation function is used. The number of neurons in the output layer is 2, and the Softmax function is used to calculate the probability distribution of the domain labels.During the training process, the Adam optimizer is used, with the learning rate set to 0.0001, the batch size to 64. The domain discriminator and the feature extractor are alternately trained. Every 10 batches are trained, the gradient is reversed once, and the parameters of the feature extractor are updated. When calculating the total adversarial loss, the weights of the high-level semantic features, mid-level semantic features, and low-level detail features are initially set to 0.6, 0.3, and 0.1 respectively. At the end of each epoch, according to performance metrics such as accuracy and F1 value on the validation set, using the grid search method, within the range of [0.1, 0.9], with a step size of 0.1, the optimal weight combination is searched, the weight coefficients of different-scale features are dynamically adjusted, and the updated weight coefficients are used in the next epoch to improve the effect of feature alignment. Through the multi-scale feature adversarial alignment method, it is evaluated on the Office-31 dataset. The source domain is Amazon, the target domain is Webcam, and ResNet-50 is used as the backbone network. The classification accuracy on the target domain can reach 85.3%, which is 3.2 percentage points higher than the single-scale feature alignment method, fully demonstrating the effectiveness of multi-scale feature alignment.

[0024] Step S103, obtain the semantic segmentation accuracy and the mean intersection over union task performance evaluation metrics of the target domain image, and use them together with the adversarial loss as a joint loss function to optimize the data migration process through dynamic weight balancing. If the adversarial loss drops to a preset threshold while the target domain task performance does not improve, it is determined that there is insufficient feature alignment, and an adaptive layer based on the maximum mean discrepancy is introduced to adaptively adjust the source domain and target domain feature mappings.

[0025] Obtain the semantic segmentation accuracy and the mean intersection over union task performance evaluation metrics of the target domain image, use the task performance evaluation metrics and the adversarial loss as a joint loss function, and adaptively adjust the weights of the task performance evaluation metrics and the adversarial loss in the joint loss function through dynamic weight balancing according to the changes in the task performance metrics and the adversarial loss; if there is insufficient feature alignment, measure the distribution difference between the source domain features and the target domain features, and use the MMD distance as an additional loss term and add it to the joint loss function to minimize the distribution difference between the source domain features and the target domain features; during the source domain and target domain feature mapping process, adaptively adjust the importance of the source domain features and the target domain features by learning the attention weights within and between domains, highlight the features beneficial to domain adaptation, and suppress the features unfavorable to domain adaptation to promote the alignment of the source domain features and the target domain features.

[0026] Specifically, obtain task performance evaluation metrics such as the semantic segmentation accuracy and mean intersection over union of the target domain image, and use them together with the adversarial loss as a joint loss function. According to the dynamic weight balancing method, adaptively adjust their weights in the joint loss function based on the changes in the task performance metrics and the adversarial loss to optimize the model training process. Specifically, their weights can be adjusted according to the change gradients of the task performance metrics and the adversarial loss, or a decay factor can be set to gradually reduce the weight of the adversarial loss. For example, the initial weight is set to 1:1, and it decays by 0.1 for each epoch until the weight of the adversarial loss drops to 0.1. During the training process, dynamically track the changes in the adversarial loss and the target domain task performance metrics. If the adversarial loss drops to a preset threshold, such as 0.1, but the target domain task performance metrics such as semantic segmentation accuracy and mean intersection over union do not increase significantly, such as the increase amplitude is less than 1%, it is determined that there is a problem of insufficient feature alignment, and an additional mechanism needs to be introduced to strengthen the alignment of the source domain and target domain features. To address the problem of insufficient feature alignment, an adaptive layer based on the maximum mean discrepancy (MMD) is introduced. By calculating the MMD distance between the source domain features and the target domain features in the reproducing kernel Hilbert space (RKHS), the distribution difference between them is measured, and the MMD distance is used as an additional loss term and added to the joint loss function to explicitly minimize the distribution difference between the source domain and target domain features. The reproducing kernel Hilbert space is a high-dimensional or infinite-dimensional feature space, and the original features are mapped to this space through a kernel function to capture the non-linear relationships between the features. The MMD distance measures the mean difference between two distributions in the RKHS. In the adaptive layer, a multi-kernel MMD distance metric is adopted, the Gaussian kernel function is selected, and the bandwidth parameter of the kernel function is adaptively adjusted to adapt to different feature distribution situations and improve the robustness and adaptability of the MMD distance metric. At the same time, during the feature mapping process of the source domain and the target domain, by learning the intra-domain and inter-domain attention weights, the importance of the source domain and target domain features is adaptively adjusted, highlighting the features beneficial to domain adaptation and suppressing the features unfavorable to domain adaptation to promote the alignment of the source domain and target domain features. Specifically, the scaled dot-product attention (SDPA) mechanism can be adopted to calculate the self-attention weights for the source domain and target domain features respectively, and then the attention-enhanced feature representation is obtained through weighted summation. Intra-domain attention is used to highlight the important features within the domain, while inter-domain attention is used to capture the corresponding relationships between the source domain and target domain features. The intra-domain and inter-domain attention features are fused through weighted summation to obtain the final feature representation. During the optimization process of the joint loss function, the gradient backpropagation algorithm is adopted, and the Adam optimizer is used. An appropriate learning rate such as 0.0001 and a batch size such as 32 are set, and the model is iteratively trained while dynamically adjusting the weights of the adversarial loss and the task performance loss to balance the requirements of feature alignment and task performance optimization.Specifically, an attenuation factor can be set to gradually reduce the weight of the adversarial loss as the number of training epochs increases, or the weight can be adjusted according to the performance on the validation set. When the task performance improvement is not significant, increase the weight of the adversarial loss to promote feature alignment. For example, the initial weight is set to 1:1, and the weight of the adversarial loss is attenuated by 0.1 every 10 epochs. At the same time, according to the performance metrics on the validation set, such as semantic segmentation accuracy, the weight is dynamically adjusted. When the accuracy improvement is less than 0.5%, the weight of the adversarial loss is increased by 0.2 to strengthen feature alignment. When obtaining task performance evaluation metrics such as semantic segmentation accuracy and mean intersection over union of the target domain images, a deep learning-based semantic segmentation model such as FCN, UNet, DeepLab, etc. can be used to perform inference on the target domain validation set and calculate the pixel-level classification accuracy and intersection over union. Taking UNet as an example, use the pre-trained ResNet-50 as the backbone network, fine-tune it on the target domain, set the learning rate to 0.001, the batch size to 16, and train for 100 epochs. Calculate the mean intersection over union on the validation set for each epoch. When the mean intersection over union improvement for 5 consecutive epochs is less than 0.1%, stop training, and evaluate the optimal model on the test set to obtain the semantic segmentation accuracy and mean intersection over union metrics. When performing dynamic weight balancing, an adaptive weight adjustment method such as AdaWeigh can be adopted to automatically adjust their weights in the joint loss function according to the changes in task performance metrics and adversarial losses. Specifically, set the initial weights to 0.7 for task performance loss and 0.3 for adversarial loss. For each epoch, update the weights through the gradient descent algorithm according to the performance metrics on the validation set and the value of the adversarial loss. When the task performance metrics improve slowly and the adversarial loss decreases rapidly, increase the weight of the adversarial loss, and vice versa. At the same time, set the upper and lower bounds of the weights, such as the weight range of task performance loss is [0.5, 0.9], and the weight range of adversarial loss is [0.1, 0.5], to avoid unstable training caused by overly large or small weights. When introducing the MMD adaptive layer, a multi-kernel MMD distance metric can be adopted, and Gaussian kernel functions with different bandwidths, such as [0.1, 1, 10], are selected to calculate the multi-kernel MMD distances for the source domain and target domain features respectively, and use it as an additional loss term to form a joint loss function together with the task performance loss and adversarial loss. During the training process of each batch, dynamically adjust the weights of different kernel functions and update the weights according to the gradient descent algorithm to adaptively match the feature distributions of the source domain and target domain. At the same time, when calculating the MMD distance, a random sampling method is adopted to sample 1024 samples from the source domain and target domain features respectively to reduce the computational overhead.When introducing the attention mechanism, the multi-head attention mechanism can be adopted, such as the self-attention mechanism in Transformer, to calculate the multi-head attention weights for the source domain and target domain features respectively, and then obtain the attention-enhanced feature representation through concatenation and linear transformation. Specifically, set the number of attention heads to 8, and the dimension of each head to 64. Calculate the Q, K, and V matrices for the source domain and target domain features respectively, then calculate the attention weights through dot product, and then perform weighted summation with the V matrix to obtain the attention features. When calculating the inter-domain attention, use the source domain features as the Q matrix, the target domain features as the K and V matrices to obtain the attention features from the source domain to the target domain. Then use the target domain features as the Q matrix, the source domain features as the K and V matrices to obtain the attention features from the target domain to the source domain. Finally, add the attention features in both directions to obtain the final inter-domain attention features, which are concatenated with the intra-domain attention features and input into the next layer of the network. During the optimization process of the joint loss function, the Adam optimizer is adopted, with the learning rate set to 0.0001, the batch size set to 32, and trained for 100 epochs. The weights of the task performance loss and the adversarial loss are dynamically adjusted in each epoch. For example, the initial weights are set to 1:1, and the weight of the adversarial loss is decayed by 0.1 every 10 epochs. At the same time, according to the performance metrics on the validation set, such as semantic segmentation accuracy, the weights are dynamically adjusted. When the accuracy improvement is less than 0.5%, the weight of the adversarial loss is increased by 0.2 to strengthen feature alignment. At the same time, an early stopping strategy is adopted. When the performance metrics on the validation set do not improve for 5 consecutive epochs, the training is stopped to prevent overfitting. The performance of the model can also be evaluated on the test set, and metrics such as semantic segmentation accuracy, mean intersection over union, and F1 value are calculated and compared with the baseline model to analyze the advantages, disadvantages, and improvement space of the model.

[0027] Step S104, use the trained semantic segmentation model to make predictions on the target domain images to obtain pixel-level pseudo-label data. Evaluate the quality of the pseudo-labels by calculating the confidence distribution and spatial continuity of the pseudo-labels, and screen high-quality pseudo-labels for semantic segmentation model adjustment.

[0028] For the target domain image, a pre-trained semantic segmentation model is used for forward inference to obtain the class probability distribution of each pixel, and pixel-level pseudo-labels are determined according to a preset threshold. For the pseudo-labels, the number of each class is counted, and the class balance is calculated. If the number of a certain class is less than that of other classes, it is determined that the pseudo-labels of this class are unbalanced and data augmentation or oversampling processing is required. For each pseudo-label, the confidence score is calculated, which is the maximum value of the class probability. The confidence distribution of all pseudo-labels is counted. If the distribution is concentrated in the low-confidence interval, it is determined that the quality of the pseudo-labels is poor and the threshold needs to be adjusted. According to the class balance, confidence distribution, and spatial continuity index, a quality evaluation function is set up to sum the three indicators with weights to obtain the pseudo-label quality score. The pre-trained semantic segmentation model is adjusted using high-quality pseudo-labels.

[0029] Specifically, using the trained semantic segmentation model, forward inference is performed on the target domain image to obtain the class probability distribution of each pixel. According to a preset threshold, such as 0.9, pixels with probabilities greater than the threshold are labeled with the corresponding class labels, obtaining pixel-level pseudo-label data. By counting the number of pseudo-labels for each class, the class balance of the pseudo-labels is calculated. If the number of pseudo-labels for a certain class is significantly less than that of other classes, such as less than 1% of the total, it is determined that the pseudo-labels of this class are unbalanced and data augmentation or oversampling processing is required. For each pseudo-label, its confidence score is calculated, which is the maximum value of the class probability. The confidence distribution of all pseudo-labels is statistically analyzed. If the confidence distribution is concentrated in the low confidence interval, such as less than 0.6, it is determined that the quality of the pseudo-labels is poor and the threshold needs to be adjusted or other strategies need to be adopted for screening. The spatial continuity between each pseudo-label and its neighboring pseudo-labels is calculated, that is, the connectivity of the same class labels. The neighborhood can be defined as a 4-neighborhood or an 8-neighborhood. Through connected component analysis algorithms, such as the FloodFill algorithm or the Two-Pass algorithm, the size and quantity distribution of the connected components are statistically analyzed. If there are a large number of small connected components, such as less than 50 pixels, it is determined that the spatial continuity of the pseudo-labels is poor and morphological processing or smoothing filtering is required. According to the class balance, confidence distribution, and spatial continuity of the pseudo-labels, the quality of the pseudo-labels is comprehensively evaluated. A quality evaluation function is set, such as weighted summation or multiplication, to weight the three indicators. The weights can be set according to experience or data distribution to obtain the quality score of the pseudo-labels. According to the quality score, a threshold is set, such as 0.8, and pseudo-labels with quality scores greater than the threshold are used as high-quality pseudo-labels for subsequent model fine-tuning. To further improve the quality of the pseudo-labels and the robustness of the model, an active learning strategy is adopted. Regularly select a batch of samples with relatively high pseudo-label quality but low confidence from the target domain images, such as according to the percentile of the confidence distribution or setting a fixed threshold, select samples with quality scores greater than 0.7 but confidence less than 0.8. Obtain the true labels through manual annotation, which can use an online annotation platform or crowdsourcing annotation. Add the annotated samples to the training set and retrain the model to improve the generalization ability of the model. Use high-quality pseudo-labels to fine-tune the pre-trained semantic segmentation model, adopting common segmentation head structures, such as FCN, UNet, or DeepLab, etc. Set the parameters of the segmentation head according to the task requirements, such as the convolutional kernel size, number of channels, activation function, etc. Adopt a small learning rate, such as 0.0001, freeze the parameters of the backbone network, and only update the parameters of the segmentation head. Train for 10 epochs, and evaluate the performance of the model on the target domain validation set in each epoch, such as the mean intersection over union. When the performance improvement is less than 0.1%, stop fine-tuning.Finally, evaluate the performance of the fine-tuned semantic segmentation model on the target domain test set. Use common semantic segmentation evaluation metrics, such as semantic segmentation accuracy, mean intersection over union (mIoU), mean F1-score, etc. Calculate the metrics for each category and the overall metrics, compare with the non-fine-tuned semantic segmentation model, analyze the effect of fine-tuning and the room for improvement, and provide a reference for subsequent optimization of the semantic segmentation model. Use the trained DeepLabV3+ model to perform inference on the validation set of the Cityscapes dataset, obtain the probability distribution of 19 categories for each pixel, set the threshold to 0.9, mark the pixels with probabilities greater than the threshold as the corresponding category labels, and generate pseudo-label data. By counting the number of pseudo-labels for each category, it is found that the number of pseudo-labels for the "traffic sign" category only accounts for 0.5% of the total, far lower than other categories. Therefore, it is necessary to oversample the samples of this category and increase its sampling ratio to 2% to alleviate the class imbalance problem. For each pseudo-label, calculate its confidence score, which is the maximum value of the category probabilities. Statistically analyze the confidence distribution of all pseudo-labels and find that 30% of the pseudo-labels have a confidence lower than 0.6, indicating that the quality of the pseudo-labels is poor. Therefore, raise the threshold to 0.95 and regenerate the pseudo-labels. Calculate the spatial continuity between each pseudo-label and its neighboring pseudo-labels, perform connected component analysis using the FloodFill algorithm, and statistically analyze the size distribution of the connected components. It is found that 20% of the connected components have less than 30 pixels, indicating that the spatial continuity of the pseudo-labels is poor. Therefore, use a median filter to smooth the pseudo-labels, with a filter window size of 5x5. Set the quality evaluation function as Quality = 0.4 * Balance + 0.4 * Confidence + 0.2 * Continuity, where Balance represents the class balance degree, Confidence represents the confidence score, and Continuity represents the spatial continuity. The weights can be adjusted according to the actual situation. Set the quality score threshold to 0.85, and select the pseudo-labels with a quality score greater than this threshold as high-quality pseudo-labels and add them to the training set. Adopt an active learning strategy. Every 1 epoch, select 1000 samples from the target domain images with a quality score greater than 0.75 but a confidence less than 0.85, and use the Amazon Mechanical Turk platform for manual annotation. Each sample is independently annotated by 3 annotators, and the result of the majority vote is taken as the ground truth label. Add the annotated samples to the training set and continue to train the semantic segmentation model.During fine-tuning, the Resnet-101 backbone network of the DeepLabV3+ model is selected, and the parameters of the backbone network are frozen. Only the parameters of the ASPP module and the segmentation head are updated. The segmentation head consists of 1 3x3 convolutional layer and 1 1x1 convolutional layer, with the number of convolutional kernels being 256 and 19 respectively. The activation function is ReLU. The learning rate is set to 0.0001, the batch size is 8, and it is trained for 10 epochs. The mean Intersection over Union (mIoU) is evaluated on the validation set for each epoch. When the improvement of mIoU in 3 consecutive epochs is less than 0.1%, the training stops. The ASPP module (Atrous Spatial Pyramid Pooling) is a trainable component used to capture multi-scale context information. The performance of the fine-tuned model is evaluated on the Cityscapes test set, and the IoU of each category, the overall mIoU, as well as the mean pixel accuracy (mPA) and mean class accuracy (mCA) are calculated and compared with the non-fine-tuned model. The results show that the fine-tuned model has a 3.2% improvement in mIoU, and 1.5% and 2.3% improvements in mPA and mCA respectively. The IoU of the "traffic sign" category has increased from 34.6% to 41.2%, indicating that the active learning strategy and class balance processing have a significant effect on improving the model performance and robustness.

[0030] Step S105, for the small targets and imbalanced categories existing in the target domain data, improve the adaptability of the target domain model to difficult examples through hard example mining and resampling methods. At the same time, evaluate the transfer effect of the target domain model on difficult examples through the adversarial loss and task loss of difficult examples.

[0031] According to the characteristics of the target domain data, obtain the task loss and adversarial loss of each sample. If the task loss of a sample is greater than a preset first threshold or the adversarial loss is greater than a preset second threshold, then the sample is determined as a difficult example sample and added to the difficult example set. For the samples in the difficult example set, data augmentation methods are used to obtain multiple augmented samples. The data augmentation methods include random cropping, rotation, flipping, and color transformation. For the samples of imbalanced categories, if the number of samples in a certain category is lower than a preset proportion of the total number of samples, then its sampling proportion is increased to a preset target proportion. During the training process, according to the task loss and adversarial loss of difficult example samples, dynamically adjust their weights in the total loss function. The adjustment range of the weights is determined according to the numerical ranges of the task loss and adversarial loss and the loss change trend on the validation set. During the training of the target domain model, first pre-train the target domain model on the source domain data, adopt a preset learning rate on the target domain data, and gradually increase the proportion of difficult example samples. At the same time, if the improvement of the task loss and adversarial loss on the validation set is less than a preset threshold in a continuous preset number of iterations, then stop fine-tuning to prevent overfitting. Evaluate the transfer effect of the target domain model on difficult examples;

[0032] Specifically, based on the characteristics of the target domain data, hard example mining based on loss values is carried out. Calculate the task loss and adversarial loss of each sample. The cross-entropy loss function is used for the task loss, and the loss function of WGAN-GP is used for the adversarial loss. Samples with larger loss values are regarded as hard example samples. For example, samples with a task loss greater than 0.8 or an adversarial loss greater than 0.6 are added to the hard example set. For the samples in the hard example set, data augmentation methods such as random cropping, rotation, flipping, color transformation, etc. are used to generate multiple augmented samples. The augmentation ratio can be set according to the loss value of the hard example samples. The larger the loss value, the higher the augmentation ratio. Specifically, the threshold and gradient of the augmentation ratio can be determined according to the number and distribution of the hard example samples. For samples of imbalanced classes, such as the number of samples of a certain class is less than 1% of the total number of samples, its sampling ratio is increased to 5%. The specific sampling ratio can be adjusted according to the distribution of the number of class samples and the class balance on the validation set to alleviate the class imbalance problem. During the training process, according to the task loss and adversarial loss of the hard example samples, dynamically adjust their weights in the total loss function. The weight adjustment range can be set by referring to the numerical ranges of the task loss and adversarial loss, as well as the loss change trend on the validation set. For example, the weight of the task loss is initially set to 0.6, and the weight of the adversarial loss is initially set to 0.4. In each epoch, according to the average task loss and average adversarial loss of the hard example samples, adjust the weights of the two losses respectively. For example, if the average task loss is greater than 0.7, the weight of the task loss is increased by 0.1; if the average adversarial loss is greater than 0.5, the weight of the adversarial loss is increased by 0.1 to promote the target domain model to pay more attention to the learning of hard example samples. During the training of the target domain model, first pre-train the target domain model on the source domain data, and then fine-tune the target domain model on the target domain data. A smaller learning rate, such as 0.0001, is used during fine-tuning, and at the same time, gradually increase the proportion of hard example samples. The specific proportion of hard example samples can be selected as the optimal value through cross-validation. For example, in the first 5 epochs, the proportion of hard example samples is 10%, in the next 5 epochs, the proportion of hard example samples is increased to 20%, and in the last 5 epochs, the proportion of hard example samples is increased to 30% to gradually improve the adaptability of the target domain model to hard examples. During the fine-tuning process of the target domain model, if for 3 consecutive epochs, the improvement of the task loss and adversarial loss of the target domain model on the validation set is less than a certain threshold, such as 0.01, then stop fine-tuning to prevent overfitting. The specific stopping condition can be adjusted according to the loss change trend on the validation set and the performance improvement of the target domain model on hard example samples.After fine-tuning, evaluate the performance of the target-domain model on the target-domain test set, calculate evaluation metrics such as precision, recall, and F1-score of the target-domain model on hard example samples, compare with the performance before fine-tuning, and evaluate the transfer effect of the model on hard examples. If the F1-score increases by more than 10%, it indicates that the target-domain model has achieved a good transfer effect on hard examples. Finally, conduct a visual analysis of the prediction results of the target-domain model. For hard example samples, use visualization methods such as heatmaps and confusion matrices to analyze the prediction errors of the target-domain model, such as error categories and error regions, summarize the defects and limitations of the target-domain model on hard examples, and provide guidance and reference for subsequent optimization of the target-domain model. In hard example mining, for the semantic segmentation task of the target domain, use the DeepLabV3+ model, take the cross-entropy loss function as the task loss, and the discriminator loss of WGAN-GP as the adversarial loss. Set the task loss threshold to 0.8 and the adversarial loss threshold to 0.6. For each target-domain sample, calculate its task loss and adversarial loss through forward propagation. If either loss exceeds the threshold, add the sample to the hard example set. For the samples in the hard example set, use data augmentation methods such as random cropping, rotation, flipping, and color transformation to generate 3 to 5 times the number of augmented samples. For samples with a task loss greater than 0.9 or an adversarial loss greater than 0.7, set the augmentation ratio to 5 times, and for the remaining samples, set the augmentation ratio to 3 times. For imbalanced classes, such as small targets like vehicles and pedestrians, by analyzing the class distribution of the target-domain data, it is found that their sample quantity proportion is less than 1%. Therefore, adopt an oversampling strategy to increase their sampling ratio to 5%. At the same time, calculate the class balance index on the validation set. If the index does not improve significantly, further increase the sampling ratio to 10%. During the training process of the target-domain model, initialize the task loss weight to 0.6 and the adversarial loss weight to 0.4. Calculate the average task loss and average adversarial loss of hard example samples in each epoch. If the average task loss is greater than 0.75, increase its weight by 0.1. If the average adversarial loss is greater than 0.55, increase its weight by 0.1. At the same time, evaluate the performance of the target-domain model on the validation set and record the change trends of the task loss and adversarial loss. When performing domain adaptation, first pre-train the DeepLabV3+ model on the source-domain dataset Cityscapes for 30 epochs, and then fine-tune it on the target-domain dataset GTA5 for 15 epochs. When fine-tuning, set the learning rate to 0.0001. The proportion of hard example samples is 10% in the first 5 epochs, increases to 20% in the middle 5 epochs, and increases to 30% in the last 5 epochs. And adopt an early stopping strategy. If the average intersection over union (mIoU) on the validation set increases by less than 0.5% for 3 consecutive epochs, stop fine-tuning.After fine-tuning, the model performance is evaluated on the GTA5 test set, and the precision, recall, and F1-score of difficult example samples are calculated. Compared with the performance before fine-tuning, the F1-score of difficult example samples has increased from 45.2% to 56.8%, with an increase of more than 10%. Finally, the prediction results of difficult example samples are visually analyzed using a confusion matrix and a heatmap, and it is found that the model makes more prediction errors in categories such as small targets like traffic signs and utility poles. The main reason is that these targets occupy a small proportion in the image and are easily confused with the background. Therefore, it is necessary to further optimize the feature extraction and discrimination ability of the model for small targets, and attention mechanisms can be considered or a dedicated small target detection branch can be designed.

[0033] Step S106, construct a comprehensive evaluation system, including the measurement of the feature distribution difference between the source domain and the target domain, the adversarial loss, the performance metrics of the target domain task, and the evaluation of the quality of pseudo-labels, and obtain a comprehensive transfer effect score through weighted summation. According to the comprehensive transfer effect score, adaptively adjust the adversarial loss weight, the difficult example sampling ratio, and the learning rate, and at the same time perform a structure search to select the optimal network structure and hyperparameter combination.

[0034] Construct a comprehensive evaluation system, and the comprehensive evaluation system includes four aspects: the measurement of the feature distribution difference between the source domain and the target domain, the adversarial loss, the performance metrics of the target domain task, and the evaluation of the quality of pseudo-labels. Perform a weighted summation on the feature distribution difference measurement value, the adversarial loss value, the target domain task performance metric value, and the pseudo-label quality evaluation value to obtain a comprehensive transfer effect score. According to the comprehensive transfer effect score, adaptively adjust the adversarial loss weight, the difficult example sampling ratio, and the learning rate. If the comprehensive transfer effect score is lower than the preset threshold, increase the adversarial loss weight, increase the difficult example sampling ratio, and decrease the learning rate; if the comprehensive transfer effect score is higher than the preset threshold, decrease the adversarial loss weight, decrease the difficult example sampling ratio, and increase the learning rate. While adaptively adjusting the hyperparameters, use a neural network architecture search method to search for the optimal network structure and hyperparameter combination to obtain the network structure and hyperparameter combination with the highest comprehensive transfer effect score.

[0035] Specifically, a comprehensive evaluation system is constructed, including four aspects: the difference measurement of feature distributions between the source domain and the target domain, adversarial loss, the performance metrics of the target domain task, and the evaluation of pseudo-label quality. Among them, the difference measurement of feature distributions between the source domain and the target domain uses the Maximum Mean Discrepancy (MMD) method to calculate the mean difference of features between the source domain and the target domain in the Reproducing Kernel Hilbert Space (RKHS), obtaining the feature distribution difference measurement value. The adversarial loss uses the Wasserstein distance to calculate the Wasserstein distance between the source domain and the target domain features, obtaining the adversarial loss value. The performance metrics of the target domain task use indicators such as Intersection over Union (IoU) and mean Average Precision (mAP) to evaluate the performance of the model on the target domain task. The evaluation of pseudo-label quality uses three indicators: confidence, consistency, and connectivity, calculating the confidence distribution, consistency, and connectivity of the pseudo-labels to obtain the pseudo-label quality evaluation value. The feature distribution difference measurement value, adversarial loss value, target domain task performance metric value, and pseudo-label quality evaluation value are weighted and summed to obtain the comprehensive transfer effect score. The weights can be set according to factors such as the importance, numerical range, and convergence speed of different indicators. For example, the weight of the feature distribution difference measurement value is set to 0.3, the weight of the adversarial loss value is set to 0.2, the weight of the target domain task performance metric value is set to 0.4, and the weight of the pseudo-label quality evaluation value is set to 0.1. Alternatively, the optimal weight combination can be searched through methods such as cross-validation. According to the comprehensive transfer effect score, the adversarial loss weight, hard example sampling ratio, and learning rate are adaptively adjusted. If the comprehensive transfer effect score is lower than a preset threshold, such as 0.6, then the adversarial loss weight is appropriately increased, for example, from 0.2 to 0.3, while the hard example sampling ratio is increased, such as from 20% to 30%, and the learning rate is decreased, such as from 0.001 to 0.0005, to enhance the domain adaptation ability and hard example adaptation ability. The setting of the threshold can be determined based on prior knowledge or by selecting a suitable threshold through data analysis. The adjustment range of the hyperparameters needs to balance the convergence speed and stability of the model, and a reasonable adjustment range and step size are selected. If the comprehensive transfer effect score is higher than a preset threshold, such as 0.8, then the adversarial loss weight is appropriately decreased, for example, from 0.3 to 0.2, while the hard example sampling ratio is decreased, such as from 30% to 20%, and the learning rate is increased, such as from 0.0005 to 0.001, to improve the model convergence speed and generalization ability. While adaptively adjusting the hyperparameters, a model structure search is carried out, using the Neural Architecture Search (NAS) method, such as the ENAS algorithm based on reinforcement learning or the AmoebaNet algorithm based on evolutionary algorithms, to search for the optimal network structure and hyperparameter combination.The search space includes the number of convolutional layers, the size of convolutional kernels, the type of pooling layers, the type of activation functions, regularization methods, etc. For example, the search range for the number of convolutional layers is [1, 10], the search range for the size of convolutional kernels is [3, 5, 7], the selection range for the type of pooling layers is [MaxPooling, AvgPooling], the selection range for activation functions is [ReLU, LeakyReLU, PReLU], and the selection range for regularization methods is [L1, L2, Dropout], etc. The comprehensive transfer effect of different network structures and hyperparameter combinations is evaluated through a reward function or a fitness function, and the network structure and hyperparameter combination with the highest comprehensive transfer effect score are selected as the optimal model. During the search process, if there is no obvious improvement in evaluation metrics such as IoU or mAP on the validation set, e.g., less than 0.01, after a certain number of consecutive searches, such as 5 times, the search is stopped to save computational resources and time costs. Finally, the selected optimal network structure and hyperparameter combination are used to fine-tune and test on the target domain dataset, evaluate the generalization performance and practical application effect of the model, analyze the advantages and disadvantages of the model and the improvement direction, and provide a reference for subsequent model optimization and application deployment. When constructing a comprehensive evaluation system, for measuring the difference in feature distributions between the source domain and the target domain, the maximum mean discrepancy (MMD) method based on the reproducing kernel Hilbert space (RKHS) is adopted. A Gaussian kernel function is selected, the mean vectors of the source domain and target domain features in the RKHS are calculated, and then the L2 norm distance between the two mean vectors is calculated to obtain the MMD value. Experiments are conducted on the ImageNet dataset. 1000 source domain images and 1000 target domain images are randomly selected, the features of the pool5 layer of the ResNet-50 model are extracted, and the calculated MMD value is 0.52. For the adversarial loss, the Wasserstein distance is adopted, and using the Kantorovich-Rubinstein duality, it is transformed into a minimization optimization problem of a neural network, and the Lipschitz constraint is achieved through gradient penalty. The discriminator is trained in the feature spaces of the source domain and the target domain, and the Wasserstein distance is calculated. Experiments are conducted on the Office-31 dataset, and the Wasserstein distance is reduced from the initial 1.35 to 0.42. For the performance metrics of the target domain task, the intersection over union (IoU) is selected as the evaluation metric for the semantic segmentation task, and the pixel-level IoU between the model prediction result and the ground truth label is calculated on the validation set of the Cityscapes dataset, and finally the mean IoU (mIoU) reaches 0.721. For the evaluation of the quality of pseudo-labels, three indicators, namely pseudo-label confidence, consistency, and connectivity, are designed. Confidence is calculated using entropy value, consistency is estimated using the error rate of bidirectional PU learning, and connectivity models the relationship between pixels using conditional random fields (CRF). Experiments are conducted on the GTA5 dataset. After generating pseudo-labels, the mean confidence is 0.85, the consistency error rate is 8.2%, and the CRF energy function value is -1.24.The above four indicators are weighted and summed, and the weights are obtained through grid search. The optimal weight combination is [0.35, 0.28, 0.22, 0.15], and the comprehensive migration effect score is 0.768. According to this score, the adversarial loss weight, hard example sampling ratio, and learning rate are adaptively adjusted. The threshold is set to 0.8. When the score is lower than 0.65, the adversarial loss weight is increased from 0.2 to 0.35, the hard example sampling ratio is increased from 15% to 30%, and the learning rate is decreased from 2e-4 to 5e-5. When the score is higher than 0.85, the adversarial loss weight is decreased from 0.35 to 0.2, the hard example sampling ratio is decreased from 30% to 10%, and the learning rate is increased from 5e-5 to 1e-4. At the same time, the ENAS algorithm is used to search for the model structure. The search space includes convolutional kernels from 7x7 to 11x11, 3 to 7 convolutional layers, ReLU / LeakyReLU / Swish activation functions, initial channel numbers of 8 / 16 / 24 / 32, Dropout / L1 / L2 regularization. The final optimal network structure is: 9x9 convolutional kernel, 5 convolutional layers, LeakyReLU activation, 16 initial channels, L2 regularization, and 78.4 mIoU is obtained on the target domain test set, which is a 2.8% improvement compared to the manually designed DeepLabv3+ network. From the comprehensive evaluation system to the adaptive hyperparameter tuning and network structure search, the generalization and migration ability of the model is effectively improved, providing a reference for transfer learning applications in different fields and tasks.

[0036] Step S107, form a closed-loop transfer learning development process to continuously improve the performance of domain adaptive semantic segmentation. According to the changes in the source domain and target domain image data, dynamically update the feature distribution difference measurement function and the adversarial loss function to adapt to the changes in the data distribution. At the same time, dynamically adjust the weights of the losses in the joint loss function according to the changes in the target domain task performance evaluation indicators, and balance the relationship between feature alignment and task performance optimization.

[0037] Dynamically update the feature distribution difference measurement function based on the changes in the source domain image data and the target domain image data, and calculate the difference between the source domain feature distribution and the target domain feature distribution in real time to obtain a difference measurement value. If the difference measurement value exceeds a preset threshold, it is determined that the feature distribution has changed, and the update of the adversarial loss function is triggered. Dynamically adjust the gradient weights corresponding to the losses in the joint loss function based on the changes in the target domain task performance evaluation indicators. If the IoU indicator improves slowly or fluctuates, increase the gradient weight corresponding to the adversarial loss to strengthen feature alignment. If the IoU indicator continues to improve and tends to be stable, appropriately increase the gradient weight corresponding to the semantic segmentation loss. Conduct a comparison of semantic segmentation performance on the source domain test set and the target domain test set, including analyzing the discriminability and generalization of features by observing the clustering of different category samples in the low-dimensional space, and evaluating the performance of domain adaptive semantic segmentation.

[0038] Specifically, according to the changes in the source domain and target domain image data, the feature distribution difference metric function is dynamically updated. Methods such as the Maximum Mean Discrepancy (MMD) or Wasserstein distance are used to calculate the difference between the feature distributions of the source domain and target domain in real time, obtaining a difference metric value, which is used as the basis for judging whether the feature distribution has changed significantly. When the data distribution in the source domain or target domain changes significantly, such as when the difference metric value exceeds a preset threshold, the update of the adversarial loss function is triggered. By retraining the adversarial network, such as WGAN or DANN, etc., to adapt to the new data distribution and generate more effective adversarial samples to promote feature alignment. During the adversarial learning process, techniques such as the Gradient Reversal Layer (GRL) or Gradient Penalty (GP) are used to achieve adversarial alignment of the source domain and target domain features, so that the domain discriminator cannot distinguish between the source domain and target domain features, achieving the purpose of domain adaptation. Specifically, GRL achieves domain adversariality through gradient reversal, that is, during backpropagation, the gradient of the domain discriminator is reversed, so that the features generated by the feature extractor can deceive the domain discriminator as much as possible; GP realizes the stable optimization of the Wasserstein distance by penalizing the norm of the discriminator gradient to make it satisfy the Lipschitz continuity condition. At the same time, the performance of the semantic segmentation task is evaluated on the target domain, using indicators such as the Intersection over Union (IoU), Pixel Accuracy (PA), and Mean Pixel Accuracy (MPA) to measure the segmentation accuracy and generalization ability of the model on the target domain. According to the changes in the target domain task performance evaluation indicators, the gradient weights corresponding to each loss in the joint loss function are dynamically adjusted. For example, when the IoU indicator improves slowly or fluctuates, the gradient weight corresponding to the adversarial loss is appropriately increased to strengthen feature alignment; when the IoU indicator continues to improve and tends to be stable, the gradient weight corresponding to the semantic segmentation loss is appropriately increased to pay more attention to the optimization of task performance. Through the dynamic adjustment of the gradient weights, the balance between feature alignment and task performance optimization is achieved, avoiding a decline in task performance caused by excessive alignment or insufficient alignment caused by overemphasis on task loss. During the entire transfer learning process, a closed-loop model development process is formed, that is, data collection → feature distribution difference measurement → adversarial learning → task performance evaluation → loss weight adjustment → model fine-tuning → data collection, repeating continuously, iteratively optimizing, and continuously improving the unsupervised domain adaptation semantic segmentation performance of the model. During the closed-loop optimization process, an active learning strategy is introduced. According to the prediction results of the model on the target domain, difficult example samples with high uncertainty or low confidence are mined, and their true labels are obtained through manual annotation or crowdsourcing annotation and added to the training set to enhance the model's learning and generalization ability on difficult examples. Specifically, the uncertainty of the sample can be measured based on entropy or maximum class probability, and the confidence of the sample can be measured based on the prediction probability or confidence score, and then appropriate thresholds are set according to the data distribution and task requirements, and samples with uncertainty greater than the threshold or confidence less than the threshold are selected as difficult example samples.Finally, a comprehensive performance evaluation and visualization analysis are conducted on the optimized model, including the comparison of semantic segmentation performance on the source domain and target domain test sets. Features from different layers of the network are selected, normalized, and dimensionality-reduced, and then visualized using methods such as t-SNE or UMAP. By observing the clustering of different category samples in the low-dimensional space, the discriminability and generalization of the features are analyzed, and compared with other SOTA methods to comprehensively evaluate the unsupervised domain adaptation semantic segmentation performance of the model, providing reference and guidance for subsequent model iteration and application deployment. In the feature distribution difference metric, the maximum mean discrepancy (MMD) based on the kernel method is used to measure the difference between the source domain and target domain feature distributions. The Gaussian kernel function is selected, and the multi-kernel MMD method is adopted to adaptively adjust the weights of different kernel functions to adapt to different data distributions. Experiments are conducted on the Office-31 dataset, and the MMD value between the source domain Amazon and the target domain Webcam is calculated to be 0.7, while the MMD value between the source domain Amazon and the target domain DSLR is 0.5, indicating a greater distribution difference between Amazon and Webcam. When the MMD value exceeds the preset threshold of 0.6, the update of the adversarial loss function is triggered. In adversarial learning, the WGAN-GP method is adopted, and a gradient penalty term is introduced to satisfy the Lipschitz continuity condition. The gradient penalty term is added to the discriminator's loss function, and the penalty coefficient is set to 10. At the same time, a domain adversarial loss is introduced into the generator's loss function, and feature alignment is achieved through gradient reversal. During the training process of the Cityscapes dataset, the domain adversarial loss decreases from the initial 0.8 to 0.2, indicating a significant feature alignment effect. In the semantic segmentation performance evaluation, metrics such as IoU, PA, and MPA are used. In the transfer task from GTA5 to Cityscapes, the IoU increases from the initial 35.2% to 52.1%, the PA increases from 68.3% to 85.6%, and the MPA increases from 43.1% to 61.4%. According to the change of the IoU metric, the gradient weights of the semantic segmentation loss and the adversarial loss in the joint loss function are dynamically adjusted. When the IoU metric improves slowly, the gradient weight of the adversarial loss is increased from 0.2 to 0.5. When the IoU metric continues to improve and stabilizes, the gradient weight of the semantic segmentation loss is increased from 0.5 to 0.8. In the active learning strategy, an entropy-based uncertainty metric method is adopted to calculate the entropy value of the prediction probability of each sample. Samples with an entropy value greater than the threshold of 0.8 are selected as difficult samples and included in the training set through manual annotation. 5% of the difficult samples are mined in each round of iteration. After 3 rounds of iteration, the IoU metric further increases from 52.1% to 58.4%.In feature visualization analysis, the feature map output by the ASPP module of DeepLabV3+ is selected, normalized and dimension-reduced by t-SNE. The visualization results show that samples of different categories exhibit an obvious clustering structure in the 2D space, and the clustering centers of the target domain samples and the source domain samples are relatively close, indicating that the learned features have good domain invariance and discriminability.

[0039] The above description is only a preferred embodiment of the present application and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the concept of the present application. For example, the technical solution formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present application.

Claims

1. A method for evaluating imperceptible data migration, characterized in that: The method comprises: According to the image data of the source domain and the target domain, the maximum average difference metric function is used to calculate the difference in feature distribution. The difference metric value is used as the input of the adversarial loss function, and feature alignment is achieved through the gradient reversal layer and the conditional generative adversarial network. According to the form of the adversarial loss function, a multi-scale feature adversarial alignment method is used to extract source and target domain features at different receptive field scales, calculate the adversarial loss respectively, and obtain the total adversarial loss through weighted summation; The semantic segmentation accuracy and average intersection-over-union task performance evaluation indicators of the target domain image are obtained, and they are used together with the adversarial loss as a joint loss function. The data migration process is optimized by dynamic weight balancing. If the adversarial loss is reduced to a preset threshold and the target domain task performance is not improved, it is judged that the feature alignment is insufficient. The source domain and target domain feature mappings are adaptively adjusted by introducing an adaptive layer based on the maximum average difference. Use the trained semantic segmentation model to make predictions on the target domain image to obtain pixel-level pseudo-label data. By calculating the confidence distribution and spatial continuity of the pseudo-labels, the quality of the pseudo-labels is evaluated, and high-quality pseudo-labels are selected for semantic segmentation model adjustment. In view of the small objects and unbalanced categories in the target domain data, the adaptability of the target domain model to difficult examples is improved through hard example mining and resampling. At the same time, the adversarial loss and task loss of difficult examples are used to evaluate the migration effect of the target domain model on difficult examples. Construct a comprehensive evaluation system, including source domain and target domain feature distribution difference measurement, adversarial loss, target domain task performance index, and pseudo-label quality evaluation. The comprehensive migration effect score is obtained by weighted summation. According to the comprehensive migration effect score, the adversarial loss weight, difficult example sampling ratio, and learning rate are adaptively adjusted. At the same time, structural search is performed to select the optimal network structure and hyperparameter combination. A closed-loop transfer learning development process is formed to continuously improve the performance of domain adaptive semantic segmentation. According to the changes in the image data in the source and target domains, the feature distribution difference measurement function and the adversarial loss function are dynamically updated to adapt to the changes in data distribution. At the same time, the weights of each loss in the joint loss function are dynamically adjusted according to the changes in the target domain task performance evaluation indicators to balance the relationship between feature alignment and task performance optimization.

2. The method according to claim 1, wherein: The method uses the maximum average difference metric function to calculate the feature distribution difference based on the source domain and target domain image data, uses the difference metric value as the input of the adversarial loss function, and realizes feature alignment through the gradient reversal layer and the conditional generative adversarial network, including: Acquire source domain image data and target domain image data, and use a pre-trained convolutional neural network to extract deep feature representations of the source domain image data and the target domain image data; Calculating the difference between the source domain image feature distribution and the target domain image feature distribution to obtain a feature distribution difference measurement value; Using the feature distribution difference metric as the input of the adversarial loss function, aligning the feature distribution of the source domain image with the feature distribution of the target domain image through a gradient reversal layer and a conditional generative adversarial network; According to the size of the feature distribution difference metric, dynamically adjust the weights of the domain discriminator and the feature generator in the adversarial loss function, wherein when the feature distribution difference metric is greater than a first threshold, increase the weight of the feature generator; When the feature distribution difference measure value is less than a second threshold, the weight of the domain discriminator is increased.

3. The method according to claim 1, wherein: According to the form of the adversarial loss function, a multi-scale feature adversarial alignment method is used to extract source domain and target domain features at different receptive field scales, and the adversarial losses are calculated respectively. The total adversarial loss is obtained by weighted summation, including: Based on the source domain image and the target domain image, convolution kernels with different receptive field scales are used to extract multi-scale features and obtain feature representations at different abstraction levels; For the source domain features and target domain features extracted at each scale, calculate the adversarial loss between them; Construct a domain discriminator, which uses a multi-layer fully connected network to take source domain features and target domain features as input, output the probability distribution of domain labels, and use the cross entropy loss function for training; When calculating the total adversarial loss, a weighted summation method is used to set the corresponding weight coefficient according to the importance of aligning features at different scales; The weight coefficient is set to be adjusted according to the influence of different scale feature alignment on the final performance.

4. The method according to claim 1, wherein: The semantic segmentation accuracy and average intersection-over-union task performance evaluation index of the target domain image are obtained, and the index is used together with the adversarial loss as a joint loss function, and the data migration process is optimized by a dynamic weight balancing method. If the adversarial loss is reduced to a preset threshold and the target domain task performance is not improved, it is judged that the feature alignment is insufficient, and the source domain and target domain feature mappings are adaptively adjusted by introducing an adaptive layer based on the maximum average difference, including: Acquire the semantic segmentation accuracy and mean intersection-over-union task performance evaluation index of the target domain image, use the task performance evaluation index and the adversarial loss as a joint loss function, and adaptively adjust the weights of the task performance evaluation index and the adversarial loss in the joint loss function according to changes in the task performance index and the adversarial loss in a dynamic weight balancing manner; If there is insufficient feature alignment, the distribution difference between the source domain features and the target domain features is measured, and the MMD distance is added as an additional loss term to the joint loss function to minimize the distribution difference between the source domain features and the target domain features; In the process of mapping the source domain and the target domain features, the importance of the source domain features and the target domain features is adaptively adjusted by learning the attention weights within and between domains, highlighting the features that are beneficial to domain adaptation and suppressing the features that are not beneficial to domain adaptation, so as to promote the alignment of the source domain features and the target domain features.

5. The method according to claim 1, wherein: The trained semantic segmentation model is used to make predictions on the target domain image to obtain pseudo-label data at the pixel level, and the pseudo-label quality is evaluated by calculating the confidence distribution and spatial continuity of the pseudo-label, and high-quality pseudo-labels are selected for semantic segmentation model adjustment, including: For the target domain image, a pre-trained semantic segmentation model is used for forward reasoning to obtain the category probability distribution of each pixel and determine the pixel-level pseudo label according to the preset threshold; For the pseudo labels, count the number of each category and calculate the category balance. If the number of a certain category is less than that of other categories, it is determined that the pseudo labels of this category are unbalanced and data enhancement or oversampling is required. For each pseudo-label, calculate the confidence score, that is, the maximum value of the category probability, and count the confidence distribution of all pseudo-labels. If the distribution is concentrated in the low confidence interval, the pseudo-label quality is judged to be poor and the threshold needs to be adjusted. According to the category balance, confidence distribution and spatial continuity indicators, a quality evaluation function is set, and the weighted sum of the three indicators is taken to obtain the pseudo label quality score; Fine-tune pre-trained semantic segmentation models using high-quality pseudo-labels.

6. The method according to claim 1, wherein: The method aims to improve the adaptability of the target domain model to difficult examples by mining and resampling the small objects and imbalanced categories in the target domain data, and evaluates the migration effect of the target domain model on difficult examples by using the adversarial loss and task loss of the difficult examples, including: According to the characteristics of the target domain data, the task loss and adversarial loss of each sample are obtained; If the task loss of the sample is greater than the preset first threshold or the adversarial loss is greater than the preset second threshold, the sample is determined as a difficult example and added to the difficult example set; For samples in the difficult example set, a data enhancement method is used to obtain multiple enhanced samples, wherein the data enhancement method includes random cropping, rotation, flipping and color transformation; For samples of imbalanced categories, if the number of samples in a certain category is lower than the preset ratio of the total number of samples, its sampling ratio will be increased to the preset target ratio; During the training process, according to the task loss and adversarial loss of the difficult sample, its weight in the total loss function is dynamically adjusted, and the adjustment range of the weight is determined according to the numerical range of the task loss and adversarial loss and the loss change trend on the validation set; In the process of training the target domain model, the target domain model is first pre-trained on the source domain data, and the preset learning rate is used on the target domain data, and the proportion of difficult samples is gradually increased. At the same time, if the increase in the task loss and adversarial loss on the validation set is less than the preset threshold in a continuous preset number of iterations, fine-tuning is stopped to prevent overfitting; Evaluate the transfer effect of the target domain model on difficult examples.

7. The method according to claim 1, wherein: The comprehensive evaluation system is constructed, including source domain and target domain feature distribution difference measurement, adversarial loss, target domain task performance index, pseudo label quality evaluation, and a comprehensive migration effect score is obtained by weighted summation. According to the comprehensive migration effect score, the adversarial loss weight, the difficult example sampling ratio and the learning rate are adaptively adjusted, and a structural search is performed at the same time to select the optimal network structure and hyperparameter combination, including: Construct a comprehensive evaluation system, which includes four aspects: source domain and target domain feature distribution difference measurement, adversarial loss, target domain task performance index and pseudo label quality evaluation; The feature distribution difference measurement value, adversarial loss value, target domain task performance index value and pseudo label quality evaluation value are weighted and summed to obtain a comprehensive migration effect score; Adaptively adjusting the adversarial loss weight, the hard example sampling ratio, and the learning rate according to the comprehensive transfer effect score; If the comprehensive migration effect score is lower than a preset threshold, the adversarial loss weight is increased, the difficult example sampling ratio is increased, and the learning rate is reduced; If the comprehensive migration effect score is higher than a preset threshold, the adversarial loss weight is reduced, the hard example sampling ratio is reduced, and the learning rate is increased; While adaptively adjusting the hyperparameters, the neural network architecture search method is used to search for the optimal network structure and hyperparameter combination, and obtain the network structure and hyperparameter combination with the highest comprehensive migration effect score.

8. The method according to claim 1, wherein: The closed-loop transfer learning development process continuously improves the domain adaptive semantic segmentation performance. According to the changes in the source and target domain image data, the feature distribution difference measurement function and the adversarial loss function are dynamically updated to adapt to the changes in data distribution. At the same time, according to the changes in the target domain task performance evaluation indicators, the weights of each loss in the joint loss function are dynamically adjusted to balance the relationship between feature alignment and task performance optimization, including: Dynamically update the feature distribution difference measurement function based on the changes of the source domain image data and the target domain image data, calculate the difference between the source domain feature distribution and the target domain feature distribution in real time, and obtain a difference measurement value; If the difference metric exceeds a preset threshold, it is determined that the feature distribution has changed, triggering an update of the adversarial loss function; Dynamically adjust the gradient weights corresponding to each loss in the joint loss function based on changes in the target domain task performance evaluation indicators; If the IoU indicator improves slowly or fluctuates, increase the gradient weight corresponding to the adversarial loss and strengthen feature alignment; If the IoU index continues to improve and tends to be stable, the gradient weight corresponding to the semantic segmentation loss is appropriately increased; The semantic segmentation performance is compared on the source domain test set and the target domain test set, including observing the clustering of samples of different categories in the low-dimensional space, analyzing the discriminability and generalization of features, and evaluating the domain adaptive semantic segmentation performance.

Citation Information

Cited By

  • Automobile part welding quality evaluation method and system, storage medium and computer

    CN120411096A

  • Ship engine vibration signal working condition classification method and system based on domain adaptation, medium, program and terminal

    CN120951129A