Uncertainty-guided high-frequency enhancement domain adaptive method

By constructing an adversarial high-frequency enhancement module and frequency gradient balance loss, combined with uncertainty-guided feature reconstruction module, the domain offset problem of deep learning models during inter-domain migration is solved, and the accuracy and robustness of the model in cross-domain image classification tasks are improved.

CN120355583APending Publication Date: 2025-07-22GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510521303.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The domain offset problem caused by distribution differences during existing deep learning models during inter-domain migration, especially in the absence of labels in the target domain, it is difficult to effectively use high-frequency information to adapt.

Method used

Build a high-frequency enhancement domain adaptive method with uncertainty guidance, and through adversarial high-frequency enhancement module, frequency gradient balance loss and uncertainty guidance feature reconstruction module, high-frequency components are decomposed and enhanced, low-frequency and high-frequency gradient contributions are balanced, high-frequency characteristics are suppressed, and effective mining of domain-invariant information is achieved.

Benefits of technology

It significantly improves the accuracy and robustness of the model in cross-domain image classification tasks, enhances attention and learning ability to high-frequency information, and improves the performance of unsupervised domain adaptive tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355583A_ABST
    Figure CN120355583A_ABST
Patent Text Reader

Abstract

The invention discloses an uncertainty-guided high-frequency enhancement domain self-adaption method, and aims to solve the problem of low-frequency learning deviation caused by neglecting high-frequency components in a domain self-adaption task in an existing method. The method comprises the following steps: constructing an adversarial high-frequency enhancement module, decomposing a source domain image into a low-frequency component and a high-frequency component by using two-dimensional discrete wavelet transform, and applying adversarial disturbance to the high-frequency component to guide network learning to obtain high-frequency features with higher domain invariance; frequency gradient balance loss is designed, and the attention of the network to high-frequency information is further enhanced by balancing gradient contributions of low-frequency and high-frequency components; an uncertainty-guided feature reconstruction module is introduced, the feature map is reconstructed through entropy weight, and features of a high-uncertainty region are inhibited; and repeatedly optimizing the source domain and the target domain, and continuously iterating until a preset condition is met. According to the method, the capturing capability of the model on detail information is effectively improved, and the adaptive capacity and generalization performance of the model in different fields are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and particularly relates to an uncertainty-guided high-frequency enhanced domain adaptation method. Background Art

[0002] The success of deep learning models largely depends on a large number of homogeneous training samples, which means that impressive performance can only be obtained when the training data and the test data have the same distribution. On the contrary, when a well-trained model is tested on other datasets with different distributions, satisfactory results cannot be obtained. Taking the training dataset as the source domain and other datasets as the target domain, due to the distribution difference between the source domain and the target domain, the cross-domain migration leads to significant domain shift, which severely limits the generalization ability of deep learning models across different domains, especially when there are no labels or scarce labels in the target domain.

[0003] To address the lack of labels in the target domain, unsupervised domain adaptation (UDA) methods aim to achieve knowledge transfer from the source domain to the target domain by using the labeled data in the source domain and the unlabeled data in the target domain, thereby reducing the distribution difference between domains. Methods based on statistical distance are one common approach, which perform adaptation by minimizing the global distribution difference between the source domain and the target domain and combining the update of pseudo-labels. However, when the true distribution of the target domain is complex or there are specific feature changes during domain migration, the distribution bias caused by pseudo-labels is often difficult to converge, thus affecting the training effect of the model. Methods based on adversarial learning provide an alternative approach, which train a discriminative network to distinguish the features of the source domain and the target domain, thereby encouraging the feature extractor to learn domain-invariant representations to deceive the discriminator. However, since the discriminator only provides a global feedback signal, it cannot clearly identify which features contribute to the learning of domain-invariant information, which limits its ability to further refine or utilize domain-invariant information for adaptation. Intuitively, domain-specific information is mainly dominated by the appearance of objects, while domain-invariant information mainly exists in the contours of objects. Appearance and contour information respectively correspond to low-frequency and high-frequency components in the frequency domain. Therefore, the different characteristics of the low-frequency and high-frequency components in the frequency domain provide new ideas, and using frequency domain methods for domain adaptation may become an effective way to solve the domain shift problem. Summary of the Invention

[0004] The purpose of the present invention is to provide an uncertainty-guided high-frequency enhanced domain adaptation method to solve the problems existing in the above-mentioned prior art.

[0005] To achieve the above purpose, the present invention provides an uncertainty-guided high-frequency enhanced domain adaptation method, including:

[0006] Construct an uncertainty-guided high-frequency enhancement network, which includes an adversarial high-frequency enhancement module, a frequency gradient balance loss, and an uncertainty-guided feature reconstruction module.

[0007] Decompose the source domain input data into low-frequency components and high-frequency components;

[0008] Apply gradient-based adversarial perturbations to the high-frequency components;

[0009] Perform inverse wavelet transform on the enhanced high-frequency components and the original low-frequency components to reconstruct the adversarially enhanced image;

[0010] Obtain the gradient information of the low-frequency and high-frequency components, and balance the gradient contributions of the low-frequency and high-frequency components;

[0011] Extract the feature map of the adversarially enhanced image, and perform weighted reconstruction on the feature map based on the information entropy of the feature map;

[0012] Repeat the optimization process until the preset conditions are met, and use the optimized network to implement the unsupervised domain adaptation task.

[0013] Optionally, decompose the source domain input data into low-frequency components and high-frequency components.

[0014] Optionally, the generation process of the adversarial perturbation includes: initializing the adversarial perturbation, generating the gradient direction of the adversarial perturbation, and updating the adversarial perturbation.

[0015] Optionally, the adversarial perturbation is initialized as a random noise with a standard normal distribution;

[0016] Optionally, the gradient direction of the adversarial perturbation is generated for the high-frequency components, and its direction is obtained by backpropagation of the source domain classification loss, which is opposite to the gradient direction of reducing this loss.

[0017] Optionally, based on the direction of the adversarial perturbation, multiply by a small learning rate to update the adversarial perturbation;

[0018] Optionally, calculate the cross entropy of the forward propagation of the adversarially enhanced image and the label, extract the gradient information of the low-frequency and high-frequency components in the backpropagation, and obtain the gradient magnitudes of the low-frequency and high-frequency components through the L2 norm.

[0019] Optionally, design the frequency gradient balance loss based on the ratio between the gradient magnitudes of the low-frequency and high-frequency components.

[0020] Optionally, based on the entropy value, generate corresponding weights for each element in the feature map, and the weights are negatively correlated with the entropy value. Multiply each element in the feature map by its corresponding weight to obtain the reconstructed feature map.

[0021] Optionally, the uncertainty-guided high-frequency enhancement network is initially optimized by minimizing the cross-entropy loss between the image prediction value enhanced by the source domain confrontation and the corresponding label; and is secondarily optimized by minimizing the frequency gradient balance loss.

[0022] The technical effects of the present invention are as follows:

[0023] Based on the perspective of frequency domain analysis, the present invention deeply analyzes the low-frequency and high-frequency information contained in image data, and emphasizes the key role of high-frequency components in capturing domain-invariant information. In view of the insufficient utilization of high-frequency information by traditional domain self-adaptation methods, the present invention creatively introduces an adversarial high-frequency enhancement module, and uses the idea of gradient confrontation to impose fine perturbations on the high-frequency components of the image, effectively highlighting the domain-invariant features contained in the high-frequency components. On this basis, in order to overcome the inherent low-frequency preference of deep learning models, the present invention further designs a frequency gradient balance strategy, and significantly improves the attention and learning ability of the network model to high-frequency information by dynamically adjusting the gradient contributions of low-frequency and high-frequency components during network training. In addition, aiming at the problem that high-frequency components are easily interfered by noise and outliers, the present invention also innovatively proposes an uncertainty-guided feature reconstruction module, which quantifies the uncertainty of features by means of information entropy, and adaptively adjusts feature weights based on this, thereby effectively suppressing the negative impact of high-uncertainty features and enhancing the robustness and generalization ability of the model. Through the synergistic effect of the above-mentioned multiple innovative technologies, the method of the present invention finally realizes the effective mining and utilization of domain-invariant high-frequency information, significantly improves the performance of unsupervised domain adaptation tasks, and particularly shows excellent technical advantages in tasks such as cross-domain image classification. Description of the Drawings

[0024] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0025] Figure 1 It is a schematic flowchart of the method of the embodiment of the present invention; Detailed Embodiments

[0026] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.

[0027] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0028] Example 1

[0029] As Figure 1 shown, in this embodiment, an uncertainty-guided high-frequency enhanced domain self-adaptation method is provided, including:

[0030] Dataset preparation

[0031] The method of the present invention has been comprehensively tested and evaluated on the Office-Home, Office-31, and Image_CLEF cross-domain image classification benchmarks. Among them, the Office-Home dataset, as a challenging domain adaptation dataset, contains a total of 15,500 images from four different domains, covering four domains: artistic images, clip art, product images, and real-world images, with a total of 65 categories. The Office-31 dataset is a classic domain adaptation dataset, containing 4,652 images, divided into 31 categories, and the images are from three domains: Amazon, webcam, and DSLR. The Image_CLEF dataset is another commonly used domain adaptation evaluation benchmark, containing 2,538 images, with a total of 12 categories, and the images are collected from three different datasets: Caltech-256, ImageNet, and Pascal VOC. The evaluation metric used in the experimental evaluation is classification accuracy (Accuracy) to comprehensively measure the performance of the method of the present invention in different domain adaptation scenarios.

[0032] Data preprocessing

[0033] Considering that the training of deep learning models usually requires sufficient data, and in order to further improve the generalization ability of the model, the embodiment of the present invention introduces necessary data augmentation operations in the data preprocessing stage. Specifically, the method of the present invention first uniformly adjusts the size of all input images to a resolution of 256×256. Subsequently, in order to increase the diversity of training samples and simulate the randomness that may occur during data acquisition, the resized images are randomly cropped into image patches of 224×224 size as the actual input of the model. In addition, in order to further expand the data volume and improve the robustness of the model, the present invention also adopts a data augmentation method of random horizontal flipping to perform horizontal mirror transformation on some training samples, thereby generating more diverse training data, enabling the model to learn image features more fully and improving the adaptation ability in cross-domain scenarios.

[0034] Parameter setting

[0035] The experimental verification of the method of the present invention is implemented based on the PyTorch deep learning framework and is trained and tested on an NVIDIA Tesla GV100 GPU. To ensure the effectiveness and reproducibility of the experiment, the embodiments of the present invention have refined the settings of the key hyperparameters of the model. Specifically, the mini-batch stochastic gradient descent optimization algorithm is used in the training process, the momentum parameter is set to 0.9, and the weight decay parameter is set to 0.0001. The initial learning rate is set to 0.01.

[0036] Based on the above embodiments, an uncertainty-guided high-frequency enhancement network for final training is built. On this network, the source domain images, source domain labels, and target domain images are used for collaborative training.

[0037] In the first step, the two-dimensional discrete wavelet transform technology is adopted to decompose the image signal into frequency domain sub-bands of different scales and directions, so as to realize the multi-resolution representation of the image. Specifically, the two-dimensional discrete wavelet transform decomposes the input source domain image into four sub-band images including one low-frequency component and three high-frequency components. Among them, the low-frequency component (LL sub-band) is approximately the low-resolution version of the original image, capturing the overall contour and main information of the image; while the three high-frequency components (HL, LH, HH sub-bands) capture the high-frequency detail information of the image in the horizontal, vertical, and diagonal directions respectively, such as edges, textures, and noises. Through the decomposition of the two-dimensional discrete wavelet transform, the image information is effectively separated into different frequency domain sub-bands, laying a foundation for subsequent differential processing and feature extraction for different frequency components.

[0038] In the second step, an adversarial high-frequency enhancement module is constructed. The gradient of the loss function with respect to the high-frequency component is used as the adversarial perturbation to provide variation for the high-frequency component within a controllable range. It increases the high-frequency information content, which becomes more obvious with iteration compared to the low-frequency component. This process enables the network to focus on high-frequency features and learn more robust and generalizable domain-invariant information. The enhanced high-frequency component can be expressed as:

[0039]

[0040] where \(C_i\in\{LH, HL, HH\}\), the perturbation \(\delta\) is initialized as \(\alpha\cdot N(0, 1)\), \(\alpha\) is a hyperparameter, and \(N(0, 1)\) is the standard normal distribution. \(\delta\) is updated as:

[0041]

[0042] where \(\epsilon\) is a small learning rate added in the gradient direction of the perturbation \(\delta\), and \(sign(\cdot)\) represents the direction of the gradient. L cls (·,·) is the classification loss \(L\) cls(·, ·) derivative with respect to δ, where θ is a parameter of the network. Multiply the small learning rate ∈ by the direction sign is a small vector. Adding a small vector to the perturbation δ will slightly increase its magnitude in each iteration, and the growth of the high-frequency information content in the image data remains at a consistent level in each epoch. This vector provides the gradient in the opposite direction of the network. This adjustment increases the contribution of the high-frequency components to the loss gradient, encouraging the network to pay more attention to these features to improve its generalization performance.

[0043] In the third step, after adversarial enhancement, the high-frequency components should be combined with the low-frequency components because the low-frequency components have rich appearance information that dominates the network's learning. Using the inverse wavelet transform, the high-frequency enhanced image The result of the combination can be expressed as:

[0044]

[0045] where ψ H (·, ·), ψ V (·, ·) and ψ D (·, ·) are the basic functions of the two-dimensional discrete wavelet transform. j controls the scaling factor of the basic function, and h and w control the horizontal and vertical translations of the basis function to cover the entire image area. H and W are the length and width of the image respectively. Different combinations of j, h, and w can generate a set of orthogonal or non-orthogonal basic functions for reconstructing different components of the image. These reconstructed data retain the basic global features while highlighting the domain-invariant high-frequency details. Such a process effectively supports the network to capture robust domain-invariant features, which is crucial for improving cross-domain generalization.

[0046] In the fourth step, although adversarial high-frequency enhancement improves the network's attention to domain-invariant features, it still cannot balance the high-frequency and low-frequency information content because the low-frequency components have greater energy. To balance the contributions of the high-frequency and low-frequency components, a frequency gradient balance loss is proposed. This loss dynamically adjusts the network's attention to the two frequency components by suppressing the low-frequency gradient (which usually dominates in learning). At the same time, it enhances the response to the high-frequency gradient, thereby dynamically adjusting the network's attention to the high-frequency and low-frequency components. High-frequency gradients usually carry neglected domain-invariant information. Through this balanced gradient distribution, the network can maintain its robustness in the domain adaptation task in the face of a large amount of low-frequency appearance information interference. The frequency gradient balance loss is defined as follows:

[0047]

[0048] where ||·||2 represents the L2 norm, and r m is the ratio of the low-frequency gradient to the high-frequency gradient. At r mAdd a small positive number ε to the denominator to prevent the denominator from being 0 in practice. grad LL and grad m are the gradients of the low-frequency component and the high-frequency component, respectively obtained by taking the partial derivatives of the low-frequency components LL and high-frequency components LH, HL, and HH with respect to L cls as follows:

[0049]

[0050] where w are the network parameters.

[0051] Theoretically, the ratio r m exceeds 1 because the magnitude of the low-frequency gradient is larger than that of the high-frequency gradient. This also provides a clue as to why, without the guidance of L FGB , the learning process of the network is mainly driven by the low-frequency component. Minimizing the ratio r m is to reduce the gradient difference between the high frequency and the low frequency. In other words, the network will suppress its learning on the low-frequency component and focus on learning the high-frequency component. Therefore, by adjusting the ratio r m between their contributions, the biased learning of the low-frequency and high-frequency components is balanced from the perspective of their contributions. Therefore, the network is forced to adjust its optimization direction, thereby reducing the over-reliance on the low-frequency component and enhancing the model's ability to adapt to different domains.

[0052] Fifth step, the uncertainty-guided feature reconstruction module. First, use the feature extractor of the convolutional neural network to extract the feature map of the high-frequency enhanced image, calculate the information entropy to measure the uncertainty of the feature map, and use the reciprocal of the uncertainty as the weight to reconstruct the feature map. The reconstruction process is defined as follows:

[0053] R k = F k · u k

[0054] where k is the index of the feature map element, R k is the reconstructed feature map, F k is the feature map extracted from the high-frequency enhanced image, and u k is the uncertainty-guided weight, defined as:

[0055] u k = 1 - H k + β

[0056] where β is a hyperparameter to ensure that the weight is not too small, and H k is the entropy of the feature map F k :

[0057] H k = -p(F k) logp(F k )

[0058] where p(F k ) = softmax(F k ) can be regarded as the probability distribution of F k .

[0059] In this module, the entropy H k is used to represent the uncertainty of the feature map F k . The larger the information entropy, the greater the uncertainty, and the smaller the weight learned by the network. Therefore, the reciprocal of the entropy is used as the weight to enhance the feature map. This process enables the network to suppress features with high uncertainty, thereby enhancing its robustness in domain adaptation. This tailored adjustment helps to reduce the impact of unreliable information and enables the network to focus on learning from more stable and discriminative features.

[0060] Step 6: Repeat the above training steps until the preset number of iterations, and test the image classification accuracy.

[0061] As described above, the above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An uncertainty-guided high-frequency enhancement domain adaptation method, characterized in that It includes the following steps: Construct an uncertainty-guided high-frequency enhancement network, which includes an adversarial high-frequency enhancement module, a frequency gradient balance loss, and an uncertainty-guided feature reconstruction module; Decompose the source domain input data into low-frequency components and high-frequency components; Apply gradient-based adversarial perturbations to the high-frequency components; Perform inverse wavelet transform on the enhanced high-frequency components and the original low-frequency components to reconstruct the adversarially enhanced image; Obtain the gradient information of the low-frequency and high-frequency components, and balance the gradient contributions of the low-frequency and high-frequency components; Extract the feature map of the adversarially enhanced image, and perform weighted reconstruction on the feature map based on the information entropy of the feature map; Repeat the optimization process until the preset conditions are met, and use the optimized network to implement the unsupervised domain adaptation task.

2. The uncertainty-guided high-frequency enhancement domain adaptation method according to claim 1, wherein The decomposition of the source domain input data includes: decomposing the source domain input data using two-dimensional discrete wavelet transform.

3. The uncertainty-guided high-frequency enhancement domain adaptation method according to claim 1, wherein The generation process of the adversarial perturbation includes: initializing the adversarial perturbation, generating the gradient direction of the adversarial perturbation, and updating the adversarial perturbation.

4. The uncertainty-guided high-frequency enhancement domain adaptation method according to claim 3, wherein The adversarial perturbation is initialized as a random noise with a standard normal distribution.

5. The uncertainty-guided high-frequency enhancement domain adaptation method according to claim 4, wherein The gradient direction of the adversarial perturbation is generated for the high-frequency component, and its direction is obtained by backpropagation of the source domain classification loss, which is opposite to the gradient direction of reducing this loss.

6. The uncertainty-guided high-frequency enhancement domain adaptation method according to claim 5, wherein Based on the direction of the adversarial perturbation, multiply by a small learning rate to update the adversarial perturbation.

7. The uncertainty-guided high-frequency enhancement domain adaptation method according to claim 1, wherein Forward propagate the adversarially enhanced image and calculate the cross-entropy with the label. In backpropagation, extract the gradient information of the low-frequency and high-frequency components, and obtain the gradient magnitudes of the low-frequency and high-frequency components through the L2 norm.

8. The uncertainty-guided high-frequency enhancement domain adaptation method according to claim 7, wherein Design the frequency gradient balance loss based on the ratio between the gradient magnitudes of the low-frequency and high-frequency components.

9. The uncertainty-guided high-frequency enhancement domain adaptation method according to claim 1, wherein Based on the entropy value, generate corresponding weights for each element in the feature map, and the weights are negatively correlated with the entropy value. Multiply each element in the feature map by its corresponding weight to obtain the reconstructed feature map.

10. The uncertainty-guided high-frequency enhancement domain adaptation method according to claim 1, wherein Preliminarily optimize the high-frequency enhancement network by minimizing the cross-entropy loss between the predicted value of the source domain adversarially enhanced image and the corresponding label; perform secondary optimization by minimizing the frequency gradient balance loss.

Citation Information

Cited By

  • Basic hole semi-supervised segmentation method for aviation structural part robot drilling

    CN120913210A