A Face Image Super-Resolution Method, System and Storage Medium Capable of Defending Against Adversarial Attacks
Through the frequency-aware decomposition network model based on empirical modal decomposition, combined with multi-branch structure and high-frequency noise suppressor, the problem of adversarial attack robustness and lack of high-frequency details of the super-resolution model of face images in the prior art is solved, and more robust and high-quality image recovery is achieved.
Patent Information
- Application Number
- CN202411460838.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-10-18
AI Technical Summary
The existing DNN and Transformer-based super-resolution models of face images lack the robustness of adversarial attacks, and the recovered images are too smooth and lack high-frequency details, especially under large upsampling factors.
Using a frequency-aware decomposition network model based on empirical modal decomposition, a multi-branch structure and high-frequency noise suppressor are used, combined with learnable tips, the loss function and adversarial training are designed to separate and eliminate adversarial noise and restore high-frequency details.
The adversarial attack robustness of the model is enhanced, high-frequency details are restored, and the visual quality and recognition accuracy of the image are improved.
Smart Images

Figure CN119417700B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a face image super-resolution method, system and storage medium that can defend against adversarial attacks. Background Art
[0002] In real-world scenarios, the quality of captured face images is often affected by hardware limitations, shooting angles, and device positioning. These low-quality images can seriously hinder downstream tasks such as face recognition and pedestrian tracking, making face super-resolution (FSR) crucial in almost all face-related applications. Different from general single-image super-resolution, FSR highly emphasizes the authenticity of the reconstructed image because accurate identity representation is vital for downstream tasks.
[0003] Face super-resolution is a key step in the face analysis process, focusing on reconstructing high-resolution (HR) images from low-resolution (LR) face images, and has made significant progress by applying deep neural networks (DNNs) and Transformers. In recent years, researchers have proposed various convolutional neural networks (CNNs) to enhance the performance of FSR, including methods that utilize geometric priors such as facial landmarks and heatmaps to extract prior information. However, CNN-based FSR methods rely on local convolutional operations and ignore the relationship between facial geometry and symmetry. To address this issue, Transformer-based FSR methods have been proposed, which capture long-range dependencies through self-attention mechanisms. Some methods apply self-attention to the entire image or within predefined local windows, while others design hybrid networks that combine CNNs and Transformers to fully utilize local and global features.
[0004] DNN- and Transformer-based methods have dominated the FSR field with their impressive performance. However, these methods still have unresolved problems. On the one hand, existing DNN- and Transformer-based FSR models lack robustness against adversarial attacks, where imperceptible noise may distort the reconstructed image and generate artifacts. On the other hand, the images restored by current FSR models are visually too smooth and lack high-frequency details, especially at larger upsampling factors. CNNs can be regarded as filters, and the self-attention mechanism in Transformers mainly acts as a low-pass filter. Essentially, it tends to extract low-frequency components and then gradually shift the focus to high-frequency details. Therefore, the super-resolved images are too smooth and lack high-frequency details of complex facial components, such as eyes. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes a face image super-resolution method, system and storage medium that can defend against adversarial attacks to solve the problems existing in the above-mentioned prior art.
[0006] To achieve the above object, the present invention provides a face image super-resolution method that can defend against adversarial attacks, including:
[0007] Construct a frequency-aware decomposition network model based on empirical mode decomposition;
[0008] Obtain a clean sample set, and obtain an adversarial sample set based on the clean sample set;
[0009] Perform adversarial training on the frequency-aware decomposition network model based on the clean sample set and the adversarial sample set;
[0010] Input a given face image into the trained frequency-aware decomposition network model to obtain deep features; obtain an intrinsic mode function feature map based on the deep features, and reconstruct the face image based on the intrinsic mode function feature map.
[0011] Optionally, the frequency-aware decomposition network model includes a convolutional block and several cascaded recovery branches. The recovery branch includes a Transformer block, a hint component and a frequency modulator. The Transformer block includes a Multi-Dconv head transposed attention and a Gated-Dconv feed-forward network. The convolutional block is used to extract shallow features of the network input.
[0012] Optionally, by using the intrinsic mode function as a constraint, respectively control several cascaded recovery branches to process information of specific frequencies, and introduce a high-frequency suppressor in the CNN filter bank corresponding to high frequencies.
[0013] Optionally, the high-frequency suppressor applies Fourier transform to map shallow features to the frequency domain, calculates the mask of the high-frequency part, and after performing element-wise multiplication on the mask of the high-frequency part and the shallow features mapped to the frequency domain, obtains deep features through inverse Fourier transform. The process is as follows:
[0014]
[0015] Wherein, represents Fourier transform, represents inverse Fourier transform, ⊙ is element-wise multiplication, and M is the frequency domain mask.
[0016] Optionally, the shallow features are input into the recovery branch of the frequency-aware decomposition network model and downsampled through two Transformer blocks. The obtained initial feature map passes through three groups of frequency modulators and Transformer blocks in sequence to obtain an output. Among them, the prompt component generates learnable parameters and inputs them into the frequency modulator. The frequency modulator modulates the output of the Transformer block and the learnable parameters, and the obtained modulation signal is processed by the Transformer blocks in the same group. When the initial feature map passes through the first two groups, the Transformer block performs an upsampling operation on the feature map.
[0017] Optionally, the frequency modulator predicts the input feature map as a weight parameter, multiplies the weight coefficient by the corresponding learnable parameter to obtain a prompt message, and connects the prompt message with the input feature map to obtain a modulation signal.
[0018] Optionally, the process of adversarial training includes:
[0019] Construct a unified loss function, and use the clean sample set to train the frequency-aware decomposition network model. When the model starts to converge, alternately use the clean sample set and the adversarial sample set for training.
[0020] Optionally, the formula of the unified loss function is as follows:
[0021]
[0022] Among them, is the IMF alignment loss; is the high-frequency retention loss; represents the reconstruction loss, and the L1 loss is adopted; represents the gradient loss; μ, ζ, and η are balance weights;
[0023]
[0024] Among them, i j represents the true IMF, is the predicted IMF, r n is the true low-frequency IMF, is the predicted remaining low-frequency IMF, α n is the balance weight;
[0025]
[0026] Among them, i j represents the true IMF, is the predicted IMF, σ j is the balance weight, and β and γ are relaxation values.
[0027] The present invention also provides a face image super-resolution system capable of defending against adversarial attacks, including:
[0028] A model construction module for constructing a frequency-aware decomposition network model based on empirical mode decomposition;
[0029] A dataset construction module for obtaining a clean sample set and obtaining an adversarial sample set based on the clean sample set;
[0030] A model training module for performing adversarial training on the frequency-aware decomposition network model through the clean sample set and the adversarial sample set;
[0031] A feature extraction module for obtaining deep features according to a given face image and the trained frequency-aware decomposition network model;
[0032] An image reconstruction module for obtaining an intrinsic mode function feature map according to the deep features and reconstructing a face image based on the intrinsic mode function feature map.
[0033] The present invention also provides a computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the steps of the above-mentioned face image super-resolution method capable of defending against adversarial attacks are implemented.
[0034] Compared with the prior art, the present invention has the following advantages and technical effects:
[0035] The present invention proposes a new robust face super-resolution (FSR) model, which can be used to defend against adversarial noise of CNN and Transformer. A new multi-branch structure based on empirical mode decomposition (EMD) is used for face image super-resolution. By decoupling and eliminating harmful perturbations in the frequency domain, the robustness of the FSR model is enhanced, and at the same time, the performance of restoring high-frequency details is improved. In the EMD-based multi-branch framework, the internal model structure and loss function are designed to capture global topology and fine texture details. A high-frequency noise suppressor is introduced to defend against adversarial attacks by selectively removing high-frequency components; learnable prompts are introduced and modulated by a frequency modulator to inject the high-frequency degradation information required by the main network. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0037] Figure 1The network structure diagram of the embodiment of the present invention. (a) is the architecture of a single CNN filter bank, (b) is the detailed diagram of the Transformer block, and (c) is the schematic diagram of the frequency modulator;
[0038] Figure 2 The schematic diagram of the frequency-aware decomposition network model of the embodiment of the present invention;
[0039] Figure 3 The schematic diagram of the structure of the high-frequency noise suppressor of the embodiment of the present invention. Detailed implementation manners
[0040] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.
[0041] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0042] Embodiment 1
[0043] As Figure 1 shown, this embodiment provides a face image super-resolution method that can defend against adversarial attacks, including:
[0044] CNN and Transformer are vulnerable to adversarial attacks, that is, adding imperceptible perturbations to the input to deceive the model. Adversarial attacks can be divided into two categories: black-box attacks and white-box attacks. Black-box attacks have limited access to the target data and the model, while white-box attacks can fully access the network architecture, parameters, and weights. This embodiment discusses white-box attacks, and white-box attacks generally follow the following paradigm:
[0045] Let x represent the clean image, and y represent its corresponding ground truth. A trained model is represented as f(x) = y. Then, the slight perturbation δ generated by the adversarial attack is added to χ pixel by pixel, thereby generating the adversarial sample x * , which guides the model to produce incorrect outputs. The perturbation δ is restricted within the range of [-∈, ∈] to ensure that it is visually imperceptible. Therefore, the adversarial sample can be obtained through x * = x + δ. Then, the perturbation is optimized by maximizing the deviation of the attacked sample:
[0046]
[0047] where O represents the objective for measuring the output degradation.
[0048] Compared with clean samples, the high-frequency coefficients of adversarial samples increase significantly, indicating that adversarial noise is densely distributed in the high-frequency part. Therefore, adversarial attacks can be defended from the frequency perspective, and more detailed facial information can be obtained.
[0049] Empirical Mode Decomposition (EMD) is an effective frequency-aware transformation tool. EMD decomposes the input into high-frequency and low-frequency subbands with strong adaptability. The components derived from EMD are called Intrinsic Mode Functions (IMFs), which can accurately represent global topology and texture information. The IMFs are naturally arranged, presenting characteristics from high frequency to low frequency, while adversarial noise is mainly distributed in the high-frequency IMFs, indicating that EMD can separate harmful interferences and has excellent frequency-aware characteristics.
[0050] This embodiment proposes a Frequency-Aware Decomposition Network model (FDNet) to decompose and extract features of different frequencies, eliminate harmful perturbations in the frequency domain, and enhance high-frequency information.
[0051] Specifically, a multi-branch structure based on EMD is designed. This structure is transformed into different cascaded filter banks and outputs IMFs in the order from low frequency to high frequency. IMF constraints are applied to each branch. Utilizing the frequency-aware ability of EMD, it is implicitly transformed into a filter bank that specifically extracts features from a specific frequency band. Adversarial noise is decoupled and restricted to the corresponding high-frequency branch for processing. In addition, each branch focuses on restoring unique features, thereby reducing the complexity of prediction, which also helps to obtain more accurate high-frequency information and eliminate perturbations.
[0052] A high-frequency noise suppressor is introduced in the high-frequency branch to eliminate potential subtle perturbations by selectively eliminating high-frequency features.
[0053] In addition, the model in this embodiment also combines learnable cues to provide additional high-frequency information that can be interactively injected into the model. A lightweight module (cue component) is used to generate a set of parameters, and a frequency modulator is proposed to modulate these cues as noise. By injecting the modulated high-frequency information into the main network, the learnable cues can compensate for high-frequency degradation, thereby effectively restoring detailed textures.
[0054] The present invention designs the detailed structure and loss function of the internal module of the multi-branch structure based on EMD, and conducts adversarial training on the model to achieve better results on both adversarial samples and clean samples. The specific content is as follows:
[0055] 1) Overall process
[0056] For a given degraded image x, y represents the real HR image. The goal of face super-resolution is to learn a mapping function f θ () with parameter θ to reconstruct a clear HR face image Assume n - level EMD, y is decomposed into the true IMFI=(i1, i2, …, i n ; r n ), while represents the predicted IMF. The mapping function is defined as the recovery branch, is the learnable prompt, then f θ is:
[0057]
[0058] As Figure 1 shown, in the model, it starts with shallow feature extraction using convolution; then, these extracted features are transformed into deep features through multiple cascaded recovery branches, thereby predicting a series of IMF feature maps. The recovery branch consists of Transformer blocks and learnable prompts. Among them, the Transformer blocks can be replaced by convolutional groups with the same effect. This figure takes Transformer as an example.
[0059] Finally, these feature maps are transformed back into IMF and stacked to reconstruct the HR face.
[0060] 2) EMD - guided paradigm
[0061] The principle of EMD is to decompose the input signal into frequency components (called IMF), thereby achieving adaptive frequency analysis. EMD was initially proposed for one - dimensional signals and has since been widely applied to image analysis, image compression, and other fields. IMF is extracted through the process of isolating the highest - frequency oscillation of the input data. The image can be reconstructed by superimposing all IMFs:
[0062]
[0063] where IMFi j represents the higher - frequency component, r n represents the remaining low - frequency residue. IMF usually has the characteristics of concentrated and continuous frequency distribution, reflecting the inherent characteristics of the image. It should be noted that for face images, the frequency distribution of IMF is relatively stable and independent.
[0064] The principle of the EMD - guided paradigm:
[0065] The intuitive display of EMD is as Figure 1 shown, and the obtained IMF after decomposition shows a trend of changing from high - frequency to low - frequency. It should be noted that when applying EMD decomposition to adversarial samples, it is obvious that adversarial noise mainly exists in the high - frequency IMF. To utilize the unique characteristics of IMF, we designed a multi - branch structure based on EMD, as Figure 2As shown, drawing on digital signal theory, it can be conceptualized as a complex digital filter bank. Therefore, the EMD-based multi-branch structure contains multiple filter banks connected in series. The IMFs are used as constraints to implicitly force each filter bank to predict the IMFs of specific frequencies. This design allows each branch to learn the frequency-aware decomposition ability of EMD and process features of specific frequencies. When adversarial samples are input into the network, the invisible perturbations are decoupled into high-frequency IMFs and routed to the corresponding high-frequency branches for further processing. In this way, the noise is decoupled and restricted, simplifying the problem of defending against attacks.
[0066] Furthermore, it is assumed that if the filter bank focuses on data of specific frequencies (e.g., high-frequency images) rather than the entire spectrum, the adjustment of internal parameters of the network will be simplified, thereby improving the fitting ability. The filter banks in the multi-branch structure independently predict the IMFs of each frequency. Therefore, specialization reduces the attention span of each branch and greatly reduces the prediction difficulty. In this embodiment, the structures of different branches are designed to pay more attention to the high-frequency branches, so as to obtain more high-frequency information in the results. At the same time, more accurate high-frequency information means that the adversarial noise is eliminated and a more robust model is obtained.
[0067] Details of the EMD-guided paradigm:
[0068] The self-attention mechanism in Transformers acts as a low-pass filter, and CNNs can also be regarded as filters that initially extract low-frequency features and then gradually focus on high frequencies. To solve this problem, multiple Transformer blocks / CNNs are combined and integrated into the restoration branches, and these branches are connected in series to form the overall network. Following the rules of the network, the low-frequency information is processed first and then the high-frequency information. Each branch is supervised by the IMFs in order from low frequency to high frequency. This paradigm can focus on predicting specific IMFs, thereby enhancing the high-frequency components and restoring fine details.
[0069] As Figure 1 shown, the restoration branch follows a hierarchical encoder-decoder structure, where multiple Transformer blocks are used at each level. In the encoder and decoder, the number of patches increases from top to bottom to maintain computational efficiency. Among them, the basic Transformer block consists of Multi-Dconv Head Transposed Attention (MDTA) and Gated-Dconv Feed-Forward Network (GDFN). This setting reduces the computational complexity while effectively transforming features.
[0070] Adversarial perturbations are usually encoded as high-frequency components. Therefore, in this embodiment, a high-frequency suppressor is introduced in the high-frequency branch to mitigate adversarial attacks. The architecture of the high-frequency noise suppressor is as Figure 3As shown. This module first applies the Fourier transform to map the input features to the frequency domain, randomly occludes the high-frequency part using a mask, then calculates the mask of the high-frequency part, performs element-wise multiplication with the feature map, and finally performs the inverse Fourier transform to obtain the output features, converting the masked spectrogram back to the feature space. Let \(X\) represent the input features, represent the output. The formula for this process is as follows:
[0071]
[0072] where, represents the Fourier transform, represents the inverse Fourier transform, ⊙ is element-wise multiplication, and \(M\) is the frequency-domain mask.
[0073] To effectively utilize EMD and restore the texture details of HR faces, this embodiment uses the IMF alignment loss and the high-frequency retention loss to constrain the IMF components. In addition, we use the reconstruction loss and the gradient loss to improve image reconstruction.
[0074] IMF alignment loss: Align the predicted IMF and the real IMF in the Euclidean space, and its formula is:
[0075]
[0076] where \(W = (\alpha_1,\alpha_2,\ldots,\alpha n+1 )\) is the weight matrix used to balance the importance between the IMF and the residual. In this embodiment, higher weights are assigned to the high-frequency IMF to restore the texture, \(i j represents the real IMF, is the predicted IMF \(r n is the real low-frequency IMF, is the predicted remaining low-frequency IMF, and \(\alpha n is the balancing weight.
[0077] High-frequency retention loss: Used to retain the high-frequency IMF and prevent it from degenerating to zero. Its formula is:
[0078]
[0079] where \(\sigma j is the balancing weight, and \(\beta\) and \(\gamma\) are relaxation values.
[0080] This embodiment adopts the gradient loss proposed by the CDC model to dynamically adjust the attention area, and uses the L1 loss as the reconstruction loss Finally, the unified loss function formula is as follows
[0081]
[0082] Among them, μ, ζ, and η are balancing weights.
[0083] 3) Learnable Hints
[0084] In this embodiment, learnable hints are used to add additional high-frequency information to the network. The learnable hints are a set of parameters created by lightweight modules and can be regarded as noise. The noise is modulated through a frequency modulator to interact with the features, making the hint components become the high-frequency information required by the main network and realizing dynamic compensation for the high-frequency degradation of the input surface.
[0085] Specifically, in this embodiment, hints are introduced in the restoration branch to implicitly enrich the features with information about the degradation. As Figure 1 (c) shows, these hint components are learnable parameters that are input into the frequency modulator to embed information related to the degradation. The frequency modulator first applies channel attention in the entire spatial domain. It predicts the input feature map as weight parameters and multiplies them with the hint components. This involves global average pooling, softmax operations, and other processes. Then, the interaction between the features and the hints is achieved by concatenating the generated hint components with the input along the channel dimension, which helps to restore the high-frequency details.
[0086] 4) Adversarial Training
[0087] In this embodiment, clean samples and adversarial samples are used for model training. During the training phase, the challenge is that if adversarial samples are used for training, the performance of the clean samples may decline. To solve this problem, a strategy is adopted to enhance the robustness of the model against adversarial attacks while maintaining the performance of the clean samples. The training phase starts with using clean samples. When the model converges, clean samples and adversarial samples are alternately used to train it. Specifically, the adversarial samples are generated using the current model and are iteratively updated during the training process. The following pseudocode summarizes the entire adversarial training strategy as follows:
[0088] Input: Low-resolution image x;
[0089] Output: Super-resolution image
[0090] Set the training times as T1 and T2, and the number of IMFs as N;
[0091] Step 1:
[0092]
[0093]
[0094] Step 2:
[0095]
[0096] This embodiment also provides a face image super-resolution system that can defend against adversarial attacks, including:
[0097] A model construction module for constructing a frequency-aware decomposition network model based on empirical mode decomposition;
[0098] A dataset construction module for obtaining a clean sample set and obtaining an adversarial sample set based on the clean sample set;
[0099] A model training module for performing adversarial training on the frequency-aware decomposition network model through the clean sample set and the adversarial sample set;
[0100] A feature extraction module for obtaining deep features according to a given face image and the trained frequency-aware decomposition network model;
[0101] An image reconstruction module for obtaining an intrinsic mode function feature map according to the deep features and reconstructing a face image based on the intrinsic mode function feature map.
[0102] Specifically, the content in the above method embodiments is applicable to the system embodiments of this system. The functions specifically implemented in the system embodiments of this system are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0103] The layers, modules, units, and / or platforms included in the system can be implemented or implemented by computer hardware, a combination of hardware and software, or computer instructions stored in a non-transitory computer-readable memory.
[0104] The present invention also provides a computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the steps of the above method for super-resolving face images that can defend against adversarial attacks are implemented.
[0105] The method can be implemented in a computer program using standard programming techniques - including a non-transitory computer-readable storage medium configured with the computer program, wherein the storage medium so configured causes the computer to operate in a specific and predefined manner. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if necessary, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, for this purpose, the program can run on a special integrated circuit programmed for this purpose.
[0106] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A face image super-resolution method capable of defending against adversarial attacks, characterized in that, Including the following steps: Construct a frequency-aware decomposition network model based on empirical mode decomposition; Obtain a clean sample set, and obtain an adversarial sample set based on the clean sample set; Perform adversarial training on the frequency-aware decomposition network model based on the clean sample set and the adversarial sample set; Input a given face image into the trained frequency-aware decomposition network model to obtain deep features; obtain an intrinsic mode function feature map based on the deep features, and reconstruct the face image based on the intrinsic mode function feature map; The frequency-aware decomposition network model includes a convolutional block and several cascaded recovery branches. The recovery branch includes a Transformer block, a prompt component, and a frequency modulator. The Transformer block includes a Multi-Dconv head transposed attention and a Gated-Dconv feed-forward network. The convolutional block is used to extract shallow features of the network input; The shallow features are input into the recovery branch in the frequency-aware decomposition network model and are downsampled through two Transformer blocks. The obtained initial feature map sequentially passes through three groups of frequency modulators and Transformer blocks to obtain an output. Among them, the prompt component generates learnable parameters and inputs them into the frequency modulator. The frequency modulator modulates the output of the Transformer block and the learnable parameters to obtain a modulation signal, and the obtained modulation signal is processed through the Transformer block in the same group; when the initial feature map passes through the first two groups, the Transformer block performs an upsampling operation on the feature map; The frequency modulator predicts the input feature map as a weight parameter, multiplies the weight coefficient by the corresponding learnable parameter to obtain a prompt message, and connects the prompt message with the input feature map to obtain a modulation signal.
2. The method for super-resolution of face images capable of defending against adversarial attacks according to claim 1, wherein By using the intrinsic mode function as a constraint, respectively control several cascaded recovery branches to process information of specific frequencies. Among them, a high-frequency suppressor is introduced into the CNN filter bank corresponding to high frequencies.
3. The method for super-resolution of face images capable of defending against adversarial attacks according to claim 2, wherein The high-frequency suppressor applies Fourier transform to map the shallow features to the frequency domain, calculates the mask of the high-frequency part, and performs element-wise multiplication on the mask of the high-frequency part and the shallow features mapped to the frequency domain, and then obtains the deep features through inverse Fourier transform. The process is expressed as follows: Among them, represents the Fourier transform, represents the inverse Fourier transform, ⊙ is the element-wise multiplication, and M is the frequency-domain mask.
4. The method for super-resolution of face images capable of defending against adversarial attacks according to claim 1, wherein The process of performing adversarial training includes: Construct a unified loss function, and use the clean sample set to train the frequency-aware decomposition network model. When the model starts to converge, alternately use the clean sample set and the adversarial sample set for training.
5. The method for super-resolution of face images capable of defending against adversarial attacks according to claim 4, wherein The formula of the unified loss function is as follows: where, is the IMF alignment loss; is the high-frequency preservation loss; represents the reconstruction loss, using the L1 loss; represents the gradient loss; μ, ζ, and η are balancing weights; Among them, i j represents the true IMF, is the predicted IMF, r n is the true low-frequency IMF, is the predicted remaining low-frequency IMF, α n is the balancing weight; where \(i\) j represents the true IMF, is the predicted IMF, \(\sigma\) j is the balancing weight, and \(\beta\) and \(\gamma\) are relaxation values.
6. A face image super-resolution system for defensively countering attacks for implementing the method according to any one of claims 1-5, characterized in that, Including: A model construction module for constructing a frequency-aware decomposition network model based on empirical mode decomposition; A dataset construction module, configured to obtain a clean sample set and obtain an adversarial sample set based on the clean sample set; A model training module, configured to perform adversarial training on the frequency-aware decomposition network model by using the clean sample set and the adversarial sample set; A feature extraction module, configured to obtain deep features according to a given face image and the trained frequency-aware decomposition network model; An image reconstruction module, configured to obtain an intrinsic mode function feature map according to the deep features and reconstruct a face image based on the intrinsic mode function feature map.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the face image super-resolution method capable of defending against adversarial attacks according to any one of claims 1-5 are implemented.
Citation Information
Patent Citations
Image super-resolution system based on empirical mode decomposition
CN114841861A
Unknown adversarial attack-oriented face forgery recognition method and device
CN115984979A