Face flaw removal method based on frequency domain repair and multi-resolution fusion

Through the face defect removal method that integrates frequency domain repair and multi-resolution, the frequency domain perception dynamic aggregation, spatial domain projection and frequency domain selection feedforward network modules are used to solve the shortcomings of global consistency and detail recovery in the existing technology, and efficient defect removal and detail retention are achieved.

CN120339124APending Publication Date: 2025-07-18BIG DATA & INFORMATION TECH RES INST OF WENZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510339147.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing facial defect removal techniques have shortcomings in global consistency and detail recovery, making it difficult to efficiently and automatically remove facial defects and retain detailed features.

Method used

The face defect removal method based on frequency domain repair and multi-resolution fusion is adopted, and the feedforward network module is used to generate high-quality deficit-defective face images through frequency domain perception dynamic aggregation, spatial domain projection and frequency domain selection, and combined with the multi-resolution fusion module.

Benefits of technology

It effectively solves the problem of detail fidelity and global consistency in facial defect removal, improves image quality and reduces work complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339124A_ABST
    Figure CN120339124A_ABST
Patent Text Reader

Abstract

The invention discloses a face defect removal method based on frequency domain repair and multi-resolution fusion, and belongs to the field of computer vision and image processing. Obtaining a Laplacian image group, a high-frequency component and a defect soft mask of the original face image; inputting the Laplacian image group and the defect soft mask into a face defect removal network composed of an encoder and a decoder, wherein the encoder and the decoder comprise a frequency domain perception dynamic aggregation module, a spatial domain projection module and a frequency domain selection feedforward network module; and through step-by-step coding and decoding, flaw-removed face images with different resolutions are generated. And through a multi-resolution fusion module, images with different resolutions and corresponding high-frequency components are connected in a cross-layer manner from the face image with the defect removed at the lowest resolution until a fusion image with the highest resolution is obtained, namely, the final face image with the defect removed is obtained. According to the method, the original features of the human face are effectively reserved while the flaws are removed, and the problems of detail fidelity and global consistency in the process of removing the facial flaws are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer vision and image processing, and particularly relates to a method for removing facial blemishes based on frequency domain repair and multi-resolution fusion. Background Art

[0002] Facial blemish removal is an important task in the fields of image processing and computer vision, and is particularly widely applied in the digital cultural industry, such as social media, virtual character generation, and online live streaming. However, this task faces many challenges, including the deficiencies of existing models in global consistency and detail restoration. Although traditional filtering and regularization editing methods can remove blemishes and smooth the skin, they require manual parameter adjustment and often result in under-smoothing or over-smoothing, losing skin texture and details. In the field of professional photography, although manual retouching is precise, it is time-consuming and requires high skills. Therefore, developing an efficient automated facial blemish removal technology that can retain details has become a key issue.

[0003] With the rapid development of computer hardware and software technologies, the application of image processing in the field of portraits has become increasingly mature and widespread. Many scholars have conducted in-depth research on facial blemish removal technologies. These studies have not only improved the quality and aesthetics of portraits, but also provided more efficient technical means for photographers, reducing the work complexity. As one of the main means of automated portrait processing, especially the application and development of generative adversarial networks (GANs) and Transformer architectures, more reliable means are available for portrait segmentation, background replacement, and blemish removal, enabling efficient and accurate image processing effects. After the portrait is processed to remove blemishes, the blemish areas in the image can be removed and beautified with high quality, thus helping users to perform subsequent editing and beautification work more precisely.

[0004] However, when existing methods solve the problem of blemish removal, there is still room for optimization in terms of global consistency and detail restoration. Therefore, proposing an efficient method for removing facial blemishes has important practical significance and application value for promoting technological progress in this field and meeting actual application requirements. Summary of the Invention

[0005] In order to solve the problems of global consistency and detail restoration in existing facial blemish removal technologies, the present invention proposes a method for removing facial blemishes based on frequency domain repair and multi-resolution fusion, and adopts the following technical solutions:

[0006] In a first aspect, the present invention proposes a method for removing facial blemishes based on frequency domain repair and multi-resolution fusion, which is characterized by including the following steps:

[0007] Obtain the Laplacian image group of the original facial image, its corresponding high-frequency components, and the blemish soft mask;

[0008] The Laplacian image group and the defect soft mask are used as the inputs of a face defect removal network composed of an encoder and a decoder. A frequency domain perception dynamic aggregation module, a spatial domain projection module, and a frequency domain selection feed-forward network module are introduced into the encoder and the decoder. After the input Laplacian image group and defect soft mask are encoded and decoded step by step, defect-free face images with different resolutions are generated.

[0009] The multi-resolution fusion module starts from the defect-free face image with the lowest resolution, cross-layer connects the defect-free face images with different resolutions and the corresponding high-frequency components until the highest-resolution fusion image is obtained as the final defect-free face image.

[0010] Further, the defect soft mask is generated by a defect soft mask generation network based on Unet.

[0011] Further, in the face defect removal network, the first layer of the encoder is a convolutional layer, and the remaining layers all include a frequency domain perception dynamic aggregation module, a spatial domain projection module, and a frequency domain selection feed-forward network module. Starting from the third layer, each layer further includes a downsampling layer before the frequency domain perception dynamic aggregation module.

[0012] Further, in the face defect removal network, the last layer is a convolutional layer, and the remaining layers all include a frequency domain perception dynamic aggregation module, a spatial domain projection module, and a frequency domain selection feed-forward network module. Starting from the third layer from the bottom, each layer further includes an upsampling layer after the frequency domain selection feed-forward network module, and skip connections are introduced between the decoders and the encoders of each layer.

[0013] Further, the frequency domain perception dynamic aggregation module is used to dynamically select and aggregate the frequency domain features of the input feature map by using a frequency domain perception dynamic aggregation strategy, and retain key information. Its calculation process includes:

[0014] (1-1) Given the input feature map z, calculate the Fourier transform result F(z)(u,v) channel by channel, extract the real part and the imaginary part, and calculate the amplitude spectrum and the phase spectrum of each channel.

[0015] (1-2) Aggregate the amplitude spectra and phase spectra of different channels through pointwise convolution:

[0016]

[0017] where F(z)(u,v) ∈ {A(z)(u,v), Φ(z)(u,v)}, A(z)(u,v) and Φ(z)(u,v) are the amplitude spectrum and the phase spectrum respectively, F(z)(u,v) is the Fourier transform result of the feature map z, G is the GEGLU activation function, F * (z c)(u, v) is the aggregated Fourier transform result, is a convolution operation with a convolution kernel of 1 and a stride of 1, M pred is the defect soft mask;

[0018] (1 - 3) introduces a learnable quantization matrix W A and W Φ , respectively perform weighting on the amplitude spectrum and phase spectrum of the aggregated Fourier transform result, and then recombine to obtain a new Fourier transform result, which is then added to the Fourier transform result in step (1 - 1) as the filtered Fourier transform result

[0019] (1 - 4) Calculate the real part according to the following formula and the imaginary part

[0020]

[0021] where, are respectively the amplitude spectrum and phase spectrum.

[0022] (1 - 5) Perform an inverse Fourier transform operation on the real part and imaginary part in (1 - 4) to obtain the finally remapped Fourier transform result as the spectral dynamic aggregation feature.

[0023] Furthermore, the spatial domain projection module is used to perform layer normalization on the output feature map of the frequency domain perception dynamic aggregation module, and then project the normalized features to obtain the query vector Q, key vector K, and value vector V, calculate the attention and perform a residual connection with the input feature map z of the frequency domain perception dynamic aggregation module in the corresponding layer to obtain the output feature z':

[0024]

[0025] where, Λ represents the channel number of the feature, B represents the position bias matrix, and the superscript T represents the transpose.

[0026] Furthermore, the calculation formula of the frequency domain selection feed-forward network module is as follows:

[0027]

[0028] where, L(z') represents the normalization operation on z', P(·) and P -1 (·) respectively represent the patch unfolding and folding operations, F, F -1 respectively represent the Fourier transform and the inverse Fourier transform, W s represents the learnable weight, are respectively two intermediate results, z iut$O$ is the output feature of the frequency domain selective feed-forward network module, and $G$ represents the GEGLU activation function.

[0029] Furthermore, the calculation formula of the multi-resolution fusion module is:

[0030]

[0031] where, $O$ l and $O$ l-1 represent the fusion result, and $O$ l = $R$ l , $R$ l and $R$ l-1 respectively correspond to the de-noised face images with different resolutions generated by the Laplacian image $I$ l , $I$ l-1 passing through the face de-noising network. $H$ l-1 is the high-frequency component of the Laplacian image $I$ l-1 . up represents bilinear interpolation, Cat represents channel concatenation, represents the convolution operation with a convolution kernel of 3 and a stride of 3, and $\sigma$ represents the leaky ReLU function; when $l = 1$, $O$ l-1 = $O_0$, that is, the final de-noised face image.

[0032] In the second aspect, the present invention proposes a face flaw removal system based on frequency domain repair and multi-resolution fusion, which is used to implement the above-mentioned face flaw removal method based on frequency domain repair and multi-resolution fusion.

[0033] Beneficial effects of the present invention:

[0034] The method proposed by the present invention has significant advantages in the face flaw removal task, especially in terms of detail preservation and global consistency. This method combines frequency domain aware dynamic aggregation (FADA), spatial domain projection (SDP), and frequency domain selective feed-forward network (SFFN) modules. By means of the joint processing of the frequency domain and the spatial domain, it effectively solves the problems of detail fidelity and global consistency in facial flaw removal. The multi-resolution fusion module (MRF) further optimizes the flaw repair effect, and has obvious advantages in small flaw removal and large area flaw repair. Compared with other methods, the method of the present invention can effectively retain the original features of the face during the process of removing flaws, and has high practical value and research significance. Through the qualitative result comparison with different methods, it is clearly visible that the flaw removal effect of the present invention is the best, and it successfully overcomes the problems of detail fidelity and global consistency in facial flaw removal. Brief Description of the Drawings

[0035] Figure 1 is the overall network architecture diagram of the present invention;

[0036] Figure 2It is the architecture diagram of a face defect removal network composed of an encoder and a decoder;

[0037] Figure 3 It is the architecture diagram of the Frequency-domain Aware Dynamic Aggregation (FADA) module and the Spatial Domain Projection (SDP) module;

[0038] Figure 4 It is the architecture diagram of the Frequency-domain Selective Feed-forward Network (SFFN) module;

[0039] Figure 5 It is the comparison diagram of the repair effects, comparing the differences between the traditional method and the present invention in terms of detail retention and color consistency. Detailed implementation manners

[0040] The following is a detailed description of the preferred implementation manner of the present invention through selected implementation cases and in combination with the accompanying drawings, aiming to deeply explain the core idea of the present invention, rather than aiming to set any form of limitation. Various flexible innovations and adjustments are encouraged on the premise of maintaining the basic principles and spirit of the present invention, including but not limited to fine-tuning of the structure, optimization of functions, and expansion of application scenarios, etc. All these variants and innovations are regarded as within the protection scope of the present invention, reflecting the broad inclusion and respect for the boundaries of technological innovation.

[0041] Through experiments and research, the applicant proposed a method based on frequency-domain selection and repair (FSR), which combines the Frequency-domain Aware Dynamic Aggregation (FADA) module, the Spatial Domain Projection (SDP) module, and the Frequency-domain Selective Feed-forward Network (SFFN) module, and introduces a Multi-Resolution Fusion (MRF) module, effectively solving the deficiencies of the prior art in terms of global consistency and detail restoration. Through extensive experiments on the synthetic dataset HGFR and the real dataset FFHQR, the results show that the proposed method is superior to the prior art in terms of indicators such as PSNR, SSIM, and LPIPS, proving its excellent performance in the face defect removal task.

[0042] As Figure 1 shown, a face defect removal method based on frequency-domain dynamic selection and multi-resolution fusion proposed by the present invention mainly includes the following steps:

[0043] S1. Obtain the Laplacian image group of the original face image, its corresponding high-frequency components, and the defect soft mask;

[0044] Taking the input face image I as an example, its size is h×w, where h and w are the height and width of the face image respectively;

[0045] First, obtain the Laplacian image group, L I =[I0, I1,..., I l , and the corresponding high-frequency components are represented as LH = [H0, H1,..., H l-1 , where I0 and H0 represent the Laplacian image of the highest resolution and its high-frequency component, I l-1 , H l-1 represent the Laplacian image corresponding to the original face image after l - 1 times of downsampling and its high-frequency component, L H is obtained according to the Laplacian pyramid, and l is the number of downsamplings (in this embodiment, it is defaulted to l = 2), I l is used to obtain H l-1 .

[0046] S2. Take the Laplacian image group and the defect soft mask as the input of the face defect removal network composed of an encoder and a decoder. As Figure 2 shown, a frequency domain perception dynamic aggregation module, a spatial domain projection module, and a frequency domain selection feedforward network module are introduced into the encoder and decoder described above; the input Laplacian image group and the defect soft mask generate defect-free face images of different resolutions after being encoded and decoded level by level;

[0047] Here, in the face defect removal network, it includes performing a Fourier transform on the input image to extract frequency domain information; using the frequency domain perception dynamic aggregation (FADA) strategy to dynamically select and aggregate frequency domain features and retain key information; using the spatial domain projection (SDP) module to perform spatial domain processing using the inverse Fourier transform to ensure detail retention and global consistency, etc. The implementation details will be described later.

[0048] S3. Use a multi-resolution fusion module to start from the defect-free face image of the lowest resolution, cross-layer connect the defect-free face images of different resolutions and their corresponding high-frequency components until the highest resolution fusion image is obtained, which is used as the final defect-free face image. Through experimental verification, the present invention effectively retains facial details while removing face defects and improves the overall quality of the image.

[0049] The functions and implementation processes of each model are introduced separately below.

[0050] (I) Frequency Domain Perception Dynamic Aggregation (FADA) Module

[0051] (1) Given a feature map z ∈ R HxWxC , where H, W, and C represent the height, width, and number of channels of the feature map respectively. First, apply the Fourier transform DFT to each channel z c to map it from the spatial domain to the frequency domain. Specifically, the Fourier transform result F(z c )(u, v) of a single channel is as follows:

[0052]

[0053] Among them, (h, w) and (u, v) represent the coordinates in the spatial domain and the frequency domain respectively, c ∈ {0, 1,..., C - 1} is the channel index, and z C represents the feature of the c-th channel, and C represents the number of channels.

[0054] Then, the real part is extracted from the Fourier transform result F(z c )(u, v) of each channel and the imaginary part Based on this information, the amplitude spectrum A(z c )(u, v) and the phase spectrum Φ(z c )(u, v) of a single channel can be further calculated:

[0055]

[0056] By applying equations (1), (2), and (3) to each channel in turn, the amplitude spectrum A(z)(u, v) ∈ R HxWxC and the phase spectrum Φ(z)(u, v) ∈ R HxWxC can be obtained.

[0057] (2) The self-attention mechanism used in most previous transformer-based methods tends to globally focus on the feature content of the input image, which is not suitable for face modification tasks that require defect removal because these regions usually have higher similarity to each other. The dynamic spectral filter can process defects with different frequency features such as moles, freckles, and acne in the frequency domain. The amplitude spectra and phase spectra of different channels are aggregated by pointwise convolution as:

[0058]

[0059] where F(z)(u, v) ∈ {A(z)(u, v), Φ(z)(u, v)}; G is the GEGLU activation function, and F * (z c )(u, v) is the aggregated Fourier transform result, is a convolution operation with a convolution kernel of 1 and a stride of 1, and M pred is the defect soft mask. To discriminatively focus on the frequency domain information, the present invention introduces two learnable quantization matrices W A and W Φ , to weight the amplitude spectrum and phase spectrum components respectively:

[0060]

[0061] where ⊙ represents dot multiplication, A′(z c )(u, v), Φ′(z c)(u, v) represent the magnitude spectrum and phase spectrum after weighting in channel c respectively. After dynamically learning these parameters in the frequency domain, they are recombined to obtain new frequency domain features:

[0062]

[0063] Among them, Cat represents the concatenation operation along the channel. The filtering operation is performed by a residual connection, and the filtered component is obtained by the following formula:

[0064]

[0065] Then, the real part and imaginary part are obtained by the following formula:

[0066]

[0067] Among them, are respectively the magnitude spectrum and phase spectrum.

[0068] After the dynamic parameter learning in the frequency domain, the feature map is remapped to the spatial domain:

[0069]

[0070] Among them, represents the remapped feature, and F -1 represents the inverse Fourier transform.

[0071] The Fourier transform and inverse Fourier transform can be implemented using the DFT and IDFT algorithms. Here, the spectral dynamic aggregation (FADA) is defined from Eq.1 to Eq.9. For convenience, FADA is denoted as FA(·).

[0072] (II) Spatial Domain Projection (SDP) Module

[0073] After the processing in the frequency domain, continue to focus on the processing in the spatial domain.

[0074] Specifically, for the input feature z ∈ R HxWxC , it is processed by the layer normalization operation L(·) to obtain the normalized feature L(z). Then, the normalized feature L(z) is projected into Q (query), K (key), and V (value) through spatial domain projection:

[0075]

[0076] Among them, FA Q (·), FA K (·), and FA V(·) respectively represent three independent projection operations with learnable parameters. The generation processes of Q, K, and V are represented as spatial domain projection (SDP). Finally, the output features of SDP are generated as follows:

[0077]

[0078] where Λ represents the channel number of the features, B represents the position bias matrix, and the element b in the matrix ij = f(i - j), and f is a function for encoding relative position information. z′ is considered as a residual, and is added to z before feeding the result z′ into the subsequent frequency selection feed-forward network (SFFN).

[0079] The above frequency domain aware dynamic aggregation (FADA) module and spatial domain projection (SDP) module are as Figure 3 shown.

[0080] (III) Frequency Selection Feed-Forward Network (SFFN) Module

[0081] FFN (feed-forward network) enhances the feature representation ability by applying scaled dot-product attention. Therefore, it is crucial to develop an effective FFN to generate features beneficial for the potential portrait de-blemishing task. Considering that not all low-frequency and high-frequency information is significant for the removal of potential portrait blemishes, the present invention proposes an adaptive frequency selection feed-forward network (SFFN) that can adaptively determine which frequency components should be retained to the next layer. To effectively identify which frequency information is most important for the de-blemishing task, inspired by the JPEG compression algorithm, the present invention proposes to introduce a learnable quantization matrix W s , and learn through the inverse process of JPEG compression to dynamically determine which frequency information to retain.

[0082] On this basis, as Figure 4 shown, the SFFN (dynamic feed-forward network) proposed by the present invention can be represented by the following formula:

[0083]

[0084] where L(z′) represents the normalization operation on z′, P(·) and P -1 (·) respectively represent the patch unfolding and folding operations in the JPEG compression method, F, F -1 respectively represent the Fourier transform and the inverse Fourier transform, W s represents the learnable weight, are two intermediate results respectively, z out is the final output result, and G represents the GEGLU activation function.

[0085] The above-mentioned Frequency Domain Aware Dynamic Aggregation (FADA) module, Spatial Domain Projection (SDP) module, and Frequency Domain Selective Feed-Forward Network (SFFN) module constitute the Frequency Domain Selection and Restoration-based method (FSR).

[0086] (4) Multi-Resolution Fusion (MRF) module

[0087] In addition, considering that high-resolution images have advantages in dealing with small defects, while low-resolution images perform better in restoring large-area defects. Based on this consideration, the present invention proposes a Multi-Resolution Fusion module (MRF), which generates a flawless and detail-rich face image through feature fusion of multi-resolution results. It effectively complements low-resolution portraits with high-resolution portraits and compensates for the information loss caused by downsampling. Since the low-resolution portrait is obtained by downsampling the high-resolution one, it lacks the information of high-frequency components.

[0088] On this basis, the present invention uses the high-frequency components of the image as an additional input to enhance the restoration of high-frequency details. Based on the above observations, the present invention designs the MRF to efficiently restore image details and refine local details while maintaining global consistency:

[0089]

[0090] where O represents the fusion result, and O l = R l . up represents bilinear interpolation, Cat represents channel concatenation, H i (i ∈ {0, 1,..., l - 1}) is the high-frequency component of the image I i , represents a convolution operation with a convolution kernel of 3 and a stride of 3, and σ represents leakyReLU with a negative slope of 0.2. Given the input, it first passes through the Transformers T to obtain a multi-resolution output. Through continuous upsampling and refinement of the hybrid layer, a high-resolution defect-free portrait with detailed transformation information is then obtained.

[0091] For the training objective, the present invention gives the training sample as a face image with defects, and its label is the defect-free face image, and uses the L1 loss, perceptual loss, and adversarial loss to train the above model.

[0092] Among them, the L1 loss measures the error of the model by calculating the absolute difference between the generated image and the label image.

[0093] The perceptual loss measures the image quality by comparing the differences between the generated image obtained by the model of the present invention and the label image in the feature space, rather than directly comparing the pixel values. In this embodiment, a pre-trained convolutional neural network (such as VGG) is used to extract the high-level features of the image.

[0094] The adversarial loss is derived from the generative adversarial network (GAN). The discriminator network is used to distinguish between the generated images and the labeled images obtained by the model of the present invention. The loss function here is well-known in the art and will not be elaborated. By analyzing the deficiencies of the prior art, the present invention designs a two-stage face defect removal network, including a frequency domain selection and repair network (FSR) and a multi-resolution fusion module (MRF). The FSR includes a frequency domain aware dynamic aggregation (FADA), a spatial domain projection (SDP) module, and a frequency domain selection-based feedforward network (SFFN). The MRF uses a technique similar to the Laplacian pyramid to achieve efficient fusion of multi-resolution features and improve the final defect repair effect. The present invention is compared with the existing Pix2PixHD, AutoRetouch, ABPN, BPFRe, and RetouchFormer methods. Pix2PixHD is an image-to-image conversion network, and the rest are specifically designed for face retouching. The above models are trained on the datasets HGFR, FFHQR, and the combination of the two datasets (i.e., FFHQR+HGFR). Among them, HGFR contains 23,760 pairs of generated face images, and FFHQR contains 70,000 pairs of natural face images. Feature similarity (FSIM), structural similarity (SSIM), peak signal-to-noise ratio (PSNR), and learned perceptual image patch similarity (LPIPS) are used for analysis. As shown in Table 1, where "↑" and "↓" respectively indicate that the higher the value of the index, the better, and the lower the value of the index, the better. The bold index values represent the optimal results under the same dataset. It can be seen from this that the performance of the present invention is the best compared with other technologies. A more intuitive effect is as Figure 5 shown.

[0095] Table 1

[0096]

[0097] Based on the same inventive concept, in this embodiment, a face defect removal system based on frequency domain repair and multi-resolution fusion is also provided, and this system is used to implement the above embodiment. Terms such as "module" and "unit" used hereinafter can be a combination of software and / or hardware that can achieve a predetermined function. Although the system described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible.

[0098] In this embodiment, a face defect removal system based on frequency domain repair and multi-resolution fusion includes:

[0099] An image preprocessing module, which is used to obtain the Laplacian image group of the original face image, its corresponding high-frequency components, and the defect soft mask;

[0100] A face defect removal network module is composed of an encoder and a decoder, and takes a Laplacian image group and a defect soft mask as input. The encoder and the decoder introduce a frequency domain perception dynamic aggregation module, a spatial domain projection module, and a frequency domain selection feedforward network module; the input Laplacian image group and defect soft mask are encoded and decoded step by step to generate defect-free face images of different resolutions;

[0101] The multi-resolution fusion module is used to start from the lowest resolution blemish-free face image, cross-layer connect blemish-free face images of different resolutions and corresponding high-frequency components until the highest resolution fused image is obtained as the final blemish-free face image.

[0102] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment, and the implementation methods of the remaining modules will not be repeated here. The system embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.

[0103] The embodiments of the system of the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The system embodiments can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, the corresponding computer program instructions in the non-volatile memory are read into the memory by the processor of any device with data processing capabilities and run.

[0104] The above examples are only specific embodiments of the present invention. Obviously, the present invention is not limited to the above examples, and many variations are possible. All variations that can be directly derived or associated with the contents disclosed by a person skilled in the art should be considered as the protection scope of the present invention.

Claims

1. A method for removing facial blemishes based on frequency domain repair and multi-resolution fusion, characterized in that, The steps include: Obtaining a Laplacian image group of the original face image, its corresponding high-frequency components, and a defect soft mask; Using the Laplacian image group and the defect soft mask as the input of a face defect removal network composed of an encoder and a decoder, and introducing a frequency domain perception dynamic aggregation module, a spatial domain projection module, and a frequency domain selection feedforward network module into the encoder and decoder; The input Laplacian image group and defect soft mask generate defect-removed face images of different resolutions after hierarchical encoding and decoding; Using a multi-resolution fusion module to start from the defect-removed face image with the lowest resolution, cross-layer connect defect-removed face images of different resolutions and their corresponding high-frequency components until a fused image with the highest resolution is obtained as the final defect-removed face image.

2. The method for removing facial blemishes based on frequency domain repair and multi-resolution fusion according to claim 1, wherein, The Laplacian image group of the original face image is denoted as L I =[I0, I1, ..., I l , and the corresponding high-frequency components are denoted as L H =[H0, H1, ..., H l-1 , where I 0i , H0 represent the Laplacian image with the highest resolution and its high-frequency components, I l-1 , H l-1 represent the Laplacian image and its high-frequency components corresponding to the original face image after l - 1 times of downsampling, and I l is used to obtain H l-1 .

3. The method for removing facial blemishes based on frequency domain repair and multi-resolution fusion according to claim 1, characterized in that The defect soft mask is generated by a defect soft mask generation network based on Unet.

4. The method for removing facial blemishes based on frequency domain repair and multi-resolution fusion according to claim 1, wherein In the face defect removal network, the first layer of the encoder is a convolutional layer, and the remaining layers all include a frequency domain perception dynamic aggregation module, a spatial domain projection module, and a frequency domain selection feedforward network module; starting from the third layer, each layer also includes a downsampling layer before the frequency domain perception dynamic aggregation module.

5. The method for removing facial blemishes based on frequency domain repair and multi-resolution fusion according to claim 1, characterized in that In the face defect removal network, the last layer is a convolutional layer, and the remaining layers all include a frequency domain perception dynamic aggregation module, a spatial domain projection module, and a frequency domain selection feedforward network module; starting from the third layer from the bottom, each layer also includes an upsampling layer after the frequency domain selection feedforward network module, and skip connections are introduced between the decoder and encoder of each layer.

6. The method for removing facial defects based on frequency domain repair and multi-resolution fusion according to claim 4 or 5, wherein The frequency domain perception dynamic aggregation module is used to dynamically select and aggregate the frequency domain features of the input feature map using a frequency domain perception dynamic aggregation strategy, and retain key information; Its calculation process includes: (1-1) Given the input feature map z, calculate the Fourier transform result F(z)(u,v) channel by channel, extract the real part and the imaginary part, and calculate the amplitude spectrum and phase spectrum of each channel; (1-2) Aggregate the amplitude spectra and phase spectra of different channels through pointwise convolution: Among them, F(z)(u, v) ∈ {A(z)(u, v), Φ(z)(u, v)}, where A(z)(u, v) and Φ(z)(u, v) are the amplitude spectrum and the phase spectrum respectively, and F(z)(u, v) is the result of the Fourier transform of the feature map z. G is the GEGLU activation function, and F * (z c )(u, v) is the result of the aggregated Fourier transform. is a convolution operation with a convolution kernel of 1 and a stride of 1. M pred is the defect soft mask; (1-3) Introduce a learnable quantization matrix W A and W Φ , respectively perform weighting on the amplitude spectrum and phase spectrum of the aggregated Fourier transform result, then recombine to obtain a new Fourier transform result, and then add it to the Fourier transform result in step (1-1) as the filtered Fourier transform result (1-4) Calculate the real part according to the following formula and the imaginary part Among them, are respectively the amplitude spectrum and the phase spectrum; (1-5) Perform an inverse Fourier transform operation on the real part and the imaginary part in (1-4) to obtain the remapped final Fourier transform result as the spectral dynamic aggregation feature.

7. The method for removing facial blemishes based on frequency domain repair and multi-resolution fusion according to claim 4 or 5, characterized in that The spatial domain projection module is used to perform layer normalization on the output feature map of the frequency domain perception dynamic aggregation module, and then project the normalized feature through the spatial domain to obtain a query vector Q, a key vector K, and a value vector V, calculate the attention, and perform a residual connection with the input feature map z of the frequency domain perception dynamic aggregation module in the corresponding layer to obtain the output feature z'; Among them, Λ represents the channel number of the feature, B represents the position bias matrix, and the superscript T represents the transpose.

8. The method for removing facial blemishes based on frequency domain repair and multi-resolution fusion according to claim 7, characterized in that, The calculation formula of the frequency domain selection feedforward network module is as follows: Among them, L(z′) represents the normalization operation on z′, P(·) and P -1 (·) represent the patch unfolding and folding operations respectively, F, F -1 represent the Fourier transform and the inverse Fourier transform respectively, W s represents the learnable weight, are two intermediate results respectively, z out is the output feature of the frequency domain selective feed-forward network module, and G represents the GEGLU activation function.

9. The method for removing facial blemishes based on frequency domain repair and multi-resolution fusion according to claim 1, characterized in that, The calculation formula of the multi-resolution fusion module is: Among them, O l and O l-1 represent the fusion results, and O l = R l , R l and R l-1 correspond to the de - flawed face images with different resolutions generated by the Laplacian image I l and I l-1 respectively after passing through the face flaw - removal network. H l-1 is the high - frequency component of the Laplacian image I l-1 . up represents bilinear interpolation, Cat represents channel concatenation, represents a convolution operation with a convolution kernel of 3 and a stride of 3, and σ represents the leakyReLU function; when l = 1, O l-1 = O0, that is, the final de - flawed face image.

10. A face flaw removal system based on frequency domain repair and multi-resolution fusion for implementing the method described in claim 1, characterized in that, The system includes: An image preprocessing module, which is used to obtain a Laplacian image group of the original face image, its corresponding high-frequency components, and a defect soft mask; The face defect removal network module is composed of an encoder and a decoder, and takes a Laplacian image group and a defect soft mask as inputs. A frequency-domain perception dynamic aggregation module, a spatial-domain projection module, and a frequency-domain selection feedforward network module are introduced into the encoder and the decoder. The input Laplacian image group and defect soft mask are encoded and decoded step by step to generate defect-removed face images with different resolutions. The multi-resolution fusion module is used to start from the defect-removed face image with the lowest resolution, cross-layer connect defect-removed face images with different resolutions and their corresponding high-frequency components until a fused image with the highest resolution is obtained as the final defect-removed face image.