Face shadow elimination method and device based on high and low frequency attention fusion mechanism
Through the face shadow removal model based on the high and low frequency attention fusion mechanism, and using global lighting information to process the face image, the problem of difficulty in removing large areas of shadows in the existing technology is solved, and high-quality shadow removal and detail retention are achieved.
Patent Information
- Application Number
- CN202510026902.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to effectively deal with the problem of large-area shadows, especially when the face is full shadows, and it is impossible to effectively remove shadows, affecting the image quality and the accuracy of downstream tasks.
The face shadow removal model based on the high and low frequency attention fusion mechanism is adopted, and the face image is processed through the encoder and decoder of the pre-trained face shadow removal model, and the facial area is re-illuminated using global lighting information.
While retaining facial details, it effectively removes large areas of shadows, improves image quality, and enhances the accuracy of downstream tasks such as face recognition and editing.
Smart Images

Figure CN120107114A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of deep learning and image processing technology, and in particular to a method and device for removing face shadows based on a high- and low-frequency attention fusion mechanism. Background Art
[0002] When the light source is blocked by an object or some parts of the face, shadows will appear in the captured face image. The presence of facial shadows will not only reduce the image quality and affect the aesthetics, but also restrict the accuracy and effectiveness of various downstream tasks related to the face, such as face recognition, face editing, and face reconstruction. Therefore, it is very meaningful to restore the illumination of the shadow area of the face image to improve the visibility and aesthetics of the image.
[0003] At present, most face shadow removal algorithms are implemented through deep learning technology. Although existing methods have achieved good results, there are still some problems that need to be solved. For example, some methods use the symmetry of the face to improve the effect of shadow removal, but the method cannot solve the case where the shadow is also symmetrical. In addition, some methods use the lighting information of the shadow-free area to help restore the lighting of the shadow area, but for large-area shadows, especially when the face is completely shadowed, this method cannot effectively remove the shadow and needs to be solved urgently. Summary of the invention
[0004] The present application provides a method and device for removing face shadows based on a high- and low-frequency attention fusion mechanism to solve the problem that the prior art cannot effectively handle large-area shadows, thereby achieving full use of global illumination information to re-illuminate the facial area while retaining facial details.
[0005] The first aspect of the present application provides a method for removing face shadows based on a high- and low-frequency attention fusion mechanism, comprising the following steps:
[0006] Obtain the face image to be processed;
[0007] Inputting the face image to be processed into a pre-trained face shadow removal model to obtain a de-shadowed face image after being processed by an encoder and a decoder of the pre-trained face shadow removal model;
[0008] Among them, the pre-trained face shadow removal model is obtained by training a face shadow removal model with a high- and low-frequency attention fusion mechanism based on a preset joint loss function using a face shadow dataset. The encoder is composed of multiple feature processing modules and downsampling modules alternately connected, and the decoder is composed of multiple feature recovery modules and upsampling modules alternately connected.
[0009] According to one embodiment of the present application, before inputting the face image to be processed into the pre-trained face shadow removal model, the method further includes:
[0010] Acquire the face shadow dataset, wherein the face shadow dataset includes multiple groups of face shadow image-clean image pairs;
[0011] Based on the face shadow dataset, training the face shadow removal model with high- and low-frequency attention fusion mechanism to obtain an initial face shadow removal model;
[0012] Based on the preset joint loss function, when it is determined that the initial face shadow removal model meets the preset training completion condition, the initial face shadow removal model is used as the pre-trained face shadow removal model.
[0013] According to one embodiment of the present application, the step of inputting the face image to be processed into a pre-trained face shadow removal model to obtain a de-shadowed face image after being processed by an encoder and a decoder of the pre-trained face shadow removal model includes:
[0014] Based on the encoder, data preprocessing is performed on the face image to be processed, the face image to be processed is converted from a color space to a feature space, and feature extraction is performed on the face image to be processed after being converted to the feature space to obtain a feature extraction result, and feature reconstruction is performed on the feature extraction result to obtain a target feature representation;
[0015] Based on the decoder, data post-processing is performed on the target feature representation, and the target feature representation is converted from feature space to color space to obtain the de-shadowed face image.
[0016] According to one embodiment of the present application, the preset joint loss function is:
[0017] L total =α 1 L pixel +α 2 L perceptual +α 3 L mutilscale ;
[0018] Among them, L total is the preset joint loss function, α 1 is the weight of pixel-level loss, α 2 is the weight of perceptual loss, α 3 is the weight of multi-scale perceptual loss, L pixel is the pixel-level loss, L perceptual is the perceptual loss, L multiscale It is a multi-scale perceptual loss.
[0019] According to the face shadow removal method based on the high-low frequency attention fusion mechanism of the embodiment of the present application, the face image to be processed is input into the pre-trained face shadow removal model, so as to obtain the de-shadowed face image after being processed by the encoder and decoder of the pre-trained face shadow removal model. Thus, the problem that the prior art cannot effectively process large-area shadows is solved, and the face area is fully illuminated by making full use of global illumination information while retaining facial details.
[0020] The second aspect of the present application provides a face shadow removal device based on a high- and low-frequency attention fusion mechanism, comprising:
[0021] An acquisition module, used for acquiring a face image to be processed;
[0022] A processing module, used for inputting the face image to be processed into a pre-trained face shadow removal model, so as to obtain a de-shadowed face image after being processed by an encoder and a decoder of the pre-trained face shadow removal model;
[0023] Among them, the pre-trained face shadow removal model is obtained by training a face shadow removal model with a high- and low-frequency attention fusion mechanism based on a preset joint loss function using a face shadow dataset. The encoder is composed of multiple feature processing modules and downsampling modules alternately connected, and the decoder is composed of multiple feature recovery modules and upsampling modules alternately connected.
[0024] According to one embodiment of the present application, before inputting the face image to be processed into the pre-trained face shadow removal model, the processing module is further used to:
[0025] Acquire the face shadow dataset, wherein the face shadow dataset includes multiple groups of face shadow image-clean image pairs;
[0026] Based on the face shadow dataset, training the face shadow removal model with high- and low-frequency attention fusion mechanism to obtain an initial face shadow removal model;
[0027] Based on the preset joint loss function, when it is determined that the initial face shadow removal model meets the preset training completion condition, the initial face shadow removal model is used as the pre-trained face shadow removal model.
[0028] According to one embodiment of the present application, the processing module is used to:
[0029] Based on the encoder, data preprocessing is performed on the face image to be processed, the face image to be processed is converted from a color space to a feature space, and feature extraction is performed on the face image to be processed after being converted to the feature space to obtain a feature extraction result, and feature reconstruction is performed on the feature extraction result to obtain a target feature representation;
[0030] Based on the decoder, data post-processing is performed on the target feature representation, and the target feature representation is converted from feature space to color space to obtain the de-shadowed face image.
[0031] According to one embodiment of the present application, the preset joint loss function is:
[0032] L total =α 1 L pixel +α 2 L perceptual +α 3 L mutilscale ;
[0033] Among them, L total is the preset joint loss function, α 1 is the weight of pixel-level loss, α 2 is the weight of perceptual loss, α 3 is the weight of multi-scale perceptual loss, L pixel is the pixel-level loss, L perceptual is the perceptual loss, L multiscale It is a multi-scale perceptual loss.
[0034] According to the face shadow removal device based on the high-low frequency attention fusion mechanism of the embodiment of the present application, the face image to be processed is input into the pre-trained face shadow removal model, so as to obtain the de-shadowed face image after being processed by the encoder and decoder of the pre-trained face shadow removal model. Thus, the problem that the prior art cannot effectively process large-area shadows is solved, and the face area is fully illuminated by making full use of global illumination information while retaining facial details.
[0035] The third aspect of the present application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the face shadow removal method based on the high- and low-frequency attention fusion mechanism as described in the above embodiment.
[0036] The fourth aspect of the present application provides a computer-readable storage medium on which a computer program is stored, and the program is executed by a processor to implement the face shadow removal method based on the high- and low-frequency attention fusion mechanism as described in the above embodiment.
[0037] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0039] Figure 1 A flowchart of a method for removing face shadows based on a high- and low-frequency attention fusion mechanism provided according to an embodiment of the present application;
[0040] Figure 2 Schematic diagram of the overall architecture of a face shadow removal model based on high and low frequency attention according to an embodiment of the present application;
[0041] Figure 3 A schematic diagram of the structure of an encoder / decoder layer according to an embodiment of the present application;
[0042] Figure 4 A schematic diagram of a process for constructing a face shadow removal model based on a high- and low-frequency attention fusion mechanism according to an embodiment of the present application;
[0043] Figure 5 Schematic diagram of a block diagram of a face shadow removal device based on a high- and low-frequency attention fusion mechanism according to an embodiment of the present application;
[0044] Figure 6 Schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0045] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0046] The following describes the face shadow removal method and device based on the high-low frequency attention fusion mechanism according to the embodiment of the present application with reference to the accompanying drawings. In view of the problem that large-area shadows cannot be effectively processed mentioned in the above background technology, the present application provides a face shadow removal method based on the high-low frequency attention fusion mechanism. In this method, the face image to be processed is input into a pre-trained face shadow removal model to obtain a de-shadowed face image after being processed by the encoder and decoder of the pre-trained face shadow removal model. Thus, the problem that the prior art cannot effectively process large-area shadows is solved, and the global lighting information can be fully utilized to re-illuminate the facial area, thereby removing facial shadows while effectively retaining facial details.
[0047] Specifically, Figure 1 A flowchart of a method for removing face shadows based on a high- and low-frequency attention fusion mechanism is provided in an embodiment of the present application.
[0048] like Figure 1 As shown, the face shadow removal method based on the high- and low-frequency attention fusion mechanism includes the following steps:
[0049] In step S101, a face image to be processed is obtained.
[0050] The face image to be processed refers to a face image containing shadows.
[0051] Optionally, the embodiment of the present application can obtain the face image to be processed by real photography, or can select the face image to be processed from an existing database, which is not specifically limited here.
[0052] In step S102, Figure 2 As shown, the face image to be processed is input into a pre-trained face shadow removal model, so as to obtain a de-shadowed face image after being processed by the encoder and decoder of the pre-trained face shadow removal model.
[0053] Among them, the pre-trained face shadow removal model of the embodiment of the present application is a Transformer model with a high and low frequency attention fusion mechanism. The pre-trained face shadow removal model is obtained by training the face shadow removal model with a high and low frequency attention fusion mechanism based on a preset joint loss function and using a face shadow dataset. The encoder is composed of multiple feature processing modules and down-sampling modules alternately connected, and the decoder is composed of multiple feature recovery modules and up-sampling modules alternately connected.
[0054] Specifically, if Figure 2As shown, the pre-trained face shadow removal model of the embodiment of the present application is a U-shaped network with an encoder-decoder structure, the encoder is composed of K feature processing modules capable of extracting high- and low-frequency features and down-sampling modules alternately connected, and the decoder is composed of K feature recovery modules and up-sampling modules alternately connected.
[0055] Furthermore, in some embodiments, the face image to be processed is input into a pre-trained face shadow removal model to obtain a de-shadowed face image after being processed by the encoder and decoder of the pre-trained face shadow removal model, including: based on the encoder, performing data pre-processing on the face image to be processed, converting the face image to be processed from color space to feature space, and performing feature extraction on the face image to be processed after conversion to the feature space to obtain a feature extraction result, and performing feature reconstruction on the feature extraction result to obtain a target feature representation; based on the decoder, performing data post-processing on the target feature representation, converting the target feature representation from feature space to color space to obtain a de-shadowed face image.
[0056] Specifically, the pre-trained face shadow removal model of the embodiment of the present application processes the input face image with shadow (i.e., the face image to be processed) into a face image without shadow and outputs it in an end-to-end manner through data preprocessing, feature extraction, feature reconstruction, and data post-processing. Among them, the feature extraction module is composed of N feature extraction layers, which are used to extract high-frequency and low-frequency features in layers; the feature recovery module is composed of N feature recovery layers, which are used to reconstruct high-frequency and low-frequency features in layers.
[0057] Furthermore, the data preprocessing part consists of a convolutional layer and an activation function, which is used to convert the input image x with a shape of (H, W, 3) from the color space to the feature space to obtain the underlying feature embedding X 0 =(H,W,C), where C represents the embedding dimension. Data post-processing consists of convolutional layers. Data post-processing is the inverse process of data pre-processing, which is used to convert the result from feature space to color space to obtain the final output result, that is, the de-shadowed face image y.
[0058] Furthermore, if Figure 3 As shown in the figure, the feature extraction layer and feature recovery layer are composed of a normalization layer, a high- and low-frequency attention calculation module, a residual connection, a normalization layer, an MLP module, and a regularization layer. The formula is expressed as:
[0059] X=X i +DropPath(HiLo(LayerNrom 1 (X i )));
[0060] X i+1=X+DropPath(MLP(LayerNorm 2 (X));
[0061] Among them, X is the feature after attention calculation, DropPath() is the regularization layer, HiLo() is the high-low frequency attention fusion module, LayerNrom 1 () is the normalization layer, X i is the output of the i-th feature extraction layer or the output of the i-th feature recovery layer, X i+1 Represents the output of the i+1th feature extraction layer or the output of the i+1th feature recovery layer. MLP() is a multi-layer perceptron. LayerNrom 2 () is the normalization layer.
[0062] Furthermore, the high- and low-frequency attention calculation module is composed of high-frequency attention calculation, low-frequency attention calculation and gated fusion module. The high-frequency attention calculation calculates qkv through the convolution layer, calculates the attention and performs weighting, and finally outputs it through the mapping layer; the low-frequency attention calculation calculates qkv through the linear layer, and then calculates the attention matrix, and finally outputs it through the mapping layer; the gated fusion module generates the weighted representation of high-frequency and low-frequency features through the linear layer, and calculates the weighted sum, which is the final target feature representation.
[0063] Furthermore, in some embodiments, before the face image to be processed is input into a pre-trained face shadow removal model, it also includes: obtaining a face shadow dataset, wherein the face shadow dataset contains multiple groups of face shadow image-clean image pairs; based on the face shadow dataset, training a face shadow removal model with a high and low frequency attention fusion mechanism to obtain an initial face shadow removal model; based on a preset joint loss function, when it is determined that the initial face shadow removal model meets a preset training completion condition, using the initial face shadow removal model as the pre-trained face shadow removal model.
[0064] Specifically, a face shadow dataset consisting of face shadow images and their corresponding clean images is constructed, where any set of shadow / no-shadow image pairs is a set of training data, denoted as (x, y), where x represents a face image with shadows, and y is the clean image corresponding to x. The constructed face shadow dataset consists of real face shadow data and face shadow data obtained by synthesis.
[0065] Furthermore, based on the face shadow dataset, a face shadow removal model with a high- and low-frequency attention fusion mechanism is trained to obtain an initial face shadow removal model.
[0066] Furthermore, in the embodiment of the present application, when performing end-to-end training on the initial face shadow removal model, a preset joint loss is used as the total loss of the optimization model and parameters, wherein the preset joint loss function is:
[0067] L total =α 1 L pixel +α 2 L perceptual +α 3 L mutilscale ;
[0068] Among them, L total is the preset joint loss function, α 1 is the weight of pixel-level loss, α 2 is the weight of perceptual loss, α 3 is the weight of multi-scale perceptual loss, L pixel is the pixel-level loss, L perceptual is the perceptual loss, L multiscale It is a multi-scale perceptual loss.
[0069] Furthermore, the calculation formula of pixel-level loss is:
[0070] L pixel =||y 0 -y|| 1 ;
[0071] Among them, L pixel is the pixel-level loss, y 0 is the deshading result of the model, and y is the ground truth image.
[0072] Furthermore, the calculation formula of the pixel-level loss function is:
[0073]
[0074] Among them, L perceptual For the perceptual loss, k takes values from layers, where layers are the different layers of the VGG network.
[0075] Furthermore, the calculation formula of the pixel-level loss function is:
[0076]
[0077] Among them, L multiscale is a multi-scale perceptual loss, scales are different downsampling scales, i takes values from scales, y 0,i for y 0 The result of downsampling according to scale i, y i is the result of downsampling y according to scale i.
[0078] For example, assuming that the number of training rounds of the face shadow removal model with a high and low frequency attention fusion mechanism in an embodiment of the present application is 200 rounds, the total loss of each round of training is calculated based on the preset joint loss function, and the parameters of the face shadow removal model are adjusted according to the total loss of each round of training, and training continues until the face shadow removal model completes 200 rounds of training to obtain the final pre-trained face shadow removal model.
[0079] In order to facilitate those skilled in the art to more clearly and intuitively understand the construction process of the face shadow removal model based on the high- and low-frequency attention fusion mechanism in the embodiment of the present application, the following is combined with Figure 4 Provide detailed explanation.
[0080] Specifically, if Figure 4 As shown in the figure, the construction process of the face shadow removal model based on the high- and low-frequency attention fusion mechanism is as follows:
[0081] First, a dataset is constructed for training the face shadow removal model.
[0082] Secondly, build the initial face shadow removal model, including data preprocessing, feature extraction, feature reconstruction and data post-processing.
[0083] Finally, the initial face shadow removal model is trained using the data set until the preset training completion conditions are met, thereby obtaining a pre-trained face shadow removal model.
[0084] According to the face shadow removal method based on the high- and low-frequency attention fusion mechanism of the embodiment of the present application, the face image to be processed is input into the pre-trained face shadow removal model, so as to obtain the de-shadowed face image after being processed by the encoder and decoder of the pre-trained face shadow removal model. Thus, the problem that the prior art cannot effectively handle large-area shadows is solved. The face shadow removal algorithm based on the high- and low-frequency attention fusion mechanism re-illuminates the facial area by extracting high-frequency texture details and low-frequency global illumination respectively, thereby obtaining a more natural and consistent shadow removal result.
[0085] Secondly, the face shadow removal device based on the high- and low-frequency attention fusion mechanism proposed in accordance with an embodiment of the present application is described with reference to the accompanying drawings.
[0086] Figure 5 It is a block diagram of a face shadow removal device based on a high- and low-frequency attention fusion mechanism according to an embodiment of the present application.
[0087] like Figure 5 As shown, the face shadow removal device 10 based on the high- and low-frequency attention fusion mechanism includes: an acquisition module 100 and a processing module 200.
[0088] Among them, the acquisition module 100 is used to acquire the face image to be processed; the processing module 200 is used to input the face image to be processed into a pre-trained face shadow removal model, so as to obtain a de-shadowed face image after processing through the encoder and decoder of the pre-trained face shadow removal model; wherein the pre-trained face shadow removal model is trained by using a face shadow dataset based on a preset joint loss function to train a face shadow removal model with a high- and low-frequency attention fusion mechanism, the encoder is composed of a plurality of feature processing modules and down-sampling modules alternately connected, and the decoder is composed of a plurality of feature recovery modules and up-sampling modules alternately connected.
[0089] Furthermore, in some embodiments, before inputting the face image to be processed into a pre-trained face shadow removal model, the processing module 200 is also used to: obtain a face shadow dataset, wherein the face shadow dataset includes multiple groups of face shadow image-clean image pairs; based on the face shadow dataset, train a face shadow removal model with a high- and low-frequency attention fusion mechanism to obtain an initial face shadow removal model; based on a preset joint loss function, when it is determined that the initial face shadow removal model meets a preset training completion condition, use the initial face shadow removal model as the pre-trained face shadow removal model.
[0090] Furthermore, in some embodiments, the processing module 200 is used to: based on the encoder, perform data preprocessing on the face image to be processed, convert the face image to be processed from the color space to the feature space, and perform feature extraction on the face image to be processed after the conversion to the feature space to obtain a feature extraction result, perform feature reconstruction on the feature extraction result to obtain a target feature representation; based on the decoder, perform data post-processing on the target feature representation, convert the target feature representation from the feature space to the color space to obtain a de-shadowed face image.
[0091] Furthermore, in some embodiments, the preset joint loss function is:
[0092] L total =α 1 L pixel +α 2 L perceptual +α 3 L mutilscale ;
[0093] Among them, L total is the preset joint loss function, α 1 is the weight of pixel-level loss, α 2 is the weight of perceptual loss, α 3 is the weight of multi-scale perceptual loss, L pixel is the pixel-level loss, L perceptualis the perceptual loss, L multiscale It is a multi-scale perceptual loss.
[0094] It should be noted that the aforementioned explanation of the embodiment of the face shadow removal method based on the high- and low-frequency attention fusion mechanism is also applicable to the face shadow removal device based on the high- and low-frequency attention fusion mechanism of this embodiment, and will not be repeated here.
[0095] According to the face shadow removal device based on the high-low frequency attention fusion mechanism of the embodiment of the present application, the face image to be processed is input into the pre-trained face shadow removal model, so as to obtain the de-shadowed face image after being processed by the encoder and decoder of the pre-trained face shadow removal model. Thus, the problem that the prior art cannot effectively handle large-area shadows is solved, and the face area is re-illuminated by obtaining low-frequency global illumination, thereby removing the face shadow while retaining the high-frequency details of the face.
[0096] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:
[0097] A memory 601 , a processor 602 , and a computer program stored in the memory 601 and executable on the processor 602 .
[0098] When the processor 602 executes the program, the face shadow removal method based on the high- and low-frequency attention fusion mechanism provided in the above embodiment is implemented.
[0099] Furthermore, the electronic device further comprises:
[0100] The communication interface 603 is used for communication between the memory 601 and the processor 602 .
[0101] The memory 601 is used to store computer programs that can be executed on the processor 602 .
[0102] The memory 601 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0103] If the memory 601, the processor 602 and the communication interface 603 are implemented independently, the communication interface 603, the memory 601 and the processor 602 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0104] Optionally, in a specific implementation, if the memory 601, the processor 602 and the communication interface 603 are integrated on a chip, the memory 601, the processor 602 and the communication interface 603 can communicate with each other through an internal interface.
[0105] The processor 602 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0106] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for removing face shadows based on the high- and low-frequency attention fusion mechanism.
[0107] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0108] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0109] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A face shadow removal method based on high- and low-frequency attention fusion mechanism, characterized in that: The following steps are involved: Obtain the face image to be processed; Inputting the face image to be processed into a pre-trained face shadow removal model to obtain a de-shadowed face image after being processed by an encoder and a decoder of the pre-trained face shadow removal model; Among them, the pre-trained face shadow removal model is obtained by training a face shadow removal model with a high- and low-frequency attention fusion mechanism based on a preset joint loss function using a face shadow dataset. The encoder is composed of multiple feature processing modules and downsampling modules alternately connected, and the decoder is composed of multiple feature recovery modules and upsampling modules alternately connected.
2. The method according to claim 1, characterized in that Before inputting the face image to be processed into the pre-trained face shadow removal model, the method further includes: Acquire the face shadow dataset, wherein the face shadow dataset includes multiple groups of face shadow image-clean image pairs; Based on the face shadow dataset, training the face shadow removal model with high- and low-frequency attention fusion mechanism to obtain an initial face shadow removal model; Based on the preset joint loss function, when it is determined that the initial face shadow removal model meets the preset training completion condition, the initial face shadow removal model is used as the pre-trained face shadow removal model.
3. The method according to claim 1, characterized in that The step of inputting the face image to be processed into a pre-trained face shadow removal model to obtain a de-shadowed face image after being processed by an encoder and a decoder of the pre-trained face shadow removal model comprises: Based on the encoder, data preprocessing is performed on the face image to be processed, the face image to be processed is converted from a color space to a feature space, and feature extraction is performed on the face image to be processed after being converted to the feature space to obtain a feature extraction result, and feature reconstruction is performed on the feature extraction result to obtain a target feature representation; Based on the decoder, data post-processing is performed on the target feature representation, and the target feature representation is converted from feature space to color space to obtain the de-shadowed face image.
4. The method according to claim 1, characterized in that: The preset joint loss function is: L total =α1L pixel +α2L perceptual +α3L mutilscale ; Among them, L total is the preset joint loss function, α1 is the weight of pixel-level loss, α2 is the weight of perceptual loss, α3 is the weight of multi-scale perceptual loss, and L pixel is the pixel-level loss, L perceptual is the perceptual loss, L multiscale It is a multi-scale perceptual loss.
5. A face shadow removal device based on high- and low-frequency attention fusion mechanism, characterized in that: include: An acquisition module, used for acquiring a face image to be processed; A processing module, used for inputting the face image to be processed into a pre-trained face shadow removal model, so as to obtain a de-shadowed face image after being processed by an encoder and a decoder of the pre-trained face shadow removal model; Among them, the pre-trained face shadow removal model is obtained by training a face shadow removal model with a high- and low-frequency attention fusion mechanism based on a preset joint loss function using a face shadow dataset. The encoder is composed of multiple feature processing modules and downsampling modules alternately connected, and the decoder is composed of multiple feature recovery modules and upsampling modules alternately connected.
6. The apparatus according to claim 5, before inputting the face image to be processed into the pre-trained face shadow removal model, the processing module is further used to: The face shadow dataset is obtained, wherein: The face shadow dataset includes multiple groups of face shadow image-clean image pairs; Based on the face shadow dataset, training the face shadow removal model with high- and low-frequency attention fusion mechanism to obtain an initial face shadow removal model; Based on the preset joint loss function, when it is determined that the initial face shadow removal model meets the preset training completion condition, the initial face shadow removal model is used as the pre-trained face shadow removal model.
7. The device according to claim 5, characterized in that The processing module is used for: Based on the encoder, data preprocessing is performed on the face image to be processed, the face image to be processed is converted from a color space to a feature space, and feature extraction is performed on the face image to be processed after being converted to the feature space to obtain a feature extraction result, and feature reconstruction is performed on the feature extraction result to obtain a target feature representation; Based on the decoder, data post-processing is performed on the target feature representation, and the target feature representation is converted from feature space to color space to obtain the de-shadowed face image.
8. The device according to claim 5, characterized in that The preset joint loss function is: L total =α1L pixel +α2L perceptual +α3L mutilscale ; Among them, L total is the preset joint loss function, α1 is the weight of pixel-level loss, α2 is the weight of perceptual loss, α3 is the weight of multi-scale perceptual loss, and L pixel is the pixel-level loss, L perceptual is the perceptual loss, L multiscale It is a multi-scale perceptual loss.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the face shadow removal method based on the high- and low-frequency attention fusion mechanism as described in any one of claims 1 to 4.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the face shadow removal method based on the high- and low-frequency attention fusion mechanism as described in any one of claims 1 to 4.