Dual-frequency domain rule learning model training method, image processing method and device
By using affine transformation and skin masking technology in the dual-frequency domain rule learning model, face repairs high-resolution face images, solving the problem of low image processing performance caused by limited computing power, and achieving efficient and natural face repair effects.
Patent Information
- Application Number
- CN202510192266.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-16
AI Technical Summary
When computing power is limited, when face repairs high-resolution face images, the calculation amount of the prior art is large, resulting in poor image processing performance.
The training method of the dual-frequency domain rule learning model is adopted. By performing affine transformation processing on the face image sample, the resolution is reduced, the image and skin mask input model for high-frequency and low-frequency repair rules to obtain key components for face repair, and model parameters are adjusted to improve performance.
Effectively reduce the calculation overhead of high-resolution image processing, reduce interference to non-defective areas, protect the lighting and texture of the original image, improve the naturalness of the repair effect, realize natural and realistic face repair, and significantly improve the real-time and deployability of the model.
Smart Images

Figure CN120013821A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to technical fields such as image processing, deep learning, computer and data processing in the field of artificial intelligence, and in particular to a training method, an image processing method and a device for a dual-frequency domain rule learning model. Background Art
[0002] With the rapid development of social media and live video applications, the demand for collecting and processing high-resolution facial images is increasing, especially in the repair of facial blemishes (such as acne marks, spots, and fine wrinkles).
[0003] Currently, for high-resolution (such as 4K or 5K level) facial images, deep convolutional networks are usually used to perform full-resolution facial defect repair on facial images. However, with limited computing power, the current solution has a large amount of calculation, resulting in poor image processing performance. Summary of the invention
[0004] The present disclosure provides a training method for a dual-frequency domain rule learning model, an image processing method and a device for improving the performance of image processing when performing face restoration on high-resolution face images.
[0005] According to a first aspect of the present disclosure, a training method for a dual-frequency domain rule learning model is provided, comprising:
[0006] Obtaining a training sample pair, the training sample pair comprising a face image sample and a face restoration image label corresponding to the face image sample;
[0007] Performing affine transformation processing on the face image sample to obtain an affine transformed image, wherein the resolution of the affine transformed image is smaller than the resolution of the face image sample;
[0008] Inputting the affine transformed image and the first skin mask image corresponding to the affine transformed image into a dual-frequency domain rule learning model to learn high-frequency and low-frequency restoration rules, and obtaining key components corresponding to the affine transformed image, which are used for face restoration;
[0009] Applying the key component to perform face restoration processing on the face image sample to obtain a face restoration image corresponding to the face image sample;
[0010] Based on the face inpainted image and the face inpainted image label, the model parameters of the dual-frequency domain rule learning model are adjusted.
[0011] According to a second aspect of the present disclosure, there is provided an image processing method, comprising:
[0012] Obtain the face image to be processed;
[0013] Performing affine transformation post-processing on the face image to be processed to obtain a target affine transformed image, wherein the resolution of the target affine transformed image is smaller than the resolution of the face image to be processed;
[0014] Inputting the target affine transformed image and the third skin mask image corresponding to the target affine transformed image into a dual-frequency domain rule learning model to learn high-frequency and low-frequency restoration rules, and obtaining a target key component corresponding to the target affine transformed image, wherein the target key component is used for face restoration, and the dual-frequency domain rule learning model is obtained by using the training method according to the first aspect of the present disclosure;
[0015] The target key component is applied to perform face restoration processing on the face image to be processed to obtain a target face restoration image corresponding to the face image to be processed.
[0016] According to a third aspect of the present disclosure, a training device for a dual-frequency domain rule learning model is provided, comprising:
[0017] An acquisition unit, used to acquire a training sample pair, the training sample pair comprising a face image sample and a face restoration image label corresponding to the face image sample;
[0018] A processing unit, used for performing affine transformation processing on the face image sample to obtain an affine transformed image, wherein the resolution of the affine transformed image is smaller than the resolution of the face image sample;
[0019] A learning unit, used for inputting the affine transformed image and the first skin mask image corresponding to the affine transformed image into a dual-frequency domain rule learning model to perform high-frequency and low-frequency restoration rule learning, and obtaining key components corresponding to the affine transformed image, wherein the key components are used for face restoration;
[0020] A restoration unit, used for applying the key component to perform face restoration processing on the face image sample to obtain a face restoration image corresponding to the face image sample;
[0021] The adjustment unit is used to adjust the model parameters of the dual-frequency domain rule learning model based on the face restoration image and the face restoration image label.
[0022] According to a fourth aspect of the present disclosure, there is provided an image processing apparatus, comprising:
[0023] A first acquisition unit, used to acquire a face image to be processed;
[0024] An affine transformation processing unit, used for performing affine transformation post-processing on the face image to be processed to obtain a target affine transformed image, wherein the resolution of the target affine transformed image is smaller than the resolution of the face image to be processed;
[0025] A second acquisition unit is used to input the target affine transformed image and the third skin mask image corresponding to the target affine transformed image into a dual-frequency domain rule learning model to perform high-frequency and low-frequency restoration rule learning, so as to obtain a target key component corresponding to the target affine transformed image, wherein the target key component is used for face restoration, and the dual-frequency domain rule learning model is obtained by adopting the training method according to the first aspect of the present disclosure;
[0026] The restoration processing unit is used to apply the target key component to perform face restoration processing on the face image to be processed, and obtain a target face restoration image corresponding to the face image to be processed.
[0027] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the training method of the dual-frequency domain rule learning model described in the first aspect or execute the image processing method described in the second aspect.
[0028] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the training method of the dual-frequency domain rule learning model described in the first aspect or execute the image processing method described in the second aspect.
[0029] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising: a computer program, the computer program being stored in a readable storage medium, at least one processor of an electronic device being able to read the computer program from the readable storage medium, the at least one processor executing the computer program so that the electronic device executes the training method of the dual-frequency domain rule learning model described in the first aspect or executes the image processing method described in the second aspect.
[0030] The technology disclosed in the present invention solves the problem of poor image processing performance when performing face restoration on high-resolution face images in the current manner under limited computing power. The present invention performs affine transformation processing on face image samples to obtain an affine transformed image with a resolution smaller than that of the face image sample; the affine transformed image and the first skin mask image corresponding to the affine transformed image are input into a dual-frequency domain rule learning model to perform high-frequency and low-frequency restoration rule learning to obtain key components for face restoration; the key components are applied to perform face restoration processing on face image samples to obtain face restoration images corresponding to the face image samples; and the model parameters of the dual-frequency domain rule learning model are adjusted based on the face restoration image and the face restoration image label. Since the dual-frequency domain rule learning model performs high-frequency and low-frequency restoration rule learning based on the affine transformed image with a smaller resolution, and the interference of non-skin areas can be effectively shielded by the first skin mask image, the scale and amount of computation of the dual-frequency domain rule learning model can be greatly reduced, and global illumination can be effectively controlled while ensuring local details. The dual-frequency domain rule learning model obtained through training can effectively reduce the computational overhead of high-resolution image processing, reduce interference to non-defective areas, protect the lighting and texture of the original image, and improve the naturalness of the restoration effect, thereby achieving natural and realistic face restoration with a smaller computational overhead, improving image processing performance, and significantly improving the real-time and deployability of the dual-frequency domain rule learning model.
[0031] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0033] Figure 1 is a schematic diagram of an application scenario to which the present disclosure is applicable;
[0034] Figure 2 is a schematic diagram according to a first embodiment of the present disclosure;
[0035] Figure 3 is a schematic diagram according to a second embodiment of the present disclosure;
[0036] Figure 4 is a schematic diagram according to a third embodiment of the present disclosure;
[0037] Figure 5 is a schematic diagram according to a fourth embodiment of the present disclosure;
[0038] Figure 6 is a schematic diagram according to a fifth embodiment of the present disclosure;
[0039] Figure 7 is a schematic diagram according to a sixth embodiment of the present disclosure;
[0040] Figure 8 is a schematic diagram according to a seventh embodiment of the present disclosure;
[0041] Fig. 9 FIG. 9 is a schematic block diagram of an example electronic device 900 that may be used to implement embodiments of the present disclosure. DETAILED DESCRIPTION
[0042] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0043] The present disclosure provides a training method, an image processing method and a device for a dual-frequency domain rule learning model, which are applied to technical fields such as image processing, deep learning, computer and data processing in the field of artificial intelligence, so as to reduce the computational complexity of high-resolution face image processing and improve the performance of image processing.
[0044] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0045] With the rapid development of social media and live video applications, the demand for collecting and processing high-resolution facial images is increasing, especially in the repair of facial blemishes (such as acne marks, spots, and fine wrinkles).
[0046] At present, for high-resolution facial images (such as 4K, 5K and higher resolutions), deep convolutional networks are usually used to perform full-resolution facial blemish repair on facial images. Among them, performing convolution, attention, and generative adversarial network (GAN) operations on facial images at high resolution requires huge computing power and video memory. However, with limited computing power, the current solution has a large amount of computation, resulting in poor image processing performance, making it difficult to meet real-time or low-computation application scenarios, and difficult to deploy efficiently on mobile terminals or in resource-limited environments.
[0047] In addition, the current scheme has the following problems: (1) The high-resolution feature extraction and restoration process is prone to loss or destruction of the overall facial light and shadow structure; (2) The relationship between defect areas and non-defect areas needs to be learned simultaneously, which increases the difficulty and network complexity sharply. When it is impossible to finely distinguish between defect areas that need to be repaired and non-defect areas that do not need to be repaired (such as facial skin or background areas), it often leads to over-restoration or processing distortion; (3) The lighting and skin quality of different faces vary greatly, and the current scheme is insufficient in maintaining details and realism.
[0048] In addition, in one related technology, for high-resolution facial images, a large network or local filter tends to be used in full-resolution scenarios to repair facial defects, which has a speed bottleneck problem for extremely high resolutions. In another related technology, users are generally required to manually specify the defect repair area, with a low degree of automation, and the processing is time-consuming for facial images with a resolution of 4K or above. In another related technology, for example, when a mobile phone has a built-in beauty lens and performs facial defect repair in high-resolution video mode, it often simplifies the processing or uses a skin smoothing filter, which makes it impossible to finely preserve the texture.
[0049] In order to solve the above problems, the present invention aims at the face restoration needs of high-resolution face images. Based on the overall idea of small-resolution learning rules and original-resolution application rules, the face image samples in the training sample pairs are affine transformed to obtain an affine transformed image with a resolution less than the resolution of the face image samples; the affine transformed image and the first skin mask image corresponding to the affine transformed image are input into the dual-frequency domain rule learning model to learn high-frequency and low-frequency restoration rules, that is, high-frequency and low-frequency restoration rule learning is performed based on the affine transformed image with a smaller resolution, and the first skin mask image can effectively shield the interference of non-skin areas, so that the computing power requirements of the dual-frequency domain rule learning model can be effectively reduced, and the global illumination can be effectively controlled while ensuring local details; and then the key components are applied to perform face restoration on the face image samples, and the model parameters of the dual-frequency domain rule learning model are adjusted based on the obtained face restoration image and face restoration image label. The dual-frequency domain rule learning model obtained by training can effectively reduce the computational overhead of high-resolution image processing, thereby realizing natural and realistic face restoration with a smaller computational overhead, and can significantly improve the real-time and deployability of the dual-frequency domain rule learning model.
[0050] Figure 11 is a schematic diagram of an application scenario applicable to the present disclosure. The application scenario may include: a server cluster 11 and a terminal 12. The server cluster 11 includes a plurality of servers 111 and a memory 112, and the terminal 12 may be a tablet computer, a laptop computer, a desktop computer, a smart home appliance, etc. The server 111 is used to train the dual-frequency domain rule learning model, obtain data from the memory 112 during the training process, and store the generated data in the memory 112. In addition, during the training process, communication is performed with the terminal 12 via a wireless network or a wired network.
[0051] In addition, the disclosed embodiments can be applied to image processing scenarios. For example, a user uploads a face image to be processed to the server 111 through the terminal 12; the server 111 obtains a target face restoration image after face restoration processing based on the face image to be processed, and sends the target face restoration image to the terminal 12; the terminal 12 displays the target face restoration image to the user.
[0052] It should be noted that Figure 1 This is only a schematic diagram of an application scenario provided by the embodiment of the present disclosure. Figure 1 does not limit the equipment included in Figure 1 The positional relationship between the devices is limited.
[0053] It can be seen that the present disclosure can be widely applied to the following scenarios:
[0054] (1) Real-time beautification functions for various face images and video social platforms;
[0055] (2) Professional photo editing software performs high-resolution facial blemish removal in an offline environment;
[0056] (3) High-end camera devices or mobile phones can quickly achieve high-resolution face blemish removal under resource-constrained conditions;
[0057] (4) Real-time facial effects processing for high-definition video streams on live streaming platforms when resources are limited;
[0058] (5) Professional medical image analysis and post-operative skin comparison scenarios, which can remove fine defects or reconstruct skin disease scars from ultra-high-resolution facial medical images, and support specialist diagnosis and post-operative comparison;
[0059] (6) Virtual anchors and augmented reality (AR) special effects make the face effect in 4K live broadcasts more natural by repairing key skin areas or superimposing special effects.
[0060] The following specific embodiments are used to describe in detail the technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below in conjunction with the accompanying drawings.
[0061] Figure 2 Schematic diagram of the first embodiment of the present disclosure. Figure 2 As shown, the training method of the dual-frequency domain rule learning model provided in the first embodiment of the present disclosure can be applied to an electronic device, which can be a server or a server cluster, etc. The training method of the dual-frequency domain rule learning model provided in the first embodiment of the present disclosure includes:
[0062] S201. Obtain a training sample pair, where the training sample pair includes a face image sample and a face restoration image label corresponding to the face image sample.
[0063] In the disclosed embodiment, the face image sample is a high-resolution face image to be repaired (i.e., blemishes removed); the face repair image label is the face image obtained after face repairing the face image sample, for example, the face repair image label can be obtained by manual image retouching. According to the face image sample and the face repair image label corresponding to the face image sample, a plurality of training sample pairs can be obtained.
[0064] S202: Perform affine transformation on the face image sample to obtain an affine transformed image, wherein the resolution of the affine transformed image is smaller than the resolution of the face image sample.
[0065] Exemplarily, the facial key point information in the facial image sample can be obtained through the facial key point detection algorithm. The facial key points may include, for example, the corners of the eyes, the tip of the nose, and the corners of the mouth, so that affine transformation processing is performed based on the facial key point information to obtain an affine transformed image. Among them, the affine transformation processing can be a resolution reduction processing of the facial image sample, or it can be a face alignment processing of the facial image sample. The resolution of the affine transformed image is smaller than the resolution of the facial image sample, which facilitates the dual-frequency domain rule learning model to perform high-frequency and low-frequency repair rule learning based on the affine transformed image with a smaller resolution, thereby reducing the computing power requirement. For specific information on how to perform affine transformation processing on the facial image sample to obtain the affine transformed image, please refer to the subsequent embodiments.
[0066] S203, inputting the affine transformed image and the first skin mask image corresponding to the affine transformed image into a dual-frequency domain rule learning model to perform high-frequency and low-frequency restoration rule learning, and obtaining key components corresponding to the affine transformed image, wherein the key components are used for face restoration.
[0067] In this step, based on the current relevant technology, the first skin mask image corresponding to the image after affine transformation can be obtained. The first skin mask image is a single-channel image, for example, 1 represents the skin area and 0 represents the non-skin area. The first skin mask image can effectively shield the interference of the non-skin area (such as background, etc.) in the image after affine transformation. Exemplarily, the image after affine transformation and the first skin mask image corresponding to the image after affine transformation are input into the dual-frequency domain rule learning model for high-frequency and low-frequency repair rule learning to obtain the key components for face repair. The key components include, for example, a multiplicative coefficient map, an additive coefficient map and a fusion weight map, wherein the multiplicative coefficient map is used to correct the global skin color and illumination based on low-frequency information; the additive coefficient map is used to repair the details of local texture defects based on high-frequency information; and the fusion weight map is used to perform an adaptive balance between the subsequent face repair and the original image (face image sample) based on low-frequency information and high-frequency information. It can be understood that by separating low-frequency information from high-frequency information in a smaller resolution space and learning high-frequency and low-frequency restoration rules, the dual-frequency domain rule learning model can effectively control global illumination while ensuring local details and reduce computing power requirements through a lightweight model structure.
[0068] For details on how to obtain the key components corresponding to the image after affine transformation, please refer to the subsequent embodiments.
[0069] S204: Apply the key component to perform face restoration processing on the face image sample to obtain a face restoration image corresponding to the face image sample.
[0070] In this step, after obtaining the key components corresponding to the image after affine transformation, the key components can be applied to the face image sample to perform face restoration processing, thereby obtaining the face restoration image corresponding to the face image sample. Exemplarily, the key components can be subjected to inverse affine transformation processing to obtain the key components after inverse affine transformation, and the resolution of the key components after inverse affine transformation is the same as the resolution of the face image sample, and then the face restoration processing is performed on the face image sample based on the key components after inverse affine transformation to obtain the face restoration image corresponding to the face image sample.
[0071] For details on how to apply the key components to perform face restoration processing on the face image samples to obtain the face restoration images corresponding to the face image samples, please refer to the subsequent embodiments.
[0072] S205: Adjust model parameters of the dual-frequency domain rule learning model based on the face restoration image and the face restoration image label.
[0073] In this step, the face repaired image and the face repaired image label may be compared pixel by pixel to determine the loss. The specific loss may include, for example, mean square error loss and perceptual loss. Based on the loss, the model parameters of the dual-frequency domain rule learning model may be adjusted to obtain a trained dual-frequency domain rule learning model.
[0074] For details on how to adjust the model parameters of the dual-frequency domain rule learning model based on the face restoration image and the face restoration image label, please refer to the subsequent embodiments.
[0075] In the disclosed embodiment, an affine transformation is performed on a face image sample to obtain an affine transformed image having a resolution smaller than that of the face image sample; the affine transformed image and the first skin mask image corresponding to the affine transformed image are input into a dual-frequency domain rule learning model to perform high-frequency and low-frequency repair rule learning to obtain key components for face repair; the key components are applied to perform face repair on the face image sample to obtain a face repair image corresponding to the face image sample; and the model parameters of the dual-frequency domain rule learning model are adjusted based on the face repair image and the face repair image label. Since the dual-frequency domain rule learning model performs high-frequency and low-frequency repair rule learning based on an affine transformed image with a relatively small resolution, and the first skin mask image can effectively shield interference from non-skin areas, the scale and amount of computation of the dual-frequency domain rule learning model can be greatly reduced, and global illumination can be effectively controlled while ensuring local details. The dual-frequency domain rule learning model obtained through training can effectively reduce the computational overhead of high-resolution image processing, reduce interference to non-defective areas, protect the lighting and texture of the original image, and improve the naturalness of the restoration effect, thereby achieving natural and realistic face restoration with a smaller computational overhead, improving image processing performance, and significantly improving the real-time and deployability of the dual-frequency domain rule learning model.
[0076] Figure 3 FIG. 1 is a schematic diagram of a second embodiment of the present disclosure. Based on the above embodiment, the present disclosure further describes the training method of the dual-frequency domain rule learning model. Figure 3 As shown, the training method of the dual-frequency domain rule learning model provided by the second embodiment of the present disclosure may include:
[0077] S301. Obtain a training sample pair, where the training sample pair includes a face image sample and a face restoration image label corresponding to the face image sample.
[0078] The implementation principle and technical effects of S301 may refer to the aforementioned embodiments and will not be described in detail.
[0079] In the disclosed embodiment, Figure 2 The step S202 may further include the following two steps S302 and S303:
[0080] S302: Perform facial key point detection on the facial image sample to determine facial key point information in the facial image sample.
[0081] For example, a facial key point detection algorithm can be used to detect facial key points on a facial image sample, thereby determining facial key point information in the facial image sample. The specific facial key point detection algorithm can refer to the current related technology, and the embodiments of the present disclosure are not limited thereto.
[0082] S303, based on the facial key point information, obtaining an affine transformation matrix; according to the affine transformation matrix, performing face alignment and resolution reduction processing on the facial image samples to obtain an affine transformed image, wherein the resolution of the affine transformed image is smaller than the resolution of the facial image samples.
[0083] In this step, after obtaining the facial key point information in the facial image sample, the affine transformation matrix (for example, represented by M) can be obtained based on the facial key point information. The facial image sample is a high-resolution facial image, and the facial image sample can be face aligned and resolution reduced according to the affine transformation matrix to obtain an affine transformed image of uniform size. The resolution of the affine transformed image is smaller than the resolution of the facial image sample, and the resolution of the affine transformed image is, for example, 512×512 or smaller. It can be understood that through the face alignment process, the feature offset caused by different postures and angles of the face can be eliminated, so that the subsequent high-frequency and low-frequency restoration rule learning at a small resolution (small size) is more stable.
[0084] Considering that the key components include the multiplicative coefficient graph, the additive coefficient graph and the fusion weight graph, therefore, in the embodiment of the present disclosure, Figure 2 The step S203 may further include the following four steps S304 to S307:
[0085] S304. Input the affine transformed image and the first skin mask image corresponding to the affine transformed image into the feature extraction module of the dual-frequency domain rule learning model. The feature extraction module extracts features from the affine transformed image under the masking effect of the first skin mask image to obtain global features and local features of the face.
[0086] For example, Figure 4 is a schematic diagram according to the third embodiment of the present disclosure. Figure 4As shown, the affine transformed image and the first skin mask image corresponding to the affine transformed image can be input into the feature extraction module of the dual-frequency domain rule learning model. The feature extraction module extracts features from the affine transformed image under the masking effect of the first skin mask image to obtain global features and local features of the face. The affine transformed image is a red, green and blue (RGB) 3-channel image, the first skin mask image is a single-channel image, and the affine transformed image and the first skin mask image together are a 4-channel image; the feature extraction module includes multiple convolution blocks, Figure 4 In the example, the feature extraction module includes convolution blocks 1 to 4. Accordingly, two local face features can be obtained, namely, local face feature E1 obtained by convolution block 1 and local face feature E2 obtained by convolution block 2, and then two global face features can be obtained, namely, global face feature E3 obtained by convolution block 3 and global face feature E4 obtained by convolution block 4. It can be understood that the dual-frequency domain rule learning model in this step uses fewer convolution blocks (convolution kernels) and lower computational complexity to ensure the lightweight of the dual-frequency domain rule learning model.
[0087] S305 , inputting the global features of the face into the deconvolution block of the dual-frequency domain rule learning model for deconvolution processing to obtain a multiplicative coefficient map.
[0088] This step can be understood as a low-frequency branch processing step, and the global face feature is used to obtain relatively smooth low-frequency information such as overall facial illumination and skin color. Figure 4 , the global features of the face (E3 and E4) are input into the deconvolution block of the dual-frequency domain rule learning model for deconvolution processing, and a multiplicative coefficient map (for example, represented by a_low) can be obtained. The resolution of the multiplicative coefficient map is the same as that of the image after affine transformation. The multiplicative coefficient map mainly performs a large-scale correction of the brightness and hue of the defect area during the face restoration process, and is used to perform multiplicative correction on the overall skin color and illumination distribution of the defect.
[0089] S306: Input the local features of the face into the first convolution block of the dual-frequency domain rule learning model for convolution processing to obtain an additive coefficient map.
[0090] This step can be understood as a high-frequency branch processing step, where local facial features are used to obtain local high-frequency information such as texture defects. Figure 4 , the local features of the face (E1 and E2) are input into the first convolution block of the dual-frequency domain rule learning model for convolution processing, and an additive coefficient map (for example, represented by b_low) can be obtained. Among them, the first convolution block uses a smaller convolution kernel (1×1) and avoids downsampling to retain more fine information. The resolution of the additive coefficient map is the same as that of the image after affine transformation. The additive coefficient map is specifically used for detail repair such as local spots, acne marks or uneven texture.
[0091] S307: Input the global face features and the local face features into the second convolution block of the dual-frequency domain rule learning model for convolution processing to obtain a fusion weight map.
[0092] This step can be understood as a fusion weight channel processing step, which generates a single-channel adaptive fusion weight map by fusing the global features of the face and the local features of the face. The fusion weight map is represented by α_low, and the value range of α_low is between [0,1]. For example, refer to Figure 4 , the global face features and local face features (E1 to E4) are input into the second convolution block of the dual-frequency domain rule learning model for convolution processing, and a fusion weight map can be obtained. The resolution of the fusion weight map is the same as the resolution of the image after affine transformation. The fusion weight map is used to adaptively balance between the subsequent restoration results and the original image (face image sample). In the subsequent processing stage, the closer α_low is to 1, the more the corresponding position is inclined to adopt the restoration result; the closer α_low is to 0, the more the original image (face image sample) information is retained.
[0093] Based on steps S304 to S307, by separating low-frequency information and high-frequency information in a smaller resolution space, the dual-frequency domain rule learning model can effectively control global illumination while ensuring local details, and reduce computing power requirements through the lightweight convolution structure of the dual-frequency domain rule learning model.
[0094] It should be noted that the embodiment of the present disclosure does not limit the execution order of S305, S306 and S307.
[0095] In the disclosed embodiment, Figure 2 The step S204 may further include the following three steps S308 to S310:
[0096] S308. Perform inverse affine transformation on the multiplicative coefficient map, the additive coefficient map and the fusion weight map respectively to obtain the multiplicative coefficient map after inverse affine transformation, the additive coefficient map after inverse affine transformation and the fusion weight map after inverse affine transformation. The resolutions of the multiplicative coefficient map after inverse affine transformation, the additive coefficient map after inverse affine transformation and the fusion weight map after inverse affine transformation are the same as the resolution of the face image sample.
[0097] It can be understood that the inverse affine transformation is the inverse operation of the affine transformation. In the disclosed embodiment, the affine transformation is to perform face alignment and resolution reduction processing on the face image sample. Accordingly, in order to match the resolution of the face image sample, the resolution of the multiplicative coefficient map, the additive coefficient map and the fusion weight map can be enlarged to the resolution of the face image sample by the inverse affine transformation, and the face inverse alignment processing can be performed, so as to obtain the multiplicative coefficient map after the inverse affine transformation, the additive coefficient map after the inverse affine transformation and the fusion weight map after the inverse affine transformation.
[0098] Exemplarily, the multiplicative coefficient map, the additive coefficient map, and the fusion weight map can be inversely transformed by the following formulas 1 to 3, respectively, to obtain the multiplicative coefficient map after inverse affine transformation (for example, represented by a), the additive coefficient map after inverse affine transformation (for example, represented by b), and the fusion weight map after inverse affine transformation (for example, represented by α):
[0099] a=inverse affine transformation (a_low) Formula 1
[0100] b = inverse affine transformation (b_low) Formula 2
[0101] α = inverse affine transformation (α_low) Formula 3
[0102] Specifically, the inverse affine transformation can be performed through the inverse affine transformation matrix (for example, represented by M_inv). The specific method of obtaining the inverse affine transformation matrix is, for example, cv2.invertAffineTransform (a method for obtaining the inverse matrix of a given affine transformation matrix) and other methods. Among them, the resolutions of the multiplicative coefficient map, the additive coefficient map, and the fusion weight map can be enlarged to the resolution of the face image sample by methods such as bilinear interpolation. The interpolation operation does not involve deep convolution and has a small amount of calculation. The (a, b, α) at the resolution of the face image sample can be combined with the face image sample to achieve the same mapping rule application for high-resolution pixels.
[0103] S309, performing face restoration processing based on the face image sample, the multiplicative coefficient map after the inverse affine transformation, the additive coefficient map after the inverse affine transformation, and the fusion weight map after the inverse affine transformation to obtain a restoration result.
[0104] In this step, the multiplicative coefficient map after the inverse affine transformation, the additive coefficient map after the inverse affine transformation and the fusion weight map after the inverse affine transformation can be applied to the face image sample to perform face restoration processing, thereby obtaining a restoration result.
[0105] Further, optionally, performing face restoration processing based on face image samples, the multiplicative coefficient map after inverse affine transformation, the additive coefficient map after inverse affine transformation and the fusion weight map after inverse affine transformation to obtain a restoration result may include: pixel-by-pixel multiplication of pixels at the same position in the face image sample and the multiplicative coefficient map after inverse affine transformation to obtain a first restored image; pixel-by-pixel addition of pixels at the same position in the first restored image and the additive coefficient map after inverse affine transformation to obtain a second restored image; and fusing the second restored image and the face image sample according to the fusion weight map after inverse affine transformation to obtain a restoration result.
[0106] Exemplarily, the restoration result (for example, represented by dst) can be obtained at the resolution of the face image sample by the following formula 4:
[0107] dst=α⊙(a⊙I_origin+b)+(1-α)⊙I_origin Formula 4
[0108] Wherein: I_origin represents the original input image. In the embodiment of the present disclosure, I_origin is a face image sample; ⊙ represents element-by-pixel multiplication.
[0109] It can be understood that according to the above formula 4, the face image samples are adjusted by the multiplicative coefficient map (a) after the inverse affine transformation and the additive coefficient map (b) after the inverse affine transformation, and the smooth transition between the adjusted image and the face image samples is controlled by the fusion weight map (α) after the inverse affine transformation, and finally the target image (i.e., the repair result dst) is generated to achieve local and global repair of face defects.
[0110] S310: Perform fusion processing based on the restoration result, the face image sample and the second skin mask image corresponding to the face image sample to obtain a face restoration image corresponding to the face image sample.
[0111] In this step, based on the current related technology, a second skin mask image corresponding to the face image sample can be obtained in advance. The second skin mask image is a single-channel image, for example, 1 represents the skin area and 0 represents the non-skin area. Based on the skin area determined by the second skin mask image, the restoration result and the face image sample are fused to obtain the face restoration image corresponding to the face image sample, so that the face restoration process only acts on the skin area, thereby ensuring that the non-skin area maintains the original image information and avoids unnecessary interference.
[0112] Further, optionally, performing fusion processing based on the repair result, the face image sample and the second skin mask image corresponding to the face image sample to obtain a face repaired image corresponding to the face image sample can include: performing skin area fusion processing on the repair result and the face image sample according to the second skin mask image to obtain a face repaired image corresponding to the face image sample.
[0113] Exemplarily, the following formula 5 can be used to perform skin area fusion processing on the repair result and the face image sample according to the second skin mask image (for example, represented by skin_mask, where skin_mask represents a binary mask, where 1 represents a skin area and 0 represents a non-skin area) to obtain a face repair image corresponding to the face image sample (for example, represented by result):
[0114] result=skin_mask⊙dst+(1-skin_mask)⊙I_origin Formula 5
[0115] It can be understood that according to the above formula five, the use of the second skin mask image can make the face restoration processing only act on the skin area, thereby ensuring that the non-skin area (such as background or accessories, etc.) retains the original image information, avoiding unnecessary interference, and ultimately obtaining a natural and delicate face restoration image at high resolution.
[0116] In the disclosed embodiment, Figure 2 The step S205 may further include the following two steps S311 and S312:
[0117] S311. Based on the face restoration image and the face restoration image label, obtain a mean square error loss and a perceptual loss.
[0118] In this step, in order to effectively balance the overall restoration degree and the perceived visual quality, the face repair image and the face repair image label (such as represented by target) can be compared pixel by pixel, and multiple loss functions can be used for joint training. Exemplarily, the mean square error loss and the perceptual loss can be obtained based on the face repair image and the face repair image label, wherein the mean square error loss is used to constrain the global pixel error to ensure the accurate restoration of illumination and color; the perceptual loss is used to compare deep perceptual features to ensure the naturalness of texture and structure.
[0119] The overall loss function of this embodiment satisfies the following formula 6:
[0120] L total =λ mse L mse +λ vgg L vgg Formula 6
[0121] Among them, Lmse represents the mean square error loss, It is used to constrain the error between the global pixel value of the face repair image (result) and the face repair image label (target) to ensure the accurate restoration of lighting and color; L vgg represents the perceived loss, φ i represents the i-th layer feature extractor of the pre-trained VGG network (a deep convolutional neural network), which is used to measure the texture and structural details of the image; λ mse Represents the weight coefficient of mean square error loss; λ vgg Represents the weight coefficient of perceptual loss.
[0122] It can be understood that through the constraints of mean square error loss and perceptual loss, mean square error loss can constrain the error in global pixel values between the face repaired image and the face repaired image label, ensuring the accurate restoration of lighting and color; perceptual loss can measure the texture and structural details of the image based on the extraction of perceptual features by the deep network, and help the network generate real and natural repair details at a higher level, so as to obtain repair effects that meet visual needs and quantitative evaluation in high-resolution scenes.
[0123] S312. According to the mean square error loss and the perceptual loss, the model parameters of the dual-frequency domain rule learning model are adjusted through back propagation and optimization iteration.
[0124] In this step, the model parameters of the dual-frequency domain rule learning model can be continuously adjusted according to the mean square error loss and the perceptual loss through back propagation and optimization iteration, so that the generation of (a, b, α) can achieve excellent restoration effect in high-resolution scenarios, and a trained dual-frequency domain rule learning model is obtained.
[0125] In the disclosed embodiment, facial key point detection is performed on facial image samples to determine facial key point information in the facial image samples, and an affine transformation matrix is obtained based on the facial key point information; according to the affine transformation matrix, face alignment is performed on the facial image samples and resolution reduction processing is performed to obtain an affine transformed image, and the face alignment processing can eliminate feature offsets caused by different postures and angles of the face, and can make the restoration rule learning performed at a small resolution more stable; the affine transformed image and the first skin mask image corresponding to the affine transformed image are input into the dual-frequency domain rule learning model. The feature extraction module performs feature extraction to obtain global features and local features of the face; the global features of the face are input into the deconvolution block of the dual-frequency domain rule learning model for deconvolution processing to obtain a multiplicative coefficient map, and the local features of the face are input into the first convolution block of the dual-frequency domain rule learning model for convolution processing to obtain an additive coefficient map; the global features of the face and the local features of the face are input into the second convolution block of the dual-frequency domain rule learning model for convolution processing to obtain a fusion weight map; by obtaining the multiplicative coefficient map and the additive coefficient map, the dual-frequency domain rule learning model can be more easily learned for facial illumination / skin color (global) According to the repair rules of local texture (defects), the first skin mask image can effectively shield the interference areas such as the background; the multiplicative coefficient map, the additive coefficient map and the fusion weight map are respectively processed by inverse affine transformation to obtain the multiplicative coefficient map after inverse affine transformation, the additive coefficient map after inverse affine transformation and the fusion weight map after inverse affine transformation; the face image samples are adjusted by the multiplicative coefficient map after inverse affine transformation and the additive coefficient map after inverse affine transformation, and the smooth transition between the adjusted image and the face image samples is controlled by the fusion weight map after inverse affine transformation to obtain the repair result, thereby realizing the local and global repair of face defects; based on The repair result, the face image sample and the second skin mask image corresponding to the face image sample are fused to obtain the face repair image corresponding to the face image sample; wherein, the second skin mask image can make the face repair process only act on the skin area, thereby ensuring that the non-skin area maintains the original image information and avoids unnecessary interference; based on the face repair image and the face repair image label, the mean square error loss and the perceptual loss are obtained, and according to the mean square error loss and the perceptual loss, the model parameters of the dual-frequency domain rule learning model are adjusted through back propagation and optimization iteration, which can take into account the overall restoration degree and the perceived visual quality. The dual-frequency domain rule learning model obtained by training can effectively reduce the computational overhead of high-resolution image processing, minimize the model learning burden and reduce the interference to non-defective areas, protect the lighting and texture of the original image, and improve the naturalness of the repair effect, thereby achieving natural and realistic face repair with a small computational overhead, and can significantly improve the real-time performance and deployability of the dual-frequency domain rule learning model.
[0126] Based on the above embodiments, Figure 5is a schematic diagram according to the fourth embodiment of the present disclosure. Figure 5 As shown, the image processing method provided in the fourth embodiment of the present disclosure can be applied to an electronic device, which can be a server or a server cluster. The image processing method provided in the fourth embodiment of the present disclosure includes:
[0127] S501: Obtain a face image to be processed.
[0128] Exemplarily, the facial image to be processed may be input by a user to the electronic device executing the embodiment of the method, or may be sent by other devices to the electronic device executing the embodiment of the method. For example, the facial image to be processed is a facial image uploaded by a user, or a video frame collected in real time, and the facial image to be processed may include a facial region or a multi-face scene.
[0129] S502: Perform affine transformation post-processing on the face image to be processed to obtain a target affine transformed image, wherein the resolution of the target affine transformed image is smaller than the resolution of the face image to be processed.
[0130] Exemplarily, the facial key point information in the face image to be processed can be obtained through a pre-deployed facial key point detection algorithm. The facial key points may include, for example, the corners of the eyes, the tip of the nose, and the corners of the mouth, so that affine transformation processing is performed based on the facial key point information to obtain the target affine transformed image. Among them, the affine transformation processing can be a resolution reduction processing of the face image samples, or it can be a face alignment processing of the face image samples. The resolution of the target affine transformed image is smaller than the resolution of the face image to be processed. The resolution of the target affine transformed image is, for example, 512×512 or smaller, so that the dual-frequency domain rule learning model can reduce the computing power requirements.
[0131] For details on how to perform affine transformation post-processing on the face image to be processed to obtain the target affine transformed image, please refer to the subsequent embodiments.
[0132] S503. Input the target affine transformed image and the third skin mask image corresponding to the target affine transformed image into the dual-frequency domain rule learning model to perform high-frequency and low-frequency restoration rule learning to obtain the target key components corresponding to the target affine transformed image. The target key components are used for face restoration. The dual-frequency domain rule learning model is obtained by using the training method of the dual-frequency domain rule learning model in any of the above method embodiments.
[0133] In this step, based on the current relevant technology, a third skin mask image corresponding to the target affine transformed image can be obtained. The third skin mask image is a single-channel image, for example, 1 represents the skin area and 0 represents the non-skin area. The third skin mask image can effectively shield the interference of the non-skin area (such as the background, etc.) in the target affine transformed image. The dual-frequency domain rule learning model is obtained by using the training method of the dual-frequency domain rule learning model in any of the above method embodiments, and can accurately output the key components for face restoration. In this step, the target affine transformed image and the third skin mask image corresponding to the target affine transformed image are input into the dual-frequency domain rule learning model for high-frequency and low-frequency restoration rule learning, and the target key components for face restoration corresponding to the target affine transformed image are obtained. The target key components include, for example, a target multiplicative coefficient map, a target additive coefficient map and a target fusion weight map, wherein the target multiplicative coefficient map is used to correct global skin color and lighting based on low-frequency information; the target additive coefficient map is used to perform detail repair of local texture defects based on high-frequency information; the target fusion weight map is used to perform adaptive balance between subsequent face restoration and the original image (face image to be processed) based on low-frequency information and high-frequency information.
[0134] S504: Apply the target key component to perform face restoration processing on the face image to be processed, and obtain a target face restoration image corresponding to the face image to be processed.
[0135] Exemplarily, the target key component is applied to the face image to be processed to perform face restoration processing, thereby obtaining a target face restoration image corresponding to the face image to be processed. Exemplarily, the target key component can be subjected to inverse affine transformation processing to obtain a target key component after inverse affine transformation, and the resolution of the target key component after inverse affine transformation is the same as the resolution of the face image to be processed, and then the face image to be processed is subjected to face restoration processing based on the target key component after inverse affine transformation to obtain a restoration result, thereby achieving local and global restoration of facial defects. The restoration result is fused with the face image to be processed in the skin area to obtain a target face restoration image corresponding to the face image to be processed, and the target face restoration image is a natural and delicate face blemish removal result at high resolution.
[0136] For details on how to apply the target key component to perform face restoration processing on the face image to be processed and obtain a target face restoration image corresponding to the face image to be processed, reference may be made to subsequent embodiments.
[0137] Optionally, if the facial image to be processed contains multiple faces or multiple facial areas, the above steps S502 to S504 can be executed for each face or facial area through a loop iteration to obtain the target face restoration image corresponding to each face or facial area, and finally a fully restored face restoration image can be obtained on a high-resolution original image (i.e., the facial image to be processed).
[0138] In the disclosed embodiment, the face image to be processed is subjected to affine transformation post-processing to obtain a target affine transformed image, and the resolution of the target affine transformed image is less than the resolution of the face image to be processed; the target affine transformed image and the third skin mask image corresponding to the target affine transformed image are input into the dual-frequency domain rule learning model for high-frequency and low-frequency repair rule learning to obtain the target key component for face repair; the target key component is applied to perform face repair processing on the face image to be processed to obtain the target face repair image corresponding to the face image to be processed. Among them, the dual-frequency domain rule learning model is obtained by using the training method of the dual-frequency domain rule learning model in any of the above method embodiments, and can accurately output the target key component with a small computational overhead, shorten the inference time while maintaining the details of the face, and realize natural and realistic face repair. The disclosed embodiment can greatly reduce the computing power and storage requirements of high-resolution face repair, improve the performance of image processing, and the solution is simple and easy to implement, and can be quickly deployed on different hardware platforms (such as mobile terminals and embedded devices).
[0139] Figure 6 FIG. 5 is a schematic diagram according to the fifth embodiment of the present disclosure. Based on the above embodiment, the present disclosure further describes the image processing method. Figure 6 As shown, the image processing method provided by the fifth embodiment of the present disclosure may include:
[0140] S601: Obtain a face image to be processed.
[0141] The implementation principle and technical effects of S601 may refer to the aforementioned embodiments and will not be described in detail.
[0142] In the disclosed embodiment, Figure 5 The step S502 may further include the following two steps S602 and S603:
[0143] S602: Perform facial key point detection on the face image to be processed to determine facial key point information in the face image to be processed.
[0144] Exemplarily, facial key point detection may be performed on a facial image sample using a facial key point detection algorithm, thereby determining facial key point information in the facial image sample.
[0145] S603, based on the facial key point information, obtaining a target affine transformation matrix; according to the target affine transformation matrix, performing face alignment and resolution reduction processing on the face image to be processed to obtain a target affine transformed image, wherein the resolution of the target affine transformed image is smaller than the resolution of the face image to be processed.
[0146] In this step, the target affine transformation matrix can be obtained based on the facial key point information. The face image to be processed is a high-resolution face image. The face image to be processed can be aligned and the resolution can be reduced according to the affine transformation matrix to obtain a target affine transformed image of uniform size. The resolution of the target affine transformed image is smaller than the resolution of the face image to be processed, and the resolution of the target affine transformed image is, for example, 512×512 or smaller.
[0147] S604. Input the target affine transformed image and the third skin mask image corresponding to the target affine transformed image into the dual-frequency domain rule learning model to perform high-frequency and low-frequency restoration rule learning to obtain the target key components corresponding to the target affine transformed image. The target key components are used for face restoration. The dual-frequency domain rule learning model is obtained by using the training method of the dual-frequency domain rule learning model in any of the above method embodiments.
[0148] Among them, the implementation principle and technical effects of S604 can refer to the aforementioned embodiments, and the target key components may include a target multiplicative coefficient map, a target additive coefficient map and a target fusion weight map.
[0149] Considering that the target key components include the target multiplicative coefficient graph, the target additive coefficient graph and the target fusion weight graph, therefore, in the embodiment of the present disclosure, Figure 5 The step S504 may further include the following three steps S605 to S607:
[0150] S605. Perform inverse affine transformation on the target multiplicative coefficient map, the target additive coefficient map and the target fusion weight map respectively to obtain the target multiplicative coefficient map after inverse affine transformation, the target additive coefficient map after inverse affine transformation and the target fusion weight map after inverse affine transformation. The resolutions of the target multiplicative coefficient map after inverse affine transformation, the target additive coefficient map after inverse affine transformation and the target fusion weight map after inverse affine transformation are the same as the resolution of the face image to be processed.
[0151] It can be understood that the inverse affine transformation is the inverse operation of the affine transformation. In the embodiment of the present disclosure, the affine transformation is to perform face alignment and resolution reduction processing on the face image samples. Accordingly, the resolutions of the target multiplicative coefficient map, the target additive coefficient map, and the target fusion weight map can be enlarged to the resolution of the face image to be processed by the inverse affine transformation, and the face inverse alignment processing can be performed, so as to obtain the target multiplicative coefficient map after the inverse affine transformation, the target additive coefficient map after the inverse affine transformation, and the target fusion weight map after the inverse affine transformation.
[0152] S606, performing face restoration processing based on the face image to be processed, the multiplicative coefficient map after the target inverse affine transformation, the additive coefficient map after the target inverse affine transformation, and the fusion weight map after the target inverse affine transformation to obtain a target restoration result.
[0153] In this step, the multiplicative coefficient map after the target inverse affine transformation, the additive coefficient map after the target inverse affine transformation and the fusion weight map after the target inverse affine transformation can be applied to the face image to be processed to perform face restoration processing, thereby obtaining the target restoration result.
[0154] Further, optionally, performing face restoration processing based on the face image to be processed, the target multiplicative coefficient map after inverse affine transformation, the target additive coefficient map after inverse affine transformation and the target fusion weight map after inverse affine transformation to obtain the target restoration result can include: multiplying the pixels at the same position in the face image to be processed and the target multiplicative coefficient map after inverse affine transformation pixel by pixel to obtain a third restored image; adding the pixels at the same position in the third restored image and the target additive coefficient map after inverse affine transformation pixel by pixel to obtain a fourth restored image; and fusing the fourth restored image with the face image to be processed according to the fusion weight map after inverse affine transformation to obtain the target restoration result.
[0155] Exemplarily, referring to the above formula 4, the face image to be processed is taken as I_origin, and the face image sample is adjusted by the multiplicative coefficient map (a) after the target inverse affine transformation and the additive coefficient map (b) after the target inverse affine transformation, and the smooth transition between the adjusted image and the face image to be processed is controlled by the fusion weight map (α) after the target inverse affine transformation, and finally the target restoration result (dst) is obtained to achieve local and global restoration of facial defects.
[0156] S607: Perform fusion processing based on the target restoration result, the face image to be processed and the fourth skin mask image corresponding to the face image to be processed to obtain a target face restoration image corresponding to the face image to be processed.
[0157] In this step, based on the current relevant technology, the fourth skin mask image corresponding to the face image to be processed can be obtained in advance. The fourth skin mask image is a single-channel image, for example, 1 represents the skin area and 0 represents the non-skin area. Based on the skin area determined by the fourth skin mask image, the target restoration result and the face image to be processed are fused to obtain the target face restoration image corresponding to the face image to be processed, so that the face restoration process only acts on the skin area, thereby ensuring that the non-skin area maintains the original image information and avoids unnecessary interference.
[0158] Further, optionally, performing fusion processing based on the target repair result, the face image to be processed and the fourth skin mask image corresponding to the face image to be processed to obtain the target face repair image corresponding to the face image to be processed can include: performing skin area fusion processing on the target repair result and the face image to be processed according to the fourth skin mask image to obtain the target face repair image corresponding to the face image to be processed.
[0159] For example, referring to the above formula 5, the face image to be processed is taken as I_origin, the fourth skin mask image is taken as skin_mask, and the target repair result is taken as dst. Through the above formula 5, the target face repair image corresponding to the face image to be processed can be obtained. Among them, the use of the fourth skin mask image can make the face repair process only act on the skin area, thereby ensuring that the non-skin area (such as background or accessories, etc.) maintains the original image information, avoiding unnecessary interference, and finally obtaining a natural and delicate target face repair image at high resolution.
[0160] In the disclosed embodiment, facial key point detection is performed on the face image to be processed to determine the facial key point information in the face image to be processed; based on the facial key point information, a target affine transformation matrix is obtained; according to the target affine transformation matrix, face alignment is performed on the face image to be processed and resolution reduction processing is performed to obtain a target affine transformed image, and the face alignment processing can eliminate feature offsets caused by different postures and angles of the face, and can make the restoration rule learning performed at a small resolution more stable; the target affine transformed image and the third skin mask image corresponding to the target affine transformed image are input into a dual-frequency domain rule learning model to perform high-frequency and low-frequency restoration rule learning to obtain a target key component for face restoration, and the target key component includes a target multiplicative coefficient map, a target additive coefficient map and a target fusion weight Figure; perform inverse affine transformation on the target multiplicative coefficient map, the target additive coefficient map and the target fusion weight map respectively to obtain the target multiplicative coefficient map after inverse affine transformation, the target additive coefficient map after inverse affine transformation and the target fusion weight map after inverse affine transformation, and the resolutions of the target multiplicative coefficient map after inverse affine transformation, the target additive coefficient map after inverse affine transformation and the target fusion weight map after inverse affine transformation are all the same as the resolution of the face image to be processed; perform face restoration processing based on the face image to be processed, the multiplicative coefficient map after the target inverse affine transformation, the additive coefficient map after the target inverse affine transformation and the fusion weight map after the target inverse affine transformation to obtain the target restoration result; perform fusion processing based on the target restoration result, the face image to be processed and the fourth skin mask image corresponding to the face image to be processed to obtain the target face restoration image corresponding to the face image to be processed. The dual-frequency domain rule learning model is obtained by using the dual-frequency domain rule learning model training method in any of the above method embodiments, and can accurately output the target key components with a small computational overhead, thereby significantly reducing the computing power and storage requirements of high-resolution face restoration, improving image processing performance, and achieving natural and realistic face restoration. The solution is simple and easy to implement, and can be quickly deployed on different hardware platforms (such as mobile terminals and embedded devices, etc.).
[0161] Based on the above embodiments, the current related technology solutions that are completely based on full-resolution large models for direct restoration can learn complete facial pixels at one time, but they require extremely high computing resources and are difficult to apply in mobile / low computing power environments. For the current related technology solutions that use multi-task networks for joint learning (face detection, segmentation, restoration, etc.), although some goals can be achieved with ideal data and computing power support, training and deployment are more complicated and do not highlight the lightweight advantage of the small resolution of the disclosed embodiments.
[0162] In summary, the technical solution provided by the present disclosure has at least the following advantages:
[0163] (1) Dual-frequency domain (high frequency and low frequency) rule learning at small resolution can effectively reduce the huge computing power consumption of full convolution or GAN generation on high-resolution images; the splitting of low-frequency branches and high-frequency branches can make it easier for the dual-frequency domain rule learning model to learn the repair rules of facial lighting / skin color (global) and local texture (defects) respectively, reducing the coupling pressure of the unified network on all frequency components.
[0164] (2) Through face alignment processing, facial postures can be unified, reducing the learning difficulty of small-resolution models; through skin mask images, interference areas such as the background can be effectively shielded; through the fusion weight map, it is possible to flexibly switch between the repaired result and the original image, ensuring that the processing result is natural and smooth, and is not prone to the plastic feeling or blemishes of excessive beauty.
[0165] (3) During the inverse affine transformation, bilinear interpolation and other methods are used to map the output of the dual-frequency domain rule learning model from the small resolution back to the original image resolution, thereby achieving small-resolution learning and original-resolution application, which can greatly reduce the scale and amount of computation of the dual-frequency domain rule learning model. Through the inverse affine transformation, the restoration result is pasted back to the original face position, which can complete the accurate restoration and synthesis of single or multiple faces.
[0166] (4) Applicable to multi-scenario and multi-platform deployment, such as real-time or offline processing on mobile terminals, personal computers (PCs), or servers; combined with skin masks, it can be quickly integrated and applied in social media, beauty photography, virtual live broadcasts, or medical beauty.
[0167] It can be seen that the present disclosure adopts dual-frequency domain rule learning in a lightweight dual-frequency domain rule learning model. By separating the low-frequency "multiplicative correction" and the high-frequency "additive repair", and combining the fusion weight map and skin mask screening, the removal of facial defects and the retention of details can be well balanced; in addition, the idea of small-resolution learning and original-resolution application can effectively reduce the computational overhead of high-resolution image processing and significantly improve the real-time and deployability of the algorithm. Compared with the current related technologies, the present disclosure can greatly save the amount of computing, and is specialized in high-resolution 4K~5K facial defect repair. By combining face alignment and skin mask guidance, the dual-frequency domain rule learning model can focus more on defect area correction at low resolution; and the current related technologies usually directly perform general repairs in full-resolution scenes or without skin area assistance, which makes it difficult to balance speed and fidelity.
[0168] In addition, compared with the traditional high-resolution fully convolutional network, the disclosed method innovatively uses the "linear mapping" idea of bilinear interpolation and small-resolution learning to maintain facial details while shortening the inference time. Based on the alpha fusion strategy of face alignment and skin mask, it is more effective to concentrate on processing defective areas while maintaining fidelity, avoiding excessive modification of non-defective areas such as eyes, mouth, and hair. The disclosed solution is simple and easy to implement, and can greatly reduce the computing power and storage requirements for high-resolution face blemish removal. The dual-frequency domain rule learning model is lightweight, the post-processing process is reasonable and clear, and can be quickly deployed on different hardware platforms.
[0169] The following are embodiments of the device disclosed herein, which can be used to execute the method embodiments disclosed herein. For details not disclosed in the device embodiments disclosed herein, please refer to the method embodiments disclosed herein.
[0170] Figure 7 is a schematic diagram according to the sixth embodiment of the present disclosure. Figure 7 As shown, the training device 700 of the dual-frequency domain rule learning model provided in the sixth embodiment of the present disclosure includes an acquisition unit 701, a processing unit 702, a learning unit 703, a repair unit 704 and an adjustment unit 705. Among them:
[0171] The acquisition unit 701 is used to acquire a training sample pair, where the training sample pair includes a face image sample and a face restoration image label corresponding to the face image sample.
[0172] The processing unit 702 is used to perform affine transformation processing on the face image sample to obtain an affine transformed image, wherein the resolution of the affine transformed image is smaller than the resolution of the face image sample.
[0173] The learning unit 703 is used to input the affine transformed image and the first skin mask image corresponding to the affine transformed image into the dual-frequency domain rule learning model to perform high-frequency and low-frequency restoration rule learning to obtain the key components corresponding to the affine transformed image, which are used for face restoration.
[0174] The restoration unit 704 is used to apply the key component to perform face restoration processing on the face image sample to obtain a face restoration image corresponding to the face image sample.
[0175] The adjustment unit 705 is used to adjust the model parameters of the dual-frequency domain rule learning model based on the face restoration image and the face restoration image label.
[0176] In some embodiments, the key components include a multiplicative coefficient map, an additive coefficient map and a fusion weight map, and the learning unit 703 may include: an extraction module (not shown in the figure), which is used to input the affine transformed image and the first skin mask image corresponding to the affine transformed image into a feature extraction module of the dual-frequency domain rule learning model, and the feature extraction module extracts features from the affine transformed image under the masking effect of the first skin mask image to obtain global face features and local face features; a deconvolution processing module (not shown in the figure), which is used to input the global face features into the deconvolution block of the dual-frequency domain rule learning model for deconvolution processing to obtain a multiplicative coefficient map; a first convolution processing module (not shown in the figure), which is used to input the local face features into the first convolution block of the dual-frequency domain rule learning model for convolution processing to obtain an additive coefficient map; a second convolution processing module (not shown in the figure), which is used to input the global face features and the local face features into the second convolution block of the dual-frequency domain rule learning model for convolution processing to obtain a fusion weight map.
[0177] In some embodiments, the repair unit 704 may include: an inverse affine transformation processing module (not shown in the figure), which is used to perform inverse affine transformation processing on the multiplicative coefficient map, the additive coefficient map and the fusion weight map respectively, to obtain the multiplicative coefficient map after the inverse affine transformation, the additive coefficient map after the inverse affine transformation and the fusion weight map after the inverse affine transformation, and the resolution of the multiplicative coefficient map after the inverse affine transformation, the additive coefficient map after the inverse affine transformation and the fusion weight map after the inverse affine transformation are all the same as the resolution of the face image sample; a repair processing module (not shown in the figure), which is used to perform face repair processing based on the face image sample, the multiplicative coefficient map after the inverse affine transformation, the additive coefficient map after the inverse affine transformation and the fusion weight map after the inverse affine transformation to obtain a repair result; a fusion processing module (not shown in the figure), which is used to perform fusion processing based on the repair result, the face image sample and the second skin mask image corresponding to the face image sample to obtain a face repaired image corresponding to the face image sample.
[0178] In some embodiments, the repair processing module may include: a multiplication submodule (not shown in the figure), which is used to multiply the face image sample and the pixels at the same position in the multiplicative coefficient map after the inverse affine transformation pixel by pixel to obtain a first repaired image; an addition submodule (not shown in the figure), which is used to add the pixels at the same position in the first repaired image and the additive coefficient map after the inverse affine transformation pixel by pixel to obtain a second repaired image; a fusion processing submodule (not shown in the figure), which is used to fuse the second repaired image and the face image sample according to the fusion weight map after the inverse affine transformation to obtain a repair result.
[0179] In some embodiments, the fusion processing module may include: a skin area fusion processing submodule (not shown in the figure), which is used to perform skin area fusion processing on the repair result and the face image sample according to the second skin mask image to obtain a face repaired image corresponding to the face image sample.
[0180] In some embodiments, the adjustment unit 705 may include: a loss acquisition module (not shown in the figure), used to obtain the mean square error loss and the perceptual loss based on the face restoration image and the face restoration image label; an adjustment module (not shown in the figure), used to adjust the model parameters of the dual-frequency domain rule learning model through back propagation and optimization iteration according to the mean square error loss and the perceptual loss.
[0181] In some embodiments, the processing unit 702 may include: a detection module (not shown in the figure), which is used to perform facial key point detection on the face image sample and determine the facial key point information in the face image sample; an acquisition module (not shown in the figure), which is used to obtain the affine transformation matrix based on the facial key point information; a processing module (not shown in the figure), which is used to perform face alignment and resolution reduction processing on the face image sample according to the affine transformation matrix to obtain an image after affine transformation.
[0182] Figure 7 The provided dual-frequency domain rule learning model training device can execute the steps in the method embodiment corresponding to the above-mentioned dual-frequency domain rule learning model training method. Its implementation principle and technical effects are similar and will not be repeated here.
[0183] Figure 8 is a schematic diagram according to the seventh embodiment of the present disclosure. Figure 8 As shown, the image processing device 800 provided in the seventh embodiment of the present disclosure includes a first acquisition unit 801, an affine transformation processing unit 802, a second acquisition unit 803 and a restoration processing unit 804. Among them:
[0184] The first acquisition unit 801 is used to acquire a face image to be processed.
[0185] The affine transformation processing unit 802 is used to perform affine transformation post-processing on the face image to be processed to obtain a target affine transformed image, wherein the resolution of the target affine transformed image is smaller than the resolution of the face image to be processed.
[0186] The second acquisition unit 803 is used to input the target affine transformed image and the third skin mask image corresponding to the target affine transformed image into the dual-frequency domain rule learning model to perform high-frequency and low-frequency restoration rule learning to obtain the target key components corresponding to the target affine transformed image. The target key components are used for face restoration. The dual-frequency domain rule learning model is obtained by adopting the training method of the dual-frequency domain rule learning model in any of the above method embodiments.
[0187] The restoration processing unit 804 is used to perform face restoration processing on the face image to be processed by applying the target key component to obtain a target face restoration image corresponding to the face image to be processed.
[0188] In some embodiments, the target key components include a target multiplicative coefficient map, a target additive coefficient map, and a target fusion weight map. The repair processing unit 804 may include: an inverse affine transformation processing module (not shown in the figure), which is used to perform inverse affine transformation processing on the target multiplicative coefficient map, the target additive coefficient map, and the target fusion weight map, respectively, to obtain a target multiplicative coefficient map after inverse affine transformation, a target additive coefficient map after inverse affine transformation, and a target fusion weight map after inverse affine transformation, and a target multiplicative coefficient map after inverse affine transformation, a target additive coefficient map after inverse affine transformation, and a target fusion weight map after inverse affine transformation. The resolution of the fusion weight map is the same as the resolution of the face image to be processed; the repair processing module (not shown in the figure) is used to perform face repair processing based on the face image to be processed, the multiplicative coefficient map after the target inverse affine transformation, the additive coefficient map after the target inverse affine transformation and the fusion weight map after the target inverse affine transformation to obtain the target repair result; the fusion processing module (not shown in the figure) is used to perform fusion processing based on the target repair result, the face image to be processed and the fourth skin mask image corresponding to the face image to be processed to obtain the target face repaired image corresponding to the face image to be processed.
[0189] In some embodiments, the repair processing module may include: a multiplication submodule (not shown in the figure), which is used to multiply the pixels at the same position in the face image to be processed and the target multiplication coefficient map after the inverse affine transformation pixel by pixel to obtain a third repaired image; an addition submodule (not shown in the figure), which is used to add the pixels at the same position in the third repaired image and the target additive coefficient map after the inverse affine transformation pixel by pixel to obtain a fourth repaired image; a fusion processing submodule (not shown in the figure), which is used to fuse the fourth repaired image and the face image to be processed according to the fusion weight map after the target inverse affine transformation to obtain the target repair result.
[0190] In some embodiments, the fusion processing module may include: a skin area fusion processing sub-module (not shown in the figure), which is used to perform skin area fusion processing on the target repair result and the face image to be processed according to the fourth skin mask image to obtain the target face repair image corresponding to the face image to be processed.
[0191] In some embodiments, the affine transformation processing unit 802 may include: a detection module (not shown in the figure), which is used to perform facial key point detection on the face image to be processed and determine the facial key point information in the face image to be processed; an acquisition module (not shown in the figure), which is used to obtain the target affine transformation matrix based on the facial key point information; a processing module (not shown in the figure), which is used to perform face alignment and resolution reduction processing on the face image to be processed according to the target affine transformation matrix to obtain the target affine transformed image.
[0192] Figure 8 The provided image processing device can execute the steps in the method embodiment corresponding to the above-mentioned image processing method. Its implementation principle and technical effects are similar and will not be repeated here.
[0193] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the solution provided by any of the above embodiments.
[0194] According to an embodiment of the present disclosure, the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute a solution provided by any of the above embodiments.
[0195] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, which includes: a computer program, the computer program is stored in a readable storage medium, at least one processor of an electronic device can read the computer program from the readable storage medium, and at least one processor executes the computer program so that the electronic device executes the solution provided by any of the above embodiments.
[0196] Fig. 9 Schematic block diagram of an example electronic device 900 that can be used to implement an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0197] like Fig. 9As shown, the electronic device 900 includes a computing unit 901, which can store data stored in a read-only memory (ROM) ( Fig. 9 A computer program in ROM 902 is loaded from storage unit 908 to random access memory (RAM) ( Fig. 9 The computer program in the RAM 903 (for example) can be used to perform various appropriate actions and processes. The RAM 903 can also store various programs and data required for the operation of the electronic device 900. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. The input / output (I / O) interface ( Fig. 9 The I / O interface 905 (for example) is also connected to the bus 904 .
[0198] Multiple components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0199] The computing unit 901 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSP), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as the training method and image processing method of the dual-frequency domain rule learning model. For example, in some embodiments, the training method and image processing method of the dual-frequency domain rule learning model may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the training method of the dual-frequency domain rule learning model and the image processing method described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to perform the training method of the dual-frequency domain rule learning model or the image processing method by any other appropriate means (e.g., by means of firmware).
[0200] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0201] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0202] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0203] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0204] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: Local Area Networks (LANs), Wide Area Networks (WANs), and the Internet.
[0205] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server of a distributed system, or a server combined with a blockchain.
[0206] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0207] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, fusions, sub-fusions and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present disclosure should be included in the protection scope of the present disclosure.
Claims
1. A training method for a dual-frequency domain rule learning model, comprising: Acquire a training sample pair, wherein the training sample pair includes a face image sample and a face restoration image label corresponding to the face image sample; Performing affine transformation processing on the face image sample to obtain an affine transformed image, wherein the resolution of the affine transformed image is smaller than the resolution of the face image sample; Inputting the affine transformed image and the first skin mask image corresponding to the affine transformed image into a dual-frequency domain rule learning model to perform high-frequency and low-frequency restoration rule learning, and obtaining key components corresponding to the affine transformed image, wherein the key components are used for face restoration; Applying the key component to perform face restoration processing on the face image sample to obtain a face restoration image corresponding to the face image sample; Based on the face restoration image and the face restoration image label, model parameters of the dual-frequency domain rule learning model are adjusted.
2. The training method according to claim 1, wherein: The key components include a multiplicative coefficient map, an additive coefficient map, and a fusion weight map. The affine transformed image and the first skin mask image corresponding to the affine transformed image are input into a dual-frequency domain rule learning model to perform high-frequency and low-frequency restoration rule learning to obtain the key components corresponding to the affine transformed image, including: Inputting the affine transformed image and the first skin mask image corresponding to the affine transformed image into the feature extraction module of the dual-frequency domain rule learning model, wherein the feature extraction module extracts features from the affine transformed image under the masking effect of the first skin mask image to obtain global features and local features of the face; Inputting the global face feature into the deconvolution block of the dual-frequency domain rule learning model for deconvolution processing to obtain the multiplicative coefficient map; Inputting the local features of the human face into the first convolution block of the dual-frequency domain rule learning model for convolution processing to obtain the additive coefficient map; The global face features and the local face features are input into the second convolution block of the dual-frequency domain rule learning model for convolution processing to obtain the fusion weight map.
3. The training method according to claim 2, wherein: The applying the key component to perform face restoration processing on the face image sample to obtain a face restoration image corresponding to the face image sample includes: Performing inverse affine transformation processing on the multiplicative coefficient map, the additive coefficient map and the fusion weight map respectively to obtain a multiplicative coefficient map after inverse affine transformation, an additive coefficient map after inverse affine transformation and a fusion weight map after inverse affine transformation, wherein the resolutions of the multiplicative coefficient map after inverse affine transformation, the additive coefficient map after inverse affine transformation and the fusion weight map after inverse affine transformation are all the same as the resolution of the face image sample; Performing face restoration processing based on the face image sample, the multiplicative coefficient map after the inverse affine transformation, the additive coefficient map after the inverse affine transformation, and the fusion weight map after the inverse affine transformation to obtain a restoration result; A fusion process is performed based on the restoration result, the face image sample and the second skin mask image corresponding to the face image sample to obtain a face restoration image corresponding to the face image sample.
4. The training method according to claim 3, wherein: The performing face restoration processing based on the face image sample, the multiplicative coefficient map after the inverse affine transformation, the additive coefficient map after the inverse affine transformation and the fusion weight map after the inverse affine transformation to obtain the restoration result includes: Multiplying the face image sample and the pixels at the same position in the multiplicative coefficient map after the inverse affine transformation pixel by pixel to obtain a first restored image; Adding pixels at the same position in the first restored image and the additive coefficient map after inverse affine transformation pixel by pixel to obtain a second restored image; The second restored image and the face image sample are fused according to the fusion weight map after the inverse affine transformation to obtain the restoration result.
5. The training method according to claim 3, wherein: The step of performing fusion processing based on the restoration result, the face image sample and the second skin mask image corresponding to the face image sample to obtain a face restoration image corresponding to the face image sample includes: The restoration result and the face image sample are subjected to skin region fusion processing according to the second skin mask image to obtain a face restoration image corresponding to the face image sample.
6. The training method according to any one of claims 1 to 5, wherein: The adjusting the model parameters of the dual-frequency domain rule learning model based on the face restoration image and the face restoration image label includes: Based on the face restoration image and the face restoration image label, obtaining a mean square error loss and a perceptual loss; According to the mean square error loss and the perceptual loss, the model parameters of the dual-frequency domain rule learning model are adjusted through back propagation and optimization iteration.
7. The training method according to any one of claims 1 to 5, wherein: The performing affine transformation on the face image sample to obtain an affine transformed image includes: Performing facial key point detection on the facial image sample to determine facial key point information in the facial image sample; Based on the facial key point information, obtaining an affine transformation matrix; According to the affine transformation matrix, the face image samples are subjected to face alignment and resolution reduction processing to obtain the affine transformed image.
8. An image processing method, comprising: Obtain the face image to be processed; Performing affine transformation post-processing on the face image to be processed to obtain a target affine transformed image, wherein the resolution of the target affine transformed image is smaller than the resolution of the face image to be processed; Inputting the target affine transformed image and a third skin mask image corresponding to the target affine transformed image into a dual-frequency domain rule learning model to perform high-frequency and low-frequency restoration rule learning, and obtaining a target key component corresponding to the target affine transformed image, wherein the target key component is used for face restoration, and the dual-frequency domain rule learning model is obtained by using the training method according to any one of claims 1 to 7; The target key component is applied to perform face restoration processing on the face image to be processed to obtain a target face restoration image corresponding to the face image to be processed.
9. The image processing method according to claim 8, wherein: The target key components include a target multiplicative coefficient map, a target additive coefficient map, and a target fusion weight map. The method of applying the target key components to perform face restoration processing on the face image to be processed to obtain a target face restoration image corresponding to the face image to be processed includes: Respectively performing inverse affine transformation processing on the target multiplicative coefficient map, the target additive coefficient map and the target fusion weight map to obtain a target inverse affine transformed multiplicative coefficient map, a target inverse affine transformed additive coefficient map and a target inverse affine transformed fusion weight map, wherein the resolutions of the target inverse affine transformed multiplicative coefficient map, the target inverse affine transformed additive coefficient map and the target inverse affine transformed fusion weight map are the same as the resolution of the face image to be processed; Performing face restoration processing based on the face image to be processed, the target multiplicative coefficient map after inverse affine transformation, the target additive coefficient map after inverse affine transformation, and the target fusion weight map after inverse affine transformation to obtain a target restoration result; A fusion process is performed based on the target restoration result, the face image to be processed and a fourth skin mask image corresponding to the face image to be processed to obtain a target face restoration image corresponding to the face image to be processed.
10. The image processing method according to claim 9, wherein: The step of performing face restoration processing based on the face image to be processed, the target multiplicative coefficient map after inverse affine transformation, the target additive coefficient map after inverse affine transformation, and the target fusion weight map after inverse affine transformation to obtain a target restoration result includes: Multiplying pixels at the same position in the face image to be processed and the target inverse affine transformed multiplicative coefficient map pixel by pixel to obtain a third restored image; Adding pixels at the same position in the third restored image and the target inverse affine transformed additive coefficient map pixel by pixel to obtain a fourth restored image; The fourth repaired image and the face image to be processed are fused according to the target inverse affine transformed fusion weight map to obtain the target repair result.
11. The image processing method according to claim 9, wherein: The step of performing fusion processing based on the target restoration result, the face image to be processed, and a fourth skin mask image corresponding to the face image to be processed to obtain a target face restoration image corresponding to the face image to be processed includes: The target face restoration result and the face image to be processed are subjected to skin region fusion processing according to the fourth skin mask image to obtain a target face restoration image corresponding to the face image to be processed.
12. The image processing method according to any one of claims 8 to 11, wherein: The step of performing affine transformation post-processing on the face image to be processed to obtain a target affine transformed image includes: Performing facial key point detection on the face image to be processed to determine facial key point information in the face image to be processed; Based on the facial key point information, obtaining a target affine transformation matrix; According to the target affine transformation matrix, the face image to be processed is subjected to face alignment and resolution reduction processing to obtain the target affine transformed image.
13. A training device for a dual-frequency domain rule learning model, comprising: An acquisition unit, configured to acquire a training sample pair, wherein the training sample pair includes a face image sample and a face restoration image label corresponding to the face image sample; A processing unit, configured to perform affine transformation on the face image sample to obtain an affine transformed image, wherein the resolution of the affine transformed image is smaller than the resolution of the face image sample; A learning unit, used for inputting the affine transformed image and the first skin mask image corresponding to the affine transformed image into a dual-frequency domain rule learning model to perform high-frequency and low-frequency restoration rule learning, and obtaining key components corresponding to the affine transformed image, wherein the key components are used for face restoration; A restoration unit, configured to perform face restoration processing on the face image sample by applying the key component to obtain a face restoration image corresponding to the face image sample; An adjustment unit is used to adjust the model parameters of the dual-frequency domain rule learning model based on the face restoration image and the face restoration image label.
14. An image processing device, comprising: A first acquisition unit, used to acquire a face image to be processed; An affine transformation processing unit, configured to perform affine transformation post-processing on the face image to be processed to obtain a target affine transformed image, wherein the resolution of the target affine transformed image is smaller than the resolution of the face image to be processed; a second acquisition unit, configured to input the target affine transformed image and a third skin mask image corresponding to the target affine transformed image into a dual-frequency domain rule learning model to perform high-frequency and low-frequency restoration rule learning, and obtain a target key component corresponding to the target affine transformed image, wherein the target key component is used for face restoration, and the dual-frequency domain rule learning model is obtained by using the training method according to any one of claims 1 to 7; The restoration processing unit is used to apply the target key component to perform face restoration processing on the face image to be processed, so as to obtain a target face restoration image corresponding to the face image to be processed.
15. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the training method of the dual-frequency domain rule learning model according to any one of claims 1 to 7 or the image processing method according to any one of claims 8 to 12.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the training method of the dual-frequency domain rule learning model according to any one of claims 1 to 7 or the image processing method according to any one of claims 8 to 12.
17. A computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the training method of the dual-frequency domain rule learning model according to any one of claims 1 to 7 or the steps of the image processing method according to any one of claims 8 to 12.