A method, device, computer device and medium for unmasking face restoration
By designing a demask face recovery model including a hollow interpolation convolution module, a dynamic selection convolution module and a context attention module, the problem of poor mask face recovery effect in the prior art is solved, and a more efficient and diverse facial image recovery effect is achieved.
Patent Information
- Application Number
- CN202210809438.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-11
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-07-11
AI Technical Summary
The prior art has poor results in restoring masked face images, lacks diversity, and is difficult to be fidelity.
A demasked face recovery model is designed, including a hollow interpolation convolution module, a multi-layer dynamic selection convolution module, a context attention module and a feature fusion module. Through the combination of these modules and the preprocessing of training set data, efficient recovery of the masked face image is achieved.
The effect of masking the face is improved, and the facial images generated are more diverse, have good applicability and high efficiency.
Smart Images

Figure CN115223012B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of deep learning technology, and in particular to a method, apparatus, computer device and storage medium for demasking human face restoration. Background Art
[0002] Masks have become an effective way to slow the spread of disease. In public places, people's faces are covered, so facial restoration technology for masked faces has gradually developed.
[0003] Facial restoration is a small area in the field of Computer Vision (CV). Traditional methods include patch-based methods and diffusion model-based methods. The patch-based method searches and expands pixels in the intact area of the image to fill the missing area patch by patch, while the diffusion model-based method fills the missing area with content. The images restored by both methods are difficult to maintain and lack diversity. Therefore, the existing technology has the problem of poor effect and adaptability. Summary of the invention
[0004] Based on this, it is necessary to provide a method, device, computer equipment and storage medium for demasking face restoration that can improve the effect of facial restoration in response to the above technical problems.
[0005] A method for removing masked face and restoring the face, the method comprising:
[0006] Acquire training set data of face images, preprocess the training set data, and obtain corresponding face mask set data; in the face mask set data, the face and mouth of the face image are replaced by a square mask;
[0007] Input the training set data and the face mask set data into the face restoration model without mask; the face restoration model without mask includes a first path, a second path, a feature fusion module, and an image output module; the first path includes an atrous interpolation convolution module, a multi-layer dynamic selection convolution module, and a multi-layer atrous convolution module; the multi-layer dynamic selection convolution module and the multi-layer atrous convolution module form a U-shaped convolution network; the second path is a U-shaped convolution network, including a multi-layer dynamic selection convolution module and a context attention module; the atrous interpolation convolution module is used to fill holes in the face mask images in the face mask set data by filling in noise; the dynamic selection convolution modules in the first path and the second path are used to obtain high-weight features through the softmax function; the atrous convolution module is used to perform feature extraction with an enlarged receptive field; the context attention module is used to borrow effective spatial pixels for hole filling; the feature fusion module is used to fuse the features output by the first path and the second path; the image output module is used to activate the output of the feature fusion module to obtain the generated image of the face;
[0008] Train the face restoration model without mask according to the real images and the corresponding generated images in the training set data through a pre-set loss function to obtain a trained face restoration model without mask;
[0009] Obtain the face mask image to be processed, and input the face mask image into the trained face restoration model without mask to obtain the restored face image without mask.
[0010] In one embodiment, it further includes: obtaining the training set data of face images; the training set data is randomly collected from the public dataset celeba;
[0011] For each face image in the training set data, obtain 68 facial feature points through the trained dlib network, determine the square mask range, and obtain the face mask image according to the square mask range;
[0012] Further obtain the face mask set data.
[0013] In one embodiment, it further includes: the mathematical representation corresponding to the dynamic selection convolution module is:
[0014]
[0015] where Output is the output of the dynamic selection convolution module, represents the feature after convolution, and σ(·) represents the weight information obtained by the softmax function.
[0016] In one embodiment, it further includes: the hole interpolation convolution module is used to add a noise filling module on the basis of the deformable convolution module, perform feature fusion on the image features learned by the noise filling module and the deformable convolution module, and fill holes in the face mask images in the face mask set data.
[0017] In one embodiment, the processing flow of the convolution module for superimposing noise includes:
[0018] Normalize the face mask images in the face mask set data channel by channel;
[0019] Superimpose noise on the normalized image;
[0020] Perform 3×3 convolution on the image after superimposing noise;
[0021] Normalize the image after convolution channel by channel again to obtain the output of the noise filling module.
[0022] In one embodiment, it further includes: training the face restoration model without mask through a preset loss function; the loss function of the generator in the face restoration model without mask includes L1 loss function, Ltv loss function and L content loss function; the objective function to be optimized by the face restoration model without mask is WGAN loss.
[0023] In one embodiment, it further includes: the face images in the training set data are frontal face images.
[0024] A device for face restoration without mask, the device includes:
[0025] A preprocessing module, configured to obtain training set data of face images, preprocess the training set data, and obtain corresponding face mask set data; in the face mask set data, the face and mouth of the face image are replaced by square masks;
[0026] A training data input module for inputting the training set data and the face mask set data into a face restoration model without mask; the face restoration model without mask includes a first path, a second path, a feature fusion module, and an image output module; the first path includes an atrous interpolation convolution module, a multi-layer dynamic selection convolution module, and a multi-layer atrous convolution module; the multi-layer dynamic selection convolution module and the multi-layer atrous convolution module form a U-shaped convolution network; the second path is a U-shaped convolution network, including a multi-layer dynamic selection convolution module and a context attention module; the atrous interpolation convolution module is used to fill holes in the face mask images in the face mask set data by filling in noise; the dynamic selection convolution modules in the first path and the second path are used to obtain high-weight features through the softmax function; the atrous convolution module is used to perform feature extraction for expanding the receptive field; the context attention module is used to borrow effective spatial pixels for hole filling; the feature fusion module is used to perform feature fusion on the outputs of the first path and the second path; the image output module is used to activate the output of the feature fusion module to obtain a generated image of the face;
[0027] A model training module for training the face restoration model without mask according to the real images and the corresponding generated images in the training set data through a pre-set loss function to obtain a trained face restoration model without mask;
[0028] A model application module for obtaining a face mask image to be processed, inputting the face mask image into the trained face restoration model without mask, and obtaining a restored face image without mask.
[0029] A computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0030] Obtain the training set data of face images, preprocess the training set data to obtain the corresponding face mask set data; in the face mask set data, the face and mouth parts of the face images are replaced by square masks;
[0031] Input the training set data and the face mask set data into the face restoration model without mask; the face restoration model without mask includes a first path, a second path, a feature fusion module, and an image output module; the first path includes an atrous interpolation convolution module, a multi-layer dynamic selection convolution module, and a multi-layer atrous convolution module; the multi-layer dynamic selection convolution module and the multi-layer atrous convolution module form a U-shaped convolution network; the second path is a U-shaped convolution network, including a multi-layer dynamic selection convolution module and a context attention module; the atrous interpolation convolution module is used to fill holes in the face mask image in the face mask set data by filling in noise; the dynamic selection convolution modules in the first path and the second path are used to obtain high-weight features through the softmax function; the atrous convolution module is used to perform feature extraction with an enlarged receptive field; the context attention module is used to borrow effective spatial pixels for hole filling; the feature fusion module is used to fuse the features of the outputs of the first path and the second path; the image output module is used to activate the output of the feature fusion module to obtain a generated image of the face;
[0032] Train the face restoration model without mask according to the real images and the corresponding generated images in the training set data through a pre-set loss function to obtain a trained face restoration model without mask;
[0033] Obtain a face mask image to be processed, and input the face mask image into the trained face restoration model without mask to obtain a restored face image without mask.
[0034] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0035] Obtain the training set data of face images, preprocess the training set data to obtain the corresponding face mask set data; in the face mask set data, the face and mouth parts of the face images are replaced by square masks;
[0036] Input the training set data and the face mask set data into the face mask removal and restoration model; the face mask removal and restoration model includes a first path, a second path, a feature fusion module, and an image output module; the first path includes a hole interpolation convolution module, a multi-layer dynamic selection convolution module, and a multi-layer dilated convolution module; the multi-layer dynamic selection convolution module and the multi-layer dilated convolution module constitute a U-shaped convolutional network; the second path is a U-shaped convolutional network, including a multi-layer dynamic selection convolution module and a context attention module; the hole interpolation convolution module is used to fill holes in the face mask images in the face mask set data by filling in noise; the dynamic selection convolution modules in the first path and the second path are used to obtain high-weight features through the softmax function; the dilated convolution module is used to perform feature extraction with an enlarged receptive field; the context attention module is used to borrow valid spatial pixels for hole filling; the feature fusion module is used to fuse the features of the outputs of the first path and the second path; the image output module is used to activate the output of the feature fusion module to obtain a generated image of the face;
[0037] Train the face mask removal and restoration model according to the real images and the corresponding generated images in the training set data through a pre-set loss function to obtain a trained face mask removal and restoration model;
[0038] Obtain a face mask image to be processed, and input the face mask image into the trained face mask removal and restoration model to obtain a restored face image without the mask.
[0039] The above method, device, computer device, and storage medium for face mask removal and restoration can fill the pixels of the face mask holes by designing a hole interpolation convolution module in the face mask removal and restoration model, improve the model execution efficiency and the diversity of the generated faces; learn features by using the dynamic selection convolution module and extract high-weight features through attention weights, which can better learn image features; in addition, through the context attention module in the second path, information can be effectively borrowed from distant spatial positions to reconstruct locally missing pixels. Train the face mask removal and restoration model with the training set data to obtain a trained face mask removal and restoration model for removing the mask from masked faces such as those wearing masks. The present invention can improve the effect of mask removal for masked faces, has good applicability and high efficiency. Description of the Drawings
[0040] Figure 1 It is a schematic flowchart of the face mask removal and restoration method in an embodiment;
[0041] Figure 2Schematic diagram of image preprocessing in the face restoration method for removing masks in an embodiment, where (a) is the original image, (b) is the key point map, (c) is the face center map, and (d) is the preprocessing output map;
[0042] Figure 3 Overall framework diagram of the face restoration model for removing masks in an embodiment;
[0043] Figure 4 Schematic diagram of the dilated interpolation convolution module in an embodiment;
[0044] Figure 5 Schematic diagram for explaining the principle of the dilated interpolation convolution module in an embodiment, where (a) is the schematic diagram of conventional convolution, (b) is the schematic diagram of deformable convolution, (c) is the schematic diagram of the defects of deformable convolution, and (d) is the schematic diagram of the dilated interpolation convolution module filling holes;
[0045] Figure 6 Schematic diagram of the structure of the dynamic selection convolution module in an embodiment;
[0046] Figure 7 Result diagram of testing with the test set in a specific embodiment;
[0047] Figure 8 Structure block diagram of the device for face restoration with mask removal in an embodiment;
[0048] Figure 9 Internal structure diagram of a computer device in an embodiment. Specific implementation mode
[0049] In order to make the objectives, technical solutions, and advantages of this application clearer, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0050] In one embodiment, as Figure 1 shown, a method for face restoration with mask removal is provided, including the following steps:
[0051] Step 102: Obtain the training set data of the face image, preprocess the training set data, and obtain the corresponding face mask set data.
[0052] In the image preprocessing stage, first use the trained network model, the dlib network, to obtain 68 feature points of the face and create a face mask image (using a square mask as a substitute). Set the mask image through the following function:
[0053] f(x) = x * (1 - mask)
[0054] where x represents the original image and mask represents the mask image. For example, Figure 2 is a schematic diagram of obtaining a face mask image from the original image.
[0055] Step 104: Input the training set data and the face mask set data into the face restoration model without mask.
[0056] The face restoration model without mask includes a first path, a second path, a feature fusion module, and an image output module; the first path includes an atrous interpolation convolution module, a multi-layer dynamic selection convolution module, and a multi-layer atrous convolution module; the multi-layer dynamic selection convolution module and the multi-layer atrous convolution module form a U-shaped convolutional network; the second path is a U-shaped convolutional network, including a multi-layer dynamic selection convolution module and a context attention module; the atrous interpolation convolution module is used to fill holes in the face mask images in the face mask set data by filling in noise; the dynamic selection convolution modules in the first path and the second path are used to obtain high-weight features through the softmax function; the atrous convolution module is used to perform feature extraction with an enlarged receptive field; the context attention module is used to borrow valid spatial pixels for hole filling; the feature fusion module is used to fuse the features output by the first path and the second path; the image output module is used to activate the output of the feature fusion module to obtain the generated image of the face.
[0057] Among them, the context attention module borrows valid spatial pixels for hole filling. The valid pixels refer to the pixel information borrowed from distant spatial positions (non-mask regions), which can be used to reconstruct locally missing pixels.
[0058] Specifically, for example, Figure 3 is the overall framework diagram of the face restoration model without mask. In the generator of the face restoration model without mask, the modules in the first row form the first path, and the modules in the second row form the second path. The first path sequentially includes a cascaded atrous interpolation convolution module, a multi-layer dynamic selection convolution module, and a multi-layer atrous convolution module; the second path sequentially includes a cascaded multi-layer dynamic convolution module and a context attention module; the features output by the first path and the second path are fused through a feature fusion module to obtain the predicted face image output by the generator. During the training process of the face restoration model without mask, the predicted face image output by the generator and the real face image are input into the discriminator, and a scalar value is output to represent the authenticity of the generated image.
[0059] The role of the atrous interpolation convolution module in the first path proposed by the present invention is to fill the pixels of the face mask module. Traditional convolution has a slow learning efficiency for face networks with masks. By adding an atrous interpolation convolution module to the model in the present invention, the execution efficiency of the model can be improved. The atrous interpolation convolution module is as Figure 4As shown in the figure: The dilated interpolation convolution module is used to add a noise filling module on the basis of the deformable convolution module, fuse the image features learned by the noise filling module and the deformable convolution module, and fill the holes in the face mask images in the face mask set data. Among them, the processing flow of the noise filling module includes: normalizing the face mask images in the face mask set data channel by channel; superimposing noise on the normalized images; performing 3×3 convolution on the images with superimposed noise; and normalizing the convolved images channel by channel again to obtain the output of the noise filling module.
[0060] The present invention proposes to use a deformable convolution module in the dilated interpolation convolution module and improve it by filling noise on the basis of the deformable convolution module. The reason is that: conventional convolution blocks cannot learn when facing holes. As Figure 5 (a) shows, the convolution kernel cannot collect information at the holes; while using deformable convolution enables the positions of the holes to be replaced by surrounding pixel points, as Figure 5 (b); however, there is a problem with deformable convolution. As Figure 5 (c) shows, it will replace the part with pixel values with parts with a value of 0; the dilated interpolation convolution module designed by the present invention solves the above problems, as Figure 5 (d); Figure 5 In the figure, the black box represents the convolution kernel, the solid dots represent the original sampling points, and the dotted circles represent the target sampling points. In addition, the dilated interpolation convolution module can improve the diversity of the generated faces.
[0061] The present invention also proposes to use a dynamic selection convolution module to learn features. The dynamic selection convolution module runs through the entire network and acts as a convolution block. As Figure 6 shown in the figure is the structure diagram of the dynamic selection convolution module. The softmax function is used to obtain the attention weights and extract the high-weight features. The mathematical expression is:
[0062]
[0063] where Output is the output of the dynamic selection convolution module, represents the features after convolution, and σ(·) represents the weight information obtained by the softmax function.
[0064] The face restoration model without mask also includes a discriminator. The role of the discriminator is to distinguish the authenticity of the generated image and the real image, so as to punish the generator and make the generator closer to the real image. The structure of the discriminator is shown in the following figure. It uses multiple ordinary convolutions. The discriminator takes the real image and the synthetic image as inputs and finally outputs a scalar value to represent the authenticity of the generated image.
[0065] The objective function to be optimized for the entire network is the WGAN loss. WGAN uses the Earth-Mover distance (EM distance) as the loss, which is the minimum cost under optimal path planning. It calculates the expected value of the distance between sample pairs under the joint distribution:
[0066]
[0067] where x is the real sampled data and z is the noise data.
[0068] Step 106: According to the real images and the corresponding generated images in the training set data, train the face restoration model without mask through a pre-set loss function to obtain a trained face restoration model without mask.
[0069] The loss function of the generator in the face restoration model without mask includes the L1 loss function, the Ltv loss function, and the L content loss function. Specifically as follows:
[0070] L1 loss function:
[0071]
[0072] where, represents the real image of the i-th masked part, represents the image of the generated masked part
[0073]
[0074] where, represents the i-th global real image, represents the generated global image
[0075] Ltv loss function:
[0076] The total differential regularization loss: Artifacts usually exist in the images generated by the GAN model, resulting in blurred generated images and affecting recognition. Adding a total regularization term to the final generated image alleviates this problem. The mathematical expression of the local tv loss function is:
[0077]
[0078] The mathematical expression of the global tv loss function is:
[0079]
[0080] L content loss function:
[0081] The style loss uses the VGG network pre-trained on ImageNet.
[0082]
[0083] Among them, φ convi is the feature of the i-th convolutional layer of VGG-19.
[0084] Step 108: Obtain the face mask image to be processed, input the face mask image into the trained face restoration model without mask, and obtain the restored face image without mask.
[0085] The above method, device, computer device and storage medium for face restoration without mask can fill the pixels of the face mask holes by designing a dilated interpolation convolution module in the face restoration model without mask, improve the model execution efficiency and the diversity of the generated face; learn features by using the dynamic selection convolution module, and obtain high-weight features through attention weights, which can better learn the features of the image; in addition, through the context attention module in the second path, information can be effectively borrowed from distant spatial positions to reconstruct the locally missing pixels. Train the face restoration model without mask with the training set data to obtain a trained face restoration model without mask, which is used to perform mask removal processing on the masked face such as wearing a mask. The present invention can improve the effect of mask removal for masked faces, has good applicability and high efficiency.
[0086] In one embodiment, it further includes: obtaining the training set data of face images; the training set data is randomly collected from the public dataset celeba; for each face image in the training set data, obtain 68 facial feature points through the trained dlib network, determine the square mask range, and obtain the face mask image according to the square mask range; further obtain the face mask set data.
[0087] In one embodiment, the face images in the training set data are frontal face images.
[0088] In a specific embodiment, the output results are as Figure 7 shown, where the first row is the face mask image, the second row is the image generated by the generator, and the third row is the real image. This result is tested in the test set, and the face images are more random, conforming to the settings in real life. For the last image, the sunglasses on the face can be removed.
[0089] It should be understood that although Figure 1 the steps in the flowchart of Figure 1At least some of the steps may include multiple sub-steps or multiple stages, and these sub-steps or stages do not necessarily need to be executed and completed at the same moment, but can be executed at different moments. The execution order of these sub-steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least some of the sub-steps or stages of other steps or other steps.
[0090] In one embodiment, as Figure 8 shown, a device for de-masking face restoration is provided, including: a preprocessing module 802, a training data input module 804, a model training module 806, and a model application module 808, where:
[0091] The preprocessing module 802 is configured to obtain training set data of face images, preprocess the training set data to obtain corresponding face mask set data; in the face mask set data, the face and mouth of the face image are replaced by a square mask;
[0092] The training data input module 804 is configured to input the training set data and the face mask set data into a de-masking face restoration model; the de-masking face restoration model includes a first path, a second path, a feature fusion module, and an image output module; the first path includes a hole interpolation convolution module, a multi-layer dynamic selection convolution module, and a multi-layer hole convolution module; the multi-layer dynamic selection convolution module and the multi-layer hole convolution module form a U-shaped convolution network; the second path is a U-shaped convolution network, including a multi-layer dynamic selection convolution module and a context attention module; the hole interpolation convolution module is configured to fill holes in the face mask image in the face mask set data by filling in noise; the dynamic selection convolution modules in the first path and the second path are configured to obtain high-weight features through the softmax function; the hole convolution module is configured to perform feature extraction for expanding the receptive field; the context attention module is configured to borrow effective spatial pixels for hole filling; the feature fusion module is configured to perform feature fusion on the outputs of the first path and the second path; the image output module is configured to activate the generated image of the face according to the output of the feature fusion module;
[0093] The model training module 806 is configured to train the de-masking face restoration model according to the real image and the corresponding generated image in the training set data through a pre-set loss function to obtain a trained de-masking face restoration model;
[0094] The model application module 808 is configured to obtain a face mask image to be processed, input the face mask image into the trained de-masking face restoration model, and obtain a de-masked restored face image.
[0095] The preprocessing module 802 is further configured to obtain the training set data of the face images; the training set data is randomly collected from the publicly available dataset celeba; for each face image in the training set data, 68 feature points of the face are obtained through the trained dlib network, the square mask range is determined, and the face mask image is obtained according to the square mask range; further, the face mask set data is obtained.
[0096] The model training module 806 is further configured to train the face restoration model without mask through a preset loss function; the loss function of the generator in the face restoration model without mask includes L 1 loss function, L tv loss function and L content loss function; the objective function to be optimized by the face restoration model without mask is the WGAN loss.
[0097] For the specific limitations of the device for face restoration without mask, reference may be made to the limitations of the method for face restoration without mask in the foregoing text, which will not be elaborated herein. Each module in the above-mentioned device for face restoration without mask can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.
[0098] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 9 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for face restoration without mask. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0099] Those skilled in the art can understand, Figure 9The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0100] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in the above method embodiment are implemented.
[0101] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiment are implemented.
[0102] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0103] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0104] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for face restoration with mask removal, characterized in that, the method includes: Obtaining training set data of face images, preprocessing the training set data to obtain corresponding face mask set data; in the face mask set data, the face and mouth parts of the face images are replaced by square masks; Inputting the training set data and the face mask set data into a face restoration model with mask removal; the face restoration model with mask removal includes a first path, a second path, a feature fusion module, and an image output module; the first path includes an atrous interpolation convolution module, a multi-layer dynamic selection convolution module, and a multi-layer atrous convolution module; the multi-layer dynamic selection convolution module and the multi-layer atrous convolution module form a U-shaped convolution network; the second path is a U-shaped convolution network, including a multi-layer dynamic selection convolution module and a context attention module; the atrous interpolation convolution module is used to fill holes in the face mask images in the face mask set data by filling in noise; the dynamic selection convolution modules in the first path and the second path are used to extract high-weight features through the softmax function; the atrous convolution module is used to perform feature extraction with an enlarged receptive field; the context attention module is used to borrow effective spatial pixels for hole filling; the feature fusion module is used to fuse the features output by the first path and the second path; the image output module is used to activate the output of the feature fusion module to obtain a generated image of the face; Training the face restoration model with mask removal according to the real images and the corresponding generated images in the training set data through a pre-set loss function to obtain a trained face restoration model with mask removal; Obtaining a face mask image to be processed, inputting the face mask image into the trained face restoration model with mask removal to obtain a restored face image with mask removed.
2. The method according to claim 1, characterized in that, obtaining training set data of face images, preprocessing the training set data to obtain corresponding face mask set data, includes: Obtaining training set data of face images; the training set data is randomly collected from the public dataset celeba; For each face image in the training set data, obtaining 68 facial feature points of the face through the trained dlib network, determining the square mask range, and obtaining a face mask image according to the square mask range; Further obtaining the face mask set data.
3. The method according to claim 2, characterized in that, the mathematical representation corresponding to the dynamic selection convolution module is: Among them, Output is the output of the dynamic selection convolution module, representing the features after convolution, and σ(·) represents the weight information obtained by the softmax function.
4. The method according to claim 3, characterized in that, the atrous interpolation convolution module is used to fill holes in the face mask images in the face mask set data by filling in noise, includes: The atrous interpolation convolution module is used to add a noise filling module on the basis of the deformable convolution module, fuse the image features learned by the noise filling module and the deformable convolution module, and fill holes in the face mask images in the face mask set data.
5. The method according to claim 4, characterized in that, The processing flow of the noise filling module includes: Normalize the face mask images in the face mask set data channel by channel; Overlay noise on the normalized image; Perform 3×3 convolution on the image with overlaid noise; Normalize the convolved image channel by channel again to obtain the output of the noise filling module.
6. The method according to claim 5, wherein, training the face mask removal and face restoration model through a pre-set loss function, including: Training the face restoration model without mask through a pre-set loss function; the loss function of the generator in the face restoration model without mask includes L 1 loss function, L tv loss function and L content loss function; the objective function to be optimized by the face restoration model without mask is the WGAN loss.
7. The method according to any one of claims 1 to 6, wherein, the face images in the training set data are frontal face images.
8. A device for face mask removal and face restoration, wherein, the device includes: A preprocessing module, configured to obtain training set data of face images, preprocess the training set data to obtain corresponding face mask set data; in the face mask set data, the face and mouth parts of the face images are replaced by square masks; A training data input module, configured to input the training set data and the face mask set data into a face mask removal and face restoration model; the face mask removal and face restoration model includes a first path, a second path, a feature fusion module, and an image output module; the first path includes a hole interpolation convolution module, a multi-layer dynamic selection convolution module, and a multi-layer dilated convolution module; the multi-layer dynamic selection convolution module and the multi-layer dilated convolution module form a U-shaped convolution network; the second path is a U-shaped convolution network, including a multi-layer dynamic selection convolution module and a context attention module; the hole interpolation convolution module is configured to fill holes in the face mask images in the face mask set data by filling in noise; the dynamic selection convolution modules in the first path and the second path are configured to obtain high-weight features through the softmax function; the dilated convolution module is configured to perform feature extraction for expanding the receptive field; the context attention module is configured to borrow effective spatial pixels for hole filling; the feature fusion module is configured to perform feature fusion on the outputs of the first path and the second path; the image output module is configured to activate the output of the feature fusion module to obtain a generated image of the face; A model training module, configured to train the face mask removal and face restoration model through a pre-set loss function according to the real images and the corresponding generated images in the training set data to obtain a trained face mask removal and face restoration model; A model application module, configured to obtain a face mask image to be processed, input the face mask image into the trained face mask removal and face restoration model, and obtain a face mask-removed and restored face image.
9. A computer device, including a memory and a processor, the memory stores a computer program, wherein, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.