Face image processing method, live image processing method, device and electronic equipment
By identifying the region to be processed in a face image, performing deformation processing, and generating a hole region, and then using an image hole completion model to complete it, the problem of background distortion and deformation is solved, achieving a more natural deformation processing effect.
Patent Information
- Application Number
- CN202211486724.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-24
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2042-11-24
AI Technical Summary
Existing technologies, when performing deformation processing on facial images, cause other background objects in the image to be distorted and deformed, resulting in poor deformation processing effects.
By identifying the region to be processed in a face image, performing deformation processing while keeping other regions unchanged, a face image containing hole regions is generated. A mask image is then obtained and input into a trained image hole completion model for completion processing, outputting a natural and realistic deformation result.
It avoids distortion of the background area, improves the deformation processing effect of facial images, and makes the processing more realistic and natural.
Smart Images

Figure CN115908115B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence and live broadcast technology, in particular to a face image processing method, a live broadcast image processing method, device, electronic equipment and computer readable storage medium. BACKGROUND
[0002] With the rapid development of the field of artificial intelligence, the beautifying and reshaping algorithm based on face key point positioning is widely used in the field of live broadcast.
[0003] However, the scheme provided by the current technology for performing morphing processing on a face image directly stretches the whole image, causing other background objects in the image to be distorted, resulting in poor morphing processing effect on the face image. SUMMARY
[0004] Therefore, it is necessary to provide a face image processing method, a live broadcast image processing method, device, electronic equipment and computer readable storage medium in view of the above technical problems.
[0005] In a first aspect, the present application provides a face image processing method. The method comprises:
[0006] determining a face region to be processed in a face image;
[0007] performing morphing processing on the face region to be processed and keeping other regions in the face image unchanged to obtain a face image containing a hollow region; the other regions are image regions in the face image except the face region to be processed;
[0008] According to the face image containing the hollow region, a mask image corresponding to the hollow region is obtained;
[0009] inputting the face image containing the hollow region and the mask image into a trained image hole completion model, and outputting a face image after hole completion by the image hole completion model according to the face image containing the hollow region and the mask image;
[0010] obtaining the face image after hole completion output by the image hole completion model.
[0011] In a second aspect, the present application provides a live broadcast image processing method. The method comprises:
[0012] obtaining a live broadcast image;
[0013] in response to a face processing instruction for the live broadcast image, processing the live broadcast image according to the face image processing method as described above to obtain a live broadcast image after hole completion;
[0014] display the live image after the hole completion.
[0015] In a third aspect, the present application provides a face image processing apparatus. The apparatus comprises:
[0016] a region determining module configured to determine a face region to be processed in a face image;
[0017] a deformation processing module configured to perform deformation processing on the face region to be processed and keep other regions in the face image unchanged, to obtain a face image containing a hole region; the other regions are image regions in the face image except the face region to be processed;
[0018] a mask obtaining module configured to obtain a mask image corresponding to the hole region according to the face image containing the hole region;
[0019] a model processing module configured to input the face image containing the hole region and the mask image into a trained image hole completion model, and output a face image after hole completion by the image hole completion model according to the face image containing the hole region and the mask image;
[0020] an image obtaining module configured to obtain the face image after hole completion output by the image hole completion model.
[0021] In a fourth aspect, the present application provides a live image processing apparatus. The apparatus comprises:
[0022] an image obtaining module configured to obtain a live image;
[0023] an image processing module configured to perform processing on the live image by using the face image processing apparatus as described above in response to a face processing instruction for the live image, to obtain a live image after hole completion;
[0024] an image displaying module configured to display the live image after hole completion.
[0025] In a fifth aspect, the present application provides an electronic device. The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the following steps when executing the computer program:
[0026] determining a to-be-processed face region in a face image; performing morphing processing on the to-be-processed face region and keeping other regions in the face image unchanged to obtain a face image containing a hole region; the other regions are image regions in the face image except the to-be-processed face region; obtaining a mask image corresponding to the hole region according to the face image containing the hole region; inputting the face image containing the hole region and the mask image into a trained image hole completion model, and outputting a face image after hole completion by the image hole completion model according to the face image containing the hole region and the mask image; and obtaining the face image after hole completion output by the image hole completion model.
[0027] In a sixth aspect, the present application provides an electronic device. The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the following steps when executing the computer program:
[0028] obtaining a live image; in response to a face processing instruction for the live image, processing the live image according to the face image processing method as described above to obtain a live image after hole completion; and displaying the live image after hole completion.
[0029] In a seventh aspect, the present application provides a computer readable storage medium. The computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the following steps:
[0030] determining a to-be-processed face region in a face image; performing morphing processing on the to-be-processed face region and keeping other regions in the face image unchanged to obtain a face image containing a hole region; the other regions are image regions in the face image except the to-be-processed face region; obtaining a mask image corresponding to the hole region according to the face image containing the hole region; inputting the face image containing the hole region and the mask image into a trained image hole completion model, and outputting a face image after hole completion by the image hole completion model according to the face image containing the hole region and the mask image; and obtaining the face image after hole completion output by the image hole completion model.
[0031] In an eighth aspect, the present application provides a computer readable storage medium. The computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the following steps:
[0032] obtaining a live image; in response to a face processing instruction for the live image, processing the live image according to the face image processing method as described above to obtain a live image after hole completion; and displaying the live image after hole completion.
[0033] The face image processing method, the live image processing method, the device, the equipment and the storage medium determine a face region to be processed in a face image, perform morphing processing on the face region to be processed and keep other regions in the face image unchanged, obtain a face image containing a hollow region, the other regions are image regions in the face image except the face region to be processed, obtain a mask image corresponding to the hollow region according to the face image containing the hollow region, input the face image containing the hollow region and the mask image into a trained image hole completion model, output a face image after hole completion by the image hole completion model according to the face image containing the hollow region and the mask image, and obtain the face image after hole completion output by the image hole completion model. In this scheme, when the face region to be processed in the face image is morphed, the other regions are kept unchanged, so that the face image containing the hollow region is obtained, and then the face image containing the hollow region and the mask image corresponding to the hollow region are input into the trained image hole completion model, so that the face image after hole completion output by the image hole completion model is obtained, thereby avoiding distortion of the regions outside the face region in the image and filling the generated hollow region by using the trained image hole completion model, so that the morphing processing of the face image is more real and natural, and the morphing processing effect of the face image is improved. BRIEF DESCRIPTION OF DRAWINGS
[0034] Figure 1 The figure is a diagram of the application environment of the related method in the embodiment of the present application;
[0035] Figure 2 The figure is a flowchart of the face image processing method in the embodiment of the present application;
[0036] Figure 3 The figure is a flowchart of the face image processing method in the embodiment of the present application;
[0037] Figure 4 The figure is a flowchart of the step of training the image hole completion model in the embodiment of the present application;
[0038] Figure 5 The figure is a flowchart of the step of training the image hole completion model in the embodiment of the present application;
[0039] Figure 6 The figure is a structural diagram of the image hole completion model in the embodiment of the present application;
[0040] Figure 7 The figure is a structural diagram of the discriminator in the embodiment of the present application;
[0041] Figure 8 The figure is a flowchart of the live image processing method in the embodiment of the present application;
[0042] Figure 9 Fig. 1 is a structural diagram of a face image processing device in an embodiment of the present application;
[0043] Figure 10 Fig. 2 is a structural diagram of a live image processing device in an embodiment of the present application;
[0044] Figure 11 Fig. 3 is an internal structural diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0046] The face image processing method and the live image processing method provided by the present application can be applied in an application environment as shown in Figure 1 Fig. 1, which can include a terminal and a server. The terminal can communicate with the server through the Internet. The terminal can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers, and the server can be implemented by an independent server or a server cluster composed of multiple servers. Specifically, the face image processing method and the live image processing method provided by the embodiments of the present application can be executed by the terminal or the server, or part of the steps can be executed by the terminal and part of the steps can be executed by the server. When executed by the terminal, the terminal can obtain a face image captured and process the face image by executing the face image processing method provided by the present application. When executed by the server, the terminal can transmit the captured face image to the server, the server can process the face image by executing the face image processing method provided by the present application after receiving the face image, and then return the face image completed by the hole to the terminal. In the case of partially executed by the terminal and partially executed by the server, the server can execute the related steps of training the image hole completion model to be trained, then send the trained image hole completion model to the terminal, and then the terminal can apply the trained image hole completion model to process the captured face image, etc.
[0047] The face image processing method and the live image processing method of the present application are described in the following parts in combination with the embodiments and the corresponding drawings.
[0048] In one embodiment, as shown in Figure 2 Fig. 1, a face image processing method is provided, which can be executed by the terminal as shown in Figure 1 Fig. 1. The method can include the following steps:
[0049] Step S201: determining a to-be-processed face region in the face image.
[0050] In this step, the terminal can capture an image containing a face of a user such as a host through a camera, and the image is referred to as a face image. Then, the terminal determines a to-be-processed face region in the face image. The user such as the host can input a processing instruction for a face in the face image to the terminal, to instruct the terminal to perform morphing processing on the face in the face image. The morphing processing can be face slimming, face shaping, or the like. After receiving the instruction, the terminal can determine the to-be-processed face region in the face image. The to-be-processed face region refers to an image region in the face image where the face is located.
[0051] In one embodiment, step S201 specifically includes: determining each face key point corresponding to a cheek part in the face image; and determining the to-be-processed face region in the face image according to positions of the each face key point corresponding to the cheek part in the face image.
[0052] In this embodiment, with reference to Figure 3 , after obtaining the face image, the terminal can first detect each face key point in the face image. Specifically, the terminal can locate each face key point in the face image, that is, obtain position coordinates of each face key point in the face image. The face key points can include face key points of parts such as eyebrows, eyes, nose, lips, and cheeks. Then, the terminal can determine each face key point corresponding to a cheek part from the face key points, to obtain position coordinates of each face key point corresponding to the cheek part in the face image. Then, the terminal can obtain a face region formed by a line connecting cheek contour points according to the position coordinates of each face key point corresponding to the cheek part in the face image, and determine the face region as the to-be-processed face region. In this way, the morphing processing range of the face image can be accurately determined, to avoid affecting image regions other than the face.
[0053] Step S202: performing morphing processing on the to-be-processed face region and keeping other regions in the face image unchanged, to obtain a face image containing a hollow region.
[0054] In this step, in combination with Figure 3, the terminal can perform deformation processing such as face slimming, reshaping, etc. on the determined face region to be processed after receiving the processing instruction of the face in the face image input by the user. Different from the conventional deformation processing scheme, in this step, only the face region to be processed is deformed when the face region to be processed is deformed, and specifically, only the face region connected by the cheek contour points can be deformed, while other regions in the face image (i.e. the image region in the face image except the face region to be processed) remain unchanged, so that the background objects in the image will not be distorted when the face region to be processed is deformed. However, since only the face region to be processed is deformed in this step, there will be a certain hollow region between the face region to be processed after deformation and other regions, such as Figure 3 In the face image after deformation processing, the hollow region (contentless region) between the solid line and the dashed line in the cheek contour part, wherein the dashed line corresponds to the face region before deformation processing, and the solid line corresponds to the face region after deformation processing. Therefore, the face image after deformation processing is recorded as a face image containing a hollow region, and in the subsequent step, the face image containing a hollow region will be processed by image hole completion processing, which is to complete the hollow region in the image, or to complete the contentless region in the image, so that the face image after hole completion is real and natural.
[0055] Step S203, according to the face image containing a hollow region, a mask image corresponding to the hollow region is obtained.
[0056] Step S204, the face image containing a hollow region and the mask image are input into the trained image hole completion model, and the image hole completion model outputs the face image after hole completion according to the face image containing a hollow region and the mask image.
[0057] Step S205, the face image after hole completion output by the image hole completion model is obtained.
[0058] Reference Figure 3The steps S203 to S205 are related processes of using the trained image hole completion model to perform image hole completion processing on the face image containing the hole region after obtaining the face image containing the hole region through deformation processing, so as to obtain the face image after hole completion. Specifically, after obtaining the face image containing the hole region, first, in step S203, a mask image corresponding to the hole region is obtained according to the face image containing the hole region. The mask image is a single-channel image consistent with the image size of the face image, and the pixel value in the mask image is binary 0 or 1, where 0 represents the hole region in the face image and 1 represents the non-hole region in the face image. Thus, the terminal can generate a mask image corresponding to the hole region according to the face image containing the hole region. Then, in step S204, the face image containing the hole region (containing R, G and B channels) and the mask image are merged in the channel dimension to obtain an input image of four channels (R, G, B and mask). The input image is input into the trained image hole completion model, which is used to perform hole completion processing on the face image containing the hole region. Specifically, a convolutional neural network can be used. The trained image hole completion model outputs a face image after hole completion according to the face image containing the hole region and the mask image. Finally, in step S205, the face image after hole completion output by the image hole completion model is obtained, thereby realizing hole completion processing on the face image containing the hole region. In specific implementation, for the image hole completion model, the original face image sample, the corresponding face image sample containing the hole region and the mask image sample corresponding to the hole region can be used for training. For example, the corresponding face image sample containing the hole region and the mask image sample corresponding to the hole region can be used as training input data, the original face image sample can be used as training label data, and a corresponding loss function can be constructed to train the image hole completion model, so that the trained image hole completion model has natural and realistic hole completion effect on the face image sample containing the hole region.
[0059] The face image processing method of the embodiment determines a face region to be processed in the face image, performs morphing processing on the face region to be processed and keeps other regions in the face image unchanged to obtain a face image containing a hollow region, the other regions being image regions in the face image except the face region to be processed, acquires a mask image corresponding to the hollow region according to the face image containing the hollow region, inputs the face image containing the hollow region and the mask image into a trained image hole completion model, outputs a face image after hole completion by the image hole completion model according to the face image containing the hollow region and the mask image, and acquires the face image after hole completion output by the image hole completion model. The scheme keeps the other regions unchanged when performing morphing processing on the face region to be processed in the face image, thereby obtaining the face image containing the hollow region, then inputs the face image containing the hollow region and the mask image corresponding to the hollow region into the trained image hole completion model, obtains the face image after hole completion output by the image hole completion model, thereby avoiding distortion of the regions outside the face region in the image and filling the generated hollow region by using the trained image hole completion model, making the morphing processing on the face image more real and natural and improving the morphing processing effect on the face image.
[0060] In one embodiment, as shown in FIG. 1, Figure 4 The face image processing method of the present application can further include the following steps:
[0061] In step S401, a face image sample is acquired, and a mask image sample corresponding to a preset hollow region is acquired.
[0062] In step S402, a face image sample containing a preset hollow region is obtained according to the face image sample and the mask image sample.
[0063] In the steps S401 and S402, the server can construct a model training database. Specifically, the server can collect a batch of face images to obtain a batch of face image samples, and can also pre-generate a batch of mask images corresponding to various shapes of the hole regions to obtain a batch of mask image samples (i.e., mask image samples corresponding to the preset hole regions). The mask image sample is a single-channel image with the same size as the face image sample, and the pixel value is binary 0 or 1. 0 represents a hole region, and 1 represents a non-hole region. Then, a face image sample and a mask image sample corresponding to the preset hole region can be obtained from the face image sample, and the face image sample and the mask image sample are multiplied pixel by pixel to obtain a face image sample containing a preset hole region. In this way, the face image sample, the mask image sample corresponding to the preset hole region, and the face image sample containing the preset hole region can be associated and stored in the model training database for training of the image hole completion model. In a specific implementation, the face image sample containing the preset hole region (having R, G, and B channels) and the mask image sample corresponding to the preset hole region (mask channel) can be first merged in the channel dimension to obtain a model training input image (having R, G, B, and mask channels), and then the model training input image and the face image sample can be stored in the model training database as a sample pair.
[0064] In step S403, the image hole completion model to be trained is trained according to the face image sample, the mask image sample, and the face image sample containing the preset hole region, to obtain a trained image hole completion model.
[0065] In this step, the model training input image obtained by merging the face image sample containing the preset hole region and the mask image sample in the channel dimension can be used as the input data for model training, the face image sample can be used as the label data for model training, and the corresponding loss function can be constructed to train the image hole completion model to be trained, so that the trained image hole completion model has a natural and realistic hole completion effect on the face image sample containing the hole region.
[0066] In one embodiment, as shown in Figure 5 In step S403, the image hole completion model to be trained is trained according to the face image sample, the mask image sample, and the face image sample containing the preset hole region, to obtain a trained image hole completion model.
[0067] In step S501, the face image sample containing the preset hole region and the mask image sample are input into the image hole completion model to be trained, and a face image after hole completion predicted by the image hole completion model to be trained according to the face image sample containing the preset hole region and the mask image sample is obtained.
[0068] Specifically, as Figure 6 The structure of the image hole completion model to be trained is shown, where Convolution represents convolution, Max Pooling represents maximum pooling, Up Sampling represents up sampling, Skip connection represents residual connection, and Block copied represents block copy. As described above, the face image sample (with R, G, and B channels) containing the preset hole region can be combined with the mask image sample (mask channel) corresponding to the preset hole region in the channel dimension to obtain an input image (with R, G, B, and mask channels) for model training, and then the input image is input into the image hole completion model to be trained. The image hole completion model to be trained directly outputs the predicted face image after hole completion (with R, G, and B channels) according to the input image, and the face image after hole completion predicted by the image hole completion model to be trained is obtained.
[0069] In step S502, a first model loss is obtained according to the mask image sample, the face image sample, and the predicted face image after hole completion.
[0070] In this step, the first model loss is used to constrain the consistency of the image region that does not need to be completed in the predicted face image after hole completion and the face image sample, where the image region that does not need to be completed can correspond to other regions in the image except the hole region. Through the first model loss, it is constrained that the content of the image region that does not need to be completed should be consistent with the original image.
[0071] As an embodiment, step S502 can include: performing difference processing on the face image sample and the predicted face image after hole completion to obtain image difference information of the face image sample and the predicted face image after hole completion; and performing multiplication processing on the mask image sample and the image difference information to obtain the first model loss.
[0072] Specifically, the face image sample is denoted as gt, the predicted face image after hole completion is denoted as pred, and the mask image sample is denoted as mask, which corresponds to the area that needs to be completed (the pixel value is binary 0 or 1, 0 represents a hole area, and 1 represents a non-hole area). Thus, the face image sample gt and the predicted face image pred after hole completion are processed by difference, and the image difference information of the face image sample and the predicted face image after hole completion is |gt-pred|. Thus, the first model loss L1 can be represented as: L1=M⊙|gt-pred|; wherein ⊙ represents pixel-by-pixel multiplication between images, and the first model loss L1 obtained in this way can constrain the consistency of the image area that does not need to be completed in the predicted face image after hole completion and the face image sample.
[0073] In step S503, the face image sample and the predicted face image after hole completion are respectively input into the pre-trained feature extraction network, and first feature map sequences and second feature map sequences obtained by the pre-trained feature extraction network for the face image sample and the predicted face image after hole completion are obtained.
[0074] In step S504, the second model loss is obtained according to the first feature map sequences and the second feature map sequences obtained by the pre-trained feature extraction network for the face image sample and the predicted face image after hole completion.
[0075] Steps S503 to S504 are related processes for obtaining the second model loss, and the second model loss is used to constrain the semantic consistency of the predicted face image after hole completion and the face image sample, so that the completion effect is real and natural. Specifically, the face image sample gt and the predicted face image pred after hole completion are respectively input into the pre-trained feature extraction network, and the pre-trained feature extraction network can use a pre-trained model publicly disclosed in the industry, such as a VGG network. The pre-trained feature extraction network has multiple layers, wherein the feature map of the ith layer obtained by inputting the face image sample gt into the pre-trained feature extraction network is denoted as The feature map of the ith layer obtained by inputting the predicted face image pred after hole completion into the pre-trained feature extraction network is denoted as It is recorded that the pre-trained feature extraction network has N layers, so that N feature maps can be obtained for the face image sample and the predicted face image after hole completion, respectively denoted as a first feature map sequence and a second feature map sequence. The first feature map sequence contains N feature maps of the face image sample, and the second feature map sequence contains N feature maps of the predicted face image after hole completion. Then, the second model loss is obtained according to the first feature map sequence and the second feature map sequence.
[0076] As an embodiment, the step S504 can include: performing difference processing on the first feature map and the second feature map corresponding to each of the layers in the first feature map sequence and the second feature map sequence, to obtain feature map difference information corresponding to each of the layers respectively between the face image sample and the predicted face image after hole completion; and obtaining the second model loss according to the feature map difference information corresponding to each of the layers respectively between the face image sample and the predicted face image after hole completion.
[0077] Specifically, taking the i-th layer as an example, the first feature map in the first feature map sequence is and the second feature map in the second feature map sequence is performing difference processing on the first feature map and the second feature map to obtain the feature map difference information corresponding to the layer Thus, the feature map difference information corresponding to each of the layers respectively can be obtained, and the second model loss can be obtained according to the feature map difference information corresponding to each of the layers respectively. Exemplarily, the second model loss L2 can be represented as: Thus, the second model loss can be calculated in combination with the first feature map to the N-th feature map, which is used to guide the semantic consistency between the predicted face image after hole completion and the original face image sample, so that the completion effect is real and natural.
[0078] In step S505, the image hole completion model to be trained is trained according to the first model loss and the second model loss, to obtain a trained image hole completion model.
[0079] Specifically, in this step, the total loss of the image hole completion model to be trained can be obtained according to the first model loss L1 and the second model loss L2. For example, L1+L2 can be taken as the total loss L of the image hole completion model to be trained. That is, the total loss L can include two parts, one part is used to constrain the consistency between the predicted face image after hole completion and the face image sample in the image region that does not need to be completed, and the other part is used to constrain the semantic consistency between the predicted face image after hole completion and the face image sample. Then, the network parameters of the image hole completion model to be trained can be updated based on the total loss L, to train the image hole completion model to be trained. As an embodiment, when the total loss L is less than or equal to a preset loss threshold, it is judged that the training of the image hole completion model is completed, and a trained image hole completion model is obtained.
[0080] The scheme of the embodiment can train an image hole completion model that can not only complete the hole region in the face image in a real and natural manner, but also keep the region that does not need to be completed consistent with the original face image, thereby improving the deformation processing effect on the face image.
[0081] In an embodiment, the method of the present application can also obtain the trained image hole completion model through generative adversarial training, and for this, the method of the present application can further include the following steps: inputting the predicted hole-completed human face image into a discriminator to obtain a discrimination result corresponding to the predicted hole-completed human face image output by the discriminator; and obtaining a third model loss according to the discrimination result corresponding to the predicted hole-completed human face image. And the above step S505 further includes: training the image hole completion model to be trained according to the first model loss, the second model loss and the third model loss.
[0082] In the present embodiment, the image hole completion model to be trained is taken as a generator G, and the discriminator is represented as D. As an example, the structure of the discriminator D is as shown in Figure 7 wherein stride represents the sampling interval during convolution, BatchNorm represents batch normalization, and LeakyReLU represents a leaky rectified linear unit. For the discriminator D, an image is input into the discriminator D, and the discriminator D outputs a single-channel feature map, each value on the feature map representing a prediction result of a corresponding original image region. The single-channel feature map serves as a discrimination result corresponding to the image output by the discriminator D. In the present embodiment, the discriminator is trained according to the human face image sample and the predicted hole-completed human face image, so that in an ideal state, the values on the feature map output by the discriminator D should all be 1 (representing a real image) when the human face image sample is input into the discriminator D, otherwise 0 (such as an image generated by the model). In generative adversarial training, the loss of the generative adversarial network GANLoss includes two parts, respectively denoted as a first part loss DLoss of the discriminator D and a second part loss GLoss of the generator G, wherein the first part loss DLoss can be calculated according to the prediction results of the discriminator D on the human face image sample and the predicted hole-completed human face image, respectively, and the second part loss GLoss can be calculated according to the prediction result of the discriminator D on the predicted hole-completed human face image.
[0083] As an example, the above first part loss DLoss can be represented as: DLoss = (D(gt) - 1) 2 + (D(pred) - 0) 2 wherein D(gt) represents the prediction result corresponding to the human face image sample output by the discriminator D, and D(pred) represents the prediction result of the predicted hole-completed human face image output by the discriminator D; and the above second part loss GLoss can be represented as: GLoss = (D(pred) - 1) 2 .
[0084] Thus, the overall process of obtaining the trained image hole completion model in combination with the generative adversarial training in this embodiment can include: first inputting the face image sample and the predicted hole-completed face image into the discriminator D, obtaining the prediction result D(gt) corresponding to the face image sample and the prediction result D(pred) corresponding to the predicted hole-completed face image output by the discriminator D, and then obtaining the first part loss DLoss as shown above and training the discriminator D according to the first part loss DLoss. The first part loss DLoss is used to enable the discriminator D to distinguish between real images and model-generated images. Then, the second part loss GLoss as shown above can be obtained according to the discrimination result D(pred) corresponding to the predicted hole-completed face image, and the second part loss GLoss is used to train the image hole completion model so that the predicted image is as close to the real image as possible, so that the discriminator D cannot distinguish between the two. In this embodiment, unlike the traditional generative adversarial training, in addition to the second part loss GLoss, the first model loss L1 and the second model loss L2 described in the foregoing embodiments are combined to train the generator G, i.e., the image hole completion model, and the second part loss GLoss is denoted as the third model loss L3. Thus, the image hole completion model to be trained is trained according to the first model loss L1, the second model loss L2 and the third model loss L3 in this embodiment, and the trained image hole completion model obtained by training has a more realistic and natural hole completion effect.
[0085] In one embodiment, as shown in FIG. 8, a live image processing method is also provided, which can be performed by a terminal as shown in FIG. 1, and the method can include the following steps: Figure 8 Figure 1
[0086] Step S801, obtaining a live image.
[0087] Step S802, in response to a face processing instruction for the live image, processing the live image according to the face image processing method as described above to obtain a hole-completed live image.
[0088] Step S803, displaying the hole-completed live image.
[0089] In this embodiment, in a live broadcast scenario, a terminal of a host obtains a live broadcast image of the host. The host can input a face processing instruction for the live broadcast image into the terminal to instruct the terminal to perform face slimming or the like on the live broadcast image. The terminal can perform morphing processing on the live broadcast image according to the face image processing method of the present application as described in each of the above embodiments. Specifically, the terminal can first locate a cheek contour in the live broadcast image, and then only affect an image region within the cheek contour when performing morphing processing such as face slimming. After the morphing processing, a hollow region between the cheek contour and other regions in the live broadcast image is completed by a trained image hole completion model. Finally, a hole-completed live broadcast image with a real and natural morphing processing effect is obtained. Then, the terminal of the host can display the hole-completed live broadcast image, and can also send the hole-completed live broadcast image to terminals of each viewer through a server for display.
[0090] The live broadcast image processing method of this embodiment can apply the face image processing method of the present application to a live broadcast scenario. When applying a beauty shaping special effect to a host face in a live broadcast image, the method can avoid distorting background objects in regions other than the host face region by using image hole completion, so that the effect of the beauty shaping special effect is more real and natural, and the live broadcast effect and user experience are improved.
[0091] It should be understood that, although each step in the flowchart involved in each of the above embodiments is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in each of the above embodiments can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or steps or stages in other steps.
[0092] Based on the same inventive concept, the present application also provides a related device for implementing the above-mentioned related method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more related device embodiments provided below can refer to the limitations of the related method described above, which will not be repeated here.
[0093] In one embodiment, as shown in FIG. 9, a face image processing device is provided. The device 900 can include: Figure 9
[0094] The region determination module 901 is configured to determine a to-be-processed face region in the face image.
[0095] The deformation processing module 902 is configured to perform deformation processing on the to-be-processed face region and keep other regions in the face image unchanged, to obtain a face image containing a hollow region.
[0096] The mask obtaining module 903 is configured to obtain a mask image corresponding to the hollow region according to the face image containing the hollow region.
[0097] The model processing module 904 is configured to input the face image containing the hollow region and the mask image into a trained image hole completion model, and output a face image after hole completion by the image hole completion model according to the face image containing the hollow region and the mask image.
[0098] The image obtaining module 905 is configured to obtain the face image after hole completion output by the image hole completion model.
[0099] In one embodiment, the region determination module 901 is configured to determine each face key point corresponding to a cheek part in the face image, and determine the to-be-processed face region in the face image according to the position of each face key point corresponding to the cheek part in the face image.
[0100] In one embodiment, the apparatus 900 can further include a model training module configured to obtain face image samples, obtain mask image samples corresponding to preset hollow regions, obtain face image samples containing the preset hollow regions according to the face image samples and the mask image samples, and train a to-be-trained image hole completion model according to the face image samples, the mask image samples, and the face image samples containing the preset hollow regions, to obtain the trained image hole completion model.
[0101] In one embodiment, the model training module is configured to input the face image sample containing the preset hole region and the mask image sample into an image hole completion model to be trained, obtain a face image after hole completion predicted by the image hole completion model to be trained according to the face image sample containing the preset hole region and the mask image sample, obtain a first model loss according to the mask image sample, the face image sample and the predicted face image after hole completion, and constrain consistency of an image region without completion in the predicted face image after hole completion and the face image sample by using the first model loss. The face image sample and the predicted face image after hole completion are input into a pre-trained feature extraction network respectively, and first feature map sequences and second feature map sequences obtained by the pre-trained feature extraction network for the face image sample and the predicted face image after hole completion respectively are obtained. A second model loss is obtained according to the first feature map sequences and the second feature map sequences obtained by the pre-trained feature extraction network for the face image sample and the predicted face image after hole completion respectively, and semantic consistency of the predicted face image after hole completion and the face image sample is constrained by using the second model loss. The image hole completion model to be trained is trained according to the first model loss and the second model loss, and a trained image hole completion model is obtained.
[0102] In one embodiment, the model training module is configured to obtain image difference information of the face image sample and the predicted face image after hole completion by performing difference processing on the face image sample and the predicted face image after hole completion, and obtain the first model loss by performing multiplication processing on the mask image sample and the image difference information.
[0103] In one embodiment, the model training module is configured to obtain feature map difference information of the face image sample and the predicted face image after hole completion corresponding to each layer in the first feature map sequence and the second feature map sequence by performing difference processing on the first feature map and the second feature map corresponding to each layer in the first feature map sequence and the second feature map sequence, and obtain the second model loss according to the feature map difference information of the face image sample and the predicted face image after hole completion corresponding to each layer.
[0104] In an embodiment, a model training module is configured to input the predicted hole-completed face image into a discriminator to obtain a discrimination result corresponding to the predicted hole-completed face image output by the discriminator; the discriminator is trained according to the face image sample and the predicted hole-completed face image; a third model loss is obtained according to the discrimination result corresponding to the predicted hole-completed face image; and the image hole completion model to be trained is trained according to the first model loss, the second model loss, and the third model loss.
[0105] In an embodiment, as shown in FIG. 10, a live image processing apparatus is provided, which can include: Figure 10
[0106] An image acquisition module 1001 is configured to acquire a live image.
[0107] An image processing module 1002 is configured to, in response to a face processing instruction for the live image, process the live image by using the face image processing apparatus as described above to obtain a hole-completed live image.
[0108] An image display module 1003 is configured to display the hole-completed live image.
[0109] The above modules can be all or partially implemented by software, hardware, and combinations thereof. The above modules can be embedded in or independent of a processor in an electronic device in a hardware form, or stored in a memory in the electronic device in a software form, so as to be called and executed by a processor to perform operations corresponding to the above modules.
[0110] In an embodiment, an electronic device is provided, which can be a terminal, and an internal structure diagram of the electronic device can be as shown in FIG. 11. Figure 11 As shown. The electronic device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The communication interface of the electronic device is used to communicate with external terminals in a wired or wireless manner. Wireless mode can be achieved through WIFI, mobile cellular network, NFC (near field communication) or other technologies. The computer program is executed by the processor to implement a face image processing method and a live image processing method. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the electronic device, or an external keyboard, touchpad or mouse, etc.
[0111] Those skilled in the art can understand that, Figure 11 The skilled in the art can understand that,
[0112] In one embodiment, an electronic device is also provided, including a memory and a processor, the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0113] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the steps in the above method embodiments.
[0114] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. The non-volatile memory can include read-only memory (Read-Only Memory, ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (Magnetoresistive Random Access Memory, MRAM), ferroelectric memory (Ferroelectric Random Access Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. The volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0115] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are information and data authorized by the user or authorized by all parties.
[0116] The technical features of the above embodiments can be combined arbitrarily. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, but as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.
[0117] The above-described embodiments are merely illustrative of several embodiments of the present application, and the description is relatively specific and detailed, but should not be understood as a limitation on the scope of the patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.
Claims
1. A face image processing method, characterized by, The method comprises: determining a to-be-processed face region in a face image; performing morphing processing on the to-be-processed face region and keeping other regions in the face image unchanged to obtain a face image containing a hollow region; the other regions are image regions in the face image except the to-be-processed face region; obtaining a mask image corresponding to the hollow region according to the face image containing the hollow region; inputting the face image containing the hollow region and the mask image into a trained image hole completion model, and outputting a face image after hole completion by the image hole completion model according to the face image containing the hollow region and the mask image; wherein the image hole completion model training method comprises: obtaining a face image sample, and obtaining a mask image sample corresponding to a preset hollow region; obtaining a face image sample containing the preset hollow region according to the face image sample and the mask image sample; training a to-be-trained image hole completion model according to the face image sample, the mask image sample and the face image sample containing the preset hollow region to obtain the trained image hole completion model; obtaining the face image after hole completion output by the image hole completion model.
2. The method of claim 1, wherein, The determination of the to-be-processed face region in the face image comprises: determining each face key point corresponding to a cheek part in the face image; determining the to-be-processed face region in the face image according to the positions of each face key point corresponding to the cheek part in the face image.
3. The method of claim 1, wherein, The training of the to-be-trained image hole completion model according to the face image sample, the mask image sample and the face image sample containing the preset hollow region to obtain the trained image hole completion model comprises: inputting the face image sample containing the preset hollow region and the mask image sample into the to-be-trained image hole completion model, and obtaining a face image after hole completion predicted by the to-be-trained image hole completion model according to the face image sample containing the preset hollow region and the mask image sample; obtaining a first model loss according to the mask image sample, the face image sample and the predicted face image after hole completion; the first model loss is used to constrain the consistency of an image region without completion in the predicted face image after hole completion and the face image sample; inputting the face image sample and the predicted face image after hole completion into a pre-trained feature extraction network respectively, and obtaining first feature map sequences and second feature map sequences obtained by each layer of the pre-trained feature extraction network respectively for the face image sample and the predicted face image after hole completion; obtaining a second model loss according to the first feature map sequences and the second feature map sequences obtained by each layer of the pre-trained feature extraction network respectively for the face image sample and the predicted face image after hole completion; the second model loss is used to constrain the semantic consistency of the predicted face image after hole completion and the face image sample. The image hole completion model to be trained is trained according to the first model loss and the second model loss, to obtain the trained image hole completion model.
4. The method of claim 3, wherein, The first model loss is obtained according to the mask image sample, the face image sample, and the predicted face image after hole completion. The face image sample and the predicted face image after hole completion are subtracted to obtain image difference information of the face image sample and the predicted face image after hole completion. The mask image sample and the image difference information are multiplied to obtain the first model loss.
5. The method of claim 3, wherein, The second model loss is obtained according to the first feature map sequence and the second feature map sequence obtained by each layer of the pre-trained feature extraction network respectively for the face image sample and the predicted face image after hole completion. The first feature map and the second feature map corresponding to each layer of the layers in the first feature map sequence and the second feature map sequence are subtracted to obtain feature map difference information of the face image sample and the predicted face image after hole completion corresponding to each layer respectively. The second model loss is obtained according to the feature map difference information of the face image sample and the predicted face image after hole completion corresponding to each layer respectively.
6. The method of any one of claims 3 to 5, further comprising: inputting the predicted face image after hole completion into a discriminator to obtain a discrimination result corresponding to the predicted face image after hole completion output by the discriminator, the discriminator being trained according to the face image sample and the predicted face image after hole completion; obtaining a third model loss according to the discrimination result corresponding to the predicted face image after hole completion; training the image hole completion model to be trained according to the first model loss, the second model loss, and the third model loss. The method comprises: obtaining a live image; 7. A live image processing method, characterized by, in response to a face processing instruction for the live image, processing the live image according to the face image processing method of any one of claims 1 to 6 to obtain a live image after hole completion; displaying the live image after hole completion. The apparatus comprises: a region determination module configured to determine a face region to be processed in a face image; 8. A face image processing apparatus, characterized by comprising: a deformation processing module configured to perform deformation processing on the face region to be processed and keep other regions in the face image unchanged to obtain a face image containing a hole region, the other regions being image regions in the face image except the face region to be processed; a mask obtaining module configured to obtain a mask image corresponding to the hole region according to the face image containing the hole region. The model processing module is configured to input the face image containing the hole region and the mask image into a trained image hole completion model, and output a face image after hole completion by the image hole completion model according to the face image containing the hole region and the mask image; wherein the image hole completion model training method comprises: obtaining a face image sample, and obtaining a mask image sample corresponding to a preset hole region; obtaining a face image sample containing the preset hole region according to the face image sample and the mask image sample; training a to-be-trained image hole completion model according to the face image sample, the mask image sample, and the face image sample containing the preset hole region, to obtain the trained image hole completion model; The image obtaining module is configured to obtain the face image after hole completion output by the image hole completion model.
9. A live image processing apparatus, characterized by comprising: The device comprises: The image obtaining module is configured to obtain the face image after hole completion output by the image hole completion model. The image processing module is configured to perform processing on the live image by using the face image processing device according to the face processing instruction of the live image, to obtain a live image after hole completion. The image display module is configured to display the live image after hole completion.
10. An electronic device comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor executes the computer program to implement the steps of the method in any one of claims 1 to 6 or the method in claim 7.
11. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 6 or the method in claim 7.
Citation Information
Patent Citations
Image processing method and device, model training method and device, body shaping processing method and device and storage medium
CN115082384A