Face image processing method and device, readable storage medium and program product
By automatically filling in facial skin content and locating the target image area, the problem of low efficiency in traditional facial image processing is solved, and natural beard images can be efficiently generated.
Patent Information
- Application Number
- CN202510750911.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-26
Smart Images

Figure CN120707696A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a facial image processing method, device, readable storage medium, and program product. Background Art
[0002] With the development of computer technology and the advancement of image processing techniques, the demand for facial image processing has become increasingly diverse. In some scenarios, there is a need to edit the facial region of a facial image. For example, in applications such as entertainment, virtual makeup try-ons, character design, and film and television post-production, there is a need to generate beards within the facial region of a facial image. Currently, a common approach is to obtain a beard stock image and manually overlay it onto a localized region of the facial image.
[0003] However, traditional processing methods rely on manual processing experience and have the problem of low image processing efficiency. Summary of the Invention
[0004] Based on this, the present application provides a facial image processing method, device, readable storage medium and program product, which can improve image processing efficiency.
[0005] In one aspect, the present application provides a method for processing a facial image, comprising:
[0006] Acquire an initial face image and a reference beard image; the initial face image includes an initial beard image region, and the reference beard image includes a reference beard image region;
[0007] Based on the initial facial image, filling the initial beard image region with image content representing facial skin to obtain an adjusted facial image;
[0008] In the adjusted facial image, locating a target image region for generating a target beard image;
[0009] Image generation is performed based on the adjusted facial image and the reference beard image to edit at least a portion of the target image area according to the reference beard image area to generate a target facial image, wherein the target facial image includes image content of the edited target beard image.
[0010] In one aspect, the present application further provides a facial image processing device, comprising:
[0011] An acquisition module, configured to acquire an initial face image and a reference beard image; the initial face image includes an initial beard image region, and the reference beard image includes a reference beard image region;
[0012] a filling module, configured to fill, based on the initial facial image, the initial beard image region with image content representing facial skin, to obtain an adjusted facial image;
[0013] a positioning module, configured to locate a target image region for generating a target beard image in the adjusted face image;
[0014] A generation module is used to generate an image based on the adjusted facial image and the reference beard image, so as to edit at least a portion of the image area in the target image area according to the reference beard image area to generate a target facial image, wherein the target facial image includes the image content of the edited target beard image.
[0015] In one aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:
[0016] Acquire an initial face image and a reference beard image; the initial face image includes an initial beard image region, and the reference beard image includes a reference beard image region;
[0017] Based on the initial facial image, filling the initial beard image region with image content representing facial skin to obtain an adjusted facial image;
[0018] In the adjusted facial image, locating a target image region for generating a target beard image;
[0019] Image generation is performed based on the adjusted facial image and the reference beard image to edit at least a portion of the target image area according to the reference beard image area to generate a target facial image, wherein the target facial image includes image content of the edited target beard image.
[0020] In one aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:
[0021] Acquire an initial face image and a reference beard image; the initial face image includes an initial beard image region, and the reference beard image includes a reference beard image region;
[0022] Based on the initial facial image, filling the initial beard image region with image content representing facial skin to obtain an adjusted facial image;
[0023] In the adjusted facial image, locating a target image region for generating a target beard image;
[0024] Image generation is performed based on the adjusted facial image and the reference beard image to edit at least a portion of the target image area according to the reference beard image area to generate a target facial image, wherein the target facial image includes image content of the edited target beard image.
[0025] The above-mentioned facial image processing method, device, readable storage medium and program product fill the image content representing the facial skin in the initial beard image area based on the initial facial image. The adjusted facial image thus obtained, compared with the initial facial image, eliminates the original beard in the initial facial image and fills the image content representing the facial skin, which can avoid the original beard in the initial facial image from affecting the effect of the subsequent target beard image; then, the target image area is located in the adjusted facial image, and then the image is generated based on the adjusted facial image and the reference beard image, so as to edit the target image according to the reference beard image area. At least a portion of the image area in the image area can obtain a target beard image, thereby generating a target facial image, and the target facial image includes the image content of the target beard image. In this way, through the above series of steps, the initial beard image area can be automatically eliminated and filled, the target image area for generating the target beard image can be automatically located, and then at least a portion of the image area in the target image area can be automatically edited according to the reference beard image area to form the target beard image. No manual operation is required, and while ensuring accurate reference to the reference beard image to generate the target facial image including the image content in the target beard image, the image processing efficiency is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.
[0027] Figure 1 1 is a flow chart of a facial image processing method according to an embodiment;
[0028] Figure 2 A schematic diagram of a simplified image segmentation process in one embodiment;
[0029] Figure 3 A schematic diagram of a simplified image filling process in one embodiment;
[0030] Figure 4 A schematic diagram of the image generation process in one embodiment;
[0031] Figure 5A schematic diagram of a flow chart of facial image preprocessing steps in one embodiment;
[0032] Figure 6 1. A schematic diagram of a flow chart of facial image post-processing steps in one embodiment;
[0033] Figure 7 A schematic diagram of a process flow of area positioning steps in one embodiment;
[0034] Figure 8 A schematic diagram of a simplified flow chart of facial image processing steps in one embodiment;
[0035] Figure 9 is a structural block diagram of a face image processing device in one embodiment;
[0036] Figure 10 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical solutions and beneficial effects of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0038] In an exemplary embodiment, Figure 1 As shown, a facial image processing method is provided. This embodiment is described by taking the method applied to a computer device as an example. The computer device can be a terminal or a server. It is understandable that the method can also be applied to a system including a terminal and a server and implemented through the interaction between the terminal and the server. The terminal can be a personal computer, a laptop, a smart phone, a tablet computer or other. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. In this embodiment, the method includes the following steps 102 to 108:
[0039] Step 102 , obtaining an initial face image and a reference beard image; the initial face image includes an initial beard image region, and the reference beard image includes a reference beard image region.
[0040] The initial facial image is the facial image to which a beard is to be added. The initial facial image can be an input facial image or a preprocessed input facial image. The input facial image can be an original facial image uploaded by a user. The initial beard image region is the image region in the initial facial image where the beard is located.
[0041] The reference beard image provides a reference beard style for the target beard image to be generated. The reference beard image can be uploaded directly by the user or a preset beard image that matches the user's selected beard type. The computer device may pre-store multiple pre-configured beard types and preset beard images that match each beard type, and the user can select a beard type from these multiple pre-set beard types. The multiple beard types may include, for example, a goatee, sideburns, mustache, or other types. The reference beard image area is the image region in the reference beard image where the beard is located.
[0042] Exemplarily, a computer device may obtain an input facial image, preprocess the input facial image to obtain an initial facial image, and the preprocessing step may include: performing facial key point detection on the input facial image to obtain a first facial key point corresponding to the input facial image, and adjusting the image view of the input facial image based on the first facial key point.
[0043] The input facial image may be a multi-channel image, for example, an image in RGB format. The input facial image may contain a face and a complex background. The face in the input facial image may be at multiple angles, for example, the face may be at a side view, an oblique view, or other angles. Therefore, it is necessary to pre-process the input facial image to improve the efficiency of subsequent image processing steps. It is understandable that the input facial image may also be a standard facial image, for example, a solid color background, and a view when the face is at eye level. In this case, the input facial image may also be directly used as the initial facial image.
[0044] The first facial key points are the facial key points in the input facial image. Facial key points are used to locate key facial features, such as the facial contour, eyes, nose, mouth, ears, or other key features. Facial key point detection can be achieved using a face detection model. For example, face detection models can use MTCNN (Multi-task Cascaded Convolutional Networks, a cascaded framework that combines face detection and key point localization), DCNN (Deep Convolutional Network, a model proposed by Yi Sun et al. that uses cascaded CNNs to achieve facial key point localization), DAN (Deep Alignment Network, a cascaded alignment model based on heatmap feedback), or other models. The number of facial key points can be multiple, specifically 68, 98, 106, or other.
[0045] Image view adjustment is a process that adjusts at least one of the image's viewing angle and field of view. Adjusting the image's viewing angle can be achieved through an affine transformation, for example, adjusting a face in a facial image from an oblique viewing angle to a binocularly level viewing angle. Adjusting the image's field of view can be achieved through at least one of image cropping and image scaling.
[0046] In one embodiment, the computer device may perform image segmentation processing on the initial facial image to identify an initial beard image region in the initial facial image.
[0047] Image segmentation can be achieved through image segmentation networks, such as U-Net (a network consisting of an encoder and a decoder with skip connections between the encoder and the decoder), FCN (Fully Convolutional Networks), SegNet (a deep learning architecture for image segmentation proposed by a research team at the University of Cambridge), Mask R-CNN (a deep learning algorithm for object detection and instance segmentation), or others.
[0048] When using an image segmentation network, an initial facial image can be input into the network to obtain an initial mask image as output. The initial mask image can have the same image size as the initial facial image. The initial mask image can be used to indicate the relative position of the initial beard image region relative to the initial facial image. Specifically, the initial mask image can be a binary image. The relative position of the first region in the initial mask image relative to the initial mask image is the same as the relative position of the initial beard image region relative to the initial facial image. Each pixel in the first region can be a non-zero pixel. For example, a non-zero pixel can represent white, and the pixel value of this pixel can be 1, 255, or other values. Each pixel in the area outside the first region of the initial mask image can be a zero pixel, and the pixel value of this pixel can be zero, which can represent black. In some scenarios, the mask can be written as "Mask" or understood as "mask". For example, the initial beard image region can be the chin region and the area around the upper lip. In the output initial mask image, the areas corresponding to the chin region and the area around the upper lip can be white, and the remaining areas can be black.
[0049] Take the U-Net segmentation network as an example, see Figure 2As shown in the schematic diagram of the image segmentation process, the U-Net segmentation network for image segmentation may include: inputting the initial face image into the U-Net segmentation network, first performing feature extraction and downsampling processing through the image encoder of the U-Net network to capture the texture, color and boundary information of the initial beard image area in the initial face image, and then performing upsampling and feature fusion processing through the image decoder to restore spatial details, and finally outputting an accurate initial mask image to clearly identify the position and shape outline of the initial beard image area in the initial face image.
[0050] Step 104 : Based on the initial facial image, fill the initial beard image region with image content representing facial skin to obtain an adjusted facial image.
[0051] Among them, filling the initial beard image area with image content representing facial skin can achieve the effect of removing the beard in the initial facial image and filling it with smooth facial skin. In the adjusted facial image, the area corresponding to the position of the initial beard image area can be obtained by filling based on the facial skin area surrounding the initial beard image area, retaining the facial skin texture and leaving no residual hair. For example, the initial beard image area can represent a beard, and the area corresponding to the position of the initial beard image area in the adjusted facial image can eliminate the beard and form smooth, beardless skin with skin texture, such as pores, wrinkles or other skin textures.
[0052] Exemplarily, the computer device may obtain an initial mask image obtained by performing image segmentation processing on the initial facial image, and perform image filling processing based on the initial facial image and the initial mask image to fill the image content representing the facial skin in the initial beard image area to obtain an adjusted facial image.
[0053] The initial mask image provides information about the location of the initial beard image region in the initial face image, which can guide the image filling model to accurately fill the image. Image filling processing can be implemented using an image filling model. Image filling models include the DeepFill model (proposed in the paper Free-Form Image Inpainting with Gated Convolution published at ICCV 2019, a free image restoration model based on gated convolution), the CR-Fill model (Contextual Reconstruction Fill, a deep learning-based generative image restoration model proposed by Zeng et al. at the 2021 ICCV conference), the LaMa model (Large Mask Inpainting, a deep learning-based image restoration model designed to address the restoration of large areas in images), or others.
[0054] Take the DeepFill model as an example, see Figure 3 As shown in the schematic diagram of the simplified image filling process, the initial face image and the initial mask image can be input into the DeepFill model for image filling processing to fill the image content representing the facial skin in the initial beard image area to obtain an adjusted face image.
[0055] The DeepFill model can be trained using a beard filling dataset. In this dataset, the labeled data can be a facial image without a beard, and the input data corresponding to each facial image without a beard can be a facial image with a beard after a beard is generated in the facial image, as well as a beard position mask obtained by performing image segmentation on the facial image with a beard. The beard position mask can represent the relative position of the beard region in the facial image with a beard. The facial image with a beard can also be an image in which the beard region can be generated in a facial image and then blacked out to simulate the presence of a beard. Training the DeepFill model using the beard filling dataset allows the DeepFill model to learn how to use the facial skin surrounding the beard region to fill the beard region, thereby eliminating and filling the beard region and forming image content representing the facial skin.
[0056] In one embodiment, after acquiring an initial facial image, a computer device may detect whether an initial beard image region exists in the initial facial image. If no initial beard image region exists, the initial facial image may be used as the adjusted facial image for subsequent steps. If an initial beard image region exists, the initial beard image region may be filled with image content representing facial skin based on the initial facial image to obtain an adjusted facial image. The presence of the initial beard image region may indicate that the face in the initial facial image already has a beard.
[0057] Step 106 : In the adjusted facial image, locate a target image region for generating a target beard image.
[0058] The target image region is the region where the target beard image to be generated is located. The target image region can be understood as the region where the beard image can be generated. The target image region can include, for example, at least one of the chin region, the lip region, and the lower cheek region.
[0059] Exemplarily, a computer device may perform facial key point detection on the adjusted facial image to obtain third facial key points corresponding to the adjusted facial image; extract first-category key points and second-category key points from the third facial key points, and determine the target image area based on the areas formed by the first-category key points and the second-category key points.
[0060] The first type of keypoints are used to form the outer contour of the target image area, while the second type of keypoints are used to indicate the area to be removed from the area formed by the first type of keypoints. Specifically, the first type of keypoints can be the keypoints of the lower half of the face, and the second type of keypoints can be the keypoints of the mouth outline. The target image area can be determined as the area remaining after removing the area formed by the second type of keypoints from the area formed by the first type of keypoints.
[0061] Step 108 , performing image generation based on the adjusted facial image and the reference beard image, so as to edit at least a portion of the target image area according to the reference beard image area, and generate a target facial image, wherein the target facial image includes image content of the edited target beard image.
[0062] The target image area is an area where a beard image can be generated, and the target image area can be adapted to various beard types. The target image area can be understood as the union of areas used to generate various beard types. At least a portion of the image area can be a portion of the target image area, or the entire target image area. For different beard types, at least a portion of the image area edited for the target image area can be different. For example, the target image area can include the chin area, the lip area, and the lower cheek area. For a goatee, at least a portion of the image area can be the oval area in the middle of the chin area. For a beard, at least a portion of the image area can be the lower cheek area.
[0063] The target beard image can be an image block representing the target beard. It is understood that the target beard can have the same style as the beard in the reference beard image, for example, having the same beard layout and hair arrangement. However, the target beard can be adapted to the edited image. For example, the target beard image can be reasonably integrated into the edited image, giving the appearance of the target beard naturally growing within the face represented by the adjusted facial image.
[0064] The target facial image includes the edited target beard image. The target facial image may include the target beard image, or the target facial image may include a beard image that has been fine-tuned to the target beard image, where the beard image retains the complete semantics of the target beard image. The fine-tuning may include, for example, super-resolution processing or edge processing.
[0065] Exemplarily, the initial facial image is obtained by adjusting the image view of the input facial image. The computer device may generate an image based on the adjusted facial image and the reference beard image to edit at least a portion of the target image region according to the reference beard image region to obtain an intermediate facial image. Image fusion processing is then performed based on the intermediate facial image and the input facial image to obtain the target facial image. The intermediate facial image includes the edited target beard image. Image generation can be achieved using an image generation model. The image generation model may be, for example, a denoising diffusion implicit model (DDIM), a latent diffusion model (LDM), a stable diffusion model (SD), or other models. After the image fusion process, super-resolution processing may be performed on the beard region in the fused image to obtain the target image.
[0066] In one embodiment, the initial facial image may be an input facial image, and the computer device may generate an image based on the adjusted facial image and the reference beard image to edit at least a portion of the image area in the target image area according to the reference beard image area to obtain the target facial image.
[0067] In the above-mentioned facial image processing method, based on the initial facial image, the initial beard image area is filled with image content representing facial skin. The adjusted facial image thus obtained, compared to the initial facial image, has the original beard in the initial facial image eliminated and is filled with image content representing facial skin, which can prevent the original beard in the initial facial image from affecting the effect of the subsequent target beard image. Furthermore, the target image area is located in the adjusted facial image, and then image generation is performed based on the adjusted facial image and the reference beard image, so that at least a portion of the image area in the target image area is edited according to the reference beard image area, and a target beard image can be obtained, thereby generating a target facial image, and the target facial image includes the image content of the target beard image. In this way, through the above-mentioned series of steps, the initial beard image area can be automatically eliminated and filled, the target image area for generating the target beard image can be automatically located, and then at least a portion of the image area in the target image area can be automatically edited according to the reference beard image area to form the target beard image, without the need for manual operation. While ensuring that the target facial image including the image content of the target beard image is accurately generated with reference to the reference beard image, the image processing efficiency is improved.
[0068] In an exemplary embodiment, step 108 may include: obtaining a target mask image that matches the size of the adjusted facial image, the target mask image being used to indicate the relative position of the target image area relative to the adjusted facial image; obtaining semantic description text of the reference beard image; and generating an image based on the adjusted facial image, the reference beard image, the target mask image, and the semantic description text to edit at least a portion of the image area in the target image area according to the reference beard image area to generate a target facial image.
[0069] The image size of the target mask image may be the same as the image size of the adjusted facial image. The target mask image may include a second area corresponding to the position of the target image area. The relative position of the second area relative to the target mask image may be the same as the relative position of the target image area relative to the adjusted facial image. The target mask image may be obtained by setting the pixels of the target image area in the adjusted facial image to the same non-zero pixels and setting the pixels of the area other than the target image area in the adjusted facial image to zero pixels. Thus, each pixel in the second area may be the same non-zero pixel, and the non-zero pixel may be, for example, a pixel representing white, and the pixel value of the pixel may be 1, 255 or other. Each pixel in the area outside the second area in the target mask image may be a zero pixel, and the pixel value of the pixel is zero, which may represent black.
[0070] Semantic description text is used to describe the semantic content of the reference beard image. This can be achieved through image captioning. For example, the semantic description text of the reference beard image could be "This is a goatee with bright color and clear hair details."
[0071] Exemplarily, the initial facial image is obtained by adjusting the image view of the input facial image. The computer device can obtain a target mask image that matches the size of the adjusted facial image, obtain semantic description text of the reference beard image, and perform image generation through an image generation model based on the adjusted facial image, reference beard image, target mask image and semantic description text to edit at least a part of the image area in the target image area according to the reference beard image area to obtain an intermediate facial image. Image fusion processing is performed based on the intermediate facial image and the input facial image to obtain a target facial image.
[0072] In this embodiment, the target mask image can indicate the relative position of the target image area with respect to the adjusted facial image, and the semantic description text can represent the semantic content of the reference beard image. In this way, image generation is performed based on the adjusted facial image, the reference beard image, the target mask image and the semantic description text, which can guide the accurate generation of a target beard image that matches the reference beard image at the position of the target image area, thereby improving the accuracy of the target beard image generation.
[0073] In an exemplary embodiment, image generation is performed based on the adjusted facial image, the reference beard image, the target mask image, and the semantic description text to edit at least a portion of the image area in the target image area according to the reference beard image area. The step of generating the target facial image may include: performing image feature extraction on the adjusted facial image and the reference beard image respectively to obtain facial features and reference beard features; performing feature fusion processing on the facial features, the reference beard features, and the target mask image to obtain fusion features; performing text feature extraction on the semantic description text to obtain text semantic features; performing image information generation processing based on the fusion features and the text semantic features to edit at least a portion of the image area in the target image area according to the reference beard image area to generate the target facial image.
[0074] The facial features can be semantic features used to adjust the facial image. They can be used to indicate adjustments to the position of facial features, facial skin texture, or other features in the facial image. Reference beard features can be stylistic features of the beard in the reference beard image. Reference beard features can be used to indicate adjustments to the shape, density, or other features of the beard in the facial image. Feature fusion processing can be implemented using feature splicing, directly splicing the facial features, reference beard features, and the target mask image to obtain fused features.
[0075] The text semantic feature can specifically be a feature vector output by a text encoder, which can be used as a conditional feature to control the details of the beard image generated by the image generation model. For example, it can be used to control the shape of the beard in the beard image to achieve precise alignment of text to image. Text encoders include the text encoder in the CLIP model (Contrastive Language-Image Pre-training, a multimodal pre-training model), the text encoder in the BLIP model (Bootstrapping Language-Image Pre-training, a multimodal framework that improves the joint understanding and generation of images and language through self-guided bootstrapping), or other text encoders.
[0076] Exemplarily, the initial facial image is obtained by adjusting the image view of the input facial image, and the image generation can be achieved by an image generation model, which may include an image encoder, an image information generator and an image decoder; the computer device can use the image encoder to extract image features of the adjusted facial image and the reference beard image respectively to obtain facial features and reference beard features; perform feature fusion processing on the facial features, reference beard features and the target mask image to obtain fusion features; use the text encoder to extract text features of the semantic description text to obtain text semantic features; based on the fusion features and text semantic features, use the image information generator to perform image information generation processing to edit at least a part of the image area in the target image area according to the reference beard image area to obtain an intermediate facial image, and perform image fusion processing based on the intermediate facial image and the input facial image to obtain a target facial image.
[0077] The image encoder can use the encoder in a variational autoencoder (VAE), a tiny autoencoder for stable diffusion (TAESD), or other methods. The image decoder can use the decoder in a variational autoencoder (VAE), a consistency decoder, or other methods. The image information generator can be a component in a diffusion model and may include a scheduler and a noise prediction network. The image information generation process may be an iterative denoising process. The scheduler may include a computational process for adding or removing noise, such as a mathematical formula or process. The scheduler may use the Scheduler component in a stable diffusion model. The noise prediction network is used to predict noise and may be a Unified Network (UNet).
[0078] See for example Figure 4 As shown in the flowchart of the image generation steps, the computer device can input the adjusted face image and the reference beard image into the encoder of the VAE respectively to obtain the face features and the reference beard features; perform a Concat operation (feature splicing processing) on the face features, the reference beard features and the target mask image to obtain the fused features; perform text feature extraction on the semantic description text through the text encoder in the CLIP model to obtain text semantic features; input the fused features and the text semantic features into the image information generator for iterative denoising processing to perform image information generation processing, and then the output features are decoded by the decoder of the VAE to obtain the intermediate face image.
[0079] The scheduler can add noise to the fused features to obtain the noisy fused features. Iterative denoising can be achieved through the following steps: In the first denoising time step, based on the noisy fused features and text semantic features, noise prediction is performed through the UNet network to obtain the predicted noise of the current denoising time step. The scheduler removes the predicted noise of the current denoising time step from the fused features to obtain the latent features generated in the first denoising time step. From the second denoising time step to the Nth denoising time step, in each denoising time step, based on the text semantic features and the latent features generated in the previous denoising time step, noise prediction is performed through the UNet network to obtain the predicted noise of the current denoising time step. The scheduler removes the predicted noise of the current denoising time step from the latent features generated in the previous denoising time step to obtain the latent features generated in the current denoising time step. N can be 100, 200, or other values.
[0080] In this embodiment, by extracting facial features and reference beard features, performing feature fusion processing on the facial features, reference beard features and target mask image to obtain fused features, and extracting text semantic features, various information can be converted into feature space for processing, and then image information generation processing can be performed based on the fused features and text semantic features. The accuracy of the target facial image finally generated can be improved by accurately adjusting the information in the facial image, reference beard image, target mask image and semantic description text.
[0081] In an exemplary embodiment, the step of obtaining an initial facial image in step 102 may include: obtaining an input facial image, performing facial key point detection on the input facial image, and obtaining a first facial key point corresponding to the input facial image; based on the first facial key point, correcting the facial perspective of the facial area in the input facial image to obtain a transformed facial image; based on the facial area in the transformed facial image, performing image cropping and then image scaling on the transformed facial image to obtain an initial facial image.
[0082] The correction process can be to convert the side view to the front view. Alternatively, the correction process can be to convert the tilted view to the eye-level view, aligning the tilted face to the eye-level view, thus achieving posture alignment. This correction process can be implemented through affine transformation or machine learning models.
[0083] During image cropping, the facial region within the transformed facial image can be determined by transforming the facial key points within the transformed facial image. Within the transformed facial image, the circumscribed rectangular frame of the facial region within the transformed facial image is expanded outward by a preset size ratio, and the expanded region is cropped to form a cropped facial image. The preset size ratio can be, for example, between 20% and 30%, or another ratio. It will be understood that the expanded region includes both the facial region and an extended region outside the facial region. By appropriately expanding outward, the complete facial information can be retained while also including edge information.
[0084] Image scaling can be performed to reduce or enlarge the image size. Image scaling can be performed on the cropped facial image to scale the image size of the cropped facial image to a preset image size. The preset image size can be, for example, 256*256. The preset image size can be an image size suitable for the image generation model.
[0085] For example, see Figure 5 As shown in the flowchart of the facial image preprocessing steps, the computer device can obtain an input facial image, perform facial key point detection on the input facial image through the MTCNN network, obtain a first facial key point corresponding to the input facial image, and based on the first facial key point, perform facial perspective correction processing on the facial area in the input facial image according to the horizontal correction strategy of the connection line of the facial key points corresponding to both eyes to obtain a transformed facial image; after expanding the circumscribed rectangular frame of the facial area in the transformed facial image outward by a preset size ratio, perform image cropping processing on the transformed facial image to crop the expanded area to obtain a cropped facial image, and perform image scaling processing on the cropped facial image to obtain an initial facial image.
[0086] In this embodiment, facial key point detection is first performed on the input facial image, so that based on the first facial key point corresponding to the input facial image, the facial perspective of the facial area in the input facial image can be quickly corrected, and then the transformed facial image is cropped and then scaled. Through this series of processing, the image quality of the initial facial image can be improved and the information in the image can be focused on the facial area, thereby improving the efficiency of subsequent image processing.
[0087] In an exemplary embodiment, the initial facial image is obtained by adjusting the image view of the input facial image, and step 108 may include: performing image generation based on the adjusted facial image and the reference beard image to edit at least a portion of the image area in the target image area according to the reference beard image area to obtain an intermediate facial image containing the edited target beard image; extracting the target beard image from the intermediate facial image to obtain a target beard image; processing the target beard image according to the target image processing method determined by the image view adjustment, and then performing image fusion processing with the input facial image to obtain a fused facial image; the fused beard image area in the fused facial image contains the image content of the target beard image; performing image super-resolution processing on the fused beard image area in the fused facial image to obtain the target facial image.
[0088] The target image processing method may be the inverse processing method corresponding to the processing method for adjusting the face area in the image view adjustment. Specifically, the image view adjustment may include the above-mentioned correction processing, image cropping processing, and image scaling processing. The processing method for adjusting the face area may be correction processing and image scaling processing. In this case, the target image processing method may be the inverse processing method of each of the correction processing and image scaling processing. For example, the image scaling processing is a process of reducing the image size by a preset multiple. In this case, the inverse processing method of the image scaling processing is a process of enlarging the image size by a preset multiple. The preset multiple may be 20%, 30%, or other.
[0089] Extracting the target beard image can be understood as cropping the target beard image from the intermediate face image to obtain the target beard image. Image super-resolution processing is the process of increasing image resolution. Image super-resolution processing can be achieved using deep learning super-resolution models, such as the SRGAN model (an image super-resolution reconstruction model based on a generative adversarial network proposed by Christian Ledig et al.), the ESRGAN model (an image super-resolution reconstruction model based on a generative adversarial network proposed by Tencent ARC Lab, an improved version of the SRGAN model), or other deep learning super-resolution models.
[0090] In this embodiment, since the initial facial image is obtained by adjusting the image view of the input facial image, an intermediate facial image can be efficiently generated by performing image generation based on the adjusted facial image obtained based on the initial facial image and the reference beard image. Furthermore, the target beard image is extracted, and after being processed according to the target image processing method determined by the image view adjustment, image fusion processing is performed with the input facial image. The obtained fused facial image can achieve the effect of generating the target beard in the input facial image, which meets the user's intention and maintains the original resolution of the input facial image. Image super-resolution processing is then performed on the fused beard image area in the fused facial image, which can further improve the quality of the generated beard and improve the quality of the final generated image.
[0091] In an exemplary embodiment, the target beard image is processed according to the target image processing method determined by the image view adjustment, and then image fusion processing is performed with the input face image to obtain a fused face image. The step of obtaining the fused face image may include: determining the target image processing method according to the image view adjustment, processing the target beard image by the target image processing method, and obtaining a processed target beard image; obtaining a first facial key point corresponding to the input face image, and determining a second facial key point in the processed target beard image; aligning the facial key points of the processed target beard image with the input face image according to the first facial key point and the second facial key point, and then performing image fusion processing based on a preconfigured fusion method to obtain a fused face image.
[0092] The second facial key points are facial key points in the processed target beard image. The second facial key points may correspond to some of the key points in the first facial key points. Aligning the facial key points of the processed target beard image with the input facial image may involve overlapping the second facial key points with the key points corresponding to the positions in the first facial key points. This allows the processed target beard image to be accurately fused to the input facial image. For example, the second facial key points may include key points in the chin region. The second facial key points may correspond to the key points in the chin region in the first facial key points.
[0093] After aligning the facial key points of the processed target beard image with the input face image, the image fusion process can be performed. The processed target beard image can retain the beard area and remove the remaining area before being fused with the input face image. Pre-configured fusion methods include Poisson fusion, pyramid fusion, and other methods. The core algorithm for image fusion can use either a gradient domain fusion algorithm or a mask-guided fusion algorithm. The former utilizes the seamlessClone function in OpenCV (a cross-platform computer vision library), using the beard area of the target beard image as the foreground and the input face image as the background, minimizing the gradient difference at the intersection. The latter uses a mask image corresponding to the beard area of the target beard image and defines fusion weights using the alpha channel to preserve the details of the beard area while blurring the edge transition area.
[0094] In this embodiment, after the target beard image is processed according to the target image processing method, conditions are created for the subsequent target beard image to be able to align the facial key points with the input face image. Furthermore, the facial key points of the processed target beard image are aligned with the input face image, which can improve the position accuracy of the beard finally generated. When image fusion processing is performed based on a preconfigured fusion method to obtain a fused face image, while retaining the original resolution of the input face image, the fused face image forms a new beard compared to the input face image.
[0095] In an exemplary embodiment, see Figure 6 The flowchart of the facial image post-processing steps shown in the figure shows that the computer device can generate an intermediate facial image, extract a target beard image from the intermediate facial image, adjust the target image processing method according to the image view, obtain the first facial key point corresponding to the input facial image, and determine the second facial key point in the processed target beard image; process the target beard image by the target image processing method to obtain a processed target beard image, align the facial key points of the processed target beard image with the input facial image according to the first facial key point and the second facial key point, and then attach the processed target beard image to the input facial image; perform image fusion processing based on Poisson fusion to obtain a fused facial image; perform image super-resolution processing on the fused beard image area in the fused facial image to obtain the target facial image.
[0096] In an exemplary embodiment, step 106 may include: performing facial key point detection on the adjusted facial image to obtain third facial key points corresponding to the adjusted facial image; extracting, from the third facial key points, earlobe key points representing the earlobe position, nose key points representing the nose bottom position, facial contour key points representing the facial contour position, and mouth contour key points representing the mouth contour position; forming a first local area based on the earlobe key points, nose key points, and facial contour key points, and forming a second local area based on the mouth contour key points; and determining the area remaining after removing the second local area from the first local area as the target image area for generating the target beard image.
[0097] The third facial key point is a facial key point in the adjusted facial image. The earlobe position can specifically be the bottom of the earlobe. The facial contour position can include the cheek contour and the chin contour. The first local area can be understood as the image area formed by the lower half of the face. The second local area can represent the area occupied by the mouth of the face.
[0098] Using Dlib Taking the 68-point model (a pre-trained model that uses 68 key points to define facial structure) as an example, the third facial key point may include 68 facial key points. The earlobe key points may be points 3 and 15. The facial contour key points are the outer contour key points below the earlobe key points, which can be 10 key points from left to right, namely points 4 to 14. The nose key points can be points 32 to 36 from left to right. The mouth contour key points can be points 49 to 60. In this way, when connecting to form the first local area, the connection can be performed in a clockwise direction. The specific connection order can be: point 3 → point 32 → point 33 → point 34 → point 35 → point 36 → point 15 → point 14... (the remaining mouth contour key points 4 to point 13 are connected in descending order) → point 3. When connecting to form the second local area, the mouth contour key points are connected to each point in sequence from point 49 until point 60, and then point 60 is connected to point 49.
[0099] For example, see Figure 7In the flowchart of the region positioning steps shown, a computer device can detect facial landmarks on the adjusted facial image to obtain third facial landmarks corresponding to the adjusted facial image. The earlobe, nose, and facial contour landmarks within the third facial landmarks are connected to form a first local region (denoted as m1). The mouth contour landmarks within the third facial landmarks are connected to form a second local region (denoted as m2). The remaining region (denoted as m1 - m2) after removing the second local region from the first local region is determined as the target image region for generating the target beard image. A target mask image corresponding to the target beard image can be generated using the mask formula M_beard = M_m1 - M_m2, where M_beard represents the target mask image, M_m1 represents the mask image corresponding to m1, and M_m2 represents the mask image corresponding to m2.
[0100] In this embodiment, by extracting the earlobe key points, nose key points and facial contour key points, connecting them to form a first local area, and extracting the mouth contour key points, connecting them to form a second local area, and then eliminating the second local area from the first local area, the remaining area is determined as the target image area for generating the target beard image. This can form a target image area that may be used to generate the beard image to a large extent, creating conditions for subsequent more accurate generation of the beard image.
[0101] In a specific embodiment, see Figure 8 The face image processing method may specifically include the following steps: a computer device may obtain an input face image, pre-process the input face image, and obtain an initial face image. The pre-processing steps may include sequentially performing face key point detection, correction processing, image cropping processing, and image scaling processing.
[0102] The computer device can perform image segmentation processing on the initial face image through an image segmentation network to identify an initial beard image region in the initial face image and output an initial mask image, which indicates the relative position of the initial beard image region with respect to the initial face image.
[0103] The computer device can use the image filling model to fill the image content representing the facial skin in the initial beard image area based on the initial facial image to obtain an adjusted facial image.
[0104] The computer device can locate the target image area for generating the target beard image in the adjusted facial image and output the target mask image.
[0105] The computer device can obtain a reference beard image, and generate an image based on the adjusted facial image, the reference beard image, the target mask image and the semantic description text through a diffusion model, so as to edit at least a part of the image area in the target image area according to the reference beard image area, and obtain an intermediate facial image containing the edited target beard image.
[0106] The computer device can perform post-processing steps based on the intermediate facial image and the input facial image, specifically including: extracting a target beard image from the intermediate facial image, processing the target beard image according to the target image processing method, and then performing image fusion processing with the input facial image to obtain a fused facial image; performing image super-resolution processing on the fused beard image area in the fused facial image to obtain the target facial image.
[0107] Compared to traditional approaches that directly overlay a beard image on a facial image, this facial image processing method incorporates beard detection (implemented through image segmentation) and beard filling (implemented through image filling) on the initial facial image. This two-stage process completely removes the existing beard from the initial facial image, avoiding issues such as color clashes (such as gray spots), texture interference (double hair flow), and shape misalignment (residual hair) caused by the overlay of old and new beards. Measured data: On a test set containing beards, the fusion naturalness (MOS score) improved from 3.2 for traditional approaches to 4.6 (out of a possible 5.0). Traditional approaches ignore the presence of beards in the original image input by the user and directly overlay the generated image, resulting in the distortion of beards overlapping each other in the chin area.
[0108] Compared to using GANs for image generation, this facial image processing method utilizes diffusion models (such as the Stable Diffusion model, optionally overlaid with a ControlNet architecture (which controls the generation process of the diffusion model through conditional inputs)). Leveraging the diffusion model's multi-step iterative generation capabilities, this method achieves hair-level detail control: hair edges are sharp, generating sub-pixel outlines of individual beards (improving edge sharpness by 40%). It also supports the generation of complex beard structures, supporting non-uniform beard styles such as curls and braids, overcoming the pattern collapse limitations of GANs. Furthermore, it achieves aesthetic adaptation by aligning user text descriptions (e.g., "retro gentleman mustache") using the CLIP model to generate stylized results. Traditional image generation relies on GANs, which can produce results with low-frequency blurring (foggy edges) and a single structure (limited to trained beard styles), forcing users to choose from a limited number of preset beard styles.
[0109] Compared to traditional methods that pre-configure reference beard images and train on them, this application's facial image processing method utilizes reference image-guided generation technology, enabling a single model to adapt to all beard styles. New beard styles require no retraining, requiring only a single reference image (reducing annotation costs by 99%), thus reducing training costs. It also minimizes storage requirements: 10 beard styles require only one 2GB (gigabyte) base model (a pre-trained diffusion model) plus ten 50MB (megabyte) reference images (for fine-tuning training), for a total of 2.5GB, compared to 15GB required by traditional methods. It also offers strong customization capabilities, allowing users to upload hand-drawn sketches or images of unusual beard shapes to generate corresponding styles. Traditional methods utilize a discrete model architecture, requiring separate training for each beard type (e.g., a full beard and a goatee, each requiring a separate model). This results in: explosive data requirements: 100 beard styles require 100,000 labeled samples (while this application only requires 100 samples for fine-tuning the base model); and an inability to handle unknown beard styles: users cannot generate designs outside the training set (e.g., a dragon beard or a cyberpunk beard).
[0110] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0111] Based on the same inventive concept, embodiments of the present application also provide a facial image processing device for implementing the aforementioned facial image processing method. The implementation solution provided by this device is similar to the implementation solution described in the aforementioned method. Therefore, the specific limitations in one or more facial image processing device embodiments provided below can be found in the above-mentioned limitations on the facial image processing method and will not be further elaborated here.
[0112] In an exemplary embodiment, Figure 9 As shown, a face image processing device 900 is provided, comprising: an acquisition module 910, a filling module 920, a positioning module 930 and a generation module 940, wherein:
[0113] The acquisition module 910 is used to acquire an initial face image and a reference beard image; the initial face image includes an initial beard image area, and the reference beard image includes a reference beard image area.
[0114] The filling module 920 is configured to fill the initial beard image region with image content representing facial skin based on the initial facial image to obtain an adjusted facial image.
[0115] The positioning module 930 is used to locate the target image area for generating the target beard image in adjusting the face image.
[0116] A generation module 940 is used to generate an image based on the adjusted facial image and the reference beard image, so as to edit at least a portion of the image area in the target image area according to the reference beard image area to generate a target facial image, wherein the target facial image includes the image content of the edited target beard image.
[0117] In an exemplary embodiment, the generation module 940 is also used to obtain a target mask image that matches the size of the adjusted facial image, the target mask image being used to indicate the relative position of the target image area relative to the adjusted facial image; obtain semantic description text of the reference beard image; and perform image generation based on the adjusted facial image, the reference beard image, the target mask image, and the semantic description text to edit at least a portion of the image area in the target image area according to the reference beard image area to generate a target facial image.
[0118] In an exemplary embodiment, the generation module 940 is also used to perform image feature extraction on the adjusted facial image and the reference beard image respectively to obtain facial features and reference beard features; perform feature fusion processing on the facial features, reference beard features and the target mask image to obtain fusion features; perform text feature extraction on the semantic description text to obtain text semantic features; perform image information generation processing based on the fusion features and text semantic features to edit at least a part of the image area in the target image area according to the reference beard image area to generate a target facial image.
[0119] In an exemplary embodiment, the initial facial image is obtained by adjusting the image view of the input facial image, and the generation module 940 is further used to generate an image based on the adjusted facial image and the reference beard image, so as to edit at least a part of the image area in the target image area according to the reference beard image area, and obtain an intermediate facial image containing the edited target beard image; extract the target beard image from the intermediate facial image to obtain a target beard image; process the target beard image according to the target image processing method determined by the image view adjustment, and then perform image fusion processing with the input facial image to obtain a fused facial image; the fused beard image area in the fused facial image contains the image content of the target beard image; and perform image super-resolution processing on the fused beard image area in the fused facial image to obtain the target facial image.
[0120] In an exemplary embodiment, the generation module 940 is also used to determine the target image processing method according to the image view adjustment, process the target beard image using the target image processing method, and obtain the processed target beard image; obtain the first facial key point corresponding to the input face image, and determine the second facial key point in the processed target beard image; according to the first facial key point and the second facial key point, align the facial key points of the processed target beard image with the input face image, and perform image fusion processing based on the preconfigured fusion method to obtain a fused face image.
[0121] In an exemplary embodiment, the acquisition module 910 is also used to obtain an input facial image, perform facial key point detection on the input facial image, and obtain a first facial key point corresponding to the input facial image; based on the first facial key point, correct the facial perspective of the facial area in the input facial image to obtain a transformed facial image; based on the facial area in the transformed facial image, perform image cropping and then image scaling on the transformed facial image to obtain an initial facial image.
[0122] In an exemplary embodiment, the positioning module 930 is also used to perform facial key point detection on the adjusted facial image to obtain third facial key points corresponding to the adjusted facial image; from the third facial key points, extract the earlobe key point representing the earlobe position, the nose key point representing the nose bottom position, the facial contour key point representing the facial contour position, and the mouth contour key point representing the mouth contour position; form a first local area based on the earlobe key point, the nose key point, and the facial contour key point, and form a second local area based on the mouth contour key point; determine the area remaining after removing the second local area from the first local area as the target image area for generating the target beard image.
[0123] Each module in the facial image processing device described above may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0124] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 10As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store data that needs to be stored when executing the above-mentioned facial image processing method. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a facial image processing method is implemented.
[0125] Those skilled in the art will understand that Figure 10 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0126] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0127] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0128] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0129] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0130] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.
[0131] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0132] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A facial image processing method, characterized in that: The method comprises: Acquire an initial face image and a reference beard image; the initial face image includes an initial beard image region, and the reference beard image includes a reference beard image region; Based on the initial facial image, filling the initial beard image region with image content representing facial skin to obtain an adjusted facial image; In the adjusted facial image, locating a target image region for generating a target beard image; Image generation is performed based on the adjusted facial image and the reference beard image to edit at least a portion of the target image area according to the reference beard image area to generate a target facial image, wherein the target facial image includes image content of the edited target beard image.
2. The method according to claim 1, characterized in that The step of generating an image based on the adjusted facial image and the reference beard image, so as to edit at least a portion of the target image region according to the reference beard image region to generate a target facial image, includes: Acquire a target mask image that matches the size of the adjusted facial image, where the target mask image is used to indicate a relative position of the target image region relative to the adjusted facial image; Obtaining semantic description text of the reference beard image; Image generation is performed based on the adjusted facial image, the reference beard image, the target mask image, and the semantic description text, so as to edit at least a portion of the image area in the target image area according to the reference beard image area to generate a target facial image.
3. The method according to claim 2, characterized in that The step of generating an image based on the adjusted facial image, the reference beard image, the target mask image, and the semantic description text, so as to edit at least a portion of the target image region according to the reference beard image region to generate a target facial image, includes: Performing image feature extraction on the adjusted face image and the reference beard image respectively to obtain face features and reference beard features; Performing feature fusion processing on the facial features, the reference beard features and the target mask image to obtain fused features; Performing text feature extraction on the semantic description text to obtain text semantic features; Image information generation processing is performed based on the fusion feature and the text semantic feature to edit at least a portion of the target image region according to the reference beard image region to generate a target face image.
4. The method according to claim 1, wherein The initial facial image is obtained by adjusting an input facial image through image view adjustment, and the image generation is performed based on the adjusted facial image and the reference beard image to edit at least a portion of the target image region according to the reference beard image region to generate the target facial image, including: generating an image based on the adjusted facial image and the reference beard image, so as to edit at least a portion of the target image region according to the reference beard image region, and obtain an intermediate facial image including the edited target beard image; Extracting the target beard image from the intermediate face image to obtain a target beard image; After processing the target beard image according to the target image processing method determined by the image view adjustment, image fusion processing is performed with the input face image to obtain a fused face image; the fused beard image region in the fused face image contains the image content of the target beard image; Perform image super-resolution processing on the fused beard image region in the fused face image to obtain a target face image.
5. The method according to claim 4, characterized in that The step of processing the target beard image according to the target image processing mode determined by adjusting the image view, and then performing image fusion processing on the target beard image with the input face image to obtain a fused face image includes: adjusting and determining a target image processing method according to the image view, processing the target beard image using the target image processing method to obtain a processed target beard image; Obtaining first facial key points corresponding to the input facial image, and determining second facial key points in the processed target beard image; According to the first facial key points and the second facial key points, the processed target beard image and the input facial image are aligned with each other in facial key points, and then image fusion processing is performed based on a preconfigured fusion method to obtain a fused facial image.
6. The method according to claim 1, wherein The obtaining of the initial face image comprises: Acquire an input face image, perform face key point detection on the input face image, and obtain a first face key point corresponding to the input face image; Based on the first facial key points, correcting the facial perspective of the face area in the input facial image to obtain a transformed facial image; Based on the face area in the transformed face image, the transformed face image is subjected to image cropping processing and then image scaling processing to obtain an initial face image.
7. The method according to any one of claims 1 to 6, characterized in that The step of locating a target image region for generating a target beard image in the adjusted face image includes: Performing facial key point detection on the adjusted facial image to obtain third facial key points corresponding to the adjusted facial image; Extracting, from the third facial key points, an earlobe key point representing an earlobe position, a nose key point representing a nose bottom position, a facial contour key point representing a facial contour position, and a mouth contour key point representing a mouth contour position; Connecting the earlobe key point, the nose key point, and the face contour key point to form a first local area, and connecting the mouth contour key point to form a second local area; The remaining area after removing the second local area from the first local area is determined as the target image area for generating the target beard image.
8. A facial image processing device, characterized in that: The device comprises: An acquisition module, configured to acquire an initial face image and a reference beard image; the initial face image includes an initial beard image region, and the reference beard image includes a reference beard image region; a filling module, configured to fill, based on the initial facial image, the initial beard image region with image content representing facial skin, to obtain an adjusted facial image; a positioning module, configured to locate a target image region for generating a target beard image in the adjusted face image; A generation module is used to generate an image based on the adjusted facial image and the reference beard image, so as to edit at least a portion of the image area in the target image area according to the reference beard image area to generate a target facial image, wherein the target facial image includes the image content of the edited target beard image.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.