A method, device and electronic device for generating a face-swapped portrait photo
By obtaining portrait images and style template images for training and mask processing, the target face change portrait photos are generated, which solves the problem of the inability to take into account the similarity and style consistency of portrait face in the prior art, and achieves high-quality face change portrait photos generation.
Patent Information
- Application Number
- CN202410941370.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-15
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-07-15
AI Technical Summary
The existing face-changing portrait photo generation solution cannot take into account the similarity and style of portrait face similarity and style.
By acquiring portrait images and style template images, performing portrait training to generate portrait models, constructing face mask images, and generating target face-changing portrait photos based on target portrait template images and face mask images to ensure the similarity of face areas and style consistency.
It realizes the effect of taking into account the similarity and style of the portrait face in the face change portrait, and generates high-quality anime face change portraits.
Smart Images

Figure CN119027301B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of image vision generation, and in particular, to a method, apparatus, and electronic device for generating a face-swapped portrait photo. Background Art
[0002] Currently, artificial intelligence (AI) portrait photos can generate highly imaginative images according to specified roles, bringing great convenience to content creation, art design, and image creation. Among them, the generation of AI portrait photos with an anime style (i.e., anime face-swapped portrait photos) has a wide range of applications.
[0003] In the process of implementing the present invention, the inventors found the following technical problems in the prior art: The existing face-swapped portrait photo generation scheme cannot balance the similarity of the human face and the consistency of the style. Summary of the Invention
[0004] Embodiments of the present invention provide a method, apparatus, and electronic device for generating a face-swapped portrait photo, which solve the problem of inability to balance the similarity of the human face and the consistency of the style.
[0005] According to one aspect of the present invention, there is provided a method for generating a face-swapped portrait photo, which may include:
[0006] In response to a face-swapped portrait photo generation instruction, obtain a portrait image and a style template image, where the face-swapped portrait photo generation instruction is an instruction for indicating to replace the human face area in the portrait image with the style face area in the style template image to generate a target face-swapped portrait photo;
[0007] Perform portrait training based on the portrait image to obtain a portrait model, and generate a target portrait template image based on the portrait model and the style template image;
[0008] Construct a face mask image corresponding to the style template image, and generate a target face-swapped portrait photo based on the target portrait template image and the face mask image.
[0009] According to another aspect of the present invention, there is provided a face-swapped portrait photo generation apparatus, which may include:
[0010] An image acquisition module, configured to obtain a portrait image and a style template image in response to a face-swapped portrait photo generation instruction, where the face-swapped portrait photo generation instruction is an instruction for indicating to replace the human face area in the portrait image with the style face area in the style template image to generate a target face-swapped portrait photo;
[0011] A target portrait template image generation module, configured to perform portrait training based on a portrait image to obtain a portrait model, and generate a target portrait template image based on the portrait model and a style template image;
[0012] A target face-swapped portrait photo generation module, configured to construct a face mask image corresponding to the style template image, and generate a target face-swapped portrait photo based on the target portrait template image and the face mask image.
[0013] According to another aspect of the present invention, there is provided an electronic device, which may include:
[0014] At least one processor; and
[0015] A memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is caused to implement the face-swapped portrait photo generation method provided in any embodiment of the present invention.
[0017] The technical solution of the embodiment of the present invention is directed to a face-swapped portrait photo generation instruction for instructing to replace the portrait face region in the portrait image with the style face region in the style template image to generate a target face-swapped portrait photo. By responding to the face-swapped portrait photo generation instruction, a portrait image and a style template image are obtained; then, portrait training is performed based on the portrait image to obtain a portrait model, and further, a target portrait template image is generated based on the portrait model and the style template image. The face region in the target portrait template image is highly similar to the portrait face region, but the remaining regions in the target portrait template image have changed compared with the style template image, that is, the style of the target portrait template image is no longer the same as the style of the style template image; further, a face mask image corresponding to the style template image is constructed, and the face mask image can represent the region to be replaced (i.e., the face region to ensure portrait face similarity) and the region not to be replaced (i.e., the remaining regions to ensure style consistency) related to the style template image, and then a target face-swapped portrait photo is generated based on the target portrait template image and the face mask image. The above technical solution replaces the face region in the target portrait template image with the style template image through the face mask image, thereby effectively generating a target face-swapped portrait photo that takes into account both portrait face similarity and style consistency.
[0018] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0020] Figure 1 It is a schematic diagram of two template images in anime style;
[0021] Figure 2a It is a schematic diagram of a generation example in a currently applied method for generating a face-swapped portrait photo;
[0022] Figure 2b It is a schematic diagram of a generation example in another currently applied method for generating a face-swapped portrait photo;
[0023] Figure 3 It is a flowchart of a method for generating a face-swapped portrait photo according to an embodiment of the present invention;
[0024] Figure 4 It is a schematic diagram of a mask example for processing an image and a facial mask image in a method for generating a face-swapped portrait photo according to an embodiment of the present invention;
[0025] Figure 5 It is in a method for generating a face-swapped portrait photo according to an embodiment of the present invention Figure 2a and Figure 2b Based on this, it is a schematic diagram of an anime face-swapped portrait photo generated;
[0026] Figure 6 It is a flowchart of another method for generating a face-swapped portrait photo according to an embodiment of the present invention;
[0027] Figure 7 It is a schematic diagram of a style template image and a corresponding facial template image in another method for generating a face-swapped portrait photo according to an embodiment of the present invention;
[0028] Figure 8 It is a flowchart of yet another method for generating a face-swapped portrait photo according to an embodiment of the present invention;
[0029] Figure 9 It is a schematic diagram of the generation process of a target portrait template image in yet another method for generating a face-swapped portrait photo according to an embodiment of the present invention;
[0030] Figure 10 It is a flowchart of still another method for generating a face-swapped portrait photo according to an embodiment of the present invention;
[0031] Figure 11It is a schematic diagram of the generation process of a pre-alignment face-swapped portrait photo in yet another face-swapped portrait photo generation method provided by an embodiment of the present invention;
[0032] Figure 12 It is a flowchart of an optional example in yet another face-swapped portrait photo generation method provided by an embodiment of the present invention;
[0033] Figure 13 It is a structural block diagram of a face-swapped portrait photo generation device provided by an embodiment of the present invention;
[0034] Figure 14 It is a schematic structural diagram of an electronic device for implementing the face-swapped portrait photo generation method of an embodiment of the present invention. Detailed implementation manners
[0035] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0036] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. The same is true for "target", "original", etc., which will not be elaborated here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0037] It should be noted that in the technical solution of the present invention, in terms of the collection, collection, update, analysis, processing, use, transmission, storage, etc. of user personal information, all comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. Necessary measures are taken for user personal information to prevent illegal access to user personal information data and maintain the security of user personal information and network security.
[0038] Before introducing the embodiments of the present invention, an exemplary description of the application scenarios that the embodiments of the present invention may involve will be given first, so as to better understand the process of generating a face-swapped portrait photo described in the following embodiments.
[0039] Exemplarily, taking the process of generating an anime face-swapped portrait photo as an example, several portrait images input by the user are obtained, and a certain anime style template image selected by the user from a variety of preset anime style template images is obtained. Among them, Figure 1 Two anime style template images are exemplarily shown. On this basis, an anime face-swapped portrait photo can be generated based on the portrait image and the selected anime style template image. The anime face-swapped portrait photo can be understood as a portrait photo obtained by replacing the portrait face area in the portrait image with the style face area in the anime style template image.
[0040] However, the currently applied anime face-swapped portrait photo generation scheme cannot take into account both the similarity of the portrait face and the style consistency. For example, Figure 2a The example given (from left to right are the anime style template image, the portrait image, and the anime face-swapped portrait photo) performs better in terms of style consistency, but performs poorly in terms of the similarity of the portrait face; another example is Figure 2b The example given (each image from left to right has the same meaning as Figure 2a ) performs better in terms of the similarity of the portrait face, but performs poorly in terms of style consistency, and urgent improvement is needed.
[0041] In addition, before introducing the embodiments of the present invention, some technical terms that the embodiments of the present invention may involve will also be explained first, so as to better understand the process of generating a face-swapped portrait photo described below.
[0042] 1. AIGC: The abbreviation of AI-Generated Content, also known as Generative AI, is a new content creation method following Professional-generated Content (PGC) and User-generated Content (UGC), and can create new digital content generation and interaction forms in aspects such as dialogue, story, image, video, and music production.
[0043] 2. Stable Diffusion: A text-to-image generation model based on Latent Diffusion Models, which can generate high-quality, high-resolution, and highly realistic images according to any text input.
[0044] 3. Magic Avatar: A technology for generating images based on AIGC text generation. In this scenario, the user uploads multiple single-person portraits, and Magic Avatar generates multiple imaginative portraits based on these portraits using the StableDiffusion text-to-image technology.
[0045] Figure 3 FIG. is a flowchart of a method for generating a face-swapped portrait photo provided by an embodiment of the present invention. This embodiment is applicable to the situation of generating a face-swapped portrait photo, especially applicable to the situation of generating a face-swapped portrait photo that takes into account both the similarity of the human face in the portrait and the consistency of the style. This method can be executed by a face-swapped portrait photo generation device provided by an embodiment of the present invention. The device can be implemented in software and / or hardware, and can be integrated into an electronic device, which can be various user terminals or servers.
[0046] See Figure 3 , the method of the embodiment of the present invention specifically includes the following steps:
[0047] S110. In response to a face-swapped portrait photo generation instruction, obtain a portrait image and a style template image, where the face-swapped portrait photo generation instruction is an instruction for instructing to replace the human face area in the portrait image with the style face area in the style template image to generate a target face-swapped portrait photo.
[0048] Among them, the portrait image can be understood as an image that can represent a human figure, and further can be understood as an image that includes a human face area, and the human face area is the area where the face is located in the portrait image.
[0049] The style template image can be understood as an image that can represent a certain style and is used as a template. The style template image at least includes a style face area, and the style face area is the area where the face is located in the style template image. Exemplarily, the style template image can be an anime style template image, a professional ID photo style template image, a wedding photo style template image, or a lazy coffee style template image, etc., which is related to the actual situation and is not specifically limited here.
[0050] The face-swapped portrait photo generation instruction can be understood as an instruction for instructing to replace the human face area with the style face area to generate a target face-swapped portrait photo, that is, an instruction for instructing to replace the style face area in the style template image with the human face area to generate a target face-swapped portrait photo. Exemplarily, the instruction can be an instruction triggered by the user after uploading the portrait image and selecting the style template image by clicking the corresponding control. In response to the face-swapped portrait photo generation instruction, obtain the portrait image and the style template image.
[0051] S120. Perform portrait training based on the portrait image to obtain a portrait model, and generate a target portrait template image based on the portrait model and the style template image.
[0052] Among them, performing portrait training based on the portrait image to obtain a portrait model, this portrait model can reflect the character image represented by this portrait image, especially the portrait face area in this portrait image. Exemplarily, performing portrait training based on the portrait image, training the character image in the portrait image onto a portrait label (such as the word sequence ABC), to obtain a portrait model, so that the portrait label can be obtained from the portrait model subsequently, and then the character image represented by this portrait image, especially the portrait face area in this portrait image, can be obtained based on the portrait label. In practical applications, optionally, portrait training can be performed based on solutions such as Text Inversion, DreamBooth, or Lora Training, etc., which are not limited here.
[0053] Generate a target portrait template image based on the portrait model and the style template image. Compared with the style template image, the target portrait template image is more similar to the character image represented by the portrait model, especially the portrait face area in this portrait image, which also makes the above target portrait template image can also be called a real person template image. In addition, it should be noted that the remaining areas in the target portrait template image may change compared with the style template image, which makes the style of the target portrait template image no longer consistent with the style of the style template image, and this will be processed in the subsequent steps. In practical applications, optionally, the above remaining areas are areas other than the face area, such as at least one of the background area, clothing area, and hair area, etc., which is related to the actual situation and is not specifically limited here.
[0054] S130. Construct a face mask image corresponding to the style template image, and generate a target face-swapped portrait based on the target portrait template image and the face mask image.
[0055] Among them, the face mask image can be an image corresponding to the style template image used to distinguish the face area and the remaining areas. Constructing the face mask image, for example, the face mask image can be directly constructed based on the style template image, that is, this face mask image can be used to distinguish the face area and the remaining areas in the style template image; or the style template image can be processed first, and then the face mask image can be constructed based on the obtained processed image, that is, this face mask image can be used to distinguish the face area and the remaining areas in the processed image; etc., which are not specifically limited here. Exemplarily, Figure 4 The left side shows a mask example of the processed image and Figure 4The right side shows the face mask image. It can be understood that the face area is the area that needs to be replaced in relation to the style template image to ensure the similarity of the portrait face, while the remaining areas are the areas that do not need to be replaced in relation to the style template image to ensure style consistency.
[0056] As described above, the face area in the target portrait template image is relatively similar to the portrait face area, but the style of the target portrait template image is inconsistent with the style of the style template image. On this basis, in order to generate a target face-swapped portrait photo that can take into account both the similarity of the portrait face and style consistency, based on the fact that the face mask image can represent the areas that need to be replaced and the areas that do not need to be replaced in relation to the style template image, a target face-swapped portrait photo can be generated based on the target portrait template image and the face mask image. Further, a target face-swapped portrait photo can be generated based on the target portrait template image, the face mask image, and the construction image (such as the style template image or the processed image) used to construct the face mask image. Exemplarily, the area 1 that needs to be replaced in the construction image can be determined based on the face mask image, and then the area 2 corresponding to the area 1 in the target portrait template image is replaced onto the area 1. The area 1 is the face area in the construction image, and the area 2 is the face area in the target portrait template image that is relatively similar to the portrait face area, so as to ensure that only the face area changes in the replaced construction image and the remaining areas remain unchanged, thereby ensuring that the finally generated target face-swapped portrait photo can take into account both the similarity of the portrait face and style consistency.
[0057] Exemplarily, here Figure 2a and Figure 2b the anime style template image and the portrait image shown in are taken as examples. After processing the anime style template image and the portrait image using the above face-swapped portrait photo generation process, the Figure 5 shown anime face-swapped portrait photo can be obtained. According to Figure 5 it can be seen that the style of the anime face-swapped portrait photo is consistent with the style of the anime style template image, and the face area in the anime face-swapped portrait photo is highly similar to the portrait face area in the portrait image, achieving the simultaneous consideration of the similarity of the portrait face and style consistency. In addition, the above style consistency is mainly for the remaining areas, and the style of the face area in the anime face-swapped portrait photo is also consistent with the style of the face area in the anime style template image, thereby further ensuring style consistency; on this basis, the face area in the anime face-swapped portrait photo is complete and normal, thus achieving high-quality anime face swapping.
[0058] The technical solution of the embodiment of the present invention is directed to a face-swapped portrait photo generation instruction for indicating to replace the portrait face area in a portrait image with the style face area in a style template image to generate a target face-swapped portrait photo. By responding to the face-swapped portrait photo generation instruction, a portrait image and a style template image are obtained; then, portrait training is performed based on the portrait image to obtain a portrait model, and further, based on the portrait model and the style template image, a target portrait template image is generated. The face area in the target portrait template image is highly similar to the portrait face area, but the remaining areas in the target portrait template image have changed compared to the style template image, that is, the style of the target portrait template image is no longer the same as the style of the style template image; further, a face mask image corresponding to the style template image is constructed. The face mask image can represent the area to be replaced (i.e., the face area to ensure the similarity of the portrait face) and the area that does not need to be replaced (i.e., the remaining areas to ensure the consistency of the style) related to the style template image. Then, based on the target portrait template image and the face mask image, a target face-swapped portrait photo is generated. In the above technical solution, through the face mask image, the face area in the target portrait template image is replaced with the face area in the style template image, thereby effectively generating a target face-swapped portrait photo that takes into account both the similarity of the portrait face and the consistency of the style.
[0059] Figure 6 It is a flowchart of another method for generating a face-swapped portrait photo provided in the embodiment of the present invention. This embodiment is optimized based on the above technical solutions. In this embodiment, optionally, generating a target portrait template image based on the portrait model and the style template image may include: processing the style template image based on the style face area to obtain a face template image and obtaining alignment data for aligning the face template image to the style template image; generating a target portrait template image based on the portrait model and the face template image; the face mask image corresponding to the style template image includes the face mask image of the face template image. Then, constructing the face mask image corresponding to the style template image and generating a target face-swapped portrait photo based on the target portrait template image and the face mask image includes: constructing the face mask image of the face template image, generating a pre-alignment face-swapped portrait photo based on the target portrait template image, the face template image, and the face mask image; and aligning the pre-alignment face-swapped portrait photo to the style template image based on the alignment data to obtain the target face-swapped portrait photo. Among them, the explanations of the same or corresponding terms as those in the above embodiments are not repeated here.
[0060] See Figure 6 , the method of this embodiment may specifically include the following steps:
[0061] S210. In response to a face-swapped portrait generation instruction, obtain a portrait image and a style template image, where the face-swapped portrait generation instruction is an instruction for indicating to replace the portrait face region in the portrait image with the style face region in the style template image to generate a target face-swapped portrait.
[0062] S230. Based on the style face region, process the style template image to obtain a face template image and obtain alignment data for aligning the face template image to the style template image.
[0063] S230. Process the style template image based on the style face region to obtain a face template image and obtain alignment data for aligning the face template image to the style template image.
[0064] Among them, processing the style template image based on the style face region to obtain a face template image that can highlight the style face region in the style template image. In practical applications, optionally, based on retaining the style face region, perform a cropping process on the style template image, and on this basis, obtain the face template image. Exemplarily, for the cropped template image obtained after the cropping process, the cropped template image can be directly used as the face template image; or the cropped template image can be processed again to obtain the face template image. For example, with the goal of placing the style face region in the middle region, fill the cropped template image to obtain the face template image, such as Figure 7 The style template image shown on the left and the face template image shown on the right, where the face template image can be a square image with a resolution of 512x512; etc., which are not specifically limited here.
[0065] The alignment data can be understood as data for aligning the face template image to the style template image. In practical applications, optionally, the alignment data can be represented by an affine transformation matrix. Obtain the alignment data, and the alignment data is applied in the following steps to align a generated face-swapped portrait to the face template image to obtain a target face-swapped portrait that matches the style template image.
[0066] Combined with the application scenarios that the embodiments of the present invention may involve, the cropping and alignment of the above style template image can utilize Modelscope; of course, it can also be implemented based on other solutions, which are not specifically limited here.
[0067] S240. Based on the portrait model and the face template image, generate a target portrait template image.
[0068] Among them, compared with directly generating a target portrait template image based on the style template image, generating a target portrait template image based on the face template image in this step can pay more attention to the face region during the generation process. Combining this with the portrait model helps to improve the similarity between the face region in the target portrait template image and the portrait face region.
[0069] S250. Construct a facial mask image for the facial template image, and generate a pre-alignment face-swapped portrait based on the target portrait template image, the facial template image, and the facial mask image.
[0070] Among them, construct a facial mask image that can be used to distinguish the facial area in the facial template image from the remaining areas, and then generate a pre-alignment face-swapped portrait based on the target portrait template image, the facial template image, and the facial mask image. Exemplarily, the area 1 in the facial template image that needs to be replaced can be determined based on the facial mask image, and then the area 2 in the target portrait template image corresponding to the area 1 is replaced onto the area 1. The area 1 is the facial area in the facial template image, and the area 2 is the facial area in the target portrait template image that is relatively similar to the portrait facial area, so as to ensure that only the facial area in the replaced facial template image changes and the remaining areas remain unchanged, thereby ensuring that the subsequent generated pre-alignment face-swapped portrait can take into account both the similarity of the portrait facial area and the style consistency.
[0071] S260. Based on the alignment data, align the pre-alignment face-swapped portrait to the style template image to obtain the target face-swapped portrait.
[0072] Among them, since the pre-alignment face-swapped portrait is generated based on the facial template image, that is, it matches the facial template image and does not match the style template image, the alignment data can be used to align (i.e., paste) the pre-alignment face-swapped portrait into the style template image, so as to obtain the target face-swapped portrait that matches the style template image.
[0073] The technical solution of the embodiment of the present invention processes the style template image based on the style facial area to obtain a facial template image that can highlight the style facial area, and then uses the facial template image and the portrait model to generate the target portrait template image, which helps to improve the similarity between the facial area in the target portrait template image and the portrait facial area, and further helps to improve the portrait facial similarity of the finally generated target face-swapped portrait.
[0074] Figure 8It is a flowchart of another method for generating a face-swapped portrait photo provided in an embodiment of the present invention. This embodiment is optimized based on the above technical solutions. In this embodiment, optionally, based on a portrait model and a face template image, generating a target portrait template image may include: generating an intermediate portrait template image based on the portrait model and the face template image; determining whether the similarity between the intermediate portrait template image and the target portrait represented by the portrait model meets a preset first iteration end condition; if not, regenerating the intermediate portrait template image based on the portrait model, the face template image, and the intermediate portrait template image, and returning to execute the step of determining whether the similarity between the intermediate portrait template image and the target portrait meets the preset iteration end condition; if it meets, using the intermediate portrait template image as the target portrait template image. Among them, the explanations of the same or corresponding terms as those in the above embodiments will not be repeated here.
[0075] See Figure 8 , the method of this embodiment may specifically include the following steps:
[0076] S310. In response to a face-swapped portrait photo generation instruction, obtain a portrait image and a style template image, where the face-swapped portrait photo generation instruction is an instruction for indicating to replace the portrait face area in the portrait image with the style face area in the style template image to generate a target face-swapped portrait photo.
[0077] S320. Perform portrait training based on the portrait image to obtain a portrait model.
[0078] S330. Process the style template image based on the style face area to obtain a face template image, and obtain alignment data for aligning the face template image to the style template image.
[0079] S340. Generate an intermediate portrait template image based on the portrait model and the face template image.
[0080] Among them, to ensure that the face area in the finally generated target portrait template image is highly similar to the portrait face area, the target portrait template image can be generated by an iterative method.
[0081] Specifically, in the first round of iteration, an intermediate portrait template image can be generated based on the portrait model and the face template image, and this intermediate portrait template image can be understood as an intermediate result in the iteration process.
[0082] It can be understood that this step is the first round of iteration process.
[0083] S350. For the target portrait represented by the portrait model, determine whether the similarity between the intermediate portrait template image and the target portrait meets a preset first iteration end condition.
[0084] Among them, the target portrait can be understood as the character image represented by the portrait model. The end condition of the first iteration can be understood as the condition preset to indicate that the iterative process of generating the target portrait template image can end. For example, it can be that the similarity between the intermediate portrait template image and the target portrait is greater than a preset first similarity threshold. Determine whether the similarity meets the end condition of the first iteration.
[0085] S360. If not satisfied, regenerate the intermediate portrait template image based on the portrait model, the face template image, and the intermediate portrait template image, and return to execute S350.
[0086] Among them, if not satisfied, this indicates that the similarity between the latest generated intermediate portrait template image and the target portrait is not high enough. Specifically, the similarity between the face region in this intermediate portrait template image and the face region of the portrait is not high enough. It is necessary to regenerate the intermediate portrait template image based on the portrait model, the face template image, and the intermediate portrait template image, and return to execute S350 to re-judge, that is, continue the iteration.
[0087] S370. If satisfied, use the intermediate portrait template image as the target portrait template image.
[0088] Among them, if satisfied, this indicates that the similarity between the latest generated intermediate portrait template image and the target portrait is high enough. Specifically, the similarity between the face region in this intermediate portrait template image and the face region of the portrait is high enough, then this intermediate portrait template image can be used as the target portrait template image.
[0089] Thus, the iterative process of generating the target portrait template image is completed.
[0090] S380. Construct a face mask image of the face template image, and generate a pre-alignment face-swapped portrait based on the target portrait template image, the face template image, and the face mask image.
[0091] S390. Align the pre-alignment face-swapped portrait to the style template image based on the alignment data to obtain the target face-swapped portrait.
[0092] The technical solution of the embodiment of the present invention generates the target face-swapped portrait in an iterative manner, and in each round of the iterative process, it determines whether the iteration ends based on the similarity between the generated intermediate face-swapped portrait and the target portrait, improving the similarity between the finally generated target face-swapped portrait and the face region of the portrait.
[0093] An optional technical solution for generating an intermediate portrait template image based on a portrait model and a face template image includes:
[0094] In the image-to-image mode, use a portrait model and a facial template image, and, in the ControlNet, use the facial template image to generate an intermediate portrait template image;
[0095] Based on the portrait model, the facial template image, and the intermediate portrait template image, regenerate the intermediate portrait template image, including:
[0096] In the image-to-image mode, use the portrait model and the intermediate portrait template image, and, in the ControlNet, use the intermediate portrait template image and the facial template image to regenerate the intermediate portrait template image.
[0097] Among them, Stable Diffusion not only supports images generated in the image-to-image (Img2Img) mode, but also supports various control methods such as ControlNet to customize the generated images. Therefore, in Stable Diffusion, the Img2Img mode and ControlNet can be used to generate an intermediate portrait template image.
[0098] Exemplarily, in the Img2Img mode, use a portrait model and a first initialization image, and in ControlNet, use the first initialization image and the facial template image, and the two cooperate to generate an intermediate portrait template image. Among them, in the first round of iteration, the first initialization image uses the facial template image, and in subsequent rounds of iteration, the first initialization image uses the intermediate portrait template image generated in the previous round of iteration. On this basis, optionally, parameters such as the denoising strength can be set in the Img2Img mode; parameters such as lineart_anime, Depth, and IPAdpaterPlus can be set in ControlNet. Among them, the input images of lineart_anime and Depth are the same as the first initialization image, and the input image of IPAdpaterPlus is the same as the facial template image. Exemplarily, Figure 9 Three intermediate portrait template images sequentially generated during the iteration are shown.
[0099] Figure 10It is a flowchart of yet another method for generating a face-swapped portrait photo provided in an embodiment of the present invention. This embodiment is optimized based on the above technical solutions. In this embodiment, optionally, based on the target portrait template image, the face template image, and the face mask image, generating a pre-alignment face-swapped portrait photo may include: generating an intermediate face-swapped portrait photo based on the target portrait template image, the face template image, and the face mask image; determining whether the similarity between the intermediate face-swapped portrait photo and the target portrait represented by the portrait model meets a preset second iteration end condition; if not, regenerating the intermediate face-swapped portrait photo based on the intermediate face-swapped portrait photo, the target portrait template image, the face template image, and the face mask image, and returning to execute the step of determining whether the similarity between the intermediate face-swapped portrait photo and the target portrait meets the preset second iteration end condition; if so, using the intermediate face-swapped portrait photo as the pre-alignment face-swapped portrait photo. Among them, the explanations of the same or corresponding terms as those in the above embodiments will not be elaborated here.
[0100] See Figure 10 , the method of this embodiment may specifically include the following steps:
[0101] S410. In response to a face-swapped portrait photo generation instruction, obtain a portrait image and a style template image, where the face-swapped portrait photo generation instruction is an instruction for indicating to replace the portrait face area in the portrait image with the style face area in the style template image to generate a target face-swapped portrait photo.
[0102] S420. Perform portrait training based on the portrait image to obtain a portrait model.
[0103] S430. Process the style template image based on the style face area to obtain a face template image, and obtain alignment data for aligning the face template image to the style template image.
[0104] S440. Generate a target portrait template image based on the portrait model and the face template image.
[0105] S450. Construct a face mask image of the face template image, and generate an intermediate face-swapped portrait photo based on the target portrait template image, the face template image, and the face mask image.
[0106] Among them, to ensure that the pre-alignment face-swapped portrait photo generated subsequently is highly similar to the target portrait, the pre-alignment face-swapped portrait photo can be generated by an iterative method.
[0107] Specifically, in the first-round iteration process, an intermediate face-swapped portrait photo can be generated based on the target portrait template image, the face template image, and the face mask image, and this intermediate face-swapped portrait photo can be understood as an intermediate result in the iteration process. It can be understood that this step is the first-round iteration process.
[0108] In practical applications, in addition to using the above three images, a portrait model can also be combined to jointly generate an intermediate face-swapped portrait photo, thereby improving the similarity between the intermediate face-swapped portrait photo and the target portrait.
[0109] S460. For the target portrait represented by the portrait model, determine whether the similarity between the intermediate face-swapped portrait photo and the target portrait meets a preset second iteration end condition.
[0110] Among them, the second iteration end condition can be understood as a condition for ending the iterative process of generating the face-swapped portrait photo before alignment, for example, it can be that the similarity between the intermediate face-swapped portrait photo and the target portrait is greater than a preset second similarity threshold. Determine whether the similarity meets the second iteration end condition.
[0111] S470. If not, based on the intermediate face-swapped portrait photo, the target portrait template image, the face template image, and the face mask image, regenerate the intermediate face-swapped portrait photo, and return to execute S460.
[0112] Among them, if not, this indicates that the similarity between the latest generated intermediate face-swapped portrait photo and the target portrait is not high enough. It is necessary to regenerate the intermediate face-swapped portrait photo based on the intermediate face-swapped portrait photo, the target portrait template image, the face template image, and the face mask image, and return to execute S460 for continuous iteration.
[0113] In practical applications, in addition to using the above four images, a portrait model can also be combined to jointly regenerate the intermediate face-swapped portrait photo, thereby improving the similarity between the intermediate face-swapped portrait photo and the target portrait.
[0114] S480. If it is satisfied, use the intermediate face-swapped portrait photo as the face-swapped portrait photo before alignment.
[0115] Among them, if it is satisfied, it means that the similarity between the latest generated intermediate face-swapped portrait photo and the target portrait is high enough, and then the intermediate face-swapped portrait photo can be used as the face-swapped portrait photo before alignment.
[0116] Thus, the iterative process of generating the face-swapped portrait photo before alignment is completed.
[0117] S490. Based on the alignment data, align the face-swapped portrait photo before alignment to the style template image to obtain the target face-swapped portrait photo.
[0118] The technical solution of the embodiment of the present invention generates the face-swapped portrait photo before alignment in an iterative manner, and in each round of the iterative process, it determines whether the iteration ends based on the similarity between the generated intermediate face-swapped portrait photo and the target portrait, which helps to improve the similarity between the face-swapped portrait photo before alignment generated at the end of the iteration and the target portrait.
[0119] An optional technical solution generates an intermediate face-swap portrait photo based on a target portrait template image, a face template image, and a face mask image, including:
[0120] In the image generation mode, the face template image and the face mask image are used, and in the control network, the target portrait template image and the face template image are used to generate an intermediate face-swapped portrait photo;
[0121] Accordingly, based on the intermediate face-changing portrait photo, the target portrait template image, the face template image and the face mask image, the intermediate face-changing portrait photo is regenerated, including:
[0122] In the image-generated-image mode, the face mask image and the intermediate face-swapped portrait photo are used, and in the control network, the target portrait template image and the face template image are used to regenerate the intermediate face-swapped portrait photo.
[0123] Among them, in Stable Diffusion, the Img2Img mode and ControlNet can be used to generate an intermediate face-changing portrait photo. Exemplarily, in the Img2Img mode, the portrait model, the face mask image and the second initialization image can be used, and in ControlNet, the target portrait template image and the face template image can be used, and the two are coordinated to generate an intermediate face-changing portrait photo. Among them, in the first round of iteration, the second initialization image uses the face template image, and in the subsequent rounds of iteration, the second initialization image uses the intermediate face-changing portrait photo generated by the previous round of iteration. On this basis, optionally, parameters such as denoising strength can be set in the Img2Img mode; parameters such as lineart_anime, SoftEdge, IPAdpaterPlusFaceV2 and IPAdpaterPlus can be set in ControlNet, wherein the input images of lineart_anime, SoftEdge and IPAdpaterPlusFaceV2 are the same as the target portrait template image, and the input image of IPAdpaterPlus is the same as the face template image. Exemplarily, Figure 11 Two intermediate face-swapped portraits generated sequentially during the iterative process are shown.
[0124] In order to better understand the above-mentioned technical solutions as a whole, the following is an exemplary description of them in combination with specific examples. Figure 12 The process of generating anime face-changing portraits is shown below:
[0125] Step 1: Perform portrait training based on multiple portrait images uploaded by users to obtain a portrait model.
[0126] Step 2: Obtain an anime-style template image, and perform face cropping on the anime-style template image to obtain a face template image. This face template image can be a square image containing a front face (i.e., the template front-face square image). For example, Figure 7 the shown anime-style template image and face template image are presented; then, obtain an affine transformation matrix that can align the face template image to the anime-style template image.
[0127] Step 3: Generate a target portrait template image with high similarity (i.e., the real-person template image). Specifically, adopt the StableDiffusion Img2Img mode, combine with the portrait model trained in Step 1, set the denoising strength to 0.4 - 0.6, and loop multiple times until the similarity > 0.6 to stop. Among them, for the first initialization image, use the template front-face square image output in Step 2 for the first time, and use the real-person template image output in the previous round of iteration for subsequent ones. Configure ControlNet with lineart_anime (weight 0.4, input image is the same as the first initialization image), Depth (weight 0.5, input image is the same as the first initialization image), and IPAdpaterPlus (weight 0.3, input image is the template front-face square image output in Step 2). Figure 9 The real-person template images generated during three iterative processes are shown.
[0128] Step 4: Construct a face mask image for the template front-face square image output in Step 2. Refer to Figure 4 where the white area represents the area that needs to be replaced and the black area represents the area that does not need to be replaced.
[0129] Step 5: Generate a pre-alignment face-swapped portrait (i.e., the face-swapped result image). Adopt the Stable Diffusion Img2Img mode, combine with the portrait model output in Step 1, set the denoising strength to 0.4 - 0.6, use the face mask image output in Step 4 as the mask, and loop multiple times until the similarity > 0.5 to stop. Among them, for the second initialization image, use the template front-face square image output in Step 2 for the first time, and use the face-swapped result image output in the previous round of iteration for subsequent ones. Configure ControlNet with lineart_anime (weight 0.75, input image is the real-person template image output in Step 3), SoftEdge (weight 0.75, input image is the real-person template image output in Step 3), IPAdpaterPlusFaceV2 (weight 1.0, input image is the real-person template image output in Step 3), and IPAdpaterPlus (weight 0.5, input image is the template front-face square image output in Step 2). Figure 11 The face-swapped result images generated during two iterative processes are shown.
[0130] Step 6: Use the affine transformation matrix output in Step 2 to paste the face-swapped result image back onto the anime-style template image to obtain the anime face-swapped portrait photo (i.e., the final result image), as Figure 5 shown.
[0131] The above example provides a generation process for anime face-swapped portrait photos based on StableDiffusion, which can take into account both the similarity of the human face and the consistency of the anime style.
[0132] Figure 13 This is the structural block diagram of the face-swapped portrait photo generation device provided in the embodiments of the present invention. The device is used to execute the face-swapped portrait photo generation method provided in any of the above embodiments. The device and the face-swapped portrait photo generation methods in the above embodiments belong to the same inventive concept. For the details not described in detail in the embodiments of the face-swapped portrait photo generation device, reference can be made to the embodiments of the face-swapped portrait photo generation method. Refer to Figure 13 and the device may specifically include: an image acquisition module 510, a target portrait template image generation module 520, and a target face-swapped portrait photo generation module 530. Among them,
[0133] The image acquisition module 510 is configured to acquire a portrait image and a style template image in response to a face-swapped portrait photo generation instruction, where the face-swapped portrait photo generation instruction is an instruction for indicating to replace the portrait face area in the portrait image with the style face area in the style template image to generate a target face-swapped portrait photo;
[0134] The target portrait template image generation module 520 is configured to perform portrait training based on the portrait image to obtain a portrait model, and generate a target portrait template image based on the portrait model and the style template image;
[0135] The target face-swapped portrait photo generation module 530 is configured to construct a face mask image corresponding to the style template image, and generate a target face-swapped portrait photo based on the target portrait template image and the face mask image.
[0136] Optionally, the target portrait template image generation module 520 may include:
[0137] The sub-module for obtaining alignment data is configured to process the style template image based on the style face area to obtain a face template image, and obtain alignment data for aligning the face template image to the style template image;
[0138] The target portrait template image generation sub-module is configured to generate a target portrait template image based on the portrait model and the face template image;
[0139] If the face mask image corresponding to the style template image includes the face mask image of the face template image, the target face-swapped portrait photo generation module 530 may include:
[0140] The pre - alignment face - swapped portrait generation sub - module is used to construct a face mask image of the face template image, and generate a pre - alignment face - swapped portrait based on the target portrait template image, the face template image, and the face mask image;
[0141] The target face - swapped portrait obtaining sub - module is used to align the pre - alignment face - swapped portrait into the style template image based on the alignment data to obtain the target face - swapped portrait.
[0142] On this basis, optionally, the target portrait template image generation sub - module may include:
[0143] The intermediate portrait template image generation unit is used to generate an intermediate portrait template image based on the portrait model and the face template image;
[0144] The first similarity judgment unit is used to judge whether the similarity between the intermediate portrait template image and the target portrait represented by the portrait model meets a preset first iteration end condition;
[0145] The first return and execution unit is used to, if not met, regenerate the intermediate portrait template image based on the portrait model, the face template image, and the intermediate portrait template image, and return to execute the step of judging whether the similarity between the intermediate portrait template image and the target portrait meets the preset iteration end condition;
[0146] The target portrait template image obtaining unit is used to, if met, use the intermediate portrait template image as the target portrait template image.
[0147] On this basis, optionally, the intermediate portrait template image generation unit may include:
[0148] The intermediate portrait template image generation subunit is used to generate an intermediate portrait template image in the image - to - image mode using the portrait model and the face template image, and using the face template image in the control network;
[0149] The first return and execution unit may include:
[0150] The intermediate portrait template image regeneration subunit is used to regenerate the intermediate portrait template image in the image - to - image mode using the portrait model and the intermediate portrait template image, and using the intermediate portrait template image and the face template image in the control network.
[0151] Another optionally, the pre - alignment face - swapped portrait generation sub - module may include:
[0152] The intermediate face - swapped portrait generation unit is used to generate an intermediate face - swapped portrait based on the target portrait template image, the face template image, and the face mask image;
[0153] A second similarity judgment unit, which is used to judge whether the similarity between the intermediate face-swapped portrait photo and the target portrait meets a preset second iteration end condition for the target portrait represented by the portrait model;
[0154] A second return and execution unit, which is used to, if not satisfied, regenerate the intermediate face-swapped portrait photo based on the intermediate face-swapped portrait photo, the target portrait template image, the face template image, and the face mask image, and return to execute the step of judging whether the similarity between the intermediate face-swapped portrait photo and the target portrait meets a preset second iteration end condition;
[0155] A unit for obtaining the pre-alignment face-swapped portrait photo, which is used to, if satisfied, use the intermediate face-swapped portrait photo as the pre-alignment face-swapped portrait photo.
[0156] On this basis, optionally, the intermediate face-swapped portrait photo generation unit may include:
[0157] An intermediate face-swapped portrait photo generation subunit, which is used to generate an intermediate face-swapped portrait photo in the image-to-image mode by using the face template image and the face mask image, and, in the control network, by using the target portrait template image and the face template image;
[0158] The second return and execution unit may include:
[0159] An intermediate face-swapped portrait photo regeneration subunit, which can be used to regenerate the intermediate face-swapped portrait photo in the image-to-image mode by using the face mask image and the intermediate face-swapped portrait photo, and, in the control network, by using the target portrait template image and the face template image.
[0160] Another optional alignment data obtaining sub-module may include:
[0161] A face template image obtaining unit, which can be used to crop the style template image based on the retained style face area to obtain the face template image.
[0162] On the basis of any of the above devices, optionally, the style template image includes an anime style template image.
[0163] The face-swapped portrait generation device provided by an embodiment of the present invention, through an image acquisition module, in response to a face-swapped portrait generation instruction, acquires a portrait image and a style template image; then, through a target portrait template image generation module, conducts portrait training based on the portrait image to obtain a portrait model, and further generates a target portrait template image based on the portrait model and the style template image; furthermore, through a target face-swapped portrait generation module, constructs a facial mask image corresponding to the style template image, and then generates a target face-swapped portrait based on the target portrait template image and the facial mask image. The above device, through the facial mask image, can replace the facial area in the target portrait template image into the style template image, realizing the effective generation of a target face-swapped portrait that takes into account both the similarity of the human face and the consistency of the style.
[0164] The face-swapped portrait generation device provided by an embodiment of the present invention can execute the face-swapped portrait generation method provided by any embodiment of the present invention, and has functional modules and beneficial effects corresponding to the execution of the method.
[0165] It should be noted that in the embodiments of the above face-swapped portrait generation device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.
[0166] Figure 14 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0167] As Figure 14As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0168] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0169] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the method for generating a face-swapped portrait photo.
[0170] In some embodiments, the method for generating a face-swapped portrait photo can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the method for generating a face-swapped portrait photo described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the method for generating a face-swapped portrait photo in any other suitable manner (e.g., by means of firmware).
[0171] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0172] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0173] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0174] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0175] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0176] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0177] The various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0178] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for generating a face-changing portrait, characterized in that: include: In response to a face-swapped portrait photo generation instruction, a portrait image and a style template image are acquired, wherein the face-swapped portrait photo generation instruction is an instruction for instructing to replace a portrait face region in the portrait image with a style face region in the style template image to generate a target face-swapped portrait photo; Performing portrait training based on the portrait image to obtain a portrait model, and generating a target portrait template image based on the portrait model and the style template image, wherein the target portrait template image is more similar to the portrait face area than the style template image; and the style of the target portrait template image is inconsistent with the style template image; A facial mask image corresponding to the style template image is constructed, and based on the target portrait template image and the facial mask image, the target face-swapped portrait photo is generated, wherein the facial mask image represents that the remaining areas related to the style template image and unrelated to the face do not need to be replaced.
2. The method according to claim 1, characterized in that The step of generating a target portrait template image based on the portrait model and the style template image includes: Based on the style face region, the style template image is processed to obtain a face template image, and alignment data for aligning the face template image to the style template image is obtained; Based on the portrait model and the face template image, generating a target portrait template image; The facial mask image corresponding to the style template image includes the facial mask image of the facial template image, and the facial mask image corresponding to the style template image is constructed, and the target face-swapped portrait is generated based on the target portrait template image and the facial mask image, including: Constructing a face mask image of the face template image, and generating a face-swapped portrait photo before alignment based on the target portrait template image, the face template image and the face mask image; Based on the alignment data, the pre-alignment face-swapped portrait photo is aligned to the style template image to obtain the target face-swapped portrait photo.
3. The method according to claim 2, characterized in that The step of generating a target portrait template image based on the portrait model and the face template image comprises: Based on the portrait model and the face template image, generating an intermediate portrait template image; For the target portrait represented by the portrait model, determining whether the similarity between the intermediate portrait template image and the target portrait satisfies a preset first iteration end condition; If not, regenerate the intermediate portrait template image based on the portrait model, the face template image and the intermediate portrait template image, and return to the step of determining whether the similarity between the intermediate portrait template image and the target portrait satisfies the preset iteration end condition; If the conditions are met, the intermediate portrait template image is used as the target portrait template image.
4. The method according to claim 3, characterized in that: The step of generating an intermediate portrait template image based on the portrait model and the face template image comprises: In the image-to-image mode, using the portrait model and the face template image, and, in the control network, using the face template image, generating an intermediate portrait template image; The regenerating the intermediate portrait template image based on the portrait model, the face template image and the intermediate portrait template image comprises: In the image-to-image mode, the intermediate portrait template image is regenerated using the portrait model and the intermediate portrait template image, and in the control network, the intermediate portrait template image and the face template image are used.
5. The method according to claim 2, characterized in that: The step of generating the face-swapped portrait photo before alignment based on the target portrait template image, the facial template image and the facial mask image comprises: Generate an intermediate face-swapped portrait photo based on the target portrait template image, the facial template image and the facial mask image; For the target portrait represented by the portrait model, determining whether the similarity between the intermediate face-swapped portrait and the target portrait satisfies a preset second iteration end condition; If not, regenerate the intermediate face-swapped portrait based on the intermediate face-swapped portrait, the target portrait template image, the facial template image and the facial mask image, and return to the step of determining whether the similarity between the intermediate face-swapped portrait and the target portrait satisfies the preset second iteration end condition; If the conditions are met, the intermediate face-changing portrait photo is used as the face-changing portrait photo before alignment.
6. The method according to claim 5, characterized in that The step of generating an intermediate face-swapped portrait photo based on the target portrait template image, the facial template image, and the facial mask image comprises: In the image-generated-image mode, using the facial template image and the facial mask image, and, in the control network, using the target portrait template image and the facial template image, generating an intermediate face-swapped portrait photo; Accordingly, the regenerating the intermediate face-swapped portrait photo based on the intermediate face-swapped portrait photo, the target portrait template image, the facial template image and the facial mask image includes: In the image-generating mode, the intermediate face-changing portrait photo is regenerated using the facial mask image and the intermediate face-changing portrait photo, and in the control network, the target portrait template image and the facial template image are used.
7. The method according to claim 2, characterized in that The step of processing the style template image based on the style face region to obtain the face template image includes: Based on retaining the style face region, the style template image is cropped to obtain a face template image.
8. The method according to any one of claims 1 to 7, characterized in that: The style template image includes an anime style template image.
9. A face-changing portrait generation device, characterized in that: include: an image acquisition module, configured to acquire a portrait image and a style template image in response to a face-swapped portrait photo generation instruction, wherein the face-swapped portrait photo generation instruction is an instruction for instructing to replace a portrait face region in the portrait image with a style face region in the style template image to generate a target face-swapped portrait photo; a target portrait template image generation module, configured to perform portrait training based on the portrait image to obtain a portrait model, and to generate a target portrait template image based on the portrait model and the style template image, wherein the target portrait template image is more similar to the portrait face area than the style template image; and the style of the target portrait template image is inconsistent with the style template image; The target face-swapped portrait generation module is used to construct a facial mask image corresponding to the style template image, and generate the target face-swapped portrait based on the target portrait template image and the facial mask image, wherein the facial mask image represents the remaining areas related to the style template image and unrelated to the face and do not need to be replaced.
10. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor executes the face-changing portrait generation method as described in any one of claims 1-8.
Citation Information
Patent Citations
Image processing method and device, storage medium and electronic equipment
CN114913061A
Image generation method and device, electronic equipment and medium
CN117764812A