Image face-changing method and system based on personalized generation technology
By generating virtual tags and face detection, combining diffusion model and random differential editing method, two personalized image generation is performed, which solves the problem of low accuracy when image face change, and achieves high-quality face change image generation.
Patent Information
- Application Number
- CN202411857976.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-12-17
AI Technical Summary
The existing personalized image generation technology has the problem of low accuracy when changing faces of images, and it is difficult to generate images that are highly consistent with the target person.
By obtaining the target face image to generate virtual tags, face detection is performed on the target editing image to determine the editing area, and two personalized image generations are performed by combining the diffusion model and the random differential editing method to ensure the reasonable splicing of face features and backgrounds.
It significantly improves the naturalness and reality of the face-changing image, ensuring that the facial features and expressions of the generated image are highly consistent with the original image, and the background rationality and face similarity achieve ideal results.
Smart Images

Figure CN119323512B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of personalized image generation technology, and more specifically, to an image face-changing method and system based on personalized generation technology. Background Art
[0002] In recent years, the Stable Diffusion generation model has made significant progress in the field of image personalization. Personalized generation refers to providing a small number of specific concept images to the diffusion model so that it can learn and master this new concept. After such training, the model will be able to generate new images related to the concept based on text prompts, including images of different scenes and styles. Among them, the obvious advantage of Textual Inversion over other text-image generation is that only 3-5 images are needed to accurately learn the "semantic" nature of the concept provided by the user.
[0003] At present, text-to-image diffusion models have made significant progress in creating high-fidelity images, and their generation capabilities are very powerful. However, when only text prompts are used to generate the required images, there are problems with low efficiency and accuracy of personalized image generation due to the complex prompt engineering involved. To improve the efficiency of personalized image generation, it is currently mainly fine-tuned directly from the pre-trained model, but it requires a lot of computing resources and is incompatible with other basic models, text prompts, and structural control. At present, a lightweight adapter IP-Adapter for implementing the image prompt function of the pre-trained text-to-image diffusion model has been proposed, which chooses to use image prompts instead of text prompts. Among them, IP-Adapter adopts a decoupled cross-attention mechanism for text features and image features. For each cross-attention layer in the U-Net of the diffusion model, it only adds an additional cross-attention layer for image features. During the training phase, only the parameters of the new cross-attention layer are trained, and the original U-Net model remains frozen. Although IP-Adapter is flexible, effective and lightweight, it can only generate images that are similar to the reference image in content and style. It cannot synthesize images that are highly consistent with the subject of a given image like existing methods such as text inversion and DreamBooth. When used for face swapping, the generated face may not be very similar to the target person provided, that is, the image accuracy is not ideal; and when using IP-Adapter to edit images, the generated images may not necessarily meet the requirements of text prompts, that is, the editability is not ideal. Summary of the invention
[0004] In order to overcome the defect of low precision when the above-mentioned personalized image generation technology is applied to image face swapping, the present invention provides an image face swapping method and system based on personalized generation technology.
[0005] In order to solve the above technical problems, the technical solution of the present invention is as follows:
[0006] An image face-changing method based on personalized generation technology comprises the following steps:
[0007] Obtain the target face image and generate the corresponding virtual label;
[0008] Obtaining a marked editing area for the target editing image based on a face detection algorithm;
[0009] The target edited image with the marked edited area is used as an image input, and the virtual label corresponding to the target face is used as a text input, and a first face-changing image is generated based on a random difference editing method in a diffusion model;
[0010] The non-marked editing area in the target editing image is spliced with the marked editing area in the first face-changing image as image input, the virtual label is used as text input, and a second face-changing image is generated and output based on a random difference editing method in a diffusion model.
[0011] Furthermore, the present invention also proposes an image face-swapping system based on personalized generation technology, and applies the image face-swapping method proposed in the present invention. The system comprises:
[0012] A text processing module is used to obtain the target face image and generate the corresponding virtual label;
[0013] A region processing module, used for obtaining a marked editing region for a target editing image based on a face detection algorithm;
[0014] A first face-changing module, on which a diffusion model is mounted, is used to input the target edited image with the marked edited area as an image, input the virtual label corresponding to the target face as a text, and generate a first face-changing image based on a random difference editing method in the diffusion model;
[0015] An image stitching module, used for stitching the unmarked edited area in the target edited image with the marked edited area in the first face-swapped image to obtain a stitched image;
[0016] The second face-changing module is equipped with a diffusion model, which is used to take the spliced image as an image input and the virtual label as a text input, generate and output a second face-changing image based on a random difference editing method in the diffusion model.
[0017] Furthermore, the present invention also proposes a device, including a memory and a processor, wherein the memory stores computer-readable instructions, wherein when the computer-readable instructions are executed by the processor, the processor executes all or part of the steps of the image face-changing method based on personalized generation technology as described in the present invention.
[0018] Furthermore, the present invention also proposes a storage medium on which computer-readable instructions are stored, wherein the computer-readable instructions, when executed by a processor, implement all or part of the steps of the image face-changing method based on personalized generation technology as described in the present invention.
[0019] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:
[0020] The present invention generates corresponding labels for the target face image, determines the target editing area for the target editing image based on the face detection algorithm, and realizes high-precision face replacement in conjunction with the diffusion model, so that the replaced face is highly consistent with the original image in facial features and expressions, greatly improving the naturalness and realism of the face-changing image.
[0021] The present invention performs secondary personalized image generation, and in the second stage, splices the non-marked editing area in the target editing image with the marked editing area of the first face-changing image as image input, and inputs the virtual label as text, and performs secondary personalized image generation based on the random difference editing method in the diffusion model, so as to obtain high-quality face-changing pictures with more ideal background rationality and face similarity. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 The present invention is a flowchart of an image face-swapping method based on personalized generation technology according to an embodiment of the present invention.
[0023] Figure 2 is a face-changing image shown according to an embodiment of the present invention; wherein, Figure 2 Part (a) is the target edited image. Figure 2 Part (b) is a mask image marking the editing area. Figure 2 Part (c) is the first face-changing image generated in the first stage. Figure 2 Part (d) is the second face-swapped image generated in the second stage.
[0024] Figure 3 FIG. 1 is a schematic diagram of a mark editing area before and after an erosion operation according to an embodiment of the present invention; wherein: Figure 3 Part (a) is a schematic diagram of the marked edited area before the corrosion operation. Figure 3 Part (b) is a schematic diagram of the marked edited area after the corrosion operation.
[0025] Figure 4 The figure is a comparison chart of face-changing results using the present invention and IP-Adapter.
[0026] Figure 5 Another comparison diagram of face-changing results using the present invention and IP-Adapter is shown below.
[0027] Figure 6 Comparison of face-swapping results with sad expression description text added.
[0028] Figure 7 Comparison chart of face swapping results with added description text for closed-eye features.
[0029] Figure 8 The present invention is an architectural diagram of an image face-swapping system based on personalized generation technology according to one embodiment of the present invention. DETAILED DESCRIPTION
[0030] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the accompanying drawings. It should be emphasized that the faces appearing in the accompanying drawings are all virtual cartoon characters, and do not involve real human faces.
[0031] The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0032] The terms used in the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the" and "the" used in the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0033] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0034] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0035] Example 1
[0036] This embodiment proposes an image face-changing method based on personalized generation technology, such as Figure 1 , which is a flow chart of the image face-changing method based on the personalized generation technology of this embodiment.
[0037] The image face-swapping method based on the personalized generation technology proposed in this embodiment includes the following steps:
[0038] S1, obtain the target face image and generate the corresponding virtual label;
[0039] S2, obtaining a marked editing area for the target editing image based on a face detection algorithm;
[0040] S3, using the target edited image with the marked edited area as an image input, using the virtual label corresponding to the target face as a text input, and generating a first face-changing image based on a random difference editing method in a diffusion model;
[0041] S4. Splice the unmarked editing area in the target editing image with the marked editing area in the first face-changing image as image input, input the virtual label as text, generate and output a second face-changing image based on a random difference editing method in a diffusion model.
[0042] This embodiment first generates a corresponding label for the target facial image, and determines the target editing area of the target editing image based on the face detection algorithm, and cooperates with the diffusion model to achieve high-precision face replacement, so that the replaced face is highly consistent with the original image in facial features and expressions, greatly improving the naturalness and realism of the face-changing image.
[0043] For example, Figure 2 As shown, Figure 2 Part (a) is the target edited image. Figure 2 Part (b) is a mask image marking the editing area. Figure 2 Part (c) is the first face-changing image generated in the first stage. Figure 2 Part (d) is the second face-swapped image generated in the second stage.
[0044] In addition, in the second stage, this embodiment splices the non-marked editing area in the target editing image with the marked editing area of the first face-changing image as image input, and uses the virtual label as text input, and performs secondary personalized image generation based on the random difference editing method in the diffusion model, so as to obtain high-quality face-changing pictures with more ideal background rationality and face similarity.
[0045] Among them, Stochastic Differential Editing (SDEdit) is an image synthesis and editing method based on diffusion model generation priors. Its core idea is to use the diffusion model to inject Gaussian noise into the input data and convert it into a known prior distribution, and then gradually restore high-quality images through iterative denoising of stochastic differential equations (SDE), which can naturally achieve a balance between realism and fidelity. This method is not only suitable for image synthesis, but also can be used for various editing tasks, such as stroke-based image generation.
[0046] In an optional embodiment, the virtual label includes a surname word-gram and a name word-gram obtained by mapping the target face image through cross initialization.
[0047] In this embodiment, the target face image is mapped into two tokens based on the cross-initialization generation technology, and the first name token and the last name token are obtained respectively and used as unique virtual name tags of the target face, which are used as prompt text input of the diffusion model to guide the generation of personalized images.
[0048] Exemplarily, when a face-changing image is generated based on the random differential editing method, the text input prompt is "a photo of {first name} {last name} person", where {first name} and {last name} are the surname word and the name word.
[0049] The cross-initialization generation technology selected in this embodiment is a technology used for personalized generation. It uses the output of the text encoder to initialize text embedding, thereby effectively solving the overfitting problem in textual inversion.
[0050] In an optional embodiment, in the generation of the first face-changing image and the second face-changing image based on the random difference editing method in the diffusion model, for the diffusion model with a total number of iterations of T, it is configured to perform denoising starting from the 2 / T step in the inverse diffusion process.
[0051] Since in the diffusion model, the inverse diffusion process of generating a picture from Gaussian noise is a process from t=T to t=0, during which the model gradually learns to generate high-quality images from random noise through a large number of iterations and optimizations. This embodiment is based on the random difference editing method, and chooses to generate the content of the target face in the target editing area from the intermediate step, which can maintain the structural information of the input image, that is, maintain the structural information of the background area, so as to improve the accuracy of the face-changing image.
[0052] Further optionally, for a diffusion model with a total number of iterations T=1000, it is configured to perform denoising starting from the number of iterations t=400-600 in the inverse diffusion process.
[0053] Exemplarily, for a diffusion model with a total number of iterations T=1000, it is configured to perform denoising starting from the number of iterations t=500 in the inverse diffusion process.
[0054] In an optional embodiment, in step S2, when splicing the unmarked edited area in the target edited image with the marked edited area of the first face-changing image, the following steps are included:
[0055] Perform an erosion operation on the edge of the marked editing area;
[0056] The target edited image is stripped of the corroded marked edited area, and is concatenated with the corroded marked edited area in the first face-changing image to obtain a second-stage input image.
[0057] In this embodiment, in order to improve the accuracy of face-changing, a second stage of personalized image generation is selected, and the first face-changing image generated in the first stage is used as the image input again. At this time, during the diffusion process, the background content is reconstructed twice, and noise is introduced twice, so there is a serious blurring problem. In this regard, this embodiment splices the non-marked editing area in the target editing image with the marked editing area of the first face-changing image as the image input for the second stage. However, there will be obvious edge traces at the connection of the spliced images. In order to effectively eliminate the obvious splicing traces of the spliced images, this embodiment performs an erosion operation on the edge of the marked editing area in the second stage to expand the marked editing area.
[0058] For example, Figure 3 As shown, Figure 3 The white area in (a) is the marked editing area in the first stage. Figure 3 The gray and white areas in (b) are the marked edited areas that have been eroded in the second stage.
[0059] Further optionally, when the edge of the mark editing area is eroded, the range is expanded to 10% to 20%.
[0060] Furthermore, the target edited image is stripped of the corroded marked edited area, and is spliced with the corroded marked edited area in the first face-changing image as the image input, and the virtual label is input as the text, and the second face-changing image is generated based on the random difference editing method in the diffusion model. Among them, since the splicing position of the target face and the background image is located in the corroded area, the splicing position is also processed by the personalized generation process in the second stage of personalized image generation, so that the traces of the splicing can be weakened, effectively improving the accuracy and authenticity of the generated image.
[0061] Further, in an optional embodiment, when the second face-changing image is generated based on the random difference editing method in the diffusion model, the text input of the diffusion model also includes facial expression, posture and / or feature description text input by the user.
[0062] In this embodiment, the text describing the facial expression, posture and / or features input by the user is processed based on the cross-initialization technology, and is combined with the virtual label as the text input of the diffusion model to achieve the editability of the personalized image.
[0063] For example, when only the face-changing operation is performed, the text input prompt in the second stage is "a photo of {first name} {last name} person"; when the user enters the facial expression description text "sad expression", the text input prompt in the second stage is "a photo of {first name} {last name} personwith sad expression".
[0064] The diffusion model can accurately control various parameters in the face-changing process based on the text input instructions, so that the generated image is more in line with the user's intention and the fidelity of the edited image is greatly improved.
[0065] As an exemplary explanation, this embodiment uses the above-mentioned image face-changing method based on personalized generation technology to compare with IP-Adapter, and performs face-changing on two different input images respectively, such as Figure 4 , 5 As shown, the face-changing image generated by using IP-Adapter is not as ideal in terms of face similarity and authenticity as that of the present invention.
[0066] Further, in Figure 4 Based on the above, we add the expression description text "a photo of a {} person with sad expression" to get the following Figure 6The face-changing image with a sad expression added is shown. The expression in the image generated by the present invention changes accordingly, while the face-changing image generated by the IP-Adapter does not show a sad expression according to the text prompt, and the editability is not strong.
[0067] Further, in Figure 5 Based on the above, we add the feature description text "a photo of a {} person closing eyes" to get Figure 7 The face-swapped image with the eye-closing feature added is shown. In the image generated by the present invention, the eye-closing representation is accurately generated at the target face position, while the face-swapped image generated by the IP-Adapter does not completely generate the eye-closing feature according to the text prompt, and has the defect of low image authenticity.
[0068] It can be seen that the image face-changing method proposed in the present invention has ideal editability. The picture prompts can be converted into text prompts through cross-initialization technology, and used together with the text input by the user as the text prompts of the random differential editing SDEdit method, so as to accurately control various parameters in the face-changing process and greatly improve the accuracy and authenticity of the edited image.
[0069] Example 2
[0070] This embodiment applies the image face-changing method based on personalized generation technology proposed in Embodiment 1, and proposes an image face-changing system based on personalized generation technology. Figure 8 , which is an architecture diagram of the image face-changing system based on personalized generation technology of this embodiment.
[0071] The image face-changing system based on the personalized generation technology proposed in this embodiment includes:
[0072] A text processing module is used to obtain the target face image and generate the corresponding virtual label;
[0073] A region processing module, used for obtaining a marked editing region for a target editing image based on a face detection algorithm;
[0074] A first face-changing module, on which a diffusion model is mounted, is used to input the target edited image with the marked edited area as an image, input the virtual label corresponding to the target face as a text, and generate a first face-changing image based on a random difference editing method in the diffusion model;
[0075] An image stitching module, used for stitching the unmarked edited area in the target edited image with the marked edited area in the first face-swapped image to obtain a stitched image;
[0076] The second face-changing module is equipped with a diffusion model, which is used to take the spliced image as an image input and the virtual label as a text input, generate and output a second face-changing image based on a random difference editing method in the diffusion model.
[0077] It can be understood that the system of this embodiment corresponds to the method of the above-mentioned embodiment 1, and the options in the above-mentioned embodiment 1 are also applicable to this embodiment, so they are not described repeatedly here.
[0078] Example 3
[0079] This embodiment proposes a computer device, including a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes all or part of the steps of the image face-changing method based on personalized generation technology proposed in Example 1.
[0080] Example 4
[0081] This embodiment proposes a storage medium on which computer-readable instructions are stored, wherein the computer-readable instructions, when executed by a processor, implement all or part of the steps of the image face-changing method based on personalized generation technology proposed in Embodiment 1.
[0082] Exemplarily, the storage medium includes, but is not limited to, a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0083] Exemplarily, the instructions, programs, code sets or instruction sets may be implemented using conventional programming languages.
[0084] Exemplarily, the processor includes but is not limited to a smart phone, a personal computer, a server, a network device, etc., and is used to execute all or part of the steps of the image face-changing method based on personalized generation technology described in Example 1.
[0085] Each embodiment of the present invention is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is merely exemplary, in which the modules described as separate components may or may not be physically separated, and the functions of each module can be implemented in the same or one or more software and / or hardware when implementing the scheme of the present invention. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the scheme of this embodiment.
[0086] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.
Claims
1. An image face-changing method based on personalized generation technology, characterized in that: The following steps are involved: Acquire a target face image and generate a corresponding virtual label; the virtual label includes a surname word and a name word obtained by mapping the target face image through cross initialization; Obtaining a marked editing area for the target editing image based on a face detection algorithm; The target edited image with the marked edited area is used as an image input, and the virtual label corresponding to the target face is used as a text input, and a first face-changing image is generated based on a random difference editing method in a diffusion model; The non-marked editing area in the target editing image is spliced with the marked editing area in the first face-changing image as image input, the virtual label is used as text input, and a second face-changing image is generated and output based on a random difference editing method in a diffusion model.
2. The image face-swapping method according to claim 1, characterized in that: In the generation of the first face-changing image and the second face-changing image based on the random difference editing method in the diffusion model, for the diffusion model with a total number of iterations of T, it is configured to perform denoising starting from the 2 / T step in the inverse diffusion process.
3. The image face-swapping method according to claim 2, characterized in that: The diffusion model is configured to perform denoising starting from iteration number t=400~600 in the inverse diffusion process.
4. The image face-swapping method according to claim 1, characterized in that: When splicing the unmarked editing area in the target editing image with the marked editing area of the first face-changing image, the following steps are included: Perform an erosion operation on the edge of the marked editing area; The target edited image is stripped of the corroded marked edited area, and is concatenated with the corroded marked edited area in the first face-changing image to obtain a second-stage input image.
5. The image face-swapping method according to claim 4, characterized in that: When the edge of the mark editing area is eroded, the range is expanded to 10%~20%.
6. The image face-swapping method according to any one of claims 1 to 5, characterized in that: When the second face-changing image is generated based on the random difference editing method in the diffusion model, the text input of the diffusion model also includes facial expressions, postures and / or feature description texts input by the user.
7. An image face-changing system based on personalized generation technology, applying the image face-changing method based on personalized generation technology according to any one of claims 1 to 6, characterized in that: include: A text processing module is used to obtain the target face image and generate the corresponding virtual label; A region processing module, used for obtaining a marked editing region for a target editing image based on a face detection algorithm; A first face-changing module, on which a diffusion model is mounted, is used to input the target edited image with the marked edited area as an image, input the virtual label corresponding to the target face as a text, and generate a first face-changing image based on a random difference editing method in the diffusion model; An image stitching module, used for stitching the unmarked edited area in the target edited image with the marked edited area in the first face-swapped image to obtain a stitched image; The second face-changing module is equipped with a diffusion model, which is used to take the spliced image as an image input and the virtual label as a text input, generate and output a second face-changing image based on a random difference editing method in the diffusion model.
8. A device comprising a memory and a processor, wherein the memory stores computer-readable instructions, characterized in that: When the computer-readable instructions are executed by the processor, the processor executes all or part of the steps of the image face-changing method based on personalized generation technology as described in any one of claims 1 to 6.
9. A storage medium having computer-readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by the processor, all or part of the steps of the image face-changing method based on personalized generation technology as described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Text-guided controllable portrait generation method, system and equipment based on diffusion model
CN118114124A
Text-driven image editing via image-specific finetuning of diffusion models
WO2024086598A1