Head and face changing method based on segmentation and staged diffusion model

Through the head-changing and face-changing method based on the segmentation and phased diffusion model, the face area of ​​the image is identified and redrawn, and the denoising processing is performed using the lora weight, the problems of blurred texture and distorted structure of the face-changing image in the prior art are solved, and a more efficient and high-precision face-changing effect is achieved.

CN120070674AActive Publication Date: 2025-05-30SHENZHEN LINGTU SHINE TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510545872.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

In the existing head and face change technology, the face change images generated by the adversarial network have problems with blurred texture and distorted structure, and generating high-definition images requires extremely deep networks and a large amount of video memory, and the training time is long.

Method used

The head and face change method based on segmentation and phased diffusion model is adopted. The face area of ​​the image to be replaced is identified through the human body analysis model and the face analysis model, and the face image is redrawn multiple times through the diffusion model, the preset redraw amplitude is adjusted, and the denoising process is used using the lora weight, and the target face part is finally fused into the image to be replaced.

Benefits of technology

It improves face swap efficiency and image accuracy, reduces the loss rate of key areas of the face, improves high-definition image processing speed, and makes face boundary transitions more natural.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070674A_ABST
    Figure CN120070674A_ABST
Patent Text Reader

Abstract

The invention discloses a head and face changing method based on a segmentation and staged diffusion model, and the method comprises the steps: recognizing a to-be-replaced image through a human body analysis model and a face analysis model, obtaining a face region of the to-be-replaced image and a mask pattern of the face region, cutting the face region of the to-be-replaced image, and obtaining a mask pattern of the face region of the to-be-replaced image; obtaining a face image containing a face part; according to the mask graph of the face area, multiple times of redrawing is conducted on the face image through a diffusion model, a target image containing a target face part is obtained, the preset redrawing amplitude is adjusted every time redrawing is conducted, and the diffusion model comprises a lora weight corresponding to the target face part; and performing fusion processing on the target face part of the target image and the face region of the to-be-replaced image to obtain a replaced image. According to the method and the device, the face information of the target face part can be stored in the lora weight, and when a specific target face needs to be replaced, the diffusion model can directly load the corresponding lora weight, so that the face replacement efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of image redrawing, and particularly to a method for face swapping based on segmentation and a staged diffusion model. Background Art

[0002] Existing face swapping technologies mainly use generative adversarial networks (GANs) for global image generation. It outputs fake images through a generator, and discriminates the real probabilities of the fake images and the input images through a discriminator. Based on adversarial training, the generator and the discriminator are trained, and finally the generator is used to generate face-swapped images.

[0003] However, the face-swapped images generated by generative adversarial networks often have problems such as blurred textures and distorted structures. At the same time, generating high-definition images (such as 1024x1024) requires extremely deep networks and a large amount of video memory, and the training time is long. Summary of the Invention

[0004] The purpose of this application is to provide a method for face swapping based on segmentation and a staged diffusion model to solve the technical problems of low face swapping efficiency and accuracy in existing face swapping methods. The many technical effects that can be produced by the preferred technical solutions provided in this application are described in detail below.

[0005] To achieve the above purpose, this application provides the following technical solutions: In a first aspect, a method for face swapping based on segmentation and a staged diffusion model provided by this application includes: identifying a to-be-replaced image through a human body parsing model and a face parsing model to obtain the face region of the to-be-replaced image and the mask image of the face region, cropping the face region of the to-be-replaced image to obtain a face image containing the face part; according to the mask image of the face region, redrawing the face image multiple times through a diffusion model to obtain a target image containing the target face part, where a preset redrawing amplitude is adjusted each time of redrawing, and the diffusion model includes the lora (Low-Rank Adaptation) weight corresponding to the target face part; fusing the target face part of the target image with the face region of the to-be-replaced image to obtain a replaced image.

[0006] In some embodiments, the method for face swapping based on segmentation and a staged diffusion model further includes: obtaining any target model image, and training the lora adapter of the diffusion model using the target model image to obtain the lora weight corresponding to the target face part.

[0007] In some embodiments, training the LoRA adapter of the diffusion model using the target model image includes: fine-tuning the linear layer in the attention layer of the diffusion model, decomposing the original large matrix parameters of the diffusion model into two low-rank matrices, and updating the two low-rank matrices based on the loss value of the target model image to obtain the LoRA adapter of the diffusion model.

[0008] In some embodiments, redrawing the face image multiple times by the diffusion model according to the mask map of the face region includes: adding random Gaussian noise and original image noise to the face image to obtain a noise image, and then denoising the noise image in combination with the LoRA weights to obtain an initial denoised image; resetting the noise of the initial denoised image and then denoising it in combination with the LoRA weights, repeating the noise reset operation and the denoising operation multiple times until the target image is obtained.

[0009] In some embodiments, the face swapping method based on the segmentation and phased diffusion model further includes: cropping the mask map of the face region to obtain a face swapping frame, and using the face swapping frame as the face region of the image to be replaced, where the face swapping frame is the largest square frame in the mask map of the face region.

[0010] In some embodiments, fusing the target face part of the target image with the face region of the image to be replaced includes: obtaining the weight matrix of the face swapping frame, and performing weighted mixing of the weight matrix with the target face part of the target image and the face region of the image to be replaced respectively, where the weights of the weight matrix decrease gradually from the center point to the edge of the face swapping frame.

[0011] In some embodiments, identifying the image to be replaced by the human body parsing model and the face parsing model includes: identifying the head region of the person in the image to be replaced by the human body parsing model, and identifying the head region by the face parsing model to remove the hair part and the accessory part of the head region.

[0012] In a second aspect, a computer-readable storage medium provided by the present application stores a computer program, and when the computer program is executed, it implements the face swapping method based on the segmentation and phased diffusion model as described above.

[0013] In a third aspect, a processing device provided by the present application includes: one or more processors; a memory for storing one or more computer programs, and the one or more processors are configured to execute the one or more computer programs stored in the memory, so that the one or more processors execute the face swapping method based on the segmentation and phased diffusion model as described above.

[0014] In a fourth aspect, a computer program product provided by the present application is stored on a data carrier and is designed to execute the face swapping method based on the segmentation and phased diffusion model as described above.

[0015] Implementing one of the above technical solutions of the present application has the following advantages or beneficial effects: In the present application, the face region of the image to be replaced is recognized and cropped, and the face image is redrawn multiple times through a diffusion model. The preset redrawing amplitude is adjusted each time of redrawing. At the same time, the diffusion model has lora weights corresponding to the target face part. Finally, an image with the target face replaced is obtained. In this case, the face information of the target face part can be stored in the form of lora weights. When it is necessary to replace a specific target face, the diffusion model can directly and timely load the corresponding lora weights, thereby improving the face swapping efficiency; at the same time, by redrawing multiple times and adjusting the preset redrawing amplitude in the present application, the control of face swapping details can be strengthened, and the loss rate of key regions of the face part can be reduced; in addition, by cropping for local processing in the present application, the processing speed of high-precision images can be improved; in addition, through multiple redrawings and background fusion, the transition of the face boundary can be made more natural. Description of the Drawings

[0016] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings: Figure 1 is a flowchart of a face swapping method based on the segmentation and phased diffusion model according to an embodiment of the present application; Figure 2 is a schematic diagram of a mask map of the face region according to an embodiment of the present application; Figure 3 is a schematic diagram of the target face part according to an embodiment of the present application; Figure 4A is a schematic diagram of an image to be replaced according to an embodiment of the present application, Figure 4B is a schematic diagram of a replaced image according to an embodiment of the present application; Figure 5AIt is another schematic diagram of the image to be replaced in the embodiment of the present application. Figure 5B It is another schematic diagram of the replaced image in the embodiment of the present application; Figure 6 It is a structural block diagram of the processing device in the embodiment of the present application.

[0017] In the figure: 1. Processing device; 10. Memory; 11. Processor. Detailed implementation manners

[0018] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, various exemplary embodiments to be described below will refer to the corresponding drawings, which form a part of the exemplary embodiments and describe various exemplary embodiments that may be adopted to implement the present application. Unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. It should be understood that they are only examples of processes, methods, devices, etc. consistent with some aspects of the present application disclosed in detail in the appended claims. Other embodiments may also be used, or structural and functional modifications may be made to the embodiments listed herein without departing from the scope and essence of the present application.

[0019] In the description of the present application, it should be understood that terms such as "center", "longitudinal", "lateral", etc. indicate the orientation or positional relationship based on the orientation shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the elements referred to must have a specific orientation, be constructed and operated in a specific orientation. Terms such as "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. The meaning of the term "plurality" is two or more. The terms "connected" and "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, an integral connection, a mechanical connection, an electrical connection, a communication connection, a direct connection, an indirect connection through an intermediate medium, and may be the connection inside two elements or the interaction relationship between two elements. The term "and / or" includes any and all combinations of one or more of the related listed items. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.

[0020] In order to illustrate the technical solutions described in the present application, the following will be described through specific embodiments, and only the parts related to the embodiments of the present application are shown.

[0021] As Figure 1 shown, the present application provides a head and face swapping method based on a segmentation and staged diffusion model, including the following steps (step S1 to step S3): S1. Use a human parsing model and a face parsing model to recognize the image to be replaced, obtain the face region of the image to be replaced and the mask image of the face region, and crop the face region of the image to be replaced to obtain a face image containing the face part.

[0022] In some embodiments, using a human parsing model and a face parsing model to recognize the image to be replaced may include: using the human parsing model to recognize the head region of the person in the image to be replaced, and using the face parsing model to recognize the head region to remove the hair part and accessory part of the head region.

[0023] In some embodiments, the face parsing model and the human parsing model can be trained using publicly available datasets. For example, for the face parsing model, the publicly available dataset can be the CelebAMask-HQ dataset or the Helen dataset, and for the human parsing model, the publicly available dataset can be the LIP dataset or the PASCAL-Person-Part dataset. The face parsing model and the human parsing model can also be trained using some publicly uploaded images on the Internet.

[0024] In some embodiments, the human parsing model can be used to segment a human body image into different semantic regions, such as the head, upper limbs, lower limbs, clothing, etc.; the face parsing model can be used to segment a face image into multiple semantic regions, such as the skin, eyes, lips, hair, etc. Further, through the cooperation between the face parsing model and the human parsing model, the face part of the image to be replaced can be recognized and located, and the occluders on the face can be recognized. Figure 2 The shown one is the mask image of the face region in the embodiment of the present application.

[0025] In some embodiments, the face parsing model and the human parsing model can be built based on the mask2former framework or the U-Net framework. The Mask2Former framework is a general image segmentation framework based on the Transformer (global attention) architecture, which supports panoramic, instance, and semantic segmentation tasks, and can uniformly process various segmentation tasks, especially object cutting in complex scenarios; the U-Net framework is also an image segmentation framework, which includes an encoder and a decoder. The encoder can extract multi-scale features through convolution and downsampling, and the decoder can restore the spatial resolution through upsampling and convolution, and combine skip connections to fuse the high-resolution details of the encoder and the semantic information of the decoder.

[0026] In some embodiments, the head and face swapping method based on the segmentation and phased diffusion model may further include: cropping the mask map of the face region to obtain a face swapping frame, and using the face swapping frame as the face region of the image to be replaced, where the face swapping frame is the largest square frame in the mask map of the face region. By further cropping to obtain the face swapping frame, the actual range of face swapping can be reduced, the computing power required for face swapping can be saved, and the face swapping efficiency can be improved.

[0027] S2. According to the mask map of the face region, the face image is redrawn multiple times through the diffusion model to obtain a target image including the target face part, where the preset redrawing amplitude is adjusted each time of redrawing, and the diffusion model includes the lora weight corresponding to the target face part.

[0028] In some embodiments, the head and face swapping method based on the segmentation and phased diffusion model may further include: obtaining any target model image, and training the lora adapter of the diffusion model by using the target model image to obtain the lora weight corresponding to the target face part. Figure 3 The target face part of the present application is shown as above.

[0029] Specifically, the diffusion model can be used for image generation, that is, it can be used to generate the image of the target face part. Lora can be used to fine-tune the parameters of the diffusion model, that is, the diffusion model can generate an image conforming to the target face part according to the lora weight.

[0030] In some embodiments, training the lora adapter of the diffusion model by using the target model image may include: fine-tuning the linear layer in the attention layer of the diffusion model, decomposing the original large matrix parameters of the diffusion model into two low-rank matrices, updating the two low-rank matrices based on the loss value of the target model image, and obtaining the lora adapter of the diffusion model. Specifically, the diffusion model can insert the above two low-rank matrices into the target layer, fix the original parameters of the diffusion model, and fine-tune the two low-rank matrices to obtain the lora adapter. The lora adapter can be independent, that is, the target model image can correspond to the lora adapter, and the lora adapter can be used to store the face information of the target model image, that is, the target face part. The diffusion model can generate different target face parts by loading different lora adapters.

[0031] In some embodiments, repeatedly redrawing a face image through a diffusion model according to a mask image of a face region may include: adding random Gaussian noise and original image noise to the face image to obtain a noise image, and then denoising the noise image in combination with LoRA weights to obtain an initial denoised image; resetting the noise of the initial denoised image and then performing denoising in combination with LoRA weights, repeating the noise reset operation and the denoising operation multiple times until a target image is obtained. Specifically, the redrawing amplitude can be used to control the intensity of image restoration. For example, when the redrawing amplitude is equal to 1, the face image can be completely replaced, and when the redrawing amplitude is equal to 0.5, some information of the face image can be retained.

[0032] Specifically, the random Gaussian noise and the original image noise can be composed by weighting, and the original image noise can be obtained by calculating the face image. The process of adding noise can refer to the forward diffusion process of the diffusion model, which can generate a noise image corresponding to the time step t. The denoising process can refer to the reverse sampling process of the diffusion model, which can restore the noise image to a face image. By combining LoRA weights, it can be restored to an image including the target face part until a target image meeting the requirements is obtained.

[0033] S3. Fuse the target face part of the target image with the face region of the image to be replaced to obtain a replaced image. The replaced image can refer to the final image in which the target face part has been replaced into the face region of the image to be replaced.

[0034] In some embodiments, fusing the target face part of the target image with the face region of the image to be replaced may include: obtaining a weight matrix of a face-swapping frame, and performing weighted mixing of the weight matrix with the target face part of the target image and the face region of the image to be replaced respectively, where the weights of the weight matrix decrease gradually from the center point to the edge of the face-swapping frame.

[0035] Specifically, the weight of the center point of the face-swapping frame can be 1, and the weight of the edge of the face-swapping frame can be 0. The distance from each point of the face-swapping frame to the edge can be calculated, the distance can be normalized, and finally the normalized distance can be converted into a weight value by using a linear function or a cosine function. This can enable the target face part to gradually transition to the background part, making the replaced image more natural.

[0036] Figure 4A 、 Figure 4B The following shows a group of images to be replaced and replaced images of the present application. Figure 5A 、 Figure 5B The following shows another group of images to be replaced and replaced images of the present application. Both groups of images use Figure 3 as the target face part to be replaced. It should be noted that the Figure 3 、 Figure 4A 、Figure 4B , Figure 5A and Figure 5B The characters and human faces in are all computer-simulated and not real people.

[0037] In this application, the face region of the image to be replaced is recognized and cropped, and the face image is redrawn multiple times through a diffusion model. The preset redrawing amplitude is adjusted each time of redrawing. At the same time, the diffusion model has lora weights corresponding to the target face part. Finally, an image with the target face replaced is obtained. In this case, the face information of the target face part can be stored in the form of lora weights. When it is necessary to replace a specific target face, the diffusion model can directly and timely load the corresponding lora weights, thereby improving the face replacement efficiency; at the same time, by redrawing multiple times and adjusting the preset redrawing amplitude in this application, the control of face replacement details can be strengthened, and the loss rate of key regions in the face part can be reduced; in addition, by performing local processing through cropping in this application, the processing speed of high-precision images can be improved; in addition, through multiple redrawings and background fusion, the transition of the face boundary can be made more natural.

[0038] Those of ordinary skill in the art can understand that all or part of the features / steps of implementing the above method embodiments can be realized by a method, a data processing system or a computer program. These features can be implemented without using hardware, all using software, or using a combination of hardware and software. The aforementioned computer program can be stored in one or more computer-readable storage media. When the computer program stored on the storage medium is executed (such as by a processor), it executes the steps of the method embodiments of the face and head replacement method based on segmentation and staged diffusion model as described above.

[0039] The aforementioned storage media that can store program codes include: a static hard disk, a solid-state drive, a random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), an optical storage device, a magnetic storage device, a flash memory, a magnetic disk or an optical disc, and / or a combination of the above devices, that is, it can be implemented by any type of volatile or non-volatile storage device or a combination thereof.

[0040] As Figure 6 shown, this application also provides an embodiment of a processing device 1, including one or more processors 11 and a memory 10; wherein, the memory 10 is used to store one or more computer programs, and one or more processors 11 are used to execute the one or more computer programs stored in the memory 10, so that the processor 11 executes the features / steps of the method embodiments of the face and head replacement method based on segmentation and staged diffusion model as described above.

[0041] The present application also provides a computer program product. The computer program product is stored on a data carrier and is designed to execute the face swapping method based on the segmentation and phased diffusion model as described above. Therefore, the computer program product according to the present application has the same advantages as those described in detail with reference to the device according to the present application. The computer program product can be executed as computer-readable instruction codes in each suitable programming language such as JAVA, C++, etc. In addition, the computer program product can be provided on a network, such as the Internet, or a network user can download the computer program product from a network, such as the Internet, when needed. The computer program product can be implemented either by means of a computer program, i.e., software, or by means of one or more dedicated electronic circuits, i.e., hardware, or in any hybrid form, i.e., by means of software components and hardware components, or in a form of software, hardware, or a mixture of software and hardware.

[0042] The above are only the preferred embodiments of the present application. Those skilled in the art will understand that various changes or equivalent replacements can be made to these features and embodiments without departing from the spirit and scope of the present application. Additionally, under the teaching of the present application, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present application. Therefore, the present application is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of the present application belong to the protection scope of the present application.

Claims

1. A head and face replacement method based on segmentation and phased diffusion model, characterized in that: include: The image to be replaced is identified by using a human body analysis model and a face analysis model to obtain a face region of the image to be replaced and a mask image of the face region, and the face region of the image to be replaced is cropped to obtain a face image containing a face portion; According to the mask map of the face area, the face image is redrawn multiple times through a diffusion model to obtain a target image containing a target face part, wherein a preset redrawing amplitude is adjusted during each redrawing, and the diffusion model includes a lora weight corresponding to the target face part; The target face portion of the target image is fused with the face region of the image to be replaced to obtain a replaced image.

2. The head and face replacement method based on segmentation and phased diffusion model according to claim 1 is characterized in that: The head and face changing method based on segmentation and staged diffusion model also includes: obtaining any target model image, using the target model image to train the Lora adapter of the diffusion model, and obtaining the Lora weight corresponding to the target face part.

3. The head and face replacement method based on segmentation and phased diffusion model according to claim 2 is characterized in that: The use of the target model image to train the LoRa adapter of the diffusion model includes: fine-tuning the linear layer in the attention layer of the diffusion model, decomposing the original large matrix parameters of the diffusion model into two low-rank matrices, updating the two low-rank matrices based on the loss value of the target model image, and obtaining the LoRa adapter of the diffusion model.

4. The head and face replacement method based on segmentation and phased diffusion model according to claim 1 is characterized in that: The face image is redrawn multiple times through a diffusion model based on the mask map of the face area, including: adding random Gaussian noise and original image noise to the face image to obtain a noise image, and then denoising the noise image in combination with the lora weight to obtain an initial denoised image; resetting the noise of the initial denoised image and then denoising it in combination with the lora weight, repeating the noise resetting operation and denoising operation multiple times until the target image is obtained.

5. The head and face replacement method based on segmentation and phased diffusion model according to claim 1 is characterized in that: The head and face changing method based on the segmentation and staged diffusion model also includes: cropping the mask image of the face area to obtain a face changing frame, and using the face changing frame as the face area of ​​the image to be replaced, wherein the face changing frame is the square frame with the largest area in the mask image of the face area.

6. The head and face replacement method based on segmentation and phased diffusion model according to claim 5 is characterized in that: The method of fusing the target face part of the target image with the face area of ​​the image to be replaced includes: obtaining a weight matrix of the face-changing frame, and making the weight matrix perform weighted mixing with the target face part of the target image and the face area of ​​the image to be replaced, respectively, wherein the weights of the weight matrix gradually decrease from the center point to the edge of the face-changing frame.

7. The head and face replacement method based on segmentation and phased diffusion model according to claim 1 is characterized in that: The identifying of the image to be replaced by using a human body analysis model and a face analysis model includes: identifying the head area of ​​the person in the image to be replaced by using the human body analysis model, identifying the head area by using the face analysis model, and removing the hair part and the accessories part of the head area.

8. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, which, when executed, implements the head and face changing method based on the segmentation and staged diffusion model as described in any one of claims 1 to 7.

9. A processing device, characterized in that: include: one or more processors; A memory for storing one or more computer programs, and one or more processors for executing one or more computer programs stored in the memory, so that one or more processors execute the head and face changing method based on the segmentation and staged diffusion model as described in any one of claims 1-7.

10. A computer program product, characterized in that The computer program product is stored on a data carrier and is designed to execute the head and face replacement method based on segmentation and staged diffusion model as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Multi-person group photo synthesis method and device

    CN117372239A

  • Virtual hairstyle changing method based on diffusion model

    CN119130783A

  • Method for processing images and electronic device

    US11488293B1

  • Local refreshing method, system and apparatus for chart, and device and medium

    WO2024156191A1