Live real-time face replacement method and electronic device based on deep learning
Through a live live real-time face replacement method based on deep learning, combined with face detection, feature point detection and mask model, the training loss function and 3D reconstruction loss function are solved, and the problem of difficult to balance the fidelity and generalization ability of the face replacement method in the existing technology is achieved, achieving a real-time face replacement effect with high fidelity and high generalization.
Patent Information
- Application Number
- CN202411080344.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-08
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-08-08
AI Technical Summary
Existing image or video face replacement methods are difficult to balance between fidelity and generalization ability, and cannot effectively deal with posture differences, expression differences, and skin tone or lighting differences.
The live real-time face replacement method based on deep learning is adopted to obtain the facial feature data set of the source face through face detection, feature point detection and mask model, and input the target face sample into the face exchange model. The training loss function, reconstruction loss function and 3D reconstruction loss function are used to evaluate and adjust the accuracy and attribute information of face features.
Real-time face swaps during live broadcasting is achieved, providing a high-fidelity and high generalization face swap framework, which can adapt to different face swap goals without retraining, and has good robustness.
Smart Images

Figure CN119052568B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and particularly relates to a live real-time face replacement method and an electronic device based on deep learning. Background Art
[0002] AI artificial intelligence is gradually infiltrating and integrating into various fields. The face replacement technology aims to replace the face in an image or video with a target face, so that the generated image is similar to the target face and has the appearance features of the face in the image or video. As one of the more popular applications in the fields of computer vision and graphics in recent years, it has been widely used in scenarios such as interactive entertainment, portrait replacement, advertising promotion, and movie post-production.
[0003] Currently, face replacement methods can be divided into two types: target attribute guidance and source identity guidance.
[0004] Target attribute guidance edits the source face and then mixes it into the target background, that is, directly deforms the source face according to the target facial key points. Therefore, it cannot solve the problems of large pose differences and expression differences. The method based on 3DMM performs face replacement through 3D fitting and re-rendering. However, these methods usually cannot handle skin color or lighting differences, and have low fidelity. The method based on GAN improves the fidelity of the generated face. Deepfakes transfers the target attributes to the source face through an encoder-decoder structure. The cyclic training of Gan can significantly reduce the authenticity generated by directly decoding the target face with the source face decoder.
[0005] Source identity guidance methods usually use feature vectors representing identity or latent expressions through StyleGAN2 to represent the source identity and inject it into the target face. SimSwap introduces a weak feature matching loss to help retain target attributes. At the same time, an ID injection module (IIM) is proposed, which transfers the identity information of the source face to the target face at the feature level. By using this module, SimSwap extends the architecture of the identity-specific face swapping algorithm to the framework of arbitrary face swapping.
[0006] Currently, there are still the following difficulties in face replacement: face replacement frameworks with high fidelity (such as deepfakes) lack strong generalization ability and cannot be used for all faces; while Simswap and FaceShifter with strong generalization ability need to balance the identity of the source face and the attributes of the target face when embedding the feature vector of the source identity. Therefore, they are weak in constraining the attributes of the target face (such as expression, pose, lighting, etc.). The 3D method can reconstruct coefficients such as expression and pose, can provide richer attribute information, and has better adaptability to various pose faces. The combination of the two methods has also become an intuitive optimization idea.
[0007] The information disclosed in this background section is only intended to enhance the overall understanding of the present invention and should not be regarded as an admission or any form of suggestion that this information constitutes prior art already known to those of ordinary skill in the art. Summary of the Invention
[0008] To overcome the deficiencies existing in the prior art, there is provided a live real-time face replacement method and an electronic device based on deep learning to solve the problem that the existing image or video face replacement methods cannot balance fidelity and generalization ability.
[0009] To achieve the above object, there is provided a live real-time face replacement method based on deep learning, including the following steps:
[0010] Obtain the source video and extract frames to obtain frame images;
[0011] Determine the position of the source face in the frame image through a face detection model;
[0012] Obtain the facial feature dataset of the source face in the frame image through a feature point detection model;
[0013] Based on the frame image, infer the face mask of the source face through a mask model, crop the face contour of the source face and determine the face fusion range;
[0014] Input the target face sample into the face swapping model to decouple the face feature parameters and fuse them onto the source face in the frame image to obtain a face conversion image;
[0015] Merge the face conversion images to generate a video stream and output it.
[0016] Further, the face swapping model is the deepfacelab face model.
[0017] Further, the face swapping model is the FaceSwap face swapping framework.
[0018] Further, when inputting the target face sample into the face swapping model to decouple the face feature parameters and fuse them onto the source face in the frame image to obtain a face conversion image, a training loss function is used to evaluate the accuracy of face feature modification, and the loss function is:
[0019]
[0020] where L id is the identity loss, X s is the image sample data of the source face, Y is the face swapping result, and Z id is the identity feature vector extracted after the image data passes through the face recognition network [ArcFace];
[0021] The training reconstruction loss function is used to determine whether the target face and the source face are from the same identity. The reconstruction loss function is as follows:
[0022]
[0023] Among them, L r is the reconstruction loss, and X s is the image sample data of the source face;
[0024] The training 3D reconstruction loss is used to constrain the 3D face reconstruction result to be consistent with the face swapping result. The 3D reconstruction loss function is as follows:
[0025] L r3d = ||R kp (F 3d (Y)) - R kp (F 3d (X t , X s ))||1,
[0026] Among them, L r3d is the 3D reconstruction loss, R kp is to extract the corresponding key points from the 3D reconstruction matrix, and F 3d is to perform 3D reconstruction on the image data. When two images are input, it means to perform 3D reconstruction on each image separately, and then obtain a new 3D model through parameter fusion conversion.
[0027] Furthermore, by extracting multi-level features from the ground truth (i.e., the reference truth, which is the target face X t ) and the face swapping result, the multi-level features are as follows:
[0028]
[0029] Among them, L oFM is the multi-level feature loss, D (i) represents the i-th layer feature extractor of the discriminator D, N i represents the sum of the i-th layer elements, and M represents the total number of layers;
[0030] Based on the multi-level features, the identity loss is corrected to obtain the final loss function to avoid network overfitting and only generating positive images with the identity of the source face while losing all the attributes of the target face. The final loss function:
[0031] L total = ω1L id + ω2L r + ω3L r3d + ω4L oFM ,
[0032] Among them, L total is the final loss, ω1 is the weighting coefficient of L id weighting coefficient, ω2 is the weighting coefficient of L r weighting coefficient, ω3 is the weighting coefficient of L r3d weighting coefficient, ω4 is the weighting coefficient of L oFM weighting coefficient.
[0033] Furthermore, the face detection model is a centerface model, an S3FD model, or a yolov5 model.
[0034] Furthermore, the feature point detection model is an LBF model, a FaceMesh model, or an Insightface 2D model.
[0035] The present invention provides an electronic device, including:
[0036] at least one processor;
[0037] a memory communicatively connected to the at least one processor, the memory storing a computer program executable by the at least one processor, and the computer program being executed by the at least one processor to enable the at least one processor to execute a live real-time face replacement method based on deep learning.
[0038] The present invention provides a computer-readable storage medium storing computer instructions for causing a processor to implement a live real-time face replacement method based on deep learning when executed.
[0039] The beneficial effects of the present invention are as follows. The live real-time face replacement method based on deep learning of the present invention realizes real-time face replacement during live broadcasts and provides a face replacement framework with high fidelity and high generalization. For different face replacement targets, there is no need to retrain, and it has good robustness. In order to run visually, the live real-time face replacement method of the present invention uses PyQT to draw a UI page, reducing the running difficulty. The system mainly performs frame transmission and processing through various services in the Backend (background files under the system directory). In the UI directory, it is connected to the external holder of the backend through controlsheet (controller, an external holder that links to the backend through the controller). When the backend processes, it calls models such as onnxruntime in modelhub for corresponding processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present application will become more apparent:
[0041] Figure 1 This is a schematic structural diagram of the live real-time face replacement method based on deep learning according to an embodiment of the present invention.
[0042] Figure 2 This is a schematic structural diagram of the face replacement network according to an embodiment of the present invention. Detailed implementation manners
[0043] The following further elaborates the present application in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the convenience of description, only parts related to the invention are shown in the drawings.
[0044] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The following will detail the present application with reference to the accompanying drawings and embodiments.
[0045] Referring to Figure 1 and Figure 2 As shown, the present invention provides a live real-time face replacement method based on deep learning, including the following steps:
[0046] S1. Obtain the source video and extract frames to obtain frame images.
[0047] In this embodiment, for the configuration of the source video, its file source input is a local file input or a local camera input. The local file is an image folder or a video file as the input.
[0048] S2. Determine the position of the source face in the frame image through a face detection model.
[0049] As a preferred implementation manner, the face detection model is a centerface model, an S3FD model, or a yolov5 model.
[0050] Identify the specific position of the face in the frame image of the source video through the above face detection model.
[0051] S3. Obtain the facial feature data set of the source face in the frame image through a feature point detection model.
[0052] As a preferred implementation manner, the feature point detection model is an LBF model, a FaceMesh model, or an Insightface 2D model.
[0053] Generate the facial feature data set through the above feature point detection model.
[0054] S4. Based on the frame image, infer the face mask of the source face through the mask model, crop the face contour of the source face, and determine the face fusion range.
[0055] In this embodiment, the mask model is the Xseg mask model.
[0056] Specifically, by inferring the face mask through the Xseg mask model, the feathering degree of the mask edge can be dynamically adjusted.
[0057] Determine the subsequent face fusion part by cropping the face contour, and determine the distance between the fusion edge and the face edge by setting the face fusion edge parameters.
[0058] In this embodiment, the facial feature dataset generated in step S3 contains face mesh information, and the bounding box of the face can be determined from the face mesh information.
[0059] Step S4 is to crop based on the face contour generated in step S3 to determine the face fusion range.
[0060] The workflow of S3 - S4 is to determine the face contour through the face mesh information in the facial feature dataset, remove the mask (such as occlusions like glasses) based on the face contour, and finally determine the part that needs to be fused.
[0061] S5. Input the target face sample into the face swapping model to decouple the face feature parameters and fuse them onto the source face of the frame image to obtain the face conversion image.
[0062] In actual use, face replacement is a switchable function option. When it is turned off, the frame image is only output through the beauty module. After it is turned on, the frame image sequentially executes steps S2, S3, and S4 and then performs face replacement.
[0063] As a preferred embodiment, the face swapping model is the deepfacelab face model or the FaceSwap face swapping framework.
[0064] The face replacement process is that the face frame image passes through the face replacement network to generate the inference result. The user can import their own pre - trained deepfacelab face model, or convert the AMP and SAEHD models into the DFM format. At the same time, it also supports a face swapping framework FaceSwap with strong generalization ability. Without collecting target face data and retraining, relying on a single source face picture, the replacement of the target face image can be achieved.
[0065] In this embodiment, the FaceSwap face swapping framework has strong generalization ability. The FaceSwap face swapping framework includes a 3D prior deformation estimation module, an encoder Encoder, a decoder Decoder, and an identity injection module.
[0066] The 3D method can reconstruct coefficients such as expression and pose, providing richer attribute information and better adaptability to various poses of human faces. With the help of the 3D deformable human face model (3DMM), coefficients such as ID (identity), Color (skin color), Expression (expression), Pose (pose), and Light (light sense) of the template face and the target face are extracted respectively, and after coefficient replacement and synthesis, they are used as additional information to be input into the generative framework for face swapping.
[0067] In the case where the source face shape and the target face shape differ greatly, with the help of the 3D deformable human face model 3DMM, the 3DMM coefficients of the source and target faces are extracted, and the reconstruction of the target face is obtained through the fusion of human face coefficients. At this time, the movement displacement obtained from the key point information of the reconstructed target face and the original human face indicates the movement change of the image pixels in space.
[0068] The encoder Encoder is used to extract the features of the target human face to be replaced. The identity injection module is dedicated to modifying the target human face features, keeping the attribute information of the target human face unchanged (such as expression, pose, etc.), and replacing the identity information of the target human face with the identity information of the source human face. The identity information of the source human face is extracted through the face recognition network ArcFace and injected into the target human face features using the image style transfer algorithm AdaIN. Finally, the image is restored through the decoder Decoder to obtain the target human face with the identity features of the source human face and unchanged attribute information.
[0069] In the identity injection module, the identity information and attribute information of the human face are highly coupled. When injecting the identity information of the source human face into the target human face, the attribute information may also be affected by the injection of the identity information.
[0070] As a preferred implementation, when the target human face sample is input into the face swapping model to decouple the human face feature parameters and fuse them into the source human face of the frame image to obtain the face conversion image, the training loss function is used to evaluate the accuracy of the modification of the human face features. The loss function is:
[0071]
[0072] Among them, L id is the identity loss, X s is the image sample data of the source human face, Y is the face swapping result, and Z id is the identity feature vector extracted after the image data passes through the face recognition network [ArcFace];
[0073] The training reconstruction loss function is used to determine whether the target human face and the source human face are from the same identity. The reconstruction loss function is:
[0074]
[0075] Among them, L r is the reconstruction loss, and X s is the image sample data of the source face;
[0076] The training 3D reconstruction loss is used to constrain the 3D face reconstruction result to be consistent with the face swapping result. The 3D reconstruction loss function is:
[0077] L r3d = ||R kp (F 3d (Y)) - R kp (F 3d (X t , X s ))||1,
[0078] Among them, L r3d is the 3D reconstruction loss, R kp is to extract the corresponding key points from the 3D reconstruction matrix, F 3d is to perform 3D reconstruction on the image data. When two images are input, it means performing 3D reconstruction on each image separately and then obtaining a new 3D model through parameter fusion conversion.
[0079] As a preferred implementation, by extracting multi-level features from the ground truth and the face swapping result, the multi-level features are:
[0080]
[0081] Among them, L oFM is the multi-level feature, D (i) represents the i-th layer feature extractor of the discriminator D, N i represents the sum of the i-th layer elements, and M represents the total number of layers;
[0082] Based on the multi-level features, the identity loss is corrected to obtain the final loss function to avoid network overfitting and only generating positive images with the identity of the source face while losing all target face attributes. The final loss function:
[0083] L total = ω1L id + ω2L r + ω3L r3d + ω4L oFM ,
[0084] Among them, L total is the final loss, ω1 is the weighting coefficient of L id , ω2 is the weighting coefficient of L r , ω3 is the weighting coefficient of L r3d , ω4 is the weighting coefficient of L oFM weighting coefficient.
[0085] S6. Combine the face conversion images to generate a video stream and output it.
[0086] Fuse the face replacement result with the face template to generate the final face conversion image, and add an operation to smooth the front and back frames through key point detection to ensure the stability of the key points in consecutive frames.
[0087] The present invention provides a live real-time face replacement method based on deep learning
[0088] This system has successfully applied AI face replacement to the live broadcast field, realizing real-time replacement of the host's image. At the same time, according to the differences in application scenarios, different types of face replacement models are provided for different user groups.
[0089] The present invention provides an electronic device, including: at least one processor and a memory.
[0090] The memory is communicatively connected to at least one processor. The memory stores a computer program executable by at least one processor. The computer program is executed by at least one processor so that at least one processor can execute the aforementioned live real-time face replacement method based on deep learning.
[0091] As can be understood from the above, the live real-time face replacement method based on deep learning provided by the embodiments of the present application can be implemented by various types of electronic devices with processing capabilities, such as being implemented and executed by the processor of an electronic device or being implemented and executed by other devices with computing and processing capabilities. Other devices with computing and processing capabilities can be intelligent terminals or servers communicatively connected to the electronic device, etc.
[0092] The processor is wirelessly connected to the server or the server through a wireless signal without a communication module. The processor is configured to support the electronic device to execute the corresponding functions in the live real-time face replacement method based on deep learning.
[0093] The processor can be a central processing unit, a network processor, a hardware chip, or any combination thereof. The hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof.
[0094] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the live real-time face replacement method based on deep learning in the embodiments of the present application. By running the non-transitory software programs, instructions, and modules stored in the memory, the processor can implement the live real-time face replacement method in this embodiment.
[0095] The memory may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk or solid state drive; the memory may further include a combination of the above types of memory.
[0096] The present invention provides a computer-readable storage medium storing computer instructions for causing a processor to implement the foregoing live real-time face replacement method based on deep learning when executed.
[0097] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0098] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a dedicated computer, or other programmable data processing device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0099] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0100] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0101] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0102] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0103] The live real-time face replacement method based on deep learning of the present invention realizes real-time face replacement during live broadcasts and provides a high-fidelity and highly generalized face replacement framework. For different face replacement targets, there is no need for retraining, and it has good robustness. In order to run visually, the live real-time face replacement method based on deep learning of the present invention uses PyQT (PyQt is a toolkit for creating GUI applications) to draw the UI page, reducing the operation difficulty. The system mainly performs frame transmission and processing through various services in the Backend directory. In the UI directory, it is connected to the external holder of the backend through the control sheet. When the backend processes, it calls models such as onnxruntime in the modelhub for corresponding processing.
[0104] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present application.
Claims
1. A live broadcast real-time face replacement method based on deep learning, characterized in that: The following steps are involved: Get the source video and extract frames to obtain frame images; Determine the source face position in the frame image by using a face detection model; Acquire a facial feature dataset of the source face of the frame image through a feature point detection model; Based on the frame image, inferring a face mask of the source face through a mask model, cropping the face contour of the source face and determining a face fusion range; Inputting the target face sample into the face exchange model to decouple the face feature parameters and fuse them to the source face of the frame image to obtain a face conversion image; Merging the face conversion images to generate a video stream and outputting it; The face swap model is the FaceSwap face swap framework; When the target face sample is input into the face swap model to decouple the face feature parameters and fused to the source face of the frame image to obtain the face conversion image, the training loss function is used to evaluate the accuracy of the face feature modification, and the loss function is: ; in, For identity loss, is the image sample data of the source face, For the face-changing results, It is the identity feature vector extracted after the image data passes through the face recognition network; The training reconstruction loss function is used to determine whether the target face is from the same identity as the source face. The reconstruction loss function is: ; in, To rebuild the losses, is the image sample data of the source face; The training 3D reconstruction loss is used to constrain the 3D face reconstruction result to be consistent with the face swap result. The 3D reconstruction loss function is: ; in, is the 3D reconstruction loss, To extract the corresponding key points of the 3D reconstruction matrix, In order to perform 3D reconstruction on image data, when two images are input, it means that 3D reconstruction is performed on each image separately, and then a new 3D model is obtained through parameter fusion transformation.
2. The live broadcast real-time face replacement method based on deep learning according to claim 1 is characterized in that: The face exchange model is the deepfacelab face model.
3. The live broadcast real-time face replacement method based on deep learning according to claim 1 is characterized in that: By extracting multi-layer features from the benchmark truth and the face-swapping result, the multi-layer features are: ; in, is a multi-layer feature. represents the i-th layer feature extractor of the discriminator D, represents the sum of elements in the i-th layer, and M represents the total number of layers; Based on the multi-layer features, the identity loss is modified to obtain a final loss function to avoid network overfitting and only generate a frontal image with the source face identity while losing all target face attributes. The final loss function is: ; in, For the final loss, for Weighting coefficient, for Weighting coefficient, for Weighting coefficient, for Weighting factor.
4. The live broadcast real-time face replacement method based on deep learning according to claim 1, characterized in that: The face detection model is a centerface model, an S3FD model, or a yolov5 model.
5. The live broadcast real-time face replacement method based on deep learning according to claim 1, characterized in that: The feature point detection models are LBF model, FaceMesh model, and Insightface 2D model.
6. An electronic device, characterized in that: include: at least one processor; A memory communicatively connected to the at least one processor, the memory storing a computer program executable by the at least one processor, the computer program being executed by the at least one processor so that the at least one processor can execute the live broadcast real-time face replacement method based on deep learning as described in any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the live broadcast real-time face replacement method based on deep learning as described in any one of claims 1 to 5 when executed.
Citation Information
Patent Citations
High-definition face replacement video generation method and system
CN112446364A
Method and system for supporting simultaneous real-time replacement of multiple faces
CN117710190A