Video face changing method and device, electronic equipment and storage medium

By obtaining a three-dimensional face model and detecting the status information of the image to be replaced, the model is adjusted to generate a replacement image with a high degree of fit, solving the problem of poor video face-changing effects and achieving a natural and smooth video face-changing experience.

CN120689196APending Publication Date: 2025-09-23BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410330186.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-21
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The existing technology has poor video face-swapping effects, and the faces are unnatural, resulting in a poor visual experience for users.

Method used

By obtaining a three-dimensional face model, detecting the facial status information of each frame image to be replaced in the processed video, and adjusting the three-dimensional face model according to the status information to generate a replacement image with a high degree of fit, a natural video face-changing effect can be achieved.

Benefits of technology

It improves the fit and naturalness of face-swapping in videos, enhances the user's visual experience, and ensures smoothness and immersiveness in video viewing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689196A_ABST
    Figure CN120689196A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, artificial intelligence, image processing and the like, in particular to a video face changing method and device, electronic equipment and a storage medium. According to the specific implementation scheme, a face three-dimensional model is obtained; determining each frame of to-be-processed image of a to-be-changed face in the to-be-processed video; detecting current face state information of a face to be replaced in each frame of image to be processed; performing image acquisition on the face three-dimensional model according to the current face state information to obtain a first face image; and replacing a to-be-replaced face in the to-be-processed image according to the first face image to obtain a second face image after face replacement. According to the method and the device, image acquisition is performed on the face three-dimensional model according to the state information of the to-be-replaced face to obtain the first face image, the first face image is close to the to-be-replaced face in state, the first face image can be better fused with the to-be-processed image, and the video after face replacement brings more natural and smooth visual experience to a user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to technical fields such as computer vision, artificial intelligence, and image processing, and in particular to a video face-swapping method, device, electronic device, and storage medium. Background Art

[0002] When watching videos, readers often try to immerse themselves in the characters themselves, immersing themselves in their emotions, in order to enhance the viewing experience. With the rapid development of artificial intelligence (AI) technology, more and more applications are now offering face-swapping features. Once a user activates this feature, the application replaces the face in the video with another selected face, achieving a better viewing experience. However, existing technologies still suffer from poor fit and unnatural facial expressions, resulting in a poor visual experience for users. Summary of the Invention

[0003] The present disclosure provides a video face-swapping method, device, electronic device, and storage medium.

[0004] According to one aspect of the present disclosure, a video face-swapping method is provided, comprising:

[0005] Obtain a 3D face model;

[0006] Determine each frame of the image to be processed in the video to be processed;

[0007] Detecting current facial status information of the face to be replaced in each frame of the image to be processed;

[0008] Performing image acquisition on the three-dimensional face model according to the current face state information to obtain a first face image;

[0009] The face to be replaced in the image to be processed is replaced according to the first face image to obtain a second face image after the face is replaced.

[0010] According to another aspect of the present disclosure, a video face-swapping device is provided, comprising:

[0011] An acquisition module is configured to acquire a three-dimensional face model;

[0012] A determination module is configured to determine each frame of an image to be processed in a video to be processed and to be subjected to face swapping;

[0013] A detection module is configured to detect current facial state information of the face to be replaced in each frame of the image to be processed;

[0014] a first image acquisition module configured to acquire an image of the three-dimensional face model according to the current face state information to obtain a first face image;

[0015] The image replacement module is configured to replace the face to be replaced in the image to be processed according to the first face image to obtain a second face image after the face is replaced.

[0016] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0017] at least one processor; and

[0018] a memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any one of the methods described in the above technical solutions.

[0020] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any one of the methods described above.

[0021] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements any one of the methods described above when executed by a processor.

[0022] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.

[0024] Figure 1 1 is a schematic diagram of the steps of the video face-swapping method according to an embodiment of the present disclosure;

[0025] Figure 2 1 is a schematic diagram of the steps of the second video face-swapping method according to an embodiment of the present disclosure;

[0026] Figure 3 This is a schematic diagram of the specific steps of three-dimensional modeling in step S202 in the embodiment of the present disclosure;

[0027] Figure 4 This is a schematic diagram of the specific steps of detecting the current facial state information of the face to be replaced in step S103 in the embodiment of the present disclosure;

[0028] Figure 5 This is a schematic diagram of the specific steps of acquiring the first facial image in step S104 in the embodiment of the present disclosure;

[0029] Figure 6 1 is a schematic diagram of the steps of the third video face-swapping method in the embodiment of the present disclosure;

[0030] Figure 7 1 is a schematic diagram of the overall process of the video face-swapping method in an embodiment of the present disclosure;

[0031] Figure 8 This is a principle block diagram of the video face-swapping device in an embodiment of the present disclosure;

[0032] Figure 9 This is a principle block diagram of a second video face-swapping device according to an embodiment of the present disclosure;

[0033] Figure 10 is a principle block diagram of a creation module in an embodiment of the present disclosure;

[0034] Figure 11 is a principle block diagram of the detection module in an embodiment of the present disclosure;

[0035] Figure 12 is a principle block diagram of the first image acquisition module in an embodiment of the present disclosure;

[0036] Figure 13 This is a principle block diagram of a third video face-swapping device according to an embodiment of the present disclosure;

[0037] Figure 14 3 is a block diagram of an electronic device used to implement the video face-swapping method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0038] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0039] In view of the technical problem of poor video face-swapping effect in the prior art, the present disclosure provides a video face-swapping method, such as Figure 1 As shown, including:

[0040] Step S101: Acquire a three-dimensional face model.

[0041] The three-dimensional face model in this embodiment can be a face model pre-built based on the face image by capturing images of the user's face at different angles through a camera or other equipment with the user's authorization. This three-dimensional face model can simulate the real face shape and support scrolling and switching of the center point to view the face at various angles. For example, if the front face is 18 degrees to the left, the three-dimensional face model is actually rotated 18 degrees to the left on the vertical coordinate (Y coordinate). If you need to see the side face of the head, you need to switch the required angle from the vertical axis and the horizontal axis. Therefore, when rendering the video, you can obtain the required face image at any angle.

[0042] Step S102, determining each frame of an image to be processed in the video to be processed in which the face is to be swapped.

[0043] Not every frame of the video to be processed requires face-changing processing. For example, if the face of character A needs to be replaced, it is possible to determine which image frames contain the face of character A and identify these image frames as images to be processed.

[0044] Step S103: detecting the current facial status information of the face to be replaced in each frame of the image to be processed.

[0045] The facial status information detected in this embodiment may include but is not limited to facial angle, facial proportion, facial color, etc., so as to determine the current status of the face to be replaced.

[0046] Step S104: performing image acquisition on the three-dimensional face model according to the current face state information to obtain a first face image.

[0047] After the state of the face to be replaced is detected, a facial image with a state close to that of the face to be replaced can be collected based on the facial state information, thereby obtaining a facial image that matches the current image to be processed.

[0048] Step S105 , replacing the face to be replaced in the image to be processed according to the first face image to obtain a second face image after face replacement.

[0049] The second facial image in this embodiment refers to the image obtained after the face swap. As previously mentioned, the first facial image is obtained by capturing an image of a 3D facial model based on the state information of the face to be replaced. This image closely matches the state of the face to be replaced. Therefore, using the first facial image allows for better integration with the processed image, resulting in a higher degree of fit with the character to be replaced. The resulting face-swapped video provides the user with a more natural and smooth visual experience.

[0050] As an optional implementation, in step S101, before obtaining the three-dimensional face model, Figure 2 As shown, it also includes:

[0051] Step S201 : collecting facial images of a target user from various angles to obtain a plurality of third facial images at different angles.

[0052] Specifically, when performing video face replacement, the character's front face or side face is generally replaced. Therefore, when collecting images, facial images of the user within 0 to 180 degrees can be collected.

[0053] Step S202 : performing three-dimensional modeling based on the third facial images at various angles to obtain a three-dimensional facial model corresponding to the target user.

[0054] Three-dimensional reconstruction is to build a three-dimensional model of the human face. It has one more dimension than the two-dimensional face image. The three-dimensional face model is a three-dimensional face image, which is equivalent to restoring the three-dimensional image of the user's face based on the collected two-dimensional face image (that is, the third face image). This makes it convenient to obtain face replacement images of different angles or different proportions at any time during video rendering without the need to repeatedly collect the user's face.

[0055] As an optional implementation, step S202, three-dimensional modeling is performed based on the third face image at each angle to obtain a three-dimensional face model corresponding to the target user, such as Figure 3 As shown, including:

[0056] Step S202a: annotate key points of the third face image to obtain corresponding key points.

[0057] When constructing a 3D facial model, in addition to the length and width of each facial organ collected from the front, facial height data, such as the height of the nose bridge, needs to be collected. Therefore, it is necessary to scan a third facial image collected at different angles. The purpose is to obtain data for any key point on the face and convert it into X, Y, and Z coordinates. Therefore, when collecting facial data in this embodiment, key points on the face are annotated from different angles, and the various facial organs are converted into model information composed of numerous 3D points. Taking the nose as an example, when collecting frontal facial information, the width and height of the nose can be obtained, and the X and Y coordinates of various points such as the tip of the nose can be obtained. When collecting profile information, the corresponding Z coordinate is generated, thereby obtaining the X, Y, and Z coordinates of each key point.

[0058] Step S202b: construct a three-dimensional face model based on the position information of each key point.

[0059] According to the X, Y, and Z coordinates of each key point, a concave-convex full data model of the face can be constructed as the three-dimensional face model in this embodiment.

[0060] As an optional implementation, after performing three-dimensional modeling based on the third facial image at each angle to obtain a three-dimensional facial model corresponding to the target user in step S202, the method further includes:

[0061] Cache the 3D face model and generate a corresponding face model list.

[0062] When using the video face swap feature, users can select a 3D face model from a list of face models and update the list at any time. Users can share the 3D face models in the list with friends, sharing the key data of the 3D face model with them, thus enabling the sharing of 3D face models.

[0063] As an optional implementation, Figure 4 As shown, in step S103, detecting the current facial state information of the face to be replaced in each frame of the image to be processed includes at least one of the following:

[0064] Step S103a: detecting the current face ratio of the face to be replaced in each frame of the image to be processed.

[0065] Specifically, the size of the face to be replaced may vary in each frame. In this case, it is necessary to detect the proportions of the face to be replaced so that a facial image of the same or similar proportions can be collected later for replacement. For example, if a frame showing the main character's face is replaced with the user's face, the 3D facial model needs to be enlarged or reduced so that the proportions of the first facial image and the face to be replaced in this frame are the same or similar, ensuring a natural-looking face.

[0066] Step S103b: detecting the current facial angle of the face to be replaced in each frame of the image to be processed.

[0067] Specifically, the angle of the face to be replaced may vary in each frame. In this case, it is necessary to detect the angle of the face to be replaced so that subsequent facial images with the same or similar angles can be collected for replacement. For example, if the face of the protagonist in a certain frame of the video is tilted 18 degrees to the left, then an image of the 3D face model tilted 18 degrees to the left based on this angle can be collected as the first facial image.

[0068] Step S103c: detecting the current facial color of the face to be replaced in each frame of the image to be processed.

[0069] Specifically, the scene or lighting intensity of the face to be replaced in each frame of the processed image may vary. For example, at night, the light is relatively dim, so the facial color of the face to be replaced is darker. Therefore, in this case, the color of the face to be replaced can be detected, which facilitates the subsequent collection of facial images with the same or similar color for replacement. This ensures that the face replaced in different environments can be more seamlessly integrated with the original video.

[0070] As an optional implementation, step S103b, detecting the current facial angle of the face to be replaced in each frame of the image to be processed, includes:

[0071] The symmetry axis of the shoulders of the character corresponding to the face to be replaced is used as the reference axis, and the deflection angle of the face to be replaced relative to the reference axis is detected as the current face angle of the face to be replaced.

[0072] As an optional implementation, step S104, according to the current face state information, the face three-dimensional model is image-captured to obtain a first face image, such as Figure 5 As shown, including:

[0073] Step S104a: adjusting the proportions of the three-dimensional face model according to the current proportions of the face to be replaced.

[0074] As mentioned above, when the facial proportions of the face to be replaced are known, the 3D facial model selected by the user can be enlarged or reduced according to the facial proportions so that the proportions of the 3D facial model are consistent with the facial proportions of the face to be replaced.

[0075] Step S104b: adjusting the angle of the three-dimensional face model according to the current face angle of the face to be replaced.

[0076] When the facial angle of the face to be replaced is known, the three-dimensional facial model selected by the user can be rotated according to the facial angle so that the shooting angle of the three-dimensional facial model is similar or consistent with the facial angle of the face to be replaced. For example, when the side face of the face to be replaced is deflected 18 degrees to the left, a facial image of the three-dimensional facial model with its side face deflected 18 degrees to the left can be collected as the first facial image.

[0077] Step S104c: adjusting the color value of each key point of the 3D facial model according to the current facial color of the face to be replaced.

[0078] When the facial color of the face to be replaced is known, the color values ​​of each key point of the three-dimensional facial model can be adjusted according to the facial color so that the facial color of the three-dimensional facial model is similar or consistent with the facial color of the face to be replaced.

[0079] Step S104d: performing image acquisition on the adjusted three-dimensional face model to obtain a corresponding first face image.

[0080] After adjusting the angle, proportion and color of the three-dimensional facial model through steps S104a to S104c, the first facial image is collected again, and a replacement image that is more consistent with the face to be replaced can be obtained, so that it can be more naturally integrated into the image to be processed after replacement, ensuring that the user's visual experience is natural, smooth and without any sense of incongruity.

[0081] As an optional implementation, step S104, performing image acquisition on the three-dimensional face model according to the current face state information to obtain the first face image includes:

[0082] Collect multiple first face images at different angles within a preset angle range.

[0083] Specifically, the replacement angle of the first facial image we generated may not be perfectly matched. When generating a substitute face, we will additionally generate facial images with facial angles similar to those of the face to be replaced. For example, when the side face of the face to be replaced is deflected 18 degrees to the left, we can collect several facial images with the side face of the three-dimensional facial model deflected 15-20 degrees to the left as the first facial image. When the first facial image with the side face deflected 18 degrees to the left cannot be matched, we can choose the first facial image with the side face deflected 17, 19, etc. to replace it, so as to find the first facial image with the highest degree of fit when rendering the video.

[0084] As an optional implementation, Figure 6 As shown, in step S105, after replacing the face to be replaced in the image to be processed according to the first face image to obtain a second face image after face replacement, the method further includes:

[0085] Step S106 , verifying the second facial image based on the image to be processed to determine whether there is an error between the image to be processed and the second facial image.

[0086] In response to an error between the image to be processed and the second facial image, the process returns to step S105 and selects a first facial image at another angle from the plurality of first facial images at different angles for replacement.

[0087] As mentioned earlier, the replacement angles we generate may not always be a perfect match. Therefore, when generating the first face image, we also generate additional images with similar angles to the face to be replaced. During verification, we compare the pre- and post-replacement images. If any discrepancies are found, we fine-tune the replacement angles based on the comparison results, ultimately finding the most compatible replacement angle. This ensures that users experience no awkwardness throughout the video and experience a greater sense of immersion.

[0088] In combination with the above embodiments, the overall face-changing process of the system disclosed in the present invention is as follows: Figure 7As shown, it may include: step S701, the user selects a video; step S702, selecting to turn on the face-changing mode or the normal viewing mode; step S703, generating a list of characters that support face-changing in the video; step S704, performing facial image acquisition according to the user's instructions to acquire a third facial image; step S705, performing three-dimensional modeling based on the third facial image to obtain a three-dimensional facial model; step S706, caching the three-dimensional facial model to a face model list; step S707, the user selects the corresponding three-dimensional facial model from the face model list; step S708, the user selects the character whose face needs to be changed; step S709, performing face replacement based on the three-dimensional facial model.

[0089] The embodiment of the present disclosure also provides a video face-changing device 800, such as Figure 8 As shown, including:

[0090] The acquisition module 801 is configured to acquire a three-dimensional face model.

[0091] The three-dimensional face model in this embodiment can be a face model pre-built based on the face image by capturing images of the user's face at different angles through a camera or other equipment with the user's authorization. This three-dimensional face model can simulate the real face shape and support scrolling and switching of the center point to view the face at various angles. For example, if the front face is 18 degrees to the left, the three-dimensional face model is actually rotated 18 degrees to the left on the vertical coordinate (Y coordinate). If you need to see the side face of the head, you need to switch the required angle from the vertical axis and the horizontal axis. Therefore, when rendering the video, you can obtain the required face image at any angle.

[0092] The determination module 802 is configured to determine each frame of the image to be processed in the video to be processed in which the face is to be swapped.

[0093] Not every frame of the video to be processed requires face-changing processing. For example, if the face of character A needs to be replaced, it is possible to determine which image frames contain the face of character A and identify these image frames as images to be processed.

[0094] The detection module 803 is configured to detect the current facial state information of the face to be replaced in each frame of the image to be processed.

[0095] The facial status information detected in this embodiment may include but is not limited to facial angle, facial proportion, facial color, etc., so as to determine the current status of the face to be replaced.

[0096] The first image acquisition module 804 is configured to perform image acquisition on the three-dimensional face model according to the current face state information to obtain a first face image.

[0097] After the state of the face to be replaced is detected, a facial image with a state close to that of the face to be replaced can be collected based on the facial state information, thereby obtaining a facial image that matches the current image to be processed.

[0098] The image replacement module 805 is configured to replace the face to be replaced in the image to be processed according to the first face image to obtain a second face image after the face is replaced.

[0099] In this embodiment, the second facial image refers to the image obtained after the face swap. As previously mentioned, the first facial image is obtained by capturing an image of the 3D facial model based on the state information of the face to be replaced. The first facial image is close to the state of the face to be replaced. Therefore, using the first facial image can better integrate with the processed image, and the fit with the character to be replaced is also higher, giving the user a more natural and smooth visual experience.

[0100] As an optional implementation, Figure 9 As shown, the apparatus 800 further includes:

[0101] The second image acquisition module 806 is configured to acquire facial images of the target user from various angles before acquiring the three-dimensional facial model, to obtain a plurality of third facial images at different angles.

[0102] Specifically, when performing video face replacement, the character's front face or side face is generally replaced. Therefore, when collecting images, facial images of the user within 0 to 180 degrees can be collected.

[0103] The creation module 807 is configured to perform three-dimensional modeling based on the third facial images at various angles to obtain a three-dimensional facial model corresponding to the target user.

[0104] Three-dimensional reconstruction is to build a three-dimensional model of the human face. It has one more dimension than the two-dimensional face image. The three-dimensional face model is a three-dimensional face image, which is equivalent to restoring the three-dimensional image of the user's face based on the collected two-dimensional face image (that is, the third face image). This makes it convenient to obtain face replacement images of different angles or different proportions at any time during video rendering without the need to repeatedly collect the user's face.

[0105] As an optional implementation, Figure 10 As shown, the creation module 807 includes:

[0106] The key point labeling unit 807a is configured to label key points of the third face image to obtain corresponding multiple key points.

[0107] When constructing a 3D facial model, in addition to the length and width of each facial organ collected from the front, facial height data, such as the height of the nose bridge, needs to be collected. Therefore, it is necessary to scan a third facial image collected at different angles. The purpose is to obtain data for any key point on the face and convert it into X, Y, and Z coordinates. Therefore, when collecting facial data in this embodiment, key points on the face are annotated from different angles, and the various facial organs are converted into model information composed of numerous 3D points. Taking the nose as an example, when collecting frontal facial information, the width and height of the nose can be obtained, and the X and Y coordinates of various points such as the tip of the nose can be obtained. When collecting profile information, the corresponding Z coordinate is generated, thereby obtaining the X, Y, and Z coordinates of each key point.

[0108] The construction unit 807b is configured to construct a three-dimensional face model based on the position information of each key point.

[0109] According to the X, Y, and Z coordinates of each key point, a concave-convex full data model of the face can be constructed as the three-dimensional face model in this embodiment.

[0110] As an optional implementation, the apparatus 800 further includes:

[0111] The cache module is configured to perform three-dimensional modeling based on the facial images of the target user at various angles, cache the three-dimensional facial model after obtaining the three-dimensional facial model corresponding to the target user, and generate a corresponding facial model list.

[0112] When using the video face swap feature, users can select a 3D face model from a list of face models and update the list at any time. Users can share the 3D face models in the list with friends, sharing the key data of the 3D face model with them, thus enabling the sharing of 3D face models.

[0113] As an optional implementation, Figure 11 As shown, the detection module 803 includes at least one of the following:

[0114] The first detection unit 803a is configured to detect the current face ratio of the face to be replaced in each frame of the image to be processed.

[0115] Specifically, the size of the face to be replaced may vary in each frame. In this case, it is necessary to detect the proportions of the face to be replaced so that a facial image of the same or similar proportions can be collected later for replacement. For example, if a frame showing the main character's face is replaced with the user's face, the 3D facial model needs to be enlarged or reduced so that the proportions of the first facial image and the face to be replaced in this frame are the same or similar, ensuring a natural-looking face.

[0116] The second detection unit 803b is configured to detect the current face angle of the face to be replaced in each frame of the image to be processed.

[0117] Specifically, the angle of the face to be replaced may vary in each frame. In this case, it is necessary to detect the angle of the face to be replaced so that subsequent facial images with the same or similar angles can be collected for replacement. For example, if the face of the protagonist in a certain frame of the video is tilted 18 degrees to the left, then an image of the 3D face model tilted 18 degrees to the left based on this angle can be collected as the first facial image.

[0118] The third detection unit 803c is configured to detect the current facial color of the face to be replaced in each frame of the image to be processed.

[0119] Specifically, the scene or lighting intensity of the face to be replaced in each frame of the processed image may vary. For example, at night, the light is relatively dim, so the facial color of the face to be replaced is darker. Therefore, in this case, the color of the face to be replaced can be detected, which facilitates the subsequent collection of facial images with the same or similar color for replacement. This ensures that the face replaced in different environments can be more seamlessly integrated with the original video.

[0120] As an optional implementation manner, the second detection unit 803b detects the current facial angle of the face to be replaced in each frame of the image to be processed, including:

[0121] The symmetry axis of the shoulders of the character corresponding to the face to be replaced is used as the reference axis, and the deflection angle of the face to be replaced relative to the reference axis is detected as the current face angle of the face to be replaced.

[0122] As an optional implementation, Figure 12 As shown, the first image acquisition module 804 includes:

[0123] The first adjusting unit 804a is configured to adjust the proportion of the three-dimensional face model according to the current face proportion of the face to be replaced.

[0124] As mentioned above, when the facial proportions of the face to be replaced are known, the 3D facial model selected by the user can be enlarged or reduced according to the facial proportions so that the proportions of the 3D facial model are consistent with the facial proportions of the face to be replaced.

[0125] The second adjusting unit 804b is configured to adjust the angle of the three-dimensional face model according to the current face angle of the face to be replaced.

[0126] When the facial angle of the face to be replaced is known, the three-dimensional facial model selected by the user can be rotated according to the facial angle so that the shooting angle of the three-dimensional facial model is similar or consistent with the facial angle of the face to be replaced. For example, when the side face of the face to be replaced is deflected 18 degrees to the left, a facial image of the three-dimensional facial model with its side face deflected 18 degrees to the left can be collected as the first facial image.

[0127] The third adjustment unit 804c is configured to adjust the color value of each key point of the 3D facial model according to the current facial color of the face to be replaced.

[0128] When the facial color of the face to be replaced is known, the color values ​​of each key point of the three-dimensional facial model can be adjusted according to the facial color so that the facial color of the three-dimensional facial model is similar or consistent with the facial color of the face to be replaced.

[0129] The image acquisition unit 804d is configured to perform image acquisition on the adjusted three-dimensional face model to obtain a corresponding first face image.

[0130] After adjusting the angle, proportion and color of the three-dimensional facial model through steps S104a to S104c, the first facial image is collected again, and a replacement image that is more consistent with the face to be replaced can be obtained, so that it can be more naturally integrated into the image to be processed after replacement, ensuring that the user's visual experience is natural, smooth and without any sense of incongruity.

[0131] As an optional implementation, the first image acquisition module 804 acquires an image of the three-dimensional face model according to the current face state information, and obtains the first face image including:

[0132] Collect multiple first face images at different angles within a preset angle range.

[0133] Specifically, the replacement angle of the first facial image we generated may not be perfectly matched. When generating a substitute face, we will additionally generate facial images with facial angles similar to those of the face to be replaced. For example, when the side face of the face to be replaced is deflected 18 degrees to the left, we can collect several facial images with the side face of the three-dimensional facial model deflected 15-20 degrees to the left as the first facial image. When the first facial image with the side face deflected 18 degrees to the left cannot be matched, we can choose the first facial image with the side face deflected 17, 19, etc. to replace it, so as to find the first facial image with the highest degree of fit when rendering the video.

[0134] As an optional implementation, Figure 13 As shown, the apparatus 800 further includes:

[0135] The verification module 808 is configured as an image replacement module to replace the face to be replaced in the image to be processed according to the first face image, and after obtaining the second face image after the face change, verify the second face image according to the image to be processed to determine whether there is any error between the image to be processed and the second face image.

[0136] In response to an error between the image to be processed and the second facial image, the image replacement module 805 selects a first facial image at another angle from the plurality of first facial images at different angles for replacement.

[0137] As mentioned earlier, the replacement angles we generate may not always be a perfect match. Therefore, when generating the first face image, we also generate additional images with similar angles to the face to be replaced. During verification, we compare the pre- and post-replacement images. If any discrepancies are found, we fine-tune the replacement angles based on the comparison results, ultimately finding the most compatible replacement angle. This ensures that users experience no awkwardness throughout the video and experience a greater sense of immersion.

[0138] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0139] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0140] Figure 14 A schematic block diagram of an example electronic device 1400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0141] like Figure 14As shown, device 1400 includes a computing unit 1401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1402 or a computer program loaded from a storage unit 1408 into a random access memory (RAM) 1403. Various programs and data required for the operation of device 1400 can also be stored in RAM 1403. Computing unit 1401, ROM 1402, and RAM 1403 are connected to each other via a bus 1404. An input / output (I / O) interface 1405 is also connected to bus 1404.

[0142] Various components in device 1400 are connected to I / O interface 1405, including: an input unit 1406, such as a keyboard, mouse, etc.; an output unit 1407, such as various types of displays, speakers, etc.; a storage unit 1408, such as a magnetic disk, optical disk, etc.; and a communication unit 1409, such as a network card, modem, wireless communication transceiver, etc. Communication unit 1409 allows device 1400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0143] Computing unit 1401 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of computing unit 1401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning objective function algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. Computing unit 1401 performs the various methods and processes described above, such as the video face-swapping method. For example, in some embodiments, the video face-swapping method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as storage unit 1408. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1400 via ROM 1402 and / or communication unit 1409. When the computer program is loaded into RAM 1403 and executed by computing unit 1401, one or more steps of the video face-swapping method described above can be performed. Alternatively, in other embodiments, the computing unit 1401 may be configured to perform the video face-swapping method in any other appropriate manner (eg, by means of firmware).

[0144] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0145] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0146] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0147] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0148] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0149] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0150] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions of this disclosure can be achieved, and this document is not limited here.

[0151] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A video face swapping method, comprising: Obtain a 3D face model; Determine each frame of the image to be processed in the video to be processed; Detecting current facial status information of the face to be replaced in each frame of the image to be processed; Performing image acquisition on the three-dimensional face model according to the current face state information to obtain a first face image; The face to be replaced in the image to be processed is replaced according to the first face image to obtain a second face image after the face is replaced.

2. The method according to claim 1, wherein Before obtaining the three-dimensional face model, the method further includes: Collect the target user's facial image from various angles to obtain multiple third-party facial images at different angles; Three-dimensional modeling is performed based on the third facial image at each angle to obtain the three-dimensional facial model corresponding to the target user.

3. The method according to claim 2, wherein: The performing three-dimensional modeling based on the third facial image at each angle to obtain the three-dimensional facial model corresponding to the target user includes: Marking key points of the third face image to obtain a plurality of key points; The three-dimensional face model is constructed based on the position information of each key point.

4. The method according to claim 2 or 3, wherein: After performing three-dimensional modeling based on the third facial image at each angle to obtain the three-dimensional facial model corresponding to the target user, the method further includes: The three-dimensional face model is cached and a corresponding face model list is generated.

5. The method according to claim 1, wherein The detecting of the current facial status information of the face to be replaced in each frame of the image to be processed includes at least one of the following: Detecting the current face ratio of the face to be replaced in each frame of the image to be processed; Detecting the current facial angle of the face to be replaced in each frame of the image to be processed; Detect the current facial color of the face to be replaced in each frame of the image to be processed.

6. The method according to claim 5, wherein: The detecting the current facial angle of the face to be replaced in each frame of the image to be processed includes: The symmetry axis of the shoulders of the character corresponding to the face to be replaced is used as a reference axis, and the deflection angle of the face to be replaced relative to the reference axis is detected as the current face angle of the face to be replaced.

7. The method according to claim 5, wherein: The step of performing image acquisition on the three-dimensional face model according to the current face state information to obtain a first face image includes: Adjusting the proportion of the three-dimensional face model according to the current facial proportions of the face to be replaced; Adjusting the angle of the three-dimensional face model according to the current face angle of the face to be replaced; Adjusting the color value of each key point of the three-dimensional face model according to the current facial color of the face to be replaced; Perform image acquisition on the adjusted three-dimensional face model to obtain the corresponding first face image.

8. The method according to any one of claims 1 to 7, wherein: The performing image acquisition on the three-dimensional face model according to the current face state information to obtain the first face image includes: Collect multiple first face images at different angles within a preset angle range.

9. The method according to claim 8, wherein After replacing the face to be replaced in the image to be processed according to the first facial image to obtain a second facial image after the face is replaced, the method further includes: verifying the second facial image based on the image to be processed to determine whether there is an error between the image to be processed and the second facial image; In response to an error between the image to be processed and the second facial image, the first facial image at another angle is selected from a plurality of first facial images at different angles for replacement.

10. A video face-swapping device, comprising: An acquisition module is configured to acquire a three-dimensional face model; A determination module is configured to determine each frame of an image to be processed in a video to be processed and to be subjected to face swapping; A detection module is configured to detect current facial state information of the face to be replaced in each frame of the image to be processed; a first image acquisition module configured to acquire an image of the three-dimensional face model according to the current face state information to obtain a first face image; The image replacement module is configured to replace the face to be replaced in the image to be processed according to the first face image to obtain a second face image after the face is replaced.

11. The device according to claim 10, wherein Also includes: The second image acquisition module is configured to acquire facial images of the target user from various angles before acquiring the three-dimensional facial model, so as to obtain a plurality of third facial images at different angles; The creation module is configured to perform three-dimensional modeling based on the third facial image at each angle to obtain the three-dimensional facial model corresponding to the target user.

12. The device according to claim 11, wherein The creation module includes: a key point labeling unit, configured to label the third face image with key points to obtain a plurality of corresponding key points; The construction unit is configured to construct the three-dimensional face model according to the point information of each key point.

13. The device according to claim 11 or 12, wherein: Also includes: The cache module is configured to perform three-dimensional modeling based on the third facial image at each angle, obtain the three-dimensional facial model corresponding to the target user, cache the three-dimensional facial model, and generate a corresponding facial model list.

14. The device according to claim 10, wherein The detection module includes at least one of the following: a first detection unit, configured to detect a current face ratio of a face to be replaced in each frame of the image to be processed; a second detection unit configured to detect a current facial angle of a face to be replaced in each frame of the image to be processed; The third detection unit is configured to detect the current facial color of the face to be replaced in each frame of the image to be processed.

15. The device according to claim 14, wherein The second detection unit detects the current face angle of the face to be replaced in each frame of the image to be processed, including: The symmetry axis of the shoulders of the character corresponding to the face to be replaced is used as a reference axis, and the deflection angle of the face to be replaced relative to the reference axis is detected as the current face angle of the face to be replaced.

16. The device according to claim 14, wherein The first image acquisition module includes: a first adjusting unit, configured to adjust the proportion of the three-dimensional face model according to the current facial proportion of the face to be replaced; a second adjusting unit, configured to adjust the angle of the three-dimensional face model according to the current face angle of the face to be replaced; a third adjustment unit, configured to adjust the color value of each key point of the three-dimensional face model according to the current facial color of the face to be replaced; The image acquisition unit is configured to perform image acquisition on the adjusted three-dimensional face model to obtain the corresponding first face image.

17. The device according to any one of claims 10 to 16, wherein: The first image acquisition module acquires an image of the three-dimensional face model according to the current face state information to obtain a first face image, which includes: Collect multiple first face images at different angles within a preset angle range.

18. The device according to claim 17, wherein Also includes: a verification module configured to, after the image replacement module replaces the face to be replaced in the image to be processed based on the first facial image, obtain a second facial image after the face replacement, and then verify the second facial image based on the image to be processed to determine whether there is any error between the image to be processed and the second facial image; In response to an error between the image to be processed and the second facial image, the image replacement module selects the first facial image at another angle from a plurality of the first facial images at different angles for replacement.

19. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.

20. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.

21. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9.