A virtual object face changing method and device, computer equipment and storage medium

By acquiring the mapping relationship between two-dimensional human face images and the human face shape of three-dimensional virtual objects, and using a trained facial feature prediction model to predict the target texture and shape feature information, the problem of low efficiency of 3D character face swapping in existing technologies for 3D games and virtual social interactions is solved, and fast and accurate 3D character face swapping is achieved.

CN116168177BActive Publication Date: 2026-03-24GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-23
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing face-swapping methods are mainly based on 2D images, which cannot be efficiently applied to 3D character face-swapping in 3D games and virtual social interactions, as they require manual sculpting, resulting in low efficiency.

Method used

By acquiring the mapping relationship between the face shape of a two-dimensional face image and a three-dimensional virtual object, the texture and shape feature information of the target are predicted using a trained face feature prediction model to determine the target three-dimensional dynamic face model, and face swapping is performed based on this.

Benefits of technology

It improves the efficiency of face swapping for 3D virtual objects, enabling fast and accurate face swapping for 3D characters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116168177B_ABST
    Figure CN116168177B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a virtual object face changing method and device, computer equipment and a storage medium, which can quickly determine the target face shape of the face image based on the face shape mapping relationship set of the virtual object to be changed and the target three-dimensional dynamic face model corresponding to the face shape feature information of the face image, so that the face changing of the virtual object to be changed can be performed based on the target face shape and the target texture, and the face changing efficiency of the three-dimensional virtual object is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, specifically to a method, apparatus, computer device, and storage medium for virtual object face swapping, wherein the storage medium is a computer-readable storage medium. Background Technology

[0002] With the rise of 3D games and virtual social networking, replacing the faces of 3D characters with the user's own face is a development trend in 3D games and virtual social networking.

[0003] Existing face-swapping methods are all based on 2D images. These methods cannot be applied to face-swapping of 3D characters in 3D games or cartoon characters in virtual social interactions. Currently, the faces of 3D characters in 3D games and cartoon characters in virtual social interactions are mainly created manually by 3D modelers. If a character's face needs to be changed, it needs to be recreated, making this method inefficient for face-swapping 3D characters. Summary of the Invention

[0004] This application provides a method for face-swapping virtual objects, a computer device, and a computer-readable storage medium, which can improve the efficiency of face-swapping for three-dimensional virtual objects.

[0005] A method for face swapping of virtual objects, comprising:

[0006] Acquire a two-dimensional face image and acquire a set of face shape mapping relationships for a three-dimensional virtual object to be face-swapped. The set of face shape mapping relationships includes the mapping relationship between the preset face shape of the virtual object before face swapping and the preset three-dimensional dynamic face model.

[0007] The trained face feature prediction model is used to predict the face image to obtain the target texture and face shape feature information of the face image;

[0008] Based on the facial shape feature information, a target three-dimensional dynamic facial model based on the facial shape feature information is determined;

[0009] Based on the target 3D dynamic face model, a preset face shape of a preset 3D dynamic face model that matches the target 3D dynamic face model is determined from the face shape mapping relationship set, so as to use the preset face shape as the target face shape of the face image.

[0010] The face of the virtual object to be swapped is swapped based on the target face shape and the target texture.

[0011] Accordingly, embodiments of this application provide a virtual object face-swapping device, including:

[0012] The acquisition unit is used to acquire a two-dimensional face image and a set of face shape mapping relationships for a three-dimensional virtual object to be face-swapped. The set of face shape mapping relationships includes the mapping relationship between the preset face shape of the virtual object to be face-swapped before face swapping and the preset three-dimensional dynamic face model.

[0013] The prediction unit is used to predict the face image using a trained face feature prediction model to obtain the target texture and face shape feature information of the face image.

[0014] The first determining unit is used to determine the target three-dimensional dynamic face model based on the face shape feature information.

[0015] The second determining unit is used to determine, based on the target three-dimensional dynamic face model, a preset face shape of a preset three-dimensional dynamic face model that matches the target three-dimensional dynamic face model from the face shape mapping relationship set, so as to use the preset face shape as the target face shape of the face image.

[0016] The face-swapping unit is used to swap the face of the virtual object to be swapped based on the target face shape and the target texture.

[0017] In some embodiments, the prediction unit may be specifically used to perform texture prediction on the face image using the trained face texture prediction model to obtain the target texture of the face image; and to perform shape prediction on the face image using the trained face shape prediction model to obtain the face shape feature information of the face image.

[0018] In some embodiments, the trained face shape prediction model includes a trained initial shape prediction sub-model and a trained shape transfer sub-model; the prediction unit can be specifically used to perform initial shape prediction on the face image using the trained initial shape prediction sub-model to obtain candidate face shape feature information of the face image; and to encode the candidate face shape feature information using the trained shape transfer sub-model to obtain the face shape feature information.

[0019] In some embodiments, the virtual object face-swapping device further includes a training unit, which can be used to acquire face sample images and acquire real face shape feature information; use an initial shape prediction sub-model to be trained to predict the face sample images to obtain reference face shape feature information; and train the initial shape prediction sub-model to be trained based on the real face shape feature information and the reference face shape feature information to obtain a trained initial shape prediction sub-model.

[0020] In some embodiments, the training unit may specifically be used to determine an initial three-dimensional dynamic face model based on the reference face shape feature information; extract features from the face sample image to obtain an initial texture of the face sample image; determine a candidate face image based on the initial three-dimensional dynamic face model and the initial texture; and train the initial shape prediction sub-model to be trained based on the real face shape feature information, the reference face shape feature information, the face sample image, and the candidate face image to obtain a trained initial shape prediction sub-model.

[0021] In some embodiments, the training unit may be used to acquire labeled face shape feature information; encode the reference face shape feature information using a shape transfer sub-model to be trained to obtain encoded face shape feature information; and train the shape transfer sub-model to be trained based on the encoded face shape feature information and the labeled face shape feature information to obtain the trained shape transfer sub-model.

[0022] In some embodiments, the training unit may specifically be used to acquire a face sample image and acquire a first texture, the texture type of which matches the texture type of the virtual object to be swapped; use the face texture prediction model to be trained to extract features from the face sample image to obtain a second texture, the texture type of which is different from that of the first texture; use the face texture prediction model to be trained to encode the texture type of the first texture to obtain an encoded first texture, the texture type of which does not match that of the second texture; use the face texture prediction model to be trained to encode the texture type of the second texture to obtain an encoded second texture, the texture type of which does not match that of the second texture; and train the face texture prediction model to be trained based on the first texture, the second texture, the encoded first texture, and the encoded second texture to obtain the trained face texture prediction model.

[0023] In some embodiments, the training unit may specifically be used to encode the texture type of the first texture using the face texture prediction model to be trained, to obtain an encoded third texture, wherein the texture type of the encoded third texture matches the texture type of the first texture; to encode the texture type of the second texture using the face texture prediction model to be trained, to obtain an encoded fourth texture, wherein the texture type of the encoded fourth texture matches the texture type of the second texture; and to train the face texture prediction model to be trained based on the first texture, the second texture, the encoded first texture, the encoded second texture, the encoded third texture, and the encoded third texture, to obtain the trained face texture prediction model.

[0024] In some embodiments, the training unit may specifically be used to calculate a first loss information of the face texture prediction model to be trained based on the first texture and the encoded first texture; calculate a second loss information of the face texture prediction model to be trained based on the first texture and the encoded third texture; calculate a third loss information of the face texture prediction model to be trained based on the second texture and the encoded second texture; calculate a fourth loss information of the face texture prediction model to be trained based on the second texture and the encoded fourth texture; and train the face texture prediction model to be trained using the first loss information, the second loss information, the third loss information, and the fourth loss information to obtain the trained face texture prediction model.

[0025] In some embodiments, the acquisition unit may be specifically used to acquire key points of the face of the virtual object to be face-swapped and a preset face shape of the virtual object before face-swapping; based on the key points, determine the object face shape feature information of the virtual object to be face-swapped; based on the object face shape feature information, determine a preset three-dimensional dynamic face model of the virtual object to be face-swapped; and based on the preset face shape and the preset three-dimensional dynamic face model, generate the face shape mapping relationship set.

[0026] In some embodiments, the acquisition unit may be specifically used to acquire the mesh vertices of a preset face shape and the mesh vertices of the preset three-dimensional dynamic face model; and to determine the face shape mapping relationship set based on the mesh vertices of the preset face shape and the mesh vertices of the preset three-dimensional dynamic face model.

[0027] Furthermore, this application also provides a computer device, including a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute any of the virtual object face-swapping methods provided in this application.

[0028] Furthermore, embodiments of this application also provide a computer-readable storage medium storing a computer program adapted for loading by a processor to execute any of the virtual object face-swapping methods provided in embodiments of this application.

[0029] This application embodiment can acquire a two-dimensional face image and a set of three-dimensional face shape mapping relationships for a virtual object to be face-swapped. The set of face shape mapping relationships includes the mapping relationship between a preset face shape of the virtual object before face-swapping and a preset three-dimensional dynamic face model. A trained face feature prediction model is used to predict the face image to obtain the target texture and face shape feature information of the face image. Based on the face shape feature information, a target three-dimensional dynamic face model is determined. Based on the target three-dimensional dynamic face model, a preset face shape of the preset three-dimensional dynamic face model matching the target three-dimensional dynamic face model is determined from the set of face shape mapping relationships, so that the preset face shape is used as the target face shape of the face image. Based on the target face shape and the target texture, the face of the virtual object to be face-swapped is swapped. Since the embodiments of this application can quickly determine the target face shape of the face image based on the set of face shape mapping relationships of the virtual object to be face-swapping and the target three-dimensional dynamic face model corresponding to the face shape feature information of the face image, the face-swapping of the virtual object to be face-swapping can be performed based on the target face shape and target texture, thereby improving the face-swapping efficiency of three-dimensional virtual objects. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 This is a schematic diagram of a scenario for the virtual object face-swapping method provided in an embodiment of this application;

[0032] Figure 2 This is a flowchart illustrating the virtual object face-swapping method provided in this application embodiment;

[0033] Figure 3 This is a flowchart illustrating the process of obtaining a set of facial shape mapping relationships for a three-dimensional virtual object to be face-swapped, as provided in an embodiment of this application.

[0034] Figure 4 These are two schematic flowcharts illustrating the virtual object face-swapping method provided in this application embodiment;

[0035] Figure 5 This is a schematic diagram of the process for training the initial shape prediction sub-model to be trained, provided in an embodiment of this application.

[0036] Figure 6 This is a schematic diagram of the process for training the shape transfer sub-model to be trained, provided in an embodiment of this application.

[0037] Figure 7 This is a schematic diagram of the process for training the shape transfer sub-model to be trained, provided in an embodiment of this application.

[0038] Figure 8 This is a schematic diagram of the structure of the virtual object face-swapping device provided in the embodiments of this application;

[0039] Figure 9 This is a schematic diagram of the structure of the computer device provided in the embodiments of this application. Detailed Implementation

[0040] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0041] This application provides an image feature processing method and related apparatus. The related apparatus includes an image feature processing device, a computer device, and a computer-readable storage medium. The image feature processing device can be integrated into the computer device, which can be a server or a terminal, etc.

[0042] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein.

[0043] For example, see Figure 1Taking the integration of a virtual object face-swapping device into a computer device as an example, the process involves acquiring a two-dimensional face image and a set of three-dimensional face shape mapping relationships for the virtual object to be face-swapped; using a trained face feature prediction model to predict the face image, obtaining the target texture and face shape feature information of the face image; determining the target three-dimensional dynamic face model based on the face shape feature information; determining the preset face shape of a preset three-dimensional dynamic face model that matches the target three-dimensional dynamic face model from the set of face shape mapping relationships, using the preset face shape as the target face shape of the face image; and performing face-swapping on the virtual object to be face-swapped based on the target face shape and target texture.

[0044] The face shape mapping relationship set includes the mapping relationship between the preset face shape of the virtual object to be face-swapped and the preset 3D dynamic face model.

[0045] The face image can be a real-life face image or a cartoon face image.

[0046] The preset face shape can be represented in the form of a grid.

[0047] Among them, the 3D dynamic face model, also known as the 3D deformable face model or 3DMM model, includes a preset 3D dynamic face model and a target 3D dynamic face model.

[0048] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.

[0049] This embodiment will be described from the perspective of a virtual object face-swapping device, which can be integrated into a computer device, such as a server or a terminal. The terminal can include tablet computers, laptops, personal computers (PCs), wearable devices, virtual reality devices, or other smart devices that can acquire data.

[0050] like Figure 2 As shown, the specific process of this virtual object face-swapping method is as follows:

[0051] S101. Obtain a two-dimensional face image and a set of face shape mapping relationships for the three-dimensional virtual object to be swapped.

[0052] The face shape mapping relationship set includes the mapping relationship between the preset face shape of the virtual object to be face-swapped and the preset 3D dynamic face model.

[0053] The preset face is represented in the form of a mesh. The preset 3D dynamic face model refers to the preset 3DMM model.

[0054] like Figure 3 As shown in the embodiments of this application, the process of obtaining the set of facial shape mapping relationships of the three-dimensional virtual object to be face-swapping can be as follows:

[0055] A1. Obtain the key points of the face of the virtual object to be face-swapped and the preset face shape of the virtual object before face-swapping.

[0056] Among them, such as Figure 4 As shown, this embodiment of the application derives the index numbers of all vertices of the mesh of the head of the virtual object to be face-swapping, and also extracts the index numbers corresponding to the 68 key points of the face of the virtual object. The index numbers of all vertices of the mesh and the index numbers corresponding to the 68 key points can be exported from the software that creates the virtual object to be face-swapping.

[0057] Based on this, the computer device can extract the mesh of the head of the virtual object to be face-swapped, as well as the coordinates of 68 key points, according to the index number. The mesh of the head of the virtual object to be face-swapped forms the preset face shape of the virtual object before face-swapping.

[0058] A2. Based on key points, determine the facial shape feature information of the virtual object to be face-swapped.

[0059] This application embodiment solves for the object's facial shape feature information through 68 key points. The object's facial feature information includes a rotation angle θ. * and deformation coefficient β * The facial shape feature information of the object can be obtained by formula (1) in this embodiment.

[0060] θ * ,β * =argmin||Kps character -R(θ)J kps (M(β))|| 2 Formula (1)

[0061] Where, θ * This refers to the rotation angle, β * θ refers to the deformation coefficient; θ refers to the initial rotation angle, and β refers to the initial deformation coefficient. θ and β are generally set to 0; Kps character This refers to the coordinates of 68 key points; R refers to the function for constructing the rotation posture of the 3D dynamic face model. Inputting the initial rotation angle into R yields R(θ). Based on this, R(θ) refers to the rotation posture of the 3D dynamic face model corresponding to θ, that is, the rotation posture of the virtual object to be swapped into the 3D dynamic face model; J kpsM refers to the regression function of 68 key points; M refers to the modeling function of the 3D dynamic face model. The initial deformation coefficient β is input into M, which is M(β). Based on this, M(β) refers to the mesh of the 3D dynamic face model corresponding to the initial deformation coefficient β, that is, the face shape of the 3D dynamic face model, which is also the face shape of the virtual object to be face swapped into the 3D dynamic face model.

[0062] Furthermore, embodiments of this application also address the rotation angle θ. * and deformation coefficient β * Optimization is performed. In this embodiment, 3000 vertices are uniformly sampled from the mesh of the virtual object to be face-swapped, and 3000 vertices are uniformly sampled from the mesh of the 3D dynamic face model. Since the number and order of vertices in the mesh of the virtual object to be face-swapped and the mesh of the 3D dynamic face model may not be the same, this embodiment ensures smooth subsequent calculations by uniformly sampling the same number of vertices. Since the order relationship between the vertices in the mesh of the virtual object to be face-swapped and the mesh of the 3D dynamic face model has not yet been determined, this embodiment uses an ICP (Iterative Nearest Point) alternating iteration method to adjust the rotation angle θ. * and deformation coefficient β * Optimization is performed to obtain the optimized rotation angle θ1. * and the optimized deformation coefficient β * The optimized formula is shown in formula (2):

[0063]

[0064] Where, θ1 * This refers to the optimized rotation angle, β1 * This refers to the optimized deformation coefficient; This refers to the mesh composed of 3000 vertices of the virtual object to be face-swapped; M 3000 This refers to the modeling function for 3000 vertices of a 3D dynamic face model, in M. 3000 The deformation coefficient β was input. * That is, M 3000 (β * Based on this, M 3000 (β * ) refers to the deformation coefficient β * The mesh of the corresponding 3D dynamic face model.

[0065] A3. Based on the facial shape feature information of the object, determine the preset three-dimensional dynamic facial model of the virtual object to be face-swapped.

[0066] In this embodiment of the application, the rotation angle θ *The input is fed into the function R that constructs the rotation pose of the 3D dynamic face model to obtain the rotation pose of the 3D dynamic face model; the deformation coefficient β is then used to... * The input is fed into the modeling function M of the 3D dynamic face model to obtain the face shape of the 3D dynamic face model, which is also the mesh of the face shape of the 3D dynamic face model. Based on the mesh of the face shape of the 3D dynamic face model and the rotation posture of the 3D dynamic face model, the preset 3D dynamic face model of the virtual object to be swapped is determined.

[0067] Of course, the optimized rotation angle θ1 can also be used in the embodiments of this application. * and the optimized deformation coefficient β1 * The rotation posture and face shape of the 3D dynamic face model are determined, as detailed above, and will not be repeated here.

[0068] A4. Generate a set of face shape mapping relationships based on preset face shapes and preset 3D dynamic face models.

[0069] In this context, it can be understood that the set of face shape mapping relationships includes the mapping relationship between preset face shapes and preset 3D dynamic face models.

[0070] This application embodiment establishes a mapping relationship between the mesh vertices of a preset human face shape and the mesh vertices of a preset three-dimensional dynamic human face model.

[0071] Specifically, for each grid vertex of the preset face shape, the three grid vertices that are closest to the grid vertices of the preset 3D dynamic face model and the preset face shape are determined.

[0072] In this embodiment, the three closest grid vertices of the preset 3D dynamic face model and the preset face shape can be determined by Euclidean distance, and weights are assigned to the three closest grid vertices of the preset 3D dynamic face model and the preset face shape, and the index numbers of these three grid vertices are recorded.

[0073] In this embodiment of the application, the mapping relationship between each grid vertex of the preset face shape and the grid vertex of the preset three-dimensional dynamic face model can be represented by formula (3), which is as follows:

[0074]

[0075] Among them, P character This refers to the coordinates of the vertices of the grid representing the pre-defined face shape; w i This refers to the weight of the i-th grid vertex in the preset 3D dynamic face model, w i =exp(-d i ), d iIt refers to the Euclidean distance from the i-th grid vertex of the preset 3D dynamic face model to the grid vertex of the preset face shape; It refers to the coordinates of the i-th grid vertex of the preset 3D dynamic face model; i is a positive integer.

[0076] Based on the above formula (3), the embodiments of this application can determine the mapping relationship between the preset face shape and the preset three-dimensional dynamic face model. That is, the computer obtains the mesh vertices of the preset face shape and the mesh vertices of the preset three-dimensional dynamic face model; based on the mesh vertices of the preset face shape and the mesh vertices of the preset three-dimensional dynamic face model, the set of face shape mapping relationships is determined.

[0077] S102. The trained face feature prediction model is used to predict the face image to obtain the target texture and face shape feature information of the face image.

[0078] In this application embodiment, the trained face feature prediction model can be a single model, which predicts the target texture and face shape feature information.

[0079] The trained face feature prediction model in this application embodiment may include multiple trained models, such as a trained face texture prediction model and a trained face shape prediction model.

[0080] In this embodiment, a trained face texture prediction model is used to predict the texture of a face image to obtain the target texture of the face image; a trained face shape prediction model is used to predict the shape of a face image to obtain the face shape feature information of the face image.

[0081] For example, such as Figure 4 As shown, in this embodiment of the application, a trained face texture prediction model can be used to extract features from a face image to obtain a face texture. Here, the texture type of the face texture matches the type of the face image. For example, if the face image is a real person type face image, the texture type of the face texture is a real person texture type. In this case, it can be said that the texture type of the face texture matches the type of the face image. The trained face texture prediction model can then be used to predict the texture type of the face texture to obtain the target texture of the face image. The texture type of the target texture is different from the texture type of the face texture. For example, if the texture type of the target texture is a cartoon texture type, and the texture type of the face texture is a real person texture type, in this case, it can be said that the texture type of the target texture is different from the texture type of the face texture.

[0082] In this embodiment, other face texture extraction models can also be used to extract features from face images to obtain face textures, where the texture type of the face texture matches the type of the face image.

[0083] In this embodiment of the application, the trained face shape prediction model can be a single model, which predicts face shape feature information.

[0084] In this application embodiment, the trained face shape prediction model may include multiple trained models. For example, the trained face shape prediction model includes a trained initial shape prediction sub-model and a trained shape transfer sub-model.

[0085] The process of predicting the shape of a face image to obtain the face shape feature information in this embodiment of the application can be as follows:

[0086] Specifically, for example, such as Figure 4 As shown, the computer device uses the initial shape prediction sub-model after training to predict the face image and obtain the candidate face shape feature information of the face image; it then uses the shape transfer sub-model after training to encode the candidate face shape feature information and obtain the face shape feature information.

[0087] In this embodiment of the application, before using a trained initial shape prediction sub-model to perform initial shape prediction on a face image and obtain the candidate face shape feature information, the initial shape prediction sub-model to be trained can be trained to obtain a trained initial shape prediction sub-model. The training process of the initial shape prediction sub-model to be trained in this embodiment of the application can be as follows:

[0088] For example, a computer device acquires a face sample image and obtains real face shape feature information; the initial shape prediction sub-model to be trained is used to predict the face sample image to obtain reference face shape feature information; based on the real face shape feature information and the reference face shape feature information, the initial shape prediction sub-model to be trained is trained to obtain the trained initial shape prediction sub-model.

[0089] The initial shape prediction sub-model to be trained can be a convolutional neural network model.

[0090] Among them, the face sample images are real-life face sample images.

[0091] like Figure 5 As shown, in this embodiment of the application, the Backbone Net of the convolutional neural network model is used to extract features from the face sample image, and the output layer processes the features to obtain reference face shape feature information.

[0092] The output layer is as follows Figure 5 The βNet (i.e., the β network) references facial shape features, including the deformation coefficient β2. The actual facial shape features are... Figure 5 The beta labels in the data are the beta labels of real people.

[0093] This application embodiment calculates the loss value between the deformation coefficient β2 and the β label. Based on this loss value, the initial shape prediction sub-model to be trained is trained until it converges, thus obtaining the trained initial shape prediction sub-model. This application embodiment minimizes the loss value between the deformation coefficient β2 and the β label during training.

[0094] Furthermore, in this embodiment of the application, after predicting the face sample image using the initial shape prediction sub-model to be trained to obtain the reference face shape feature information, it further includes:

[0095] The computer device determines an initial three-dimensional dynamic face model based on the reference face shape feature information; extracts features from the face sample image to obtain the initial texture of the face sample image; and determines candidate face images based on the initial three-dimensional dynamic face model and the initial texture.

[0096] For example, such as Figure 5 As shown, in this embodiment, the predicted deformation coefficient β2 is input into the modeling function of the 3D dynamic face model to determine the initial 3D dynamic face model, i.e., the mesh of the initial 3D dynamic face model, which is the reference face shape feature information. This embodiment extracts features from the face sample image to obtain an initial texture, which is a real-person texture type face texture. The computer device inputs the initial texture and the mesh of the initial 3D dynamic face model into a Diff Renderer (i.e., a differential renderer or a differentiable renderer) to render a 2D candidate face image. The type of the candidate face image is the same as the type of the face sample image; for example, both the candidate face image and the face sample image are real-person type images.

[0097] Based on this, in addition to training the initial shape prediction sub-model using the loss value between the deformation coefficient β2 and the β label, this embodiment of the application also calculates the loss value between the candidate face image and the face sample. Based on this loss value, the initial shape prediction sub-model is trained to obtain the trained initial shape prediction sub-model. In this process, this embodiment of the application minimizes the loss value between the candidate face image and the face sample.

[0098] In this embodiment of the application, before encoding the candidate's face shape feature information using a trained shape transfer sub-model to obtain the face shape feature information, the shape transfer sub-model to be trained can be trained to obtain a trained shape transfer sub-model. The training process of the shape transfer sub-model to be trained in this embodiment of the application can be as follows:

[0099] For example, a computer device acquires the shape feature information of a labeled face; the shape transfer sub-model to be trained encodes the shape feature information of a reference face to obtain the encoded face shape feature information; based on the encoded face shape feature information and the labeled face shape feature information, the shape transfer sub-model to be trained is trained to obtain the trained shape transfer sub-model.

[0100] Among them, the face shape feature information of the label can be the cartoon deformation coefficient β'.

[0101] Among them, the shape transfer sub-model to be trained can be a generative adversarial network model.

[0102] like Figure 6 As shown, in this embodiment, the facial shape feature information, i.e., the deformation coefficient β2, is input into the Encoder (i.e., the encoder), encoded into the latent space, and then decoded by the Decoder (i.e., the decoder) to obtain the encoded facial shape feature information. This encoded facial shape feature information is cartoon-type encoded facial shape feature information. In this embodiment, the encoded facial shape feature information and the labeled facial shape feature information are input into the Discriminator (i.e., the discriminator) for discrimination. In this embodiment, the encoded facial shape feature information and the labeled facial shape feature information are trained to be close enough that the Discriminator cannot distinguish them. The labeled facial shape feature information is cartoon-type facial shape feature information.

[0103] In this embodiment of the application, before using a trained face texture prediction model to predict the texture of a face image and obtain the target texture of the face image, the face texture prediction model to be trained can be trained to obtain a trained face texture prediction model. The process of training the shape transfer sub-model to be trained in this embodiment of the application can be as follows:

[0104] For example, a computer device acquires a face sample image and a first texture, the texture type of which matches the texture type of the virtual object to be swapped; a face texture prediction model to be trained is used to extract features from the face sample image to obtain a second texture, the texture type of which is different from that of the first texture; the face texture prediction model to be trained is used to encode the texture type of the first texture to obtain an encoded first texture, the texture type of which does not match that of the first texture; the face texture prediction model to be trained is used to encode the texture type of the second texture to obtain an encoded second texture, the texture type of which does not match that of the second texture; based on the first texture, the second texture, the encoded first texture, and the encoded second texture, the face texture prediction model to be trained is trained to obtain a trained face texture prediction model.

[0105] Furthermore, in this embodiment, the face texture prediction model to be trained is used to encode the texture type of the first texture to obtain the encoded third texture, and the texture type of the encoded third texture matches the texture type of the first texture; the face texture prediction model to be trained is used to encode the texture type of the second texture to obtain the encoded fourth texture, and the texture type of the encoded fourth texture matches the texture type of the second texture.

[0106] Texture types can include cartoon texture types and realistic texture types. When two textures have the same texture type, they are said to have a matching texture type; when two textures have different texture types, they are said to have a mismatched texture type.

[0107] For example, such as Figure 7 As shown, the first texture is a cartoon texture type, or simply a cartoon texture. The second texture is a realistic texture type, or simply a realistic texture.

[0108] During training, this embodiment uses EncoderA to encode the real-life texture into the latent space, and also uses EncoderB to encode the cartoon texture into the latent space. DecoderA receives the encoded information from the latent space and decodes the real-life texture; DecoderB receives the encoded information from the latent space and decodes the cartoon texture.

[0109] During the training process, the face texture prediction model to be trained in this application embodiment performs the following encoding: (1) The face texture prediction model to be trained encodes the texture type of the first texture to obtain the encoded first texture. The texture type of the encoded first texture does not match the texture type of the first texture. For example, EncoderB encodes the cartoon texture into the latent space, and DecoderA receives the encoding information of the latent space and decodes the real texture; (2) The face texture prediction model to be trained encodes the texture type of the second texture to obtain the encoded second texture. The texture type of the encoded second texture does not match the texture type of the second texture. For example, EncoderA encodes the real texture into the latent space, and DecoderB receives the encoding information of the latent space and decodes the cartoon texture; (3) The face texture prediction model to be trained encodes the texture type of the first texture to obtain the encoded third texture. The texture type of the encoded third texture matches the texture type of the first texture. For example, EncoderB encodes the cartoon texture into the latent space, and DecoderB receives the latent space encoding information and decodes the cartoon texture; (3) Encode the space information and decode the cartoon texture; (4) Use the face texture prediction model to be trained to encode the texture type of the second texture and obtain the encoded fourth texture. For example, use EncoderA to encode the real texture into the latent space, and DecodeA receives the encoding information of the latent space and decodes the real texture.

[0110] Based on this, the loss information is calculated in the following ways in the embodiments of this application: (1) Based on the first texture and the encoded first texture, the first loss information of the face texture prediction model to be trained is calculated, for example, the first loss information between cartoon texture and real texture is calculated; (2) Based on the first texture and the encoded third texture, the second loss information of the face texture prediction model to be trained is calculated, for example, the second loss information between cartoon texture and cartoon texture is calculated; (3) Based on the second texture and the encoded second texture, the third loss information of the face texture prediction model to be trained is calculated, for example, the third loss information between real texture and cartoon texture is calculated; (4) Based on the second texture and the encoded fourth texture, the fourth loss information of the face texture prediction model to be trained is calculated, for example, the fourth loss information between real texture and real texture is calculated.

[0111] The embodiments of this application can train the face texture prediction model to be trained based on the first loss information and the third loss information until the face texture prediction model to be trained converges, thereby obtaining the trained face texture prediction model.

[0112] The embodiments of this application can train the face texture prediction model to be trained based on the second loss information and the fourth loss information until the face texture prediction model to be trained converges, thereby obtaining the trained face texture prediction model.

[0113] In the training process of this application embodiment, the real texture encoded by DecoderA is supervised by the real texture label of real texture type, so that Discriminator A (i.e., the discriminator A) cannot distinguish it; the cartoon texture encoded by DecoderB is supervised by the cartoon texture label of cartoon texture type, so that Discriminator B (i.e., the discriminator B) cannot distinguish it.

[0114] S103. Based on facial shape feature information, determine the target three-dimensional dynamic facial model.

[0115] In this embodiment of the application, the facial shape feature information, i.e. the deformation coefficient, is input into the modeling function of the three-dimensional dynamic facial model to determine the target three-dimensional dynamic facial model with the facial shape feature information.

[0116] Since the rotation angle and deformation coefficient can be obtained through formula (1) and / or formula (2) in the embodiments of this application, that is, the rotation angle and deformation coefficient are mapped to each other. Based on this, the embodiments of this application obtain the rotation angle corresponding to the face shape feature information based on the face shape feature information, i.e., the deformation coefficient, and based on the mapping relationship between the rotation angle and the deformation coefficient.

[0117] Of course, a default rotation angle can also be set in the embodiments of this application.

[0118] In this embodiment, the rotation angle is input into the function R for constructing the rotation posture of the 3D dynamic face model to obtain the rotation posture of the 3D dynamic face model; the deformation coefficient is input into the modeling function M of the 3D dynamic face model to obtain the face shape of the 3D dynamic face model, that is, the mesh of the face shape of the 3D dynamic face model. The target 3D dynamic face model is determined based on the mesh of the face shape of the 3D dynamic face model and the rotation posture of the 3D dynamic face model.

[0119] S104. Based on the target three-dimensional dynamic face model, determine the preset face shape of the preset three-dimensional dynamic face model that matches the target three-dimensional dynamic face model from the face shape mapping relationship set, so as to use the preset face shape as the target face shape of the face image.

[0120] When the preset 3D dynamic face model and the target 3D dynamic face model are the same, it can be said that the preset 3D dynamic face model and the target 3D dynamic face model are matched.

[0121] The embodiments of this application can determine the grid vertices of a preset face shape based on the grid vertices of a preset three-dimensional dynamic face model, see formula (3).

[0122] S105. Based on the target face shape and target texture, perform face swapping on the virtual object to be swapped.

[0123] In this embodiment, the target face shape is used as the face shape of the virtual object to be face-swapping. Since both the target texture and the target face shape are obtained through face image extraction processing, there is a mapping relationship between the coordinates of the key points of the target texture and the coordinates of the key points of the target face shape. Based on this, this embodiment uses this mapping relationship to attach the target texture to the face of the virtual object to be face-swapping, thereby achieving face swapping of the virtual object.

[0124] This application embodiment can acquire a two-dimensional face image and a three-dimensional face shape mapping relationship set for a virtual object to be face-swapped. The face shape mapping relationship set includes the mapping relationship between the preset face shape of the virtual object before face-swapping and the preset three-dimensional dynamic face model. A trained face feature prediction model is used to predict the face image to obtain the target texture and face shape feature information of the face image. Based on the face shape feature information, the target three-dimensional dynamic face model is determined. Based on the target three-dimensional dynamic face model, the preset face shape of the preset three-dimensional dynamic face model that matches the target three-dimensional dynamic face model is determined from the face shape mapping relationship set, so as to use the preset face shape as the target face shape of the face image. Based on the target face shape and the target texture, the face of the virtual object to be face-swapped is swapped. Since the embodiments of this application can quickly determine the target face shape of the face image based on the set of face shape mapping relationships of the virtual object to be face-swapping and the target three-dimensional dynamic face model corresponding to the face shape feature information of the face image, the face-swapping of the virtual object to be face-swapping can be performed based on the target face shape and target texture, thereby improving the face-swapping efficiency of three-dimensional virtual objects.

[0125] To better implement the above methods, this application also provides a virtual object face-swapping device, which can be integrated into a computer device, such as a server or terminal. The terminal may include a tablet computer, a laptop computer, and / or a personal computer.

[0126] For example, such as Figure 8 As shown, the virtual object face-swapping device may include an acquisition unit 301, a prediction unit 302, a first determination unit 303, a second determination unit 304, a face-swapping unit 305, and a training unit 306, as follows:

[0127] (1) Obtain unit 301;

[0128] The acquisition unit 301 can be used to acquire a two-dimensional face image and a set of face shape mapping relationships for a three-dimensional virtual object to be face-swapped. The set of face shape mapping relationships includes the mapping relationship between the preset face shape of the virtual object to be face-swapped before face swapping and the preset three-dimensional dynamic face model.

[0129] In some embodiments, the acquisition unit 301 can be specifically used to acquire key points of the face of the virtual object to be face-swapped and the preset face shape of the virtual object before face-swapping; based on the key points, determine the object face shape feature information of the virtual object to be face-swapped; based on the object face shape feature information, determine the preset three-dimensional dynamic face model of the virtual object to be face-swapped; and based on the preset face shape and the preset three-dimensional dynamic face model, generate a set of face shape mapping relationships.

[0130] In some embodiments, the acquisition unit 301 can be specifically used to acquire the mesh vertices of a preset face shape and the mesh vertices of a preset three-dimensional dynamic face model; and to determine a set of face shape mapping relationships based on the mesh vertices of the preset face shape and the mesh vertices of the preset three-dimensional dynamic face model.

[0131] (2) Prediction unit 302;

[0132] The prediction unit 302 can be used to predict face images using a trained face feature prediction model to obtain the target texture and face shape feature information of the face image.

[0133] In some embodiments, the prediction unit 302 can be specifically used to perform texture prediction on a face image using a trained face texture prediction model to obtain the target texture of the face image; and to perform shape prediction on a face image using a trained face shape prediction model to obtain the face shape feature information of the face image.

[0134] In some embodiments, the prediction unit 302 can be specifically used to perform initial shape prediction on a face image using a trained initial shape prediction sub-model to obtain candidate face shape feature information of the face image; and to encode the candidate face shape feature information using a trained shape transfer sub-model to obtain face shape feature information.

[0135] (3) First determining unit 303;

[0136] The first determining unit 303 can be used to determine a target three-dimensional dynamic face model based on face shape feature information.

[0137] (4) Second determining unit 304;

[0138] The second determining unit 304 can be used to determine the preset face shape of the preset three-dimensional dynamic face model that matches the target three-dimensional dynamic face model from the set of face shape mapping relationships based on the target three-dimensional dynamic face model, so as to use the preset face shape as the target face shape of the face image.

[0139] (5) Face-swapping unit 305;

[0140] The face-swapping unit 305 can be used to swap the face of a virtual object to be swapped based on the target face shape and target texture.

[0141] (6) Training Unit 306;

[0142] Training unit 306 can be used to acquire face sample images and real face shape feature information; use the initial shape prediction sub-model to be trained to predict the face sample images to obtain reference face shape feature information; and train the initial shape prediction sub-model to be trained based on the real face shape feature information and the reference face shape feature information to obtain the trained initial shape prediction sub-model.

[0143] The training unit 306 can be used to determine an initial 3D dynamic face model based on reference face shape feature information; extract features from face sample images to obtain the initial texture of the face sample images; determine candidate face images based on the initial 3D dynamic face model and the initial texture; and train the initial shape prediction sub-model to be trained based on real face shape feature information, reference face shape feature information, face sample images, and candidate face images to obtain the trained initial shape prediction sub-model.

[0144] Training unit 306 can be used to acquire the shape feature information of the labeled face; the shape transfer sub-model to be trained is used to encode the shape feature information of the reference face to obtain the encoded face shape feature information; based on the encoded face shape feature information and the labeled face shape feature information, the shape transfer sub-model to be trained is trained to obtain the trained shape transfer sub-model.

[0145] Training unit 306 is specifically used to acquire face sample images and acquire a first texture, the texture type of which matches the texture type of the virtual object to be swapped; to extract features from the face sample images using the face texture prediction model to be trained, resulting in a second texture, the texture type of which is different from that of the first texture; to encode the texture type of the first texture using the face texture prediction model to be trained, resulting in an encoded first texture, the texture type of which does not match that of the second texture; to encode the texture type of the second texture using the face texture prediction model to be trained, resulting in an encoded second texture, the texture type of which does not match that of the second texture; and to train the face texture prediction model to be trained based on the first texture, the second texture, the encoded first texture, and the encoded second texture, resulting in a trained face texture prediction model.

[0146] Training unit 306 can be specifically used to encode the texture type of the first texture using the face texture prediction model to be trained, to obtain the encoded third texture, the texture type of the encoded third texture matching the texture type of the first texture; to encode the texture type of the second texture using the face texture prediction model to be trained, to obtain the encoded fourth texture, the texture type of the encoded fourth texture matching the texture type of the second texture; and to train the face texture prediction model to be trained based on the first texture, the second texture, the encoded first texture, the encoded second texture, the encoded third texture, and the encoded third texture, to obtain the trained face texture prediction model.

[0147] The training unit 306 can be specifically used to calculate the first loss information of the face texture prediction model to be trained based on the first texture and the encoded first texture; calculate the second loss information of the face texture prediction model to be trained based on the first texture and the encoded third texture; calculate the third loss information of the face texture prediction model to be trained based on the second texture and the encoded second texture; calculate the fourth loss information of the face texture prediction model to be trained based on the second texture and the encoded fourth texture; and train the face texture prediction model to be trained using the first loss information, the second loss information, the third loss information, and the fourth loss information to obtain the trained face texture prediction model.

[0148] As can be seen from the above, the acquisition unit 301 of this application embodiment can acquire a two-dimensional face image and a three-dimensional face shape mapping relationship set of the virtual object to be face-swapped. The face shape mapping relationship set includes the mapping relationship between the preset face shape of the virtual object to be face-swapped before face swapping and the preset three-dimensional dynamic face model. The prediction unit 302 can use a trained face feature prediction model to predict the face image and obtain the target texture of the face image and the face shape feature information of the face image. The first determination unit 303 can determine the target three-dimensional dynamic face model of the face shape feature information based on the face shape feature information. The second determination unit 304 can determine the preset face shape of the preset three-dimensional dynamic face model that matches the target three-dimensional dynamic face model from the face shape mapping relationship set based on the target three-dimensional dynamic face model, so as to use the preset face shape as the target face shape of the face image. The face-swapping unit 305 can perform face swapping on the virtual object to be face-swapped based on the target face shape and the target texture. Since the embodiments of this application can quickly determine the target face shape of the face image based on the set of face shape mapping relationships of the virtual object to be face-swapping and the target three-dimensional dynamic face model corresponding to the face shape feature information of the face image, the face-swapping of the virtual object to be face-swapping can be performed based on the target face shape and target texture, thereby improving the face-swapping efficiency of three-dimensional virtual objects.

[0149] This application also provides a computer device, such as... Figure 9 As shown, it illustrates a structural schematic diagram of the computer device involved in the embodiments of this application, specifically:

[0150] The computer device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art will understand that... Figure 9 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0151] The processor 401 is the control center of the computer device. It connects various parts of the computer device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 402, and by calling data stored in the memory 402, thereby providing overall monitoring of the computer device. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and computer programs, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 401.

[0152] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, computer programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0153] The computer device also includes a power supply 403 that supplies power to the various components. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0154] The computer device may also include an input unit 404, which can be used to receive input digital or character information communication, and to generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0155] Although not shown, the computer device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the computer device loads the executable files corresponding to the processes of one or more computer programs into the memory 402 according to the following instructions, and the processor 401 runs the computer programs stored in the memory 402 to realize various functions, as follows:

[0156] The process involves acquiring a 2D face image and a set of 3D face shape mapping relationships for the virtual object to be face-swapped. This set includes the mapping relationship between the preset face shape of the virtual object before face-swapping and a preset 3D dynamic face model. A trained face feature prediction model is used to predict the face image, obtaining the target texture and face shape feature information. Based on the face shape feature information, a target 3D dynamic face model is determined. Based on the target 3D dynamic face model, a preset face shape matching the target 3D dynamic face model is determined from the face shape mapping relationship set, and this preset face shape is used as the target face shape of the face image. Finally, face-swapping is performed on the virtual object to be face-swapped based on the target face shape and target texture.

[0157] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0158] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by a computer program, or by a computer program controlling related hardware. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0159] Therefore, embodiments of this application provide a computer-readable storage medium storing a computer program that can be loaded by a processor to execute any of the virtual object face-swapping methods provided in embodiments of this application.

[0160] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0161] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0162] Since the instructions stored in the computer-readable storage medium can execute the steps of any of the virtual object face-swapping methods provided in the embodiments of this application, the beneficial effects that any of the virtual object face-swapping methods provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.

[0163] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations of the above embodiments.

[0164] The foregoing has provided a detailed description of a virtual object face-swapping method, computer device, and computer-readable storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for face swapping of virtual objects, characterized in that, include: Acquire a two-dimensional face image and acquire a set of face shape mapping relationships for a three-dimensional virtual object to be face-swapped. The set of face shape mapping relationships includes the mapping relationship between the preset face shape of the virtual object before face swapping and the preset three-dimensional dynamic face model. The trained face feature prediction model is used to predict the face image to obtain the target texture and face shape feature information of the face image; Based on the facial shape feature information, a target three-dimensional dynamic facial model based on the facial shape feature information is determined; Based on the target three-dimensional dynamic face model, a preset face shape of a preset three-dimensional dynamic face model that matches the target three-dimensional dynamic face model is determined from the face shape mapping relationship set, so as to use the preset face shape as the target face shape of the face image. The face of the virtual object to be swapped is swapped based on the target face shape and the target texture.

2. The virtual object face-swapping method according to claim 1, characterized in that, The trained face feature prediction model includes a trained face texture prediction model and a trained face shape prediction model; the step of using the trained face feature prediction model to predict the face image to obtain the target texture and face shape feature information of the face image includes: The trained face texture prediction model is used to predict the texture of the face image to obtain the target texture of the face image; The trained face shape prediction model is used to predict the shape of the face image to obtain the face shape feature information of the face image.

3. The virtual object face-swapping method according to claim 2, characterized in that, The trained face shape prediction model includes a trained initial shape prediction sub-model and a trained shape transfer sub-model; the step of using the trained face shape prediction model to predict the shape of the face image and obtain the face shape feature information of the face image includes: The trained initial shape prediction sub-model is used to predict the face image to obtain the candidate face shape feature information of the face image; The face shape feature information of the candidate is encoded using a trained shape transfer sub-model to obtain the face shape feature information.

4. The virtual object face-swapping method according to claim 3, characterized in that, Before using the trained initial shape prediction sub-model to predict the face image and obtain the candidate face shape feature information of the face image, the method further includes: Acquire sample images of faces and obtain real facial shape feature information; The face sample image is predicted using the initial shape prediction sub-model to be trained, and reference face shape feature information is obtained. Based on the real face shape feature information and the reference face shape feature information, the initial shape prediction sub-model to be trained is trained to obtain the trained initial shape prediction sub-model.

5. The virtual object face-swapping method according to claim 4, characterized in that, After using the initial shape prediction sub-model to be trained to predict the face sample image and obtain the reference face shape feature information, the method further includes: Based on the reference face shape feature information, an initial three-dimensional dynamic face model based on the reference face shape feature information is determined; Feature extraction is performed on the face sample image to obtain the initial texture of the face sample image; Based on the initial 3D dynamic face model and the initial texture, candidate face images are determined; The step of training the initial shape prediction sub-model to be trained based on the real face shape feature information and the reference face shape feature information to obtain the trained initial shape prediction sub-model includes: training the initial shape prediction sub-model to be trained based on the real face shape feature information, the reference face shape feature information, the face sample image, and the candidate face image to obtain the trained initial shape prediction sub-model.

6. The virtual object face-swapping method according to claim 3, characterized in that, Before encoding the candidate face shape feature information using a trained shape transfer sub-model to obtain the face shape feature information, the method further includes: Obtain the facial shape feature information of the tagged person; The shape transfer sub-model to be trained is used to encode the shape feature information of the reference face to obtain the encoded face shape feature information; Based on the encoded face shape feature information and the labeled face shape feature information, the shape transfer sub-model to be trained is trained to obtain the trained shape transfer sub-model.

7. The virtual object face-swapping method according to claim 2, characterized in that, Before using the trained face texture prediction model to predict the texture of the face image and obtain the target texture of the face image, the method further includes: Acquire a face sample image and acquire a first texture, wherein the texture type of the first texture matches the texture type of the virtual object to be swapped; The face sample image is used to extract features using a face texture prediction model to be trained, and a second texture is obtained. The texture type of the second texture is different from that of the first texture. The first texture is encoded using the face texture prediction model to be trained, resulting in an encoded first texture. The texture type of the encoded first texture does not match the texture type of the first texture. The face texture prediction model to be trained is used to encode the texture type of the second texture to obtain the encoded second texture. The texture type of the encoded second texture does not match the texture type of the second texture. Based on the first texture, the second texture, the encoded first texture, and the encoded second texture, the face texture prediction model to be trained is trained to obtain the trained face texture prediction model.

8. The virtual object face-swapping method according to claim 7, characterized in that, After extracting features from the face sample image using the face texture prediction model to be trained to obtain the second texture, the process further includes: The first texture is encoded using the face texture prediction model to be trained to obtain an encoded third texture, wherein the texture type of the encoded third texture matches the texture type of the first texture. The face texture prediction model to be trained is used to encode the texture type of the second texture to obtain the encoded fourth texture, and the texture type of the encoded fourth texture matches the texture type of the second texture. The step of training the face texture prediction model to be trained based on the first texture, the second texture, the encoded first texture, and the encoded second texture to obtain the trained face texture prediction model includes: training the face texture prediction model to be trained based on the first texture, the second texture, the encoded first texture, the encoded second texture, the encoded third texture, and the encoded third texture to obtain the trained face texture prediction model.

9. The virtual object face-swapping method according to claim 8, characterized in that, The step of training the face texture prediction model to be trained based on the first texture, the second texture, the encoded first texture, the encoded second texture, the encoded third texture, and the encoded third texture to obtain the trained face texture prediction model includes: Based on the first texture and the encoded first texture, calculate the first loss information of the face texture prediction model to be trained; Based on the first texture and the encoded third texture, calculate the second loss information of the face texture prediction model to be trained; Based on the second texture and the encoded second texture, the third loss information of the face texture prediction model to be trained is calculated; Based on the second texture and the encoded fourth texture, calculate the fourth loss information of the face texture prediction model to be trained; The first loss information, the second loss information, the third loss information, and the fourth loss information are used to train the face texture prediction model to obtain the trained face texture prediction model.

10. The virtual object face-swapping method according to claim 1, characterized in that, The process of obtaining the set of facial shape mapping relationships for the three-dimensional virtual object to be face-swapped includes: Obtain the key points of the face of the virtual object to be face-swapped and the preset face shape of the virtual object before face-swapping; Based on the aforementioned key points, the facial shape feature information of the virtual object to be face-swapped is determined; Based on the facial shape feature information of the object, a preset three-dimensional dynamic facial model of the virtual object to be face-swapped is determined; Based on the preset face shape and the preset 3D dynamic face model, the set of face shape mapping relationships is generated.

11. The virtual object face-swapping method according to claim 10, characterized in that, The step of generating the face shape mapping relationship set based on the preset face shape and the preset 3D dynamic face model includes: Obtain the mesh vertices of the preset face shape, and obtain the mesh vertices of the preset three-dimensional dynamic face model; Based on the mesh vertices of the preset face shape and the mesh vertices of the preset 3D dynamic face model, the set of face shape mapping relationships is determined.

12. A virtual object face-swapping device, characterized in that, include: The acquisition unit is used to acquire a two-dimensional face image and a set of face shape mapping relationships for a three-dimensional virtual object to be face-swapped. The set of face shape mapping relationships includes the mapping relationship between the preset face shape of the virtual object to be face-swapped before face swapping and the preset three-dimensional dynamic face model. The prediction unit is used to predict the face image using a trained face feature prediction model to obtain the target texture and face shape feature information of the face image. The first determining unit is used to determine the target three-dimensional dynamic face model based on the face shape feature information. The second determining unit is used to determine, based on the target three-dimensional dynamic face model, a preset face shape of a preset three-dimensional dynamic face model that matches the target three-dimensional dynamic face model from the face shape mapping relationship set, so as to use the preset face shape as the target face shape of the face image. The face-swapping unit is used to swap the face of the virtual object to be swapped based on the target face shape and the target texture.

13. A computer device, characterized in that, It includes a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to perform the operations in the virtual object face-swapping method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted for loading by a processor to execute the virtual object face-swapping method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Image processing method and apparatus

    CN105096377A

  • Model head portrait creation method and device, electronic equipment and storage medium

    CN112669447A