Expression migration method and device, electronic equipment and computer readable storage medium
Through the expression feature extraction and migration model, ordinary shooting equipment is used to obtain facial images and train network models, the problem of poor facial expression migration effect in the existing technology is solved, efficient and accurate expression migration effect is achieved, and labor and time costs are reduced.
Patent Information
- Application Number
- CN202410172138.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-06
- Publication Date
- 2025-08-08
AI Technical Summary
In the prior art, when facial expressions are transferred to virtual characters' faces, the expression migration effect is quite different from the facial expressions of the characters, and it requires a lot of manual and time costs for manual correction and repair.
The expression feature extraction model and expression migration model are used to obtain facial images through ordinary shooting equipment, and the network model is trained using multiple sample triplets and label information, and fine-grained continuous expression feature vectors are extracted, and input them into the expression migration model to generate facial images of the target virtual character.
It improves the effect of facial expression migration, reduces labor and time costs, and realizes accurate expression migration from any object to any virtual character, avoiding specific device and scene restrictions.
Smart Images

Figure CN120452037A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an expression migration method, device, electronic device and computer-readable storage medium. Background Art
[0002] With the continuous development of image processing technology, facial expression and motion capture technology has emerged. Facial expression and motion capture technology refers to the process of using sensors, cameras, and other imaging equipment to record a person's facial expressions and movements and convert them into a series of expression parameters. These expression parameters can then be used to restore facial expressions, and based on these expression parameters, facial images of virtual characters with corresponding facial expressions can be created.
[0003] Related technologies require users to wear a specific helmet in a specific scenario, record their facial expressions through the helmet, and then use specific software to interpret the recorded facial expression data, transfer the user's facial expressions to the face of a virtual character, and finally output a facial image of the virtual character with the user's facial expression. For example, if the user's facial expression is recorded as laughing, the final output is a facial image of the virtual character with a laughing expression.
[0004] However, due to the limitations of the accuracy and recording effect of the above-mentioned specific software, the above-mentioned related technologies result in a significant difference between the facial expression migration effect to the virtual character's face and the facial expression of the person himself. It is necessary to further manually correct and repair the facial expressions in the virtual character image, which requires a lot of manpower and time costs. Summary of the Invention
[0005] The present application provides an expression transfer method, device, electronic device and computer-readable storage medium to improve the effect of facial expression transfer and reduce labor and time costs.
[0006] In a first aspect, an embodiment of the present application provides an expression migration method, the method comprising:
[0007] Acquire a first facial image containing a target facial expression;
[0008] Inputting the first facial image into an expression feature extraction model to obtain a target expression feature vector output by the expression feature extraction model for characterizing the target facial expression; wherein the expression feature extraction model is obtained by training a first network model based on a plurality of sample triplets and annotation information corresponding to each of the sample triplets, each sample triplet including three first sample images, and the annotation information being used to describe the facial expressions corresponding to the three first sample images in the sample triplet;
[0009] The target expression feature vector is input into an expression transfer model to obtain a second facial image of the target virtual character having the target facial expression output by the expression transfer model.
[0010] In a second aspect, an embodiment of the present application provides an expression migration device, the device comprising:
[0011] an acquisition module, configured to acquire a first facial image containing a target facial expression;
[0012] a feature extraction module, configured to input the first facial image into an expression feature extraction model to obtain a target expression feature vector output by the expression feature extraction model and used to characterize the target facial expression; wherein the expression feature extraction model is obtained by training a first network model based on a plurality of sample triplets and annotation information corresponding to each of the sample triplets, each sample triplet including three first sample images, and the annotation information being used to describe the facial expressions corresponding to the three first sample images in the sample triplet;
[0013] The processing module is used to input the expression feature vector into an expression transfer model to obtain a second facial image of the target virtual character having the target facial expression output by the expression transfer model.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, comprising:
[0015] A memory and a processor, wherein the memory and the processor are coupled;
[0016] The memory is used to store one or more computer instructions;
[0017] The processor is used to execute the one or more computer instructions to implement the expression migration method described in any one of the first aspects above.
[0018] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having one or more computer instructions stored thereon, characterized in that the instruction is executed by a processor to implement the expression migration method described in any one of the first aspects above.
[0019] In a fifth aspect, an embodiment of the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the expression migration method described in any one of the first aspects above.
[0020] Compared with the prior art, this application has the following advantages:
[0021] The expression transfer method provided in the present application first obtains an arbitrary first facial image containing a target facial expression. Subsequently, the first facial image is input into an expression feature extraction model to obtain a target expression feature vector output by the expression feature extraction model for characterizing the target facial expression. The first facial image used in the present application can be recorded by an ordinary shooting device, without the need to use specific equipment such as a specific helmet to capture the expression of a real person, and is not subject to any scene restrictions. This greatly improves the diversity and concurrency of facial expression image acquisition and avoids a large amount of labor and time costs. The first facial image is input into the expression feature extraction model to obtain a target expression feature vector for characterizing the target facial expression in the first facial image. The expression feature extraction model is obtained by training a first network model based on multiple sample triplets and the annotation information corresponding to each sample triplet. Each sample triplet includes three facial images, and the annotation information is used to describe the facial expressions of the three first sample images in the corresponding sample triplet. The expression feature extraction model in this application can extract any fine-grained, continuous expression features. Moreover, since the annotation information corresponding to the sample triples only contains expression information and does not contain any irrelevant information (such as identity information, background, light, etc.), the expression feature vector extracted by the expression feature extraction model is only related to the expression, thereby achieving the decoupling of the expression feature vector from irrelevant information such as identity information, which can accurately perceive the fine-grained expressions between different identities and realize the expression transfer from any object to any virtual character. Finally, the target expression feature vector is input into the expression transfer model to obtain the second facial image of the target virtual character with the target facial expression output by the expression transfer model. This application can greatly improve the effect of expression transfer. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0023] Figure 1 A flowchart of an expression transfer method according to one embodiment of the present application;
[0024] Figure 2 A schematic diagram of the visualization of the facial expression feature space provided in one embodiment of the present application;
[0025] Figure 3 A schematic diagram of the expression migration process provided in one embodiment of the present application;
[0026] Figure 4 This is a flowchart of an algorithm iteration based on human-machine collaboration and data closed loop provided in one embodiment of the present application;
[0027] Figure 5 A schematic diagram of the structure of an expression transfer device provided in one embodiment of the present application;
[0028] Figure 6 A schematic diagram of the hardware structure of an electronic device provided in one embodiment of the present application.
[0029] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0030] To make the purposes, advantages, and features of this application more clear, the following clearly and completely describes this application in conjunction with the accompanying drawings and specific embodiments. In the following description, many specific details are set forth to facilitate a full understanding of this application. However, the described embodiments are only some of the embodiments of this application, not all of them. All other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.
[0031] It should be noted that, in the description of this application, the terms "first", "second", etc. are used for descriptive purposes only and are not to be understood as indicating or implying relative importance, or a specific order or precedence. For those skilled in the art, the specific meanings of the above terms in this application can be understood in specific circumstances. In addition, in the description of this application, unless otherwise specified, the term "plurality" refers to two or more. The term "and / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. The terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.
[0032] In order to facilitate understanding of the technical solution of this application, the relevant concepts involved in this application are first introduced.
[0033] With the rapid development of computer technology, the production of character images in animation, games, metaverse and other fields has been fully transferred from manual frame-by-frame drawing to the use of professional 3D animation production software such as Maya and 3Ds Max.
[0034] Facial motion capture technology, also known as facial expression motion capture technology, is a part of motion capture technology. It refers to the process of using sensors, video cameras, and other equipment to record facial expressions and movements of people and convert them into a series of expression parameters. The converted expression parameters can be used in the production of computer graphics (CG) and animation, so that the facial expressions and movements of people can be restored on the face of virtual characters. Facial expression motion capture technology is widely used in virtual reality, game development, film production, human-computer interaction and other fields. Through facial expression motion capture technology, the facial expressions of virtual characters can be produced based on the facial expressions of real people. Compared with manually created virtual character facial expression frames, since the facial expressions of computer graphics CG characters are derived from the facial expressions of real people, the facial expression effects of the CG characters will be more realistic and delicate, and the efficiency of expression production will be greatly improved.
[0035] Below, the prior art involved in this application and the problems existing in the prior art are described:
[0036] The 3D facial expression motion capture solutions in related technologies mainly rely on commercial systems such as Faceware or Dynamicxyz, and require wearing a special helmet as a hardware device. Facial expression animation production includes the following steps: facial movement recording, data solution and manual repair. Among them, facial movement recording: the actors need to wear a special helmet as a hardware camera system in a specific scene to perform and record facial expressions. Data solution: after the recording is completed, specific commercial software is needed to solve the captured expression data and output facial animation. Manual repair: due to the limitations of equipment accuracy and actual recording, the computer graphics or animations obtained through solution cannot be directly applied, and professional and experienced animators are required to correct, repair and iterate parameters multiple times in the later stage.
[0037] However, this related technical solution still has the following defects: (1) The diversity and concurrency of facial expression images are limited by the number of hardware and commercial software copyrights; (2) Specific helmets and scenes limit the recording flexibility and data richness; (3) The equipment accuracy and recording conditions limit the solved computer graphics or animation, resulting in a large difference between the facial expressions of the virtual character and the facial expressions of the real person after facial expression migration, that is, the actual migration effect of facial expressions is difficult to achieve the expected effect and needs to be repaired later; (4) Animation repair requires a lot of manual post-repair by professional animators, which consumes a lot of manpower and time costs.
[0038] To address at least some of the above-mentioned issues and improve the effectiveness of facial expression transfer, this application provides an expression transfer method, an expression transfer device corresponding to this method, an electronic device capable of implementing this expression transfer method, and a computer-readable storage medium. The following examples provide detailed descriptions of the aforementioned method, device, electronic device, and computer-readable storage medium.
[0039] In order to make the purpose and technical solution of the present application clearer and more intuitive, the method provided by the embodiment of the present application will be described in detail with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. It is understood that the following embodiments may exist separately, and the following embodiments and features in the embodiments may be combined with each other when there is no conflict between the embodiments provided in the present application. For the same or similar content, the description will not be repeated in different embodiments. In addition, the step sequence in the following method embodiments is only an example and is not strictly limited. In some cases, the steps shown or described may be performed in a different order from this.
[0040] The present application provides an expression migration method, device, electronic device and computer-readable storage medium. Specifically, the expression migration method of one embodiment of the present application can be executed by a computer device, wherein the computer device can be a terminal or server device. The terminal can be a terminal device such as a smart phone, a tablet computer, a laptop computer, a touch screen, etc. The terminal can also include a client, which can be a game application client, a browser client with a game program, or an instant messaging client, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and basic cloud computing services such as big data and artificial intelligence platforms.
[0041] Next, combine Figure 1 , an expression transfer method provided in one embodiment of the present application is described, Figure 1 A flowchart of an expression transfer method provided in one embodiment of the present application.
[0042] like Figure 1 As shown, the expression migration method includes steps S10-S30:
[0043] S10: Acquire a first facial image containing a target facial expression.
[0044] The first facial image described above can be a facial image of a person, a facial image of another animal, or a facial image of another virtual object (such as a virtual human character or a virtual animal). This is merely an example, and this application does not impose any limitation thereto. In other words, the expression transfer method provided in the embodiments of this application can be applied not only to the transfer of facial expressions of human faces, but also to the transfer of facial expressions of any other object having facial expressions.
[0045] As mentioned above, the first facial image may be obtained by scanning or photographing a two-dimensional image, and the two-dimensional image may be a facial image of a person or other animal captured by a monocular camera. The person or other animal here may be real or may be a virtual image (for example, a virtual character or virtual animal in a 3D game).
[0046] As mentioned above, the target facial expression can be any facial expression, such as happiness, sadness, anger, fatigue, etc., which is only an example and is not limited in this embodiment of the present application.
[0047] S20: Input the first facial image into an expression feature extraction model to obtain a target expression feature vector output by the expression feature extraction model for representing the target facial expression. The expression feature extraction model is obtained by training a first network model based on a plurality of sample triplets and annotation information corresponding to each sample triplet, where each sample triplet includes three first sample images, and the annotation information is used to describe the facial expressions of the three first sample images in the corresponding sample triplet.
[0048] The first network model described above can be a convolutional neural network. The first network model is trained using multiple sample triplets and the annotation information corresponding to each sample triplet. After training, the first network model becomes an expression feature extraction model. Each sample triplet includes three first sample images with facial expressions, and the annotation information describes the facial expressions of the three facial images in the sample triplet.
[0049] The annotation information described above is only used to describe the facial expression of the first sample image and is unrelated to the background, lighting, identity information, and other information in the facial image. Training the first network model based on this annotation information enables the expression feature extraction model to extract facial expression information from any facial image while removing irrelevant, redundant, and interfering information (such as the person's identity, background, and lighting).
[0050] An optional implementation method of training the first network model to obtain the expression feature extraction model is described in detail, including the following steps S201-S204:
[0051] S201. Obtain multiple sample triplets and labeling information corresponding to each sample triplet, where the sample triplets comprise three first sample images, namely, an anchor sample image, a positive sample image, and a negative sample image. The anchor sample image and the positive sample image are two facial images with the same or similar facial expressions, and the anchor sample image and the negative sample image are two facial images with different or dissimilar facial expressions.
[0052] S202. Input the anchor sample image, positive sample image, negative sample image and their respective annotation information in the sample triplet into the first network model for feature extraction to obtain the third expression feature vector corresponding to the anchor sample image, positive sample image and negative sample image.
[0053] S203. Based on the third loss function, calculate the third loss function value according to the third expression feature vectors corresponding to the anchor sample image, the positive sample image, and the negative sample image.
[0054] S204: Adjusting parameters of the first network model based on the third loss function value to obtain an expression feature extraction model. The parameter adjustment is used to shorten the distance between the anchor sample image and the positive sample image in the expression feature space, and to increase the distance between the anchor sample image and the negative sample image in the expression feature space.
[0055] As mentioned above, the sample triplet (A, P, N) includes three first sample images, which are divided into an anchor sample image (Anchor, A), a positive sample image (Positive, P) and a negative sample image (Negative, N). Among them, the anchor sample image A is a sample image to be classified, that is, a sample image that requires the first network model to classify its facial expression. The positive sample image P is a facial image that belongs to the same facial expression category as the anchor sample, that is, a sample image with a similar or identical facial expression to the anchor sample. The negative sample image N is a facial image that belongs to a different or dissimilar facial expression than the anchor sample image, that is, a facial image with a different or dissimilar facial expression than the anchor sample image. For example, the facial expression in the anchor sample image in the sample triplet is laughing, the facial expression in the positive sample image is smiling, and the facial expression in the negative sample image is crying. The annotation information corresponding to the sample triplet is used to describe the facial expressions of the three first sample images in the sample triplet.
[0056] In an embodiment of the present application, the anchor sample image, positive sample image, negative sample image and their respective annotation information in the sample triplet are input into the first network model for feature extraction to obtain the third expression feature vector corresponding to the anchor sample image, positive sample image and negative sample image.
[0057] In an embodiment of the present application, the purpose of training the first network model based on the sample triples is to train the first network model so that it can achieve, in the expression feature space, the distance between the expression feature vectors of the first sample images with the same or similar facial expressions is shortened, and the distance between the expression feature vectors of the first sample images with different or dissimilar facial expressions is increased. In other words, the goal of optimizing the training of the first network model is to shorten the distance between the anchor sample image and the positive sample image in the expression feature space, and to increase the distance between the anchor sample image and the negative sample image in the expression feature space. For example, if the facial expression in the anchor sample image in the sample triples is laughing, the facial expression in the positive sample image is smiling, and the facial expression in the negative sample image is crying, then after training the first network model, the distance between the expression feature vector corresponding to laughing and the expression feature vector corresponding to smiling in the expression feature space is smaller, and the distance between the expression feature vector corresponding to laughing and the expression feature vector corresponding to crying in the expression feature space is larger.
[0058] In this embodiment of the present application, a third loss function is used to calculate a third loss function value based on the third expression feature vectors corresponding to the anchor sample image, the positive sample image, and the negative sample image. The parameters of the first network model are adjusted based on the third loss function value to obtain an expression feature extraction model. This parameter adjustment is used to reduce the distance between the anchor sample image and the positive sample image in the expression feature space, and to increase the distance between the anchor sample image and the negative sample image in the expression feature space.
[0059] As mentioned above, the third loss function can be a triplet loss function, a contrast loss function, etc., which is only an example and the embodiments of the present application do not impose any limitations on this.
[0060] An optional implementation method is to adjust the parameters of the first network model through the stochastic gradient descent (SGD) method according to the third loss function value until the loss function converges (such as the third loss function value is less than the preset loss threshold), and then determine the trained first network model as the above-mentioned expression feature extraction model.
[0061] In an embodiment of the present application, a first facial image is input into an expression feature extraction model, and a target expression feature vector for representing a target facial expression is obtained as output by the expression feature extraction model. The distance (e.g., Euclidean distance) between the expression feature vectors of two different first facial images reflects the degree of similarity between the facial expressions of the two first facial images, i.e., a smaller distance indicates a greater degree of similarity between the first facial expressions of the two facial images, and a larger distance indicates a lesser degree of similarity between the first facial expressions of the two facial images.
[0062] For example, in combination Figure 2 , the expression feature vector is illustrated as follows. Figure 2 A schematic diagram of the visualization of the expression feature space provided in one embodiment of the present application.
[0063] Similar expressions tend to cluster in the same area, while the differences in expressions between different areas are quite large. Figure 2 As shown in the figure, the black point cloud is the visual distribution of the expression feature vector. Figure 2 Four regions are marked above, along with three facial images of examples of each region. Figure 2 As shown, the facial expression in region 1 is calm (i.e., expressionless), the facial expression in region 2 is surprised (i.e., mouth wide open, eyes wide open, eyebrows raised); the facial expression in region 3 is a normal smile (i.e., mouth open, mouth corners raised, eyes curved); and the facial expression in region 4 is a big laugh (i.e., mouth open wider, mouth corners raised higher, eyes curved more significantly). It can be seen that the closer the distance between regions 3 and 4, the greater the similarity in facial expression between the facial images in region 3 and 4; the farther the distance between regions 1 and 4, the less similarity in facial expression between the facial images in region 3 and 4.
[0064] Currently, existing facial expression feature extraction models require the pre-defined categories of several expressions, such as smiling, laughing, calm, and crying. These categories are often discrete. Many ambiguous expressions fall between different categories and cannot be accurately categorized into a single category. Consequently, when processing complex expressions, the facial expression feature extraction models are prone to inaccurate classification, resulting in the extraction of incorrect facial expression feature vectors.
[0065] Compared with the prior art, the training method adopted by the expression feature extraction model provided by the present application does not require the definition of expression categories. It mainly uses the first network model to train any facial expression image in the form of sample triples to obtain the expression feature extraction model. Therefore, the expression feature extraction model can extract any fine-grained, continuous expression features. Moreover, since the annotation information corresponding to the sample triples only contains expression information and does not contain any irrelevant information (such as identity information, background, light, etc.), the expression feature vector extracted by the expression feature extraction model is only related to the expression, thereby realizing the decoupling of the expression feature vector from irrelevant information such as identity information, so that fine-grained expressions between different identities can also be accurately perceived, and expression migration from any object to any virtual character can be realized.
[0066] S30: Input the target expression feature vector into the expression transfer model to obtain a second facial image of the target virtual character having the target facial expression output by the expression transfer model.
[0067] The expression transfer model is used to output a second facial image of the target virtual character with the target facial expression based on the expression feature vector representing the target facial expression.
[0068] As mentioned above, the model structure of the expression transfer model can be a generative adversarial network, which is only an example and the embodiments of the present application do not impose any limitation on this.
[0069] The expression migration method provided in the embodiment of the present application first obtains an arbitrary first facial image containing a target facial expression. Subsequently, the first facial image is input into an expression feature extraction model to obtain a target expression feature vector output by the expression feature extraction model for characterizing the target facial expression. The first facial image used in the present application can be recorded by an ordinary shooting device. There is no need to use specific equipment such as a specific helmet to capture the expression of a real person, and it is not subject to any scene restrictions. This greatly improves the diversity and concurrency of facial expression image acquisition and avoids a large amount of manpower and time costs. The first facial image is input into the expression feature extraction model to obtain a target expression feature vector for characterizing the target facial expression in the first facial image. Among them, the expression feature extraction model is obtained by training a first network model based on multiple sample triplets and the annotation information corresponding to each sample triplet. Each sample triplet includes three facial images, and the annotation information is used to describe the facial expressions of the three first sample images in the corresponding sample triplet. The expression feature extraction model in this application can extract any fine-grained, continuous expression features. Moreover, since the annotation information corresponding to the sample triples only contains expression information and does not contain any irrelevant information (such as identity information, background, light, etc.), the expression feature vector extracted by the expression feature extraction model is only related to the expression, thereby achieving the decoupling of the expression feature vector from irrelevant information such as identity information, which can accurately perceive the fine-grained expressions between different identities and realize the expression transfer from any object to any virtual character. Finally, the target expression feature vector is input into the expression transfer model to obtain the second facial image of the target virtual character with the target facial expression output by the expression transfer model. This application can greatly improve the effect of expression transfer.
[0070] Based on the above embodiments, the expression migration method provided in the embodiments of the present application is further described below.
[0071] In an optional implementation manner, the first sample image includes a facial image of a real person and a facial image of a virtual character.
[0072] Compared with the expression migration scheme based on facial expression motion capture technology, the facial image used in the expression migration method provided in the embodiment of the present application can be a facial expression image taken with any ordinary camera, without the need for specific equipment such as wearable devices for facial expression motion capture, and is no longer restricted by specific equipment and recording scenes. Therefore, the present application greatly improves the flexibility and diversity of expression migration, and avoids the high equipment cost.
[0073] In an optional implementation manner, the expression migration method provided in the embodiment of the present application further includes steps S40-S50:
[0074] S40. Input the second facial image into an expression parameter determination model to obtain target expression driving parameters output by the expression parameter determination model, wherein the expression parameter determination model is used to generate expression driving parameters that cause the face of the three-dimensional model to present the target facial expression.
[0075] The expression parameter determination model is configured to output target expression driving parameters based on the target facial expression in the second facial image. This expression parameter determination model is configured to generate expression driving parameters that cause the face of the three-dimensional model to exhibit the target facial expression. In other words, based on these target expression driving parameters, the face of the three-dimensional model corresponding to the target virtual character can exhibit the target facial expression.
[0076] S50: Input the target expression driving parameters into the three-dimensional model corresponding to the target virtual character, so that the face of the three-dimensional model presents the target facial expression.
[0077] The 3D model corresponding to the target virtual character is a 3D model that includes the target virtual character's face. It can also be a 3D facial model of the target virtual character. This is for illustrative purposes only and is not limited in this application. As long as the 3D model includes the target virtual character's face, it is sufficient. The facial expression of the 3D model corresponding to the target virtual character can be any expression, such as a neutral face, laughter, smile, or tears. A neutral face is one without any expression, also known as a base face or standard face.
[0078] For example, in combination Figure 3 , the expression transfer method provided in the embodiment of the present application is exemplarily described, Figure 3 A schematic diagram of the expression migration process provided in one embodiment of the present application.
[0079] like Figure 3 As shown, a first facial image having a facial expression is input into an expression feature extraction model to obtain a target expression feature vector for representing the facial expression in the first facial image. The target expression feature vector is input into an expression transfer model to obtain a second facial image of a target virtual character having a target facial expression, which is output by the expression transfer model. The second facial image is input into an expression parameter determination model to obtain target expression driving parameters output by the expression parameter determination model. The target expression driving parameters are input into a three-dimensional model corresponding to the target virtual character via a three-dimensional model renderer, so that the face of the three-dimensional model exhibits the same facial expression as the second facial image.
[0080] In an embodiment of the present application, a second facial image of a target virtual character having a target facial expression is first input into an expression parameter determination model to obtain target expression driving parameters output by the expression parameter determination model. Subsequently, the target expression driving parameters are input into a three-dimensional model corresponding to the target virtual character, and the face of the three-dimensional model exhibits the target facial expression. The three-dimensional model renderer can use the generated target expression driving parameters to reconstruct the facial expression of the three-dimensional character, thereby completing the migration of facial expressions. This method does not require matching data between a real person's face and a virtual character, nor does it require the collection or annotation of feature point positions on the face, greatly reducing the time cost and process complexity of three-dimensional expression migration.
[0081] In an optional implementation manner, the expression migration method provided in the embodiment of the present application further includes steps S301-S302:
[0082] S301: Construct an image generation model and an image discrimination model.
[0083] In an embodiment of the present application, the image generation model may be an adversarial generation model, and the image discrimination model may be a convolutional neural network or a recurrent neural network.
[0084] S302. For each second sample image in the plurality of second sample images, execute the first step to train the image generation model and the image discrimination model until the first convergence condition is met, and determine the trained image generation model as an expression transfer model, and the second sample image is a facial image containing a facial expression.
[0085] The first step includes steps S3021-S3024:
[0086] S3021. Input the second sample image into the expression feature extraction model to obtain a first expression feature vector output by the expression feature extraction model.
[0087] S3022: Input the first facial expression feature vector into an image generation model to obtain a third facial image of the target virtual character having a facial expression output by the image generation model.
[0088] S3023. Input the third facial image and the preset rendered facial image as two input images into the image discrimination model for discrimination, and obtain a discrimination result output by the image discrimination model; wherein the preset rendered facial image is a two-dimensional rendered image obtained by rendering the three-dimensional face model, and the discrimination result is used to indicate the probability value of each input image belonging to the rendered image and the generated image, respectively.
[0089] S3024. Based on the discrimination result and the first loss function, calculate the first loss function value and adjust the parameters of the image generation model and the image discrimination model according to the first loss function value.
[0090] Below, the above steps S3021-S3024 are described in detail.
[0091] As mentioned above, the training process of the expression transfer model involves two models: an image generation model (Generator) and an image discriminator model (Discriminator). The input of the image generation model is the first expression feature vector extracted from the second sample image, and the output is a third facial image of the target virtual character with facial expressions. The goal of the image generation model is to generate images that are as realistic as possible in an attempt to deceive the image discriminator model. The input of the image discriminator model is the third facial image and a preset rendered facial image. The goal of the image discriminator model is to distinguish between the rendered image and the generated image as accurately as possible.
[0092] In the embodiment of the present application, the training steps of the expression transfer model are described below:
[0093] Step 1: Initialize the image generation model and image discrimination model.
[0094] Step 2: Input the second sample image into the expression feature extraction model to obtain a first expression feature vector output by the expression feature extraction model.
[0095] Step 3: Input the first facial expression feature vector into an image generation model to obtain a third facial image of the target virtual character with facial expression output by the image generation model.
[0096] Step 4: Input the third facial image and the preset rendered facial image as input images into the image discrimination model, and obtain a discrimination result output by the image discrimination model. The preset rendered facial image is a two-dimensional rendered image obtained by rendering a three-dimensional face model. The discrimination result indicates the probability of each input image belonging to the rendered image and the generated image, respectively.
[0097] As described above, the preset rendered facial image is a two-dimensional rendered image obtained by rendering a three-dimensional facial model corresponding to any virtual character. The virtual character in the preset rendered facial image can be any virtual character, and the facial expression in the preset rendered facial image can be any facial expression. This application does not impose any restrictions on the virtual character in the preset rendered facial image or the facial expression of the virtual character, as long as the preset rendered facial image is a two-dimensional rendered image obtained by rendering a three-dimensional facial model.
[0098] As described above, the discrimination results output by the image discrimination model are used to indicate the probability values of each input image belonging to a rendered image and a generated image, respectively. This probability value can be a probability value between 0 and 1, used to indicate the likelihood that each input image belongs to a rendered image or a generated image. The larger the probability value of belonging to a rendered image, the more likely the corresponding input image is a rendered image; the larger the probability value of belonging to a generated image, the more likely the corresponding input image is a generated image (i.e., a pseudo image) output by the image generation model.
[0099] For example, in this application, the image input to the image discrimination model includes a third facial image and a preset rendered facial image. The image discrimination model outputs the following discrimination results: the third facial image has a probability value of 0.2 for being a rendered image and a probability value of 0.8 for being a generated image; the preset rendered facial image has a probability value of 0.9 for being a rendered image and a probability value of 0.1 for being a generated image.
[0100] Step 5: Based on the discrimination result and the first loss function, calculate the first loss function value and adjust the parameters of the image generation model and the image discrimination model according to the first loss function value.
[0101] In the above, the first loss function can be a cross entropy loss function, which can be specifically referred to Formula 1:
[0102]
[0103] Among them, N is the number of input images of the input image discrimination model, y i is the value of the category of the input image i. When the category is a rendered image, y i The value of is 1; when the category is generated image, y i The value of p is 0. i is the probability value of judging that the input image i is a rendered image, p' i To determine the probability value of the input image i being the generated image, the logarithmic base of log can be e, and the specific value of the first loss function L is the first loss function value.
[0104] In an embodiment of the present application, based on the discrimination result and the first loss function, the first loss function value is calculated and the parameters (such as weights) of the image generation model and the image discrimination model are adjusted according to the first loss function value. Specifically, back propagation and an optimization algorithm (such as the Adam optimization algorithm) can be used to achieve this.
[0105] Repeat steps 2-5 until the first convergence condition is met, i.e., the trained model converges. The image generation model with adjusted parameters is used as the expression transfer model. Specifically, the generated image output by the trained expression transfer model is sufficiently close to the rendered image. The first convergence condition can be, for example, that the first loss function value is less than a preset loss threshold.
[0106] In an embodiment of the present application, by training the image generation model and the image discrimination model, the trained image generation model is capable of outputting a third facial image (i.e., a generated image) of the target virtual character with the target facial expression based on the expression feature vector, and the image quality of the generated image is sufficiently close to that of the rendered image. The trained image generation model is determined as the expression transfer model.
[0107] Compared to the prior art, the expression transfer method provided in the embodiment of the present application inputs the expression feature vector into the expression transfer model, and the expression transfer model can output a third facial image (i.e., a generated image) of the target virtual character with the target facial expression. The expression transfer effect of the present application is good, which can save a large number of manual repair processes of relevant personnel, save a lot of manpower costs and greatly improve the efficiency of expression animation production. Furthermore, the picture quality of the generated image is close enough to the rendered image, so that the expression transfer method provided by the present application can create facial expressions for virtual characters in animation, game development and other projects, which greatly improves the efficiency of creating facial expressions of virtual characters and extremely reduces the threshold for producing facial expressions of virtual characters.
[0108] In an optional implementation manner, before step S40 "inputting the second facial image into the expression parameter determination model to obtain the target expression driving parameters output by the expression parameter determination model", the expression transfer method provided in the embodiment of the present application further includes steps S401-S402:
[0109] S401: Obtain a second network model and a differentiable rendering network model.
[0110] The second network model is configured such that its input is a facial image of a target virtual character with a target facial expression, and its output is a 3D model expression driving parameter corresponding to the target facial expression. The network structure of the second network model can be a convolutional neural network, for example only.
[0111] The aforementioned differentiable neural network, also known as a differentiable neural renderer, is a pre-trained convolutional neural network whose input is expression-driven parameters and whose output is a facial image of an avatar's facial expression. Because the entire convolutional neural network is differentiable and is used to replace analog rendering engines for rendering avatar facial images, it is called a differentiable neural renderer.
[0112] S402. For each virtual character facial image in the plurality of virtual character facial images, repeatedly perform the second step to train the second network model until the first convergence condition is met, and determine the trained second network model as the expression parameter determination model.
[0113] The second step includes S4021-S4023:
[0114] S4021. Input the virtual character's facial image into the second network model to obtain the first expression driving parameter output by the second network model.
[0115] S4022: Input the first expression driving parameter into the differentiable rendering network model to obtain a predicted facial image output by the differentiable rendering network model.
[0116] S4023. Calculate a loss value based on the virtual character's facial image, the predicted facial image, and the second loss function, and adjust parameters of the second network model based on the loss value.
[0117] As mentioned above, the second loss function may be a mean square error loss function (ie, MSE loss function).
[0118] In one optional embodiment, the second loss function includes a parameter constraint sub-function and an image pixel constraint sub-function. The parameter constraint sub-function is used to constrain the first expression driving parameter to be consistent with the second expression driving parameter, where the second expression driving parameter is the expression driving parameter of the three-dimensional model corresponding to the virtual character in the virtual character's facial image. The image pixel constraint sub-function is used to constrain the pixels of the virtual character's facial image and the predicted facial image to be consistent.
[0119] In this embodiment of the present application, the second loss function consists of two parts: one is the animation parameter constraint, which is used to constrain the generated second expression driving parameters to be consistent with the actual expression driving parameters; the other is the image pixel constraint, which is used to constrain the input image and the image generated by the differentiable renderer to be consistent. This step can be completed using backpropagation and optimization algorithms (such as the Adam optimizer).
[0120] In this embodiment, a differentiable rendering network model is used to train an expression parameter determination model whose input is a virtual character's facial image and whose output is expression driving parameters. Therefore, by inputting the virtual character's facial image into the expression parameter determination model, expression driving parameters are obtained that can drive the 3D model to produce the facial expression conveyed by the virtual character's facial image.
[0121] In an optional implementation manner, before step S202 "inputting the anchor sample image, positive sample image, negative sample image, and their respective annotation information in the sample triple into the first network model for feature extraction to obtain third expression feature vectors corresponding to the anchor sample image, positive sample image, and negative sample image respectively", the expression transfer method provided in the embodiment of the present application further includes steps S501-S503:
[0122] S501: Perform labeling processing on the sample triples based on a preset labeling model to obtain first labeling information corresponding to the sample triples output by the preset labeling model.
[0123] S502. Obtain target labeling information corresponding to the sample triplet. If the first labeling information and the target labeling information are inconsistent, train the preset labeling model. Repeat the above steps until a second convergence condition is met, and determine the preset labeling model after training as the target labeling model.
[0124] S503 : labeling multiple sample triples according to the target labeling model to obtain labeling information corresponding to each of the sample triples.
[0125] The above steps S501-S503 are described in detail below.
[0126] In an embodiment of the present application, a large amount of sample data (i.e., sample triples and corresponding annotation information) is provided through human-computer collaboration to continuously support the iteration and update of the expression feature extraction model and the expression migration model. A large number of sample triples can be manually annotated to obtain the annotation information of the sample triples. In a crowdsourcing scenario (i.e., a large annotation task is divided into multiple small sub-annotation tasks, and the sub-tasks are assigned to different user subjects to complete the annotation tasks), considering that the reliability of the annotation information of a single user is low, multiple users are introduced to participate in the annotation, and the number of annotations is automatically controlled by relying on interval estimation and truth inference algorithms to obtain more reliable annotation information recognized by the public, so as to ensure the accuracy and reliability of the annotation information corresponding to the sample triples.
[0127] In an embodiment of the present application, after a large number of expression task annotations, portrait data of each user can be obtained. This portrait data can not only be used for the access, automatic quality inspection and other links of the current expression sample triple task to improve the accuracy of the annotation information, but can also be migrated to more abundant expression-related annotation tasks in the future, and a group of expression-related field experts can be obtained from crowdsourcing users.
[0128] Furthermore, considering that labeling a large number of sample triples requires a lot of manpower and time, in order to reduce the labeling cost and improve the labeling efficiency, this application introduces a preset labeling model (i.e., regarded as an AI worker) to label the sample triples and obtain the first labeling information corresponding to the sample triples output by the preset labeling model.
[0129] In an embodiment of the present application, before manual labeling, a preset labeling model can be used to label the sample triples to obtain the first labeling information corresponding to the sample triples. After manually labeling the sample triples, the target labeling information corresponding to the sample triples is obtained. The target labeling information of the sample triples is sent to the preset labeling model. The preset labeling model is trained based on the first labeling information and the target labeling information to gradually improve the labeling accuracy of the preset labeling model. The preset labeling model after training is determined to be the target labeling model. The target labeling model can then be used directly to label the sample triples. In this application, labeling the sample triples using the target labeling model can reduce the cost of manual labeling and improve the efficiency of labeling the sample triples. For example, if a piece of data originally requires 6 labelers to label, if the labeling results are consistent, it is a valid piece of data. If the first labeling information after pre-labeling by the target labeling model can be labeled by 3 labelers first, if the target labeling information labeled by the 3 labelers is consistent with the first labeling information, it can be considered a valid labeling information, thus saving 50% of the manual labeling cost.
[0130] In the embodiments of the present application, manual labeling can be considered a process of correcting and improving the preset labeling model. For example, if the preset labeling model labels a sample triple as '1' while the manual labeling result is '0', then the preset labeling model can be considered to have an inaccurate labeling result for that sample triple. Such inaccurately labeled sample triplets can be collected and used to train the preset labeling model to improve its labeling accuracy. Subsequently, the preset labeling model with improved labeling accuracy can be used for pre-labeling again. This process can be repeated, iteratively, to continuously improve its accuracy.
[0131] For example, in combination Figure 4 , an exemplary explanation of the algorithm iteration based on human-machine collaboration and data closed loop, Figure 4 This is a flowchart of an algorithm iteration based on human-machine collaboration and data closure provided in one embodiment of the present application.
[0132] like Figure 4As shown, the algorithm iteration process includes a data closed-loop process and a business processing process. In the data closed-loop process, the expression feature extraction model and the expression transfer model are mainly trained based on the sample triples and the corresponding annotation information. Subsequently, in the business processing process, a face image with a target facial expression (such as laughing) is input, and the expression feature extraction model and the expression transfer model trained in the data closed-loop process are used to obtain the face image of the virtual character with the target facial expression (such as laughing) output by the expression transfer model. It is judged whether the face image of the virtual character with the target facial expression output by the expression transfer model meets the requirements. If not, the input face image is collected. Subsequently, in the data closed-loop process, sample triples are constructed using this type of face image, and the expression feature extraction model and the expression transfer model are iteratively trained.
[0133] The following is a detailed description of the data closed-loop process.
[0134] A data cold start refers to the initial phase of a business, when, without the first batch of data, it is impossible to use the target annotation model or facial expression motion capture results to screen facial images. At this time, constructing sample triplets requires completely random sampling, and the annotation process relies solely on manual labeling. During the data cold start phase, manual labeling is used to annotate sample triplets to obtain annotation information. During the data hot start phase, sample triplets can be annotated using a collaborative approach of manual labeling and the target annotation model, or they can be annotated solely using the target annotation model.
[0135] The sample triples and corresponding annotation information are input into an expression feature extraction model, which is then trained to enable it to extract facial expression features. The expression feature vectors output by the expression feature extraction model are input into an expression transfer model, which is then trained to enable it to transfer expressions, i.e., to output a facial image of a target virtual character with a target facial expression.
[0136] The expression transfer device provided in the present application is described below. The expression transfer device described below and the expression transfer method described above can be referenced to each other.
[0137] Figure 5 This is a schematic diagram of the structure of the expression transfer device provided in one embodiment of the present application. Figure 5 As shown, the expression migration device includes: an acquisition module 501, a feature extraction module 502 and a processing module 503.
[0138] an acquisition module, configured to acquire a first facial image containing a target facial expression;
[0139] A feature extraction module is used to input the first facial image into an expression feature extraction model to obtain a target expression feature vector output by the expression feature extraction model for characterizing the target facial expression; wherein the expression feature extraction model is obtained by training a first network model based on multiple sample triplets and annotation information corresponding to each of the sample triplets, each sample triplet includes three first sample images, and the annotation information is used to describe the facial expressions corresponding to the three first sample images in the sample triplet.
[0140] The processing module is used to input the expression feature vector into an expression transfer model to obtain a second facial image of the target virtual character having the target facial expression output by the expression transfer model.
[0141] Optionally, the device further includes a driving module, and the driving module is specifically configured to:
[0142] Inputting the second facial image into an expression parameter determination model to obtain target expression driving parameters output by the expression parameter determination model, wherein the expression parameter determination model is used to generate expression driving parameters that cause the face of the three-dimensional model to present the target facial expression;
[0143] The target expression driving parameters are input into the three-dimensional model corresponding to the target virtual character, so that the face of the three-dimensional model presents the target facial expression.
[0144] Optionally, the device further includes a first training module, wherein the first training module is specifically configured to:
[0145] Build image generation model and image discrimination model;
[0146] For each second facial image in a plurality of second sample images, executing the first step to train the image generation model and the image discrimination model until a first convergence condition is satisfied, and determining the trained image generation model as the expression transfer model, wherein the second sample image is a facial image containing a facial expression; the first step includes:
[0147] Inputting the second sample image into the expression feature extraction model to obtain a first expression feature vector output by the expression feature extraction model;
[0148] Inputting the first expression feature vector into the image generation model to obtain a third facial image of the target virtual character having a facial expression output by the image generation model;
[0149] The third facial image and the preset rendered facial image are input as two input images into the image discrimination model for discrimination, thereby obtaining a discrimination result output by the image discrimination model; wherein the preset rendered facial image is a two-dimensional rendered image obtained by rendering a three-dimensional face model, and the discrimination result is used to indicate the probability value of each input image belonging to the rendered image and the generated image respectively;
[0150] Based on the discrimination result and the first loss function, a first loss function value is calculated and parameters of the image generation model and the image discrimination model are adjusted according to the first loss function value.
[0151] Optionally, the device further includes a second training module, wherein the second training module is specifically configured to:
[0152] Obtain a second network model and a differentiable rendering network model;
[0153] For each of the plurality of virtual character facial images, repeatedly performing the second step to train the second network model until a second convergence condition is met, and determining the trained second network model as the expression parameter determination model;
[0154] The second step includes:
[0155] Inputting the virtual character's facial image into the second network model to obtain a first expression driving parameter output by the second network model;
[0156] Inputting the first expression driving parameter into the differentiable rendering network model to obtain a predicted facial image output by the differentiable rendering network model;
[0157] According to the virtual character facial image, the predicted facial image and the second loss function, a second loss function value is calculated and parameters of the second network model are adjusted according to the second loss function value.
[0158] Optionally, the second loss function includes a parameter constraint sub-function and an image pixel constraint sub-function;
[0159] Among them, the parameter constraint sub-function is used to constrain the first expression driving parameter and the second expression driving parameter to be consistent, and the second expression driving parameter is the expression driving parameter of the three-dimensional model corresponding to the virtual character in the virtual character facial image; the image pixel constraint sub-function is used to constrain the pixels of the virtual character facial image and the predicted facial image to be consistent.
[0160] Optionally, the device further includes a third training module, wherein the third training module is specifically configured to:
[0161] Obtaining the multiple sample triplets and annotation information corresponding to each of the sample triplets, wherein the three first sample images included in the sample triplets are respectively an anchor sample image, a positive sample image, and a negative sample image, the anchor sample image and the positive sample image are two facial images with the same or similar facial expressions, and the anchor sample image and the negative sample image are two facial images with different or dissimilar facial expressions;
[0162] Inputting the anchor sample image, the positive sample image, the negative sample image and their respective annotation information in the sample triplet into the first network model for feature extraction, and obtaining third expression feature vectors corresponding to the anchor sample image, the positive sample image and the negative sample image respectively;
[0163] Calculating a third loss function value based on a third loss function according to the third expression feature vectors corresponding to the anchor sample image, the positive sample image, and the negative sample image;
[0164] The parameters of the first network model are adjusted based on the third loss function value to obtain the expression feature extraction model; the parameter adjustment is used to shorten the distance between the anchor point sample image and the positive sample image in the expression feature space, and to increase the distance between the anchor point sample image and the negative sample image in the expression feature space.
[0165] Optionally, the device further includes a fourth training module, wherein the fourth training module is specifically configured to:
[0166] Performing labeling processing on the sample triples based on a preset labeling model to obtain first labeling information corresponding to the sample triples output by the preset labeling model;
[0167] Obtain target labeling information corresponding to the sample triplet; if the first labeling information and the target labeling information are inconsistent, train the preset labeling model, repeat the above steps until a second convergence condition is met, and determine the trained preset labeling model as the target labeling model;
[0168] A labeling process is performed on multiple sample triples according to the target labeling model to obtain labeling information corresponding to each of the sample triples.
[0169] Optionally, the first sample images include facial images of real people and facial images of virtual characters.
[0170] The expression migration device provided in this embodiment can be used to implement the technical solution of the above-mentioned expression migration method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.
[0171] Figure 6 A schematic diagram of the hardware structure of an electronic device provided in one embodiment of the present application is shown in FIG. Figure 6 As shown, the electronic device 600 of this embodiment includes: a processor 601 and a memory 602;
[0172] Memory 602, for storing computer-executable instructions;
[0173] The processor 601 is configured to execute the computer-executable instructions stored in the memory to implement the various steps of the expression transfer method in the above embodiment. For details, please refer to the relevant description in the above method embodiment.
[0174] Optionally, the memory 602 may be independent or integrated with the processor 601 .
[0175] When the memory 602 is independently provided, the electronic device further includes a bus 603 for connecting the memory 602 and the processor 601 .
[0176] One embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the technical solution corresponding to the expression migration method in any of the above embodiments executed by the electronic device is implemented.
[0177] One embodiment of the present application also provides a computer program product, which includes: a computer program, which is stored in a readable storage medium, and at least one processor of an electronic device can read the computer program from the readable storage medium, and at least one processor executes the computer program so that the electronic device executes the technical solution corresponding to the expression migration method in any of the above embodiments.
[0178] Although the present application is disclosed as above with the preferred embodiments, it is not intended to limit the present application. Any person skilled in the art may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.
[0179] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or module, which can be electrical, mechanical or other forms.
[0180] The above-mentioned integrated module implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above-mentioned software functional module is stored in a storage medium and includes a number of instructions for causing an electronic device (which can be a personal computer, server, or network device, etc.) or a processor to perform some steps of the method described in various embodiments of the present application.
[0181] It should be understood that the processor described above may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), or application-specific integrated circuits (ASICs). A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.
[0182] The memory may include a high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk or an optical disk.
[0183] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be classified into address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.
[0184] The storage medium may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0185] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0186] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for expression transfer, characterized in that: The method comprises: Acquire a first facial image containing a target facial expression; Inputting the first facial image into an expression feature extraction model to obtain a target expression feature vector output by the expression feature extraction model for characterizing the target facial expression; wherein the expression feature extraction model is obtained by training a first network model based on a plurality of sample triplets and annotation information corresponding to each of the sample triplets, each sample triplet including three first sample images, and the annotation information being used to describe the facial expressions corresponding to the three first sample images in the sample triplet; The target expression feature vector is input into an expression transfer model to obtain a second facial image of the target virtual character having the target facial expression output by the expression transfer model.
2. The method according to claim 1, characterized in that The method further comprises: Inputting the second facial image into an expression parameter determination model to obtain target expression driving parameters output by the expression parameter determination model, wherein the expression parameter determination model is used to generate expression driving parameters that cause the face of the three-dimensional model to present the target facial expression; The target expression driving parameters are input into the three-dimensional model corresponding to the target virtual character, so that the face of the three-dimensional model presents the target facial expression.
3. The method according to claim 1, characterized in that The method further comprises: Build image generation model and image discrimination model; For each of the plurality of second sample images, executing the first step to train the image generation model and the image discrimination model until a first convergence condition is satisfied, and determining the trained image generation model as the expression transfer model, wherein the second sample image is a facial image containing a facial expression; the first step comprises: Inputting the second sample image into the expression feature extraction model to obtain a first expression feature vector output by the expression feature extraction model; Inputting the first expression feature vector into the image generation model to obtain a third facial image of the target virtual character having a facial expression output by the image generation model; The third facial image and the preset rendered facial image are input as two input images into the image discrimination model for discrimination, thereby obtaining a discrimination result output by the image discrimination model; wherein the preset rendered facial image is a two-dimensional rendered image obtained by rendering a three-dimensional face model, and the discrimination result is used to indicate the probability value of each input image belonging to the rendered image and the generated image respectively; Based on the discrimination result and the first loss function, a first loss function value is calculated and parameters of the image generation model and the image discrimination model are adjusted according to the first loss function value.
4. The method according to claim 2, characterized in that Before inputting the second facial image into the expression parameter determination model to obtain the target expression driving parameters output by the expression parameter determination model, the method further includes: Obtain a second network model and a differentiable rendering network model; For each of the plurality of virtual character facial images, repeatedly performing the second step to train the second network model until a second convergence condition is met, and determining the trained second network model as the expression parameter determination model; The second step includes: Inputting the virtual character's facial image into the second network model to obtain a first expression driving parameter output by the second network model; Inputting the first expression driving parameter into the differentiable rendering network model to obtain a predicted facial image output by the differentiable rendering network model; According to the virtual character facial image, the predicted facial image and the second loss function, a second loss function value is calculated and parameters of the second network model are adjusted according to the second loss function value.
5. The method according to claim 4, characterized in that The second loss function includes a parameter constraint sub-function and an image pixel constraint sub-function; Among them, the parameter constraint sub-function is used to constrain the first expression driving parameter and the second expression driving parameter to be consistent, and the second expression driving parameter is the expression driving parameter of the three-dimensional model corresponding to the virtual character in the virtual character facial image; the image pixel constraint sub-function is used to constrain the pixels of the virtual character facial image and the predicted facial image to be consistent.
6. The method according to claim 1, characterized in that The method further comprises: Obtaining the multiple sample triplets and annotation information corresponding to each of the sample triplets, wherein the three first sample images included in the sample triplets are respectively an anchor sample image, a positive sample image, and a negative sample image, the anchor sample image and the positive sample image are two facial images with the same or similar facial expressions, and the anchor sample image and the negative sample image are two facial images with different or dissimilar facial expressions; Inputting the anchor sample image, the positive sample image, the negative sample image and their respective annotation information in the sample triplet into the first network model for feature extraction, and obtaining third expression feature vectors corresponding to the anchor sample image, the positive sample image and the negative sample image respectively; Calculating a third loss function value based on a third loss function according to the third expression feature vectors corresponding to the anchor sample image, the positive sample image, and the negative sample image; The parameters of the first network model are adjusted based on the third loss function value to obtain the expression feature extraction model; the parameter adjustment is used to shorten the distance between the anchor point sample image and the positive sample image in the expression feature space, and to increase the distance between the anchor point sample image and the negative sample image in the expression feature space.
7. The method according to claim 6, characterized in that Before inputting the anchor sample image, the positive sample image, the negative sample image, and their respective annotation information in the sample triplet into the first network model for feature extraction to obtain third expression feature vectors corresponding to the anchor sample image, the positive sample image, and the negative sample image, the method further includes: Performing labeling processing on the sample triples based on a preset labeling model to obtain first labeling information corresponding to the sample triples output by the preset labeling model; Obtain target labeling information corresponding to the sample triplet; if the first labeling information and the target labeling information are inconsistent, train the preset labeling model, repeat the above steps until a second convergence condition is met, and determine the trained preset labeling model as the target labeling model; A labeling process is performed on multiple sample triples according to the target labeling model to obtain labeling information corresponding to each of the sample triples.
8. The method according to claim 1, characterized in that The first sample images include a facial image of a real person and a facial image of a virtual object.
9. An expression transfer device, characterized in that: The device comprises: an acquisition module, configured to acquire a first facial image containing a target facial expression; a feature extraction module, configured to input the first facial image into an expression feature extraction model to obtain a target expression feature vector output by the expression feature extraction model and used to characterize the target facial expression; wherein the expression feature extraction model is obtained by training a first network model based on a plurality of sample triplets and annotation information corresponding to each of the sample triplets, each sample triplet including three first sample images, and the annotation information being used to describe the facial expressions corresponding to the three first sample images in the sample triplet; The processing module is used to input the expression feature vector into an expression transfer model to obtain a second facial image of the target virtual character having the target facial expression output by the expression transfer model.
10. An electronic device, characterized in that: The electronic device comprises: processor; and The memory is used to store a data processing program. After the electronic device is powered on and the program is run by the processor, the expression migration method according to any one of claims 1 to 8 is executed.
11. A computer-readable storage medium, characterized in that A data processing program is stored, and the program is run by a processor to execute the expression migration method according to any one of claims 1 to 8.