Facial expression generation method and device of virtual human, electronic equipment and storage medium
By obtaining the feature vectors of the virtual person's facial expressions and using the expression degree grading and classification model to generate the virtual person's facial expressions, the problem of incoherent expressions is solved, more coordinated and accurate facial expression generation is achieved, and the user experience is improved.
Patent Information
- Application Number
- CN202510745227.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-23
AI Technical Summary
In existing virtual human facial expression generation methods, inconsistent expressions lead to a poor user experience. The expression coefficients directly obtained based on the feature vector of the facial expression to be generated may be inconsistent with other expression coefficients, resulting in inconsistent expressions in different parts of the face.
By obtaining the feature vector of the facial expression to be generated, using the pre-trained expression degree grading model and expression classification model, the expression degree level and category results are determined, the input vector is constructed and input into the expression coefficient generation model to generate the facial expression of the virtual person, and more feature mapping dimensions are used to represent the facial expression to be generated.
It improves the coordination and accuracy of generated facial expressions, enhances the robustness of virtual human facial expression generation, and improves the user experience of 3D virtual human technology.
Smart Images

Figure CN120689473A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device, electronic device and storage medium for generating facial expressions of a virtual human. Background Art
[0002] With the rapid development of artificial intelligence (AI), virtual digital human technology has been integrated into various industries, such as virtual digital lobby managers in banks, virtual digital human anchors in e-commerce live broadcasts, and virtual digital human voice assistants in television. To meet users' emotional empathy during interactions and achieve an immersive interactive experience, it is necessary to provide users with realistic generation effects in multiple aspects such as virtual human face, lip generation, and facial expressions.
[0003] At present, the commonly used method for generating facial expressions of virtual humans is to directly input the feature vector of the facial expression to be generated into a coefficient generation model based on a deep learning model to generate the expression coefficient of the facial expression to be generated, and then combine the expression coefficient and the human face template to generate the facial expression of the virtual human.
[0004] However, if the expression coefficient is obtained directly based on the feature vector of the facial expression to be generated, it may be that the expression generated by the tendency of individual expression coefficients is inconsistent with that of other expression coefficients, which may lead to the problem of incoordination of the expressions of various parts of the face in the final generated facial expression of the virtual person, resulting in a poor user experience. Summary of the Invention
[0005] The present invention provides a method, device, electronic device and storage medium for generating facial expressions of virtual humans, which are used to solve the defects of the prior art such as inconsistent facial expressions generated by virtual humans and poor user experience, and realize a solution that can generate coordinated facial expressions for virtual humans and improve user experience.
[0006] The present invention provides a method for generating facial expressions of a virtual human, comprising: Obtaining a feature vector of the facial expression to be generated; Determining the expression level and expression category result of the facial expression to be generated based on the feature vector; Generate a facial expression of a virtual person based on the expression degree level and the expression category result.
[0007] According to a method for generating facial expressions of a virtual human provided by the present invention, determining the expression degree level and expression category result of the facial expression to be generated based on the feature vector includes: The feature vector is input into a pre-trained expression degree grading model to obtain the expression degree level output by the expression degree grading model; the expression degree grading model is pre-trained based on multiple feature vector samples and their corresponding expression degree level labels; the expression degree level label is determined based on at least two expression level sub-labels; the expression level sub-label is an expression intensity sub-label, an expression duration sub-label or an expression accompanying reaction sub-label.
[0008] According to a method for generating facial expressions of a virtual human provided by the present invention, determining the expression degree level and expression category result of the facial expression to be generated based on the feature vector includes: The feature vector is input into a pre-trained expression classification model to obtain the expression category result output by the expression classification model; the expression classification model is pre-trained based on multiple feature vector samples and their corresponding expression category labels; the expression classification model is constructed based on a multi-layer perceptron model.
[0009] According to a method for generating facial expressions of a virtual person provided by the present invention, generating the facial expressions of the virtual person based on the expression degree level and the expression category result includes: Based on the expression degree level and the expression category result, an input vector is constructed; the input vector is input into a pre-trained expression coefficient generation model to obtain the expression coefficient output by the expression coefficient generation model; the expression coefficient generation model is trained based on multiple training vector samples and their corresponding expression coefficient labels; based on the expression coefficient, the facial expression is generated.
[0010] According to a method for generating facial expressions of a virtual human provided by the present invention, constructing an input vector based on the expression degree level and the expression category result includes: The input vector is constructed based on the feature vector, the expression degree level and the expression category result.
[0011] According to a method for generating facial expressions of a virtual human provided by the present invention, obtaining a feature vector of a facial expression to be generated includes: Based on user portrait information, a portrait feature vector is extracted; the user portrait information includes at least one of user personal information, user behavior information or user preference information; based on situational awareness information, a situational feature vector is extracted; the situational awareness information includes at least one of weather information, location information, festival information or event information; based on generated voice information, a voice feature vector is extracted; based on the portrait feature vector, the situational feature vector and the voice feature vector, the feature vector is determined.
[0012] According to a method for generating facial expressions of a virtual human provided by the present invention, the method of extracting speech feature vectors based on generated speech information includes: The speech feature vector of the generated speech information is input into a pre-trained semantic parsing model to obtain a semantic parsing vector output by the semantic parsing model; the speech feature vector is input into a pre-trained sentiment analysis model to obtain a sentiment analysis vector output by the sentiment analysis model; the speech feature vector is input into a pre-trained intention parsing model to obtain an intention parsing vector output by the intention parsing model; and the speech feature vector is determined based on the semantic parsing vector, the sentiment analysis vector and the intention parsing vector.
[0013] The present invention also provides a device for generating facial expressions of a virtual human, comprising: A feature acquisition module, used to obtain a feature vector of the facial expression to be generated; A feature processing module, configured to determine an expression level and an expression category result of the facial expression to be generated based on the feature vector; The expression generation module is used to generate facial expressions of a virtual person based on the expression degree level and the expression category result.
[0014] The present invention also provides an electronic device comprising a memory, a processor and a computer program stored in the memory and running on the processor, wherein when the processor executes the computer program, any of the above-mentioned methods for generating facial expressions of a virtual person is implemented.
[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the method for generating facial expressions of a virtual human as described above is implemented.
[0016] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned methods for generating facial expressions of a virtual human.
[0017] The virtual human facial expression generation method, device, electronic device and storage medium provided by the present invention first determine the expression degree level and expression category result of the facial expression to be generated based on the feature vector, and then generate the facial expression of the virtual human based on the expression degree level and expression category result, use more feature mapping dimensions to better characterize the facial expression to be generated, and ensure that the facial expression of the virtual human finally generated conforms to the corresponding expression degree level and expression category result, thereby improving the coordination and accuracy of the generated facial expression, and also improving the robustness of the facial expression generation of the virtual human, bringing users a better 3D virtual human technology usage experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 This is one of the flow charts of the method for generating facial expressions of a virtual human provided by the present invention.
[0020] Figure 2 This is the second flow chart of the method for generating facial expressions of a virtual human provided by the present invention.
[0021] Figure 3 It is a structural schematic diagram of the facial expression generation device for a virtual human provided by the present invention.
[0022] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0023] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0024] It should be noted that, in the description of the present invention, the term "comprise" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or also includes elements inherent to such process, method, article, or apparatus. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0025] The terms "first," "second," and so forth, used herein are used to distinguish similar objects, not to describe a specific order or precedence. It should be understood that such terms are interchangeable where appropriate, allowing embodiments of the present invention to be implemented in an order other than that illustrated or described herein. Furthermore, the terms "first," "second," and so forth generally distinguish objects of a single type, and do not limit the number of objects. For example, the first object may be one or more. Furthermore, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the connected objects.
[0026] The following combination Figure 1-Figure 4 The present invention describes a method, device, electronic device and storage medium for generating facial expressions of a virtual human.
[0027] It should be noted that all actions of acquiring signals, information or data in the present invention are performed in compliance with the corresponding data protection laws and policies of the region where the device is located and with the authorization given by the owner of the corresponding device.
[0028] Figure 1 This is one of the flow charts of the method for generating facial expressions of a virtual person provided by the present invention. Figure 1 As shown, the method for generating facial expressions of a virtual person includes but is not limited to steps 101 to 103.
[0029] It should be noted that the execution subject of the virtual human facial expression generation method provided by the present invention is the corresponding virtual human facial expression generation device, which can specifically be a server or computer equipment, such as a mobile phone, tablet computer, laptop computer, PDA, vehicle-mounted electronic equipment, wearable device, ultra-mobile personal computer (UMPC), netbook or personal digital assistant (PDA), etc.
[0030] Step 101: Obtain a feature vector of the facial expression to be generated.
[0031] The facial expression to be generated refers to the facial expression of a virtual person to be generated.
[0032] Specifically, relevant information of the facial expression to be generated is collected, or input information of the facial expression of the virtual person generated by the user is received, the relevant information or input information is processed, and features are extracted as feature vectors of the facial expression to be generated.
[0033] It is understandable that the feature vector of the facial expression to be generated can be extracted and determined from text information, voice information, image information, video information or multimodal information, and the present invention does not impose any limitation on this.
[0034] For example, when a user wants to generate a virtual human facial expression that is similar to his or her own image, the user's facial image information and facial video information are collected using the camera of the user terminal, and pre-set parameter text information can be combined to jointly extract and determine the feature vector.
[0035] Step 102: Based on the feature vector, determine the expression level and expression category result of the facial expression to be generated.
[0036] Among them, the expression degree level is used to represent the intensity of the expression. The higher the level, the more intense the expression. The expression type result is the expression category of the facial expression to be generated determined based on the feature vector.
[0037] For example, the expression intensity levels of the facial expression to be generated are defined as level 1, level 2, and level 3. When the expression intensity is level 1, the expression intensity is the lowest; when the expression intensity is level 3, the expression intensity is the highest.
[0038] It is understandable that the number of levels defined in the expression degree level is not limited to 3, but is determined based on a comprehensive combination of factors such as the user's facial expression generation requirements and the computing performance of the expression generation device. The present invention does not impose any restrictions on this.
[0039] For another example, the expression categories of the facial expressions to be generated are defined as including but not limited to interest, happiness, surprise, sadness, fear, shyness, contempt, anger, etc.
[0040] Specifically, based on the feature vector of the facial expression to be generated, the expression degree level and expression category result of the facial expression to be generated are generated.
[0041] Optionally, different expression degree level systems are defined for different expression category results; determining the expression degree level and expression category result of the facial expression to be generated based on the feature vector includes: first determining the expression category result of the facial expression to be generated based on the feature vector; determining the target expression degree level system based on the expression category result; and determining the expression degree level of the facial expression to be generated in combination with the feature vector and the target expression degree level system.
[0042] For example, when determining the expression level and expression category of a facial expression to be generated based on a machine learning or deep learning model, different expression level grading models are pre-built and pre-trained for different expression categories. During actual inference, the expression category of the facial expression to be generated is first determined based on the feature vector and the expression classification model. Based on the expression category, a corresponding target expression level grading model is then determined. The feature vector is then input into the target expression level grading model to obtain the expression level of the facial expression to be generated.
[0043] Step 103: Generate facial expressions of the virtual person based on the expression degree level and the expression category result.
[0044] Specifically, an expression coefficient is generated according to the expression degree level and expression category of the facial expression to be generated, and then the obtained expression coefficient is input into the virtual human model. By adjusting and applying the corresponding Blender Shape, the 3D virtual human facial expression can be generated to generate the final virtual human facial expression.
[0045] Generally speaking, the facial expression generation method of a virtual human adopts the idea of single-dimensional mapping generation, that is, the facial expression of the virtual human is directly generated according to the feature vector of the facial expression to be generated.
[0046] The facial expression generation method of a virtual human provided by the present invention first determines the expression degree level and expression category result of the facial expression to be generated based on the feature vector, and then generates the facial expression of the virtual human based on the expression degree level and expression category result. It uses more feature mapping dimensions to better characterize the facial expression to be generated, and ensures that the facial expression of the virtual human finally generated conforms to the corresponding expression degree level and expression category result, thereby improving the coordination and accuracy of the generated facial expression, and also improving the robustness of the facial expression generation of the virtual human, bringing users a better 3D virtual human technology usage experience.
[0047] Based on the above embodiment, as an optional embodiment, determining the expression degree level and expression category result of the facial expression to be generated based on the feature vector includes: Inputting the feature vector into a pre-trained expression degree grading model to obtain the expression degree level output by the expression degree grading model; The expression degree grading model is obtained by pre-training based on multiple feature vector samples and their corresponding expression degree level labels; the expression degree level label is determined based on at least two expression level sub-labels; the expression level sub-label is an expression intensity sub-label, an expression duration sub-label or an expression accompanying reaction sub-label.
[0048] Among them, facial expressions are accompanied by reactions, such as physical reactions, sound reactions, etc. that accompany facial expressions, such as laughter, crying, and changes in body movements.
[0049] Specifically, an expression degree grading model is pre-constructed and training data is collected, and the training data is processed and data annotation is performed, including data annotation of expression intensity, expression duration, expression accompanying reaction and expression level, to obtain multiple feature vector samples and their corresponding expression degree level labels, and each expression degree level label is determined based on at least two expression level sub-labels among the expression intensity sub-label, the expression duration sub-label and the expression accompanying reaction sub-label.
[0050] For example, the expression degree level label is determined based on the expression intensity sub-label and the expression duration sub-label.
[0051] For another example, the expression degree level label is determined based on the expression intensity sub-label, the expression duration sub-label, and the expression accompanying reaction sub-label.
[0052] The expression degree grading model is pre-trained using multiple feature vector samples and their corresponding expression degree level labels until the expression degree grading model meets the iteration termination conditions such as a certain accuracy, thus completing the pre-training of the expression degree grading model. In the inference stage of the virtual human's facial expression generation, the feature vector of the facial expression to be generated can be input into the pre-trained expression degree grading model, and the expression degree grading model processes the feature vector to obtain the expression degree level output by the expression degree grading model. .
[0053] Optionally, the expression degree grading model is constructed based on a multi-layer perceptron model (Multi Layer Perceptron, MLP); the expression degree grading model includes 4 hidden unit layers, and the nodes included in each hidden unit layer are 64, 128, 64 and 32 respectively, and a ReLU activation function is used; the output layer of the expression degree grading model uses a Softmax activation function.
[0054] The expression degree level is not an isolated expression representation. Taking a happy expression as an example, in terms of expression intensity, the corners of the mouth are slightly raised to form a slight smile, and the duration of the expression is short, just a momentary reaction. There is no obvious accompanying reaction in the expression, or the eyes are just slightly brighter. At this time, the expression degree level can be determined as level 1; in terms of expression intensity, the corners of the mouth are more obviously raised to form an obvious smile or a grin, and the duration of the expression is longer, lasting for a few seconds or longer. The accompanying reaction of the expression is accompanied by a slight laugh, bright eyes, and slight body movements. At this time, the expression degree level can be determined as level 2; in terms of expression intensity, the mouth is wide open to form an expression of laughing or laughing heartily; the duration of the expression is longer, lasting for several seconds or longer; the accompanying reaction of the expression is accompanied by laughter, the laughter is loud or continuous, the eyes are very bright, and the body movements are obvious. At this time, the expression degree level can be determined as level 3.
[0055] The virtual human facial expression generation method provided by the present invention sets specific expression degree grading through observation data such as the intensity, duration, and accompanying reactions of facial expressions, and utilizes the feature vectors of the facial expressions to be generated to be mapped to corresponding expression degree levels through a neural network, which can further improve the accuracy of facial expression generation processing and thereby improve the coordination of the facial expressions of the generated virtual human.
[0056] Based on the above embodiment, as an optional embodiment, determining the expression degree level and expression category result of the facial expression to be generated based on the feature vector includes: Inputting the feature vector into a pre-trained expression classification model to obtain the expression category result output by the expression classification model; The expression classification model is obtained by pre-training based on multiple feature vector samples and their corresponding expression category labels; the expression classification model is constructed based on a multi-layer perceptron model.
[0057] Specifically, an expression classification model is constructed based on a multi-layer perceptron model, and training data is collected. The training data is then processed and labeled with expression categories to obtain multiple feature vector samples and their corresponding expression category labels. The expression classification model is pre-trained using these feature vector samples and their corresponding expression category labels until the model meets certain accuracy and other iteration termination criteria, completing the pre-training process.
[0058] In the inference stage of facial expression generation of virtual people, the feature vector of the facial expression to be generated can be input into the pre-trained expression classification model, and the expression classification model processes the feature vector to obtain the expression category result output by the expression classification model. .
[0059] Optionally, the expression classification model includes 4 hidden unit layers, each hidden unit layer includes 64, 128, 64 and 32 nodes respectively, and uses a ReLU activation function; the output layer of the expression classification model uses a Softmax activation function.
[0060] The virtual human facial expression generation method provided by the present invention can further improve the accuracy of facial expression generation processing and thus improve the coordination of the generated virtual human's facial expression by mapping the feature vector of the facial expression to be generated to the corresponding expression degree level through the MLP model.
[0061] Based on the above embodiment, as an optional embodiment, generating the facial expression of the virtual person based on the expression degree level and the expression category result includes: constructing an input vector based on the expression degree level and the expression category result; Inputting the input vector into a pre-trained expression coefficient generation model to obtain an expression coefficient output by the expression coefficient generation model; the expression coefficient generation model is trained based on multiple training vector samples and their corresponding expression coefficient labels; Based on the expression coefficient, the facial expression is generated.
[0062] Optionally, the expression coefficient is determined based on the 52 expression coefficients of the ARkit standard Blender Shape standard.
[0063] Specifically, an expression coefficient generation model is pre-built and training data is collected. The training data is then processed and labeled, including labeling each coefficient according to different expression coefficient standards to obtain multiple feature vector samples and their corresponding expression coefficient labels. The expression coefficient generation model is pre-trained using these multiple feature vector samples and their corresponding expression coefficient labels until the expression coefficient generation model meets certain accuracy and other iteration termination conditions, completing the pre-training of the expression coefficient generation model.
[0064] In the inference stage of generating the facial expression of the virtual person, the expression level obtained according to the above embodiment can be and expression category results , construct the input vector, and input the input vector into the pre-trained coefficient generation model, which processes the input vector to obtain the expression coefficient output by the coefficient generation model .
[0065] Further use of expression coefficient , combined with the preset 3D virtual human model, the expression coefficient Input the virtual human model and adjust the corresponding Blender Shape to generate 3D virtual human facial expressions and generate the final virtual human facial expressions.
[0066] Optionally, the expression coefficient generation model includes an input layer, an encoder network, a decoder network and a linear activation layer; the encoder network is used to extract the features of the input vector to obtain the encoded low-dimensional feature vector; the decoder network is used to generate an output sequence based on the low-dimensional feature vector; the linear activation layer is used to generate the expression coefficient based on the output sequence mapping using a linear activation method.
[0067] Optionally, the encoder network includes multiple layers of first fully connected layers, and the activation function uses a ReLU activation function; the decoder network includes multiple layers of second fully connected layers, and the number of nodes in the second fully connected layers is the same as the number of nodes in the first fully connected layers corresponding to the encoder network; Optionally, the decoder network includes a multi-head self-attention network layer, a feedforward neural network layer and a fully connected layer network connected in sequence.
[0068] Optionally, an input vector is constructed based on the expression degree level and the expression category result; the input vector is input into a pre-trained expression coefficient generation model, and an expression for the expression coefficient output by the expression coefficient generation model is obtained as follows: ; in, is the expression coefficient; Generate a model for the expression coefficient; This is the result of expression category; The expression level.
[0069] The virtual human facial expression generation method provided by the present invention can obtain various expression coefficients that tend to generate the same expression degree level and expression category results by simultaneously using the expression degree level and expression category results to map the expression coefficients, thereby improving the coordination of facial expressions generated using the expression coefficients, and also improving the robustness of the virtual human facial expression generation, bringing users a better 3D virtual human technology usage experience.
[0070] Based on the above embodiment, as an optional embodiment, constructing an input vector based on the expression level and the expression category result includes: The input vector is constructed based on the feature vector, the expression degree level and the expression category result.
[0071] Specifically, in the inference stage of generating the facial expression of the virtual person, the Expression level and expression category results , and the feature vector of the facial expression to be generated , construct the input vector, and input the input vector into the pre-trained coefficient generation model, which processes the input vector to obtain the expression coefficient output by the coefficient generation model .
[0072] Further use of expression coefficient , combined with the preset 3D virtual human model, the expression coefficient Input the virtual human model and adjust the corresponding Blender Shape to generate 3D virtual human facial expressions and generate the final virtual human facial expressions.
[0073] Optionally, the input vector is constructed based on the feature vector, the expression degree level and the expression category result; the input vector is input into a pre-trained expression coefficient generation model, and the expression coefficient output by the expression coefficient generation model is obtained as follows: ; in, is the expression coefficient; Generate a model for the expression coefficient; is the eigenvector; This is the result of expression category; The expression level.
[0074] The virtual human facial expression generation method provided by the present invention can obtain expression coefficients that tend to generate the same expression level and expression category results by simultaneously using feature vectors, expression degree levels and expression category results, while not missing the original features of the facial expressions to be generated. It can improve the coordination of facial expressions generated using expression coefficients, generate virtual human facial expressions that better meet user needs, improve the robustness of virtual human facial expression generation, and bring users a better 3D virtual human technology usage experience.
[0075] Based on the above embodiment, as an optional embodiment, obtaining a feature vector of the facial expression to be generated includes: Extracting a portrait feature vector based on user portrait information; the user portrait information includes at least one of user personal information, user behavior information, or user preference information; Extracting a scenario feature vector based on scenario perception information, wherein the scenario perception information includes at least one of weather information, location information, festival information, or event information; Extracting speech feature vectors based on generated speech information; The feature vector is determined based on the portrait feature vector, the scenario feature vector and the voice feature vector.
[0076] Among them, user personal information includes but is not limited to at least one of information such as age, gender, occupation, etc.; user behavior information includes but is not limited to at least one of information such as user browsing history, search history, etc.; user preference information includes but is not limited to at least one of information such as preferred food, preferred travel, etc.
[0077] The generated voice information is voice information input by the user and used to instruct the generation of facial expressions of a virtual person.
[0078] Specifically, when obtaining the feature vector of the facial expression to be generated, on the one hand, the user portrait information is determined according to at least one of the user personal information, user behavior information or user preference information, and the portrait feature vector is extracted from the user portrait information. On the other hand, the situational awareness information is determined based on at least one of weather information, location information, festival information or event information, and the situational feature vector is extracted from the situational awareness information. On the other hand, according to the generated voice information input by the user, the voice feature vector is extracted from the generated voice information .
[0079] Finally, the image feature vector , scenario feature vector and speech feature vector Normalize them separately, then concatenate and fuse them to get the feature vector .
[0080] Optionally, the expression for determining the feature vector based on the portrait feature vector, the scenario feature vector, and the voice feature vector is as follows: ; in, is the eigenvector; is the portrait feature vector; is the scenario feature vector; is the speech feature vector.
[0081] Optionally, the extraction of portrait feature vectors based on user portrait information includes: preprocessing user personal information, user behavior information and user preference information to obtain portrait text information; performing word segmentation and preprocessing operations on the portrait text information to obtain a token sequence vector; and inputting the token sequence vector into a pre-trained portrait feature extraction model to obtain a portrait feature vector output by the portrait feature extraction model.
[0082] Optionally, the portrait feature extraction model is constructed based on any one of the following models, including but not limited to the Embeddings from Language Model (ELMo), the BERT (Bidirectional Encoder Representations form Transformers) model, the GPT (Generative Pretrained Transformer) model, and the XLNet (eXtensible Language Net) model.
[0083] Optionally, when the portrait feature extraction model is built based on the ELMo model, the bidirectional LSTM layer of the portrait feature extraction model is connected to an attention mechanism layer to focus on important parts of the text and improve the performance and accuracy of user portrait feature extraction.
[0084] Optionally, the extracting of scenario feature vectors based on scenario perception information includes: preprocessing weather information, location information, festival information and / or event information to obtain scenario text information; inputting the scenario text information into a pre-trained scenario feature extraction model, and having the scenario feature extraction model embed and encode the scenario text information to obtain a scenario feature vector output by the scenario feature extraction model.
[0085] Optionally, the scenario feature extraction model is constructed based on any one of including but not limited to an ELMo model, a BERT model, a GPT model and an XLNet model.
[0086] Optionally, when the scenario feature extraction model is constructed based on the ELMo model, the bidirectional LSTM layer of the scenario feature extraction model is connected to an attention mechanism layer to focus on important parts of the text and improve the performance and accuracy of scenario-aware feature extraction.
[0087] Optionally, extracting the speech feature vector based on the generated speech information includes: obtaining speech text information based on the generated speech information; inputting the speech text information into a pre-trained speech feature extraction model to obtain the speech feature vector output by the speech feature extraction model.
[0088] Optionally, the speech feature extraction model is constructed based on any one of the following models, including but not limited to the wav2vec2.0 model, the HuBERT model, and the WavLM model.
[0089] The method for generating facial expressions of virtual humans provided by the present invention parses and extracts features of user portrait information, user voice information, and situational perception information in user interaction information, and determines feature vectors of facial expressions to be generated based on the obtained multi-dimensional features. This method can make the extracted and determined feature vectors more consistent with user interaction intentions and environmental context information, thereby generating facial expressions of virtual humans that are more consistent with user interaction intentions, situations, and more accurate and coordinated, thereby improving the user's sense of immersion and interactive experience.
[0090] Based on the above embodiment, as an optional embodiment, the step of extracting a speech feature vector based on generated speech information includes: Inputting the speech feature vector of the generated speech information into a pre-trained semantic parsing model to obtain a semantic parsing vector output by the semantic parsing model; Inputting the speech feature vector into a pre-trained sentiment analysis model to obtain a sentiment analysis vector output by the sentiment analysis model; Inputting the speech feature vector into a pre-trained intention parsing model to obtain an intention parsing vector output by the intention parsing model; The speech feature vector is determined based on the semantic analysis vector, the sentiment analysis vector, and the intention analysis vector.
[0091] Among them, the semantic parsing model is obtained by pre-training using multiple speech feature vector samples and their corresponding semantic parsing labels; the sentiment analysis model is obtained by pre-training using multiple speech feature vector samples and their corresponding sentiment semantic labels; and the intention parsing model is obtained by pre-training using multiple speech feature vector samples and their corresponding intention parsing labels.
[0092] Intent parsing tags include but are not limited to office needs, online shopping, wedding purchases, reason consultation, emotional release, stress regulation, service complaints, etc. Optionally, the intent parsing tags in the intent parsing model can be dynamically added or deleted.
[0093] Specifically, a semantic parsing model, a sentiment analysis model and an intent parsing model are pre-constructed, and the semantic parsing model is pre-trained using multiple speech feature vector samples and their corresponding semantic parsing labels; the sentiment analysis model is pre-trained using multiple speech feature vector samples and their corresponding sentiment semantic labels; and the intent parsing model is pre-trained using multiple speech feature vector samples and their corresponding intent parsing labels.
[0094] When extracting the third feature vector from the user's generated voice information, the voice signal sampling rate of the generated voice information is first converted to 16kHz to obtain the generated voice information after signal conversion, and the initial voice vector is obtained by feature extraction.
[0095] On the one hand, the speech feature vector is input into the pre-trained semantic parsing model, and the semantic parsing model extracts the deep semantic features in the speech feature vector, thereby obtaining the semantic parsing vector output by the semantic parsing model. .
[0096] On the other hand, the speech feature vector is input into the pre-trained sentiment analysis model, and the sentiment analysis model extracts the emotional semantic features in the speech feature vector, thereby obtaining the sentiment analysis vector output by the sentiment analysis model. .
[0097] On the other hand, the speech feature vector is input into the pre-trained intention parsing model, and the intention parsing model extracts the intention features in the speech feature vector, thereby obtaining the intention parsing vector output by the intention parsing model. .
[0098] Finally, the semantic parsing vector , the sentiment analysis vector and the intent resolution vector Perform operations such as weighted summation of feature vectors to obtain the speech feature vector that represents the user's speech analysis features .
[0099] Optionally, the calculation formula for determining the speech feature vector based on the semantic analysis vector, the sentiment analysis vector, and the intention analysis vector is as follows: ; in, is the speech feature vector; is the semantic parsing vector; is the sentiment analysis vector; Parse vectors for intents; 、 、 Semantic parsing vectors , sentiment analysis vector , intent resolution vector The weight of .
[0100] For example, =0.7, =0.2, =0.1.
[0101] Optionally, the semantic parsing model includes 7 feature encoders; the encoders are used to obtain a first feature based on a speech feature vector; the semantic parsing model also includes 24 Transformer blocks; the Transformer blocks are used to obtain a semantic parsing vector based on the first feature.
[0102] Furthermore, each of the feature encoders includes a temporal convolutional network with 512 channels.
[0103] Optionally, the sentiment analysis model and the intent parsing model are both constructed based on a wav2vec model.
[0104] The virtual human facial expression generation method provided by the present invention can fully extract the deep features of the user's voice by simultaneously performing semantic analysis, sentiment analysis and intention analysis on the voice information of the user interaction, thereby generating a more multi-dimensional feature vector, and more accurate expression coefficients can be obtained by mapping and fusing the multi-dimensional features, so that the subsequently generated facial expressions are more in line with the user interaction context, and the user interaction experience is better.
[0105] Figure 2 This is the second flow chart of the method for generating facial expressions of a virtual person provided by the present invention. Figure 2As shown, in one embodiment, on the one hand, user personal information, user behavior information, and user preference information are used as user portrait information, and preprocessed to obtain portrait text information. The portrait text information is then segmented and preprocessed to obtain a token sequence vector. The token sequence vector is then input into a pretrained portrait feature extraction model, which then performs user portrait analysis to obtain an output portrait feature vector. On the other hand, weather information, location information, holiday information, and event information are used as context-aware information, and context-aware text information is then processed to obtain context-aware text information. The context-aware text information is then input into a pretrained context feature extraction model, which then performs context-aware analysis to obtain an output context feature vector. Furthermore, an initial speech vector is extracted from the generated speech information, and the initial speech vector is input into a speech parsing model, a sentiment analysis model, and an intent parsing model, respectively, for semantic parsing, sentiment analysis, and intent parsing, respectively, to obtain output semantic parsing vectors, sentiment analysis vectors, and intent parsing vectors. A weighted sum operation is then performed on the semantic parsing vectors, sentiment analysis vectors, and intent parsing vectors to obtain a speech feature vector.
[0106] The portrait feature vector, scene feature vector and speech feature vector are normalized separately, and then concatenated and fused to obtain the feature vector.
[0107] Afterwards, on the one hand, the feature vector is input into the pre-trained expression degree grading model to obtain the expression degree level output by the expression degree grading model; on the other hand, the feature vector is input into the pre-trained expression classification model to obtain the expression category result output by the expression classification model.
[0108] The input vector is constructed using the expression level, expression category, and feature vector of the facial expression to be generated. This input vector is then fed into a pre-trained coefficient generation model, which processes the input vector and generates the expression coefficients. The expression coefficients are then combined with a pre-set 3D virtual human model and input into the model. By adjusting and applying the corresponding Blender Shape, 3D virtual human facial expressions can be generated, resulting in the final virtual human facial expression.
[0109] The facial expression generation method for a virtual human provided by the present invention comprehensively combines user interaction information and situational perception information, extracts feature vectors and fuses them, maps expression category labels and expression degree levels, generates expression coefficients in combination with the fused feature vectors, and finally realizes the facial expression generation of a 3D virtual human. This expression coefficient mapping network that combines fusion features, expression classification, and expression grading maps and fuses multi-dimensional feature information, making the obtained expression coefficients more accurate and coordinated, making the subsequently generated facial expressions more consistent with the user interaction context, making the user interaction experience and immersion stronger, and improving the robustness and user experience of the virtual human facial expression generation.
[0110] Figure 3 Schematic diagram of the structure of the virtual human facial expression generation device provided by the present invention, such as Figure 3 As shown, the facial expression generation device of the virtual human includes but is not limited to a feature acquisition module 301 , a feature processing module 302 and an expression generation module 303 .
[0111] The feature acquisition module 301 is used to acquire the feature vector of the facial expression to be generated.
[0112] The feature processing module 302 is used to determine the expression level and expression category result of the facial expression to be generated based on the feature vector.
[0113] The expression generation module 303 is configured to generate facial expressions of a virtual person based on the expression degree level and the expression category result.
[0114] It should be noted that the virtual human facial expression generation device provided by the present invention can execute the virtual human facial expression generation method described in any of the above embodiments during specific operation, which will not be described in detail in this embodiment.
[0115] The facial expression generation device for a virtual human provided by the present invention first determines the expression degree level and expression category result of the facial expression to be generated based on the feature vector, and then generates the facial expression of the virtual human based on the expression degree level and expression category result. It uses more feature mapping dimensions to better characterize the facial expression to be generated, and ensures that the facial expression of the virtual human finally generated conforms to the corresponding expression degree level and expression category result, thereby improving the coordination and accuracy of the generated facial expression, and also improving the robustness of the facial expression generation of the virtual human, bringing users a better 3D virtual human technology usage experience.
[0116] Figure 4 Schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 4As shown, the electronic device may include: a processor (Processor) 410, a communication interface (Communications Interface) 420, a memory (Memory) 430, and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute the virtual human facial expression generation method provided in any of the above embodiments, the virtual human facial expression generation method including but not limited to the following steps: obtaining a feature vector of a facial expression to be generated; determining the expression level and expression category result of the facial expression to be generated based on the feature vector; and generating the facial expression of the virtual human based on the expression level and the expression category result.
[0117] Furthermore, the logic instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0118] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the virtual human facial expression generation method provided in any of the above embodiments. The virtual human facial expression generation method includes but is not limited to the following steps: obtaining a feature vector of the facial expression to be generated; based on the feature vector, determining the expression degree level and expression category result of the facial expression to be generated; and generating the virtual human's facial expression based on the expression degree level and the expression category result.
[0119] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for generating facial expressions of a virtual person provided in any of the above embodiments is implemented. The method for generating facial expressions of a virtual person includes but is not limited to the following steps: obtaining a feature vector of a facial expression to be generated; determining an expression degree level and an expression category result of the facial expression to be generated based on the feature vector; and generating a facial expression of the virtual person based on the expression degree level and the expression category result.
[0120] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0121] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for generating facial expressions of a virtual human, characterized in that: include: Obtaining a feature vector of the facial expression to be generated; Determining the expression level and expression category result of the facial expression to be generated based on the feature vector; Generate a facial expression of a virtual person based on the expression degree level and the expression category result.
2. The method for generating facial expressions of a virtual human according to claim 1, wherein: The step of determining the expression level and expression category of the facial expression to be generated based on the feature vector includes: Inputting the feature vector into a pre-trained expression degree grading model to obtain the expression degree level output by the expression degree grading model; The expression degree grading model is obtained by pre-training based on multiple feature vector samples and their corresponding expression degree level labels; the expression degree level label is determined based on at least two expression level sub-labels; the expression level sub-label is an expression intensity sub-label, an expression duration sub-label or an expression accompanying reaction sub-label.
3. The method for generating facial expressions of a virtual human according to claim 1, wherein: The step of determining the expression level and expression category of the facial expression to be generated based on the feature vector includes: Inputting the feature vector into a pre-trained expression classification model to obtain the expression category result output by the expression classification model; The expression classification model is obtained by pre-training based on multiple feature vector samples and their corresponding expression category labels; the expression classification model is constructed based on a multi-layer perceptron model.
4. The method for generating facial expressions of a virtual human according to claim 1, wherein: Generating the facial expression of the virtual person based on the expression degree level and the expression category result includes: constructing an input vector based on the expression degree level and the expression category result; Inputting the input vector into a pre-trained expression coefficient generation model to obtain an expression coefficient output by the expression coefficient generation model; the expression coefficient generation model is trained based on multiple training vector samples and their corresponding expression coefficient labels; Based on the expression coefficient, the facial expression is generated.
5. The method for generating facial expressions of a virtual human according to claim 4, wherein: The step of constructing an input vector based on the expression degree level and the expression category result includes: The input vector is constructed based on the feature vector, the expression degree level and the expression category result.
6. The method for generating facial expressions of a virtual human according to claim 1, wherein: The step of obtaining a feature vector of a facial expression to be generated includes: Extracting a portrait feature vector based on user portrait information; the user portrait information includes at least one of user personal information, user behavior information, or user preference information; Extracting a scenario feature vector based on scenario perception information, wherein the scenario perception information includes at least one of weather information, location information, festival information, or event information; Extracting speech feature vectors based on generated speech information; The feature vector is determined based on the portrait feature vector, the scenario feature vector and the voice feature vector.
7. The method for generating facial expressions of a virtual human according to claim 6, wherein: The method of extracting a speech feature vector based on generated speech information includes: Inputting the speech feature vector of the generated speech information into a pre-trained semantic parsing model to obtain a semantic parsing vector output by the semantic parsing model; Inputting the speech feature vector into a pre-trained sentiment analysis model to obtain a sentiment analysis vector output by the sentiment analysis model; Inputting the speech feature vector into a pre-trained intention parsing model to obtain an intention parsing vector output by the intention parsing model; The speech feature vector is determined based on the semantic analysis vector, the sentiment analysis vector, and the intention analysis vector.
8. A device for generating facial expressions of a virtual person, characterized in that: include: A feature acquisition module, used to obtain a feature vector of the facial expression to be generated; A feature processing module, configured to determine an expression level and an expression category result of the facial expression to be generated based on the feature vector; The expression generation module is used to generate facial expressions of a virtual person based on the expression degree level and the expression category result.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method for generating facial expressions of a virtual human as described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating facial expressions of a virtual human as claimed in any one of claims 1 to 7 is implemented.
11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for generating facial expressions of a virtual human as claimed in any one of claims 1 to 7 is implemented.