Virtual character driving method, device and equipment based on facial expression recognition
Real-time expression recognition technology drives virtual characters to respond to user expressions, solving the problem that traditional virtual characters cannot understand expressions, and achieving smoother and smarter interaction.
Patent Information
- Application Number
- CN202210567627.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-23
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-05-23
AI Technical Summary
Traditional virtual characters cannot understand and respond to user expressions, resulting in poor interaction process and unsmart interaction.
By obtaining the user's face image in real time, using the basic model of facial expression recognition and the multimodal alignment model for expression classification, and driving virtual characters to perform response behaviors based on the classification results.
The degree of anthropomorphism of virtual characters is improved, allowing them to respond to user expressions in a timely manner, and enhance the smoothness and intelligence of interaction.
Smart Images

Figure CN114821744B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence, deep learning, machine learning, virtual reality, etc. in computer technology, and particularly relates to a virtual character driving method, device, and equipment based on facial expression recognition. Background Art
[0002] In traditional interactions between virtual characters and humans, voice is mainly used as the carrier. The interaction between virtual characters and humans only stays at the voice level and does not have the ability to understand visual information such as human facial expressions. Virtual characters cannot make corresponding feedback based on human facial expressions. For example, during the virtual character's broadcast, if the content currently broadcast by the virtual character is not the information that the person as the interaction object wants to obtain, the person will make impatient or even angry expressions. If it were a real person interaction, they would actively ask questions to facilitate the smooth and effective progress of the current conversation, but virtual characters do not have this ability; when there is no voice from the user to interrupt the virtual character's broadcast, but there is an obvious intention to interrupt in the expression, the virtual character cannot make corresponding interruption actions, and the anthropomorphic degree of the virtual character is low, resulting in an unsmooth and unintelligent interaction process. Summary of the Invention
[0003] This application provides a virtual character driving method, device, and equipment based on facial expression recognition to solve the problem that the low anthropomorphic degree of traditional virtual characters leads to an unsmooth and unintelligent communication process.
[0004] On the one hand, this application provides a virtual character driving method based on facial expression recognition, including:
[0005] Obtain the three-dimensional image rendering model of the virtual character to provide interactive services to the user using the virtual character;
[0006] During a round of conversation between the virtual character and the user, obtain the face image of the user in real time;
[0007] Input the face image of the user into the base model and the multi-modal alignment model for facial expression recognition respectively. Determine the first expression classification result through the base model and determine the second expression classification result through the multi-modal alignment model;
[0008] Determine the target classification of the user's current expression according to the first expression classification result and the second expression classification result;
[0009] If it is determined that the target classification belongs to a preset expression classification and the response trigger condition for the current target classification is satisfied, then determine the corresponding driving data according to the response strategy corresponding to the target classification;
[0010] Drive the virtual character to execute the corresponding response behavior according to the driving data and the three-dimensional image rendering model of the virtual character.
[0011] On the other hand, the present application provides a virtual character driving device based on facial expression recognition, including:
[0012] A rendering model acquisition module, configured to acquire a three-dimensional image rendering model of a virtual character, so as to provide an interaction service to a user by using the virtual character;
[0013] A real-time data acquisition module, configured to acquire a face image of the user in real time during a round of conversation between the virtual character and the user;
[0014] A real-time facial expression recognition module, configured to input the face image of the user into a base model and a multi-modal alignment model for facial expression recognition respectively, determine a first facial expression classification result through the base model, and determine a second facial expression classification result through the multi-modal alignment model; determine a target classification of the current facial expression of the user according to the first facial expression classification result and the second facial expression classification result;
[0015] A decision-making driving module, configured to, if it is determined that the target classification belongs to a preset facial expression classification and the response triggering condition corresponding to the target classification is currently satisfied, determine corresponding driving data according to the response strategy corresponding to the target classification; drive the virtual character to execute a corresponding response behavior according to the driving data and the three-dimensional image rendering model of the virtual character.
[0016] On the other hand, the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0017] The memory stores computer-executable instructions;
[0018] The processor executes the computer-executable instructions stored in the memory to implement the method described above.
[0019] On the other hand, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method described above.
[0020] The virtual character driving method, device, and equipment provided by this application can, during a round of conversation between the virtual character and the user, obtain the user's face image in real time, input the user's face image into the base model for facial expression recognition and the multi-modal alignment model respectively, determine the first expression classification result through the base model, and determine the second expression classification result through the multi-modal alignment model; based on the first expression classification result and the second expression classification result, determine the target classification of the user's current expression, so as to accurately identify the expression classification of the user's facial expression in real time; based on the expression classification of the user's expression, when it is determined that the target classification belongs to the preset expression classification and the response trigger condition of the target classification is currently met, determine the corresponding driving data according to the response strategy corresponding to the expression classification of the user's expression, and drive the virtual character to execute the corresponding response behavior according to the driving data and the three-dimensional image rendering model of the virtual character, so that the virtual character in the output video stream makes the corresponding response behavior, enhancing the user's expression recognition ability, and driving the virtual character to make a timely response to the user's facial expression, improving the anthropomorphic degree of the virtual character and making the interaction between the virtual character and the human more smooth and intelligent. Description of the Drawings
[0021] The drawings here are incorporated into the specification and form a part of this specification, showing the embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0022] Figure 1 It is a system framework diagram of the virtual character driving method provided by this application;
[0023] Figure 2 It is a flowchart of the virtual character driving method based on expression recognition provided by an embodiment of this application;
[0024] Figure 3 It is a framework diagram of the expression recognition method provided by an exemplary embodiment of this application;
[0025] Figure 4 It is a flowchart of the virtual character driving method based on expression recognition provided by another embodiment of this application;
[0026] Figure 5 It is a flowchart of the virtual character driving method based on expression recognition provided by another embodiment of this application;
[0027] Figure 6 It is a flowchart of the virtual character driving method provided by another embodiment of this application;
[0028] Figure 7 It is a structural schematic diagram of the virtual character driving device based on expression recognition provided by an exemplary embodiment of this application;
[0029] Figure 8Schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application.
[0030] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners
[0031] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numerals in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0032] First, the terms involved in the present application are explained:
[0033] Multimodal interaction: Users can communicate with virtual characters through means such as text, voice, and expressions. The virtual characters can understand user information such as text, voice, and expressions, and can in turn communicate with users through means such as text, voice, and expressions.
[0034] Duplex interaction: A real-time and two-way interaction method where users can interrupt virtual characters at any time, and virtual characters can also interrupt themselves while speaking when necessary.
[0035] Static expression recognition: Separating a person's specific expression state from a given static image and giving a judgment on the expression type.
[0036] The virtual character driving method based on expression recognition provided by the present application relates to fields such as artificial intelligence, deep learning, machine learning, and virtual reality in computer technology, and can be specifically applied to scenarios of interaction between virtual characters and humans.
[0037] Exemplarily, common scenarios of interaction between virtual characters and humans include: intelligent customer service, government affairs consultation, life services, intelligent transportation, virtual companions, virtual anchors, virtual teachers, online games, and so on.
[0038] Aiming at the problem that the low anthropomorphic degree of traditional virtual characters leads to an unsmooth and unintelligent communication process, this application provides a virtual character driving method based on facial expression recognition. During a round of conversation between the virtual character and the user, the user's face image is obtained in real time; the user's face image is respectively input into the base model for facial expression recognition and the multimodal alignment model. The first expression classification result is determined through the base model, and the second expression classification result is determined through the multimodal alignment model; according to the first expression classification result and the second expression classification result, the target classification of the user's current expression is determined, so as to accurately and real-time recognize the user's facial expression. When it is determined that the target classification belongs to the preset expression classification and the current response trigger condition of the target classification is satisfied, the corresponding driving data is determined according to the response strategy corresponding to the target classification; according to the driving data and the three-dimensional image rendering model of the virtual character, the virtual character is driven to execute the corresponding response behavior, so that the virtual character can timely make corresponding responses to the user's expressions, improve the anthropomorphic degree of the virtual character, and make the communication process between the virtual character and people smoother and more intelligent.
[0039] Figure 1 It is the system framework diagram of the virtual character driving method based on facial expression recognition provided by this application, as Figure 1As shown, the system framework includes the following four modules: a user face image acquisition module, a real-time expression recognition module, a duplex decision-making module, and a virtual character driving module. Among them, the user face image acquisition module is used to: during the interaction between the virtual character and the user, monitor the video stream on the user side in real time, obtain the video frames on the user side, perform face detection on the video frames through a face detection algorithm, and obtain the user's face image. The real-time expression recognition module is used to: use the trained base model and multi-modal alignment model for face expression recognition to perform expression recognition on the user's face image, and recognize the expression classification and confidence of the user's current expression in the face image, so as to accurately recognize the user's facial expression in real time. The duplex decision-making module is used to: preset the preset expression classifications that need to make responses and the response strategies corresponding to each preset expression classification, and make a decision on whether the virtual character responds to the user's current expression and what response behavior to take based on the target classification and confidence of the user's current expression, as well as the current dialogue context information. Specifically, based on the target classification of the user's current expression, determine whether the target classification of the user's current expression belongs to the preset expression classification, and whether the current meets the response trigger condition of the target classification. When it is determined that the target classification of the user's current expression belongs to the preset expression classification and the current meets the response trigger condition of the target classification, determine the response strategy corresponding to the target classification. Among them, the response strategies include making expressions, announcing words, making actions, etc. The virtual character driving module is used to determine the corresponding driving data according to the response strategy corresponding to the target classification; drive the virtual character to execute the corresponding response behavior according to the driving data and the three-dimensional image rendering model of the virtual character, so that the virtual character can make a corresponding response behavior in a timely manner to the user's current expression, improving the anthropomorphic degree of the virtual character and making the communication process between the virtual character and people smoother and more intelligent.
[0040] Exemplarily, during the dialogue between the virtual character and the person, the response behaviors executed by the virtual character in response to the user's expression may include duplex interaction strategies such as interrupting the current announcement, emotional care behaviors, and taking over and assisting the dialogue process.
[0041] The following uses specific embodiments to elaborate in detail on the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems. These specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0042] Figure 2The flowchart of a virtual character driving method based on facial expression recognition provided by an embodiment of the present application. The virtual character driving method based on facial expression recognition provided by this embodiment can be specifically applied to an electronic device with the function of using a virtual character to implement human-computer interaction. The electronic device can be a dialogue robot, a terminal, a server, etc. In other embodiments, the electronic device can also be implemented by other devices, which are not specifically limited herein.
[0043] As Figure 2 shown, the specific steps of this method are as follows:
[0044] Step S201: Obtain the three-dimensional image rendering model of the virtual character to provide interactive services to users using the virtual character.
[0045] Among them, the three-dimensional image rendering model of the virtual character includes the rendering data required to implement the rendering of the virtual character. Based on the three-dimensional image rendering model of the virtual character, the bone data of the virtual character can be rendered into the three-dimensional image of the virtual character presented to the user.
[0046] The method provided by this embodiment can be applied to the scenario of interaction between a virtual character and a human. By using a virtual character with a three-dimensional image, the real-time interaction function between a machine and a human can be realized to provide intelligent services to humans.
[0047] Step S202: In a round of conversation between the virtual character and the user, obtain the user's face image in real time.
[0048] Generally, in the process of interaction between a virtual character and a human, the virtual character can have multiple rounds of conversations with the human. In each round of conversation, the video stream from the user can be monitored in real time, video frames can be sampled at a preset frequency, and face detection can be performed on the video frames to obtain the face part in the video frames, thereby obtaining the user's face image.
[0049] Among them, face detection can be performed on the video frames to obtain the user's face image, which can be implemented using common face detection algorithms and is not specifically limited herein.
[0050] Generally, the interaction scenario between a virtual character and a human is usually a one-on-one conversation scenario, that is, a scenario where one user interacts with a virtual character. If there are multiple faces in the video frame, the face images of each face in the video frame can be detected, and according to the distance between the center point of the face image and the center point of the video frame, as well as the area of the face image, one of the face images can be used as the face image of the current user.
[0051] Exemplarily, the face image with the largest area and closest to the center of the video frame can be used as the face image of the current user.
[0052] Exemplarily, if there is only one face image with the largest area, then use this face image with the largest area as the face image of the current user. If there are multiple face images with the largest area, that is, the areas of multiple face images are the same and the largest, then according to the distance between the center point of the face image with the largest area and the center point of the video frame, use the face image with the largest distance from the center point of the video frame as the face image of the current user.
[0053] After obtaining the user's face image in real time, through steps S203 - S204, perform real-time facial expression recognition processing on the user's face image to determine the expression classification of the user's current expression in the face image.
[0054] Step S203: Input the user's face image into the base model and the multi-modal alignment model for facial expression recognition respectively. Determine the first expression classification result through the base model, and determine the second expression classification result through the multi-modal alignment model.
[0055] In this embodiment, in order to improve the accuracy of expression recognition, a method of combining the base model of facial expression recognition and the multi-modal alignment model is adopted for expression recognition.
[0056] Specifically, input the user's face image into the trained base model for facial expression recognition, perform expression recognition on the user's face image through this base model to obtain the first expression classification result; and input the user's face image into the trained multi-modal alignment model, perform expression recognition on the user's face image through the multi-modal alignment model to obtain the second expression classification result.
[0057] Among them, the base model for facial expression recognition is a model trained based on the expression classification task, and an expression recognition model with better performance on the public dataset in the field of facial expression recognition can be used. For example, face alignment algorithms (Deep Alignment Network, abbreviated as DAN), MobileNet, ResNet, etc.
[0058] In order to alleviate the problem that insufficient training data affects the expression recognition effect, this solution introduces a multi-modal alignment model that has been pre-trained on a public dataset of a large amount of text and image data, and fine-tunes the pre-trained multi-modal alignment model on a small amount of training data with expression classification annotations for expression recognition, and a trained multi-modal alignment model can be obtained. This multi-modal alignment model is used for expression recognition to determine the expression classification.
[0059] Exemplarily, the multimodal alignment model may adopt any one of the models such as CoOp (Context Optimization), CLIP (Contrastive Language-Image Pre-training), Prompt Ensembling, PET (Pattern-Exploting Training), etc., which is not specifically limited here.
[0060] Step S204: Determine a target category for the user's current expression based on the first expression classification result and the second expression classification result.
[0061] After performing expression recognition on the user's facial image using a base model for facial expression recognition and a multimodal alignment model to obtain a first expression classification result and a second expression classification result, the first expression classification result and the second expression classification result are combined to determine a target classification of the user's current expression to improve the accuracy of expression recognition.
[0062] Step S205: determine whether the target category belongs to a preset expression category.
[0063] In this embodiment, preset expression categories can be configured, and the response strategy corresponding to each preset expression category can be configured. The preset expression category is the configured expression category that requires the virtual character to perform corresponding behavior. The preset expression category can include one or more expression categories.
[0064] The response strategy corresponding to the preset expression classification may include one or more response strategies. The response strategy includes the specific implementation method of the virtual character responding to the expression of the corresponding preset expression classification. The type of response strategy corresponding to each preset expression classification and the specific implementation method of each response strategy are configured independently, and the response strategies corresponding to different preset expression classifications may be different.
[0065] For example, in an optional implementation, the preset expression classification may include at least one of the following: neutral, sad, angry, happy, afraid, disgusted, and surprised. Among them, the response strategies for the three preset expression classifications of "sad", "fearful", and "disgusted" may only include the interruption strategy; the response strategy for the preset expression classification of "happy" may only include the follow-up strategy; the response strategies for the two preset expression classifications of "angry" and "surprised" may include both the interruption strategy and the follow-up strategy.
[0066] In addition, the number and type of preset expression categories, as well as the response strategy for each preset expression category, can be set differently according to different specific application scenarios, and are not specifically limited here. For example, in other optional implementations, more other preset expression categories can be set, and preset expression categories such as "sad", "fearful", and "disgusted" can also set a takeover strategy, and "angry" and "surprised" can also not set an interruption strategy, etc.
[0067] After determining the target category of the user's current expression, in this step, it is determined whether the target category of the user's current expression belongs to a preset expression category.
[0068] If the target category belongs to the preset expression category, the virtual character may need to respond to the user's current expression and continue to execute the subsequent step S206.
[0069] If the target category does not belong to the preset expression category, the virtual character does not need to respond to the user's current expression, and the virtual character continues the current processing.
[0070] Step S206: Determine whether the response trigger condition of the target classification is currently met.
[0071] In this embodiment, each preset expression classification also has a corresponding response trigger condition. Only when the corresponding response trigger condition is met can the virtual character be driven to perform a corresponding response behavior based on the response strategy corresponding to the preset expression classification.
[0072] Among them, the response trigger conditions of the preset expression classification can be configured and adjusted according to the specific application scenario, and are not specifically limited here.
[0073] In this step, it is determined whether the response triggering condition of the target classification is currently met. If the response triggering condition of the target classification is currently met, steps S207-S208 are executed to drive the virtual character to perform the corresponding response behavior according to the response strategy corresponding to the target classification.
[0074] If the response trigger condition of the target category is not met at present, the virtual character does not need to respond to the user's current expression, and the virtual character continues the current processing.
[0075] Step S207: Determine corresponding driving data according to the response strategy corresponding to the target classification.
[0076] The response strategy includes a specific implementation method for the virtual character to respond to expressions corresponding to the preset expression classification.
[0077] Exemplarily, the response strategies corresponding to the preset expression categories may include: expressions, words, actions, etc. made by the virtual character.
[0078] When it is determined that the target classification belongs to the preset expression classification and the current response trigger condition for the target classification is met, according to the response strategy corresponding to the target classification, the corresponding driving data is determined. The driving data includes all driving parameters required to drive the virtual character to execute the response strategy corresponding to the target classification.
[0079] Exemplarily, if the response strategy corresponding to the target classification includes the virtual character making a specified expression, the driving data includes expression driving parameters; if the response strategy corresponding to the target classification includes the virtual character making a specified action, the driving data includes action driving parameters; if the response strategy corresponding to the target classification includes the virtual character broadcasting a specified speech, the driving data includes voice driving parameters; if the response strategy corresponding to the target classification includes multiple response methods such as expression, speech, and action, the driving data includes the corresponding multiple driving parameters, which can drive the virtual character to execute the response behavior corresponding to the response strategy.
[0080] Step S208: According to the driving data and the three-dimensional image rendering model of the virtual character, drive the virtual character to execute the corresponding response behavior.
[0081] After determining the corresponding driving data according to the response strategy corresponding to the target classification, the skeletal model of the virtual character is driven according to the driving data to obtain the skeletal data corresponding to the response behavior, and the skeletal data is rendered according to the three-dimensional detailed rendering model of the virtual character to obtain the virtual character image data corresponding to the response behavior. By rendering the virtual character image data into the output video stream, the virtual character in the output video stream makes the corresponding response behavior, so as to realize the multi-modal duplex interaction function of the virtual character making a timely response to the user's facial expression.
[0082] In this embodiment, in a round of conversation between the virtual character and the user, the user's face image is obtained in real time, and the user's face image is respectively input into the base model for facial expression recognition and the multi-modal alignment model. The first expression classification result is determined by the base model, and the second expression classification result is determined by the multi-modal alignment model; according to the first expression classification result and the second expression classification result, the target classification of the user's current expression is determined, so as to accurately identify the expression classification of the user's facial expression in real time; based on the expression classification of the user's expression, when it is determined that the target classification belongs to the preset expression classification and the current response trigger condition for the target classification is met, according to the response strategy corresponding to the expression classification of the user's expression, the corresponding driving data is determined, and the virtual character is driven to execute the corresponding response behavior according to the driving data and the three-dimensional image rendering model of the virtual character, so that the virtual character in the output video stream makes the corresponding response behavior, increasing the real-time recognition ability of the user's expression, and driving the virtual character to make a timely response to the user's facial expression, improving the anthropomorphic degree of the virtual character, and making the interaction between the virtual character and the person smoother and more intelligent.
[0083] In an alternative embodiment, in order to improve the accuracy of facial expression recognition, a method of combining a base model for facial expression recognition with a multi-modal alignment model is used for facial expression recognition.
[0084] In the above step S203, the user's facial image is input into a trained base model for facial expression recognition, and the base model performs facial expression recognition on the user's facial image to obtain a first facial expression classification result; and the user's facial image is input into a trained multi-modal alignment model, and the multi-modal alignment model performs facial expression recognition on the user's facial image to obtain a second facial expression classification result.
[0085] Among them, the first facial expression classification result includes a set of confidence levels corresponding to all facial expression classifications, including the first confidence level that the user's current facial expression belongs to each facial expression classification. The greater the first confidence level, the higher the possibility that the user's current facial expression belongs to this facial expression classification.
[0086] Among them, the second facial expression classification result includes another set of confidence levels corresponding to all facial expression classifications, including the second confidence level that the user's current facial expression belongs to each facial expression classification. The greater the second confidence level, the higher the possibility that the user's current facial expression belongs to this facial expression classification.
[0087] In the above step S204, according to the first confidence level and the second confidence level that the user's current facial expression belongs to each facial expression classification, the target classification of the user's current facial expression and the confidence level that the user's current facial expression belongs to the target classification are determined.
[0088] Optionally, according to the first facial expression classification result and the second facial expression classification result, the facial expression classification corresponding to the maximum confidence level in the two sets of confidence levels is used as the target classification of the current user's facial expression.
[0089] Optionally, in this step, according to the first facial expression classification result and the second facial expression classification result, the mean value of the first confidence level and the second confidence level corresponding to the same facial expression classification in the two sets of confidence levels can be calculated as the third confidence level of this facial expression classification; according to the third confidence levels of each facial expression classification, the facial expression classification with the maximum third confidence level is used as the target classification of the current user's facial expression.
[0090] Exemplarily, Figure 3 is a framework diagram of the facial expression recognition method provided by an exemplary embodiment of the present application. Taking the base model using the DAN model for facial expression recognition and the multi-modal alignment model using the CoOp model for facial expression recognition as an example, as Figure 3As shown, the face images of the user obtained in real time are respectively input into the DAN model and the CoOp model. The face images are encoded by the image encoder of the DAN model, feature extraction is performed based on multiple attention modules, the features extracted by the multiple attention modules are fused based on the attention fusion module, and classification processing is performed based on the fusion result to obtain the first expression classification result. At the same time, the face images are encoded by the image encoder of the CoOp model to obtain image features, and the text information built into the model is encoded by the text encoder to obtain text features. Similarity calculation and classification processing are performed on the multimodal features (including text features and image features) to obtain the second expression classification result. The first expression classification result and the second expression classification result are combined to determine the final classification result, and the final result of the expression classification of the user's expression is obtained.
[0091] Further, based on the current application scenario, face images with various different expressions and the expression classification labels of each face image are obtained as training data, and the pre-trained CoOp model is trained until the model converges to obtain a trained CoOp model. The trained CoOp model has a high expression classification accuracy on the test set and can meet the requirements of the current application scenario. Among them, the pre-trained CoOp model refers to the CoOp model pre-trained on the public dataset.
[0092] In this embodiment, face expression recognition is performed by combining the base model for face expression recognition and the multimodal alignment model, which improves the accuracy of expression recognition. The accuracy of face expression classification on a test dataset reaches 92.9%.
[0093] Based on any of the above method embodiments, the model for face recognition in this embodiment combines the base model and the multimodal alignment model. The model has a large number of model parameters and is complex. In the case of limited hardware resources (for example, only CPU resources and no GPU resources), the inference time (respond time, abbreviated as RT) of the model is relatively high, and the efficiency of face expression recognition is relatively low.
[0094] In order to meet the real-time requirements of face expression recognition and improve the efficiency of face expression recognition, before the face images of the user are respectively input into the base model for face expression recognition and the multimodal alignment model, and the first expression classification result is determined by the base model and the second expression classification result is determined by the multimodal alignment model, the trained base model and multimodal alignment model for face expression recognition are obtained, and model distillation is performed on the base model and the multimodal alignment model to compress the model for face expression recognition, reduce the number of model parameters while ensuring that the accuracy of face expression recognition meets the requirements, reduce the model inference time, improve the efficiency of face expression recognition, and realize real-time classification recognition of face expressions.
[0095] Through model distillation, while maintaining the classification accuracy of facial expression recognition basically (reaching 91.2% on the above test dataset), the number of parameters of the model can be reduced, and the single-frame inference time on the CPU is controlled within 30 ms, achieving the effect of real-time recognition.
[0096] Optionally, model pruning technology or other model compression technologies can also be used to replace model regularization to compress the base model for face recognition and the multi-modal alignment model. Specific limitations are not provided here.
[0097] Figure 4 This is a flowchart of a virtual character driving method based on facial expression recognition provided by another embodiment of the present application. Based on any of the above method embodiments, the ability of duplex communication can be achieved, the virtual character has the ability to actively or passively interrupt its current broadcast, and can make response behaviors after interruption according to the user's facial expressions to guide the subsequent conversation process, making the interaction between the virtual character and the user more fluent and intelligent. As Figure 4 shown, the specific steps of this method are as follows:
[0098] Step S401: Obtain the three-dimensional image rendering model of the virtual character to provide interaction services to the user.
[0099] Step S402: During a round of conversation between the virtual character and the user, obtain the user's face image in real time.
[0100] Step S403: Input the user's face image into the base model for facial expression recognition and the multi-modal alignment model respectively. Determine the first expression classification result through the base model and the second expression classification result through the multi-modal alignment model.
[0101] Step S404: Determine the target classification of the user's current expression according to the first expression classification result and the second expression classification result.
[0102] The implementation manners of the above steps S401 - S404 are similar to those of the above steps S201 - S204. For specific implementation, refer to the detailed introduction in the above embodiments. Details are not repeated here in this embodiment.
[0103] Step S405: If the current conversation state is the state where the virtual character outputs and the user receives, determine whether the target classification belongs to the first preset expression classification.
[0104] Among them, the response strategy of the first preset expression classification includes an interruption strategy. When the interruption strategy is executed, it will interrupt the current processing of the virtual character and drive the virtual character to execute the response behavior corresponding to the interruption strategy.
[0105] In this embodiment, for the case where, in the dialogue state received by the virtual character output for the user, the virtual character needs to interrupt the current processing and perform a response behavior for the user's expression, a first preset expression classification is set, and an interruption strategy corresponding to each first preset expression classification is set.
[0106] The first preset expression classification and the interruption strategy corresponding to each first preset expression classification can be set and adjusted according to the needs of the actual application scenario, and no specific limitation is made here.
[0107] Exemplarily, the first preset expression classification and the interruption strategy corresponding to the first preset expression classification can be set as shown in Table 1 below.
[0108] Table 1
[0109]
[0110] The above Table 1 is only an example. The interruption strategy may not simultaneously include response behaviors such as broadcasting a speech, making a specified expression, and making a specified action, and may only include any one or any two of them. For example, it may not broadcast a speech and only make a specified expression and action; or broadcast a specified speech and make a specified expression, but not do any action, etc.
[0111] After determining the target classification of the user's current expression, it is judged in this step whether the target classification of the user's current expression belongs to the first preset expression classification.
[0112] If the target classification belongs to the first preset expression classification, the virtual character may need to perform a response process for the user's current expression, and continue to execute the subsequent step S406.
[0113] If the target classification does not belong to the first preset expression classification, the virtual character does not need to perform a response process for the user's current expression, and the virtual character continues the current processing.
[0114] Exemplarily, based on the first preset expression classification and the corresponding interruption strategy in Table 1 above, assuming that the current virtual character is broadcasting a reply message based on a question raised by the user, if it is detected that the user's expression is angry, it is determined that the virtual character may need to respond to the user's angry expression, and continue to execute the subsequent step S406 to judge whether the current meets the interruption trigger condition corresponding to the target classification. If the current meets the interruption trigger condition corresponding to the target classification, the current output of the virtual character is interrupted, and according to the interruption strategy corresponding to the target classification, the virtual character is driven to execute the corresponding interruption response behavior.
[0115] Step S406: Judge whether the current meets the interruption trigger condition corresponding to the target classification.
[0116] Among them, the interruption trigger conditions corresponding to the target classification include at least one of the following: the confidence that the user's current expression belongs to the target classification is greater than or equal to the confidence threshold corresponding to the target classification; the number of turns between the current conversation turn and the previous conversation turn that triggered the interruption is greater than or equal to the preset number of turns.
[0117] Specifically, during the interaction between the virtual character and the user, context information can be recorded, and the context information includes information on whether an interruption has occurred. Based on the current context information, it can be determined whether the number of turns between the current conversation turn and the previous conversation turn that triggered the interruption is greater than or equal to the preset number of turns.
[0118] In this embodiment, considering that the virtual character interrupting the current output to make an interruption response behavior may interfere with the normal interaction between the virtual character and the user, the interruption trigger conditions corresponding to each first preset expression classification can be set according to the needs of the actual application scenario to avoid affecting the normal interaction between the virtual character and the user and improve the fluency and intelligence of the interaction between the virtual character and the user.
[0119] Optionally, the interruption trigger conditions corresponding to the first preset expression classification can include: the confidence that the user's current expression belongs to the first preset expression classification is greater than or equal to the confidence threshold corresponding to the first preset expression classification. In this way, only when the user's expression has a relatively high confidence of being the first preset expression classification, the execution of the interruption strategy corresponding to the first preset expression classification is triggered, which can avoid affecting the normal interaction between the virtual character and the user and improve the fluency and intelligence of the interaction between the virtual character and the user.
[0120] Optionally, the interruption trigger conditions corresponding to the first preset expression classification can include: the number of turns between the current conversation turn and the previous conversation turn that triggered the interruption is greater than or equal to the preset number of turns, so as to avoid frequently interrupting the output of the virtual character and prevent continuous multi-turn interruptions from affecting the normal interaction between the virtual character and the user, and improve the fluency and intelligence of the interaction between the virtual character and the user.
[0121] Optionally, the interruption trigger conditions corresponding to the first preset expression classification can include: the confidence that the user's current expression belongs to the first preset expression classification is greater than or equal to the confidence threshold corresponding to the first preset expression classification, and the number of turns between the current conversation turn and the previous conversation turn that triggered the interruption is greater than or equal to the preset number of turns. In this way, only when the user's expression has a relatively high confidence of being the first preset expression classification and does not frequently interrupt the output of the virtual character, the execution of the interruption strategy corresponding to the first preset expression classification is triggered, which can better avoid affecting the normal interaction between the virtual character and the user and improve the fluency and intelligence of the interaction between the virtual character and the user.
[0122] Among them, the confidence thresholds corresponding to different first preset expression classifications can be different, and can be specifically set and adjusted according to the needs of the actual application scenario. The number of preset rounds can be set and adjusted according to the needs of the actual application scenario, and no specific limitation is made here in this embodiment.
[0123] Exemplarily, taking the interruption trigger condition corresponding to the target classification as: the confidence that the user's current expression belongs to the target classification is greater than or equal to the confidence threshold corresponding to the target classification; and the number of rounds between the current conversation round and the previous conversation round that triggered the interruption is greater than or equal to the number of preset rounds, and taking these two conditions being satisfied simultaneously as an example, in this step, according to the confidence that the user's current expression belongs to the target classification and the current context information, it is determined whether the current interruption trigger condition corresponding to the target classification is satisfied.
[0124] If the current interruption trigger condition corresponding to the target classification is satisfied, steps S407 - S408 are executed, and according to the interruption strategy corresponding to the target classification, the virtual character is driven to execute the corresponding interruption response behavior.
[0125] If the current interruption trigger condition corresponding to the target classification is not satisfied, the virtual character does not need to perform interruption response processing for the user's current expression, and the virtual character continues the current processing.
[0126] Step S407: If the current interruption trigger condition corresponding to the target classification is satisfied, then interrupt the current output of the virtual character, and according to the interruption strategy corresponding to the target classification, determine the corresponding driving data, where the driving data is used to drive the virtual character to execute at least one of the following interruption response behaviors: announce the words for the corresponding expression classification, make an expression with a specific emotion, and make a specified action.
[0127] In this embodiment, the interruption strategy may include at least one of the following interruption response behaviors: announce the words for the corresponding expression classification, make an expression with a specific emotion, and make a specified action. Among them, the types and specific contents of the interruption response behaviors included in the interruption strategies corresponding to different first preset expression classifications can be different.
[0128] Exemplarily, as shown in Table 1, the "sad" and "angry" corresponding interruption strategies make the same expressions and actions, but the announced words are different; the expressions and actions made in "afraid" and "disgust" are both different, and the announced words are also different.
[0129] When it is determined that the target classification belongs to the first preset expression classification and the current interruption trigger condition of the target classification is satisfied, the corresponding driving data is determined according to the interruption strategy corresponding to the target classification. This driving data includes all the driving parameters required to drive the virtual character to execute the interruption strategy corresponding to the target classification.
[0130] In this step, any existing virtual character driving method for generating driving data of a virtual character based on a determined strategy can be adopted, and no detailed description will be given here.
[0131] Step S408: According to the driving data and the three-dimensional image rendering model of the virtual character, drive the virtual character to perform at least one of the following interruption response behaviors: announce the words corresponding to the corresponding expression classification, make an expression with a specific emotion, and make a specified action.
[0132] Among them, the three-dimensional image rendering model of the virtual character includes the rendering data required to implement the rendering of the virtual character. Based on the three-dimensional image rendering model of the virtual character, the bone data of the virtual character can be rendered into the three-dimensional image of the virtual character presented to the user.
[0133] The method provided in this embodiment can be applied to the scenario of interaction between a virtual character and a person. By using a virtual character with a three-dimensional image, the real-time interaction function between a machine and a person can be realized to provide intelligent services to people.
[0134] After determining the corresponding driving data according to the response strategy corresponding to the target classification, drive the bone model of the virtual character according to the driving data to obtain the bone data corresponding to the response behavior, and render the bone data according to the three-dimensional detailed rendering model of the virtual character to obtain the virtual character image data corresponding to the response behavior. By rendering the virtual character image data into the output video stream, the virtual character in the output video stream makes the corresponding response behavior, thereby realizing the duplex interaction function of the virtual character to respond in a timely manner to the user's facial expression.
[0135] In this embodiment, for the situation where the virtual character needs to interrupt the current process and execute the response behavior for the user's expression in the dialogue state received by the user, a first preset expression classification is set, and an interruption strategy corresponding to each first preset expression classification is set. By real-time identifying the target classification of the user's expression, when it is determined that the target classification belongs to the first preset expression classification, and according to the confidence of the user's current expression belonging to the target classification and the current context information, when it is determined that the current interruption trigger condition corresponding to the target classification is met, according to the response strategy corresponding to the target classification, determine the corresponding driving data, and drive the virtual character to perform at least one of the following interruption response behaviors: announce the words corresponding to the corresponding expression classification, make an expression with a specific emotion, and make a specified action, which can avoid affecting the normal interaction between the virtual character and the user, and at the same time improve the fluency and intelligence of the interaction between the virtual character and the user.
[0136] In an optional implementation manner, after step S408, the following steps may be included:
[0137] Step S409: If the user's voice input is received within the first preset time period and the semantic information of the user's voice input is recognized, the next round of dialogue is started and the dialogue is processed according to the semantic information of the user's voice input.
[0138] Among them, the first preset duration is generally set to a shorter duration so that the user will not feel a long pause. The first preset duration can be set and adjusted according to the needs of the actual application scenario, such as hundreds of milliseconds, 1 second, or even several seconds, etc., and is not specifically limited here.
[0139] Step S410: If no voice input from the user is received within the first preset time period, or if semantic information of the voice input from the user cannot be recognized, the current output of the interrupted virtual character is continued.
[0140] Optionally, if the user's voice input is not received within the first preset time period, or the semantic information of the user's voice input cannot be recognized, the current output of the interrupted virtual character can be continued after pausing for a third preset time period to leave enough time for the user to input voice.
[0141] Among them, the third preset time length can be several hundred milliseconds, 1 second, or even several seconds, etc., and can be set and adjusted according to the needs of the actual application scenario, and is not specifically limited here.
[0142] In this embodiment, after driving the virtual character to perform an interruption response behavior, if a voice input with semantic information from the user is received within a first preset time period, a new round of conversation is started; if no voice input with semantic information from the user is received, the previous broadcast of the virtual character can be continued after a certain period of time, so as to avoid the interruption response behavior affecting the normal interaction between the virtual character and the user, thereby improving the fluency and intelligence of the interaction between the virtual character and the user.
[0143] Figure 5 This is a flowchart of a virtual character driving method based on expression recognition provided by another embodiment of the present application. Based on the above method embodiment, the interaction scheme between the virtual character and the human can have duplex capability, and the virtual character has the function of actively taking over according to the user's expression to guide the subsequent dialogue process, making the interaction between the virtual character and the user smoother and more intelligent. Figure 5 As shown, the specific steps of this method are as follows:
[0144] Step S501: Obtain a three-dimensional image rendering model of a virtual character to provide interactive services to users using the virtual character.
[0145] Step S502: During a conversation between the virtual character and the user, a facial image of the user is obtained in real time.
[0146] Step S503: input the user's facial image into the base model and the multimodal alignment model for facial expression recognition respectively, determine the first expression classification result through the base model, and determine the second expression classification result through the multimodal alignment model.
[0147] Step S504: Determine a target category for the user's current expression based on the first expression classification result and the second expression classification result.
[0148] The implementation of the above steps S501-S504 is similar to that of the above steps S201-S204. For specific implementation, please refer to the detailed description of the above embodiment, which will not be repeated in this embodiment.
[0149] Step S505: If the current dialogue state is a state where the user input is received by the virtual character, determine whether the target category belongs to the second preset expression category.
[0150] The response strategy of the second preset expression category includes a follow-up strategy. The follow-up strategy is mainly applied to the dialogue state where the user inputs and the virtual character receives, and will not obviously interrupt the user's input. The virtual character makes a follow-up response behavior that will not affect the user's input.
[0151] In this embodiment, in a dialogue state where the user inputs something and the virtual character receives it, based on the facial expressions of the user during the input process, the virtual character can be driven to simulate the situation where a real human being responds to the other party's expression in real time during the interaction, a second preset expression category is set, and a corresponding acceptance strategy is set for each second preset expression category.
[0152] The second preset expression categories and the acceptance strategies corresponding to each second preset expression category can be set and adjusted according to the needs of the actual application scenario, and are not specifically limited here.
[0153] Exemplarily, the succession strategy corresponding to the second preset expression points and the second preset expression categories may be set as shown in Table 2 below.
[0154] Table 2
[0155]
[0156] After determining the target category of the user's current expression, in this step, it is determined whether the target category of the user's current expression belongs to the second preset expression category.
[0157] If the target category belongs to the second preset expression category, the virtual character may need to respond to the user's current expression and continue to execute subsequent steps.
[0158] If the target category does not belong to the second preset expression category, the virtual character does not need to respond to the user's current expression, and the virtual character continues the current processing.
[0159] Step S506: Determine whether the triggering condition corresponding to the target category is currently met.
[0160] The triggering condition corresponding to the target classification includes: the user's expression in at least N consecutive frames of images all belongs to the target classification, where N is a positive integer and N is a preset value corresponding to the target classification.
[0161] Exemplarily, N may be 5, and the value of N may be set and adjusted according to the needs of the actual application scenario, and is not specifically limited in this embodiment.
[0162] In this embodiment, by setting the connection triggering condition corresponding to the second preset expression classification to that the user's expressions in at least N consecutive frames of images all belong to the second preset expression classification, frequent and unnecessary connection response behaviors of the virtual character can be avoided, thereby improving the fluency and intelligence of the interaction between the virtual character and the user.
[0163] In this step, if the acceptance triggering conditions corresponding to the target classification are currently met, steps S507-S508 are executed to drive the virtual character to execute the corresponding acceptance response behavior according to the acceptance strategy corresponding to the target classification.
[0164] If the triggering condition corresponding to the target classification is not met at present, the virtual character does not need to perform a response process for the user's current expression, and the virtual character continues the current process.
[0165] In addition, step S506 is an optional step. In other embodiments, when it is determined in the above step S505 that the target category belongs to the second preset expression category, step S507 can be directly executed to determine the corresponding driving data according to the acceptance strategy corresponding to the target category, and drive the virtual character to perform the acceptance response behavior according to the driving data.
[0166] Step S507, determine the corresponding driving data according to the acceptance strategy corresponding to the target classification, and the driving data is used to drive the virtual character to perform at least one of the following acceptance response behaviors: broadcasting acceptance words with a specific tone, making expressions with specific emotions, and performing prescribed actions.
[0167] In this embodiment, the follow-up strategy may include at least one of the following follow-up response behaviors: broadcasting a follow-up speech with a specific tone, making an expression with a specific emotion, and performing a prescribed action. Among them, the follow-up strategies corresponding to different second preset expression categories may include different types and specific contents of the follow-up response behaviors.
[0168] When it is determined that the target category belongs to the second preset expression category and the acceptance triggering condition of the target category is currently met, the corresponding driving data is determined according to the acceptance strategy corresponding to the target category. The driving data includes all driving parameters required to drive the virtual character to execute the acceptance strategy corresponding to the target category.
[0169] In this step, any existing virtual character driving method for generating driving data of a virtual character based on a certain strategy may be adopted, and no detailed description will be given here.
[0170] In an optional implementation, this step can also be implemented in the following manner: obtain the voice data currently input by the user, and identify the user intention information corresponding to the language data, and determine the emotional polarity corresponding to the user intention information; determine the specific tone and specific emotion used in the continuation response behavior according to the emotional polarity and the continuation strategy corresponding to the user intention information; determine the corresponding driving data according to the continuation strategy corresponding to the target classification and the specific tone and specific emotion used in the continuation response behavior, and the driving data is used to drive the virtual character to perform at least one of the following continuation response behaviors: broadcasting a continuation speech with a specific tone, making an expression with a specific emotion, and performing a prescribed action.
[0171] Exemplarily, the voice stream of user input can be obtained in real time. When it is determined that the target category belongs to the second preset expression category, the voice data input by the user in the most recent time period can be obtained and the voice data can be converted into corresponding text information; the user intention information corresponding to the text information can be identified, and the emotional polarity corresponding to the user intention information can be determined.
[0172] Usually, users Figure 1 There are 6 types in total: "expressing command", "expressing insult", "expressing question", "positive statement", "negative statement", and "other intentions". Among them, "expressing command" and "positive statement" can be classified as positive semantics, that is, the emotional polarity is positive; "expressing question" and "other intentions" can be classified as neutral semantics, that is, the emotional polarity is neutral; "expressing insult" and "negative statement" can be classified as negative semantics, that is, the emotional polarity is negative.
[0173] Among them, the sentiment polarity corresponding to the user intent information includes: positive, negative and neutral. Identifying the user intent information corresponding to the text information can be achieved through the existing natural language understanding (NLU) neural network classification model, such as TextRCNN, which balances the effect and the computational cost of the model, has a high classification accuracy, low model complexity, and low reasoning overhead. In addition, alternative models of this model include TextCNN, Transformer, etc., which will not be repeated here in this embodiment.
[0174] Exemplarily, the follow-up strategy may also include the corresponding relationship between the emotion polarity corresponding to the user intention information and the specific tone of the broadcasting words in the follow-up response behavior, the corresponding relationship between the emotion polarity corresponding to the user intention information and the specific emotion of the expression in the follow-up response behavior, and the corresponding relationship between the emotion polarity corresponding to the user intention information and the type of action in the follow-up response behavior. Based on the emotion polarity corresponding to the user intention information and the follow-up strategy, the specific tone and specific emotion used in the follow-up response behavior can be determined.
[0175] In this implementation, by identifying the emotional polarity corresponding to the user intention information of the user's current voice input, the specific tone of voice and the specific emotion of the expression of the virtual character when making a takeover response behavior in the takeover strategy are determined, so that the virtual character can make a takeover with the corresponding tone and emotion according to the user's current emotional polarity, thereby improving the degree of personification of the virtual character, increasing the user's enthusiasm for continued interaction, and improving the fluency and intelligence of the interaction between the virtual character and the user.
[0176] It should be noted that the follow-up words announced in the follow-up strategy are usually set to short content, such as "um", "yes", "right", "um", "oh oh", etc. The announcement of the follow-up words will not affect the user's normal voice input.
[0177] Step S508, according to the driving data and the three-dimensional image rendering model of the virtual character, drive the virtual character to perform at least one of the following follow-up response behaviors: broadcasting a follow-up speech with a specific tone, making an expression with a specific emotion, and performing a prescribed action.
[0178] Among them, the three-dimensional image rendering model of the virtual character includes rendering data required to realize the rendering of the virtual character. Based on the three-dimensional image rendering model of the virtual character, the skeleton data of the virtual character can be rendered into the three-dimensional image of the virtual character displayed to the user.
[0179] The method provided in this embodiment can be applied to scenarios where virtual characters interact with humans. By using virtual characters with three-dimensional images, real-time interaction between machines and humans can be achieved to provide intelligent services to humans.
[0180] In this embodiment, in the dialogue state where the user inputs and the virtual character receives it, according to the facial expressions of the user during the input process, the virtual character can be driven to simulate the situation where a real human being responds to the other party's expression in real time during the interaction process, a second preset expression category is set, and a continuation strategy corresponding to each second preset expression category is set, and the target category of the user's expression is identified in real time. When it is determined that the target category belongs to the second preset expression category and it is determined that the continuation triggering condition corresponding to the target category is currently met, the corresponding driving data is determined according to the continuation strategy corresponding to the target category and the emotional polarity of the user intention information in the user input voice data, and the virtual character is driven to perform at least one of the following continuation response behaviors: broadcasting a continuation speech with a specific tone, making an expression with a specific emotion, and performing a prescribed action, which can avoid affecting the normal interaction between the virtual character and the user, while improving the degree of anthropomorphism of the virtual character, and improving the fluency and intelligence of the interaction between the virtual character and the user.
[0181] Figure 6 This is a flowchart of a method for driving a virtual character provided by another embodiment of the present application. Based on any of the above method embodiments, the interaction scheme between the virtual character and the human can have duplex capability, and the virtual character can actively take over according to the user's input voice to guide the subsequent dialogue process, making the interaction between the virtual character and the user smoother and more intelligent. Figure 6 As shown, the specific steps of this method are as follows:
[0182] Step S601: Obtain a three-dimensional image rendering model of a virtual character to provide interactive services to users using the virtual character.
[0183] This step is consistent with the above step S201 and will not be repeated here.
[0184] Step S602: During a conversation between the virtual character and the user, voice data input by the user is obtained in real time.
[0185] Typically, during the interaction between a virtual character and a human, the virtual character can have multiple rounds of dialogues with the human, and during each round of dialogue, the virtual character can receive a voice stream from the user, that is, voice data input by the user, in real time.
[0186] Step S603: when it is detected that the silence duration of the voice data input by the user is greater than or equal to the second preset duration, if it is determined that the voice input is not finished, the voice data is converted into corresponding text information.
[0187] In this embodiment, voice activity detection (VAD) may be performed on the voice stream input by the user in real time to obtain the silence duration (ie, VAD time) input by the user.
[0188] When it is detected that the silence duration of the user input is greater than or equal to the second preset duration, if it is determined that the current round of voice input has not ended at this time, that is, the user has generated a long pause in the voice input process, in this case, subsequent response processing is performed to enable the virtual character to make a response behavior to guide the subsequent dialogue process, so that the interaction between the virtual character and the user is smoother and more intelligent.
[0189] The second preset duration is a shorter duration than the silence duration threshold, and the silence duration threshold is the silence duration for judging whether the current round of user input has ended. When the silence duration of the user's voice input reaches the silence duration threshold, it is determined that the current round of user voice input has ended. For example, the silence duration threshold may be 800ms, and the second preset duration may be 300ms. The second preset duration may be set and adjusted according to the needs of the actual application scenario, and is not specifically limited here.
[0190] Specifically, when it is detected that the silence duration of the voice data input by the user is greater than or equal to the second preset duration, if it is determined that the voice input has not ended, that is, the silence duration is less than the silence duration threshold, the voice data is converted into corresponding text information, and subsequent processing is performed based on the text information.
[0191] Step S604: Identify user intent information corresponding to the text information, and determine the sentiment polarity corresponding to the user intent information.
[0192] Among them, the emotional polarity corresponding to the user intention information includes: positive, negative and neutral.
[0193] This step can be implemented by an existing natural language understanding (NLU) algorithm, which will not be described in detail in this embodiment.
[0194] Step S605: Determine corresponding driving data according to the emotion polarity corresponding to the user intention information.
[0195] Specifically, according to the emotion polarity corresponding to the user intention information, a corresponding undertaking strategy is determined; and according to the corresponding undertaking strategy, corresponding driving data is generated.
[0196] In this embodiment, for the emotional polarity corresponding to the user's intention information when the user's voice input pauses for a long time (reaching the preset silence duration), different acceptance strategies corresponding to different emotional polarities are set. Different emotional polarities of user intention information correspond to different acceptance strategies.
[0197] Exemplarily, in the function of actively taking over according to the user's input voice, the different emotional polarities of the user's intention information can be set as the taking over strategy shown in Table 3.
[0198] Table 3
[0199]
[0200] In this step, any existing virtual character driving method for generating driving data of a virtual character based on a certain strategy may be adopted, and no detailed description will be given here.
[0201] Step S606: according to the driving data and the three-dimensional image rendering model of the virtual character, drive the virtual character to perform at least one of the following follow-up response behaviors: broadcasting the follow-up words in the specified tone configured with the emotional polarity corresponding to the user's intention information, making an expression with a specific emotion, and performing a specified action.
[0202] Among them, the broadcast of the follow-up speech in the prescribed tone configured for the emotional polarity corresponding to the user's intention information does not affect the user's voice input.
[0203] In this embodiment, the follow-up strategy may include at least one of the following follow-up response behaviors: broadcasting a follow-up speech with a specific tone, making an expression with a specific emotion, and performing a prescribed action. Among them, the follow-up strategies corresponding to different second preset expression categories may include different types and specific contents of the follow-up response behaviors.
[0204] In this step, any existing virtual character driving method for generating driving data of a virtual character based on a certain strategy may be adopted, and no detailed description will be given here.
[0205] In this embodiment, during a conversation between a virtual character and a user, voice data input by the user is obtained in real time; when it is detected that the silence duration of the voice data input by the user is greater than or equal to a second preset duration and the voice input has not ended, the user intention information and its corresponding emotional polarity are identified according to the voice data, and a corresponding takeover strategy is determined according to the emotional polarity corresponding to the current user intention information, and the virtual character is driven to perform at least one takeover response behavior according to the corresponding takeover strategy: broadcasting a takeover speech in a prescribed tone configured for the emotional polarity corresponding to the user intention information, making an expression with a specific emotion, and performing a prescribed action, which can improve the degree of personification of the virtual character without affecting the user input, and improve the fluency and intelligence of the interaction between the virtual character and the user.
[0206] It should be noted that, during the interaction between the virtual character and the user, at least two of the above embodiments may be used in combination, so that the user can perceive the virtual character's feedback ability at the visual level and gain the feeling that the "virtual character is smarter and more intelligent."
[0207] Figure 7The structural schematic diagram of a virtual character driving device based on facial expression recognition provided by an exemplary embodiment of the present application. The virtual character driving device based on facial expression recognition provided by the embodiments of the present application can execute the processing flow provided by the embodiments of the virtual character driving method based on facial expression recognition. As Figure 7 shown, the virtual character driving device 70 based on facial expression recognition includes: a rendering model acquisition module 71, a real-time data acquisition module 72, a real-time facial expression recognition module 73, and a decision-making and driving module 74.
[0208] The rendering model acquisition module 71 is used to acquire the three-dimensional image rendering model of the virtual character, so as to provide an interaction service to the user by using the virtual character.
[0209] The real-time data acquisition module 72 is used to acquire the face image of the user in real time during a round of conversation between the virtual character and the user.
[0210] The real-time facial expression recognition module 73 is used to input the face image of the user into the base model and the multi-modal alignment model for facial expression recognition respectively, determine the first expression classification result through the base model, and determine the second expression classification result through the multi-modal alignment model; according to the first expression classification result and the second expression classification result, determine the target classification of the user's current expression.
[0211] The decision-making and driving module 74 is used to, if it is determined that the target classification belongs to the preset expression classification and the response trigger condition for the target classification is currently met, determine the corresponding driving data according to the response strategy corresponding to the target classification; according to the driving data and the three-dimensional image rendering model of the virtual character, drive the virtual character to execute the corresponding response behavior.
[0212] The device provided by the embodiments of the present application can be specifically used to execute the above Figure 2 The solutions provided by the corresponding method embodiments, and the specific functions and achievable technical effects are not described in detail here.
[0213] In an optional embodiment, the first expression classification result includes: the first confidence level of the user's current expression belonging to each expression classification, and the second expression classification result includes the second confidence level of the user's current expression belonging to each expression classification.
[0214] When determining the target classification of the user's current expression according to the first expression classification result and the second expression classification result, the real-time facial expression recognition module is further used to: determine the target classification of the user's current expression and the confidence level of the user's current expression belonging to the target classification according to the first confidence level and the second confidence level of the user's current expression belonging to each expression classification.
[0215] In an alternative embodiment, if it is determined that the target classification belongs to a preset expression classification and the response trigger condition for the target classification is currently met, when determining the corresponding driving data according to the response strategy corresponding to the target classification, the decision-making and driving module is further configured to: if the current dialogue state is a state where the virtual character outputs and the user receives, and the target classification belongs to the first preset expression classification, then determine whether the current meets the interruption trigger condition corresponding to the target classification according to the confidence that the user's current expression belongs to the target classification and the current context information. The first preset expression classification has a corresponding interruption strategy; if it is determined that the current meets the interruption trigger condition corresponding to the target classification, then interrupt the current output of the virtual character and determine the corresponding driving data according to the interruption strategy corresponding to the target classification. The driving data is used to drive the virtual character to perform at least one of the following interruption response behaviors: broadcasting the lines corresponding to the expression classification, making an expression with a specified emotion, and making a specified action.
[0216] In an alternative embodiment, the interruption trigger condition corresponding to the target classification includes at least one of the following: the confidence that the user's current expression belongs to the target classification is greater than or equal to the confidence threshold corresponding to the target classification; the number of turns between the current dialogue turn and the previous dialogue turn that triggered the interruption is greater than or equal to the preset number of turns.
[0217] In an alternative embodiment, when driving the virtual character to perform the corresponding response behavior according to the driving data and the three-dimensional image rendering model of the virtual character, the decision-making and driving module is further configured to: if a voice input from the user is received within the first preset duration and the semantic information of the user's voice input is recognized, then start the next round of dialogue and process the dialogue according to the semantic information of the user's voice input; if a voice input from the user is not received within the first preset duration, or the semantic information of the user's voice input cannot be recognized, then continue the current output of the interrupted virtual character.
[0218] In an alternative embodiment, if it is determined that the target classification belongs to a preset expression classification and the response trigger condition for the target classification is currently met, when determining the corresponding driving data according to the response strategy corresponding to the target classification, the decision-making and driving module is further configured to: if the current dialogue state is a state where the user inputs and the virtual character receives, and the target classification belongs to the second preset expression classification, then determine whether the current meets the connection trigger condition corresponding to the target classification according to the target classification. The second preset expression classification has a corresponding connection strategy; if it is determined that the current meets the connection trigger condition corresponding to the target classification, then determine the corresponding driving data according to the connection strategy corresponding to the target classification. The driving data is used to drive the virtual character to perform at least one of the following connection response behaviors: broadcasting the connection lines with a specific tone, making an expression with a specific emotion, and making a specified action; wherein, broadcasting the connection lines with a specific tone does not affect the user's voice input.
[0219] In an optional embodiment, when the corresponding driving data is determined according to the acceptance strategy corresponding to the target classification, and the driving data is used to drive the virtual character to perform at least one acceptance response behavior, the decision and driving module is also used to: identify the user intention information corresponding to the language data according to the voice data currently input by the user, and determine the emotional polarity corresponding to the user intention information; determine the specific tone and specific emotion used in the acceptance response behavior according to the emotional polarity and the acceptance strategy corresponding to the user intention information; determine the corresponding driving data according to the acceptance strategy corresponding to the target classification and the specific tone and specific emotion used in the acceptance response behavior, and the driving data is used to drive the virtual character to perform at least one of the following acceptance response behaviors: broadcasting acceptance words with a specific tone, making expressions with specific emotions, and performing prescribed actions.
[0220] In an optional embodiment, the triggering condition corresponding to the target classification includes: the user's expression in at least N consecutive frames of images all belongs to the target classification, where N is a positive integer and N is a preset value corresponding to the target classification.
[0221] In an optional embodiment, the real-time data acquisition module is further used to: acquire voice data input by the user in real time during a conversation between the virtual character and the user.
[0222] The decision-making and driving module is also used for: when it is detected that the silence duration of the voice data input by the user is greater than or equal to the second preset duration, if it is determined that the voice input has not ended, converting the voice data into corresponding text information; identifying the user intention information corresponding to the text information, and determining the emotional polarity corresponding to the user intention information; determining the corresponding driving data according to the emotional polarity corresponding to the user intention information; and driving the virtual character to perform at least one of the following follow-up response behaviors according to the driving data and the three-dimensional image rendering model of the virtual character: broadcasting the follow-up words in the specified tone configured for the emotional polarity corresponding to the user intention information, making an expression with a specific emotion, and making a specified action.
[0223] Among them, the broadcast of the follow-up speech in the prescribed tone configured for the emotional polarity corresponding to the user's intention information does not affect the user's voice input.
[0224] In an optional embodiment, before the user's facial image is respectively input into a base model and a multimodal alignment model for facial expression recognition, a first expression classification result is determined by the base model, and a second expression classification result is determined by the multimodal alignment model, the real-time expression recognition module is also used to: obtain a trained base model and a multimodal alignment model for facial expression recognition; and perform model distillation on the base model and the multimodal alignment model.
[0225] The device provided by the embodiment of the present application can be specifically used to execute the solution provided by any of the above method embodiments. The specific functions and achievable technical effects are not described herein again.
[0226] Figure 8 It is a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application. As Figure 8 shown, the electronic device 80 includes: a processor 801, and a memory 802 communicatively connected to the processor 801, and the memory 802 stores computer-executable instructions.
[0227] Wherein, the processor executes the computer-executable instructions stored in the memory to implement the solution provided by any of the above method embodiments. The specific functions and achievable technical effects are not described herein again.
[0228] The embodiment of the present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the solution provided by any of the above method embodiments. The specific functions and achievable technical effects are not described herein again.
[0229] The embodiment of the present application also provides a computer program product, which includes: a computer program. The computer program is stored in a readable storage medium. At least one processor of the electronic device can read the computer program from the readable storage medium, and at least one processor executes the computer program to enable the electronic device to execute the solution provided by any of the above method embodiments. The specific functions and achievable technical effects are not described herein again.
[0230] In addition, in some of the processes described in the above embodiments and the accompanying drawings, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. They are only used to distinguish different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types. The meaning of "multiple" is more than two, unless otherwise specifically defined.
[0231] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include the common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0232] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A virtual character driving method based on facial expression recognition, characterized in that Including: Obtain a three-dimensional image rendering model of a virtual character to provide interactive services to users using the virtual character; During a round of conversation between the virtual character and the user, obtain the face image of the user in real time; Input the face image of the user into a base model for facial expression recognition and a multimodal alignment model respectively. Determine a first expression classification result through the base model and a second expression classification result through the multimodal alignment model; Determine the target classification of the user's current expression according to the first expression classification result and the second expression classification result; If it is determined that the target classification belongs to a preset expression classification and the response trigger condition for the target classification is currently met, determine the corresponding driving data according to the response strategy corresponding to the target classification; The preset expression classification includes a first preset expression classification with a corresponding interruption strategy and a second preset expression classification with a corresponding continuation strategy. The driving data includes all driving parameters required to drive the virtual character to execute the response strategy corresponding to the target classification; the interruption strategy is used to interrupt the current processing of the virtual character and drive the virtual character to execute the response behavior corresponding to the interruption strategy; the continuation strategy is used in the dialogue state where the user's input is received by the virtual character, without interrupting the user's input, and the virtual character makes a continuation response behavior that does not affect the user's input; Drive the virtual character to execute the corresponding response behavior according to the driving data and the three-dimensional image rendering model of the virtual character.
2. The method according to claim 1, wherein The first expression classification result includes: the first confidence level of the user's current expression belonging to each expression classification, and the second expression classification result includes the second confidence level of the user's current expression belonging to each expression classification. The determining the target classification of the user's current expression according to the first expression classification result and the second expression classification result includes: Determine the target classification of the user's current expression and the confidence level of the user's current expression belonging to the target classification according to the first confidence level and the second confidence level of the user's current expression belonging to each expression classification.
3. The method according to claim 1, characterized in that The if it is determined that the target classification belongs to a preset expression classification and the response trigger condition for the target classification is currently met, then determine the corresponding driving data according to the response strategy corresponding to the target classification includes: If the current dialogue state is the state where the virtual character outputs and the user receives, and the target classification belongs to the first preset expression classification, then determine whether the interruption trigger condition corresponding to the target classification is currently met according to the confidence level of the user's current expression belonging to the target classification and the current context information; if it is determined that the interruption trigger condition corresponding to the target classification is currently met, then interrupt the current output of the virtual character and determine the corresponding driving data according to the interruption strategy corresponding to the target classification. The driving data is used to drive the virtual character to execute at least one of the following interruption response behaviors: announce the words corresponding to the expression classification, make an expression with a specified emotion, make a specified action.
4. The method according to claim 3, characterized in that, The interruption trigger condition corresponding to the target classification includes at least one of the following: The confidence that the user's current expression belongs to the target classification is greater than or equal to the confidence threshold corresponding to the target classification; The number of turns between the current conversation turn and the previous conversation turn that triggered an interruption is greater than or equal to a preset number of turns.
5. The method according to claim 3, wherein After driving the virtual character to perform the corresponding response behavior according to the driving data and the three-dimensional image rendering model of the virtual character, it further includes: If a voice input from the user is received within the first preset duration and the semantic information of the user's voice input is recognized, then the next round of conversation is started, and conversation processing is performed according to the semantic information of the user's voice input; If a voice input from the user is not received within the first preset duration, or the semantic information of the user's voice input cannot be recognized, then the current output of the interrupted virtual character is continued.
6. The method according to claim 1, wherein If it is determined that the target classification belongs to a preset expression classification and the response trigger condition of the target classification is currently satisfied, then according to the response strategy corresponding to the target classification, the corresponding driving data is determined, including: If the current conversation state is the state where the user inputs and the virtual character receives, and the target classification belongs to the second preset expression classification, then according to the target classification, it is judged whether the current contact trigger condition corresponding to the target classification is satisfied; If it is determined that the current contact trigger condition corresponding to the target classification is satisfied, then according to the contact strategy corresponding to the target classification, the corresponding driving data is determined, and the driving data is used to drive the virtual character to perform at least one of the following contact response behaviors: broadcasting a contact speech with a specific tone, making an expression with a specific emotion, making a specified action; Among them, the broadcasting of the contact speech with a specific tone does not affect the user's voice input.
7. The method according to claim 6, wherein According to the contact strategy corresponding to the target classification, the corresponding driving data is determined, and the driving data is used to drive the virtual character to perform at least one of the contact response behaviors, including: According to the voice data currently input by the user, the user intention information corresponding to the voice data is recognized, and the emotional polarity corresponding to the user intention information is determined; According to the emotional polarity corresponding to the user intention information and the contact strategy, the specific tone and specific emotion used for the contact response behavior are determined; According to the contact strategy corresponding to the target classification and the specific tone and specific emotion used for the contact response behavior, the corresponding driving data is determined, and the driving data is used to drive the virtual character to perform at least one of the following contact response behaviors: broadcasting a contact speech with a specific tone, making an expression with a specific emotion, making a specified action.
8. The method according to claim 6, wherein The contact trigger condition corresponding to the target classification includes: The expressions of the user in at least N consecutive frames of images all belong to the target classification, where N is a positive integer and N is the preset value corresponding to the target classification.
9. The method according to any one of claims 1-8, characterized in that, The method further includes: During a round of conversation between the virtual character and the user, the voice data input by the user is obtained in real time; When it is detected that the silent duration of the voice data input by the user is greater than or equal to the second preset duration, if it is determined that the voice input has not ended, then the voice data is converted into the corresponding text information; Identifying user intent information corresponding to the text information, and determining the sentiment polarity corresponding to the user intent information; Determining corresponding driving data according to the emotion polarity corresponding to the user intention information; According to the driving data and the three-dimensional image rendering model of the virtual character, the virtual character is driven to perform at least one of the following follow-up response behaviors: broadcasting a follow-up speech with a prescribed tone configured for the emotional polarity corresponding to the user intention information, making an expression with a specific emotion, and performing a prescribed action; Among them, the broadcast of the follow-up speech with the specified tone configured for the emotional polarity corresponding to the user intention information does not affect the user's voice input.
10. The method according to any one of claims 1-8, characterized in that, Before inputting the facial image of the user into a base model and a multimodal alignment model for facial expression recognition, determining a first expression classification result by the base model, and determining a second expression classification result by the multimodal alignment model, the method further includes: Obtain the trained base model and multimodal alignment model for facial expression recognition; Model distillation is performed on the base model and the multimodal alignment model.
11. A virtual character driving device based on facial expression recognition, characterized in that, include: A rendering model acquisition module is used to acquire a three-dimensional image rendering model of a virtual character so as to provide interactive services to users using the virtual character; A real-time data acquisition module, used to acquire a face image of the user in real time during a conversation between the virtual character and the user; A real-time expression recognition module, used for inputting the user's facial image into a base model and a multimodal alignment model for facial expression recognition, determining a first expression classification result through the base model, and determining a second expression classification result through the multimodal alignment model; Determining a target classification of the user's current expression according to the first expression classification result and the second expression classification result; A decision-making driving module, for determining corresponding driving data according to a response strategy corresponding to the target classification if it is determined that the target classification belongs to a preset expression classification and the response triggering condition of the target classification is currently met; The preset expression classification includes a first preset expression classification with a corresponding interruption strategy and a second preset expression classification with a corresponding acceptance strategy. The driving data includes all driving parameters required to drive the virtual character to execute the response strategy corresponding to the target classification; the interruption strategy is used to interrupt the current processing of the virtual character and drive the virtual character to execute the response behavior corresponding to the interruption strategy; the acceptance strategy is used to not interrupt the user's input in the dialogue state received by the user input virtual character, and the virtual character performs the acceptance response behavior that will not affect the user input; According to the driving data and the three-dimensional image rendering model of the virtual character, the virtual character is driven to perform corresponding response behaviors.
12. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1-10 when executed by a processor.
Citation Information
Patent Citations
Multi-modal interaction method, device and system based on virtual character, storage medium and terminal
CN112162628A
Feature fusion and decision fusion mixed multi-modal emotion recognition method
CN112800875A