Virtual face evaluation method and device and storage medium
By extracting multimodal features and calculating weighted evaluation values from virtual facial data and user behavior data, the problem of inaccurate evaluation of virtual facial generation systems is solved, enabling more comprehensive evaluation of virtual facial data and improving the standardization and automation of evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UBTECH ROBOTICS CORP LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-05-05
AI Technical Summary
Existing virtual face generation systems suffer from inaccurate evaluations, inconsistent standards, and an inability to fully reflect the realism and naturalness of virtual face data.
By extracting multimodal features from virtual facial data and user behavior data, calculating multiple sub-item evaluation values, and weighting them, a comprehensive evaluation value for virtual facial data is obtained, which is then integrated with multimodal data for evaluation.
This improves the standardization and automation of virtual facial data evaluation, and more comprehensively reflects the realism and naturalness of virtual facial data.
Smart Images

Figure CN121982178A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of image processing technology, and in particular relates to a virtual face evaluation method, apparatus and storage medium. Background Technology
[0002] With the development of image processing technology, its application in virtual face generation is becoming increasingly widespread. For example, in fields such as film and television, games, live streaming, intelligent customer service, education, and interactive experiences, virtual faces are often needed to complete corresponding functions. Among these applications, the realism and naturalness of the virtual face's facial expressions significantly impact the user experience.
[0003] Virtual faces are typically generated by a virtual face generation system. These systems often use voice, text, or motion signals to drive the model and generate corresponding expression frames. Therefore, the performance of the virtual face generation system not only affects the visual quality of the virtual face but also directly determines the user experience.
[0004] In summary, there is an urgent need for a method to evaluate virtual faces generated by virtual face generation systems. Summary of the Invention
[0005] The purpose of this application is to provide a virtual face evaluation method, device, and storage medium, aiming to solve the problems of inaccurate evaluation and inconsistent standards in virtual face generation systems.
[0006] A first aspect of this application provides a virtual facial evaluation method, the method comprising: Acquire virtual face data generated by the virtual face generation system and user behavior data used to generate the virtual face data; Multimodal features of the virtual facial data and the user behavior data corresponding to the virtual facial data are extracted respectively; Based on the multimodal features, calculate multiple sub-item evaluation values for the virtual facial data; The evaluation values of the multiple sub-items are weighted to obtain the evaluation value of the virtual facial data.
[0007] In some embodiments, calculating multiple sub-item evaluation values of the virtual facial data based on the multimodal features includes: Based on the emotion vector of the virtual facial data and the target emotion corresponding to the user behavior data, determine the emotion matching evaluation value of the virtual facial data; and / or Based on the lip movement sequence of the virtual facial data and the audio features corresponding to the user behavior data, determine the lip-sync evaluation value of the virtual facial data; and / or Based on the generation parameters of any two adjacent virtual facial images in the virtual facial data, determine the naturalness evaluation value of the virtual facial data; and / or The acceptability rating of the virtual facial data is determined using a realism discrimination model; and / or When an image anomaly is detected in the virtual facial data, a penalty evaluation value is generated.
[0008] In some embodiments, the step of extracting multimodal features from the virtual facial data and the user behavior data corresponding to the virtual facial data includes: A first emotion vector representing the emotional features of the virtual facial data is identified using a facial emotion recognition system. Based on the semantic and image information in the user behavior data, identify the second emotion vector corresponding to the target emotion in the user behavior data; The step of determining the emotion matching evaluation value of the virtual facial data based on the emotion vector of the virtual facial data and the target emotion corresponding to the user behavior data includes: The emotional fit evaluation value of the virtual facial data is determined based on the difference between the first emotion vector and the second emotion vector.
[0009] In some embodiments, the step of extracting multimodal features from the virtual facial data and the user behavior data corresponding to the virtual facial data includes: Extract the audio features from the audio data of the user behavior data; Extract the motion sequence of key points of the mouth from the virtual facial data; The step of determining the lip-sync evaluation value of the virtual facial data based on the lip key point motion sequence of the virtual facial data and the audio features corresponding to the user behavior data includes: Determine the degree of correlation between the audio features and the changes in the mouth key point motion sequence; Based on the correlation of the change process, the lip-sync evaluation value of the virtual facial data is determined.
[0010] In some embodiments, the step of extracting multimodal features from the virtual facial data and the user behavior data corresponding to the virtual facial data includes: Extract the temporal features of the virtual face data and the generation parameters of each frame of the virtual face image; The step of determining the naturalness evaluation value of the virtual facial data based on the generation parameters of any two adjacent frames of virtual facial images in the virtual facial data includes: Based on the temporal characteristics of the virtual facial data, adjacent virtual facial images are determined from the virtual facial data; For any two adjacent virtual face images, determine the difference in the generation parameters of the adjacent virtual face images; The naturalness evaluation value of the virtual facial data is determined based on the difference in the generated parameters.
[0011] In some embodiments, before extracting the multimodal features of the virtual facial data and the user behavior data corresponding to the virtual facial data, the method further includes: The virtual facial data is aligned with the user behavior data based on the timestamp of the user behavior data corresponding to the generation of the virtual facial image.
[0012] In some embodiments, the method further includes: Based on the evaluation value of the virtual facial data, the parameters of the virtual facial generation system are adjusted.
[0013] A second aspect of this application provides a virtual facial evaluation device, the device comprising: The acquisition unit is used to acquire virtual face data generated by the virtual face generation system and user behavior data used to generate the virtual face data; The feature extraction unit is used to extract multimodal features of the virtual face data and the user behavior data corresponding to the virtual face data, respectively. The first calculation unit is used to calculate multiple sub-item evaluation values of the virtual facial data based on the multimodal features; The second calculation unit is used to weight the evaluation values of the multiple sub-items to obtain the evaluation value of the virtual facial data.
[0014] In some embodiments, the first calculation unit is configured to determine the emotion matching evaluation value of the virtual facial data based on the emotion vector of the virtual facial data and the target emotion corresponding to the user behavior data; The first calculation unit is used to determine the lip-sync evaluation value of the virtual facial data based on the motion sequence of the mouth key points of the virtual facial data and the audio features corresponding to the user behavior data. The first calculation unit is used to determine the naturalness evaluation value of the virtual face data based on the generation parameters of any two adjacent frames of virtual face images in the virtual face data. The first calculation unit is used to determine the acceptability evaluation value of the virtual facial data through a realism discrimination model; The first calculation unit is used to generate a penalty evaluation value when an image anomaly is detected in the virtual facial data.
[0015] In some embodiments, the feature extraction unit is configured to identify a first emotion vector representing the emotional features of the virtual facial data through a facial emotion recognition system; and to identify a second emotion vector corresponding to the target emotion of the user behavior data based on semantic information and image information in the user behavior data. The first calculation unit is used to determine the emotion matching evaluation value of the virtual facial data based on the difference between the first emotion vector and the second emotion vector.
[0016] In some embodiments, the feature extraction unit is used to extract audio features from the audio data of the user behavior data; and to extract the motion sequence of key points of the mouth from the virtual face data; The first calculation unit is used to determine the degree of correlation between the audio features and the change process of the mouth key point motion sequence; and based on the degree of correlation of the change process, to determine the lip-sync evaluation value of the virtual facial data.
[0017] In some embodiments, the feature extraction unit is used to extract the temporal features of the virtual face data and the generation parameters of each frame of the virtual face image; The first calculation unit is configured to determine adjacent virtual face images from the virtual face data based on the temporal characteristics of the virtual face data; for any two adjacent virtual face images, determine the difference in the generation parameters of the adjacent virtual face images; and determine the naturalness evaluation value of the virtual face data based on the difference in the generation parameters.
[0018] In some embodiments, the apparatus further includes: The data preprocessing unit is used to align the virtual facial data with the user behavior data based on the timestamp of the user behavior data corresponding to the generation of the virtual facial image.
[0019] In some embodiments, the apparatus further includes: The parameter adjustment unit is used to adjust the parameters of the virtual face generation system based on the evaluation value of the virtual face data.
[0020] A third aspect of this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the virtual face evaluation method as described above.
[0021] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the virtual face evaluation method described above.
[0022] A fifth aspect of this application provides a computer program product that, when run on an electronic device, causes the electronic device to perform the virtual face evaluation method as described above.
[0023] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows: In this embodiment, multimodal feature extraction is performed on virtual facial data and user behavior data. Based on the extracted multimodal features, the virtual facial data generated by the virtual facial generation system is evaluated from different aspects to obtain multiple sub-item evaluation values. The multiple sub-item evaluation values are weighted to obtain the evaluation value of the virtual facial data. In this way, by integrating multimodal data of virtual facial data and user behavior data to evaluate virtual facial data, the realism and naturalness of virtual facial data are more comprehensively reflected, and the standardization and automation requirements of virtual facial data evaluation are improved. Attached Figure Description
[0024] Figure 1 A schematic diagram of a virtual facial evaluation system provided in an exemplary embodiment is shown; Figure 2 A flowchart illustrating a virtual face evaluation method provided in an exemplary embodiment is shown. Figure 3 A flowchart illustrating a virtual face evaluation method provided in an exemplary embodiment is shown. Figure 4 A schematic diagram of a virtual facial evaluation device provided in an exemplary embodiment is shown; Figure 5 A schematic diagram of the structure of an electronic device provided in an exemplary embodiment is shown. Detailed Implementation
[0025] To make the technical problems, technical solutions, and beneficial effects to be solved by this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this application.
[0026] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0027] With the development of image processing technology, its application in virtual face generation is becoming increasingly widespread. For example, in fields such as film and television, games, live streaming, intelligent customer service, education, and interactive experiences, virtual faces are often needed to complete corresponding functions. Among these applications, the realism and naturalness of the virtual face's facial expressions significantly impact the user experience.
[0028] Virtual faces are typically generated by a virtual face generation system. These systems often use voice, text, or motion signals to drive the model and generate corresponding expression frames. Therefore, the performance of the virtual face generation system not only affects the visual quality of the virtual face but also directly determines the user experience.
[0029] Virtual faces are typically generated by a virtual face generation system. These systems often use voice, text, or motion signals to drive the model and generate corresponding expression frames. Therefore, the performance of the virtual face generation system not only affects the visual quality of the virtual face but also directly determines the user experience.
[0030] In some embodiments, virtual facial images are generated using a parametric model-based method. For example, facial expressions are controlled by preset BlendShape parameters, facial bones, or deformation model mapping. This type of method is computationally efficient, but it struggles to capture complex nonlinear facial expression changes, resulting in unsmooth generated images that are prone to appearing "stiff" or "plastic."
[0031] In other embodiments, virtual facial expressions are generated using a learning-based approach. For example, based on deep learning models (e.g., GAN, Diffusion, Audio2Face, etc.), semantic mappings to facial expressions are learned from speech, images, or videos to generate natural virtual facial expressions. However, the training process of deep learning models often lacks unified quantitative standards, leading to ambiguity in the direction of model optimization.
[0032] Currently, commonly used training evaluation metrics include Mean Squared Error (MSE), Structural Similarity Index Measure (SSIM), and Peak Signal-to-Noise Ratio (PSNR). However, these training evaluation metrics are more suitable for evaluating static image instructions and are difficult to comprehensively describe the realism, flow, and semantic consistency of virtual facial images.
[0033] In summary, to improve the accuracy of evaluation of face generation systems, this application provides a virtual face evaluation method. This method extracts multimodal features from virtual face data and user behavior data, evaluates the virtual face data generated by the virtual face generation system from different aspects based on the extracted multimodal features, and obtains multiple sub-item evaluation values. These sub-item evaluation values are then weighted to obtain the final evaluation value of the virtual face data. By integrating multimodal data from virtual face data and user behavior data to evaluate the virtual face data, this method more comprehensively reflects the realism and naturalness of the virtual face data, and improves the standardization and automation requirements for virtual face data evaluation.
[0034] The present application will now be described with reference to specific embodiments. See below. Figure 1 This illustrates a schematic diagram of a virtual face evaluation system involved in an exemplary embodiment of a virtual face evaluation method. For example... Figure 1 As shown, the virtual facial evaluation system includes an electronic device and a database. The electronic device is communicatively connected to the database. It should be noted that the communication connection between the electronic device and the database can be wireless, or it can be established through a hardware transmission interface.
[0035] The database stores user behavior data and virtual facial data generated based on that user behavior data. The electronic device is used to evaluate this virtual facial data.
[0036] In some embodiments, the electronic device further includes a virtual face generation system for generating corresponding virtual face data in real time based on user behavior data. Accordingly, the electronic device also includes a display screen and a behavior data acquisition device. The display screen displays the generated virtual face data. The behavior data acquisition device collects user behavior data input by the user. For example, the behavior data acquisition device includes an image acquisition unit for collecting the user's real image data; the behavior data acquisition device also includes an audio acquisition unit for collecting audio data input by the user; the behavior data acquisition device may also include a text input unit for collecting text data input by the user, etc.
[0037] It should be noted that the virtual face generation system can be a system built based on a deep learning model. Specifically, the virtual face generation system can be an incompletely trained system. In this case, the virtual face data generated by the system can be evaluated using the virtual face evaluation method provided in this application, and the evaluation result can be used as the system's loss value to adjust the system parameters. Alternatively, the virtual face generation system can be a fully trained system. In this case, the virtual face data generated by the system can be evaluated using the virtual face evaluation method provided in this application, and the evaluation result can be used as self-learning feedback to adjust the system parameters.
[0038] See Figure 2 The diagram illustrates a flowchart of a virtual face evaluation method provided by an exemplary embodiment. This method is applied to electronic devices, not as a limitation.
[0039] S201, the electronic device acquires virtual face data generated by the virtual face generation system and user behavior data used to generate the virtual face data.
[0040] The virtual facial data consists of virtual facial images generated by a virtual facial generation system based on the user behavior data. Accordingly, the virtual facial data includes at least one virtual facial image frame. This virtual facial image frame can be a three-dimensional image or a two-dimensional image; this is not specifically limited in this embodiment. In some embodiments, the electronic device obtains the user behavior data and the virtual facial data from a database. In other embodiments, the electronic device can collect user behavior data in real time using a user behavior data acquisition device and generate virtual facial data corresponding to the real-time collected user behavior data using a virtual facial generation system.
[0041] For example, during the training of a virtual face generation system, virtual face data and corresponding user behavior data for model training can be obtained from a database. When optimizing a trained virtual face generation system, electronic devices can collect user behavior data in real time through user behavior data acquisition devices.
[0042] S202, the electronic device extracts the multimodal features of the virtual facial data and the user behavior data corresponding to the virtual facial data.
[0043] The multimodal features refer to various types of features extracted from virtual facial data and user behavior data. These multimodal features can include geometric features, visual features, video features, and temporal features. For example, they include features for emotion, audio features, mouth keypoint motion sequences, and generation parameters for virtual facial data.
[0044] It should be noted that the multimodal features extracted from the virtual face data and the user behavior data can be of the same type or different types. Furthermore, when extracting the same type of features from the virtual face data and the user behavior data, the same feature extraction method or different feature extraction methods can be used. In this embodiment, no specific limitation is made in this regard.
[0045] For example, electronic devices can extract corresponding emotional features from virtual facial data and user behavior data, respectively. Specifically, the electronic device can use a facial emotion recognition system to identify a first emotion vector from the virtual facial data, which is a vector representing the emotional features of the virtual facial data. For the emotional features corresponding to the user behavior data, representing the user's target emotion, the electronic device can also use the facial emotion recognition system to identify them. Furthermore, the electronic device can use semantic and image information in the user behavior data to identify a second emotion vector corresponding to the user behavior data, which is a feature vector representing the user's target emotion.
[0046] In addition, electronic devices can also extract audio features from user behavior data through audio feature processors; identify the movement sequence of mouth key points in the virtual face data through key point recognition systems; extract the generation parameters of the virtual face data, as well as the temporal features of the virtual face data, etc.
[0047] Since both virtual facial data and user behavior data are composed of data from a specific period of time, the user behavior data and virtual facial data may differ at different points in time within that period. Therefore, when evaluating virtual facial data, it is necessary to align the timestamps of the virtual facial data and the user behavior data. Accordingly, prior to this step, the electronic device aligns the virtual facial data with the user behavior data based on the timestamp of the user behavior data corresponding to the generation of the virtual facial image.
[0048] The electronic device extracts the timestamps from the virtual facial data and user behavior data, and aligns virtual facial data and user behavior data with the same timestamp based on these timestamps. The timestamp corresponds to the time the data was generated.
[0049] It should be noted that the categories of the aforementioned multimodal features can be added or reduced as needed. In this embodiment, the categories and number of categories of the multimodal features are not specifically limited. For example, other dimensions of multimodal features, such as illumination robustness features and identity consistency features, can be added as needed. This allows for the formation of customized evaluation systems for different application scenarios of virtual face generation systems, thereby adapting virtual face generation data to virtual face generation systems corresponding to different application scenarios and improving the evaluation accuracy of virtual face data in different application scenarios.
[0050] S203, the electronic device calculates multiple sub-item evaluation values of the virtual facial data based on the multimodal features.
[0051] The electronic device evaluates the virtual facial data based on the multimodal features corresponding to the virtual data. Different evaluation sub-items can use the same or different categories of multimodal features. The number and categories of evaluation sub-items can be set as needed, and are not specifically limited in this embodiment. For example, the evaluation sub-items may include multiple evaluation sub-items such as emotion alignment, lip-sync, smoothness and realism, and avoidance of unnaturalness / distortion.
[0052] S204, the electronic device weights the evaluation values of the multiple sub-items to obtain the evaluation value of the virtual facial data.
[0053] The evaluation value of the virtual facial data is usually obtained by combining the evaluation values of multiple sub-items using a weighted summation method. For example, when the evaluation sub-items may include emotion alignment, lip-sync, smoothness and realism, and avoidance of penalty, the evaluation value of the virtual facial data can be determined by the following formula.
[0054] Formula 1:
[0055] in, The evaluation value for virtual facial data; The weights corresponding to the emotional fit evaluation sub-items are: The sub-item evaluation value corresponding to the emotional fit evaluation sub-item; Assign weights to the lip-sync evaluation sub-items. The sub-item evaluation value corresponding to the lip-sync evaluation sub-item. The weights corresponding to the naturalness / coherence evaluation sub-items. The evaluation value corresponding to the naturalness / coherence evaluation sub-item; The weights corresponding to the sub-items for evaluating authenticity. The evaluation value corresponding to the authenticity evaluation sub-item; The evaluation value corresponding to the penalty item. The weights corresponding to the penalty items.
[0056] In this embodiment, the sum of the weights of the multiple evaluation sub-items is 1. Furthermore, the weight of each evaluation sub-item can be automatically learned based on the application scenario, or set by the developers based on their experience. In this embodiment, the value of the weight of each evaluation sub-item is not specifically limited.
[0057] After determining the evaluation value of the virtual facial data, the virtual facial generation system that generates the virtual facial data can be optimized based on this evaluation value. Accordingly, the electronic device adjusts the parameters of the virtual facial generation system based on the evaluation value of the virtual facial data.
[0058] In some embodiments, the virtual face data generated by the virtual face generation system can be evaluated using the virtual face evaluation method provided in this application, and the evaluation result can be used as the system loss value to adjust the system parameters of the face generation system. In other embodiments, the virtual face generation system can be a pre-trained system, and accordingly, the virtual face data generated by the virtual face generation system can be evaluated using the virtual face evaluation method provided in this application, and the evaluation result can be used as a self-learning feedback result to adjust the system parameters of the face generation system.
[0059] In this embodiment, multimodal feature extraction is performed on virtual facial data and user behavior data. Based on the extracted multimodal features, the virtual facial data generated by the virtual facial generation system is evaluated from different aspects to obtain multiple sub-item evaluation values. The multiple sub-item evaluation values are weighted to obtain the evaluation value of the virtual facial data. In this way, by integrating multimodal data of virtual facial data and user behavior data to evaluate virtual facial data, the realism and naturalness of virtual facial data are more comprehensively reflected, and the standardization and automation requirements of virtual facial data evaluation are improved.
[0060] For different application scenarios, electronic devices can match different evaluation sub-items to this virtual face evaluation method. See also Figure 3 The diagram illustrates a flowchart of a virtual face evaluation method provided by an exemplary embodiment. This method is applied to electronic devices, not as a limitation.
[0061] S301, the electronic device acquires virtual face data generated by the virtual face generation system and user behavior data used to generate the virtual face data.
[0062] This step is based on the same principle as step S201, and will not be repeated here.
[0063] S302, the electronic device identifies a first emotion vector representing the emotional characteristics of the virtual facial data through a facial emotion recognition system.
[0064] This facial emotion recognition system is used to identify the emotions expressed by facial expressions. The first emotion vector represents the probability distribution of the emotion represented by the virtual facial data in an emotion space composed of basic emotions, including happiness, sadness, anger, surprise, disgust, and fear.
[0065] S303, the electronic device identifies the second emotion vector of the target emotion corresponding to the user behavior data based on the semantic and image information in the user behavior data.
[0066] For the emotional features corresponding to user behavior data that represent the user's target emotion, electronic devices can be identified by the facial emotion recognition system. Electronic devices can also identify the second emotion vector corresponding to the user behavior data by using semantic information and image information in the user behavior data. The second emotion vector is a feature vector used to represent the user's target emotion.
[0067] S304, the electronic device determines the emotion matching evaluation value of the virtual facial data based on the emotion vector of the virtual facial data and the target emotion corresponding to the user behavior data.
[0068] The electronic device determines the emotion rating of the virtual facial data based on the difference between the first emotion vector and the second emotion vector. Accordingly, the electronic device determines the difference between the first emotion vector and the second emotion vector, and determines the emotion fit rating of the virtual facial data based on the difference between the first emotion vector and the second emotion vector.
[0069] In some embodiments, the electronic device determines the square of the Euclidean distance between the two emotion vectors based on the difference between the first emotion vector and the second emotion vector, which is used to measure the difference between the first emotion vector and the second emotion vector.
[0070] In this embodiment, the evaluation value of the virtual facial data is used to evaluate the virtual facial generation system, and this evaluation value is used as the loss value for system optimization. A larger loss value indicates that the virtual facial data generated by the virtual facial generation system is less accurate. Therefore, to facilitate direct training of the virtual facial generation system using this evaluation value, the evaluation value of the virtual facial data is inverted, thus meeting the needs of subsequent system optimization. See Formula 2, which illustrates a method for calculating the emotion matching evaluation value provided by an exemplary embodiment.
[0071] Formula 2:
[0072] in, The sub-item evaluation value corresponding to the emotional fit evaluation sub-item; Represents the first emotion vector; This is the second emotion vector; Represents an exponential function; This is a sensitivity adjustment parameter used to control the intensity of the influence of the emotional difference on the final evaluation value. Its value is greater than 0 and can be set as needed. In this embodiment, [the parameter is used to adjust the sensitivity]. The value of is not specifically limited.
[0073] This calculation method quantifies the difference between virtual facial data and the user's actual emotional expression. The smaller the difference between the emotion expressed by the virtual facial data and the user's actual emotion, the closer the score is to 1.
[0074] S305, The electronic device extracts the audio features of the audio data of the user behavior data.
[0075] The audio feature represents features such as the user's pronunciation content and temporal information in the user behavior data. Specifically, the audio feature can be an audio feature extracted using Mel-Frequency Cepstral Coefficients (MFCC); or, it can be an audio feature extracted using wav2vec2. In this embodiment, the method for extracting the audio feature is not specifically limited.
[0076] S306, The electronic device extracts the motion sequence of the key points of the mouth in the virtual facial data.
[0077] The mouth key point motion sequence includes the coordinates or displacement of key points when the lips of the virtual face open up and down, open to the left and right, and stretch the corners of the mouth; it is used to quantify the movements of the lips when the virtual face simulation occurs, including: the opening range, the rounding of the lips, and the opening and closing of the corners of the mouth.
[0078] In this step, the electronic device can use any keypoint recognition method to determine the mouth keypoint motion sequence. In this embodiment, the method for obtaining the mouth keypoint motion sequence is not specifically limited. For example, the electronic device can obtain the generation parameters of the virtual face image in the virtual face data, and obtain the mouth keypoint motion sequence from the generation parameters.
[0079] S307, the electronic device determines the lip-sync evaluation value of the virtual facial data based on the motion sequence of the key points of the mouth in the virtual facial data and the audio features corresponding to the user behavior data.
[0080] In this step, the electronic device determines the lip-sync evaluation value based on whether the lip shape represented by the motion sequence of the mouth key points of the virtual face at the same timestamp is correlated with the audio features of the user's actual voice. Accordingly, the electronic device determines the degree of correlation between the audio features and the change process of the mouth key point motion sequence; based on this degree of correlation, it determines the lip-sync evaluation value of the virtual facial data.
[0081] For example, an electronic device can represent the lip-sync evaluation value by calculating the dynamic time-warped distance or cross-correlation coefficient between audio features and the motion sequence of the lip key points. See Formula 3, which illustrates a method for calculating the lip-sync evaluation value provided by an exemplary embodiment.
[0082] Formula 3:
[0083] in, This indicates the lip-sync evaluation value; Indicates audio characteristics; This represents the sequence of motion at key points of the mouth. The covariance between the feature sequence corresponding to the audio feature and the mouth key point motion sequence is used to measure the co-change trend between the audio feature and the mouth key point motion sequence. Representing audio features The standard deviation of represents the intensity of change in audio features; Represents the motion sequence of key points in the mouth The standard deviation represents the intensity of change in the motion sequence of key points in the mouth; correspondingly, Formula 3 is used to calculate audio features. Movement sequence of key points of the mouth The correlation coefficient between the audio feature and the mouth key point motion sequence is calculated, with a value of [-1, 1]. The larger the correlation coefficient, the higher the correlation between the audio feature and the mouth key point motion sequence.
[0084] The evaluation value of the virtual facial data is used to evaluate the virtual facial generation system and is also used as the loss value for system optimization. A larger loss value indicates that the virtual facial data generated by the system is less accurate. Therefore, to facilitate direct training of the virtual facial generation system using this evaluation value, the evaluation value of the virtual facial data is inverted, thus meeting the needs of subsequent system optimization. Therefore, in this step, the result of subtracting the correlation coefficient from 1 is taken as the value of the lip-sync evaluation value.
[0085] S308, the electronic device extracts the temporal features of the virtual face data and the generation parameters of each frame of the virtual face image.
[0086] The electronic device parses each frame of the virtual face data to obtain the generation parameters for each frame. These generation parameters refer to the relevant parameters used when rendering the virtual face image, including the positional information of key points in the virtual face image, feature vectors, and other parameters that can constitute the pose, expression, and shape of the virtual face image.
[0087] S309, the electronic device determines the naturalness evaluation value of the virtual face data based on the generation parameters of any two adjacent frames of virtual face images in the virtual face data.
[0088] The electronic device determines the similarity of generation parameters between multiple consecutive pairs of virtual facial images in the virtual facial data based on the temporal characteristics of the virtual facial data. By measuring the similarity of generation parameters between adjacent virtual facial images, it determines whether there are significant jumps between adjacent virtual facial images, thus preventing problems such as jumps in the generated virtual facial data. Accordingly, the electronic device determines adjacent virtual facial images from the virtual facial data based on the temporal characteristics of the virtual facial data; for any two adjacent virtual facial images, it determines the difference in generation parameters between the adjacent virtual facial images; and based on the difference in generation parameters, it determines the naturalness evaluation value of the virtual facial data. Accordingly, see Formula 4, which illustrates a method for calculating the naturalness evaluation value provided by an exemplary embodiment.
[0089] Formula 4:
[0090] in, This represents the naturalness rating. This represents the generation parameters of the virtual face image in frame t; This represents the square of the Euclidean distance used to calculate the difference. This is a weighting coefficient used to control the influence of the naturalness evaluation value on the final evaluation value. Its value is greater than 0 and can be set as needed. In this embodiment, it is used for... The value of is not specifically limited. Furthermore, the evaluation value of this virtual facial data is used to evaluate the virtual facial generation system, and this evaluation value is used as the loss value for system optimization. The larger the loss value, the less accurate the virtual facial data generated by the virtual facial generation system is. Therefore, in order to facilitate the direct use of this evaluation value to train the virtual facial generation system, the evaluation value of the virtual facial data is inverted, thus meeting the needs of subsequent system optimization.
[0091] S3010, the electronic device determines the acceptability evaluation value of the virtual facial data through a realism discrimination model.
[0092] This realism discrimination model is used to score the difference between a real facial image and a virtual facial image in the virtual facial data. In some embodiments, a human scoring model can be used to score the acceptability. Accordingly, the acceptability rating can be... ,in, This represents a model for scoring the degree of authenticity.
[0093] S311 When an image abnormality is detected in the virtual facial data, the electronic device generates a penalty evaluation value.
[0094] To prevent facial distortion, loss of key points, or non-physical jumps in virtual facial data, a penalty item score has been added. When anomalies such as facial distortion, loss of key points, or non-physical jumps are detected, the stability of the position system is evaluated through this penalty item score.
[0095] S312, the electronic device weights the evaluation values of the multiple sub-items to obtain the evaluation value of the virtual facial data.
[0096] This step is based on the same principle as step S204, and will not be repeated here.
[0097] In this embodiment, multimodal feature extraction is performed on virtual facial data and user behavior data. Based on the extracted multimodal features, the virtual facial data generated by the virtual facial generation system is evaluated from different aspects to obtain multiple sub-item evaluation values. The multiple sub-item evaluation values are weighted to obtain the evaluation value of the virtual facial data. In this way, by integrating multimodal data of virtual facial data and user behavior data to evaluate virtual facial data, the realism and naturalness of virtual facial data are more comprehensively reflected, and the standardization and automation requirements of virtual facial data evaluation are improved.
[0098] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0099] See Figure 4 The diagram illustrates the structure of a virtual facial evaluation device provided in this application, including various units used to perform the steps in the above embodiments. See also... Figure 4 The virtual facial evaluation device includes: The acquisition unit 401 is used to acquire virtual face data generated by the virtual face generation system and user behavior data used to generate the virtual face data. Feature extraction unit 402 is used to extract multimodal features of the virtual face data and the user behavior data corresponding to the virtual face data, respectively; The first calculation unit 403 is used to calculate multiple sub-item evaluation values of the virtual face data based on the multimodal features; The second calculation unit 404 is used to weight the evaluation values of the multiple sub-items to obtain the evaluation value of the virtual facial data.
[0100] In some embodiments, the device further includes: The data preprocessing unit is used to align the virtual facial data with the user behavior data based on the timestamp of the user behavior data corresponding to the generation of the virtual facial image.
[0101] In some embodiments, the device further includes: The parameter adjustment unit is used to adjust the parameters of the virtual face generation system based on the evaluation value of the virtual face data.
[0102] The functions implemented by the above-mentioned units are the same as those in the above-described virtual face evaluation method embodiment, and will not be repeated here.
[0103] In this embodiment, multimodal feature extraction is performed on virtual facial data and user behavior data. Based on the extracted multimodal features, the virtual facial data generated by the virtual facial generation system is evaluated from different aspects to obtain multiple sub-item evaluation values. The multiple sub-item evaluation values are weighted to obtain the evaluation value of the virtual facial data. In this way, by integrating multimodal data of virtual facial data and user behavior data to evaluate virtual facial data, the realism and naturalness of virtual facial data are more comprehensively reflected, and the standardization and automation requirements of virtual facial data evaluation are improved.
[0104] Figure 5 This is a schematic diagram of an electronic device provided in an exemplary embodiment of this application. (As shown...) Figure 5 As shown, the electronic device 5 in this embodiment includes: a processor 50, a memory 51, and a computer program 52 stored in the memory 51 and executable on the processor 50, such as a virtual face evaluation program. When the processor 50 executes the computer program 52, it implements the steps described in the various virtual face evaluation method embodiments above, for example... Figure 2 Steps S201 to S204 are shown. Alternatively, when the processor 50 executes the computer program 52, it implements the functions of each unit in the above-described device embodiments, for example... Figure 4 The functions of units 401 to 404 are shown.
[0105] For example, the computer program 52 can be divided into one or more units, which are stored in the memory 51 and executed by the processor 50 to complete this application. The one or more units can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program 52 in the electronic device 5. For example, the computer program 52 can be divided into an acquisition unit 401, a feature extraction unit 402, a first calculation unit 403, and a second calculation unit 404, with the specific functions of each module as follows: The acquisition unit 401 is used to acquire virtual face data generated by the virtual face generation system and user behavior data used to generate the virtual face data. Feature extraction unit 402 is used to extract multimodal features of the virtual face data and the user behavior data corresponding to the virtual face data, respectively; The first calculation unit 403 is used to calculate multiple sub-item evaluation values of the virtual face data based on the multimodal features; The second calculation unit 404 is used to weight the evaluation values of the multiple sub-items to obtain the evaluation value of the virtual facial data.
[0106] The electronic device 5 can be any electronic device with control functions. The electronic device 5 may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art will understand that... Figure 5 This is merely an example of electronic device 5 and does not constitute a limitation on electronic device 5. It may include more or fewer components than shown, or combine certain components, or different components. For example, electronic device 5 may also include input / output devices, network access devices, buses, etc.
[0107] The processor 50 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0108] The memory 51 can be an internal storage unit of the electronic device 5, such as a hard disk or memory. The memory 51 can also be an external storage device of the electronic device 5, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. Furthermore, the memory 51 can include both internal and external storage units of the electronic device 5. The memory 51 is used to store the computer program and other programs and data required by the terminal device. The memory 51 can also be used to temporarily store data that has been output or will be output.
[0109] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0110] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0111] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0112] In the embodiments provided in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0113] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0114] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0115] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0116] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the above method embodiments.
[0117] This application also provides a computer program product that, when run on a mobile terminal, enables the mobile terminal to implement the steps described in the various method embodiments above.
[0118] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A virtual facial evaluation method, characterized in that, The method includes: Acquire virtual face data generated by the virtual face generation system and user behavior data used to generate the virtual face data; Multimodal features of the virtual facial data and the user behavior data corresponding to the virtual facial data are extracted respectively; Based on the multimodal features, calculate multiple sub-item evaluation values for the virtual facial data; The evaluation values of the multiple sub-items are weighted to obtain the evaluation value of the virtual facial data.
2. The method as described in claim 1, characterized in that, The calculation of multiple sub-item evaluation values of the virtual facial data based on the multimodal features includes: Based on the emotion vector of the virtual facial data and the target emotion corresponding to the user behavior data, determine the emotion matching evaluation value of the virtual facial data; and / or Based on the lip movement sequence of the virtual facial data and the audio features corresponding to the user behavior data, determine the lip-sync evaluation value of the virtual facial data; and / or Based on the generation parameters of any two adjacent virtual facial images in the virtual facial data, determine the naturalness evaluation value of the virtual facial data; and / or The acceptability rating of the virtual facial data is determined using a realism discrimination model; and / or When an image anomaly is detected in the virtual facial data, a penalty evaluation value is generated.
3. The method as described in claim 2, characterized in that, The step of extracting multimodal features from the virtual facial data and the user behavior data corresponding to the virtual facial data includes: A first emotion vector representing the emotional features of the virtual facial data is identified using a facial emotion recognition system. Based on the semantic and image information in the user behavior data, identify the second emotion vector corresponding to the target emotion in the user behavior data; The step of determining the emotion matching evaluation value of the virtual facial data based on the emotion vector of the virtual facial data and the target emotion corresponding to the user behavior data includes: The emotional fit evaluation value of the virtual facial data is determined based on the difference between the first emotion vector and the second emotion vector.
4. The method as described in claim 2, characterized in that, The step of extracting multimodal features from the virtual facial data and the user behavior data corresponding to the virtual facial data includes: Extract the audio features from the audio data of the user behavior data; Extract the motion sequence of key points of the mouth from the virtual facial data; The step of determining the lip-sync evaluation value of the virtual facial data based on the lip key point motion sequence of the virtual facial data and the audio features corresponding to the user behavior data includes: Determine the degree of correlation between the audio features and the changes in the mouth key point motion sequence; Based on the correlation of the change process, the lip-sync evaluation value of the virtual facial data is determined.
5. The method as described in claim 2, characterized in that, The step of extracting multimodal features from the virtual facial data and the user behavior data corresponding to the virtual facial data includes: Extract the temporal features of the virtual face data and the generation parameters of each frame of the virtual face image; The step of determining the naturalness evaluation value of the virtual facial data based on the generation parameters of any two adjacent frames of virtual facial images in the virtual facial data includes: Based on the temporal characteristics of the virtual facial data, adjacent virtual facial images are determined from the virtual facial data; For any two adjacent virtual face images, determine the difference in the generation parameters of the adjacent virtual face images; The naturalness evaluation value of the virtual facial data is determined based on the difference in the generated parameters.
6. The method according to any one of claims 1-5, characterized in that, Before extracting the multimodal features of the virtual facial data and the user behavior data corresponding to the virtual facial data, the method further includes: The virtual facial data is aligned with the user behavior data based on the timestamp of the user behavior data corresponding to the generation of the virtual facial image.
7. The method according to any one of claims 1-5, characterized in that, The method further includes: Based on the evaluation value of the virtual facial data, the parameters of the virtual facial generation system are adjusted.
8. A virtual facial evaluation device, characterized in that, The device includes: The acquisition unit is used to acquire virtual face data generated by the virtual face generation system and user behavior data used to generate the virtual face data; The feature extraction unit is used to extract multimodal features of the virtual face data and the user behavior data corresponding to the virtual face data, respectively. The first calculation unit is used to calculate multiple sub-item evaluation values of the virtual facial data based on the multimodal features; The second calculation unit is used to weight the evaluation values of the multiple sub-items to obtain the evaluation value of the virtual facial data.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the virtual face evaluation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the virtual face evaluation method as described in any one of claims 1 to 7.