Virtual avatar facial expression generation system and virtual avatar facial expression generation method

Through the multi-data source sensing and emotional decision-making merging technology, virtual avatar facial expressions that are consistent with user emotions are generated, solving the problem of inaccurate expression simulation in the virtual environment and improving the social experience of the virtual environment.

CN113313795BActive Publication Date: 2025-08-19XRSPACE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010709804.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-27
Filing Date
2020-07-22
Publication Date
2025-08-19
Estimated Expiration
2040-07-22

AI Technical Summary

Technical Problem

The prior art is difficult to accurately capture the user's facial expressions and generate virtual avatar facial expressions consistent with the real environment in the virtual environment, especially because the head-mounted display covers the facial features, resulting in inaccurate expression simulation.

Method used

Through user data sensing from multiple data sources, virtual avatar facial expressions are generated using tracking devices, memory and processors, including conflict analysis and merging of emotional decisions, ensuring that the virtual avatar expression is consistent with user emotions.

Benefits of technology

It realizes the generation of virtual avatar facial expressions consistent with the user's emotions in the virtual environment, and improves the quality of social experience in the virtual environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113313795B_ABST
    Figure CN113313795B_ABST
Patent Text Reader

Abstract

The present invention provides a virtual avatar facial expression generation system and a virtual avatar facial expression generation method. In the method, multiple user data are obtained, and the multiple user data are related to user sensing results from multiple data sources. Multiple first emotional decisions are determined based on each piece of user data. A determination is made as to whether an emotional conflict occurs between the first emotional decisions. The emotional conflict is related to a mismatch between the corresponding emotional groups of the first emotional decisions. Based on the determination of the emotional conflict, a second emotional decision is determined from one or more emotional groups. The first emotional decision or the second emotional decision is related to an emotional group. A virtual avatar facial expression is generated based on the second emotional decision. Thus, an appropriate virtual avatar facial expression can be presented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to virtual avatar simulation, and more particularly, to a virtual avatar facial expression generation system and a virtual avatar facial expression generation method. Background Art

[0002] Technologies for simulating sensations, perceptions, and / or environments, such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and extended reality (XR), are currently popular. These technologies are being applied in a variety of fields, such as gaming, military training, healthcare, and remote work.

[0003] In order to enable the user to perceive the virtual environment and the real environment, the movements of the user's body parts or the user's facial expressions in the real environment will be tracked so that the facial expressions of the virtual avatar presented on the VR display, AR display, MR display or XR display can change in response to the user's movements or facial expressions, and the social influence in the virtual environment can be improved.

[0004] Regarding facial expression simulation, in conventional methods, a camera is positioned to capture the user's face using a head-mounted display (HMD), and simulated facial expressions are generated based on the facial features in the captured image. However, since part of the face is covered by the HMD, it is difficult to recognize facial features and facial expressions, and the facial expressions of the virtual avatar may not be the same as the user's facial expressions in the real environment. Summary of the Invention

[0005] It is indeed difficult to predict the correct facial expression by relying solely on a camera. Therefore, the present disclosure relates to a virtual avatar facial expression generation system and a virtual avatar facial expression generation method to simulate the facial expression of a virtual avatar with emotions in a virtual environment.

[0006] In one exemplary embodiment, a method for generating facial expressions for a virtual avatar includes, but is not limited to, the following steps: A plurality of user data are obtained. Each piece of user data is associated with user sensing results from multiple data sources. A plurality of first emotion decisions are determined based on each piece of user data. A determination is made as to whether an emotion conflict occurs between the first emotion decisions. The emotion conflict is associated with a mismatch between the corresponding emotion groups of the first emotion decisions. A second emotion decision is determined from one or more emotion groups based on the determination of the emotion conflict. The first emotion decision or the second emotion decision is associated with an emotion group. A facial expression for the virtual avatar is generated based on the second emotion decision.

[0007] In one exemplary embodiment, a facial expression generation system includes, but is not limited to, one or more tracking devices, a memory, and a processor. The tracking device obtains a plurality of user data. Each piece of user data is associated with a user's sensing result from one of a plurality of data sources. The memory stores program code. The processor is coupled to the memory and loaded with the program code to perform the following steps. The processor determines a plurality of first emotion decisions based on each piece of user data, determines whether an emotion conflict occurs between the first emotion decisions, determines a second emotion decision from one or more emotion groups based on the determined emotion conflict, and generates a facial expression for an avatar based on the second emotion decision. The emotion conflict is associated with a mismatch between the corresponding emotion groups of the first emotion decisions. The first emotion decision or the second emotion decision is associated with one emotion group.

[0008] Based on the above, according to the avatar facial expression generation system and method according to embodiments of the present invention, when an emotional conflict occurs between these first emotional decisions, a second emotional decision is further determined from two or more emotional groups, and (only) a single emotional group is selected for generating the facial emotion of the avatar. This allows for the presentation of appropriate facial expressions for the avatar.

[0009] However, it should be understood that this summary may not contain all aspects and embodiments of the present disclosure, is not intended to be restrictive or limiting in any way, and the invention as disclosed herein is and will be understood by those skilled in the art to encompass obvious improvements and modifications thereto. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings are included to provide a further understanding of the disclosure and are incorporated in and constitute a part of this specification. The drawings illustrate embodiments of the disclosure and together with the description serve to explain the principles of the disclosure.

[0011] Figure 1 is a block diagram illustrating a virtual avatar facial expression generation system according to one of the exemplary embodiments of the present disclosure;

[0012] Figure 2 is a flow chart illustrating a method for generating facial expressions of a virtual avatar according to one of the exemplary embodiments of the present disclosure;

[0013] Figure 3 is a flow chart illustrating user data generation according to one of the exemplary embodiments of the present disclosure;

[0014] Figure 4 is a schematic diagram illustrating types of emotion groups according to one of the exemplary embodiments of the present disclosure;

[0015] Figure 5is a schematic diagram illustrating a relationship between user data and a first emotion decision according to one of the exemplary embodiments of the present disclosure;

[0016] Figure 6 is a schematic diagram illustrating a first stage according to one of the exemplary embodiments of the present disclosure;

[0017] Figure 7 is a flow chart illustrating second emotion decision generation according to one of the exemplary embodiments of the present disclosure;

[0018] Figure 8 is a flowchart illustrating user data transformation according to one of the exemplary embodiments of the present disclosure.

[0019] Explanation of Figure Numbers

[0020] 100: Virtual Avatar Facial Expression Generation System;

[0021] 110: tracking device;

[0022] 115: sensor;

[0023] 120: display;

[0024] 130: memory;

[0025] 150: processor;

[0026] 231: First emotion classifier;

[0027] 232: distances related to facial features;

[0028] 233: semantic analysis;

[0029] 234: Image analysis;

[0030] 401, 402: type;

[0031] S210, S211, S213, S215, S230, S250, S255, S260, S261, S262, S263, S270, S290: steps. DETAILED DESCRIPTION

[0032] Reference will now be made in detail to the preferred embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings. Whenever possible, the same reference numerals are used in the drawings and the description to refer to the same or like parts.

[0033] Figure 1 is a block diagram illustrating a virtual avatar facial expression generation system according to one exemplary embodiment of the present disclosure. Figure 1The virtual avatar facial expression generation system 100 includes but is not limited to: one or more tracking devices 110, a display 120, a memory 130, and a processor 150. The virtual avatar facial expression generation system 100 is suitable for VR, AR, MR, XR, or other reality simulation related technologies.

[0034] The tracking device 110 is a handheld controller, a wearable device (e.g., a wearable controller, a smartwatch, an ankle sensor, a head-mounted display (HMD), or the like), or a sensing device (e.g., a camera, an inertial measurement unit (IMU), a heart rate monitor, an infrared (IR) transmitter / receiver, an ultrasonic sensor, a voice recorder, a strain gauge) for obtaining user data. The user data relates to sensing results of the user from one or more data sources. The tracking device 110 may include one or more sensors 115 to sense corresponding target parts of the user and generate a sequence of sensing data based on the sensing results (e.g., camera images, sensed intensity values) at multiple time points within a time period. These data sources differ in the user's target parts or sensing technology. For example, the target part can be a body part (e.g., a portion or the entire face, hand, head, ankle, leg, waist), organ (e.g., brain, heart, eye), or tissue (e.g., muscle, neural tissue) of the user. The sensing technology of the sensor 115 may be related to the following: images, sound waves, ultrasonic waves, current, electric potential, IR, force, motion sensing data related to displacement and rotation of human body parts, etc.

[0035] In one embodiment, the data source may be: facial muscle movement, speech, an image of part or all of the face, arm, leg, or head movement, cardiac activity, or brain activity. In some embodiments, the data source may be real-time data detection from sensor 115 or pre-configured data generated by processor 150.

[0036] The display 120 may be a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, or other display. In embodiments of the present disclosure, the display 120 is used to display images, such as a virtual environment. It should be noted that in some embodiments, the display 120 may be a display of an external device (e.g., a smartphone, tablet, or the like), and the external device may be placed on the main body of the HMD.

[0037] The memory 130 may be any type of fixed or removable random-access memory (RAM), read-only memory (ROM), flash memory, similar devices, or a combination thereof. The memory 130 records program code, device configuration, buffered data, or persistent data (e.g., user data, training data, emotion classifiers, emotion decisions, emotion configurations, weighted relationships, linear relationships, and emotion groups), and this data will be described later.

[0038] The processor 150 is coupled to the tracking device 110, the display 120, and the memory 130. The processor 150 is configured to load the program code stored in the memory 130 to execute the program of the exemplary embodiment of the present disclosure.

[0039] In some embodiments, processor 150 may be a central processing unit (CPU), a microprocessor, a microcontroller, a digital signal processing (DSP) chip, or a field-programmable gate array (FPGA). The functions of processor 150 may also be implemented by an independent electronic device or an integrated circuit (IC), and the operations of processor 150 may also be implemented by software.

[0040] It should be noted that the processor 150 may not be located in the same device as the tracking device 110 and the display 120. However, the devices equipped with the tracking device 110, the display 120, and the processor 150 may further include communication transceivers or physical transmission lines with compatible communication technologies, such as Bluetooth, Wi-Fi, and IR wireless communication, to transmit or receive data with each other. For example, the display 120 and the processor 150 may be located in the HMD, while the sensor 115 is located outside the HMD. For another example, the processor 150 may be located in the computing device, while the tracking device 110 and the display 120 are located outside the computing device.

[0041] To better understand the operational processes provided in one or more embodiments of the present disclosure, several examples are provided below to explain in detail the operational process of the virtual avatar facial expression generation system 100. The following examples utilize the devices and modules of the virtual avatar facial expression generation system 100 to illustrate the virtual avatar facial expression generation method provided herein. Each step of the method can be adjusted based on actual implementation circumstances and should not be limited to the content described herein.

[0042] Figure 2is a flow chart illustrating a method for generating a virtual avatar facial expression according to one exemplary embodiment of the present disclosure. Figure 2 The processor 150 may obtain a plurality of user data via the tracking device 110 (step S210). Specifically, the user data is obtained from a plurality of data sources. The processor 150 uses more data sources to improve the accuracy of emotion estimation. Figure 3 FIG. 1 is a flow chart illustrating user data generation according to one of the exemplary embodiments of the present disclosure. Figure 3 In one embodiment, the processor 150 may obtain sensing results such as raw data or pre-processed data from each sensor 115 (which may be processed by, for example, filtering, amplification, and analog-to-digital conversion) to generate user data in real time (i.e., the aforementioned real-time data detection) (step S211). For example, the user data may be raw data collected from one or more images of a portion or the entire face of the user, the user's facial feature movements (such as movements of the eyebrows, eyes, nose, and mouth), and / or the user's voice. In another embodiment, the processor 150 may perform feature extraction on each sensing result from each sensor 115 (step S213) to generate pre-configured data (step S215). Feature extraction is used to obtain informative and non-redundant derived values (features) from the sensing results, thereby facilitating subsequent analysis steps. For example, independent component analysis (ICA), isomap, and principal component analysis (PCA) are examples. Feature extraction can collect one or more specific actions / activities of the user corresponding to the target part, one or more specific keywords or key phrases within a predetermined time interval, or any features defined in machine learning techniques (such as neural networks (NN), K-means (K-means), and support vector machines (SVM)). For example, pre-configured data can be facial features such as blinking or nodding within a predetermined time interval, or randomly generated facial features. For another example, pre-configured data can be the user's voice content or voice intonation. In some embodiments, processor 150 can obtain a combination of real-time data and pre-configured data.

[0043] The processor 150 may determine a plurality of first emotion decisions based on each user data (step S230 ). Specifically, the processor 150 may predefine a plurality of emotion groups. Figure 4 is a diagram illustrating types of emotion groups according to one of the exemplary embodiments of the present disclosure. Figure 4In one embodiment, such as type 401, an emotion group includes only one emotion category, such as happiness, sadness, worry, disgust, anger, surprise, or excitement. In another embodiment, such as type 402, an emotion group includes multiple emotion categories, and each category can be a positive emotion or a negative emotion. Positive emotions can include, for example, happiness, excitement, and surprise. Negative emotions can include, for example, sadness, worry, and anger. In some embodiments, some emotion groups can include only one emotion category, while other emotion groups can include multiple emotion categories.

[0044] It should be noted that each first emotion decision is associated with (only) one emotion group. Figure 5 is a schematic diagram illustrating a relationship between user data and a first emotion decision according to one exemplary embodiment of the present disclosure. Figure 5 In the first stage, step S230, the processor 150 may determine a corresponding emotion group for each piece of user data from multiple data sources, and generate a first emotion decision based on the user data. In one embodiment, each first emotion decision is a specific emotion. For example, the first data source is an image of the user's eyes, and the second data source is speech. The first emotion decisions are happiness and sadness from the first and second data sources, respectively. In another embodiment, each first emotion decision is a weighted combination of emotions from two or more emotion categories. The emotion weight of the weighted combination can be in the form of a percentage or intensity level (i.e., a hierarchy of emotions). For example, the third data source is facial muscle movement. The first emotion decision is 60% happiness and 40% surprise from the third data source, where the emotion weight of happiness is 0.6 and the emotion weight of surprise is 0.4. The emotion weight can be the ratio of the corresponding emotion category to all corresponding emotion categories. In some embodiments, each emotion may further include multiple levels. For example, happiness includes three levels, where the first level represents the minimum intensity level of happiness and the third level represents the maximum intensity level. Thus, the emotion weight may be the intensity level of the corresponding emotion category. Therefore, the processor 150 may further determine the level of emotion of each first emotion decision.

[0045] Figure 6 is a schematic diagram illustrating a first stage of one of the exemplary embodiments according to the present disclosure. Figure 6In one embodiment, the processor 150 may determine each first emotion decision using a first emotion classifier 231 based on a machine learning technique (NN, K-means, SVM, etc.) or a tree-based classification method (boosted tree, bootstrap aggregating decision tree, etc.). In machine learning techniques, a classifier or model is used to identify to which of a set of categories an observation belongs. In this embodiment, the observation of the first emotion is user data, and the category of the first emotion corresponds to the second emotion decision. That is, the first emotion classifier is used to identify to which of the emotion groups each user data belongs. In other words, each user data can be input data to the first emotion classifier, and each first emotion decision is output data from the first emotion classifier. Taking artificial neural networks (ANNs) as an example, an ANN is composed of artificial neurons that receive input data or the output of preneurons. The network is composed of connections, each connection providing the output of a preneuron and its connections as a weighted sum, where each connection corresponds to an input weight. During the learning phase of the ANN, the input weights can be adjusted to improve the accuracy of the classifier's results. It should be noted that during the learning phase, the processor 150 may train a first emotion classifier for each data source based on a plurality of previously obtained first training emotions and training sensory data. These first training emotions encompass all emotion groups. This means that the output data of the first emotion classifier can be any of the emotion groups. Furthermore, the training sensory data is obtained from each data source and corresponds to a specific emotion (which will become the first training emotion).

[0046] In another embodiment, the processor 150 may determine the first emotion decision based on one or more distances 232 associated with facial features. For example, the presence of a wrinkle on the root of the nose, the shape of the eyes, or the presence of teeth, tongue, or nose. If the distance between the upper eyelid and the eyebrow is less than a threshold, the first emotion decision may be happiness or surprise. Furthermore, if the size of the mouth opening is greater than another threshold, the first emotion decision may be surprise.

[0047] In another embodiment, the processor 150 may identify words in the user data from speech and perform semantic analysis 233 on the identified words. During the semantic analysis, the processor 150 may determine whether the identified words in the user data match specific keywords or specific key phrases to determine whether the specific keywords or specific key phrases are detected in the user data. The processor 150 may predefine a plurality of keywords and / or key phrases, with each predefined keyword or predefined key phrase corresponding to a specific emotion, a specific level of emotion, a specific weighted combination of two or more emotion categories, or a specific weighted combination of two or more emotions at a specific level. For example, the user data relates to the sentence "I am very happy," and the keyword "very happy" corresponds to the fifth level of the happy emotion. If the identified words match a predefined keyword or a predefined key phrase (i.e., the detected predefined keyword or phrase), the processor 150 determines that the corresponding first emotion decision is the fifth level of the happy emotion.

[0048] In another embodiment, processor 150 may analyze user data from camera images or motion sensing data. Processor 150 may perform image analysis 234 to determine whether a predefined action or predefined facial expression is detected in the image. For example, if processor 150 detects a raised corner of the mouth in the camera image, processor 150 may consider a happy emotion to have been detected. For another example, if processor 150 detects a user raising both hands in the motion sensing data, processor 150 may consider a happy emotion to have been detected.

[0049] It should be noted that, depending on the data source, there may be multiple other methods for determining the first emotion decision, and the embodiments are not limited thereto. Furthermore, in some embodiments, the processor 150 may select one or more data sources from all data sources to determine which data sources correspond to the first emotion decision. The selected data sources may have a more accurate decision regarding emotion estimation than other data sources.

[0050] After determining the first emotion decisions of the plurality of user data (or data sources), the processor 150 may determine whether an emotion conflict occurs between the first emotion decisions (step S250). Specifically, an emotion conflict is associated with the corresponding emotion groups of the first emotion decisions not matching each other. For example, if the first emotion decision of the fourth data source (e.g., eye features) is positive and the first emotion decision of the fifth data source (e.g., mouth features) is negative, then an emotion conflict occurs. For another example, if the first emotion decision of the sixth data source (e.g., electrocardiography (ECG)) is happy and the first emotion decision of the seventh data source (e.g., electromyography (EMG)) is sad, then an emotion conflict occurs.

[0051] In one embodiment, the processor 150 may use a confidence level for decision-making regarding emotional conflicts. The confidence level represents the degree of reliability of the first emotional decision. Specifically, the processor 150 may determine emotion values for each of these first emotional decisions. The emotion value is related to the degree of reliability or reliability level of the first emotional decision. The larger the emotion value, the more reliable the first emotional decision and the higher the reliability level. The smaller the emotion value, the less reliable the first emotional decision and the lower the reliability level. The emotion value may be determined based on the output of the first emotional classifier or another algorithm related to the confidence level. Next, the processor 150 determines a weighted combination of the emotion values and compares the weighted combination of the emotion values with a confidence threshold. The processor 150 may assign a corresponding emotion weight to each emotion value of the first emotional decision and perform a weight calculation on the emotion values using the corresponding emotion weight. If the weighted combination of the emotions is greater than the confidence threshold, then an emotional conflict has not occurred. On the other hand, if the weighted combination is not greater than the confidence threshold, then an emotional conflict has occurred. It should be noted that if the first emotion decision is a weighted combination of emotions of multiple emotion categories, then the emotion value can also be a weighted combination of emotions of multiple emotion categories, and the corresponding reliable threshold will be equal to or similar to a linear equation, a curve equation or another equation in the coordinate system in which the emotion value is located.

[0052] In some embodiments, the processor 150 may select one or more first emotion decisions with higher reliability to determine whether an emotion conflict occurs. For example, the processor 150 selects two first emotion decisions from facial muscle activity and voice, and compares whether the first emotion decisions belong to the same emotion group.

[0053] Next, the processor 150 may determine a second emotion decision from one or more emotion groups based on the determined result of the emotion conflict (step S255). The determined result may be that an emotion conflict has occurred, and the occurrence of the emotion conflict may be that the emotion conflict has not occurred. The processor 150 may fuse the one or more emotion groups to generate a second emotion decision related to (only) one emotion group.

[0054] In one embodiment, if an emotional conflict occurs, the processor 150 may determine a second emotional decision from at least two emotional groups (step S260). Specifically, if an emotional conflict occurs, the first emotional decision may include two or more emotional groups. In the second stage, the processor 150 may further determine a second emotional decision from the emotional group to which the first emotional decision belongs or from all emotional groups, where the second emotional decision is related to (only) one emotional group.

[0055] Figure 7 FIG. 1 is a flow chart illustrating the generation of a second emotion decision according to one of the exemplary embodiments of the present disclosure. Figure 7 In one embodiment, the processor 150 may use one or more first emotion decisions to determine the second emotion decision (step S261). This means that the first emotion decision will be a reference for the second emotion decision. In one embodiment, the processor 150 may determine a weighted decision combination of two or more first emotion decisions, and determine the second emotion decision based on the weighted decision combination. The processor 150 may perform a weight calculation on the first emotion decision, and the calculated result is related to the second emotion decision. The second emotion decision may be a real number, a specific emotion category, a specific level of a specific emotion category, or a weighted combination of emotions of multiple emotion categories. In another embodiment, the second emotion decision is determined via machine learning technology or a tree-based classification method, where the first emotion decision is the input data of the decision model.

[0056] It should be noted that, in some embodiments, the processor 150 may select two or more first emotion decisions having higher reliability or different emotion groups to determine the second emotion decision.

[0057] In another embodiment, the processor 150 may use one or more user data from one or more data sources to determine the second emotion decision (step S263). This means that the user data will be a reference for the second emotion decision. In one embodiment, the processor 150 may determine the second emotion decision by using a second emotion classifier based on machine learning technology or a tree-based classification method. The second emotion classifier is used to identify which of the emotion groups these user data belong to. The user data may be the input data of the second emotion classifier, and the second emotion decision is the output data of the second emotion classifier. It should be noted that the processor 150 may train the second emotion classifier based on multiple previous second training emotions and training sensory data. These second training emotions include two or more emotion groups. This means that the output data of the second emotion classifier may be only one of the selected emotion groups. In addition, the training sensory data is obtained separately from each data source and corresponds to a specific emotion (which will become the second training emotion). The processor 150 may select a second emotion classifier trained by the emotion group of the first emotion decision or all emotion groups for the second emotion decision.

[0058] It should be noted that raw data, pre-processed data, or pre-configured data from multiple data sources may not have the same amount, unit, or collection time interval. Figure 8 FIG. 1 is a flow chart illustrating a user data transformation according to one of the exemplary embodiments of the present disclosure. Figure 8 In one embodiment, the processor 150 may further combine two or more user data to be input to generate a user data combination. For example, after feature extraction, the user data from the first data source is a 40×1 matrix, the user data from the second data source is an 80×2 matrix, and the user data combination will be a 120×1 matrix. The processor 150 may further perform a linear transformation on the facial features extracted from the user data combination (step S262). The linear transformation is designed based on a specific machine learning technique or a specific tree-based classification method. Then, the data after the linear transformation will become the input of the second emotion classifier.

[0059] On the other hand, in one embodiment, if an emotional conflict does not occur, the processor 150 may determine a second emotional decision from (only) one emotional group (step S270). Specifically, if an emotional conflict does not occur, the first emotional decision may include only one emotional group. In one embodiment, an emotional group includes only one emotional category, and the processor 150 may determine any one of the first emotional decisions as the second emotional decision.

[0060] However, in some embodiments, an emotion group may include multiple emotion categories, and an emotion category may include multiple levels. The processor 150 may further determine a second emotion decision based on the emotion category to which the first emotion decision belongs, and the second emotion decision is related (only) to a specific emotion category with a specific level or a specific weighted combination of emotions of the emotion categories.

[0061] In one embodiment, the processor 150 may determine the second emotion decision by using a third emotion classifier based on machine learning technology or a tree-based classification method. The third emotion classifier is used to identify which of the emotion groups the user data or the first emotion decision belongs to. The user data or one or more first emotion decisions are input data of the third emotion classifier, and the second emotion decision is output data of the third emotion classifier. It should be noted that, compared with the first emotion classifier and the second emotion classifier, the processor 150 trains the third emotion classifier based on the third training emotion, and the third training emotion includes only one emotion group. The processor 150 may select the third emotion classifier trained by the emotion group of the first emotion decision for the second emotion decision. In another embodiment, the processor 150 may determine a weighted decision combination of two or more first emotion decisions, and determine the second emotion decision based on the weighted decision combination.

[0062] Next, the processor 150 may generate a facial expression for the avatar based on the second emotion decision (step S290). Specifically, the avatar's face may include multiple facial features (e.g., the shape or movement of the face, eyes, nose, and eyebrows). The avatar's facial expression may include geometric parameters and texture parameters (collectively referred to as facial expression parameters). Each geometric parameter is used to indicate the 2D coordinates or 3D coordinates of a vertex of the avatar's face. In some embodiments, each texture parameter is used to indicate the position of the face to which the facial image corresponding to the second emotion decision (e.g., a specific emotion, a specific level of a specific emotion, or a specific weighted combination of emotions from multiple emotion categories) is applied.

[0063] Processor 150 can utilize features of facial expressions to generate, merge, or replace a second emotion decision to generate a facial expression corresponding to a specific emotion. In one embodiment, processor 150 can select a facial expression from a corresponding facial expression group based on a probability distribution (e.g., normal distribution, geometric distribution, or Bernoulli distribution). Each expression group includes multiple facial expressions. Each emotion or each level of an emotion corresponds to a specific expression group. For example, if there are 10 facial expressions for a specific second emotion decision, processor 150 can randomly select one of the 10 facial expressions.

[0064] In some embodiments, processor 150 may generate facial features for each second emotion decision. Each second emotion decision may be configured with specific restrictions on facial feature parameters (e.g., length, angle, color, size), and corresponding facial features may be generated based on these restrictions. For example, when the second emotion decision has a happy emotion and the emotion weight of the happy emotion is greater than 0.1, the lip length may have a certain range.

[0065] In some embodiments, each second emotion decision corresponds to a face template, and the face template corresponds to a specific image or a specific animation. The processor 150 may paste the face template at a specific position of the face model.

[0066] In summary, in the avatar facial expression generation system and method according to embodiments of the present invention, multiple emotional decisions are determined in the first stage based on multiple data sources. If an emotional conflict arises between these emotional decisions in the first stage, an appropriate emotional decision is determined in the second stage. This allows the avatar to display normal facial expressions accompanied by appropriate emotions, reducing uncertain facial expression parameters. Furthermore, social interaction in virtual environments is improved by the presence of vivid facial expressions.

[0067] It will be apparent to those skilled in the art that various modifications and variations can be made to the structure of the present disclosure without departing from the scope or spirit of the present disclosure. In view of the above, it is hoped that the present disclosure covers modifications and variations of the present disclosure as long as the modifications and variations fall within the scope of the appended claims and their equivalents.

Claims

1. A method for generating facial expressions of a virtual avatar, comprising: obtaining a plurality of user data, wherein each of the user data is related to a sensing result of a user from one of a plurality of data sources; determining a plurality of first emotion decisions based on each of the user data, wherein each of the first emotion decisions is associated with one of a plurality of emotion groups; Determining whether an emotional conflict occurs between the first emotional decisions, wherein the emotional conflict is related to corresponding emotional groups of the first emotional decisions not matching each other, and the step of determining whether an emotional conflict occurs between the first emotional decisions comprises: respectively determining a plurality of emotion values of the first emotion decision; determining an emotion-weighted combination of the emotion values by performing a weight calculation on the emotion values and corresponding emotion weights, wherein each of the emotion values is related to a reliability of a corresponding first emotion decision; and comparing the emotion-weighted combination of the emotion values to a reliability threshold, wherein responsive to the emotion-weighted combination being greater than the reliability threshold, determining that the emotion conflict does not occur, and responsive to the emotion-weighted combination being not greater than the reliability threshold, determining that the emotion conflict occurs; Determining a second emotion decision from at least one of the emotion groups according to the determined result of the emotion conflict, wherein the second emotion decision is related to one of the emotion groups, and the step of determining the second emotion decision from at least one of the emotion groups according to the determined result of the emotion conflict comprises: determining the second emotion decision from at least two of the emotion groups in response to the emotion conflict occurring; and A facial expression of the avatar is generated based on the second emotion decision.

2. The method for generating facial expressions of a virtual avatar according to claim 1 , wherein the step of determining the second emotion decision from at least two of the emotion groups in response to the occurrence of the emotion conflict comprises: The second emotion decision is determined using at least one of the first emotion decisions.

3. The method for generating facial expressions of a virtual avatar according to claim 1 , wherein the step of determining the second emotion decision from at least two of the emotion groups in response to the occurrence of the emotion conflict comprises: The second emotion decision is determined using at least one of the user data.

4. The method for generating facial expressions of a virtual avatar according to claim 2, wherein the step of using at least one of the first emotional decisions to determine the second emotional decision comprises: determining a weighted decision combination of at least two of the first emotion decisions; as well as The second emotion decision is determined based on the weighted decision combination.

5. The method for generating facial expressions of a virtual avatar according to claim 3, wherein the step of using at least one of the user data to determine the second emotional decision comprises: combining at least two of the user data to generate a user data combination; as well as A linear transformation is performed on the combination of user data to extract facial features from the at least two of the user data.

6. The method for generating facial expressions of a virtual avatar according to claim 3, wherein the step of using at least one of the user data to determine the second emotional decision comprises: The second emotion decision is determined based on machine learning technology by using a first emotion classifier, wherein the first emotion classifier is used to identify which of the emotion groups the at least one of the user data belongs to, the at least one of the user data is input data of the first emotion classifier, and the second emotion decision is output data of the first emotion classifier, and the first emotion classifier is trained based on a plurality of first training emotions including at least two of the emotion groups.

7. The method for generating facial expressions of a virtual avatar according to claim 1 , wherein the step of determining the first emotional decision based on each of the user data comprises: Each of the first emotion decisions is determined separately by using a second emotion classifier based on machine learning technology, wherein the second emotion classifier is used to identify which of the emotion groups each of the user data belongs to, each of the user data is input data of the second emotion classifier, each of the first emotion decisions is output data of the second emotion classifier, and the second emotion classifier is trained based on a plurality of second training emotions including all of the emotion groups. 8 . The method for generating facial expressions of a virtual avatar according to claim 1 , wherein each of the first emotion decision or the second emotion decision is a weighted combination of emotions of a plurality of emotion categories.

9. The method for generating facial expressions of a virtual avatar according to claim 1 , wherein the step of determining the second emotion decision from at least one of the emotion groups according to the determined result of the emotion conflict comprises: The second emotion decision is determined from one of the emotion groups in response to the emotion conflict not occurring.

10. The method for generating facial expressions of a virtual avatar according to claim 9, wherein the step of determining the second emotion decision from one of the emotion groups comprises: The second emotion decision is determined based on machine learning technology by using a third emotion classifier, wherein the third emotion classifier is used to identify which of the emotion groups the user data or the first emotion decision belongs to, at least one of the user data or at least one of the first emotion decisions is input data of the third emotion classifier, the second emotion decision is output data of the third emotion classifier, and the third emotion classifier is trained based on a third training emotion that only includes one of the emotion groups.

11. The method for generating facial expressions of an avatar according to claim 1, wherein the data sources differ in target parts or sensing technologies of the user.

12. A virtual avatar facial expression generation system, comprising: at least one tracking device that obtains a plurality of user data, wherein each of the user data is related to a sensing result of a user from one of a plurality of data sources; Memory, which stores program code; as well as A processor, coupled to the memory and loaded with the program code to execute: determining a plurality of first emotion decisions based on each of the user data, wherein each of the first emotion decisions is associated with one of a plurality of emotion groups; determining whether an emotion conflict occurs between the first emotion decisions, wherein the emotion conflict is related to corresponding emotion groups of the first emotion decisions not matching each other; determining a second emotion decision from at least one of the emotion groups based on the determined result of the emotion conflict, wherein the second emotion decision is related to one of the emotion groups; as well as generating a facial expression of the avatar based on the second emotion decision, wherein the processor further performs: respectively determining a plurality of emotion values of the first emotion decision; determining an emotion-weighted combination of the emotion values by performing a weight calculation on the emotion values and corresponding emotion weights, wherein each of the emotion values is related to a reliability of a corresponding first emotion decision; comparing the emotion-weighted combination of the emotion values to a reliability threshold, wherein responsive to the emotion-weighted combination being greater than the reliability threshold, determining that the emotion conflict does not occur, and responsive to the emotion-weighted combination being not greater than the reliability threshold, determining that the emotion conflict occurs; as well as The second emotion decision is determined from at least two of the emotion groups in response to the emotion conflict occurring.

13. The avatar facial expression generation system according to claim 12, wherein the processor further executes: The second emotion decision is determined using at least one of the first emotion decisions.

14. The avatar facial expression generation system according to claim 12, wherein the processor further executes: The second emotion decision is determined using at least one of the user data.

15. The avatar facial expression generation system according to claim 13, wherein the processor further executes: determining a weighted decision combination of at least two of the first emotion decisions; and The second emotion decision is determined based on the weighted decision combination.

16. The avatar facial expression generation system according to claim 14, wherein the processor further executes: combining at least two of the user data to generate a user data combination; and A linear transformation is performed on the combination of user data to extract facial features from the at least two of the user data.

17. The avatar facial expression generation system according to claim 14, wherein the processor further executes: The second emotion decision is determined based on machine learning technology by using a first emotion classifier, wherein the first emotion classifier is used to identify which of the emotion groups the at least one of the user data belongs to, the at least one of the user data is input data of the first emotion classifier, and the second emotion decision is output data of the first emotion classifier, and the first emotion classifier is trained based on a plurality of first training emotions including at least two of the emotion groups.

18. The avatar facial expression generation system according to claim 12, wherein the processor further executes: Each of the first emotion decisions is determined separately by using a second emotion classifier based on machine learning technology, wherein the second emotion classifier is used to identify which of the emotion groups each of the user data belongs to, each of the user data is input data of the second emotion classifier, each of the first emotion decisions is output data of the second emotion classifier, and the second emotion classifier is trained based on a plurality of second training emotions including all of the emotion groups.

19. The avatar facial expression generation system according to claim 12, wherein each of the first emotion decision or the second emotion decision is an emotion-weighted combination of a plurality of emotion categories.

20. The avatar facial expression generation system according to claim 12, wherein the processor further executes: The second emotion decision is determined from one of the emotion groups in response to the emotion conflict not occurring.

21. The avatar facial expression generation system according to claim 20, wherein the processor further executes: The second emotion decision is determined based on machine learning technology by using a third emotion classifier, wherein the third emotion classifier is used to identify which of the emotion groups the user data or the first emotion decision belongs to, at least one of the user data or at least one of the first emotion decisions is input data of the third emotion classifier, the second emotion decision is output data of the third emotion classifier, and the third emotion classifier is trained based on a third training emotion that only includes one of the emotion groups.

22. The avatar facial expression generation system of claim 12, wherein the data sources differ in target parts or sensing techniques of the user.

Citation Information

Patent Citations

  • Emotion recognition method and system thereof

    US20110141258A1

  • Method for detecting facial expressions and emotions of users

    US20190138096A1