A method and system for intelligently generating facial expressions based on facial motion units
By obtaining multimodal facial samples to create unit data sets of facial motion units, using machine learning and physical constraints to generate expression animations, the accuracy and controllability of expression generation in the existing technology are solved, and the expression generation effect with high precision and high reality is achieved.
Patent Information
- Application Number
- CN202510764181.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Existing expression generation technologies are difficult to achieve high-precision, highly realistic and controllable expression generation. The keyframe method is large in work, the expression capture method is costly and the environment is demanding, while deep learning methods are difficult to accurately control the changes in expression details.
By obtaining multimodal facial samples of individuals in specific situations, establishing unit data sets of facial movement units, using machine learning algorithms to learn mapping relationships and neutral relationships, combining physical constraints of facial muscle movement and natural transition constraints of expression changes, expression animations are generated for interaction.
It realizes high-precision, high-reality and controllable expression generation, adapts to different individuals and expression states, and provides real-time and natural expression animation interaction effects.
Smart Images

Figure CN120318890B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent generation technology, and in particular to a facial expression intelligent generation method and system based on facial motion units. Background Art
[0002] Among existing expression generation technologies, common methods include keyframe-based expression animation, expression capture, and deep learning-based expression generation. Keyframe-based methods require manual setting of expression keyframes, which is labor-intensive and makes it difficult to generate natural and smooth expression animations. While expression capture-based methods can obtain realistic and natural expression data, they are expensive, require stringent operating environments, and suffer from issues such as data transmission and processing delays. While deep learning-based expression generation methods can achieve automated generation, they often struggle to precisely control the subtle changes in expressions, resulting in deficiencies in realism and controllability.
[0003] Facial Action Units (FAs) are the basic units of facial muscle movement. Different FA combinations can produce a rich variety of human expressions. However, there is currently no mature technology that effectively integrates FAUs and applies them to intelligent expression generation to achieve high-precision, highly realistic, and controllable expression generation.
[0004] Therefore, the present invention proposes a facial expression intelligent generation method and system based on facial motion units. Summary of the Invention
[0005] The present invention provides a facial expression intelligent generation method and system based on facial motion units, which are used to solve the above-mentioned technical problems.
[0006] The present invention provides a facial expression intelligent generation method based on facial motion units, comprising:
[0007] Step 1: Obtain audio samples and multimodal facial samples of each individual in corresponding set text and set scene, and determine a unit data set of each facial motion unit of the corresponding individual based on the multimodal facial samples;
[0008] Step 2: performing phoneme encoding on the audio sample and inputting the encoded data into an expression predictor to obtain a prediction data set for each facial motion unit;
[0009] Step 3: establishing a mapping relationship based on the unit data set of each facial motion unit in a variety of expression states for different individuals, and establishing a neutral relationship between each facial motion unit and the corresponding expression state in combination with the deviation set determined based on the predicted data set and the unit data set;
[0010] Step 4: Use machine learning algorithms to learn mapping relationships and neutral relationships and establish facial movement unit expression models;
[0011] Step 5: Receive the user's voice interaction instructions, search for the required motion features and parameter settings from the established facial motion unit expression model, and dynamically adjust based on the physical constraints of facial muscle movement and the natural transition constraints of expression changes to generate expression animation for interaction.
[0012] Preferably, the unit data set includes: motion amplitude, motion direction and duration based on each modality.
[0013] Preferably, the multimodal facial samples include: RGB image facial samples, depth image facial samples and electromyographic signal facial samples.
[0014] Preferably, a mapping relationship is established based on the unit data set of each facial motion unit according to different individuals in various expression states, including:
[0015] Sequentially obtain a unit data set based on the same facial motion unit in each expression state;
[0016] The expression state is associated with the unit data set of the same facial motion unit as a mapping relationship.
[0017] Preferably, establishing a neutral relationship between each facial movement unit and the corresponding expression state includes:
[0018] Establishing the prediction data set and the unit data set to determine the spatial occupancy of the corresponding facial motion unit based on the subset under each modality in a preset coordinate system;
[0019] Taking the spatial occupancy of the predicted dataset as a benchmark, the deviation vector is constructed based on the spatial occupancy of each subset in turn. The local expression of the facial error between the corresponding individual audio and the corresponding facial motion unit is obtained, and the global expression of the corresponding individual is established.
[0020] Determine the first extreme condition related to the phoneme and the second extreme condition related to the expression state under different expression states, and combine the boundary conditions of different facial motion units to solve the reference amplitude matrix, reference direction matrix, and reference time matrix of each facial motion unit in the global matrix composed of the global expression under the same expression state under the influence of the first extreme condition and the second extreme condition and the boundary condition constraints;
[0021] The reference amplitude matrix, reference direction matrix and reference time matrix are simplified respectively to obtain reference sets of corresponding facial motion units and associate them as neutral relationships.
[0022] Preferably, obtaining a reference set of corresponding facial motion units includes:
[0023] Numerical cluster analysis was performed on each reference matrix;
[0024] Determine whether the corresponding reference matrix satisfies a preset form. If so, perform a first simplification on the corresponding reference matrix to remove redundant reference elements.
[0025] Otherwise, determine the order of the corresponding reference matrix, and evenly decompose the reference matrix according to the order / 2, and perform numerical cluster analysis on each reference matrix to count the types of clusters contained in each submatrix;
[0026] If the cluster type is one, performing a second simplification on the submatrix according to a preset form;
[0027] If there are multiple cluster types, determine whether a cluster center falls within the sub-matrix; if not, perform a third simplification on the sub-matrix according to a mean processing method;
[0028] If it exists, performing a fourth simplification on the submatrix according to the location and the radiation coefficient;
[0029] Determine the mean value of each column in the simplified matrix to obtain the reference set.
[0030] Preferably, dynamic adjustment is performed based on the physical constraints of facial muscle movement and the natural transition constraints of expression changes, including:
[0031] Constructing a physical model of facial muscle movement and determining physical constraints of facial muscle movement, wherein the physical model of facial muscle movement is implemented based on the anatomical structure and mechanical properties of human facial muscles;
[0032] The initial parameters obtained from the expression model are input into the physical model to check and optimize the parameters of each facial movement unit;
[0033] Based on the emotional trends contained in the user's voice interaction commands and the parameters optimized by physical constraints, an interpolation algorithm is used to dynamically adjust the parameters of each facial motion unit at the start, change, and end stages of the expression animation.
[0034] Drive the graphics rendering engine to generate expression animation.
[0035] Preferably, after generating the facial expression animation for interaction, the method further includes:
[0036] Determining scene attributes of an actual application scene, wherein the scene attributes include formal attributes and informal attributes;
[0037] Calculate the position, displacement, and speed differences of corresponding feature points between the facial animation and the real animation frame by frame, as well as the differences in facial muscle deformation;
[0038] The facial feature point data of real animations and physiological data are aligned in time series. The aligned real animation data and physiological data are then fused to form a high-dimensional feature vector. The fused high-dimensional feature vector is trained using a deep neural network model to explore the micro-correlations between physiological states and real animations, and determine the facial expression change patterns corresponding to different heart rate ranges under corresponding scene attributes.
[0039] Based on the fine-grained features obtained from the change vector of the specified difference at consecutive moments of the same feature point, the specified difference is divided into a main difference and an auxiliary difference in combination with the scene attributes, thereby constructing a main feedback mechanism and an auxiliary feedback mechanism;
[0040] determining a first influence of the facial expression change pattern based on the primary difference and a second influence based on the secondary difference;
[0041] If the first influence is greater than or equal to the second influence, a primary feedback mechanism is constructed based on the primary difference and the facial expression change pattern, and a secondary feedback mechanism is constructed based on the secondary difference;
[0042] If the first influence is smaller than the second influence, a primary feedback mechanism is constructed based on the primary difference, and a secondary feedback mechanism is constructed based on the secondary difference and the facial expression change pattern;
[0043] Based on the model parameter adjustment information calculated by the main feedback mechanism, the facial movement unit expression model is optimized. If the optimization result meets the target setting, the optimization of the model is stopped.
[0044] Otherwise, the auxiliary feedback mechanism is continued to optimize the facial motion unit expression model.
[0045] The present invention provides an expression intelligent generation system based on facial motion units, comprising:
[0046] A data set construction module is used to obtain audio samples and multimodal facial samples of each individual in corresponding set texts and set scenes, and determine a unit data set of each facial motion unit of the corresponding individual based on the multimodal facial samples;
[0047] A prediction module, configured to perform phoneme encoding on the audio sample and input the encoded data into an expression predictor to obtain a prediction data set for each facial motion unit;
[0048] a relationship establishment module for establishing a mapping relationship based on a unit data set of each facial motion unit according to different individuals in multiple expression states, and establishing a neutral relationship between each facial motion unit and the corresponding expression state in combination with a deviation set determined based on the predicted data set and the unit data set;
[0049] A machine learning module is used to learn mapping relationships and neutral relationships using machine learning algorithms to establish a facial movement unit expression model;
[0050] The expression generation module is used to receive the user's voice interaction instructions, search for the required motion features and parameter settings from the established facial motion unit expression model, and dynamically adjust based on the physical constraints of facial muscle movement and the natural transition constraints of expression changes to generate expression animation for interaction.
[0051] Compared with the prior art, the present invention has the following advantages:
[0052] Acquire rich multimodal data, comprehensively describe the characteristics of an individual's facial motor units in specific situations, establish a connection between audio and facial motor units, and through phoneme encoding and expression predictors, predict facial expression changes from voice information, achieve a preliminary association between voice and expression, clarify the intrinsic connection between facial motor units and expression states, correct prediction errors through deviation sets, make the established relationship more in line with actual conditions, improve the adaptability and accuracy of the expression model for different individuals and different expression states, construct a model that can accurately predict the state of facial motor units and expression changes, and generate and interact with real-time, natural expression animations based on user voice commands to achieve high-precision, high-realism and controllable expression generation effects.
[0053] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.
[0054] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0056] Figure 1 Flowchart of a method for intelligently generating facial expressions based on facial motion units in an embodiment of the present invention;
[0057] Figure 2This is a structural diagram of an intelligent expression generation system based on facial motion units in an embodiment of the present invention. DETAILED DESCRIPTION
[0058] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0059] The present invention provides a method for intelligently generating facial expressions based on facial motion units. Figure 1 Shown, including:
[0060] Step 1: Obtain audio samples and multimodal facial samples of each individual in corresponding set text and set scene, and determine a unit data set of each facial motion unit of the corresponding individual based on the multimodal facial samples;
[0061] Step 2: performing phoneme encoding on the audio sample and inputting the encoded data into an expression predictor to obtain a prediction data set for each facial motion unit;
[0062] Step 3: establishing a mapping relationship based on the unit data set of each facial motion unit in a variety of expression states for different individuals, and establishing a neutral relationship between each facial motion unit and the corresponding expression state in combination with the deviation set determined based on the predicted data set and the unit data set;
[0063] Step 4: Use machine learning algorithms to learn mapping relationships and neutral relationships and establish facial movement unit expression models;
[0064] Step 5: Receive the user's voice interaction instructions, search for the required motion features and parameter settings from the established facial motion unit expression model, and dynamically adjust based on the physical constraints of facial muscle movement and the natural transition constraints of expression changes to generate expression animation for interaction.
[0065] Preferably, the unit data set includes: motion amplitude, motion direction and duration based on each modality.
[0066] Preferably, the multimodal facial samples include: RGB image facial samples, depth image facial samples and electromyographic signal facial samples.
[0067] Preferably, a mapping relationship is established based on the unit data set of each facial motion unit according to different individuals in various expression states, including:
[0068] Sequentially obtain a unit data set based on the same facial motion unit in each expression state;
[0069] The expression state is associated with the unit data set of the same facial motion unit as a mapping relationship.
[0070] In this embodiment, the individual is a specific person, and the main purpose is to obtain facial information to obtain a facial sample. The facial motion unit expression model established based on step 4 acts on the face of the robot that interacts with the individual, so that in step 5 the robot can make corresponding facial expression animations for interaction after receiving the user's voice interaction instructions.
[0071] In this embodiment, the audio sample is generated by an individual in a corresponding scenario, for example, a recording of individual A1 saying "Wow, I'm so happy" (set text) when receiving a gift while queuing for his birthday (set scenario).
[0072] In this embodiment, the RGB image facial sample is a facial photo taken by an ordinary color camera; the depth image facial sample can be obtained by a depth camera and can reflect the three-dimensional structural information of the face, such as the degree of protrusion of the nose; the electromyography signal facial sample is data on the electrical activity of the facial muscles collected by sensors, which can understand which muscles are moving.
[0073] In this embodiment, the facial motion unit is a basic unit of facial muscle movement, such as a mouth corner raising unit, a frowning unit, etc. Taking the mouth corner raising unit as an example, the unit data set includes the motion amplitude (the height of the mouth corner raised), motion direction (movement diagonally upward), and duration (the time the mouth corner is kept raised) of the unit under the RGB image; the three-dimensional motion amplitude under the depth image; the duration of the change in the intensity of the relevant muscle electrical activity under the electromyographic signal (electromyographic sensor attached to the face), etc. based on relevant data of each modality.
[0074] In this embodiment, speech processing techniques, such as hidden Markov models (HMMs), are used to analyze audio samples and convert them into phoneme codes. A pre-trained expression predictor can be used to train a deep learning model (such as a recurrent neural network (RNN) or long short-term memory (LSTM) network) using a large number of samples of labeled phoneme codes and corresponding facial movement unit data. The phoneme codes are input into the trained expression predictor to generate a predicted dataset for each facial movement unit. The predicted dataset includes the movement amplitude, direction, and duration of the corresponding facial movement unit.
[0075] In this embodiment, the mapping relationship is the correspondence between different expression states and facial motion unit data sets, for example, the correspondence between the "happy" expression state and related unit data sets such as the mouth corner raising unit and the orbicularis oculi muscle contraction unit.
[0076] In this embodiment, the difference data set between the predicted data set and the unit data set is the deviation set, that is, the deviation set contains a difference subset of each modality information under the predicted data set and the unit data set.
[0077] In this embodiment, the neutral relationship is a relatively accurate relationship between facial motion units and expression states determined based on the mapping relationship and the deviation set.
[0078] In this embodiment, the facial motion unit expression model is constructed by learning mapping relationships and neutral relationships, and can predict the state of the facial motion unit based on input information, and then generate a model of the corresponding expression.
[0079] In this embodiment, the voice interaction instruction is an instruction issued by the user through voice, such as "show a happy expression".
[0080] Motion features: relevant features of facial motion units, such as movement amplitude, direction, speed, etc.
[0081] Parameter settings: parameters related to expression generation, such as expression duration, transition speed, etc.
[0082] Physical constraints: restrictions based on the physiological characteristics of facial muscle movement, such as the upward angle of the mouth corners cannot exceed a certain range. Natural transition constraints: restrictions to ensure that facial expressions change naturally and smoothly, such as the change from smiling to laughing cannot be too abrupt.
[0083] Expression animation: Dynamic facial expression display generated based on the state of facial motor units, presented in the form of video or animation sequence.
[0084] The beneficial effects of the above technical solution are: obtaining rich multimodal data, comprehensively describing the characteristics of the facial movement units of an individual in a specific situation, establishing a connection between audio and facial movement units, and through phoneme encoding and expression predictors, being able to predict facial expression changes from voice information, realizing a preliminary association between voice and expression, clarifying the intrinsic connection between facial movement units and expression states, correcting prediction errors through deviation sets, making the established relationship more in line with actual conditions, improving the adaptability and accuracy of the expression model to different individuals and different expression states, constructing a model that can accurately predict the state of facial movement units and expression changes, and achieving real-time, natural expression animation generation and interaction based on user voice commands, to achieve high-precision, high-realism and controllable expression generation effects.
[0085] The present invention provides an intelligent expression generation method based on facial motion units, which establishes a neutral relationship between each facial motion unit and the corresponding expression state, including:
[0086] Establishing the prediction data set and the unit data set to determine the spatial occupancy of the corresponding facial motion unit based on the subset under each modality in a preset coordinate system;
[0087] Taking the spatial occupancy of the predicted dataset as a benchmark, the deviation vector is constructed based on the spatial occupancy of each subset in turn. The local expression of the facial error between the corresponding individual audio and the corresponding facial motion unit is obtained, and the global expression of the corresponding individual is established.
[0088] Determine the first extreme condition related to the phoneme and the second extreme condition related to the expression state under different expression states, and combine the boundary conditions of different facial motion units to solve the reference amplitude matrix, reference direction matrix, and reference time matrix of each facial motion unit in the global matrix composed of the global expression under the same expression state under the influence of the first extreme condition and the second extreme condition and the boundary condition constraints;
[0089] The reference amplitude matrix, reference direction matrix and reference time matrix are simplified respectively to obtain reference sets of corresponding facial motion units and associate them as neutral relationships.
[0090] In this embodiment, a preset coordinate system is a coordinate system artificially set to describe the position and direction of facial motion unit data, similar to the coordinate system on a map, which facilitates the unified measurement of the state of facial motion units in space. For example, a coordinate system is established with the tip of the nose as the origin, the horizontal rightward direction is the x-axis, the vertical upward direction is the y-axis, and the vertical screen outward direction is the z-axis. The RGB image modality subset: the movement amplitude is 3mm in the x-axis direction, 4mm in the y-axis direction, and the duration is 0.5 seconds, which can be expressed as {(3,4,0),0.5}; the depth image modality subset: the coordinates in the three-dimensional space are (3,3,0), the duration is 0.6 seconds, that is, {(3,3,0),0.6}; the electromyography signal modality subset: the relevant muscle electrical signal intensity variation range is [10,20]mV, the duration is 0.4 seconds, which can be expressed as {[10,20],0.4}.
[0091] For example, the predicted data set is: the movement amplitude is 4mm in the x-axis direction, 4mm in the y-axis direction, and the duration is 0.45 seconds, that is, {(4,4,0),0.45}. The corresponding EMG signal prediction result is: the signal strength is in the range of [12,22]mV, and the duration is 0.45 seconds, that is, {[12, 22], 0.45}. The above series of results are the spatial occupancy in the preset coordinate system, which is mainly reflected in the coordinates.
[0092] In this embodiment, taking the RGB image as an example, the distance deviation between {(4,4,0),0.45} and {(3,4,0),0.5} is: , the time deviation is: |0.4-0.5|, at this time, the deviation vector is (1, 0.05) as the local expression, and the global expression is all local expressions.
[0093] Since the global expression is the local expression under each modality, and there are several individuals and facial motion units, the global matrix based on the global expression is: QJ= , where n represents the total number of individuals. Therefore, by setting constraints on this matrix, the reference amplitude matrix, reference direction matrix, and reference time matrix can be directly extracted.
[0094] In this embodiment, assuming that the expression is "happy", the first extreme condition related to the phoneme is that when making certain specific sounds, the maximum movement amplitude of the mouth corner raising unit in the x-axis direction cannot exceed 8mm; the second extreme condition related to the expression state is that when extremely happy, the maximum movement amplitude in the y-axis direction cannot exceed 6mm; the boundary condition is that the movement amplitude range in the z-axis direction is [-1,1]mm. The global matrix based on the global expression is solved by mathematical algorithms such as linear programming. For the reference amplitude matrix, the reference amplitude range of the mouth corner raising unit in the x-axis direction is [2,6]mm, the y-axis direction is [3,5]mm, and the z-axis direction is [-1,1]mm under the influence of extreme conditions and meeting the boundary conditions, which can be expressed as: , at this time, it is assumed that the actual amplitude matrix obtained is: ,at this time, Corresponding to the x-axis, Corresponding to the y-axis, Corresponding to the z-axis, since 8 is not between 2 and 6, 8 is modified by 6 to obtain the reference amplitude matrix: .
[0095] Since the positive direction of the x, y, and z axes is 1, the negative direction is -1, and the axis is 0, the reference direction matrix is: At this time, the reference time matrix obtained is: [a1, a2], where a1 and a2 are the reference ranges of duration obtained by statistics of multiple experiments, which are obtained based on the boundary average value.
[0096] The first extreme condition is a special case related to phonemes, such as the facial movement units reaching their limit when pronouncing certain phonemes, such as the violent movement of the lip muscles when pronouncing explosive sounds.
[0097] Second extreme condition: special circumstances related to the expression state, such as the extreme contraction state of the eyebrows and muscles around the eyes in an extremely angry expression.
[0098] Boundary conditions: The restrictions on the movement of the facial movement unit itself, based on the physiological structure of the facial muscles, such as the upward range of the eyebrows cannot exceed a certain angle.
[0099] In this embodiment, the simplified matrix results are associated to obtain a reference set corresponding to the mouth corner raising unit as {amplitude reference: [5, 5, 0], direction reference: [1, 1, 0], time reference: 0.5}, which serves as the neutral relationship between the mouth corner raising unit and the "happy" expression state.
[0100] The beneficial effects of the above technical solution are: by establishing a neutral relationship, it can more accurately describe the relationship between facial motion units and expression states, and use spatial analysis and extreme condition constraints to effectively correct prediction errors, so that the expression generation model can more accurately simulate the actual movement of facial motion units of different individuals in various expression states, thereby improving the accuracy and authenticity of expression generation.
[0101] The present invention provides an intelligent expression generation method based on facial motion units, which obtains a reference set of corresponding facial motion units, including:
[0102] Numerical cluster analysis was performed on each reference matrix;
[0103] Determine whether the corresponding reference matrix satisfies a preset form. If so, perform a first simplification on the corresponding reference matrix to remove redundant reference elements.
[0104] Otherwise, determine the order of the corresponding reference matrix, and evenly decompose the reference matrix according to the order / 2, and perform numerical cluster analysis on each reference matrix to count the types of clusters contained in each submatrix;
[0105] If the cluster type is one, performing a second simplification on the submatrix according to a preset form;
[0106] If there are multiple cluster types, determine whether a cluster center falls within the sub-matrix; if not, perform a third simplification on the sub-matrix according to a mean processing method;
[0107] If it exists, performing a fourth simplification on the submatrix according to the location and the radiation coefficient;
[0108] Determine the mean value of each column in the simplified matrix to obtain the reference set.
[0109] In this embodiment, performing a fourth simplification on the submatrix includes:
[0110] Obtain the variance of the first column after excluding each cluster center in the submatrix and the variance of the second column of the submatrix, and calculate the absolute value of the difference between each variance of the first column and the variance of the second column;
[0111] A discrete analysis is performed on each column vector in the submatrix in turn, and a second number N2 of discrete points of each column vector is counted. At the same time, a first number N1 of difference values whose absolute value is greater than a preset value is counted. When a difference value has an absolute value greater than the preset value, the corresponding position is calibrated for significance, and the number of significance calibrations is the first number N1. The preset value is obtained by matching from a matrix type-value comparison table, which contains direction, amplitude, and time matrix types, and the preset value corresponding to each matrix is pre-set, which is convenient for direct retrieval and use.
[0112] Determine whether there are the same points in each column vector based on the first number N1 and the second number N2. If so, the number of the same points is used as the radiation number N3, and The radiation coefficient is obtained by multiplying the average of the absolute values of the differences of the corresponding column vectors, where N is the total number of elements in the column vector;
[0113] The fourth simplification is achieved by averaging the corresponding columns of the submatrix and adding or subtracting the average from the corresponding radiation coefficient.
[0114] It should be noted that the rules for determining addition or subtraction are as follows: when the number of positive differences in each first column variance and second column variance is greater than the number of negative differences, addition is determined; otherwise, subtraction is determined.
[0115] By calculating variance, we can determine the degree of dispersion of the submatrix data. Variance reflects the fluctuation of the data around the mean. The absolute value of the difference between the variances of the first and second columns measures the difference in dispersion between the two columns. Discrete analysis counts the number of discrete points (N2) and the number of absolute differences greater than a preset value (N1), providing further insight into the data distribution characteristics. Different thresholds are preset for different matrix types (directional, amplitude, time matrices, etc.) because the data characteristics of different matrices vary. Setting these thresholds as needed allows analysis to be tailored to each matrix type, providing more accurate data analysis.
[0116] The first number N1 and the second number N2 are identified, and their number is used as the radiation number N3. Based on this, the radiation coefficient is calculated. This process considers the correlation between different statistical quantities, integrates multiple aspects of information, and avoids the one-sidedness of single-dimensional analysis. The radiation coefficient obtained in this way can more comprehensively reflect the characteristics of the submatrix column vector data, providing a more reasonable basis for subsequent matrix simplification.
[0117] The fourth simplification, achieved by averaging the matrix columns and combining it with the emissivity calculation, is based on the data characteristics analyzed previously. Averaging provides a general understanding of the data, while the emissivity factor incorporates characteristics such as data dispersion and significant deviations. Combining these two methods allows for a reasonable simplification of the matrix while preserving important data features. The resulting matrix is both concise and reflects the key information of the original matrix.
[0118] The original matrix data is redundant and complex, and direct processing is inefficient and easy to conceal key information. Through the above series of operations, the key features of the data can be refined, unnecessary details can be removed, the data dimension can be reduced, the efficiency of subsequent data processing and analysis can be improved, and the data with key features can be highlighted, which helps the model to better learn the rules and improve the model accuracy and generalization ability.
[0119] In this embodiment, a new matrix is constructed according to the simplified result of each sub-matrix, and the new matrix is the simplified matrix.
[0120] In this embodiment, a decomposition scale of order / 2 can balance "local detail preservation" and "global trend extraction." If the decomposition scale is too small (such as order / 4), the submatrix may contain data across muscle groups (such as the junction area between the corners of the mouth and the cheeks), resulting in mixed clustering results. If the scale is too large, the submatrix may contain data from multiple muscle groups (such as the area around the eyes and mouth at the same time), making it impossible to distinguish the characteristics of different motor units. Assuming that the residual of the original matrix is R, the residual of each submatrix after decomposition at order / 2 is R1~R4, and the total residual R_total=R1+R2+R3+R4. Experiments show that when the decomposition scale is order / 2, R_total≈R, that is, the decomposition loses almost no information. If the scale is order / 3, R_total may increase by 15%~20% (due to increased clustering error caused by interference from cross-regional data).
[0121] In this embodiment, a clustering algorithm (such as a K-means clustering algorithm) is used to perform numerical clustering analysis on each reference matrix, and the data in the matrix are classified according to numerical features.
[0122] Cluster center: the average value of the values related to the x-axis, y-axis, and z-axis in the corresponding clustering results.
[0123] In this embodiment, for example, for the reference amplitude matrix, the preset form is that the number of matrix rows does not exceed 2 and the difference between elements does not exceed 1.5 mm. At this time, the original matrices of the reference amplitude matrix and the reference direction matrix are 3 rows, which do not meet the preset form. The original matrix of the reference time matrix is 1 row, which meets the preset form and does not need to be simplified.
[0124] In this embodiment, it is assumed that there is a reference amplitude matrix B with 4 rows and 3 columns. In this case, the preset form is not satisfied, and the matrix B is split into 2 sub-matrices according to the result of order / 2, that is, each sub-matrix corresponds to two rows. It should be noted that the number of rows of matrix B is greater than 100. For the convenience of calculation, 4 rows are used here.
[0125] Assume that the matrix , the number of clusters obtained is k=2, and the matrix The cluster centers are: ,at this time, Not present The simplified result is: .
[0126] The second simplification is that when there is only one type of cluster in the submatrix, it is simplified according to the preset form to remove repeated or redundant information; the third simplification is that when there are multiple types of clusters in the submatrix and there is no cluster center falling in the submatrix, it is simplified according to the number of elements under each cluster type and the main information is retained; the fourth simplification is that when there are multiple types of clusters in the submatrix and there is a cluster center falling in the submatrix, it is simplified according to the location and radiation range of the cluster center to highlight the key information.
[0127] The beneficial effects of the above technical solution are: strictly following the order / 2 decomposition rule, demonstrating the difference between the two simplified logics through specific matrix operations, and obtaining a reference set through a series of analysis and simplification processing of the reference matrix, which can remove redundant and interfering information in the reference matrix and retain key features, so that the reference set can more concisely and accurately reflect the core motion characteristics of the facial motion unit, providing a more efficient and accurate reference basis for the expression generation model, and further improving the quality and efficiency of expression generation.
[0128] The present invention provides an intelligent expression generation method based on facial motion units, which performs dynamic adjustment based on the physical constraints of facial muscle movement and the natural transition constraints of expression changes, including:
[0129] Constructing a physical model of facial muscle movement and determining physical constraints of facial muscle movement, wherein the physical model of facial muscle movement is implemented based on the anatomical structure and mechanical properties of human facial muscles;
[0130] The initial parameters obtained from the expression model are input into the physical model to check and optimize the parameters of each facial movement unit;
[0131] Based on the emotional trends contained in the user's voice interaction commands and the parameters optimized by physical constraints, an interpolation algorithm is used to dynamically adjust the parameters of each facial motion unit at the start, change, and end stages of the expression animation.
[0132] Drive the graphics rendering engine to generate expression animation.
[0133] In this embodiment, medical imaging technology (such as MRI and CT) is used to obtain anatomical data of human facial muscles, and information such as the starting point, insertion point, and shape of the muscles is recorded; and biomechanical experiments are used to measure mechanical property data such as the strength and elastic coefficient of facial muscles in different contraction states.
[0134] Model construction: Based on the collected data, computer graphics and mechanics algorithms are used to construct a physical model of facial muscle movement in a virtual environment. For example, finite element analysis can be used to divide facial muscles into multiple tiny units to simulate muscle deformation under stress. Alternatively, physics-based modeling methods can be used to describe the connection between muscles, bones, and skin using mathematical equations.
[0135] Constraint determination: Combining anatomical knowledge and experimental data, we analyze the limits of facial muscle movement and determine physical constraints. For example, by observing a large number of real-life facial expressions, we found that in normal facial expressions, the muscles around the eyes contract to a degree that closes the eyes but does not completely block vision. This allows us to set angular constraints for eye muscle movement based on this.
[0136] Initial parameters: Parameters of facial motion units obtained from the facial expression model, such as the amplitude, direction, speed, and duration of each facial motion unit. For example, in a "happy" expression, the initial amplitude of the upward movement of the corner of the mouth is 10 mm, and the movement direction is 45 degrees upward.
[0137] Check and optimize: Check whether the initial parameters meet the physical constraints. If not, adjust and optimize to make the parameters reasonable. For example, if the initial parameters for the upward angle of the mouth corner are 20mm, which exceeds the upper limit of 15mm of the physical constraint, adjust it to 15mm.
[0138] The initial stage is the stage when the expression begins to appear, the changing stage is the process of the expression changing from the initial state to the target state, and the ending stage is the stage when the expression becomes stable or disappears.
[0139] In facial expression animation, linear interpolation algorithms are used to interpolate parameters between the start and end states of an expression, making the transition more natural. For example, in the initial phase, an interpolation algorithm (such as linear interpolation) is used to smoothly transition from the neutral state parameters to the optimized initial expression parameters. In the transition phase, the parameters of the facial movement units are dynamically adjusted based on the trend of emotional changes. If the emotion becomes increasingly excited, the values of parameters such as the upward angle of the mouth corners and the angle of the eyebrow rise are increased. In the final phase, the interpolation algorithm is also used to smoothly transition the expression from the current state to the final state (such as returning to the neutral state).
[0140] Graphics rendering engine: A software or hardware system used to convert virtual 3D models and animation data into visual images or animations. Common graphics rendering engines include Unity's rendering engine and Unreal Engine.
[0141] The beneficial effects of the above technical solution are: establishing a physical model and constraints of facial muscle movement that conforms to the physiological characteristics of the human body, providing a scientific basis for subsequent facial expression animation generation, ensuring that the parameters of the facial expression animation conform to the physical laws of human facial muscle movement by obtaining parameters from the model and checking whether the parameters meet the constraints, and enabling the facial expression animation to be dynamically and naturally adjusted according to the user's emotional changes through linear interpolation, converting the processed facial expression animation data into intuitive and visible animation effects, providing users with a vivid and realistic visual experience.
[0142] The present invention provides a facial expression intelligent generation method based on facial motion units, which, after generating facial expression animation for interaction, further comprises:
[0143] Determining scene attributes of an actual application scene, wherein the scene attributes include formal attributes and informal attributes;
[0144] Calculate the position, displacement, and speed differences of corresponding feature points between the facial animation and the real animation frame by frame, as well as the differences in facial muscle deformation;
[0145] The facial feature point data of real animations and physiological data are aligned in time series. The aligned real animation data and physiological data are then fused to form a high-dimensional feature vector. The fused high-dimensional feature vector is trained using a deep neural network model to explore the micro-correlations between physiological states and real animations, and determine the facial expression change patterns corresponding to different heart rate ranges under corresponding scene attributes.
[0146] Based on the fine-grained features obtained from the change vector of the specified difference at consecutive moments of the same feature point, the specified difference is divided into a main difference and an auxiliary difference in combination with the scene attributes, thereby constructing a main feedback mechanism and an auxiliary feedback mechanism;
[0147] determining a first influence of the facial expression change pattern based on the primary difference and a second influence based on the secondary difference;
[0148] If the first influence is greater than or equal to the second influence, a primary feedback mechanism is constructed based on the primary difference and the facial expression change pattern, and a secondary feedback mechanism is constructed based on the secondary difference;
[0149] If the first influence is smaller than the second influence, a primary feedback mechanism is constructed based on the primary difference, and a secondary feedback mechanism is constructed based on the secondary difference and the facial expression change pattern;
[0150] Based on the model parameter adjustment information calculated by the main feedback mechanism, the facial movement unit expression model is optimized. If the optimization result meets the target setting, the optimization of the model is stopped.
[0151] Otherwise, the auxiliary feedback mechanism is continued to optimize the facial motion unit expression model.
[0152] In this embodiment, scenario attributes refer to characteristic classifications of actual application scenarios, including formal attributes (such as business meetings and speeches) and informal attributes (such as chatting and games).
[0153] In this embodiment, for example, the determined feature point differences and muscle deformation differences are as follows:
[0154] Table 1 Differences in feature points and muscle deformation
[0155]
[0156] Quantifying expression similarity provides accurate error indicators for model optimization, making subsequent training objectives clearer.
[0157] Time series alignment: Synchronize facial landmark data with physiological data (such as heart rate and electromyography) in the time dimension.
[0158] High-dimensional feature vector: A composite feature vector formed by fusing multimodal data, such as [position, heart rate, electromyography].
[0159] Micro-correlations: Weak but meaningful associations between physiological states and facial expressions, such as increased blinking rate when heart rate increases.
[0160] The DTW (Dynamic Time Warping) algorithm is used to align data, and the LSTM network is used to mine micro-correlations.
[0161] Input data:
[0162] Facial landmark sequence:
[0163] Heart rate sequence: [72,75,78,...,85]
[0164] The fused feature vector:
[0165] By classifying heart rate zones (e.g. 60-70bpm, 70-80bpm), we can discover facial expression patterns in specific scenarios:
[0166] Formal scenario: When the heart rate is >80bpm, pupil dilation increases by 15%.
[0167] Informal scenario: When the heart rate is >80bpm, the speed at which the corners of the mouth rise increases by 20%.
[0168] Analyze the contribution of each difference through SHAP value and classify the main and auxiliary differences:
[0169] Formal scenes: primary differences (mouth corner position, pupil size), secondary differences (eyebrow height).
[0170] Informal scenario: primary differences (eyebrow dynamics, blinking frequency), secondary differences (nasal flare).
[0171] The primary-auxiliary feedback mechanism is a hierarchical optimization strategy based on the importance of different differences, and a dynamic weight distribution mechanism is constructed. The weight of the difference in the corner of the mouth position in formal scenes is 0.7, and the dynamic weight of the eyebrows in informal scenes is increased to 0.6.
[0172] The influence coefficients were calculated using structural equation modeling (SEM):
[0173] Formal scenario: First effect = 0.82, Second effect = 0.56 → The primary feedback mechanism is dominated by the main difference.
[0174] Informal scenario: First influence = 0.45, Second influence = 0.68 → The primary feedback mechanism integrates auxiliary differences and expression patterns.
[0175] Optimization of the main feedback mechanism: Adjustment of the parameters for the upward amplitude of the mouth corners to reduce the error from 1.2mm to 0.5mm.
[0176] Optimized the auxiliary feedback mechanism: A mapping relationship between heart rate and blink frequency was added, reducing the blink frequency error from 15% to 8%.
[0177] The model optimization iteration effect is shown in the table:
[0178] Table 2 Model optimization iteration effect
[0179]
[0180] Scenario adaptability: In formal and informal scenarios, the naturalness of expression scores increased by 25% and 31%, respectively.
[0181] Physiological authenticity: The correlation between heart rate and facial expressions increased by 35%, and the accuracy of capturing micro-expressions increased by 42%.
[0182] Iteration efficiency: Through the primary and secondary feedback mechanism, the model convergence speed is increased by 40% and the training time is reduced by 32%.
[0183] The beneficial effects of the above technical solution are: adaptively adjusting the feedback mechanism, adding the association rules between heart rate and eyebrow dynamics to the scene to make expressions more immersive, and significantly improving the realism and adaptability of expression animation through the layered optimization strategy of scene perception.
[0184] The present invention provides an expression intelligent generation system based on facial motion units, such as Figure 2 Shown, including:
[0185] A data set construction module is used to obtain audio samples and multimodal facial samples of each individual in corresponding set texts and set scenes, and determine a unit data set of each facial motion unit of the corresponding individual based on the multimodal facial samples;
[0186] A prediction module, configured to perform phoneme encoding on the audio sample and input the encoded data into an expression predictor to obtain a prediction data set for each facial motion unit;
[0187] a relationship establishment module for establishing a mapping relationship based on a unit data set of each facial motion unit according to different individuals in multiple expression states, and establishing a neutral relationship between each facial motion unit and the corresponding expression state in combination with a deviation set determined based on the predicted data set and the unit data set;
[0188] A machine learning module is used to learn mapping relationships and neutral relationships using machine learning algorithms to establish a facial movement unit expression model;
[0189] The expression generation module is used to receive the user's voice interaction instructions, search for the required motion features and parameter settings from the established facial motion unit expression model, and dynamically adjust based on the physical constraints of facial muscle movement and the natural transition constraints of expression changes to generate expression animation for interaction.
[0190] The beneficial effects of the above technical solution are: obtaining rich multimodal data, comprehensively describing the characteristics of the facial movement units of an individual in a specific situation, establishing a connection between audio and facial movement units, and through phoneme encoding and expression predictors, being able to predict facial expression changes from voice information, realizing a preliminary association between voice and expression, clarifying the intrinsic connection between facial movement units and expression states, correcting prediction errors through deviation sets, making the established relationship more in line with actual conditions, improving the adaptability and accuracy of the expression model to different individuals and different expression states, constructing a model that can accurately predict the state of facial movement units and expression changes, and achieving real-time, natural expression animation generation and interaction based on user voice commands, to achieve high-precision, high-realism and controllable expression generation effects.
[0191] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. A facial expression intelligent generation method based on facial motion units, characterized in that: include: Step 1: Obtain audio samples and multimodal facial samples of each individual in corresponding set text and set scene, and determine a unit data set of each facial motion unit of the corresponding individual based on the multimodal facial samples; Step 2: performing phoneme encoding on the audio sample and inputting the encoded data into an expression predictor to obtain a prediction data set for each facial motion unit; Step 3: establishing a mapping relationship based on the unit data set of each facial motion unit in a variety of expression states for different individuals, and establishing a neutral relationship between each facial motion unit and the corresponding expression state in combination with the deviation set determined based on the predicted data set and the unit data set; Step 4: Use machine learning algorithms to learn mapping relationships and neutral relationships and establish facial movement unit expression models; Step 5: Receive the user's voice interaction command, search for the required motion features and parameter settings from the established facial motion unit expression model, and dynamically adjust based on the physical constraints of facial muscle movement and the natural transition constraints of expression changes to generate expression animation for interaction; Among them, establishing a neutral relationship between each facial movement unit and the corresponding expression state includes: Establishing the prediction data set and the unit data set to determine the spatial occupancy of the corresponding facial motion unit based on the subset under each modality in a preset coordinate system; Taking the spatial occupancy of the predicted dataset as a benchmark, the deviation vector is constructed based on the spatial occupancy of each subset in turn. The local expression of the facial error between the corresponding individual audio and the corresponding facial motion unit is obtained, and the global expression of the corresponding individual is established. Determine the first extreme condition related to the phoneme and the second extreme condition related to the expression state under different expression states, and combine the boundary conditions of different facial motion units to solve the reference amplitude matrix, reference direction matrix, and reference time matrix of each facial motion unit in the global matrix composed of the global expression under the same expression state under the influence of the first extreme condition and the second extreme condition and the boundary condition constraints; The reference amplitude matrix, reference direction matrix and reference time matrix are simplified respectively to obtain the reference set of corresponding facial motion units and correlate them as neutral relations; Among them, the reference set of corresponding facial motion units is obtained, including: Numerical cluster analysis was performed on each reference matrix; Determine whether the corresponding reference matrix satisfies a preset form. If so, perform a first simplification on the corresponding reference matrix to remove redundant reference elements. Otherwise, determine the order of the corresponding reference matrix, and evenly decompose the reference matrix according to the order / 2, and perform numerical cluster analysis on each reference matrix to count the types of clusters contained in each submatrix; If the cluster type is one, performing a second simplification on the submatrix according to a preset form; If there are multiple cluster types, determine whether a cluster center falls within the sub-matrix; if not, perform a third simplification on the sub-matrix according to a mean processing method; If it exists, performing a fourth simplification on the submatrix according to the location and the radiation coefficient; Determine the mean value of each column in the simplified matrix to obtain the reference set.
2. The method for intelligently generating facial expressions based on facial motion units according to claim 1, wherein: The unit data set includes: motion amplitude, motion direction, and duration based on each modality.
3. The method for intelligently generating facial expressions based on facial motion units according to claim 1, wherein: The multimodal facial samples include: RGB image facial samples, depth image facial samples and electromyographic signal facial samples.
4. The method for intelligently generating facial expressions based on facial motion units according to claim 1, wherein: A mapping relationship is established based on the unit data set of each facial motion unit for different individuals in various expression states, including: Sequentially obtain a unit data set based on the same facial motion unit in each expression state; The expression state is associated with the unit data set of the same facial motion unit as a mapping relationship.
5. The method for intelligently generating facial expressions based on facial motion units according to claim 1, wherein: Dynamic adjustments are made based on the physical constraints of facial muscle movement and the natural transition constraints of expression changes, including: Constructing a physical model of facial muscle movement and determining physical constraints of facial muscle movement, wherein the physical model of facial muscle movement is implemented based on the anatomical structure and mechanical properties of human facial muscles; The initial parameters obtained from the expression model are input into the physical model to check and optimize the parameters of each facial movement unit; Based on the emotional trends contained in the user's voice interaction commands and the parameters optimized by physical constraints, an interpolation algorithm is used to dynamically adjust the parameters of each facial motion unit at the start, change, and end stages of the expression animation. Drive the graphics rendering engine to generate expression animation.
6. The method for intelligently generating facial expressions based on facial motion units according to claim 1, wherein: After generating expression animation for interaction, it also includes: Determining scene attributes of an actual application scene, wherein the scene attributes include formal attributes and informal attributes; Calculate the position, displacement, and speed differences of corresponding feature points between the facial animation and the real animation frame by frame, as well as the differences in facial muscle deformation; The facial feature point data of real animations and physiological data are aligned in time series. The aligned real animation data and physiological data are then fused to form a high-dimensional feature vector. The fused high-dimensional feature vector is trained using a deep neural network model to explore the micro-correlations between physiological states and real animations, and determine the facial expression change patterns corresponding to different heart rate ranges under corresponding scene attributes. Based on the fine-grained features obtained from the change vector of the specified difference at consecutive moments for the same feature point, the specified difference is divided into a main difference and an auxiliary difference in combination with the scene attributes, thereby constructing a main feedback mechanism and an auxiliary feedback mechanism; determining a first influence of the facial expression change pattern based on the primary difference and a second influence based on the secondary difference; If the first influence is greater than or equal to the second influence, a primary feedback mechanism is constructed based on the primary difference and the facial expression change pattern, and a secondary feedback mechanism is constructed based on the secondary difference; If the first influence is smaller than the second influence, a primary feedback mechanism is constructed based on the primary difference, and a secondary feedback mechanism is constructed based on the secondary difference and the facial expression change pattern; Based on the model parameter adjustment information calculated by the main feedback mechanism, the facial movement unit expression model is optimized. If the optimization result meets the target setting, the optimization of the model is stopped. Otherwise, the auxiliary feedback mechanism is continued to optimize the facial motion unit expression model.
7. An intelligent expression generation system based on facial motion units, characterized in that: include: A data set construction module is used to obtain audio samples and multimodal facial samples of each individual in corresponding set texts and set scenes, and determine a unit data set of each facial motion unit of the corresponding individual based on the multimodal facial samples; A prediction module, configured to perform phoneme encoding on the audio sample and input the encoded data into an expression predictor to obtain a prediction data set for each facial motion unit; a relationship establishment module for establishing a mapping relationship based on a unit data set of each facial motion unit according to different individuals in multiple expression states, and establishing a neutral relationship between each facial motion unit and the corresponding expression state in combination with a deviation set determined based on the predicted data set and the unit data set; A machine learning module is used to learn mapping relationships and neutral relationships using machine learning algorithms to establish a facial movement unit expression model; The expression generation module is used to receive the user's voice interaction instructions, find the required movement features and parameter settings from the established facial movement unit expression model, and dynamically adjust based on the physical constraints of facial muscle movement and the natural transition constraints of expression changes to generate expression animation for interaction; Among them, establishing a neutral relationship between each facial movement unit and the corresponding expression state includes: Establishing the prediction data set and the unit data set to determine the spatial occupancy of the corresponding facial motion unit based on the subset under each modality in a preset coordinate system; Taking the spatial occupancy of the predicted dataset as a benchmark, the deviation vector is constructed based on the spatial occupancy of each subset in turn. The local expression of the facial error between the corresponding individual audio and the corresponding facial motion unit is obtained, and the global expression of the corresponding individual is established. Determine the first extreme condition related to the phoneme and the second extreme condition related to the expression state under different expression states, and combine the boundary conditions of different facial motion units to solve the reference amplitude matrix, reference direction matrix, and reference time matrix of each facial motion unit in the global matrix composed of the global expression under the same expression state under the influence of the first extreme condition and the second extreme condition and the boundary condition constraints; The reference amplitude matrix, reference direction matrix and reference time matrix are simplified respectively to obtain the reference set of corresponding facial motion units and correlate them as neutral relations; Among them, the reference set of corresponding facial motion units is obtained, including: Numerical cluster analysis was performed on each reference matrix; Determine whether the corresponding reference matrix satisfies a preset form. If so, perform a first simplification on the corresponding reference matrix to remove redundant reference elements. Otherwise, determine the order of the corresponding reference matrix, and evenly decompose the reference matrix according to the order / 2, and perform numerical cluster analysis on each reference matrix to count the types of clusters contained in each submatrix; If the cluster type is one, performing a second simplification on the submatrix according to a preset form; If there are multiple cluster types, determine whether a cluster center falls within the sub-matrix; if not, perform a third simplification on the sub-matrix according to a mean processing method; If it exists, performing a fourth simplification on the submatrix according to the location and the radiation coefficient; Determine the mean value of each column in the simplified matrix to obtain the reference set.
Citation Information
Patent Citations
Robot emotion recognition method and system based on AIGC and storage medium
CN118626966A
Voice-driven three-dimensional face animation generation method and device based on reinforcement learning
CN119027557A