Intelligent expression generation method and system based on facial motion unit
The integration of facial action units with audio encoding and machine learning addresses the limitations of existing expression generation methods, enabling high-precision, realistic, and controllable expression synthesis through user-driven facial animation.
Patent Information
- Application Number
- CN202510764181.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Existing expression generation technologies are difficult to achieve high-precision, highly realistic and controllable expression generation. The keyframe-based method is work-intensive, the expression-based method is costly and the environment is harsh. It is difficult to accurately control the changes in expression details based on deep learning methods.
By obtaining the individual's audio and multimodal facial samples in a specific situation, a unit data set of facial movement units is established, and a machine learning algorithm is used to learn mapping relationships and neutral relationships, combining physical constraints of facial muscle movement and natural transition constraints of expression changes, expression animations are generated for interaction.
It realizes high-precision, high-reality and controllable expression generation, adapts to different individuals and expression states, and can respond to user voice commands in real time and naturally to generate realistic expression animations.
Smart Images

Figure CN120318890A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent generation technology, and particularly relates to an intelligent expression generation method and system based on facial action units. Background Art
[0002] In existing expression generation technologies, common methods include expression animation production based on key frames, expression capture-based methods, and deep learning-based expression generation. The key-frame-based method requires manual setting of expression key frames, which is labor-intensive and difficult to generate natural and smooth expression animations; although the expression capture-based method can obtain real and natural expression data, it has high equipment costs, strict requirements for the use environment, and problems such as data transmission and processing delays; although the deep learning-based expression generation method can achieve automated generation, it is often difficult to precisely control the detailed changes of expressions, and the generated expressions are insufficient in terms of realism and controllability.
[0003] Facial Action Units (FAs) are the basic units of facial muscle movement, and different FA combinations can produce a rich variety of human expressions. However, there is currently no mature technology to effectively integrate facial action units and apply them to intelligent expression generation to achieve high-precision, high-realism, and controllable expression generation effects.
[0004] Therefore, the present invention proposes an intelligent expression generation method and system based on facial action units. Summary of the Invention
[0005] The present invention provides an intelligent expression generation method and system based on facial action units to solve the above-mentioned technical problems.
[0006] The present invention provides an intelligent expression generation method based on facial action units, including: Step 1: Obtain the audio samples and multi-modal facial samples of each individual under the corresponding set text and set scenario, and determine the unit dataset of each facial action unit of the corresponding individual based on the multi-modal facial samples; Step 2: Perform phoneme encoding on the audio samples and input them into an expression predictor to obtain the prediction dataset of each facial action unit; Step 3: Establish a mapping relationship based on the unit datasets of each facial action unit of different individuals in multiple expression states, and combine with the deviation set determined based on the prediction dataset and the unit dataset to establish a neutral relationship between each facial action unit and the corresponding expression state; Step 4: Use machine learning algorithms to learn the mapping relationship and the neutral relationship to establish a facial action unit expression model; Step 5: Receive the user's voice interaction instruction, search for the required motion features and parameter settings from the established facial motion unit expression model, and perform dynamic adjustment based on the physical constraints of facial muscle movement and the natural transition constraints of expression changes to generate an expression animation for interaction.
[0007] Preferably, the unit data set includes: the motion amplitude, motion direction, and duration based on each modality.
[0008] Preferably, the multi-modal facial samples include: RGB image facial samples, depth image facial samples, and electromyogram signal facial samples.
[0009] Preferably, establish a mapping relationship based on the unit data set of each facial motion unit in multiple expression states of different individuals, including: Sequentially obtain the unit data set of the same facial motion unit in each expression state; Associate the expression state with the unit data set of the same facial motion unit as the mapping relationship.
[0010] Preferably, establish a neutral relationship between each facial motion unit and the corresponding expression state, including: Establish the spatial occupancy of the subset of the corresponding facial motion unit based on each modality in the preset coordinate system by the prediction data set and the unit data set; Based on the spatial occupancy of the prediction data set as a reference, sequentially construct a deviation vector with the spatial occupancy of each subset as the target to obtain the local expression of the facial error between the corresponding individual audio and the corresponding facial motion unit, and establish the global expression of the corresponding individual; Determine the first extreme condition related to the phoneme and the second extreme condition related to the expression state in different expression states, and combine the boundary conditions of different facial motion units to solve the reference amplitude matrix, reference direction matrix, and reference time matrix of each facial motion unit in the global matrix composed of the global expressions in the same expression state under the influence of the first extreme condition and the second extreme condition and the constraint of the boundary condition; Respectively perform simplification processing on the reference amplitude matrix, reference direction matrix, and reference time matrix to obtain the reference set of the corresponding facial motion unit and perform association as the neutral relationship.
[0011] Preferably, obtaining the reference set of the corresponding facial motion unit includes: Perform numerical clustering analysis on each reference matrix; Judge whether the corresponding reference matrix meets the preset form. If so, perform the first simplification on the corresponding reference matrix to remove redundant reference elements; Otherwise, determine the order of the corresponding reference matrix, uniformly decompose the reference matrix according to the order / 2, and perform numerical clustering analysis on each reference matrix to count the types of clusters included in each sub-matrix; If the type of cluster is 1, perform a second simplification on the sub-matrix according to a preset form; If the types of clusters are multiple, determine whether there is a cluster center falling within the sub-matrix. If not, perform a third simplification on the sub-matrix according to the mean processing method; If there is, perform a fourth simplification on the sub-matrix according to the falling position and the radiation coefficient; Determine the average value of each column in the simplified matrix to obtain a reference set.
[0012] Preferably, perform dynamic adjustment based on the physical constraints of facial muscle movement and the natural transition constraints of expression changes, including: Construct a physical model of facial muscle movement, and determine the physical constraint conditions of facial muscle movement, where the physical model of facial muscle movement is implemented based on the anatomical structure and mechanical properties of human facial muscles; Input the initial parameters obtained from the expression model into the physical model to check and optimize the parameters of each facial movement unit; According to the trend of emotional changes contained in the user's voice interaction instructions and the parameters optimized by the physical constraints, respectively use interpolation algorithms to dynamically adjust the parameters of each facial movement unit in the starting, changing, and ending stages of the expression animation; Drive the graphics rendering engine to generate an expression animation.
[0013] Preferably, after generating the expression animation for interaction, it further includes: Determine the scene attributes of the actual application scenario, where the scene attributes are formal attributes and informal attributes; Frame by frame, calculate the position difference, displacement difference, speed difference, and the difference in facial muscle deformation between the expression animation and the corresponding feature points in the real animation; Align the facial feature point data and physiological data of the real animation in time series, fuse the aligned real animation data and physiological data to form a high-dimensional feature vector, use a deep neural network model to train the fused high-dimensional feature vector, mine the micro-correlation between the physiological state and the real animation, and determine the facial expression change patterns corresponding to different heart rate intervals under the corresponding scene attributes; According to the fine-grained features obtained from the change vectors of the same feature point at consecutive moments based on the specified differences, divide the specified differences into main differences and auxiliary differences in combination with the scene attributes to construct a main feedback mechanism and an auxiliary feedback mechanism; Determine that the facial expression change pattern is based on the first influence under the main difference and the second influence under the auxiliary difference; If the first influence is greater than or equal to the second influence, construct a main feedback mechanism based on the main difference and the facial expression change pattern, and construct an auxiliary feedback mechanism based on the auxiliary difference; If the first influence is less than the second influence, construct a main feedback mechanism based on the main difference, and construct an auxiliary feedback mechanism based on the auxiliary difference and the facial expression change pattern; Based on the model parameter adjustment information calculated by the main feedback mechanism, optimize the facial movement unit expression model. If the optimization result meets the target setting, at this time, stop optimizing the model; Otherwise, continue to optimize the facial movement unit expression model using the auxiliary feedback mechanism.
[0014] The present invention provides an expression intelligent generation system based on facial movement units, including: A dataset construction module, configured to obtain audio samples and multi-modal facial samples of each individual under corresponding set texts and set scenarios, and determine the unit dataset of each facial movement unit of the corresponding individual based on the multi-modal facial samples; A prediction module, configured to perform phoneme encoding on the audio samples and input them into an expression predictor to obtain a prediction dataset of each facial movement unit; A relationship establishment module, configured to establish a mapping relationship based on the unit datasets of each facial movement unit of different individuals in multiple expression states, and combine the deviation set determined based on the prediction dataset and the unit dataset to establish a neutral relationship between each facial movement unit and the corresponding expression state; A machine learning module, configured to use machine learning algorithms to learn the mapping relationship and the neutral relationship to establish a facial movement unit expression model; An expression generation module, configured to receive a voice interaction instruction from a user, search for required motion features and parameter settings from the established facial movement unit expression model, and perform dynamic adjustment based on the physical constraints of facial muscle movement and the natural transition constraints of expression changes to generate an expression animation for interaction.
[0015] Compared with the prior art, the beneficial effects of the present application are as follows: Obtain rich multi-modal data to comprehensively describe the characteristics of an individual's facial movement units in a specific context, establish the connection between audio and facial movement units, and through phoneme encoding and an expression predictor, be able to predict facial expression changes from speech information, achieve the preliminary association between speech and expression, clarify the internal connection between facial movement units and expression states, correct the prediction error through a deviation set, make the established relationship more in line with the actual situation, improve the adaptability and accuracy of the expression model for different individuals and different expression states, construct a model that can accurately predict the states of facial movement units and expression changes, and realize real-time and natural expression animation generation and interaction based on the user's voice commands, achieving high-precision, high-fidelity, and controllable expression generation effects.
[0016] Other features and advantages of the present invention will be described in the following specification, and part of them will be obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the written specification and the accompanying drawings.
[0017] The technical solution of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings
[0018] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the accompanying drawings: Figure 1 is a flowchart of an expression intelligent generation method based on facial movement units in an embodiment of the present invention; Figure 2 is a structural diagram of an expression intelligent generation system based on facial movement units in an embodiment of the present invention. Detailed Embodiments
[0019] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to explain and illustrate the present invention, and are not used to limit the present invention.
[0020] The present invention provides an expression intelligent generation method based on facial movement units, as Figure 1 shown, including: Step 1: Obtain the audio samples and multi-modal facial samples of each individual under the corresponding set text and set scene, and determine the unit data set of each facial movement unit of the corresponding individual based on the multi-modal facial samples; Step 2: Perform phoneme encoding on the audio samples and input them into an expression predictor to obtain the prediction data set of each facial movement unit; Step 3: Establish a mapping relationship based on the unit data sets of each facial movement unit in multiple expression states of different individuals, and establish a neutral relationship between each facial movement unit and the corresponding expression state in combination with the deviation set determined based on the prediction data set and the unit data set; Step 4: Use machine learning algorithms to learn the mapping relationship and the neutral relationship, and establish a facial movement unit expression model; Step 5: Receive the voice interaction instructions of the user, search for the required motion features and parameter settings from the established facial movement unit expression model, and perform dynamic adjustment based on the physical constraints of facial muscle movement and the natural transition constraints of expression changes, and generate an expression animation for interaction.
[0021] Preferably, the unit data set includes: the movement amplitude, movement direction, and duration based on each modality.
[0022] Preferably, the multi-modal facial samples include: RGB image facial samples, depth image facial samples, and electromyogram signal facial samples.
[0023] Preferably, establishing a mapping relationship based on the unit data sets of each facial movement unit in multiple expression states of different individuals includes: Sequentially obtain the unit data sets based on the same facial movement unit in each expression state; Associate the expression state with the unit data set of the same facial movement unit as the mapping relationship.
[0024] In this embodiment, the individual is a specific person. Mainly, facial information is obtained to get facial samples. The facial movement unit expression model established in step 4 acts on the face of the robot interacting with the individual, which facilitates the robot in step 5 to make corresponding expression animations for interaction after receiving the voice interaction instructions of the user.
[0025] In this embodiment, the audio sample is generated by the individual in the corresponding scenario. For example, it is the recording of individual A1 saying "Wow, so happy" (set text) when receiving a gift in a birthday queue (set scenario).
[0026] In this embodiment, the RGB image facial sample is a facial photo taken by an ordinary color camera; the depth image facial sample can be obtained by a depth camera and can reflect the three-dimensional structure information of the face, such as the convexity of the nose; the electromyogram signal facial sample is the data collected by a sensor for the electrical activity of facial muscles, and can understand which muscles are in motion.
[0027] In this embodiment, a facial motion unit is the basic unit of facial muscle movement, such as the mouth-corner raising unit, the frowning unit, etc. Taking the mouth-corner raising unit as an example, the unit dataset includes the movement amplitude (the height of the raised mouth corner), the movement direction (moving obliquely upward), and the duration (the time the mouth corner remains raised) of the unit in the RGB image; the three-dimensional movement amplitude in the depth image; the duration of the change in the intensity of the relevant muscle electrical activity under the electromyogram signal (facial electromyogram sensors are pasted). And other relevant data based on each modality.
[0028] In this embodiment, speech processing technologies, such as the Hidden Markov Model (HMM), etc., are used to analyze audio samples and convert them into phoneme encodings. An expression predictor is pre-trained. A large number of samples with labeled phoneme encodings and corresponding facial motion unit data can be used to train a deep learning model (such as a Recurrent Neural Network RNN, a Long Short-Term Memory Network LSTM, etc.). Input the phoneme encoding into the trained expression predictor to obtain the prediction dataset of each facial motion unit. Among them, the prediction dataset is the movement amplitude, movement direction, and duration of the corresponding facial motion unit.
[0029] In this embodiment, the mapping relationship is the corresponding connection between different expression states and the facial motion unit dataset. For example, the corresponding relationship between the "happy" expression state and the relevant unit datasets such as the mouth-corner raising unit and the orbicularis oculi muscle contraction unit.
[0030] In this embodiment, the difference data set between the prediction dataset and the unit dataset is the deviation set, that is, the deviation set contains the difference subsets of each modality information under the prediction dataset and the unit dataset.
[0031] In this embodiment, the neutral relationship is the relatively accurate relationship between the facial motion unit and the expression state determined based on the mapping relationship and the deviation set.
[0032] In this embodiment, the facial motion unit expression model is constructed by learning the mapping relationship and the neutral relationship, and can predict the facial motion unit state according to the input information, and then generate the corresponding expression model.
[0033] In this embodiment, the voice interaction instruction: an instruction issued by the user through voice, such as "display a happy expression".
[0034] Motion characteristics: The relevant characteristics of the facial motion unit, such as movement amplitude, direction, speed, etc.
[0035] Parameter setting: Parameters related to expression generation, such as expression duration, transition speed, etc.
[0036] Physical constraints: Limitations formed based on the physiological characteristics of facial muscle movements, such as the upward curvature of the corners of the mouth not exceeding a certain range. Natural transition constraints: Limiting conditions to ensure natural and smooth expression changes, for example, the expression change from smiling to laughing should not be too abrupt.
[0037] Expression animation: A dynamic facial expression display generated based on the states of facial movement units, presented in the form of a video or animation sequence.
[0038] The beneficial effects of the above technical solutions are as follows: Obtain rich multi-modal data, comprehensively describe the characteristics of individual facial movement units in a specific context, establish the connection between audio and facial movement units. Through phoneme encoding and an expression predictor, it is possible to predict facial expression changes from speech information, realize the preliminary association between speech and expression, clarify the internal connection between facial movement units and expression states, correct the prediction error through a deviation set, make the established relationship more in line with the actual situation, improve the adaptability and accuracy of the expression model for different individuals and different expression states, construct a model that can accurately predict the states of facial movement units and expression changes, generate and interact with real-time and natural expression animations based on user voice commands, and achieve high-precision, high-fidelity, and controllable expression generation effects.
[0039] The present invention provides an intelligent expression generation method based on facial movement units, which establishes a neutral relationship between each facial movement unit and the corresponding expression state, including: Establish the prediction data set and the unit data set to determine the spatial occupancy of the corresponding facial movement unit based on the subsets under each modality in a preset coordinate system; Based on the spatial occupancy of the prediction data set, construct deviation vectors with the spatial occupancy of each subset as the target in turn, obtain the local expression of the facial error between the corresponding individual audio and the corresponding facial movement unit, and establish the global expression of the corresponding individual; Determine the first extreme conditions related to phonemes and the second extreme conditions related to expression states under different expression states, and combine the boundary conditions of different facial movement units to solve the reference amplitude matrix, reference direction matrix, and reference time matrix of each facial movement unit in the global matrix composed of the global expressions under the same expression state under the influence of the first extreme conditions and the second extreme conditions and the constraints of the boundary conditions; Simplify the reference amplitude matrix, reference direction matrix, and reference time matrix respectively, obtain the reference set of the corresponding facial movement unit and make an association, which is used as the neutral relationship.
[0040] In this embodiment, the preset coordinate system is a coordinate system artificially set to describe the position and direction of facial movement unit data. Similar to the coordinate system on a map, it is convenient to uniformly measure the state of facial movement units in space. For example, with the tip of the nose as the origin, the horizontal right direction as the x-axis, the vertical upward direction as the y-axis, and the direction perpendicular to the screen and outward as the z-axis to establish a coordinate system. The RGB image modality subset: the movement amplitude is 3 mm in the x-axis direction, 4 mm in the y-axis direction, and the duration is 0.5 seconds, which can be expressed as {(3, 4, 0), 0.5}; the depth image modality subset: the coordinates in the three-dimensional space are (3, 3, 0), and the duration is 0.6 seconds, that is, {(3, 3, 0), 0.6}; the electromyogram signal modality subset: the change range of the relevant muscle electromyogram signal intensity is [10, 20] mV, and the duration is 0.4 seconds, which can be expressed as {[10, 20], 0.4}.
[0041] The prediction data set is, for example: the movement amplitude is 4 mm in the x-axis direction, 4 mm in the y-axis direction, and the duration is 0.45 seconds, that is, {(4, 4, 0), 0.45}. The prediction result of the corresponding electromyogram signal is: the signal intensity ranges from [12, 22] mV, and the duration is 0.45 seconds, that is, {[12, 22], 0.45}. The above series of results are the space occupancy under the preset coordinate system, mainly the embodiment of coordinates.
[0042] In this embodiment, taking the RGB image as an example, the distance deviation between {(4, 4, 0), 0.45} and {(3, 4, 0), 0.5} is: , and the time deviation is: |0.4 - 0.5|. At this time, the deviation vector is (1, 0.05) as the local expression, and the global expression is all local expressions.
[0043] Since the global expression is the local expression in each modality, and there are several individuals and facial movement units, the global matrix composed of the global expression is: QJ = , where n represents the total number of individuals. Therefore, by setting constraints on this matrix, the reference amplitude matrix, the reference direction matrix, and the reference time matrix can be directly extracted.
[0044] In this embodiment, assuming that in the "happy" expression state, the first extreme condition related to the phoneme is that when making certain specific sounds, the maximum movement amplitude of the corner of the mouth upward unit in the x-axis direction cannot exceed 8mm; the second extreme condition related to the expression state is that when extremely happy, the maximum movement amplitude in the y-axis direction cannot exceed 6mm; the boundary condition is that the movement amplitude range in the z-axis direction is [-1,1]mm. The global matrix based on the global expression is solved by mathematical algorithms such as linear programming. For the reference amplitude matrix, the reference amplitude range of the corner of the mouth upward unit in the x-axis direction is [2,6]mm, the y-axis direction is [3,5]mm, and the z-axis direction is [-1,1]mm under the influence of extreme conditions and meeting boundary conditions, which can be expressed as: , at this time, it is assumed that the actual amplitude matrix obtained is: ,at this time, Corresponding to the x-axis, Corresponding to the y-axis, Corresponding to the z-axis, since 8 is not between 2 and 6, 8 is modified by 6 to obtain the reference amplitude matrix: .
[0045] Since the positive direction of the x, y, and z axes is 1, the negative direction is -1, and the axis is 0, the reference direction matrix is: At this time, the reference time matrix obtained is: [a1, a2], where a1 and a2 are the reference ranges of duration obtained by statistics of multiple experiments, which are obtained based on the boundary average value.
[0046] The first extreme condition: special situations related to phonemes, such as the facial motor units reaching their limit when pronouncing certain phonemes, such as the violent movement of the lip muscles when pronouncing plosive sounds.
[0047] Second extreme condition: special circumstances related to the expression state, such as the extreme contraction state of the eyebrows and muscles around the eyes under extremely angry expressions.
[0048] Boundary conditions: The restrictions on the movement of the facial motor unit itself, based on the physiological structure of the facial muscles, such as the upward range of the eyebrows cannot exceed a certain angle.
[0049] In this embodiment, the simplified matrix results are associated to obtain a reference set corresponding to the mouth corner raising unit, which is {amplitude reference: [5, 5, 0], direction reference: [1, 1, 0], time reference: 0.5}, which serves as the neutral relationship between the mouth corner raising unit and the "happy" expression state.
[0050] The beneficial effects of the above technical solution are as follows: By establishing a neutral relationship, the relationship between facial movement units and expression states can be described more accurately. Using spatial analysis and extreme condition constraints, the prediction error can be effectively corrected, enabling the expression generation model to more accurately simulate the real movement of facial movement units of different individuals in various expression states, and improving the accuracy and authenticity of expression generation.
[0051] The present invention provides an intelligent expression generation method based on facial movement units, and obtains a reference set corresponding to the facial movement units, including: Performing numerical clustering analysis on each reference matrix; Judging whether the corresponding reference matrix meets a preset form. If it meets, performing a first simplification on the corresponding reference matrix to remove redundant reference elements; Otherwise, determining the order of the corresponding reference matrix, uniformly disassembling the reference matrix according to the order / 2, performing numerical clustering analysis on each reference matrix to count the types of clusters included in each sub-matrix; If the type of cluster is 1, performing a second simplification on the sub-matrix according to the preset form; If the type of cluster is multiple, judging whether there is a cluster center falling in the sub-matrix. If not, performing a third simplification on the sub-matrix according to the mean processing method; If there is, performing a fourth simplification on the sub-matrix according to the falling position and the radiation coefficient; Determining the average value of each column in the simplified matrix to obtain the reference set.
[0052] In this embodiment, performing the fourth simplification on the sub-matrix includes: Obtaining the first column variance of the sub-matrix after removing each cluster center and obtaining the second column variance of the sub-matrix, and according to the absolute value of the difference between each first column variance and the second column variance; Successively performing discrete analysis on each column vector in the sub-matrix, counting the second number N2 of discrete points of each column vector, and at the same time, counting the first number N1 of the absolute values of the differences greater than a preset value. Wherein, when there is an absolute value of the difference greater than the preset value, performing a significant calibration on the corresponding falling position, and the significant calibration number is the first number N1. The preset value is obtained by matching from a matrix type - value comparison table, which includes direction, amplitude, and time matrix types, and the preset value corresponding to each matrix is preset, which is convenient for direct retrieval and use.
[0053] Determining whether there are the same points based on the first number N1 and the second number N2 in each column vector. If there are, taking the number of the same points as the radiation number N3, and through Multiply by the average of the absolute values of the differences from the corresponding column vectors to obtain the radiation coefficient, where N is the total number of elements present in the column vectors; Perform averaging on the corresponding columns of the sub-matrix and add or subtract the corresponding radiation coefficients to achieve the fourth simplification.
[0054] It should be noted that the judgment rule for addition or subtraction is as follows: When the number of positive differences of each first-column variance and second-column variance is greater than the number of negative differences, addition is determined; otherwise, subtraction is determined.
[0055] Obtain the degree of dispersion of the sub-matrix data by calculating the variance. The variance can reflect the fluctuation of the data around the mean value. The absolute value of the difference between the first-column variance and the second-column variance can measure the difference in the degree of dispersion of the two columns of data. Analyzing the number of discrete points (N2) and the number of absolute values of differences greater than the preset value (N1) statistically can further insight into the data distribution characteristics. Different matrix types (direction, amplitude, time matrix, etc.) have different preset thresholds because the data characteristics of different types of matrices are different, and setting according to needs makes the analysis fit various matrices and more accurately analyze the data.
[0056] Determine the same points between the first quantity N1 and the second quantity N2, and use their quantity as the radiation quantity N3, and calculate the radiation coefficient based on this. This process takes into account the correlation between different statistics, integrates information from multiple aspects, and avoids the one-sidedness of single-dimensional analysis. The radiation coefficient obtained in this way can more comprehensively reflect the data characteristics of the sub-matrix column vectors, and provide a more reasonable basis for subsequent matrix simplification.
[0057] Performing averaging on the matrix columns and combining with the radiation coefficient operation to achieve the fourth simplification is based on the data characteristics analyzed previously. Averaging can obtain the general level of the data, and the radiation coefficient incorporates characteristics such as data dispersion and significant deviation. The combination of the two can reasonably simplify the matrix while retaining important data characteristics, making the processed matrix both concise and able to reflect the key information of the original matrix.
[0058] The original matrix data is redundant and complex. Direct processing has low efficiency and is prone to masking key information. Through the above series of operations, the key features of the data can be refined, unnecessary details can be removed, the data dimension can be reduced, and the efficiency of subsequent data processing and analysis can be improved. Highlighting the data with key features helps the model better learn the rules and improve the accuracy and generalization ability of the model.
[0059] In this embodiment, a new matrix is formed according to the simplification results of each sub-matrix, and the new matrix is the simplified matrix.
[0060] In this embodiment, the disassembly scale of order / 2 can balance "local detail preservation" and "global trend extraction". If the disassembly scale is too small (such as order / 4), the sub-matrix may contain data across muscle groups (such as the junction area between the corners of the mouth and the cheeks), resulting in mixed clustering results; if the scale is too large, the sub-matrix may contain data of multiple muscle groups (such as including both the eye area and the mouth area at the same time), and the characteristics of different motor units cannot be distinguished. Assume that the residual of the original matrix is R, and the residuals of each sub-matrix after disassembly by order / 2 are R1~R4, and the total residual R_total = R1 + R2 + R3 + R4. Experiments show that when the disassembly scale is order / 2, R_total ≈ R, that is, the disassembly hardly loses information; if the scale is order / 3, R_total may increase by 15%~20% (due to the increase in clustering error caused by cross-regional data interference).
[0061] In this embodiment, for each reference matrix, a clustering algorithm (such as the K-means clustering algorithm) is used for numerical clustering analysis to classify the data in the matrix according to numerical characteristics.
[0062] Clustering center: The average value of the numerical values corresponding to the x-axis, y-axis, and z-axis respectively in the clustering result.
[0063] In this embodiment, for example, for the reference amplitude matrix, its preset form is that the number of rows of the matrix does not exceed 2 and the difference between elements does not exceed 1.5 mm. At this time, the original matrices of the reference amplitude matrix and the reference direction matrix are 3 rows, which do not meet the preset form, and the original matrix of the reference time matrix is 1 row, which meets the preset form and does not need to be simplified.
[0064] In this embodiment, assume that there is a reference amplitude matrix B which is a 4-row and 3-column matrix. At this time, it does not meet the preset form. Then, matrix B is split into 2 sub-matrices according to the result of order / 2, that is, each sub-matrix corresponds to two rows. It should be noted that the number of rows of matrix B is greater than 100. For the convenience of calculation, 4 rows are used here.
[0065] Assume the matrix , the obtained number of clusters k = 2, and the clustering center of the matrix is:[[]] , at this time, is not in , and the simplified result is:[[]] .
[0066] The second simplification is that when there is only one type of cluster in the submatrix, it is simplified according to the preset form to remove repeated or redundant information; the third simplification is that when there are multiple types of clusters in the submatrix and there is no cluster center falling in the submatrix, it is simplified according to the number of elements under each cluster type and the main information is retained; the fourth simplification is that when there are multiple types of clusters in the submatrix and there is a cluster center falling in the submatrix, it is simplified according to the location and radiation range of the cluster center to highlight the key information.
[0067] The beneficial effects of the above technical solution are: strictly following the decomposition rule of order / 2, demonstrating the difference between the two simplified logics through specific matrix operations, and obtaining a reference set through a series of analysis and simplification processing of the reference matrix, which can remove redundant and interfering information in the reference matrix and retain key features, so that the reference set can more concisely and accurately reflect the core motion characteristics of the facial motion unit, provide a more efficient and accurate reference basis for the expression generation model, and further improve the quality and efficiency of expression generation.
[0068] The present invention provides an expression intelligent generation method based on facial motion unit, which is dynamically adjusted based on the physical constraints of facial muscle movement and the natural transition constraints of expression change, including: Constructing a physical model of facial muscle movement and determining physical constraints of facial muscle movement, wherein the physical model of facial muscle movement is implemented based on the anatomical structure and mechanical properties of human facial muscles; The initial parameters obtained from the expression model are input into the physical model to check and optimize the parameters of each facial motion unit; According to the emotional change trend contained in the user's voice interaction instructions and the parameters optimized by physical constraints, the interpolation algorithm is used to dynamically adjust the parameters of each facial motion unit at the start, change and end stages of the expression animation; Drive the graphics rendering engine to generate expression animation.
[0069] In this embodiment, medical imaging technology (such as MRI, CT) is used to obtain anatomical data of human facial muscles, and information such as the starting point, end point, and shape of the muscles are recorded; biomechanical experiments are used to measure mechanical property data such as the strength and elastic coefficient of facial muscles in different contraction states. Model construction: Based on the collected data, computer graphics and mechanics-related algorithms are used to construct a physical model of facial muscle movement in a virtual environment. For example, the finite element analysis method is used to divide the facial muscles into multiple tiny units to simulate the deformation of the muscles when they are subjected to force; or a physics-based modeling method is used to describe the connection between muscles, bones, and skin using mathematical equations. Constraint determination: Combining anatomical knowledge and experimental data, analyze the extreme cases of facial muscle movement to determine physical constraints. For example, by observing a large number of real expressions, it is found that when a person has a normal expression, the degree of eye closure caused by the contraction of the muscles around the eyes will not completely block the line of sight. Based on this, set the angular constraint for the movement of the eye muscles.
[0070] Initial parameters: Parameters about facial motor units obtained from the expression model, such as data on the movement amplitude, direction, speed, duration, etc. of each facial motor unit. For example, in the "happy" expression, the initial movement amplitude of the mouth-corner raising unit is 10 mm, and the movement direction is 45 degrees obliquely upward. Inspection and optimization: Check whether the initial parameters meet the physical constraint conditions. If not, adjust and optimize them to make the parameters reasonable. For example, if the mouth-corner raising amplitude in the initial parameters is 20 mm, which exceeds the 15 mm upper limit of the physical constraint, then adjust it to 15 mm.
[0071] The starting stage is the stage when the expression begins to be generated. The changing stage is the process of the expression changing from the starting state to the target state. The ending stage is the stage when the expression reaches stability or disappears.
[0072] In expression animation, the parameters of the intermediate state are inserted between the parameters of the starting and ending states of the expression through the linear interpolation algorithm to make the expression change more natural. For example, in the starting stage, use the interpolation algorithm (such as linear interpolation) to smoothly transition from the parameters of the expressionless state to the optimized initial expression parameters; in the changing stage, according to the trend of emotional change, dynamically adjust the parameters of the facial motor units. If the emotion gradually becomes more intense, increase the parameter values such as the mouth-corner raising amplitude and the eyebrow raising angle; in the ending stage, also use the interpolation algorithm to make the expression smoothly transition from the current state to the ending state (such as returning to the expressionless state).
[0073] Graphics rendering engine: A software or hardware system used to convert virtual 3D models, animation data, etc. into visual images or animations. Common graphics rendering engines include the rendering engine of Unity, Unreal Engine, etc.
[0074] The beneficial effects of the above technical solutions are: Establish a physical model and constraint conditions of facial muscle movement that conform to human physiological characteristics, provide a scientific basis for subsequent expression animation generation, ensure that the parameters of the expression animation conform to the physical laws of human facial muscle movement by obtaining parameters from the model and checking whether the parameters meet the constraint conditions, enable the expression animation to be dynamically and naturally adjusted according to the user's emotional changes through linear interpolation, and convert the processed expression animation data into an intuitive and visible animation effect, providing users with a vivid and realistic visual experience.
[0075] The present invention provides an intelligent expression generation method based on facial motion units. After generating an expression animation for interaction, it further includes: Determine the scene attributes of the actual application scenario, where the scene attributes are formal attributes and informal attributes; Frame by frame, calculate the position difference, displacement difference, speed difference between corresponding feature points in the expression animation and the real animation, and the difference in facial muscle deformation; Align the facial feature point data and physiological data of the real animation in time series, fuse the aligned real animation data and physiological data to form a high-dimensional feature vector, and use a deep neural network model to train the fused high-dimensional feature vector to mine the micro-correlation between the physiological state and the real animation, and determine the facial expression change patterns corresponding to different heart rate intervals under the corresponding scene attributes; Based on the fine-grained features obtained from the change vectors of the same feature point at consecutive moments based on the specified difference, divide the specified difference into a main difference and an auxiliary difference according to the scene attributes, and construct a main feedback mechanism and an auxiliary feedback mechanism; Determine the first influence of the facial expression change pattern based on the main difference and the second influence based on the auxiliary difference; If the first influence is greater than or equal to the second influence, construct a main feedback mechanism based on the main difference and the facial expression change pattern, and construct an auxiliary feedback mechanism based on the auxiliary difference; If the first influence is less than the second influence, construct a main feedback mechanism based on the main difference, and construct an auxiliary feedback mechanism based on the auxiliary difference and the facial expression change pattern; Based on the model parameter adjustment information calculated by the main feedback mechanism, optimize the facial motion unit expression model. If the optimization result meets the target setting, stop optimizing the model at this time; Otherwise, continue to optimize the facial motion unit expression model using the auxiliary feedback mechanism.
[0076] In this embodiment, the scene attributes: the characteristic classification of the actual application scenario, including formal attributes (such as business meetings, speeches) and informal attributes (such as chatting, games).
[0077] In this embodiment, for example, the determined feature point differences and muscle deformation differences are as follows in the table: Table 1 Feature Point Differences, Muscle Deformation Differences
[0078] Quantify the expression similarity to provide an accurate error index for model optimization and make the subsequent training objectives more clear.
[0079] Time series alignment: Synchronize facial feature point data with physiological data (such as heart rate, electromyogram) in the time dimension.
[0080] High-dimensional feature vector: A composite feature vector formed by fusing multi-modal data, such as [position, heart rate, electromyogram].
[0081] Micro-correlation: A weak but meaningful association between physiological states and facial expressions, such as an increase in blink frequency when heart rate rises.
[0082] Use the DTW (Dynamic Time Warping) algorithm to align data and use the LSTM network to mine micro-correlations.
[0083] Input data: Facial feature point sequence:
[0084] Heart rate sequence: [72, 75, 78,..., 85] Fused feature vector:
[0085] Classify through heart rate intervals (such as 60 - 70 bpm, 70 - 80 bpm) to discover expression patterns in specific scenarios: Formal scenario: When heart rate > 80 bpm, the pupil dilation amplitude increases by 15%.
[0086] Informal scenario: When heart rate > 80 bpm, the upward speed of the corners of the mouth increases by 20%.
[0087] Analyze the contribution degree of each difference through SHAP values to divide the main / auxiliary differences: Formal scenario: Main differences (corner of the mouth position, pupil size), auxiliary difference (eyebrow height).
[0088] Informal scenario: Main differences (eyebrow dynamics, blink frequency), auxiliary difference (nasal flare).
[0089] The main - auxiliary feedback mechanism is a hierarchical optimization strategy constructed based on the importance of different differences, constructing a dynamic weight allocation mechanism. The weight of the corner of the mouth position difference in the formal scenario is 0.7, and the weight of eyebrow dynamics in the informal scenario is increased to 0.6.
[0090] Calculate the influence coefficient through the Structural Equation Model (SEM): Formal scenario: First influence = 0.82, second influence = 0.56 → The main feedback mechanism is dominated by the main differences.
[0091] Informal scenario: First influence = 0.45, second influence = 0.68 → The main feedback mechanism integrates auxiliary differences and expression patterns.
[0092] Optimization of the main feedback mechanism: Adjust the parameter of the upward curvature of the corners of the mouth to reduce the error from 1.2 mm to 0.5 mm.
[0093] Optimization of the auxiliary feedback mechanism: Incorporate the mapping relationship between heart rate and blink frequency to reduce the blink frequency error from 15% to 8%.
[0094] The effects of the model optimization iteration are shown in the table: Table 2 Effects of the model optimization iteration
[0095] Scene adaptability: In formal and informal scenarios, the naturalness scores of expressions are increased by 25% and 31% respectively.
[0096] Physiological authenticity: The correlation between heart rate and expression is increased by 35%, and the accuracy of micro-expression capture is increased by 42%.
[0097] Iteration efficiency: Through the main and auxiliary feedback mechanisms, the model convergence speed is increased by 40%, and the training time is reduced by 32%.
[0098] The beneficial effects of the above technical solutions are: adaptively adjust the feedback mechanism, incorporate the correlation rules between heart rate and eyebrow dynamics into the scene to make the expression more immersive, and significantly improve the authenticity and adaptability of the expression animation through the hierarchical optimization strategy of scene perception.
[0099] The present invention provides an intelligent expression generation system based on facial movement units, as Figure 2 shown, including: A dataset construction module, configured to obtain audio samples and multi-modal facial samples of each individual under corresponding set texts and set scenes, and determine the unit dataset of each facial movement unit of the corresponding individual based on the multi-modal facial samples; A prediction module, configured to perform phoneme encoding on the audio samples and input them into an expression predictor to obtain the prediction dataset of each facial movement unit; A relationship establishment module, configured to establish a mapping relationship based on the unit datasets of each facial movement unit of different individuals in multiple expression states, and establish a neutral relationship between each facial movement unit and the corresponding expression state in combination with the deviation set determined based on the prediction dataset and the unit dataset; A machine learning module, configured to learn the mapping relationship and the neutral relationship using a machine learning algorithm to establish a facial movement unit expression model; An expression generation module, configured to receive a voice interaction instruction from a user, search for the required motion features and parameter settings from the established facial movement unit expression model, and perform dynamic adjustment based on the physical constraints of facial muscle movement and the natural transition constraints of expression changes to generate an expression animation for interaction.
[0100] The beneficial effects of the above technical solution are as follows: obtaining rich multi-modal data, comprehensively describing the characteristics of an individual's facial motor units in a specific scenario, establishing the connection between audio and facial motor units, being able to predict facial expression changes from speech information through phoneme encoding and an expression predictor, achieving the preliminary association between speech and expression, clarifying the internal connection between facial motor units and expression states, correcting prediction errors through a deviation set, making the established relationship more in line with the actual situation, improving the adaptability and accuracy of the expression model for different individuals and different expression states, constructing a model that can accurately predict the states of facial motor units and expression changes, generating and interacting with real-time and natural expression animations based on user speech commands, and achieving a high-precision, high-fidelity and controllable expression generation effect.
[0101] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. An intelligent expression generation method based on facial movement units, characterized in that, Including: Step 1: Obtain the audio samples and multi-modal facial samples of each individual under the corresponding set text and set scenario, and determine the unit dataset of each facial action unit of the corresponding individual based on the multi-modal facial samples; Step 2: Perform phoneme encoding on the audio samples and input them into an expression predictor to obtain the prediction dataset of each facial action unit; Step 3: Establish a mapping relationship based on the unit datasets of each facial action unit of different individuals in multiple expression states, and combine the deviation set determined based on the prediction dataset and the unit dataset to establish a neutral relationship between each facial action unit and the corresponding expression state; Step 4: Use a machine learning algorithm to learn the mapping relationship and the neutral relationship to establish a facial action unit expression model; Step 5: Receive the voice interaction instruction of the user, search for the required motion features and parameter settings from the established facial action unit expression model, and perform dynamic adjustment based on the physical constraints of facial muscle movement and the natural transition constraints of expression changes to generate an expression animation for interaction.
2. The method for intelligent expression generation based on facial movement units according to claim 1, characterized in that, The unit dataset includes: the motion amplitude, motion direction, and duration based on each modality.
3. The method for intelligent expression generation based on facial movement units according to claim 1, wherein The multi-modal facial samples include: RGB image facial samples, depth image facial samples, and electromyogram signal facial samples.
4. The method for intelligent expression generation based on facial movement units according to claim 1, wherein Establishing a mapping relationship based on the unit datasets of each facial action unit of different individuals in multiple expression states includes: Sequentially obtain the unit datasets of the same facial action unit under each expression state; Associate the expression state with the unit dataset of the same facial action unit as the mapping relationship.
5. The method for intelligent expression generation based on facial movement units according to claim 1, wherein Establishing a neutral relationship between each facial action unit and the corresponding expression state includes: Establish the spatial occupancy of the subsets based on each modality of the corresponding facial action unit determined by the prediction dataset and the unit dataset in a preset coordinate system; Taking the spatial occupancy of the prediction dataset as a benchmark, sequentially construct deviation vectors with the spatial occupancy of each subset as the target to obtain the local expression of the facial error between the corresponding individual's audio and the corresponding facial action unit, and establish the global expression of the corresponding individual; Determine the first extreme condition related to phonemes and the second extreme condition related to the expression state under different expression states, and combine the boundary conditions of different facial action units to solve the reference amplitude matrix, reference direction matrix, and reference time matrix of each facial action unit in the global matrix composed of the global expressions under the same expression state under the influence of the first extreme condition and the second extreme condition and the constraint of the boundary condition; Perform simplification processing on the reference amplitude matrix, reference direction matrix, and reference time matrix respectively to obtain the reference set of the corresponding facial action unit and perform association as the neutral relationship.
6. The method for intelligent expression generation based on facial movement units according to claim 5, wherein Obtaining the reference set of the corresponding facial action unit includes: Perform numerical clustering analysis on each reference matrix; Judge whether the corresponding reference matrix meets the preset form. If it meets, perform the first simplification on the corresponding reference matrix to remove redundant reference elements; Otherwise, determine the order of the corresponding reference matrix, uniformly decompose the reference matrix according to the order / 2, and perform numerical clustering analysis on each reference matrix to count the types of clusters included in each sub-matrix; If the type of cluster is one, perform a second simplification on the sub-matrix according to a preset form; If there are multiple types of clusters, determine whether there is a cluster center falling within the sub-matrix. If not, perform a third simplification on the sub-matrix according to the mean processing method; If so, perform a fourth simplification on the sub-matrix according to the falling position and the radiation coefficient; Determine the average value of each column in the simplified matrix to obtain a reference set.
7. The method for intelligent expression generation based on facial movement units according to claim 1, characterized in that Perform dynamic adjustment based on the physical constraints of facial muscle movement and the natural transition constraints of expression changes, including: Construct a physical model of facial muscle movement, and determine the physical constraint conditions of facial muscle movement, where the physical model of facial muscle movement is implemented based on the anatomical structure and mechanical properties of human facial muscles; Input the initial parameters obtained from the expression model into the physical model to check and optimize the parameters of each facial movement unit; According to the emotional change trend contained in the user's voice interaction instruction and the parameters optimized by the physical constraints, for the starting, changing, and ending stages of the expression animation, respectively use interpolation algorithms to dynamically adjust the parameters of each facial movement unit; Drive the graphics rendering engine to generate an expression animation.
8. The method for intelligent expression generation based on facial movement units according to claim 1, wherein After generating the expression animation for interaction, it also includes: Determine the scene attributes of the actual application scenario, where the scene attributes are formal attributes and informal attributes; Frame by frame, calculate the position difference, displacement difference, speed difference, and facial muscle deformation difference between the corresponding feature points in the expression animation and the real animation; Align the facial feature point data and physiological data of the real animation in time series, fuse the aligned real animation data and physiological data to form a high-dimensional feature vector, use a deep neural network model to train the fused high-dimensional feature vector, mine the micro-correlation between the physiological state and the real animation, and determine the facial expression change patterns corresponding to different heart rate intervals under the corresponding scene attributes; According to the fine-grained features obtained from the change vectors of the same feature point based on the specified differences at consecutive moments, combine the scene attributes to divide the specified differences into main differences and auxiliary differences, and construct a main feedback mechanism and an auxiliary feedback mechanism; Determine the first influence of the facial expression change pattern based on the main difference and the second influence based on the auxiliary difference; If the first influence is greater than or equal to the second influence, construct a main feedback mechanism based on the main difference and the facial expression change pattern, and construct an auxiliary feedback mechanism based on the auxiliary difference; If the first influence is less than the second influence, construct a main feedback mechanism based on the main difference, and construct an auxiliary feedback mechanism based on the auxiliary difference and the facial expression change pattern; Based on the model parameter adjustment information calculated by the main feedback mechanism, optimize the facial movement unit expression model. If the optimization result meets the target setting, stop optimizing the model at this time; Otherwise, continue to optimize the facial movement unit expression model using the auxiliary feedback mechanism.
9. An intelligent expression generation system based on facial movement units, characterized in that, Including: The dataset construction module is used to obtain the audio samples and multi-modal facial samples of each individual under the corresponding set text and set scenario, and determine the unit dataset of each facial action unit of the corresponding individual based on the multi-modal facial samples; The prediction module is used to perform phoneme encoding on the audio samples and input them into the expression predictor to obtain the prediction dataset of each facial action unit; The relationship establishment module is used to establish a mapping relationship based on the unit datasets of each facial action unit of different individuals in multiple expression states, and combine the deviation set determined based on the prediction dataset and the unit dataset to establish a neutral relationship between each facial action unit and the corresponding expression state; The machine learning module is used to learn the mapping relationship and the neutral relationship using machine learning algorithms to establish a facial action unit expression model; The expression generation module is used to receive the voice interaction instruction of the user, search for the required motion features and parameter settings from the established facial action unit expression model, and perform dynamic adjustment based on the physical constraints of facial muscle movement and the natural transition constraints of expression changes to generate an expression animation for interaction.
Citation Information
Patent Citations
Establishment method of pig face facial expression recognition framework based on multi-task cascade
CN113065460A
Robot emotion recognition method and system based on AIGC and storage medium
CN118626966A
Voice-driven three-dimensional face animation generation method and device based on reinforcement learning
CN119027557A
Virtual image expression generation method and system based on real feeling technology
CN119295683A
Face image clustering method and system based on localized simple multiple kernel k-means
US20240331351A1
Cited By
Multi-modal expression generation system and dynamic optimization method in virtual-real fusion scene
CN121999100A