A real-time closed-loop compensation method for facial expression of a bionic human head

By employing a single-model architecture based on self-supervised learning and Jacobian matrix analytical compensation, the problems of high resource consumption and response latency in the facial expression control of bionic robots are solved, achieving real-time closed-loop compensation and efficient correction.

CN122110754BActive Publication Date: 2026-07-07SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUTH CHINA UNIV OF TECH
Filing Date
2026-04-30
Publication Date
2026-07-07

AI Technical Summary

Technical Problem

Existing methods for controlling the facial expressions of bionic robots rely on the coupled operation of two models, which consumes a lot of computational resources, has a delayed correction process and insufficient response capability, and cannot compensate for sudden errors in real time.

Method used

A single inverse model is constructed using self-supervised learning. Instantaneous analytical compensation of facial expression residuals is achieved through the Jacobian matrix. Combined with error accumulation monitoring and model reconstruction, real-time closed-loop compensation is realized.

Benefits of technology

It reduces computing resource requirements, improves response speed and control precision, and ensures long-term stability and low latency in transient response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122110754B_ABST
    Figure CN122110754B_ABST
Patent Text Reader

Abstract

The application discloses a face expression real-time closed-loop compensation method for a bionic human head, constructs a forward model through self-supervised learning and constrains a reverse model, only uses the reverse model to output a reference angle vector during operation, simultaneously extracts a Jacobian matrix in real time by using automatic differentiation, directly analyzes instantaneous expression residuals of visual feedback into actuator angle compensation amounts, generates driving instructions after superposition, realizes real-time closed-loop compensation of the actuator based on the Jacobian matrix, and triggers collaborative reconstruction of a mapping model when the residuals continuously do not decrease, and adapts to long-term physical drift. The application reduces calculation overhead by single model operation, and compensation calculation is only once matrix multiplication, so that response is rapid, and real-time accuracy and long-term stable control of face expression of the bionic human head can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of bionic robot motion control and deep learning, and in particular to a real-time closed-loop compensation method for facial expressions in a bionic human head. Background Technology

[0002] Precise control of facial expressions in bionic robots is one of the key technologies for achieving highly natural human-computer interaction. Traditional control methods typically employ an "offline calibration + open-loop control" strategy, which involves pre-establishing a mapping table between actuator angles and facial expression parameters (BlendShape). During runtime, actuator commands are obtained by looking up the table or interpolating the target facial expression parameters (BlendShape). However, this method heavily relies on the precision of mechanical assembly and static assumptions of the physical model. It cannot cope with time-varying nonlinear factors that occur during long-term operation, such as nonlinear deformation of silicone skin, servo motor hysteresis, and environmental temperature drift. This results in a significant decrease in expression execution accuracy over time.

[0003] To address the aforementioned issues, existing technologies have begun to incorporate visual feedback and data-driven methods. For example, the patent application by Li Bo et al., "A Bionic Expression Control System and Method Based on 3D Self-Supervised Learning," discloses a method for simultaneously training a forward mapping network and a backward mapping network, and jointly fine-tuning the two networks when error accumulation is detected. Although the above methods represent a significant improvement over traditional open-loop control, the following technical shortcomings still exist in practical applications:

[0004] 1. Existing technologies require the simultaneous maintenance of two networks, a forward model and a reverse model, both of which participate in real-time computation, consuming twice the computing resources and storage space, thus placing a significant burden on the platform.

[0005] 2. Existing solutions rely on mechanisms such as "error accumulation triggering model retraining" or "gradual iterative optimization". The correction process requires asynchronous triggering of model weight fine-tuning or multiple iterative calculations, rather than real-time response within the control period, which causes the system to be in a non-ideal state for most of the running time.

[0006] 3. Existing technologies essentially only contain a slow correction loop based on model retraining, lacking a fast loop that can compensate residuals in real time within the inner control cycle, resulting in insufficient system response to sudden and short-term errors. Summary of the Invention

[0007] The purpose of this invention is to address the problems of decreased control accuracy and delayed correction in existing bionic robot facial expression control due to skin material aging, mechanical lag, and environmental interference. To address the high latency issues caused by existing technologies relying on coupled "forward-inverse" dual models and requiring fine-tuning of model weights or iterative optimization for correction, this invention proposes a real-time closed-loop compensation method for facial expressions in bionic heads. This method utilizes a single inverse model derived from self-supervised learning and a novel control architecture that achieves instantaneous analytical compensation of facial expression residuals by extracting its internal Jacobian matrix in real time. The aim is to achieve real-time closed-loop compensation for facial expressions in bionic heads without modifying model parameters or performing iterative optimization. Simultaneously, through error accumulation monitoring, forward model recalibration, and inverse model reconstruction, adaptive correction and reconstruction of the mapping model are achieved.

[0008] To achieve the above objectives, the technical solution provided by this invention is: a real-time closed-loop compensation method for facial expressions in a bionic head, wherein the surface of the bionic head is covered with bionic skin, and its facial expressions are controlled by multiple actuators. The Blendshape coefficient vector of the facial expressions changes with the change in the angle vector of the actuators. The method includes the following steps:

[0009] S1: Construct and train a forward model for mapping the angle vector of the executor to the Blendshape coefficient vector of the facial expression and an inverse model for mapping the Blendshape coefficient vector of the facial expression to the angle vector of the executor.

[0010] S2: In each control cycle of the actuator, the Blendshape coefficient vector corresponding to the target expression is input into the trained inverse model to obtain the angle vector corresponding to the actuator as the reference angle vector, and the instantaneous gradient of the inverse model at the Blendshape coefficient vector of the target expression. Then, a Jacobian matrix is ​​constructed to characterize the local linear mapping relationship between the small changes in facial expression and the changes in the actuator vector space.

[0011] S3: Real-time acquisition of the facial expressions of the bionic human head and their corresponding Blendshape coefficient vectors; vector difference between the Blendshape coefficient vector corresponding to the target expression and the Blendshape coefficient vector of the current actual expression to obtain the current instantaneous expression residual;

[0012] S4: Multiply the Jacobian matrix with the instantaneous facial expression residual to obtain the instantaneous angle compensation amount of the actuator;

[0013] S5: Linearly superimpose the reference angle vector and the instantaneous angle compensation amount to generate the final drive command and send it to the actuator to complete the execution and correction of the reference action within the adjacent control cycle;

[0014] S6: Continuously monitor the instantaneous facial expression residual within each control cycle of the actuator. If it is detected that the instantaneous facial expression residual has not decreased within multiple consecutive control cycles, the data is expanded using the Blendshape coefficient vector of the bionic human head's facial expression over historical time and the actual angle vector of the actuator. The forward model is then corrected, and the inverse model is reconstructed using the corrected forward model. The mapping relationship between the Blendshape coefficient vector of the facial expression and the angle vector of the actuator is updated using the reconstructed inverse model, thereby achieving real-time closed-loop accurate compensation for facial expressions.

[0015] Furthermore, the specific steps of step S1 are as follows:

[0016] S1.1: Perform self-supervised data acquisition, i.e., control all actuators to perform random exploratory motion. Through randomly generated or preset excitation signal sequences, drive each actuator to independently or collaboratively traverse its entire motion range, including different combinations of amplitude, frequency, and phase. During the motion, simultaneously acquire two data streams at a fixed sampling frequency as training datasets: record the real-time angle values ​​of all actuators to obtain the angle vectors of the actuators. ,in For the set of real numbers, A positive integer, representing the number of actuators in the bionic head. This represents the actuator control space; the facial expressions of the bionic head and their corresponding BlendShape coefficient vectors are synchronously acquired through a vision sensor. ,in A positive integer, representing the semantic feature dimension of the facial expression state. The training dataset represents the facial expression semantic space; it covers the complete mapping relationship between the actuator control space and the facial expression semantic space.

[0017] S1.2: Based on the above training dataset, a multilayer perceptron is trained as a forward model. This forward model contains several fully connected operation layers to realize high-dimensional linear mapping of features between the facial expression semantic space and the actuator control space. Each fully connected operation layer is followed by a non-linear activation function layer to accurately fit the non-linear physical deformation law of bionic skin under different tensions. The forward model takes the angle vector of the actuator as input and outputs the BlendShape coefficient vector of the predicted expression. During training, the mean square error between the BlendShape coefficient vector of the predicted expression and the BlendShape coefficient vector measured by the visual sensor is minimized to determine the associated weight parameters inside the forward model, thus determining how the angle change of the actuator is non-linearly mapped to the final facial BlendShape coefficient.

[0018] S1.3: After the forward model training converges and the parameters are fixed, a multilayer perceptron is constructed as the inverse model. The input of the inverse model is the BlendShape coefficient vector, and the output is the angle vector. The training of the inverse model is divided into two stages. In the first stage, the input and output are directly swapped in supervised training using the above training dataset, so that the inverse model initially establishes the mapping relationship from the facial expression semantic space to the actuator control space. In the second stage, a physical consistency constraint based on the forward model is introduced. First, the BlendShape coefficient vector corresponding to the target expression is input into the inverse model to obtain the corresponding predicted angle vector. Then, the predicted angle vector is input into the fixed forward model in real time to obtain the BlendShape coefficient vector of the restored expression. By calculating the cycle consistency loss between the BlendShape coefficient vector corresponding to the target expression and the BlendShape coefficient vector of the restored expression, the weights of the inverse model are finely tuned, and finally a trained inverse model is obtained.

[0019] Furthermore, the specific steps of step S2 are as follows:

[0020] S2.1: Denote the Blendshape coefficient vector corresponding to the target expression as... In each control cycle of the actuator, The input is fed into the trained inverse model, and the output angle vector serves as the reference angle vector for subsequent compensation, denoted as... ;

[0021] S2.2: Without updating the weight parameters of the inverse model, the automatic differentiation interface of the inverse model is called to calculate the instantaneous gradient of the output of the inverse model with respect to the input. This calculation follows the chain rule, propagating backward from the output layer to the input layer, multiplying the local gradients of each layer, and finally obtaining a gradient of size [value missing]. The Jacobian matrix is ​​denoted as The Jacobian matrix is ​​the partial derivative matrix of the actuator's angle vector with respect to the blendshape coefficient vector of the facial expression. It is used to characterize the local linear mapping relationship between small changes in facial expression and changes in the actuator vector space.

[0022] ;

[0023] The specific form of this Jacobian matrix is ​​as follows:

[0024] ;

[0025] In the formula, Indicates the first An angle of the actuator Indicates the first There are BlendShape coefficients, among which ; matrix elements The physical meaning is: at the current working point, i.e. Nearby, when the When the first Blendshape coefficient changes slightly, the... The sensitivity coefficient needs to be adjusted accordingly for the angle of each actuator.

[0026] Furthermore, in step S3, a visual capture device is used to acquire the facial expressions of the bionic head in real time at a fixed frame rate, and the Blendshape coefficient vector of the current actual expression is extracted, denoted as... The instantaneous expression residual for the current frame is obtained by subtracting the Blendshape coefficient vector corresponding to the target expression from the Blendshape coefficient vector of the current actual expression. .

[0027] Furthermore, in step S4, the Jacobian matrix is ​​used. As a linear mapping operator from the expression residual space to the actuator control space; multiplying the instantaneous expression residual with the Jacobian matrix directly yields the instantaneous angle compensation of the actuator, denoted as . ,Right now .

[0028] Furthermore, in step S5, the reference angle vector and the instantaneous angle compensation amount are linearly superimposed to generate the final drive command, denoted as... ,Right now The instruction is then sent to each actuator to complete the execution and correction of the baseline action within the adjacent control cycle.

[0029] Furthermore, the specific steps of step S6 are as follows:

[0030] S6.1: Continuously monitor the instantaneous expression residual in each control cycle of the actuator. If the instantaneous expression residual does not decrease in multiple consecutive control cycles, it is determined that the forward model has deviated from the current physical entity. At this time, the reverse model is reconstructed.

[0031] S6.2: Data augmentation is performed using the Blendshape coefficient vectors of facial expressions of the bionic head over historical time periods and the actual angle vectors of the actuators to obtain an augmented training dataset. An offline fine-tuning method is used to perform a small number of iterations on the weights of the original forward model, so that the Blendshape coefficient vectors corresponding to the predicted expressions output by the model approximate the measured Blendshape coefficient vectors again. The weight parameters of the forward model are then corrected so that the predicted expressions output by the model are now in line with the physical entity representation of the current bionic head, resulting in an updated forward model.

[0032] S6.3: Using the updated forward model, the inverse model is reconstructed according to the training process described in step S1.3: In the first stage, supervised training with input and output swapping is performed using the expanded training dataset to initially establish the mapping relationship between the facial expression semantic space and the actuator control space of the inverse model; In the second stage, physical consistency constraints based on the updated forward model are introduced. First, the Blendshape coefficient vector corresponding to the target expression is input into the inverse model to obtain the corresponding predicted angle vector. Then, the predicted angle vector is input into the updated forward model in real time to obtain the Blendshape coefficient vector of the restored expression. By calculating the cycle consistency loss between the Blendshape coefficient vector corresponding to the target expression and the Blendshape coefficient vector of the restored expression, the weights of the inverse model are finely adjusted to obtain a reconstructed inverse model. The subsequent real-time accurate compensation of facial expressions will be based on this reconstructed inverse model. During the reconstruction process, the old inverse model can be temporarily used without affecting the continuous operation of real-time accurate compensation of the actuator controlling facial expressions.

[0033] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0034] 1. This invention ensures the physical consistency of the control logic by constraining the design of the inverse model through the forward model; and in the real-time operation phase, only the inverse model needs to participate in the work. Through the integrated solution of "static prediction" and "dynamic compensation of derivative", the demand for computing resources and storage space is reduced without sacrificing accuracy, thus realizing the lightweight control architecture.

[0035] 2. This invention decentralizes the complex residual correction logic from the traditional "parameter iteration layer" to the "control quantity linear compensation layer." The process of extracting the Jacobian matrix requires only one backpropagation operation, and subsequent compensation involves only a single matrix multiplication. This analytical computational logic achieves synchronous feedback, accelerating the response speed of residual compensation.

[0036] 3. Unlike the iterative method of "gradually optimizing the servo control signal" in existing technologies, this invention adopts additive compensation, which can complete the compensation and correction of the actuator angle in a single step without performing multiple forward calculations of the model. This greatly improves the following speed and naturalness of the bionic head under dynamic expression switching.

[0037] 4. This invention constructs a closed-loop control architecture consisting of "real-time compensation based on the Jacobian matrix" and "model collaborative reconstruction." Within each control cycle, the Jacobian matrix is ​​used to perform analytical compensation on the facial expression residuals, achieving real-time compensation correction. Error accumulation monitoring triggers forward model fine-tuning and inverse model reconstruction, solving the long-term physical drift problem. This approach ensures both low latency in transient response and long-term operational stability, achieving precise control across the entire time domain. Attached Figure Description

[0038] Figure 1 This is a flowchart illustrating the overall process framework of the method of the present invention.

[0039] Figure 2 This is a schematic diagram illustrating the real-time compensation principle based on the Jacobian matrix of the present invention.

[0040] Figure 3 This is a flowchart illustrating the collaborative reconstruction process of the mapping model (i.e., the forward model and the reverse model) in the method of this invention. Detailed Implementation

[0041] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto.

[0042] This embodiment discloses a real-time closed-loop compensation method for facial expressions in a bionic head. The experimental platform used is a 33-DOF bionic head prototype, the specific structure of which is as follows:

[0043] 1. Mechanical structure: The head skeleton adopts a composite structure of metal and resin 3D printed parts, and the face is covered with biomimetic silicone skin. It has 33 independent degrees of freedom of movement, covering the main facial movement areas such as eyelids, eyeballs, eyebrows, lips, jaw, and neck.

[0044] 2. Actuator Unit: Each degree of freedom is driven independently by a servo motor. The 33 servo motors are controlled by three PCA9685 servo driver boards (each driver board provides 16 PWM outputs, the three boards can drive a total of 48 channels, but 33 channels are actually used). The Raspberry Pi communicates with the three PCA9685 boards via the I2C bus, sending PWM pulse width commands to each board, which then generates the corresponding PWM signals to drive the servos.

[0045] 3. Visual perception unit: A monocular camera is fixedly installed in front of the bionic human head and connected to a computer via a USB interface to collect the robot's facial expressions in real time.

[0046] 4. Computing and Control Unit: A computer is used as the main controller, running a Linux operating system, deploying a Python environment, a deep learning framework, and an OpenCV vision processing library. It also reads images captured by a monocular camera via a USB interface. The Raspberry Pi controls three PCA9685 driver boards via an I2C bus to drive the servos.

[0047] like Figures 1 to 3 As shown, the specific implementation of the real-time closed-loop compensation method for facial expressions includes the following steps:

[0048] S1: Collect the angle vectors of all servos and their corresponding blendshape coefficient vectors of facial expressions through self-supervised learning to form a dataset. Train a multilayer perceptron as a forward model based on the training dataset to map the servo angle vectors to the blendshape coefficient vectors of facial expressions. Based on the physical consistency constraints provided by the training dataset and the trained forward model, train a multilayer perceptron as an inverse model to map the blendshape coefficient vectors of facial expressions to the servo angle vectors. The specific details are as follows:

[0049] S1.1: First, it is necessary to establish the mapping relationship between the bionic human head / face and the expression coefficient from the servo motor angle. The specific operation is as follows:

[0050] S1.1.1: Write a Python program on a Raspberry Pi to control the servo motors. Send PWM commands to three PCA9685 driver boards via the I2C bus. Each PCA9685 board has a different address (e.g., 0x40, 0x41, 0x42). The program configures and sends commands to the three boards sequentially. The script also designs a random motion generator capable of generating random angle sequences covering the full range of the servo motors, including:

[0051] a) Single servo sweep motion: Each servo moves independently, with the amplitude covering the main range of its stroke;

[0052] b. Multi-servo combination motion: Randomly select multiple servos to perform combined motion, simulating the coordinated mode of natural facial expressions;

[0053] c. Resting state: After the motion sequence ends, collect static data when the servo motor returns to center;

[0054] S1.1.2: Record the real-time angle values ​​of all servos to obtain the servo angle vectors. ,in For the set of real numbers, The value is 33, indicating the number of servo motors in the bionic head. Represents the servo control space; synchronously acquires facial expressions of the bionic head and their corresponding BlendShape coefficient vectors via a monocular camera. ,in The value is 52, representing the semantic feature dimension of facial expression states. Representing the semantic space of facial expressions;

[0055] S1.1.3: Store the above data in a computer as a training dataset; the entire process runs continuously for a certain period of time, covering the main motion space of the bionic human head and face, to ensure that sufficient training data is obtained, and the training dataset covers the complete mapping relationship between the servo control space and the facial expression semantic space.

[0056] S1.2: After data acquisition, a multilayer perceptron is trained on a computer based on the above training dataset as a forward model. This forward model contains several fully connected operation layers to realize high-dimensional feature linear mapping between the facial expression semantic space and the servo control space. Each fully connected operation layer is followed by a non-linear activation function layer to accurately fit the non-linear physical deformation law of bionic skin under different tensions. The forward model takes the 33-dimensional angle vector of the servo as input and outputs the 52-dimensional BlendShape coefficient vector of the predicted expression. During training, the mean square error between the BlendShape coefficient vector of the predicted expression and the BlendShape coefficient vector measured by the monocular camera is minimized to determine the correlation weight parameters inside the forward model, thus determining how the angle change of the servo is non-linearly mapped to the final facial BlendShape coefficient.

[0057] S1.3: After the forward model training converges and the parameters are fixed, a multilayer perceptron is constructed as the inverse model. The input of the inverse model is a 52-dimensional BlendShape coefficient vector, and the output is a 33-dimensional angle vector. The training of the inverse model is divided into two stages. In the first stage, the input and output are directly swapped for supervised training using the training dataset, so that the inverse model initially establishes the mapping relationship from the facial expression semantic space to the servo control space. In the second stage, a physical consistency constraint based on the forward model is introduced. First, the BlendShape coefficient vector corresponding to the target expression is input into the inverse model to obtain the corresponding predicted angle vector. Then, the predicted angle vector is input into the fixed forward model in real time to obtain the BlendShape coefficient vector of the restored expression. By calculating the cycle consistency loss between the BlendShape coefficient vector corresponding to the target expression and the BlendShape coefficient vector of the restored expression, the weights of the inverse model are finely tuned, and finally a trained inverse model is obtained.

[0058] S2: In each control cycle of the servo motor, the blendshape coefficient vector corresponding to the target facial expression is input into the trained inverse model. After forward computation, the angle vector corresponding to the servo motor is obtained as the reference angle vector. Simultaneously, the instantaneous gradient of the inverse model at the blendshape coefficient vector of the target facial expression is calculated using the automatic differentiation interface of the inverse model, thereby constructing a Jacobian matrix. This Jacobian matrix is ​​the partial derivative matrix of the servo motor angle vector with respect to the blendshape coefficient vector of the facial expression, used to characterize the local linear mapping relationship between small changes in facial expression and changes in the servo motor vector space. The specific details are as follows:

[0059] S2.1: Denote the Blendshape coefficient vector corresponding to the target expression as... In each control cycle of the servo motor, Input to the reverse model;

[0060] S2.2: Will The angle vector output after inputting into the inverse model is used as the reference angle vector for subsequent compensation, and this reference angle vector is denoted as... ;

[0061] S2.3: Without updating the weight parameters of the inverse model, the automatic differentiation interface of the inverse model is called to calculate the instantaneous gradient of the output of the inverse model with respect to the input. This calculation follows the chain rule, propagating backward from the output layer to the input layer, multiplying the local gradients of each layer, and finally obtaining a Jacobian matrix of size 33×52, denoted as . The Jacobian matrix is ​​the partial derivative matrix of the servo angle vector with respect to the blendshape coefficient vector of the facial expression, used to characterize the local linear mapping relationship between small changes in facial expression and changes in the servo vector space:

[0062] ;

[0063] The specific form of this Jacobian matrix is ​​as follows:

[0064] ;

[0065] In the formula, Indicates the first The angle of each servo motor Indicates the first There are BlendShape coefficients, among which ; matrix elements The physical meaning is: at the current working point, i.e. Nearby, when the When the first Blendshape coefficient changes slightly, the... The sensitivity coefficient needs to be adjusted accordingly for the angle of each servo motor.

[0066] S3: Utilize a monocular camera to capture facial expressions of the bionic head in real-time at a fixed frame rate and extract the Blendshape coefficient vector of the current actual expression, denoted as... The instantaneous expression residual for the current frame is obtained by subtracting the Blendshape coefficient vector corresponding to the target expression from the Blendshape coefficient vector of the current actual expression. .

[0067] S4: Using the Jacobian matrix obtained in step S2 As a linear mapping operator from the expression residual space to the servo control space; the instantaneous expression residual obtained in step S3 is multiplied by the Jacobian matrix to directly obtain the instantaneous angle compensation of the servo, denoted as . ,Right now This calculation involves only one matrix multiplication and can be completed within a single control cycle.

[0068] S5: Linearly superimpose the reference angle generated by the inverse model in step S2 with the instantaneous compensation amount to generate the final drive command, denoted as... ,Right now ;Will The 33 angle values ​​are converted into corresponding PWM pulse width values, which are then written to the corresponding channel registers of the three PCA9685 driver boards via the I2C bus to generate PWM signals with corresponding duty cycles, driving the connected servos to perform actions and complete the execution and correction of the reference action.

[0069] The above steps S2 to S5 are repeated in each control cycle to achieve real-time compensation of the servo motor based on the Jacobian matrix, thereby correcting the facial expressions of the bionic head.

[0070] S6: Continuously monitor the instantaneous facial expression residual within each control cycle of the servo motor. If the instantaneous facial expression residual does not decrease within multiple consecutive control cycles, data augmentation is performed using the Blendshape coefficient vector of the bionic human head's facial expression over historical time periods and the actual angle vector of the servo motor. This data is then used to correct the forward model. The corrected forward model is then used to reconstruct the inverse model, and the reconstructed inverse model is used to update the mapping relationship between the Blendshape coefficient vector of the facial expression and the angle vector of the servo motor, achieving real-time closed-loop precise compensation for facial expressions. The specific details are as follows:

[0071] S6.1: Continuously monitor the instantaneous facial expression residuals in each control cycle of the servo motor; after the bionic head has been running continuously for a long time, due to the decline in elasticity of the silicone skin or wear of the servo motor transmission, even if the closed-loop real-time accurate compensation in steps S2 to S5 takes effect, there may be a situation where the instantaneous facial expression residuals have not decreased in multiple consecutive control cycles. At this time, it is determined that the forward model has deviated from the current physical entity, and the reverse model is reconstructed.

[0072] S6.2: The data is augmented using the Blendshape coefficient vector of the human face expression of the bionic head in historical time and the actual angle vector of the servo motor. The augmented training dataset is obtained. The offline fine-tuning method is used to perform a small number of iterations on the weights of the original forward model so that the Blendshape coefficient vector corresponding to the predicted expression output is close to the measured Blendshape coefficient vector. The weight parameters of the forward model are corrected so that the predicted expression output is aligned with the physical entity of the current bionic head, and the updated forward model is obtained.

[0073] S6.3: Reconstruct the inverse model using the updated forward model following the training process described in step S1.3: In the first stage, supervised training with input and output swapping is performed using the expanded training dataset from step S6.2, enabling the inverse model to initially establish a mapping relationship between the facial expression semantic space and the servo control space. In the second stage, a physical consistency constraint based on the updated forward model is introduced. First, the Blendshape coefficient vector corresponding to the target expression is input into the inverse model to obtain the corresponding predicted angle vector. Then, the predicted angle vector is input into the updated forward model in real time to obtain the Blendshape coefficient vector of the reconstructed expression. The dshape coefficient vector is used to fine-tune the weights of the inverse model by calculating the cycle consistency loss between the Blendshape coefficient vector corresponding to the target expression and the Blendshape coefficient vector of the reconstructed expression, resulting in a reconstructed inverse model. Subsequent real-time accurate compensation for facial expressions will be based on this reconstructed inverse model. During the reconstruction process, the old inverse model is temporarily used without affecting the continued operation of closed-loop real-time accurate compensation for facial expressions, thus achieving adaptive correction and reconstruction of the mapping model (i.e., the forward model and the inverse model) of the bionic human head and facial expressions under long-term operation.

[0074] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.

Claims

1. A real-time closed-loop compensation method for facial expressions in a bionic head, wherein the surface of the bionic head is covered with bionic skin, and its facial expressions are controlled by multiple actuators, wherein changes in the angle vector of the actuators will change the blendshape coefficient vector of the facial expressions, characterized in that... Includes the following steps: S1: Construct and train a forward model for mapping the angle vector of the executor to the Blendshape coefficient vector of the facial expression and an inverse model for mapping the Blendshape coefficient vector of the facial expression to the angle vector of the executor. S2: In each control cycle of the actuator, the Blendshape coefficient vector corresponding to the target expression is input into the trained inverse model to obtain the angle vector corresponding to the actuator as the reference angle vector, and the instantaneous gradient of the inverse model at the Blendshape coefficient vector of the target expression. Then, a Jacobian matrix is ​​constructed to represent the local linear mapping relationship between small changes in facial expression and changes in the actuator vector space. The specific operation steps are as follows: S2.1: Denote the Blendshape coefficient vector corresponding to the target expression as... In each control cycle of the actuator, The input is fed into the trained inverse model, and the output angle vector serves as the reference angle vector for subsequent compensation, denoted as... ; S2.2: Without updating the weight parameters of the inverse model, the automatic differentiation interface of the inverse model is called to calculate the instantaneous gradient of the output of the inverse model with respect to the input. This calculation follows the chain rule, propagating backward from the output layer to the input layer, multiplying the local gradients of each layer, and finally obtaining a gradient of size [value missing]. The Jacobian matrix is ​​denoted as The Jacobian matrix is ​​the partial derivative matrix of the actuator's angle vector with respect to the blendshape coefficient vector of the facial expression. It is used to characterize the local linear mapping relationship between small changes in facial expression and changes in the actuator vector space. ; The specific form of this Jacobian matrix is ​​as follows: ; In the formula, Indicates the first An angle of the actuator Indicates the first There are BlendShape coefficients, among which ; matrix elements The physical meaning is: at the current working point, i.e. Nearby, when the When the first Blendshape coefficient changes slightly, the... The sensitivity coefficient of each actuator needs to be adjusted accordingly at its angle; S3: Real-time acquisition of the facial expressions of the bionic human head and their corresponding Blendshape coefficient vectors; vector difference between the Blendshape coefficient vector corresponding to the target expression and the Blendshape coefficient vector of the current actual expression to obtain the current instantaneous expression residual; S4: Multiply the Jacobian matrix with the instantaneous facial expression residual to obtain the instantaneous angle compensation amount of the actuator; S5: Linearly superimpose the reference angle vector and the instantaneous angle compensation amount to generate the final drive command and send it to the actuator to complete the execution and correction of the reference action within the adjacent control cycle; S6: Continuously monitor the instantaneous facial expression residual within each control cycle of the actuator. If it is detected that the instantaneous facial expression residual has not decreased within multiple consecutive control cycles, the data is expanded using the Blendshape coefficient vector of the bionic human head's facial expression over historical time and the actual angle vector of the actuator. The forward model is then corrected, and the inverse model is reconstructed using the corrected forward model. The mapping relationship between the Blendshape coefficient vector of the facial expression and the angle vector of the actuator is updated using the reconstructed inverse model, thereby achieving real-time closed-loop accurate compensation for facial expressions.

2. The real-time closed-loop compensation method for facial expressions in a bionic head according to claim 1, characterized in that, The specific steps for step S1 are as follows: S1.1: Perform self-supervised data acquisition, i.e., control all actuators to perform random exploratory motion. Through randomly generated or preset excitation signal sequences, drive each actuator to independently or collaboratively traverse its entire motion range, including different combinations of amplitude, frequency, and phase. During the motion, simultaneously acquire two data streams at a fixed sampling frequency as training datasets: record the real-time angle values ​​of all actuators to obtain the angle vectors of the actuators. ,in For the set of real numbers, A positive integer, representing the number of actuators in the bionic head. This represents the actuator control space; the facial expressions of the bionic head and their corresponding BlendShape coefficient vectors are synchronously acquired through a vision sensor. ,in A positive integer, representing the semantic feature dimension of the facial expression state. The training dataset represents the facial expression semantic space; it covers the complete mapping relationship between the actuator control space and the facial expression semantic space. S1.2: Based on the above training dataset, a multilayer perceptron is trained as a forward model. This forward model contains several fully connected operation layers to realize high-dimensional linear mapping of features between the facial expression semantic space and the actuator control space. Each fully connected operation layer is followed by a non-linear activation function layer to accurately fit the non-linear physical deformation law of bionic skin under different tensions. The forward model takes the angle vector of the actuator as input and outputs the BlendShape coefficient vector of the predicted expression. During training, the mean square error between the BlendShape coefficient vector of the predicted expression and the BlendShape coefficient vector measured by the visual sensor is minimized to determine the associated weight parameters inside the forward model, thus determining how the angle change of the actuator is non-linearly mapped to the final facial BlendShape coefficient. S1.3: After the forward model training converges and the parameters are fixed, a multilayer perceptron is constructed as the inverse model. The input of the inverse model is the BlendShape coefficient vector and the output is the angle vector. The training of the inverse model is divided into two stages. In the first stage, the input and output are directly swapped in the above training dataset for supervised training, so that the inverse model can initially establish the mapping relationship from the facial expression semantic space to the actuator control space. The second stage introduces a physical consistency constraint based on the forward model. First, the BlendShape coefficient vector corresponding to the target expression is input into the inverse model to obtain the corresponding predicted angle vector. Then, the predicted angle vector is input into the fixed forward model in real time to obtain the BlendShape coefficient vector of the restored expression. By calculating the cycle consistency loss between the BlendShape coefficient vector corresponding to the target expression and the BlendShape coefficient vector of the restored expression, the weights of the inverse model are finely adjusted, and finally a well-trained inverse model is obtained.

3. The real-time closed-loop compensation method for facial expressions in a bionic head according to claim 2, characterized in that, In step S3, a visual capture device is used to acquire the facial expressions of the bionic head in real time at a fixed frame rate, and the Blendshape coefficient vector of the current actual expression is extracted, denoted as... The instantaneous expression residual for the current frame is obtained by subtracting the Blendshape coefficient vector corresponding to the target expression from the Blendshape coefficient vector of the current actual expression. .

4. The real-time closed-loop compensation method for facial expressions in a bionic head according to claim 3, characterized in that, In step S4, the Jacobian matrix is ​​used. As a linear mapping operator from the expression residual space to the actuator control space; multiplying the instantaneous expression residual with the Jacobian matrix directly yields the instantaneous angle compensation of the actuator, denoted as . ,Right now .

5. The real-time closed-loop compensation method for facial expressions in a bionic head according to claim 4, characterized in that, In step S5, the reference angle vector and the instantaneous angle compensation amount are linearly superimposed to generate the final drive command, denoted as... ,Right now The instruction is then sent to each actuator to complete the execution and correction of the baseline action within the adjacent control cycle.

6. The real-time closed-loop compensation method for facial expressions in a bionic head according to claim 5, characterized in that, The specific steps for step S6 are as follows: S6.1: Continuously monitor the instantaneous expression residual in each control cycle of the actuator. If the instantaneous expression residual does not decrease in multiple consecutive control cycles, it is determined that the forward model has deviated from the current physical entity. At this time, the reverse model is reconstructed. S6.2: Data augmentation is performed using the Blendshape coefficient vectors of facial expressions of the bionic head over historical time periods and the actual angle vectors of the actuators to obtain an augmented training dataset. An offline fine-tuning method is used to perform a small number of iterations on the weights of the original forward model, so that the Blendshape coefficient vectors corresponding to the predicted expressions output by the model approximate the measured Blendshape coefficient vectors again. The weight parameters of the forward model are then corrected so that the predicted expressions output by the model are now in line with the physical entity representation of the current bionic head, resulting in an updated forward model. S6.3: Reconstruct the inverse model using the updated forward model according to the training process described in step S1.3: In the first stage, supervised training with input and output swapping is performed using the expanded training dataset, so that the inverse model initially establishes the mapping relationship from the facial expression semantic space to the actuator control space. The second stage introduces physical consistency constraints based on the updated forward model. First, the Blendshape coefficient vector corresponding to the target expression is input into the inverse model to obtain the corresponding predicted angle vector. Then, the predicted angle vector is input into the updated forward model in real time to obtain the Blendshape coefficient vector of the reconstructed expression. By calculating the cycle consistency loss between the Blendshape coefficient vector corresponding to the target expression and the Blendshape coefficient vector of the reconstructed expression, the weights of the inverse model are finely adjusted to obtain a reconstructed inverse model. Subsequent real-time accurate compensation for facial expressions will be based on this reconstructed inverse model. During the reconstruction process, the old inverse model can be temporarily used without affecting the continuous operation of real-time accurate compensation for the actuator controlling facial expressions.

Citation Information

Patent Citations

  • Metacosm virtual digital human manufacturing method and system

    CN114565696A

  • Digital human 3D face reconstruction method, computer equipment and readable storage medium

    CN120388114A