A robot expression online driving method

By adopting the RobotFELNet method based on Transformer encoding structure in robot facial expression learning, the problem that the existing technology cannot generate space-time information in various areas of the face is solved, and the continuous space-time coordinated information of facial muscles and driving motors is reflected, improving the fidelity of robot expression learning and the smoothness of motor movement.

CN115122345BActive Publication Date: 2025-06-06ANQING NORMAL UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210673780.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-15
Publication Date
2025-06-06
Estimated Expiration
2042-06-15

AI Technical Summary

Technical Problem

The prior art cannot generate space-time information in various areas of the face, and cannot reflect the continuous space-time information of facial muscles and driving motors, resulting in low vividness in the expressions learned by robots.

Method used

The RobotFELNet method based on Transformer encoding structure is adopted, including a facial deformation extraction subnet, a face-motor cross-coordinated subnet and a driving sequence generation subnet. Through the regional spatiotemporal attention and cross-attention mechanism, the spatiotemporal characteristics of the facial deformation sequence are extracted and the face-motor collaborative mapping is realized.

Benefits of technology

Generate spatiotemporal semantic representation vectors for each area of ​​the face, which can reflect the continuous spatiotemporal synergistic information between facial muscles and the driving motor, improving the fidelity of robot facial expression learning and the smoothness of motor movement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115122345B_ABST
    Figure CN115122345B_ABST
Patent Text Reader

Abstract

The invention discloses an online driving method for robot expression, which comprises: inputting a facial driving sequence of a performer into a facial deformation extraction subnet based on regional spatiotemporal attention to generate spatiotemporal semantic representation vectors of various facial regions; inputting a driving sequence of a driving motor of a robot and the spatiotemporal semantic representation vectors of various facial regions of the performer into a facial-motor cross-cooperative subnet; inputting an output result of the facial-motor cross-cooperative subnet into a driving sequence generation subnet based on B-spline smoothing constraints, and realizing rolling generation and regularization of future motor driving sequences based on multi-layer LSTM and B-spline smoothing constraints; realizing online expression learning of the robot based on RobotFELNet; the invention has the advantages of providing more accurate target driving data for robot facial expression learning, and finally making the facial expression learned by the robot more realistic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robot facial expression learning, and more specifically to a robot expression online driving method. Background Art

[0002] As humanoid robots with human appearance and action capabilities are gradually applied to human-machine interaction scenarios such as service guides, intelligent reception, and elderly care, how to make "similar" robots have "similar" facial expressions has become a new research hotspot in the field of natural human-machine interaction. Affected by the increase in the number of motors and the degree of freedom of movement, manually choreographed robot expressions are difficult to meet the requirements of "similar" in interactive scenarios, and their realism and motion smoothness need to be further improved. Therefore, under the constraints of hardware structure and mechanical inertia, exploring the automatic learning method of robots for human facial expressions has become a key issue that needs to be solved in the field of natural human-machine interaction.

[0003] The robot facial expression learning based on human teaching is based on the idea of ​​performance drive (or demonstration teaching), mapping the performer's expression sequence to the motor drive sequence and driving the robot to present the same expression. In recent years, researchers have proposed a variety of robot automatic expression learning methods based on key frame technology, polynomial fitting and neural network. The above methods can capture the single-frame static mapping relationship between facial features and motor drive vectors, but cannot reflect the multi-frame spatiotemporal coordination information of facial muscles and drive motors, so the robot learning fidelity is not high. For example, Chinese patent publication number CN108908353A discloses a robot expression imitation method based on a smooth constraint inverse mechanical model, the method comprising: A: extracting robot facial feature vectors: B: constructing a smooth constraint inverse mechanical model based on facial feature sequence to motor control sequence: C: taking the performer's real-time facial features as the target, generating the optimal motor control sequence based on the smooth constraint inverse mechanical model to drive the robot facial motor so that the robot presents an expression corresponding to the performer's face. This method constructs a smooth constrained inverse mechanical model based on a neural network. It can capture the multi-frame mapping relationship between facial features and motor drive vectors, but cannot generate the spatiotemporal information of each facial area, and cannot reflect the continuous spatiotemporal coordination information of facial muscles and drive motors, so the robot learning fidelity is not high. Summary of the invention

[0004] The technical problem to be solved by the present invention is that the existing robot expression learning method cannot generate the spatiotemporal information of each facial area, and cannot reflect the continuous spatiotemporal coordination information of facial muscles and drive motors, so the robot learning fidelity is not high.

[0005] The present invention solves the above technical problems through the following technical means: a robot expression online driving method, the method is implemented based on RobotFELNet, the RobotFELNet includes a facial deformation extraction subnet built based on the Transformer encoding structure, a facial-motor cross-cooperation subnet and a driving sequence generation subnet, the method includes:

[0006] Step 1: Input the performer's facial driving sequence into the facial deformation extraction subnetwork based on regional spatiotemporal attention to generate spatiotemporal semantic representation vectors for each facial region;

[0007] Step 2: The robot's facial expressions are driven by multiple drive motors. The drive sequence of the robot's drive motors and the spatiotemporal semantic representation vectors of each facial region of the performer are input into the facial-motor cross-cooperative subnetwork to achieve the mapping of the spatiotemporal semantics of facial deformation to the motor drive sequence.

[0008] Step 3: The output results of the facial-motor cross-cooperative subnetwork are input into the drive sequence generation subnetwork based on B-spline smoothness constraints, and the rolling generation and regularization of future motor drive sequences are realized based on multi-layer LSTM and B-spline smoothness constraints;

[0009] Step 4: Construct a minimization objective function and use the gradient descent method to solve the optimal parameters of RobotFELNet;

[0010] Step 5: Implement online expression learning of robots based on RobotFELNet under optimal parameters.

[0011] The present invention constructs a facial deformation extraction subnet based on the Transformer encoding structure, extracts the spatiotemporal features of facial deformation sequences, generates spatiotemporal semantic representation vectors of various facial regions, characterizes the spatiotemporal semantic information of different levels and granularities of facial deformation sequences, and can reflect the continuous spatiotemporal coordination information of facial muscles and drive motors, thereby providing more accurate target driving data for robot facial expression learning, and ultimately making the facial expressions learned by the robot more realistic.

[0012] Furthermore, the step 1 comprises:

[0013] Step 101: inputting the facial deformation sequence of the performer into an encoder to obtain an encoded representation of the facial deformation sequence;

[0014] Step 102: The encoded representation of the facial deformation sequence is grouped according to regions to obtain the facial deformation sequence of each region, and the deformation components of the facial deformation sequence of each region are temporally embedded using a fully connected layer to obtain a deformation embedding feature. The deformation embedding feature is used as a token, and encoding and feature cascading are performed based on an L-layer Transformer encoder to obtain the deformation feature of each region;

[0015] Step 103: Using the deformation features of each region as a token, an inter-domain collaborative attention module is constructed, and the inter-domain collaborative attention module outputs the spatiotemporal semantic representation vector of each region.

[0016] Furthermore, the step 101 includes:

[0017] The encoding module of the Transformer architecture is used to extract deep feature semantics of facial deformation sequences and motor drive sequences. The encoder in the encoding module is represented by Y=Endcoder(X+Positional_Encoding,L), where X is the input sequence; Positional_Encoding is the position encoding; Y represents the output sequence of the encoder, and L is the number of layers of the encoder;

[0018] Input the performer's facial deformation sequence into the encoder to obtain F = Encoder (X k (T)+Positional_Encoding,L), where X k (T)∈R N×k is the embedding matrix composed of the facial deformation sequence of k frames before time T, F = [F T-k+1 ,F T-k+2 ,…,F T ] T ∈R N×k It is the encoded representation of the facial deformation sequence of k frames before time T.

[0019] Furthermore, the step 102 includes:

[0020] The encoding representation of the facial deformation sequence of k frames before time T is divided into P groups according to the region. The facial deformation sequence of the pth region is N p represents the number of feature points in the pth region, and represents the time series of the i-th deformation component of the P-th region between T-k+1 and T; represents the amplitude value of the i-th deformation component of the P-th region at time u;

[0021] By formula i∈[1,N p ] The fully connected layer is used to realize the temporal embedding of the deformation component, and FNN(·) is a fully connected layer; F i p Deformation embedding features of

[0022] Embedding features with deformation Token, based on L-layer Transformer encoder, encoding and feature concatenation to get Z p =Concate(Encoder(R p +Positional_Encoding,L)), where is the embedding matrix; N p The P-th regional deformation feature is obtained by cascading the output vectors. for The encoding representation of ; Concate(·) represents the vector cascade function.

[0023] Furthermore, the step 103 includes:

[0024] With P regional deformation features Token, the inter-domain collaborative attention module is constructed through the formula A = Encoder (Z + Positional_Encoding, L), where is the embedding matrix of the inter-domain collaborative attention module; is the output matrix of the inter-domain collaborative attention module, It's Z p The encoding representation of , which means the spatiotemporal semantic representation vector of region p.

[0025] Furthermore, the step 2 includes:

[0026] Step 201: using the driving sequence of the driving motor as the query vector and the spatiotemporal semantics of the region deformation as the key vector and the value vector to construct a cross-attention module to obtain a cross-semantic representation of the driving sequence of the driving motor and the spatiotemporal semantics of the region p;

[0027] Step 202: rearrange the output of the cross-attention module according to the influence of the driving motor, take the deformation semantics of the area affected by the driving motor as a token, perform embedding, Transformer encoding, cascading and fully connected layer mapping to obtain the cross-semantics of the driving motor, take the cross-semantics of the driving motor as a token, and realize motor collaborative representation based on the self-attention mechanism.

[0028] Furthermore, the step 201 includes:

[0029] The intersection semantics of the driving sequence of the driving motor and the spatiotemporal semantics of the p region is expressed as

[0030] Where, Q = Y(T-1)W Q , K=A p W K , V = A p WV , Y(T-1)=[Y 1 ,…,Y j ,…,Y M ] T ∈R M×h is the query vector composed of the historical motion trajectories of M drive motors, is the historical motion trajectory of the jth driving motor; is the spatiotemporal semantic representation of the face in the pth region; W is the cross-semantic representation of the historical control sequence of the j-th driving motor and the spatiotemporal semantics of the face in the p-th region; Q ∈R h ×q , W K ∈R s×q , W V ∈R s×q are weight matrices; s and q represent the embedding dimensions of the cross-attention module.

[0031] Furthermore, the step 202 includes:

[0032] The output of the criss-cross attention module Rearrange according to the influence of the driving motor: j∈[1,M], Indicates the influence of the j-th drive motor on the deformation timing of the p-th region;

[0033] The deformation semantics of P regions affected by the jth driving motor As a Token, through formula C j =FNN(Concate(E j )) and E j =Endcoder(B j +Positional_Encoding, L) for embedding, Transformer encoding, cascading and fully connected layer mapping, where For input The encoding representation of C j ∈R H is the cross semantics of the j-th drive motor;

[0034] by Token, based on the self-attention mechanism to achieve motor collaborative representation, that is, D = Endcoder (C + Positional_Encoding, L), where C = [C 1 ,…,C j ,…C M ] T ∈RM×H is the embedding matrix of the cross attention module; D = [D 1 ,…,D j ,…D M ] T ∈R M×H , D j C j The encoding representation of .

[0035] Furthermore, the step three includes:

[0036] Step 301: Use the cross semantics D of the j-th drive motor j By formula Predict its future d-frame driving sequence, where j∈[1,M], is the driving value of the jth motor at the uth moment, u∈[T,T+d-1], LSTM(·) represents the l-layer LSTM encoder. After parallel calculation and reorganization of M LSTM encoders, the driving vector of the robot at the uth moment is obtained.

[0037] Step 302: Using the driving value of the j-th driving motor for d+1 consecutive moments Construct d-2 cubic uniform B-spline curve segments, the parametric equation of the curve segment is in, is the parametric equation of the i-th segment curve of the j-th drive motor; v∈[0,1] is the node vector, i∈[0,d-3], A i is the coefficient matrix;

[0038] Step 303: Based on the parameter equation of the constructed curve, sample the smoothing values ​​of the M motors respectively and form a smoothing constraint vector in, q = (d-2) / d*(u+1);

[0039] Step 304: The motor drive vector after regularization at time u is expressed as Among them, γ∈[0,1] is the smoothing factor.

[0040] Furthermore, the step 4 includes:

[0041] By formula

[0042]

[0043] Construct a minimization objective function and use the gradient descent method to solve the optimal parameters of RobotFELNet, where: is the real facial deformation vector of the pth region u∈[T,T+d-1] at the moment; is the output vector of the RobotFELNet model at time u; For the current The facial deformation vector of the pth region to be presented after being sent to the robot hardware controller, FK(·) is the robot forward mechanical model, the model parameters are known, and the model is expressed as X = FK(Y), X = (x 1 ,x 2 …,x N )∈R N is the robot's facial deformation vector, Y=(y 1 ,…,y j ,…,y n ) is the facial driving vector of the robot, λ p is the importance of the pth region, and 1≤p≤P,

[0044] Furthermore, the step five includes:

[0045] Under the optimal parameters, the facial deformation sequence of the performer T-k+1~T frames And the robot Th~T-1 frame motor historical drive sequence Y h (T-1) = (Y T-h ,Y T-h+1 ,…,Y T-1 ) is used as input, and the optimal motor drive sequence of the robot at time T~T+d-1 is output based on RobotFELNet:

[0046] Among them, θ is the optimal parameter set of RobotFELNet,

[0047] At time T, The robot is transmitted to the hardware controller and drives the relevant motors to present the robot's expression at time T; at time T+1, the facial deformation sequence of the performer's frame T-k+2~T+1 is used And the robot T-h+1~T frame motor historical drive sequence For input, prediction And drive the motor to present the robot's expression at time T+1.

[0048] The advantages of the present invention are:

[0049] (1) The present invention constructs a facial deformation extraction subnet based on the Transformer encoding structure, extracts the spatiotemporal features of the facial deformation sequence, generates spatiotemporal semantic representation vectors of each facial region, and characterizes the spatiotemporal semantic information of different levels and granularities of the facial deformation sequence. It can reflect the continuous spatiotemporal coordination information of facial muscles and drive motors, thereby providing more accurate target driving data for robot facial expression learning, and ultimately making the facial expressions learned by the robot more realistic.

[0050] (2) Based on the idea of ​​performance-driven learning, the present invention proposes an end-to-end RobotFELNet framework to realize the frame-to-frame migration of the performer’s rigid head posture and non-rigid facial expressions by the robot, thereby improving the realism of the robot’s facial expression learning and the smoothness of the motor movement.

[0051] (3) In order to learn the interactive relationship between the two mutually prime modes of facial spatiotemporal semantics and motor drive sequence, the present invention designs a facial-motor cross-cooperative analysis subnet, and realizes the mapping of facial spatiotemporal semantics to motor drive sequence by embedding cross-attention and multi-motor cooperative analysis modules (the process of step 202 is the implementation process of the multi-motor cooperative analysis module). The obtained motor drive sequence is used to drive the motor movement so that the robot can realize facial expression learning, providing a more accurate reverse driving mechanism for robot expression learning.

[0052] (4) In order to suppress the motor jitter problem of the robot, the present invention constructs a drive sequence generation subnet based on B-spline smoothness constraints, and realizes the regularity of the motor prediction sequence by constructing a B-spline constraint function, thereby providing a smoother motor drive sequence for the robot facial expression learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 A schematic diagram of the RobotFELNet model framework in a robot expression online driving method provided by an embodiment of the present invention;

[0054] Figure 2 A schematic diagram of a robot head driving motor in a robot expression online driving method provided by an embodiment of the present invention;

[0055] Figure 3 A facial topology structure composed of 48 feature points of a robot in an online driving method for robot expression provided by an embodiment of the present invention;

[0056] Figure 4 A schematic diagram of a facial deformation extraction subnet in a robot expression online driving method provided by an embodiment of the present invention;

[0057] Figure 5 A schematic diagram of a facial-motor cross-cooperative subnetwork in an online driving method for robot expression provided by an embodiment of the present invention;

[0058] Figure 6 A schematic diagram of a driving sequence generation subnet in a robot expression online driving method provided by an embodiment of the present invention;

[0059] Figure 7 A schematic diagram of a cubic quasi-uniform B-spline curve segment in an online driving method for robot expression provided by an embodiment of the present invention;

[0060] Figure 8 A schematic diagram of motor drive deviations of RobotFELNet at different times in a robot expression online driving method provided by an embodiment of the present invention;

[0061] Fig. 9 A schematic diagram of average motor drive deviations of RobotFELNet at different times in a robot expression online driving method provided by an embodiment of the present invention;

[0062] Fig.10 A schematic diagram of the deformation preservation result of the robot facial area in a robot expression online driving method provided by an embodiment of the present invention;

[0063] Fig.11 A schematic diagram of the fidelity of robot facial deformation results in a robot expression online driving method provided by an embodiment of the present invention;

[0064] Fig.12 A schematic diagram of the motor motion smoothness results in a robot expression online driving method provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0065] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described in combination with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0066] like Figure 1As shown in the figure, in view of the dynamics of human expressions and the natural advantages of the Transformer model (the Transformer model is an existing model proposed by Google) in sequence processing, the present invention introduces the Transformer architecture into robot expression learning and proposes a robot facial expression learning network (RobotFELNet: Robot Facial Expression Learning Network) based on human teaching. RobotFELNet consists of three parts: facial deformation extraction subnet, facial-motor cross-cooperative subnet and drive sequence generation subnet. The facial deformation extraction subnet mainly completes the spatiotemporal feature extraction of facial deformation sequence, thereby providing accurate drive data for robot expression learning; the facial-motor cross-cooperative subnet mainly realizes the mapping of facial deformation spatiotemporal semantics to motor drive sequence under multi-motor collaborative analysis, thereby providing an inverse kinematics driving mechanism for robot expression learning; the drive sequence generation subnet realizes the rolling generation and regularization of future motor drive sequences according to the facial-motor collaborative cross-semantics, thereby providing a smooth drive sequence for robot expression learning. In the facial deformation extraction subnet, the facial deformation features are first constructed with the head posture, facial motion unit, 3D face mesh and other data obtained from Kinect API, and the facial deformation features are divided into 4 regions according to the head, eyebrows, eyes, mouth, etc. Then, the two mechanisms of intra-domain deformation attention and inter-domain collaborative attention are introduced in the Transformer encoding structure to realize the spatiotemporal feature extraction of the 4 expression regions. In the face-motor cross-cooperative subnet, the facial spatiotemporal features of the 4 regions and the motor drive sequence are used as input to construct a cross-attention module to realize the association learning of the two modalities, and the temporal representation of the face-motor cooperative cross-semantics is realized from the perspectives of single motor temporal attention and multi-motor cooperative attention. In the drive sequence generation subnet, the face-motor cooperative cross-semantics are used as input to realize the rolling prediction of 11 motor drive sequences based on the multi-layer LSTM module, and the cubic B-spline curve constraint is used to realize the regularity and smoothness of the rolling sequence. The following sections introduce in detail the construction process of each subnet of the RobotFELNet model and the process of realizing robot expression learning based on the RobotFELNet model.

[0067] 1. Introduction to robots

[0068] 1.1 Robot selection

[0069] The robot that learns facial expressions in the present invention is a high-imitation robot "Ren Xiangxiang" developed by Hefei University of Technology with 47 control motors. In order to make the designed robot closer to humans, the robot is driven by air pressure to imitate the muscle movements of the human head, shoulders, arms, wrists, waist, legs, etc., and elastic silicone is used to simulate the human skin color and blood vessel texture. In view of the fact that facial expressions are not only the most important carrier of emotional expression, but also the most effective form of reflecting subjective intentions in human-computer interaction, the present invention only takes the head and facial motors related to facial expressions as the research object. The 11 control motors and degrees of freedom in the head are as follows: Figure 2 shown.

[0070] 1.2 Robot facial drive vector representation

[0071] Assume that the facial driving vector Y composed of 11 motors is:

[0072] Y=(y 1 ,…,y j ,…,y n )(n=11) (1)

[0073] In formula (1), y j ∈[0,1] represents the normalized driving value of the jth motor. Similar to human muscles, driving 11 motors can not only allow the robot to show basic expressions such as happiness, surprise, and sadness, as well as delicate emotional details such as blinking, frowning, and raising the corners of the mouth; it can also allow the robot to accompany gestures such as shaking its head and nodding to express subjective intentions.

[0074] 1.3 Robot facial deformation vector representation

[0075] In order to express the mapping relationship between the motor drive vector and the facial details it presents, this paper first uses KinectAPI 2.0 to obtain the robot's 17 facial motion units. 3 head posture rotation angles {R pitch ,R yaw ,R roll} and a parametric face mesh composed of 1347 feature points. Then, among the 1347 feature points, 46 key feature points related to motor and facial deformation were selected. In addition, considering the important role of gaze and eye movement in human-computer interaction, two left and right eye feature points were added. The topological structure and position number of the 48 feature points are as follows: Figure 3 shown.

[0076] exist Figure 3 Based on this, the present invention measures seven facial geometric features related to the drive motor:

[0077]

[0078] In formula (2), line(a,b) and dis(a,b) represent the line segment and Euclidean distance formed by two points a and b respectively; d 0 = dis(27,31) is the distance between the two eye corners; d 1 d 2 Indicates the curvature of the left and right eyebrows; d 3 d 4 Indicates the horizontal and vertical displacement of the left eyeball; d 5 d 6 Indicates the horizontal and vertical displacement of the right eyeball; d 7 In order to accurately describe the facial expression changes and head posture, the present invention uses the rotation angles of the three axes of the head {R pitch ,R yaw ,R roll}、17 facial movement units and 7 facial geometric features As the facial deformation vector X:

[0079] X=(x 1 ,x 2 …,x N )∈R N (3)

[0080] In formula (3), N = 27 is the dimension of the facial deformation vector. 1 ~x 3 is the head posture parameter, x 4 ~x 20 is the facial motion unit parameter, x 21 ~x 27 is the expression geometric feature parameter. From the construction process of the facial deformation vector, we can see that X not only includes the nonlinear deformation of facial muscles, but also incorporates the rigid changes of head posture. Compared with 3D position coordinate information, the representation method based on deformation features can better capture facial nonlinear deformation and describe the changes in facial expressions caused by motor drive.

[0081] 1.4 Problem Description and Modeling

[0082] The facial expressions presented by the robot are directly driven by the motor, and its forward kinematics model FK (Forward Kinematics) can be expressed as:

[0083] X=FK(Y) (4)

[0084] The robot forward mechanical model depends on the internal hardware structure and motor drive capability. When the robot hardware system is developed, the equivalent samples of Y and X under random motor drive states can be collected, and the forward mechanical model can be constructed based on the BP neural network. Since the motor drive degrees of freedom and amplitude no longer change, the forward mechanical model can be considered known and the parameters remain unchanged. However, in the field of human-computer interaction such as demonstration teaching, expression transfer, and action choreography, people are more concerned about how to reverse the motor drive vector through predetermined facial features, that is, the inverse mechanical model (Inverse Kinematics):

[0085] Y=IK(X) (5)

[0086] In formula (5), IK(·) is the mapping function from the robot facial feature space to the mechanical control space. The study shows that although the inverse mechanical model based on frame-to-frame mapping can be constructed in a similar way to the forward mechanical model, the frame-to-frame mapping method, on the one hand, does not consider the dynamics of the input facial expression (i.e., the expression has a median-peak-median change), which affects the fidelity of the driven expression; on the other hand, it does not consider the coherence of the motor-driven motion trajectory, which affects the smoothness of the driven expression. At the same time, a facial feature may be driven by multiple motors, such as the corner of the mouth being raised by the cheek lift (CH 2 )、Mouth opening and closing (CH 6 ), mouth corner contraction (CH 7 ) and other motor control; a drive motor will also affect multiple facial deformations, such as mouth opening and closing (CH 6 ) will affect the opening and closing range of the mouth and the degree of cheek deformation. This many-to-many relationship also increases the complexity of solving IK(·). In addition, in the learning of robot facial expressions, it is necessary to ensure that the facial features presented by the robot are consistent with those of the performer in time and space. This requires not only that the extracted facial features have temporal sequence, but also that the constructed inverse mechanical model has the ability to process and predict temporal information. Therefore, combined with previous research, the present invention improves formula (5) IK(·) into a sequence-sequence inverse mechanical model:

[0087]

[0088] In formula (6), X k (T) = [X T-k+1 ,X T-k+2 ,…,X T ] T ∈R k×N Y is the facial deformation sequence of T-k+1~T robots, which provides target data for robot expression learning; h (T-1) = [Y T-h ,Y T-h+1 ,…,Y T-1 ]T ∈R h×M The motor drive sequence of the robot Th~T-1 provides the historical motion trajectory of the motor for the smooth constraint of the robot; is the motor drive sequence T to T+d-1 predicted by the inverse mechanical model, where Y T Transmitted to the hardware controller to drive the motor to show the expression; For smooth constraint control. Formula (7) IK(·) is the facial deformation sequence X of k frames before time T k (T) and the motor drive sequence Y of the robot h frames before time T-1 h (T-1), predict the motor drive sequence of the robot for the next d frames In order to solve the inverse mechanical model IK(·), the present invention constructs a RobotFELNet robot facial expression learning network based on the Transformer coding structure.

[0089] 2. RobotFELNet robot expression learning network

[0090] The Transformer architecture based on the self-attention mechanism consists of an encoding module and a decoding module. The basic encoding module mainly includes a multi-head attention layer (Multi-HeadAttention), a fully connected feedforward network FNN, and an Add&Norm layer. This invention only uses the Transformer encoding module to extract deep feature semantics of facial deformation sequences and motor drive sequences. For the convenience of the following expression, the encoder containing L layers of Transformer is represented as:

[0091] Y=Endcoder(X+Positional_Encoding,L) (7)

[0092] In formula (7), X is the input sequence; Positional_Encoding is the position encoding; and Y represents the output sequence of the encoder.

[0093] 2.1 Facial Deformation Extraction Subnetwork Based on Regional Spatiotemporal Attention

[0094] From equation (6), we can see that the robot inverse mechanical model takes the facial deformation sequence as input. Therefore, how to achieve the temporal semantic representation of the facial deformation sequence is the main problem to be solved in this section. Intuitively, we can directly convert the facial deformation sequence X from T-k+1 to T k (T) Input to Transformer encoder:

[0095] F=Encoder(X k (T)+Positional_Encoding,L) (8)

[0096] In formula (8), X k (T)∈R N×k is the embedding matrix composed of k frames of facial deformation sequence, F = [F T-k+1 ,F T-k+2 ,…,F T ] T ∈R N×k is the encoded representation of the k-frame facial deformation sequence. However, the Transformer encoder in formula (8) uses the deformation vector X of each frame u (T-k+1≤u≤T) is used as a token, which only focuses on the temporal dependency of the facial sequence, while ignoring the spatial structure and deformation information of each frame of expression. Research related to expression recognition shows that facial shape features and deformation information play a vital role in expression perception and understanding. To solve the problem of missing spatial information, a feasible strategy is to use the deformation component of the deformation vector of each frame as a token and input it to the Transformer encoder. Although this method takes into account the facial structure and deformation information, the number of input tokens increases from k to k*N, and the amount of self-attention calculation will also increase from O(k 3 ) increases suddenly to O(N 3 k 3 ). Therefore, this strategy is difficult to meet the real-time task requirements of robot online expression driving. To address the above problems, this paper introduces two mechanisms, intra-domain deformation attention and inter-domain collaborative attention, based on temporal attention, and proposes a facial deformation extraction subnet based on regional spatiotemporal attention. It mainly includes regional deformation feature grouping, inter-domain deformation attention module and embedded intra-domain deformation attention module, such as Figure 4 shown.

[0097] 2.1.1 Representation of Region Deformation Sequence

[0098] Considering that human expressions are mainly determined by head posture and key areas such as mouth, eyes, and eyebrows, combined with the forward mechanical characteristics of the robot drive motor, the present invention divides N deformation components into P (P=4) groups according to the area. The sequence number and quantity of deformation components in each area are shown in Table 1.

[0099] Table 1 Regional grouping of facial deformation vectors

[0100] Expression area The sequence number of the facial deformation component The number of regional deformation components head 1,2,3,27 4 Eyebrow 17,18,21,22 4 Eye 15,16,23-26 6 mouth 4-14,19,20 13

[0101] Under the premise of regional grouping, reorganize the facial deformation sequence X from T-k+1 to T k (T). Let the facial deformation sequence of the pth region be:

[0102]

[0103] In formula (9), p∈[1,P], N p represents the number of feature points in the pth region, and represents the time series of the i-th deformation component of the p-th region between T-k+1 and T; Represents the amplitude value of the i-th deformation component of the p-th region at time u.

[0104] 2.1.2 In-Domain Deformation Attention Module

[0105] The dynamic change of facial expression depends on the change of details in the local area, and the details of the regional expression are closely related to the topological structure and deformation trajectory information of the regional feature points. Different from Transformer in Transformer and Poseformer, which first embed spatially and then embed temporally, the present invention first measures the spatial and temporal relationships between deformation components in the domain. p The motion sequence of the deformation components For Token, construct an intra-domain deformation attention module, such as Figure 4 As shown. First, the fully connected layer is used to achieve the temporal embedding of the deformation component:

[0106]

[0107] In formula (10), FNN(·) is a fully connected layer; F i p The deformation embedding feature.

[0108] Then, embed the features with deformation For Token, encoding and feature concatenation are performed based on L-layer Transformer encoder:

[0109] Z p =Concate(Encoder(R p +Positional_Encoding,L)) (11)

[0110] In formula (11), is the embedding matrix; N p The concatenation of output vectors, for The encoding representation of ; Concate(·) represents the vector cascade function. p From the calculation process, it can be seen that it integrates N p According to the same strategy, the facial deformation semantics of P regions can be obtained in parallel Compared with the direct cascade deformation sequence, the intra-domain deformation attention module highlights the importance and relevance of different deformation components in the description of expression details; in addition, the intra-domain deformation attention module is limited to the similarity measurement of deformation time sequence in the region, overcoming the inhibitory effect of large deformation regions on small deformation regions. Therefore, the intra-domain deformation attention module can better characterize the deformation process of regional expressions.

[0111] 2.1.3 Inter-domain Collaborative Attention Module

[0112] The regional facial deformation semantics output by the intra-domain deformation attention module reflects the dynamic changes of local expression details. However, the feature information of key areas such as eyebrows, eyes and mouth has certain ambiguity and uncertainty. In addition, facial expression changes are not only closely related to the expression details of key areas, but also closely related to the overall deformation presented by the cooperation between key areas. In view of this, the present invention uses P key area deformation features to For Token, we build an inter-domain collaborative attention module, such as Figure 4 As shown:

[0113] A=Encoder(Z+Positional_Encoding,L) (12)

[0114] Formula (12), is the embedding matrix; is the output matrix, For TokenZ p The encoding representation of .

[0115] At this point, the facial deformation extraction subnet transforms the facial deformation sequence into Transformation into spatiotemporal semantics for describing facial region deformation and collaboration It not only maintains the temporal correlation and global dependency of different frames, but also includes the spatiotemporal information within the expression region of each frame and the spatial collaboration information between regions.

[0116] 2.2 Collaborative analysis subnetwork based on facial-motor cross-attention

[0117] The robot's expression is driven by motor hardware. In the forward mechanical model, the motor drive sequence directly determines the deformation of the robot's face, while in the reverse mechanical model, the facial deformation sequence determines the motion trajectory and trend of the motor drive sequence, that is, there is a high degree of interaction and correlation between the two. However, the motor drive sequence determined by mechanical motion and the facial deformation sequence measured by geometric deformation are essentially coprime. In order to characterize the interactive relationship between different modalities, a facial-motor synergy analysis subnet based on cross-attention is proposed, such as Figure 5As shown, it mainly includes a cross-attention module guided by a motor drive sequence and a multi-motor collaborative analysis module.

[0118] 2.2.1 Motor-driven sequence-guided cross-attention module

[0119] One motor will drive multiple expression areas, and one expression area will be affected by multiple driving motors. In order to accurately describe the many-to-many relationship between the two, a motor drive sequence guided cross attention mechanism is proposed. Its structure is as follows: Figure 5 Different from the self-attention mechanism, the present invention uses the motor drive sequence as the query vector and the region deformation spatiotemporal semantics as the key vector and value vector to construct a cross-attention mechanism. Taking the influence of the motor drive sequence on the pth region as an example, the relevant calculation process is as follows:

[0120]

[0121] Formula (13), Y h (T-1) = [Y 1 ,…,Y j ,…,Y M ] T ∈R M×h is the query vector composed of the historical motion trajectories of M drive motors, is the historical motion trajectory of the jth driving motor; is the spatiotemporal semantic representation of the face in the pth region, which can be obtained by Figure 4 Module obtained; W is the cross-semantic representation of the historical control sequence of the jth driving motor and the spatiotemporal semantics of the face in the pth region. Q ∈R h×q , W K ∈R s ×q , W V ∈R s×q is the trainable weight matrix; s and q represent the embedding dimensions of the cross attention. The cross attention mechanism guided by the motor drive sequence can calculate the historical motor motion trajectory Y in parallel h (T-1) and P regions of facial spatiotemporal semantics of relevance.

[0122] 2.2.2 Multi-motor drive collaborative analysis module

[0123] To present a certain expression, it is necessary to use a single motor to drive the local details independently, and to use multiple motors to coordinate and control the global expression. Therefore, the driving motor not only needs to maintain the timing of its own movement, but also needs to maintain coordination with other motors. This invention constructs a multi-motor drive coordination analysis module from the two perspectives of single motor timing attention and multi-motor coordination attention. Figure 5 shown.

[0124] First, the output of the criss-cross attention module Rearranged according to the effect of the drive motor:

[0125]

[0126] Indicates the influence of the j-th drive motor on the deformation timing of the p-th region.

[0127] Then, in order to highlight the influence of the driving motor on different regions, a single motor temporal attention mechanism is introduced. As a Token, Figure 4 Similar structures are used for embedding, Transformer encoding, cascading, and fully connected layer mapping, namely:

[0128] C j =FNN(Concate(E j ))

[0129] E j =Endcoder(B j +Positional_Encoding,L) (15)

[0130] In formula (15), For input The encoding representation of ; is the cross-semantics of the j-th driving motor, which reflects the correlation and interaction between the motor's historical motion trajectory and the spatiotemporal semantics of the expression area.

[0131] A single motor temporal attention mechanism is used to obtain the driving semantics of M driving motors. Finally, considering the cooperation between different motors, a multi-motor collaborative attention mechanism is introduced, such as Figure 5 As shown. Token, based on the self-attention mechanism to achieve motor collaborative representation, that is:

[0132] D=Endcoder(C+Positional_Encoding,L) (16)

[0133] In formula (16), C = [C 1 ,…,C j ,…C M ] T ∈R M×H is the embedding matrix; D = [D 1 ,…,D j,…D M ] T ∈R M×H , D j TokenC j The encoding representation realizes the cross-fusion and multi-motor coordination of the two modalities: motor drive sequence and spatiotemporal semantics of the affected area deformation.

[0134] 2.3 Driving sequence generation subnetwork based on B-spline smoothness constraints

[0135] The drive sequence generation subnet takes the facial-motor synergy cross semantics as input and realizes the rolling prediction and regularization of the motor drive sequence, thereby providing a smooth drive sequence for robot expression learning. The present invention proposes a drive sequence generation subnet based on B-spline smoothness constraints, such as Figure 6 As shown, it mainly includes two modules: motor drive sequence rolling prediction and motor drive sequence smoothing and regularization.

[0136] 2.3.1 Rolling Prediction Based on LSTM Driven Sequence

[0137] From the above facial-motor coordination analysis subnetwork construction process, we can see that its output is coordinated cross-semantic D j (j∈[1,M]) not only contains the fusion information of the motor drive sequence at time Td~T-1 and the spatiotemporal semantics of the affected area T-k+1~T, but also contains the coordination information between M motors. Therefore, the drive sequence generation subnet only uses the coordinated cross-semantics D of the jth motor. j Predict its future d-frame drive sequence, such as Figure 6 As shown:

[0138]

[0139] In formula (17), is the driving value of the jth motor at the uth moment (u∈[T,T+d-1]), LSTM(·) represents the l-layer LSTM encoder [1] After parallel calculation and reorganization of M LSTM encoders, the driving vector of the robot at time u (u∈[T,T+d-1]) can be obtained.

[0140] 2.3.2 Sequence regularization based on B-spline smoothness constraints

[0141] Unlike the virtual human expression animation based on three-dimensional deformation technology, the driving sequence based on multi-layer LSTM prediction has jumps and jitters due to the influence of mechanical motion inertia and delay. This not only affects the recognition of robot expression learning, but also damages the silicone skin and hardware system due to the large jump of the driving motor. In order to suppress the motor jump in expression learning, the present invention uses a cubic B-spline curve to regularize the motor drive sequence output by LSTM. First, the driving value of the j-th driving motor for d+1 consecutive moments is used. Construct d-2 cubic quasi-uniform B-spline curve segments, such as Figure 7 shown.

[0142] The parametric equations of the d-2 curve segments can be expressed as:

[0143]

[0144] In formula (18), The parametric equation of the i-th segment curve of the j-th drive motor; v∈[0,1] is the node vector. i∈[0,d-3]. According to the geometric properties of the cubic quasi-uniform B-spline curve and the absence of jump constraints at the connection of the curve segments, the coefficient matrix of the d-2 curve segments can be derived as follows:

[0145]

[0146] Then, based on the constructed curve equation, the smoothed values ​​of the M motors at time u (u∈[T,T+d-1]) are sampled and the smoothed constraint vector is formed

[0147]

[0148] In formula (19), q = (d-2) / d*(u+1), i is the integer part of q, and v is the decimal part of q.

[0149] Finally, in order to balance the fidelity of expression learning and the smoothness of motor drive, the motor drive vector after regularization at time u is It can be expressed as:

[0150]

[0151] In formula (20), γ∈[0,1] is the smoothing factor. The larger the γ is, the more likely the regularized motor drive sequence is to maintain the fidelity of expression learning; the smaller the γ is, the more likely the regularized motor drive vector is to maintain the smoothness of the motor motion.

[0152] 2.4 RobotFELNet network parameter optimization

[0153] Considering the importance of different expression regions, the present invention maximizes the fidelity and temporal consistency of the target expression and the driving expression by minimizing the regional deformation deviation:

[0154]

[0155] Formula (21), J is the minimization objective function, is the real facial deformation vector of the pth region u∈[T,T+d-1] at the moment; is the output vector of the RobotFELNet model at time u; For the current The facial deformation vector of the pth region to be presented after being sent to the robot hardware controller, FK(·) is the robot forward mechanical model, and the model parameters are known. p (1≤p≤P) is the importance of the pth region, and λ p The larger the value is, the higher the attention of the pth region in expression learning is, and vice versa. According to the objective function of formula (21), the gradient descent method can be used to solve the optimal parameters of the RobotFELNet model. Under the optimal parameters, the robot facial deformation sequence X k (T) and the historical motion sequence Y of the driving motor h (T-1) is the input, and RobotFELNet can output the motor control sequence with smooth constraints That is, (6) the inverse mechanical model can be expressed as:

[0156]

[0157] In formula (22), θ is the optimal parameter set of RobotFELNet. So far, the proposed RobotFELNet realizes the solution of the time-series-based inverse mechanical model IK(·) under the smoothness constraint and completes the robot facial deformation sequence To the motor drive sequence End-to-end mapping.

[0158] 3. Robot online expression learning based on RobotFELNet

[0159] Robot expression learning is to capture the performer's facial movements online, reverse solve the drive sequence in real time, and drive the motor to make the expression it presents consistent with the performer. And the robot Th~T-1 frame motor historical drive sequence Y h (T-1) = (Y T-h ,Y T-h+1 ,…,Y T-1) is used as input, and the optimal motor drive sequence of the robot at time T~T+d-1 is output based on RobotFELNet.

[0160]

[0161] The difference between formula (22) and formula (23) is that formula (22) takes the robot facial deformation sequence X k (T) is used as input to optimize the parameters of the RobotFELNet network; Formula (23) takes the performer’s current facial deformation sequence The goal is to inversely solve the optimal actuation sequence to keep the robot’s expressions similar to those of the performer.

[0162] In order to maintain the synchronization of the robot's online expression learning, a rolling method is used to achieve frame-by-frame imitation of the expression. The robot is transmitted to the hardware controller and drives the relevant motors to present the robot's expression at time T; at time T+1, the facial deformation sequence of the performer's frame T-k+2~T+1 is used And the robot T-h+1~T frame motor historical drive sequence For input, prediction And drive the motor to present the robot's expression at time T+1; and so on. This input rolling strategy allows the robot to learn expressions in real time with only one frame delay from the performer, ensuring the real-time nature of the robot's expression learning.

[0163] 4. Experimental results and analysis

[0164] 4.1 Dataset and Experimental Details

[0165] In order to solve the parameters of RobotFELNet and verify the effect of online expression learning based on RobotFELNet, this paper uses two methods, manual arrangement and automatic learning, to construct a data set. The manual arrangement method is that the robot animator manually arranges 60 motor drive sequences containing different facial expressions and head postures; the automatic learning method is based on the previous algorithm. [1,2] , allowing the robot to automatically imitate the performer's expression sequence of 20 different postures. Both action sequences lasted for 90 seconds and included non-rigid changes in expression details such as eyeball movement, mouth corner contraction, eyebrow stretching, and rigid changes in head posture. Then, the robot's facial data was captured at a frame rate of 30 per second through the Kinect 2.0 camera and the expression variation features were extracted; finally, the robot's facial feature sequence X of the k frames before time T (400≤T≤2400) was converted into a real-time facial feature sequence. k (T), h frame historical drive sequence Y h (T-1) and the next d frame motor drive sequence Y d (T) constitutes a sample set In the sample set, 120,000 groups of samples are randomly selected for training RobotFELNet, and the remaining Q = 4000 groups of samples are used for performance testing. The Transfromer, LSTM, FFN and other modules in Pytorch are used to build and train RobotFELNet. The relevant parameters are shown in Table 2, and other parameters are default.

[0166] Table 2 Network parameter settings

[0167]

[0168] 4.2 Performance Analysis of RobotFELNet

[0169] 4.2.1 RobotFELNet offline drive deviation evaluation

[0170] To evaluate the effectiveness of RobotFELNet, the driving deviation of the jth motor at time u∈[T,T+d-1] is calculated according to formula (24):

[0171]

[0172] In formula (24), They represent the true value of the jth drive motor at time u for the rth (r∈[1,4000])th group of samples and the predicted value of RobotFELNet. The average motor drive deviation at each moment and the motor drive deviation at different moments of the 4000 groups of samples are Figure 8 , Fig. 9 shown.

[0173] Depend on Figure 8 , Fig. 9The motor drive deviation and average motor drive deviation of the RobotFELNet network d (d = 7) step prediction are shown. From the perspective of timing, the average motor drive deviation shows an upward trend with the cumulative error caused by multi-step prediction, among which the average motor drive deviation at time T is the smallest (3.2), while the average motor drive deviation at time T+d-1 is the largest (6.36); from the perspective of motors, the drive deviations of each motor also show an upward trend with multi-step prediction, and there are certain differences. In general, the motors that have a greater impact on facial expressions, such as gaze (eyeball left and right / eyeball up and down), posture deflection (head left / right), and expression changes (eyebrow up and down, cheek lifting, mouth opening and closing), have lower drive deviations; and due to the influence of the air pressure drive mode, the drive deviation of the upper and lower motors of the robot head is large, but the drive deviation at time T does not exceed 5, and the control deviation at time T+d-1 does not exceed 10. Compared with previous work, RobotFELNet further reduces the drive deviations of motors such as eyebrows, mouth, and eyes through regional grouping and attention mechanism, and improves the overall motor drive accuracy of facial expressions. This shows that the proposed RobotFELNet network better reflects the intrinsic relationship between the robot hardware system and the facial deformation presented by its drive, and more accurately realizes the inverse solution of the robot's facial deformation sequence to the motor drive sequence.

[0174] 4.2.2 Evaluation of RobotFELNet Online Expression Learning

[0175] The above motor drive deviation indicators evaluate the control accuracy of the inverse solution of the motor drive sequence by the RobotFELNet network. However, the motor drive deviation based on offline sequence measurement can only reflect the effectiveness and generalization ability of the RobotFELNet network in the offline state. The dynamic process of robot online expression learning requires not only the evaluation of the realism of expression transfer but also the evaluation of the smoothness of motor movement. In order to evaluate these timing indicators, we first used Kinect Studio V2.0 to record 50 facial action sequences of different performers (each sequence lasted 60 seconds, including neutral-peak-neutral expression intensity changes and head posture changes), and extracted the facial deformation sequence according to the method in Section 2.2. Then, according to equation (22), the optimal motor drive sequence for frame d is solved: And the optimal driving vector Y at time T T The facial expressions are transmitted to the robot hardware control system via serial communication. Simultaneously, the dynamic expressions learned by the robot are captured in real time through the Kinect camera and the facial deformation sequence is extracted based on the same method as in Section 2.2. Finally, based on the expression deformation sequences of the two, the degree of preservation of the robot's facial deformation is counted according to the expression area. On the one hand, the fidelity of robot expression learning is measured based on facial deformation features

[0176]

[0177] In formula (25), They represent the deformation features that need to be learned and the deformation features presented by the robot’s p-th region at time u respectively; They represent the motion amplitude that the robot needs to learn and the motion amplitude presented at time u. S = 50 × 1791 is the total number of frames for the robot to learn online expressions. The fitting function Sim(x1,x2) = 1-|x2-x1| / x1 converts facial deformation deviation or facial motion deviation into a similarity of 0 to 1. Fig.10 , Fig.11 The regional deformation preservation degree of the robot's real-time migration of 50 facial action sequences and the expression migration fidelity when the smoothing factor γ=0.4 are shown respectively.

[0178] Fig.10 It shows that the preservation degree of the eyes and mouth regions is greater than that of the eyebrows and head. As the smoothing factor increases, the deformation preservation degree of each region decreases, but still maintains a high level (greater than 85%). This shows that although RobotFELNet corrects the motor drive sequence through B-spline constraints, it has little effect on the deformation characteristics of the expression area; Fig.11 The results show that the fidelity of each facial motion unit exceeds 85%, especially the facial details that are more sensitive to human senses, such as eye closure, cheek bulging, and mouth corner contraction, which maintain a high fidelity. This shows that the RobotFELNet network can better maintain the similarity of expression motion units through multi-motor collaborative control and smooth constraints. Fig.10 and Fig.11 It can be seen that RobotFELNet has good retention and fidelity for the eye and mouth regions, followed by the eyebrow region, and relatively low retention and fidelity for the head region. This is related to the accuracy of the inverse solution drive sequence of the RobotFELNet network, and is also related to the lack of associated motors in the robot hardware system. Intuitively, the eye and mouth regions contain more drive motors, so they can accurately reproduce the deformation and motion information of the target expression, so these regions have high retention and fidelity; the eyebrow region only contains the upper and lower eyebrow motors, and its reproduced regional deformation information is limited, so the retention and fidelity are low; due to the periodic change of air pressure caused by the pneumatic method, the jitter of the robot head affects the capture accuracy of the head posture data, thereby restricting the learning effect of the head posture.

[0179] From the calculation process of formula (25), it can be seen that the expression deformation retention measures the consistency of the facial deformation amplitude of the robot and the performer, and evaluates the similarity of the robot's expression transfer; the expression fidelity measures the temporal consistency of the facial movement units of the two, and evaluates the rhythm of the robot's expression transfer. However, whether there is a jump in the motor value also has an important impact on the expression learning effect. Therefore, the present invention further measures the motor movement smoothness of online expression learning:

[0180]

[0181] In formula (26): τ is the jump threshold, and in the experiment, τ = 10 / 256. To illustrate the advantages of introducing B-spline smoothing constraints in equation (20), the smoothness effects of motor motion under different smoothing factors are statistically analyzed, as follows: Fig.12 shown.

[0182] Fig.12 It shows that without smoothness constraints (γ = 0), the smoothness of motor motion remains above 80%. This is mainly because the training samples manually arranged by animators have a certain degree of visual smoothness. In addition, the previous robot expression learning methods all have a certain degree of smoothness constraint ability, so that the expression sequences they generate have fewer jumps. This also shows that the proposed RobotFELNet network can maintain the motion characteristics of the training set and has the ability to predict and regularize the timing of motor drive sequences. As γ increases, the smoothness of each drive motor shows an upward trend. Among them, the smoothness of motors such as eyeball left and right, mouth opening and closing, eyebrows up and down, and cheek lifting is better, which shows that the RobotFELNet network has a good motion capture and migration ability for dynamic expression details such as mouth corner lifting and mouth opening. Compared with the motors that drive non-rigid facial deformation, the smoothness of the motors that drive rigid head motion is relatively low. The main reasons are: (1) The mechanical properties of the head posture drive motor are difficult to match the amplitude and speed of the performer's head movement; (2) The facial deformation features obtained based on the Kinect API are more conducive to the description and representation of non-rigid facial semantics.

[0183] Through the above technical scheme, in order to improve the spatiotemporal consistency of robot facial expression learning and reduce the influence of mechanical motion constraints, the present invention proposes a RobotFELNet to realize the real-time migration of the robot to the performer's rigid head posture and non-rigid facial expression. Under the regional grouping strategy, RobotFELNet realizes the reverse mapping of facial deformation sequence to motor drive sequence by introducing regional deformation attention mechanism and facial-motor cross attention mechanism, and maintains the spatiotemporal consistency of expression migration; under the smoothness constraint, the RobotFELNet network realizes the regularization of motor drive sequence based on cubic quasi-uniform B-spline curve, and improves the motion smoothness of robot expression learning. The experiment verifies the effectiveness and generalization ability of RobotFELNet, and discusses the influence of regional weight allocation and smoothing factor on timing indicators such as motor drive deviation, expression learning fidelity, and motor motion smoothness. The experimental results show that the proposed method can reduce the motor drive deviation while further improving the performance of timing indicators that are more sensitive and concerned to human senses.

[0184] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A robot expression online driving method, It is characterized in that The method is implemented based on RobotFELNet, which includes a facial deformation extraction subnet, a facial-motor cross-coordination subnet, and a drive sequence generation subnet built based on a Transformer encoding structure. The method includes: Step 1: Input the performer's facial driving sequence into the facial deformation extraction subnet based on regional spatiotemporal attention to generate spatiotemporal semantic representation vectors for each facial region; Step 1 includes: Step 101: inputting the facial deformation sequence of the performer into an encoder to obtain an encoded representation of the facial deformation sequence; Step 102: The encoded representation of the facial deformation sequence is grouped according to regions to obtain the facial deformation sequence of each region, and the deformation components of the facial deformation sequence of each region are temporally embedded using a fully connected layer to obtain a deformation embedding feature. The deformation embedding feature is used as a token, and encoding and feature cascading are performed based on an L-layer Transformer encoder to obtain the deformation feature of each region; Step 103: Using the deformation features of each region as a token, an inter-domain collaborative attention module is constructed, and the inter-domain collaborative attention module outputs a spatiotemporal semantic representation vector of each region; Step 2: The robot's facial expressions are driven by multiple drive motors. The drive sequence of the robot's drive motors and the spatiotemporal semantic representation vectors of each facial region of the performer are input into the facial-motor cross-cooperative subnetwork to achieve the mapping of the spatiotemporal semantics of facial deformation to the motor drive sequence. Step 3: The output results of the facial-motor cross-cooperative subnetwork are input into the drive sequence generation subnetwork based on B-spline smoothness constraints, and the rolling generation and regularization of future motor drive sequences are realized based on multi-layer LSTM and B-spline smoothness constraints; Step 4: Construct a minimization objective function and use the gradient descent method to solve the optimal parameters of RobotFELNet; Step 5: Implement online expression learning of robots based on RobotFELNet under optimal parameters.

2. A robot expression online driving method according to claim 1, It is characterized in that The step 101 comprises: The encoding module of the Transformer architecture is used to extract deep feature semantics of facial deformation sequences and motor drive sequences. The encoder in the encoding module is represented by Y=Endcoder(X+Positional_Encoding,L), where X is the input sequence; Positional_Encoding is the position encoding; Y represents the output sequence of the encoder, and L is the number of layers of the encoder; Input the performer's facial deformation sequence into the encoder to obtain F = Encoder (X k (T)+Positional_Encoding,L), where X k (T)∈R N×k is the embedding matrix composed of the facial deformation sequence of k frames before time T, F = [F T-k+1 ,F T-k+2 ,…,F T ] T ∈R N×k It is the encoded representation of the facial deformation sequence of k frames before time T.

3. A robot expression online driving method according to claim 2, It is characterized in that The step 102 includes: The encoding representation of the facial deformation sequence of k frames before time T is divided into P groups according to the region. The facial deformation sequence of the pth region is N p represents the number of feature points in the pth region, and represents the time series of the i-th deformation component of the P-th region between T-k+1 and T; represents the amplitude value of the i-th deformation component of the P-th region at time u; By formula i∈[1,N p ] The fully connected layer is used to realize the temporal embedding of the deformation component, and FNN(·) is a fully connected layer; F i p Deformation embedding features of Embedding features with deformation Token, based on L-layer Transformer encoder, encoding and feature concatenation to get Z p =Concate(Encoder(R p +Positional_Encoding,L)), where is the embedding matrix; N p The P-th regional deformation feature is obtained by cascading the output vectors. for The encoding representation of ; Concate(·) represents the vector cascade function.

4. A robot expression online driving method according to claim 3, It is characterized in that The step 103 comprises: With P regional deformation features Token, the inter-domain collaborative attention module is constructed through the formula A = Encoder (Z + Positional_Encoding, L), where is the embedding matrix of the inter-domain collaborative attention module; is the output matrix of the inter-domain collaborative attention module, It's Z p The encoding representation of , which means the spatiotemporal semantic representation vector of region p.

5. A robot expression online driving method according to claim 4, It is characterized in that The second step comprises: Step 201: using the driving sequence of the driving motor as the query vector and the spatiotemporal semantics of the region deformation as the key vector and the value vector to construct a cross-attention module to obtain a cross-semantic representation of the driving sequence of the driving motor and the spatiotemporal semantics of the region p; Step 202: rearrange the output of the cross-attention module according to the influence of the driving motor, take the deformation semantics of the area affected by the driving motor as a token, perform embedding, Transformer encoding, cascading and fully connected layer mapping to obtain the cross-semantics of the driving motor, take the cross-semantics of the driving motor as a token, and realize motor collaborative representation based on the self-attention mechanism.

6. A robot expression online driving method according to claim 5, It is characterized in that The step 201 includes: The intersection semantics of the driving sequence of the driving motor and the spatiotemporal semantics of the p region is expressed as Where, Q = Y(T-1)W Q , K=A p W K , V = A p W V , Y(T-1)=[Y 1 ,…,Y j ,…,Y M ] T ∈R M×h is the query vector composed of the historical motion trajectories of M drive motors, is the historical motion trajectory of the jth driving motor; is the spatiotemporal semantic representation of the face in the pth region; W is the cross-semantic representation of the historical control sequence of the j-th driving motor and the spatiotemporal semantics of the face in the p-th region; Q ∈R h×q , W K ∈R s×q , W V ∈R s×q are weight matrices; s and q represent the embedding dimensions of the cross-attention module.

7. A robot expression online driving method according to claim 6, It is characterized in that The step 202 includes: The output of the criss-cross attention module Rearrange according to the influence of the driving motor: j∈[1,M],B j p Indicates the influence of the j-th drive motor on the deformation timing of the p-th region; The deformation semantics of P regions affected by the jth driving motor As a Token, through formula C j =FNN(Concate(E j )) and E j =Endcoder(B j +Positional_Encoding, L) for embedding, Transformer encoding, cascading and fully connected layer mapping, where For input The encoding representation of C j ∈R H is the cross semantics of the j-th drive motor; by Token, based on the self-attention mechanism to achieve motor collaborative representation, that is, D = Endcoder (C + Positional_Encoding, L), where C = [C 1 ,…,C j ,…C M ] T ∈R M×H is the embedding matrix of the cross attention module; D = [D 1 ,…,D j ,…D M ] T ∈R M×H , D j C j The encoding representation of .

8. A robot expression online driving method according to claim 7, It is characterized in that The step three comprises: Step 301: Use the cross semantics D of the j-th drive motor j By formula Predict its future d-frame driving sequence, where j∈[1,M], is the driving value of the jth motor at the uth moment, u∈[T,T+d-1], LSTM(·) represents the l-layer LSTM encoder. After parallel calculation and reorganization of M LSTM encoders, the driving vector of the robot at the uth moment is obtained. Step 302: Using the driving value of the j-th driving motor for d+1 consecutive moments Construct d-2 cubic uniform B-spline curve segments, the parametric equation of the curve segment is in, is the parametric equation of the i-th segment curve of the j-th drive motor; v∈[0,1] is the node vector, i∈[0,d-3], A i is the coefficient matrix; Step 303: Based on the parameter equation of the constructed curve, sample the smoothing values ​​of the M motors respectively and form a smoothing constraint vector in, q = (d-2) / d*(u+1); Step 304: The motor drive vector after regularization at time u is expressed as Among them, γ∈[0,1] is the smoothing factor.

9. A robot expression online driving method according to claim 8, It is characterized in that The step five comprises: Under the optimal parameters, the facial deformation sequence of the performer T-k+1~T frames And the robot Th~T-1 frame motor historical drive sequence Y h (T-1) = (Y T-h ,Y T-h+1 ,…,Y T-1 ) is used as input, and the optimal motor drive sequence of the robot at time T~T+d-1 is output based on RobotFELNet: Among them, θ is the optimal parameter set of RobotFELNet, At time T, The robot is transmitted to the hardware controller and drives the relevant motors to present the robot's expression at time T; at time T+1, the facial deformation sequence of the performer's frame T-k+2~T+1 is used And the robot T-h+1~T frame motor historical drive sequence For input, prediction And drive the motor to present the robot's expression at time T+1.

Citation Information

Patent Citations

  • Robot expression imitating method and device based on smoothness constraint reverse mechanical model

    CN108908353A