A method for facial expression recognition and imitation of virtual characters or robots

By collecting face images in a loop and identifying facial key points using the depth model mediapipe, the Euler angle of the head position of the computer robot solves the problem of robot expression imitation data with difficult to obtain micro-expressions in the prior art, and realizes accurate imitation and expression of expressions with small dynamic amplitudes.

CN119380394BActive Publication Date: 2025-06-13SHANGHAI DROIDUP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411953666.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-06-13
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively obtain robot expression imitation data with small dynamic amplitude micro-expressions, and it is difficult to achieve precise control of the robot servo in expressions with large dynamic amplitude.

Method used

By collecting face images in a loop, using the depth model mediapipe to identify facial key points and head pose parameters, the Euler angle of the head pose of the computer robot, and map it to the rotation angle of the robot face control servo according to the predicted facial expression value, realizing dynamic control of the robot head pose and facial servo.

Benefits of technology

The robot expression imitation data of micro-expressions with smaller dynamic amplitude is achieved, and the accuracy of expression expression is improved by linking the overall head posture, and it is convenient to map from a large number of video expressions to robot entities to obtain a large number of robot recognition and imitation data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119380394B_ABST
    Figure CN119380394B_ABST
Patent Text Reader

Abstract

An algorithm for facial expression recognition and imitation of virtual characters or robots. The present invention belongs to the field of image recognition and imitation. The algorithm for facial expression recognition and imitation of robots mainly includes: inputting the cropped face image into a deep model, identifying facial key points in the face image through the deep model, and obtaining head pose parameters, and then predicting the value of facial expression from the facial key points; calculating the head pose of the robot: calculating the rotation matrix through the head pose parameters, and converting the rotation matrix into the Euler angles of the robot's head pose; mapping the predicted value of facial expression into the rotation angle of the robot's facial control servo; thereby controlling the operation of the robot. The present invention can obtain robot expression imitation data of micro-expressions with a small dynamic amplitude, and is convenient to accurately map a large number of video expressions into the robot entity to obtain a large amount of robot recognition and imitation data, and link the overall head pose, making the expression more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition imitation, and particularly relates to a method for facial expression recognition and imitation of virtual characters or robots. Background Art

[0002] With the development of artificial intelligence technology, intelligent human-computer interaction technology has become increasingly interesting, and thus is widely loved by users. However, traditional intelligent human-computer interaction technology is only limited to voice and image interaction, and it has been difficult to meet the diversified needs of users. Therefore, emerging expression recognition and interaction technology and eye movement tracking interaction technology have emerged. However, these technologies all require a carrier to better present. The traditional presentation carrier is generally an AI virtual character, but currently, the simulation anthropomorphic expression robot is undoubtedly the best carrier for human-computer interaction. It can not only integrally present various human-computer interaction technologies in a diversified manner, but also is the most natural and friendly interaction carrier for users. Therefore, mature expression robot products can be widely applied to fields such as education assistance, medical companion care, customer guidance service, entertainment demonstration, and social interaction in the future.

[0003] However, at present, the expression robot technology still requires a large amount of expression recognition and imitation data to make the robot's expression interaction more natural and vivid. In the prior art, the patent document with the publication number of CN116597484A discloses a method and device for real-time expression imitation of a robot. This method can obtain the positions of key points on all human faces in the image by recognizing the face image to be detected, determine the face to be imitated in the face image to be detected according to these key point positions, and compare the current frame of the face image to be detected with the previous frame in real time, so as to obtain the transformation trend and transformation ratio of the face to be imitated. The transformation trend is used to adjust the rotation angle of the robot's head, and the transformation ratio is used to calculate the rotation angles and running times of all the robot's servos, so that the robot's servos can control the robot's actions and perform expression imitation on the face to be imitated. Although the above solution can provide robot expression imitation, it is difficult to provide a large amount of expression recognition and imitation data by using the real-time expression imitation method. By obtaining the transformation trend and transformation ratio of all the key point coordinates of the face to be imitated, and then adjusting the robot according to the transformation trend, the imitation data collection can only be realized in expressions with a large dynamic range. However, for some micro-expressions with a small dynamic range, the above solution is difficult to reflect on the robot for imitation actions, and thus it is difficult to obtain robot expression imitation data of micro-expressions with a small dynamic range. Summary of the Invention

[0004] In order to make up for the deficiencies in the existing image recognition and imitation technologies, the present invention proposes a robot expression imitation data that can obtain micro-expressions with a small dynamic amplitude, and is convenient for accurately mapping a large number of video expressions to robot entities to obtain a large number of robot recognition and imitation data, and linking the overall head posture to make the expression more accurate. At the same time, the present invention can also associate virtual characters to perform synchronous expression recognition and imitation, conduct joint interaction, or conduct comparative reference observation experiments, thereby realizing a facial expression recognition and imitation method for virtual characters or robots.

[0005] The specific technical solutions are as follows:

[0006] A method for imitating facial expressions of virtual characters or robots is provided, comprising the following steps:

[0007] S1, cyclically collect face images and perform cropping preprocessing on the face images;

[0008] S2, input the cropped face image into the deep model mediapipe, identify the facial key points in the face image through the deep model mediapipe, and obtain the head pose parameters, and then predict the value of the facial expression based on the facial key points;

[0009] S3. Calculate the robot head pose: establish three basic rotation matrices through the head pose parameters, the product of the three basic rotation matrices is the rotation matrix, and then transform the rotation matrix into three Euler angles corresponding to the robot head pose;

[0010] The calculation of the robot head pose is as follows:

[0011] According to the depth model mediapipe, the head pose parameters are α, β and γ, and the three basic rotation matrices are established as follows:

[0012]

[0013] Then, the product of the three basic rotation matrices is the rotation matrix RM = R (α) x ·R(β) y ·R(γ) z ,Finally, the rotation matrix is ​​transformed into three Euler angles corresponding to the robot head pose;

[0014] The rotation matrix is ​​converted into three Euler angles corresponding to the robot head posture as follows:

[0015] If the rotation matrix is ​​set to:

[0016]

[0017] Then calculate the shaking angle for:

[0018] If r 33 = 0, then it is necessary to determine the value according to the specific rotation situation;

[0019] The pitch angle θ is calculated as:

[0020] θ = arcsin(-r 31 ), and its value range is

[0021] The yaw angle ω is calculated as:

[0022] If r 11 = 0, then it is necessary to determine the value of ω according to the specific rotation situation;

[0023] S4. Map the position of the predicted facial expression value within its predicted original data set range to the target data set for controlling the rotation of the robot's facial servo, so as to obtain the rotation angle of the robot's facial servo;

[0024] S5. Finally, control the pose of the robot's head and the rotation of the robot's facial servo according to the Euler angles and the rotation angles.

[0025] Furthermore, in S1, the face images are cyclically collected by the image acquisition device.

[0026] Furthermore, the numerical values of the Euler angles and the rotation angles obtained in S3 and S4 respectively are subjected to filtering processing.

[0027] Furthermore, the filtering processing method is mean filtering processing. By averaging the angle values generated within the adjacent time of the occurrence time of the currently processed angle value, the obtained average value replaces the currently processed angle value to adjust and control the pose of the robot's head or the rotation of the robot's facial servo. Specifically:

[0028]

[0029] Among them, the obtained O(x i ) is the stable angle value obtained after filtering processing, x i is the currently processed angle value, W is the size of the neighborhood window selected according to x i , and x i+j is the angle value obtained within the adjacent time of x i .

[0030] Furthermore, the value of W is 5 - 10.

[0031] Further, 52 facial key points are recognized in the face image through the depth model mediapipe, 3 head pose parameters are obtained, and then 52 facial expression values are predicted from the facial key points.

[0032] Further, 52 shape keys and 3 head shape poses of the human face are selected in the blender software, and are respectively bound to the corresponding positions of the virtual character's head. Then, a mapping relationship is formed between the 3 head pose parameters obtained by recognizing the face image through the depth model mediapipe and the 3 head shape poses in the blender software, and a mapping relationship is formed between the 52 facial expression values predicted by the facial key points and the 52 shape keys selected in the blender software, so as to realize the binding control of 52 shape keys and 3 head shape poses of the virtual character's facial expression.

[0033] The beneficial effects of the present invention are as follows: It can obtain robot expression imitation data of micro-expressions with a small dynamic range, and it is convenient to accurately map a large number of video expressions to the robot entity to obtain a large amount of robot recognition and imitation data. Moreover, it can link the overall head pose, making the expression more accurate. At the same time, it can also be associated with virtual characters for synchronous expression recognition and imitation, for joint interaction or for comparative reference observation experiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a schematic flow chart of the whole of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] The following elaborates on the preferred embodiments of the present invention in conjunction with the drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.

[0036] Embodiment:

[0037] As Figure 1 shown: A method for virtual character or robot facial expression recognition and imitation includes the following steps:

[0038] S1. The face image is cyclically collected through an image acquisition device, or can be directly extracted from a video picture, and the face image is subjected to cropping preprocessing.

[0039] S2. The cropped face image is input into the depth model mediapipe (an open-source cross-platform framework developed by Google, mainly used for building multimedia processing and machine learning applications). Through the depth model mediapipe, the facial key points in the face image are recognized, and the head pose parameters are obtained. Then, the facial expression values are predicted from the facial key points.

[0040] Among them, 52 facial key points are recognized in the face image through the deep model mediapipe, 3 head pose parameters are obtained, and then 52 facial expression values are predicted from the facial key points.

[0041] S3. Calculate the head pose of the robot: Establish three basic rotation matrices through the head pose parameters. The product of the three basic rotation matrices is the rotation matrix. Then, convert the rotation matrix into three Euler angles corresponding to the head pose of the robot, and link the overall head pose to make the expression of the facial expression more accurate and conducive to the precise mapping of the subsequent facial expression values.

[0042] In S3, the specific method for calculating the head pose of the robot is as follows:

[0043] According to the head pose parameters α, β, and γ obtained from the deep model mediapipe, the three basic rotation matrices are established as follows:

[0044]

[0045] Then, according to the product of the three basic rotation matrices being the rotation matrix RM = R(α) x ·R(β) y ·R(γ) z , finally, convert the rotation matrix into three Euler angles corresponding to the head pose of the robot.

[0046] The specific method for converting the rotation matrix into three Euler angles corresponding to the head pose of the robot is generally as follows:

[0047] If the obtained rotation matrix is set as:

[0048]

[0049] Then calculate the head shake angle as:

[0050] If r 33 = 0, then it is necessary to determine the value of according to the specific rotation situation.

[0051] Calculate the pitch angle θ as:

[0052] θ = arcsin(-r 31 ), and its value range is

[0053] Calculate the yaw angle ω as:

[0054] If r 11 = 0, then it is necessary to determine the value of ω according to the specific rotation situation.

[0055] If the rotation order is different, the formula for calculating Euler angles will also be different. The above calculation order is the Z-Y-X order. For the X-Y-Z order, the calculation process and formula will change; and in practical applications, due to the gimbal lock problem of Euler angles, that is, when the pitch angle is , it will cause the loss of a degree of freedom, and special attention needs to be paid to this problem when performing conversions and using.

[0056] S4. According to the position of the predicted facial expression value within the range of its predicted original data set, map it to the target data set for the rotation of the robot's facial control servo, so as to obtain the rotation angle of the robot's facial control servo. For example, generally, the minimum and maximum values of the facial expression value predicted by identifying the face image through the depth model mediapipe are 0 and 1 respectively, that is, the limit value range is [0,1]. Assuming that the rotation limit range of the servo is 0° to 90°, when the actually predicted facial expression value is 0.1, the rotation angle of the robot's facial control servo obtained is 9°; when the actually predicted facial expression value is 0.5, the rotation angle of the robot's facial control servo obtained is 45°, and so on. Such a mapping method can obtain robot expression imitation data with relatively small dynamic amplitude for microexpressions, and has low requirements for the picture quality, which is convenient for extracting face images from a large number of existing videos or film and television dramas and other rich expressions, accurately mapping them into the robot entity to obtain a large amount of robot recognition imitation data, and based on the prior determination of the head pose, it can more accurately extract and express facial expressions.

[0057] For the numerical values of Euler angles and rotation angles obtained in S3 and S4 respectively, filtering processing is performed to avoid excessive jitter in the movement amplitude of the robot. The filtering processing method is mean filtering processing. By averaging the angle values generated within the adjacent time of the occurrence time of the currently processed angle value, the obtained average value replaces the currently processed angle value to adjust and control the head pose of the robot or the rotation of the robot's facial control servo. Specifically:

[0058]

[0059] Among them, the obtained O(x i ) is the stable angle value obtained after filtering processing, x i is the currently processed angle value, W is the size of the neighborhood window selected according to x i , and its value of W is 5 - 10. Taking the most common 28 frames per second for video processing as an example, the value of W is 5. Because it involves taking values from subsequent times, there is a certain delay, but the delay should not be too large, otherwise it is easy to give an obvious feeling of slowness to the human-machine interaction. x i+j is the angle value obtained within the adjacent time of x i .

[0060] S5. Finally, according to the Euler angles and the rotation angles, control the pose of the robot head and the rotation of the robot face control servo.

[0061] In addition, by selecting 52 shape keys of the human face and 3 head shape poses in the blender software (3D computer graphics software), and binding them to the corresponding positions of the virtual character's head respectively, then a mapping relationship is formed between the 3 head pose parameters obtained by identifying the face image through the depth model mediapipe and the 3 head shape poses in the blender software. The mapping relationship is relatively simple, generally direct input. A mapping relationship is formed between the 52 facial expression values predicted by the facial key points and the 52 shape keys selected in the blender software, so as to realize the binding control of 52 shape keys of the virtual character's facial expression and 3 head shape poses. The virtual character and the robot entity participate in the expression recognition and imitation in parallel at the same time. Furthermore, the virtual character can be associated to perform synchronous expression recognition and imitation, for joint interaction or for comparative reference observation experiments.

[0062] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims.

Claims

1. A method for imitating facial expression recognition of a virtual character or robot, characterized in that: The following steps are involved: S1, cyclically collect face images and perform cropping preprocessing on the face images; S2, input the cropped face image into the deep model mediapipe, identify the facial key points in the face image through the deep model mediapipe, and obtain the head pose parameters, and then predict the value of the facial expression based on the facial key points; S3. Calculate the robot head pose: establish three basic rotation matrices through the head pose parameters, the product of the three basic rotation matrices is the rotation matrix, and then transform the rotation matrix into three Euler angles corresponding to the robot head pose; The calculation of the robot head pose is as follows: According to the depth model mediapipe, the head pose parameters are α, β and γ, and the three basic rotation matrices are established as follows: Then, the product of the three basic rotation matrices is the rotation matrix RM = R (α) x ·R(β) y ·R(γ) z ,Finally, the rotation matrix is ​​transformed into three Euler angles corresponding to the robot head pose; The way to convert the rotation matrix into three Euler angles corresponding to the robot head posture is as follows: If the rotation matrix is ​​set to: Then calculate the shaking angle for: If r 33 = 0, it is necessary to determine according to the specific rotation situation The value of Calculate the pitch angle θ as: θ=arcsin(-r 31 ), whose value range is Calculate the yaw head angle ω as: If r 11 =0, the value of ω needs to be determined according to the specific rotation situation; S4, mapping the predicted facial expression value within the predicted original data set to the target data set of the robot facial control steering gear rotation, thereby obtaining the rotation angle of the robot facial control steering gear; S5. Finally, according to the Euler angle and the rotation angle, the robot head posture and the robot face control servo rotation are controlled.

2. The method for imitating facial expression recognition of a virtual character or robot according to claim 1, characterized in that: In S1, face images are cyclically acquired through an image acquisition device.

3. The method for imitating facial expression recognition of a virtual character or robot according to claim 1 or 2, characterized in that: The values ​​of the Euler angle and the rotation angle obtained in S3 and S4 are filtered.

4. The method for imitating facial expression recognition of a virtual character or robot according to claim 3, characterized in that: The filtering method is mean filtering, which calculates the average value of the angle values ​​generated in the adjacent time of the current angle value, and replaces the current angle value with the average value to adjust the robot head posture or the robot face to control the steering gear rotation. Specifically: Among them, the obtained O(x i ) is the stable angle value obtained after filtering, x i is the angle value currently being processed, and W is the value based on x i The selected neighborhood window size, x i+j For x i The angle values ​​obtained in adjacent times.

5. The method for imitating facial expression recognition of a virtual character or robot according to claim 4, characterized in that: The value of W is 5-10.

6. The method for imitating facial expression recognition of a virtual character or robot according to any one of claims 1, 2, 4 or 5, characterized in that: The deep model mediapipe recognizes 52 facial key points in the face image, and 3 head pose parameters are obtained. Then, the facial expression values ​​predicted by the facial key points are also 52.

7. The method for imitating facial expression recognition of a virtual character or robot according to claim 1, characterized in that: By selecting 52 shape keys and 3 head shape poses related to the human face in the blender software, and binding them to the corresponding positions of the virtual character's head respectively, then mapping the 3 head pose parameters obtained by recognizing the face image through the deep model mediapipe with the 3 head shape poses in the blender software, and mapping the values ​​of the 52 facial expressions predicted by the facial key points with the 52 shape keys selected in the blender software.

Citation Information

Patent Citations

  • Robot real-time expression simulation method and device

    CN116597484A

  • Lightweight robust face alignment method and system based on multi-task learning

    CN115205926A

  • Method for realizing real-time mirroring behavior of robot based on lightweight neural network

    CN115648203A