A facial expression capture system and method
By using facial expression capture devices and methods, combined with multiple modules, accurate positioning of facial markers and expression calculation are achieved, solving the problems of mismatched markers and excessive manual operation, and improving the accuracy and convenience of facial expression capture.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- LANZHOU FUTURE NEW FILM CULTURE & TECH GRP CO LTD
- Filing Date
- 2022-10-18
- Publication Date
- 2026-05-22
AI Technical Summary
Existing facial expression capture systems suffer from problems such as mismatch of marker points, loss of matches, and excessive manual operation, resulting in low capture accuracy and complex operation.
The device employs a facial expression capture system, including an image acquisition sensor, a facial feature point detection module, a facial target point tracking module, a facial template update module, a self-localization calculator, a facial feature point localization module, a facial expression editor, and a facial expression solver. Through various tracking, matching, and verification correction strategies, it achieves accurate localization of facial markers and expression calculation, reducing human interaction.
It improves the accuracy, convenience, stability and applicability of facial expression capture, greatly reduces manual operation, and enhances the system's practicality and accuracy.
Smart Images

Figure CN115497148B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human motion capture technology, specifically to a facial expression capture system and method. Background Technology
[0002] Facial expression capture is a part of motion capture technology, referring to the process of using mechanical devices, cameras, and other equipment to record human facial expressions and movements, and converting them into a series of parametric data. Compared to motion capture, which involves joints and relatively stable human movements, facial expressions are more subtle and complex, thus requiring higher data accuracy. Currently, almost all facial expression capture devices are optical, and can be divided into two main categories based on data sources—two-dimensional data-based and three-dimensional data-based.
[0003] Two-dimensional data-based methods use optical lenses to capture facial two-dimensional data and its changes, then process and label the captured data using algorithms to understand facial expressions. The advantages of this method are low cost, easy access, and ease of use; the disadvantage is lower capture accuracy. Three-dimensional data-based methods, while acquiring two-dimensional data through optical lenses, simultaneously use certain methods or equipment to obtain the depth of the image.
[0004] There are two main methods based on 3D data. The first is the camera array method, where cameras are arranged at certain intervals and according to rules. Camera arrays designed for facial expression capture are usually circular, and the actor needs to be positioned at the center point for filming. The purpose is to obtain different images from different perspectives, thereby acquiring 3D data of facial expressions. The advantages of camera arrays are high precision and good results, but the disadvantages are high filming difficulty, high equipment cost, and the inability to use them during a performance because the actor cannot move. The second method is the structured light method, which uses an infrared lens in addition to an optical lens. Sometimes, auxiliary equipment such as floodlights and floodlight sensors are also needed to obtain depth information of the face. Another commonly used method is to draw or paste markers on the face to help with positioning and infer facial expressions based on changes in these points.
[0005] In terms of face shooting conditions, there are two methods: with markers and without markers.
[0006] Marker-based facial expression capture systems are quite common. The number of markers is determined by the accompanying equipment and system. The equipment is usually head-mounted and can be used in conjunction with human motion and posture capture systems. The actor's performance is not affected. Depending on the shooting scene and implementation method, sometimes handheld shooting is assisted by others. Markerless facial expression capture systems usually rely on landmarks such as nostrils, corners of the eyes, lips, and dimples to determine facial expressions. This method was first implemented by institutions such as CMU, IBM, and the University of Manchester using models and techniques such as Subjective Expression Model (AAM) and Principal Component Analysis (PCA). In recent years, facial reconstruction methods based on deep learning have also emerged in large numbers.
[0007] Markerless systems are highly adaptable and have low equipment costs, but due to their lower system accuracy, they are currently mostly only applicable to the entertainment and consumer market. Marker systems are usually more complex and require more manual operation, but due to their high system accuracy, they are still the preferred solution in fields such as film and television industrialization and virtual characters. Summary of the Invention
[0008] The purpose of this invention is to provide a facial expression capture system and method to solve the technical problems of mismatch, loss of match, and excessive manual operation in the prior art.
[0009] To achieve the above objectives, the present invention provides the following technical solution:
[0010] The present invention provides a facial expression capture system, comprising:
[0011] A facial expression capture device, which is used to collect two-dimensional marker points, key points and facial video information of a human face;
[0012] The facial feature point detection module is used to detect the marker points and feature points to obtain a two-dimensional marker point feature point set and a two-dimensional key point feature point set;
[0013] The facial target point tracking module is used to determine, verify, and correct the specific physical meaning of the marker point feature point set at different times, thereby obtaining a marker point tracking feature point set with accurate physical meaning;
[0014] The facial template update module is used to generate two-dimensional dynamic facial template sets and three-dimensional dynamic facial template sets at different times.
[0015] The self-positioning calculator is used for positioning calculation of the facial expression capture device, that is, to calculate the three-dimensional spatial coordinates and attitude angles of the facial expression capture device in the second coordinate system.
[0016] The facial feature point localization module is used to calculate, verify and correct the three-dimensional coordinates of each point in the marker point tracking feature point set and each point in the key point feature point set in the first coordinate system, thereby generating the three-dimensional feature point set of the facial marker point tracking feature point set and the three-dimensional coordinates of each point in the facial key point feature point set in the first coordinate system.
[0017] A facial expression editor is used to edit and mark the basic facial expression set attributes of the facial video information collected by the facial expression capture device, thereby obtaining the basic expression attribute set of a specific face collected by the facial expression capture device.
[0018] The facial expression solver is used to transform the three-dimensional feature point set of the marker point tracking feature point set and the three-dimensional coordinates of each point in the facial key point feature point set from the first coordinate system to the second coordinate system. Then, based on the basic expression attribute set obtained in the facial expression editor, it solves the expression of a specific face and obtains the expression weight parameters.
[0019] As an optional implementation, the facial expression capture device includes at least two image acquisition sensors, which are used to acquire two-dimensional marker points, key points and facial video information of the face;
[0020] The facial feature point detection module includes at least one facial marker detector and at least one facial key point detector. The facial marker detector is used to detect dot-shaped markers in the facial video information acquired by the image acquisition sensor, thereby obtaining a two-dimensional marker feature point set of the face, wherein all dot-shaped markers are masked. The facial key point detector is used to detect key points in the facial video information acquired by the image acquisition sensor, thereby obtaining a two-dimensional key point feature point set of the face.
[0021] The facial target point tracking module includes at least one facial marker point tracker and at least one marker point tracking verification and correction unit. The facial marker point tracker is used to track the two-dimensional marker point feature point set obtained by the facial marker point detector and determine the specific physical meaning of the two-dimensional marker point feature point set at different times of the face. The marker point tracking verification and correction unit is used to verify and correct the specific physical meaning of the two-dimensional marker point feature point set, thereby obtaining a two-dimensional marker point tracking feature point set with accurate physical meaning.
[0022] The facial template update module includes at least one two-dimensional facial template updater and at least one three-dimensional facial template updater. The two-dimensional facial template updater dynamically updates the two-dimensional facial template to generate a set of dynamic two-dimensional facial templates at different times. The three-dimensional facial template updater dynamically updates the three-dimensional facial template to generate a set of dynamic three-dimensional facial templates at different times.
[0023] The facial feature point localization module includes at least one facial marker point locator, at least one facial key point locator, and at least one facial marker point localization verification and correction unit. The facial marker point locator is used to calculate the three-dimensional coordinates of each point in the marker point tracking feature point set in a first coordinate system. The facial key point locator is used to calculate the three-dimensional coordinates of each point in the key point feature point set in the first coordinate system, thereby obtaining the three-dimensional coordinates of each point in the facial key point feature point set in the first coordinate system. The facial marker point localization verification and correction unit is used to verify and correct the three-dimensional coordinates of each point in the marker point tracking feature point set obtained by the facial marker point locator in the first coordinate system, thereby generating a three-dimensional feature point set of the facial marker point tracking feature point set.
[0024] As an optional implementation, the facial expression capture device further includes a light source, an inertial sensor, a depth image sensor, a wireless communication device for data exchange, and one or more electronic chip processors for analysis and calculation.
[0025] As an optional implementation, the light source is used to assist the image acquisition sensor in image acquisition; the inertial sensor and the depth image sensor are used to assist the facial expression capture device in self-localization.
[0026] A method for capturing facial expressions includes the following steps:
[0027] S1. Acquire at least two image sequences using no fewer than two image acquisition sensors installed within the facial expression capture device;
[0028] S2. The facial feature point detection module detects the marker points and key points of the facial region in each frame of the image sequence, thereby obtaining a two-dimensional marker point feature point set and a two-dimensional key point feature point set of the facial region.
[0029] S3. The facial target point tracking module tracks the two-dimensional marker feature point set of the face to obtain the specific physical meaning of the two-dimensional marker feature point set of the face at different times and from different perspectives. After obtaining the specific physical meaning of the two-dimensional marker feature point set of the face, the specific physical meaning of the two-dimensional marker feature point set is verified and corrected to obtain a two-dimensional marker tracking feature point set of the face with accurate physical meaning.
[0030] The facial marker tracker tracks the two-dimensional marker feature point set and the two-dimensional key point feature point set obtained by the facial feature point detection module from different perspectives, thereby obtaining the same-named marker points at the same time under different perspectives.
[0031] S4. The facial template update module updates the two-dimensional marker point tracking feature point set and the two-dimensional key point feature point set to obtain a two-dimensional dynamic facial template set.
[0032] S5. The facial feature point localization module performs spatial localization calculations on the two-dimensional marker point tracking feature point set to obtain the three-dimensional coordinates of the two-dimensional marker point tracking feature point set in the first coordinate system; at the same time, the two-dimensional key point feature point set performs spatial localization calculations to obtain the three-dimensional coordinates of the two-dimensional key point tracking feature point set in the first coordinate system; and the facial marker point localization verification and correction device verifies and corrects the three-dimensional coordinates of the two-dimensional marker point tracking feature point set in the first coordinate system.
[0033] The method for tracking a set of two-dimensional marker feature points on a human face includes the following steps:
[0034] S51. The first two frames use keyframe block tracking. The blocks are processed according to rigid body transformation. The basic template and reference keyframes are used, eliminating the need for manual interaction.
[0035] S52. All subsequent frames use the two most recent short-term marker reference frames for constant-speed tracking by default. The blocks are processed by affine transformation, and the tracking results are determined by the nearest neighbor criterion and correlation matching to determine whether the tracking is successful.
[0036] S53. For points that are not tracked by the short-term reference frame, use the long-term reference frame for tracking.
[0037] S6. The facial template update module updates the three-dimensional coordinates of the facial two-dimensional marker point tracking feature point set in the first coordinate system and the three-dimensional coordinates of the facial two-dimensional key point tracking feature point set in the first coordinate system to obtain a facial three-dimensional dynamic template set.
[0038] During the first update, the facial template update module generates dynamic template sets of facial two-dimensional marker point tracking feature points and dynamic template sets of facial two-dimensional key point tracking feature points by means of built-in preset templates, manual editing, or automatic template matching;
[0039] S7. Using the self-positioning calculator, perform spatial self-positioning calculation on the facial video information image acquired by the facial expression capture device to obtain the spatial three-dimensional coordinates and attitude angle of the facial expression capture device in the second coordinate system.
[0040] S8. Using the facial expression editor, the image sequence acquired by the facial expression capture device is edited and marked with basic facial expression set attributes, thereby obtaining the basic expression attribute set of a specific face acquired by the facial expression capture device.
[0041] The method for editing and labeling the basic facial expression set attributes of the image sequence acquired by the facial expression capture device includes manual calculation or preset templates;
[0042] S9. Using the facial expression solver, the three-dimensional feature point set of the marker point tracking feature point set and the three-dimensional coordinates of each point in the facial key point feature point set in the first coordinate system are transformed from the first coordinate system to the second coordinate system. Then, based on the basic expression attribute set obtained in the facial expression editor, the expression of a specific face is solved to obtain the expression weight parameters.
[0043] As an optional implementation, in S2, key points of the facial region are set at the facial features, and multiple key points divide the facial region into multiple areas.
[0044] The marker points are selected symmetrically on the left and right sides of the human face, and at least four are set in each area;
[0045] As an optional implementation, the method for tracking the two-dimensional marker feature point set of a human face in S3 includes the following steps:
[0046] S51. The first two frames use keyframe block tracking. The blocks are processed according to rigid body transformation. The basic template and reference keyframes are used, eliminating the need for manual interaction.
[0047] S52. All subsequent frames use the two most recent short-term marker reference frames for constant-speed tracking by default. The blocks are processed by affine transformation, and the tracking results are determined by the nearest neighbor criterion and correlation matching to determine whether the tracking is successful.
[0048] S53. For points that are not tracked by the short-term reference frame, use the long-term reference frame for tracking.
[0049] Based on the above technical solution, the embodiments of the present invention can produce at least the following technical effects:
[0050] This invention provides a facial expression capture system and method. By combining facial markers and key points through a facial feature point detection module and employing multiple tracking, matching, and verification correction strategies through a facial target point tracking module, it can accurately achieve spatial positioning of facial markers and expression calculation. This greatly reduces manual interaction and ensures the accuracy of facial expression capture results, significantly improving the practicality, convenience, accuracy, stability, and applicability of facial expression capture technology. Attached Figure Description
[0051] Figure 1 This is a block diagram of the facial expression capture system of the present invention.
[0052] 10. Facial expression capture device; 11. Image acquisition sensor; 20. Facial feature point detection module; 21. Facial marker point detector; 22. Facial key point detector; 30. Facial target point tracking module; 31. Facial marker point tracker; 32. Marker point tracking verification and correction device; 40. Facial template update module; 41. Facial 2D template updater; 42. Facial 3D template updater; 50. Self-localization calculator; 60. Facial feature point localization module; 61. Facial marker point locator; 62. Facial key point locator; 63. Facial marker point localization verification and correction device; 70. Facial expression editor; 80. Facial expression solver; 90. Result output device; 100. Terminal display. Detailed Implementation
[0053] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0054] Please see Figure 1 A facial expression capture system, comprising:
[0055] A facial expression capture device 10 includes two image acquisition sensors 11, which are used to acquire two-dimensional marker points, key points, and facial video information of a human face. The device also includes a light source, an inertial sensor, a depth image sensor, a wireless communication device for data exchange, and one or more electronic chip processors for analysis and calculation. The light source assists the image acquisition sensors 11 in image acquisition. The inertial sensor and depth image sensor assist the device in self-localization. The depth image sensor performs self-localization calculations by registering 3D features to a second coordinate system.
[0056] The facial feature point detection module 20 includes a facial marker detector 21 and a facial key point detector 22. The facial marker detector 21 is used to detect dot-shaped markers in the facial video information acquired by the image acquisition sensor 11, thereby obtaining a two-dimensional marker feature point set of the face, wherein all dot-shaped markers are masked. The facial key point detector 22 is used to detect key points in the facial video information acquired by the image acquisition sensor 11, thereby obtaining a two-dimensional key point feature point set of the face.
[0057] Two image acquisition sensors 11 synchronously acquire and transmit facial video information to the facial feature point detection module 20. Each facial video information has its own projection information and represents facial information with identical data. Relevant information within the facial video information is generated by the projected image from the image acquisition sensors 11 reflected in the facial region, facial key point features, and facial region marker features. These features are used to calculate the spatial position of facial feature points and also to calculate the relative position between the image acquisition sensors 11. Because all images in a given frame are captured simultaneously and include facial key point features and facial region marker features, synchronization of these features is inherent. Facial region marker features are placed on the facial region, allowing the facial region marker features to remain stationary while moving in space relative to the object coordinate system according to facial expressions. This enables the facial region marker features to move in space while being scanned by the image acquisition sensor devices on their surface.
[0058] The facial marker detector 21 and facial keypoint detector 22 extract facial marker feature points and facial key feature points from each frame of the image. For each frame of facial marker feature points and facial key feature points, a two-dimensional marker feature point set and a two-dimensional key feature point set are output. These points and features are identified in the image based on their inherent characteristics. Facial keypoint features are pixel representations of facial features and contour areas, and are descriptions of the physical features of the facial region in image space. Facial marker features are marker points that are fixedly drawn or pasted on the entire facial region. They are auxiliary information used in facial expression capture technology to mark muscle movements and expression changes at different positions on the face. The two-dimensional positions of facial marker feature points in the image can be obtained using target detection and feature point extraction methods based on image geometric features and grayscale features.
[0059] The facial target point tracking module 30 includes a facial marker point tracker 31 and a marker point tracking verification and correction unit 32. The facial marker point tracker 31 is used to track the two-dimensional marker point feature point set obtained by the facial marker point detector 21 and determine the specific physical meaning of the two-dimensional marker point feature point set at different times of the face. The marker point tracking verification and correction unit 32 is used to verify and correct the specific physical meaning of the two-dimensional marker point feature point set, thereby obtaining a two-dimensional marker point tracking feature point set with accurate physical meaning.
[0060] The facial template update module 40 includes a two-dimensional facial template updater 41 and a three-dimensional facial template updater 42. The two-dimensional facial template updater 41 dynamically updates the two-dimensional facial template. The initial template generates basic data using a combination of manual annotation and automatic generation through binocular matching. Based on the tracking situation within different regions of the face and the correlation matching score, the short-term reference dynamic template, long-term reference dynamic template, key feature point dynamic template, marked feature point dynamic template, and correlation matching dynamic template are updated to generate a set of two-dimensional facial templates at different times. The three-dimensional facial template updater 42 updates the short-term reference dynamic template, long-term reference dynamic template, key feature point dynamic template, and marked feature point dynamic template based on the three-dimensional reconstruction results within different regions of the face to generate a set of three-dimensional facial templates at different times.
[0061] The facial 2D template updater 41 generates a set of facial 2D dynamic templates. Based on the motion consistency of local facial expression regions, it tracks the facial expressions by intelligently combining various templates, such as local and overall, key points and marker points, long-term and short-term, motion constraints and grayscale correlation constraints.
[0062] The facial target point tracking module 30 obtains two-dimensional marker point tracking feature point sets from different viewpoints and at different times by tracking the two-dimensional marker point feature point set. Based on the facial two-dimensional dynamic template set output by the two-dimensional template updater 41, and considering the different facial expressions and the local consistency and temporal continuity of the segmented movement of facial muscle groups, it selects a constant velocity reference dynamic template, a short-term reference dynamic template, a long-term reference dynamic template, a key feature point dynamic template, a marker feature point dynamic template, and a correlation matching dynamic module for intelligent combination tracking in the facial two-dimensional dynamic template set. The module performs segmented tracking and matching of the facial region. Temporal tracking follows the nearest neighbor method and the correlation matching method. Spatial domain tracking uses a method based on epipolar constraints and dynamic basic templates for automatic matching to determine all facial 2D marker point feature point sets of the same face at different viewpoints and all times. Finally, the two-dimensional marker point tracking feature point set is obtained through the marker point tracking verification and correction module 32.
[0063] The self-positioning calculator 50 calculates the pose relationship between the image acquisition sensors 11 in the expression capture device based on stable natural texture feature points or artificially marked tracking points in the image acquired by the image acquisition sensor 11, or based on the dual-target calibration method of the calibration board. Based on fixed natural texture feature points in the background of the image acquired by the image acquisition sensor, or key features such as the tip of the nose and forehead in the complete three-dimensional result set of facial feature points of the facial region, a facial front view coordinate system is constructed, and the spatial position and posture of the expression capture device in the facial front view coordinate system are calculated.
[0064] The facial feature point localization module 60 includes a facial marker point locator 61, a facial key point locator 62, and a facial marker point localization verification and correction unit 63. The facial marker point locator 61 performs three-dimensional reconstruction of the facial region marker points using at least two image acquisition sensors 11 simultaneously. The facial region is equipped with marker points of one or more spatial geometric shapes, used to obtain or assist in obtaining the three-dimensional coordinates of each point in the marker point tracking feature point set in the first coordinate system. The facial key point locator 62 is used to calculate the three-dimensional coordinates of each point in the key point feature point set in the first coordinate system, thereby obtaining the three-dimensional coordinates of each point in the facial key point feature point set in the first coordinate system. The facial marker point localization verification and correction unit 63 is used to verify and correct the three-dimensional coordinates of each point in the marker point tracking feature point set obtained by the facial marker point locator 61 in the first coordinate system, thereby generating a three-dimensional feature point set of the facial marker point tracking feature point set.
[0065] The facial marker locator 62 uses at least two image acquisition sensors 11 at the same time to perform three-dimensional reconstruction of the facial region markers to obtain the spatial position of the markers.
[0066] The facial expression editor 70 is used to edit and mark the basic facial expression set attributes of the facial video information collected by the facial expression capture device 10, thereby obtaining the basic expression attribute set of a specific face collected by the facial expression capture device 10.
[0067] The facial expression solver 80 is used to transform the three-dimensional feature point set of the marker point tracking feature point set and the three-dimensional coordinates of each point in the facial key point feature point set in the first coordinate system to the second coordinate system. The self-positioning calculator 50 transforms the three-dimensional feature point set of the complete marker point tracking feature point set of the facial region and the facial key point feature point set to the face front view coordinate system. Then, based on the basic expression attribute set obtained in the facial expression editor 70, the expression of a specific face is solved to obtain the expression weight parameters, and finally output through the result output device and the terminal display.
[0068] A facial expression capture system includes:
[0069] S1. At least two image sequences are acquired by at least two image acquisition sensors 11 installed within the facial expression capture device 10;
[0070] S2. The facial feature point detection module 20 detects the marker points and key points of the facial region in each frame of the image sequence, thereby obtaining the two-dimensional marker point feature point set and the two-dimensional key point feature point set of the facial region.
[0071] S3. The facial target point tracking module 30 tracks the two-dimensional marker point feature set of the face to obtain the specific physical meaning of the two-dimensional marker point feature set of the face at different times and from different perspectives. After obtaining the specific physical meaning of the two-dimensional marker point feature set of the face, the specific physical meaning of the two-dimensional marker point feature set of the face is verified and corrected to obtain a two-dimensional marker point tracking feature set of the face with accurate physical meaning.
[0072] The facial marker tracker 31 tracks the two-dimensional marker feature set and the two-dimensional key point feature set obtained by the facial feature point detection module 20 from different perspectives, thereby obtaining the same-name marker points at the same time under different perspectives.
[0073] S4. The facial template update module 40 updates the two-dimensional marker point tracking feature point set and the two-dimensional key point feature point set to obtain a two-dimensional dynamic facial template set.
[0074] S5. The facial feature point localization module 60 performs spatial localization calculations on the two-dimensional marker point tracking feature point set to obtain the three-dimensional coordinates of the two-dimensional marker point tracking feature point set in the first coordinate system; at the same time, it performs spatial localization calculations on the two-dimensional key point feature point set to obtain the three-dimensional coordinates of the two-dimensional key point tracking feature point set in the first coordinate system; and the facial marker point localization verification and correction device 63 verifies and corrects the three-dimensional coordinates of the two-dimensional marker point tracking feature point set in the first coordinate system.
[0075] S6. The facial template update module 40 updates the three-dimensional coordinates of the facial two-dimensional marker point tracking feature point set in the first coordinate system and the three-dimensional coordinates of the facial two-dimensional key point tracking feature point set in the first coordinate system to obtain a facial three-dimensional dynamic template set.
[0076] During the first update, the facial template update module 40 generates a dynamic template set of facial two-dimensional marker point tracking feature points and a dynamic template set of facial two-dimensional key point tracking feature points by means of built-in preset templates, manual editing, or automatic template matching;
[0077] S7. Using the self-positioning calculator 50, perform spatial self-positioning calculation on the face video information image collected by the face facial expression capture device 10 to obtain the spatial three-dimensional coordinates and attitude angle of the face facial expression capture device 10 in the second coordinate system.
[0078] S8. The facial expression editor 70 edits and marks the basic facial expression set attributes of the image sequence collected by the facial expression capture device 10, thereby obtaining the basic expression attribute set of a specific face collected by the facial expression capture device 10.
[0079] The method for editing and labeling the basic facial expression set attributes of the image sequence acquired by the facial expression capture device 10 includes manual calculation or preset templates;
[0080] S9. Using the facial expression solver 80, the three-dimensional feature point set of the marker point tracking feature point set and the three-dimensional coordinates of each point in the facial key point feature point set in the first coordinate system are transformed from the first coordinate system to the second coordinate system. Then, based on the basic expression attribute set obtained in the facial expression editor 70, the expression of a specific face is solved to obtain the expression weight parameters.
[0081] like Figure 1As shown, the discrete component groups that communicate with each other via different digital signal connections are readily understood by those skilled in the art. Since some components perform their functions through functions or operations of a given hardware or software system, and the numerous data communication channels shown perform their functions through data communication systems in a computer operating system or application, the preferred embodiment is constituted by a combination of hardware and software components. Therefore, the illustrated structure is provided to effectively illustrate this preferred embodiment.
[0082] Those skilled in the art will understand that marked feature points, which are marked or affixed to the facial area for scanning, can also be provided in other ways as targets to be detected by sensor devices.
[0083] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installed," "equipped with," "sleeved / connected," "connected," etc., should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be a connection within two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0084] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A facial expression capture system, characterized in that, include: A facial expression capture device (10) is used to collect two-dimensional marker points, key points and facial video information of a human face; The facial feature point detection module (20) is used to detect the marker points and feature points to obtain a two-dimensional marker point feature point set and a two-dimensional key point feature point set; The facial target point tracking module (30) is used to determine, verify and correct the specific physical meaning of the marker point feature point set at different times based on the facial two-dimensional dynamic template set generated by the facial template update module (40), so as to obtain a marker point tracking feature point set with accurate physical meaning. The facial template update module (40) is used to generate a two-dimensional dynamic template set of the face at different times and a three-dimensional dynamic template set of the face at different times; the two-dimensional dynamic template set of the face includes a constant velocity reference dynamic template, a short-term reference dynamic template, a long-term reference dynamic template, a key feature point dynamic template, a marked feature point dynamic template, and a correlation matching dynamic template. The self-positioning calculator (50) is used for positioning calculation of the facial expression capture device (10), that is, to calculate the three-dimensional spatial coordinates and attitude angles of the facial expression capture device (10) in the second coordinate system. The facial feature point localization module (60) is used to calculate, verify and correct the three-dimensional coordinates of each point in the marker point tracking feature point set and each point in the key point feature point set in the first coordinate system based on the facial three-dimensional dynamic template set, thereby generating the three-dimensional feature point set of the facial marker point tracking feature point set and the three-dimensional coordinates of each point in the facial key point feature point set in the first coordinate system. A facial expression editor (70) is used to edit and mark the basic facial expression set attributes of the facial video information collected by the facial expression capture device (10), thereby obtaining the basic expression attribute set of a specific face collected by the facial expression capture device (10). The facial expression solver (80) is used to transform the three-dimensional feature point set of the marker point tracking feature point set and the three-dimensional coordinates of each point in the facial key point feature point set in the first coordinate system to the second coordinate system. Then, based on the basic expression attribute set obtained in the facial expression editor (70), it solves the expression of a specific face and obtains the expression weight parameters.
2. The facial expression capture system according to claim 1, characterized in that, The facial expression capture device (10) includes at least two image acquisition sensors (11), which are used to acquire two-dimensional marker points, key points and facial video information of the face; The facial feature point detection module (20) includes at least one facial marker detector (21) and at least one facial key point detector (22). The facial marker detector (21) is used to detect point markers in the facial video information acquired by the image acquisition sensor (11) to obtain a two-dimensional marker feature point set of the face. The point markers are all masked. The facial key point detector (22) is used to detect key points in the facial video information acquired by the image acquisition sensor (11) to obtain a two-dimensional key point feature point set of the face. The facial target point tracking module (30) includes at least one facial marker point tracker (31) and at least one marker point tracking verification and correction unit (32); the facial marker point tracker (31) is used to track the two-dimensional marker point feature point set obtained by the facial marker point detector (21) and determine the specific physical meaning of the two-dimensional marker point feature point set of the face at different times; The marker tracking verification and correction unit (32) is used to verify and correct the specific physical meaning of the two-dimensional marker feature point set, thereby obtaining a two-dimensional marker tracking feature point set with accurate physical meaning; The facial template update module (40) includes at least one two-dimensional facial template updater (41) and at least one three-dimensional facial template updater (42). The two-dimensional facial template updater (41) dynamically updates the two-dimensional facial template to generate a set of dynamic two-dimensional facial templates at different times. The three-dimensional facial template updater (42) dynamically updates the three-dimensional facial template to generate a set of dynamic three-dimensional facial templates at different times. The facial feature point localization module (60) includes at least one facial marker point locator (61), at least one facial key point locator (62), and at least one facial marker point localization verification and correction device (63). The facial marker point locator (61) is used to calculate the three-dimensional coordinates of each point in the marker point tracking feature point set in the first coordinate system. The facial key point locator (62) is used to calculate the three-dimensional coordinates of each point in the key point feature point set in the first coordinate system, thereby obtaining the three-dimensional coordinates of each point in the facial key point feature point set in the first coordinate system. The facial marker point localization verification and correction device (63) is used to verify and correct the three-dimensional coordinates of each point in the marker point tracking feature point set obtained by the facial marker point locator (61) in the first coordinate system, thereby generating a three-dimensional feature point set of the facial marker point tracking feature point set.
3. The facial expression capture system according to claim 2, characterized in that, The facial expression capture device (10) also includes a light source, an inertial sensor, a depth image sensor, a wireless communication device for data exchange, and one or more electronic chip processors for analysis and calculation.
4. The facial expression capture system according to claim 3, characterized in that, The light source is used to assist the image acquisition sensor (11) in image acquisition; the inertial sensor and the depth image sensor are used to assist the facial expression capture device (10) in self-localization.
5. A method for capturing facial expressions, characterized in that, A method for capturing facial expressions based on the facial expression capture system according to any one of claims 1-4, comprising the following steps: S1. Acquire at least two image sequences using at least two image acquisition sensors (11) within the facial expression capture device (10); S2. The facial feature point detection module (20) detects the marker points and key points of the facial region of each frame in the image sequence, thereby obtaining the two-dimensional marker point feature point set and the two-dimensional key point feature point set of the facial region. S3. The facial target point tracking module (30) tracks the two-dimensional marker point feature set of the face by combining the two-dimensional dynamic template set of the face, thereby obtaining the specific physical meaning of the two-dimensional marker point feature set of the face at different times and from different perspectives; after obtaining the specific physical meaning of the two-dimensional marker point feature set of the face, the specific physical meaning of the two-dimensional marker point feature set of the face is verified and corrected, thereby obtaining the two-dimensional marker point tracking feature set of the face with accurate physical meaning; S4. The facial template update module (40) updates the two-dimensional marker point tracking feature point set and the two-dimensional key point feature point set to obtain a two-dimensional dynamic facial template set. S5. The facial feature point positioning module (60) performs spatial positioning calculation on the two-dimensional marker point tracking feature point set to obtain the three-dimensional coordinates of the facial two-dimensional marker point tracking feature point set in the first coordinate system; at the same time, it performs spatial positioning calculation on the two-dimensional key point feature point set to obtain the three-dimensional coordinates of the facial two-dimensional key point tracking feature point set in the first coordinate system, and performs verification and correction based on the facial three-dimensional dynamic template set. S6. The facial template update module (40) updates the three-dimensional coordinates of the facial two-dimensional marker point tracking feature point set in the first coordinate system and the three-dimensional coordinates of the facial two-dimensional key point feature point set in the first coordinate system, thereby obtaining a facial three-dimensional dynamic template set. S7. Using the self-positioning calculator (50), perform spatial self-positioning calculation on the face video information image collected by the face facial expression capture device (10) to obtain the three-dimensional spatial coordinates and attitude angle of the face facial expression capture device (10) in the second coordinate system. S8. The facial expression editor (70) edits and marks the basic facial expression set attributes of the image sequence collected by the facial expression capture device (10) to obtain the basic expression attribute set of a specific face collected by the facial expression capture device (10). S9. Using the facial expression solver (80), the three-dimensional feature point set of the marker point tracking feature point set and the three-dimensional coordinates of each point in the facial key point feature point set in the first coordinate system are transformed from the first coordinate system to the second coordinate system. Then, based on the basic expression attribute set obtained in the facial expression editor (70), the expression of a specific face is solved, and the expression weight parameters are obtained.
6. The facial expression capture method according to claim 5, characterized in that, In S2, key points for the facial region are set at the facial features, and multiple key points divide the facial region into multiple areas. The markers are selected symmetrically on the left and right sides of the face, and at least four are set in each area.
7. The facial expression capture method according to claim 5, characterized in that, The method for tracking a set of two-dimensional marker feature points on a human face in S3 includes the following steps: S51. The first two frames use keyframe block tracking. The blocks are processed according to rigid body transformation. The basic template and reference keyframes are used, eliminating the need for manual interaction. S52. All subsequent frames use the two most recent short-term marker reference frames for constant-speed tracking by default. The blocks are processed by affine transformation, and the tracking results are determined by the nearest neighbor criterion and correlation matching to determine whether the tracking is successful. S53. For points that are not tracked by the short-term reference frame, use the long-term reference frame for tracking.