A face key point recognition and tracking method and system applied to cross-platform
The cross-platform face key point recognition and tracking method enhances accuracy and efficiency by pre-processing images and using layered Kalman filtering to stabilize and expand key point detection, addressing recognition speed and diversity issues in existing systems.
Patent Information
- Application Number
- CN201910688610.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-07-29
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2039-07-29
AI Technical Summary
Existing face detection and tracking systems face issues with slow recognition speed, limited key point detection, low accuracy due to third-party library vulnerabilities, inflexibility in key point expansion, and reduced recognition rates for diverse skin tones and occlusions, along with computational inefficiencies and delays in traditional methods.
A cross-platform face key point recognition and tracking method using pre-processing of images, multi-task convolutional neural networks, and layered Kalman filtering to enhance robustness and expand key point detection, leveraging pre-frame key point information for efficient tracking and robust face key point expansion.
Improves recognition accuracy and efficiency by pre-processing images, stabilizes key point detection through layered Kalman filtering, and expands key point recognition to 108 points without retraining, addressing computational delays and diverse recognition challenges.
Smart Images

Figure CN110399844B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method and system for face key point recognition and tracking applied to cross-platforms. Background Art
[0002] Existing face detection and tracking systems are either based on traditional algorithms and implemented using open source libraries such as OpenCV or MTCNN, or are face key point detection and tracking developed by some enterprises themselves. They are based on machine learning methods and use labeled data to train the labeled data to obtain a model for realizing detection and tracking.
[0003] Disadvantages of current technologies:
[0004] Algorithms that only rely on open source libraries such as OpenCV or MTCNN have disadvantages such as slow recognition speed and few tracked key points. In addition, the security performance of third-party open source libraries is low, with many backdoors, which is likely to cause threats. In the face region detection of such systems, there are problems such as low accuracy of face region detection due to issues such as the brightness, exposure, and contrast of images in the video, or long time consumption and low efficiency of face region detection due to the size of video images.
[0005] Only based on machine learning methods, the number of recognized face key points is limited to a certain value affected by the pre-trained model, and it is impossible to calculate more face key point data from the recognized key points, resulting in inflexibility and inability to expand.
[0006] Based on the training algorithm of machine learning, the accuracy of the dataset will affect the accuracy of the trained model. Currently in the industry, the original dataset is manually labeled, and the error rate is very high.
[0007] Currently, the face recognition and tracking systems on the market will have a recognition rate reduction of more than 50% in the case of people of different ethnic groups, black skin, wearing glasses, hats, etc., and tracking cannot be achieved.
[0008] In traditional tracking systems, the method for improving robustness is limited to the use of Kalman filtering, which will cause a delay in the face key points recognized in the current frame after being processed by the Kalman filter. Summary of the Invention
[0009] The object of the present invention is to provide a cross-platform face key point recognition and tracking method and system, which performs image preprocessing before detection to improve the accuracy; unifies face region detection and face key point recognition to avoid precision loss caused by data transmission between face region detection and face key point recognition, and improves the accuracy; only needs to perform face region detection on the first frame during the tracking process, and uses the key point information of the previous frame as the input of the current frame during the subsequent face key point recognition process, saving calculation time and improving system efficiency; performs robustness enhancement and face key point extension calculation after recognition or tracking calculation to improve the stability of face key point recognition and increase the number of recognized face key points.
[0010] To achieve the above object, the technical solution of the present invention is as follows:
[0011] A cross-platform face key point recognition and tracking method includes the following steps:
[0012] Step 1: Collect face images, mark the key points of each face image, and make a face image training sample set;
[0013] Step 2: Based on the multi-task convolutional neural network algorithm, perform face region detection and key point recognition training on each image to obtain a trained multi-task convolutional neural network model;
[0014] Step 3: Collect face images and preprocess the face images, load the multi-task convolutional neural network model, read the current frame image and synchronously obtain the face region and the corresponding face key point position information;
[0015] Step 4: Use the key point information of the first frame face image as the input of the current frame, calculate the face key point information of the current frame through the multi-task convolutional neural network model, and determine whether the key points of the current frame face image are in a successfully tracked state;
[0016] Step 5: After accumulating the key point information of the expected number of tracked faces, calculate the Euler angle of the face through a pre-trained face deflection intersection calculation model to complete face pose estimation.
[0017] Among them, the face image preprocessing in Step 3 includes reducing the image proportionally, collecting the image brightness, exposure, clarity and contrast and adjusting them to the optimal values.
[0018] Further, the multi-task convolutional network model is divided into three network structures: P-Net, R-Net, and O-Net; the P-Net network is used to detect face image regions and quickly generate face candidate boxes; the R-Net network is used to filter out face candidate boxes with relatively poor effects and perform bounding box regression on the selected candidate boxes and facial key point locators to perform bounding box regression and key point localization of the face region, and output face regions with high credibility; the O-Net network performs face region bounding box regression and face feature localization again, and outputs the upper left coordinate, lower right coordinate of the face region, and five feature points of the face region.
[0019] In step four, after identifying the key points of the current frame face image, the hierarchical Kalman filtering algorithm is used to perform robust enhancement processing on the face key points, and then the SVM support vector machine method is used to determine whether the key points of the current frame face image are in a successfully tracked state.
[0020] In step five, when the cumulative tracking of face key points reaches 68, 68 key points are associated based on the triangulation algorithm and the triangular edges formed by extending specific key points are formed into a focus. The position where the focus is located is the key point estimated for the face organs or the positions around the face, until the face key points are extended and calculated to 108.
[0021] Further, the specific method for performing robust enhancement processing on face key points using the hierarchical Kalman filtering algorithm is: divide the face key points into the 0-16th points of the face outer contour and the 17-67th points of the face inner contour. The face outer contour uses two memory spaces that are m and n times the size of the face shape respectively to store the coordinates of the face shapes of the most recent m and n frames with successful tracking, and set the start flag bit, where 1 ≤ n < m ≤ 100; use the stored valid n and m frame face shape coordinate information and the Kalman filter to perform filtering processing on the currently obtained shape coordinates and output the filtered face shape coordinates as the true coordinates of the current frame.
[0022] In step two, before training the multi-task convolutional neural network model, the linear SVM algorithm is added to store the judgment criteria for whether the face key points are aligned with the face in the model in file format, which is used to determine whether the current frame is successfully tracked when performing face key point tracking.
[0023] A cross-platform face key point recognition and tracking system, characterized in that it includes a face detection and key point recognition model training module, a face image processing module, and a face information calculation module;
[0024] The face detection and key point recognition model training module, based on the collected face images and annotated key points, uses the multi-task convolutional neural network algorithm for training to obtain a multi-task convolutional neural network model;
[0025] The face image processing module includes an image restoration and preprocessing unit, a face region and key point recognition unit, and a face region and key point tracking unit;
[0026] The image restoration and preprocessing unit scales down the face image to be detected proportionally, collects the brightness, exposure, clarity, and contrast of the face image, and adjusts them to the optimal values;
[0027] The face region and key point recognition unit reads the current frame of the face image to be detected input into the multi-task convolutional neural network model, and initially obtains the face region and the corresponding position information of the face key points;
[0028] The face region and key point tracking unit uses the key point information of the previous frame as the input of the current frame to determine whether the face key points of the current frame are in a successfully tracked state until the key points of the expected number of faces are cumulatively tracked;
[0029] The face information calculation module calculates the Euler angles of the face through a pre-trained face deflection intersection calculation model to complete face pose estimation.
[0030] Furthermore, the face image processing module further includes a face key point robustness enhancement unit, which uses a hierarchical Kalman filtering algorithm to perform robustness enhancement processing on the key points of the current frame face image.
[0031] Furthermore, the face image processing module includes a face key point calculation unit, which associates each key point based on the triangulation algorithm and extends the triangular edges formed by specific key points to form a focus. The position where the focus is located is the key point of the estimated face organ or the position around the face, expanding the face key points.
[0032] The cross-platform face key point recognition and tracking method and system of the present invention have the following advantages:
[0033] Improved recognition accuracy: By performing image correction and preprocessing before detection, detecting the parameter information of the image to be detected including size, brightness, exposure, and contrast, and repairing these parameters of the image to the most suitable values for detection, the recognition accuracy is improved by using preprocessing;
[0034] Secondly, the face region detection and face key point recognition are unified to avoid the accuracy loss caused by data transmission between face region detection and face key point recognition, and the accuracy is improved.
[0035] Improve system efficiency, stability, and availability: During the tracking process, only face region detection is required in the first frame. In the subsequent face key point recognition process, the key point information of the previous frame is used as the input for the current frame, saving calculation time and improving system efficiency. After recognition or tracking calculations, robustness enhancement and face key point expansion calculations are performed to improve the stability of face key point recognition, and the number of face key points recognized can be increased to 108 without training a model, thereby enhancing system availability;
[0036] Using a hierarchical Kalman filter to perform filtering operations on face key point information at different positions can reduce the computational amount of the filter, improve the operation speed, eliminate delays while ensuring jitter elimination, and solve the problem of delays caused by Kalman filtering in traditional methods. Brief Description of the Drawings
[0037] The drawings forming a part of the specification depict embodiments of the present invention and, together with the description, are used to explain the principles of the present invention. Referring to the drawings, the present invention can be more clearly understood:
[0038] Figure 1 is a flowchart of a method for cross-platform face key point recognition and tracking according to an embodiment of the present invention;
[0039] Figure 2 is a flowchart of face image preprocessing according to an embodiment of the present invention;
[0040] Figure 3 is a schematic block diagram of a cross-platform face key point recognition and tracking system according to an embodiment of the present invention;
[0041] Figure 4 is an output effect diagram of using a multi-task convolutional neural network model according to an embodiment of the present invention;
[0042] Figure 5 is an output effect diagram of expanding key points using a triangulation algorithm according to an embodiment of the present invention. Detailed Embodiments
[0043] The technical solutions of the present invention will be further described in detail below with reference to the drawings and embodiments.
[0044] As Figure 1 , a method for cross-platform face key point recognition and tracking according to the present invention includes the following steps:
[0045] Step 1: Collect face images, mark the key points of each face image, and make a face image training sample set;
[0046] Step 2: Based on the multi-task convolutional neural network algorithm, perform face region detection and training for identifying face key points on each image to obtain a trained multi-task convolutional neural network model;
[0047] Step 3: Collect face images and preprocess the face images. Load the multi-task convolutional neural network model, read the current frame image, and synchronously obtain the face region and the corresponding face key point position information;
[0048] Step 4: Use the key point information of the first frame of the face image as the input of the current frame. Calculate the face key point information of the current frame through the multi-task convolutional neural network model, and determine whether the key points of the current frame face image are in a successfully tracked state by the SVM (Support Vector Machine) method;
[0049] Step 5: After accumulating the predicted number of face key point information for tracking, calculate the Euler angles of the face through a pre-trained face deflection intersection calculation model to complete face pose estimation.
[0050] Before training the multi-task convolutional neural network model, add the linear SVM algorithm to store the judgment criteria for whether the face key points are aligned with the face in the model in a file format, which is used to determine whether the current frame is successfully tracked when performing face key point tracking.
[0051] As Figure 2 shown, the face image preprocessing includes reducing the image proportionally, collecting the image brightness, exposure, clarity, and contrast, and adjusting them to the optimal values.
[0052] As Figure 4 shown, the multi-task convolutional network model is divided into three-layer network structures: P-Net, R-Net, and O-Net; the P-Net network is used to detect the face image region and quickly generate face candidate boxes; the R-Net network is used to filter out the face candidate boxes with relatively poor effects and perform border regression on the selected candidate boxes and perform border regression and key point localization of the face region by the facial key point locator to output a face region with high credibility; the O-Net network performs face region border regression and face feature localization again, and outputs the upper left coordinate, lower right coordinate of the face region, and five feature points of the face region.
[0053] Multi-task Convolutional Neural Network (MTCNN) processes face region detection and face key point detection in parallel. Its main framework is generally divided into three network structures: P-Net, R-Net, and O-Net. It is a multi-task neural network model for face detection tasks, and adopts the idea of candidate boxes plus classifiers for fast and efficient face detection. These three cascaded networks are P-Net for quickly generating candidate windows, R-Net for filtering and selecting candidate windows with high precision, and O-Net for generating the final bounding boxes and face key points. At the same time, the multi-task convolutional neural network model also uses technologies such as image pyramid, bounding box regression, and non-maximum suppression. The specific process is as follows:
[0054] Build an image pyramid: First, transform the image at different scales to build an image pyramid to adapt to the detection of faces of different sizes.
[0055] P-Net, full name Proposal Network, is a fully connected network. For the image pyramid constructed in the previous step, it conducts preliminary feature extraction and calibration of bounding boxes through a FCN, and uses bounding box regression to adjust the windows and non-maximum suppression to filter most of the windows.
[0056] P-Net is a region detection network for face regions. After the network inputs the features into three convolutional layers, it uses a face classifier to determine whether the region is a face. At the same time, it uses bounding box regression and facial key point locators to preliminarily detect the face region. Finally, it outputs many face regions where faces may exist and inputs them into R-Net for further processing.
[0057] R-Net, full name Refine Network, is a convolutional neural network. Compared with the first-layer P-Net, it adds a fully connected layer, so the screening of input data is more strict. After the face image passes through P-Net, many prediction windows will be left. The R-Net network is used to filter out face candidate boxes with poor effects and perform bounding box regression and facial key point locators on the selected candidate boxes for bounding box regression and key point localization of the face region, and outputs face regions with high credibility.
[0058] The output of P-Net is only face regions that may have a certain credibility. Through R-Net, the input will be refined and selected, and most of the wrong inputs will be discarded. Then, bounding box regression and facial key point locators are used again for bounding box regression and key point localization of the face region. Finally, the output will be more credible face regions for O-Net to use. Compared with the 1x1x32 features output by P-Net using a fully convolutional layer, R-Net uses a fully connected layer of 128 after the last convolutional layer, retaining more image features, and its accuracy performance is also better than P-Net.
[0059] The O-Net, whose full name is Output Network, is a relatively complex convolutional neural network. Compared with the R-Net, it has one more convolutional layer. The difference in the effect between the O-Net and the R-Net lies in that this layer structure can identify the facial region through more supervision and perform regression on the facial feature points of a person, and finally output five facial feature points of a human face.
[0060] A fully connected layer with 256 is adopted, which retains more image features. At the same time, face discrimination, face region bounding box regression, and face feature localization are performed, and finally the upper left coordinate and the lower right coordinate of the face region and five feature points of the face region are output. The O-Net has more feature-rich inputs and a more complex network structure, and also has better performance. The output of this layer is used as the output of the final network model.
[0061] In order to balance performance and accuracy and avoid the huge performance consumption brought by traditional ideas such as sliding windows plus classifiers, the MTCNN first uses a small model to generate candidate bounding boxes of target regions with a certain possibility, and then uses a more complex model for fine classification and higher-precision region bounding box regression, and recursively executes to construct a three-layer network. At the input layer, an image pyramid is used for scale transformation of the initial image, and the P-Net is used to generate a large number of candidate target region bounding boxes. Then, the R-Net is used to perform the first selection and bounding box regression on these target region bounding boxes, excluding most of the negative examples. Then, a more complex and higher-precision network, the O-Net, is used to perform discrimination and region bounding box regression on the remaining target region bounding boxes.
[0062] After identifying the key points of the face image in the current frame, the hierarchical Kalman filter algorithm is used to perform robustness enhancement processing on the face key points, and then the SVM (Support Vector Machine) method is used to determine whether the key points of the face image in the current frame are in a successfully tracked state.
[0063] Traditional methods perform filtering operations on the positions of all face key points through filtering algorithms such as Kalman filters to eliminate the jitter of the key points, but this will cause the face key point information calculated in the current frame to be the key point information under the face positions of the previous n frames, resulting in delays. The hierarchical Kalman filter performs filtering operations on the face key point information at different positions.
[0064] The specific method for robustly enhancing facial key points using a hierarchical Kalman filtering algorithm is as follows: The facial key points are divided into points 0 - 16 of the facial outer contour and points 17 - 67 of the facial inner contour. The facial outer contour uses two memory spaces that are m and n times the size of the facial shape respectively to store the coordinates of the facial shapes of the most recent m and n frames with successful tracking. A starting flag is set, where 1 ≤ n < m ≤ 100. Using the stored valid coordinate information of n and m frames of facial shapes and a Kalman filter to perform filtering on the currently obtained shape coordinates and output the filtered facial shape coordinates as the true coordinates of the current frame. By using hierarchical operations, the computational amount of the filter can be reduced, the operation speed can be increased, and the delay can be eliminated while ensuring the elimination of jitter.
[0065] In the module for robustly enhancing facial key points, the least squares SVM algorithm is used to correct the alignment of the facial key points in the current frame, making the facial key points in the current frame closer to the true facial positions and reducing the jitter caused by excessive error offsets of the facial key points between adjacent frames.
[0066] As Figure 5 shown, when the facial key points are cumulatively tracked up to 68, 68 key points are associated based on the triangulation algorithm and the triangular edges formed by extending specific key points are formed into a focus. The position of the focus is the key point of the estimated facial organ or the position around the face. This is continued until the facial key points are extended and calculated to 108. By using pre - set determination parameters, information related to facial expressions is calculated, such as whether the mouth is open, whether the eyes are closed, whether the eyebrows are raised, whether the lips are pursed, whether the tongue is sticking out, etc., to achieve the estimation of facial expression recognition.
[0067] A cross - platform facial key point recognition and tracking system, as Figure 3 shown, includes a facial detection and key point recognition model training module, a facial image processing module, and a facial information calculation module. Both the facial detection and key point recognition model training module and the facial information calculation module are connected to the facial image processing module.
[0068] Among them, the facial detection and key point recognition model training module, based on the collected facial images with key points annotated, uses a multi - task convolutional neural network algorithm for training to obtain a multi - task convolutional neural network model;
[0069] The facial image processing module includes an image restoration and pre - processing unit and a facial region and key point recognition unit, a facial region and key point tracking unit, a facial key point robustness enhancement unit, and a facial key point calculation unit that are connected to the image restoration and pre - processing unit in sequence.
[0070] The image restoration and pre - processing unit reduces the size of the face image to be detected proportionally and collects the brightness, exposure, clarity, and contrast of the face image and adjusts them to the optimal values.
[0071] The face region and key point recognition unit reads the current frame of the face image to be detected input into the multi-task convolutional neural network model, and preliminarily obtains the face region and the corresponding face key point position information.
[0072] The face region and key point tracking unit uses the key point information of the previous frame as the input of the current frame to determine whether the face key points of the current frame are in a successfully tracked state until the key points of the expected number of faces are accumulated.
[0073] The face key point robustness enhancement unit uses a hierarchical Kalman filtering algorithm to perform robustness enhancement processing on the key points of the current frame face image.
[0074] The face key point calculation unit associates each key point based on the triangulation algorithm and extends the triangular edges formed by specific key points to form a focus. The position where the focus is located is the key point of the estimated face organ or the position around the face, expanding the face key points.
[0075] The face information calculation module calculates the Euler angle of the face through a pre-trained face deflection intersection calculation model to complete the face pose estimation.
[0076] The main functions provided by the face key point recognition and tracking system of the present invention applied to cross-platform include: it can be used across platforms (windows, mac, IOS, Android, Linux); it can detect the image of each frame of the video in real time, recognize the face region in the image and the corresponding face key point position information; 40 key point positions of the face around and other corresponding organs are calculated from the face region and the corresponding 68 face key point positions; finally, the face pose and expression detected in each frame of the video are estimated.
[0077] The above specific implementation manners have further detailed the purpose, technical solution and beneficial effects of the present invention. It should be understood that the above are only the specific implementation manners of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for face key point recognition and tracking applied to cross-platforms, characterized in that, It includes the following steps: Step 1: Collect face images, mark the key points of each face image, and make a training sample set of face images; Step 2: Based on the multi-task convolutional neural network algorithm, detect the face region of each image and train to identify the key points of the face, obtaining a trained multi-task convolutional neural network model; before training the multi-task convolutional neural network model, add the linear SVM algorithm to store the determination criteria for whether the face key points are aligned with the face in the model in a file format, which is used to determine whether the current frame is successfully tracked when tracking the face key points; Step 3: Collect face images and preprocess the face images, load the multi-task convolutional neural network model, read the current frame image and synchronously obtain the face region and the corresponding face key point position information; Step 4: Use the key point information of the first frame of the face image as the input of the current frame, calculate the face key point information of the current frame through the multi-task convolutional neural network model, and determine whether the key points of the current frame face image are in a successfully tracked state. After identifying the key points of the current frame face image, perform a hierarchical Kalman filter algorithm to strengthen the robustness of the face key points. After the robustness enhancement process, use the SVM support vector machine method to determine whether the key points of the current frame face image are in a successfully tracked state; the specific method of using the hierarchical Kalman filter algorithm to strengthen the robustness of the face key points is: divide the face key points into points 0-16 of the face outer contour and points 17-67 of the face inner contour. The face outer contour uses two memory spaces that are m and n times the size of the face shape respectively to store the coordinates of the face shapes of the most recent m and n frames that have been successfully tracked, and set a starting flag bit, where 1≤n<m≤100; use the stored effective n and m frame face shape coordinate information and the Kalman filter to filter the currently obtained shape coordinates and output the filtered face shape coordinates as the true coordinates of the current frame; Step 5: After accumulating the key point information of the expected number of tracked faces, calculate the Euler angle of the face through a pre-trained face deflection intersection calculation model to complete face pose estimation; when accumulating the face key points to 68, based on the triangulation algorithm, associate the 68 key points and extend the triangular edges formed by specific key points to form a focus, and the position where the focus is located is the key point of the estimated face organ or the position around the face, until the face key points are extended and calculated to 108.
2. The method for cross-platform face key point recognition and tracking according to claim 1, characterized in that: In Step 3, the face image preprocessing includes reducing the image proportionally, collecting the image brightness, exposure, sharpness, and contrast and adjusting them to the optimal values.
3. The method for face key point recognition and tracking applied to cross-platform according to claim 1, wherein: The multi-task convolutional network model is divided into three network structures: P-Net, R-Net, and O-Net. The P-Network is used to detect face image regions and quickly generate face candidate boxes. The R-Network is used to filter out face candidate boxes with relatively poor effects and perform bounding box regression and facial key point localization on the selected candidate boxes to perform bounding box regression and key point localization of the face region, and output a face region with high credibility. The O-Network performs face region bounding box regression and face feature localization again, and outputs the upper left coordinate, lower right coordinate of the face region, and five feature points of the face region.
4. A face key point recognition and tracking system applied to cross-platform, based on the method described in any one of claims 1-3, characterized in that: It includes a face detection and key point recognition model training module, a face image processing module, and a face information calculation module. The face detection and key point recognition model training module is connected to the face image processing module, and the face image processing module is connected to the face information calculation module; The face detection and key point recognition model training module is trained based on the collected face images with key points marked, using the multi-task convolutional neural network algorithm to obtain a multi-task convolutional neural network model; The face image processing module includes an image restoration and preprocessing unit, and a face region and key point recognition unit and a face region and key point tracking unit that are sequentially connected to the image restoration and preprocessing unit; The image restoration and preprocessing unit scales down the face image to be detected proportionally, and collects the brightness, exposure, clarity, and contrast of the face image and adjusts them to the optimal values; The face region and key point recognition unit reads the current frame of the face image to be detected input into the multi-task convolutional neural network model, and initially obtains the face region and the corresponding face key point position information; The face region and key point tracking unit uses the key point information of the previous frame as the input of the current frame to determine whether the face key points of the current frame are in a successfully tracked state until the key points of the expected number of faces are cumulatively tracked; The face information calculation module calculates the Euler angle of the face through a pre-trained face deflection intersection calculation model to complete face pose estimation.
5. The face key point recognition and tracking system applied to cross-platform according to claim 4, wherein: The face image processing module further includes a face key point robustness enhancement unit, which is connected to the face region and key point tracking unit, and uses a hierarchical Kalman filter algorithm to perform robustness enhancement processing on the key points of the current frame face image.
6. The cross-platform face key point recognition and tracking system according to claim 4, wherein: The face image processing module includes a face key point calculation unit, which is connected to the face key point robustness enhancement unit, and associates each key point based on the triangulation algorithm and extends the triangular edges formed by specific key points to form a focus. The position where the focus is located is the key point for estimating the face organs or the positions around the face, and the face key points are extended.
Citation Information
Patent Citations
Face key point tracking system and method applied to mobile device
CN106909888A
Method and device for image merging
CN108898551A