Facial expression tracking method, model training method, device, medium and product

A facial expression tracking model that segments and performs 3D convolution operations on multiple frames of facial images solves the problem of inaccurate facial expression features and achieves accurate facial expression tracking under conditions of occlusion or partial area acquisition.

CN121392925BActive Publication Date: 2026-07-24GEER TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GEER TECH CO LTD
Filing Date
2025-10-11
Publication Date
2026-07-24

Smart Images

  • Figure CN121392925B_ABST
    Figure CN121392925B_ABST
Patent Text Reader

Abstract

The application discloses a facial expression tracking method, a facial expression tracking model training method, an electronic device, a storage medium and a computer program product, and relates to the technical field of image processing. The method comprises the following steps: taking each frame of facial image in a plurality of frames of facial images continuously collected for a target object as a first target image, respectively, and obtaining a first facial image sequence corresponding to the first target image; inputting each frame of facial image in the first facial image sequence into a preset facial contour segmentation model for segmentation, to obtain a binary first facial contour image; inputting a first contour image sequence comprising the first facial contour image into a preset facial expression tracking model for prediction, to obtain a facial expression feature corresponding to the first target image, wherein the convolution operation in the facial expression tracking model is a 3D convolution operation in three dimensions of time, width and height. The application improves the accuracy of the facial expression feature extracted in the facial expression tracking technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a facial expression tracking method, a facial expression tracking model training method, an electronic device, a storage medium, and a computer program product. Background Technology

[0002] Facial expression tracking technology is a technique that captures and analyzes a user's facial expressions. In recent years, with technological advancements, facial expression tracking has found increasingly widespread applications in human-computer interaction, virtual reality (VR), and other fields. For example, in virtual social environments, a user's facial expressions can be mirrored onto their avatar in real time for a more immersive social experience; it can also be applied to micro-expression analysis and other scenarios. However, current facial expression tracking technologies suffer from several drawbacks. The extracted facial expression features are not always accurate enough, leading to poor performance in specific applications, such as discontinuous reconstructed facial expressions and inaccurate micro-expression analysis.

[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main objective of this application is to provide a facial expression tracking method, a facial expression tracking model training method, an electronic device, a storage medium, and a computer program product, which aim to improve the accuracy of facial expression features extracted by facial expression tracking technology.

[0005] To achieve the above objectives, this application proposes a facial expression tracking method, which includes: Each frame of facial images continuously acquired from multiple frames of facial images of the target object is taken as the first target image, and a first facial image sequence corresponding to the first target image is obtained. The first facial image sequence includes the first target image and at least one frame of facial image acquired before and / or after the first target image. Each frame of the facial image in the first facial image sequence is input into a preset facial contour segmentation model for segmentation to obtain a binarized first facial contour image of each frame. The facial contour segmentation model is trained in advance using a first training dataset. Each training sample in the first training dataset includes a frame of facial image and corresponding first label data. The first label data is a binarized facial contour image labeled for the corresponding facial image. A first contour image sequence, including each of the first facial contour images, is input into a preset facial expression tracking model for prediction to obtain facial expression features corresponding to the first target image. The convolution operation in the facial expression tracking model is a 3D convolution operation performed in three dimensions: time, width, and height. The facial expression tracking model is pre-trained using a second training dataset. Each training sample in the second training dataset includes a set of second contour image sequences and corresponding second label data. The second contour image sequence includes each frame of second facial contour image. Each second facial contour image is obtained by segmenting each frame of facial image in the second facial image sequence by inputting it into the trained facial contour segmentation model. The second facial image sequence includes a second target image and at least one frame of facial image acquired before and / or after the second target image. The second label data is the full-face expression features of the same object acquired synchronously with the second target image.

[0006] Optionally, each frame of the first facial image sequence includes a sub-image obtained by simultaneously capturing the face of the target object from at least two different angles, and the first facial contour image obtained by segmenting each frame of the facial image includes each binarized facial contour sub-image obtained by segmenting each of the sub-images. The step of inputting a sequence of first contour images, including each of the first facial contour images, into a preset facial expression tracking model for prediction to obtain facial expression features corresponding to the first target image includes: Each facial contour sub-image corresponding to the same angle in the first contour image sequence is used as input data for one channel. The input data of each channel is input into the facial expression tracking model for prediction to obtain facial expression features corresponding to the first target image.

[0007] Optionally, the full-face expression features are a Blendshape parameter set, and the facial expression tracking model includes a first deep residual network and a regression network. The convolution and batch normalization operations in the first deep residual network are 3D convolution and 3D batch normalization operations. The step of inputting a first contour image sequence including each of the first facial contour images into a preset facial expression tracking model for prediction to obtain facial expression features corresponding to the first target image includes: The first contour image sequence is input into the first depth residual network for feature extraction to obtain a feature vector; The feature vector is input into the regression network for prediction to obtain the Blendshape parameter set corresponding to the first target image, which serves as the facial expression feature corresponding to the first target image.

[0008] Optionally, the facial contour segmentation model includes a second depth residual network and a segmentation network. The step of inputting each frame of the facial image in the first facial image sequence into the preset facial contour segmentation model for segmentation to obtain binarized first facial contour images of each frame includes: Each frame of the facial image in the first facial image sequence is used as the fourth target image. The fourth target image is input into the second deep residual network for feature extraction to obtain a feature map. The feature map is input into the segmentation network for segmentation to obtain a binarized first facial contour image corresponding to the fourth target image.

[0009] Optionally, the facial expression tracking method further includes: Acquire multiple frames of facial images of the target object; The multi-frame facial images are sent to a training device so that the training device can generate at least one set of third facial image sequences based on the multi-frame facial images. The facial contour segmentation model trained with the first training dataset is used to segment each frame of the facial images in the third facial image sequence to obtain a third contour image sequence. The third contour image sequence and the Blendshape parameter group of the target object, which is synchronously acquired with the third target image in the third facial image sequence, are used to fine-tune the facial expression tracking model trained with the second training dataset. The third facial image sequence includes the third target image and at least one frame of facial image acquired before and / or after the third target image. The system receives the fine-tuned facial expression tracking model sent by the training device and uses the fine-tuned facial expression tracking model as the preset facial expression tracking model.

[0010] Furthermore, to achieve the above objectives, this application also proposes a facial expression tracking model training method, which includes: Obtain a first training dataset, wherein each training sample in the first training dataset includes a frame of facial image and corresponding first label data, wherein the first label data is a binarized facial contour image annotated for the corresponding facial image. The first training dataset is used to train the facial contour segmentation model to be trained; Acquire at least one set of second facial image sequences and second label data corresponding to the second facial image sequences, wherein the second facial image sequence includes a second target image and at least one frame of facial image acquired before and / or after the second target image, and the second label data is the full-face expression features of the same object acquired synchronously with the second target image; Each frame of the facial image in the second facial image sequence is input into the trained facial contour segmentation model for segmentation to obtain a second contour image sequence, wherein the second contour image sequence includes each frame of the second facial contour image in binarization. The second contour image sequence and the corresponding second label data are used as a training sample. A second training dataset is obtained based on multiple training samples. The facial expression tracking model to be trained is trained using the second training dataset. The convolution operation in the facial expression tracking model is a 3D convolution operation performed in three dimensions: time, width, and height.

[0011] Optionally, the step of training the facial expression tracking model using the second training dataset includes: Each time, a data set containing multiple training samples is extracted from the second training dataset; Based on the proportion of the first type of training samples in the data set, the first weight corresponding to the first type of training samples and the second weight corresponding to the second type of training samples are calculated. The first type of training samples are training samples with facial expressions in the corresponding second target image, and the second type of training samples are training samples without facial expressions in the corresponding second target image. The first weight is greater than the second weight, and the larger the proportion of the first and second weights, the larger they are. Substitute each training sample from the data set into the facial expression tracking model to be trained to calculate the original loss; The original loss of each training sample is weighted and averaged according to the corresponding first weight and second weight to obtain the total loss, and the facial expression tracking model is updated according to the total loss; After iteratively updating the facial expression tracking model using at least one set of the data set in multiple rounds, a trained facial expression tracking model is obtained.

[0012] In addition, to achieve the above objectives, this application also proposes an electronic device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the facial expression tracking method described above, or to implement the steps of the facial expression tracking model training method described above.

[0013] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the facial expression tracking method described above, or implements the steps of the facial expression tracking model training method described above.

[0014] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the facial expression tracking method described above, or implements the steps of the facial expression tracking model training method described above.

[0015] One or more technical solutions proposed in this application have at least the following technical effects: By acquiring multiple consecutive frames of facial images of a target object, a data foundation is provided for continuous facial expression tracking. Based on the objective law that facial expressions typically change gradually over time when a user makes consecutive facial expressions, for each frame, a facial image sequence is generated, containing that frame and at least one previous and / or subsequent frame. This sequence serves as the data basis for extracting facial expression features from that frame, utilizing the temporal changes in facial expressions. This approach yields more accurate facial expression features compared to extracting features from a single frame. The facial images are segmented using a pre-trained facial contour segmentation model to obtain binarized facial contour images. This simplifies the data volume, reducing the computational burden of subsequent facial expression feature extraction, while retaining the facial contour information used for feature extraction, thus ensuring the accuracy of the extracted facial expression features. Facial expression features are predicted by inputting a sequence of contour images from each frame into a pre-trained facial expression tracking model. Furthermore, the convolution operation in the facial expression tracking model is designed to use 3D convolution operations, enabling the prediction of facial expressions based on the temporal changes in facial expressions contained in multiple frames of facial contour images. The feature extraction method employs 3D convolution operations across three dimensions (time, width, and height) to simultaneously and effectively capture temporal and spatial information between image frames. This allows for the accurate prediction of facial expression features using the captured temporal and spatial information. Furthermore, the facial expression tracking model is pre-trained on a training dataset labeled with full-face expression features. This ensures that even if the facial images in the training dataset only include a portion of the face, the model can learn, through the supervision of full-face expression features, how to extract facial expression information from facial contour images containing only a portion of the face. The facial expression information predicts accurate full-face expression features, enabling accurate prediction of full-face expression features even when the target's facial area is occluded or when the camera can only capture facial images including a portion of the face. Furthermore, the training samples consist of contour image sequences including multiple frames of facial contour images. Under the supervision of full-face expression features, the facial expression tracking model learns during training how to extract temporal changes in facial expressions and predict more accurate facial expression features based on these changes, ensuring the continuity of the predicted facial expressions.First, a facial contour segmentation model is trained. Then, the facial contour images segmented by the facial contour segmentation model are used as training data for the facial expression tracking model. This training method also improves the synergy between the facial contour segmentation model and the facial expression tracking model, enabling the facial expression tracking model to more accurately understand the facial contour images segmented by the facial contour segmentation model, and thus predict more accurate full-face expression features based on the contour image sequence. Attached Figure Description

[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating the first embodiment of the facial expression tracking method of this application. Figure 2 This is a schematic diagram of a facial contour image according to one embodiment of this application; Figure 3 This is a schematic diagram of an image preprocessing process provided in one embodiment of this application; Figure 4 This is a flowchart illustrating the fourth embodiment of the facial expression tracking model training method of this application. Figure 5 A schematic diagram of the model architecture and training framework provided in one embodiment of this application; Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the facial expression tracking method in the embodiments of this application.

[0019] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0020] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0021] It should be noted that in the description of this application and the appended claims, the terms "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0022] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0023] Facial expression tracking technology is a technique that captures and analyzes a user's facial expressions. In recent years, with technological advancements, facial expression tracking has found increasingly widespread applications in human-computer interaction, virtual reality (VR), and other fields. For example, in virtual social environments, a user's facial expressions can be mirrored onto their avatar in real time for a more immersive social experience; it can also be applied to micro-expression analysis and other scenarios. However, current facial expression tracking technologies suffer from several drawbacks. The extracted facial expression features are not always accurate enough, leading to poor performance in specific applications, such as discontinuous reconstructed facial expressions and inaccurate micro-expression analysis.

[0024] The facial expression tracking method proposed in this application provides a data foundation for continuous facial expression tracking of a target object by acquiring multiple frames of facial images continuously captured. Based on the objective law that facial expressions typically change gradually over time when a user makes continuous facial expression movements, for each frame of facial image, a facial image sequence is generated, containing that frame of facial image and at least one frame of facial image captured before and / or after it. The facial image sequence is used as the data foundation for extracting facial expression features from that frame of facial image. By utilizing the temporal change information of facial expressions, it can extract more accurate facial expression features compared to extracting facial expression features based on only a single frame of facial image. By segmenting each frame of a facial image sequence using a pre-trained facial contour segmentation model, binarized facial contour images are obtained. This simplifies the data volume, reducing the computational burden of subsequent facial expression feature extraction, while preserving the facial contour information used for feature extraction, thus ensuring the accuracy of the extracted facial expression features. Facial expression features are then predicted by inputting the contour image sequence, including each frame of facial contour images, into a pre-trained facial expression tracking model. Furthermore, the convolution operation in the facial expression tracking model is designed to employ 3D convolution operations, enabling the utilization of temporal changes in facial expressions contained within multiple frames of facial contour images. This method uses information to predict facial expression features. By employing 3D convolution operations in three dimensions (time, width, and height), it effectively captures both temporal and spatial information between image frames, thus predicting accurate facial expression features. Furthermore, the facial expression tracking model is pre-trained on a training dataset labeled with full-face expression features. This allows the model to learn how to extract facial expression information from partial facial contours even when the training dataset only includes a portion of the face. Based on this facial expression information, accurate full-face expression features can be predicted, thus enabling accurate prediction of the full-face expression features of the target object even when the target object's facial area is occluded or when the camera can only capture facial images including a portion of the facial area. Furthermore, the training samples use a sequence of contour images including multiple frames of facial contour images. Under the supervision of full-face expression features, the facial expression tracking model can also learn how to extract temporal changes in facial expressions during the training phase and predict more accurate facial expression features based on this change information, ensuring the continuity of the predicted facial expression movements.First, a facial contour segmentation model is trained. Then, the facial contour images segmented by the facial contour segmentation model are used as training data for the facial expression tracking model. This training method also improves the synergy between the facial contour segmentation model and the facial expression tracking model, enabling the facial expression tracking model to more accurately understand the facial contour images segmented by the facial contour segmentation model, and thus predict more accurate full-face expression features based on the contour image sequence.

[0025] The following presents a first embodiment of the facial expression tracking method of this application. The executing entity of this embodiment can be an electronic device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone. It can also be a wearable device such as headphones or glasses. The following description uses an "expression tracking device" as an example. (Refer to...) Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the facial expression tracking method of this application. In this embodiment, the facial expression tracking method includes steps S10 to S30: Step S10: Take each frame of the multi-frame facial images continuously acquired from the target object as the first target image, and obtain the first facial image sequence corresponding to the first target image. The first facial image sequence includes the first target image and at least one frame of facial image acquired before and / or after the first target image.

[0026] The target object refers to the person whose facial expression tracking needs to be performed, such as a user wearing headphones. Facial images of the target object can be captured using a built-in or external camera in the expression tracking device. The camera can be adjusted to capture a full-face image or a portion of the target object's face; that is, the captured facial image can include the entire or a portion of the target object's face. The camera can be configured to capture facial images at a certain acquisition frequency. The expression tracking device acquires multiple frames of facial images continuously captured by the camera, extracts the corresponding facial expression features from each frame, and thus achieves continuous facial expression tracking.

[0027] In specific implementations, facial expression tracking can be performed in real-time or non-real-time. That is, the expression tracking device can acquire facial images captured by a camera in real time and predict facial expression features frame by frame in real time. Alternatively, the expression tracking device can acquire a pre-captured video segment and predict facial expression features for each frame of the video. In this embodiment, it is not limited to real-time or non-real-time tracking.

[0028] In this embodiment, to improve the accuracy of the extracted facial expression features, for each frame of a facial image, the facial expression features corresponding to that frame are predicted by combining at least one frame of facial images captured before and / or after that frame. Specifically, when a user makes continuous facial expression movements, the expression usually changes gradually over time. By combining at least one frame of facial images captured before and / or after that frame to predict the facial expression features, the temporal change information of facial expressions is utilized, thereby extracting more accurate facial expression features compared to extracting facial expression features based solely on a single frame of facial images.

[0029] Since the process of predicting the corresponding facial expression features for each frame of a facial image is the same, we will use a single frame of a facial image as an example, and refer to this frame as the first target image for distinction. In other words, the expression tracking device uses each frame of a series of continuously acquired facial images as the first target image, extracts the facial expression features corresponding to the first target image, and thus obtains the facial expression features corresponding to each frame of the facial image.

[0030] An image sequence including a first target image and at least one frame of facial image captured before and / or after the first target image is referred to as the first facial image sequence for distinction. It is understood that the first facial image sequence includes, in addition to the first target image, one or more frames of facial image before the first target image, or one or more frames of facial image after the first target image, or one or more frames of facial image before and after the first target image. The specific number of frames included in the first facial image sequence is not limited in this embodiment. For example, one frame of facial image captured before the first target image and one frame of facial image captured after it can be acquired. If the time sequence number of the first target image is T, and P represents the image, then the first facial image sequence corresponding to the first target image can be represented as [P]. T-1 P T P T+1 ].

[0031] In a specific implementation, the number of frames of facial images in the first facial image sequence corresponding to each frame of facial image is the same, for example, 3 frames; the frames of facial images in the first facial image sequence are ordered in time sequence, and the temporal position of each frame of facial image in its corresponding first facial image sequence is the same, for example, the 2nd frame; the number of frames of facial images in the first facial image sequence is the same as the number of time dimensions specified in the input data format of the facial expression tracking model in subsequent steps.

[0032] In a specific implementation, when the first facial image sequence needs to include one or more frames of facial images preceding the first target image, for the first frame of facial image acquired, when generating the first facial image sequence corresponding to that frame, the first frame of facial image can be copied and used as the facial image acquired before that frame, or a pre-set fixed facial image can be used as the facial image acquired before that frame. Similarly, when the first facial image sequence needs to include one or more frames of facial images following the first target image, for the last frame of facial image acquired, when generating the first facial image sequence corresponding to that frame, the last frame of facial image can be copied and used as the facial image acquired after that frame, or a pre-set fixed facial image can be used as the facial image acquired after that frame.

[0033] Step S20: Input each frame of the face image in the first face image sequence into a preset face contour segmentation model for segmentation to obtain a binarized first face contour image of each frame. The face contour segmentation model is trained in advance using a first training dataset. Each training sample in the first training dataset includes a face image and corresponding first label data. The first label data is a binarized face contour image labeled for the corresponding face image.

[0034] The facial expression tracking device can be pre-configured with a facial contour segmentation model for image segmentation of facial images. The facial contour segmentation model can be implemented using a deep learning model structure; this embodiment does not limit the specific model structure used for the facial expression tracking model. The training dataset of the facial contour segmentation model (hereinafter referred to as the first training dataset for distinction) includes multiple training samples. Each training sample includes a frame of facial image and corresponding label data (ground truth), hereinafter referred to as the first label data for distinction. The first label data is a binarized facial contour image labeled for the frame of facial image. The facial contour image has the same size as the frame of facial image. The frame of facial image is a three-channel color image, while the facial contour image is a single-channel image. Its pixels are binarized, for example, they can be either 0 and 1 pixel values ​​or 0 and 255 pixel values. These two pixel values ​​are used to represent whether the image belongs to the facial contour region or not, that is, to distinguish between the facial contour region and the background region in the facial image. By training the facial contour segmentation model using this first training dataset, the trained model can learn how to segment a facial image to obtain a binarized facial contour image of the same size as the original facial image. For example... Figure 2The image shown illustrates a facial contour, where the black area represents the background and the white area represents the facial contour, distinguished by two different pixel values. It should be noted that... Figure 2 The example used here is a facial image containing the entire face region to illustrate facial contour images. This does not mean that the acquired facial image can only be a facial image containing the entire face region.

[0035] The facial expression tracking device first inputs each frame of the first facial image sequence into a pre-trained facial contour segmentation model for segmentation. Each frame of the facial image is segmented to obtain a corresponding binary facial contour image (hereinafter referred to as the first facial contour image for distinction). It should be noted that when a person makes facial expressions, their facial contours will change accordingly. Different facial expressions correspond to different changes in facial contours. By segmenting the facial images to obtain binary facial contour images, the amount of data is reduced on the one hand, and the facial contour information that can be used to extract facial expression features is preserved on the other hand. This ensures the accuracy of the extracted facial expression features while reducing the amount of input data for the facial expression tracking model in subsequent steps, thereby reducing the computational burden of facial expression feature extraction. In other words, a model structure with lower complexity and fewer model parameters can be used to implement the facial expression tracking model.

[0036] The facial contour segmentation model can be pre-trained using a first training dataset, and then the trained facial contour segmentation model can be deployed in an expression tracking device. The training process of the facial contour segmentation model can be performed in the expression tracking device or other devices; this embodiment does not impose any restrictions on this. The method of obtaining the first training dataset is not limited in this embodiment.

[0037] Step S30: Input the first contour image sequence, including each of the first facial contour images, into a preset facial expression tracking model for prediction to obtain facial expression features corresponding to the first target image. The convolution operation in the facial expression tracking model is a 3D convolution operation performed in three dimensions: time, width, and height. The facial expression tracking model is pre-trained using a second training dataset. Each training sample in the second training dataset includes a set of second contour image sequences and corresponding second label data. The second contour image sequence includes each frame of second facial contour image. Each second facial contour image is obtained by segmenting each frame of facial image in the second facial image sequence by inputting it into the trained facial contour segmentation model. The second facial image sequence includes a second target image and at least one frame of facial image acquired before and / or after the second target image. The second label data is the full-face expression features of the same object acquired synchronously with the second target image.

[0038] In this embodiment, the expression tracking device can assemble the first facial contour images corresponding to each frame of the first facial image sequence into an image sequence (hereinafter referred to as the first contour image sequence for distinction). The first contour image sequence corresponding to the first target image is then input into a pre-trained facial expression tracking model for prediction to obtain facial expression features. These facial expression features are then used as the facial expression features corresponding to the first target image. It is understood that by using the same method to obtain the corresponding facial expression features for multiple consecutively acquired facial images, continuous facial expression tracking can be achieved.

[0039] The facial expression tracking model can be implemented using a deep learning model structure; however, this embodiment does not limit the specific model structure used. To achieve facial expression prediction based on a sequence of contour images containing multiple frames of facial contour images, the convolution operation in the facial expression tracking model can be a 3D convolution operation performed in three dimensions: time, width, and height. That is, compared to 2D convolution operations performed in the width and height dimensions of the image, this embodiment, since it needs to process a sequence of contour images containing multiple frames of facial contour images and needs to utilize the temporal change information of facial expressions contained in the multiple frames of facial contour images, uses a 3D convolution operation in three dimensions: time, width, and height. This allows for the simultaneous and effective capture of temporal and spatial information between each image frame, thereby using the captured temporal and spatial information to predict accurate facial expression features. In one feasible implementation, if the structure of the facial expression tracking model also includes a batch normalization (BN) operation, then the batch normalization operation can also be a 3D batch normalization operation, which is suitable for batch normalizing data with three dimensions: time, width, and height.

[0040] The facial expression tracking model can be pre-trained using a training dataset, and then the trained model can be deployed in an expression tracking device. The training process for the facial expression tracking model can be performed on the expression tracking device or on other devices; this embodiment does not impose any restrictions on this.

[0041] The method of obtaining the training dataset is not limited in this embodiment. The training dataset includes multiple training samples. Each training sample includes a set of contour image sequences (hereinafter referred to as the second contour image sequence for distinction) and corresponding label data (ground truth), hereinafter referred to as the second label data for distinction. The second contour image sequence includes at least a second facial contour image. Each frame of the second facial contour image is obtained by inputting each frame of the facial image in the second facial image sequence into the trained facial contour segmentation model for segmentation. The second facial image sequence includes at least two frames of facial images collected for the same object. The second label data is the full-face expression features of the same object collected synchronously with one of the facial images (referred to as the second target image for distinction). It is understood that the other facial images in the second facial image sequence are collected before and / or after the second target image. The second training dataset may include training samples obtained by collecting data for at least one object. The second contour image sequence in each training sample includes the same number of frames of the second facial contour image, for example, 3 frames. Since the number of frames of the second facial contour image in the second contour image sequence is the same as the number of frames of the facial image in the second facial image sequence, the number of frames of the facial image included in each second facial image sequence is also the same. The facial images in each frame of the second facial image sequence are ordered chronologically, and the temporal position of the second target image in each second facial image sequence is the same, for example, the 2nd frame. The number of frames of the facial images included in the second facial image sequence is the same as the number of frames of the facial images included in the first facial image sequence. The temporal position of the second target image in the second facial image sequence is the same as the temporal position of the first target image in the first facial image sequence, for example, the 2nd frame.

[0042] It should be noted that facial expression features refer to features that reflect facial expressions. In this embodiment, there is no limitation on the specific form of features used; for example, facial key points, textures, and expression parameters can be used. Full-face expression features are a special type of facial expression feature that can reflect the expression of the entire facial region of the subject. The method of obtaining full-face expression features is not limited in this embodiment. For example, full-face expression features can be extracted from a full-face image of the same subject (containing the entire facial region of the subject) acquired synchronously with the second target image. Extraction can be done manually or through other automated methods; this is not limited in this embodiment.

[0043] Each frame of the second facial image sequence can contain the entire or a portion of the subject's face. Since facial muscles stretch and contract when a user makes facial expressions, causing changes in facial contours, even images capturing only a portion of the face contain significant facial expression information. Based on this, by using full-face expression features as label data, even if the facial images in the second facial image sequence only include a portion of the face, the facial expression tracking model can learn how to extract facial expression information from the partial facial contour image through the supervision of the full-face expression features. Based on this facial expression information, it can predict accurate full-face expression features. In other words, the facial expression features predicted by the facial expression tracking device using the deployed facial expression tracking model from the first contour image sequence are as close as possible to the accurate full-face expression features of the target object. This allows for accurate prediction of the target object's full-face expression features even when the target object's face is occluded during facial expression tracking, or when the camera of the expression tracking device can only capture facial images including a portion of the face. In addition, the second contour image sequence includes multiple frames of second facial contour images. Under the supervision of full-face expression features, the facial expression tracking model can learn how to understand the evolution of facial expressions over time during the training phase, and thus predict more accurate facial expression features.

[0044] In one feasible implementation, each training sample in the second training dataset can come from multiple collected objects, thereby enabling the trained facial expression tracking model to be applicable to facial expression tracking of different objects. In another feasible implementation, the second training dataset may include training samples collected under various environmental conditions, such as different ambient lighting or different degrees of occlusion, so that the trained facial expression tracking model can accurately predict facial expression features for facial image sequences collected under various environmental conditions.

[0045] In this embodiment, the specific training method of the facial expression tracking model is not limited.

[0046] The predicted facial expression features have many applications, and this embodiment does not limit their uses. For example, in virtual reality (VR), continuous facial tracking technology can enhance the user's interactive experience. By accurately tracking the user's facial expressions, it can improve the user's immersion, which is particularly important for applications such as virtual character video calls and virtual meetings. In the field of silent voice interfaces, continuous facial tracking technology can predict user facial expressions and convert them into text or commands, which can assist people with speech impairments in communication or achieve silent control in noisy environments. In the field of mental health monitoring, continuous facial tracking technology can be combined with continuously extracted facial expression features to identify the user's emotional fluctuations, and can be used for early screening and long-term monitoring of mental illnesses such as depression and anxiety.

[0047] For example, an expression tracking device can reconstruct facial expressions in real time based on predicted facial expression features. The object of expression reconstruction can be a virtual avatar, thus achieving the effect that whatever expression the target object makes, the virtual avatar also makes the same expression. Alternatively, the object of expression reconstruction can also be the target object itself. For instance, in scenarios where video of the target object needs to be captured (such as mobile video call scenarios), if a full-face image cannot be captured due to obstruction or the camera's shooting angle, real-time facial expression tracking can be performed on the target object. Based on the predicted facial expression features, expression reconstruction can be performed on the target object's image, so that the target object in the captured video has an accurate full-face expression.

[0048] In this embodiment, by acquiring multiple frames of facial images continuously captured from the target object, a data foundation is provided for continuous facial expression tracking of the target object. Based on the objective law that facial expressions usually change gradually over time when a user makes continuous facial expression movements, for each frame of facial image, a facial image sequence is generated, containing that frame of facial image and at least one frame of facial image captured before and / or after it. The facial image sequence is used as the data basis for extracting facial expression features from that frame of facial image. By utilizing the temporal change information of facial expressions, more accurate facial expression features can be extracted compared to extracting facial expression features based on only a single frame of facial image. Each frame of facial images in the sequence is segmented using a pre-trained facial contour segmentation model to obtain binarized facial contour images. This simplifies the data volume, reducing the computational burden of subsequent facial expression feature extraction, while retaining the facial contour information used for feature extraction, thus ensuring the accuracy of the extracted facial expression features. Facial expression features are predicted by inputting the contour image sequence, including each frame of facial contour images, into a pre-trained facial expression tracking model. The convolution operation in the facial expression tracking model is designed to use 3D convolution operations, enabling the prediction of facial expression features by utilizing the temporal changes in facial expressions contained in multiple frames of facial contour images. The facial expression tracking model employs 3D convolution operations across three dimensions (time, width, and height) to simultaneously and effectively capture temporal and spatial information between image frames. This allows for the accurate prediction of facial expression features. Furthermore, the model is pre-trained on a training dataset labeled with full-face expression features. This ensures that even if the training dataset contains only partial facial images, the model can learn how to extract facial expression information from partial facial contours through the supervision of full-face expression features. Facial expression information predicts accurate full-face expression features, enabling accurate prediction of full-face expression features even when the target's facial area is occluded or when the camera can only capture facial images including a portion of the face. Furthermore, the training samples consist of contour image sequences containing multiple frames of facial contour images. Under the supervision of full-face expression features, the facial expression tracking model learns during training how to extract temporal changes in facial expressions and predict more accurate facial expression features based on these changes, ensuring the continuity of the predicted facial expressions.First, a facial contour segmentation model is trained. Then, the facial contour images segmented by the facial contour segmentation model are used as training data for the facial expression tracking model. This training method also improves the synergy between the facial contour segmentation model and the facial expression tracking model, enabling the facial expression tracking model to more accurately understand the facial contour images segmented by the facial contour segmentation model, and thus predict more accurate full-face expression features based on the contour image sequence.

[0049] Based on the first embodiment described above, a second embodiment of the facial expression tracking method of this application is proposed. In this embodiment, content that is the same as or similar to that in the first embodiment can be referred to the above description and will not be repeated hereafter. In this embodiment, to further improve the accuracy of the extracted facial expression features, facial images of the target object can be captured from at least two angles simultaneously. Therefore, each continuously acquired frame of facial image can include images (hereinafter referred to as sub-images) captured simultaneously from at least two different angles of the target object's face. To achieve capturing from at least two angles, at least two cameras can be used to capture from different angles. Therefore, the at least two sub-images captured at the same time can be captured separately by each camera. The camera system can be configured to stitch the sub-images into one image and output it to the expression tracking device, which then splits it to obtain individual sub-images. Alternatively, it can be configured to output each sub-image separately to the expression tracking device, which then matches the sub-images captured at the same time based on the timestamp. Accordingly, each frame of the first facial image sequence includes sub-images of the target object's face captured simultaneously from at least two different angles. Therefore, the first facial contour image obtained by segmenting each frame of the first facial image sequence also includes binarized facial contour images (referred to as facial contour sub-images for distinction) obtained by segmenting each sub-image. For example, if a frame of facial image includes two sub-images, facial contour segmentation is performed on the two sub-images to obtain two facial contour sub-images. The first facial contour image corresponding to this frame of facial image includes these two facial contour sub-images.

[0050] For example, when the first facial image sequence includes a first target image, and one frame of facial image acquired before and one frame of facial image acquired after the first target image, the first facial image sequence can be represented as [P T-1 P T P T+1 If each frame of a facial image includes a sub-image taken from the left side of the face and a sub-image taken from the right side of the face, then P T-1 P T and P T+1 Each image consists of two sub-images. The first facial image sequence can be represented as [[P] T-1Left, P T-1 [Right], [P] T Left, P T [Right], [P] T+1 Left, P T+1 _Right]], after performing facial contour segmentation on each sub-image, we obtain each facial contour sub-image. The first contour image sequence including each facial contour sub-image can be represented as [[P]] T-1 _Left 轮廓 P T-1 _right 轮廓 ], [P T _Left 轮廓 P T _right 轮廓 ], [P T+1 _Left 轮廓 P T+1 _right 轮廓 ]).

[0051] In one feasible implementation, step S30 includes S311: taking each of the facial contour sub-images corresponding to the same angle in the first contour image sequence as input data for one channel, and inputting the input data of each channel into the facial expression tracking model for prediction to obtain facial expression features corresponding to the first target image.

[0052] If each facial contour sub-image corresponding to the same angle in the first contour image sequence is taken as input data for one channel, then the number of channels in the input data of the facial expression tracking model is the number of shooting angles. For example, if each frame of facial image includes sub-images taken from two angles, then the number of input data channels is 2; if each frame of facial image includes sub-images taken from three angles, then the number of input data channels is 3. It can be understood that the input data of the facial expression tracking model is a 4D tensor, whose dimensions can be represented as [H, W, T, C], where C represents the number of channels, T represents the number of frames (the time dimension), W represents the image width, and H represents the image height.

[0053] For example, if the grayscale sub-images obtained by processing the first facial image sequence are represented as [[P] T-1 _Left 轮廓 P T-1 _right 轮廓 ], [P T _Left 轮廓 P T _right 轮廓 ], [P T+1 _Left 轮廓 P T+1 _right 轮廓 Then, the input data for the facial expression tracking model can be represented as [[P]]. T-1 _Left轮廓 P T _Left 轮廓 P T+1 _Left 轮廓 ], [P T-1 _right 轮廓 P T _right 轮廓 P T+1 _right 轮廓 Its dimensions are [H, W, 3, 2], [P] T-1 _Left 轮廓 P T _Left 轮廓 P T+1 _Left 轮廓 [P] represents a channel. T-1 _right 轮廓 P T _right 轮廓 P T+1 _right 轮廓 [This is another channel.] Figure 3 As shown, the left and right facial images captured at times T-1, T, and T+1 are segmented using a facial contour segmentation model to obtain individual facial contour sub-images. Each facial contour sub-image is then divided into channels according to the shooting angle and combined to obtain a 2-channel image sequence, which serves as the input data for the facial expression tracking model.

[0054] It should be noted that each frame of the facial image sequence also includes sub-images of the subject's face captured simultaneously from at least two angles, and the number of sub-images in each frame is the same as the number of sub-images in the facial images of the first facial image sequence, for example, two sub-images. Accordingly, each frame of the second facial contour image sequence also includes at least two facial contour sub-images.

[0055] In one feasible implementation, the expression tracking device can be an eyeglass device; that is, the facial expression tracking method can be applied to the eyeglass device. The eyeglass device has two cameras with different shooting angles. One camera can be used to capture an image of the user's left side of their face from the left, and the other can be used to capture an image of the user's right side of their face from the right. Therefore, it can be understood that each frame of the facial image sequence includes two sub-images, which are simultaneously captured by the two cameras located in the eyeglass device.

[0056] When collecting the training dataset, the subject can also wear glasses and the two cameras on the glasses can capture images of the subject's left and right sides of the face. In the second facial image sequence generated based on the images captured by the glasses, each frame of the facial image also includes two sub-images, which are captured simultaneously by the two cameras set on the glasses.

[0057] Traditional facial tracking systems utilize computer vision (CV) technology to reconstruct facial expressions by analyzing the user's face captured by a front-facing camera. This method offers excellent performance when the camera is unobstructed. However, the setup requirements of this method are quite demanding in real-world scenarios. It severely restricts the user's range of motion, requiring the user to remain constantly in front of the camera, leading to excessive restriction and compromising user privacy by placing the user in a state of constant surveillance. Furthermore, it may not function properly if the user is moving or outdoors, as it is difficult to set up a front-facing camera. In this embodiment, a camera mounted on a glasses device captures partial facial images of the user, and a facial expression tracking model pre-trained with full-face expression features as labeled data predicts facial expression features based on the partial facial images. This allows facial expression tracking without requiring the user to remain constantly in front of the camera, protecting user privacy and expanding the application scenarios, enabling normal operation even when the user is moving or outdoors. For example, to provide facial feedback during mobile video calls, users currently have to hold their phones and point the camera at their faces. In this embodiment, however, in mobile video call applications, the glasses device can predict full-face expression features based on images of a portion of the user's face, and reconstruct a virtual avatar or the user's own full-face expression based on these features. This is equivalent to automatically recording the user's facial expressions and presenting them in a hands-free manner. Therefore, the communication experience can be greatly improved when the user is carrying food, washing dishes, jogging, etc.

[0058] In another feasible embodiment, the facial expression tracking device can also be an earphone device. The earphone device has two cameras with different shooting angles. The left and right earpieces can each have a camera to capture images (sub-images) of the user's face from two different angles. Since the earpieces are worn on the ears, the two sub-images captured at the same time are the image of the user's left and right face, respectively. The specific implementation of the facial expression tracking method on the earphone device in the various embodiments of this application can be referred to the specific implementation on the glasses device, and will not be repeated here.

[0059] Based on the first and / or second embodiments described above, a third embodiment of the facial expression tracking method of this application is proposed. In this embodiment, content that is the same as or similar to the first and second embodiments described above can be referred to the above description and will not be repeated hereafter. In this embodiment, the form of facial expression features can be a Blendshape parameter group. The Blendshape parameter group includes multiple Blendshape parameters, for example, 52 Blendshape parameters. The Blendshape parameter group is a set of scalar values ​​from 0 to 1. Each Blendshape parameter precisely controls the degree of interpolation (Blend / Interpolate) of a specific target shape relative to the base shape. By blending the target shape and the base shape through parameters, complex facial expressions can be synthesized. Correspondingly, the full-face expression features as the second label data can be the Blendshape parameter group of the frontal face of the same object, which is synchronously acquired with the second target image. In a feasible implementation, a mobile phone with Blendshape parameter group acquisition function can be used to synchronously acquire the Blendshape parameter group of the same object with the device acquiring the second target image.

[0060] This embodiment provides a specific structure for a facial expression tracking model when using the Blendshape parameter set for facial expression features. The facial expression tracking model can be designed to include a deep residual network (ResNet) and a regression network (i.e., a regression head). This deep residual network is referred to as the first deep residual network for distinction. The convolutional operations in the first deep residual network can be 3D convolutional operations, and the batch normalization operation can be 3D batch normalization. Since the input is a sequence of contour images, ordinary 2D convolutions cannot simultaneously extract and process additional temporal information. Therefore, all 2D convolutions in the first deep residual network are replaced with 3D convolutions, and the BN layers are replaced with 3D BN layers. 3D convolutions can effectively capture both temporal and spatial information between consecutive image frames, and the ResNet architecture itself can also extract image features well, thereby further improving the accuracy of predicted facial expression features.

[0061] In specific implementations, when the facial expression tracking method is applied to wearable devices, considering the limitations of the computing resources of wearable devices, the first deep residual network can be ResNet34. This adapts to the computing resources of wearable devices, reduces the computational burden and shortens the computation time, while also maximizing the accuracy of facial expression tracking. When applied to other types of devices, other models of ResNet, such as ResNet18 and ResNet50, can also be used.

[0062] When designing the first deep residual network, the last fully connected layer of ResNet can be removed, that is, only the feature vector output by the penultimate layer of ResNet is needed to be input into the regression network.

[0063] The regression network may include at least one fully connected layer to convert the dimension of the feature vector output by the first deep residual network to the same number of parameters as in the Blendshape parameter set. That is, the dimension of the vector output by the last fully connected layer of the regression network is the same as the number of Blendshape parameters in the Blendshape parameter set, for example, both being 52. The regression network also includes an activation function (such as Sigmoid) following the last fully connected layer to map the values ​​output by the last fully connected layer to the range 0-1, thereby obtaining the various Blendshape parameters.

[0064] In one feasible implementation, a first deep residual network can be designed to output a 512-dimensional feature vector; the regression network includes a fully connected layer to reduce the 512-dimensional feature vector output by the first deep residual network to 52 dimensions. In this way, the regression network only needs to include a fully connected layer, which reduces the overall number of parameters of the facial expression tracking model and avoids the problem of feature loss caused by directly reducing from a higher dimension to 52 dimensions.

[0065] In one feasible embodiment, step S30 includes S321~S322: Step S321: Input the first contour image sequence into the first depth residual network for feature extraction to obtain a feature vector.

[0066] Step S322: Input the feature vector into the regression network for prediction to obtain the Blendshape parameter group corresponding to the first target image, which serves as the facial expression feature corresponding to the first target image.

[0067] The facial expression tracking device can first input the first contour image sequence into the first deep residual network for feature extraction to obtain feature vectors, and then input the feature vectors into the regression network for prediction to obtain multiple regression values ​​in the range of 0-1. The regression values ​​are then combined into a Blendshape parameter group as the facial expression features corresponding to the first target image.

[0068] In one feasible implementation, the facial contour segmentation model may include a deep residual network (hereinafter referred to as the second deep residual network for distinction) and a segmentation network (i.e., a segmentation head). The first and second deep residual networks can be implemented using the same type of deep residual network (ResNet), or different types of ResNet, depending on the requirements. Since the tasks of the facial expression tracking model and the facial contour segmentation model are different, the output data of the first deep residual network differs in format: the output data of the first deep residual network is in the form of feature vectors, while the output data of the second deep residual network is in the form of feature maps.

[0069] In specific implementations, when the facial expression tracking method is applied to wearable devices, considering the limitations of the computing resources of wearable devices, the second deep residual network can be ResNet18. This adapts to the computing resources of wearable devices, reducing the computational burden and shortening the computation time, while also maximizing the accuracy of facial contour segmentation. For applications in other types of devices, other ResNet models, such as ResNet34 and ResNet50, can also be used.

[0070] In designing a second deep residual network, the original output layer of ResNet can be removed, meaning only the feature maps output by each convolutional layer or the last convolutional layer of ResNet are needed as input to the segmentation network.

[0071] The segmentation network upsamples the feature map output by the second deep residual network back to the original input size, i.e., the size of the facial image, and predicts the category of each pixel (e.g., 1 indicates belonging to the facial contour region, and 0 indicates not belonging to the facial contour region), resulting in a binarized facial contour image. The pixel value of each point in this binarized facial contour image can directly use the predicted category, such as 0 or 1; or, alternatively, the two different predicted categories can be mapped to two different grayscale values, for example, category 1 mapped to grayscale value 255, and category 0 mapped to grayscale value 0, resulting in a binarized facial contour image where the facial contour region is black and the background region is white. In one feasible implementation, the segmentation network can be implemented using a Feature Pyramid Network (FPN).

[0072] In one feasible embodiment, step S20 includes S211~S212: Step S211: Each frame of the face image in the first face image sequence is taken as the fourth target image, and the fourth target image is input into the second deep residual network for feature extraction to obtain a feature map.

[0073] Step S212: Input the feature map into the segmentation network for segmentation to obtain a binarized first facial contour image corresponding to the fourth target image.

[0074] The steps for facial contour segmentation are the same for each frame of the first facial image sequence. Therefore, we take one frame of the facial image as an example and refer to it as the fourth target image for distinction. The expression tracking device inputs the fourth target image into the second deep residual network for feature extraction to obtain a feature map. Then, the feature map is input into the segmentation network for segmentation to obtain a binarized first facial contour image corresponding to the fourth target image.

[0075] In one feasible embodiment, the facial expression tracking method further includes steps S40-S60: Step S40: Acquire multiple frames of facial images of the target object.

[0076] In this embodiment, the facial expression tracking model deployed in the expression tracking device can be a personalized facial expression tracking model that is finely tuned for a specific user (i.e., the target object), thereby improving the accuracy of the facial expression tracking model in predicting the facial expression features of new users, thus giving new users a good experience.

[0077] To enable personalized fine-tuning of the facial expression tracking model, the expression tracking device can acquire multiple frames of facial images of the target object. It should be noted that the multiple frames of facial images acquired at this time are for personalized fine-tuning and are different from the facial images acquired in step S10 for facial expression feature prediction.

[0078] In this embodiment, the number of frames of facial images acquired for personalized fine-tuning is not limited. For example, it can be set to continuously acquire a preset duration (e.g., 1 minute). Then, the number of frames of facial images acquired = preset duration × acquisition frame rate.

[0079] Step S50: The multi-frame facial images are sent to the training device so that the training device can generate at least one set of third facial image sequences based on the multi-frame facial images. The facial contour segmentation model trained with the first training dataset is used to segment each frame of the facial images in the third facial image sequence to obtain a third contour image sequence. The third contour image sequence and the Blendshape parameter set of the target object, which is synchronously acquired with the third target image in the third facial image sequence, are used to fine-tune the facial expression tracking model trained with the second training dataset. The third facial image sequence includes the third target image and at least one frame of facial image acquired before and / or after the third target image.

[0080] The facial expression tracking device can send multiple frames of captured facial images to a training device. The training device can be a device with model training capabilities, such as a PC (Personal Computer) or a cloud server. The training device generates at least one set of facial image sequences (hereinafter referred to as the third facial image sequence) based on the multiple frames of facial images. Each set of third facial image sequences includes at least two frames of facial images, one of which is called the third target image. The other facial images are those captured before and / or after the third target image. The training device inputs each frame of the third facial image sequence into a trained facial contour segmentation model for segmentation, obtaining a binary facial contour image (hereinafter referred to as the third facial contour image) corresponding to each frame. These third facial contour images are then combined into an image sequence (hereinafter referred to as the third contour image sequence for distinction). The training device obtains the Blendshape parameter set of the target object captured synchronously with the third target image, using it as the label data for the third contour image sequence (hereinafter referred to as the third label data), and fine-tunes the facial expression tracking model trained using a second training dataset. It is understandable that if the Blendshape parameter set of the target object is simultaneously acquired while the facial image is being captured by the expression tracking device, then a third set of facial image sequences can be generated for each frame of facial image captured by the expression tracking device, and the Blendshape parameter set acquired synchronously with each frame of facial image can be used as the corresponding label data.

[0081] It should be noted that the number of frames of the third facial contour images included in each third contour image sequence is the same, for example, 3 frames. Since the number of frames of the third facial contour images in the third contour image sequence is the same as the number of frames of the facial images in the third facial image sequence, the number of frames of the facial images included in each third facial image sequence is also the same. The facial images in each frame of the third facial image sequence are ordered in chronological order, and the chronological position of the third target image in each third facial image sequence is the same, for example, the second frame. The number of frames of the facial images included in the third facial image sequence is the same as the number of frames of the facial images included in the first facial image sequence. The chronological position of the third target image in the third facial image sequence is the same as the chronological position of the first target image in the first facial image sequence, for example, the second frame. In the case that each frame of the face image in the first face image sequence includes multiple sub-images, each frame of the face image in the third face image sequence also includes sub-images obtained by simultaneously capturing the face of the target object from at least two angles, and the number of sub-images included in each frame of the face image is the same as the number of sub-images included in the face image in the first face image sequence, for example, 2 sub-images.

[0082] There are many methods for fine-tuning facial expression tracking models, and this implementation does not impose any limitations. For example, LoRa (Low-Rank Adaptation) can be used for fine-tuning, which can significantly improve the effect of personalized fine-tuning while introducing a small number of parameters, providing a good experience for new users. In one feasible implementation, fine-tuning can be performed only on the regression network of the facial expression tracking model, that is, keeping the parameters in the first deep residual network unchanged and updating the parameters of the fully connected layers in the regression network.

[0083] It should be noted that the loss calculation method used in the fine-tuning process of the facial expression tracking model can be the same as the loss calculation method used in the training process.

[0084] The Blendshape parameter set of the target object can be obtained through a device that supports Blendshape parameter set extraction. This device extracts the Blendshape parameter set from an image of the target object's frontal face, reflecting the facial expressions of the entire target object. In one feasible implementation, when the expression tracking device is a pair of glasses, the Blendshape parameter set can be obtained through other devices (such as a mobile phone). An interactive data acquisition program can be developed in the training device. The user can wear the glasses and adjust the camera angle of the mobile phone to face their frontal face. In the training device (such as a PC), a command to start data acquisition is triggered through the interactive window of the data acquisition program. The training device responds to this command, controlling the glasses and the mobile phone to start data acquisition simultaneously. The image acquisition frame rate of the glasses is the same as the frequency at which the mobile phone acquires the Blendshape parameter set. The glasses will capture each frame... Each frame of facial image (with timestamp) is sent to the training device, and the mobile phone sends the collected Blendshape parameter sets (with timestamps) to the training device. The user can trigger multiple acquisition commands, and the training device responds to each command triggered by the user, controlling the glasses device and the mobile phone to perform multiple data acquisitions. The training device receives the data transmitted by the glasses device and the mobile phone, matches facial images with the same timestamp with Blendshape parameter sets, and generates a third facial image sequence corresponding to each frame of facial image. The Blendshape parameter sets that match the timestamps of each frame of facial image are used as the corresponding label data for fine-tuning the facial expression tracking model.

[0085] Step S60: Receive the fine-tuned facial expression tracking model sent by the training device, and use the fine-tuned facial expression tracking model as the preset facial expression tracking model.

[0086] After fine-tuning the facial expression tracking model, the training device can send the fine-tuned facial expression tracking model to the expression tracking device. The expression tracking device saves the fine-tuned facial expression tracking model and uses it to predict facial expression features of subsequent facial images collected for the target object.

[0087] Based on the first, second, and / or third embodiments described above, a fourth embodiment of the facial expression tracking model training method of this application is proposed. In this embodiment, content that is the same as or similar to the first, second, and third embodiments described above can be referred to the above description and will not be repeated hereafter. The executing entity of this embodiment can be an electronic device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, server, etc. The following description uses a "training device" as an example. In this embodiment, refer to... Figure 4 The facial expression tracking model training method includes A10~A50: Step A10: Obtain the first training dataset, wherein each training sample in the first training dataset includes a frame of facial image and corresponding first label data, and the first label data is a binarized facial contour image annotated for the corresponding facial image.

[0088] Step A20: Train the facial contour segmentation model to be trained using the first training dataset.

[0089] Step A30: Obtain at least one set of second facial image sequences and second label data corresponding to the second facial image sequences, wherein the second facial image sequence includes a second target image and at least one frame of facial image acquired before and / or after the second target image, and the second label data is the full-face expression features of the same object acquired synchronously with the second target image.

[0090] Step A40: Input each frame of the facial image in the second facial image sequence into the trained facial contour segmentation model for segmentation to obtain a second contour image sequence, wherein the second contour image sequence includes binarized frames of the second facial contour image.

[0091] Step A50: The second contour image sequence and the corresponding second label data are used as a training sample. A second training dataset is obtained based on multiple training samples. The facial expression tracking model to be trained is trained using the second training dataset. The convolution operation in the facial expression tracking model is a 3D convolution operation performed in three dimensions: time, width, and height.

[0092] The specific implementation of steps A10 to A50 in this embodiment can refer to the specific implementation of steps S10 to S30 above, and will not be repeated here.

[0093] In one feasible implementation, due to factors such as individual differences in wearing position, the model's performance will vary depending on the wearing position. To improve the model's robustness, data augmentation strategies such as rotation, translation, and scaling can be employed. The numerical range of all data augmentation strategies (e.g., rotation angle, translation amount, scaling ratio, etc.) can conform to a pre-defined Gaussian distribution or follow pre-defined fixed values. When training with each batch of training samples, the application of a data augmentation strategy can be determined according to a pre-defined probability, for example, a probability of 50% or 60%. When performing data augmentation on each frame of the facial image sequence, the same strategy is used, such as rotating them all by 90°, to maintain the temporal and spatial correlation between the frames of facial images and avoid data augmentation disrupting this correlation.

[0094] In one feasible implementation, the step A50 of training the facial expression tracking model to be trained using the second training dataset includes A511~A515: Step A511: Take out a data group containing multiple training samples from the second training dataset each time.

[0095] The training process for a facial expression tracking model can involve iterative training using multiple batches of training data, each batch being called a data set. A data set containing multiple training samples can be extracted from a second training dataset each time, and one data set is used to train the facial expression tracking model each time. Each training iteration builds upon the previous training iteration, hence the term iterative training. The following explanation uses the process of training a facial expression tracking model using a single data set as an example.

[0096] Step A512: Based on the proportion of the first type of training samples in the data group, calculate the first weight corresponding to the first type of training samples and the second weight corresponding to the second type of training samples. The first type of training samples are training samples with facial expressions in the corresponding second target image, and the second type of training samples are training samples without facial expressions in the corresponding second target image. The first weight is greater than the second weight, and the larger the proportion of the first and second weights, the larger they are.

[0097] Each training sample in the second training dataset can be pre-labeled with a category, i.e., divided into a first category and a second category. The first category consists of training samples in the corresponding second target image that exhibit facial expressions, while the second category consists of training samples in the corresponding second target image that do not exhibit facial expressions. The second target image corresponding to a training sample refers to the second target image in the second facial image sequence included in that training sample. Facial expressions mean that the subject in the image makes at least one facial expression, such as "smiling" or "raising an eyebrow," while no facial expressions mean that the subject in the image does not make any facial expressions. The category labeling of each training sample can be done manually, or a small number of training samples can be manually labeled, and then a binary classification model can be trained using these labeled training samples. The trained binary classification model can then be used to classify each training sample to obtain its category.

[0098] The proportion of the first type of training samples in the acquired data set can be calculated, that is, the ratio of the number of first type training samples to the total number of training samples in the data set. Based on this proportion, the weights corresponding to the first type of training samples (hereinafter referred to as the first weights for distinction) and the weights corresponding to the second type of training samples (hereinafter referred to as the second weights for distinction) are calculated.

[0099] The quantity percentage and the first and second weights must satisfy the following relationship: the larger the quantity percentage, the larger the first weight; and the larger the quantity percentage, the larger the second weight; at the same time, the first weight is greater than the second weight. This implementation method does not limit the specific calculation method of the first and second weights based on the quantity percentage, as long as the calculated first and second weights conform to the above relationship with the quantity percentage.

[0100] For example, when training a facial expression tracking model that includes a first-level deep residual network and a regression network, Huber loss can be used to ensure a robust regression. Huber loss combines the advantages of mean squared error (MSE) and mean absolute error (MAE), providing some robustness to outliers. The facial expression tracking model... Loss can be defined as:

[0101]

[0102]

[0103]

[0104] Where B is the batch size, i.e. the number of training samples in the data set, y is the second label data, and f(x) is the predicted value of the facial expression tracking model. This represents the first weight, and K is the number of all facial expression types in the second training dataset. It is the number of training samples of the first type. This represents the number of training samples of the second class. When the i-th training sample is a training sample of the first class... The value is 1. The value is 0 when the i-th training sample is a training sample of the second class. The value is 0. The value is 1. This is a hyperparameter used to control the switching point of Huber loss between MSE and MAE. It occurs when the absolute value of the error between the predicted and actual values ​​is less than or equal to... When the absolute value of the error is greater than 1, use mean squared error; when the absolute value of the error is greater than 1. When using linear functions, the impact of outliers on the loss can be reduced.

[0105] For example, such as Figure 5 As shown, the facial contour segmentation model can be implemented using ResNet18 and FPN, i.e., ResNet18-FPN in the figure; the facial expression tracking model can include ResNet34-3D and a regression network, where ResNet34-3D is ResNet34 using 3D convolution operations and 3D batch normalization operations; when training the facial expression tracking model, each frame of the facial image sequence is input into the pre-trained facial contour segmentation model for segmentation to obtain a contour image sequence, the contour image sequence is input into the facial expression tracking model to predict the Blendshape parameter set, the Huber loss is calculated using the label data corresponding to the facial image sequence, the model parameters in the facial expression tracking model are updated according to the Huber loss, and the model parameters in the facial expression tracking model are updated in multiple rounds by using multiple training samples to obtain the trained facial expression tracking model.

[0106] Step A513: Substitute each training sample in the data set into the facial expression tracking model to be trained to calculate the original loss.

[0107] For each training sample in the data set, the second contour image sequence in the training sample is input into the facial expression tracking model to be trained for prediction, and the facial expression features corresponding to the second target image are obtained. Based on the facial expression features and the second label data in the training sample, the loss corresponding to the training sample (hereinafter referred to as the original loss) is calculated according to the preset loss calculation method.

[0108] Step A514: The original loss of each training sample is weighted and averaged according to the corresponding first weight and second weight to obtain the total loss, and the facial expression tracking model is updated according to the total loss.

[0109] The raw losses of each training sample are weighted and averaged according to their respective weights. Specifically, the raw losses of the first type of training samples are weighted using their first weights, and the raw losses of the second type of training samples are weighted using their second weights. The facial expression tracking model is then updated based on the total loss obtained from the weighted average; specifically, the model parameters within the facial expression tracking model are updated.

[0110] Using a dataset allows for at least one round of iterative updates to the facial expression tracking model.

[0111] Step A515: After iteratively updating the facial expression tracking model using at least one set of the data set for multiple rounds, a trained facial expression tracking model is obtained.

[0112] In this embodiment, the loss weighting weights corresponding to training samples with facial expressions and those without are calculated based on the proportion of training samples with facial expressions in the data set. When the proportion of training samples with facial expressions is high, the loss weights of each training sample in the data set are higher, increasing the contribution of the data set to the training process of the facial expression tracking model. At the same time, the loss weighting weights corresponding to training samples with facial expressions are greater than those corresponding to training samples without facial expressions, further increasing the contribution of training samples with facial expressions to the training process of the facial expression tracking model. As a result, the facial expression tracking model can more fully learn and understand various facial expressions during the entire training process, thereby predicting more accurate facial expression features.

[0113] In one feasible embodiment, step A30 includes A311 to A314: Step A311: Send a synchronous data acquisition command to the glasses device and the target mobile phone so that the glasses device and the target mobile phone can synchronously acquire facial images and Blendshape parameter groups for the same object, wherein the target mobile phone is a mobile phone that supports the Blendshape parameter group extraction function.

[0114] The facial expression tracking device can be a pair of glasses; that is, the trained facial expression tracking model can be deployed in a pair of glasses. Therefore, during the training dataset acquisition phase, facial images can be acquired through the glasses. A mobile phone supporting Blendshape parameter set extraction (hereinafter referred to as the target phone for distinction) can be used to acquire the Blendshape parameter set. For example, a mobile phone equipped with a TrueDepth camera and ARKit can be used as the target phone. The TrueDepth camera is a highly integrated front-facing camera and sensor system whose core capability is to provide the device with accurate facial depth perception and 3D modeling capabilities. ARKit is an augmented reality (AR) development platform.

[0115] A data acquisition program can be developed within the training device to control the synchronous data acquisition of the glasses device and the target mobile phone, and to control the duration of synchronous data acquisition, such as continuous acquisition for 1 minute. The method of controlling the synchronous data acquisition of the glasses device and the target mobile phone can be, for example, establishing a communication connection between the training device, the glasses device, and the target mobile phone, such as connecting to the same wireless network; the training device can send a synchronous data acquisition command to the glasses device and the target mobile phone, so that the glasses device and the target mobile phone simultaneously begin acquiring facial images and Blendshape parameter sets at the same frequency according to the data acquisition command, thereby achieving synchronous acquisition of facial images and Blendshape parameter sets.

[0116] During data acquisition, the subject can wear glasses and the target phone's camera angle can be adjusted to focus on the subject's face. This allows both the glasses and the target phone to capture the same facial image and Blendshape parameter set, ensuring that the Blendshape parameter set reflects the facial expressions of the entire subject's face.

[0117] Step A312: Receive timestamped facial images of each frame sent by the glasses device, and receive timestamped Blendshape parameter groups sent by the target mobile phone.

[0118] The glasses device sends each frame of facial image captured, along with a timestamp, to the training device. The target phone sends each set of Blendshape parameters captured, along with a timestamp, to the training device. For example, the captured data can be transmitted to the training device via UDP (User Datagram Protocol). When collecting data from multiple subjects, the facial images and Blendshape parameter sets sent by the glasses device and the target phone can also include identification information for each subject to distinguish them.

[0119] Step A313: Take each frame of the facial images in each frame as the second target image and generate the second facial image sequence corresponding to the second target image.

[0120] After acquiring each frame of facial images collected for the same target, the training device generates a facial image sequence (second facial image sequence) corresponding to each frame of facial images. That is, each frame of facial images is used as a second target image to generate a second facial image sequence corresponding to the second target image.

[0121] Specifically, the training device can combine at least one frame of facial image before and / or after the second target image with the second target image to form a corresponding second facial image sequence.

[0122] Step A314: The Blendshape parameter group that has the same timestamp as the second target image is used as the second label data corresponding to the second facial image sequence.

[0123] The training device matches facial images and Blendshape parameter groups with the same timestamp based on the timestamp of the facial image and the timestamp of the Blendshape parameter group. The Blendshape parameter group with the same timestamp as the second target image is used as the second label data of the second facial image sequence corresponding to the second target image, which is also used as the second label data of the second contour image sequence obtained subsequently.

[0124] In one feasible implementation, when performing personalized fine-tuning of the facial expression tracking model, the specific implementation method for obtaining the third facial image sequence and the corresponding third label data can refer to the specific implementation method for obtaining the second facial image sequence and the corresponding second label data in steps A311 to A314 above, and simply replace the second facial image sequence and the second label data with the third facial image sequence and the third label data.

[0125] In one feasible implementation, during the personalized fine-tuning of the facial expression tracking model, the first type of training samples and the second type of training samples are also distinguished. When different weights are used for loss weighting, the training device can obtain the category label corresponding to each third facial image sequence, that is, obtain the category label indicating whether the third target image in the third facial image sequence has facial expression action. An interactive data acquisition program is developed for the training device. This program includes a customizable facial expression / action function, supporting users to define various types of facial expressions / actions. Users can input or select the type of facial expression / action (if no type is input or selected, the default is no facial expression / action) through the interactive window of the data acquisition program in the training device to trigger the custom facial expression / action command. The training device responds to the custom command by generating a prompt message and outputting the prompt message through the interactive window to prompt the user to perform that type of facial expression / action or not to perform any facial expression / action. After outputting the prompt message, the training device can control the headset and mobile phone to simultaneously start acquiring data to obtain a third facial image sequence and corresponding third label data. This third facial image sequence and corresponding third label data are used as a training sample for fine-tuning, and the corresponding type of facial expression / action is labeled or labeled as no facial expression / action. The training device can determine whether the training sample belongs to the first type or the second type of training sample based on the labeling of the training sample.

[0126] This application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the facial expression tracking method or facial expression tracking model training method in the above embodiments.

[0127] The following is for reference. Figure 6 This document illustrates a structural diagram of an electronic device suitable for implementing embodiments of this application. The electronic devices in these embodiments may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), wearable devices, and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0128] like Figure 6As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. The communication device 1009 allows the electronic device to communicate wirelessly or wiredly with other devices to exchange data. Although the diagrams show electronic devices with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.

[0129] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0130] The electronic device provided in this application adopts the facial expression tracking method or facial expression tracking model training method in the above embodiments. Compared with the prior art, the beneficial effects of the electronic device provided in this application are the same as those of the facial expression tracking method or facial expression tracking model training method provided in the above embodiments. Furthermore, the other technical features of the electronic device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.

[0131] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0132] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0133] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the facial expression tracking method or facial expression tracking model training method in the above embodiments.

[0134] The computer-readable storage medium provided in this application embodiment may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0135] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.

[0136] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by an electronic device, cause the electronic device to perform the functions defined in the methods of the embodiments disclosed in this application.

[0137] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0138] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0139] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0140] The readable storage medium provided in this application embodiment is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-described facial expression tracking method or facial expression tracking model training method. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application embodiment are the same as the beneficial effects of the facial expression tracking method or facial expression tracking model training method provided in the above-described embodiments, and will not be repeated here.

[0141] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the facial expression tracking method or facial expression tracking model training method described above.

[0142] Compared with the prior art, the beneficial effects of the computer program product provided in this application embodiment are the same as the beneficial effects of the facial expression tracking method or facial expression tracking model training method provided in the above embodiments, and will not be repeated here.

[0143] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A facial expression tracking method, characterized in that, The facial expression tracking method includes: Each frame of facial images continuously acquired from multiple frames of facial images of the target object is taken as the first target image, and a first facial image sequence corresponding to the first target image is obtained. The first facial image sequence includes the first target image and at least one frame of facial image acquired before and / or after the first target image. Each frame of the facial image in the first facial image sequence is input into a preset facial contour segmentation model for segmentation to obtain a binarized first facial contour image of each frame. The facial contour segmentation model is trained in advance using a first training dataset. Each training sample in the first training dataset includes a frame of facial image and corresponding first label data. The first label data is a binarized facial contour image labeled for the corresponding facial image. A first contour image sequence, including each of the first facial contour images, is input into a preset facial expression tracking model for prediction to obtain facial expression features corresponding to the first target image. The convolution operation in the facial expression tracking model is a 3D convolution operation performed in three dimensions: time, width, and height. The facial expression tracking model is pre-trained using a second training dataset. Each training sample in the second training dataset includes a set of second contour image sequences and corresponding second label data. The second contour image sequence includes each frame of second facial contour image. Each second facial contour image is obtained by segmenting each frame of facial image in the second facial image sequence by inputting it into the trained facial contour segmentation model. The second facial image sequence includes a second target image and at least one frame of facial image acquired before and / or after the second target image. The second label data is the full-face expression features of the same object acquired synchronously with the second target image.

2. The facial expression tracking method as described in claim 1, characterized in that, Each frame of the first facial image sequence includes a sub-image obtained by simultaneously capturing the face of the target object from at least two different angles. The first facial contour image obtained by segmenting each frame of the facial image includes each binarized facial contour sub-image obtained by segmenting each of the sub-images. The step of inputting a sequence of first contour images, including each of the first facial contour images, into a preset facial expression tracking model for prediction to obtain facial expression features corresponding to the first target image includes: Each facial contour sub-image corresponding to the same angle in the first contour image sequence is used as input data for one channel. The input data of each channel is input into the facial expression tracking model for prediction to obtain facial expression features corresponding to the first target image.

3. The facial expression tracking method as described in claim 1, characterized in that, The full-face expression features are Blendshape parameter sets. The facial expression tracking model includes a first deep residual network and a regression network. The convolution and batch normalization operations in the first deep residual network are 3D convolution and 3D batch normalization operations. The step of inputting a first contour image sequence including each first facial contour image into a preset facial expression tracking model for prediction to obtain facial expression features corresponding to the first target image includes: The first contour image sequence is input into the first depth residual network for feature extraction to obtain a feature vector; The feature vector is input into the regression network for prediction to obtain the Blendshape parameter set corresponding to the first target image, which serves as the facial expression feature corresponding to the first target image.

4. The facial expression tracking method as described in claim 1, characterized in that, The facial contour segmentation model includes a second depth residual network and a segmentation network. The step of inputting each frame of the facial image in the first facial image sequence into the preset facial contour segmentation model for segmentation to obtain binarized first facial contour images of each frame includes: Each frame of the facial image in the first facial image sequence is used as the fourth target image. The fourth target image is input into the second deep residual network for feature extraction to obtain a feature map. The feature map is input into the segmentation network for segmentation to obtain a binarized first facial contour image corresponding to the fourth target image.

5. The facial expression tracking method as described in claim 1, characterized in that, The facial expression tracking method also includes: Acquire multiple frames of facial images of the target object; The multi-frame facial images are sent to a training device so that the training device can generate at least one set of third facial image sequences based on the multi-frame facial images. The facial contour segmentation model trained with the first training dataset is used to segment each frame of the facial images in the third facial image sequence to obtain a third contour image sequence. The third contour image sequence and the Blendshape parameter group of the target object, which is synchronously acquired with the third target image in the third facial image sequence, are used to fine-tune the facial expression tracking model trained with the second training dataset. The third facial image sequence includes the third target image and at least one frame of facial image acquired before and / or after the third target image. The system receives the fine-tuned facial expression tracking model sent by the training device and uses the fine-tuned facial expression tracking model as the preset facial expression tracking model.

6. A method for training a facial expression tracking model, characterized in that, The training method for the facial expression tracking model includes: Obtain a first training dataset, wherein each training sample in the first training dataset includes a frame of facial image and corresponding first label data, wherein the first label data is a binarized facial contour image annotated for the corresponding facial image. The first training dataset is used to train the facial contour segmentation model to be trained; Acquire at least one set of second facial image sequences and second label data corresponding to the second facial image sequences, wherein the second facial image sequence includes a second target image and at least one frame of facial image acquired before and / or after the second target image, and the second label data is the full-face expression features of the same object acquired synchronously with the second target image; Each frame of the facial image in the second facial image sequence is input into the trained facial contour segmentation model for segmentation to obtain a second contour image sequence, wherein the second contour image sequence includes each frame of the second facial contour image in binarization. The second contour image sequence and the corresponding second label data are used as a training sample. A second training dataset is obtained based on multiple training samples. The facial expression tracking model to be trained is trained using the second training dataset. The convolution operation in the facial expression tracking model is a 3D convolution operation performed in three dimensions: time, width, and height.

7. The facial expression tracking model training method as described in claim 6, characterized in that, The step of training the facial expression tracking model using the second training dataset includes: Each time, a data set containing multiple training samples is extracted from the second training dataset; Based on the proportion of the first type of training samples in the data set, the first weight corresponding to the first type of training samples and the second weight corresponding to the second type of training samples are calculated. The first type of training samples are training samples with facial expressions in the corresponding second target image, and the second type of training samples are training samples without facial expressions in the corresponding second target image. The first weight is greater than the second weight, and the larger the proportion of the first and second weights, the larger they are. Substitute each training sample from the data set into the facial expression tracking model to be trained to calculate the original loss; The original loss of each training sample is weighted and averaged according to the corresponding first weight and second weight to obtain the total loss, and the facial expression tracking model is updated according to the total loss; After iteratively updating the facial expression tracking model using at least one set of the data set in multiple rounds, a trained facial expression tracking model is obtained.

8. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the facial expression tracking method as described in any one of claims 1 to 5, or configured to implement the steps of the facial expression tracking model training method as described in any one of claims 6 to 7.

9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the facial expression tracking method as described in any one of claims 1 to 5, or the steps of the facial expression tracking model training method as described in any one of claims 6 to 7.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the facial expression tracking method as described in any one of claims 1 to 5, or implements the steps of the facial expression tracking model training method as described in any one of claims 6 to 7.