Model training, action mapping method and device, electronic equipment and storage medium

By training a target mapping model and combining sample facial data with a motion capture model, adaptive binding of digital human facial data is achieved, solving the high cost problem of facial binding under different motion capture models and reducing the binding workload.

CN115457174BActive Publication Date: 2026-07-31SHANGHAI PUDONG DEVELOPMENT BANK
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI PUDONG DEVELOPMENT BANK
Filing Date
2022-08-31
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Because different motion capture models may differ during the motion capture process, the cost and workload of face binding remain high. Existing technologies require face binding to be redone for new motion capture models.

Method used

By acquiring sample facial data that has been bound to a digital human face, the digital human face is driven to perform sample actions. Sample videos are collected and input into the motion capture model to obtain sample action data. Based on the sample action data and facial data, the original mapping model is trained to obtain the target mapping model, thus achieving adaptive binding of facial data.

Benefits of technology

Only one face binding is required for the digital human, and it can then be directly applied to any motion capture model, reducing the cost and workload of face binding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115457174B_ABST
    Figure CN115457174B_ABST
Patent Text Reader

Abstract

This invention discloses a model training and motion mapping method, apparatus, electronic device, and storage medium. The model training method includes: acquiring sample facial data that has been facially bound to a digital human face; driving the digital human face to perform sample actions using the sample facial data, and capturing the digital human face performing the sample actions to obtain a sample video; inputting the sample video into a motion capture model to obtain sample motion data; and training an original mapping model based on the sample motion data as actual input data and the sample facial data as expected output data to obtain a target mapping model. The technical solution of this invention can reduce the cost and workload of facial binding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of digital humans, and more particularly to a model training, motion mapping method, apparatus, electronic device, and storage medium. Background Technology

[0002] In the field of digital humans, facial rigging can be used to define the movement capabilities of different parts of a digital human's face. Specifically, motion capture models are used to process motion videos of real human faces to capture motion data. This motion data is then manually rigged onto the digital human's face, allowing the digital human's face to perform the same movements as a real human face.

[0003] It's important to note that different motion capture models may differ during the motion capture process. This means that facial rigging performed for one motion capture model may not be applicable to another. In other words, when a new motion capture model is added, facial rigging needs to be performed again for that new model. Obviously, this leads to high costs and workload for facial rigging. Summary of the Invention

[0004] This invention provides a model training, motion mapping method, device, electronic device, and storage medium, which solves the problem of high cost and workload in face binding.

[0005] According to one aspect of the present invention, a model training method is provided, which may include:

[0006] Obtain sample facial data that has been facially bound to a digital human face;

[0007] The sample facial data is used to drive the digital human face to perform sample actions, and the digital human face performing the sample actions is captured to obtain sample videos.

[0008] The sample video is input into the motion capture model to obtain sample motion data;

[0009] Based on sample action data as actual input data and sample facial data as expected output data, the original mapping model is trained to obtain the target mapping model.

[0010] According to another aspect of the present invention, an action mapping method is provided, which may include:

[0011] When a target person performs a target action on their face, the target video obtained after capturing the target person's face is acquired and input into the motion capture model to obtain target action data.

[0012] Obtain the target mapping model trained according to the model training method provided in any embodiment of the present invention;

[0013] The target motion data is input into the target mapping model, and the target facial data that can be mapped onto the digital human face is obtained based on the output of the target mapping model.

[0014] Among them, the digital human face and motion capture model is associated with the target mapping model.

[0015] According to another aspect of the present invention, a model training apparatus is provided, which may include:

[0016] The sample facial data acquisition module is used to acquire sample facial data that has been facially bound to the digital human face;

[0017] The sample video acquisition module is used to drive digital human faces to perform sample actions through sample facial data, and to capture digital human faces performing sample actions to obtain sample videos.

[0018] The sample motion data acquisition module is used to input sample videos into the motion capture model to obtain sample motion data;

[0019] The target mapping model acquisition module is used to train the original mapping model based on sample action data as actual input data and sample facial data as expected output data, to obtain the target mapping model.

[0020] According to another aspect of the present invention, a motion mapping device is provided, which may include:

[0021] The target motion data acquisition module is used to acquire the target video obtained after capturing the target person's face when the target person performs the target motion, and input the target video into the motion capture model to obtain the target motion data.

[0022] The target mapping model acquisition module is used to acquire the target mapping model trained according to the model training method provided in any embodiment of the present invention;

[0023] The target facial data acquisition module is used to input target motion data into the target mapping model, and based on the output of the target mapping model, to map the target facial data that can be used for facial binding on the digital human face;

[0024] Among them, the digital human face and motion capture model is associated with the target mapping model.

[0025] According to another aspect of the present invention, an electronic device is provided, which may include:

[0026] At least one processor; and

[0027] A memory that is communicatively connected to at least one processor; wherein,

[0028] The memory stores a computer program that can be executed by at least one processor, which is executed by at least one processor to cause the at least one processor to perform the model training method or the action mapping method provided in any embodiment of the present invention.

[0029] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement the model training method or the action mapping method provided in any embodiment of the present invention.

[0030] The technical solution of this invention involves acquiring sample facial data that has been facially bound to a digital human face; then driving the digital human face to perform sample actions using the sample facial data, and capturing the digital human face performing the sample actions to obtain sample video; inputting the sample video into a motion capture model to obtain sample motion data; and training the original mapping model based on the sample motion data as actual input data and the sample facial data as expected output data to obtain a target mapping model. This technical solution allows for adaptive learning and successful facial binding of the target mapping model through training, and only requires facial binding between the digital human and the motion capture model once, eliminating the need for other motion capture models to be facially bound to the digital human, thereby reducing the cost and workload of facial binding.

[0031] It should be understood that the description in this section is not intended to identify key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0033] Figure 1 This is a flowchart of a model training method provided in Embodiment 1 of the present invention;

[0034] Figure 2 This is a schematic diagram of the digital human eye under different degrees of opening and closing in a model training method provided in Embodiment 1 of the present invention;

[0035] Figure 3 This is a flowchart of a model training method provided in Embodiment 2 of the present invention;

[0036] Figure 4 This is a schematic diagram illustrating the mapping of sample action data to sample images in a model training method provided in Embodiment 2 of the present invention;

[0037] Figure 5 This is a flowchart of an optional example of a model training method provided in Embodiment 2 of the present invention;

[0038] Figure 6 This is a structural diagram of a convolutional neural network provided in Embodiment 2 of the present invention;

[0039] Figure 7 This is a flowchart of a model training method provided in Embodiment 3 of the present invention;

[0040] Figure 8 This is a flowchart of an optional example of a model training method provided in Embodiment 3 of the present invention;

[0041] Figure 9 This is a flowchart of another optional example of a model training method provided in Embodiment 3 of the present invention;

[0042] Figure 10 This is a flowchart of an action mapping method provided in Embodiment 4 of the present invention;

[0043] Figure 11 This is a flowchart of an optional example of an action mapping method provided in Embodiment 4 of the present invention;

[0044] Figure 12 This is a flowchart of another optional example of an action mapping method provided in Embodiment 4 of the present invention;

[0045] Figure 13 This is a structural block diagram of a model training device provided in Embodiment 5 of the present invention;

[0046] Figure 14 This is a structural block diagram of an action mapping device provided in Embodiment Six of the present invention;

[0047] Figure 15 This is a schematic diagram of the structure of an electronic device that implements the model training method or action mapping method of the present invention. Detailed Implementation

[0048] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0049] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. The same applies to "target," "original," etc., and will not be repeated here. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0050] Example 1

[0051] Figure 1 This is a flowchart of a model training method provided in Embodiment 1 of the present invention. This embodiment is applicable to training a target mapping model that matches the motion capture model, thereby enabling the direct mapping of target motion data captured by the motion capture model onto a digital human face based on the target mapping model. This method can be executed by the model training device provided in this embodiment of the invention. This device can be implemented in software and / or hardware, and can be integrated into an electronic device, which can be various user terminals or servers.

[0052] See Figure 1 The method of this invention specifically includes the following steps:

[0053] S110. Obtain sample facial data that has been facially bound to the digital human face.

[0054] Digital humans are virtual humans that use information science methods to virtually simulate the human body at different levels of form and function. Sample facial data can be understood as data that has been bound to the digital human's face and can drive the digital human to perform facial movements, such as data that drives the digital human to perform non-linear facial movements. In practical applications, optionally, the sample facial data can be sequence data with multiple dimensions, such as sequence data with 52 dimensions, each dimension can represent a facial morphology, such as the eye opening and closing dimension, the mouth opening and closing dimension, or the eyebrow raising dimension, etc.; on this basis, optionally, the value range of the data under each dimension can be [0,1]. Further optionally, the number of sample facial data obtained can be one, two, or more, without specific limitation.

[0055] S120. Drive the digital human face to perform sample actions through sample facial data, and collect the digital human face performing the sample actions to obtain sample video.

[0056] Here, sample actions can be understood as facial movements performed by a digital human face based on sample facial data. The tools used to capture the digital human face performing sample actions can be cameras, screen recording tools, or Web Real-Time Communications (WebRTC) streaming recording tools, etc., without specific limitations.

[0057] In this embodiment of the invention, sample facial data can be used to drive a digital human face to form facial surfaces of different shapes to perform sample actions, and the digital human face performing the sample actions can be captured to obtain a sample video. As can be seen from the above examples, the sample video can be offline video or real-time video.

[0058] For example, see Figure 2 One dimension of the sample facial data is the eye opening and closing dimension. The sample facial data can include data corresponding to five eye opening and closing dimensions: 0.0, 0.2, 0.5, 0.8, and 1.0. The sample facial data can be used to drive the digital human's eyes to perform sample movements with different degrees of opening and closing.

[0059] S130. Input the sample video into the motion capture model to obtain sample motion data.

[0060] In this embodiment of the invention, the motion capture model can be understood as a motion capture model that needs to be mapped to a digital human face. Sample motion data can be understood as facial motion data obtained by capturing sample movements of a digital human face in a sample video through the motion capture model.

[0061] It should be noted that, in light of the application scenarios that may be involved in the embodiments of the present invention, the sample facial data mentioned above can be understood as facial data obtained by binding the bound motion data captured by the bound capture model to the digital human face. The bound capture model is a different motion capture model from the motion capture model in this step, that is, the motion capture model in this step is a motion capture model that has not been bound accordingly.

[0062] In this embodiment of the invention, sample videos can be input into a motion capture model to capture sample actions in the sample videos based on the motion capture model, thereby obtaining sample action data.

[0063] It should be noted that, due to the special nature of digital human binding, the dimensions of sample motion data and sample facial data may be different or different, and no specific limitation is made here.

[0064] S140. Based on the sample action data as actual input data and the sample facial data as expected output data, the original mapping model is trained to obtain the target mapping model.

[0065] Here, the original mapping model can be understood as the mapping model before it has been trained. The target mapping model can be understood as the mapping model that, after training, can directly map the sample motion data captured by the motion capture model onto the digital human face. The actual input data can be understood as the input data fed into the original mapping model for training the original mapping model. The expected output data can be understood as the data that the original mapping model is expected to output after the actual input data has been fed into it.

[0066] In this embodiment of the invention, sample action data can be used as actual input data, and sample facial data can be used as expected output data. The original mapping model can be trained accordingly to obtain the trained target mapping model.

[0067] The technical solution of this invention involves acquiring sample facial data that has been facially bound to a digital human face; then driving the digital human face to perform sample actions using the sample facial data, and capturing the digital human face performing the sample actions to obtain sample video; inputting the sample video into a motion capture model to obtain sample motion data; and training the original mapping model based on the sample motion data as actual input data and the sample facial data as expected output data to obtain a target mapping model. This technical solution only requires one binding operation of the digital human face based on any motion capture model to obtain sample facial data. When adding a new motion capture model, a target mapping model matching the motion capture model can be directly trained based on the sample facial data. Therefore, the sample motion data captured by the motion capture model can be directly mapped onto the digital human face based on the target mapping model, without needing to re-bind the face for the motion capture model, thereby reducing the cost and workload of face binding.

[0068] An optional technical solution for obtaining sample facial data that has been facially bound to a digital human face includes: obtaining a facial data set, wherein the facial data set includes multiple candidate facial data, each candidate facial data having been facially bound to a digital human face; selecting at least two candidate facial data from the facial data set; interpolating based on each selected candidate facial data, and using the interpolated facial data and each selected candidate facial data as sample facial data that has been facially bound to a digital human face.

[0069] It should be noted that, due to limitations such as the low sampling rate when collecting facial data, the continuity of the sample actions represented by the candidate facial data in the facial dataset may not be very good. Therefore, at least two candidate facial data can be selected from the facial dataset, interpolation can be performed between the selected candidate facial data, and the interpolated facial data and each selected candidate facial data can be used as sample facial data that has been facially bound to the digital human face.

[0070] For example, suppose one dimension of the candidate facial data is the eye opening dimension. If the data of two candidate facial data selected from the facial dataset are 0.2 and 0.8 respectively in the eye opening dimension, then 0.3, 0.4, 0.5, 0.6 and 0.7 can be interpolated from these two data, thus obtaining the interpolated facial data corresponding to 0.3, 0.4, 0.5, 0.6 and 0.7 respectively.

[0071] For example, suppose N groups of candidate facial data are selected from the facial dataset. Each group of candidate facial data contains 2 candidate facial data. For each group of candidate facial data, t-2 interpolated facial data are obtained by interpolation in the group of candidate facial data, that is, N*t sample facial data can be obtained. Each sample facial data is a sequence of data with a blendshapeDim1 dimension. At this time, all sample facial data can be represented as a matrix of [N*t, blendshapeDim1], or it can be represented as data of N*t frames.

[0072] The advantage of setting interpolation in candidate facial data is that it can obtain sample facial data that represents more continuous sample actions, thereby improving the accuracy of model training.

[0073] Example 2

[0074] Figure 3 This is a flowchart of a model training method provided in Embodiment 2 of the present invention. This embodiment is based on and optimized from the above-described technical solutions. In this embodiment, optionally, after obtaining the sample action data, the above-described model training method further includes: performing position encoding on the sample action data to obtain position features, and combining and arranging the position features to obtain a sample image; placing the sample action data in the sample image, and using the sample image obtained after placement as the sample action data. The explanations of terms that are the same as or corresponding to those in the above embodiments are not repeated here.

[0075] See Figure 3 The method in this embodiment may specifically include the following steps:

[0076] S210. Obtain sample facial data that has been facially bound to the digital human face.

[0077] S220. Drive the digital human face to perform sample actions using sample facial data, and capture the digital human face performing the sample actions to obtain sample video.

[0078] S230. Input the sample video into the motion capture model to obtain sample motion data.

[0079] S240. The sample action data is encoded to obtain position features, and the position features are combined and arranged to obtain the sample image.

[0080] In this embodiment of the invention, considering the non-linear relationship between the sample action data and the sample facial data sequences, the subsequent model training process is precisely the process of learning this non-linear relationship. However, during model training, on the one hand, directly inputting sequence data such as sample action data into the original mapping model may result in poor model training performance; on the other hand, if the sample action data is processed using traditional feature extraction and selection methods before being input into the original mapping model, such a processing procedure may waste a significant amount of time and human resources.

[0081] Therefore, in order to ensure the model training effect while reducing the consumption of time and human resources, in this embodiment of the invention, positional encoding can be performed on sample action data without a concept of order to obtain positional features, thereby mapping the sample action data to a higher-dimensional space; then, the positional features are combined and arranged to form a set of two-dimensional data, namely a black and white sample image, so that the complex nonlinear relationship between positional features can be expressed through their positional information in the sample image. Such a sample image helps to reduce the training difficulty of the original mapping model, thereby improving the mapping accuracy of the trained target mapping model.

[0082] S250. Place the sample action data into the sample image, and use the sample image obtained after placement as the sample action data.

[0083] In order to maximize the utilization of sample action data by allowing the original mapping model to directly utilize the sample action data, in this embodiment of the invention, the sample action data can be placed in the sample image, such as above, in the middle, or below the sample image, without specific limitations. Furthermore, the sample image containing the sample action data is used as or updated as the sample action data, thereby enabling model training based on the sample action data represented by the sample image.

[0084] For example, sample motion data can be mapped onto a two-dimensional sample image using methods such as positional encoding and combination permutation. Then, the sample motion data is placed in the center of this sample image, resulting in a sample image as shown below. Figure 4 As shown, sample motion data captured from sample videos can be represented based on feature images such as sample images.

[0085] S260. Based on the sample action data as actual input data and the sample facial data as expected output data, the original mapping model is trained to obtain the target mapping model.

[0086] The technical solution of this invention involves encoding the position of sample action data to obtain positional features, combining and arranging these features to obtain a sample image, placing the sample action data within the sample image, and using the resulting sample image as the sample action data. This technical solution reduces the training difficulty of the original mapping model, thereby improving the mapping accuracy of the trained target mapping model.

[0087] To better understand the technical solutions of the above embodiments of the present invention, an optional example is provided herein. For example, see... Figure 5 In conjunction with the technical solutions of the above embodiments of the present invention, the obtained sample action data can be converted to obtain sample images, and model training can be performed based on the sample images to obtain target mapping models. Then, the sample images can be input into the target mapping models to obtain sample facial data.

[0088] Based on the above scheme, an optional technical solution includes a convolutional neural network in the original mapping model used for regression; the original mapping model is trained based on sample action data as actual input data and sample facial data as expected output data to obtain a target mapping model, including: inputting the sample action data as actual input data into the convolutional neural network to extract features from the sample action data through the convolutional neural network, and regressing the regression facial data corresponding to the sample facial data based on the feature extraction results; comparing the sample facial data as expected output data and the regression facial data as actual output data, and adjusting the network parameters in the original mapping model based on the comparison results to train the target mapping model.

[0089] Convolutional Neural Networks (CNNs) are a type of feedforward neural network that incorporates convolutional computations and has a deep structure. CNNs possess representation learning capabilities, enabling them to perform translation-invariant classification of input information according to their hierarchical structure. The original mapping model can be understood as a model with regression functionality that incorporates a CNN. Regressed facial data can be understood as the facial data actually output by the original mapping model after inputting sample action data.

[0090] In this embodiment of the invention, since the sample action data is represented by sample images, and the original mapping model includes a convolutional neural network with image feature extraction capabilities, the sample action data can be input into the convolutional neural network. The convolutional neural network then extracts features from the sample action data represented by the sample images, extracting relevant features of the sample action data. Based on the feature extraction results, regression facial data corresponding to the sample facial data is then regressed. In this way, the sample facial data and the regression facial data can be compared, and the network parameters in the original mapping model can be adjusted based on the comparison results, thereby training the target mapping model.

[0091] For example, see Figure 6 In this embodiment of the invention, the original mapping model can be a regression model, which may include an input layer, a convolutional neural network, a global average pooling layer, a dropout layer, a dense fully connected layer, and an output layer connected in sequence. The convolutional neural network may include a two-dimensional convolution (conv2D), a batch normalization layer, an activation function, and a max pooling layer. Based on this, since the regression model needs to predict the specific numerical values ​​of the sample facial data, the activation function of the output layer can be set to empty, allowing it to output real numbers.

[0092] The original mapping model includes a convolutional neural network, which can be used in conjunction with sample action data represented by sample images. This allows the convolutional neural network to quickly and accurately extract relevant features from the sample action data, thereby ensuring the mapping accuracy of the target mapping model obtained through subsequent training and reducing the time and manpower required in the feature extraction process.

[0093] Example 3

[0094] Figure 7This is a flowchart of a model training method provided in Embodiment 3 of the present invention. This embodiment is based on and optimized from the above-mentioned technical solutions. In this embodiment, optionally, the number of sample facial data is multiple frames. After obtaining sample facial data that has been facially bound to a digital human face, the above-mentioned model training method further includes: for each video segment composed of multiple video segments of sample facial data, setting a target sound signal for the first frame of sample facial data in the video segment; after obtaining the sample video, the model training method further includes: detecting the sound components in the sample video, obtaining sample video frames containing the target sound signal in the sample video, and determining the video frame position of the sample video frame containing the target sound signal in the sample video; combining the sample video frames corresponding to each sample action data in the sample video, determining the position action data corresponding to the video frame position from each sample action data; according to each position action data and the segment duration of each video segment, cropping the sample action data that has a cropping requirement in each sample action data, and updating the captured sample action data based on the sample action data retained after cropping. The explanations of terms that are the same as or corresponding to those in the above embodiments are not repeated here.

[0095] See Figure 7 The method in this embodiment may specifically include the following steps:

[0096] S310. Obtain multi-frame sample facial data that has been facially bound to the digital human face.

[0097] This involves acquiring multiple frames of sample facial data, each of which can drive the digital human face to perform corresponding sample actions.

[0098] S320. For each video segment composed of multiple video segments with multi-frame sample facial data, set the target audio signal for the first frame sample facial data in the video segment.

[0099] In this context, a video segment can be understood as a video composed of at least two frames of sample facial data from multiple frames of sample facial data. Specifically, it involves driving a digital human face to perform sample actions based on these at least two frames of sample facial data, and then capturing the digital human face performing the sample actions to obtain a video segment. For each video segment, the first frame of sample facial data can be understood as the sample facial data corresponding to the first sample video frame among all sample video frames in the video segment, and a target audio signal is set for this first frame of sample facial data. For example, combining the facial data interpolation example above, assuming that each group of candidate facial data includes two candidate facial data, for each group of candidate facial data, t-2 interpolated facial data are obtained by interpolation in the group of candidate facial data. Thus, the two candidate facial data in the group of candidate facial data and the interpolated t-2 interpolated facial data can be used to construct a video segment with t frames of sample facial data, and the first candidate facial data among the two candidate facial data is directly used as the first frame of sample facial data in this video segment.

[0100] Based on this, optionally, in addition to setting a target audio signal for the first frame of sample facial data, other audio signals different from the target signal can also be set for each frame of sample facial data corresponding to the video segment, excluding the first frame. No specific limitations are made here. For example, for a video segment, a corresponding audio signal 'a' can be set for each frame of sample facial data corresponding to the video segment. For instance, a=1 can be set for the first frame of sample facial data (i.e., the target audio signal is 1), and a=0 can be set for other frames of sample facial data (i.e., the remaining audio signals are 0). In practical applications, optionally, 'a' can be an analog signal. When a=1, the sample video frame corresponding to that frame of sample facial data has a modulated high-pitched tone, while the sample video frames corresponding to other sample facial data are silent.

[0101] It is important to note that the purpose of the above steps is to take into account the possible errors that may occur during video recording and motion capture, which may cause the sample facial data and sample motion data to not be 100% aligned. Therefore, the first frame of sample facial data in each video segment can be marked to ensure that the multi-frame sample facial data and multi-frame sample motion data corresponding to each video segment can be aligned one by one.

[0102] S330: Drive the digital human face to perform sample actions through multi-frame sample facial data, and capture the digital human face performing the sample actions to obtain sample video.

[0103] Since the above steps have already set the target audio signal for the first frame of the sample facial data in each of the multiple video segments composed of multiple frames of sample facial data, the sample video obtained based on these sample facial data will also contain the corresponding target audio signal.

[0104] S340. Input the sample video into the motion capture model to obtain sample motion data.

[0105] Since the sample video frames in the sample video have a sequential order in the time dimension, the sample motion data captured by the motion capture model also have a sequential order in the time dimension.

[0106] S350. Detect the sound components in the sample video, obtain the sample video frame containing the target sound signal, and determine the video frame position of the sample video frame containing the target sound signal in the sample video.

[0107] The audio component can be understood as the audio-related component in the sample video. The sample video can consist of multiple sample video frames. As mentioned above, since the sample video also contains the target audio signal, in this embodiment of the invention, the audio component in the sample video can be detected to obtain sample video frames containing the target audio signal. The number of sample video frames containing the target audio signal can be the same as the number of video segments. Furthermore, for each sample video frame containing the target audio signal, its position within the sample video can be determined, facilitating subsequent cropping of the sample motion data based on that position.

[0108] It should be noted that in practical applications, S340 and S350 can be executed serially or in parallel. If S340 and S350 are executed serially, S340 can be executed first and then S350, or S350 can be executed first and then S340; the execution order is not specifically limited.

[0109] S360. Combining the sample video frames corresponding to each sample action data in the sample video, determine the position action data corresponding to the video frame position from each sample action data.

[0110] Since the sample video consists of multiple sample video frames, and the sample motion data is obtained by motion capture of the sample video, each sample motion data has a corresponding sample video frame in the sample video. Based on this, since the video frame position is the location of the sample video frame containing the target sound signal within the sample video, the positional motion data corresponding to the video frame position can be determined from each sample motion data by combining the corresponding sample video frames. In other words, the positional motion data can be understood as the sample motion data corresponding to the sample motion data represented by that video frame position.

[0111] S370. Based on the motion data at each location and the duration of each video segment, the sample motion data that requires cropping is cropped, and the captured sample motion data is updated based on the sample motion data retained after cropping.

[0112] The duration of each video segment can be obtained in advance, and the durations of different video segments can be the same or different, without specific limitations. Since the video segment to which each position's motion data belongs can be obtained directly, and the segment duration can reflect the number of sample motion data corresponding to the corresponding video segment, the sample motion data that requires trimming can be determined based on the position's motion data and the duration of each video segment.

[0113] For example, for each positional motion data, the number of sample motion data starting from that positional motion data can be determined to belong to the same video segment based on the corresponding segment duration. The next sample motion data of the last sample motion data in these sample motion data is taken as the first motion data, and the previous sample motion data of the next positional motion data of that positional motion data is taken as the second motion data. Thus, the first motion data, the second motion data, and all sample motion data between the two can be taken as sample motion data with cropping requirements.

[0114] Furthermore, these sample motion data that require cropping are cropped, and the cropped sample motion data is used as the captured sample motion data. That is, the sample motion data used in subsequent applications is aligned one by one with the sample facial data.

[0115] S380. Based on the sample action data as actual input data and the sample facial data as expected output data, the original mapping model is trained to obtain the target mapping model.

[0116] The technical solution of this invention sets a target sound signal for the first frame sample facial data in each video segment. This allows for the detection of sound components in the sample video, resulting in a sample video frame containing the target sound signal, and the determination of the video frame position of this sample video frame within the sample video. Furthermore, by combining the sample video frames corresponding to each sample action data in the sample video, positional action data corresponding to the video frame position is determined from each sample action data. Based on the positional action data and the duration of each video segment, sample action data requiring trimming is trimmed. The resulting multi-frame sample action data is aligned one-to-one with the directly acquired multi-frame sample facial data, thereby improving the training accuracy of the original mapping model.

[0117] Based on the above scheme, an optional technical solution, before combining the sample video frames corresponding to each sample action data in the sample video, the model training method further includes: resampling each sample action data, and updating the captured sample action data based on the resampling results.

[0118] It should be noted that when performing motion capture on sample videos based on the motion capture model, the sampling rate mentioned above may differ from the sampling rate used in digital human-driven rendering because the motion capture model has a certain sampling rate in its calculation. This is because the motion data of each sample can be resampled here, thereby ensuring the effectiveness of subsequent sample motion data cropping.

[0119] In this embodiment of the invention, the process of resampling each sample action data can be as follows: linearly traverse the data of all sample video frames in the sample video, use the time point position of each frame to calculate the data closest to the time point, and use the finally obtained data as the sample action data.

[0120] To better understand the technical solutions of the above embodiments of the present invention, another optional example is provided here. For example, see... Figure 8 The process involves acquiring sample facial data and setting a target sound signal for the sample facial data; then, acquiring sample videos based on the sample facial data, performing motion capture on the sample videos to obtain sample motion data, and training the original mapping model based on the sample motion data and the sample facial data with the set target sound signal to obtain the target mapping model.

[0121] To better understand the technical solutions of the above embodiments of the present invention, another optional example is provided here. For example, see... Figure 9The process involves: acquiring sample videos; capturing motion in the sample videos using a motion capture model to obtain sample motion data X; resampling X to obtain resampled sample motion data X'; detecting the audio components of the sample videos to obtain the positions of sample video frames containing the target audio signal; cropping X' based on the sample video frame positions to obtain sample motion data X'” as the actual input data; and training the original mapping model based on X'”.

[0122] Example 4

[0123] Figure 10 This is a flowchart of a motion mapping method provided in Embodiment 4 of the present invention. This embodiment is applicable to motion mapping situations where facial binding costs and workloads are relatively low. The method can be executed by the motion mapping device provided in this embodiment of the invention. This device can be implemented in software and / or hardware, and can be integrated into an electronic device, which can be various user terminals or servers.

[0124] See Figure 10 The method in this embodiment may specifically include the following steps:

[0125] S410. When the target person performs a target action on their face, acquire the target video obtained after capturing the target person's face, and input the target video into the motion capture model to obtain target action data.

[0126] In this context, the target person can be understood as the individual whose facial movements need to be mapped onto the digital human's face. The target person can be a real person or someone in a video; the target person can perform the target movements as a real person or as a cartoon character, without any specific limitations. The target movement can be understood as the movement that needs to be mapped onto the digital human's face.

[0127] In this embodiment of the invention, when a target person is performing a target action on their face, video capture of the target action is performed to obtain a target video after capturing the target person's face. The capture method can be offline recording or real-time shooting by an RGB camera, and no specific limitation is made here. Then, the target video is input into a motion capture model, thereby capturing the target video based on the motion capture model to obtain target action data.

[0128] S420. Obtain the target mapping model trained by the model training method provided in any embodiment of the present invention, wherein the target mapping model is associated with the motion capture model.

[0129] It should be noted that a target mapping model corresponding to the above-mentioned motion capture model can be pre-trained according to the model training method provided in any embodiment of the present invention. This target mapping model can directly map the target motion data captured by the motion capture model onto the digital human face.

[0130] S430. Input the target motion data into the target mapping model, and map the target facial data that can be bound to the digital human face according to the output of the target mapping model, wherein the digital human face is associated with the target mapping model.

[0131] The target facial data can be understood as the data that can drive the digital human face to perform actions according to the target action. This digital human face is the digital human face involved in the training process of the target mapping model.

[0132] In this embodiment of the invention, target motion data can be input into the target mapping model. Since the digital human face and motion capture model are associated with the target mapping model, the target mapping model can output a result that can drive the digital human face to perform actions according to the target motion. In other words, based on the output of the target mapping model, target facial data that can be mapped onto the digital human face can be obtained.

[0133] In practical applications, optionally, since the motion capture model has a certain sampling rate in its calculations, this sampling rate may differ from the sampling rate used during digital human-driven rendering. Therefore, the target facial data can be resampled. In this embodiment of the invention, the process of resampling the target facial data can be as follows: linearly traversing the data of all target video frames in the target video, using the time point position of each frame, calculating the data closest to that time point, and using the finally obtained data as the target facial data. Such target facial data can be sent to the digital human program for driving and rendering, thereby obtaining a digital human animated video.

[0134] The technical solution of this invention involves acquiring a target video of the target person's face after capturing the target action, and inputting the target video into a motion capture model to obtain target action data. Then, a target mapping model trained according to the model training method provided in any embodiment of this invention is acquired. Finally, the target action data is input into the target mapping model, and based on the output of the target mapping model, target facial data capable of facial binding on a digital human face is obtained. This technical solution allows for direct mapping of the target action data captured by the motion capture model to a digital human face based on a target mapping model that matches the motion capture model, eliminating the need for facial binding of the motion capture model. This achieves motion mapping with low cost and workload for facial binding.

[0135] To better understand the technical solutions of the above embodiments of the present invention, an optional example is provided herein. For example, see... Figure 11 The system acquires offline recorded video or target video captured by an RGB camera, and inputs the target video into a motion capture model to capture motion and obtain target motion data; the target motion data is then input into a target mapping model to obtain target facial data.

[0136] To better understand the technical solutions of the above embodiments of the present invention, another optional example is provided here. For example, see... Figure 12 The system acquires target videos recorded offline or captured via an RGB camera, performs motion capture on the target videos based on a motion capture model to obtain target motion data, and inputs the target motion data into a target mapping model to map the target motion data into target facial data; the target facial data is then resampled to obtain resampled target facial data, and the digital human face is driven to perform actions based on the target facial data.

[0137] Example 5

[0138] Figure 13 This is a structural block diagram of the model training apparatus provided in Embodiment 5 of the present invention. This apparatus is used to execute the model training method provided in any of the above embodiments. This apparatus and the model training methods of the above embodiments belong to the same inventive concept. Details not described in detail in the embodiments of the model training apparatus can be found in the embodiments of the above model training methods. See also... Figure 13 The device may specifically include: a sample facial data acquisition module 510, a sample video acquisition module 520, a sample motion data acquisition module 530, and a target mapping model acquisition module 540.

[0139] Among them, the sample facial data acquisition module 510 is used to acquire sample facial data that has been facially bound to the digital human face;

[0140] The sample video acquisition module 520 is used to drive the digital human face to perform sample actions through sample facial data, and to capture the digital human face performing sample actions to obtain sample video.

[0141] The sample motion data acquisition module 530 is used to input sample videos into the motion capture model to obtain sample motion data;

[0142] The target mapping model is obtained by module 540, which is used to train the original mapping model based on sample action data as actual input data and sample facial data as expected output data to obtain the target mapping model.

[0143] Optionally, the model training apparatus may also include:

[0144] The sample image acquisition module is used to encode the sample action data after obtaining the sample action data, obtain position features, and combine and arrange the position features to obtain the sample image.

[0145] The sample motion data placement module is used to place sample motion data into a sample image and use the resulting sample image as the sample motion data.

[0146] Optionally, based on the above scheme, the original mapping model used for regression includes a convolutional neural network;

[0147] The target mapping model yields module 540, which includes:

[0148] The facial data regression unit is used to input sample action data, which is the actual input data, into the convolutional neural network, so that the convolutional neural network can extract features from the sample action data and regress the regression facial data corresponding to the sample facial data based on the feature extraction results.

[0149] The target mapping model obtains a unit, which is used to compare the sample facial data, which is the expected output data, with the regressed facial data, which is the actual output data, and adjusts the network parameters in the original mapping model based on the comparison results to train the target mapping model.

[0150] Optionally, the number of sample facial data is multiple frames, and the model training device also includes:

[0151] The target audio signal setting module is used to set the target audio signal for the first frame of sample facial data in each video segment composed of multiple video segments composed of multiple frames of sample facial data after acquiring sample facial data that has been facially bound to the digital human face.

[0152] The model training device also includes:

[0153] The video frame position determination module is used to detect the sound components in the sample video after obtaining the sample video, obtain the sample video frame containing the target sound signal, and determine the video frame position of the sample video frame containing the target sound signal in the sample video.

[0154] The position and motion data determination module is used to combine the sample video frames corresponding to each sample motion data in the sample video to determine the position and motion data corresponding to the video frame position from each sample motion data.

[0155] The first sample motion data update module is used to crop the sample motion data that needs to be cropped in each sample motion data according to the motion data of each location and the segment length of each video segment, and update the captured sample motion data based on the sample motion data retained after cropping.

[0156] Optionally, based on the above scheme, the model training device further includes:

[0157] The second sample motion data update module is used to resample each sample motion data before combining the sample video frames corresponding to each sample motion data in the sample video, and update the captured sample motion data based on the resampling results.

[0158] Optionally, the sample facial data acquisition module 510 includes:

[0159] A facial data set acquisition unit is used to acquire a facial data set, wherein the facial data set includes multiple candidate facial data, and each candidate facial data has been facially bound to a digital human face;

[0160] A candidate facial data selection unit is used to select at least two candidate facial data from a facial data set;

[0161] The sample facial data binding unit is used to interpolate based on each selected candidate facial data, and to use the interpolated facial data and each selected candidate facial data as sample facial data that has been facially bound to the digital human face.

[0162] The model training device of this invention acquires sample facial data that has been facially bound to a digital human face through a sample facial data acquisition module; then, a sample video acquisition module drives the digital human face to perform sample actions based on the sample facial data and captures the digital human face performing the sample actions to obtain sample videos; the sample action data acquisition module inputs the sample videos into a motion capture model to obtain sample action data; and the target mapping model acquisition module trains the original mapping model based on the sample action data as actual input data and the sample facial data as expected output data to obtain a target mapping model. The model training device of this invention can adaptively learn and complete facial binding by training the target mapping model, and only requires facial binding between the digital human and the motion capture model once, without needing to bind other motion capture models to the digital human face, thereby reducing the cost and workload of facial binding.

[0163] The model training apparatus provided in this embodiment of the invention can execute the model training method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0164] It is worth noting that in the embodiments of the above-mentioned model training device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.

[0165] Example 6

[0166] Figure 14 This is a structural block diagram of the motion mapping device provided in Embodiment Six of the present invention. This device is used to execute the motion mapping method provided in any of the above embodiments. This device and the motion mapping methods of the above embodiments belong to the same inventive concept. Details not described in detail in the embodiments of the motion mapping device can be found in the embodiments of the above motion mapping methods. See also... Figure 14 The device may specifically include: a target motion data acquisition module 610, a target mapping model acquisition module 620, and a target facial data acquisition module 630.

[0167] Among them, the target motion data acquisition module 610 is used to acquire the target video obtained after capturing the target person's face when the target person performs the target motion, and input the target video into the motion capture model to obtain the target motion data.

[0168] The target mapping model acquisition module 620 is used to acquire the target mapping model trained according to the model training method provided in any embodiment of the present invention;

[0169] The target facial data acquisition module 630 is used to input target motion data into the target mapping model, and based on the output of the target mapping model, to map target facial data that can be used for facial binding on a digital human face.

[0170] Among them, the digital human face and motion capture model is associated with the target mapping model.

[0171] The motion mapping device of this invention, in the context of a target person performing a target action, acquires a target video of the target person's face after facial capture by a target action data acquisition module, and inputs the target video into a motion capture model to obtain target action data. Then, a target mapping model acquisition module acquires a target mapping model trained according to the model training method provided in any embodiment of this invention. Finally, a target facial data acquisition module inputs the target action data into the target mapping model and, based on the output of the target mapping model, maps it to target facial data that can be used for facial binding on a digital human face. The motion mapping device of this invention can perform motion mapping with low cost and workload for facial binding.

[0172] The motion mapping device provided in the embodiments of the present invention can execute the motion mapping method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0173] It is worth noting that in the embodiments of the above-mentioned motion mapping device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.

[0174] Example 7

[0175] Figure 15 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0176] like Figure 15 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded into the RAM 13 from storage unit 18. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0177] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0178] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as model training methods or action mapping methods.

[0179] In some embodiments, the model training method or action mapping method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the model training method or action mapping method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to execute the model training method or action mapping method by any other suitable means (e.g., by means of firmware).

[0180] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0181] Computer programs used to implement the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs can be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0182] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0183] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0184] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0185] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0186] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0187] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A model training method, characterized in that, include: Obtain sample facial data that has been facially bound to a digital human face; The sample facial data is used to drive the digital human face to perform sample actions, and the digital human face performing the sample actions is captured to obtain sample videos; The sample video is input into the motion capture model to obtain sample motion data; Based on the sample action data as actual input data and the sample facial data as expected output data, the original mapping model is trained to obtain the target mapping model. The sample facial data consists of multiple frames. After acquiring the sample facial data that has been facially bound to the digital human face, the method further includes: For each of the multiple video segments composed of the sample facial data of multiple frames, a target audio signal is set for the first frame of the sample facial data in the video segment; After obtaining the sample video, the method further includes: The sound components in the sample video are detected to obtain sample video frames containing the target sound signal, and the video frame position of the sample video frame containing the target sound signal in the sample video is determined. By combining the sample video frames corresponding to each of the sample motion data in the sample video, position motion data corresponding to the position of the video frame is determined from each of the sample motion data; Based on the positional motion data and the duration of each video segment, the sample motion data that requires cropping is cropped, and the captured sample motion data is updated based on the cropped sample motion data.

2. The method of claim 1, wherein, After obtaining the sample action data, the method further includes: The sample action data is positionally encoded to obtain positional features, and the positional features are combined and arranged to obtain a sample image; The sample action data is placed in the sample image, and the resulting sample image is used as the sample action data.

3. The method of claim 2, wherein, The original mapping model used for regression includes a convolutional neural network; Based on the sample action data as actual input data and the sample facial data as expected output data, the original mapping model is trained to obtain the target mapping model, including: The sample action data, which is used as actual input data, is input into the convolutional neural network to extract features from the sample action data through the convolutional neural network, and regression facial data corresponding to the sample facial data is regressed based on the feature extraction results. The sample facial data, which is the expected output data, and the regression facial data, which is the actual output data, are compared, and the network parameters in the original mapping model are adjusted according to the comparison results to train the target mapping model.

4. The method according to claim 1, characterized in that, Before combining the sample motion data of each of the sample videos into their respective corresponding sample video frames in the sample video, the method further includes: The sample action data is resampled, and the captured sample action data is updated based on the resampling results.

5. The method according to claim 1, characterized in that, The acquisition of sample facial data that has been facially bound to a digital human face includes: Obtain a set of facial data, wherein the set of facial data includes multiple candidate facial data, and each candidate facial data has been facially bound to a digital human face; Select at least two candidate facial data from the facial dataset; Interpolation is performed on each of the selected candidate facial data, and the interpolated facial data and each of the selected candidate facial data are used as sample facial data that have been facially bound to the digital human face.

6. An action mapping method, characterized in that, include: When a target person performs a target action on their face, acquire the target video obtained after capturing the target person's face, and input the target video into the motion capture model to obtain target action data; Obtain the target mapping model trained according to the model training method of any one of claims 1-5; The target motion data is input into the target mapping model, and the target facial data that can be mapped onto the digital human face is obtained based on the output of the target mapping model. The digital human face and the motion capture model are associated with the target mapping model.

7. A model training device, characterized in that, include: The sample facial data acquisition module is used to acquire sample facial data that has been facially bound to the digital human face; The sample video acquisition module is used to drive the digital human face to perform sample actions through the sample facial data, and to capture the digital human face performing the sample actions to obtain sample videos; The sample motion data acquisition module is used to input the sample video into the motion capture model to obtain sample motion data; The target mapping model acquisition module is used to train the original mapping model based on the sample action data as actual input data and the sample facial data as expected output data to obtain the target mapping model. The sample facial data consists of multiple frames, and the model training device further includes: The target audio signal setting module is used to set a target audio signal for the first frame of the sample facial data in each of the multiple video segments composed of multiple frames of the sample facial data after the sample facial data has been acquired and facially bound to the digital human face. The model training device further includes: The video frame position determination module is used to detect the sound components in the sample video, obtain the sample video frame containing the target sound signal in the sample video, and determine the video frame position of the sample video frame containing the target sound signal in the sample video. The position motion data determination module is used to combine the sample motion data of each sample video with the sample video frames corresponding to the sample video, and determine the position motion data corresponding to the position of the video frame from the sample motion data of each sample video. The first sample motion data update module is used to crop the sample motion data that requires cropping based on the positional motion data and the duration of each video segment, and to update the captured sample motion data based on the cropped sample motion data.

8. A motion mapping device, characterized in that, include: The target motion data acquisition module is used to acquire a target video obtained after capturing the target person's face when the target person performs a target action, and input the target video into the motion capture model to obtain target motion data; The target mapping model acquisition module is used to acquire the target mapping model trained according to the model training method of any one of claims 1-5; The target facial data acquisition module is used to input the target motion data into the target mapping model, and map the target facial data that can be bound to the digital human face according to the output of the target mapping model, wherein the digital human face and the motion capture model are associated with the target mapping model.

9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to cause the at least one processor to perform the model training method as described in any one of claims 1-5, or the action mapping method as described in claim 6.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute and implement the model training method as described in any one of claims 1-5, or the action mapping method as described in claim 6.