Mapping network training method, face driving parameter mapping method, device and equipment

CN122821606APending Publication Date: 2026-09-25SUZHOU ZHIJUXINLIAN MICROELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611176998.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-05
Publication Date
2026-09-25

AI Technical Summary

Benefits of technology

[0017]本申请实施例提供的技术方案中,映射网络包括并行的多个分支神经网络。在训练阶段,通过从样本人脸图像中提取N维面部形变特征和M维面部驱动参数,将M维面部驱动参数作为样本标签与N维面部形变特征配对以构建训练样本集。通过将每个训练样本的样本标签按面部属性进行维度拆分,得到各分支神经网络对应的参数分量标签,并对于每一个分支神经网络,基于该分支神经网络预测的驱动参数分量与对应的参数分量标签计算损失,并利用独立的优化器根据损失更新参数,且各分支神经网络在反向传播时梯度相互隔离。如此,各分支神经网络之间是物理隔离的,且在训练过程中仅受其对应面部属性的监督信号驱动,参数更新相互独立,避免了不同面部属性之间的梯度干扰;同时,样本标签按面部属性进行维度拆分后独立监督各分支神经网络,使得各分支神经网络能够更准确地学习对应面部属性的低维参数到高维参数的映射关系。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821606A_ABST
    Figure CN122821606A_ABST
Patent Text Reader

Abstract

The application provides a mapping network training method, a face driving parameter mapping method, a device and equipment. The mapping network training method comprises: obtaining an image sequence composed of multiple frames of face images from a sample video; for each frame of face image in the image sequence, extracting N-dimensional face deformation features and M-dimensional face driving parameters, pairing the M-dimensional face driving parameters as sample labels with the N-dimensional face deformation features to form a training sample, and constructing a training sample set; for each training sample, inputting the N-dimensional face deformation features into a mapping network, and dimensionally splitting the M-dimensional face driving parameters as sample labels according to face attributes to obtain corresponding parameter component labels; for each branch neural network in the mapping network, calculating a loss according to the predicted driving parameter component and the corresponding parameter component label, and updating parameters according to the loss by using an independent optimizer to obtain a trained mapping network, and gradient isolation when each branch neural network is back propagated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a mapping network training method, a face driving parameter mapping method, an apparatus and device. Background Technology

[0002] With the development of metaverse, virtual reality (VR), embodied intelligence, and high-fidelity avatar technologies, real-time driving of 3D facial animation has become a crucial link connecting real users and virtual avatars. In rendering pipelines based on 3D Gaussian sputtering (3DGS), to ensure the anatomical rigor and physical rendering quality of digital human facial deformation, 3D parametric face models such as FLAME (Faces Learned with an Articulated Model and Expressions) or SMPL-X (Skinned Multi-Person Linear model with eXpressive features) are typically used as the underlying control platform for the digital human. These models drive the facial deformation of the digital human with high precision through high-dimensional nonlinear parameters (e.g., 109-dimensional parameters).

[0003] However, in practical applications, facial expression capture of real users (such as using ordinary smartphones, monocular RGB cameras combined with algorithms for extraction), and most digital human assets modeled manually in a traditional way, often use facial blending shape coefficients as the data exchange protocol (e.g., the 52-dimensional facial blending shape (BS) standard). This results in two situations: on the one hand, easily accessible, low-dimensional consumer-grade facial blending shape-driven data; on the other hand, high-dimensional industrial-grade 3D parametric face model-driven parameters. How to convert low-dimensional facial blending shape coefficients into high-dimensional 3D parametric face model-driven parameters is a pressing technical problem that needs to be solved in the current digital human industrial pipeline. Summary of the Invention

[0004] In view of this, embodiments of this application provide a mapping network training method, a face driving parameter mapping method, an apparatus, and a device to solve at least one problem existing in the background art.

[0005] Firstly, a method for training a mapping network is provided, the mapping network being used to generate facial driving parameters, the mapping network comprising multiple parallel branch neural networks; the training method includes: Obtain an image sequence consisting of multiple frames of face images from the sample video; For each frame of face image in the image sequence, N-dimensional facial deformation features and M-dimensional facial driving parameters are extracted. The M-dimensional facial driving parameters are used as sample labels and paired with the N-dimensional facial deformation features to form a training sample to construct a training sample set. Here, M and N are both positive integers, and M is greater than N. For each training sample, the N-dimensional facial deformation features are input into the mapping network, and the M-dimensional facial driving parameters, which serve as sample labels, are split according to facial attributes to obtain the corresponding parameter component labels. For each branch neural network in the mapping network, the corresponding loss is calculated based on the driving parameter components predicted by the branch neural network and the corresponding parameter component labels. The parameters of the branch neural network are updated based on the loss using an independent optimizer to obtain the trained mapping network. The gradients of each branch neural network are isolated from each other during backpropagation.

[0006] In some embodiments, the M-dimensional facial driving parameters include X-dimensional expression parameters, Y-dimensional jaw posture parameters, and Z-dimensional eye posture parameters, where X+Y+Z=M; the plurality of branch neural networks include: An expression branch neural network is used to predict the X-dimensional expression parameters based on the N-dimensional facial deformation features. A mandibular branch neural network is used to predict the Y-dimensional mandibular posture parameters based on the N-dimensional facial deformation features. An eye-branching neural network is used to predict the Z-dimensional eye pose parameters based on the N-dimensional facial deformation features.

[0007] In some embodiments, the facial expression branch neural network, the mandibular branch neural network, and the ocular branch neural network are all multilayer perceptrons, and each branch neural network includes a hidden layer, a batch normalization layer, and a ReLU activation layer; wherein, the feature dimensions of the hidden layers used for dimensionality enhancement in the mandibular branch neural network and the ocular branch neural network are smaller than the feature dimensions of the hidden layers used for dimensionality enhancement in the facial expression branch neural network.

[0008] In some embodiments, the facial expression branch neural network includes a first hidden layer, a first batch of normalized layers, a first ReLU activation layer, a second hidden layer, a second batch of normalized layers, a second ReLU activation layer, and a first output layer connected in sequence; wherein, the first hidden layer receives the N-dimensional facial deformation features and increases the input dimension to 256 dimensions, the second hidden layer outputs 128 dimensions, and the first output layer outputs 100-dimensional facial expression parameters. The mandibular branch neural network includes a third hidden layer, a third batch normalization layer, a third ReLU activation layer, a fourth hidden layer, a fourth batch normalization layer, a fourth ReLU activation layer, and a second output layer connected in sequence; wherein, the third hidden layer receives the N-dimensional facial deformation features and increases the input dimension to 128-dimensional features, the fourth hidden layer outputs 64-dimensional features, and the second output layer outputs 3-dimensional mandibular posture parameters. The eye branch neural network includes a fifth hidden layer, a fifth batch normalization layer, a fifth ReLU activation layer, a sixth hidden layer, a sixth batch normalization layer, a sixth ReLU activation layer, and a third output layer connected in sequence; wherein, the fifth hidden layer receives the N-dimensional facial deformation features and increases the input dimension to 128-dimensional features, the sixth hidden layer outputs 64-dimensional features, and the third output layer outputs 6-dimensional eye pose parameters.

[0009] In some embodiments, for each frame of a face image in the image sequence, N-dimensional facial deformation features and M-dimensional facial driving parameters are extracted, including: A facial blending shape model is used to extract blending shape coefficients from the face image, each dimension of which is used to characterize the activation intensity of a facial muscle group; Facial driving parameters are extracted from the face image using a three-dimensional deformable face model. These facial driving parameters include expression parameters, jaw pose parameters, and eye pose parameters.

[0010] Secondly, a facial driving parameter mapping method is provided, the method comprising: Obtain the target face image; N-dimensional facial deformation features are extracted from the target face image, and each dimension of the N-dimensional facial deformation features corresponds to a facial blending shape coefficient. The N-dimensional facial deformation features are input into a mapping network, and multiple parallel branch neural networks in the mapping network predict the driving parameter components used to describe different facial attributes. The driving parameter components predicted by the multiple branch neural networks are concatenated according to their dimensions to obtain M-dimensional facial driving parameters, where M and N are both positive integers and M is greater than N. The M-dimensional facial driving parameters are used to drive the facial deformation of the three-dimensional virtual object. The mapping network is obtained through supervised training based on a training sample set; each training sample in the training sample set includes N-dimensional facial deformation features extracted from the sample face image and M-dimensional facial driving parameters as sample labels.

[0011] In some embodiments, the training process of the mapping network includes: Obtain an image sequence consisting of multiple frames of face images from the sample video; For each frame of a face image in the image sequence, N-dimensional facial deformation features and M-dimensional facial driving parameters are extracted. The extracted M-dimensional facial driving parameters are used as sample labels and paired with the N-dimensional facial deformation features of the face image to form a training sample, so as to construct a training sample set. For each training sample, the N-dimensional facial deformation features are input into the mapping network, and the M-dimensional facial driving parameters, which serve as sample labels, are split according to facial attributes to obtain the corresponding parameter component labels. For each branch neural network in the mapping network, the corresponding loss is calculated based on the driving parameter components predicted by the branch neural network and the corresponding parameter component labels. The parameters of the branch neural network are updated based on the loss using an independent optimizer to obtain the trained mapping network. The gradients of each branch neural network are isolated from each other during backpropagation.

[0012] In some embodiments, the M-dimensional facial driving parameters include X-dimensional expression parameters, Y-dimensional jaw posture parameters, and Z-dimensional eyeball posture parameters, where X+Y+Z=M; The plurality of branched neural networks include: An expression branch neural network is used to predict the X-dimensional expression parameters based on the N-dimensional facial deformation features. A mandibular branch neural network is used to predict the Y-dimensional mandibular posture parameters based on the N-dimensional facial deformation features. An eye-branching neural network is used to predict the Z-dimensional eye pose parameters based on the N-dimensional facial deformation features.

[0013] Thirdly, a mapping network training apparatus is provided, the mapping network being used to generate facial driving parameters, the mapping network comprising multiple parallel branch neural networks; the apparatus includes: The first acquisition module is used to acquire an image sequence consisting of multiple frames of face images from the sample video; The first extraction module is used to extract N-dimensional facial deformation features and M-dimensional facial driving parameters for each frame of face image in the image sequence. The sample construction module is used to pair the M-dimensional facial driving parameters as sample labels with the N-dimensional facial deformation features to form a training sample, so as to construct a training sample set, where M and N are both positive integers, and M is greater than N; Supervised training module, used for: For each training sample, the N-dimensional facial deformation features are input into the mapping network, and the M-dimensional facial driving parameters, which serve as sample labels, are split according to facial attributes to obtain the corresponding parameter component labels. For each branch neural network in the mapping network, the corresponding loss is calculated based on the driving parameter components predicted by the branch neural network and the corresponding parameter component labels. The parameters of the branch neural network are updated based on the loss using an independent optimizer to obtain the trained mapping network. The gradients of each branch neural network are isolated from each other during backpropagation.

[0014] Fourthly, a facial driving parameter mapping device is provided, the device comprising: The second acquisition module is used to acquire the target face image; The second extraction module is used to extract N-dimensional facial deformation features from the target face image, where each dimension of the N-dimensional facial deformation features corresponds to a facial blending shape coefficient. The parameter prediction module is used to map the N-dimensional facial deformation feature network, and the multiple branch neural networks in the mapping network predict the driving parameter components used to describe different facial attributes respectively. The parameter splicing module is used to splice the driving parameter components predicted by the multiple branch neural networks according to the dimensions to obtain M-dimensional facial driving parameters, where M and N are both positive integers and M is greater than N. The M-dimensional facial driving parameters are used to drive the facial deformation of the three-dimensional virtual object. The mapping network is obtained through supervised training based on a training sample set; each training sample in the training sample set includes N-dimensional facial deformation features extracted from the sample face image and M-dimensional facial driving parameters as sample labels.

[0015] Fifthly, an electronic device is provided, including a processor, the processor being configured to invoke instructions to cause the electronic device to perform the mapping network training method as described in the first aspect, or the face driving parameter mapping method as described in the second aspect.

[0016] In a sixth aspect, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, which, when executed by a processor, implements the mapping network training method of the first aspect, or the face driving parameter mapping method of the second aspect.

[0017] In the technical solution provided in this application embodiment, the mapping network includes multiple parallel branch neural networks. During the training phase, N-dimensional facial deformation features and M-dimensional facial driving parameters are extracted from sample face images. The M-dimensional facial driving parameters are used as sample labels and paired with the N-dimensional facial deformation features to construct a training sample set. By splitting the sample labels of each training sample according to facial attributes, parameter component labels corresponding to each branch neural network are obtained. For each branch neural network, the loss is calculated based on the driving parameter components predicted by that branch neural network and the corresponding parameter component labels. An independent optimizer is used to update the parameters according to the loss, and the gradients of each branch neural network are isolated during backpropagation. Thus, the branch neural networks are physically isolated and are driven only by the supervision signals of their corresponding facial attributes during training. Parameter updates are independent, avoiding gradient interference between different facial attributes. Simultaneously, the sample labels, after being split according to facial attributes, independently supervise each branch neural network, enabling each branch neural network to more accurately learn the mapping relationship from low-dimensional parameters to high-dimensional parameters of the corresponding facial attributes.

[0018] During the inference phase, for the N-dimensional facial deformation features extracted from the target face image, the driving parameter components are predicted by each branch neural network in the trained mapping network, and then concatenated according to the dimensions to form M-dimensional facial driving parameters (M>N). This achieves the mapping from low-dimensional facial deformation features to high-dimensional facial driving parameters, which can be adapted to various virtual objects and rendering pipelines. Attached Figure Description

[0019] Figure 1 A flowchart illustrating the mapping network training method provided in this application embodiment; Figure 2 for Figure 1 The flowchart of step S102 is shown below; Figure 3 A flowchart illustrating the face driving parameter mapping method provided in this application embodiment; Figure 4 A flowchart illustrating the mapping method from BS parameters to FLAME parameters based on a feature-decoupled cross-domain distillation neural network provided in this application embodiment; Figure 5a Two sets of rendering comparison diagrams of the mandibular closed state provided for embodiments of this application; Figure 5b Two sets of rendering comparison diagrams of the mandibular open state provided in the embodiments of this application; Figure 5c The training loss convergence curves of the three independent branch networks provided in the embodiments of this application; Figure 5d A schematic diagram of the branch network test results provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of the mapping network training device provided in the embodiments of this application; Figure 7 This is a schematic diagram of the structure of the face driving parameter mapping device provided in the embodiments of this application. Detailed Implementation

[0020] To make the technical solution and beneficial effects of this application more apparent and understandable, a detailed description is provided below by listing specific embodiments. The accompanying drawings are not necessarily drawn to scale, and local features may be enlarged or reduced to more clearly show the details of the local features; unless otherwise defined, the technical and scientific terms used herein have the same meanings as those in the technical field to which this application pertains.

[0021] The embodiments in this application are not exhaustive, but merely illustrative of some embodiments, and are not intended to limit the scope of protection of this disclosure. Unless otherwise specified, each step in a particular embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a particular embodiment can also be implemented as an independent embodiment, and the order of the steps in a particular embodiment can be arbitrarily interchanged. Furthermore, the optional implementation methods in a particular embodiment can be arbitrarily combined; moreover, the embodiments can be arbitrarily combined, for example, some or all steps of different embodiments can be arbitrarily combined, and a particular embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.

[0022] In each embodiment of this application, unless otherwise specified or in case of logical conflict, the terminology and / or descriptions of the embodiments are consistent and can be referenced by each other. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.

[0023] In the description of the embodiments of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0024] To facilitate understanding, a brief explanation of some of the terms used in this application will be provided first.

[0025] Blendshape (BS for short): a common data exchange standard for consumer-grade facial capture (such as Apple ARKit, ordinary camera facial capture). It synthesizes the final expression by linearly superimposing the weight values ​​(0.0~1.0) of preset local expression bases (such as opening the mouth, blinking, etc., a total of 52 dimensions).

[0026] FLAME (Articulated Expression-Controlled 3D Parametric Head Model): A high-precision statistical face model. It precisely controls a complex 3D mesh through dimensionality-reduced nonlinear parameters (including expression parameters, jaw pose, and eye pose, totaling 109 dimensions).

[0027] 3DGS (3D Gaussian Splatting): A cutting-edge high-fidelity 3D scene and digital human rendering technology that does not rely on traditional polygonal meshes, but instead renders using Gaussian ellipsoids.

[0028] MediaPipe: A lightweight, open-source machine learning algorithm framework that can regress 3D facial keypoints from 2D images and output 52-dimensional BS parameters.

[0029] VHAP (Video-based High-fidelity Avatar Parameterization): A high-precision video parameterization algorithm that can extract extremely high-precision FLAME parameters from video frames through synthesis analysis and complex energy optimization.

[0030] Cross-linking phenomenon (feature coupling distortion): In facial animation, when a physiological part (such as the jaw opening) moves violently, due to the lack of decoupling of algorithm features, it incorrectly drives another unrelated physiological part (such as eyelid twitching).

[0031] For cross-domain facial driving parameter mapping problems (e.g., mapping from low-dimensional facial blending shape coefficients to high-dimensional 3D parametric face model driving parameters), the following technical solutions exist: Option 1: Direct linear superposition based on 3D mesh and deformation matrix. This option does not involve transformation of the parameter latent space. It directly extracts the 52D deformation basis matrix from the standard ARKit (usually a pre-calculated vertex offset matrix, with dimensions such as 52*5023*3), performs linear matrix multiplication with the input 52D BS coefficients, and directly superimposes the 3D offset onto the base 3D mesh with zero facial expression to generate animation. Because this option is based on 3D matrix multiplication, the number of vertices in the deformation matrix and the base mesh must be strictly consistent (e.g., they must correspond precisely to 5023 points). However, to improve the realism of the digital human, when the number of vertices changes due to the addition of modules such as teeth or eyeballs (e.g., changing to an irregular number such as 5118), the matrix multiplication algorithm will fail due to dimensional mismatch. At the same time, it cannot be used for 3D Gaussian sputtering digital humans with non-mesh structures.

[0032] Option 2: Nonlinear Optimization Fitting Method Based on Energy Function. This method constructs a target energy function, calculates the geometric projection error between the 3D keypoints generated by the input BS expression and those generated by the FLAME model, and uses iterative optimization algorithms such as L-BFGS (Limited-memory Broyden Fletcher Goldfarb Shanno) or gradient descent to solve for the corresponding FLAME parameters frame by frame. For example, the SMPLify-X optimization framework constructs a target energy function that includes the 3D keypoint reprojection error and prior penalty terms for expression / pose, and uses an L-BFGS optimizer to forcibly find the FLAME facial parameters and SMPL (Skinned Multi-Person Linear Model) body parameters that minimize the energy function after hundreds of iterations. This method involves a nonlinear inverse solution process, requiring high-dimensional iterative optimization in each frame, which consumes significant CPU / GPU computing resources. Excessive computing power overhead directly leads to a convergence latency of tens or even hundreds of milliseconds for single-frame processing, making it difficult to meet the high concurrency and low latency requirements of 30fps (within about 33 milliseconds) for game engines or VTuber (Virtual YouTuber) live streaming.

[0033] Option 3: A single end-to-end mapping network with shared representations. As a standard baseline approach in deep learning for parameter regression using fully connected layers, this method constructs a multilayer perceptron (MLP) or a single decoder. Its input layer receives images of a person's face, extracts features through shared hidden layers, and finally outputs a unified set of mixed parameters (directly concatenated 109-dimensional FLAME parameters) through a single output layer. The network is trained by backpropagation by calculating the overall mean squared error. However, the inventors found that this approach lacks explicit decoupling of physiologically independent joints in the hidden layers, easily leading to feature crosstalk. Facial muscle tension and jaw opening / closing are physiologically independent motion systems. Because the single network uses a parameter-sharing hybrid fitting mechanism in the hidden layers, this results in gradient interference during backpropagation. For example, when processing a complex expression like "laughing with slight eye movement," the dominant gradient generated by a large jaw opening may directly mask or deflect subtle gradients in the eyes or micro-expressions. This lack of decoupling at the algorithm level in the extraction of physiological features makes it easy for cross-linking phenomena to occur during model inference. For example, when a person opens their mouth wide, their eyelids will exhibit unnatural physiological twitching.

[0034] To achieve the mapping from low-dimensional facial blending shape coefficients to high-dimensional 3D parametric face model driving parameters, this application provides a mapping network training method that can be applied to digital human-driven scenarios, such as real-time facial animation driving of 3D Gaussian sputtering (3DGS) digital humans, real-time expression capture and redirection of virtual anchors, and facial animation control of virtual characters in game engines. This method can be executed by computer devices, such as servers, desktop computers, laptops, tablets, or smartphones.

[0035] Figure 1 This is a flowchart illustrating the mapping network training method provided in an embodiment of this application. The mapping network is used to generate facial driving parameters, and the mapping network includes multiple parallel branch neural networks. Figure 1 As shown, the method includes: S101: Obtain an image sequence consisting of multiple frames of face images from the sample video; S102: For each frame of face image in the image sequence, extract N-dimensional facial deformation features and M-dimensional facial driving parameters, and pair the M-dimensional facial driving parameters as sample labels with the N-dimensional facial deformation features to form a training sample, so as to construct a training sample set, where M and N are both positive integers, and M is greater than N; S103: For each training sample, input the N-dimensional facial deformation features into the mapping network, and split the M-dimensional facial driving parameters, which serve as sample labels, according to facial attributes to obtain the corresponding parameter component labels. S104: For each branch neural network in the mapping network, calculate the corresponding loss based on the driving parameter components predicted by the branch neural network and the corresponding parameter component labels, and use an independent optimizer to update the parameters of the branch neural network according to the loss to obtain the trained mapping network. The gradients of each branch neural network are isolated from each other during backpropagation.

[0036] In some examples, the sample video can be a real person's monocular RGB video containing continuous and rich facial expressions, or it can be a high-fidelity virtual person video generated by a 3D rendering engine. This embodiment does not limit the source of the sample video.

[0037] Extracting the sample video frame by frame into a temporally continuous image sequence ensures that the N-dimensional facial deformation features extracted from each frame correspond strictly with the M-dimensional facial driving parameters in the time dimension, thereby ensuring frame-level matching between the N-dimensional facial deformation features and the M-dimensional facial driving parameter labels in the training samples.

[0038] In some examples, the N-dimensional facial deformation features include blending shape coefficients extracted from a face image, where each dimension of the blending shape coefficients represents the activation intensity of a specific facial muscle group or muscle combination. The value of N can be 52, resulting in 52-dimensional facial blending shape coefficients. In other possible implementations, N can also take other values, such as 40 or 50, depending on the number of blending shape bases used.

[0039] In some examples, the M-dimensional face driving parameters include three-dimensional parameterized face model parameters (e.g., FLAME parameters) extracted from the face image. M is equal to the sum of the output dimensions of each branch of the neural network in the mapping network.

[0040] For example, the M-dimensional facial driving parameters include 109 dimensions of facial driving parameters, of which expression parameters are 100 dimensions, jaw pose parameters are 3 dimensions, and eye pose parameters are 6 dimensions. Alternatively, in other examples, a lightweight driving parameter configuration can be used, which includes 80 dimensions of expression parameters, 4 dimensions of jaw pose parameters (including displacement), and 3 dimensions of eye pose parameters (only gaze direction). In this example, the lightweight driving parameter configuration is 87 dimensions.

[0041] In this embodiment, the M-dimensional facial driving parameters extracted from the sample face image are used as sample labels. In this way, the labels of the training samples can be obtained automatically by the model without relying on manually labeled high-dimensional parameters.

[0042] In the process of constructing the training sample set, data augmentation can also be introduced, such as adding Gaussian noise or random masking of some dimensions to the sample images, in order to improve the generalization ability of the mapping network.

[0043] Different branch neural networks are used to output driving parameter components that represent different facial attributes. The number of branch neural networks can be determined according to the types of facial attributes that need to be driven. Here, facial attributes refer to categories of physiological or postural features related to facial movements, such as expression attributes, jaw posture attributes, and / or eye posture attributes.

[0044] In the mapping network, multiple branch neural networks are structurally parallel and do not share parameters. For example, the number of branch neural networks can be three, including an expression branch neural network, a jaw branch neural network, and an eye branch neural network, which output expression parameters, jaw pose parameters, and eye pose parameters, respectively. Alternatively, the number of branch neural networks can be two, including an expression branch and a head pose branch, etc. This application does not limit the number of branches.

[0045] In some examples, each branch of the neural network can be implemented based on a multilayer perceptron (MLP) or a residual network (such as a ResNet-style MLP) to alleviate the gradient vanishing problem in deep network training; alternatively, a one-dimensional convolutional neural network can be used, utilizing convolutional kernels to capture the correlation between adjacent B / S dimensions; furthermore, a Transformer encoder can be used, with self-attention to capture long-range dependencies, suitable for extreme expressions. This embodiment does not impose specific limitations on these aspects.

[0046] In some examples, in step S104 above, the loss of each branch neural network can be mean squared error loss, mean absolute error (MAE), Huber loss, or smoothed L1 loss; or a perceptual loss can be introduced, for example, calculating the chamfer distance between the FLAME grid vertices corresponding to the driving parameters of the mapping network output and the ground truth grid vertices as an additional supervision signal to enhance the mapping accuracy at the geometric level.

[0047] Each branch of the neural network uses an independent optimizer for parameter updates during training, and the gradients of each branch are isolated during backpropagation. For example, each branch uses an independent Adam (Adaptive Moment Estimation) optimizer for parameter updates. Each branch's Adam optimizer only receives the gradient signal of its own branch's loss function and updates only the network parameters of its own branch, completely isolated from the optimization processes of other branches. Alternatively, each branch could also use an AdamW (Adaptive Moment Estimation Decoupled Weight Decay) optimizer or a stochastic gradient descent with momentum (SGD+Momentum) optimizer.

[0048] In this embodiment, each branch neural network takes N-dimensional facial deformation features as input and performs parameter mapping independently, predicting the driving parameter components used to describe different facial attributes. The branch neural networks do not share parameters, nor are they cascaded or feature-transferred. Furthermore, each branch uses an independent optimizer during training, and gradients are isolated from each other, ensuring complete decoupling of the parameter mapping process for each facial attribute. This improves the mapping accuracy of each branch neural network.

[0049] In the aforementioned mapping network training method, the neural networks of each branch are physically isolated and driven only by the supervision signals of their corresponding facial attributes during training. Parameter updates are independent of each other, avoiding gradient interference between different facial attributes. Simultaneously, the sample labels are split dimensionally according to facial attributes and then independently supervise each branch neural network, enabling each branch to more accurately learn the mapping relationship from low-dimensional parameters to high-dimensional parameters of the corresponding facial attributes. Therefore, the trained mapping network can decouple the input low-dimensional N-dimensional facial deformation features and map them into driving parameter components corresponding to each facial attribute, which are then concatenated to obtain high-dimensional M-dimensional facial driving parameters.

[0050] In some embodiments, such as Figure 2 As shown, in step S102 above, for each frame of face image in the image sequence, N-dimensional facial deformation features and M-dimensional facial driving parameters are extracted, which may include steps S201 to S202.

[0051] S201: Extract blend shape coefficients from a face image using a facial blend shape model, where each dimension of the blend shape coefficient is used to characterize the activation intensity of a facial muscle group.

[0052] For example, for an input single-frame RGB image This involves calling a deep learning-based facial feature extraction model (e.g., using the MediaPipe machine learning algorithm framework). The underlying data flow of this framework comprises two stages: facial landmark regression and Black-Scholes weight calculation. The input is the t-th frame of the RGB image, whose dimensions are represented as H×W×3; H and W represent the height and width of the image in pixels, respectively.

[0053] Phase 1: Inputting the facial region into a high-density facial mesh regression network 478 three-dimensional facial landmarks were predicted. The 478 three-dimensional facial landmarks constitute a dense facial mesh, including contours, sparse key points of facial features, and dense key points of the cheeks, forehead, and inside the lips. Each key point contains three spatial coordinate components, wherein: ; Phase Two: The system introduces predefined neutral face topology templates. Calculate the three-dimensional geometric displacement vector of the current detection point. Subsequently, the pre-trained Blendshape projection matrix is ​​used. Bias term of the linear regression model Spatial displacement is mapped to a 52-dimensional feature space using a linear regression model with a truncated activation function: ; in, The function ensures that the output is a floating-point weight vector. Strictly constrained Within the interval, each component (j=1,2,...,52) represents the normalized activation intensity of specific facial muscle groups (such as jaw opening and closing, orbicularis oculi muscle contraction, etc.). This weight vector serves as the input feature of the mapping network.

[0054] Taking the 52-dimensional facial blending shape coefficient as an example, some of its dimensions and their corresponding facial movements are as follows: jaw opening represents the amplitude of jaw opening and closing; mouth smile left represents the activation intensity of the left side of the smiling action; eye blink left represents the degree of left eyelid closure; brow down represents the intensity of frowning action; cheek puff represents the degree of inflating of both cheeks; eye look up left represents the coordinated movement of the left eye in the vertical and horizontal directions; mouth pucker represents the action of lips closing and protruding forward; and nose sneeze left represents the activation intensity of the left side of the nose sniffling or disgust expression.

[0055] It is understood that, in other possible implementations, the facial blending shape model can also be a model trained based on the action units defined by the Facial Action Coding System (FACS), which maps the intensity of each action unit to the face image and then converts it into the corresponding facial blending shape coefficients through a preset mapping relationship; or it can be an end-to-end regression model based on a convolutional neural network or Transformer architecture, which takes the face image as input and directly outputs N blending shape coefficient values.

[0056] In some examples, considering that facial blending shape models such as MediaPipe, when outputting the aforementioned N-dimensional (e.g., 52-dimensional) blending shape coefficients, may carry at least one additional auxiliary state bit (e.g., a state bit for representing gender), forming an original feature vector with a dimension greater than N. To ensure strict matching of the feature dimensions of the input mapping network, the system can uniformly truncate or mask that dimension to zero during preprocessing. For example, a binary mask vector can be introduced. The original feature vector is cleaned in terms of dimensions, where, This is the dimension of the returned list. For invalid features representing neutral baseline expressions (such as "_neutral"), the corresponding mask value is 0; for other valid physiological features, the mask value is 1. The cleaned input tensor The calculation is as follows: ; in, This represents the Hadamard product (i.e., element-wise multiplication). For a 52-dimensional BS input feature tensor, It is a binary mask vector.

[0057] Thus, by performing dimensionality cleaning on the original feature vector, 52-dimensional hybrid shape coefficients that effectively represent the activation intensity of facial muscle groups can be retained. According to the preset feature arrangement order (such as alphabetical order), the 52-dimensional hybrid shape coefficients are organized into batch input tensors and used as inputs to the mapping network.

[0058] S202: Extract facial driving parameters from the face image using a three-dimensional deformable face model, the facial driving parameters including expression parameters, jaw pose parameters, and eye pose parameters.

[0059] For example, a 3D deformable face model (such as the FLAME model) can invoke the VHAP algorithm to process the exact same image frame. VHAP uses an analysis-by-synthesis multi-constraint optimization architecture for offline preprocessing.

[0060] Taking the FLAME model as an example, the system defines the generation function of the FLAME 3D parametric mesh. It calculates the 3D coordinates of 5023 vertices using a linear blend skinning (LBS) algorithm. : ; in, With zero facial expression base, For head shape parameters, For facial expression parameters, These are the posture parameters (including jaw posture and eye posture). This is a shape basis transformation function used to define the basic facial contours and skeletal features of a digital human through parameter combinations. This is a facial expression base transformation function used to superimpose parameter-controlled global facial muscle deformations onto a base contour. This is the attitude basis transformation function, used to dynamically compensate for physical distortions such as mesh volume collapse caused by large rotations of the jaw or eyeballs. For skinning functions, For skin weights.

[0061] To solve for the precise parameters of the current frame, VHAP constructs a global energy objective function that includes three core constraints. : ; in, This indicates the reprojection error of the landmark. This indicates the photometric consistency error. This represents the prior regularization loss. , , These represent the weighting coefficients for each loss item.

[0062] Landmark reprojection error The calculation formula is: ; in, This represents a predefined set of keypoint indices, such as facial feature points on a FLAME model; This represents the 3D spatial coordinates of the k-th keypoint in the FLAME mesh. Detected in face images The corresponding two-dimensional feature point coordinates; This is a perspective projection matrix used to project three-dimensional spatial points onto a two-dimensional image plane; and To represent the global rotation matrix and translation vector of the head, describing the pose of the head in the camera coordinate system; is the confidence weight of the k-th key point.

[0063] Photometric consistency error The calculation formula is: ;in, This represents a differentiable renderer used to generate 2D rendered images based on 3D geometry, texture, and lighting parameters; V represents the 3D mesh vertices, and Albedo represents the surface albedo texture parameter. This represents lighting model parameters, such as spherical harmonic lighting coefficients; This represents the face image in frame t. This represents the L1 norm, which is the sum of the absolute differences between pixel values ​​in the rendered image and the target face image within the facial region, calculated pixel by pixel. Photometric consistency error is addressed by utilizing a differentiable renderer. Combined with spherical harmonic lighting model Render the generated image and calculate its similarity to the original face image. In the face mask area within Pixel-level color error was obtained.

[0064] Prior regularization loss The calculation formula is: ;in, , , These are the prior regularization weights for shape, expression, and posture parameters, respectively. By using the Gaussian prior assumption to restrict the parameter value space, non-physiological deformations can be avoided.

[0065] In each frame of the image, the VHAP algorithm uses the L-BFGS iterative optimizer to calculate the loss. After iterative convergence, the corresponding pose and expression parameters are extracted and concatenated into a precise 109-dimensional ground truth tensor for the target. ;in, This is the transpose of the facial expression parameters of the FLAME model. This is the transpose of the mandibular pose parameters of the FLAME model. This is the transpose of the eye pose parameters in the FLAME model. The concatenated complete target ground truth tensor is used as the pseudo-ground truth for training the mapping network (i.e., forming a 109-dimensional label vector).

[0066] In this embodiment, the execution order of steps S201 and S202 is not specifically limited. For example, steps S201 and S202 can be executed in parallel.

[0067] In some embodiments, the M-dimensional facial driving parameters include X-dimensional expression parameters, Y-dimensional jaw pose parameters, and Z-dimensional eye pose parameters, where X+Y+Z=M; the plurality of branch neural networks include: An expression branch neural network is used to predict X-dimensional expression parameters based on N-dimensional facial deformation features. A mandibular branch neural network is used to predict Y-dimensional mandibular posture parameters based on N-dimensional facial deformation features. An eye branch neural network is used to predict Z-dimensional eye pose parameters based on N-dimensional facial deformation features.

[0068] For example, X is greater than Z, and Z is greater than Y.

[0069] This embodiment uses three structurally independent branch neural networks to independently predict three different facial attributes: expression, jaw pose, and eye pose. This decouples the parameter mapping process of each facial attribute from each other in structure, effectively avoiding feature crosstalk that may be caused by a single hybrid network.

[0070] In some embodiments, the facial expression branch neural network, the mandibular branch neural network, and the ocular branch neural network are all multilayer perceptrons, and each branch neural network includes a hidden layer, a batch normalization layer, and a ReLU activation layer; wherein, the feature dimensions of the hidden layers used for dimensionality enhancement in the mandibular branch neural network and the ocular branch neural network are smaller than the feature dimensions of the hidden layers used for dimensionality enhancement in the facial expression branch neural network.

[0071] This embodiment, by configuring a larger up-dimensional hidden layer for the facial expression branch neural network, can fully capture the high-dimensional nonlinear features of complex facial expressions; while the mandibular branch neural network and the eye branch neural network use a relatively small up-dimensional dimension, which reduces parameter redundancy while ensuring the accuracy of pose estimation and improves the overall computational efficiency of the network.

[0072] In some embodiments, the facial expression branch neural network includes a first hidden layer, a first batch of normalized layers, a first ReLU (Rectified Linear Unit) activation layer, a second hidden layer, a second batch of normalized layers, a second ReLU activation layer, and a first output layer connected in sequence; wherein, the first hidden layer receives the N-dimensional facial deformation features and increases the input dimension to 256 dimensions, the second hidden layer outputs 128 dimensions, and the first output layer outputs 100-dimensional facial expression parameters.

[0073] Specifically, the first batch normalization layer performs batch normalization on the 256-dimensional features; the first ReLU activation layer performs nonlinear activation on the normalized features; the second batch normalization layer performs batch normalization on the 128-dimensional features; and the second ReLU activation layer performs nonlinear activation on the normalized features.

[0074] In this embodiment, the expression branch neural network, through its structure design of first increasing the dimensionality and then decreasing it, can fully extract high-dimensional nonlinear information related to facial deformation features. In this way, the expression branch neural network can have a strong nonlinear mapping capability, which helps to meet the high-dimensional expression needs of complex facial expressions.

[0075] In some embodiments, the mandibular branch neural network includes a third hidden layer, a third batch normalization layer, a third ReLU activation layer, a fourth hidden layer, a fourth batch normalization layer, a fourth ReLU activation layer, and a second output layer connected in sequence; wherein, the third hidden layer receives N-dimensional facial deformation features and increases the input dimension to 128-dimensional features, the fourth hidden layer outputs 64-dimensional features, and the second output layer outputs 3-dimensional mandibular posture parameters.

[0076] The third batch normalization layer performs batch normalization on the 128-dimensional features; the third ReLU activation layer performs nonlinear activation on the normalized features; the fourth batch normalization layer performs batch normalization on the 64-dimensional features; and the fourth ReLU activation layer performs nonlinear activation on the normalized features.

[0077] In this embodiment, the dimensionality of the mandibular branch neural network is lower than that of the facial expression branch neural network, making the structure of the mandibular branch neural network lighter and the number of parameters relatively small.

[0078] In some embodiments, the eye branch neural network includes a fifth hidden layer, a fifth batch normalization layer, a fifth ReLU activation layer, a sixth hidden layer, a sixth batch normalization layer, a sixth ReLU activation layer, and a third output layer connected in sequence; wherein, the fifth hidden layer receives N-dimensional facial deformation features and increases the input dimension to 128-dimensional features, the sixth hidden layer outputs 64-dimensional features, and the third output layer outputs 6-dimensional eye pose parameters.

[0079] The fifth batch normalization layer is used to perform batch normalization on the 128-dimensional features; the fifth ReLU activation layer is used to perform nonlinear activation on the normalized features; the sixth batch normalization layer is used to perform batch normalization on the 64-dimensional features; and the sixth ReLU activation layer is used to perform nonlinear activation on the normalized features.

[0080] In this embodiment, the dimensionality of the eye branch neural network is lower than that of the facial expression branch neural network, making the structure of the eye branch neural network more concise and the number of parameters relatively small.

[0081] In some embodiments, the batch normalization in each branch of the neural network described above can be replaced with layer normalization, instance normalization, or group normalization. These normalization methods behave consistently during single-frame inference and are suitable for different deployment scenarios. Furthermore, while the ReLU activation function is used as an example in the above embodiments, it can also be replaced with other activation functions such as LeakyReLU, Gaussian Error Linear Unit (GELU), Si-shaped Linear Unit (SiLU), or Exponential Linear Unit (ELU).

[0082] In some embodiments, for parallel, parameter-free facial expression branch neural networks, mandibular branch neural networks, and eyeball branch neural networks, the first branch neural network in each branch neural network... There are 1 hidden layer, and its forward propagation expression is: ;in, For the first The weight matrix of each hidden layer For the first The bias vectors of each hidden layer For batch normalization operations, To modify the activation function of the linear unit; Indicates the first The output feature vectors of each hidden layer; Indicates the first The output feature vectors of each hidden layer.

[0083] For example, the output layer of the expression branch neural network uses a linear mapping without activation functions (no activation functions to allow negative values). Its output layer forward propagation expression is: ;in, The facial expression parameters are output by the facial expression branch neural network. This is the weight matrix of the output layer of the facial expression branch neural network. This is the bias vector of the output layer of the facial expression branch neural network. The output feature vector of the second hidden layer of the facial expression branch neural network is used as the input to its output layer.

[0084] For example, the forward propagation expression of the output layer of the mandibular branch neural network is as follows: ;in, The facial expression parameters output by the mandibular branch neural network. This is the weight matrix of the output layer of the mandibular branch neural network. This is the bias vector for the output layer of the mandibular branch neural network. The output feature vector of the second hidden layer of the mandibular branch neural network is used as the input to its output layer.

[0085] For example, the forward propagation expression of the output layer of the eye branch neural network is: ;in, These are the facial expression parameters output by the eye branch neural network. This is the weight matrix of the output layer of the eye branch neural network. This is the bias vector for the output layer of the eye branch neural network. The output feature vector of the second hidden layer of the eye branch neural network is used as the input to its output layer.

[0086] In this embodiment, the global total error is not calculated during the training phase; instead, the mean squared error (MSE Loss) is calculated independently for each of the three networks. The loss functions of the three branch neural networks are calculated based on their respective output driving parameter components and corresponding parameter component labels, and the backpropagation gradients of each branch neural network are isolated from each other.

[0087] Loss of the expression branch neural network The expression is: ;in, This represents the predicted value of the facial expression parameters output by the facial expression branch neural network for the i-th training sample (i.e., the driving parameter components predicted by the facial expression branch neural network). This represents the label of the facial expression parameter component corresponding to the i-th training sample.

[0088] Loss of the mandibular branch neural network The expression is: ;in, This represents the predicted mandibular posture parameters output by the mandibular branch neural network for the i-th training sample (i.e., the driving parameter components predicted by the mandibular branch neural network). This represents the label of the mandibular pose parameter component corresponding to the i-th training sample.

[0089] Loss of the eye branch neural network The expression is: ;in, This represents the predicted eye pose parameters output by the eye branch neural network for the i-th training sample (i.e., the driving parameter components predicted by the eye branch neural network). This represents the label of the eye pose parameter component corresponding to the i-th training sample.

[0090] Among the loss functions above, This indicates the number of samples in each training batch. This represents the square of the L2 norm, which is the sum of the squares of the components of the difference between the predicted and labeled values.

[0091] In this embodiment, the backpropagation gradients of each branch neural network are isolated from each other. Specifically, the weight update gradient of each branch neural network depends only on the partial derivative of its own loss function with respect to its own parameters, and is independent of the loss functions of other branch neural networks. For example, the weight update gradient of the mandibular branch neural network... That is, the weight parameters of the mandibular branch neural network. The gradient is equal to the loss of the mandibular branch neural network. right The partial derivative of the weight, which updates the gradient, depends on... , but not including or Error signal.

[0092] By isolating gradients, the gradient calculation paths of each branch of the neural network are independent during backpropagation, and error signals between different facial attributes will not cross-propagate, thereby suppressing feature crosstalk that may occur when a single hybrid network is used for processing.

[0093] Figure 3 This is a flowchart illustrating the face driving parameter mapping method provided in an embodiment of this application. Figure 3 As shown, the method includes steps S301 to S304.

[0094] S301: Obtain the target face image.

[0095] In some examples, the target face image can be an RGB image captured in real time by an image acquisition device, such as a single frame face image captured by a monocular RGB camera, or it can be a frame image extracted from a video stream, or it can be a pre-stored face photo or a face photo uploaded by the user. This embodiment does not specifically limit this.

[0096] S302: Extract N-dimensional facial deformation features from the target face image, where each dimension of the N-dimensional facial deformation features corresponds to a facial blending shape coefficient.

[0097] N-dimensional facial deformation features may include hybrid shape coefficients extracted from the target face image. Each dimension of the hybrid shape coefficient can characterize the activation intensity of a specific facial muscle group or muscle combination, such as raising the left eyebrow or stretching the right corner of the mouth.

[0098] The value of N can be 52, which represents a 52-dimensional facial blending shape coefficient. In other possible implementations, N can also take other values, such as 40 or 50, depending on the number of blending shape bases used.

[0099] In some examples, step S302 above, extracting N-dimensional facial deformation features from the target face image, includes: A facial blending shape model is used to extract blending shape coefficients from the face image, each dimension of which is used to characterize the activation intensity of a facial muscle group.

[0100] In some examples, the extraction method of N-dimensional facial deformation features can refer to the optional implementation of step S201 in the aforementioned embodiments, which will not be repeated here.

[0101] S303: Input N-dimensional facial deformation features into a mapping network, where multiple parallel branch neural networks predict driving parameter components used to describe different facial attributes.

[0102] The mapping network is obtained through supervised training based on a training sample set. Each training sample in the training sample set includes N-dimensional facial deformation features and M-dimensional facial driving parameters extracted from the sample face image. The M-dimensional facial driving parameters are used as sample labels.

[0103] In some examples, the training process of the mapping network can be described with reference to the exemplary description of the mapping network training method provided in the foregoing embodiments, and will not be repeated here.

[0104] S304: The driving parameter components predicted by multiple branch neural networks are concatenated according to their dimensions to obtain M-dimensional facial driving parameters, where M and N are both positive integers, and M is greater than N. The M-dimensional facial driving parameters are used to drive the facial deformation of the three-dimensional virtual object.

[0105] M-dimensional facial driving parameters can be directly used to drive any 3D virtual object that supports parametric control. The 3D virtual object can be a virtual digital human built based on statistical parametric face models such as FLAME or SMPL-X, or it can be a custom virtual character.

[0106] In some examples, driving facial deformation of a 3D virtual object based on M-dimensional facial driving parameters may include: The M-dimensional parameters are input into the facial controller of the target virtual object, and its facial mesh or point cloud is updated in real time to generate corresponding expressions and pose changes.

[0107] When driving digital humans based on 3D Gaussian sputtering (3DGS), the M-dimensional facial driving parameters can control the position, transparency, color and other attributes of the Gaussian ellipsoid to achieve high-fidelity facial deformation.

[0108] For example, during the inference phase, the application processes the video stream captured by the camera to extract the 52-dimensional facial blending shape coefficients of the user's face, and inputs them into the trained mapping network. Through a single forward propagation, each branch of the neural network independently predicts its own driving parameter components, which are then concatenated along the feature dimensions to form the final 109-dimensional driving parameters. : ; in, This is the transpose of the facial expression parameters predicted by the facial expression branch neural network. This is a transpose of the mandibular posture parameters predicted by the mandibular branch neural network. This is the transpose of the eye pose parameters predicted by the eye branch neural network.

[0109] These 109-dimensional driving parameters can be directly fed into the 3DGS rendering engine or the Mesh rendering engine to drive facial deformation of 3D virtual objects. Since the driving process only involves mapping in the parameter space and does not rely on the topological calculations of the underlying 3D vertices, it can be adapted to various digital human assets.

[0110] This application provides a facial driving parameter mapping method. By inputting N-dimensional facial deformation features into a mapping network containing multiple branch neural networks, different branch neural networks predict driving parameter components corresponding to different facial attributes. This ensures that the parameter mapping process for each facial attribute is structurally independent, avoiding feature crosstalk between different facial attributes when using a single network. Furthermore, by concatenating the driving parameter components predicted by each branch neural network according to their dimensions to form M-dimensional facial driving parameters (M>N), a mapping from low-dimensional facial blending shape coefficients to high-dimensional 3D parametric face model driving parameters is achieved, adaptable to various virtual objects and rendering pipelines.

[0111] In some embodiments, the M-dimensional facial driving parameters include X-dimensional expression parameters, Y-dimensional jaw pose parameters, and Z-dimensional eye pose parameters, where X+Y+Z=M; the multiple branch neural networks include: An expression branch neural network is used to predict X-dimensional expression parameters based on N-dimensional facial deformation features. A mandibular branch neural network is used to predict Y-dimensional mandibular posture parameters based on N-dimensional facial deformation features. An eye branch neural network is used to predict Z-dimensional eye pose parameters based on N-dimensional facial deformation features.

[0112] In some embodiments, the facial expression branch neural network, the mandibular branch neural network, and the ocular branch neural network are all multilayer perceptrons, and each branch neural network includes a hidden layer, a batch normalization layer, and a ReLU activation layer; wherein, the feature dimensions of the hidden layers used for dimensionality enhancement in the mandibular branch neural network and the ocular branch neural network are smaller than the feature dimensions of the hidden layers used for dimensionality enhancement in the facial expression branch neural network.

[0113] In some embodiments, the facial expression branch neural network includes a first hidden layer, a first batch of normalized layers, a first ReLU activation layer, a second hidden layer, a second batch of normalized layers, a second ReLU activation layer, and a first output layer connected in sequence; wherein, the first hidden layer receives N-dimensional facial deformation features and increases the input dimension to 256 dimensions, the second hidden layer outputs 128 dimensions, and the first output layer outputs 100-dimensional facial expression parameters. And / or, the mandibular branch neural network includes a third hidden layer, a third batch normalization layer, a third ReLU activation layer, a fourth hidden layer, a fourth batch normalization layer, a fourth ReLU activation layer, and a second output layer connected in sequence; wherein, the third hidden layer receives N-dimensional facial deformation features and increases the input dimension to 128-dimensional features, the fourth hidden layer outputs 64-dimensional features, and the second output layer outputs 3-dimensional mandibular posture parameters. And / or, the eye branch neural network includes a fifth hidden layer, a fifth batch normalization layer, a fifth ReLU activation layer, a sixth hidden layer, a sixth batch normalization layer, a sixth ReLU activation layer, and a third output layer connected in sequence; wherein, the fifth hidden layer receives N-dimensional facial deformation features and increases the input dimension to 128-dimensional features, the sixth hidden layer outputs 64-dimensional features, and the third output layer outputs 6-dimensional eye pose parameters.

[0114] The various embodiments or implementation methods described in this specification are presented in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other.

[0115] It is worth noting that, without contradiction, different embodiments can be arbitrarily combined. For example, some or all of the steps of different embodiments can be arbitrarily combined, and one embodiment can be arbitrarily combined with the optional implementations of other embodiments.

[0116] Next, the technical solutions provided in the embodiments of this application will be further explained with reference to specific examples.

[0117] Figure 4 This is a flowchart illustrating the mapping method from BS parameters to FLAME parameters based on a feature-decoupled cross-domain distillation neural network provided in this application embodiment. The data flow is a non-linear mapping process from the input space to the output space: the face image is processed by facial blending shape coefficient extraction to obtain a 52-dimensional BS feature vector; this vector is input in parallel to three structurally independent branch neural networks, which are mapped to 100-dimensional expression parameters, 3-dimensional jaw pose parameters, and 6-dimensional eye pose parameters, respectively; the three sets of parameters are concatenated along the feature dimensions to form 109-dimensional FLAME parameters, which are output to the rendering engine to drive the facial deformation of the 3D virtual object.

[0118] like Figure 4 As shown, the mapping method includes the following steps: Step S1: Dual-path parsing of monocular video input and cross-domain distillation.

[0119] Obtain a real-life RGB monocular video of a person containing continuous and rich facial expressions. Decompose the video into an image sequence frame by frame and simultaneously send it into two independent parsing paths, path A and path B, under the constraint of timestamp alignment.

[0120] Path A: On each frame of the RGB image, call MediaPipe to capture and parse 478 points to obtain 52-dimensional BS parameters.

[0121] The process consists of two stages: facial landmark regression and BS weighting.

[0122] Path B: On each frame of the RGB image, the VHAP algorithm is called to perform L-BFGS iterative optimization to obtain 109-dimensional FLAME parameters as the target ground truth tensor.

[0123] Step S2: Time alignment and data preprocessing to generate a cross-domain distillation training set.

[0124] Data preprocessing includes iterating through and removing invalid frames that are not strictly paired. Each training sample in the cross-domain distillation training set consists of 52-dimensional BS parameters extracted from the same frame of face image and 109-dimensional FLAME parameters used as sample labels.

[0125] Step S3: Input the cross-domain distillation training set into three branch neural networks based on facial attribute decoupling. Each branch neural network is trained under the supervision of an independent optimizer, and the backpropagation gradients are isolated from each other. The three branch neural networks include: expression network, jaw network, and eyeball network.

[0126] Three parallel, parameter-free branch neural networks, each employing a multilayer perceptron (MLP), constitute the mapping network described in the previous embodiment. During training, the global total error is not calculated; instead, the mean squared error is calculated independently for each of the three networks. Three independent Adam optimizers are instantiated. During backpropagation, the differentiation process is strictly physically isolated. This gradient isolation mechanism addresses the feature cross-bonding phenomenon.

[0127] Step S4: Real-time inference driven by high-fidelity digital human.

[0128] The application obtains a 52-dimensional BS weight vector of the user's face by processing the video sequence captured by the camera. This vector is then input into each branch network. After one forward propagation, the branch networks output three sets of predicted values, which are then concatenated along the feature dimensions to form the final 109-dimensional driving parameters.

[0129] The 109-dimensional driving parameters are fed into the 3DGS rendering engine or the Mesh rendering engine. Since the driving process only involves parameter space mapping and avoids the topological calculations of the underlying 3D vertex coordinates, it can seamlessly adapt to and drive various digital human assets.

[0130] It is understood that in the above embodiments, the M-dimensional driving parameters are illustrated using the 109-dimensional FLAME parameters extracted by the VHAP algorithm as an example. In other possible implementations, the VHAP algorithm can be replaced by monocular image-based 3D face reconstruction methods such as DECA (Detailed Expression Capture and Animation) and EMOOCA (Emotion Capture and Animation). The input dimension is also not limited to 52 dimensions; for example, it can be extended to 63-dimensional hybrid shape coefficients or any custom facial hybrid shape system, requiring only corresponding adjustment of the input layer dimension of the mapping network.

[0131] Furthermore, in the above embodiments, FLAME is used as an example for illustrative purposes. In other possible implementations, the facial portion of SMPL-X (Skinned Multi-Person Linear model with eXpressive features), the FaceVerse 3D face model, the FaceScape 3D face model, or a self-developed parametric face model can also be used, as long as it has a differentiable mesh generation function.

[0132] In the specific training, a cross-domain distillation training set was constructed, containing 15 acquisition sequences and a total of 31,276 frames, covering a wide distribution from natural speech to extreme facial expressions. Path A outputs a 52-dimensional Blendshape JSON per frame, and path B outputs a 109-dimensional FLAME JSON per frame.

[0133] Regarding network specifications: the facial expression branch structure is 52→256→128→100, approximately 152K parameters; the mandibular branch structure is 52→128→64→3, approximately 38K parameters; and the eyeball branch structure is 52→128→64→6, approximately 38K parameters. Each layer contains BatchNorm1d (Batch Normalization, or BN) and ReLU, with no activation in the output layer. Training hyperparameters were set as follows: batch size 128, training epochs 145, learning rate 0.001, Adam optimizer (β1=0.9, β2=0.999), mean squared error (MSE) loss, Kaiming initialization (a weight initialization method for the ReLU activation function), and the training process was completed on a single graphics processor (e.g., a GPU with 24GB of VRAM), resulting in a relatively short total training time for all 145 epochs.

[0134] Figure 5a Two sets of rendering comparison diagrams of the mandibular closed state provided in the embodiments of this application. Figure 5bTwo sets of rendering comparison diagrams of the mandible open state provided in the embodiments of this application. Figure 5a and Figure 5b The left side of the image represents the rendered ground truth, while the right side represents the rendered predicted values. As can be seen from the comparison, this application can accurately reproduce facial deformation in both the closed and open jaw states. The predicted results are visually almost identical to the ground truth, and no unnatural deformations of other facial areas (such as eyelids and eyebrows) occur when the jaw is wide open.

[0135] Figure 5c The training loss convergence curves for the three independent branch networks provided in this embodiment are shown. The horizontal axis represents the number of training epochs, totaling 145 epochs; the vertical axis represents the loss value, using a logarithmic scale. Figure 5c From left to right, the training loss convergence curves of the facial expression branch network, the mandibular branch network, and the eyeball branch network are shown in sequence. After 145 rounds of training, the loss of each of the three branch networks converged smoothly without overfitting, indicating that the training process of each branch network was stable.

[0136] Figure 5d This is a schematic diagram of the branch network test results provided in an embodiment of this application. Figure 5d The left side shows the frame-by-frame prediction MAE error of the facial expression and chin branches under time-series testing, while the right side shows the coefficient of determination quantification of each branch and the overall prediction accuracy of each branch.

[0137] Test results show that the MSE of the facial expression branch network is approximately 0.0016 and the MAE (mean absolute error) is approximately 0.024; the MSE of the mandibular branch network is approximately 0.0010 and the MAE is approximately 0.017; and the MSE of the ocular branch network is approximately 0.0006 and the MAE is approximately 0.012.

[0138] It is understandable that BS parameters and FLAME parameters differ fundamentally in space. BS parameters are linear superposition models, with each dimension independently controlling the local expression base, ranging from [0,1], and lacking the ability to express three-dimensional pose and eye movement. FLAME parameters, on the other hand, are nonlinear statistical models, with expression parameters encoded by PCA (Principal Component Analysis), not directly corresponding to specific muscle groups. Jaw and eye parameters are represented using Euler angles or rotation vectors, and mesh generation requires nonlinear deformation mixing and linear skinning, resulting in a highly nonlinear overall structure. Therefore, the mapping from BS parameters to FLAME parameters is essentially a cross-space mapping from a linear interpretable semantic space to a nonlinear statistical latent space, requiring simultaneous numerical correspondence learning, missing information completion (eye movement and pose), and complex decoupling of the PCA dimension. This application's embodiment solves the cross-domain mapping problem from BS parameters to FLAME parameters through a joint scheme of dual-path cross-domain distillation data construction, multi-branch network design based on physiological feature decoupling, and independent gradient optimization.

[0139] In summary, the technical solution provided by the embodiments of this application has at least the following advantages: (1) The mapping is completely confined to the latent space of the parameters and is defined as a pure function mapping. The input relies on only 52 scalar BS coefficients, and the output consists of only 109 scalar FLAME parameters. It does not involve vertex coordinate calculations throughout the process, and can be directly driven as long as the underlying layer uses the FLAME parameterization protocol, regardless of changes in the target mesh topology. Furthermore, the 3DGS pipeline also accepts FLAME parameters and is compatible with both traditional Mesh and cutting-edge 3DGS pipelines, enabling cross-pipeline reuse.

[0140] (2) Replacing online optimization with "offline training + online inference" allows for the use of VHAP to annotate a large number of video frames at once as training data, and the offline training time is not included in the inference time. Online inference degenerates into pure forward propagation, that is, the input of three lightweight MLPs of 52 dimensions each performs matrix multiplication once and outputs a 109-dimensional single frame, which greatly improves the speed and realizes high-concurrency real-time driving. Consumer-grade hardware can drive multiple digital humans at the same time.

[0141] (3) The three physiological facial attributes correspond one-to-one with three physically isolated MLP branches. Each branch weight only receives its own loss gradient signal, realizing gradient physical isolation and feature decoupling. Furthermore, the three independent optimizers each maintain three types of features in the implicit representation space of the adaptive learning rate network: facial expression, jaw, and eyeballs. This avoids feature crosstalk between different attributes and enables high-fidelity, natural, and distortion-free facial animation.

[0142] Figure 6This is a schematic diagram of a mapping network training device provided in an embodiment of this application. The mapping network is used to generate facial driving parameters, and the mapping network includes multiple parallel branch neural networks. Figure 6 As shown, the mapping network training device 100 includes: The first acquisition module 101 is used to acquire an image sequence consisting of multiple frames of face images from the sample video; The first extraction module 102 is used to extract N-dimensional facial deformation features and M-dimensional facial driving parameters for each frame of face image in the image sequence. The sample construction module 103 is used to pair the M-dimensional facial driving parameters as sample labels with the N-dimensional facial deformation features to form a training sample, so as to construct a training sample set, wherein M and N are both positive integers, and M is greater than N; Supervised training module 104 is used for: For each training sample, the N-dimensional facial deformation features are input into the mapping network, and the M-dimensional facial driving parameters, which serve as sample labels, are split according to facial attributes to obtain the corresponding parameter component labels. For each branch neural network in the mapping network, the corresponding loss is calculated based on the driving parameter components predicted by the branch neural network and the corresponding parameter component labels. The parameters of the branch neural network are updated based on the loss using an independent optimizer to obtain the trained mapping network. The gradients of each branch neural network are isolated from each other during backpropagation.

[0143] In some embodiments, the M-dimensional facial driving parameters include X-dimensional expression parameters, Y-dimensional jaw posture parameters, and Z-dimensional eye posture parameters, where X+Y+Z=M; the plurality of branch neural networks include: An expression branch neural network is used to predict the X-dimensional expression parameters based on the N-dimensional facial deformation features. A mandibular branch neural network is used to predict the Y-dimensional mandibular posture parameters based on the N-dimensional facial deformation features. An eye-branching neural network is used to predict the Z-dimensional eye pose parameters based on the N-dimensional facial deformation features.

[0144] In some embodiments, the facial expression branch neural network, the mandibular branch neural network, and the eyeball branch neural network are all multilayer perceptrons, and each branch neural network includes a hidden layer, a batch normalization layer, and a ReLU activation layer. The feature dimensions of the hidden layers used for dimensionality enhancement in both the mandibular branch neural network and the eye branch neural network are smaller than the feature dimensions of the hidden layers used for dimensionality enhancement in the facial expression branch neural network.

[0145] In some embodiments, the facial expression branch neural network includes a first hidden layer, a first batch of normalized layers, a first ReLU activation layer, a second hidden layer, a second batch of normalized layers, a second ReLU activation layer, and a first output layer connected in sequence; wherein, the first hidden layer receives the N-dimensional facial deformation features and increases the input dimension to 256 dimensions, the second hidden layer outputs 128 dimensions, and the first output layer outputs 100-dimensional facial expression parameters. And / or, the mandibular branch neural network includes a third hidden layer, a third batch normalization layer, a third ReLU activation layer, a fourth hidden layer, a fourth batch normalization layer, a fourth ReLU activation layer, and a second output layer connected in sequence; wherein, the third hidden layer receives the N-dimensional facial deformation features and increases the input dimension to 128-dimensional features, the fourth hidden layer outputs 64-dimensional features, and the second output layer outputs 3-dimensional mandibular posture parameters; And / or, the eye branch neural network includes a fifth hidden layer, a fifth batch normalization layer, a fifth ReLU activation layer, a sixth hidden layer, a sixth batch normalization layer, a sixth ReLU activation layer, and a third output layer connected in sequence; wherein, the fifth hidden layer receives the N-dimensional facial deformation features and increases the input dimension to 128-dimensional features, the sixth hidden layer outputs 64-dimensional features, and the third output layer outputs 6-dimensional eye pose parameters.

[0146] In some embodiments, the first extraction module 102 is used to: A facial blending shape model is used to extract blending shape coefficients from the face image, each dimension of which is used to characterize the activation intensity of a facial muscle group; Facial driving parameters are extracted from the face image using a three-dimensional deformable face model. These facial driving parameters include expression parameters, jaw pose parameters, and eye pose parameters.

[0147] Figure 7 This is a schematic diagram of the structure of a face driving parameter mapping device provided in an embodiment of this application, as shown below. Figure 7 As shown, the facial driving parameter mapping device 200 includes: The second acquisition module 201 is used to acquire the target face image; The second extraction module 202 is used to extract N-dimensional facial deformation features from the target face image, where each dimension of the N-dimensional facial deformation features corresponds to a facial mixing shape coefficient. The parameter prediction module 203 is used to predict the driving parameter components used to describe different facial attributes by the N-dimensional facial deformation feature mapping network, which is composed of multiple parallel branch neural networks in the mapping network. The parameter splicing module 204 is used to splice the driving parameter components predicted by the multiple branch neural networks according to the dimensions to obtain M-dimensional facial driving parameters, where M and N are both positive integers and M is greater than N. The M-dimensional facial driving parameters are used to drive the facial deformation of the three-dimensional virtual object. The mapping network is obtained through supervised training based on a training sample set; each training sample in the training sample set includes N-dimensional facial deformation features extracted from the sample face image and M-dimensional facial driving parameters as sample labels.

[0148] In some embodiments, the apparatus further includes a training module, the training module being used for: Obtain an image sequence consisting of multiple frames of face images from the sample video; For each frame of a face image in the image sequence, N-dimensional facial deformation features and M-dimensional facial driving parameters are extracted. The extracted M-dimensional facial driving parameters are used as sample labels and paired with the N-dimensional facial deformation features of the face image to form a training sample, so as to construct a training sample set. For each training sample, the N-dimensional facial deformation features are input into the mapping network, and the M-dimensional facial driving parameters, which serve as sample labels, are split according to facial attributes to obtain the corresponding parameter component labels. For each branch neural network in the mapping network, the corresponding loss is calculated based on the driving parameter components predicted by the branch neural network and the corresponding parameter component labels. The parameters of the branch neural network are updated based on the loss using an independent optimizer to obtain the trained mapping network. The gradients of each branch neural network are isolated from each other during backpropagation.

[0149] In some embodiments, the M-dimensional facial driving parameters include X-dimensional expression parameters, Y-dimensional jaw posture parameters, and Z-dimensional eye posture parameters, where X+Y+Z=M; the plurality of branch neural networks include: An expression branch neural network is used to predict the X-dimensional expression parameters based on the N-dimensional facial deformation features. A mandibular branch neural network is used to predict the Y-dimensional mandibular posture parameters based on the N-dimensional facial deformation features. An eye-branching neural network is used to predict the Z-dimensional eye pose parameters based on the N-dimensional facial deformation features.

[0150] In some embodiments, the facial expression branch neural network, the mandibular branch neural network, and the ocular branch neural network are all multilayer perceptrons, and each branch neural network includes a hidden layer, a batch normalization layer, and a ReLU activation layer; wherein, the feature dimensions of the hidden layers used for dimensionality enhancement in the mandibular branch neural network and the ocular branch neural network are smaller than the feature dimensions of the hidden layers used for dimensionality enhancement in the facial expression branch neural network.

[0151] In some embodiments, the facial expression branch neural network includes a first hidden layer, a first batch of normalized layers, a first ReLU activation layer, a second hidden layer, a second batch of normalized layers, a second ReLU activation layer, and a first output layer connected in sequence; wherein, the first hidden layer receives the N-dimensional facial deformation features and increases the input dimension to 256 dimensions, the second hidden layer outputs 128 dimensions, and the first output layer outputs 100-dimensional facial expression parameters. And / or, the mandibular branch neural network includes a third hidden layer, a third batch normalization layer, a third ReLU activation layer, a fourth hidden layer, a fourth batch normalization layer, a fourth ReLU activation layer, and a second output layer connected in sequence; wherein, the third hidden layer receives the N-dimensional facial deformation features and increases the input dimension to 128-dimensional features, the fourth hidden layer outputs 64-dimensional features, and the second output layer outputs 3-dimensional mandibular posture parameters; And / or, the eye branch neural network includes a fifth hidden layer, a fifth batch normalization layer, a fifth ReLU activation layer, a sixth hidden layer, a sixth batch normalization layer, a sixth ReLU activation layer, and a third output layer connected in sequence; wherein, the fifth hidden layer receives the N-dimensional facial deformation features and increases the input dimension to 128-dimensional features, the sixth hidden layer outputs 64-dimensional features, and the third output layer outputs 6-dimensional eye pose parameters.

[0152] It should be understood that each module in the mapping network training device and each module in the face driving parameter mapping device can be implemented in the form of processor calling software, or each module can be implemented in the form of hardware circuit. The functions of some or all modules can be realized through the design of the hardware circuit, which can be understood as one or more processors.

[0153] This application also provides an electronic device, including a processor, which is configured to invoke instructions to cause the electronic device to execute a mapping network training method or a face driving parameter mapping method provided in any of the foregoing embodiments.

[0154] This application also provides a computer-readable storage medium storing an executable program thereon. When the executable program is executed by a processor, it implements the mapping network training method or the face driving parameter mapping method provided in any of the foregoing embodiments.

[0155] For ease of understanding, the following focuses on explaining the terminology used in this embodiment: In this application embodiment, the processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction read and execute capabilities, such as a Central Processing Unit (CPU), a microprocessor, a Graphics Processing Unit (GPU) (which can be understood as a type of microprocessor), or a Digital Signal Processor (DSP). In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. The logical relationships of the aforementioned hardware circuits are fixed or reconfigurable. For example, the processor is a hardware circuit implemented using an Application-Specific Integrated Circuit (ASIC) or a Programmable Logic Device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units or modules. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a Neural Network Processing Unit (NPU), a Tensor Processing Unit (TPU), a Deep Learning Processing Unit (DPU), etc.

[0156] The computer-readable storage medium provided in this embodiment can execute the mapping network training method or the face driving parameter mapping method of the above embodiments. Its implementation principle and technical effect are similar to those of the above embodiments, and will not be repeated here.

[0157] The aforementioned computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0158] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in application-specific integrated circuits (ASICs). Alternatively, the processor and the readable storage medium can exist as discrete components in an electronic device or a host device.

[0159] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0160] The various embodiments or implementation methods described in this specification are presented in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other.

[0161] In the description of this specification, references to "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A method for training a mapping network, characterized in that, The mapping network is used to generate facial driving parameters, and the mapping network includes multiple parallel branch neural networks; the training method includes: Obtain an image sequence consisting of multiple frames of face images from the sample video; For each frame of face image in the image sequence, N-dimensional facial deformation features and M-dimensional facial driving parameters are extracted. The M-dimensional facial driving parameters are used as sample labels and paired with the N-dimensional facial deformation features to form a training sample to construct a training sample set. Here, M and N are both positive integers, and M is greater than N. For each training sample, the N-dimensional facial deformation features are input into the mapping network, and the M-dimensional facial driving parameters, which serve as sample labels, are split according to facial attributes to obtain the corresponding parameter component labels. For each branch neural network in the mapping network, the corresponding loss is calculated based on the driving parameter components predicted by the branch neural network and the corresponding parameter component labels. The parameters of the branch neural network are updated based on the loss using an independent optimizer to obtain the trained mapping network. The gradients of each branch neural network are isolated from each other during backpropagation.

2. The mapping network training method according to claim 1, characterized in that, The M-dimensional facial driving parameters include X-dimensional expression parameters, Y-dimensional jaw posture parameters, and Z-dimensional eyeball posture parameters, where X+Y+Z=M; The plurality of branched neural networks include: An expression branch neural network is used to predict the X-dimensional expression parameters based on the N-dimensional facial deformation features. A mandibular branch neural network is used to predict the Y-dimensional mandibular posture parameters based on the N-dimensional facial deformation features. An eye-branching neural network is used to predict the Z-dimensional eye pose parameters based on the N-dimensional facial deformation features.

3. The mapping network training method according to claim 2, characterized in that, The facial expression branch neural network, the mandibular branch neural network, and the eyeball branch neural network are all multilayer perceptrons, and each branch neural network includes a hidden layer, a batch normalization layer, and a ReLU activation layer. The feature dimensions of the hidden layers used for dimensionality enhancement in both the mandibular branch neural network and the eye branch neural network are smaller than the feature dimensions of the hidden layers used for dimensionality enhancement in the facial expression branch neural network.

4. The mapping network training method according to claim 3, characterized in that, The facial expression branch neural network includes a first hidden layer, a first batch of normalized layers, a first ReLU activation layer, a second hidden layer, a second batch of normalized layers, a second ReLU activation layer, and a first output layer connected in sequence; wherein, the first hidden layer receives the N-dimensional facial deformation features and increases the input dimension to 256 dimensions, the second hidden layer outputs 128 dimensions, and the first output layer outputs 100-dimensional facial expression parameters. And / or, the mandibular branch neural network includes a third hidden layer, a third batch normalization layer, a third ReLU activation layer, a fourth hidden layer, a fourth batch normalization layer, a fourth ReLU activation layer, and a second output layer connected in sequence; wherein, the third hidden layer receives the N-dimensional facial deformation features and increases the input dimension to 128-dimensional features, the fourth hidden layer outputs 64-dimensional features, and the second output layer outputs 3-dimensional mandibular posture parameters; And / or, the eye branch neural network includes a fifth hidden layer, a fifth batch normalization layer, a fifth ReLU activation layer, a sixth hidden layer, a sixth batch normalization layer, a sixth ReLU activation layer, and a third output layer connected in sequence; wherein, the fifth hidden layer receives the N-dimensional facial deformation features and increases the input dimension to 128-dimensional features, the sixth hidden layer outputs 64-dimensional features, and the third output layer outputs 6-dimensional eye pose parameters.

5. The mapping network training method according to any one of claims 1 to 4, characterized in that, For each frame of a face image in the image sequence, N-dimensional facial deformation features and M-dimensional facial driving parameters are extracted, including: A facial blending shape model is used to extract blending shape coefficients from the face image, each dimension of which is used to characterize the activation intensity of a facial muscle group; Facial driving parameters are extracted from the face image using a three-dimensional deformable face model. These facial driving parameters include expression parameters, jaw pose parameters, and eye pose parameters.

6. A facial driving parameter mapping method, characterized in that, The method includes: Acquire the target face image; N-dimensional facial deformation features are extracted from the target face image, and each dimension of the N-dimensional facial deformation features corresponds to a facial blending shape coefficient. The N-dimensional facial deformation features are input into a mapping network, and multiple parallel branch neural networks in the mapping network predict the driving parameter components used to describe different facial attributes. The driving parameter components predicted by the multiple branch neural networks are concatenated according to their dimensions to obtain M-dimensional facial driving parameters, where M and N are both positive integers and M is greater than N. The M-dimensional facial driving parameters are used to drive the facial deformation of the three-dimensional virtual object. The mapping network is obtained through supervised training based on a training sample set; each training sample in the training sample set includes N-dimensional facial deformation features extracted from the sample face image and M-dimensional facial driving parameters as sample labels.

7. The facial driving parameter mapping method according to claim 6, characterized in that, The training process of the mapping network includes: Obtain an image sequence consisting of multiple frames of face images from the sample video; For each frame of a face image in the image sequence, N-dimensional facial deformation features and M-dimensional facial driving parameters are extracted. The extracted M-dimensional facial driving parameters are used as sample labels and paired with the N-dimensional facial deformation features of the face image to form a training sample, so as to construct a training sample set. For each training sample, the N-dimensional facial deformation features are input into the mapping network, and the M-dimensional facial driving parameters, which serve as sample labels, are split according to facial attributes to obtain the corresponding parameter component labels. For each branch neural network in the mapping network, the corresponding loss is calculated based on the driving parameter components predicted by the branch neural network and the corresponding parameter component labels. The parameters of the branch neural network are updated based on the loss using an independent optimizer to obtain the trained mapping network. The gradients of each branch neural network are isolated from each other during backpropagation.

8. The facial driving parameter mapping method according to claim 6 or 7, characterized in that, The M-dimensional facial driving parameters include X-dimensional expression parameters, Y-dimensional jaw posture parameters, and Z-dimensional eyeball posture parameters, where X+Y+Z=M; The plurality of branched neural networks include: An expression branch neural network is used to predict the X-dimensional expression parameters based on the N-dimensional facial deformation features. A mandibular branch neural network is used to predict the Y-dimensional mandibular posture parameters based on the N-dimensional facial deformation features. An eye-branching neural network is used to predict the Z-dimensional eye pose parameters based on the N-dimensional facial deformation features.

9. A mapping network training device, characterized in that, The mapping network is used to generate facial driving parameters, and the mapping network includes multiple parallel branch neural networks. The device includes: The first acquisition module is used to acquire an image sequence consisting of multiple frames of face images from the sample video; The first extraction module is used to extract N-dimensional facial deformation features and M-dimensional facial driving parameters for each frame of face image in the image sequence. The sample construction module is used to pair the M-dimensional facial driving parameters as sample labels with the N-dimensional facial deformation features to form a training sample, so as to construct a training sample set, where M and N are both positive integers, and M is greater than N; Supervised training module, used for: For each training sample, the N-dimensional facial deformation features are input into the mapping network, and the M-dimensional facial driving parameters, which serve as sample labels, are split according to facial attributes to obtain the corresponding parameter component labels. For each branch neural network in the mapping network, the corresponding loss is calculated based on the driving parameter components predicted by the branch neural network and the corresponding parameter component labels. The parameters of the branch neural network are updated based on the loss using an independent optimizer to obtain the trained mapping network. The gradients of each branch neural network are isolated from each other during backpropagation.

10. A facial driving parameter mapping device, characterized in that, The device includes: The second acquisition module is used to acquire the target face image; The second extraction module is used to extract N-dimensional facial deformation features from the target face image, where each dimension of the N-dimensional facial deformation features corresponds to a facial blending shape coefficient. The parameter prediction module is used to map the N-dimensional facial deformation feature network, and the multiple branch neural networks in the mapping network predict the driving parameter components used to describe different facial attributes respectively. The parameter splicing module is used to splice the driving parameter components predicted by the multiple branch neural networks according to the dimensions to obtain M-dimensional facial driving parameters, where M and N are both positive integers and M is greater than N. The M-dimensional facial driving parameters are used to drive the facial deformation of the three-dimensional virtual object. The mapping network is obtained through supervised training based on a training sample set; each training sample in the training sample set includes N-dimensional facial deformation features extracted from the sample face image and M-dimensional facial driving parameters as sample labels.

11. An electronic device, characterized in that, Includes a processor, the processor being configured to invoke instructions to cause the electronic device to perform the mapping network training method as described in any one of claims 1 to 5, or the face driving parameter mapping method as described in any one of claims 6 to 8.