Virtual digital human expression recognition method and system based on multi-scale expansion convolution

By adopting innovative designs such as multi-scale expansion convolution and fusion attention mechanisms in expression recognition technology, the problem of insufficient recognition capabilities of the existing technology in complex scenarios is solved, and a more efficient and robust expression recognition effect is achieved.

CN120220213APending Publication Date: 2025-06-27YANTAI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510344433.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing expression recognition technology faces the problems of insufficient multi-scale feature perception ability, poor dynamic environment robustness, insufficient attention mechanism fusion and weak detailed feature retention ability in complex scenarios.

Method used

A neural network design based on multi-scale expansion convolution is adopted, combining the fusion attention mechanism, residual mask mechanism and coordinate attention mechanism to enhance the model's ability to extract and classify expression features.

Benefits of technology

The accuracy and robustness of expression recognition are improved, especially under low resolution input, which improves the recognition accuracy of "surprised" expressions, and the effect of virtual digital human expression driven is verified through the MEAD dataset.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220213A_ABST
    Figure CN120220213A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of expression recognition, in particular to a virtual digital human expression recognition method and system based on multi-scale expansion convolution. The method comprises the following steps: preprocessing acquired image data; constructing a neural network model based on multi-scale expansion convolution based on the preprocessed data, and training the neural network model by using the preprocessed image data; applying the trained model to image processing, and outputting an expression recognition result; and driving the virtual digital human to generate an expression animation according to an identification result. According to the method, efficient extraction and fusion of multi-level expression features are realized through a multi-scale expansion convolution fusion attention module (MDFA) and a coordinate attention collaboration mechanism, and through global average pooling and a 1 * 1 convolution dimension reduction strategy, the recognition accuracy of the surprising expression with low-resolution input is improved by 7.8%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of facial expression recognition, and in particular, to a virtual digital human facial expression recognition method and system based on multi-scale dilated convolution. Background Art

[0002] Facial expression recognition is an important research direction in the field of computer vision, aiming to recognize and analyze facial expressions through a computer to achieve an understanding of an individual's emotional state.

[0003] With the development of artificial intelligence and deep learning technologies, facial expression recognition methods based on convolutional neural networks (CNNs) have made remarkable progress and have gradually become a research hotspot. However, existing technologies still face many challenges in complex scenarios, mainly reflected in the following aspects: 1. Insufficient multi-scale feature perception ability: Traditional CNNs use convolutional kernels of fixed size, making it difficult to effectively capture the differential expressions of facial expression features at different spatial scales. Especially in complex scenarios, the correlation between subtle facial expression changes (such as micro-expressions) and macroscopic facial muscle movements is easily ignored, resulting in the loss of key features.

[0004] 2. Poor robustness in dynamic environments: Existing methods have insufficient adaptability to sudden changes in lighting, head pose offsets, and partial occlusions. The main reason is the lack of a directional enhancement mechanism for key regions during the feature extraction process, and redundant background information interferes with the model's judgment.

[0005] 3. Inadequate integration of attention mechanisms: Although some studies have introduced channel or spatial attention, they have failed to achieve the collaborative optimization of multi-dimensional features. For example, channel attention tends to ignore the spatial position correlation, while a single spatial attention is difficult to distinguish the subtle differences between similar facial expressions (such as anger and disgust), resulting in a decrease in feature discriminability.

[0006] 4. Weak ability to preserve detailed features: Traditional pooling operations and fixed-step convolutions are prone to the loss of high-frequency information, affecting the extraction efficiency of micro-facial expression features. Especially in low-resolution images, the degradation of key cues such as the shape of eyebrows and the curvature of lips will directly lead to misclassification.

[0007] Currently, there are still the following deficiencies in facial expression recognition of virtual digital humans: 1. Insufficient multi-scale feature perception ability: Due to the fixed-size convolutional kernels and limited receptive fields of traditional convolutional neural networks, it is difficult to jointly capture the local details (such as micro-expression lines) and global context (such as facial contours) of expression features, resulting in the lack of cross-scale feature correlation modeling; 2. Insufficient robustness in dynamic environments: Existing models have poor adaptability to changes in illumination, head pose offsets, and partial occlusions. The main reason is the lack of an adaptive enhancement mechanism for key expression regions, and redundant background information interferes with classification decisions; 3. Inadequate fusion of attention mechanisms: Existing single-dimensional attention (such as channel or spatial attention) fails to achieve the collaborative optimization of multi-modal features, resulting in the difficulty of distinguishing the subtle differences between similar expressions (such as anger and disgust), and the discriminability of features is reduced; 4. Weak ability to maintain detailed features: Traditional pooling operations and fixed-step convolutions cause the loss of high-frequency information. Especially in the case of low-resolution inputs, key expression cues such as the shape of eyebrows and the curvature of lips degenerate severely, affecting the recognition accuracy of micro-expressions. Summary of the Invention

[0008] To solve the above-mentioned problems, the present invention provides a virtual digital human expression recognition method and system based on multi-scale dilated convolutions. It aims to solve the deficiencies of the prior art in feature extraction, attention mechanism, robustness, and generalization ability through innovative network design and module optimization. By introducing multi-scale dilated convolutions and a fused attention mechanism, the present invention can simultaneously capture the detailed information and global context information in the image, enhancing the model's ability to recognize complex expressions.

[0009] In the first aspect, a virtual digital human expression recognition method based on multi-scale dilated convolutions provided by the present invention adopts the following technical solutions: A virtual digital human expression recognition method based on multi-scale dilated convolutions includes: Obtain image data; Preprocess the obtained image data; Construct a neural network model based on multi-scale dilated convolutions using the preprocessed data, and train the neural network model using the preprocessed image data; Apply the trained model to image processing and output the expression recognition result; Drive the virtual digital human to generate an expression animation according to the recognition result.

[0010] Further, the preprocessing of the obtained image data includes converting the image data for data diversity using a geometric-illumination joint enhancement matrix. Among them, the matrix generates diverse training samples through random parameter combinations to simulate complex scenarios such as side faces, occlusions, and light and shade changes to improve the recognition accuracy and generalization ability of the model. The enhancement matrix is expressed as: where is the random rotation angle, controls the scale scaling, , is the translation amount, is the brightness scaling factor, is the Gaussian noise perturbation term.

[0011] Furthermore, the neural network model based on multi-scale dilated convolution is constructed based on the preprocessed data, including using ResNet34 as the backbone structure, and embedding multi-scale dilated convolution fusion attention mechanism, residual mask mechanism and coordinate attention mechanism to construct the neural network model. Among them, the residual mask mechanism is embedded after each residual layer, and its structure is: where and represent the input and output feature maps respectively, σ is the Sigmoid function, represents element-wise multiplication, and a 3×3 convolutional layer is followed by batch normalization.

[0012] Furthermore, training the neural network model using the preprocessed image data includes using the Layer1 layer of the neural network model to extract features from the 224×224×3 normalized image data, outputting a 56×56×64 feature map, then performing feature enhancement through the residual mask mechanism, and after passing through the Layer2 layer, outputting a 28×28×128 feature map, superimposing the RMB mechanism and introducing a channel dropout strategy to prevent overfitting. The feature enhancement of the Layer1 layer is expressed as: where the output channels of the 3×3 convolution remain 64, σ is the Sigmoid function, and Sigmoid generates a spatial mask to enhance the response of the key expression regions.

[0013] Furthermore, training the neural network model using the preprocessed image data also includes inputting the data processed by the Layer2 layer into the Layer3 layer, and after passing through the Layer3 layer, outputting a 14×14×256 feature map, and then connecting to the coordinate attention mechanism. Among them, first perform bidirectional pooling to generate H / W direction feature descriptors, reduce the dimension to C / 8 channels through 1×1 convolution, separate and generate spatial attention weights after ReLU activation, and then use feature reweighting to enhance the spatial correlation of the expression key points of the eyebrows and the corners of the mouth. The H / W direction feature descriptor is expressed as: where the height is h and the width is w.

[0014] Further, when training the neural network model using the preprocessed image data, it further includes inputting the data processed by Layer 3 into Layer 4. After being processed by Layer 4, a 7×7×512 feature map is output, and then a multi-scale dilated convolution fusion attention mechanism is connected. Among them, parallel five-way processing is adopted, and the input data passes through a 1×1 standard convolution of Branch 1 in parallel, outputting 512 channels; 3×3 dilated convolutions of Branch 2 to 4, each outputting 512 channels; a global average pooling layer and a 1×1 convolution of Branch 5. Then, the five-way data features are concatenated and weighted by channel attention to achieve feature fusion of micro-expression lines and facial contours, which is expressed as: 。

[0015] Further, when applying the trained model to image processing and outputting an expression recognition result, it includes compressing the fused feature map by using adaptive average pooling with the trained model, and performing Softmax probability calculation. The corresponding expression with the largest value is the output expression, and the Softmax function converts the probability distribution and is expressed as: Output probability vector ,satisfying 。

[0016] Further, driving the virtual digital human to generate an expression animation according to the recognition result includes extracting the category index corresponding to the maximum probability according to the expression probability vector , parsing it into a semantic label according to a preset expression coding table, and retrieving a pre-constructed expression-parameter mapping library according to to obtain the corresponding facial action unit AU weight vector y and geometric deformation parameters ,so as to match the parameter library, where K is the number of AUs and N is the number of facial mesh vertices.

[0017] Further, driving the virtual digital human to generate an expression animation according to the recognition result further includes inputting the current frame expression parameters into a pre-trained LSTM network to predict the parameter change amount at the next time step, constructing a motion equation for the m-th facial key point for geometric-appearance joint driving, calculating through a deformation field and discretizing the driving equation into frame-level deformation to generate a real-time animation, and the driving equation is expressed as: In the formula, the appearance feature generation residual term ​; The global image features of the current frame ; The AU weight vector of the current frame .

[0018] In a second aspect, a virtual digital human expression recognition system based on multi-scale dilated convolution includes: A data acquisition module, configured to acquire image data; A preprocessing module, configured to preprocess the acquired image data; A model training module, configured to construct a neural network model based on multi-scale dilated convolution based on the preprocessed data, and train the neural network model using the preprocessed image data; A recognition module, configured to apply the trained model to image processing and output an expression recognition result; An animation module, configured to drive a virtual digital human to generate an expression animation according to the recognition result.

[0019] In summary, the present invention has the following beneficial technical effects: Through the multi-scale dilated convolution fusion attention module (MDFA) and the coordinate attention cooperation mechanism, the present invention realizes the efficient extraction and fusion of multi-level expression features. Through the global average pooling and 1×1 convolution dimensionality reduction strategy, the recognition accuracy of the "surprised" expression with low-resolution input is increased by 7.8%. Based on the MEAD dataset, the expression driving effect of the virtual digital human is verified, the matching degree of the key muscle movement trajectories reaches 89.7%, and the smoothness of the animation transition is improved by 3%-5%, effectively solving the problems of cross-style adaptation and rendering noise suppression. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a schematic diagram of a virtual digital human expression recognition method based on multi-scale dilated convolution according to Embodiment 1 of the present invention; Figure 2 is an architecture diagram of a multi-scale dilated convolution fusion attention residual network according to Embodiment 1 of the present invention; Figure 3 is a comparison diagram of gradient class activation functions of a convolutional neural network according to Embodiment 1 of the present invention; Figure 4 is a schematic diagram of the virtual digital human expression driving process according to Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The present invention will be further described in detail below with reference to the accompanying drawings.

[0022] Embodiment 1 Referring to Figure 1 , a virtual digital human expression recognition method based on multi-scale dilated convolution according to this embodiment includes: Acquire image data; Preprocess the acquired image data; Build a neural network model based on multi-scale dilated convolution using the preprocessed data, and train the neural network model with the preprocessed image data; Apply the trained model to image processing and output the facial expression recognition result; Drive the virtual digital human to generate facial expression animations according to the recognition result.

[0023] Specifically: S1: Input an image or video: Real-time acquire the facial image or video stream of the virtual digital human through the virtual digital human platform or video device, ensure that the resolution is not less than 224×224 pixels, support multiple formats and decompose the video into image frames.

[0024] S2: Data preprocessing and enhancement: Standardize the acquired image to 224×224 pixels, apply enhancement techniques to simulate facial expression changes, and locate and crop the facial area to improve data quality.

[0025] S3: Neural network model training: Based on the preprocessed data, train the model using the multi-scale dilated convolution network (MREmoNet), and embed innovative components to optimize facial expression classification.

[0026] S4: Output the facial expression recognition result: Apply the trained model to image or video processing, calculate the facial expression probability and output the result of the highest category.

[0027] S5: Drive the virtual digital human to generate facial expressions: Drive the virtual digital human to generate facial expression animations according to the recognition result, adjust facial features and optimize transitions to enhance the immersion and authenticity of the interaction Preferably, a further technical solution of the present invention is: For the input image or video processing in S1, the system acquires the facial image or video stream of the virtual digital human in real time through the virtual digital human generation platform or video capture device to ensure image quality and model compatibility. It is required that the resolution of the input image or video frame is not less than 224×224 pixels, and common image formats such as JPEG and PNG and MP4 video format are supported. For video input, the system will automatically decompose it into continuous image frames to provide a stable data basis for subsequent virtual digital human facial expression recognition.

[0028] For the data preprocessing and enhancement in S2, after image acquisition, data diversity is achieved through a geometric-illumination joint enhancement matrix to improve the recognition accuracy and generalization ability of the model. Define the enhancement matrix as: where is the random rotation angle, controls the scale scaling, , is the translation amount, is the brightness scaling factor, is the Gaussian noise perturbation term. This matrix generates diverse training samples through random parameter combinations, and can simulate complex scenarios such as side faces, occlusions, and light and dark changes, improving the robustness of the model under low-quality inputs.

[0029] The image is first adjusted to a unified 224×224 pixel RGB format, and the influence of illumination changes is eliminated through gray normalization. Subsequently, data augmentation techniques such as random rotation (-15° to 15°), horizontal flipping (50% probability), brightness adjustment (±20%), and contrast enhancement (±10%) are applied to simulate the expression changes of the virtual digital human in different scenarios. At the same time, face detection algorithms such as MTCNN are used to locate and crop the facial area of the virtual digital human, ensuring that the data input into the network focuses on the key parts of the expression, providing high-quality input data for model training.

[0030] For the neural network model training in S3, all preprocessed virtual digital human facial image data needs to be trained with a multi-scale dilated convolutional fusion attention residual network (MREmoNet) to optimize the expression classification performance. This network uses ResNet34 as the backbone structure, and embeds innovative components such as a multi-scale dilated convolutional fusion attention module (MDFA), a residual mask block (RMB), and a coordinate attention module (CoordAtt) to enhance the ability to extract and classify virtual digital human facial features. The RAdam optimizer is used during training, with an initial learning rate of 0.0001, a batch size of 48, and the ReduceLROnPlateau strategy is used to dynamically adjust the learning rate to improve efficiency.

[0031] The model construction process reforms the architecture and embeds modules based on ResNet34 as the backbone network. First, the basic structure of ResNet34 is retained, including the initial 7×7 convolutional layer (stride 2), batch normalization layer (BN), ReLU activation function, and 3×3 max pooling layer (stride 2). The number of input channels is adjusted to 3, the output feature map size is normalized to 56×56×64, and the four residual layers (Layer1-Layer4) of the original network are retained, with output feature map sizes of 56×56, 28×28, 14×14, and 7×7 in sequence. On the basis of the backbone network, innovative functional modules are embedded in stages: 1. The residual mask block (RMB) is embedded after each residual layer, and its structure is defined as: Among them, and represent the input and output feature maps respectively, and σ is the Sigmoid function. Denotes element-wise multiplication, followed by batch normalization after a 3×3 convolutional layer.

[0032] 2. The Coordinate Attention Module (CoordAtt) is inserted at the output end of Layer3, and its calculation process includes the following steps: First, global average pooling is performed separately along the height ( H ) and width ( W ) directions of the feature map to generate direction-sensitive feature vectors: where represents the feature value of the input feature map at position ( h , w ).

[0033] Then, the pooling results are concatenated and reduced in dimension through a 1×1 convolution, and passed through the ReLU activation: where represents the ReLU function.

[0034] After that, the channels are separated and the Sigmoid function is applied to generate spatial attention weights and , and then feature weighting is achieved through the outer product operation ( ): 3. The Multi-scale Dilated Convolution Fusion Attention Module (MDFA) is integrated at the end of Layer4 and contains five parallel processing branches: Branch1: 1×1 standard convolution (output channels 512) Branch2-4: 3×3 dilated convolutions with dilation rates of 6, 12, and 18 respectively (output channels 512 each) Branch5: Global average pooling (GAP) followed by a 1×1 convolution (output channels 512) The feature fusion process is achieved through channel concatenation (Concat) and attention weighting: where is the concatenated multi-scale feature (size 7×7×2560), is the channel attention weight, and ⊙ represents channel-wise multiplication. This design fuses local details and global context, enhancing the model's ability to express cross-scale expression features.

[0035] During the model training phase, S3 is executed according to the following steps: 1. Data input and feature extraction, The 224×224×3 normalized images output by S2 are input into the improved ResNet34 backbone network. The first layer uses a 7×7 convolutional kernel (stride = 2, padding = 3) for initial feature extraction, outputting a 56×56×64 feature map, which enters the residual layer stack after 3×3 max pooling (stride = 2).

[0036] 2. Multi-level residual feature processing, After being processed by Layer1, a 56×56×64 feature map is output, and then feature enhancement is performed through a Residual Mask Block (RMB): Among them, the output channels of the 3×3 convolution remain 64, σ is the Sigmoid function, and Sigmoid generates a spatial mask to enhance the response of the key expression regions.

[0037] Then the data after feature enhancement is input into Layer2. After being processed by Layer2, a 28×28×128 feature map is output, and then the RMB module is stacked and a channel dropout strategy (DropPath = 0.2) is introduced to prevent overfitting.

[0038] 3. Cross-dimensional attention fusion, The data processed by Layer2 is input into Layer3. After being processed by Layer3, a 14×14×256 feature map is output, and then the Coordinate Attention module (CoordAtt) is connected: First, bidirectional pooling is performed to generate H / W direction feature descriptors: It is reduced to C / 8 channels through a 1×1 convolution, and after ReLU activation, spatial attention weights are separated and generated: Then feature reweighting is used to enhance the spatial correlation of expression key points such as eyebrows and corners of the mouth: .

[0039] 4. Multi-scale feature fusion, The data processed by Layer3 is input into Layer4. After being processed by Layer4, a 7×7×512 feature map is output, and then the Multi-scale Dilated Convolution Fusion Attention module (MDFA) is connected: First, perform parallel five-way processing. The input data passes through the 1×1 standard convolution of Branch1 in parallel, and 512 channels are output; the 3×3 dilated convolutions (dilation = 6, 12, 18) of Branch2 to 4, each outputting 512 channels; the global average pooling layer and 1×1 convolution of Branch5.

[0040] Then, after concatenating the five-way data features, perform channel attention weighting to achieve the feature fusion of micro-expression wrinkles (small dilation rate) and facial contours (large dilation rate): 5. Optimize the classifier, The final feature map is sent to the fully connected layer after adaptive average pooling, and the label smoothing cross-entropy loss function is used: where = 0.1, C is the number of expression categories, to alleviate overfitting caused by the rendering style differences of virtual digital humans.

[0041] After the S4 model training is completed, the expression probability calculation and classification decision are achieved through the following steps: 1. The output feature map obtained from S3 is compressed to 1×1×2560 through adaptive average pooling: Then, output the feature vector: 2. Perform Softmax probability calculation. First, input the feature vector into the fully connected classification layer: Output the unnormalized category (C is the number of expression categories), and then, convert the probability distribution through the Softmax function: Output the probability vector , satisfying .

[0042] The expression corresponding to the largest value is the output expression.

[0043] The drive of S5 generates expressions for the virtual digital human, and drives it to generate corresponding expression animations according to the recognized expression categories. 3D models and animation parameters of seven basic expressions (such as raised eyebrows, upturned mouth corners, widened eyes, etc.) are designed in advance for the virtual digital human, and the facial features are dynamically adjusted according to the recognized expression categories. The specific steps are as follows: 1. Decode the expression category: Receive the expression probability vector output by S4 , extract the category index corresponding to the maximum probability , and parse it into semantic labels (such as "happy", "angry", etc.) according to the preset expression coding table.

[0044] 2. Match the parameter library: According to y Retrieve the pre-constructed expression-parameter mapping library to obtain the corresponding facial action unit (AU) weight vector (K is the number of AUs) and geometric deformation parameters (N is the number of facial mesh vertices).

[0045] The method for constructing the parameter library is as follows: First, define 52 basic AU weights based on the FACS system. Then, collect the blend shapes of seven basic expressions through 3D scanning. After that, establish a non-linear mapping relationship table between expression labels and parameter groups.

[0046] (1) Construction and preprocessing of the parameter library First, based on the FACS (Facial Action Coding System), define 52 basic AUs (facial action units), including action units such as raised eyebrows and upturned mouth corners, and collect data of seven basic expressions (happy, angry, sad, fear, surprise, disgust, neutral), as well as micro-expressions (such as slight twitching of the mouth corners) and blended expressions (such as "surprise + happy") through 3D scanning technology. The dataset is extended to include at least 100,000 groups of samples, covering different lighting, pose, and occlusion conditions to improve the generalization ability of the model. Each group of data contains an expression label, an AU weight vector (52-dimensional, corresponding to 52 AUs) and geometric deformation parameters , where N is the number of facial mesh vertices (the typical value is 6890, based on the SMPL-X model). Next, use the pre-trained BERT model (specifically using bert-base-uncased, with a 768-dimensional output) to perform semantic encoding on the expression labels to generate embedding vectors . At the same time, design a multi-layer perceptron (MLP) to map the AU weight vector and geometric deformation parameters to the same embedding space. The MLP structure is a three-layer fully connected network (the input layer dimension is 52 + N, the hidden layer is 512-dimensional, the output layer is 768-dimensional, and the activation function is ReLU). The specific mapping process is as follows: Concatenate into a vector, input it into the MLP, and output the embedding vector . To support efficient retrieval, the Faiss library (specifically using the IndexHNSWFlat index) is used to build a vector index. The HNSW parameters are set to M = 32 (number of connections) and efConstruction = 200 (search depth during construction) to balance retrieval speed and accuracy. The parameter library storage format is a list of tuples , and it is stored partitioned by expression categories, supporting dynamic updates (e.g., when adding a new expression category, only need to append a new tuple and update the index).

[0047] (2) Expression category decoding and semantic embedding, First, from the expression probability vector obtained in S4 (C is the number of expression categories), the category index corresponding to the maximum probability is extracted. Suppose y = [0.1, 0.6, 0.05, 0.05, 0.1, 0.05, 0.05], then argmax(y) = 1, corresponding to the label "happy". The index is mapped to the semantic label "happy", and the mapping table is predefined in dictionary form (e.g., {0: "angry", 1: "happy",...}). Then, the same BERT model (bert-base-uncased) as used for building the parameter library is used to encode it to generate an embedding vector . To enhance robustness, the global features of the input image are combined: the global features of the input image (size 224×224×3) are extracted using a pre-trained ResNet50 model (the first 50 layers are frozen, only the last layer is fine-tuned) , and passed through a dimensionality reduction MLP (from 2048 dimensions to 512 dimensions, with a structure of two fully connected layers, a hidden layer of 1024 dimensions, and an activation function ReLU) to reduce the dimensionality to . Then, a fusion MLP (input layer of 1280 dimensions, hidden layer of 1024 dimensions, output layer of 768 dimensions, activation function ReLU) is designed to and be fused into , and the formula is: To handle boundary conditions (such as overly smooth probability distributions), if the maximum probability is less than 0.3, then take the weighted label combination of the top two probabilities (e.g., 0.4 "happy" + 0.3 "surprised"), and generate a mixed embedding vector through linear interpolation. The fusion embedding vector of the predicted expression is output , which is used for subsequent parameter retrieval.

[0048] (3) Parameter retrieval, Using the Faiss index (IndexHNSWFlat), calculate using cosine similarity with all in the parameter library The similarity, with the formula as follows: Set the retrieval parameter efSearch = 50 (search depth), select the top k candidates (k = 3) with the highest similarity, and obtain the corresponding , and map it back to the AU weight through the parameter library and the geometric deformation parameter . For smoothing the results, perform weighted fusion on the top k candidates, and the weights are based on similarity normalization, with the formula as follows: To handle abnormal situations (such as too low similarity), if the highest similarity is lower than 0.5, return the default neutral expression parameters and record the log for subsequent optimization of the parameter library.

[0049] Output the AU weight vector obtained from the preliminary retrieval and the geometric deformation parameter , which serves as the input for subsequent fine-tuning.

[0050] (4)Dynamic parameter fine-tuning, Introduce a parameter fine-tuning module based on the variational autoencoder (VAE). The VAE structure includes an encoder (three-layer fully connected, with the input layer dimension of 52 + N + D, the hidden layer dimension of 1024, and the output layer dimension of 256, generating the latent distribution parameters ) and a decoder (three-layer fully connected, with a reverse structure, outputting , ). The input is the retrieval parameter and the facial key point features , where extracts 68 key points (2D coordinates for each point, D = 136) through OpenFace. The encoder encodes the input into the latent distribution , and the decoder samples from the latent distribution to generate the fine-tuning parameters. During training, the VAE optimization objectives include the reconstruction loss (mean squared error), KL divergence, and adversarial loss (ensuring the authenticity of the parameters through a discriminator), with the formula as follows: The hyperparameters are set to , .

[0051] Output the fine-tuned AU weight vector and the geometric deformation parameter , which is more suitable for the input image features.

[0052] (5)Caching and real-time optimization, Set a cache for high-frequency expression categories (such as "happy" and "neutral"), with a cache capacity of 1000 entries, and manage it using the LRU (Least Recently Used) strategy: If the cache is hit, directly return the stored , , with a hit rate of approximately 60% (statistically based on user expression distribution). When a miss occurs, perform retrieval and fine-tuning, and store the results in the cache. To further accelerate, use an asynchronous computing pipeline: Retrieval and fine-tuning are executed in background threads (Python thread pool, 4 threads), and the main thread focuses on animation rendering and user interaction. The asynchronous threads communicate with the main thread through a queue (Python multiprocessing.Queue) to ensure data synchronization.

[0053] Output the matching parameters optimized for real-time and , which are directly used for virtual digital human animation driving.

[0054] 3. Use LSTM for temporal modeling: Input the current frame expression parameters into a pre-trained LSTM network to predict the change amount of parameters at the next time step : where is the time interval between frames, is the hidden state, and the output dimension is the same as the number of AUs.

[0055] 4. Perform geometric-appearance joint driving: Construct a motion equation for the m th facial key point: (1) Data preparation, providing initial input data for the motion equation, including neutral state coordinates, AU deformation vectors, and initial parameters. First, extract the neutral state coordinates. Use the 3D face model SMPL-X to define the facial mesh, which contains N vertices (N = 6890). Extract the vertex coordinates when the neutral expression (without any AU activation) is present, denoted as , where m is the vertex index. Generate a neutral state model through a pre-trained model DECA to ensure accurate coordinates. Then construct the AU deformation vectors. Based on FACS, define 52 types of AUs (such as AU1: Inner Brow Raiser), and collect the vertex deformations when each AU is activated through 3D scanning. For each AU, calculate its deformation vector for each vertex, and the formula is: where is the vertex coordinate when the kth AU is activated.

[0056] Normalized deformation vector to ensure (e.g., ), to avoid excessive deformation. After that, obtain the fine-tuned AU weight vector from the parameter library matching step , each dimension , representing the activation intensity of the k-th AU.

[0057] (2) Parameter estimation, estimating the dynamic parameters in the motion equation and , and generating the residual term through appearance features . First, use LSTM to predict the decay constant and vibration frequency. Take the AU weight vector of the current frame, the AU weight sequence of the previous few frames (e.g., the previous 5 frames), and the global image feature of the current frame (extracted by ResNet, dimension 2048) as the input. Use a two-layer LSTM network (each layer with 256 hidden units, input dimension of 52 + 2048 52 + 2048 52+2048, output dimension of 2×52 2 \times 52 2×52) to predict the change in the decay constant and vibration frequency of each AU: where , is the predicted change.

[0058] Perform parameter update, initial value , preset based on muscle type (e.g., for the AU at the corner of the mouth: ), and the update formula is: where the constraint , , to ensure numerical stability.

[0059] Then, calculate the appearance feature-driven residual term. First, calculate and extract the facial texture gradient of the current frame image based on the RGB channels of the input image using the Sobel operator . Second, use a small MLP (input layer 2D, hidden layer 64D, output layer 3D) to map the texture gradient to the residual deformation: where the constraint e.g., ), to avoid excessive offset.

[0060] (3) Assemble all parameters into a complete motion equation to describe the dynamic deformation of each facial key point.

[0061] First, calculate the base deformation. For each AU, calculate its deformation contribution to the m-th key point: where t is the current time step (e.g., , is the frame interval).

[0062] Then, sum up the deformation contributions of all AUs: Output the preliminary motion equation .

[0063] (4) Constraint optimization. Perform constraint optimization on the motion equation to prevent unnatural deformations or those exceeding the reasonable range.

[0064] First, constrain the total displacement of each key point: where .

[0065] If it exceeds the range, scale it proportionally: Then, perform smoothness constraint. Impose smoothness constraint on the displacement change between adjacent frames and calculate the frame difference: If (e.g., ), then adjust it through exponential smoothing: Output the optimized motion equation : 5. Perform real-time animation generation and optimization. First, perform deformation field calculation. The goal of deformation field calculation is to generate a discrete frame-level deformation field from the continuous motion equation , where represents the displacement vector of the m-th vertex at the n-th frame.

[0066] The following are the detailed algorithm steps for deformation field calculation, which are divided into four stages: discretization calculation, difference smoothing, time interpolation, and constraint optimization.

[0067] (1) Discretization calculation. Discretize the motion equation into frame-level vertex coordinates to generate the initial deformation field.

[0068] First, perform time discretization: Given the frame rate (e.g., 30fps), the frame interval . For the nth frame, calculate the corresponding time . Use the motion equation to calculate the coordinates of each vertex at this time step: Then, calculate the initial deformation field: Calculate the deformation vector of each vertex: where is the reference coordinate in the neutral state.

[0069] After that, use GPU parallel acceleration: Allocate vertex calculations to be executed in parallel on the GPU, using CUDA or OpenGL shaders (e.g., GLSL). Assume there are N vertices (e.g., N = 6890), allocate a thread for each vertex, and calculate its and .

[0070] Output the initial deformation field , representing the displacement of each vertex in the nth frame.

[0071] (2) Difference smoothing, based on the mesh topology, smooth the deformation field to prevent excessive local deformation or mesh distortion. Detailed process: First, construct the neighborhood relationship: Use the topology of the facial mesh (e.g., triangular mesh) to construct a set of neighborhood vertices m for each vertex , which contains all the vertices connected to it (usually each vertex has 5 - 7 neighbors).

[0072] Secondly, Laplacian smoothing: Calculate the Laplacian deformation (difference deformation) of each vertex, defined as the average difference between the vertex deformation and the deformation of its neighborhood: Apply Laplacian smoothing to the deformation field and update the deformation vector: where is the smoothing coefficient, controlling the smoothing intensity.

[0073] Then, iterative smoothing: Repeat Laplacian smoothing 2 - 3 times to ensure local smoothing of the deformation field without losing the main features. Check the energy of the deformation field after each iteration: If the energy change is less than the threshold (e.g., ), stop the iteration.

[0074] Output the smoothed deformation field , the local shape change is natural and the mesh distortion is reduced.

[0075] (3) Temporal interpolation, which smooths the deformation field in the time dimension, ensures continuous inter-frame deformation, and avoids animation jitter. Detailed process: First, Catmull-Rom spline interpolation: Use Catmull-Rom spline interpolation to generate intermediate frame deformations. For each vertex m , take the current frame , the previous two frames and the predicted value of the next frame (predicted by LSTM or linearly extrapolated). Define the Catmull-Rom spline curve: where represents the time offset within the current frame.

[0076] Second, in-frame interpolation calculation: For the in-frame time , calculate the interpolated deformation: where is dynamically determined by the vertical synchronization signal of the rendering engine (for example represents the midpoint of the frame).

[0077] Then, GPU acceleration: The interpolation calculation is also executed in parallel on the GPU, using shaders or CUDA kernel functions, with one thread per vertex. The interpolation time for a single frame is about 2ms (based on NVIDIA RTX 4090, N = 6890).

[0078] Output the deformation field with smoothed time , and the inter-frame transition is smoother.

[0079] (4) Constraint optimization, which imposes geometric and physical constraints on the deformation field to ensure that the deformation conforms to the face structure and does not distort. Detailed process: First, check the deformation amplitude of each vertex: where . If it exceeds the range, then scale it proportionally: Second, perform physical constraints (minimization of deformation energy), define the deformation energy function, based on the elastic potential energy of the mesh: where is the elastic coefficient, represents the mesh edge.

[0080] Use gradient descent to optimize the deformation field and minimize , iterate 5 times, step size . Update the vertex coordinates after optimization: (3) Boundary processing. If the deformation of some vertices is abnormal (for example, texture gradient noise caused by occlusion), use the average deformation of neighboring vertices to fill: Output the optimized deformation field , which conforms to geometric and physical constraints and is suitable for real-time animation.

[0081] After that, perform rendering optimization. The steps are as follows: (1) Batch process vertex transformations through GPU Instancing; (2) Enable motion blur compensation for high-frequency vibration AUs; (3) Use an asynchronous computing pipeline to separate parameter updates and rendering threads.

[0082] (1) Batch process vertex transformations through GPU Instancing. The detailed process: First, prepare the vertex data. Obtain the frame-level deformation field from the deformation field calculation step , which represents the displacement of each vertex in the nth frame. Store the neutral state mesh vertices and the deformation field as vertex buffer objects (VBOs). Each vertex contains a position (3 floating-point numbers) and a deformation (3 floating-point numbers), for a total of 6 floating-point numbers. Use a texture buffer object (TBO) to store the AU weights , which is convenient for fast access by the GPU.

[0083] Second, implement GPU Instancing using OpenGL or Vulkan. Create a single mesh instance and batch process all vertices. Perform deformation transformations in the Vertex Shader. Read the neutral state position and deformation vector of the vertex, calculate the new position, and then apply the model-view-projection matrix (MVP matrix) to generate the final screen coordinates. Each vertex is processed through a single Draw Call to reduce the CPU-GPU communication overhead.

[0084] Then, optimize the Compute Shader. Use the Compute Shader to pre-compute the deformation field transformation to further improve performance. First, pass the deformation field and the neutral mesh into the Compute Shader, and then use the Compute Shader to parallelly calculate the new vertex positions. The output is stored in a shared buffer (Shader Storage Buffer Object, SSBO) for direct use by the rendering pipeline.

[0085] Then, optimize the performance by enabling vertex buffer object streaming updates (GL_STREAM_DRAW) to support dynamic deformation field updates. Use multi-sampling anti-aliasing (MSAA, 4x sampling) to reduce edge jaggedness and improve visual quality. The optimized vertex positions are directly used for rendering.

[0086] (2)Enable motion blur compensation for high-frequency vibration AUs. The detailed process is as follows: First, identify high-frequency vibration AUs. Extract the vibration frequency of each AU from the motion equation parameters. If the frequency exceeds a threshold (e.g., 3 Hz), it is marked as a high-frequency vibration AU. Calculate the maximum displacement amplitude of the AU on the vertices: obtain the maximum displacement through the dot product of the AU weight and the deformation vector. If it exceeds a threshold (e.g., 0.05 mm), motion blur is required.

[0087] Second, perform adaptive motion blur calculation. Based on the frequency and amplitude, adaptively calculate the blur intensity: the blur intensity is calculated proportionally according to the frequency and displacement amplitude, restricted between 0 and 1. Implement temporal accumulation blur in the fragment shader: record the vertex positions of the previous two frames and calculate the velocity vector. According to the velocity and blur intensity, sample multiple pixels (e.g., 5 sampling points) along the velocity direction and weight-average the colors. Use Gaussian weights for weighting to ensure smooth blur.

[0088] Then, optimize the performance: Apply motion blur only to the vertex regions affected by high-frequency vibration AUs (such as the corners of the mouth and eyelids) to reduce the computational overhead. Use a low-resolution buffer (half-resolution) to pre-calculate the blur effect and then up-sample and merge it into the main frame buffer to reduce the GPU load.

[0089] Output the rendered frame with motion blur compensation, which is visually smoother and has no obvious flickering in the high-frequency vibration regions.

[0090] (3)Adopt an asynchronous computing pipeline to separate parameter updates and rendering threads. The detailed process is as follows: First, thread separation and task division: Parameter update thread: The parameter update thread is responsible for tasks such as deformation field calculation, AU weight update, and LSTM prediction, running on the CPU or GPU computing threads. The rendering thread is responsible for vertex transformation, motion blur, and final drawing, running on the GPU rendering pipeline. Use a multi-threaded framework to manage these two threads to ensure that they can run independently without interference.

[0091] Second, double-buffer mechanism: Create two vertex buffer objects (VBOs) as double buffers: one as the front buffer and the other as the back buffer. The parameter update thread writes the calculated deformation field into the current back buffer, while the rendering thread reads the vertex data from the current front buffer for rendering. After each frame ends, the front and back buffers are swapped, and a mutex is used to ensure thread safety and avoid data competition issues.

[0092] Then, for the asynchronous computing pipeline: The parameter update thread runs on an independent computing queue (such as Vulkan's Compute Queue or CUDA Stream), while the rendering thread runs on the rendering queue (Vulkan Graphics Queue). Through an event synchronization mechanism (such as Vulkan events or OpenGL fence synchronization), it is ensured that the rendering thread starts working only after the computing task is completed, thus avoiding unnecessary waiting or data inconsistency issues.

[0093] After that, for latency optimization: The parameter update thread calculates one frame in advance. For example, when rendering the nth frame, it calculates the deformation field of the (n + 1)th frame and uses LSTM to predict the parameters of the next frame. If the computing thread lags, the rendering thread will use the deformation field of the previous frame, while recording the latency and dynamically adjusting the computing load (such as reducing the number of smoothing iterations). In this way, the average inter-frame latency can be controlled within 1 ms (based on RTX 4090, 30 fps).

[0094] The output parameter update is completely separated from the rendering thread, and the overall frame rate is increased by about 20%, stably reaching 30 - 60 fps (RTX4090, 1080p resolution).

[0095] Finally, for the rendering output, it realizes the real-time synchronization of the virtual digital human's expression and the user's emotion, enhancing the immersion and realism of human-computer interaction.

[0096] As Figure 2 shown, for the architecture of the multi-scale dilated convolution fusion attention residual network, the upper layer is the overall structure of the network, the two sides of the lower layer are the specific network levels of the residual layer, and the middle is the structure diagram of the multi-scale dilated convolution fusion attention module. Among them, Residaul Layer represents the residual layer, coordAtt represents the coordinate attention module, RMB represents the residual mask block, MDFA represents the multi-scale dilated convolution fusion attention module, conv represents the convolution operation, conv2d represents the two-dimensional convolution operation, rate represents the dilation rate of the convolution, AvgPool2d represents the average pooling operation on two-dimensional image data, GlobalAvgPooling represents the global pooling operation, and Sigmoid, Softmax, and ReLU are commonly used activation functions in neural networks.

[0097] Figure 3 In Figure 3Shows the attention heatmaps of different deep learning models when recognizing facial expressions. The figure contains six expressions: anger, disgust, fear, happy, neutral, sad, and surprised. Each row corresponds to an expression, and each column represents a different model or method: Original Image: The original image without any processing.

[0098] VGG19: The attention heatmap generated using the VGG19 model.

[0099] GoogLeNet: The attention heatmap generated using the GoogLeNet model.

[0100] ResNet34: The attention heatmap generated using the ResNet34 model.

[0101] Ours: The attention heatmap generated using the multi-scale dilated convolutional fusion attention residual network.

[0102] The colors in the heatmap represent the regions that the model focuses on when recognizing expressions. The brighter the color (such as red and yellow), the more the model focuses on that region. These heatmaps help the experimenter understand which parts of the face the model pays more attention to when making predictions. For example, when recognizing the "anger" expression, the model usually focuses on the areas around the eyebrows and eyes because these areas change significantly when expressing anger. By comparing the heatmaps of different models, the performance differences and characteristics in the expression recognition task can be analyzed.

[0103] Figure 4 Shows a schematic diagram of the expression driving process of a virtual digital human, which is mainly divided into three stages: feature extraction, prediction, and interaction. First, an input image of a person is processed by the feature extraction network "MREmoNet" to identify and extract key expression features. Then, the extracted features are further processed through average pooling operations and fully connected layers, and are converted into a probability distribution of different expression categories through the Softmax function to complete expression prediction. In the interaction stage, the system performs "empathy" processing on the prediction result and the expression state of the original virtual digital human, that is, adjusts the expression of the virtual digital human to match the expression of the input image. This adjustment is achieved through the driving module to ensure the synchronization of the expression of the virtual digital human with the expression of the input image. Finally, the virtual digital human image after expression adjustment is output, achieving consistency with the expression of the input image.

[0104] Example 2 This example provides a virtual digital human expression recognition system based on multi-scale dilated convolution, including: A data acquisition module, configured to: A computer-readable storage medium storing multiple instructions, the instructions being adapted to be loaded and executed by a processor of a terminal device for the virtual digital human expression recognition method based on multi-scale dilated convolution.

[0105] A terminal device, comprising a processor and a computer-readable storage medium, the processor being configured to implement each instruction; the computer-readable storage medium being configured to store multiple instructions, the instructions being adapted to be loaded and executed by the processor for the virtual digital human expression recognition method based on multi-scale dilated convolution.

[0106] The above are all preferred embodiments of the present invention, and the protection scope of the present invention is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention shall be covered within the protection scope of the present invention.

Claims

1. A virtual digital human expression recognition method based on multi-scale dilated convolution, characterized in that: include: Get image data; Preprocessing the acquired image data; A neural network model based on multi-scale dilated convolution is constructed based on the preprocessed data, and the neural network model is trained using the preprocessed image data; Apply the trained model to image processing and output expression recognition results; Drive the virtual digital human to generate facial expression animation based on the recognition results.

2. The method for virtual digital human expression recognition based on multi-scale dilated convolution according to claim 1, characterized in that: The preprocessing of the acquired image data includes converting the image data into data diversity using a geometry-illumination joint enhancement matrix, wherein the matrix is ​​used to generate diversified training samples by randomly combining parameters, simulating complex scenes of side faces, occlusions, and light and dark changes to improve the recognition accuracy and generalization ability of the model. The enhancement matrix is ​​expressed as: in is a random rotation angle, Control the scale. , is the translation amount, is the brightness scaling factor, is the Gaussian noise disturbance term.

3. The method for virtual digital human expression recognition based on multi-scale dilated convolution according to claim 2, characterized in that: The neural network model based on multi-scale dilated convolution is constructed based on the preprocessed data, including taking ResNet34 as the backbone structure, embedding multi-scale dilated convolution fusion attention mechanism, residual mask mechanism and coordinate attention mechanism to construct the neural network model, wherein the residual mask mechanism is embedded in each residual layer, and its structure is: in, and Respectively represent the input and output feature maps, σ is the Sigmoid function, represents element-wise multiplication, a 3×3 convolutional layer followed by batch normalization.

4. The method for virtual digital human expression recognition based on multi-scale dilated convolution according to claim 3, characterized in that: The method of training the neural network model using the preprocessed image data includes extracting features from the 224×224×3 standardized image data using the Layer1 layer of the neural network model, outputting a 56×56×64 feature map, performing feature enhancement through a residual mask mechanism, and then outputting a 28×28×128 feature map after processing by the Layer2 layer, superimposing the RMB mechanism and introducing a channel discarding strategy to prevent overfitting. The feature enhancement of the Layer1 layer is expressed as: The 3×3 convolution output channel remains at 64, σ is the Sigmoid function, and Sigmoid generates a spatial mask to enhance the response of key expression areas.

5. The method for virtual digital human expression recognition based on multi-scale dilated convolution according to claim 4, characterized in that: The method of training the neural network model using the pre-processed image data also includes inputting the data processed by Layer2 into Layer3, outputting a 14×14×256 feature map after processing by Layer3, and then connecting to the coordinate attention mechanism, wherein firstly, bidirectional pooling is performed to generate an H / W direction feature descriptor, and the dimension is reduced to C / 8 channels through 1×1 convolution, and then separated and generated spatial attention weights after ReLU activation, and then feature re-weighting is used to enhance the spatial correlation of the expression key points of the eyebrows and the corners of the mouth. The H / W direction feature descriptor is expressed as: Where the height is h and the width is w.

6. The method for virtual digital human expression recognition based on multi-scale dilated convolution according to claim 5, characterized in that: The method of training the neural network model using the pre-processed image data also includes inputting the data processed by Layer3 into Layer4, outputting a 7×7×512 feature map after being processed by Layer4, and then connecting to a multi-scale dilated convolution fusion attention mechanism, wherein a parallel five-way processing is adopted, the input data is parallelly passed through the 1×1 standard convolution of Branch1, and 512 channels are output; the 3×3 hole convolutions of Branch2 to 4, each outputting 512 channels; the global average pooling layer and 1×1 convolution of Branch5, and then the five-way data features are spliced ​​and weighted by channel attention to achieve the feature fusion of micro-expression lines and facial contours, which is expressed as: 。 7. The method for virtual digital human expression recognition based on multi-scale dilated convolution according to claim 6, characterized in that: The trained model is applied to image processing to output the expression recognition result, including using the trained model to compress the fused feature map through adaptive average pooling, and performing Softmax probability calculation. The corresponding expression is the output expression, and the Softmax function conversion probability distribution is expressed as: Output probability vector ,satisfy .

8. The method for virtual digital human expression recognition based on multi-scale dilated convolution according to claim 7, characterized in that: The method drives the virtual digital human to generate expression animation according to the recognition result, including: Extract the category index corresponding to the maximum probability , according to the preset expression encoding table, it is parsed into semantic tags. y Retrieve the pre-built expression-parameter mapping library and obtain the corresponding facial action unit AU weight vector and geometric deformation parameters , thereby matching the parameter library, where K is the number of AUs and N is the number of facial mesh vertices.

9. The method for virtual digital human expression recognition based on multi-scale dilated convolution according to claim 8, characterized in that: The method of driving the virtual digital human to generate expression animation according to the recognition result also includes inputting the expression parameters of the current frame into the pre-trained LSTM network, predicting the parameter change amount in the next time step, constructing the motion equation for the mth facial key point to perform geometry-appearance joint driving, calculating the deformation field and discretizing the driving equation into frame-level deformation to generate real-time animation. The driving equation is expressed as: In the formula, the appearance feature generates the residual term ; Global image features of the current frame ; AU weight vector of the current frame .

10. A virtual digital human expression recognition system based on multi-scale dilated convolution, comprising: A data acquisition module is configured to acquire image data; A preprocessing module is configured to preprocess the acquired image data; A model training module is configured to construct a neural network model based on multi-scale dilated convolution based on the preprocessed data, and train the neural network model using the preprocessed image data; The recognition module is configured to apply the trained model to image processing and output expression recognition results; The animation module is configured to drive the virtual digital human to generate expression animation according to the recognition result.

Citation Information

Cited By

  • Fine-grained facial expression digital human head portrait generation method and system based on consistent identity recognition

    CN121600139A