Speech-driven face video generation method and device based on depth perception fusion
By fusing information from RGB images and depth maps and combining cross-modal attention learning networks, the problem of ignoring 3D facial geometric structure in the prior art is solved, and realistic and natural facial videos are generated.
Patent Information
- Application Number
- CN202510311256.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-17
AI Technical Summary
The prior art ignores the geometric relationship between the 3D face of the image in the voice-driven face video generation, resulting in the facial details of the generated video that are not vivid, limiting its application in actual scenes.
A face video generation model is constructed, and the semantic information of RGB images and the depth information of the depth map is fused through the cross-reference module, and combined with a cross-modal attention learning network to improve facial structure accuracy and fine-grained details of motion.
The generated video is more realistic and natural, providing more application flexibility.
Smart Images

Figure CN119832929B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision and image processing, and particularly to a voice-driven face video generation method and device based on depth perception fusion. Background Art
[0002] The voice-driven face video generation technology is a cutting-edge technology that combines speech processing and computer vision. Its main goal is to generate high-quality human videos that match the speech input by analyzing and understanding the speech input. In this field, researchers not only need to overcome the semantic gap between speech and video, but also need to solve the problems of temporal synchronization and action expression matching between speech and video.
[0003] Currently, most research mainly relies on learning 2D facial representations (such as appearance and motion) from input images, while ignoring the use of 3D facial geometric structure relationships in images, lacking perception and exploration of facial depth information, and lacking corresponding 2D-3D feature fusion optimization research, resulting in the generated video facial details being not vivid, which limits the popularization and application of face video generation in practical scenarios. Summary of the Invention
[0004] In view of the problems raised in the above background art, the present invention proposes a voice-driven face video generation method and device based on depth perception fusion, constructs a face video generation model, constructs a cross-reference module according to facial motion attributes, mines rich semantic information and texture information in RGB images and key depth information in depth maps, and combines a cross-modal attention learning network, effectively improving the facial structure accuracy of the face video while taking into account the fine-grained details of motion, making it more realistic and natural.
[0005] On the one hand, a voice-driven face video generation method based on depth perception fusion is as follows:
[0006] S1, obtain a face speaking video dataset with audio segments and reference images, and after preprocessing the dataset, divide it into a training dataset and a test dataset;
[0007] S2, construct a face video generation model; the face video generation model includes an audio encoder, an image encoder, a depth encoder, a cross-reference module, and a cross-modal attention module; the audio encoder extracts the Mel spectrogram features of the audio in the dataset; the image encoder extracts the RGB features of the images in the dataset; the depth encoder extracts the depth map features of the images in the dataset; the cross-reference module fuses the depth map features and RGB features; the cross-modal attention module introduces a self-attention mechanism to enhance the facial structure retention ability;
[0008] S3. Use the training data set to train the face video generation model to obtain a trained face video generation model;
[0009] S4. Input the test data set into the trained face video generation model and output a face video combining audio and video.
[0010] Preferably, the cross-reference module is specifically as follows:
[0011] S21. Perform global average pooling on the depth map features and RGB features respectively to obtain an RGB feature vector and a depth map feature vector; among them, the input feature generated by the th convolutional block of the RGB feature is expressed as , and the input feature generated by the th convolutional block of the depth map feature is expressed as ;
[0012] S22. Input the RGB feature vector and the depth map feature vector into the fully connected layer and the Softmax activation function respectively to obtain two-channel attention vectors. The formula is as follows:
[0013] ;
[0014] ;
[0015] Among them, and represent the parameters of the fully connected layer of the th feature; represents the global average pooling operation; represents the channel attention vector of the RGB feature; represents the channel attention vector of the depth map feature; represents the Softmax activation function;
[0016] S23. Multiply the two-channel attention vectors by the corresponding input features channel by channel to generate channel-enhanced features, as follows:
[0017] ;
[0018] ;
[0019] Among them, represents multiplication by channel; represents the channel-enhanced feature of the RGB feature; represents the channel-enhanced feature of the depth map feature;
[0020] S24. The channel attention vectors and Aggregate through the maximum function and perform a normalization operation to obtain the fused channel attention vector, denoted as:
[0021] ;
[0022] wherein, denotes the fused channel attention vector; denotes the normalization operation; denotes the maximum function aggregation;
[0023] S25, based on the fused channel attention vector , use to and perform feature enhancement to obtain the enhanced features and ; further concatenate the two enhanced features and feed them into a 1×1 convolutional layer to generate the cross-modal fusion feature , and this process is denoted as:
[0024] ;
[0025] ;
[0026] ;
[0027] wherein, denotes the enhanced feature of the RGB feature; denotes the enhanced feature of the depth map feature; denotes the 1×1 convolutional layer; denotes the concatenation connection; denotes the cross-modal fusion feature.
[0028] Preferably, the cross-modal attention module is specifically as follows:
[0029] Take the source depth map and the source image feature after warping deformation as the input; the source depth map is obtained from the reference image in the dataset; the is generated by the motion field generator; the motion field generator predicts the dynamic motion information of the face in the video in the face video generation model;
[0030] Input the source depth map into the depth encoder to obtain the depth feature map , and then perform a linear projection on through a 1×1 convolutional layer to the latent feature map ; perform linear projections on through another two 1×1 convolutional layersPerform a linear projection to obtain the latent feature maps respectively and ; Use as the query parameter in the self-attention mechanism, use as the key parameter in the self-attention mechanism, and use as the value parameter in the self-attention mechanism; The final feature is expressed by the following formula:
[0031] ;
[0032] where Softmax represents the Softmax normalization function; represents the linear projection of the query parameter; represents the linear projection of the key parameter; represents the linear projection of the value parameter.
[0033] Preferably, the final objective function of the face video generation model is expressed as:
[0034] ;
[0035] where represents the loss function of the generative adversarial network, represents the loss weight of the loss of the generative adversarial network; represents the perceptual reconstruction loss function, represents the loss weight of the perceptual reconstruction loss; represents the key point loss function, represents the loss weight of the key point loss function; represents the input-driven frame, represents the corresponding reconstructed frame.
[0036] Preferably, the loss function of the generative adversarial network adopts the least squares loss; The loss function of the generative adversarial network is expressed as follows:
[0037] ;
[0038] where represents adjusting the parameters of the discriminator D to make the loss function of the discriminator reach the minimum value; represents adjusting the generator to make the loss function of the generator reach the minimum value; represents the function of the generator; represents the function of the discriminator; represents the noise, following a normalized or Gaussian distribution; Represents real data of the probability distribution; Represents the probability distribution under of the expected value; Represents of the probability distribution; Represents the probability distribution under of the expected value; Represents the probability distribution under of the expected value; Represents the label of the generated sample; Represents the target label of the generated sample; Represents the label of the real sample.
[0039] Preferably, the perceptual reconstruction loss uses a pre-trained VGG-19 network as the network structure for evaluating the loss; the perceptual reconstruction loss is expressed as:
[0040] ;
[0041] wherein, Represents the th channel feature extracted from a specific VGG-19 layer, Represents the number of feature channels in this layer; Represents the absolute value.
[0042] Preferably, the key point loss function is expressed as:
[0043] ;
[0044] wherein, Represents the intermediate representation of the key points of the reference image in the dataset; Represents the intermediate representation of the key points of the generated image; Represents the intermediate representation of the Jacobian matrix of the reference image in the dataset; Represents the intermediate representation of the Jacobian matrix of the generated image; T represents time; Represents the L1 loss.
[0045] On the other hand, a speech-driven face video generation device based on depth perception fusion includes the following:
[0046] A data acquisition and preprocessing module, which is used to acquire a face talking video dataset with audio segments and reference images, preprocess the dataset, and divide it into a training dataset and a test dataset according to a ratio;
[0047] A model construction module for constructing a face video generation model; the face video generation model includes an audio encoder, an image encoder, a depth encoder, a cross-reference module, and a cross-modal attention module; the audio encoder extracts the Mel spectrogram features of the audio in the dataset; the image encoder extracts the RGB features of the images in the dataset; the depth encoder extracts the depth map features of the images in the dataset; the cross-reference module fuses the depth map features and the RGB features; the cross-modal attention module introduces a self-attention mechanism to enhance the facial structure retention ability;
[0048] A model training module for training the face video generation model using a training dataset to obtain a trained face video generation model;
[0049] A model testing module for inputting a test dataset into the trained face video generation model and outputting a generated face video combining audio and video.
[0050] Compared with the prior art, the present invention has the following beneficial effects:
[0051] The face video generation model of the present invention models the facial motion representation using the two-dimensional facial representation and the three-dimensional facial geometric structure relationship of the reference image, constructs a cross-reference module according to the facial motion attributes, mines the rich semantic information and texture information in the RGB image and the key depth information in the depth map, and combines a cross-modal attention learning network, which not only effectively improves the facial structure accuracy of the face video but also takes into account the fine-grained details of the motion, making it more realistic and natural, thus providing more flexibility for various application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The following further describes the present invention in detail with reference to the drawings;
[0053] Figure 1 It is a flowchart of a voice-driven face video generation method based on depth perception fusion according to an embodiment of the present invention;
[0054] Figure 2 It is an overall model diagram of a voice-driven face video generation method based on depth perception fusion according to an embodiment of the present invention;
[0055] Figure 3 It is a structural diagram of a cross-reference module of a voice-driven face video generation method based on depth perception fusion according to an embodiment of the present invention;
[0056] Figure 4 It is a structural block diagram of a voice-driven face video generation device based on depth perception fusion according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0057] The following further describes the present invention through specific embodiments.
[0058] See Figure 1 As shown, the speech-driven face video generation method based on depth perception fusion, the specific steps include:
[0059] S1. Obtain a face speaking video dataset with audio segments and reference images, and after preprocessing the dataset, divide it into a training dataset and a test dataset.
[0060] Specifically, extract various different audio-visual segments such as speakers' speeches, lectures, interviews, etc. from existing public videos to construct the required dataset. The extracted content is intercepted to a length of 4 seconds, and at the same time, the video frame rate is converted to 25 FPS and frame division is performed, and the corresponding speech content sampling rate is converted to 16k. Use the librosa audio processing library to preprocess and extract features from the speech data. The obtained audio features altogether include 41 acoustic features, including 13 Mel-frequency cepstral coefficients, 26 Mel-filter bank energy features, pitch, and unvoiced parts.
[0061] In addition, the dataset also includes a validation dataset for subsequent performance evaluation of the face video generation model.
[0062] S2. Construct a face video generation model; the face video generation model includes an audio encoder, an image encoder, a depth encoder, a cross-reference module, and a cross-modal attention module; the audio encoder extracts the MelP features of the audio in the dataset; the image encoder extracts the RGB features of the images in the dataset; the depth encoder extracts the depth map features of the images in the dataset; the cross-reference module fuses the depth map features and RGB features; the cross-modal attention module introduces a self-attention mechanism to enhance the facial structure retention ability. The overall model diagram of the face video generation model is shown in Figure 2 as shown.
[0063] See Figure 3 As shown, the cross-reference module (CRM) aims to mine and combine the most discriminative feature channels in the depth map features and RGB features and generate more informative features. Specifically, given two input features produced by the th and convolutional blocks of the RGB feature and the depth map feature respectively, first use global average pooling (Global Average Pooling, GAP) to obtain the global information in the RGB image and the depth view. Then, input the two feature vectors into a fully connected layer (FC layer) and a Softmax activation function to obtain the channel attention vectors and , which respectively reflect the importance of RGB features and depth features. Then, the attention vector is multiplied with the input features channel by channel to dynamically adjust the weights of each feature channel. In this way, the Cross-Reference Module (CRM) will automatically identify the feature channels closely related to the current task while suppressing those redundant or irrelevant feature channels, thereby improving the quality and discriminative power of feature representation. The mathematical formula for this process is as follows:
[0064] ;
[0065] ;
[0066] where, and represent the parameters of the FC layer of the th feature, while represents the global average pooling operation.
[0067] Then, the channel-enhanced features , ; where, represents channel-wise multiplication. In addition, the attention vectors and are aggregated through the max function to retain the useful feature channels from the RGB image and the depth map, and then fed into the normalization operation to normalize the output to the range from 0 to 1. Thus, the cross-referenced fused channel attention vector is obtained. This process is expressed as:
[0068] ;
[0069] Based on the fused channel attention vector , adding and to enhanced features, the enhanced features and can be obtained. The enhanced features of the RGB branch and the depth branch are further concatenated and fed into a 1×1 convolutional layer to generate cross-modal fusion features , and this process is expressed as:
[0070] ;
[0071] ;
[0072] ;
[0073] where, represents the enhanced feature of RGB features; represents the enhanced feature of depth map features; Denotes a 1×1 convolutional layer; Denotes a concatenation connection.
[0074] Cross-modal attention module. A cross-modal attention module is used to generate a dense depth-aware attention map to guide the source image features for face generation. The source image features are obtained through a motion field generator. The motion field generator in the face video generation model predicts the dynamic motion information of the face in the video based on audio features. The cross-modal attention learning network in the cross-modal attention module aims to enhance the model's ability to retain facial structures and generate subtle facial movements closely related to expressions. The introduction of depth information is based on its ability to provide detailed 3D geometric information for the model, which is crucial for maintaining the integrity of facial structures and identifying key expression movements.
[0075] Specifically, the depth encoder takes the source depth map of the input reference image as input and encodes it to obtain a depth feature map , and then performs a linear projection on and the source image features obtained by warping and deformation to three latent feature maps , and . , and denote the query, key, and value in the self-attention mechanism respectively. Therefore, the geometry-related query features generated from the depth map can be fused with the appearance-related key features to generate a dense motion field for face generation. And the final feature is used to generate facial images, and its formula is expressed as follows:
[0076] ;
[0077] where, denotes the linear projection of the query parameter; denotes the linear projection of the key parameter; denotes the linear projection of the value parameter; Softmax denotes the Softmax normalization function, which outputs a dense depth-aware attention map containing important 3D geometric information for generating a face with more fine-grained facial structures and motion details. Finally, the decoder takes the refined warped features as input to generate the final synthesized image.
[0078] The final objective function of the face video generation model is:
[0079] ;
[0080] Among them, , and are hyperparameters during training, representing the loss weights for the generative adversarial network loss, perceptual reconstruction loss, and keypoint loss used to constrain the model to generate videos, respectively. Specifically, the least squares loss is introduced to replace the binary cross-entropy loss in the traditional generative adversarial network. The loss function of the generative adversarial network in this method is expressed as follows:
[0081] ;
[0082] Among them, G represents the generator; D represents the discriminator, represents adjusting the parameters of the discriminator D to minimize the loss function of the discriminator ; represents adjusting the generator to minimize the loss function of the generator ; represents the function of the generator; represents the function of the discriminator; it is used to evaluate the synchronization degree between the lip movement in the generated video frame and the input audio; represents noise, following a normalized or Gaussian distribution; represents the probability distribution of the real data ; represents ; is the expected value, is also the expected value; the perceptual reconstruction loss uses the pre-trained VGG-19 network as the network structure for evaluating the loss. For the input driving frame and the corresponding reconstructed frame , the perceptual reconstruction loss is expressed as:
[0083] ;
[0084] Among them, is the th channel feature extracted from a specific VGG-19 layer, is the number of feature channels in this layer; represents the absolute value. The keypoint loss function constrains the keypoints and Jacobian matrix for generating the motion field in this method, and its mathematical formula is expressed as follows:
[0085] ;
[0086] Among them, T represents time, represents the L1 loss, and respectively represent the intermediate representations of the key points of the input reference image and the generated image, and respectively represent the intermediate representations of the Jacobian matrices of the input reference image and the generated image, both of which are obtained by the key point detection network.
[0087] S3. Use the training data set to train the face video generation model to obtain a trained face video generation model.
[0088] S4. Input the test data set into the trained face video generation model and output the generated face video combining audio and video.
[0089] As Figure 4 shown, the present invention also discloses a method and device for speech-driven face video generation based on depth perception fusion, including:
[0090] A data acquisition and preprocessing module 401, configured to acquire a face speaking video data set with audio segments and reference images, preprocess the data set, and divide it into a training data set and a test data set according to a ratio.
[0091] A model construction module 402, configured to construct a face video generation model; the face video generation model includes an audio encoder, an image encoder, a depth encoder, a cross-reference module, and a cross-modal attention module; the audio encoder extracts the Melp features of the audio in the data set; the image encoder extracts the RGB features of the images in the data set; the depth encoder extracts the depth map features of the images in the data set; the cross-reference module fuses the depth map features and the RGB features; the cross-modal attention module introduces a self-attention mechanism to enhance the facial structure retention ability.
[0092] A model training module 403, configured to use the training data set to train the face video generation model to obtain a trained face video generation model.
[0093] A model testing module 404, configured to input the test data set into the trained face video generation model and output the generated face video combining audio and video.
[0094] The specific implementation of the speech-driven face video generation device based on depth perception fusion is the same as that of the speech-driven face video generation method based on depth perception fusion, and will not be repeated in this embodiment.
[0095] The above is only the specific implementation manner of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantive modification made to the present invention using this concept shall fall within the scope of infringement of the protection scope of the present invention.
Claims
1. A method for generating a speech-driven face video based on depth perception fusion, characterized in that, It includes the following steps: S1. Obtain a face speaking video dataset with audio segments and reference images, and after preprocessing the dataset, divide it into a training dataset and a test dataset; S2. Construct a face video generation model; the face video generation model includes an audio encoder, an image encoder, a depth encoder, a cross-reference module, and a cross-modal attention module; the audio encoder extracts the MelP features of the audio in the dataset; the image encoder extracts the RGB features of the images in the dataset; the depth encoder extracts the depth map features of the images in the dataset; the cross-reference module fuses the depth map features and the RGB features; the cross-modal attention module introduces a self-attention mechanism to enhance the facial structure retention ability; S3. Use the training dataset to train the face video generation model to obtain a trained face video generation model; S4. Input the test dataset into the trained face video generation model to output a face video combining audio and video; The cross-reference module is specifically as follows: S21, perform global average pooling on the depth map features and RGB features respectively to obtain an RGB feature vector and a depth map feature vector; among them, the input feature generated by the i-th convolutional block of the RGB feature is represented as F i RGB , and the input feature generated by the i-th convolutional block of the depth map feature is represented as F i Depth ; S22. Input the RGB feature vector and the depth map feature vector into a fully connected layer and a Softmax activation function respectively to obtain a two-channel attention vector, and the formula is as follows: where, w i and b i represent the parameters of the fully connected layer of the i-th feature; AvgPooling(·) represents the global average pooling operation; represents the channel attention vector of the RGB feature; represents the channel attention vector of the depth map feature; δ(·) represents the Softmax activation function; S23. Multiply the two-channel attention vector with the corresponding input features channel by channel to generate channel-enhanced features, as follows: Among them, represents multiplication by channels; represents the channel enhancement feature of RGB features; represents the channel enhancement feature of depth map features; S24, aggregate the channel attention vectors and through the maximum function and perform a normalization operation to obtain the fused channel attention vector of the cross-reference, expressed as: Among them, represents the fused channel attention vector; represents the normalization operation; Max(·) represents the maximum function aggregation; S25, based on the fused channel attention vector Use To And Perform feature enhancement to obtain enhanced features And Further concatenate the two enhanced features and feed them into a 1×1 convolutional layer to generate the cross-modal fusion feature F i , and this process is expressed as: Among them, represents the enhanced feature of RGB features; represents the enhanced feature of depth map features; Conv 1×1 represents a 1×1 convolutional layer; Concat represents a concatenation connection; F i represents the cross-modal fusion feature.
2. The method for generating a voice-driven face video based on depth perception fusion according to claim 1, wherein The cross-modal attention module is specifically as follows: With the source depth map D s and the distorted source image feature F w as the input; the source depth map D s is obtained from the reference image in the dataset; the F w is generated by a motion field generator; the motion field generator predicts the dynamic motion information of the face in the video in a face video generation model; Input the source depth map D s into the depth encoder to obtain the depth feature map F d . Then, perform a linear projection on F d through a 1×1 convolutional layer to the latent feature map F q ; perform linear projections on F w through another two 1×1 convolutional layers to respectively obtain the latent feature maps F k and F v ; use F q as the query parameter in the self-attention mechanism, use F k as the key parameter in the self-attention mechanism, and use F v as the value parameter in the self-attention mechanism; the final obtained feature F g is expressed by the following formula: F g = Softmax((W q F d )(W k F w )) × (W T ) × (W v F w )); Among them, Softmax represents the Softmax normalization function; W q represents the linear projection of the query parameter; W k represents the linear projection of the key parameter; W v represents the linear projection of the value parameter.
3. The method for generating a speech-driven face video based on depth perception fusion according to claim 1, wherein The final objective function of the face video generation model is expressed as: Among them, represents the final objective function; represents the generative adversarial network loss function, λ G represents the loss weight of the generative adversarial network loss; represents the perceptual reconstruction loss function, λ rec represents the loss weight of the perceptual reconstruction loss; represents the key point loss function, λ kp represents the loss weight of the key point loss function; represents the input-driven frame, represents the corresponding reconstructed frame.
4. The method for generating a voice-driven face video based on depth perception fusion according to claim 3, wherein The loss function of the generative adversarial network adopts the least squares loss; the loss function of the generative adversarial network is expressed as follows: Among them, means adjusting the parameters of discriminator D to minimize the loss function V(D) of discriminator D; means adjusting generator G to minimize the loss function V(G) of generator G; G(·) represents the function of the generator; D(·) represents the function of the discriminator; z represents noise, following a normalized or Gaussian distribution; represents the probability distribution of the real data x; represents the probability distribution the expected value of x under; represents the probability distribution of z; represents the probability distribution the expected value of z under; represents the probability distribution the expected value of x under; a represents the label of the generated sample; b represents the target label of the generated sample; c represents the label of the real sample.
5. The method for generating a speech-driven face video based on depth perception fusion according to claim 3, wherein The perceptual reconstruction loss uses a pre-trained VGG-19 network as the network structure for evaluating the loss; the perceptual reconstruction loss is expressed as: Among them, VGG m (·) represents the m-th channel feature extracted from a specific VGG-19 layer, M represents the number of feature channels in the specific VGG-19 layer; |·| represents the absolute value.
6. The method for generating a speech-driven face video based on depth perception fusion according to claim 3, wherein The key point loss function is expressed as: Among them, represents the intermediate representation of the key points of the reference image in the dataset; represents the intermediate representation of the key points of the generated image; represents the intermediate representation of the Jacobian matrix of the reference image in the dataset; represents the intermediate representation of the Jacobian matrix of the generated image; T represents time; ||·||1 represents the L1 loss.
7. A speech-driven face video generation device based on depth perception fusion, including the following: A data acquisition and preprocessing module, which is used to obtain a face speaking video dataset with audio segments and reference images, and after preprocessing the dataset, divide it into a training dataset and a test dataset according to a ratio; A model construction module, which is used to construct a face video generation model; the face video generation model includes an audio encoder, an image encoder, a depth encoder, a cross-reference module, and a cross-modal attention module; the audio encoder extracts the MelP features of the audio in the dataset; the image encoder extracts the RGB features of the images in the dataset; the depth encoder extracts the depth map features of the images in the dataset; the cross-reference module fuses the depth map features and the RGB features; The cross-modal attention module introduces a self-attention mechanism to enhance the facial structure retention ability; A model training module, which is used to use the training dataset to train the face video generation model to obtain a trained face video generation model; A model testing module, which is used to input the test dataset into the trained face video generation model and output the generated face video combining audio and video; The cross-reference module is specifically as follows: S21, perform global average pooling on the depth map features and RGB features respectively to obtain an RGB feature vector and a depth map feature vector; among them, the input feature generated by the i-th convolutional block of the RGB feature is represented as F i RGB , and the input feature generated by the i-th convolutional block of the depth map feature is represented as F i Depth ; S22. Input the RGB feature vector and the depth map feature vector into a fully connected layer and a Softmax activation function respectively to obtain a two-channel attention vector, and the formula is as follows: where w i and b i represent the parameters of the fully connected layer of the i-th feature; AvgPooling(·) represents the global average pooling operation; represents the channel attention vector of the RGB feature; represents the channel attention vector of the depth map feature; δ(·) represents the Softmax activation function; S23. Multiply the two-channel attention vectors with the corresponding input features channel by channel to generate channel-enhanced features as follows: Among them, represents multiplication by channels; represents the channel enhanced feature of RGB features; represents the channel enhanced feature of depth map features; S24, Aggregate the channel attention vector and through the maximum function, and perform a normalization operation to obtain the fused channel attention vector of the cross-reference, expressed as: Among them, represents the fused channel attention vector; represents the normalization operation; Max(·) represents the maximum function aggregation; S25, based on the fused channel attention vector Use to and perform feature enhancement to obtain enhanced features and Further concatenate the two enhanced features and feed them into a 1×1 convolutional layer to generate the cross-modal fusion feature F i , and this process is expressed as: Among them, represents the enhanced feature of RGB features; represents the enhanced feature of depth map features; Conv 1×1 represents a 1×1 convolutional layer; Concat represents a concatenated connection; F i represents the cross-modal fusion feature.
Citation Information
Patent Citations
Voice-driven speaking face video generation method based on teacher-student network
CN113628635A
Speaking face video generation method and system based on adaptive region occlusion
CN117153195A