A method and system for face video style transfer

By constructing a generator network model, using technical means such as conditional reversible residual blocks and deformation protection layers, the problems of insufficient facial details and structural deformation in facial video style migration in the existing technology are solved, and high-quality facial video style migration is achieved to ensure style uniformity and content fidelity.

CN119762331BActive Publication Date: 2025-05-27BEIJING JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510259447.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-05-27
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

The existing facial video style transfer methods have problems such as insufficient facial details, structural deformation and discontinuity of video frames, and cannot fully retain and reconstruct facial details of the face, affecting style consistency and content fidelity.

Method used

Build a generator network model, including an encoder and multiple conditional reversible residual blocks, introduce content-driven multi-scale convolution strategies and spatiotemporal attention mechanisms, and introduce an affine transformation layer and a deformation protection layer in the conditional reversible residual block. Through these technical means, the target style is naturally applied to the face and generate realistic stylized videos.

Benefits of technology

The high-quality facial video style transfer is achieved, ensuring that the generated video performs well in terms of unified style, rich facial details and smooth timeliness, solving the problems of insufficient facial details and structural deformation, and improving style consistency and content fidelity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119762331B_ABST
    Figure CN119762331B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for face video style transfer, belonging to the field of computer vision technology. The acquired face video frame images are input into the constructed generator network model, and the feature sequence of the face video frame images output by the encoder of the model is subjected to an affine transformation, and the feature sequence is fused into target style features through a transformation function; in the deformation protection layer, the features of the facial region of the target style features are constrained according to the face mask to obtain the stylized face frame image features generated by the current conditional invertible residual block; and the stylized face frame image features generated by the current conditional invertible residual block are sequentially input into the next conditional invertible residual block for cyclic execution, and the stylized face frame image after the fusion of multiple conditional invertible residual blocks is obtained through decoding by the decoder. The video generated by this method has significant improvements in terms of style consistency, facial detail retention, and temporal coherence, realizing high-quality face video style transfer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and more particularly to a method and system for face video style transfer. Background Art

[0002] The face video style transfer technology aims to convert the input face video into a video with a target style while preserving the features and expressions of the original face. It has broad application prospects in fields such as film and television production, virtual reality, and social media.

[0003] Currently, traditional face video style transfer methods mainly include those based on generative adversarial networks (GANs) and recurrent neural networks (RNNs). By learning the mapping relationship between the input video and the target style, style transfer is achieved. Traditional encoders usually adopt a stacked network structure, consisting of a series of convolutional layers, which are used to gradually extract the features of the input image.

[0004] To enhance the nonlinear expression ability of the model and gradually compress the spatial size of the feature map, activation functions and pooling layers are usually also included in the encoder. Since traditional encoders use convolutional kernels of fixed size and treat all regions equally for feature extraction. This method cannot adapt to the feature requirements of different regions in the image, resulting in the inability to fully capture details in regions with rich details and complex textures, and overprocessing regions with simple textures; at the same time, its fixed downsampling method (pooling layer or strided convolution) is used to compress the size of the feature map, which will cause the loss of detail information, especially in image regions that require fine processing, such as the expressions and facial features of the face; furthermore, traditional encoders cannot capture global and local features simultaneously, and thus it is difficult to balance the overall style and detail performance of the image.

[0005] Due to the high computational complexity of the models constructed in existing face video style transfer methods, there are problems such as insufficient facial details in face video style transfer, deformation of the face structure, and discontinuity of video frames, and it is unable to fully preserve and reconstruct the facial details of the face, affecting the style consistency and content fidelity of face video transfer. Summary of the Invention

[0006] Aiming at the problems existing in the above fields, the present invention proposes a method and system for face video style transfer. By constructing a generator network model, the target style can be naturally applied to the face, generating a realistic stylized video and achieving high-quality face style transfer.

[0007] To solve the above technical problems, the present invention discloses a method for face video style transfer, including the following steps:

[0008] Obtain a face video frame image and decompose it to generate a face mask;

[0009] Construct a generator network model; the generator network model includes an encoder and multiple conditional reversible residual blocks; introduce a content-driven multi-scale convolution strategy and a spatio-temporal attention mechanism in the encoder, and introduce an affine transformation layer and a deformation protection layer in each conditional reversible residual block respectively;

[0010] Input the obtained face video frame image into the encoder to extract a feature sequence with multi-scale and spatio-temporal information; in the affine transformation layer, perform an affine transformation on the feature sequence, and fuse the feature sequence into a target style feature through a transformation function; in the deformation protection layer, constrain the features of the facial area of the target style feature according to the face mask, convert the target style feature into a facial area and a non-facial area, retain the original features of the facial features in the facial area, and perform style conversion on the non-facial area through the residual block in the conditional reversible residual block, and fuse the original features with the features of the style conversion to obtain the fused stylized face frame image features generated by the current conditional reversible residual block;

[0011] Input the fused stylized face frame image features generated by the current conditional reversible residual block into the next conditional reversible residual block in sequence and execute in a loop, and obtain the stylized face frame image comprehensively stylized by multiple conditional reversible residual blocks through decoding by the decoder.

[0012] Preferably, the fusion of the feature sequence into the target style feature through the transformation function specifically includes:

[0013] Input the feature sequence fused with multi-scale and spatio-temporal information output by the encoder into multiple conditional reversible residual blocks;

[0014] Input the input feature and evenly divide it into two parts according to the channel dimension ;

[0015] In the affine transformation layer, perform an affine transformation on the input feature and fuse the two parts of the feature into the target style feature through the transformation function;

[0016] Define the transformation function as:

[0017] ;

[0018] where f ( x 2 , c ) and g ( y 1 , c ) respectively represent two branches of the affine transformation, x1 and y 2 respectively represent the features of two groups of inputs; is the activation function ReLU; W f and W g are learnable weight matrices respectively; ⊙ represents element-wise multiplication, that is, each channel in the feature is scaled by the scaling parameters and respectively, and then the offset parameter or is added to the corresponding channel;

[0019] Branch f and branch g are complementary transformations. Branch f focuses on low-level texture injection, and branch g tends to high-level semantics. At the residual connection, the two are added or concatenated to obtain the synthesized stylized output;

[0020] Use the pre-trained style encoder VGG network to extract the style conditions of the style image ;

[0021] Calculate the affine transformation parameters through the extracted style conditions .

[0022] Preferably, the use of the pre-trained style encoder VGG network to extract the style conditions of the style image includes the following steps:

[0023] Through the VGG-19 network, extract features at different levels in multiple convolutional layers and multiple pooling layers, where:

[0024] The extraction feature process and description of multiple convolutional layers include:

[0025] In the Conv1_1 and Conv1_2 layers: capture the low-level features of the face, including edges, textures, and skin color details;

[0026] In the Conv2_2 and Conv2_2 layers: capture the intermediate-level features of the face, including identifying the basic shapes of the five facial features of eyes, nose, and mouth;

[0027] In the Conv3_1, Conv3_2, and Conv3_3 layers: capture the high-level features of the face, including details of the five facial features, expression features, and facial texture; In the Conv4_1 layer: capture the overall shape, pose, and spatial layout of the face;

[0028] Extract the feature map according to the output of each convolutional layer, where Denote the th convolutional layer;

[0029] For each feature map , calculate its corresponding Gram matrix , which is used to represent the style information of this convolutional layer;

[0030] Weightedly sum the Gram matrices of different convolutional layers according to the set weight coefficients to obtain the comprehensive style condition :

[0031] ;

[0032] Input the style condition into the conditional network to calculate the affine transformation parameters:

[0033] ;

[0034] ;

[0035] Among them, , are respectively the scaling parameters of a group of branches c related to this style output by the conditional network CondNet with the conditional feature f as the input, and the branch g ; , are the corresponding offset parameters.

[0036] Preferably, the obtaining of the fused stylized face frame image features generated by the current conditional reversible residual block specifically includes:

[0037] Constraining the features of the facial region of the target style feature according to the face mask:

[0038] ;

[0039] Among them, represents the face mask, represents the output of the encoder, represents being responsible for balancing the retention degree of the facial region features and the effect of style conversion, whose value is determined by the hyperparameter tuning method; y represents the output of the residual connection;

[0040] Represent the facial region with the mask being 1. For the facial region, retain the original features of the facial features;

[0041] The mask of 0 represents the non-face area. For the non-face area, the style conversion is freely performed through the residual block to generate the target style features;

[0042] The original features output by the encoder that retain the facial features in the facial area are fused with the target style features generated by the residual block to obtain the fused stylized face frame image features.

[0043] Preferably, the decoder includes multiple deconvolution layers, each of which combines the features of the encoder in a jump connection manner, gradually restores the spatial size of the features through multi-layer upsampling and convolution operations, maps the high-dimensional features back to the pixel space, and generates multiple conditionally reversible residual blocks to comprehensively stylize the face frame image.

[0044] Preferably, the generator network model also includes a discriminator, specifically including:

[0045] The generated multiple conditionally reversible residual blocks are used to synthesize the stylized face frame images and input them into the multi-task discriminator. Multiple discriminators and generators are trained in parallel. After one forward propagation, each discriminator calculates the discrimination loss and back-propagates it to the generator together with the corresponding weights to guide the generator optimization.

[0046] The discriminator includes a global discriminator, a local discriminator and a temporal discriminator, wherein:

[0047] Global Discriminator Determine whether the generated image is realistic and conforms to the target style;

[0048] Local Discriminator Focus on key facial areas including eyes and mouth, judging the authenticity and consistency of details;

[0049] Timing Discriminator Evaluate the temporal consistency and continuity between generated video frames;

[0050] By extracting the intermediate feature representation of each discriminator , and , and perform feature fusion:

[0051] ;

[0052] By fusion features Make the final judgment decision.

[0053] Preferably, the generator network model further comprises constructing a comprehensive loss function; the comprehensive loss function comprises adversarial loss, feature matching loss, temporal consistency loss, reconstruction loss and key point loss;

[0054] The adversarial loss, feature matching loss, temporal consistency loss, reconstruction loss, and keypoint loss are weighted to obtain the comprehensive loss function as follows:

[0055] ;

[0056] Among them, represents the weight coefficient of each loss term;

[0057] The calculation of the keypoint loss specifically includes:

[0058] By defining the keypoint loss and deformation protection loss, constraints are imposed on the parameter update of the generator:

[0059] Based on the face frame image after comprehensive stylization by multiple conditional reversible residual blocks and the original face frame image , use the pre-trained face keypoint detection model Dlip model to perform face keypoint detection, and obtain the keypoint coordinates and ;

[0060] Calculate the keypoint loss used to measure the difference between the generated image and the real image at the face keypoint positions:

[0061] ;

[0062] Through the edge detection algorithm, obtain the edge images of the face frame image after comprehensive stylization by multiple conditional reversible residual blocks and the original face frame image :

[0063] ;

[0064] Among them, E s represents the proportion of edge pixels in the total pixels, h and w represent the height and width of the edge image respectively.

[0065] Preferably, the generation of the face mask includes the following steps:

[0066] Decompose the input face video into a frame sequence, where represents the th frame;

[0067] Use the face detection algorithm MTCNN to detect the face region in each frame and generate the face mask , obtaining the face frame column .

[0068] Preferably, the output fuses the feature sequence of multi-scale and spatio-temporal information, including the following steps:

[0069] Perform key point detection on the face frame column to obtain the key point coordinates ;

[0070] Align the face according to the key point coordinates to obtain the aligned face frame by unifying the scale and pose ;

[0071] Input the obtained aligned face frame into the content-driven multi-scale convolutional network module in the encoder, perform convolutional operations on the input image, and obtain the feature maps corresponding to the scales;

[0072] Perform weighted fusion on the obtained multi-scale feature maps to obtain a comprehensive feature representation;

[0073] Input the obtained comprehensive feature representation into the global average pooling layer in the network module of the spatio-temporal attention mechanism to obtain the feature vector after dimensionality reduction;

[0074] Input the feature vector after dimensionality reduction into the fully connected layer to calculate the temporal attention weights;

[0075] Multiply the temporal attention weights by the corresponding comprehensive features to obtain the features after temporal weighting;

[0076] Perform convolutional operations on the features after temporal weighting through the region adaptive spatial attention module to obtain the spatial feature map;

[0077] Obtain the spatial attention weights, multiply the spatial attention weights by the spatial feature map to obtain the features after spatial weighting, and use it as the feature sequence that fuses multi-scale and spatio-temporal information output after encoding by the encoder .

[0078] Preferably, it further includes a face video style transfer system, including:

[0079] A video acquisition and preprocessing module, which is used to acquire face video frame images and decompose them to generate face masks;

[0080] A generator network model construction module, which is used to construct a generator network model; the generator network model includes an encoder and multiple conditional reversible residual blocks; a content-driven multi-scale convolutional strategy and a spatio-temporal attention mechanism are introduced into the encoder, and an affine transformation layer and a deformation protection layer are respectively introduced into each conditional reversible residual block;

[0081] The style transfer module is used to input the obtained face video frame images into the encoder to extract a feature sequence with multi-scale and spatio-temporal information; in the affine transformation layer, perform an affine transformation on the feature sequence, and fuse the feature sequence into the target style features through a transformation function; in the deformation protection layer, constrain the features of the facial area of the target style features according to the face mask, convert the target style features into the facial area and the non-facial area, retain the original features of the facial features in the facial area, perform style conversion on the non-facial area through the residual block in the conditional invertible residual block, and fuse the original features with the features after style conversion to obtain the fused stylized face frame image features generated by the current conditional invertible residual block; sequentially input the fused stylized face frame image features generated by the current conditional invertible residual block into the next conditional invertible residual block for cyclic execution, and obtain the face frame images comprehensively stylized by multiple conditional invertible residual blocks through decoding by the decoder.

[0082] Compared with the prior art, the present invention has the following beneficial effects:

[0083] The face video style transfer method proposed by the present invention can efficiently convert the input face video into a stylized video through the constructed generator network model, ensuring that the generated video performs excellently in terms of unified style, rich facial details, and smooth timing, and realizing high-quality face video style transfer. The generator set in the network model includes multiple conditional invertible residual blocks, which can convert the input face video sequence into a video with a specific style. The introduced learnable affine transformation layer performs an affine transformation on the input video features by determining the affine transformation parameters and fuses them into the target style features, making the calculation of the residual block controlled by the style features and enhancing the expression ability of the model; the introduced deformation protection layer constrains the facial features of the target style features fused by the affine transformation layer through the face mask, retains more facial feature information and moderately styles it, and reduces the deformation caused by style conversion. This method can naturally apply the target style to the face and generate a realistic stylized video. Description of the Drawings

[0084] Figure 1 It is a flowchart of the face video style transfer method proposed by the present invention;

[0085] Figure 2 It is the network model architecture constructed by the present invention;

[0086] Figure 3 It is a flowchart of the feature extraction by the encoder of the network model constructed by the present invention;

[0087] Figure 4 It is the architecture of the content-driven multi-scale convolutional feature extraction network in the encoder of the network model constructed by the present invention;

[0088] Figure 5 The network architecture based on the spatio-temporal attention mechanism in the network model encoder constructed for the present invention. Specific embodiments

[0089] Next, in combination with the accompanying drawings in the embodiments of the present invention Figures 1 - 5 , the technical solutions in the embodiments of the present invention will be clearly and completely described. It should be understood that the terms described in the present invention are only for describing specific embodiments and are not used to limit the present invention.

[0090] As Figure 1 shown, the present invention proposes a method for face video style transfer, including the following steps:

[0091] S1: Obtain a face video frame image and decompose it to generate a face mask;

[0092] S2: Construct a generator network model; the generator network model includes an encoder and a plurality of conditional reversible residual blocks; a content-driven multi-scale convolution strategy and a spatio-temporal attention mechanism are introduced in the encoder, and an affine transformation layer and a deformation protection layer are respectively introduced in each conditional reversible residual block;

[0093] S3: Input the obtained face video frame image into the encoder, dynamically adjust the scale of the convolution kernel through the multi-scale convolution strategy, capture the temporal dependence between image frames through the temporal attention mechanism, and output a feature sequence that combines multi-scale and spatio-temporal information; in the affine transformation layer, perform an affine transformation on the feature sequence, and fuse the feature sequence into the target style features through a transformation function; in the deformation protection layer, constrain the features of the facial area of the target style features according to the face mask, convert the style into the facial area and the non-facial area, retain the original features of the facial features in the facial area, perform style conversion on the non-facial area through the residual block, and fuse the original features with the features of the style conversion to obtain the fused stylized face frame image features generated by the current conditional reversible residual block;

[0094] S4: Sequentially input the fused stylized face frame image features generated by the current conditional reversible residual block into the next conditional reversible residual block and execute in a loop, and obtain the face frame image comprehensively stylized by a plurality of conditional reversible residual blocks through decoding by a decoder.

[0095] Specifically, in step S1, generating a face mask includes the following steps:

[0096] Decompose the input face video into a frame sequence, where represents the th frame;

[0097] Use the face detection algorithm MTCNN to detect the face area in each frame and generate a face mask , obtain a sequence of face frames .

[0098] Face key point detection and alignment:

[0099] Key point detection is used to identify the key positions of the face, ensuring that facial features can be processed standardly under different poses; alignment ensures that the faces in each frame are in the same standard space, which is crucial for video continuity and consistency.

[0100] For each face frame perform key point detection to obtain the key point coordinates ;

[0101] Align the face according to the key point coordinates, unify the scale and pose, and obtain the aligned face frame ;

[0102] Input the obtained aligned face frame into the content-driven multi-scale convolutional network module in the encoder, perform convolutional operations on the input image, and obtain the feature maps of corresponding scales;

[0103] Perform weighted fusion on the obtained multi-scale feature maps to obtain a comprehensive feature representation;

[0104] Input the obtained comprehensive feature representation into the global average pooling layer in the network module of the spatio-temporal attention mechanism to obtain the feature vector after dimensionality reduction;

[0105] Input the feature vector after dimensionality reduction into the fully connected layer, and calculate the temporal attention weights;

[0106] Multiply the temporal attention weights by the corresponding comprehensive features to obtain the features after temporal weighting;

[0107] Perform convolutional operations on the features after temporal weighting through the region adaptive spatial attention module to obtain the spatial feature map;

[0108] Obtain the spatial attention weights, multiply the spatial attention weights by the spatial feature map, obtain the features after spatial weighting, and use it as the feature sequence that fuses multi-scale and spatio-temporal information output after encoding by the encoder .

[0109] Image normalization processing:

[0110] Perform pixel normalization processing on the aligned face frames, and scale the pixel values to the range of [-1, 1].

[0111] Traditional encoders usually adopt a stacked network structure, consisting of a series of convolutional layers, such as Figure 3As shown, these convolutional layers are used to gradually extract the features of the input image. To enhance the non-linear expression ability of the model and gradually compress the spatial size of the feature map, activation functions and pooling layers are usually included in the encoder.

[0112] However, since the traditional encoder uses convolutional kernels of a fixed size and treats all regions equally for feature extraction. This method cannot adapt to the feature requirements of different regions in the image, resulting in the inability to fully capture details in regions with rich details and complex textures, and over-processing of regions with simple textures; at the same time, its fixed downsampling method (pooling layer or strided convolution) is used to compress the size of the feature map, which will cause the loss of detail information, especially in image regions that require fine processing, such as the expressions and facial features of a face; furthermore, the traditional encoder cannot capture global and local features simultaneously, and thus it is difficult to balance the overall style and detail performance of the image.

[0113] In step S2, in view of the above, the present invention designs an encoder based on content-driven multi-scale convolutional feature extraction and fusion of spatio-temporal attention mechanisms to improve the effect of face video style transfer. The specific network structure of this encoder includes:

[0114] A content-driven multi-scale convolutional feature extraction network, as Figure 4 shown; a network based on spatio-temporal attention mechanisms, as Figure 5 shown.

[0115] In step S2, the process of feature extraction by the encoder proposed by the present invention is as follows:

[0116] Step1: Calculate the scale selection weight ;

[0117] For each aligned face frame , calculate its texture complexity :

[0118] Through operator to calculate the gradient magnitude :

[0119] ;

[0120] By calculating the sum of squares of the gradient magnitude within the local region of the image, the texture complexity is obtained:

[0121] ;

[0122] Normalize the texture complexity to the interval.

[0123] Calculate edge intensity ;

[0124] Use the edge detection algorithm to obtain the edge image , and by calculating the proportion of edge pixels in the total pixels, the following can be obtained :

[0125] ;

[0126] Weightedly sum the texture complexity and the edge intensity to obtain the comprehensive calculation scale selection weight :

[0127] ;

[0128] Among them, and are the weight systems, and are obtained through methods such as grid search algorithm or cross-validation to test the overall effect of style transfer, such as facial detail retention and style consistency, and finally select the optimal combination to meet . The test results show that when and sum to 1 and are both non-zero, the model can better balance both "high texture details" and "high edge information".

[0129] Calculate the weight :

[0130] ;

[0131] Among them, is the central value corresponding to the th convolution kernel scale( ), is the adjustment parameter used to control the steepness of the weight distribution. Combining with actual face video data for experiments, it is found that when takes a value near 10, it can balance the detail retention and computational efficiency while distinguishing different scale convolution kernels. When is too small, the weights of the multi-scale convolution branches tend to be averaged, losing the "adaptive" effect; when is too large, the model is prone to over-bias towards a certain scale at the initial stage of training, resulting in unstable convergence. Therefore, in the present invention takes 10, is the number of convolution kernel scales.

[0132] In the initial stage of the present invention, with a relatively wide range of hyperparameters( ∈ {1, 5, 10, 15}, ∈ {0.1, 0.3, 0.5}, = 1 - ) Perform a grid search, using the average structural similarity (SSIM), style consistency score, and temporal smoothness of the model on the validation set as the measurement metrics. The final results show that ≈ 0.4, ≈ 0.6, The combination of = 10 has the best comprehensive performance in terms of facial detail fidelity and overall style smoothness.

[0133] , and The values of and are not fixed. In different application scenarios, such as higher-resolution videos or different types of face textures, they can be appropriately adjusted according to actual needs and experimental results , and of to achieve the best migration effect.

[0134] Step2: Input the aligned face frames into the content-driven multi-scale convolutional network module in the encoder, perform convolutional operations on the input image, and obtain the feature maps corresponding to the corresponding scales;

[0135] Prepare a set of convolutional kernels with different scales, For each scale of convolutional kernel, initialize the corresponding convolutional layer .

[0136] Then set the size of the aligned face frame image to , and at the same time input it into the above 3 convolutional branches.

[0137] Each convolutional branch uses the corresponding convolutional kernel to perform convolutional operations on the input image, and obtains the feature maps corresponding to the corresponding scales, with the size of .

[0138] After being processed by the activation function ReLu, 3 feature maps with different scales are obtained.

[0139] According to the weights obtained in Step1, perform weighted fusion on the 3 feature maps with different scales to obtain the comprehensive feature representation , with the size of .

[0140] Step3: Input the above comprehensive feature representation into the network module based on the spatio-temporal attention mechanism;

[0141] Input the comprehensive feature representation obtained in Step 2 , into the global average pooling layer in the temporal attention module, and obtain the feature vector after dimensionality reduction :

[0142] ;

[0143] Input the feature vector into the fully connected layer, so as to map it to a scalar , which is used to represent the temporal score of the current frame;

[0144] Apply the Softmax function to all to obtain the temporal attention weights , and satisfy , and:

[0145] ;

[0146] ;

[0147] Among them, , are all learnable parameters in the fully connected layer.

[0148] Multiply by the corresponding comprehensive feature , and obtain the temporally weighted feature .

[0149] Input the obtained temporally weighted feature into the region adaptive spatial attention module, that is, based on face key point detection, locate key regions such as eyes and mouth, and assign higher spatial attention weights to improve the detail performance of these regions.

[0150] Perform a convolution operation on to obtain the spatial feature map to fuse the region mask:

[0151] ;

[0152] Among them, is the region mask corresponding to the video frame image, is the control parameter.

[0153] Through map the obtained spatial feature map to obtain the spatial attention weights ;

[0154] Multiply element-wise by , and obtain the spatially weighted feature 。

[0155] Step4: Output the encoded feature sequence 。

[0156] In a generative adversarial network (GAN), the role of the generator is to receive input data and generate a target output through network mapping. In the task of face video style transfer, the generator needs to convert the input face video sequence into a video with a specific style while maintaining the structure and details of the face.

[0157] Traditional generator networks usually adopt an encoder-decoder. The overall architecture consists of an encoder and a decoder. The input image first passes through the encoder to extract multi-level features of the input image and map them to a low-dimensional latent space representation. The decoder includes multiple transposed convolutional layers. Subsequently, these latent features are reconstructed through the transposed convolutional layers of the decoder to generate a high-resolution output image.

[0158] Currently, to address the possible loss of detailed information during downsampling and upsampling, skip connections are introduced between the encoder and the decoder. These connections directly transfer feature maps between corresponding layers, enabling the decoder to obtain the high-resolution features of the encoder and enhancing the effects of detail retention and feature fusion. However, even so, there are still problems such as the inability to fully retain and reconstruct facial details, the inability to simultaneously meet style consistency and content fidelity, deformation of the face structure, and high model complexity.

[0159] Therefore, the present invention improves the network architecture of the generator, specifically manifested in: when constructing the network model, a conditional reversible residual block is designed in the generator, and a learnable affine transformation layer and a deformation protection module are introduced into the conditional reversible residual block, which can effectively solve the above existing problems.

[0160] As Figure 2 shown, it is the overall network architecture of the generator designed by the present invention. The specific network structure of the designed generator includes: an encoder, multiple conditional reversible residual blocks, and a decoder.

[0161] The composition of each conditional reversible residual block includes: an affine transformation layer and a deformation protection layer.

[0162] The present invention introduces a "learnable affine transformation layer" into the residual block, and linearly maps the input features through and these two learnable parameters, which can not only flexibly inject target style features but also retain a certain amount of original content.

[0163] The process of generating high-quality style-transferred video frame images through multiple conditionally reversible residual blocks of the generator is as follows:

[0164] Step 1: Encode through the encoder to obtain the encoded feature sequence ;

[0165] Step 2: Sequence the features Input into multiple conditional residual blocks, and gradually transfer the style to the target style in each residual block;

[0166] Compared with the traditional one-time overall affine, the present invention can control the fusion of style and content in a finer granularity through a two-branch structure, reducing problems such as excessive local facial deformation or style loss caused by a "one-size-fits-all" approach.

[0167] The input features Divide into two parts evenly according to the channel dimension , corresponding to different semantic channels or spatial information respectively. This segmentation enables the network to perform richer transformations within the residual block, and then splice and fuse them at the end to achieve more flexible style injection.

[0168] The input features Perform an affine transformation and transform Fusion target style features, the transformation function is defined as:

[0169] ;

[0170] in, f ( x 2 , c )and g ( y 1 , c ) represent the two branches of affine transformation, x 1 and y 2 Represent the characteristics of the two sets of inputs respectively; is the activation function ReLU; W f and W g are respectively learnable weight matrices; ⊙ represents element-by-element multiplication, that is, each channel in the feature is multiplied by the scaling parameter and Scaling is performed and then the offset parameters are or Add to the corresponding channel.

[0171] In this way, f (x 2 , c ) and g ( y 1 , c ) can both adjust the input features according to the style conditions c so that the output features better match the target style (such as the color, texture, and style atmosphere of the face).

[0172] The obtained f ( x 2 , c ) and g ( y 1 , c ) will merge the output features in the branch y , and further fuse with the input features x through the residual connection. At this time, to prevent excessive deformation of the key areas of the face, a deformation protection layer is introduced, and interpolation fusion between the original features and the style features is performed in the facial area with the help of the face mask M.

[0173] Branch f and branch g can be regarded as complementary transformations in function. Branch f focuses on low-level texture injection, and branch g is more inclined to high-level semantics; finally, the two are added or concatenated at the residual connection to obtain a comprehensive stylized output.

[0174] Use the pre-trained style encoder VGG network to extract the style conditions of the style image ;

[0175] Through the extracted style conditions , calculate the affine transformation parameters.

[0176] Through the VGG-19 network, extract features at different levels in multiple convolutional layers and multiple pooling layers, where:

[0177] The process and description of feature extraction in multiple convolutional layers include:

[0178] In the Conv1_1 and Conv1_2 layers: capture the low-level features of the face, including edges, textures, and skin color details;

[0179] In the Conv2_2 and Conv2_2 layers: capture the intermediate-level features of the face, including identifying the basic shapes of the five facial features of eyes, nose, and mouth;

[0180] In the Conv3_1, Conv3_2, and Conv3_3 layers: capture the high-level features of the face, including the details of facial features, expression features, and facial texture; in the Conv4_1 layer: capture the overall shape, pose, and spatial layout of the face;

[0181] Extract feature maps based on the outputs of each convolutional layer , where represents the th convolutional layer;

[0182] For each feature map , calculate its corresponding Gram matrix , which is used to represent the style information of this convolutional layer;

[0183] Weightedly sum the Gram matrices of different convolutional layers according to the set weight coefficients to obtain the comprehensive style condition :

[0184] ;

[0185] Input the style condition into the conditional network to calculate the affine transformation parameters:

[0186] ;

[0187] ;

[0188] The above formula represents a small fully-connected network called the "conditional network" (Conditon Network, denoted as CondNet), which takes the conditional feature c as input and outputs a set of scaling parameters , and offset parameters , related to this style.

[0189] These scaling parameters and offset parameters can be understood as "learnable style adjustment coefficients". When the input conditional feature c is different, this network will automatically output different scaling parameters and offset parameters for stylizing and adjusting the input features.

[0190] Traditional affine layers are mostly applied to overall image style transfer. However, in the present invention, there is a deformation protection strategy for the facial area, which divides the style into the facial area and the non-facial area. The original features of the facial features are retained in the facial area, and style transfer is performed on the non-facial area through residual blocks, and the original features are fused with the features of the style transfer to obtain the stylized face frame image generated by the current conditional reversible residual block, so that the facial features structure remains faithful after style transfer.

[0191] After the residual connection, the features of the facial area of the target style features are constrained according to the face mask, and the output of the residual connection is passed through the deformation protection layer to perform the operation of the following formula:

[0192] ;

[0193] where, represents the face mask; represents the output of the encoder; is responsible for balancing the retention degree of the facial area features and the effect of style transfer, and is determined by the hyperparameter tuning method, y represents the output of the residual connection.

[0194] The mask of 1 represents the facial area. For the facial area, the original features of the facial features are retained;

[0195] The mask of 0 represents the non-facial area. For the non-facial area, style transfer is freely performed through residual blocks to generate target style features;

[0196] And the original features of the facial features retained in the facial area output by the encoder are fused with the target style features generated by the residual blocks to obtain the features of the fused stylized face frame image.

[0197] In this embodiment, is set to 0.4. Specifically, the value of

[0198] is obtained by testing the influence of different values on the facial distortion degree and visual style effect in the validation set, which can not only avoid obvious distortion of facial features, but also inject enough style features into the facial area.

[0199] At the same time, in each conditional reversible residual block, a deformation protection layer is also added. This layer constrains the features of the facial area and restricts the generator from over-modifying the facial area. Specifically:

[0200] Through this operation, it is ensured that for the facial region (the part with a mask value of 1), more of the original features of the facial features are retained and moderately stylized, reducing the deformation caused by style conversion.

[0201] For the non-facial region (the part with a mask value of 0), the residual block is allowed to freely perform style conversion.

[0202] In this way, the generator is constrained by the deformation protection layer during the generation process, avoiding excessive deformation of the facial region, and ensuring the structural consistency of the face and the integrity of facial features.

[0203] The is input into the next residual block, and the above operations are repeated.

[0204] The generator network model constructed by the present invention includes multiple conditional reversible residual blocks. Each conditional reversible residual block utilizes the idea of a reversible network to reduce model parameters and memory consumption. In the residual block, effective fusion of styles is achieved by applying style features to perform affine transformation on intermediate features.

[0205] Step3: The features processed by multiple conditional reversible residual blocks are input into the decoder, and are gradually upsampled to restore to the resolution of the original image, and the number of feature channels is reduced to 3.

[0206] The role of the decoder is to map the high-dimensional features back to the pixel space to generate the final stylized face image. The decoder gradually restores the spatial size of the features through multiple upsampling and convolution operations. At the same time, the decoder can combine the features of the encoder (through skip connection) to enhance the details and quality of the generated image.

[0207] The generator network model also includes a discriminator, specifically including:

[0208] The comprehensively stylized face frame images generated by multiple conditional reversible residual blocks are input into the multi-task discriminator. Through parallel adversarial training between multiple discriminators and the generator, each discriminator calculates the discriminant loss respectively after one forward propagation, and together with the corresponding weights, backpropagates to the generator to guide the generator to optimize;

[0209] The discriminator includes a global discriminator, a local discriminator, and a temporal discriminator, where:

[0210] Global discriminator Judges whether the generated image is real and whether it conforms to the target style;

[0211] Local discriminator Focuses on key facial regions including eyes and mouth, and judges the authenticity and consistency of details;

[0212] Temporal discriminator Evaluate the temporal consistency and continuity between the generated video frames;

[0213] By extracting the intermediate feature representations of each discriminator 、 and ,and perform feature fusion:

[0214] ;

[0215] Through the fused features Make the final discriminative decision to improve the discriminative ability and collaborative effect of the discriminator.

[0216] The construction of the network model in step S2 also includes: constructing a comprehensive loss function

[0217] The comprehensive loss function includes adversarial loss, feature matching loss, temporal consistency loss, reconstruction loss, and key point loss, where:

[0218] Adversarial loss ;

[0219] Function: Make the images generated by the generator (Generator) as realistic as possible to deceive the discriminator (Discriminator), and be used to drive the generator to learn the data distribution, so as to generate high-quality images.

[0220] For the discriminator :

[0221] ;

[0222] For the generator :

[0223] ;

[0224] Where: Represents the real image (i.e., the face image of the target style); Represents the image generated by the generator according to the input feature ; Represents the edge image of the image generated by the generator according to the input feature ; Represents the output of the discriminator for the input image, with a value between [0,1], estimating the probability that the input image is real.

[0225] Feature matching loss ;

[0226] Function: Make the generated image match the real image in the high-level feature space, so as to improve the quality and stability of the generated image.

[0227] ;

[0228] Among them, represents the edge image between the real image and the image generated by the generator according to the input features ; represents the feature map extracted by the discriminator at the layer; is the number of discriminator layers used to calculate the feature matching loss; is the number of elements of the layer features; represents norm, that is, the sum of the absolute values of the feature differences is taken.

[0229] By minimizing the difference between the generated image and the real image in the discriminator feature space, the generator is guided to generate high-level features similar to the real image, thereby improving the structural consistency and visual quality of the image.

[0230] Temporal consistency loss ;

[0231] Function: The temporal consistency loss is used to ensure smooth transitions between generated video frames, reduce jitter and discontinuities, and improve the temporal consistency of the video.

[0232] ;

[0233] Among them, is the th frame image generated; is the th frame image generated; is the optical flow estimation function from the th frame to the th frame, representing the movement of pixels.

[0234] Optical flow estimation captures the pixel movement information between frames, transforms the previous frame image according to the optical flow, and obtains the predicted current frame. By minimizing the difference between the current frame generated image and the predicted image, the coherence of the content between frames is ensured.

[0235] Reconstruction loss ;

[0236] Function: The reconstruction loss is used to ensure that the generated image is close to the real image in the pixel space and retain the details of the original content.

[0237] ;

[0238] Among them, is the real target image; It is an image generated by the generator.

[0239] By minimizing the difference between the generated image and the real image at the pixel level, the generator is encouraged to retain more detailed information and generate results closer to the real image.

[0240] Key point loss ;

[0241] Function: To maintain the integrity of the face structure.

[0242] The calculation of the key point loss specifically includes:

[0243] By defining the key point loss and the deformation protection loss, constraints are imposed on the parameter update of the generator:

[0244] According to the face frame image after comprehensive stylization by multiple conditional reversible residual blocks and the original face frame image , use the pre-trained face key point detection model Dlip model to perform face key point detection, and obtain the key point coordinates respectively and ;

[0245] Calculate the key point loss used to measure the difference between the generated image and the real image at the face key point positions:

[0246] ;

[0247] Through The edge images of the face frame image after comprehensive stylization by multiple conditional reversible residual blocks and the original face frame image obtained by the edge detection algorithm :

[0248] ;

[0249] Among them, E s represents the proportion of edge pixels to the total pixels, h and w represent the height and width of the edge image respectively.

[0250] Weight the adversarial loss, feature matching loss, temporal consistency loss, reconstruction loss, and key point loss to obtain the comprehensive loss function as:

[0251] ;

[0252] Among them, are the weight coefficients of each loss term, used to balance the influence of different losses on model training.

[0253] Finally, the trained generator network model can convert the input face video into a stylized face video. The stylized frame sequence output by the generator is combined in chronological order to form a coherent video. The generated video has significant improvements in style consistency, facial detail retention, and temporal coherence, achieving high-quality face video style transfer.

[0254] The model is optimized through a comprehensive loss function, including adversarial loss, feature matching loss, temporal consistency loss, reconstruction loss, and keypoint loss, etc. The generator adjusts its own parameters according to the feedback of the discriminator and the guidance of the loss function, gradually improving the quality and authenticity of the generated results.

[0255] The entire model adopts an adversarial training strategy of a generator and a discriminator. Multiple discriminators and the generator perform parallel adversarial training. After each discriminator's forward propagation, it calculates its own discriminant loss and backpropagates it to the generator together with the corresponding weights to guide the optimization of the generator.

[0256] This application also proposes a face video style transfer system, including:

[0257] A video acquisition and preprocessing module, used to acquire face video frame images and decompose them to generate face masks;

[0258] An improved generator network model construction module, used to construct a generator network model; the generator network model includes an encoder and multiple conditional reversible residual blocks; a content-driven multi-scale convolution strategy and a spatio-temporal attention mechanism are introduced in the encoder, and an affine transformation layer and a deformation protection layer are respectively introduced in each conditional reversible residual block;

[0259] A style transfer module, used to input the acquired face video frame images into the encoder to extract a feature sequence of multi-scale and spatio-temporal information; in the affine transformation layer, perform an affine transformation on the feature sequence, and fuse the feature sequence into target style features through a transformation function; in the deformation protection layer, constrain the features of the facial area of the target style features according to the face mask, convert the target style features into facial areas and non-facial areas, retain the original features of the facial features in the facial area, perform style conversion on the non-facial area through the residual blocks in the conditional reversible residual blocks, and fuse the original features with the style-converted features to obtain the fused stylized face frame image features generated by the current conditional reversible residual block; sequentially input the fused stylized face frame image features generated by the current conditional reversible residual block into the next conditional reversible residual block for cyclic execution, and obtain the comprehensively stylized face frame images of multiple conditional reversible residual blocks through decoding by the decoder.

[0260] The method proposed by the present invention can efficiently convert the input face video into a stylized video, ensuring that the generated video performs excellently in terms of unified style, rich facial details, and smooth temporal sequence, achieving high-quality face video style transfer.

[0261] Embodiment

[0262] In this embodiment, a face video style transfer method based on deep learning is taken as an example to verify the method proposed by the present invention.

[0263] Feature extraction:

[0264] The face frames obtained after preprocessing are input into the encoder. The encoder, through a content-driven multi-scale convolutional feature extraction network, dynamically adjusts the scale of the convolutional kernel according to the texture complexity and edge information of the image to capture features at different levels. At the same time, through a network based on a spatio-temporal attention mechanism, a spatio-temporal attention mechanism is introduced: temporal attention is used to capture the temporal dependence between frames, emphasizing the importance of key frames; spatial attention is based on face key points, focusing on key facial regions such as eyes and mouth, and enhancing the feature expression of these regions. The encoder finally outputs a feature sequence that fuses multi-scale and spatio-temporal information .

[0265] The feature sequence output by the encoder is input into the conditional invertible residual generator. At the same time, the affine transformation parameters of the style features are input into the affine transformation layer of the conditional invertible residual block of the generator for fusing style information during the generation process.

[0266] Specifically, the style conditions of the face style video are extracted , and the affine transformation parameters are calculated:

[0267] The style features of the style image are extracted using a pre-trained style encoder (VGG network) . .

[0268] The VGG-19 network is selected to extract features at different levels (including 19 convolutional layers and 5 pooling layers, and only 8 of the 19 convolutional layers are applied in the present invention)

[0269] The features extracted by 5 convolutional layers and their descriptions are as follows:

[0270] Conv1_1 and Conv1_2: Capture low-level features of the face: edges, textures, and skin color details.

[0271] Conv2_2 and Conv2_2: Extract intermediate-level features of the face: identify the basic shapes of facial features such as eyes, nose, and mouth.

[0272] Conv3_1, Conv3_2, and Conv3_3: Capture high-level features of the face: details of facial features, expression features, and facial texture.

[0273] Conv4_1: Capture the overall shape, pose, and spatial layout of the face, fuse facial features at different levels, and obtain style features. The detailed steps are as follows:

[0274] Step1. Extract feature maps from the outputs of the above 8 convolutional layers. where represents the th convolutional layer.

[0275] Step2. Calculate the Gram matrix:

[0276] For each feature map , calculate its Gram matrix to represent the style information of this layer. .

[0277] Step3. Weightedly sum the Gram matrices of different layers according to the set weight coefficients to obtain a comprehensive style feature representation :

[0278] ;

[0279] According to the importance of features at different levels, the values of the weight coefficients are as follows:

[0280] , , , , , , .

[0281] Step4. Input the style condition into the conditional network to calculate the affine transformation parameters:

[0282] ;

[0283] ;

[0284] The generator consists of multiple conditional invertible residual blocks. Each block uses the idea of an invertible network to reduce model parameters and memory consumption. In the residual block, the style features are used to perform an affine transformation on the intermediate features to achieve effective style fusion.

[0285] The feature sequence processed by the conditional invertible residual block is input into the decoder. The role of the decoder is to map the high-dimensional features back to the pixel space to generate the final stylized face image. The decoder gradually restores the spatial dimensions of the features through multiple layers of upsampling and convolution operations. At the same time, the decoder can combine the features of the encoder (through skip connections) to enhance the details and quality of the generated image.

[0286] At the same time, in each conditional invertible residual block, a deformation protection layer is also added. This layer constrains the features in the facial area and restricts the generator from overly modifying the facial area.

[0287] Specifically, after the output of each residual block fuses the features, the original features of the encoder and the generated features of the residual block are fused according to the face mask:

[0288] For the facial area (the part with a mask of 1), more key structural information of the facial features is retained and moderately stylized to reduce the deformation caused by the style conversion. For the non-facial area (the part with a mask of 0), the residual block is allowed to freely perform style conversion.

[0289] In this way, the generator is constrained by the deformation protection layer during the generation process, avoiding excessive deformation of the facial area and ensuring the structural consistency of the face and the integrity of the facial features.

[0290] During the training process, a face structure preservation module is introduced. By defining the key point loss and the deformation protection loss, constraints are imposed on the parameter update of the generator. The specific process is as described in the improved generator network model above.

[0291] Finally, the trained model can convert the input face video into a stylized face video. The stylized frame sequence output by the generator is combined in chronological order to form a coherent video. The generated video has significant improvements in terms of style consistency, facial detail retention, and temporal coherence, achieving high-quality face video style transfer.

[0292] In summary, the method proposed in the present invention first performs data preprocessing on the input video, decomposes it into single-frame images, and standardizes the face area of each frame through face detection and key point alignment technology. This step ensures the consistency and accuracy in subsequent processing.

[0293] Extract the multi-level features of the style image through the pre-trained VGG-19 network, fuse the style information of different levels by calculating the Gram matrix, and generate a comprehensive style condition.

[0294] The preprocessed face frames are input into the encoder. The encoder adopts multi-scale convolution and spatio-temporal attention mechanism to extract a feature sequence that fuses multi-level and spatio-temporal information. These feature sequences are passed together with the calculated style affine transformation parameters into the conditional invertible residual blocks of the generator. The generator effectively fuses the style information into the face features through multiple residual blocks, and through the deformation protection layer, ensures that the integrity of the facial structure is not overly modified.

[0295] The stylized features output by the generator are gradually restored to high-quality stylized face images through the decoder. To ensure the coherence and realism of the generated video, the generated frame sequence is fed into a multi-task discriminator. The discriminator consists of three parts: global, local and temporal, which evaluate the overall authenticity of the image, the details of the key facial regions, and the temporal consistency between frames respectively.

[0296] During the training process, the generator and the discriminator are optimized with each other through an adversarial training strategy. The generator is committed to generating high-quality images that can deceive the discriminator, while the discriminator continuously improves its ability to distinguish between generated images and real images. At the same time, the comprehensive loss function, including adversarial loss, feature matching loss, temporal consistency loss, reconstruction loss and key point loss, guides the generator to improve the visual quality and temporal coherence of the image while maintaining the consistency of the face structure.

[0297] Finally, the model optimized by the above adversarial training can efficiently convert the input face video into a stylized video, ensuring that the generated video performs excellently in terms of unified style, rich facial details and smooth temporal sequence, achieving high-quality face video style transfer.

[0298] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered within the protection scope of the present invention.

[0299] In addition, unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the art to which the present invention belongs. All documents mentioned in this specification are incorporated by reference to disclose and describe the methods related to the said documents. In case of conflict with any incorporated document, the content of this specification shall prevail.

Claims

1. A face video style transfer method, characterized in that: The following steps are involved: Obtain a face video frame image and decompose it to generate a face mask; A generator network model is constructed; the generator network model includes an encoder and multiple conditional reversible residual blocks; a content-driven multi-scale convolution strategy and a spatiotemporal attention mechanism are introduced into the encoder, and an affine transformation layer and a deformation protection layer are introduced into each conditional reversible residual block; The acquired face video frame image is input into the encoder to extract the feature sequence of multi-scale and spatiotemporal information; in the affine transformation layer, the feature sequence is affine transformed, and the feature sequence is fused into the target style feature through the transformation function; in the deformation protection layer, the features of the facial area of ​​the target style feature are constrained according to the face mask, and the target style feature is converted into the facial area and the non-facial area, and the original features of the facial features are retained in the facial area. In the non-facial area, the residual block in the conditionally reversible residual block is used to perform style conversion, and the original features are fused with the style conversion features to obtain the fused stylized face frame image features generated by the current conditionally reversible residual block; The fused stylized face frame image features generated by the current conditional reversible residual block are sequentially input into the next conditional reversible residual block for loop execution, and the face frame image after comprehensive stylization of multiple conditional reversible residual blocks is obtained through decoding by the decoder; The method of fusing the feature sequence into the target style feature through the transformation function specifically includes: The feature sequence output by the encoder combines multi-scale and spatiotemporal information Input into multiple conditionally reversible residual blocks; Split the input feature x into two parts x=[x1,x2] evenly according to the channel dimension; In the affine transformation layer, the input feature x is affine transformed, and the two parts of the feature are fused into the target style feature through the transformation function; The transformation function is defined as: f(x2,c)=σ((W f x2)⊙γ f (c)+β f (c)) g(y1,c)=σ((W g y1)⊙γ g (c)+β g (c)); Among them, f(x2, c) and g(y1, c) represent the two branches of affine transformation, x1 and y2 represent the features of two sets of inputs, σ(·) is the activation function ReLU; W f and W g are respectively learnable weight matrices; ⊙ represents element-by-element multiplication, that is, each channel in the feature is multiplied by the scaling parameter γ f (c) and γ g (c) is scaled and then the offset parameter β is f (c) or β g (c) added to the corresponding channel; Branch f and branch g are complementary transformations. Branch f focuses on low-level texture injection, while branch g is biased towards high-level semantics. The two are added or concatenated at the residual connection to obtain a comprehensive stylized output. Use the pre-trained style encoder VGG network to extract the style condition c of the style image; Using the extracted style condition c, affine transformation parameters are calculated.

2. The face video style transfer method according to claim 1, characterized in that: The method of using the pre-trained style encoder VGG network to extract the style condition c of the style image includes the following steps: Through the VGG-19 network, features at different levels in multiple convolutional layers and multiple pooling layers are extracted, including: The feature extraction process of multiple convolutional layers and its description include: In Conv1_1 and Conv1_2 layers: capture low-level features of the face, including edges, textures, and skin color details; In Conv2_2 and Conv2_2 layers: capture intermediate features of the face, including identifying the basic shapes of eyes, nose, and mouth; In Conv3_1, Conv3_2 and Conv3_3 layers: capture high-level features of the face, including details of facial features, expression characteristics and facial texture; In the Conv4_1 layer: the overall shape, pose and spatial layout of the face are captured; According to the output of each convolutional layer, extract the feature map F l , where l represents the lth convolutional layer; For each feature map F l , calculate its corresponding Gram matrix G l =F l F l T , used to represent the style information of the convolutional layer; The Gram matrices of different convolutional layers are transformed according to the set weight coefficient α l Perform weighted summation to obtain the comprehensive style condition c: Input the style condition c into the conditional network and calculate the affine transformation parameters: Among them, γ f (c), γ g (c) The conditional network CondNet takes the conditional feature c as input and outputs a set of scaling parameters of the branch f and branch g related to the style; β f (c), β g (c) is the corresponding offset parameter.

3. The face video style transfer method according to claim 2, characterized in that: The step of obtaining the fused stylized face frame image features generated by the current conditionally reversible residual block specifically includes: Constrain the features of the facial region of the target style feature based on the face mask: y protected =M⊙(λ·x)+(1-M)☉y; Where M represents the face mask, x represents the output of the encoder, λ represents the balance between the degree of retention of face region features and the effect of style conversion, and the value of λ is determined by the hyperparameter tuning method; y represents the output of the residual connection; The mask of 1 represents the facial area, and for the facial area, the original features of the facial features are retained; The mask of 0 represents the non-face area. For the non-face area, the style conversion is freely performed through the residual block to generate the target style features; The original features output by the encoder that retain the facial features in the facial area are fused with the target style features generated by the residual block to obtain the fused stylized face frame image features.

4. The face video style transfer method according to claim 3, characterized in that: The decoder includes multiple deconvolution layers, each of which combines the features of the encoder in a jump connection manner, gradually restores the spatial size of the features through multi-layer upsampling and convolution operations, maps the high-dimensional features back to the pixel space, and generates a face frame image after comprehensive stylization by multiple conditionally reversible residual blocks.

5. The method for face video style transfer according to claim 4, characterized in that: The generator network model also includes a discriminator, specifically including: The generated multiple conditionally reversible residual blocks are used to synthesize the stylized face frame images and input them into the multi-task discriminator. Multiple discriminators and generators are trained in parallel. After one forward propagation, each discriminator calculates the discrimination loss and back-propagates it to the generator together with the corresponding weights to guide the generator optimization. The discriminator includes a global discriminator, a local discriminator and a temporal discriminator, wherein: Global Discriminator D global Determine whether the generated image is realistic and conforms to the target style; Local Discriminator D local Focus on key facial areas including eyes and mouth, judging the authenticity and consistency of details; Timing Discriminator D temporal Evaluate the temporal consistency and continuity between generated video frames; By extracting the intermediate feature representation F of each discriminator global 、F local and F temporal , and perform feature fusion: F fusion =Concat(F global ,F local ,F temporal ); By fusion feature F fusion Make the final judgment decision.

6. The method for face video style transfer according to claim 5, characterized in that: The generator network model also includes constructing a comprehensive loss function; the comprehensive loss function includes adversarial loss, feature matching loss, temporal consistency loss, reconstruction loss and key point loss; The adversarial loss, feature matching loss, temporal consistency loss, reconstruction loss and key point loss are weighted to obtain the comprehensive loss function: Among them, λ GAN , FM , TC , Recon , key Represents the weight coefficient of each loss item; The calculation of the key point loss specifically includes: By defining key point loss and deformation protection loss, constraints are imposed on the parameter update of the generator: The stylized face frame image is synthesized according to multiple conditionally reversible residual blocks and the original face frame image Use the pre-trained face key point detection model Dlip model to detect face key points and obtain the key point coordinates respectively and Calculate the keypoint loss used to measure the difference between the generated image and the real image in the location of the facial keypoints: The face frame image after stylization is synthesized by multiple conditionally reversible residual blocks obtained by the Canny edge detection algorithm and the original face frame image The edge image E t : Among them, E s It represents the ratio of edge pixels to total pixels, and h and w represent the height and width of the edge image respectively.

7. The method for face video style transfer according to claim 1, characterized in that: The generating of the face mask comprises the following steps: The input face video Decomposed into a frame sequence, where I t represents the tth frame; Use the face detection algorithm MTCNN to detect the face area in each frame and generate a face mask M t (x,y), get the face frame sequence 8. The method for face video style transfer according to claim 1, characterized in that: The output is a feature sequence that integrates multi-scale and spatiotemporal information, and includes the following steps: Face frame sequence Perform key point detection and obtain the key point coordinates K t ; According to the key point coordinates K t Align faces and obtain aligned face frames by unifying scale and posture The aligned face frame The content-driven multi-scale convolutional network module input into the encoder performs convolution operations on the input image to obtain feature maps of corresponding scales; Perform weighted fusion on the obtained multi-scale feature maps to obtain a comprehensive feature representation; The obtained comprehensive feature representation is input into the global average pooling layer in the network module of the spatiotemporal attention mechanism to obtain the feature vector after dimensionality reduction; The reduced feature vector is input into the fully connected layer to calculate the temporal attention weight; Multiply the temporal attention weight by the corresponding comprehensive feature to obtain the temporal weighted feature; Through the regional adaptive spatial attention module, the time-weighted features are convolved to obtain a spatial feature map; Get the spatial attention weight, multiply the spatial attention weight with the spatial feature map to get the spatial weighted feature, and use it as the feature sequence that integrates multi-scale and spatiotemporal information after encoding by the encoder 9. A face video style transfer system, characterized in that: include: The video acquisition and preprocessing module is used to acquire the face video frame image and decompose it to generate the face mask; A generator network model building module is used to build a generator network model; the generator network model includes an encoder and multiple conditional reversible residual blocks; a content-driven multi-scale convolution strategy and a spatiotemporal attention mechanism are introduced into the encoder, and an affine transformation layer and a deformation protection layer are introduced into each conditional reversible residual block; The style transfer module is used to input the acquired face video frame image into the encoder to extract the feature sequence of multi-scale and spatiotemporal information; in the affine transformation layer, the feature sequence is affine transformed, and the feature sequence is fused into the target style feature through the transformation function; in the deformation protection layer, the features of the facial area of ​​the target style feature are constrained according to the face mask, and the target style feature is converted into the facial area and the non-facial area, the original features of the facial information are retained in the facial area, and the style is converted in the non-facial area through the residual block in the conditional reversible residual block, and the original features are fused with the style conversion features to obtain the fused stylized face frame image features generated by the current conditional reversible residual block; the fused stylized face frame image features generated by the current conditional reversible residual block are sequentially input into the next conditional reversible residual block for loop execution, and the face frame image after comprehensive stylization of multiple conditional reversible residual blocks is obtained through decoding by the decoder.

Citation Information

Patent Citations

  • Chinese opera makeup migration method based on UV space mapping

    CN117196935A

  • Facial expression style migration method and system based on deep learning

    CN117237184A