Data processing method and device, electronic equipment and storage medium

By performing multimodal data processing and optimization of the video generation model, the problem of low model limitation and driving accuracy in face drivers is solved, and higher content quality and driving effect are achieved.

CN119992644APending Publication Date: 2025-05-13CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411978907.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

During the face driving process, the model can only process specific input information, resulting in a decrease in the driving accuracy of the language changes and the content quality is not high.

Method used

By obtaining the video generation model and its training data, including audio and video, using discriminators, action encoders, and appearance encoders to perform feature extraction and model training on audio and video, the video generation model is optimized to generate target videos consistent with the input video lip shape and have the appearance of the speaker video face.

Benefits of technology

It improves the accuracy of model processing, improves the authenticity and fluency of the generated target video, can capture the speaker's facial dynamics more comprehensively, and improves the driving effect of their respective modalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992644A_ABST
    Figure CN119992644A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method and device, electronic equipment and a storage medium, and relates to the technical field of computer vision, and the method comprises the steps: obtaining a video generation model and training data for the video generation model, and the training data comprises audios and videos; performing feature extraction on the audio and the video to obtain a driving signal corresponding to the audio and the video and an original coordinate of a sampling point corresponding to each frame of driving image in the audio and the video; inputting the original coordinates corresponding to each sampling point and the driving signal into a video generation model for model training to obtain a first loss function corresponding to a discriminator and a second loss function corresponding to an action encoder and an appearance encoder; and parameter tuning is performed on the video generation model based on the first loss function and the second loss function, and the trained video generation model is obtained, so that the accuracy of model processing is improved, the facial dynamics of the speaker can be captured more comprehensively, and the driving effect of each modal is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a data processing method, a data processing device, an electronic device and a computer-readable storage medium. Background Art

[0002] Face driving refers to driving the target face through a voice or video to produce expressions, movements, and postures that are consistent with the source voice / source video. With the development of face driving technology, it can be widely used in animation production, virtual anchors, video conferencing, and human-computer interaction. However, in the process of face driving, there is still a situation where the model can only process specific input information. At the same time, as the language changes, the driving accuracy of the language may also decrease, resulting in low quality of the generated content. Summary of the invention

[0003] The embodiments of the present invention provide a data processing method, device, electronic device and computer-readable storage medium to solve or partially solve the problems of model limitations, language limitations and low driving accuracy in the process of driving faces to generate corresponding videos.

[0004] The embodiment of the present invention discloses a data processing method, including:

[0005] Acquire a video generation model and training data for the video generation model, wherein the training data includes at least audio and video, and the video generation model includes a discriminator, an action encoder, and an appearance encoder;

[0006] Extracting features of the audio and video to obtain a driving signal corresponding to the audio and video and original coordinates of sampling points corresponding to each frame of a driving image in the audio and video;

[0007] Inputting the original coordinates corresponding to each of the sampling points and the driving signal into the video generation model for model training, and obtaining a first loss function corresponding to the discriminator and a second loss function corresponding to the action encoder and the appearance encoder;

[0008] Based on the first loss function and the second loss function, the parameters of the video generation model are tuned until the first loss function and the second loss function both meet preset conditions, thereby obtaining a trained video generation model. The video generation model is used to generate a target video with a lip shape consistent with the speaker video, the audio or the driving video and with the facial appearance of the speaker video for a given input speaker video or audio or driving video.

[0009] The embodiment of the present invention further discloses a data processing device, including:

[0010] A data acquisition module, used to acquire a video generation model and training data for the video generation model, wherein the training data includes at least audio and video, and the video generation model includes a discriminator, an action encoder, and an appearance encoder;

[0011] A feature extraction module is used to extract features from the audio and video to obtain a driving signal corresponding to the audio and video and original coordinates of a sampling point corresponding to each frame of a driving image in the audio and video;

[0012] A function determination module, used for inputting the original coordinates corresponding to each of the sampling points and the driving signal into the video generation model for model training, and obtaining a first loss function corresponding to the discriminator and a second loss function corresponding to the action encoder and the appearance encoder;

[0013] A training module is used to tune the parameters of the video generation model based on the first loss function and the second loss function until the first loss function and the second loss function both meet preset conditions, thereby obtaining a trained video generation model. The video generation model is used to generate a target video with a lip shape consistent with the speaker video or the audio or the driving video and with the facial appearance of the speaker video for a given input speaker video or audio or driving video.

[0014] The embodiment of the present invention further discloses an electronic device, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;

[0015] The memory is used to store computer programs;

[0016] The processor is used to implement the method described in the embodiment of the present invention when executing the program stored in the memory.

[0017] The embodiment of the present invention further discloses a computer-readable storage medium having instructions stored thereon, which, when executed by one or more processors, enables the processors to execute the method described in the embodiment of the present invention.

[0018] The embodiments of the present invention include the following advantages:

[0019] In an embodiment of the present invention, a video generation model and training data for the video generation model are obtained, the training data includes at least audio and video, and the video generation model includes at least a discriminator, an action encoder, and an appearance encoder. Then, feature extraction is performed on the audio and video to obtain a driving signal corresponding to the audio and video and the original coordinates of the sampling points corresponding to each frame of the driving image in the audio and video. Then, the original coordinates corresponding to each sampling point and the driving signal are input into the video generation model for model training to obtain a first loss function corresponding to the discriminator and a second loss function corresponding to the action encoder and the appearance encoder. Finally, based on the first loss function and the second loss function, parameter tuning is performed on the video generation model until both the first loss function and the second loss function meet preset conditions, and a trained video generation model is obtained. The video generation model is used to generate a target video that is consistent with the lip shape of the speaker video or audio or driving video and has the appearance of the speaker's face for a given input speaker video or audio or driving video, thereby improving the accuracy of model processing by processing data of multiple modalities, improving the authenticity and fluency of the target video generated based on the video generation model, and at the same time, based on multi-modal signal processing, it is possible to more comprehensively capture the speaker's facial dynamics and improve the driving effects of each modality. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a flowchart of a method for processing data provided in an embodiment of the present invention;

[0021] Figure 2 is a flow chart of a model structure provided in an embodiment of the present invention;

[0022] Figure 3 is a schematic diagram of the structure of an action encoder provided in an embodiment of the present invention;

[0023] Figure 4 is a schematic diagram of the structure of an appearance encoder provided in an embodiment of the present invention;

[0024] Figure 5 It is a structural block diagram of a data processing device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0025] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0026] As an example, in the process of implementing face and voice driving through Neural Radiance Field (NeRF), there are corresponding shortcomings:

[0027] 1) The model can only receive one mode of driving signal input (voice or image) to generate a driving video; 2) In the voice-driven task, the Chinese speech text content recognition (Automatic Speech Recognition, ASR) model extracts a higher dimension and often uses a feature extraction model trained under an English dataset for Chinese speech, resulting in a low accuracy rate for Chinese driving.

[0028] In this regard, in the present invention, by obtaining a video generation model and training data for the video generation model, the training data includes at least audio and video, and the video generation model includes at least a discriminator, an action encoder, and an appearance encoder, then feature extraction is performed on the audio and video, and the driving signal corresponding to the audio and video and the original coordinates of the sampling points corresponding to each frame of the driving image in the audio and video are obtained, and then the original coordinates corresponding to each sampling point and the driving signal are input into the video generation model for model training, and a first loss function corresponding to the discriminator and a second loss function corresponding to the action encoder and the appearance encoder are obtained. Finally, based on the first loss function and the second loss function, the video generation model is parameter tuned until the first loss function and the second loss function both meet the preset conditions, and a trained video generation model is obtained. The video generation model is used to generate a target video that is consistent with the lip shape of the speaker video or audio or driving video and has the appearance of the speaker's face for a given input speaker video or audio or driving video, thereby improving the accuracy of model processing by processing data of multiple modes, and improving the authenticity and fluency of the target video generated based on the video generation model. At the same time, based on multi-modal signal processing, the speaker's facial dynamics can be captured more comprehensively, and the driving effect of each mode can be improved.

[0029] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, some technical features involved in the embodiments of the present invention are explained and illustrated below:

[0030] Neural Radiance Field: Neural Radiance Field (NeRF) is an implicit method for representing 3D scenes using neural networks. Unlike traditional explicit geometric models (such as triangular meshes or voxels), NeRF uses a multi-layer perceptron (MLP) to represent the radiance field in the scene. It takes 3D coordinates and viewpoint as input and outputs the color and density of each point in the scene.

[0031] NeRF generates images from synthetic perspectives through volume rendering equations. The principle of NeRF relies on the following key ideas: (1) Implicit representation: NeRF maps coordinates (x, y, z) and directions (θ, φ) in three-dimensional space to colors (RGB, Reg, Green, Blue) and volume density (σ) by learning a neural network. This method reduces the storage requirements in traditional three-dimensional mesh model representations; (2) Volume rendering: NeRF generates images from new perspectives by solving the volume rendering equations and combining the learned volume density and color information; (3) Multi-view supervision: NeRF needs to use multi-view images for training, and supervises the network to learn the geometry and lighting information of the scene through image information from different perspectives.

[0032] Face driving: usually refers to driving the target face through a voice or video to make it produce expressions, movements and postures consistent with the source voice / source video.

[0033] Specifically, refer to Figure 1 , shows a flowchart of a data processing method provided in an embodiment of the present invention, which may specifically include the following steps:

[0034] Step 101, obtaining a video generation model and training data for the video generation model, wherein the training data includes at least audio and video, and the video generation model includes a discriminator, an action encoder, and an appearance encoder;

[0035] In the embodiment of the present invention, based on the neural radiation field, voice and video can be used simultaneously to generate face driving video. Based on this, in the process of training the video generation model, corresponding training data can be obtained, and the training data can at least include audio and video and corresponding driving images. Among them, audio and video can include corresponding driving audio and several frames of driving images, etc., and the present invention does not limit this.

[0036] Optionally, for the video generation model, it may include a discriminator, an action encoder and an appearance encoder. The discriminator can be used to output a driving latent vector under different modes; the action encoder can be used to extract corresponding action information from the driving signal to generate corresponding coordinate offsets, thereby realizing the action generation of the character; the appearance encoder can be used to determine the corresponding density and color, thereby realizing the appearance generation of the character.

[0037] Step 102, extracting features from the audio and video to obtain a driving signal corresponding to the audio and video and original coordinates of sampling points corresponding to each frame of a driving image in the audio and video;

[0038] After the training data is determined, features can be extracted from the training data first so that model training can be performed based on the extracted features. Specifically, features can be extracted from the audio and video in the training data to obtain the driving signals corresponding to the audio and video and the original coordinates of the sampling points corresponding to each frame of the driving image in the audio and video, etc., so as to perform model training based on the driving signals and the original coordinates.

[0039] In some feasible implementations, for feature extraction of audio and video, feature extraction can be first performed on each frame of the driving image in the audio and video to obtain a deformable parameterized model corresponding to each frame of the driving image and a 64-dimensional coefficient associated with the expression base, and the deformable parameterized model and the 64-dimensional coefficient are used as video driving signals corresponding to the audio and video. The driving audio in the audio and video is then time-sequentially segmented at 25 frames per second to obtain a switched audio signal, and feature extraction is performed on the segmented audio signal to obtain a corresponding multi-dimensional feature vector. The multi-dimensional feature vector is used to characterize the probability distribution of each Chinese character in the audio frame, and then the target Chinese characters whose occurrence probability is greater than or equal to a preset threshold are extracted from the multi-dimensional feature vector, and calculations are performed based on the pinyin information corresponding to each target Chinese character to obtain the audio driving signal corresponding to the audio and video.

[0040] Among them, the pinyin information can at least include the initial consonants and finals of the target Chinese characters. For the audio driving signal, the first occurrence probability of each initial consonant in all target Chinese characters and the second occurrence probability of each final in all target Chinese characters can be calculated, and then the first occurrence probability and the second occurrence probability can be used for calculation to obtain the joint distribution probability corresponding to the initial consonants and the finals, and the joint distribution probability can be used as the audio driving signal corresponding to the audio and video.

[0041] In addition, for the driving image, the original coordinates corresponding to each sampling point can be obtained by determining the light projected from each pixel point on the driving image to the screen space, randomly sampling several target rays from the rays, and sampling several sampling points on each target ray.

[0042] In one example, for video driving signal extraction, the Deep3DFace method [6] can be used to extract 64-dimensional coefficients related to the expression basis in the 3D Morphable Model (3DMM) of each frame image as the video driving signal.

[0043] For audio drive signal extraction, first, for the input audio signal, the time sequence can be segmented at a frame rate of 25 frames per second, and then the segmented audio signal is sent to the pre-trained Chinese ASR model (Automatic Speech Recognition) based on wav2vec for feature extraction. Each frame of audio signal is converted into a 3503-dimensional feature vector (i.e., a multidimensional feature vector), which represents the probability distribution of each Chinese character in the frame. Among them, the 3503-dimensional feature corresponds to the probability value of 3503 commonly used Chinese characters, ranging from 0 to 1, indicating the probability of each Chinese character being pronounced in the frame. Then, for each 3503-dimensional feature vector, the K Chinese characters with the largest probability value are extracted. Optionally, the K round library is taken as 20, so as to find the Chinese characters that are most likely to be pronounced in the current audio frame among all Chinese characters, and then focus on the most important features. Decompose the selected K Chinese characters into their corresponding initials and finals. Initials may include "b", "p", "m", "f", "d", "t", "n", "l", "g", "k", "h", "j", "q", "x", "zh", "ch", "sh", "r", "z", "c", "s", and finals may include "a", "o", "e", "i", "u", "v", "ai", "ei", "ui", "ao", "ou", "iu", "ie", "ve", "er", "an", "en", "in", "un", "vn", "ang", "eng", "ing", "ong", etc., so that the model can better capture the detailed information of the speech by refining the pronunciation features of Chinese characters. After obtaining the corresponding initials and finals, the probability of occurrence of each initial and final in the K Chinese characters can be calculated respectively, and then these probability values ​​can be superimposed to form a new joint distribution probability of initials and finals.

[0044] For example, if there are multiple Chinese characters among K Chinese characters that share the same initial consonant or final consonant, the probability will increase accordingly. Finally, a 45-dimensional feature vector is constructed using the joint distribution probability of initial consonants and final consonants. This vector contains the key information extracted from the high-dimensional 3503-dimensional features, retains the main speech features, and removes redundant information, thereby effectively reducing the 3503-dimensional audio features to 45 dimensions and optimizing the feature representation. By extracting the probability distribution of initial consonants and final consonants, the details of Chinese speech can be captured more accurately, thereby improving the model's driving accuracy for Chinese speech, and this dimensionality reduction method makes the model more efficient and reduces the consumption of computing resources.

[0045] In addition, for the driving image, assume that each pixel in the image is the projection of a ray on the screen. Randomly sample K pixels, that is, K rays. Sample N points on each ray on the K rays. For example, assuming that the driving image includes 225 pixels, 50 pixels can be randomly sampled from it, that is, 50 rays are determined, and then 15 points are randomly sampled on the rays, so that 50*15, that is, 750 sampling points can be obtained, and the original coordinates corresponding to each sampling point can be obtained.

[0046] Step 103, inputting the original coordinates corresponding to each of the sampling points and the driving signal into the video generation model for model training, and obtaining a first loss function corresponding to the discriminator and a second loss function corresponding to the action encoder and the appearance encoder;

[0047] After obtaining the corresponding original coordinates and driving signals through feature extraction, the original coordinates corresponding to each sampling point and the driving signal can be input into the video generation model for model training to obtain a first loss function corresponding to the discriminator and a second loss function corresponding to the action encoder and the appearance encoder, so as to perform parameter tuning on the video generation model based on the first loss function and the second loss function.

[0048] In some feasible implementations, the second loss function includes at least a color loss function and a feature loss function. During the model training process, the audio driving signal can be calculated first to obtain the corresponding audio features, and then the video driving signal can be calculated to obtain the corresponding video features. Then, the original coordinates, audio features, and video features corresponding to each sampling point are input into the discriminator for adversarial learning and comparative learning to obtain the corresponding target feature information and the first loss function corresponding to the target feature information. The target feature information includes the video latent vector of the video feature in the latent space and the audio latent vector of the audio feature in the same latent space. After mapping to the latent vector in the latent space, the video latent vector, the audio latent vector and the original coordinates corresponding to each sampling point can be input into the action encoder for encoding to obtain the first multidimensional feature and the coordinate offset information corresponding to the original coordinates. The coordinate offset information, the first multidimensional feature and the original coordinates are input into the appearance encoder for encoding to obtain the corresponding second multidimensional feature, and the second multidimensional feature is used for calculation to obtain the density and point color corresponding to each sampling point; then the density and color corresponding to the sampling point are used for calculation to obtain the target color corresponding to the light to which the sampling point belongs, and the color loss function corresponding to the target color is determined, and the feature loss function between the audio feature and the video feature is calculated, so as to use the voice and the video as the driving signal, combine the audio feature and the video feature, and jointly optimize the audio feature and the video feature in the same latent space by using the adversarial learning and contrast learning method, so as to improve the authenticity and fluency of the generated face video, and after the training is completed, the face speaking video can be generated based on the video and audio drive at the same time, which has a larger application range, and the design based on the dual driving signal can more comprehensively capture the facial dynamics of the speaker and improve the driving effect of each modality.

[0049] In addition, for the first loss function, which includes a contrast loss function and an adversarial loss function, the original coordinates, audio features, and video features corresponding to each sampling point can be input into the discriminator for adversarial learning and contrast learning to obtain corresponding target feature information;

[0050] And the adversarial loss function corresponding to the target feature information is calculated by the following formula:

[0051] L GAN =E[logD(z1)]+E[log(1-D(z2)]

[0052] And, the contrast loss function corresponding to the target feature information is calculated by the following formula:

[0053]

[0054] Among them, z1 is the audio feature; z2 is the video feature; sim(·) is the cosine similarity measure; τ is the temperature coefficient, which is used to adjust the sharpness of the distribution; N is the number of negative samples.

[0055] In one example, for two driving signals, the feature z is calculated by a fully connected layer (MLP, Multilayer Perceptron):

[0056] z1=MLP1(d1)

[0057] z2=MLP2(d2)

[0058] Among them, d1 and d2 are audio and video driving signals. The driving signals under various modes are uniformly mapped to features in a latent space to obtain audio and video features z1 and z2 in the same latent space.

[0059] In order to realize the information exchange and complementarity between modalities of NeRF driving and improve the accuracy of driving, in the embodiment of the present invention, the consistency of features in the two domains can be jointly optimized by adversarial generation. Specifically, the discriminator inputs the driving latent vector under different modalities, determines whether the latent vector comes from a certain modality, and adds adversarial loss to the network training (training the discriminator and adjusting the corresponding parameters of the discriminator according to the adversarial loss function):

[0060] L GAN =E[logD(z1)]+E[log(1-D(z2)]

[0061] In order to further enhance the similarity of features between modalities, a contrastive learning method is introduced to improve the similarity between modalities and optimize the final driving effect of the convergence of the model.

[0062] Specifically, for audio feature z1 and video feature z2, define positive sample pairs (z 1:t , z 2:t ), and negative sample pairs (z 1:t , z 2:t’ ).

[0063] By improving the consistency of different modal features from the same time frame and expanding the distinguishability of different modal features from different time frames, the representation ability of the same facial dynamics in different modalities is improved. The contrast loss is defined as (adjusting the corresponding parameters of the latent space through the contrast loss function):

[0064]

[0065] Where sim(·) is the cosine similarity measure, τ is the temperature coefficient, which adjusts the sharpness of the distribution, and N is the number of negative samples.

[0066] Randomly select a feature z mapped from a driving signal and the coordinates X of all points and input them into the motion encoder, output the corresponding 128-dimensional features, and output the coordinate offset m through an MLP:

[0067] m=MotionEncoder(P,z)

[0068] Add the motion encoder output to the original coordinates to get the actual sampling point coordinates:

[0069] P′=P+m

[0070] Where P' is the offset point coordinate set.

[0071] Input the actual coordinates P' and feature z into the face 3D appearance encoder (AppearanceEncoder), and output the corresponding 128-dimensional feature f:

[0072] f = AppearanceEncoder(P′,z)

[0073] The density σ and color c of each point are calculated by MLP:

[0074] (σ,c)=MLP(f)

[0075] In terms of loss function, the color of each light is calculated according to the volume rendering formula, and the L1 loss is calculated with GT:

[0076] L color =||C pred -C GT ||

[0077] At the same time, the L1 loss between the features zi generated by the two driving signals is calculated:

[0078] L latent =‖z1-z2‖

[0079] The final loss function is:

[0080] L color =L color +μ1L latent +μ2L GAN +μ3L contrast

[0081] Among them, μ1, μ2 and μ3 are adjustment parameters, which are set to 0.1, 0.5 and 0.1, etc. The correlation between modes can be dynamically improved by joint optimization based on adversarial learning and contrastive learning of latent space, which can effectively improve the driving generation effect of each mode.

[0082] Step 104, tuning the parameters of the video generation model based on the first loss function and the second loss function until the first loss function and the second loss function both meet preset conditions, thereby obtaining a trained video generation model, wherein the video generation model is used to generate a target video that is consistent with the lip shape of the speaker video or the audio or the driving video and has the facial appearance of the speaker video for a given input speaker video or audio or driving video.

[0083] In an embodiment of the present invention, after obtaining the corresponding loss function through the above process, the parameters of the video generation model can be tuned according to the corresponding parameter tuning strategy and combined with the loss function to ensure the model prediction effect of the video generation model. Specifically, the parameters of the video generation model can be tuned based on the first loss function and the second loss function until the first loss function and the second loss function both meet the preset conditions to obtain a trained video generation model. The video generation model is used to generate a target video that is consistent with the lip shape of the speaker video or audio or driving video and has the facial appearance of the speaker video for a given input speaker video or audio or driving video, thereby improving the accuracy of model processing by processing data of multiple modalities, and improving the authenticity and fluency of the target video generated based on the video generation model. At the same time, based on multi-modal signal processing, it is possible to more comprehensively capture the facial dynamics of the speaker and improve the driving effect of each modality.

[0084] In one example, based on the trained model, the video generation model can be applied to the virtual anchor, which may specifically include the following process:

[0085] Step 1: Input the basic video material of the virtual anchor

[0086] Select a 3-5 minute basic speaking video of a virtual anchor and prepare the corresponding voice input.

[0087] Step 2: Feature extraction and preprocessing

[0088] The Deep3DFace and ASR models are used to extract video and audio features, and feature dimensionality reduction is performed according to the method described in the present invention.

[0089] Step 3: Model training and optimization

[0090] The basic video material and voice input are fed into the training model for multiple rounds of iterative optimization to ensure that the model can generate facial expressions and movements consistent with the virtual anchor.

[0091] Step 4: Real-time Inference

[0092] During the live broadcast of the virtual anchor, the anchor's audio and video input are sampled in real time. Thanks to the low-dimensional speech features and the efficient action encoder and appearance encoder used in this model, this model can achieve real-time video generation, thereby using the trained model for dynamic generation and updating the virtual anchor's facial expressions and lip shape in real time.

[0093] It should be noted that the embodiments of the present invention include but are not limited to the above examples. It is understandable that those skilled in the art can also make settings according to actual needs under the guidance of the ideas of the embodiments of the present invention, and the present invention is not limited to this.

[0094] In an embodiment of the present invention, a video generation model and training data for the video generation model are obtained, the training data includes at least audio and video, and the video generation model includes at least a discriminator, an action encoder, and an appearance encoder. Then, feature extraction is performed on the audio and video to obtain a driving signal corresponding to the audio and video and the original coordinates of the sampling points corresponding to each frame of the driving image in the audio and video. Then, the original coordinates corresponding to each sampling point and the driving signal are input into the video generation model for model training to obtain a first loss function corresponding to the discriminator and a second loss function corresponding to the action encoder and the appearance encoder. Finally, based on the first loss function and the second loss function, parameter tuning is performed on the video generation model until both the first loss function and the second loss function meet preset conditions, and a trained video generation model is obtained. The video generation model is used to generate a target video that is consistent with the lip shape of the speaker video or audio or driving video and has the appearance of the speaker's face for a given input speaker video or audio or driving video, thereby improving the accuracy of model processing by processing data of multiple modalities, improving the authenticity and fluency of the target video generated based on the video generation model, and at the same time, based on multi-modal signal processing, it is possible to more comprehensively capture the speaker's facial dynamics and improve the driving effects of each modality.

[0095] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the following examples are used for exemplary description:

[0096] Reference Figure 2 , a flow chart of the model structure provided in an embodiment of the present invention is shown. In the process of model training, the training data may include a certain length of audio and video, which may be composed of corresponding driving audio and several frames of driving images. Then, the driving audio and driving images may be feature extracted to obtain the corresponding driving signal. For the extraction process of the video driving signal, the Deep3DFace method may be used to extract the 64-dimensional coefficients related to the expression base in the 3D deformable parameterized model (3D Morphable Model, 3DMM) of each frame of the image as the video driving signal.

[0097] For the extraction of audio driving signals, the input audio signal can be time-sequenced at a frame rate of 25 frames per second. The segmented audio signal is sent to the pre-trained Chinese ASR model based on wav2vec for feature extraction. Each frame of audio signal is converted into a 3503-dimensional feature vector, which represents the probability distribution of each Chinese character in the frame. Specifically, these 3503-dimensional features correspond to the probability values ​​of 3503 commonly used Chinese characters, ranging from 0 to 1, indicating the probability of each Chinese character being pronounced in the frame. And, for each 3503-dimensional feature vector, extract the K Chinese characters with the largest probability value, in the present invention, K is 20, etc., and then decompose the selected K Chinese characters into their corresponding initial consonants and finals, wherein the initial consonants include "b", "p", "m", "f", "d", "t", "n", "l", "g", "k", "h", "j", "q", "x", "zh", "ch", "sh", "r", "z", "c", "s", and the finals include "a", "o", "e", "i", "u", "v", "ai", "ei", "ui", "ao", "ou", "iu", "ie", "ve", "er", "an", "en", "in", "un", "vn", "ang", "eng", "ing", "ong". For each initial consonant and final, calculate its probability of occurrence in the K Chinese characters, superimpose these probability values, and form a new joint distribution probability of the initial consonant and final.

[0098] For example, if there are multiple Chinese characters among K Chinese characters that share the same initial consonant or final consonant, the probability will increase accordingly. Finally, a 45-dimensional feature vector is constructed using the joint distribution probability of initial consonants and final consonants. This vector contains the key information extracted from the high-dimensional 3503-dimensional features, retains the main speech features, and removes redundant information, thereby achieving a more accurate Chinese driving effect through feature dimensionality reduction: the 3503-dimensional audio features are effectively reduced to 45 dimensions, optimizing the feature representation. By extracting the probability distribution of initial consonants and final consonants, the details of Chinese speech can be captured more accurately, thereby improving the accuracy of the model's driving of Chinese speech. Moreover, this dimensionality reduction method makes the model more efficient and reduces the consumption of computing resources.

[0099] After the training data is processed through the above process, during the model training process, input a frame of image, assuming that each pixel in the image is the projection of a ray on the screen. Randomly sample K pixels, that is, K rays. Sample N points on each ray on the K rays. The input of the model is the coordinates P of each point and two driving signals. For the two driving signals, the feature z is calculated through the fully connected layer (MLP):

[0100] z1=MLP1(d1)

[0101] z2=MLP2(d2)

[0102] Among them, d1 and d2 are audio and video driving signals. The driving signals under various modes are uniformly mapped to features in a latent space to obtain audio and video features z1 and z2 in the same latent space.

[0103] In order to enable NeRF's driving to achieve information exchange and complementarity between modalities and improve the accuracy of driving, this patent also proposes to jointly optimize the consistency of features in the two domains through adversarial generation.

[0104] The discriminator D inputs the driving latent vector under different modes, determines whether the latent vector comes from a certain mode, and adds adversarial loss to the network training (training the discriminator and adjusting the corresponding parameters of the discriminator according to the adversarial loss function):

[0105] L GAN =E[logD(z1)]+E[log(1-D(z2)]

[0106] In order to further enhance the similarity of features between modalities, a contrastive learning method is introduced to improve the similarity between modalities and optimize the final driving effect of the convergence of the model.

[0107] Specifically, for audio feature z1 and video feature z2, define positive sample pairs (z 1:t , z 2:t ), and negative sample pairs (z 1:t , z 2:t’ ).

[0108] By improving the consistency of different modal features from the same time frame and expanding the distinguishability of different modal features from different time frames, the representation ability of the same facial dynamics in different modalities is improved. The contrast loss is defined as (adjusting the corresponding parameters of the latent space through the contrast loss function):

[0109]

[0110] Where sim(·) is the cosine similarity measure, τ is the temperature coefficient, which adjusts the sharpness of the distribution, and N is the number of negative samples.

[0111] Randomly select a feature z mapped from a driving signal and the coordinates X of all points and input them into the motion encoder (MotionEncoder), the structure is as follows Figure 3 As shown. Output 128-dimensional features and pass an MLP output coordinate offset m:

[0112] m=MotionEncoder(P,z)

[0113] Add the motion encoder output to the original coordinates to get the actual sampling point coordinates:

[0114] P′=P+m

[0115] Where P' is the offset point coordinate set.

[0116] The actual coordinates P' and feature z are input into the face 3D appearance encoder (AppearanceEncoder), the structure is as follows Figure 4 As shown. Output 128-dimensional feature f:

[0117] f = AppearanceEncoder(P′,z)

[0118] The density σ and color c of each point are calculated by MLP:

[0119] (σ,c)=MLP(f)

[0120] In terms of loss function, the color of each light is calculated according to the volume rendering formula, and the L1 loss is calculated with GT:

[0121] L color =||C pred -C GT ||

[0122] At the same time, the L1 loss between the features zi generated by the two driving signals is calculated:

[0123] L latent =‖z1-z2‖

[0124] The final loss function is:

[0125] L color =L color +μ1L latent +μ2L GAN +μ3L contrast

[0126] Among them, μ1, μ2 and μ3 are adjustment parameters, which are set to 0.1, 0.5 and 0.1. The joint optimization of latent space based on adversarial learning and contrastive learning can dynamically improve the correlation between modes, thereby improving the driving generation effect of each mode.

[0127] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0128] Reference Figure 5 , shows a structural block diagram of a data processing device provided in an embodiment of the present invention, which may specifically include the following modules:

[0129] A data acquisition module 501 is used to acquire a video generation model and training data for the video generation model, wherein the training data includes at least audio and video, and the video generation model includes a discriminator, an action encoder, and an appearance encoder;

[0130] A feature extraction module 502 is used to extract features from the audio and video to obtain a driving signal corresponding to the audio and video and original coordinates of a sampling point corresponding to each frame of a driving image in the audio and video;

[0131] A function determination module 503 is used to input the original coordinates corresponding to each of the sampling points and the driving signal into the video generation model for model training, so as to obtain a first loss function corresponding to the discriminator and a second loss function corresponding to the action encoder and the appearance encoder;

[0132] The training module 504 is used to tune the parameters of the video generation model based on the first loss function and the second loss function until the first loss function and the second loss function both meet preset conditions, thereby obtaining a trained video generation model. The video generation model is used to generate a target video with a lip shape consistent with the speaker video or the audio or the driving video and with the facial appearance of the speaker video for a given input speaker video or audio or driving video.

[0133] In some feasible implementations, the feature extraction module 502 is specifically used for:

[0134] Performing feature extraction on each frame of the driving image in the audio and video, obtaining a deformable parameterized model corresponding to each frame of the driving image and a 64-dimensional coefficient associated with an expression base, and using the deformable parameterized model and the 64-dimensional coefficient as a video driving signal corresponding to the audio and video;

[0135] Sequentially segment the driving audio in the audio and video according to 25 frames per second to obtain a switched audio signal;

[0136] Performing feature extraction on the segmented audio signal to obtain a corresponding multi-dimensional feature vector, wherein the multi-dimensional feature vector is used to characterize the probability distribution of each Chinese character in the audio frame;

[0137] Target Chinese characters whose occurrence probability is greater than or equal to a preset threshold are extracted from the multidimensional feature vector, and calculation is performed based on the pinyin information corresponding to each of the target Chinese characters to obtain an audio driving signal corresponding to the audio and video.

[0138] In some feasible implementations, the pinyin information includes at least the initial consonant and the final vowel of the target Chinese character, and the feature extraction module 502 is specifically used for:

[0139] Calculate the first appearance probability of each of the initial consonants in all the target Chinese characters;

[0140] Calculate the second occurrence probability of each of the finals in all the target Chinese characters;

[0141] The first occurrence probability and the second occurrence probability are used to perform calculations to obtain a joint distribution probability corresponding to the initial consonant and the final consonant, and the joint distribution probability is used as an audio driving signal corresponding to the audio and video.

[0142] In some feasible implementations, the feature extraction module 502 is specifically used for:

[0143] Determine the light projected from each pixel point on the driving image to the screen space;

[0144] A number of target light rays are randomly sampled from the light rays, and a number of sampling points are sampled on each of the target light rays to obtain the original coordinates corresponding to each of the sampling points.

[0145] In some feasible implementations, the second loss function includes at least a color loss function and a feature loss function, and the function determination module 503 is specifically used to:

[0146] Calculating the audio driving signal to obtain corresponding audio features;

[0147] Calculating the video drive signal to obtain corresponding video features;

[0148] Inputting the original coordinates, the audio features, and the video features corresponding to each of the sampling points into the discriminator for adversarial learning and contrastive learning to obtain corresponding target feature information and a first loss function corresponding to the target feature information, wherein the target feature information includes a video latent vector of the video feature in a latent space and an audio latent vector of the audio feature in the same latent space;

[0149] Inputting the video latent vector, the audio latent vector and the original coordinates corresponding to each sampling point into the action encoder for encoding, to obtain a first multi-dimensional feature and coordinate offset information corresponding to the original coordinates;

[0150] The coordinate offset information, the first multidimensional feature and the original coordinate are input into the appearance encoder for encoding to obtain a corresponding second multidimensional feature;

[0151] The second multidimensional feature is used to perform calculation to obtain density and point color corresponding to each sampling point;

[0152] Calculating using the density and color corresponding to the sampling point to obtain a target color corresponding to the light to which the sampling point belongs, and determining a color loss function corresponding to the target color;

[0153] A feature loss function between the audio feature and the video feature is calculated.

[0154] In some feasible implementations, the first loss function includes at least a contrast loss function and an adversarial loss function, and the function determination module 503 is specifically used to:

[0155] Inputting the original coordinates, the audio features, and the video features corresponding to each of the sampling points into the discriminator for adversarial learning and contrastive learning to obtain corresponding target feature information;

[0156] The adversarial loss function corresponding to the target feature information is calculated by the following formula:

[0157] L GAN =E[logD(z1)]+E[log(1-D(z2)]

[0158] And, the contrast loss function corresponding to the target feature information is calculated by the following formula:

[0159]

[0160] Wherein, z1 is the audio feature; z2 is the video feature; sim(·) is the cosine similarity measure; τ is the temperature coefficient used to adjust the sharpness of the distribution; and N is the number of negative samples.

[0161] In some feasible implementations, the model training module 504 is specifically used to:

[0162] Obtaining a first adjustment parameter corresponding to the first loss function and a second adjustment coefficient corresponding to the feature loss function;

[0163] Based on the first loss function and the first adjustment coefficient, the color loss function, the feature loss function and the second adjustment coefficient, the parameters of the video generation model are tuned until the first loss function and the second loss function both meet preset conditions, thereby obtaining a trained video generation model.

[0164] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0165] In addition, an embodiment of the present invention further provides an electronic device, including: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the various processes of the above-mentioned data processing method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0166] The embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, each process of the above-mentioned data processing method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0167] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0168] It should be understood by those skilled in the art that the embodiments of the embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, EEPROM, Flash, and eMMC, etc.) containing computer-usable program codes.

[0169] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0170] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0171] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0172] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0173] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0174] The data processing method and the data processing device provided by the present invention are introduced in detail above. The principle and implementation mode of the present invention are explained by using specific examples in this article. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation mode and the scope of application. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A data processing method, characterized in that: include: Acquire a video generation model and training data for the video generation model, wherein the training data includes at least audio and video, and the video generation model includes a discriminator, an action encoder, and an appearance encoder; Extracting features of the audio and video to obtain a driving signal corresponding to the audio and video and original coordinates of sampling points corresponding to each frame of a driving image in the audio and video; Inputting the original coordinates corresponding to each of the sampling points and the driving signal into the video generation model for model training, and obtaining a first loss function corresponding to the discriminator and a second loss function corresponding to the action encoder and the appearance encoder; Based on the first loss function and the second loss function, the parameters of the video generation model are tuned until the first loss function and the second loss function both meet preset conditions, thereby obtaining a trained video generation model. The video generation model is used to generate a target video with a lip shape consistent with the speaker video, the audio or the driving video and with the facial appearance of the speaker video for a given input speaker video or audio or driving video.

2. The method according to claim 1, characterized in that The extracting features of the audio and video to obtain a driving signal corresponding to the audio and video includes: Performing feature extraction on each frame of the driving image in the audio and video, obtaining a deformable parameterized model corresponding to each frame of the driving image and a 64-dimensional coefficient associated with an expression base, and using the deformable parameterized model and the 64-dimensional coefficient as a video driving signal corresponding to the audio and video; Sequentially segment the driving audio in the audio and video according to 25 frames per second to obtain a switched audio signal; Performing feature extraction on the segmented audio signal to obtain a corresponding multi-dimensional feature vector, wherein the multi-dimensional feature vector is used to characterize the probability distribution of each Chinese character in the audio frame; Target Chinese characters whose occurrence probability is greater than or equal to a preset threshold are extracted from the multidimensional feature vector, and calculations are performed based on the pinyin information corresponding to each of the target Chinese characters to obtain audio driving signals corresponding to the audio and video.

3. The method according to claim 2, characterized in that The pinyin information at least includes the initial consonant and the final vowel of the target Chinese character, and the calculation is performed according to the pinyin information corresponding to each of the target Chinese characters to obtain the audio driving signal corresponding to the audio and video, including: Calculate the first appearance probability of each of the initial consonants in all the target Chinese characters; Calculate the second occurrence probability of each of the finals in all the target Chinese characters; The first occurrence probability and the second occurrence probability are used to perform calculations to obtain a joint distribution probability corresponding to the initial consonant and the final consonant, and the joint distribution probability is used as an audio driving signal corresponding to the audio and video.

4. The method according to any one of claims 1 to 3, characterized in that: The extracting features of the audio and video to obtain the original coordinates of the sampling points corresponding to each frame of the driving image in the audio and video includes: Determine the light projected from each pixel point on the driving image to the screen space; A number of target light rays are randomly sampled from the light rays, and a number of sampling points are sampled on each of the target light rays to obtain the original coordinates corresponding to each of the sampling points.

5. The method according to claim 2 or 3, characterized in that: The second loss function at least includes a color loss function and a feature loss function, and the original coordinates corresponding to each of the sampling points and the driving signal are input into the video generation model for model training to obtain the first loss function corresponding to the discriminator and the second loss functions corresponding to the action encoder and the appearance encoder, including: Calculating the audio driving signal to obtain corresponding audio features; Calculating the video drive signal to obtain corresponding video features; Inputting the original coordinates, the audio features, and the video features corresponding to each of the sampling points into the discriminator for adversarial learning and contrastive learning to obtain corresponding target feature information and a first loss function corresponding to the target feature information, wherein the target feature information includes a video latent vector of the video feature in a latent space and an audio latent vector of the audio feature in the same latent space; Inputting the video latent vector, the audio latent vector and the original coordinates corresponding to each sampling point into the action encoder for encoding, to obtain a first multi-dimensional feature and coordinate offset information corresponding to the original coordinates; The coordinate offset information, the first multidimensional feature and the original coordinate are input into the appearance encoder for encoding to obtain a corresponding second multidimensional feature; The second multidimensional feature is used to perform calculation to obtain density and point color corresponding to each sampling point; Calculating using the density and color corresponding to the sampling point to obtain a target color corresponding to the light to which the sampling point belongs, and determining a color loss function corresponding to the target color; A feature loss function between the audio feature and the video feature is calculated.

6. The method according to claim 5, characterized in that The first loss function at least includes a contrast loss function and an adversarial loss function, and the original coordinates corresponding to each of the sampling points, the audio features, and the video features are input into the discriminator for adversarial learning and contrastive learning to obtain corresponding target feature information and a first loss function corresponding to the target feature information, including: Inputting the original coordinates, the audio features, and the video features corresponding to each of the sampling points into the discriminator for adversarial learning and contrastive learning to obtain corresponding target feature information; The adversarial loss function corresponding to the target feature information is calculated by the following formula: L GAN =E[logD(z1)]+E[log(1-D(z2)] And, the contrast loss function corresponding to the target feature information is calculated by the following formula: Wherein, z1 is the audio feature; z2 is the video feature; sim(·) is the cosine similarity measure; τ is the temperature coefficient used to adjust the sharpness of the distribution; and N is the number of negative samples.

7. The method according to claim 5, characterized in that The step of tuning the parameters of the video generation model based on the first loss function and the second loss function until both the first loss function and the second loss function meet preset conditions to obtain a trained video generation model includes: Obtaining a first adjustment parameter corresponding to the first loss function and a second adjustment coefficient corresponding to the feature loss function; Based on the first loss function and the first adjustment coefficient, the color loss function, the feature loss function and the second adjustment coefficient, the parameters of the video generation model are tuned until the first loss function and the second loss function both meet preset conditions, thereby obtaining a trained video generation model.

8. A data processing device, characterized in that: include: A data acquisition module, used to acquire a video generation model and training data for the video generation model, wherein the training data includes at least audio and video, and the video generation model includes a discriminator, an action encoder, and an appearance encoder; A feature extraction module is used to extract features from the audio and video to obtain a driving signal corresponding to the audio and video and original coordinates of a sampling point corresponding to each frame of a driving image in the audio and video; A function determination module, used for inputting the original coordinates corresponding to each of the sampling points and the driving signal into the video generation model for model training, and obtaining a first loss function corresponding to the discriminator and a second loss function corresponding to the action encoder and the appearance encoder; A training module is used to tune the parameters of the video generation model based on the first loss function and the second loss function until the first loss function and the second loss function both meet preset conditions, thereby obtaining a trained video generation model. The video generation model is used to generate a target video with a lip shape consistent with the speaker video or the audio or the driving video and with the facial appearance of the speaker video for a given input speaker video or audio or driving video.

9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; The memory is used to store computer programs; The processor is used to implement the method according to any one of claims 1 to 7 when executing the program stored in the memory.

10. A computer-readable storage medium having instructions stored thereon, which, when executed by one or more processors, cause the processors to perform the method according to any one of claims 1 to 7.