A fast voice-driven facial video editing and generation method and system
By extracting the features of facial video data and using the teacher and student rectifier model to generate inverse sampling direction vectors for video denoising, the problem of insufficient generation quality and speed in traditional technology is solved, and fast and high-quality facial video editing and generation is achieved.
Patent Information
- Application Number
- CN202510260392.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-06
AI Technical Summary
Traditional voice-driven facial video generation technology has shortcomings in terms of generation quality and generation stability. Although the generation technology based on diffusion model improves the quality, the generation speed is slow and cannot be applied to video editing.
A fast voice-driven facial video editing and generation method is proposed. By obtaining facial video data, facial expression text description and video speech, video features, text features and speech features are extracted, and the teacher rectification model and student rectification model are trained using noise-added video features to generate inverse sampling direction vectors for video denoising and editing.
It greatly shortens the time for video generation or editing, improves the quality of target videos, and realizes fast voice-driven facial video generation and editing.
Smart Images

Figure CN119785270B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of face generation, and in particular relates to a fast voice-driven face video editing and generation method and system. Background Art
[0002] With the continuous development of speech-driven face generation technology (Talking Face Generation, TFG), this technology has shown broad potential in application scenarios such as virtual character generation, video conferencing, and short video advertisement generation. However, traditional face generation technology still needs to be improved in terms of generation quality and generation stability.
[0003] Traditional technologies often construct a mapping from speech to lip movements, while keeping the area outside the lips still, which limits the naturalness of the generation. In addition, the low-resolution image generated and the unnatural connection caused by pasting the generated lip area back to the image result in limited generation quality. When the generation technology based on the diffusion model emerged, although it improved the generation quality of facial video generation, it often required dozens of diffusion generation steps to generate a video, which was very time-consuming and not suitable for video editing. Summary of the invention
[0004] In order to overcome the problems of slow generation speed and poor generation quality of existing voice-driven facial video generation models, the present invention provides a fast voice-driven facial video editing and generation method and system, which can greatly shorten the video generation or editing time and improve the quality of the target video.
[0005] In order to achieve the above object, the specific technical solution adopted by the present invention is:
[0006] In a first aspect, the present invention provides a fast voice-driven facial video editing and generation method, comprising:
[0007] Obtain facial video data, facial expression text description and video voice, and extract video features, text features and voice features as feature data sets respectively;
[0008] The video features and Gaussian noise in the feature data set are masked by random time periods to generate noisy video features, and a teacher rectification model is trained using the noisy video features, text features and speech features. The teacher rectification model can generate an inverse sampling direction vector for gradually denoising the noisy video features.
[0009] A student rectification model is trained using the feature data set. In each training batch, the following steps are performed: the teacher rectification model generates an inverse sampling direction vector for the noisy video feature, and denoises the noisy video feature to generate a clean video feature based on the inverse sampling process of the ordinary differential equation of the inverse sampling direction vector; the clean video feature and its corresponding noisy video feature at time step 0 are used as feature sample pairs to regenerate the noisy video feature based on a random time step; the student rectification model runs the rectification learning objective for training based on the feature sample pairs and the regenerated noisy video feature;
[0010] The trained student rectification model is used to generate a video from a given facial image, or to edit a given facial video.
[0011] Furthermore, the step of performing random time period masking on the video features in the feature data set and the Gaussian noise to generate the noisy video features includes:
[0012] Create a video time segment mask, where the length of the blocked area in the mask is random, perform linear interpolation on the video features of the blocked area and Gaussian noise at the sampling time step, and generate noisy video features;
[0013] The length of the random occluded region of the video time segment mask is less than the video feature time length, and the sampling time step is sampled from a uniform [0, 1] distribution.
[0014] Furthermore, the video features of the blocked area and the Gaussian noise are linearly interpolated at the sampling time step, which is expressed as:
[0015] ;
[0016] in, represents the sampling time step, represents the sampling time step The noise-added video features under Represents the video time period mask, and the area where the value of M is 1 is the blocked area; represents the original video features, represents Gaussian noise, Represents dot product.
[0017] Furthermore, in each training batch, the generation process of feature sample pairs is as follows:
[0018] Initialize sampling time step ;
[0019] Create a video time segment mask with random length of the occluded area, perform linear interpolation on the video features of the occluded area and Gaussian noise under the setting of the initial sampling time step, and generate the initialized noisy video features ;
[0020] The teacher rectification model uses the Euler ordinary differential equation solver to run the inverse sampling process to denoise the parts blocked by the mask, and continuously iterates and updates until ,at this time ,in Represents clean video features;
[0021] The inverse sampling process of the Euler ordinary differential equation solver is expressed as follows:
[0022] ;
[0023] ;
[0024] in, represents the teacher rectification model, represents the inverse sampling direction vector output by the teacher rectification model, represents the noise-added video feature, Represents the voice features, Represents text features, Indicates the generation stride;
[0025] Clean Video Features And the noisy video features when time step is 0 As feature sample pairs for reflow training of student model.
[0026] Furthermore, the regeneration of the noisy video features based on random time steps refers to the regeneration of the clean video features. The noised video features at time step 0 Regenerate the noisy video features after masking at random time periods:
[0027] .
[0028] Furthermore, the teacher rectification model and the student rectification model are rectification models with the same structure.
[0029] Furthermore, the teacher rectification model and the student rectification model include:
[0030] The speech feature preprocessing layer is used to interpolate the speech features to make them have the same time dimension length as the flattened noisy video features;
[0031] The first fully connected layer is used to map the interpolated speech features to a channel dimension that is half of the latent dimension of the Transformer block;
[0032] The second fully connected layer is used to map the video features to a channel dimension that is half of the hidden dimension of the Transformer block;
[0033] The third fully connected layer is used to map text features to the latent dimension of the Transformer block with a channel dimension;
[0034] A concatenation layer is used to concatenate the output results of the first fully connected layer and the second fully connected layer in the channel dimension, and then concatenate the output results of the third fully connected layer in the time dimension;
[0035] N stacked Transformer blocks, which are used to extract a coding sequence from the output of the concatenated layer, where the channel dimension of the coding sequence is the same as the channel dimension of the noisy video features;
[0036] The reverse sampling direction vector output layer is used to crop the part corresponding to the noisy video feature from the encoded sequence as the reverse sampling direction vector.
[0037] Furthermore, the method of generating a video of a given facial image or editing a given facial video using the trained student rectification model includes:
[0038] Given a single-frame facial reference image, a text description of facial expressions, and a video voice, video features, text features, and voice features are extracted respectively; Gaussian noise corresponding to the length of the video voice is spliced after the video features, and the trained student rectification model takes the text features, voice features, and video features after splicing Gaussian noise as inputs to gradually generate inverse sampling direction vectors; the video features after splicing Gaussian noise are subjected to an inverse sampling process based on ordinary differential equations of direction vectors to gradually obtain clean video features, and the video is generated after decoding;
[0039] Alternatively, given a facial reference video to be edited, a text description of facial expressions, and edited video speech, video features, text features, and edited speech features are extracted respectively; noisy video features are generated after masking the given editing time period of the video features, and the trained student rectification model takes the text features, the edited speech features, and the noisy video features as inputs to gradually generate an inverse sampling direction vector; the noisy video features are subjected to an inverse sampling process based on an ordinary differential equation of the direction vector to gradually obtain clean video features, and the edited video is obtained after decoding.
[0040] In a second aspect, the present invention provides a fast voice-driven facial video editing and generation system for implementing the above-mentioned fast voice-driven facial video editing and generation method.
[0041] The beneficial effects of the present invention are:
[0042] The present invention first trains a teacher rectification model that can generate an inverse sampling direction vector for gradually denoising noisy video features. Video masks of random length and position are used in the training process, and the video of the occluded position is predicted based on the video of the unoccluded part, laying a foundation for occluding part of the video time period during reasoning to achieve video generation and video editing. In addition, a student rectification model with the same structure and function as the teacher rectification model is obtained after running rectification learning target training based on the pairing of new noise generated by the teacher rectification model with generated features. The training method can improve the straightness of the ordinary differential equation trajectory, so the student rectification model can use a larger sampling step size without causing excessive errors, and can generate high-quality facial videos with a lower number of generation steps than the teacher rectification model to achieve fast voice-driven facial video generation and editing. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a flow chart of a fast voice-driven facial video editing and generation method proposed by the present invention;
[0044] Figure 2 It is a training schematic diagram of the teacher rectification model of the present invention;
[0045] Figure 3 is a schematic diagram of the present invention using the student rectification model to achieve the video generation task;
[0046] Figure 4 It is a schematic diagram of the present invention using the student rectification model to achieve the video editing task. DETAILED DESCRIPTION
[0047] The present invention is further described and illustrated below in conjunction with specific embodiments. The embodiments are merely exemplary of the present disclosure and do not define the scope of limitation. The technical features of each embodiment of the present invention may be combined accordingly without conflicting with each other.
[0048] The accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0049] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the steps. For example, some steps may be decomposed, while some steps may be combined or partially combined, so the actual execution order may change according to the actual situation.
[0050] refer to Figure 1The present invention proposes a fast voice-driven facial video editing and generation method, which mainly includes the following steps:
[0051] S1, obtain facial video data, facial expression text description, and video speech as a dataset.
[0052] The facial video data refers to a video containing a speaker's facial image; the facial expression text description refers to a description of the speaker's facial expression, such as "seriously speaking"; and the video voice refers to the voice corresponding to the facial video data.
[0053] In a specific implementation of the present invention, the step of resampling the facial video and the video voice is also included. The facial video is resampled to 25 frames per second as a reference video frame, and the voice data is resampled to 16000 Hz. The resampled data is used for the following calculations.
[0054] S2, extract video features using the encoder part of the pre-trained video variational autoencoder , extract text features using pre-trained text encoder , using pre-trained speech encoder to extract speech features .
[0055] In a specific implementation of the present invention, the pre-trained video encoder is implemented using the encoder part of the video variational autoencoder. Here, the video variational autoencoder structure is common knowledge in the art and will not be described in detail. In addition, the pre-trained text encoder is implemented using the existing Flan-T5-Large model, and the pre-trained speech encoder is implemented using the existing Wav2Vec2 model.
[0056] S3, creates a video time segment mask with random length of the occluded area, samples random time steps, linearly interpolates the video features of the occluded area with Gaussian noise, and generates noisy video features.
[0057] In a specific implementation of the present invention, the random occlusion length of the video time segment mask is lower than the video feature time length, and the sampling time step is sampled from a uniform [0, 1] distribution.
[0058] Specifically, the process of generating noisy video features is as follows:
[0059] Sample times from a uniform [0,1] distribution , , sample Gaussian noise from a Gaussian distribution with mean 0 and variance I , where I is the identity matrix.
[0060] Create a video time segment mask , The value of the randomly occluded area is 1, and the value of the unoccluded area is 0. Indicates the feature time length of the video. Indicates the video feature height, Indicates the width of the frequency feature; using mask Video features obtained by encoding the video encoder The corresponding occluded area is subjected to noise interpolation, namely:
[0061]
[0062] in, represents the sampling time step, represents the sampling time step The noise-added video features under Indicates the video time period mask. represents the original video features, represents Gaussian noise, represents dot product, Represents the video feature channel dimension.
[0063] S4, trains a teacher rectification model using noisy video features, text features, and speech features.
[0064] Here, the teacher rectification model consists of several fully connected layers, N stacked Transformer blocks and other main structures. The teacher rectification model can generate an inverse sampling direction vector for gradually denoising the noisy video features.
[0065] In a specific implementation of the present invention, the training process of the teacher rectification model is as follows Figure 2 As shown in the figure, some calculations are simplified for the convenience of representation.
[0066] First, two learnable fully connected layers are used to transform the speech features and noisy video features Map them to the same shape, connect the mapped speech features and noisy video features in the channel dimension, and the connection result is called speech and video features, whose channel dimension is the same as the hidden dimension of the Transformer block; use a learnable fully connected layer to map the text features to the hidden features of the same dimension as the Transformer block, and connect the mapped text features with the speech, video and text features in the time dimension to obtain the total features.
[0067] Specifically, for the obtained speech features, the nearest neighbor interpolation is first used to make them have the same time dimension length as the flattened noisy video features. The flattened noisy video features are recorded as , the length of the flattened time dimension is , then the nearest interpolated speech feature is:
[0068]
[0069] Two learnable fully connected layers are used to map the speech features and video features to the same shape:
[0070]
[0071]
[0072] The channel dimensions of the mapped speech features and the noisy video features are both half of the hidden dimension of the Transformer block, and after connecting them on the channel dimension, we get:
[0073]
[0074] The text features encoded by the text encoder are also mapped to the same hidden dimension as the Transformer block through the fully connected layer:
[0075]
[0076] Total features after time dimension connection: ;
[0077] in, represents the channel dimension of speech features, Indicates the length of the text feature sequence, represents the fully connected layer, represents the noise-added video feature, represents the flattened noisy video features, and represents the speech and noisy video features after being mapped by the fully connected layer, Represents text features, represents the text embedding feature after mapping by the fully connected layer, d represents the hidden dimension of the Transformer block, Represents the total characteristics after connection.
[0078] Secondly, the total features are input into the Transformer block for self-attention calculation.
[0079] Here, the calculation process of a layer of Transformer block is as follows:
[0080] Based on the total characteristics , the QKV matrix is obtained and self-attention is calculated as follows:
[0081]
[0082]
[0083]
[0084]
[0085] Then use the learnable feedforward network layer to convert the self-attention vector Perform linear mapping to keep the channel dimension unchanged;
[0086] The mapping result is used as the total feature of the next layer of Transformer block , after N attention calculations, the final Transformer block output is obtained.
[0087] Finally, since the output of the Transformer block has the dimension First, the channel dimension is mapped to the same dimension of the video feature channel through linear mapping, that is, it is mapped to , and then cut out the part corresponding to the length of the video feature sequence as the inverse sampling direction vector.
[0088] In one specific implementation of the present invention, the loss function for training the teacher rectification model is:
[0089]
[0090] in, represents the teacher rectification model loss, represents the teacher rectification model, represents the inverse sampling direction vector output by the teacher rectification model, represents the sampling time step, represents the original video features, represents the sampling time step The noise-added video features under Represents the voice features, Represents text features, Indicates the video time period mask. represents Gaussian noise, represents dot product, Represents the norm.
[0091] S5, define a student rectification model with the same structure as the teacher rectification model, initialize the student rectification model with the teacher rectification model parameters, and train the student rectification model with the reflux algorithm.
[0092] In a specific implementation of the present invention, the reflow algorithm training process is as follows:
[0093] The parameters of the student rectification model are recorded as , the parameters of the teacher rectification model obtained after training are recorded as , use the teacher rectification model parameters to initialize the student rectification model parameters, that is, .
[0094] Define the dataset as , which is composed of the video features, text features and speech features obtained in step S2; and, the teacher rectification model is defined as , the student rectification model is , the number of inverse sampling steps of the teacher rectification model is , then the inverse sampling generates the stride ;
[0095] Iterating over the dataset , taking out a small batch of samples each time, , execute the following process ae until the student rectification model converges.
[0096] a. Randomly set the video time period mask each time , cover part of the video, The length of the value is 1 in the uniform distribution uniform sampling, where Indicates the size of the time dimension of video features.
[0097] b. Sampling from Gaussian distribution and video features Noise of the same shape .
[0098] c. Initialization , then The masked parts are replaced with Gaussian noise:
[0099]
[0100] The teacher rectification model uses an Euler ordinary differential equation solver to denoise the parts occluded by the mask. The solver will run times until , which is expressed as follows:
[0101]
[0102]
[0103] in, represents the teacher rectification model, represents the inverse sampling direction vector output by the teacher rectification model, represents the noise-added video feature, Represents the voice features, Represents text features, Indicates the generation stride.
[0104] when When , the generation steps of the teacher rectification model are completed, and the original noisy video features After the inverse sampling process based on the ordinary differential equation of the direction vector, the clean video features are gradually obtained. , where the part blocked by the mask has been denoised.
[0105] d. Clean video features And the noisy video features at the predefined time step 0 as feature sample pairs.
[0106] Use the previously defined noise And the prediction results of the teacher rectification model As a new pairing, it is used to train the student rectification model, which can make the ordinary differential equation trajectory of the student rectification model straighter, thereby ensuring that the student rectification model can use a larger step size for low step number generation.
[0107] e. Train the student rectifier model:
[0108] After getting the clean video features predicted by the teacher rectification model Then, from the uniform distribution Medium Sampling ;
[0109] The occluded part of the video feature is Interpolate using noise:
[0110]
[0111] Calculate the training loss of the student rectifier model as the rectifier loss:
[0112]
[0113] in, represents the student rectification model loss, Represents the inverse sampling direction vector of the output of the student rectification model. Calculate the gradient according to the loss function and use the optimizer to adjust the parameters of the student rectification model to update.
[0114] Since the student rectification model uses the teacher rectification model parameters The student rectifier model parameters can converge quickly after initialization. Moreover, the trajectory of the ordinary differential equation of the student rectifier model is straighter after reflux training, so the video frames of the same quality can be generated with fewer steps than the teacher rectifier model, thereby realizing fast voice-driven face video editing and generation.
[0115] S6, generates a video based on a given facial image using the trained student rectification model, or edits a given facial video.
[0116] The facial video editing and generation method here refers to extracting video features based on a single-frame image or a video of the target object, splicing Gaussian noise or masking the position to be edited, combining audio features and text features of the text description of facial expressions, and using the trained rectification model to generate video features of the occluded area, and finally decoding it into the target video. The trained rectification model is the student rectification model, which is obtained after running the rectification learning target training based on the teacher rectification model.
[0117] like Figure 3 As shown in the figure, for the video generation task, the user provides a single-frame facial reference image, a text description of facial expressions, and video speech, and extracts video features, text features, and speech features respectively; Gaussian noise corresponding to the length of the video speech is spliced after the video features, and the trained student rectification model takes the text features, speech features, and video features after splicing Gaussian noise as input, and gradually generates an inverse sampling direction vector; the video features after splicing Gaussian noise are subjected to an inverse sampling process based on an ordinary differential equation of the direction vector, and clean video features are gradually obtained, and the video is generated after decoding.
[0118] like Figure 4 As shown in the figure, for the video editing task, the user provides a facial reference video to be edited, a text description of facial expressions and the edited video voice, and extracts the video features, text features and edited voice features respectively; after masking the given editing time period of the video features, the noisy video features are generated, where the given editing time period refers to the video segment to be edited specified by the user, and the corresponding video mask is obtained according to the video segment to be edited to generate the noisy video features; the trained student rectification model takes the text features, the edited voice features and the noisy video features as inputs, and gradually generates the inverse sampling direction vector; the noisy video features are subjected to the inverse sampling process of the ordinary differential equation based on the direction vector, and the clean video features are gradually obtained, and the edited video is obtained after decoding.
[0119] Based on the same inventive concept, a fast voice-driven facial video editing and generation system is also provided in this embodiment, which is used to implement the above-mentioned embodiment. The terms "module", "unit", etc. used below can implement a combination of software and / or hardware for predetermined functions. Although the system described in the following embodiments is preferably implemented in software, it is also possible to implement hardware, or a combination of software and hardware.
[0120] In this embodiment, a fast voice-driven facial video editing and generation system includes:
[0121] A data acquisition module is used to acquire facial video data, facial expression text description and video voice, and extract video features, text features and voice features as feature data sets respectively;
[0122] A teacher rectification model training module is used to generate noisy video features after masking the video features in the feature data set for random time periods, and to train a teacher rectification model using the noisy video features, text features, and speech features. The teacher rectification model can generate an inverse sampling direction vector for gradually denoising the noisy video features.
[0123] The student rectification model training module is used to train a student rectification model using a feature data set. In each training batch, the following steps are performed: the teacher rectification model generates an inverse sampling direction vector for the noisy video feature, and denoises the noisy video feature to generate a clean video feature based on the inverse sampling process of the ordinary differential equation of the inverse sampling direction vector. The clean video feature and its corresponding noisy video feature at a time step of 0 are used as feature sample pairs to regenerate the noisy video feature based on a random time step; the student rectification model runs the rectification learning objective for training based on the feature sample pairs and the regenerated noisy video feature;
[0124] The facial video editing and generation module is used to generate a video from a given facial image or to edit a given facial video using the trained student rectification model.
[0125] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment, and the implementation methods of the remaining modules will not be repeated here. The system embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.
[0126] The embodiments of the system of the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The system embodiments can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, the corresponding computer program instructions in the non-volatile memory are read into the memory by the processor of any device with data processing capabilities and run.
[0127] The above examples are only specific embodiments of the present invention. Obviously, the present invention is not limited to the above examples, and many variations are possible. All variations that can be directly derived or associated with the contents disclosed by a person skilled in the art should be considered as the protection scope of the present invention.
Claims
1. A fast voice-driven facial video editing and generation method, characterized in that: include: Obtain facial video data, facial expression text description and video voice, and extract video features, text features and voice features as feature data sets respectively; The video features and Gaussian noise in the feature data set are masked by random time periods to generate noisy video features, and a teacher rectification model is trained using the noisy video features, text features and speech features. The teacher rectification model can generate an inverse sampling direction vector for gradually denoising the noisy video features. A student rectification model is trained using the feature data set. In each training batch, the following steps are performed: the teacher rectification model generates an inverse sampling direction vector for the noisy video feature, and performs an inverse sampling process of an ordinary differential equation based on the inverse sampling direction vector to denoise the noisy video feature and generate a clean video feature; The clean video feature and its corresponding noisy video feature when the time step is 0 are used as feature sample pairs to regenerate the noisy video feature based on random time steps; The student rectification model is trained by running the rectification learning objective based on the feature sample pairs and the regenerated noisy video features; The trained student rectification model is used to generate a video from a given facial image, or to edit a given facial video.
2. The fast voice-driven facial video editing and generation method according to claim 1, characterized in that: The method of generating a noisy video feature after masking the video feature and Gaussian noise in the feature data set at random time periods includes: Create a video time segment mask with random length of the occluded area, perform linear interpolation on the video features of the occluded area and Gaussian noise at the sampling time step to generate noisy video features; The length of the random occluded region of the video time segment mask is less than the video feature time length, and the sampling time step is sampled from a uniform [0, 1] distribution.
3. The fast voice-driven facial video editing and generation method according to claim 2, characterized in that: The video features of the blocked area and the Gaussian noise are linearly interpolated at the sampling time step, which is expressed as: ; in, represents the sampling time step, represents the sampling time step The noise-added video features under Indicates the video time period mask. represents the original video features, represents Gaussian noise, Represents dot product.
4. The fast voice-driven facial video editing and generation method according to claim 3, characterized in that: In each training batch, the generation process of feature sample pairs is as follows: Initialize sampling time step ; Create a video time segment mask with random length of the occluded area, perform linear interpolation on the video features of the occluded area and Gaussian noise under the setting of the initial sampling time step, and generate the initialized noisy video features ; The teacher rectification model uses the Euler ordinary differential equation solver to run the inverse sampling process to denoise the parts blocked by the mask, and continuously iterates and updates until ,at this time ,in Represents clean video features; The inverse sampling process of the Euler ordinary differential equation solver is expressed as follows: ; ; in, represents the teacher rectification model, represents the inverse sampling direction vector output by the teacher rectification model, represents the noise-added video feature, Represents the voice features, Represents text features, Indicates the generation stride; Clean Video Features And the noisy video features when time step is 0 As feature sample pairs for reflow training of student model.
5. The fast voice-driven facial video editing and generation method according to claim 4, characterized in that: The regeneration of the noise-added video features based on random time steps refers to the clean video features The noised video features at time step 0 The noisy video features are regenerated after random time period masking.
6. The fast voice-driven facial video editing and generation method according to claim 4, characterized in that: The loss functions for training the teacher rectification model and the student rectification model are: ; ; in, represents the teacher rectification model loss, represents the student rectification model loss, represents the student rectifier model, represents the inverse sampling direction vector output by the student rectification model, represents the student model sampling time step The noise-added video features under Represents the norm.
7. The fast voice-driven facial video editing and generation method according to claim 1, characterized in that: The teacher rectification model and the student rectification model are rectification models with the same structure.
8. The fast voice-driven facial video editing and generation method according to claim 7, characterized in that: The teacher rectification model and the student rectification model include: The speech feature preprocessing layer is used to interpolate the speech features to make them have the same time dimension length as the flattened noisy video features; The first fully connected layer is used to map the interpolated speech features to a channel dimension that is half of the latent dimension of the Transformer block; The second fully connected layer is used to map the video features to a channel dimension that is half of the hidden dimension of the Transformer block; The third fully connected layer is used to map text features to the latent dimension of the Transformer block with a channel dimension; A concatenation layer is used to concatenate the output results of the first fully connected layer and the second fully connected layer in the channel dimension, and then concatenate the output results of the third fully connected layer in the time dimension; N stacked Transformer blocks, which are used to extract a coding sequence from the output of the concatenated layer, where the channel dimension of the coding sequence is the same as the channel dimension of the noisy video features; The reverse sampling direction vector output layer is used to crop the part corresponding to the noisy video feature from the encoded sequence as the reverse sampling direction vector.
9. The fast voice-driven facial video editing and generation method according to claim 1, characterized in that: The method of generating a video from a given facial image using the trained student rectification model, or editing a given facial video, includes: Given a single-frame facial reference image, a text description of facial expressions, and a video voice, video features, text features, and voice features are extracted respectively; Gaussian noise corresponding to the length of the video voice is spliced after the video features, and the trained student rectification model takes the text features, voice features, and video features after splicing Gaussian noise as inputs to gradually generate inverse sampling direction vectors; the video features after splicing Gaussian noise are subjected to an inverse sampling process based on ordinary differential equations of direction vectors to gradually obtain clean video features, and the video is generated after decoding; Alternatively, given a facial reference video to be edited, a text description of facial expressions, and edited video speech, video features, text features, and edited speech features are extracted respectively; noisy video features are generated after masking the given editing time period of the video features, and the trained student rectification model takes the text features, the edited speech features, and the noisy video features as inputs to gradually generate an inverse sampling direction vector; the noisy video features are subjected to an inverse sampling process based on an ordinary differential equation of the direction vector to gradually obtain clean video features, and the edited video is obtained after decoding.
10. A fast voice-driven facial video editing and generation system, used to implement the method of claim 1; characterized in that: The system comprises: A data acquisition module is used to acquire facial video data, facial expression text description and video voice, and extract video features, text features and voice features as feature data sets respectively; A teacher rectification model training module is used to generate noisy video features after masking the video features in the feature data set for random time periods, and to train a teacher rectification model using the noisy video features, text features, and speech features. The teacher rectification model can generate an inverse sampling direction vector for gradually denoising the noisy video features. The student rectification model training module is used to train a student rectification model using a feature data set. In each training batch, the following steps are performed: the teacher rectification model generates an inverse sampling direction vector for the noisy video feature, and denoises the noisy video feature to generate a clean video feature based on the inverse sampling process of the ordinary differential equation of the inverse sampling direction vector. The clean video feature and its corresponding noisy video feature at a time step of 0 are used as feature sample pairs to regenerate the noisy video feature based on a random time step; the student rectification model runs the rectification learning objective for training based on the feature sample pairs and the regenerated noisy video feature; The facial video editing and generation module is used to generate a video from a given facial image or to edit a given facial video using the trained student rectification model.
Citation Information
Patent Citations
Inverse transform sampling by ray tracing
CN115731118A
Knowledge distillation based neural network training method, device, and storage medium
WO2023212997A1