Model training method, electronic device and program product
Through the method of training the training video set and the synthetic video set, the problem that each target character in the prior art needs to independently fine-tune the diffusion model, improve the video generation efficiency, and realize the universality of the diffusion model.
Patent Information
- Application Number
- CN202510002420.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, each target character needs to independently fine-tune the diffusion model, resulting in the need to repeat the entire fine-tune training process when facing a new target character, and the video generation efficiency is low.
By obtaining the training video set, each training video contains only a single character, the training image frames and training audio features of the training video are obtained, the video is generated, and the diffusion model is trained based on the training video set and the synthetic video set to obtain the trained diffusion model.
The problem that each target character needs to independently fine-tune the diffusion model, which improves the efficiency of video generation. The trained diffusion model is universal and can adapt to the video generation of different characters.
Smart Images

Figure CN119940460A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to technical fields such as artificial intelligence, and in particular to a model training method, electronic equipment, and program product. Background Art
[0002] In recent years, generative artificial intelligence technology has made remarkable progress. As one of the important generative technologies, the diffusion model has attracted widespread attention for its high controllability and stability in generating images and videos, and has shown great application potential in many fields such as film and television production, virtual reality, and advertising creativity.
[0003] In the prior art, for each specific target person, a large amount of image data containing the person needs to be collected as a target picture set, and these pictures are used to train the diffusion model so that the diffusion model can learn the unique characteristics and style of the target person in order to generate video content related to the task.
[0004] However, since the diffusion model needs to be fine-tuned independently for each target person, it means that every time a new target person is encountered, the entire fine-tuning training process needs to be repeated, resulting in low video generation efficiency. Summary of the invention
[0005] The present disclosure provides a model training method, an electronic device, and a program product.
[0006] According to one aspect of the present disclosure, a model training method is provided, comprising:
[0007] Obtain a training video set, wherein each training video in the training video set contains only a single person;
[0008] Respectively obtain training image frames and training audio features of each training video in the training video set;
[0009] Generate videos according to the training image frames and training audio features of each training video to obtain a synthetic video set;
[0010] The diffusion model is trained according to the training video set and the synthetic video set to obtain a trained diffusion model.
[0011] According to the model training method of at least one embodiment of the present disclosure, for any training video in the training video set, respectively obtaining the training image frame and the training audio feature of each training video in the training video set includes:
[0012] Obtaining training image frames and training audio files from the training video;
[0013] Feature extraction is performed on the training audio file to obtain training audio features.
[0014] According to the model training method of at least one embodiment of the present disclosure, extracting features from the training audio file to obtain training audio features includes:
[0015] Preprocessing the training audio file to obtain processed audio;
[0016] Performing feature extraction on the processed audio to obtain extracted features;
[0017] The extracted features are post-processed to obtain training audio features.
[0018] According to the model training method of at least one embodiment of the present disclosure, the training of the diffusion model according to the training video set and the synthetic video set includes:
[0019] Respectively obtain text features and training video hidden features of each training video in the training video set;
[0020] Respectively obtain the synthetic video features of each synthetic video in the synthetic video set;
[0021] Adding noise to the latent features of the training video to obtain noisy latent features;
[0022] The diffusion model is trained according to the noisy latent features, the synthetic video features, the text features and the training audio features.
[0023] According to the model training method of at least one embodiment of the present disclosure, the step of respectively obtaining the text features and the hidden features of each training video in the training video set includes:
[0024] Obtaining a video description text for each training video in the training video set;
[0025] The text features of each training video are obtained according to the video description text of each training video.
[0026] According to the model training method of at least one embodiment of the present disclosure, the keys and values of the cross-frame attention module in the diffusion model are feature representations generated by cross-video associated frames of the training video set and the synthetic video set.
[0027] According to the model training method of at least one embodiment of the present disclosure, the loss function used in the process of training the diffusion model is a loss function corresponding to the target area of the human body.
[0028] The model training method according to at least one embodiment of the present disclosure further includes:
[0029] Acquire a target person image, a noise image, and a driving audio;
[0030] Generate an initial video according to the target person image and the driving audio;
[0031] Acquire initial text features, initial video features, and initial voice features of the initial video;
[0032] The noise image, initial text features, initial video features and initial speech features are input into the trained diffusion model to obtain a target video.
[0033] According to another aspect of the present disclosure, an electronic device is provided, comprising: a memory storing execution instructions; and a processor executing the execution instructions stored in the memory, so that the processor executes the model training method of any embodiment of the present disclosure.
[0034] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the model training method of any embodiment of the present disclosure is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings illustrate exemplary embodiments of the present disclosure and together with the description serve to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification.
[0036] Figure 1 This is a process of a model training method according to an embodiment of the present disclosure. Figure 1 .
[0037] Figure 2 yes Figure 1 Flowchart of the feature acquisition method in the model training method shown.
[0038] Figure 3 yes Figure 2 Flowchart of the feature processing method in the feature acquisition method shown.
[0039] Figure 4 yes Figure 1 Flowchart of the diffusion training method in the model training method shown.
[0040] Figure 5 yes Figure 4 Flowchart of the text feature acquisition method in the diffusion training method shown.
[0041] Figure 6 This is a process of a model training method according to an embodiment of the present disclosure. Figure 2 .
[0042] Figure 7It is a schematic block diagram of the structure of a model training device according to an embodiment of the present invention.
[0043] Figure 8 It is a schematic block diagram of the structure of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0044] The present disclosure is further described in detail below in conjunction with the accompanying drawings and examples. It is understood that the specific examples described herein are only used to explain the relevant content, rather than to limit the present disclosure. It should also be noted that, for ease of description, only the parts related to the present disclosure are shown in the accompanying drawings.
[0045] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in the present disclosure can be combined with each other. The technical solution of the present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0046] The present disclosure proposes a model training method, an electronic device, a readable storage medium and a computer program product. The present disclosure can be implemented by setting a model training software on electronic devices such as high-performance servers (such as rack servers, blade servers, etc.), workstation computers (such as professional graphics workstations, ordinary high-performance workstations, etc.).
[0047] For the convenience of description and to make the technical solutions of the specific implementation methods of the present disclosure easier to understand, before describing the model training method implemented in the present disclosure, the technical terms involved in the specific implementation methods of the present disclosure are explained as follows:
[0048] Diffusion Model is a type of generative model based on probability model.
[0049] The Cross-Frame Attention Module is an important component in the diffusion model for modeling the temporal dependencies between video frames.
[0050] Cross-video associated frames refer to a pair of frames from different videos that are associated through semantic or temporal information.
[0051] Figure 1 FIG. 1 is a schematic diagram showing the overall process of a model training method M100 according to an embodiment of the present disclosure. Figure 1 The model training method shown includes steps S110 to S140. The method can be executed by an electronic device such as a server.
[0052] Specifically, Figure 1 The model training method shown includes:
[0053] Step S110, obtaining a training video set.
[0054] In some embodiments of the present disclosure, each training video in the training video set in step S110 contains only a single person, and the training videos in the training video set can be obtained based on manual shooting, or can be obtained from a video platform, a public data set, etc. based on a screening function.
[0055] In particular, in order to improve the universality of the trained diffusion model, different videos in the training video set may contain different characters.
[0056] Step S120, respectively obtaining training image frames and training audio features of each training video in the training video set.
[0057] In some embodiments of the present disclosure, step S120 may use a random frame extraction method to obtain a training image frame of a training video, that is, randomly select a frame from the training video for extraction; step S120 may also use a time point extraction method to obtain a training image frame of a training video, that is, select a specific time point (such as the tth second or a specific time mark) of the training video to extract a frame; step S120 may also use a key frame extraction method to obtain a training image frame of a training video, that is, extract a key frame from the training video, etc. In particular, in order to improve the training effect of the diffusion model, the training image frame extracted by step S120 should be an image frame including a person.
[0058] The training audio features obtained through step S120 refer to various data representations extracted from the audio track in the training video, which can describe different aspects of the audio content, and generally include: time domain features (including waveform, zero crossing rate, amplitude, etc.), frequency domain features (including spectrum, Mel-frequency cepstral coefficients, etc.), pitch and rhythm features, speech features (including resonance peaks, speech speed and fluency, etc.), emotion and mood features, timing features (including rhythm and speed, dynamic range, etc.), noise and background features (including signal-to-noise ratio, environmental noise features, etc.).
[0059] Step S130 , generating videos according to the training image frames and training audio features of each training video to obtain a synthetic video set.
[0060] In some embodiments of the present disclosure, step S130 may generate an initial video using an existing video generation model that can generate a video that matches the facial expression of the target person through audio drive, and the video generation model may be SadTalker or the like.
[0061] Step S140: training the diffusion model according to the training video set and the synthetic video set to obtain a trained diffusion model.
[0062] In some embodiments of the present disclosure, a diffusion model is trained based on a training video set and a synthetic video set, which can effectively learn the visual differences and similarities between the training video set and the synthetic video set. The trained diffusion model can not only capture the temporal consistency and visual features in the training video, but also adapt to changes in the synthetic video, ensuring that the model can handle noise, distortion or artifacts in the synthesis process, so that higher realism and consistency can be maintained when generating new videos through the trained diffusion model.
[0063] The model training method provided by the present invention has universal applicability because the training video set is not set for the target person. This solves the problem in the prior art that the diffusion model needs to be fine-tuned independently for each target person, which means that the entire fine-tuning training process needs to be repeated every time a new target person is encountered, resulting in low video generation efficiency.
[0064] Regarding step S120, in some embodiments of the present disclosure, for any training video in the training video set, the following may be included: Figure 2 Steps S121 to S122 are shown.
[0065] Step S121, obtaining a training image frame and a training audio file from the training video.
[0066] In some embodiments of the present disclosure, the training image frame obtained from the training video in step S121 may be the first frame of the training video. The method of obtaining the training audio file from the training video in step S121 may be to extract the training audio file from the training video using an audio device such as a recording card, an audio separator, etc.; or to extract the training audio file from the training video using a signal separation method; or to extract the training audio file from the training video using an audio and video processing tool such as ffmpeg.
[0067] The training audio file obtained in step S121 may be in various formats, such as .wav, .mp3, etc.
[0068] Step S122: extract features from the training audio file to obtain training audio features.
[0069] In some embodiments of the present disclosure, step S122 may select a suitable feature extraction strategy according to the analysis objective. For example, when the training audio features include time domain features, the process of extracting features from the training audio file in step S120 may include time domain analysis; when the training audio features include frequency domain features, the process of extracting features from the training audio file in step S120 may include frequency domain analysis, etc.
[0070] By extracting training audio features from the training audio file through steps S121 and S122, the core information of the training audio file can be effectively summarized, which is convenient for further processing and analysis.
[0071] Regarding step S122, in some embodiments of the present disclosure, it may include the following steps: Figure 3 Steps S1221 to S1223 are shown.
[0072] Step S1221, pre-processing the training audio file to obtain processed audio.
[0073] In some embodiments of the present disclosure, the purpose of preprocessing the training audio file in step S1221 is to standardize, denoise and transform the training audio file to ensure that the processed audio is suitable for subsequent feature extraction. The preprocessing in step S1221 may include denoising and filtering, framing and windowing, normalization and other processes.
[0074] Step S1222, performing feature extraction on the processed audio to obtain extracted features.
[0075] Step S1223, post-processing the extracted features to obtain training audio features.
[0076] In some embodiments of the present disclosure, the purpose of post-processing in step S1223 is to organize, optimize and format the extracted features to adapt to different application scenarios; the post-processing in step S1223 may include feature selection and dimensionality reduction, feature normalization, data smoothing, feature encoding and other processes.
[0077] By performing feature extraction from step S1221 to step S1223, it can be ensured that the training audio features extracted from the training audio file are efficient and reliable, and are suitable for a variety of application scenarios.
[0078] Regarding step S140, in some embodiments of the present disclosure, it may include: Figure 4 Steps S141 to S144 are shown.
[0079] Step S141, respectively obtaining text features and training video latent features of each training video in the training video set.
[0080] In some embodiments of the present disclosure, the text features of the training video in step S141 refer to features in text form that can be analyzed and represented based on the training video; the text features can be obtained through subtitles, scene text, audio of the training video, etc. Taking the acquisition of text features through audio as an example, the acquisition process may include: first extracting audio from the training video; secondly transcribing the audio into text; again preprocessing the text (such as correcting wrong words, etc.) to obtain processed text; and finally encoding the processed text into text features. Among them, the encoding process is implemented using term frequency-inverse document frequency (TermFrequency-Inverse Document Frequency, TF-IDF), text vector model, deep learning model (such as BERT, GPT, etc.), etc. Step S141 can also obtain the text features of the training video through a text feature extraction model, etc.
[0081] The hidden features of the training video in step S141 refer to deep features learned from the training video that cannot be directly observed; step S141 can be obtained by processing the training video through the encoding structure of a pre-trained variational autoencoder (VAE).
[0082] Step S142, respectively obtaining the synthetic video features of each synthetic video in the synthetic video set.
[0083] In some embodiments of the present disclosure, the method of obtaining the synthetic video features in step S142 may adopt feature extraction based on frames, audio or actions, etc. Taking the frame-based feature extraction method as an example, the process of obtaining the synthetic video features in step S142 may include: extracting image features from each frame of the synthetic video through a convolutional neural network (such as a multimodal model CLIP); obtaining the temporal features between video frames according to the image features; and summarizing the temporal features to obtain the synthetic video features.
[0084] Step S143, adding noise to the latent features of the training video to obtain noisy latent features.
[0085] In some embodiments of the present disclosure, step S143 can randomly or gradually add noise to the latent features of the training video, so that the diffusion model can learn how to recover the original information of the video from the noise. The noise added to the latent features of the video is usually Gaussian noise.
[0086] Step S144, training the diffusion model according to the noisy latent features, the synthetic video features, the text features and the training audio features.
[0087] In some embodiments of the present disclosure, in step S144, during the process of training the diffusion model according to the noisy latent features, synthetic video features, text features and training audio features, a loss function can be calculated according to the predicted noise and the actually added noise to control the training of the diffusion model.
[0088] During the training of the diffusion model, the encoder module is used to encode the video into latent features in the training phase, the decoder module is used to convert the latent features into videos in the inference phase, and the U-net module is used to predict and denoise to obtain high-quality videos. Specifically, the fine-tuning process is to feed a batch of data from noisy latent features, synthetic video features, text features, and training audio features into the U-net module to predict the added noise; calculate the loss function based on the predicted results and the actual noise; perform reverse gradient propagation of the loss function in the U-net module, update the parameters of the U-net module, and finally complete the training of the diffusion model.
[0089] Since most lip movements are highly correlated with speech signals, training the diffusion model using the training audio features as control conditions through steps S141 to S144 can improve the generation quality and authenticity of the trained model.
[0090] Regarding step S141, in some embodiments of the present disclosure, it may include the following: Figure 5 Steps S1411 and S1412 are shown.
[0091] Step S1411, obtaining the video description text of each training video in the training video set.
[0092] In some embodiments of the present disclosure, step S1411 can obtain video description text through an image annotation model, Transformer, etc., or can obtain video description text by using metadata (such as title, tags, etc.) of a training video, or can obtain video description text from a database, where the video description text stored in the database is pre-generated by the above method or manually.
[0093] Step S1412, obtaining text features of each training video according to the video description text of each training video.
[0094] In some embodiments of the present disclosure, step S1412 may use TF-IDF, a text vector model, a deep learning model, etc. to obtain text features.
[0095] Acquiring text features based on the video description text through steps S1411 and S1412 can significantly improve the video analysis and generation capabilities.
[0096] In some embodiments of the present disclosure, the keys and values of the cross-frame attention module in the diffusion model are feature representations generated by cross-video associated frames of the training video set and the synthetic video set.
[0097] Taking the calculation of the feature representation of the i-th frame as an example, its key and value specifically include two parts, one part is the key and value obtained by processing the features of the i-1th frame of a training video in the training video set; the other part is the key and value obtained by processing the features of the i-th frame of the corresponding synthetic video in the synthetic video set. Among them, the corresponding synthetic video in the synthetic video set is the synthetic video generated based on the training video.
[0098] This method uses the frames corresponding to the synthetic video as generation conditions, which improves the consistency between frames.
[0099] In some embodiments of the present disclosure, the loss function used in the process of training the diffusion model is a loss function corresponding to the target area of the human body.
[0100] In the above method, the human target area can be pre-set; since the face and the mouth shape are the key focus areas in the audio-driven target person video generation task, the face and the mouth shape can be used as the human target area.
[0101] The formula for calculating the loss function corresponding to the human target area is as follows:
[0102]
[0103] Among them, M represents the segmentation mask of the human target area, and (1-M) represents the area opposite to M. λ1 and λ2 are model hyperparameters that control the weights of the human target area and the background area.
[0104] The above method can enable the diffusion model to better focus on areas that are more correlated with speech features, thereby allowing the diffusion model to learn better.
[0105] Further, after step S140, the model training method provided by the present disclosure may also include: Figure 6 Steps S150 to S180 are shown.
[0106] Step S150, obtaining a target person image, a noise image and a driving audio.
[0107] In some embodiments of the present disclosure, the driving audio in step S150 refers to the audio used to control or drive the target person's performance in the video in the audio-driven video generation task; the driving audio usually includes voice information, intonation, rhythm, pauses, emotions and other elements, which are used to determine the target person's mouth shape, facial expressions, movements and other speech-related dynamic changes in the video.
[0108] In step S150, the noisy image refers to an image without a clear target or containing random, irrelevant information, and these images do not contain features that are helpful for the task.
[0109] Step S160, generating an initial video according to the target person image and the driving audio.
[0110] In some embodiments of the present disclosure, step S160 may generate an initial video using an existing video generation model that can be driven by audio to generate a video that matches the facial expression of the target person.
[0111] Step S170, obtaining initial text features, initial video features, and initial voice features of the initial video.
[0112] In some embodiments of the present disclosure, the method for obtaining the initial text features in step S170 may be based on text prompts; the text prompts may be constructed manually, by models, etc., based on the initial video; the method for obtaining the initial voice features in step S170 may be to extract features from the driving audio to obtain the initial voice features. The feature extraction process is similar to Figure 2 The method of obtaining the initial video features in step S170 can be based on frame, audio or action feature extraction, which is similar to step S122 shown in FIG. Figure 4 The step S142 shown is similar.
[0113] Step S180, inputting the noise image, initial text features, initial video features and initial speech features into the trained diffusion model to obtain a target video.
[0114] In some embodiments of the present disclosure, the specific process of obtaining the target video through step S180 is to analyze the noise image, initial text features, initial video features and initial speech features through the encoding structure of the trained diffusion model, and after obtaining the latent features of the initial video, decode the latent features of the initial video into the target video through the decoding structure of the pre-trained VAE.
[0115] The model training method provided in the present disclosure can be used in audio-driven target person video generation tasks after the training model.
[0116] Based on any of the above embodiments, the present disclosure also provides a model training device.
[0117] Figure 7 It is a schematic block diagram of the structure of a model training device according to an embodiment of the present invention.
[0118] like Figure 7 As shown, the model training device includes:
[0119] The training set acquisition module 110 is used to acquire a training video set, where each training video in the training video set contains only a single person.
[0120] The feature acquisition module 120 is used to respectively acquire the training image frame and training audio features of each training video in the training video set.
[0121] The video synthesis module 130 is used to generate videos according to the training image frames and training audio features of each training video to obtain a synthesized video set.
[0122] The model training module 140 is used to train the diffusion model according to the training video set and the synthetic video set to obtain a trained diffusion model.
[0123] The above-mentioned model training device can be in the form of computer software, and each module of the above-mentioned model training device can be implemented by a computer software module.
[0124] The executor of the model training method in the specific implementation manner of the present disclosure may be an electronic device such as a server.
[0125] Therefore, based on any of the above embodiments, the present disclosure also provides an electronic device, which can execute the model training method of any of the embodiments described above in the present disclosure.
[0126] Figure 8 1 is a schematic block diagram of the structure of an electronic device 1000 according to an embodiment of the present disclosure.
[0127] The hardware structure of the electronic device 1000 can be implemented using a bus architecture. The bus architecture can include any number of interconnected buses and bridges, depending on the specific application and overall design constraints of the hardware. The bus 1100 connects various circuits including one or more processors 1200, memory 1300 and / or hardware modules together. The bus 1100 can also connect various other circuits 1400 such as peripherals, voltage regulators, power management circuits, external antennas, etc.
[0128] The bus 1100 may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Component (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the figure only uses one connecting line, but does not mean that there is only one bus or one type of bus.
[0129] The present disclosure also provides a readable storage medium, in which a computer program is stored, and the computer program is used to implement the above method when executed by a processor. "Readable storage medium" can be any device that can contain, store, communicate, propagate or transmit a program for use in an instruction execution system, device or equipment or in combination with these instruction execution systems, devices or equipment. More specific examples of readable storage media include the following: an electrical connection portion (electronic device) with one or more wirings, a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and editable read-only memory (EPROM or flash memory), an optical fiber device, and a portable read-only memory (CDROM), etc.
[0130] The present disclosure also provides a computer program product. The method of the present disclosure can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instruction is loaded and executed, the process or function of the present disclosure is executed in whole or in part.
[0131] The computer program or instructions may be stored in a readable storage medium or transmitted from one readable storage medium to another readable storage medium, for example, the computer program or instructions may be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired or wireless means. The readable storage medium may be any available medium that can be accessed or a data storage device such as a server, data center, etc. that integrates one or more available media. The available medium may be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it may also be an optical medium, such as a digital video disk; it may also be a semiconductor medium, such as a solid state drive. The computer readable storage medium may be a volatile or non-volatile storage medium, or may include both volatile and non-volatile types of storage media.
[0132] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0133] The present disclosure is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0134] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0135] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0136] In the description of this specification, the description with reference to the terms "one embodiment / method", "some embodiments / methods", "example", "specific example", or "some examples" etc. means that the specific features, structures, or characteristics described in conjunction with the embodiment / method or example are included in at least one embodiment / method or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment / method or example. Moreover, the specific features, structures, or characteristics described may be combined in a suitable manner in any one or more embodiments / methods or examples. In addition, those skilled in the art may combine and combine different embodiments / methods or examples described in this specification and features of different embodiments / methods or examples, unless they are contradictory.
[0137] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of the present disclosure, "plurality" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0138] Those skilled in the art should understand that the above embodiments are only for the purpose of clearly illustrating the present disclosure, and are not intended to limit the scope of the present disclosure. For those skilled in the art, other changes or modifications may be made based on the above disclosure, and these changes or modifications are still within the scope of the present disclosure.
Claims
1. A model training method, characterized in that: include: Obtain a training video set, wherein each training video in the training video set contains only a single person; Respectively obtain training image frames and training audio features of each training video in the training video set; Generate videos according to the training image frames and training audio features of each training video to obtain a synthetic video set; The diffusion model is trained according to the training video set and the synthetic video set to obtain a trained diffusion model.
2. The model training method according to claim 1, characterized in that: For any training video in the training video set, respectively obtaining the training image frame and the training audio feature of each training video in the training video set includes: Obtaining training image frames and training audio files from the training video; Feature extraction is performed on the training audio file to obtain training audio features.
3. The model training method according to claim 2, characterized in that: The step of extracting features from the training audio file to obtain training audio features includes: Preprocessing the training audio file to obtain processed audio; Performing feature extraction on the processed audio to obtain extracted features; The extracted features are post-processed to obtain training audio features.
4. The model training method according to claim 1, characterized in that: The step of training the diffusion model according to the training video set and the synthetic video set includes: Respectively obtain text features and training video hidden features of each training video in the training video set; Respectively obtain the synthetic video features of each synthetic video in the synthetic video set; Adding noise to the latent features of the training video to obtain noisy latent features; The diffusion model is trained according to the noisy latent features, the synthetic video features, the text features and the training audio features.
5. The model training method according to claim 4, characterized in that: The step of respectively obtaining the text features and hidden features of each training video in the training video set includes: Obtaining a video description text for each training video in the training video set; The text features of each training video are obtained according to the video description text of each training video.
6. The model training method according to any one of claims 1 to 5, characterized in that: The keys and values of the cross-frame attention module in the diffusion model are feature representations generated by cross-video correlation frames of the training video set and the synthetic video set.
7. The model training method according to any one of claims 1 to 5, characterized in that: The loss function used in the process of training the diffusion model is the loss function corresponding to the human target area.
8. The model training method according to any one of claims 1 to 5, characterized in that: Also includes: Acquire a target person image, a noise image, and a driving audio; Generate an initial video according to the target person image and the driving audio; Acquire initial text features, initial video features, and initial voice features of the initial video; The noise image, initial text features, initial video features and initial speech features are input into the trained diffusion model to obtain a target video.
9. An electronic device, characterized in that: include: A memory storing execution instructions; as well as A processor, wherein the processor executes the execution instructions stored in the memory so that the processor executes the model training method according to any one of claims 1 to 8.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the model training method described in any one of claims 1 to 8 is implemented.