Audio and video joint generation model training method and device and audio and video joint generation method and device

By employing a training method for a joint audio-video generation model, and utilizing a diffusion transformer architecture and a cross-attention mechanism to fuse video and audio tag information, this approach addresses the issues of existing models lacking audio generation capabilities and having temporal misalignment, thereby achieving efficient and synchronized audio-video generation.

CN121888048APending Publication Date: 2026-04-17BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2025-12-10
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing video generation models lack audio generation capabilities, which limits the generation scenarios. Furthermore, models with audio and video co-generation capabilities suffer from increased processing time and timing misalignment due to their serial execution method, resulting in reduced generation quality.

Method used

An audio-video joint generation model training method is adopted. By acquiring video and audio training data, the audio branch is initially trained and then jointly trained. The audio-video interaction is realized by using a diffusion transformer architecture, and cross-attention mechanism fusion is performed by combining video and audio label information.

Benefits of technology

It enables the simultaneous generation of high-quality audio and video, reduces overall processing time, avoids audio and video timing misalignment, and improves generation quality and consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121888048A_ABST
    Figure CN121888048A_ABST
Patent Text Reader

Abstract

The invention provides an audio and video joint generation model training method and device and an audio and video joint generation method and device, relates to the artificial intelligence field of computer vision, deep learning, large models, natural language processing and the like, and can be applied to scenes of digital human, artificial intelligence-based content generation and the like. The method comprises the following steps: acquiring video training data and audio training data; performing preliminary training on the audio branches according to the audio training data, wherein the audio and video joint generation model comprises the audio branches and the video branches; and in response to determining that the preliminary training is completed, performing joint training on the audio branch and the video branch according to the video training data and the audio training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, particularly to computer vision, deep learning, large models and natural language processing, and especially to audio and video joint generation model training and audio and video joint generation methods and apparatus. Background Technology

[0002] Current video generation models typically lack audio generation capabilities, severely limiting their application scenarios. Models capable of jointly generating audio and video often have redundant structures, such as generating audio after video generation, or vice versa. Regardless of the form, the serial execution method increases overall processing time and may lead to audio-video timing misalignment, thus reducing generation quality. Summary of the Invention

[0003] This disclosure provides training of audio-video joint generation models and methods and apparatus for audio-video joint generation.

[0004] A method for training an audio-video joint generation model includes:

[0005] Acquire video and audio training data;

[0006] The audio branch is initially trained based on the audio training data, and the audio-video joint generation model includes the audio branch and the video branch.

[0007] In response to determining that the initial training is complete, the audio branch and the video branch are jointly trained based on the video training data and the audio training data.

[0008] An audio-video joint generation method, comprising:

[0009] Obtain the audio / video generation requirement description information and the fourth reference audio;

[0010] The audio and video generation requirement description information and the fourth reference audio input audio and video are combined to form a joint generation model to obtain the output target audio and video content. The audio in the target audio and video content conforms to the acoustic characteristics of the fourth reference audio. The joint audio and video generation model is trained using the above-described training method.

[0011] An audio-video joint generation model training device includes: a data acquisition module, a first training module, and a second training module;

[0012] The data acquisition module is used to acquire video training data and audio training data;

[0013] The first training module is used to perform preliminary training on the audio branch based on the audio training data, and the audio-video joint generation model includes the audio branch and the video branch;

[0014] The second training module is configured to, in response to determining that the initial training is complete, jointly train the audio branch and the video branch based on the video training data and the audio training data.

[0015] An audio-visual co-generation device includes: an information acquisition module and a content generation module;

[0016] The information acquisition module is used to acquire audio and video generation requirement description information and fourth reference audio;

[0017] The content generation module is used to combine the audio and video generation requirement description information and the fourth reference audio input audio and video to jointly generate a model, thereby obtaining the output target audio and video content. The audio in the target audio and video content conforms to the acoustic characteristics of the fourth reference audio. The audio and video joint generation model is trained using the above-mentioned training method.

[0018] An intelligent agent includes: an input module, a processing module, and an output module;

[0019] The input module is used to receive audio and video generation requirement description information and fourth reference audio;

[0020] The processing module is used to combine the audio and video generation requirement description information and the fourth reference audio input audio and video joint generation model to obtain the generated target audio and video content. The audio in the target audio and video content conforms to the acoustic characteristics of the fourth reference audio. The audio and video joint generation model is trained using the above-mentioned training method.

[0021] The output module is used to output the target audio and video content.

[0022] An electronic device, comprising:

[0023] At least one processor; and

[0024] A memory communicatively connected to the at least one processor; wherein,

[0025] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described above.

[0026] A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the methods described above.

[0027] A computer program product includes a computer program / instructions that, when executed by a processor, implement the method described above.

[0028] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0029] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0030] Figure 1 This is a flowchart illustrating an embodiment of the audio-video joint generation model training method described in this disclosure;

[0031] Figure 2 This is a flowchart illustrating an embodiment of the method for generating video training data and audio data training as described in this disclosure;

[0032] Figure 3 This is a schematic diagram illustrating the process of generating the first fusion result as described in this disclosure;

[0033] Figure 4 This is a flowchart of an embodiment of the method for generating third input information as described in this disclosure;

[0034] Figure 5 This is a flowchart of an embodiment of the audio and video joint generation method described in this disclosure;

[0035] Figure 6 This is a schematic diagram of the input and output of the audio and video joint generation model 600 described in this disclosure;

[0036] Figure 7 This is a schematic diagram of the composition structure of Embodiment 700 of the audio-video joint generation model training device described in this disclosure;

[0037] Figure 8 This is a schematic diagram of the composition structure of Embodiment 800 of the audio and video joint generation device described in this disclosure;

[0038] Figure 9 This is a schematic diagram of the composition structure of the intelligent agent embodiment 900 described in this disclosure;

[0039] Figure 10 A schematic block diagram of an electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0040] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0041] Furthermore, it should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0042] Figure 1 This is a flowchart illustrating an embodiment of the audio-video joint generation model training method described in this disclosure. Figure 1 As shown, the specific implementation methods are as follows.

[0043] In step 101, video training data and audio training data are acquired.

[0044] In step 102, the audio branch is initially trained based on the audio training data. The audio-video joint generation model includes an audio branch and a video branch.

[0045] In step 103, in response to determining that the initial training is complete, the audio branch and the video branch are jointly trained based on the video training data and the audio training data.

[0046] By employing the scheme described in the above-described method embodiments, an audio-video joint generation model capable of simultaneously generating audio and video can be trained. This model is based on a dual-tower architecture for audio and video generation, thereby expanding the application scenarios of the model. Furthermore, by initially training the audio branch, the audio branch already possesses independent and high-quality audio generation capabilities before entering joint training, thus laying a solid foundation for subsequent joint training. Moreover, through joint training, the model can simultaneously process and learn from video and audio training data, enabling the model to generate audio and video in parallel during inference, rather than sequentially, thereby reducing overall time consumption and minimizing the problem of audio-video timing misalignment, thus improving the quality of generated audio and video content.

[0047] To train the audio-video joint generation model, it is necessary to first obtain video training data and audio training data.

[0048] In some embodiments of this disclosure, an original video and a first original audio can be obtained, and the original video and the first original audio can be preprocessed. For example, the preprocessing may include: discarding the original video and the first original audio with a duration less than a predetermined first duration, then extracting the second original audio from the preprocessed original video, and then generating video training data based on the preprocessed original video, and generating audio training data based on the preprocessed first original audio and the second original audio.

[0049] The sources of the original video and the first original audio are not restricted. For example, original video and original audio can be collected from predetermined websites and open-source communities. To distinguish them from other audio, the collected original audio is referred to as the first original audio. In addition, each original video and each first original audio typically corresponds to only one speaker.

[0050] For the collected raw video and first raw audio, preprocessing can be performed first. This includes discarding raw video and first raw audio with a duration shorter than a predetermined first duration, the specific value of which can be determined according to actual needs, such as 5 seconds. Additionally, the preprocessing may include other steps, such as discarding raw video with audio-visual desynchronization, discarding raw video with a duration longer than a second duration, and discarding raw video and first raw audio containing sensitive content, where the second duration is longer than the first. Preprocessing reduces the workload of subsequent processing and improves the quality of the subsequently generated training data.

[0051] In addition, audio can be extracted from the preprocessed original video (which can be called the second original audio), thereby supplementing the amount of audio training data in the initial training phase and providing the necessary audio training data for subsequent joint training.

[0052] Based on the preprocessed original video, video training data can be generated. In some embodiments of this disclosure, for any original video, in response to determining that the duration of the original video is equal to a first duration, the original video can be determined as video training data; in response to determining that the duration of the original video is greater than the first duration, the original video can be segmented in a manner that segments every first duration, and the segmented video segments of the first duration can be determined as video training data.

[0053] For example, assuming the first duration is 5 seconds, for a certain original video, if the original video duration is 5 seconds, then the original video can be directly identified as video training data. If the original video duration is 10 seconds, then the original video can be divided into two 5-second video segments, and the two 5-second video segments can be identified as video training data respectively. If the original video duration is 13 seconds, then the original video can be divided into two 5-second video segments and one 3-second video segment (less than 5 seconds), and the two 5-second video segments can be identified as video training data.

[0054] The above processing can standardize the duration of video training data and effectively increase the number of video training samples.

[0055] Audio training data can be generated based on the preprocessed first and second original audio files. In some embodiments of this disclosure, the first and second original audio files with a duration equal to a first duration can be determined as audio training data. Alternatively, for any second original audio file with a duration greater than the first duration, the second original audio file can be segmented at intervals of the first duration, and the resulting audio segments of the first duration can be determined as audio training data. Furthermore, for any first original audio file with a duration greater than the first duration, the first original audio file can be directly determined as audio training data. Or, in response to the determination that the first original audio file needs to be segmented, the first original audio file can be segmented at intervals of the first duration, and the resulting audio segments of the first duration can be determined as audio training data.

[0056] In other words, for a given second original audio file, if its duration is the same as the first duration, then the second original audio file can be directly identified as audio training data. Otherwise, it can be segmented into multiple audio segments, and the segmented audio segments with a duration equal to the first duration can be identified as audio training data. Similarly, for a given first original audio file, if its duration is the first duration, then the first original audio file can be directly identified as audio training data. Otherwise, the first original audio file can be identified as audio training data. Alternatively, in practical applications, to supplement the number of short audio files mentioned later, some first original audio files with a duration greater than the first duration can also be segmented. Correspondingly, if the first original audio file is a selected (e.g., randomly selected) first original audio file that needs to be segmented, then it can be segmented into multiple audio segments, and the segmented audio segments with a duration equal to the first duration can be identified as audio training data.

[0057] In some embodiments of this disclosure, video tag information corresponding to video training data can also be obtained. The video tag information may include: video description information of the corresponding video training data, audio description information of the audio in the corresponding video training data, and text content information after the audio is converted into text. Audio tag information corresponding to audio training data can also be obtained. The audio tag information may include: audio description information of the corresponding audio training data and text content information after the audio is converted into text.

[0058] For example, for video training data, multimodal large model inference can be used to obtain the corresponding video label information, and for audio training data, multimodal large model inference can be used to obtain the corresponding audio label information, so as to assist in the subsequent training of audio and video branches.

[0059] Video description information refers to a summary textual description of the video content, which can cover core visual elements such as the subject, actions, and scenes in the video. Audio description information refers to descriptions of the acoustic features of the audio, such as the speaker's timbre, tone, and speech rate. Text content information refers to the text content recognition results obtained after performing speech recognition on the audio.

[0060] Based on the above introduction, Figure 2 This is a flowchart illustrating an embodiment of the method for generating video training data and audio data training as described in this disclosure. Figure 2 As shown, the specific implementation methods are as follows.

[0061] In step 201, the original video and the first original audio are acquired.

[0062] In step 202, the original video and the first original audio are preprocessed.

[0063] In step 203, the second original audio is extracted from the preprocessed original video.

[0064] In step 204, video training data is generated based on the preprocessed original video, and audio training data is generated based on the preprocessed first and second original audio.

[0065] In step 205, the video label information corresponding to the video training data is obtained, and the audio label information corresponding to the audio training data is obtained.

[0066] Once video and audio training data are generated, the audio-video joint generation model can be trained using these data.

[0067] First, the audio branches can be initially trained based on the audio training data. In some embodiments of this disclosure, each audio training data can correspond to a speaker. Accordingly, the audio training data can be clustered, and the audio training data in the same clustering result correspond to the same speaker. Then, the clustering result that meets the following requirements can be determined as the target result: the clustering result includes at least two audio training data, and the clustering result includes both long audio and short audio. The long audio is audio training data with a duration greater than a first duration, and the short audio is audio training data with a duration equal to the first duration. Then, the audio branches can be initially trained based on the long audio and short audio.

[0068] There are no restrictions on how clustering is performed; for example, clustering can be achieved through voiceprint feature extraction and clustering algorithms.

[0069] From the multiple clustering results obtained, the target result that meets the above requirements can be selected. The target result must include at least two audio training data, and must include both long and short audio.

[0070] Accordingly, in some embodiments of this disclosure, the initial training of the audio branch may include: a first-stage training and a second-stage training, wherein the first-stage training of the audio branch may be performed based on long audio clips, and the second-stage training of the audio branch may be performed based on short audio clips.

[0071] As can be seen, the initial training can be divided into two stages, and each stage can use different types of audio training data, thereby improving the targeting of the training, and thus improving the training efficiency and training effect.

[0072] In some embodiments of this disclosure, the method of performing a first-stage training of an audio branch based on a long audio segment may include: extracting a long audio segment from any target result and determining it as a first target audio segment; extracting an audio training data segment different from the first target audio segment from the same target result and determining it as a first reference audio segment; determining first input information based on the first target audio segment and the first reference audio segment; inputting the first input information into the audio branch to obtain a first output result; determining a first loss based on the first output result; and updating the parameters of the audio branch based on the first loss.

[0073] The first target audio and the first reference audio come from the same target result. The first target audio can be a randomly selected long audio, and the first reference audio can be a long audio or a short audio, but it needs to be different from the first target audio.

[0074] The first input information can be determined based on the first target audio and the first reference audio. In some embodiments of this disclosure, the first target audio can be encoded to obtain a first target encoding result, and the first target encoding result can be superimposed with first noise to obtain a first noisy encoding result. The first reference audio can also be encoded to obtain a first reference encoding result. Then, the first noisy encoding result and the first reference encoding result can be fused to obtain a first fusion result, and the first input information can be determined based on the first fusion result. For example, the first noisy encoding result and the first reference encoding result can be fused by addition.

[0075] Figure 3 This is a schematic diagram illustrating the process of generating the first fusion result as described in this disclosure. Figure 3 As shown, for the first target audio and the first reference audio, if necessary, they can be length-uniformed by a patch operation. Then, they can be encoded separately using an audio encoder. For example, patch embedding can be used to encode the first target audio and the first reference audio separately, thus obtaining the first target encoding result and the first reference encoding result respectively. For the first target encoding result, first noise can be superimposed to it to obtain the first noisy encoding result. The first noise can be randomly generated. Furthermore, the first noisy encoding result and the first reference encoding result can be fused by addition to obtain the first fused result. Compared with the traditional concatenation method, using addition to fuse the first noisy encoding result and the first reference encoding result can avoid changing the network input dimension due to the concatenation method, thus causing changes in the model parameter structure, and also ensuring the symmetry of the audio and video branches.

[0076] In some embodiments of this disclosure, the method of determining the first input information based on the first fusion result may include: obtaining a first auxiliary encoding result, the first auxiliary encoding result including at least: the result obtained after encoding the audio tag information corresponding to the first target audio, and determining the first fusion result and the first auxiliary encoding result as the first input information.

[0077] For example, the audio tag information corresponding to the first target audio can be encoded using a text encoder to obtain a first auxiliary encoding result. In practical applications, if necessary, the first auxiliary encoding result may also include the encoded result of the audio tag information corresponding to the first reference audio.

[0078] In this way, the first fusion result and the first auxiliary coding result can be determined as the first input information, input into the audio branch, and the first output result can be obtained. The first loss can be determined based on the first output result, and then the parameters of the audio branch can be updated based on the first loss.

[0079] For example, the audio branch can generate a first predicted noise (first output result) based on the first input information. Accordingly, the first predicted noise can be compared with the first noise. Based on the comparison result, a first loss can be generated through a predetermined loss function. Then, the parameters of the audio branch can be updated based on the first loss.

[0080] In the first stage of training, the audio branch can learn to predict the first noise superimposed on the target audio under the "guidance" of the reference audio. Through multiple training sessions with a large amount of data, the audio branch learns to extract the speaker's acoustic features from a short reference audio and apply them to the generation of new content. In addition, by training the audio branch with long audio, the audio branch can be equipped with the ability to generate long audio, thus improving its performance.

[0081] Furthermore, by inputting the first auxiliary encoding result into the audio branch, the model can be provided with clear semantic and acoustic feature guidance. This enables the model to not only imitate the acoustic features of the reference audio when learning to generate audio, but also to make the generated content consistent with the text description, thereby effectively improving the accuracy and naturalness of the generated audio.

[0082] In response to the determination that the predetermined first-stage training termination condition has been met, the first-stage training can be terminated, and the second-stage training can be performed on the audio branch based on the short audio clip until the predetermined second-stage training termination condition is met.

[0083] In some embodiments of this disclosure, a short audio segment can be extracted from any target result and determined as the second target audio segment. Alternatively, an audio training data segment different from the second target audio segment can be extracted from the same target result and determined as the second reference audio segment. Then, the second input information can be determined based on the second target audio segment and the second reference audio segment. The second input information can be input into the audio branch to obtain the second output result. Furthermore, the second loss can be determined based on the second output result, and the parameters of the audio branch can be updated based on the second loss.

[0084] The second target audio and the second reference audio come from the same target result. The second target audio can be a short audio file that is randomly selected, and the second reference audio can be a long audio file or a short audio file, but it needs to be different from the second target audio file.

[0085] In some embodiments of this disclosure, the method for determining the second input information based on the second target audio and the second reference audio may include: encoding the second target audio to obtain a second target encoding result; superimposing second noise on the second target encoding result to obtain a second noisy encoding result; encoding the second reference audio to obtain a second reference encoding result; fusing the second noisy encoding result and the second reference encoding result to obtain a second fusion result; and determining the second input information based on the second fusion result. For example, the second noisy encoding result and the second reference encoding result may be fused by addition.

[0086] The second noise can be randomly generated. The process of generating the second fusion result is similar to... Figure 3 The process of generating the first fusion result is similar and will not be described in detail here.

[0087] In some embodiments of this disclosure, the method of determining the second input information based on the second fusion result may include: obtaining a second auxiliary encoding result, the second auxiliary encoding result including at least: the result obtained after encoding the audio tag information corresponding to the second target audio, and determining the second fusion result and the second auxiliary encoding result as the second input information.

[0088] For example, the audio tag information corresponding to the second target audio can be encoded using a text encoder to obtain a second auxiliary encoding result. In practical applications, if necessary, the second auxiliary encoding result can also include the encoded result of the audio tag information corresponding to the second reference audio.

[0089] In this way, the second fusion result and the second auxiliary coding result can be determined as the second input information and input to the audio branch. The audio branch can generate the second predicted noise (the second output result) based on the second input information. Furthermore, the second predicted noise can be compared with the second noise. Based on the comparison result, a second loss can be generated through a predetermined loss function. Then, the parameters of the audio branch can be updated based on the second loss.

[0090] In the scheme described in this disclosure, the duration of the target audio and video content generated by the trained audio and video joint generation model is the first duration. Correspondingly, the audio branch is trained in the second stage using short audio clips, which can improve the timing synchronization and overall consistency of audio and video in the subsequently generated target audio and video content.

[0091] After completing the initial training for the audio branch, the audio and video branches can be jointly trained based on the video training data and the audio training data, thus enabling audio-video joint training.

[0092] In some embodiments of this disclosure, training data pairs may be obtained, which may include: a target video extracted from video training data and a third target audio that matches the target video extracted from audio training data. Then, the audio branch and the video branch may be jointly trained based on the training data pairs.

[0093] The matching third target audio may include: audio training data that is included in the same original video as the target video and is time-aligned with the target video.

[0094] For example, a video training data can be extracted and identified as the target video, and the corresponding audio training data can be identified as the third target audio. The corresponding audio training data refers to the audio training data that is included in the same original video as the target video and is time-aligned. An audio training data different from the third target audio can be extracted from the target result where the third target audio is located and identified as the third reference audio. Then, the third input information can be determined based on the target video, the third target audio, and the third reference audio. The third input information can then be input into the audio-video joint generation model to obtain the third output result. The third loss can be determined based on the third output result, and the parameters of the audio branch and the video branch can be updated based on the third loss.

[0095] For example, a 10-second original video is split into two video training data sets, namely video training data a and video training data b, in chronological order. The original audio corresponding to the original video is also split into two audio training data sets, namely audio training data a and audio training data b, in chronological order. Assuming that video training data a is extracted as the target video, then audio training data a is the third target audio.

[0096] It can be seen that the third target audio must be a short audio. In addition, a third reference audio needs to be extracted from the target result where the third target audio is located. The third reference audio can be a long audio or a short audio, but it needs to be different from the third target audio.

[0097] By using naturally aligned, source-specific audio and video training data and introducing reference audio from the same speaker, the model can be effectively driven to simultaneously grasp the temporal correlation between audio and video and the generalization and cloning ability of acoustic features, thereby improving the generation quality and consistency of subsequent target audio and video content.

[0098] In some embodiments of this disclosure, the method for determining the third input information based on the target video, the third target audio, and the third reference audio may include: encoding the target video to obtain a video encoding result, and superimposing third noise on the video encoding result to obtain a video with added noise encoding result; encoding the third target audio to obtain a third target encoding result, and superimposing fourth noise on the third encoding result to obtain a third with added noise encoding result; encoding the third reference audio to obtain a third reference encoding result; fusing the third with added noise encoding result and the third reference encoding result to obtain a third fusion result; and determining the third input information based on the third fusion result and the video with added noise encoding result. For example, the third with added noise encoding result and the third reference encoding result may be fused by addition.

[0099] Specifically, in some embodiments of this disclosure, a third auxiliary coding result can be obtained, which includes at least: the result obtained after encoding the audio tag information corresponding to the third target audio, and the result obtained after encoding the video tag information corresponding to the target video. Then, the video noise-adding coding result, the third fusion result, and the third auxiliary coding result can be determined as the third input information.

[0100] Based on the above introduction, Figure 4 This is a flowchart illustrating an embodiment of the method for generating third input information as described in this disclosure. Figure 4 As shown, it includes:

[0101] In step 401, a video training data is extracted and identified as the target video, and the audio training data corresponding to the target video is identified as the third target audio.

[0102] In step 402, an audio training data that is different from the third target audio is extracted from the target result where the third target audio is located, and is determined as the third reference audio.

[0103] In step 403, the target video is encoded to obtain a video encoding result, and a third noise is superimposed on the video encoding result to obtain a video noise-enhanced encoding result.

[0104] For example, a video encoder can be used to encode the target video to obtain the video encoding result, and a third noise can be superimposed on the video encoding result to obtain a video with added noise encoding result. The third noise can be randomly generated.

[0105] In step 404, the third target audio is encoded to obtain the third target encoding result, and the fourth noise is superimposed on the third encoding result to obtain the third noise-added encoding result.

[0106] In step 405, the third reference audio is encoded to obtain the third reference encoding result.

[0107] In step 406, the third noise-adding coding result and the third reference coding result are fused by addition to obtain the third fusion result.

[0108] In step 407, a third auxiliary encoding result is obtained, which includes at least: the result obtained after encoding the audio tag information corresponding to the third target audio, and the result obtained after encoding the video tag information corresponding to the target video.

[0109] In step 408, the video noise-adding coding result, the third fusion result, and the third auxiliary coding result are determined as the third input information.

[0110] The third input information can be input into the audio and video joint generation model to obtain the third output result. The third loss can be determined based on the third output result, and the parameters of the audio branch and the video branch can be updated based on the third loss.

[0111] For example, the audio-video joint generation model can generate a third prediction noise and a fourth prediction noise based on the third input information. The third prediction noise corresponds to the video branch, and the fourth prediction noise corresponds to the audio branch. Accordingly, a third loss can be generated based on the third prediction noise and the third noise through a predetermined loss function, and a fourth loss can be generated based on the fourth prediction noise and the fourth noise through a predetermined loss function. Then, the third loss and the fourth loss can be combined to jointly update the parameters of the video branch and the audio branch.

[0112] The audio-video joint generation model can be implemented using a Diffusion Transformer (DiT) architecture. In the scheme described in this disclosure, due to the symmetrical network structure, a cross-attention mechanism can be easily designed at symmetrical positions to facilitate audio-video interaction. Simultaneously, video and audio tag information can also undergo cross-attention with the audio and video branches respectively in DiT. Therefore, audio and video can not only be fused together but also incorporate relevant text information, thereby improving the audio-video consistency of the generated content and ensuring that the generation follows text instructions.

[0113] Once the predetermined joint training termination conditions are met, the joint training process can be terminated, and the trained audio-video joint generation model can be used for practical inference applications.

[0114] Accordingly, Figure 5 This is a flowchart illustrating an embodiment of the audio / video joint generation method described in this disclosure. Figure 5 As shown, the specific implementation methods are as follows.

[0115] In step 501, the audio and video generation requirement description information and the fourth reference audio are obtained.

[0116] In step 502, the audio and video generation requirement description information and the fourth reference audio input audio and video are used to jointly generate a model to obtain the output target audio and video content. The audio in the target audio and video content conforms to the acoustic characteristics of the fourth reference audio.

[0117] Among them, the audio and video joint generation model can be adopted Figure 1 The method shown is used for training.

[0118] The audio and video generation requirement description information may include video description information, audio description information, and corresponding text content information. The fourth reference audio may be the audio of the target person, such as the audio of an entertainment star.

[0119] It should be noted that the various text, audio, and video information in the embodiments described in this disclosure are not targeted at any specific user and are not intended to reflect the personal information of any specific user. Furthermore, the entity implementing the solution described in this disclosure can obtain the text, audio, and video information through various public, legal, and compliant means, such as obtaining the audio of an entertainment star after authorization. In summary, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution of this disclosure all comply with relevant laws and regulations and do not violate public order and good morals.

[0120] The audio / video generation requirement description information and the fourth reference audio / video input input can be jointly generated into a model to obtain the output target audio / video content. As the name suggests, the target audio / video content includes both audio and video, and the audio conforms to the acoustic characteristics of the fourth reference audio, such as its timbre. Accordingly, Figure 6 This is a schematic diagram of the input and output of the audio and video joint generation model 600 described in this disclosure.

[0121] use Figure 5 The scheme described in the illustrated embodiment can leverage a pre-trained audio-video joint generation model to jointly generate target audio and video content that includes both audio and video from end to end. This improves the generation efficiency of the target audio and video content and minimizes the problem of audio-video timing misalignment, thereby enhancing the generation quality of the target audio and video content.

[0122] Moreover, in traditional methods, achieving timbre control inevitably requires calling additional models, making the implementation complex. However, using... Figure 5The method described in the embodiment shown allows users to quickly and easily generate target audio and video content that matches the timbre and other acoustic characteristics of the input reference audio simply by providing a reference audio clip, thereby meeting the diverse needs of users.

[0123] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this disclosure. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this disclosure. Furthermore, for parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0124] The above is an introduction to the method embodiments. The following describes the solution described in this disclosure further through device embodiments.

[0125] Figure 7 This is a schematic diagram of the structural composition of an embodiment 700 of the audio-video joint generation model training device described in this disclosure. Figure 7 As shown, it includes: a data acquisition module 701, a first training module 702, and a second training module 703.

[0126] The data acquisition module 701 is used to acquire video training data and audio training data.

[0127] The first training module 702 is used to perform preliminary training on the audio branch based on the audio training data. The audio and video joint generation model includes an audio branch and a video branch.

[0128] The second training module 703 is used to perform joint training on the audio branch and the video branch based on the video training data and the audio training data in response to determining that the preliminary training is complete.

[0129] In some embodiments of this disclosure, the data acquisition module 701 can acquire the original video and the first original audio, and can preprocess the original video and the first original audio. Then, it can extract the second original audio from the preprocessed original video, and then generate video training data based on the preprocessed original video, and generate audio training data based on the preprocessed first original audio and the second original audio.

[0130] The preprocessing may include: discarding original video with a duration shorter than a predetermined first duration, discarding original audio with a duration shorter than the first duration, etc.

[0131] In some embodiments of this disclosure, for any original video, the data acquisition module 701, in response to determining that the duration of the original video is equal to a first duration, can determine the original video as video training data; in response to determining that the duration of the original video is greater than the first duration, can segment the original video in a manner that segments once every first duration, and can determine the segmented video segments of the first duration as video training data.

[0132] In some embodiments of this disclosure, the data acquisition module 701 can determine a first original audio and a second original audio with a duration equal to a first duration as audio training data. In addition, for any second original audio with a duration greater than the first duration, the second original audio can be segmented according to the method of segmenting once every first duration, and the segmented audio segments of the first duration can be determined as audio training data. Furthermore, for any first original audio with a duration greater than the first duration, the first original audio can be directly determined as audio training data, or, in response to determining that the first original audio needs to be segmented, the first original audio can be segmented according to the method of segmenting once every first duration, and the segmented audio segments of the first duration can be determined as audio training data.

[0133] In some embodiments of this disclosure, the data acquisition module 701 can also acquire video tag information corresponding to the video training data. The video tag information may include: video description information of the corresponding video training data, audio description information of the audio in the corresponding video training data, and text content information after the audio is converted into text. It can also acquire audio tag information corresponding to the audio training data. The audio tag information may include: audio description information of the corresponding audio training data and text content information after the audio is converted into text.

[0134] Once video and audio training data are generated, the audio-video joint generation model can be trained using these data.

[0135] First, the first training module 702 can perform preliminary training on the audio branches based on the audio training data. In some embodiments of this disclosure, each piece of audio training data can correspond to a speaker. Accordingly, the first training module 702 can cluster the audio training data. Audio training data in the same clustering result correspond to the same speaker. Then, the clustering result that meets the following requirements can be determined as the target result: the clustering result includes at least two pieces of audio training data, and the clustering result includes both long audio and short audio. The long audio is audio training data with a duration greater than a first duration, and the short audio is audio training data with a duration equal to the first duration. Then, the audio branches can be preliminarily trained based on the long audio and short audio.

[0136] In some embodiments of this disclosure, the initial training of the audio branch by the first training module 702 may include: a first-stage training and a second-stage training, wherein the first-stage training of the audio branch may be performed based on long audio and the second-stage training of the audio branch may be performed based on short audio.

[0137] In some embodiments of this disclosure, the first training module 702 may perform a first-stage training of the audio branch based on a long audio segment by: extracting a long audio segment from any target result and determining it as a first target audio segment; extracting a different audio training data segment from the same target result and determining it as a first reference audio segment; determining first input information based on the first target audio segment and the first reference audio segment; inputting the first input information into the audio branch to obtain a first output result; determining a first loss based on the first output result; and updating the parameters of the audio branch based on the first loss. The second-stage training of the audio branch based on a short audio segment may include: extracting a short audio segment from any target result and determining it as a second target audio segment; extracting a different audio training data segment from the same target result and determining it as a second reference audio segment; determining second input information based on the second target audio segment and the second reference audio segment; inputting the second input information into the audio branch to obtain a second output result; determining a second loss based on the second output result; and updating the parameters of the audio branch based on the second loss.

[0138] In some embodiments of this disclosure, the method by which the first training module 702 determines the first input information based on the first target audio and the first reference audio may include: encoding the first target audio to obtain a first target encoding result; superimposing first noise on the first target encoding result to obtain a first noisy encoding result; encoding the first reference audio to obtain a first reference encoding result; fusing the first noisy encoding result and the first reference encoding result by addition to obtain a first fusion result; and determining the first input information based on the first fusion result. Determining the second input information based on the second target audio and the second reference audio includes: encoding the second target audio to obtain a second target encoding result; superimposing second noise on the second target encoding result to obtain a second noisy encoding result; encoding the second reference audio to obtain a second reference encoding result; fusing the second noisy encoding result and the second reference encoding result by addition to obtain a second fusion result; and determining the second input information based on the second fusion result.

[0139] In some embodiments of this disclosure, the first training module 702 determines the first input information based on the first fusion result by: obtaining a first auxiliary encoding result, the first auxiliary encoding result including at least: the result obtained after encoding the audio tag information corresponding to the first target audio, and determining the first fusion result and the first auxiliary encoding result as the first input information; the method for determining the second input information based on the second fusion result may include: obtaining a second auxiliary encoding result, the second auxiliary encoding result including at least: the result obtained after encoding the audio tag information corresponding to the second target audio, and determining the second fusion result and the second auxiliary encoding result as the second input information.

[0140] In some embodiments of this disclosure, the second training module 703 may perform joint training of the audio branch and the video branch based on video training data and audio training data in the following manner: acquiring a training data pair, wherein the training data pair includes: a target video extracted from the video training data and a third target audio that matches the target video extracted from the audio training data; and performing joint training of the audio branch and the video branch based on the training data pair.

[0141] The matching third target audio may include: audio training data that is included in the same original video as the target video and is time-aligned with the target video.

[0142] For example, a video training data can be extracted and identified as the target video, and a third target audio can be identified. Then, an audio training data different from the third target audio can be extracted from the target results where the third target audio is located and identified as the third reference audio. The third input information is determined based on the target video, the third target audio, and the third reference audio. The third input information is input into the audio-video joint generation model to obtain the third output result. The third loss is determined based on the third output result, and the parameters of the audio branch and the video branch are updated based on the third loss.

[0143] In some embodiments of this disclosure, the second training module 703 determines the third input information based on the target video, the third target audio, and the third reference audio in the following ways: encoding the target video to obtain a video encoding result, and superimposing third noise on the video encoding result to obtain a video noise-added encoding result; encoding the third target audio to obtain a third target encoding result, and superimposing fourth noise on the third encoding result to obtain a third noise-added encoding result; encoding the third reference audio to obtain a third reference encoding result; fusing the third noise-added encoding result and the third reference encoding result by adding them together to obtain a third fusion result; and determining the third input information based on the third fusion result and the video noise-added encoding result.

[0144] In some embodiments of this disclosure, the second training module 703 determines the third input information based on the third fusion result and the video noise-adding coding result in the following ways: obtaining a third auxiliary coding result, which includes at least: the result obtained after encoding the audio tag information corresponding to the third target audio, and the result obtained after encoding the video tag information corresponding to the target video; and determining the video noise-adding coding result, the third fusion result, and the third auxiliary coding result as the third input information.

[0145] Figure 8 This is a schematic diagram of the structural composition of embodiment 800 of the audio and video joint generation device described in this disclosure. Figure 8 As shown, it includes: an information acquisition module 801 and a content generation module 802.

[0146] The information acquisition module 801 is used to acquire the audio and video generation requirement description information and the fourth reference audio.

[0147] The content generation module 802 is used to combine the audio and video generation requirement description information and the fourth reference audio input audio and video to jointly generate a model and obtain the output target audio and video content. The audio in the target audio and video content conforms to the acoustic characteristics of the fourth reference audio.

[0148] Among them, the audio and video joint generation model can be adopted Figure 1 The method shown is used for training.

[0149] The present disclosure also discloses an intelligent agent. Figure 9 This is a schematic diagram of the composition structure of the intelligent agent embodiment 900 described in this disclosure. Figure 9 As shown, it includes: an input module 901, a processing module 902, and an output module 903.

[0150] Input module 901 is used to receive audio and video generation requirement description information and fourth reference audio.

[0151] The processing module 902 is used to combine the audio and video generation requirement description information and the fourth reference audio input audio and video to jointly generate a model, thereby obtaining the generated target audio and video content. The audio in the target audio and video content conforms to the acoustic characteristics of the fourth reference audio.

[0152] Output module 903 is used to output the target audio and video content.

[0153] Among them, the audio and video joint generation model can be adopted Figure 1 The method shown is used for training.

[0154] The specific workflow of each of the above device embodiments can be found in the relevant descriptions in the foregoing method embodiments, and will not be repeated here.

[0155] The solutions described in this disclosure can be applied to the field of artificial intelligence, particularly computer vision, deep learning, large-scale models, and natural language processing. Artificial intelligence is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It involves both hardware and software technologies. Artificial intelligence hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. Artificial intelligence software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0156] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0157] Figure 10 A schematic block diagram of an electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0158] like Figure 10 As shown, the electronic device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. The RAM 1003 may also store various programs and data required for the operation of the electronic device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0159] Multiple components in electronic device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of displays, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows electronic device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0160] The computing unit 1001 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as those described in this disclosure. For example, in some embodiments, the methods described in this disclosure can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the methods described in this disclosure can be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to perform the methods described herein by any other suitable means (e.g., by means of firmware).

[0161] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0162] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0163] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM), flash memory, optical fiber, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0164] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0165] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0166] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0167] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0168] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for training an audio-video joint generation model, comprising: Acquire video and audio training data; The audio branch is initially trained based on the audio training data, and the audio-video joint generation model includes the audio branch and the video branch. In response to determining that the initial training is complete, the audio branch and the video branch are jointly trained based on the video training data and the audio training data.

2. The method according to claim 1, wherein, The acquisition of video training data and audio training data includes: Acquire the original video and the first original audio; The original video and the first original audio are preprocessed; Extract the second original audio from the preprocessed original video; The video training data is generated based on the preprocessed original video, and the audio training data is generated based on the preprocessed first original audio and the second original audio.

3. The method according to claim 2, wherein, The preprocessing includes: discarding the original video whose duration is less than a predetermined first duration; The generation of the video training data includes: For any original video, in response to determining that the duration of the original video is equal to the first duration, the original video is determined as the video training data; In response to determining that the duration of the original video is greater than the first duration, the original video is segmented in a manner that segments once every first duration, and the video segments of the first duration obtained from the segmentation are determined as the video training data.

4. The method according to claim 3, wherein, The preprocessing further includes: discarding the first original audio with a duration shorter than the first duration; The generation of the audio training data includes: The first original audio and the second original audio, both with a duration equal to the first duration, are determined as the audio training data; For any second original audio segment whose duration is longer than the first duration, the second original audio segment is segmented according to the method of segmenting once every first duration, and the segmented audio segments of the first duration are determined as the audio training data; For any first original audio segment whose duration is longer than the first duration, the first original audio segment is directly determined as the audio training data. Alternatively, in response to determining that the first original audio segment needs to be segmented, the first original audio segment is segmented once every first duration, and the segmented audio segments of the first duration are determined as the audio training data.

5. The method according to claim 4, further comprising: Obtain the video tag information corresponding to the video training data. The video tag information includes: video description information of the corresponding video training data, audio description information of the audio in the corresponding video training data, and text content information after the audio is converted into text. Obtain the audio tag information corresponding to the audio training data. The audio tag information includes: the audio description information and the text content information of the corresponding audio training data.

6. The method according to claim 5, wherein, Each audio training data point corresponds to a speaker; The preliminary training of the audio branch based on the audio training data includes: The audio training data is clustered, and audio training data in the same clustering result correspond to the same speaker; Clustering results that meet the following requirements are identified as target results: the clustering results include at least two audio training data, and the clustering results include both long audio and short audio, wherein the long audio is audio training data with a duration greater than the first duration, and the short audio is audio training data with a duration equal to the first duration; The audio branch is initially trained based on the long audio and the short audio.

7. The method according to claim 6, wherein, The preliminary training includes: a first-stage training and a second-stage training; The initial training of the audio branch based on the long audio and the short audio includes: The first stage of training is performed on the audio branch based on the long audio; The second phase of training is performed on the audio branch based on the short audio clip.

8. The method according to claim 7, wherein, The first-stage training of the audio branch based on the long audio includes: Extract a long audio segment from any target result and determine it as the first target audio segment. Extract an audio training data segment that is different from the first target audio segment from the same target result and determine it as the first reference audio segment. Determine the first input information based on the first target audio segment and the first reference audio segment. Input the first input information into the audio branch to obtain the first output result. Determine the first loss based on the first output result. Update the parameters of the audio branch based on the first loss. The second-stage training of the audio branch based on the short audio clip includes: Extract a short audio segment from any target result and designate it as the second target audio segment. Extract a different audio training data segment from the same target result and designate it as the second reference audio segment. Determine the second input information based on the second target audio segment and the second reference audio segment. Input the second input information into the audio branch to obtain the second output result. Determine the second loss based on the second output result. Update the parameters of the audio branch based on the second loss.

9. The method according to claim 8, wherein, The step of determining the first input information based on the first target audio and the first reference audio includes: The first target audio is encoded to obtain a first target encoding result. First noise is added to the first target encoding result to obtain a first noisy encoding result. The first reference audio is encoded to obtain a first reference encoding result. The first noisy encoding result and the first reference encoding result are fused to obtain a first fusion result. The first input information is determined based on the first fusion result. The step of determining the second input information based on the second target audio and the second reference audio includes: The second target audio is encoded to obtain a second target encoding result. The second target encoding result is superimposed with second noise to obtain a second noisy encoding result. The second reference audio is encoded to obtain a second reference encoding result. The second noisy encoding result and the second reference encoding result are fused to obtain a second fusion result. The second input information is determined based on the second fusion result.

10. The method according to claim 9, wherein, The step of determining the first input information based on the first fusion result includes: Obtain a first auxiliary encoding result, which includes at least: the result obtained after encoding the audio tag information corresponding to the first target audio, and determine the first fusion result and the first auxiliary encoding result as the first input information; The step of determining the second input information based on the second fusion result includes: Obtain a second auxiliary encoding result, which includes at least the result obtained by encoding the audio tag information corresponding to the second target audio. Determine the second fusion result and the second auxiliary encoding result as the second input information.

11. The method according to claim 5, wherein, The step of jointly training the audio branch and the video branch based on the video training data and the audio training data includes: Acquire training data pairs, wherein the training data pairs include: a target video extracted from the video training data, and a third target audio that matches the target video extracted from the audio training data; Based on the training data pair, the audio branch and the video branch are jointly trained.

12. The method according to claim 11, wherein, The matching third target audio includes: audio training data that is included in the same original video as the target video and is time-aligned with the target video; The joint training of the audio branch and the video branch includes: Extract an audio training data point that is different from the third target audio from the target results where the third target audio is located, and determine it as the third reference audio; A third input information is determined based on the target video, the third target audio, and the third reference audio. The third input information is then input into the audio-video joint generation model to obtain a third output result. A third loss is determined based on the third output result, and the parameters of the audio branch and the video branch are updated based on the third loss.

13. The method according to claim 12, wherein, The step of determining the third input information based on the target video, the third target audio, and the third reference audio includes: The target video is encoded to obtain a video encoding result, and a third noise is superimposed on the video encoding result to obtain a video noisy encoding result; The third target audio is encoded to obtain a third target encoding result, and a fourth noise is superimposed on the third encoding result to obtain a third noisy encoding result; The third reference audio is encoded to obtain the third reference encoding result; The third noisy coding result and the third reference coding result are fused to obtain the third fusion result; The third input information is determined based on the third fusion result and the video noise-enhancing coding result.

14. The method according to claim 13, wherein, The step of determining the third input information based on the third fusion result and the video noise-enhancing coding result includes: Obtain a third auxiliary encoding result, which includes at least: the result obtained by encoding the audio tag information corresponding to the third target audio, and the result obtained by encoding the video tag information corresponding to the target video; The video noise-adding encoding result, the third fusion result, and the third auxiliary encoding result are determined as the third input information.

15. A method for jointly generating audio and video, comprising: Obtain the audio / video generation requirement description information and the fourth reference audio; The audio and video generation requirement description information and the fourth reference audio input audio and video are combined to generate a joint generation model to obtain the output target audio and video content. The audio in the target audio and video content conforms to the acoustic characteristics of the fourth reference audio. The joint audio and video generation model is trained using the method described in any one of claims 1 to 14.

16. A training device for a joint audio-video generation model, comprising: Data acquisition module, first training module, and second training module; The data acquisition module is used to acquire video training data and audio training data; The first training module is used to perform preliminary training on the audio branch based on the audio training data, and the audio-video joint generation model includes the audio branch and the video branch; The second training module is configured to, in response to determining that the initial training is complete, jointly train the audio branch and the video branch based on the video training data and the audio training data.

17. An audio-visual co-generation apparatus, comprising: Information acquisition module and content generation module; The information acquisition module is used to acquire audio and video generation requirement description information and fourth reference audio; The content generation module is used to combine the audio and video generation requirement description information and the fourth reference audio input audio and video to form a joint generation model to obtain the output target audio and video content. The audio in the target audio and video content conforms to the acoustic characteristics of the fourth reference audio. The joint audio and video generation model is trained using the method described in any one of claims 1 to 14.

18. An intelligent agent, comprising: Input module, processing module, and output module; The input module is used to receive audio and video generation requirement description information and fourth reference audio; The processing module is used to combine the audio and video generation requirement description information and the fourth reference audio input audio and video joint generation model to obtain the generated target audio and video content. The audio in the target audio and video content conforms to the acoustic characteristics of the fourth reference audio. The audio and video joint generation model is trained using the method described in any one of claims 1 to 14. The output module is used to output the target audio and video content.

19. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-15.

20. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-15.

21. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the method of any one of claims 1-15.

Citation Information

Patent Citations

  • Audio and video synthesis method

    CN110728971A

  • Image display method

    JP2004056286A

  • Text to synchronized joint video and audio generation

    WO2025178789A1