Audio and video generation method and device, electronic device, readable storage medium

By using a multi-stage hybrid generation framework, audio and video are generated using the first frame of the video and dynamic descriptions or editing instructions. This solves the problem of the single generation method in existing audio and video generation methods, realizes diversified audio and video generation, and supports audio and video generation of multiple themes and diverse scenes in the real world.

CN122120572APending Publication Date: 2026-05-29BEIJING UNIV OF POSTS & TELECOMM

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING UNIV OF POSTS & TELECOMM
Filing Date
2026-02-27
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing audio and video generation methods are too simplistic and cannot cover the diverse scenarios of the real world. The lack of comprehensive generation methods results in the generated audio and video failing to reflect cross-modal relationships and diverse spatiotemporal patterns in complex scenarios.

Method used

It adopts a multi-stage hybrid generation framework, which uses a proprietary model to plan generation tasks and an expert model to execute precise generation. It combines the first frame of the video with dynamic descriptions or editing instructions to generate synthetic audio and video and edit audio and video, covering multiple themes and diverse scenes in the real world.

Benefits of technology

It generates diverse audio and video content covering multiple themes in the real world, supports diverse generation types, makes up for the shortcomings of existing technologies in scene diversity and generation capabilities, and provides a systematic and scalable technical solution for building audio and video generation data for real-world scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122120572A_ABST
    Figure CN122120572A_ABST
Patent Text Reader

Abstract

The application provides an audio and video generation method and device, electronic equipment and readable storage medium, belonging to the field of audio and video processing. The method comprises the following steps: obtaining an audio and video generation task, the generation task comprising a synthesis task or an editing task; the synthesis task comprising a first frame of video and corresponding dynamic description; the editing task comprising editing instructions; generating a fake audio and video based on the generation task, comprising: generating a synthesized video based on the first frame of video and the corresponding dynamic description, and generating a synthesized audio based on the synthesized video to obtain a synthesized audio and video; or, performing an editing operation on selected frames of a real video based on the editing instructions, generating an edited video based on the edited selected frames, and generating an edited audio based on the edited video to obtain an edited audio and video. The application makes up for the deficiencies of existing audio and video generation methods in terms of scene diversity and generation capability, and provides a systematic and expandable technical solution for constructing audio and video generation data for real-world scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of audio and video processing technology, and more specifically, relates to an audio and video generation method and apparatus, electronic device, and readable storage medium. Background Technology

[0002] In recent years, audio and video generation technology has developed rapidly, especially cross-modal generation systems represented by models such as Sora and KLING, which have achieved significant breakthroughs in dynamic scene construction, semantically consistent generation, and high-fidelity content generation. However, despite the continuous improvement of generation model capabilities, the current field of audio and video generation still suffers from at least the following core shortcomings: First, there is a lack of audio and video generation datasets that cover multiple themes and contexts in the real world. Existing publicly available datasets are mostly focused on specific human subjects, and the audio and video content is mainly facial feature videos and speaker audio, which is difficult to cover the diverse themes that are widely present in the real world, such as natural landscapes, animal activities, and traffic environments. As a result, existing audio and video generation data cannot reflect cross-modal correlations and diverse spatiotemporal patterns in complex scenes.

[0003] Secondly, the generation method is singular and lacks a comprehensive generation approach.

[0004] Therefore, the generative models trained on the current dataset have a limited range of generation methods, and the generated audio and video cannot cover the rich scenes of the real world. Summary of the Invention

[0005] The purpose of this application is to provide an audio and video generation method, device, electronic device, and readable storage medium, which solves the problem that existing audio and video generation methods have a single generation mode and the generated audio and video cannot cover multiple scenarios.

[0006] A first aspect of this application provides an audio / video generation method, including: Obtain audio and video generation tasks, which include either compositing or editing tasks. Compositing tasks include the first frame of the video and its corresponding dynamic description; editing tasks include editing instructions. The generation task generates fake audio and video, which includes synthesized audio and video and edited audio and video. The generation task generates fake audio and video, which includes: generating a synthesized video based on the first frame of the video and the corresponding dynamic description, and generating synthesized audio based on the synthesized video to obtain synthesized audio and video; or, performing editing operations on selected frames of the real video based on editing instructions, generating an edited video based on the edited selected frames, and generating edited audio based on the edited video to obtain edited audio and video.

[0007] In one embodiment, when the generation task is a synthesis task, an audio / video generation task is obtained, including: Generate dynamic descriptions based on the first frame of real video; The task involves combining the first frame of a real video with its dynamic description.

[0008] In one embodiment, when the generation task is a synthesis task, an audio / video generation task is obtained, including: Generate static images based on preset real-world scenes; Use a static image as the first frame of the video to generate a dynamic description based on the first frame of the video; The first frame of the video and the dynamic description are combined into a synthesis task.

[0009] In one embodiment, when the generation task is an editing task, obtaining the audio / video generation task includes: A predetermined number of sampled frames are obtained by uniformly sampling real video frames. Based on sampling frame generation tampering scheme; Structured editing instructions are generated based on the tampering scheme, and these instructions are used as editing tasks.

[0010] In one embodiment, the editing instructions include: a target region, a time window, an editing operation, and a target object. The editing operation is performed on selected frames of a real video based on the editing instructions, including: Selected frames are obtained from real video based on a time window; Editing operations are performed on target objects in the selected frame based on the target region and the editing operation.

[0011] In one embodiment, the editing instructions include: a target region, a time window, an editing operation, and a target object. The editing operation is performed on selected frames of a real video based on the editing instructions, including: Selected frames are obtained from real video based on a time window; The target object in the selected frame is segmented to obtain the masked video; Editing operations are performed on the target object in the selected frame based on the masked video.

[0012] In one embodiment, The method also includes: arranging and combining fake audio and video with real audio and video to obtain target audio and video, including: real audio and real video, real audio and edited video, real audio and synthesized video, edited audio and real video, edited audio and edited video, synthesized audio and real video, and synthesized audio and synthesized video.

[0013] A second aspect of this application provides an audio / video generation apparatus, comprising: The task acquisition module is used to obtain audio and video generation tasks, which include: compositing tasks or editing tasks; compositing tasks include: the first frame of the video and its corresponding dynamic description; editing tasks include: editing instructions. The fake audio and video generation module is used to generate fake audio and video based on the generation task. The fake audio and video includes synthesized audio and video and edited audio and video. It includes: generating synthesized video based on the first frame of the video and the corresponding dynamic description, and generating synthesized audio based on the synthesized video to obtain synthesized audio and video; or, performing editing operations on selected frames of real video based on editing instructions, generating edited video based on the edited selected frames, and generating edited audio based on the edited video to obtain edited audio and video.

[0014] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.

[0015] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.

[0016] A fifth aspect of this application provides a computer program product, including a computer program or computer executable instructions, wherein when the computer program or computer executable instructions are executed by a processor, the steps of the above-described method are implemented.

[0017] The beneficial effects of the audio and video generation method and apparatus, electronic device, and readable storage medium provided in this application are as follows: This application first obtains an audio / video generation task, which includes either a compositing task or an editing task. The compositing task includes the first frame of a video and its corresponding dynamic description. The editing task includes editing instructions. Then, based on the generation task, a forged audio / video is generated, which includes compositing audio / video and edited audio / video. Generating forged audio / video based on the generation task includes generating a compositing video based on the first frame of the video and its corresponding dynamic description, and generating compositing audio based on the compositing video. Alternatively, editing operations are performed on selected frames of a real video based on editing instructions, and an edited video is generated based on the edited selected frames, and edited audio is generated based on the edited video.

[0018] This application addresses the compositing task by obtaining a dynamic description corresponding to the first frame of a video. Similarly, it generates editing instructions for editing tasks. Since the first frame can be of any theme, and the editing instructions can perform arbitrary editing operations on the original video, this application achieves rich generative semantics, thereby generating diverse audio and video content covering multiple real-world themes and supporting diverse generation types. This application effectively overcomes the shortcomings of existing audio and video generation methods in terms of scene diversity and generation capabilities, providing a systematic and scalable technical solution for constructing audio and video generation data for real-world scenes. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A flowchart illustrating an audio / video generation method provided in an embodiment of this application; Figure 2 A flowchart of a synthesis task provided in an embodiment of this application; Figure 3 An editing task flowchart provided for one embodiment of this application; Figure 4 This is a schematic diagram of a single generation type in existing technologies; Figure 5 A schematic diagram illustrating the diverse generation types provided in one embodiment of this application; Figure 6 This is a structural block diagram of an audio / video generation apparatus provided in an embodiment of this application; Figure 7 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0021] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0022] To make the objectives, technical solutions, and advantages of this application clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.

[0023] To address the limitation of existing audio and video generation methods in supporting large-scale, batch generation of audio and video content across multiple real-world scenarios, this application provides a multi-stage hybrid generation framework. This framework enables different types of generation models (e.g., video compositing models, image-driven video editing models, audio generation models, etc.) to work collaboratively within the same framework, thereby supporting large-scale, batch generation of multi-themed, multi-type audio and video content. This application divides the audio and video generation process into three steps: a proprietary model planning the generation task, an expert model executing precise generation, and further includes cross-modal combination. The details are as follows: Please refer to Figure 1 , Figure 1 This is a flowchart illustrating an audio / video generation method provided in an embodiment of this application. The method can be executed by an electronic device and may include: S11: Obtain audio and video generation tasks. Generation tasks include: compositing tasks or editing tasks; compositing tasks include: the first frame of the video and its corresponding dynamic description; editing tasks include: editing instructions.

[0024] In this embodiment, the synthesis task aims to generate a semantically coherent, dynamic, and audio-visual synchronized real-world scene; the editing task is to perform localized spatiotemporal operations while maintaining global visual and auditory realism.

[0025] This step is the proprietary model planning and generation task, and this application covers two generation methods: synthesis and editing.

[0026] The synthesis task includes the first frame of the video and its corresponding dynamic description, which serves as textual guidance consistent with the motion during video synthesis.

[0027] In one embodiment, such as Figure 2 As shown, when the generation task is a compositing task, the resulting audio / video generation task includes: Generating dynamic descriptions from the first frame of real video footage involves compositing the first frame of the real video footage with the dynamic descriptions. For example... Figure 2 The first frame of the real video is an image containing a train.

[0028] In one embodiment, an LMM model can be used to predict the reasonable dynamic evolution of a real video based on the first frame, thereby generating a dynamic description. That is, this embodiment uses frame-driven planning for the synthesis task. Figure 2 As shown, the dynamic description generated for the first frame of the video is "In the distance, a train passes through a small mountain village...".

[0029] In another embodiment, such as Figure 2 As shown, when the generation task is a compositing task, the resulting audio / video generation task includes: S110. Generate static images based on preset real-world scenes.

[0030] This application divides real-world scenarios into human-themed scenarios and general-themed scenarios. Human-themed scenarios correspond to video generation and speaker audio generation based on facial features in existing benchmarks. General-themed scenarios cover 10 common scenarios in the real world, including natural landscapes, social activities, animals, music, transportation, daily life, sports, industry, alarm signals, and science.

[0031] This step can generate a still image based on one of the themes mentioned above. For example, a still image can be generated based on a static description of any real-world scene using a T21 model (such as Midjourney). Figure 2 The static image generated using Midjourney based on a traffic scene is a picture containing a train.

[0032] S111. Use the static image as the first frame of the video to generate a dynamic description based on the first frame of the video, and combine the first frame of the video and the dynamic description into a synthesis task.

[0033] Similar to the previous embodiment, an LMM model can be used to predict the reasonable dynamic evolution of the video based on the first frame, thereby generating a dynamic description. Unlike the previous embodiment, this embodiment uses scene-driven planning for the synthesis task.

[0034] Editing tasks include editing instructions. In one embodiment, when the generation task is an editing task, an audio / video generation task is obtained, including: S110. Perform uniform sampling based on real video frames to obtain a predetermined number of sampled frames; In one embodiment, such as Figure 3 As shown, a predetermined number of sampled frames can be obtained from a real video through uniform sampling. Uniform sampling means sampling one sampled frame at set intervals, for example, sampling one sampled frame at intervals of 3 frames. This predetermined number of sampled frames can be, for example, 8 frames, or other numbers set according to user needs.

[0035] S111, A tampering scheme based on sampled frames; In one embodiment, an existing proprietary model can be used to generate a tampering scheme based on the sampled frame, for example, for... Figure 3 The sampled frames generated a tampering scheme (editing scheme) that "removes the wooden stake in the center of the video within a time period of 3-5 seconds".

[0036] S112. Generate structured editing instructions based on the tampering scheme, and treat the editing instructions as editing tasks.

[0037] In one embodiment, an LMM model can be used to generate structured editing instructions from the tampering scheme. These structured editing instructions include: target region, time window, editing operation, and target object. These structured editing instructions provide clear editing guidance for subsequent editing.

[0038] In one embodiment, for the tampering scheme "delete the wooden stake in the center of the video within 3-5 seconds", the generated editing instruction has the following corresponding time window: 3-5 seconds; target area: center of the video; target object: wooden stake; and editing operation: delete.

[0039] S12: Generate fake audio and video based on the generation task. The fake audio and video includes: synthesized audio and video and edited audio and video; including: generating synthesized video based on the first frame of the video and the corresponding dynamic description and generating synthesized audio based on the synthesized video to obtain synthesized audio and video; or, performing editing operations on selected frames of the real video based on editing instructions, generating edited video based on the edited selected frames, and generating edited audio based on the edited video to obtain edited audio and video.

[0040] This step involves the expert model performing precise generation, which generates fake audio and video, including synthesized audio and video and edited audio and video.

[0041] In one embodiment, for a synthesis task, when generating fake audio and video based on a generation task to generate synthesized audio and video, such as... Figure 2 As shown, it specifically includes: Video synthesis models (e.g., KLING or QingYing) can be used to synthesize videos based on the first frame and dynamic descriptions. Then, an audio processing model is used to generate synthesized audio based on the synthesized video, resulting in a synthesized audio-visual product. This image-to-video model uses the first frame of the video as a visual anchor and the dynamic description as a motion cue to generate a synthesized video. Furthermore, an audio generation model is used to generate synthesized audio that is temporally and semantically consistent with the video. This two-step synthesis method ensures consistency between the visuals, actions, and ambient sound effects.

[0042] In one embodiment, for an editing task, a forged audio / video file is generated based on a generation task to generate edited audio / video, such as... Figure 3 As shown, the process includes: performing editing operations on selected frames of a real video based on editing instructions, generating an edited video based on the edited selected frames, and generating edited audio based on the edited video, resulting in edited audio and video.

[0043] In one embodiment, performing an editing operation on selected frames of a real video based on editing instructions includes: Selected frames are obtained from real video based on a time window; this real video is uniformly sampled when the editing task is generated. The time window is the time window included in the editing instruction. Continuing the previous example, the time window in the editing instruction is 3-5 seconds; the target region is the center of the video; the target object is a wooden stake; and the editing operation is deletion. Therefore, obtaining selected frames from real video based on the time window means obtaining 3-5 seconds of video frames from the real video as the selected frames. In one embodiment, the selected frames and editing instructions can be input into an image-driven video editing model (e.g., KLING), which performs editing operations on the selected frames based on the target region (which can be constrained by the user) to obtain the edited video.

[0044] Editing operations are performed on target objects within selected frames based on the target region and the editing operation. Continuing from the previous example, the wooden stake in the center of the video in selected frames 3-5 is deleted.

[0045] In another embodiment, performing editing operations on selected frames of a real video based on editing instructions includes: Selected frames are obtained from real video based on time windows; as described in the previous embodiments, they will not be repeated here.

[0046] The target object in the selected frame is segmented to obtain a masked video; one implementation of this masked video is as follows: Figure 3 As shown in the diagram. In one embodiment, SAM2 can be used to segment the target object.

[0047] Editing operations are performed on target objects in selected frames based on masked video. Since the white areas in the masked video represent the regions where the target objects are located, editing operations can be performed on the target objects in the selected frames based on the white areas in the masked video. In one embodiment, the masked video and the selected frames can be input together into an image-driven video editing model (e.g., KLING), which performs editing operations on the selected frames based on the masked video to obtain the edited video.

[0048] In one embodiment, to ensure cross-modal consistency, the edited video is further processed by an audio generation model to generate edited audio with temporal consistency, ultimately resulting in edited audio-video. Furthermore, the edited audio-video is inserted back into the corresponding time segment of the original real video.

[0049] In one embodiment, such as Figure 2 and Figure 3As shown, the method also includes: arranging and combining forged audio and video with real audio and video to obtain target audio and video, which also includes a forged combination step. The target audio and video types obtained after combination include the following seven types: real audio and real video, real audio and edited video, real audio and synthesized video, edited audio and real video, edited audio and edited video, synthesized audio and real video, and synthesized audio and synthesized video.

[0050] like Figure 4 and Figure 5 The diagrams shown illustrate a single generation type in existing technologies and a diversified generation type in this application, respectively. A comparison reveals that the audio / video generation method of this application achieves richer generation semantics compared to existing technologies, thereby generating diverse audio / video content covering multiple real-world themes and supporting diverse generation types. This application effectively overcomes the shortcomings of existing audio / video generation methods in terms of scene diversity and generation capabilities, providing a systematic and scalable technical solution for constructing audio / video generation data for real-world scenes.

[0051] To verify the authenticity of the audio and video generated using the method proposed in this application, two existing audio and video generation data detection methods and 11 AV-LMMs were selected. Four detection tasks were set from easy to difficult: binary judgment, generation type classification, generation detail selection, and interpretability reasoning. Binary judgment aims to determine whether the overall audio and video data is generated or real data; generation type classification aims to explore whether the model can detect fine-grained cross-modal generation combinations; generation detail selection aims to explore whether the model can detect generation traces; and interpretability reasoning aims to explore whether the model can correctly interpret the judgment criteria. For binary judgment, generation type classification, and generation detail selection, this application uses accuracy and macro F1 score as evaluation metrics to measure overall correctness and performance under class imbalance conditions. Specific detection results are shown in Table 1 below.

[0052] Table 1 As shown in Table 1, even the most powerful detection model struggles to distinguish between real and generated data in the dataset, strongly demonstrating that the audio and video generated using the audio and video generation method of this application has high fidelity.

[0053] This application constructs a multi-stage hybrid generation framework, dividing the audio and video generation process into a proprietary model planning generation task and an expert model executing precise generation. Furthermore, based on the differences in generation mechanisms between compositing and editing methods, this application constructs targeted task planning and modification processes for each type of generation. This application can generate audio and video data covering 11 real-world scenes including human and general themes, and 7 fine-grained audio and video combination types.

[0054] This application addresses the compositing task by obtaining a dynamic description corresponding to the first frame of a video. Similarly, it generates editing instructions for editing tasks. Since the first frame can be of any theme, and the editing instructions can perform arbitrary editing operations on the original video, this application achieves rich generative semantics, thereby generating diverse audio and video content covering multiple real-world themes and supporting diverse generation types. This application effectively overcomes the shortcomings of existing audio and video generation methods in terms of scene diversity and generation capabilities, providing a systematic and scalable technical solution for constructing audio and video generation data for real-world scenes.

[0055] Corresponding to the audio and video generation method in the above embodiments, Figure 6 This is a structural block diagram of an audio / video generation apparatus provided according to an embodiment of this application. For ease of explanation, only the parts relevant to the embodiment of this application are shown. References Figure 6 The audio and video generation device 80 includes: a generation task acquisition module 801 and a fake audio and video generation module 802. The generation task acquisition module 801 is used to acquire audio and video generation tasks, which include: compositing tasks or editing tasks; compositing tasks include: the first frame of the video and its corresponding dynamic description; editing tasks include: editing instructions. The fake audio and video generation module 802 is used to generate fake audio and video based on the generation task. The fake audio and video includes synthesized audio and video and edited audio and video. It includes: generating a synthesized video based on the first frame of the video and the corresponding dynamic description, and generating synthesized audio based on the synthesized video to obtain synthesized audio and video; or, performing editing operations on selected frames of the real video based on editing instructions, generating an edited video based on the edited selected frames, and generating edited audio based on the edited video to obtain edited audio and video.

[0056] The method for generating audio and video by coordinating the various modules of the device is described in the method embodiment and will not be repeated here.

[0057] See Figure 7 , Figure 7 This is a schematic block diagram of an electronic device provided according to an embodiment of this application. Figure 7 The electronic device 300 shown in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304.

[0058] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0059] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.

[0060] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory. For example, the memory 304 may also store device type information.

[0061] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation methods described in the audio and video generation methods provided in the embodiments of this application, or they can execute the implementation methods of the electronic devices described in the embodiments of this application, which will not be repeated here.

[0062] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0063] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0064] This application provides a computer program product, which includes computer-executable instructions or a computer program. The computer-executable instructions or computer program are stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the audio and video generation method described above in this application.

[0065] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0066] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the electronic devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0067] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces or units, or it may be an electrical, mechanical, or other form of connection.

[0068] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.

[0069] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0070] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for generating audio and video, characterized in that, include: Obtain an audio / video generation task, wherein the generation task includes: a compositing task or an editing task; the compositing task includes: the first frame of the video and its corresponding dynamic description; the editing task includes: editing instructions; Based on the generation task, a fake audio and video is generated, which includes synthesized audio and video and edited audio and video; the generation of fake audio and video based on the generation task includes: generating a synthesized video based on the first frame of the video and the corresponding dynamic description, and generating synthesized audio based on the synthesized video to obtain synthesized audio and video; or, performing an editing operation on selected frames of the real video based on the editing instructions, generating an edited video based on the edited selected frames, and generating edited audio based on the edited video to obtain edited audio and video.

2. The method as described in claim 1, characterized in that, When the generation task is a synthesis task, obtaining the audio / video generation task includes: Generate dynamic descriptions based on the first frame of real video; The synthesis task is composed of the first frame of the real video and the dynamic description.

3. The method as described in claim 1, characterized in that, When the generation task is a synthesis task, obtaining the audio / video generation task includes: Generate static images based on preset real-world scenes; The static image is used as the first frame of the video to generate a dynamic description based on the first frame of the video. The synthesis task is composed of the first frame of the video and the dynamic description.

4. The method as described in claim 1, characterized in that, When the generation task is an editing task, obtaining the audio / video generation task includes: A predetermined number of sampled frames are obtained by uniformly sampling real video frames. A tampering scheme is generated based on the aforementioned sampled frames; Structured editing instructions are generated based on the aforementioned tampering scheme, and these editing instructions are used as editing tasks.

5. The method as described in claim 1 or 4, characterized in that, The editing instructions include: target area, time window, editing operation, and target object. The step of performing editing operations on selected frames of the real video based on the editing instructions includes: Selected frames are obtained from the real video based on the time window; The editing operation is performed on the target object in the selected frame based on the target region and the editing operation.

6. The method as described in claim 1 or 4, characterized in that, The editing instructions include: target area, time window, editing operation, and target object. The step of performing editing operations on selected frames of the real video based on the editing instructions includes: Selected frames are obtained from the real video based on the time window; The target object in the target region of the selected frame is segmented to obtain the masked video; The editing operation is performed on the target object in the selected frame based on the masked video.

7. The method as described in claim 1, characterized in that, The method further includes: arranging and combining the forged audio and video with real audio and video to obtain target audio and video, wherein the target audio and video includes: real audio and real video, real audio and edited video, real audio and synthesized video, edited audio and real video, edited audio and edited video, synthesized audio and real video, and synthesized audio and synthesized video.

8. An audio / video generation device, characterized in that, include: The generation task acquisition module is used to obtain audio and video generation tasks, which include: compositing tasks or editing tasks; the compositing task includes: the first frame of the video and its corresponding dynamic description; the editing task includes: editing instructions; A fake audio / video generation module is used to generate fake audio / video based on the generation task. The fake audio / video includes synthesized audio / video and edited audio / video. It includes: generating a synthesized video based on the first frame of the video and the corresponding dynamic description, and generating synthesized audio based on the synthesized video to obtain synthesized audio / video; or, performing an editing operation on selected frames of a real video based on the editing instructions, generating an edited video based on the edited selected frames, and generating edited audio based on the edited video to obtain edited audio / video.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.