A video generation method, apparatus, device, and storage medium

By introducing a space-time enhancement network into video generation, combining text and image information, the problem of inaccurate video generation in the prior art is solved, and more efficient and accurate video generation is achieved.

CN117857896BActive Publication Date: 2025-06-24SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410026820.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-08
Publication Date
2025-06-24
Estimated Expiration
2044-01-08

AI Technical Summary

Technical Problem

When the prior art generates video directly through text, the video content is only guided from the spatial dimension, and the time dimension cannot be effectively utilized, resulting in inaccurate videos.

Method used

The space-time enhancement network is adopted, combining text prompt information and image information to guide video generation in the time dimension, and input noise video, text prompt information and mask video through the target model to obtain the target denoising video.

Benefits of technology

It improves the accuracy and efficiency of video generation, and can more accurately capture the time dimension information of video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117857896B_ABST
    Figure CN117857896B_ABST
Patent Text Reader

Abstract

The present invention discloses a video generation method, device, equipment and storage medium. The method includes: obtaining a current mode, a noisy video and text prompt information; inputting the noisy video, the text prompt information and a mask video corresponding to the current mode into a target model to obtain a target denoised video, wherein the target model is obtained by iteratively training a first model with a target sample set, and the target sample set includes: video samples and text annotations in the video samples. Through the technical solution of the present invention, the accuracy and efficiency of the generated video can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of computer technology, and in particular, to a video generation method, apparatus, device, and storage medium. Background Art

[0002] With the continuous development of computer technology, video generation technology is also constantly updated; currently, in order to improve the efficiency of video creation, users can directly generate videos through text, and can obtain videos without having to search for and edit materials, thereby reducing the time required for video production.

[0003] Generating videos directly through text only directly guides the content generated by the video in the spatial dimension. However, for video generation, using text guidance only in the spatial dimension is often insufficient because videos also have a temporal dimension, and the semantics of the text can be confused in some cases, resulting in inaccurate generated videos. Summary of the Invention

[0004] Embodiments of the present invention provide a video generation method, apparatus, device, and storage medium to improve the accuracy and efficiency of generated videos.

[0005] According to one aspect of the present invention, there is provided a video generation method, including:

[0006] Obtain the current mode, a noise video, and text prompt information;

[0007] Input the noise video, the text prompt information, and the mask video corresponding to the current mode into a target model to obtain a target denoised video, where the target model is obtained by iteratively training a first model with a target sample set, and the target sample set includes: video samples and text annotations in the video samples.

[0008] According to another aspect of the present invention, there is provided a video generation apparatus, which includes:

[0009] An obtaining module, configured to obtain the current mode, a noise video, and text prompt information;

[0010] A target denoised video determination module, configured to input the noise video, the text prompt information, and the mask video corresponding to the current mode into a target model to obtain a target denoised video, where the target model is obtained by iteratively training a first model with a target sample set, and the target sample set includes: video samples and text annotations in the video samples.

[0011] According to another aspect of the present invention, there is provided an electronic device, where the electronic device includes:

[0012] At least one processor; and

[0013] A memory communicatively connected to the at least one processor; wherein,

[0014] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the video generation method according to any embodiment of the present invention.

[0015] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the video generation method according to any embodiment of the present invention when executed.

[0016] In the embodiments of the present invention, by obtaining a current mode, a noisy video, and text prompt information; inputting the noisy video, the text prompt information, and a mask video corresponding to the current mode into a target model to obtain a target denoised video, wherein the target model is obtained by iteratively training a first model with a target sample set, and the target sample set includes: video samples and text annotations in the video samples, the accuracy and efficiency of the generated video can be improved.

[0017] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Description of the Drawings

[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 is a flowchart of a video generation method in an embodiment of the present invention;

[0020] Figure 2 is a schematic structural diagram of a first model in an embodiment of the present invention;

[0021] Figure 3 is a schematic structural diagram of a spatio-temporal enhancer sub-network in an embodiment of the present invention;

[0022] Figure 4 is a schematic diagram of hybrid probability mode selection in an embodiment of the present invention;

[0023] Figure 5It is a schematic structural diagram of a video generation device in an embodiment of the present invention;

[0024] Figure 6 It is a schematic structural diagram of an electronic device in an embodiment of the present invention. Detailed implementation manners

[0025] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to users and the authorization of users should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0028] Embodiment 1

[0029] Figure 1 It is a flowchart of a video generation method provided in an embodiment of the present invention. This embodiment is applicable to the situation of video generation. This method can be executed by the video generation device in the embodiment of the present invention, and the device can be implemented in a software and / or hardware manner, such as Figure 1 As shown, the method specifically includes the following steps:

[0030] S110, obtain the current mode, the noise video, and the text prompt information.

[0031] Among them, the current mode can be a generation mode or a prediction mode. The noisy video can be at least two video frames with noise, and the text prompt information can be scene description information related to the noisy video. For example, the text prompt information can be: XX comes to XX city.

[0032] S120, input the noisy video, the text prompt information, and the mask video corresponding to the current mode into the target model to obtain a target denoised video.

[0033] Among them, the target model is obtained by iteratively training the first model with a target sample set, and the target sample set includes: video samples and text annotations in the video samples.

[0034] Among them, the target model may include: an encoder, a spatio-temporal enhancement network, and a decoder.

[0035] Among them, the mask video corresponding to the current mode can be the mask video corresponding to the generation mode or the mask video corresponding to the prediction mode, and the mask probability corresponding to the generation mode is greater than the mask probability corresponding to the prediction mode. The mask video can be a mask video obtained by masking the video input by the user, and the mask video can also be a mask video obtained by masking the video output by the model last time. The embodiment of the present invention does not limit the acquisition method of the mask video.

[0036] Specifically, the method of inputting the noisy video, the text prompt information, and the mask video corresponding to the current mode into the target model to obtain a target denoised video can be: if the current mode is the generation mode, input the noisy video, the text prompt information, and the mask video corresponding to the generation mode into the encoder to obtain the feature information output by the encoder, input the feature information output by the encoder into the spatio-temporal enhancement network to obtain the feature information output by the spatio-temporal enhancement network, and input the feature information output by the spatio-temporal enhancement network into the decoder to obtain the target denoised video. If the current mode is the prediction mode, input the noisy video, the text prompt information, and the mask video corresponding to the prediction mode into the encoder to obtain the feature information output by the encoder, input the feature information output by the encoder into the spatio-temporal enhancement network to obtain the feature information output by the spatio-temporal enhancement network, and input the feature information output by the spatio-temporal enhancement network into the decoder to obtain the target denoised video.

[0037] It should be noted that the spatio-temporal enhancement network includes at least two spatio-temporal enhancement sub-networks. The spatio-temporal enhancement sub-network sequentially includes, from the input to the output direction: a spatial convolutional layer, a spatial self-attention layer, a spatial image interaction attention layer, a spatial text interaction attention layer, a temporal self-attention layer, and a temporal text interaction attention layer.

[0038] Optionally, iteratively training the first model with a target sample set includes:

[0039] Obtain a target sample set, where the target sample set includes: video samples and text annotations in the video samples;

[0040] Add noise to the video samples to obtain noisy video samples;

[0041] Mask the video samples to obtain masked video samples;

[0042] Select any frame from the video samples as visual cue information;

[0043] Input the noisy video samples, the masked video samples, the text annotations in the video samples, and the visual cue information into the first model to obtain a predicted denoised video, where the masked video samples include: the first masked video samples or the second masked video samples, and the probability that the masked video samples are the first masked video samples is greater than the probability that they are the second masked video samples;

[0044] Train the parameters of the first model according to the objective function generated from the predicted denoised video and the video samples;

[0045] Return to perform the operation of inputting the noisy video samples, the masked video samples, the text annotations in the video samples, and the visual cue information into the first model to obtain a predicted denoised video until the target model is obtained.

[0046] Specifically, the method of selecting any frame from the video samples as visual cue information can be: randomly select a frame from the video samples as visual cue information.

[0047] It should be noted that in order to improve the accuracy of the generated videos, when selecting visual cue information, video frames with a proportion of the target object in the video samples greater than a set threshold can be selected as visual cue information.

[0048] Among them, the first masked video samples are the masked video samples corresponding to the generation mode, and the second masked video samples are the masked video samples corresponding to the prediction mode. Since there are multiple types of prediction modes, the second masked video samples include: the masked video samples corresponding to any prediction mode. It should be noted that multiple prediction modes can be included. For example: the prediction modes include: the prediction mode of splicing one frame, the prediction mode of splicing two frames, and the prediction mode of splicing three frames, and so on. Without further elaboration here. That is to say, the prediction modes include: the prediction modes of splicing different frames. The masking probability corresponding to the generation mode is greater than the masking probability corresponding to the prediction mode of splicing any number of frames.

[0049] In addition, during the training of the first model, the number of samples corresponding to the generation mode is greater than the number of samples corresponding to the prediction mode of splicing any number of frames. That is to say, during the model training process, the probability that the masked video sample is the masked video sample corresponding to the generation mode is greater than the probability that the masked video sample is the masked video sample corresponding to the prediction mode of splicing any number of frames.

[0050] Optionally, masking the video sample to obtain a masked video sample includes:

[0051] Performing a first masking on the video sample according to a first masking probability to obtain a first masked video sample;

[0052] Performing a second masking on the video sample according to a second masking probability to obtain a second masked video sample, where the first masking probability is greater than the second masking probability.

[0053] Wherein, the first masking probability is the masking probability corresponding to the generation mode, and the second masking probability is the masking probability corresponding to the prediction mode.

[0054] It should be noted that since the prediction mode includes prediction modes of splicing different frames, the masking probabilities corresponding to the prediction modes of splicing different frames are all different. For example, the first masking probability corresponding to the generation mode may be greater than the masking probability corresponding to the prediction mode of splicing one frame, the masking probability corresponding to the prediction mode of splicing one frame is greater than the masking probability corresponding to the prediction mode of splicing two frames, and the masking probability corresponding to the prediction mode of splicing two frames is greater than the masking probability corresponding to the prediction mode of splicing three frames.

[0055] Optionally, the first model sequentially includes, from input to output: an encoder, a spatio-temporal enhancement network, and a decoder.

[0056] In a specific example, the first model is as Figure 2 shown. The first model sequentially includes, from input to output: an encoder, a spatio-temporal enhancement network, and a decoder. The spatio-temporal enhancement network includes 4 spatio-temporal enhancement sub-networks.

[0057] The first model provided by the embodiments of the present invention can simultaneously input text prompt information and image prompt information. The two types of prompts and the model interact in the spatio-temporal enhancement network, so that the generated video content is aligned with the two types of prompt information at the same time.

[0058] It should be noted that, as Figure 2 shown, the first model further includes: a visual encoder and a text encoder.

[0059] Optionally, input the noisy video sample, masked video sample, text annotation in the video sample, and visual prompt information into the first model to obtain a predicted video frame, including:

[0060] Input the noisy video sample, masked video sample, text annotation in the video sample, and visual prompt information into the encoder to obtain the feature information output by the encoder;

[0061] Determine the time step according to the noise added to the video sample;

[0062] Input the feature information output by the encoder and the time step into the spatio-temporal enhancement network to obtain target feature information;

[0063] Input the target feature information into the decoder to obtain a predicted denoised video.

[0064] Specifically, the way to input the feature information output by the encoder and the time step into the spatio-temporal enhancement network to obtain target feature information can be: input the feature information output by the encoder and the time step into the first spatio-temporal enhancement sub-network to obtain the feature information output by the first spatio-temporal enhancement sub-network, input the feature information output by the first spatio-temporal enhancement sub-network into the second spatio-temporal enhancement sub-network to obtain the feature information output by the second spatio-temporal enhancement sub-network, and repeat the above process until the feature information output by the last spatio-temporal enhancement sub-network, that is, the target feature information, is obtained.

[0065] In a specific example, as Figure 2 shown, input the feature information output by the encoder and the time step into the first spatio-temporal enhancement sub-network to obtain the feature information output by the first spatio-temporal enhancement sub-network, input the feature information output by the first spatio-temporal enhancement sub-network into the second spatio-temporal enhancement sub-network to obtain the feature information output by the second spatio-temporal enhancement sub-network, input the feature information output by the second spatio-temporal enhancement sub-network into the third spatio-temporal enhancement sub-network to obtain the feature information output by the third spatio-temporal enhancement sub-network, and input the feature information output by the third spatio-temporal enhancement sub-network into the fourth spatio-temporal enhancement sub-network to obtain the target feature information.

[0066] In the prior art, in the process of generating a video based on text, usually the text content only directly guides the content generated by the video in the spatial dimension. However, for video generation, using text guidance only in the spatial dimension is often insufficient because the video also has a time dimension, and the semantics of the text may be confused in some cases. Therefore, the embodiments of the present invention propose a spatio-temporal enhancement network, which adds text guidance for generating the video in the time dimension, and in order to enable the model to refer to pictures while generating the video, an image is added to guide the generation of the video in the spatial dimension.

[0067] Considering the limitations of computing efficiency and the video memory of devices, the current general training scheme for text-video generation models is to train using videos within 20 frames. The video generation models obtained in this way can only generate videos within 3 seconds. To generate longer videos, the embodiments of the present invention propose a new hybrid training strategy to simultaneously train the text-video generation ability and the text-video prediction ability of the video generation model, so that after video generation, video prediction can continue to be used to generate videos with a longer time span.

[0068] Optionally, the spatio-temporal enhancement network includes at least two spatio-temporal enhancement sub-networks. The spatio-temporal enhancement sub-network sequentially includes, from the input to the output direction: a spatial convolutional layer, a spatial self-attention layer, a spatial image interaction attention layer, a spatial text interaction attention layer, a temporal self-attention layer, and a temporal text interaction attention layer.

[0069] In a specific example, as Figure 3 shown, the spatio-temporal enhancement sub-network sequentially includes, from the input to the output direction: a spatial convolutional layer, a spatial self-attention layer, a spatial image interaction attention layer, a spatial text interaction attention layer, a temporal self-attention layer, and a temporal text interaction attention layer.

[0070] The text prompt information interacts with the generated content from both the spatial dimension and the temporal dimension. The interaction method is cross-attention, that is, extracting corresponding information from the text to the spatial dimension and the temporal dimension, while the image information interacts with the generated content from the spatial dimension, and the interaction method is also cross-attention, that is, extracting corresponding information from the image to the spatial dimension.

[0071] Optionally, inputting the feature information output by the encoder and the time step into the spatio-temporal enhancement network to obtain target feature information includes:

[0072] Inputting the feature information output by the encoder and the time step into the first spatio-temporal enhancement sub-network to obtain the feature information output by the first spatio-temporal enhancement sub-network;

[0073] Inputting the feature information output by the Nth spatio-temporal enhancement sub-network into the (N + 1)th spatio-temporal enhancement sub-network to obtain the feature information output by the (N + 1)th spatio-temporal enhancement sub-network, where N is a positive integer greater than 1;

[0074] Determining the feature information output by the last spatio-temporal enhancement sub-network as the target feature information.

[0075] Specifically, if the spatio-temporal enhancer network is the first spatio-temporal enhancer network, the feature information output by the encoder and the time step are input into the first spatio-temporal enhancer network to obtain the feature information output by the first spatio-temporal enhancer network. If the spatio-temporal enhancer network is not the first spatio-temporal enhancer network, the feature information output by the Nth spatio-temporal enhancer network is input into the (N + 1)th spatio-temporal enhancer network to obtain the feature information output by the (N + 1)th spatio-temporal enhancer network.

[0076] Optionally, inputting the feature information output by the encoder and the time step into the first spatio-temporal enhancer network to obtain the feature information output by the first spatio-temporal enhancer network includes:

[0077] Inputting the feature information output by the encoder and the time step into a spatial convolutional layer to obtain first feature information;

[0078] Inputting the first feature information into a spatial self-attention layer to obtain second feature information;

[0079] Inputting visual cue information into a visual encoder to obtain visual features;

[0080] Inputting the text annotation in the video sample into a text encoder to obtain text features;

[0081] Inputting the second feature information and the visual features into a spatial image interaction attention layer to obtain third feature information;

[0082] Inputting the second feature information and the text features into a spatial text interaction attention layer to obtain fourth feature information;

[0083] Inputting the superimposed third feature information and fourth feature information into a temporal self-attention layer to obtain fifth feature information;

[0084] Inputting the fifth feature information and the text features into a temporal text interaction attention layer to obtain the feature information output by the first spatio-temporal enhancer network.

[0085] Optionally, inputting the feature information output by the Nth spatio-temporal enhancer network into the (N + 1)th spatio-temporal enhancer network to obtain the feature information output by the (N + 1)th spatio-temporal enhancer network includes:

[0086] Inputting the feature information output by the Nth spatio-temporal enhancer network and the time step into a spatial convolutional layer to obtain the feature information output by the spatial convolutional layer;

[0087] Inputting the feature information output by the spatial convolutional layer into a spatial self-attention layer to obtain the feature information output by the spatial self-attention layer;

[0088] Inputting visual cue information into a visual encoder to obtain visual features;

[0089] Input the text annotation in the video sample into the text encoder to obtain text features;

[0090] Input the feature information output by the spatial self-attention layer and the visual features into the spatial image interaction attention layer to obtain the feature information output by the spatial image interaction attention layer;

[0091] Input the feature information output by the spatial self-attention layer and the text features into the spatial text interaction attention layer to obtain the feature information output by the spatial text interaction attention layer;

[0092] Overlay the feature information output by the spatial image interaction attention layer and the feature information output by the spatial text interaction attention layer and input them into the temporal self-attention layer to obtain the feature information output by the temporal self-attention layer;

[0093] Input the feature information output by the temporal self-attention layer and the text features into the temporal text interaction attention layer to obtain the feature information output by the (N + 1)-th spatio-temporal enhancement sub-network.

[0094] It should be noted that the mode selection mechanism for video generation can select the training of generation and prediction according to a certain probability distribution during training, and it is necessary to balance the training ratio of the two mechanisms, because generation is more a process from scratch, while prediction requires extracting more information from the previous frames.

[0095] In a specific example, as Figure 4 shown, the mode selection mechanism regulates the first K frames of pictures that the model can see during training through a masking method under a certain probability mechanism, ensuring that the model is endowed with the ability to generate text-to-video and the ability to predict during the training process.

[0096]

[0097] Among them, α is a preset parameter, k is the masking probability, and m is the maximum masking probability.

[0098] The technical solution of this embodiment, by obtaining the current mode, the noisy video, and the text prompt information; inputting the noisy video, the text prompt information, and the masked video corresponding to the current mode into the target model to obtain the target denoised video, where the target model is obtained by iteratively training the first model with a target sample set, and the target sample set includes: video samples and text annotations in the video samples, can improve the accuracy and efficiency of generating videos.

[0099] Embodiment 2

[0100] Figure 5The figure is a schematic structural diagram of a video generation device provided by an embodiment of the present invention. This embodiment is applicable to the situation of video generation. The device can be implemented in software and / or hardware, and can be integrated into any device that provides video generation functions, such as Figure 5 As shown, the video generation device specifically includes: an acquisition module 210 and a target denoised video determination module 220.

[0101] Among them, the acquisition module is used to acquire the current mode, the noisy video, and the text prompt information;

[0102] The target denoised video determination module is used to input the noisy video, the text prompt information, and the mask video corresponding to the current mode into the target model to obtain the target denoised video, where the target model is obtained by iteratively training the first model with a target sample set, and the target sample set includes: video samples and text annotations in the video samples.

[0103] The above product can execute the method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0104] Embodiment III

[0105] Figure 6 The figure shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0106] As Figure 6 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0107] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0108] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the video generation method.

[0109] In some embodiments, the video generation method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the video generation method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the video generation method in any other suitable manner (e.g., by means of firmware).

[0110] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, the one or more computer programs being executable and / or interpretable on a programmable system including at least one programmable processor, the programmable processor can be a special or general programmable processor, can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0111] A computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0112] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0113] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, voice input, or tactile input).

[0114] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0115] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0116] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0117] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A video generation method, characterized in that: include: Get the current mode, noise video and text prompt information; Inputting the noisy video, text prompt information, and the mask video corresponding to the current mode into a target model to obtain a target denoised video, wherein the target model is obtained by iteratively training a first model with a target sample set, and the target sample set includes: video samples and text annotations in the video samples; The first model includes, from input to output, an encoder, a spatiotemporal enhancement network, and a decoder; The spatiotemporal enhancement network includes at least two spatiotemporal enhancement sub-networks, which include, from input to output, a spatial convolution layer, a spatial self-attention layer, a spatial image interactive attention layer, a spatial text interactive attention layer, a temporal self-attention layer, and a temporal text interactive attention layer.

2. The method according to claim 1, characterized in that: The first model is iteratively trained using the target sample set, including: Acquire a target sample set, wherein the target sample set includes: video samples and text annotations in the video samples; Adding noise to the video sample to obtain a noisy video sample; Masking the video sample to obtain a masked video sample; Select any frame from the video sample as visual cue information; Inputting the noisy video sample, the masked video sample, the text annotation in the video sample, and the visual prompt information into the first model to obtain the predicted denoised video, wherein the masked video sample includes: the first masked video sample or the second masked video sample, and the probability that the masked video sample is the first masked video sample is greater than the probability that it is the second masked video sample; Training parameters of the first model according to an objective function generated by predicting denoised videos and video samples; Return to execute the operation of inputting the noisy video sample, the masked video sample, the text annotation in the video sample, and the visual prompt information into the first model to obtain the predicted denoised video until the target model is obtained.

3. The method according to claim 1, characterized in that Masking the video sample to obtain the masked video sample includes: Performing a first mask on the video sample according to a first mask probability to obtain a first masked video sample; Performing a second mask on the video sample according to a second mask probability to obtain a second masked video sample, wherein the first mask probability is greater than the second mask probability.

4. The method according to claim 1, characterized in that: The noisy video sample, the masked video sample, the text annotation in the video sample, and the visual prompt information are input into the first model to obtain a predicted video frame, including: Inputting the noisy video samples, the masked video samples, the text annotations in the video samples, and the visual prompt information into the encoder to obtain the feature information output by the encoder; determining a time step based on the noise added to the video sample; Inputting the feature information output by the encoder and the time step into a spatiotemporal enhancement network to obtain target feature information; The target feature information is input into a decoder to obtain a predicted denoised video.

5. The method according to claim 4, characterized in that Inputting the feature information output by the encoder and the time step into a spatiotemporal enhancement network to obtain target feature information, including: Inputting the feature information output by the encoder and the time step into a first spatiotemporal enhancement sub-network to obtain the feature information output by the first spatiotemporal enhancement sub-network; Input the feature information output by the Nth spatiotemporal enhancer subnetwork into the N+1th spatiotemporal enhancer subnetwork to obtain the feature information output by the N+1th spatiotemporal enhancer subnetwork, where N is a positive integer greater than 1; The feature information output by the last spatiotemporal enhancement sub-network is determined as the target feature information.

6. The method according to claim 5, characterized in that Inputting the feature information output by the encoder and the time step into the first spatiotemporal enhancement sub-network to obtain the feature information output by the first spatiotemporal enhancement sub-network, including: Inputting the feature information output by the encoder and the time step into a spatial convolution layer to obtain first feature information; Input the first feature information into the spatial self-attention layer to obtain the second feature information; Input the visual cue information into the visual encoder to obtain visual features; Input the text annotations in the video sample into the text encoder to obtain text features; Input the second feature information and the visual feature into the spatial image interactive attention layer to obtain the third feature information; Input the second feature information and the text feature into the spatial text interaction attention layer to obtain the fourth feature information; The third feature information and the fourth feature information are superimposed and input into the temporal self-attention layer to obtain the fifth feature information; The fifth feature information and text features are input into the temporal text interaction attention layer to obtain the feature information output by the first spatiotemporal enhancement sub-network.

7. A video generating device, characterized in that: include: The acquisition module is used to obtain the current mode, noise video and text prompt information; A target denoised video determination module is used to input the noisy video, text prompt information and the mask video corresponding to the current mode into a target model to obtain a target denoised video, wherein the target model is obtained by iteratively training a first model with a target sample set, and the target sample set includes: video samples and text annotations in the video samples; The first model includes, from input to output, an encoder, a spatiotemporal enhancement network, and a decoder; The spatiotemporal enhancement network includes at least two spatiotemporal enhancement sub-networks, which include, from input to output, a spatial convolution layer, a spatial self-attention layer, a spatial image interactive attention layer, a spatial text interactive attention layer, a temporal self-attention layer, and a temporal text interactive attention layer.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the video generating method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the video generation method according to any one of claims 1 to 6 when executed.

Citation Information

Patent Citations

  • Long video generation method based on background transition

    CN117354443A

  • Systems and methods for video and language pre-training

    US20230154188A1