A video generation method, apparatus, electronic device, computer-readable storage medium, and computer program product

By extracting semantic features from the global video and segment description text during the video generation process and using iterative generation of variable block sequences, the problem of insufficient correlation between video segments and the global video is solved, thereby improving video quality.

CN119316683BActive Publication Date: 2025-11-21TENCENT DIGITAL TIANJIN
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411577669.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-06
Publication Date
2025-11-21
Estimated Expiration
2044-11-06

AI Technical Summary

Technical Problem

In existing technologies, there is a lack of correlation between the latent variable subsequences of video clips and the descriptive text of the complete video, resulting in low correlation between the generated video clips and the original video footage, and thus low video quality.

Method used

Through iterative processing, semantic features are determined from the global video and the segment description text, respectively. The output variable block sequence is generated using the intermediate variable block sequence and the position sequence. Combining the time interval and semantic features, the target video is generated step by step, realizing the association between the video segment and the global video.

Benefits of technology

It improves the quality of generated videos, making the transitions between video clips and the complete video more natural, thus enhancing the overall quality of the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119316683B_ABST
    Figure CN119316683B_ABST
Patent Text Reader

Abstract

The application provides a video generation method and device, electronic equipment, computer readable storage medium and computer program product; the application embodiment can be applied to the video generation scene of game, virtual reality, video processing and the like; the method comprises: determining a first semantic feature from a first description text describing a whole video, and determining a second semantic feature from a second description text describing a video segment; the following processing is performed through iteration i: generating an i-th intermediate variable block sequence by using an i-th input variable block sequence, the first semantic feature and a position sequence; determining an i-th output variable block sequence based on the i-th intermediate variable block sequence, a time interval corresponding to the video segment, the position sequence and the second semantic feature; and determining a target video based on an N-th output variable block sequence and the position sequence. Through the application, the quality of the generated video can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a video generation method and device, electronic equipment, computer readable storage medium and computer program product. BACKGROUND

[0002] Video generation is an important application direction of artificial intelligence. This technology can automatically generate a matching video according to the description text input by a user. In order to meet the requirements of the user as much as possible through the description text during video generation, the video segments need to be finely controlled during the video generation process. However, in the related art, there is a lack of association between the latent variable subsequence of the video segment and the description text of the complete video, which leads to a low correlation between the picture of the generated video segment and the picture of the original video, and ultimately results in a low quality of the generated video. SUMMARY

[0003] The embodiments of the present application provide a video generation method and device, electronic equipment, computer readable storage medium and computer program product, which can improve the quality of the generated video.

[0004] The technical solutions of the embodiments of the present application are as follows:

[0005] The embodiments of the present application provide a video generation method, which comprises:

[0006] determining a first semantic feature from a first description text describing a global video, and determining a second semantic feature from a second description text describing a video segment;

[0007] perform the following processing through iteration i, 1≤i≤N, N is a positive integer:

[0008] generating an intermediate variable block sequence of the i-th time using the input variable block sequence of the i-th time, the first semantic feature and the position sequence; when i = 1, the input variable block sequence of the i-th time includes a latent variable block obtained by blocking the latent variables of an initial latent variable sequence, and when i > 1, the input variable block sequence of the i-th time includes the output variable block sequence of the i-1-th time, and the position sequence includes position information of each latent variable block;

[0009] determining an output variable block sequence of the i-th time based on the intermediate variable block sequence of the i-th time, a time interval corresponding to the video segment, the position sequence and the second semantic feature;

[0010] determining a target video indicated by the first description text and the second description text based on the output variable block sequence of the N-th time and the position sequence.

[0011] The embodiment of the application provides a video generation device, comprising:

[0012] The feature determination module is configured to determine a first semantic feature from a first description text describing a video globally, and determine a second semantic feature from a second description text describing a video segment;

[0013] The sequence determination module is configured to generate an ith intermediate variable block sequence by using an ith input variable block sequence, the first semantic feature and a position sequence; when i = 1, the ith input variable block sequence comprises a latent variable block obtained by blocking latent variables of an initial latent variable sequence, when i > 1, the ith input variable block sequence comprises an (i-1)th output variable block sequence, and the position sequence comprises position information of each latent variable block; and determine an ith output variable block sequence based on the ith intermediate variable block sequence, a time interval corresponding to the video segment, the position sequence and the second semantic feature;

[0014] The video determination module is configured to determine a target video indicated by the first description text and the second description text based on the Nth output variable block sequence and the position sequence.

[0015] In the above scheme, the sequence determination module is further configured to determine a first sub-sequence from the ith intermediate variable block sequence according to the time interval and the position sequence, and determine a position sub-sequence corresponding to the first sub-sequence from the position sequence; generate a second sub-sequence based on the second semantic feature, the first sub-sequence and the position sub-sequence; and determine the ith output variable block sequence by using the second sub-sequence and the ith intermediate variable block sequence.

[0016] In the above scheme, the second semantic feature comprises a multi-level text feature of the second description text and a single-level text feature of the second description text; and the sequence determination module is further configured to fuse the multi-level semantic feature of the second description text into each latent variable block of the first sub-sequence, and generate a third sub-sequence by using the fused latent variable block; fuse the multi-level text feature of the second description text and the single-level text feature of the second description text to obtain a fusion feature; and generate the second sub-sequence based on the third sub-sequence and the fusion feature.

[0017] In the above scheme, the sequence determining module is further configured to determine a control identifier for the latent variable block in the i th intermediate variable block sequence according to the time interval and the time position in the position sequence, where the control identifier indicates whether the latent variable block needs to be regenerated according to the second description text; extract the latent variable block indicated by the control identifier as needing to be regenerated according to the second description text from the i th intermediate variable block sequence, and determine a sequence generated by using the extracted latent variable block as the first subsequence.

[0018] In the above scheme, the sequence determining module is further configured to replace the first subsequence in the i th intermediate variable block sequence by using the second subsequence to obtain an i th to-be-optimized variable block sequence; and perform smoothing on the i th to-be-optimized variable block to obtain the i th output variable block sequence.

[0019] In the above scheme, the feature determining module is further configured to extract K second text features from the second description text, where K is a positive integer; and determine the second semantic feature based on the K second text features.

[0020] In the above scheme, the K second text features include multi-level text features of the second description text; the feature determining module is further configured to encode the second description text to obtain second encoding features; extract second text sub-features at multiple levels from the second encoding features; and determine the multi-level text features of the second description text by using the second text sub-features at multiple levels.

[0021] In the above scheme, the video generating module further includes a variable blocking module configured to divide each latent variable in the initial latent variable sequence into a plurality of latent variable blocks; and generate the 1 st input variable block sequence by using the plurality of latent variable blocks of each latent variable.

[0022] In the above scheme, the video determining module is further configured to restore the N th output variable block sequence to an output latent variable sequence according to the position sequence; decode the output latent variable sequence to obtain a video frame sequence; and determine the target video indicated by the first description text and the second description text based on the video frame sequence.

[0023] In the above scheme, i=M, and the sequence determining module is further configured to determine the M th intermediate variable block sequence as the M th output variable block sequence, where M can be any one or more values between 1 and N.

[0024] Embodiments of the present application provide an electronic device apparatus, which comprises:

[0025] a memory for storing computer executable instructions or computer programs;

[0026] a processor for implementing the video generation method provided by the embodiments of the present application when executing the computer executable instructions or computer programs stored in the memory.

[0027] The embodiments of the present application provide a computer readable storage medium storing computer programs or computer executable instructions for implementing the video generation method provided by the embodiments of the present application when executed by a processor.

[0028] The embodiments of the present application provide a computer program product comprising computer programs or computer executable instructions, which, when executed by a processor, implement the video generation method provided by the embodiments of the present application.

[0029] The embodiments of the present application have the following beneficial effects: in each iteration, first, the first semantic feature of the first description text describing the whole video, the input variable block sequence and the position sequence of each iteration are used to generate the intermediate variable block sequence, so as to realize the injection of the first semantic feature, and then, on the basis of the intermediate variable block sequence, the time interval of the video segment and the second semantic feature of the second description text describing the video segment are combined to generate the output variable block sequence of each iteration, so as to realize the injection of the second semantic feature, so that the latent variable block corresponding to the video segment can be associated with the first description text, and the process is repeatedly iterated, so that the semantic features of the first description text and the second description text can be fully injected into the latent variable block, so as to make the picture of the video segment and the picture of the complete video transition naturally, and thus the quality of the generated video is improved. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 is an architecture schematic diagram of the video generation system provided by the embodiments of the present application;

[0031] Figure 2 is a structure schematic diagram of the server in Figure 1 provided by the embodiments of the present application;

[0032] Figure 3 is a flow schematic diagram of the video generation method provided by the embodiments of the present application Figure 1 ;

[0033] Figure 4 is a flow schematic diagram of the video generation method provided by the embodiments of the present application Figure 2 ;

[0034] Figure 5 is a flow schematic diagram of the video generation method provided by the embodiments of the present application Figure 3;

[0035] Figure 6A is a short video generation schematic provided by an embodiment of the present application Figure 1 ;

[0036] Figure 6B is a short video generation schematic provided by an embodiment of the present application Figure 2 ;

[0037] Figure 7 is a video generation system architecture diagram provided by an embodiment of the present application

[0038] Figure 8 is a schematic diagram of overall prompt word encoding and segment prompt word encoding provided by an embodiment of the present application

[0039] Figure 9 is a model architecture diagram of a video generation part provided by an embodiment of the present application

[0040] Figure 10 is a structure schematic of a video generation model provided by an embodiment of the present application Figure 1 ;

[0041] Figure 11 is a generation schematic diagram of marking information provided by an embodiment of the present application

[0042] Figure 12 is a structure schematic of a video generation model provided by an embodiment of the present application Figure 2 . DETAILED DESCRIPTION

[0043] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings, and the described embodiments should not be regarded as limiting the present application, and all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present application.

[0044] In the following description, "some embodiments" are described, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0045] In the following description, the terms "first\second\third" are only to distinguish similar objects, and do not represent a specific order of the objects, and it can be understood that "first\second\third" can be interchanged with a specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0046] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that includes the functions of the module or unit.

[0047] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as commonly understood by one of ordinary skill in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0048] The relevant data collection process in the embodiments of the present application should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of authorization of laws and regulations and the personal information subject.

[0049] Before the embodiments of the present application are further described in detail, the terms and phrases involved in the embodiments of the present application are explained, and the terms and phrases involved in the embodiments of the present application are applicable to the following explanations.

[0050] 1) Video generation refers to outputting a video matching the description text according to the input text (hereinafter referred to as description text) describing the video through artificial intelligence processing. For example, the text input by the user is "a white rabbit suddenly appears in the deep forest", and the generated video can contain a green forest picture and a picture of a white rabbit jumping out at a certain part of the forest.

[0051] 2) Description text refers to a text used to describe the content of a video. In the embodiments of the present application, there can be multiple different description texts for the same video, such as a first description text and a second description text. The first description text is a text describing the complete video, i.e., a global video, which provides information about the theme, overall style, etc. of the video to set a unified background or framework for the video. The second description text is a text describing a specific part of the video, i.e., a text describing a video segment, which can describe a specific scene, action or detail in the video in detail to finely control the local content in video generation. For example, "generate a romantic sunset beach walking video" can be used as the first description text, and "in the video, when the main character walks to the beach, show flying seagulls and floating waves" can be used as the second description text.

[0052] 3) latent sequence, is a sequence generated by latent variables, each latent variable in the latent sequence corresponds to a video frame in the video frame sequence. Here, the latent variable can be regarded as a variable that is not directly observed in the video generation process, that is, the most basic feature in the video generation process, which can be obtained by adding noise or removing noise to obtain a new video frame. In other words, the latent variable is an abstract representation of the video frame, which contains the motion pattern, texture information and other key features of the video frame, and can be used as the basis for generating new video frames (for example, adding a description text to the existing latent variable to generate a new latent variable, and decoding the latent variable to obtain a new video frame).

[0053] 4) position sequence, is a sequence generated by the position information of the latent variable block. The latent variable block is obtained by blocking the latent variable, and the position information of the latent variable block includes the spatial position of the latent variable block, that is, the position of the latent variable block in the latent variable, and also includes the time position of the latent variable block, that is, the appearance time of the latent variable block.

[0054] 5) control identifier, is an identifier used to distinguish whether the latent variable block needs to be regenerated according to the second description text of the video segment. In other words, each latent variable block has its corresponding control identifier, and the control identifier represents whether the latent variable block is a latent variable block that needs to be finely controlled.

[0055] 6) Diffusion Transformer (DIT) model, is a deep learning model combining diffusion model and transformer model, which uses the powerful sequence modeling ability of transformer and the advantage of diffusion model in data generation to generate high-quality videos. In DIT, the diffusion process starts from a simple distribution (such as Gaussian noise), and through a series of steps, the simple distribution is transformed into the target distribution (for example, the distribution of the training data), each step can be regarded as adding some noise; Then use the transformer to predict the next distribution. This process can be regarded as a reverse Markov chain, which starts from noise and gradually changes the distribution to generate complex samples.

[0056] 7) Contrastive Language Image Pre-training (CLIP) is a multi-modal deep learning model that learns the relationship between images and text through a process of contrastive learning on large-scale text-image pairs. During pre-training, the model learns to encode both images and text into a unified vector space, enabling it to understand image content and match images with the text that describes them. After pre-training, CLIP can recognize objects, scenes, actions, and other elements in images, while also understanding text related to the images, such as labels, descriptions, and captions.

[0057] 8) Text-to-Text Transfer Transformer (T5) is a pre-trained Natural Language Processing (NLP) model based on the Transformer architecture. T5 models all NLP tasks as text-to-text conversion problems, allowing it to accept any form of text input and generate corresponding text output.

[0058] 9) Cross Attention is a mechanism used in models to handle the relationship between multiple input sequences, capturing important interactions between different sequences. Cross attention helps the model better understand the relationships between different parts, leading to more accurate predictions of changes and influences during the diffusion process.

[0059] 10) Self-Attention is a mechanism that assigns different weights to each element in a sequence, representing the relevance of that element to others. This allows the model to dynamically focus on other elements in the sequence when processing each element. For example, in a diffusion model, self-attention is applied to latent variables or latent variables that have been fused with semantic features, allowing the semantic features in the latent variables to be more prominent and improving the video generation results.

[0060] Video generation is an important application direction of artificial intelligence. The technology can automatically generate a matching video according to the description text input by a user. In video generation, in order to meet the requirements of the user as much as possible through the description text, the video segments need to be finely controlled in the video generation process. In the related art, video generation can be achieved by the following processing: feature extraction is performed on the global description text of the video, and the obtained semantic features are injected into a model capable of video generation together with an initialized latent variable sequence, so that the model can generate a new latent variable sequence that fuses the semantic features of the description text, and then the model performs noise reduction on the new latent variable sequence to obtain a video that meets the requirements of the description text.

[0061] For example, a time-aware self-attention mechanism is added to a U-shaped network (UNet) of a stable diffusion model, so that the video generation capability of the stable diffusion model is improved. In this scheme, the input is the description text of the entire video and the latent variable sequence initialized by the stable diffusion model. Each latent variable in the latent variable sequence is a basis for generating a video frame. After the description text is encoded by an encoder in a CLIP model, relevant features are generated, and then the features are injected into a UNet combined with a video generation technique together with the latent variable sequence. The UNet generates a new latent variable sequence, and finally, noise reduction is performed on the new latent variable sequence to obtain a video that meets the requirements of the description text of the user.

[0062] Among them, for the video segments that need to be finely controlled, the corresponding latent variable sub-sequence can be queried in the latent variable sequence according to the time interval of the video segment, and then the features of the description text of the video segment and the latent variable sub-sequence are injected into a model (also based on the stable diffusion model structure) specially for sub-content generation. The sub-content generation model generates a new latent variable sub-sequence for the video segment, and the original latent variable sub-sequence is replaced using the latent variable sub-sequence, thereby completing the fine control of the video segment.

[0063] Alternatively, the video segments that need to be finely controlled can also be generated twice. For example, using the video generation method in the foregoing, all video frame sequences are first generated, and then the video frames in the video segments that need to be finely controlled are extracted and encoded into a new latent variable sequence using a variational autoencoder (VAE) in the stable diffusion model, so as to be input to the stable diffusion model again. The description text of the video segment is also input to the stable diffusion model. In this way, the stable diffusion model can regenerate the video frames of the video segment, and finally the generated video frames are inserted into the original time position, thereby completing the generation of the entire video.

[0064] However, the above scheme can control the video content in the specified time segment, but the generation of the latent variable sequence of the complete video and the generation of the latent variable sub-sequence of the video segment can be regarded as two independent latent variable sequence generations, so that the latent variable sub-sequence of the video segment and the description text of the complete video lack association, resulting in that the generated video segment has weak association with the original video, and the quality of the generated video is low. In addition, in the above scheme, the generated latent variable sub-sequence is directly inserted into the original latent variable sequence, or the generated video segment is directly replaced with the original video segment, which can easily cause serious picture mutation, and finally the quality of the generated video is low.

[0065] The embodiment of the present application provides a video generation method and device, electronic equipment, computer readable storage medium and computer program product, which can improve the quality of the generated video. The following describes an exemplary application of the electronic equipment provided by the embodiment of the present application. The electronic equipment provided by the embodiment of the present application can be implemented as a notebook computer, a tablet computer, a desktop computer, a set-top box, a smart phone, a smart watch, a smart television, a vehicle-mounted terminal, and various types of terminals. It can also be implemented as a server. The following describes an exemplary application when the electronic equipment is implemented as a server.

[0066] Referring to Figure 1 , Figure 1 is an architecture diagram of a video generation system provided by the embodiment of the present application. To realize a video generation application, in the video generation system 100, the terminal 400 (exemplarily shows the terminal 400-1 and the terminal 400-2) is connected to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two. It should be noted that in the video generation system 100, a database 500 is also provided to provide data support to the server 200. The database 500 can be configured in the server 200, and can also be independent of the server 200, Figure 1 It is shown that the database 500 is independent of the server 200.

[0067] The terminal 400 is used to obtain the first description text and the second description text in response to the input operation of the user in the text input interface of the graphical interface (exemplarily shows the graphical interface 400-11 and the graphical interface 400-21), and send the first description text and the second description text to the server 200 through the network 300.

[0068] The server 200 is configured to determine a first semantic feature from a first description text describing a video as a whole, and determine a second semantic feature from a second description text describing a video segment; perform the following processing through iteration i, 1≤i≤N, N being a positive integer: generate an i-th intermediate variable block sequence by using an i-th input variable block sequence, the second semantic feature and a position sequence, wherein when i=1, the i-th input variable block sequence comprises a latent variable block obtained by blocking latent variables of an initial latent variable sequence, when i>1, the i-th input variable block sequence comprises an (i-1)-th output variable block sequence, and the position sequence comprises position information of each latent variable block; determine an i-th output variable block sequence based on the i-th intermediate variable block sequence, a time interval corresponding to the video segment, the position sequence and the second semantic feature; determine a target video indicated by the first description text and the second description text based on an N-th output latent variable block sequence; and deliver the target video to the terminal 400.

[0069] The terminal 400 is further configured to display the target video in the graphical interface 400-11 and the graphical interface 400-21.

[0070] In some embodiments, the server 200 can be a stand-alone physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal and the server can be connected directly or indirectly through wired or wireless communication, which is not limited in the embodiments of the present application.

[0071] Referring to Figure 2 , Figure 2 is a structural schematic diagram of the server (one implementation of an electronic device) in Figure 1 provided by the embodiments of the present application, Figure 2 The server 200 shown in FIG. 2 includes at least one processor 210, a memory 250, and at least one network interface 220. The various components in the server 200 are coupled together by a bus system 240. It can be understood that the bus system 240 is used to realize the connection and communication between the components. In addition to the data bus, the bus system 240 also includes a power bus, a control bus and a status signal bus. However, for the purpose of clear illustration, all kinds of buses are marked as the bus system 240 in Figure 2 .

[0072] The processor 210 can be an integrated circuit chip that has the processing capability of signals, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor.

[0073] The memory 250 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 250 optionally includes one or more storage devices remotely located from the processor 210 in a physical location.

[0074] The memory 250 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), and the volatile memory can be random access memory (RAM). The memory 250 described in the embodiments of the present application is intended to include any suitable type of memory.

[0075] In some embodiments, the memory 250 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, which are exemplarily illustrated below.

[0076] The operating system 251 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0077] The network communication module 252 is used to communicate with other electronic devices via one or more (wired or wireless) network interfaces 220, and exemplary network interfaces 220 include Bluetooth, wireless compatibility certification (WiFi), and universal serial bus (USB), etc.

[0078] In some embodiments, the video generation apparatus provided by the embodiments of the present application can be realized in a software manner, Figure 2 The video generation apparatus 255 stored in the memory 250 is shown, which can be software in the form of programs and plug-ins, etc., including the following software modules: feature determination module 2551, sequence determination module 2552, video determination module 2553, and variable blocking module 2554, which are logical, and thus can be combined or further split according to the implemented functions. The functions of each module will be described below.

[0079] In some embodiments, the video generation apparatus provided by the embodiments of the present application can be implemented in a hardware manner. For example, the video generation apparatus provided by the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the video generation method provided by the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can be implemented by using one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic elements.

[0080] In some embodiments, the terminal or the server (both of which are possible implementations of electronic devices) can implement the video generation method provided by the embodiments of the present application by running various computer-executable instructions or computer programs. For example, the computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. The computer programs can be native programs or software modules in an operating system; can be native applications (APPs) that need to be installed in an operating system to run, such as video making APPs or short video APPs; or can be applets that can be embedded into any APP, that is, programs that only need to be downloaded into a browser environment to run. In summary, the above computer-executable instructions can be any form of instructions, and the above computer programs can be any form of application programs, modules, or plug-ins.

[0081] In the following, the video generation method provided by the embodiments of the present application will be described in combination with exemplary applications and implementations of electronic devices provided by the embodiments of the present application. As described above, the electronic device implementing the video generation method of the embodiments of the present application can be a terminal, a server, or a combination of the two. Therefore, the execution subject of each step will not be repeated in the following description.

[0082] Referring to Figure 3 , Figure 3 is a flowchart of the video generation method provided by the embodiments of the present application Figure 1 will be described in combination with the steps shown in Figure 3 , Figure 3 The subject of the step is an electronic device.

[0083] Step 101, determining a first semantic feature from a first description text describing a video as a whole, and determining a second semantic feature from a second description text describing a video segment.

[0084] Embodiments of the present application are implemented in the scenario of generating a matched video according to an input text of a user. The input text of the user includes a first description text describing a video as a whole, and a second description text describing a video segment in the video. An electronic device extracts semantic features from the first description text and the second description text respectively to obtain a first semantic feature reflecting the text meaning and semantic information of the first description text, and to obtain a second semantic feature reflecting the text meaning and semantic information of the second description text.

[0085] It should be noted that the video as a whole herein refers to the entire video, so that the first description text outlines information related to the overall content of the video to be generated, such as main theme, emotion, style, etc. It can be seen that the purpose of the first description text is to provide a global semantic framework for video generation. The video segment refers to a segment in the video to be generated, for example, the 2nd to 4th second of the video, so that the second description text describes the content, action or change in the time period in detail, such as a specific event in the scene, the action of a character in the video, expression, etc. It can be seen that the purpose of the second description text is to finely control the video of a specific time segment to obtain more personalized video content.

[0086] For example, the first description text input by the user can be "a group of people walking on the beach in the summer sunset", which specifies the main scene and participants of the video and provides a direction for video generation; the second description text input by the user can be "seagulls flying in the distance", which describes a specific segment in the video in more detail to guide the content of the specific segment.

[0087] It should be noted that the electronic device can extract semantic features from the first description text and the second description text in a variety of different ways. Below, first the extraction process of the second semantic feature is described.

[0088] In some embodiments of the present application, Figure 3 In step 101, the second semantic feature is determined from the second description text describing the video segment, which can be achieved by: extracting K second text features from the second description text; determining the second semantic feature based on the K second text features. Here, K is a positive integer.

[0089] The electronic device first extracts K different features for the second description text, to obtain K different second text features for the same second description text. Then, the K second text features can be directly used to generate the second semantic feature, to complete the semantic feature extraction of the second description text.

[0090] Here, the electronic device can implement N times of feature extraction on the second description text through K different text encoders, for example, respectively using a text encoder in a contrastive language image pre-training (CLIP) model and a text encoder in a text-to-text transfer transformer (T5) model to extract features from the second description text to obtain two different second text features.

[0091] The electronic device can also implement K times of feature extraction on the second description text using K different content feature analysis methods, for example, respectively using entity (such as location, item name) recognition to extract features from the second description text and using keyword extraction to extract features from the second description text, so as to obtain two different second text features, i.e., entity and keyword, for the second description text.

[0092] In some embodiments of the present application, the K second text features can include multi-level text features of the second description text. The multi-level text features refer to text features that fuse text meanings and semantic information at different levels, so that the second semantic feature obtained based on the multi-level text features of the second description text can simultaneously retain text features at different levels, for example, can simultaneously retain high-level features for overall picture description in the second description text and can simultaneously retain low-level features for picture detail description in the second description text, so as to improve the information amount of the second semantic feature.

[0093] In the case where the K second text features include multi-level text features of the second description text, the K second text features extracted from the second description text include: encoding the second description text to obtain second encoding features; extracting second text sub-features at multiple levels from the second encoding features, and determining the multi-level text features of the second description text using the second text sub-features at multiple levels.

[0094] The electronic device can encode the second description text by using a bag-of-words model or a word embedding model to map the second description text to a feature space (e.g., a vector space, a matrix space, etc.) to obtain corresponding second encoding features. Then, the electronic device can perform feature extraction on the second encoding features by a text encoder, such as a text encoder of CLIP, and take the outputs of the last several layers of the text encoder as second text sub-features, for example, the outputs of the last 6 layers of the text encoder of CLIP as the second text sub-features, respectively. Then, the electronic device can directly determine the second text sub-features of multiple levels as the multi-level text features of the second description text, or can fuse (e.g., concatenate, superimpose bit by bit, etc.) the obtained second text sub-features of multiple levels and determine the fusion result as the multi-level text features of the second description text.

[0095] Of course, the K second text features can include a single-level text feature of the second description text, where the single-level text feature refers to a text feature that only has single-level text meaning and semantic information. It can be the output of the last layer of the text encoder.

[0096] After obtaining the K second text features, the electronic device can determine the second semantic feature corresponding to the second description text by using the obtained K second text features. Here, the electronic device can fuse the K second text features to obtain the second semantic feature, or can perform dimension conversion on the K second text features respectively (i.e., to make the K second text features capable of being used as subsequent processing), and collectively determine the K second text features after completing the dimension conversion as the second semantic feature.

[0097] At this point, the generation of the second semantic feature is completed.

[0098] In some embodiments of the present application, Figure 3 The determination of the first semantic feature from the first description text describing the video as a whole in step 101 can be implemented by the following processing: extracting K first text features from the first description text; and determining the first semantic feature based on the K first text features. Similarly, K is a positive integer.

[0099] Among them, the N first text features include multi-level text features of the first description text, in which case, the extraction of the K first text features from the first description text includes: encoding the first description text to obtain first encoding features; extracting first text sub-features of multiple levels from the first encoding features, and determining the multi-level text features of the first description text by using the first text sub-features of multiple levels.

[0100] It should be noted that the manner of determining the first semantic feature from the first description text is basically similar to the manner of determining the second semantic feature from the second description text, which will not be repeated here.

[0101] After completing step 101, the electronic device can perform steps 102 to 103 through iteration i, where 1≤i≤N, N is a positive integer, and the specific value of N can be set according to actual needs.

[0102] Step 102, using the i-th input variable block sequence, the first semantic feature and the position sequence, generates the i-th intermediate variable block sequence.

[0103] After the electronic device extracts the first semantic feature and the second semantic feature from the user's input, it will start the injection of the first semantic feature and the second semantic feature in the process of obtaining the latest latent variable sequence. It should be noted that in order to make the details of all regions in the video frame more rich, in the embodiments of the present application, the injection of semantic features is performed in units of latent variable blocks of latent variables, and the injection of semantic features is performed multiple times through iteration i, so that the latent variable blocks, the first semantic feature and the second semantic feature can be fully fused. Here, when i = 1, the i-th input variable block sequence includes the latent variable blocks obtained by blocking the latent variables of the initial latent variable sequence, and when i > 1, the i-th input variable block sequence includes the (i-1)-th output variable block sequence; the position sequence records the position information of each latent variable block. The initial latent variable sequence can be initialized by random noise, for example, obtaining a latent variable sequence (which can be a historical video or a latent variable sequence corresponding to a blank video), and adding random noise to the latent variable sequence to obtain the initial latent variable sequence.

[0104] At each iteration, the electronic device will generate a new variable block sequence based on the input variable block sequence of each iteration combined with the first semantic feature and the position sequence, and the newly generated variable block sequence is recorded as the intermediate variable block sequence of each iteration, which is used to inject the second semantic feature in the subsequent. When i = 1, i.e. in the first iteration, the electronic device will take the sequence of latent variable blocks of each latent variable in the initial latent variable sequence, i.e. the variable block sequence corresponding to the initial latent variable sequence, as the input variable block sequence of the first iteration, when i > 1, i.e. in the second iteration and subsequent iterations, the electronic device can take the output variable block sequence of the (i-1) iteration, i.e. the previous iteration, as the input variable block sequence of this iteration. Of course, the electronic device can also select one of the output variable block sequences generated by the previous i-1 iterations, i.e. all iterations before the current iteration, as the input variable block sequence of the i-th iteration.

[0105] It should be noted that the position information of the latent variable block can include the spatial position of the latent variable block, i.e., the position of the latent variable block in the latent variable, or the time position of the latent variable block, i.e., the appearance time of the latent variable block. Through the position information of the latent variable block, the electronic device can determine the time position and the spatial position of the latent variable block, so as to determine what kind of processing should be performed on the latent variable block to generate a new latent variable block meeting the requirements. Here, the spatial position of the latent variable block can be given by the center coordinates or the coordinates of the upper left corner of the latent variable block, and the time position of the latent variable block can be given according to the video frame identifier of the latent variable to which the latent variable block belongs (here, each latent variable corresponds to a video frame, so each latent variable has a frame identifier) or the appearance time of the video frame corresponding to the latent variable in the video.

[0106] In some embodiments of the present application, the first semantic feature includes a multi-level text feature of the first description text and a single-level text feature of the first description text. In this case, the electronic device can first fuse the multi-level text feature of the first description text and the latent variable block in the input variable block sequence of each iteration, and fuse the multi-level text feature of the first description text and the single-level text feature of the first description text, and then operate the two fusion results and the position sequence through the attention mechanism to obtain a new variable block sequence. The obtained new variable block sequence can be processed by regularization, nonlinear transformation and dimension transformation, and the processing result is taken as the intermediate variable block sequence. The obtained new variable block sequence can also be processed by regularization, nonlinear transformation and dimension transformation, and the fusion result of the new variable block sequence and the single-level semantic feature of the first description text is also processed by regularization, nonlinear transformation and dimension transformation, and finally the two processing results are combined to obtain the intermediate variable block sequence.

[0107] It should be noted that the fusion, attention mechanism, regularization, nonlinear transformation and dimension transformation in the above processing can be implemented through network layers with corresponding functions. Among them, the electronic device can use a LayerNormZero layer to implement the fusion of the multi-level text features of the first description text and the variable blocks in the input variable block sequence, and the fusion of the multi-level text features of the first description text and the single-level text features of the first description text. The electronic device can use an attention layer (Attention Layer) to implement the attention mechanism, and the input of the attention layer is the fusion result of the multi-level text features of the first description text and the variable blocks, the fusion result of the multi-level text features of the first description text and the single-level text features of the first description text, and the position sequence. Among them, the value vector (Value) in the attention mechanism can be calculated from the fusion result of the multi-level text features of the first description text and the variable blocks, and the query vector (Query) and the key vector (Key) can be calculated from the position information of the latent variable blocks in the fusion result of the multi-level text features of the first description text and the single-level text features of the first description text, and the position sequence (the calculation parameters of the query vector (Query) and the key vector (Key) cannot be repeated). The electronic device can use a normalization layer (LayerNorm) to implement the regularization processing, and use a feedforward layer (Feedforward) to implement the nonlinear transformation and dimension transformation processing.

[0108] In some embodiments of the present application, the first semantic feature includes the multi-level text features of the first description text. In this case, the electronic device can first fuse the multi-level text features of the first description text and the variable blocks in the input variable block sequence of each iteration, and can apply the attention mechanism to the new variable block sequence obtained by fusing the multi-level text features of the first description text and the variable blocks in combination with the position sub-sequence. Then, the electronic device can perform regularization, nonlinear transformation and dimension transformation on the new variable block sequence to obtain the intermediate variable block sequence. Alternatively, the electronic device can perform regularization, nonlinear transformation and dimension transformation on the new variable block sequence, and perform regularization, nonlinear transformation and dimension transformation on the fusion result of the new variable block sequence and the multi-level semantic features of the first description text at the same time. Finally, the electronic device can combine the two processing results to obtain the intermediate variable block sequence.

[0109] Step 103, determining the output variable block sequence of the i-th time based on the intermediate variable block sequence of the i-th time, the time interval corresponding to the video segment, the position sequence and the second semantic feature.

[0110] After obtaining the intermediate variable block sequence of each iteration, the electronic device injects the second semantic feature into the intermediate variable block sequence according to the time interval and the position sequence of the video segment, to obtain a new variable block sequence, and records the obtained variable block sequence as the output variable block sequence of each iteration. The obtained output variable block sequence can be used as the input variable block sequence of the subsequent iteration, so as to generate a variable block sequence with more detailed and rich representation through iterative processing.

[0111] In the embodiments of the present application, the video segment refers to a segment that needs to be finely controlled during video generation, and thus the time interval of the video segment refers to an interval formed from the start time of the segment that needs to be finely controlled to the end time of the segment. The time interval can be specified by a user, or can be determined by the electronic device according to the logical relationship of the events described in the first description text and the second description text (for example, when the first description text is "dance performance, the audience sits down at the beginning of the performance, and the audience applauds after the end of the performance", and the second description text is "displaying the close-up of the expression of a certain dancer", at this time, the electronic device can set the time interval of the video segment to 1 / 2 to 2 / 3 of the video). Of course, the time interval can also be randomly set by the electronic device.

[0112] Referring to Figure 4 , Figure 4 is a flowchart of a video generation method provided by the embodiments of the present application Figure 2 In some embodiments of the present application, Figure 3 Step 103 in the embodiments of the present application, that is, determining the output variable block sequence of the i th iteration based on the intermediate variable block sequence of the i th iteration, the time interval corresponding to the video segment, the position sequence and the second semantic feature, can be implemented through the following processing:

[0113] Step 1031, determining a first sub-sequence from the intermediate variable block sequence of the i th iteration according to the time interval and the position sequence, and determining a position sub-sequence corresponding to the first sub-sequence from the position sequence.

[0114] The position sequence records the position information of the latent variable block, and the position information includes the temporal position and the spatial position of the latent variable block. The position sequence does not change at each iteration, and is a sequence generated from the position information of the variable block of the latent variable in the initial latent variable sequence. In this way, in each iteration, the electronic device can read the temporal position of each latent variable block in the intermediate variable block sequence from the position sequence, and extract the first sub-sequence from the intermediate variable block sequence according to the read temporal position and the time interval. After determining the first sub-sequence, the electronic device extracts the position information corresponding to each latent variable block in the first sub-sequence from the position sequence, and the sequence formed by these position information is the position sub-sequence.

[0115] In some embodiments of the present application,Figure 4 The step 1031 in the method 1000, i.e., determining the first sub-sequence from the i-th intermediate variable block sequence according to the time interval and the position sequence, can be implemented by the following processing: determining a control identifier for a latent variable block in the i-th intermediate variable block sequence according to the time position in the time interval and the position sequence, where the control identifier represents whether the latent variable block needs to be regenerated according to the second description text; extracting, from the i-th intermediate variable block sequence, the latent variable block represented by the control identifier as needing to be regenerated according to the second description text, and determining, as the first sub-sequence, a sequence generated by using the extracted latent variable block.

[0116] That is, the electronic device compares the time position read from the position sequence with the time interval, and when the time position of the latent variable block hits the time interval, it indicates that the latent variable to which the latent variable block belongs corresponds to the video frame that needs to be finely controlled, i.e., the latent variable block needs to be regenerated according to the second description text, at this time, the electronic device sets the control identifier of the latent variable block to represent that it needs to be regenerated according to the second description text (for example, sets the control identifier of the latent variable block to 1); when the time position of the latent variable block does not hit the time interval, it indicates that the latent variable to which the latent variable block belongs does not correspond to the video frame that needs to be finely controlled, so the latent variable block does not need to be regenerated according to the second description text, at this time, the electronic device sets the control identifier of the latent variable block to represent that it does not need to be regenerated according to the second description text (for example, sets the control identifier of the latent variable block to 0). Then, the electronic device extracts the latent variable block represented by the control identifier as needing to be regenerated according to the second description text from the intermediate variable block sequence of each iteration, and generates a new sequence according to the order of the extracted latent variable block in the intermediate variable block sequence, and the sequence is the first sub-sequence.

[0117] Of course, after the electronic device extracts the latent variable block that needs to be regenerated from the intermediate variable block sequence, the electronic device can also generate a new sequence according to a random order or an order opposite to the order in the intermediate variable block sequence, and the generated sequence is the first sub-sequence, which is not limited by the embodiments of the present application.

[0118] The step 1032, generating the second sub-sequence based on the second semantic feature, the first sub-sequence, and the position sub-sequence.

[0119] After the electronic device obtains the first sub-sequence that needs to be finely controlled and the position sub-sequence corresponding to the first sub-sequence, the electronic device will combine the second semantic feature and the position sub-sequence to inject the second semantic feature into the latent variable block in the first sub-sequence to regenerate the latent variable block, and the sequence formed by the regenerated latent variable block is the second sub-sequence.

[0120] In some embodiments of the present application, the second semantic feature comprises a multi-level text feature of the second description text and a single-level text feature of the second description text, in which case, Figure 4 The step 1032 in the method 1000, i.e., generating the second subsequence based on the second semantic feature, the first subsequence and the position subsequence, can be implemented by the following processing: fusing the multi-level semantic feature of the second description text into each latent variable block of the first subsequence, and generating a third subsequence using the fused latent variable block; fusing the multi-level text feature of the second description text and the single-level text feature of the second description text to obtain a fused feature; and generating the second subsequence based on the third subsequence and the fused feature.

[0121] That is, the electronic device can first fuse the multi-level text feature of the second description text and the latent variable block in the intermediate variable block sequence of each iteration, and generate a new variable block sequence from the latent variable block fused with the multi-level text feature of the second description text according to a random order or the order of the latent variable block in the intermediate variable block sequence, which is the third subsequence. Then, the electronic device fuses the multi-level text feature of the second description text and the single-level text feature of the second description text to obtain a fused feature of the text, and performs operations on the third subsequence, the fused feature and the position subsequence through an attention mechanism to obtain a new variable block sequence. The new variable block sequence can then be processed by regularization, nonlinear transformation and dimension transformation, and the processing result is taken as the second subsequence; or the new variable block sequence can be processed by regularization, nonlinear transformation and dimension transformation at the same time as the fusion result of the new variable block sequence and the single-level semantic feature of the second description text is processed by regularization, nonlinear transformation and dimension transformation, and finally the two processing results are merged to obtain the second variable block sequence.

[0122] Similar to the generation process of the intermediate variable block sequence of each iteration, the fusion, attention mechanism, regularization, nonlinear transformation and dimension transformation in the above processing can also be implemented by network layers with corresponding functions, for example, the fusion can be implemented by a LayerNormZero layer, the attention mechanism can be implemented by an attention layer, the regularization processing can be implemented by a LayerNorm layer, and the nonlinear transformation and dimension transformation can be implemented by a Feedforward layer.

[0123] In some embodiments of the present application, the second semantic feature comprises a multi-level text feature of the second description text, in which case, the electronic device can first fuse the multi-level text feature of the second description text and the latent variable block in the first sub-sequence of each iteration, and then apply an attention mechanism to the new variable block sequence obtained by fusion in combination with the position sub-sequence, and perform regularization, nonlinear transformation and dimension transformation and the like on the processing result, and take the processing result as the second sub-sequence; or the electronic device can perform regularization, nonlinear transformation and dimension transformation and the like on the new variable block sequence, and perform regularization, nonlinear transformation and dimension transformation and the like on the fusion result of the new variable block sequence and the multi-level semantic feature of the second description text, and finally combine the processing results of the two paths to obtain the second sub-sequence.

[0124] In step 1033, the second sub-sequence and the intermediate variable block sequence of the i-th iteration are used to determine the output variable block sequence of the i-th iteration.

[0125] After obtaining the second sub-sequence, the electronic device can replace the first sub-sequence in the intermediate variable block sequence of each iteration with the second sub-sequence, and directly take the intermediate variable block sequence after replacement as the output variable block sequence of each iteration. Of course, the electronic device can also take some optimization measures for the intermediate variable block sequence after replacement, so that the final output variable block sequence is smoother, so that the picture of the video obtained based on the output variable block sequence is more harmonious, and picture mutation is avoided.

[0126] In some embodiments of the present application, Figure 5 In step 1033, the second sub-sequence and the intermediate variable block sequence of the i-th iteration are used to determine the output variable block sequence of the i-th iteration, which can be implemented by the following processing: replacing the first sub-sequence in the intermediate variable block sequence of the i-th iteration with the second sub-sequence to obtain the i-th to-be-optimized variable block sequence; smoothing the i-th to-be-optimized variable block to obtain the i-th output variable block sequence.

[0127] That is, after the electronic device replaces the first sub-sequence in the intermediate variable block sequence with the second sub-sequence, it will smooth the new variable block sequence obtained, i.e., the to-be-optimized variable block sequence, to eliminate the inharmoniousness between the second sub-sequence and other parts of the intermediate variable block sequence through smoothing, and take the variable block sequence after smoothing as the output variable block sequence of each iteration. In this way, it can be ensured that the fine control latent variable block in the obtained output variable block sequence and other content latent variable blocks can be smoothly transitioned to avoid picture mutation between the fine control video segment and other segments in the entire video, further improving the quality of the generated video.

[0128] It should be noted that the electronic device can read in the sequence of variable blocks to be optimized through a self-attention mechanism, i.e., through a self-attention layer, to achieve smoothing processing through the self-attention mechanism. The electronic device can also achieve smoothing processing through linear interpolation on the sequence of variable blocks to be optimized, which is not limited in the embodiments of the present application.

[0129] Step 104, determining the target video indicated by the first description text and the second description text based on the sequence of output variable blocks and the sequence of positions in the Nth iteration.

[0130] After completing the last iteration, the electronic device will simultaneously combine the sequence of output variable blocks and the sequence of positions in the last iteration for video coding, and the obtained video is the target video matching the first description text and the second description text, i.e., the target video meeting the user's requirements.

[0131] Referring to Figure 5 , Figure 3 is a flowchart of a video generation method provided by the embodiments of the present application Figure 3 In some embodiments of the present application, Figure 6A Step 104 in the above method, i.e., determining the target video indicated by the first description text and the second description text based on the sequence of output variable blocks and the sequence of positions in the Nth iteration, can be achieved through the following processing:

[0132] Step 1041, restoring the sequence of output variable blocks in the Nth iteration to a sequence of output latent variables according to the sequence of positions.

[0133] The electronic device will read the time position and the spatial position of each latent variable block from the sequence of positions, restore the latent variable blocks with the same time position to latent variables according to the spatial positions, sort all restored latent variables according to the order of time positions, and thus obtain a sequence of latent variables. In order to distinguish from the initial sequence of latent variables, the obtained sequence of latent variables is denoted as a sequence of output latent variables.

[0134] Step 1042, decoding the sequence of output latent variables to obtain a sequence of video frames, and determining the target video indicated by the first description text and the second description text based on the sequence of video frames.

[0135] The electronic device decodes each latent variable in the output latent variable sequence, i.e., removes the noise in the latent variable, so as to obtain the video frame corresponding to each latent variable. After completing the decoding of all latent variables in the output latent variable sequence, the electronic device can obtain the video frame sequence. Then, the electronic device can determine the video parameters, such as the frame rate, resolution, height and width of the video, etc., for the video frame sequence, and then convert the video frame sequence into a video in combination with the video parameters, directly determine the obtained video as the target video; or add an audio track, clip a segment, adjust the speed, etc. for the converted video, and take the processed video as the target video.

[0136] It should be noted that the electronic device can complete the above decoding process through a decoder specially used for decoding latent variables. The decoder can implement decoding through a VAE decoder in the DIT, or through a generative adversarial network (GANs), and the embodiments of the present application do not make specific limitations here.

[0137] In some other embodiments of the present application, Figure 1 The step 104 in the above formula, i.e., determining the target video indicated by the first description text and the second description text based on the Nth output variable block sequence and the position sequence, can also be implemented through the following processing: decoding each latent variable block in the Nth output variable block sequence to obtain a picture block sequence; converting the picture block sequence into a video frame sequence in combination with the position sequence, and determining the target video indicated by the first description text and the second description text based on the video frame sequence.

[0138] That is, the electronic device can also directly decode the latent variable blocks in the output variable block sequence of the last iteration, and read the time position and the spatial position of each latent variable block from the position sequence. For the latent variable blocks with the same time position, the electronic device restores them into a video frame according to the spatial position. For the obtained video frame, the electronic device sorts it according to the time position, so as to obtain the video frame sequence, and further obtain the target video.

[0139] It can be understood that, compared with the related art, there is a lack of association between the latent variable subsequence of the video segment and the description text of the complete video, which ultimately leads to a low quality of the generated video. In the embodiments of the present application, in each iteration, the first semantic feature of the first description text describing the video as a whole, the input variable block sequence and the position sequence of each iteration are used to generate an intermediate variable block sequence to realize the injection of the first semantic feature. Then, based on the intermediate variable block sequence, the time interval of the video segment, and the second semantic feature of the second description text describing the video segment, the output variable block sequence of each iteration is generated to realize the injection of the second semantic feature, so that the latent variable block corresponding to the video segment can be associated with the first description text. Through iteration, the semantic features of the first description text and the second description text can be fully injected into the latent variable block, so that the picture of the video segment and the picture of the complete video can be smoothly transitioned, and the quality of the generated video is improved.

[0140] In some embodiments of the present application, before starting the iteration of i, the method can further include the following processing: dividing each latent variable in the initial latent variable sequence into a plurality of latent variable blocks; and generating the input variable block sequence of the first time using the plurality of latent variable blocks of each latent variable.

[0141] Wherein, the electronic device can uniformly block each latent variable in the initial latent variable sequence to obtain a plurality of latent variable blocks, or can non-uniformly block each latent variable to obtain a plurality of latent variable blocks. Then, the electronic device can generate a sequence according to a certain order using the blocking result of each latent variable, i.e. the plurality of latent variable blocks, and the sequence is the input variable block sequence of the first iteration.

[0142] It should be noted that the electronic device can form a subsequence according to the order from left to right and from top to bottom for the latent variable blocks in the same latent variable, and then form the input variable block sequence according to the time order of the video frames corresponding to the latent variables. The electronic device can also arrange all the latent variable blocks in a random order to obtain the input variable block sequence, which is not limited in the embodiments of the present application.

[0143] It should be further noted that, since the output variable block sequence is determined based on the intermediate variable block sequence, the time interval corresponding to the video segment, the position sequence and the second semantic feature, it takes a certain time and computing resources to complete the processing, so this step is performed every iteration, which may affect the efficiency of video generation. In view of this, in the embodiments of the present application, the electronic device can perform the above processing only for some rounds of iteration, and for other rounds of iteration, the intermediate variable block sequence is directly taken as the output variable block sequence.

[0144] That is, in some embodiments of the present application, after the i-th intermediate variable block sequence is generated by using the i-th input variable block sequence, the first semantic feature and the position sequence, when i = M, the method can further include the following processing: determining the M-th intermediate variable block sequence as the M-th output variable block sequence, where M can take any one or more values between 1 and N.

[0145] Thus, when the electronic device is performing the M-th iteration, the processing of determining the M-th output variable block sequence based on the M-th intermediate variable block sequence, the time interval corresponding to the video segment, the position sequence and the second semantic feature can be directly skipped, so that the processing is only performed in part of the N iterations, and thus the time required for video generation and the required computing resources can be reduced.

[0146] Next, the specific application scenarios of the video generation method of the embodiments of the present application are described.

[0147] The video generation method of the embodiments of the present application can be applied to the generation process of a game promotion video, at this time, the first description text can be used to describe the overall style of the game promotion video, and the second description text can be used to describe the display details of the game characters in the game promotion video, so that the electronic device can determine the first semantic feature from the first description text describing the overall style of the game promotion video, and determine the second semantic feature from the second description text describing the display details of the game characters; the following processing is performed by iteration i, 1≤i≤N, N is a positive integer: generate the i-th intermediate variable block sequence by using the i-th input variable block sequence, the first semantic feature and the position sequence, where the latent variable block of the first input variable block sequence includes the latent variable block obtained by blocking the latent variables of the initial latent variable sequence, and the position sequence includes the position information of each latent variable block; determine the i-th output variable block sequence based on the i-th intermediate variable block sequence, the time interval corresponding to the video segment, the position sequence and the second semantic feature; when i is iterated to N, determine the target video indicated by the first description text and the second description text based on the N-th output variable block sequence and the position sequence.

[0148] The video generation method of the embodiment of the present application can also be applied to animation video production. At this time, the first description text describes the plot and appearing characters of the entire animation video, and the second description text describes the plot and appearing characters in the shot. Therefore, the electronic device can determine the first semantic feature from the first description text describing the plot and appearing characters of the entire animation video, and determine the second semantic feature from the second description text describing the plot and appearing characters in the shot. The following processing is performed through iteration i, 1≤i≤N, N being a positive integer: the i-th intermediate variable block sequence is generated by using the i-th input variable block sequence, the first semantic feature, and the position sequence, wherein the latent variable block of the first input variable block sequence includes the latent variable blocks obtained by blocking the latent variables of the initial latent variable sequence, and the position sequence includes the position information of each latent variable block; the i-th output variable block sequence is determined based on the i-th intermediate variable block sequence, the time interval corresponding to the video segment, the position sequence, and the second semantic feature; when i is iterated to N, the target video indicated by the first description text and the second description text is determined based on the N-th output variable block sequence and the position sequence.

[0149] Of course, the video generation method of the embodiment of the present application can not be limited to the application scenarios described above.

[0150] In the following, an exemplary application of the embodiment of the present application in an actual application scenario will be described.

[0151] The embodiment of the present application is realized in the scene of generating a short video (a video with a time length less than a time length threshold) meeting the requirements of a user according to the overall prompt word and the segment prompt word of the user.

[0152] Referring to FIG. 6, Figure 6B is a short video generation schematic provided by the embodiment of the present application Figure 2 When the overall prompt word (the first description text) input by the user is “a video in an animation style depicting a girl sitting by a window, with leaves covering the wall, and the leaves slowly falling”, and the segment prompt word (the second description text) is “the girl's hair is blowing in the wind”, a matching short video 6-1 can be generated accordingly. Figure 7 is a short video generation schematic provided by the embodiment of the present application Figure 7 When the overall prompt word input by the user is “this cartoon-style video shows a scene: four children sitting on the ground, roasting a bonfire at night, and looking up at the stars in the sky”, and the segment prompt word is “the stars gradually fall”, a matching short video 6-2 can be generated accordingly.

[0153] Next, the generation of the short video described above will be described.

[0154] First, referring to Figure 8 , Figure 8Figure 7-1 is a schematic diagram of a video generation system provided by an embodiment of the present application. The video generation system architecture 7-1 includes a prompt word encoding part 7-11 and a video generation part 7-12. The prompt word encoding part 7-11 includes an overall prompt word encoding 7-111 and a segment prompt word encoding 7-112. In the overall prompt word encoding 7-111 and the segment prompt word encoding 7-112, two encoders are respectively arranged, for example, a text encoder 7-2 of CLIP and an encoder 7-3 of T5 model, and a text embedding network layer 7-113 is connected behind each encoder to perform dimension conversion on the output of the encoder. The encoded output after dimension conversion is input to the video generation part 7-12 for processing. The input of the video generation part 7-12 includes a latent variable block sequence (first input variable block sequence) 7-121, a variable block position vector 7-122 corresponding to the latent variable block sequence, and the four encoded outputs mentioned above. The latent variable block sequence 7-121 is formed by a plurality of latent variable blocks 7-14 obtained by splitting each frame in a latent variable sequence (initial latent variable sequence) 7-13 initialized by random noise. In the splitting process, the position of each latent variable block in the original latent variable and the time of the latent variable are recorded, and a vector is generated using these records, which is the variable block position vector 7-122 (position sequence). The video generation part 7-12 can use a transformer 7-4 as the core network, where the number of transformers is N. After a series of transformer calculations, the input can fully integrate the desired video content into each latent variable block in the latent variable block sequence to obtain a new latent variable block sequence 7-15 (Nth output variable block sequence). Then the new latent variable block sequence 7-15 is restored to a new latent variable sequence 7-16 (output latent variable sequence), and the new latent variable sequence is decoded 7-18 to obtain the desired video (target video). Here, the input of the transformer 7-4 also includes label information 7-17 indicating whether fine control is needed for each latent variable block in the latent variable block sequence 7-121.

[0155] As can be seen from the above, in the embodiment of the present application, the overall prompt word and the segment prompt word input by the user are subjected to staged feature extraction by using the overall prompt word encoding and the segment prompt word encoding. This not only enables the semantics of the overall prompt word to be better integrated into the entire video, but also enables the segment description word of the video segment that needs to be controlled to be separately injected, thereby improving the control capability of video generation. In order to better extract semantic features, two different encoders can be used in the embodiment of the present application to complementarily extract semantic features, thereby improving the accuracy of the generated video.

[0156] Figure 9 is a schematic diagram of the overall prompt word encoding and the segment prompt word encoding provided by the embodiment of the present application. Referring to Figure 9 In the overall prompt word encoding 8-1 and the segment prompt word encoding 8-2, the text encoder 8-3 of CLIP and the encoder 8-4 of T5 can be used to encode the overall prompt word at the same time, and the text encoder 8-3 of CLIP and the encoder 8-4 of T5 can be used to encode the segment prompt word at the same time, thereby obtaining 4 encoding results 8-5 (K second text features). Then, the electronic device will use the text embedding layer 8-6 (which can be implemented as a multi-layer perception) to perform dimension alignment on the obtained 4 encoding results respectively (the alignment results are the first semantic features and the second semantic features), so that the dimensions of the 4 encoding results can meet the dimension requirements of the input of the subsequent video generation part.

[0157] It should be noted that the text encoder of CLIP is composed of 12 layers of Transformer, and in a general encoding process, only the output of the last layer of Transformer is needed as the encoding result, while in the embodiment of the present application, the outputs of the last 6 layers of Transformer (the first text sub-features of multiple levels and the second text sub-features of multiple levels) are used together to form the encoding result (the multi-level text features). This operation is to enable the encoding result to have semantic features of different levels, so that the encoding result can simultaneously retain the description of the overall picture at the high level and the description of the details at the low level.

[0158] The T5 encoder is a multi-layer encoder formed by self-attention and feedforward neural networks. In the embodiment of the present application, the T5 encoder is added to increase the understanding of the overall prompt word and the segment prompt word, so that the obtained semantic features can be supplemented to the semantic features encoded by the text encoder of CLIP, to improve the accuracy of the finally generated video.

[0159] Figure 9 is a model architecture diagram of the video generation part provided by the embodiment of the present application. Referring to Figure 10 The video generation part of the embodiment of the present application can be divided into three parts, namely the latent variable block sequence generation 9-11, the video generation model 9-12 and the decoding calculation 9-13.

[0160] In the latent variable block sequence generation 9-11, the latent variable sequence (initial latent variable sequence) 9-111 initialized by random noise is disassembled into a latent variable block sequence (input variable block sequence for the first time) 9-112, and a variable block position vector (position sequence) 9-113 is generated by using the position of the latent variable block, and the latent variable block sequence 9-112 and the variable block position vector 9-113 are used as the input of the video generation model 9-12 for subsequent calculation. FromFigure 1 As can be seen, each frame in the latent variable sequence 9-111 is split into a series of latent variable blocks, and the latent variable blocks are stretched to form a latent variable block sequence 9-112. The variable block position vector 9-113 is provided to the video generation model 9-12, so that the video generation model 9-12 can perceive the time position and spatial position of each latent variable block in the latent variable sequence, and after the video generation model 9-12 completes the calculation, the latent variable block sequence (the Nth output variable block sequence) 9-121 output by the video generation model 9-12 is combined into a new latent variable sequence 9-131 in combination with the variable block position vector 9-113, so as to perform video frame coding 9-132 to obtain a video. It should be noted that the input of the video generation model 9-12 also includes the extracted semantic features 9-2 and the label information 9-3 indicating whether fine control is needed for each latent variable block in the latent variable block sequence 9-112.

[0161] The video generation model 9-12 is composed of a plurality of Transformers, and the number of Transformers is N, and the value of N can be freely set. Figure 10 is a structure diagram of a video generation model provided by an embodiment of the present application Figure 11 Referring to Figure 10 , the input of the video generation model has three parts, namely the latent variable block sequence 10-1, the variable block position vector 10-2, and the semantic features extracted from the overall prompt and the segment prompt. The semantic features include four parts, namely the encoding result of the overall prompt by the text encoder of CLIP 10-3, the encoding result of the overall prompt by the encoder of T5 10-4, the encoding result of the segment prompt by the text encoder of CLIP 10-5, and the encoding result of the segment prompt by T5 10-6. Among them, the video generation model includes a converter 10-7 for semantic injection of the overall video, and a converter 10-8 for semantic injection of the video segment. The structure of the converter 10-7 and the converter 10-8 is similar, and only the converter 10-7 is taken as an example for structural description.

[0162] In the converter 10-7, firstly, the encoding result 10-3 and the latent variable block sequence 10-1 are fused through a zero initial normalization layer 10-71 (for the purpose of fusing the semantics of the latent variable block and the text encoder of CLIP), and the encoding result 10-3 and the encoding result 10-4 are fused through a zero initial normalization layer 10-72 (for the purpose of fusing the semantics of the two encoding results), and after the fusion is completed, the fusion result is input into an attention layer 10-73 for cross-attention calculation, and it is noted that the input of the attention layer 10-73 also includes the variable block position vector 10-2, so that the converter 10-7 can know the time position and the space position of the latent variable block. After the attention calculation is completed, the obtained calculation result 10-74 can be divided into two paths for calculation, one of which is to directly input the calculation result 10-74 into a normalization layer 10-75 and a feedforward layer 10-76 for calculation (during which a short connection is performed), and the other of which is to firstly fuse the calculation result 10-74 and the encoding result 10-4, and then input the fusion result into the normalization layer 10-75 and the feedforward layer 10-76 for calculation (during which a short connection is also performed). Finally, the calculation results of the two paths are merged 10-77, so as to complete the calculation of the converter 10-7, and obtain a new latent variable block sequence 10-9.

[0163] After that, the electronic device determines a sub-sequence 10-11 (first sub-sequence) from the latent variable block sequence 10-9 (intermediate variable block sequence) according to the time range (time interval) 10-10 and the variable block position vector 10-2 which needs to be finely controlled. After that, the electronic device determines a sub-vector 10-12 corresponding to the sub-sequence 10-11 from the variable block position vector 10-2, and takes the sub-sequence 10-11, the sub-vector 10-12, the encoding result 10-5 and the encoding result 10-6 as inputs of the converter 10-8 to perform calculation. After the calculation of the converter 10-8 is completed, a new sub-sequence 10-13 (second sub-sequence) is obtained, and the sub-sequence 10-11 in the latent variable block sequence 10-9 is replaced by the sub-sequence 10-13, and the replaced latent variable block sequence 10-14 (to-be-optimized variable block sequence) is input into the self-attention layer 10-15 so that the sub-sequence 10-13 in the latent variable block sequence 10-14 can be harmonized with other sequences. Finally, the output of the self-attention layer 10-15 is input into the next converter to perform new fusion. In this way, after step-by-step calculation, the finally generated latent variable sequence is completely harmonized, so that the decoded video frame will not have a picture mutation. Similar to the converter 10-7, the converter 10-8 includes a zero initial normalization layer 10-81, a zero initial normalization layer 10-82, an attention layer 10-83, two processing paths of a calculation result 10-84 composed of a normalization layer 10-85 and a feedforward layer 10-86, and a merging 10-87 of the calculation result.

[0164] When determining the sub-sequence from the latent variable block sequence according to the time range and the variable block position vector, the electronic device generates corresponding marking information (control identifier) for each latent variable block in the latent variable block sequence according to the time range and the variable block position vector, wherein 1 indicates that the latent variable block needs to be finely controlled, i.e., regenerated, and 0 indicates that the latent variable block does not need to be finely controlled.

[0165] Figure 10 FIG. 11 is a schematic diagram of generation of marking information provided by an embodiment of the present application. If each latent variable in the latent variable sequence 11-1 can be divided into 16 latent variable blocks, if the 2nd frame to the 4th frame of the video, i.e., the 2nd latent variable to the 4th latent variable, need to be finely controlled, then the three frames should be extracted as a sub-sequence 11-2. At this time, the electronic device sets the marking information corresponding to the latent variable blocks of the three frames in the latent variable block sequence to 1, i.e., sets the 17th point to the 64th point to 1, and sets the marking information corresponding to other latent variable blocks to 0.

[0166] After the calculation of the plurality of Transformers, a semantic latent variable block sequence integrating the overall prompt word and the segment prompt word is obtained, and the electronic device combines the latent variable in the form of a latent variable to obtain a latent variable sequence. Finally, the electronic device inputs each latent variable in the latent variable sequence into a VAE decoder for decoding to obtain a video frame sequence, and further obtain a corresponding video (target video).

[0167] It should be noted that although the above scheme can improve the quality of the video, it also increases the calculation amount and memory of video generation. For example, when the length of the latent variable block sequence is 128 and the length of the sub-sequence that needs to be finely controlled is 64, the calculation amount will be increased by 1 / 2 and the memory occupation will also be increased in the manner of Figure 12 In order to reduce memory consumption and reduce calculation amount, in some embodiments, in the N-layer Transformer, some layers can be selected for segment prompt word processing, that is, the processing in the converter 10-8 in Figure 2 .

[0168] That is, in some embodiments, the processing in the converter 10-7 can be performed only on the latent variable block sequence. Figure 10 is a structural diagram of a video generation model provided by an embodiment of the present application Figure 2 . For the Mth converter in the N converters (M can take one or more values of 1 to N), the output of the converter 10-7 in Figure 3 can be directly processed by the self-attention layer 12-1, and the processing result is input into the next converter for operation, so that the required calculation amount for effectively reducing memory consumption can be reduced.

[0169] It can be understood that in the embodiments of the present application, user information such as data related to the first description text and the second description text is involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards.

[0170] The following continues to illustrate an example structure of the video generation apparatus 255 provided by an embodiment of the present application as a software module. In some embodiments, as shown in ​ , the software module in the video generation apparatus 255 stored in the memory 250 can include:

[0171] a feature determination module 2551 configured to determine a first semantic feature from a first description text describing a video as a whole, and determine a second semantic feature from a second description text describing a video segment;

[0172] The sequence determination module 2552 is configured to generate an i th intermediate variable block sequence by using an i th input variable block sequence, the first semantic feature, and a position sequence; when i = 1, the i th input variable block sequence includes a latent variable block obtained by blocking latent variables in an initial latent variable sequence; when i > 1, the i th input variable block sequence includes an (i-1) th output variable block sequence; and the position sequence includes position information of each latent variable block; and determine an i th output variable block sequence based on the i th intermediate variable block sequence, a time interval corresponding to the video segment, the position sequence, and the second semantic feature.

[0173] The video determination module 2553 is configured to determine a target video indicated by the first description text and the second description text based on the N th output variable block sequence and the position sequence.

[0174] In the foregoing scheme, the sequence determination module 2552 is further configured to determine a first subsequence from the i th intermediate variable block sequence and a position subsequence corresponding to the first subsequence from the position sequence according to the time interval and the position sequence; generate a second subsequence based on the second semantic feature, the first subsequence, and the position subsequence; and determine the i th output variable block sequence by using the second subsequence and the i th intermediate variable block sequence.

[0175] In the foregoing scheme, the second semantic feature includes a multi-level text feature of the second description text and a single-level text feature of the second description text; and the sequence determination module 2552 is further configured to fuse the multi-level semantic feature of the second description text into each latent variable block of the first subsequence, and generate a third subsequence by using the fused latent variable block; fuse the multi-level text feature of the second description text and the single-level text feature of the second description text to obtain a fused feature; and generate the second subsequence based on the third subsequence and the fused feature.

[0176] In the foregoing scheme, the sequence determination module 2552 is further configured to determine a control identifier for the latent variable block in the i th intermediate variable block sequence according to a time position in the position sequence and the time interval, where the control identifier indicates whether the latent variable block needs to be regenerated according to the second description text; extract, from the i th intermediate variable block sequence, the latent variable block indicated by the control identifier as needing to be regenerated according to the second description text, and determine a sequence generated by using the extracted latent variable block as the first subsequence.

[0177] In the above scheme, the sequence determining module 2552 is further configured to replace the first sub-sequence in the i th intermediate variable block sequence with the second sub-sequence to obtain an i th to-be-optimized variable block sequence; and perform smoothing on the i th to-be-optimized variable block to obtain the i th output variable block sequence.

[0178] In the above scheme, the feature determining module 2551 is further configured to extract K second text features from the second description text, where K is a positive integer; and determine the second semantic feature based on the K second text features.

[0179] In the above scheme, the K second text features include multi-level text features of the second description text; and the feature determining module 2551 is further configured to encode the second description text to obtain second encoding features; extract second text sub-features at multiple levels from the second encoding features; and determine the multi-level text features of the second description text by using the second text sub-features at multiple levels.

[0180] In the above scheme, the video generating module 255 further includes a variable blocking module 2554 configured to divide each latent variable in the initial latent variable sequence into a plurality of latent variable blocks; and generate the 1 st input variable block sequence by using the plurality of latent variable blocks of each latent variable.

[0181] In the above scheme, the video determining module 2553 is further configured to restore the N th output variable block sequence to an output latent variable sequence according to the position sequence; decode the output latent variable sequence to obtain a video frame sequence; and determine the target video indicated by the first description text and the second description text based on the video frame sequence.

[0182] In the above scheme, i = M, and the sequence determining module 2552 is further configured to determine the M th intermediate variable block sequence as the M th output variable block sequence, where M can be any one or more values between 1 and N.

[0183] The embodiment of the present application provides a computer program product, which includes a computer program or computer executable instructions stored in a computer readable storage medium. A processor of an electronic device reads the computer executable instructions from the computer readable storage medium, and the processor executes the computer executable instructions, so that the electronic device executes the video generation method in the embodiment of the present application.

[0184] The embodiment of the present application provides a computer readable storage medium, wherein computer executable instructions or computer programs are stored, and when the computer executable instructions or computer programs are executed by a processor, the processor executes the video generation method provided by the embodiment of the present application, for example, as shown in the following. ​ The video generation method is shown.

[0185] In some embodiments, the computer readable storage medium can be RAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM memory, etc.; and can also be various devices including one or any combination of the above storage.

[0186] In some embodiments, the computer executable instructions can be in the form of programs, software, software modules, scripts or codes, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and can be deployed in any form, including being deployed as independent programs or being deployed as modules, components, subroutines or other units suitable for use in a computing environment.

[0187] As an example, the computer executable instructions can but not necessarily correspond to files in a file system, can be stored in a part of a file storing other programs or data, for example, stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program in question, or stored in multiple cooperative files (for example, files storing one or more modules, subroutines or code parts).

[0188] As an example, the computer executable instructions can be deployed to be executed on one electronic device, or executed on multiple electronic devices located in one place, or executed on multiple electronic devices distributed in multiple places and interconnected through a communication network.

[0189] To sum up, by the embodiments of the present application, in each iteration, the first semantic feature of the first description text describing the whole video, the input variable block sequence and the position sequence of each iteration are used to generate the intermediate variable block sequence, so as to realize the injection of the first semantic feature, and then on the basis of the intermediate variable block sequence, the time interval of the video segment and the second semantic feature of the second description text describing the video segment are combined to generate the output variable block sequence of each iteration, so as to realize the injection of the second semantic feature, so that the latent variable block corresponding to the video segment can be associated with the first description text, and the process is repeatedly iterated, so that the semantic features of the first description text and the second description text can be fully injected into the latent variable block, so as to make the picture of the video segment and the picture of the complete video transition naturally, and thus the quality of the generated video is improved. The injection of the second semantic feature can also be performed only in part of the iterations of N iterations, so that the time required for video generation and the required computing resources can be reduced.

[0190] The above merely describes the embodiments of the present application, but is not used to limit the protection scope of the present application. Any modification, equivalent replacement and improvement made within the spirit and scope of the present application shall be included in the protection scope of the present application.

Claims

1. A method of video generation, the method comprising: The method comprises: determining a first semantic feature from a first description text describing a video as a whole, and determining a second semantic feature from a second description text describing a video segment; The following is processed by iteration i, N is a positive integer: Using the i-th input variable block sequence, the first semantic feature, and the position sequence, generate the i-th intermediate variable block sequence; when When, the i-th input variable block sequence includes latent variable blocks obtained by dividing the latent variables of the initial latent variable sequence into blocks, when At that time, the sequence of input variable blocks for the i-th iteration includes the i-th iteration. The output variable block sequence of the sequence includes the position information of each latent variable block; determining an i-th output variable block sequence based on an i-th intermediate variable block sequence, a time interval corresponding to the video segment, the position sequence, and the second semantic feature; determining a target video indicated by the first description text and the second description text based on an N-th output variable block sequence and the position sequence.

2. The method of claim 1, wherein, The determination of the i-th output variable block sequence based on the i-th intermediate variable block sequence, the time interval corresponding to the video segment, the position sequence, and the second semantic feature comprises: determining a first sub-sequence from the i-th intermediate variable block sequence according to the time interval and the position sequence, and determining a position sub-sequence corresponding to the first sub-sequence from the position sequence; generating a second sub-sequence based on the second semantic feature, the first sub-sequence, and the position sub-sequence; determining the i-th output variable block sequence by using the second sub-sequence and the i-th intermediate variable block sequence.

3. The method of claim 2, wherein, The second semantic feature comprises a multi-level text feature of the second description text and a single-level text feature of the second description text. The generation of the second sub-sequence based on the second semantic feature, the first sub-sequence, and the position sub-sequence comprises: fusing the multi-level semantic feature of the second description text into each latent variable block of the first sub-sequence, and generating a third sub-sequence by using the fused latent variable block; fusing the multi-level text feature of the second description text and the single-level text feature of the second description text to obtain a fused feature; generating the second sub-sequence based on the third sub-sequence and the fused feature.

4. The method of claim 2, wherein, The determination of the first sub-sequence from the i-th intermediate variable block sequence according to the time interval and the position sequence comprises: determining a control identifier for the latent variable block in the i-th intermediate variable block sequence according to a time position in the time interval and the position sequence, wherein the control identifier represents whether the latent variable block needs to be regenerated according to the second description text; extracting the latent variable block represented by the control identifier as needing to be regenerated according to the second description text from the i-th intermediate variable block sequence, and determining a sequence generated by using the extracted latent variable block as the first sub-sequence.

5. The method of claim 2, wherein, The determination of the i-th output variable block sequence by using the second sub-sequence and the i-th intermediate variable block sequence comprises: replacing the first sub-sequence in the i-th intermediate variable block sequence by using the second sub-sequence to obtain an i-th to-be-optimized variable block sequence; smoothing the i-th to-be-optimized variable block to obtain the i-th output variable block sequence.

6. The method according to any one of claims 1 to 5, characterized in that, The determination of the second semantic feature from the second description text describing the video segment comprises: extracting K second text features from the second description text, wherein K is a positive integer. determine the second semantic feature based on the K second text features.

7. The method of claim 6, wherein, The K second text features comprise multi-level text features of the second description text, and the K second text features are extracted from the second description text. encode the second description text to obtain second encoded features; extract a plurality of levels of second text sub-features from the second encoded features, and determine the multi-level text features of the second description text by using the plurality of levels of second text sub-features.

8. The method according to any one of claims 1 to 5, characterized in that, Before starting the iteration of i, the method further comprises: divide each latent variable in the initial latent variable sequence into a plurality of latent variable blocks; generate the first input variable block sequence by using the plurality of latent variable blocks of each latent variable.

9. The method according to any one of claims 1 to 5, characterized in that, The determining of the target video indicated by the first description text and the second description text based on the Nth output variable block sequence and the position sequence comprises: restore the Nth output variable block sequence to an output latent variable sequence according to the position sequence; decode the output latent variable sequence to obtain a video frame sequence, and determine the target video indicated by the first description text and the second description text based on the video frame sequence.

10. The method according to any one of claims 1 to 5, characterized in that, after generating the i-th intermediate variable block sequence using the i-th input variable block sequence, the first semantic feature, and the position sequence, the method further comprises: The Mth intermediate variable block sequence is determined as the Mth output variable block sequence, where M can take any one or more values between 1 and N.

11. A video generating apparatus characterized by comprising: The apparatus comprises: a feature determination module configured to determine a first semantic feature from a first description text describing a video as a whole, and determine a second semantic feature from a second description text describing a video segment; The sequence determination module is used to perform the following processing through iteration i. N is a positive integer: using the i-th input variable block sequence, the first semantic feature, and the position sequence, generate the i-th intermediate variable block sequence; when When, the i-th input variable block sequence includes latent variable blocks obtained by dividing the latent variables of the initial latent variable sequence into blocks, when At that time, the sequence of input variable blocks for the i-th iteration includes the i-th iteration. The output variable block sequence of the i-th time, wherein the position sequence includes the position information of each latent variable block; the output variable block sequence of the i-th time is determined based on the intermediate variable block sequence of the i-th time, the time interval corresponding to the video segment, the position sequence and the second semantic feature; a video determination module configured to determine a target video indicated by the first description text and the second description text based on an Nth output variable block sequence and a position sequence.

12. The apparatus of claim 11, wherein The sequence determination module is further configured to determine a first sub-sequence from the i-th intermediate variable block sequence according to the time interval and the position sequence, and determine a position sub-sequence corresponding to the first sub-sequence from the position sequence; generate a second sub-sequence based on the second semantic feature, the first sub-sequence, and the position sub-sequence, and determine an i-th output variable block sequence by using the second sub-sequence and the i-th intermediate variable block sequence.

13. The apparatus of claim 12, wherein, The second semantic feature comprises multi-level text features of the second description text and single-level text features of the second description text. The sequence determination module is further configured to fuse the multi-level semantic features of the second description text into each latent variable block of the first sub-sequence, generate a third sub-sequence by using the fused latent variable block, fuse the multi-level text features of the second description text and the single-level text features of the second description text to obtain a fused feature, and generate the second sub-sequence based on the third sub-sequence and the fused feature.

14. The apparatus of claim 11, wherein The video determination module is further configured to restore the Nth output variable block sequence to an output latent variable sequence according to the position sequence; decode the output latent variable sequence to obtain a video frame sequence, and determine the target video indicated by the first description text and the second description text based on the video frame sequence.

15. An electronic device, comprising: The electronic device comprises: a memory configured to store computer executable instructions or computer programs; a processor configured to execute the computer executable instructions or computer programs stored in the memory to implement the method of any one of claims 1 to 10.

16. A computer-readable storage medium storing computer-executable instructions or a computer program, wherein the computer-executable instructions or the computer program comprise the steps of: The computer executable instructions or computer programs are executed by the processor to implement the method of any one of claims 1 to 10.

17. A computer program product comprising computer-executable instructions or a computer program, characterized in that, The computer executable instructions or computer programs are executed by the processor to implement the method of any one of claims 1 to 10. The computer executable instructions or computer programs are executed by the processor to implement the method of any one of claims 1 to 10.

Citation Information

Patent Citations

  • Video description information generation method, video processing method and corresponding devices

    CN109960747A

  • Video generation method and system, electronic equipment and storage medium

    CN118714417A