Video generation method, device, equipment and storage medium
Through multiple rounds of image processing and dynamic adjustment of inter-frame offset sequences, the problem of unchanged reference information of image generation models in video generation is solved, and the quality and diversity of video generation are improved.
Patent Information
- Application Number
- CN202410743937.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-07
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2044-06-07
AI Technical Summary
In existing video generation methods, the image generation model uses unchanged reference information when generating video frame sequences, which makes the generation direction easily inaccurate and reduces the video quality.
An image vector sequence is generated by obtaining the reference image code and performing multiple rounds of image processing. After each round of processing, the image vector is dynamically adjusted based on the inter-frame offset sequence. The dynamically adjusted image vector sequence is input into the image generation model for the next round of processing until a video frame sequence is generated.
The generation performance of the image generation model has been improved, ensuring that the video quality of the video frame sequence is more diverse and avoiding the content confinement caused by the unchanging reference information.
Smart Images

Figure CN118784938B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of Internet technology, specifically to the field of artificial intelligence technology, and in particular to a video generation method, apparatus, device and storage medium. Background Art
[0002] With the development of artificial intelligence technology, an image generation model has been proposed. The image generation model is a model that can generate a video frame sequence based on an image. At present, when using an image generation model to generate a video frame sequence, an image vector of a reference image is usually obtained, and the image vector is directly input into the image generation model to guide the corresponding model to generate a video frame sequence. It can be seen that the existing video generation method will cause the information referenced by the model (such as image vectors) to remain unchanged throughout the video generation process. This can easily cause the model to lose its generation direction in the process of generating a video frame sequence, thereby reducing the generation performance of the entire model, resulting in the inability to accurately achieve the purpose of generating a video frame sequence, and reducing the video quality of the final generated video frame sequence. Summary of the Invention
[0003] The embodiments of the present application provide a video generation method, apparatus, device, and storage medium, which can improve the generation performance of an image generation model, thereby improving the video quality of a video frame sequence generated by the image generation model.
[0004] In one aspect, an embodiment of the present application provides a video generation method, the method comprising:
[0005] Acquire a reference image, where the reference image is used to indicate video content in a video frame sequence to be generated, where the video frame sequence includes M video frames, where M is an integer greater than 1;
[0006] Generate an image vector sequence based on the reference image encoding, the image vector sequence including M image vectors, each of which is used to represent the reference image; the m-th image vector in the image vector sequence is used to guide the generation of the m-th video frame in the video frame sequence, m∈[1,M];
[0007] Performing multiple rounds of image processing based on the image vector sequence and the image generation model, each round of image processing generating M frames of image data, the m-th frame of image data corresponding to the m-th video frame; wherein, during the multiple rounds of image processing, each time M frames of image data are generated after a round of image processing, obtaining an inter-frame offset sequence corresponding to the current M frames of image data, dynamically adjusting at least one image vector in the image vector sequence based on the inter-frame offset sequence, and inputting the dynamically adjusted image vector sequence into the image generation model for the next round of image processing; the inter-frame offset sequence is used to indicate: a difference between any two adjacent frames of image data in the corresponding M frames of image data;
[0008] The video frame sequence is generated according to the M frames of image data generated by the last round of image processing in the multiple rounds of image processing.
[0009] In another aspect, an embodiment of the present application provides a video generation device, comprising:
[0010] an acquisition unit, configured to acquire a reference image, wherein the reference image is used to indicate video content in a video frame sequence to be generated, wherein the video frame sequence includes M video frames, where M is an integer greater than 1;
[0011] a processing unit, configured to generate an image vector sequence based on the reference image encoding, the image vector sequence comprising M image vectors, each of which is used to represent the reference image; the m-th image vector in the image vector sequence is used to guide the generation of the m-th video frame in the video frame sequence, m∈[1,M];
[0012] The processing unit is further configured to perform multiple rounds of image processing based on the image vector sequence and the image generation model, each round of image processing generating M frames of image data, the m-th frame of image data corresponding to the m-th video frame; wherein, during the multiple rounds of image processing, each time M frames of image data are generated after a round of image processing, an inter-frame offset sequence corresponding to the current M frames of image data is obtained, and at least one image vector in the image vector sequence is dynamically adjusted based on the inter-frame offset sequence, and the dynamically adjusted image vector sequence is input into the image generation model for the next round of image processing; the inter-frame offset sequence is used to indicate: a difference between any two adjacent frames of image data in the corresponding M frames of image data;
[0013] The processing unit is further configured to generate the video frame sequence based on the M frames of image data generated by the last round of image processing in the multiple rounds of image processing.
[0014] In another aspect, an embodiment of the present application provides a computer device, the computer device including an input interface and an output interface, and the computer device further including:
[0015] processors and computer storage media;
[0016] The processor is suitable for implementing one or more instructions, the computer storage medium stores one or more instructions, and the one or more instructions are suitable for being loaded by the processor and executing the above-mentioned video generation method.
[0017] On the other hand, an embodiment of the present application provides a computer storage medium, which stores one or more instructions, and the one or more instructions are suitable for being loaded by a processor and executing the above-mentioned video generation method.
[0018] On the other hand, an embodiment of the present application provides a computer program product, which includes one or more instructions; when the one or more instructions in the computer program product are executed by a processor, the above-mentioned video generation method is implemented.
[0019] In an embodiment of the present application, an image vector sequence can be generated by encoding a reference image for indicating the video content in a video frame sequence to be generated, and multiple rounds of image processing are performed based on the image vector sequence and an image generation model, thereby generating a video frame sequence based on M frames of image data generated by the last round of image processing. Each time M frames of image data are generated through a round of image processing, an inter-frame offset sequence can be obtained to indicate the difference between any two adjacent frames of image data in the currently generated M frames of image data, and at least one image vector in the image vector sequence can be dynamically adjusted based on the inter-frame offset sequence. The dynamically adjusted image vector sequence is then input into the image generation model for the next round of image processing. This allows the image generation model to understand the current image generation situation based on the dynamically adjusted image vector sequence during the next round of image processing, thereby deeply perceiving the generation direction of the video frame sequence, improving the generation performance of the entire model, and accurately achieving the purpose of generating a video frame sequence. Moreover, by dynamically adjusting the image vector sequence, it is also possible to avoid the situation where the information referred to by the image generation model remains unchanged during the entire video generation process, thereby avoiding the image generation model being restricted in the content of the final generated video frame sequence due to the unchanged reference information. While maintaining the video content information indicated by the reference image, the image generation model can exert its own diverse generation capabilities, making the video content of the final generated video frame sequence more colorful, thereby improving the video quality of the video frame sequence. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] Figure 1 This is a flow chart of a video generation solution provided in an embodiment of the present application;
[0022] Figure 2 This is a flow chart of a video generation method provided in an embodiment of the present application;
[0023] Figure 3a This is a schematic diagram of a process for constructing an image vector sequence provided in an embodiment of the present application;
[0024] Figure 3b This is a schematic diagram of another process for constructing an image vector sequence provided in an embodiment of the present application;
[0025] Figure 3c This is a structural diagram of an image generation model provided in an embodiment of the present application;
[0026] Figure 3d This is a schematic diagram of the structure of a cross-attention layer provided in an embodiment of the present application;
[0027] Figure 4 is a flowchart of a video generation method provided by another embodiment of the present application;
[0028] Figure 5a is a schematic diagram of a video frame sequence provided in an embodiment of the present application;
[0029] Figure 5b Schematic diagram of the calculation principle of an inter-frame offset matrix provided in an embodiment of the present application;
[0030] Figure 5c is a schematic diagram of dynamically adjusting an image vector provided by an embodiment of the present application;
[0031] Figure 5d is a schematic diagram of adjusting an intermediate image vector based on a cross-attention network provided in an embodiment of the present application;
[0032] Figure 5e 1 is a schematic diagram of the principle of a method for generating a video sequence based on an inter-frame offset calculation mechanism provided by an embodiment of the present application;
[0033] Figure 5f 1 is a schematic diagram of a calculation flow of an inter-frame offset calculation mechanism provided in an embodiment of the present application;
[0034] Figure 6 is a structural diagram of a video generation device provided in an embodiment of the present application;
[0035] Figure 7 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0036] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.
[0037] The embodiment of the present application is based on AIGC (Artificial Intelligence Generated Content, generative artificial intelligence) technology, and proposes a video generation solution to improve the generation performance of the image generation model, thereby improving the video quality of the video frame sequence generated by the model. Among them, the core idea of AIGC technology is to use artificial intelligence (AI) technology to generate content (such as images, video frame sequences) with certain creativity and quality. AI technology refers to the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science; it mainly produces a new intelligent machine that can respond in a similar way to human intelligence by understanding the essence of intelligence, so that the intelligent machine has multiple functions such as perception, reasoning and decision-making.
[0038] Specifically, AI technologies can be broadly divided into several areas, including natural language processing (NLP) and machine learning (ML) / deep learning. Natural language processing (NLP) is a key area within both computer science and artificial intelligence. It studies theories and methods that enable effective communication between humans and computers using natural language. Natural language processing involves natural language, the language we use daily, and is closely related to linguistics. It also involves computer science and mathematics. Specifically, natural language processing technologies typically include text processing, semantic understanding, machine translation, robotic question answering, and knowledge graphs. Machine learning is a multidisciplinary field, encompassing probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of AI and is the fundamental path to computer intelligence. Its applications span all areas of artificial intelligence. Deep learning, on the other hand, is a machine learning technique that utilizes deep neural network systems.
[0039] Specifically, the general principle of the video generation solution proposed by the embodiment of the present application based on the AIGC technology is as follows: the sequence length of the video frame sequence to be generated can be obtained, and the value of the sequence length is equal to the number M of video frames in the video frame sequence, that is, the sequence length is used to indicate that the number of video frames in the video frame sequence is M, and M is an integer greater than 1. In addition, a reference image (ref image) can be obtained, and an image vector sequence is constructed based on the value of the sequence length and the reference image; wherein the reference image is used to indicate the video content in the video frame sequence to be generated, and the constructed image vector sequence includes M image vectors, each image vector can be used to represent the reference image, and the mth image vector in the image vector sequence is used to guide the generation of the mth video frame in the video frame sequence, m∈[1,M]. Furthermore, an image generation model can be obtained, and multiple rounds of image processing can be performed based on the image vector sequence and the image generation model, so as to generate a video frame sequence based on the M frames of image data generated by the last round of image processing. It can be seen that in the actual use of the video generation solution proposed in the embodiment of the present application, as long as a reference image and the sequence length of the video frame sequence to be generated are given, an image vector sequence can be constructed based on the reference image and the sequence length, and then multiple rounds of image processing are performed based on the image vector sequence and the image generation model to generate a video frame sequence. The entire system does not need to input additional image information or other conditions to complete the task of generating a video frame sequence. This can reduce the difficulty of preparing reference information in advance when generating a video frame sequence (no need to provide multiple reference images).
[0040] Furthermore, the principle of performing multiple rounds of image processing based on an image vector sequence and an image generation model can be roughly as follows: the image vector sequence is input into the image generation model for a first round of image processing to generate M frames of image data. Based on the currently generated M frames of image data (i.e., the M frames of image data generated by the first round of image processing), an inter-frame offset sequence is calculated. This inter-frame offset sequence is used to indicate the difference between any two adjacent frames of image data in the corresponding M frames of image data, and the image vector sequence is dynamically adjusted based on this inter-frame offset sequence. The dynamically adjusted image vector sequence is input into the image generation model for a second round of image processing to generate M frames of image data. Based on the currently generated M frames of image data (i.e., the M frames of image data generated by the second round of image processing), an inter-frame offset sequence is calculated, and the image vector sequence is dynamically adjusted based on this inter-frame offset sequence. The dynamically adjusted image vector sequence is input into the image generation model for a third round of image processing to generate M frames of image data, and so on, until the final round of image processing is performed to generate M frames of image data.
[0041] It can be seen that in the process of performing multiple rounds of image processing, the video generation scheme proposed in the embodiment of the present application generates M frames of image data after each round of image processing. The image vector sequence can be dynamically adjusted based on the inter-frame offset sequence corresponding to the currently generated M frames of image data, and the dynamically adjusted image vector sequence is input into the image generation model for the next round of image processing. This not only enables the image generation model to understand the current image generation situation based on the dynamically adjusted image vector sequence in the next round of image processing, thereby deeply perceiving the generation direction of the video frame sequence, improving the generation performance of the entire model, and accurately achieving the purpose of generating the video frame sequence; it can also avoid the image generation model from being confined to the content of the finally generated video frame sequence due to the unchanging reference information, so that the image generation model can exert its own diverse generation capabilities while maintaining the video content indicated by the reference image, so that the generated video frame sequence is not bound by the limited content of the reference image, and its video content can be more colorful, thereby improving the video quality of the video frame sequence.
[0042] In a specific implementation, the above-mentioned video generation scheme can be executed by a computer device, which can be a terminal or a server, that is, the video generation scheme can be executed by the terminal or the server alone. Optionally, the video generation scheme can also be executed jointly by the terminal and the server. For example, the terminal can be responsible for obtaining the sequence length of the video frame sequence to be generated and obtaining the reference image, and constructing an image vector sequence based on the value of the sequence length and the reference image, thereby sending the image vector sequence to the server, and the server is responsible for obtaining the image generation model, and performing multiple rounds of image processing based on the image vector sequence and the image generation model, thereby generating a video frame sequence based on the M frames of image data generated by the last round of image processing, such as Figure 1 Alternatively, the server may return the M frames of image data generated by the last round of image processing to the terminal, and the terminal is responsible for generating a video frame sequence based on the M frames of image data generated by the last round of image processing.
[0043] The aforementioned terminals may include smartphones, computers (such as tablets, laptops, and desktop computers), smart wearable devices (such as smartwatches and smart glasses), intelligent voice interaction devices, smart home appliances (such as smart TVs), in-vehicle terminals, or aircraft. The aforementioned servers may include standalone physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDNs (Content Delivery Networks), and big data and artificial intelligence platforms. Furthermore, the terminals and servers may be located within or outside the blockchain network, without limitation. Furthermore, the terminals and servers may upload any internally stored data (such as a new image embedded with a watermark) to the blockchain network for storage, thereby preventing tampering with the internally stored data and enhancing data security.
[0044] Based on the above description, it can be seen that the video generation solution proposed in the embodiment of the present application is mainly to improve the generation performance of the image generation model in the scenario of generating video frame sequences based on images, so that the video content of the video frame sequences generated by the image generation model can be richer and more diverse. Therefore, the video generation solution proposed in the embodiment of the present application can be applied to the implementation direction of various video frame sequence generation, providing more efficient generation capabilities for various video frame sequence generation applications. For example:
[0045] (1) The video generation solution can be applied in the promotion of games or TV series. Designers can carefully design a rich poster image for the characters and scenes in the game or TV series to be promoted, and then input the poster image as a reference image into the video generation solution to generate a complete video frame sequence. The main content of each video frame in the generated video frame sequence is consistent with the content of the poster image. At the same time, it can also include more rich content and image elements generated based on the current poster image. Then, such a video frame sequence can be used to help promote the promotion activities, thereby improving the efficiency and fun of the entire promotion.
[0046] (2) This video generation solution can be used as a creative aid in the game production process. Prepare a high-quality game character image as a reference image in advance, and then use this video generation solution to generate various video frame sequences containing multiple game character images in the game based on the reference image. These video frame sequences can contain content that is highly consistent with the desired high-quality game character image, and can also contain various other exciting and rich content (such as changes in actions, special effects backgrounds, etc.). By passing these video frame sequences to the game production team, it is possible to provide the game production team with more creative materials, thereby improving the efficiency of game production.
[0047] Based on the above description of the video generation solution, the present embodiment proposes a video generation method. In the present embodiment, the video generation method is mainly described by taking a computer device executing the video generation method as an example. Figure 2 As shown, the video generation method may include the following steps S201-S204:
[0048] S201, obtaining a reference image.
[0049] In a specific implementation, the computer device may obtain an image input or uploaded by the user as a reference image; alternatively, the computer device may obtain an image from an open source image collection or database as a reference image. The embodiments of this application do not limit the method of obtaining the reference image. It is worth emphasizing that in the embodiments of this application, when data related to user information (such as images input or uploaded by the user, etc.) is involved, when any of the method embodiments proposed in the embodiments of this application is applied to a specific product or technology, these relevant data are collected with the user's permission or consent, and the collection, use and processing of the relevant data comply with the relevant laws, regulations and standards of the relevant region.
[0050] Among them, the reference image is an image to be input into the image generation model to indicate the video content in the video frame sequence to be generated. The image generation model can use the reference image to guide subsequent prediction generation. It can be seen that the role of the reference image is to inform the image generation model what the main video content in the current video frame sequence to be generated is. It is equivalent to a reference information for the image generation model to know the specific video content in the final generated video frame sequence, so as to ensure that the video content in the entire video frame sequence ultimately generated by the image generation model is consistent with the image content of the input reference image, greatly improving the restoration of the desired video content. At the same time, the control of the reference image can ensure that there will be no content distortion in the video frame sequence ultimately generated by the image generation model.
[0051] For example, if a picture containing a game character image is obtained as a reference image, the reference image can be used to indicate that the video content in the video frame sequence to be generated includes the game character image, that is, to inform the model that the main content of the video frame sequence to be generated is the game character image, and the generated video frame sequence can be a video formed by performing various action changes, as well as changes in the environment and camera movement for the game character image. For another example, if a picture containing an actual person is obtained as a reference image, the reference image can be used to indicate that the video content in the video frame sequence to be generated includes the actual person, that is, to inform the model that the main content of the video frame sequence to be generated is the actual person, and the generated video frame sequence can be a video formed by performing various action changes and expression changes for the actual person, and so on.
[0052] It should be noted that the above is only an illustrative list of the specific contents of the reference image, which is not exhaustive and does not limit the specific contents of the reference image. The specific contents can be set according to actual business needs. In addition, the video frame sequence to be generated mentioned above can be simply referred to as video, which can include M video frames and M video frames are arranged in sequence, and M is an integer greater than 1; in actual applications, the computer device can obtain the sequence length of the video frame sequence to be generated, and thus determine the value of M by the value of the sequence length. For example, if the sequence length is 5, it can be determined that the value of M is 5, that is, it is determined that the video frame sequence to be generated includes 5 video frames. It can be understood that a video frame is equivalent to an image; based on this, a video frame can also be called an image frame.
[0053] S202, generating an image vector sequence based on the reference image encoding.
[0054] The image vector sequence may include M image vectors, each of which may be used to represent a reference image. The mth image vector in the image vector sequence is used to guide the generation of the mth video frame in the video frame sequence, where m∈[1,M]; that is, the mth image vector in the image vector sequence is used to indicate the video content of the mth video frame in the video frame sequence.
[0055] In the specific implementation of step S202, the computer device may perform M image encoding on the reference image to obtain M image vectors, where each image encoding generates one image vector, and then the M image vectors are arranged to obtain an image vector sequence, such as Figure 3a Alternatively, the computer device may perform image encoding on the reference image to obtain an image vector for representing the reference image, and replicate the image vector multiple times to obtain M image vectors, thereby arranging the M image vectors to obtain an image vector sequence, as shown in FIG. Figure 3bAs shown in . Since the copy operation of vector is simpler and requires less processing resources than image encoding, it is used Figure 3b The method shown in FIG. 4 is used to generate an image vector sequence, which can reduce the number of times image encoding is performed on a reference image. This can save processing resources required for image encoding and improve the generation efficiency of the image vector sequence.
[0056] Furthermore, the above-mentioned image encoding can be implemented by calling any image encoder. For example, a computer device can call the image encoder in the Clip model to perform image encoding on a reference image. The Clip model is an advanced deep learning model that can understand text and images and is a multimodal model. It has been trained to match text and images and has learned to recognize the content in the image and the language describing the image. The main goal of the Clip model is to learn to match images and text through contrastive learning. The Clip model can be trained to predict which image belongs to a given text, and vice versa. During the model training process, the Clip model can learn to encode images and text into a unified vector space, which enables it to understand the relationship between them in language and vision. In this way, the Clip model can recognize elements such as objects, scenes, actions, etc. in the image, and can also understand text related to the image, such as labels, descriptions, titles, etc. It is worth emphasizing that the Clip model has been shown to have excellent performance in visual and language tasks.
[0057] S203, performing multiple rounds of image processing based on the image vector sequence and the image generation model.
[0058] Each round of image processing generates M frames of image data, the mth frame of image data corresponds to the mth video frame, and each frame of image data includes image values of multiple position points. During multiple rounds of image processing, the computer device can input the image vector sequence into the image generation model to perform the first round of image processing to generate M frames of image data. After each round of image processing to generate M frames of image data, the computer device can obtain the inter-frame offset sequence corresponding to the current M frames of image data, and dynamically adjust at least one image vector in the image vector sequence based on the inter-frame offset sequence, and input the dynamically adjusted image vector sequence into the image generation model for the next round of image processing; wherein the inter-frame offset sequence is used to indicate the difference between any two adjacent frames of image data in the corresponding M frames of image data.
[0059] In a specific implementation of step S203, the computer device may perform multiple rounds of image processing based solely on the image vector sequence and the image generation model. In this case, any round of image processing is performed by the image generation model based on the currently input image vector sequence, i.e., the information referenced by the image generation model during any round of image processing only includes the currently input image vector sequence. Alternatively, the computer device may obtain content description text, which is used to indicate the display form of the video content in the video frame sequence; generate a text vector sequence based on the content description text, wherein the text vector sequence includes M text vectors, each text vector being used to represent the content description text, thereby performing multiple rounds of image processing based on the text vector sequence, the image vector sequence, and the image generation model. In this case, any round of image processing is performed by the image generation model based on the text vector sequence and the currently input image vector sequence, i.e., the information referenced by the image generation model during any round of image processing includes not only the currently input image vector sequence, but also the text vector sequence. This enriches the information referenced by the image generation model, thereby facilitating the image generation model to fully utilize its diverse generation capabilities, making the video content of the ultimately generated video frame sequence more diverse and richer, thereby improving the video quality of the video frame sequence.
[0060] It should be noted that the embodiments of the present application do not limit the specific selection of the image generation model.
[0061] For example, the image generation model mentioned above can be a model based on the Transformer architecture (a model structure based on the attention mechanism for data processing), which can be composed of multiple networks stacked together, with the output of the previous network serving as the input of the next network, so as to iteratively process the input data (such as an image vector sequence) through multiple networks, thereby generating and outputting multiple video frames through the last network. In this case, the one frame of image data mentioned above is one video frame; that is, in this case, the M frames of image data generated by any round of image processing based on the image generation model are M video frames, and the position points in each frame of image data are pixels, and the image values are the pixel values of the pixels.
[0062] For another example, the image generation model mentioned above can be a diffusion model. A diffusion model can include at least a denoising network, whose purpose is to eliminate the continuous application of Gaussian noise to the image. It can be viewed as a series of denoising autoencoders. Diffusion models, such as the Denoising Diffusion Probabilistic Model (DDPM), define two processes: forward diffusion and reverse denoising. In the forward diffusion process, noise (such as Gaussian noise) is continuously added to a given image over multiple time steps until it becomes pure noise. In the reverse denoising process, a noisy image is sampled, and the denoising network continuously performs denoising operations over multiple time steps to generate a feature map, thereby generating a high-quality image based on the feature map. In this case, the aforementioned frame of image data is a frame of feature map. That is, in this case, the M frames of image data generated by any round of image processing based on the image generation model are M frames of feature maps, with the locations in each frame of image data being feature points, and the image values being the feature values of the feature points.
[0063] For example, the image generation model mentioned above can be a Stable Diffusion model; the Stable Diffusion model is a variant of the diffusion model, which can be called a "latent diffusion model (LDM)"; DDPM performs various operations on the original image, while LDM introduces a VAE (Variational AutoEncoder) model based on the diffusion model. That is, the Stable Diffusion model can at least include a VAE and a denoising network. Optionally, the Stable Diffusion model can also include a text encoder so that the Stable Diffusion model can support not only image input with the corresponding image vector as a reference, but also text input with the corresponding text vector as a reference. The text encoder can be used to embed the input text into a latent text embedding space that can be understood by the denoising network to generate a vector for the corresponding text; this text encoder can use the Text encoder module in the Clip model, which is a simple transformer-based encoder that maps a token sequence to a latent text embedding sequence, so that good text cues can be used to obtain better expected output.
[0064] The VAE mentioned above can be composed of an encoder and a decoder. The former (i.e., the encoder) is used to convert the input image into a latent variable in a low-dimensional latent space during the model training process. The latent variable can be used as the input of the next component (i.e., the denoising network) for denoising. The latter (i.e., the decoder) will do the opposite, which can be used to convert the denoising result output by the denoising network into an image during the model training process or the model inference process. In this case, the denoising result output by the denoising network is the denoised latent variable, and the denoising network predicts the noise in the input latent variable based on reference information (such as the vector of the input image, the vector of the input text, etc.), and subtracts the predicted noise from the input latent variable to achieve denoising of the input latent variable, thereby outputting the denoised latent variable.
[0065] Among them, the latent variable is essentially a feature map, which refers to the variables (or factors) that are not directly observed in the model, that is, the most basic features in the entire model processing flow. Since in the Stable Diffusion model, the latent variables represent unobserved factors, such as individual characteristics, behaviors, style, and composition, etc., these factors may affect the forward diffusion process, so the latent variables are very important for explaining the data observed by the model and the behavior of the model. Based on this, when the image generation model includes the Stable Diffusion model, the one-frame image data mentioned above is a frame of latent vector (i.e., feature map); that is, in this case, the M frames of image data generated by any round of image processing based on the image generation model are M frames of latent vectors (i.e., M frames of feature maps), and the position points in each frame of image data are feature points, and the image values are the feature values of the feature points.
[0066] Based on the above, in the model training process of the Stable Diffusion model, the VAE encoder can be used to encode the input image to obtain the latent variables of the corresponding image in the low-dimensional latent space. In the model inference process, the VAE decoder can be used to convert the latent variables into images, such as Figure 3cAs shown. Specifically, the model training process of the Stable Diffusion model is as follows: the sample image is encoded through the VAE encoder to obtain the latent variable of the sample image. In the forward diffusion process, noise (such as Gaussian noise) is continuously added to the latent variable of the given sample image through multiple time steps to obtain the final denoising result. In the reverse denoising process, the denoising network is called to perform noise prediction on the denoised result through multiple time steps, and the loss value is calculated based on the difference between the predicted noise and the noise added in the forward diffusion process, thereby optimizing the model parameters based on the loss value. Correspondingly, the model inference process of the Stable Diffusion model is as follows: the latent vector obtained by encoding the noisy image is obtained, and the denoising network is called to predict the noise in the latent vector based on the reference information, the latent vector is denoised based on the predicted noise, the denoised latent vector is output, and the latent vector is decoded by the VAE decoder to generate an image.
[0067] Thus, in the forward diffusion process of the Stable Diffusion model, noise can be iteratively applied to the compressed latent variables generated by the VAE-based encoder. In the reverse denoising process, each denoising step is completed by the denoising network, which obtains the latent variable by denoising in the reverse direction from forward diffusion. The VAE decoder then converts the latent variable back into pixel space to generate the output image. Therefore, the Stable Diffusion model can train the VAE to convert the image into a latent variable in a low-dimensional latent space. The processes of adding noise and removing noise (such as Gaussian noise) are both applied to this latent variable, and the final denoising result (i.e., the latent variable output by the denoising network) can be decoded into the pixel space by the VAE decoder to achieve image generation.
[0068] Among them, the denoising network mentioned above can be a U-Net. The U-Net mentioned here is a network that uses a fully convolutional network for semantic segmentation, and its structure is a symmetrical U-shaped structure that uses a compression path and an expansion path. Specifically, U-Net may include an encoder and a decoder. The encoder is used to compress the image representation into a low-resolution image, which may include multiple downsampling layers, and the decoder is used to decode the low-resolution image back to a high-resolution image, which may include multiple upsampling layers; one downsampling layer corresponds to one upsampling layer, and in order to prevent U-Net from losing important information during downsampling, a shortcut connection may be added between each sampling layer in the encoder and the corresponding upsampling layer in the decoder. It can be understood that whether it is a downsampling layer or an upsampling layer, it is essentially a unet block (network block) in U-Net, and each unetblock may include at least ResNet (residual network). Whether in the model training process or in the model inference process, U-Net can support the input of latent variables, and perform a series of downsampling and upsampling processes on the input latent variables based on reference information (such as the vector of the input image, the vector of the input text, etc.), thereby predicting the noise in the latent variables and subtracting the predicted noise from the input latent variables to achieve denoising of the input latent variables.
[0069] Optional: ① A cross attention layer can be added to each unet block to adjust the output of the text embedding through the cross attention layer. Figure 3d As shown, the cross-attention layer can include three convolution modules (conv), three attention parameters (Query (Q), Key (K), Value (V)), MatMul (matrix product module) and Softmax (activation module). ② In order to make the quality of the image generated by U-Net more exquisite, a self-attention layer can be added to each unet block. ③ In order to generate smooth video frames, a motion module can be added to each unet block. This motion module uses the animatediff mechanism (global attention mechanism), that is, a motion module created by using global attention on the time sequence is embedded in the unetblock. During the training or generation process, this motion module can calculate a correlation attention for all frame data to complete the information exchange between frames, so that when the final video frame is generated, it can pay attention to the semantics of the surrounding frames, so that the frames in the generated video frame sequence are more continuous, and the video frame sequence is smoother and has no jumps.
[0070] Taking the unet block as an example, which includes resnet, self-attention layer, cross-attention layer, and motion module in sequence, its working principle is as follows: first call resnet to process the input image data, obtain the first processing result and input the first processing result into the self-attention layer, the self-attention layer processes the input first processing result to obtain the second processing result and input the second processing result into the cross-attention layer as the attention parameter Q in the cross-attention layer, and input the text vector sequence and the image vector sequence into the cross-attention layer as the attention parameter Q and attention parameter V of the cross-attention layer respectively, the cross-attention layer performs feature fusion on Q, K and V to obtain the fusion result (including multiple frames of image data), and inputs the fusion result into the motion module, the motion module adjusts each frame of image data based on the attention between multiple frames of image data to obtain the output result of the unet block.
[0071] S204 , generating a video frame sequence according to M frames of image data generated by the last round of image processing in the multiple rounds of image processing.
[0072] As can be seen from the above, the M frames of image data generated by the last round of image processing in multiple rounds of image processing can be M video frames or M frame feature maps. Then, when the M frames of image data generated by the last round of image processing are M video frames, the specific implementation method of step S204 can be: arranging the M frames of image data (i.e., M video frames) generated by the last round of image processing in multiple rounds of image processing to obtain a video frame sequence. When the M frames of image data generated by the last round of image processing are M frame feature maps, the specific implementation method of step S204 can be: calling an image decoder (such as a VAE decoder) to decode the M frames of image data (i.e., M frame feature maps) generated by the last round of image processing in multiple rounds of image processing to map the M frames of image data generated by the last round of image processing to the pixel space to obtain M video frames, thereby arranging the M video frames to obtain a video frame sequence.
[0073] In an embodiment of the present application, an image vector sequence can be generated by encoding a reference image for indicating the video content in a video frame sequence to be generated, and multiple rounds of image processing are performed based on the image vector sequence and an image generation model, thereby generating a video frame sequence based on M frames of image data generated by the last round of image processing. Each time M frames of image data are generated through a round of image processing, an inter-frame offset sequence can be obtained to indicate the difference between any two adjacent frames of image data in the currently generated M frames of image data, and at least one image vector in the image vector sequence can be dynamically adjusted based on the inter-frame offset sequence. The dynamically adjusted image vector sequence is then input into the image generation model for the next round of image processing. This allows the image generation model to understand the current image generation situation based on the dynamically adjusted image vector sequence during the next round of image processing, thereby deeply perceiving the generation direction of the video frame sequence, improving the generation performance of the entire model, and accurately achieving the purpose of generating a video frame sequence. Moreover, by dynamically adjusting the image vector sequence, it is also possible to avoid the situation where the information referred to by the image generation model remains unchanged during the entire video generation process, thereby avoiding the image generation model being restricted in the content of the final generated video frame sequence due to the unchanged reference information. While maintaining the video content information indicated by the reference image, the image generation model can exert its own diverse generation capabilities, making the video content of the final generated video frame sequence more colorful, thereby improving the video quality of the video frame sequence.
[0074] Based on the above description, the embodiment of the present application proposes another video generation method; in the embodiment of the present application, the video generation method is still described by taking a computer device executing the video generation method as an example. Figure 4 As shown, the video generation method may include the following steps S401-S404:
[0075] S401, obtaining a reference image, and generating an image vector sequence based on the reference image encoding.
[0076] The reference image is used to indicate the video content in the video frame sequence to be generated, and the video frame sequence includes M video frames, where M is an integer greater than 1. It is understood that the specific implementation of step S401 can be found in the relevant description of steps S201-S202 in the aforementioned method embodiment, and will not be repeated here.
[0077] S402: Obtain content description text, and generate a text vector sequence based on the content description text.
[0078] In a specific implementation, a computer device can obtain content description text and generate a text vector sequence based on the content description text, so that the text vector sequence can be subsequently input into an image generation model for multiple rounds of image processing to enrich the information referenced by the image generation model in each round of image processing, thereby helping the image generation model to fully exert its diversity generation capabilities, making the video content of the final generated video frame sequence more colorful, thereby improving the video quality of the video frame sequence.
[0079] Among them, the content description text can be used to indicate the display form of the video content in the video frame sequence. For example, if the reference image is used to indicate that the video content in the video frame sequence includes a game character, the content description text can be "A handsome male video game character is dancing". The content description text can be used to indicate that the display form of the video content (i.e., the game character) in the video frame sequence is a dancing form, so that the subsequently generated video frame sequence is a video about the game character dancing. For another example, if the reference image is used to indicate that the video content in the video frame sequence includes a little girl, the content description text can be "A little girl is walking with an umbrella". The content description text can be used to indicate that the display form of the video content (i.e., the actual person) in the video frame sequence is a walking form with an umbrella, so that the subsequently generated video frame sequence is a video about a little girl walking with an umbrella, such as Figure 5a shown.
[0080] The text vector sequence may include M text vectors, each representing the content description text. The mth text vector in the text vector sequence is used to guide the generation of the mth video frame in the video frame sequence, that is, the mth text vector in the text vector sequence is used to indicate the display form of the video content in the mth video frame in the video frame sequence. When generating the text vector sequence based on the content description text, the computer device may invoke a text encoder to perform text encoding on the content description text M times to obtain M text vectors, with each text encoding generating one text vector, and then arranging the M text vectors to obtain a text vector sequence. Alternatively, the computer device may invoke a text encoder to perform text encoding on the content description text once to obtain a text vector representing the content description text, and then replicate this text vector multiple times to obtain M text vectors, and then arrange the M text vectors to obtain a text vector sequence. Since the vector replication operation is simpler and requires fewer processing resources than text encoding, the encoding-then-replication operation to generate the text vector sequence can reduce the number of text encoding operations performed on the content description text, thereby saving processing resources required for text encoding and improving the efficiency of text vector sequence generation.
[0081] It should be noted that the present embodiment does not limit the order in which step S401 and step S402 are executed. For example, the computer device may execute step S401 first and then step S402; or the computer device may execute step S402 first and then step S401; or the computer device may execute step S401 and step S402 simultaneously.
[0082] S403: Perform multiple rounds of image processing based on the text vector sequence, the image vector sequence, and the image generation model.
[0083] Each round of image processing is performed by the image generation model based on the text vector sequence and the current input image vector sequence. Each round of image processing generates M frames of image data, where the mth frame of image data corresponds to the mth video frame, and each frame of image data includes image values at multiple locations. Taking the image generation model including a denoising network as an example, the process of the nth round of image processing includes: obtaining M frames of feature maps for the nth round of image processing, where n is a positive integer; when n = 1, the M frame feature maps are obtained by encoding M noise images, with one frame of feature map corresponding to one noise data; when n > 1, the M frame feature maps are the M frames of image data generated by the n-1th round of image processing. The M-frame feature maps and the current image vector sequence are input into the denoising network in the image generation model, so that the denoising network predicts the noise existing in each frame feature map in the M-frame feature map based on the input data (such as the currently input image vector sequence and text vector sequence), and obtains the predicted noise of each frame feature map; the denoising network is controlled to perform denoising on the corresponding feature map based on the predicted noise of each frame feature map, and the denoised M-frame feature map is used as the M-frame image data generated by the n-th round of image processing.
[0084] During multiple rounds of image processing, the computer device may input the image vector sequence and the text vector sequence into the image generation model for the first round of image processing to generate M frames of image data. Each time M frames of image data are generated through a round of image processing, the computer device may obtain an inter-frame offset sequence corresponding to the current M frames of image data, dynamically adjust at least one image vector in the image vector sequence based on the inter-frame offset sequence, and input the dynamically adjusted image vector sequence and text vector sequence into the image generation model for the next round of image processing. The inter-frame offset sequence is used to indicate the difference between any two adjacent frames of image data in the corresponding M frames of image data.
[0085] In a specific implementation, the specific implementation method of obtaining the inter-frame offset sequence corresponding to the current M frames of image data may include the following steps s11-s13:
[0086] s11, perform graph optimization processing on the current M-frame image data, where the graph optimization processing includes at least one of the following: eliminating the same image semantics at the same position point in the M-frame image data, and smoothing the image values of each position point in each frame image data. Among them, by eliminating the same image semantics at the same position point in the M-frame image data to optimize the current M-frame image data, it is possible to avoid the image generation model being affected by the image semantics of different image values during the subsequent processing process, thereby improving the accuracy of the subsequent processing results of the image generation model. By smoothing the image values of each position point in each frame image data to optimize the current M-frame image data, it is possible to eliminate the blurring and shifting problems of each image value in each frame image data, improve the data quality of each frame image data, and thus improve the accuracy of the subsequent processing results of the image generation model. Specifically:
[0087] (1) When the graph optimization process includes eliminating the same image semantics at the same position point in M frames of image data, since research shows that before the graph optimization process is performed, the image value of any position point in each frame of image data is used to represent the image semantics of the corresponding position point, the computer device can achieve the purpose of eliminating the same image semantics at the same position point in the M frames of image data by changing the image values of each position point in each frame of image data to destroy the image semantics represented by the corresponding image value. Specifically, the specific implementation of step s11 can be as follows: traverse multiple position points, and take the currently traversed position point as the i-th position point, i∈[1,G], G is the number of position points. Obtain the image value regularization parameter of the i-th position point, which is obtained by statistically analyzing the M image values at the i-th position point in the current M frames of image data. Further, based on the image value regularization parameter, regularize the M image values at the i-th position point in the current M frames of image data. Regularization here refers to the process of performing numerical adjustment according to preset rules.
[0088] In one embodiment, the statistical analysis mentioned above may include mean calculation; in this case, the specific method of obtaining the image value regularization parameter at the i-th position point may be: read out the M image values at the i-th position point from the current M frames of image data, perform mean calculation on the read M image values, and obtain the image value regularization parameter. Alternatively, the computer device may also use 3D convolution (three-dimensional convolution) technology to obtain the image value regularization parameter, thereby improving the acquisition efficiency and accuracy; the 3D convolution mentioned here is to form a cube by stacking multiple consecutive frames, and then apply a 3D convolution kernel in the cube. Through this structure, the frames in the convolution layer will be connected to multiple adjacent frames in the previous layer, thereby capturing motion information, thereby improving the accuracy of the convolution result. Specifically, the computer device can obtain a first convolution layer, which includes M channels of first convolution kernels, each of which has a size of 1×1, and the first convolution kernels of the M channels are used to perform convolution mean calculation on M image values at the same position point in the current M frames of image data (i.e., calculate the mean between the M image values); call the first convolution layer to perform convolution mean calculation on the current M frames of image data to obtain a first convolution map; wherein the first convolution map includes G first convolution values, and the i-th first convolution value is used to represent: the mean of the M image values at the i-th position point in the current M frames of image data. From the first convolution map, obtain the i-th first convolution value as the image value normalization parameter at the i-th position point.
[0089] Based on the above, it can be seen that when the statistical analysis includes mean calculation, the image value normalization parameter includes: the mean of the M image values at the i-th position point in the current M frames of image data. Accordingly, in this case, based on the image value normalization parameter, a specific implementation method for normalizing the M image values at the i-th position point in the current M frames of image data can be: using the mean of the M image values at the i-th position point in the current M frames of image data as the mean of the image values corresponding to the i-th position point; traversing the M image values at the i-th position point in the current M frames of image data, and using the currently traversed image value as the m-th image value; calculating the difference between the m-th image value and the mean of the image values corresponding to the i-th position point to obtain a first difference; and using the ratio between the first difference and the mean of the image values corresponding to the i-th position point to update the m-th image value, so that the updated m-th image value is the ratio between the first difference and the mean of the image values corresponding to the i-th position point.
[0090] For example, take M=4 (i.e., a total of M frames of image data) and i=1 (i.e., the first position point) as an example; assuming that the four image values at the first position point in the current M frames of image data are 3, 5, 8, and 2, respectively, the mean value can be determined to be (3+5+8+2) / 4=4.5. For the first image value of these four image values (i.e., the value 3), the difference between the first image value and the mean value (i.e., the value 4.5) can be calculated to be -1.5, so the ratio between -1.5 and the mean value (i.e., the value 4.5) (i.e., -0.33) can be used to update the first image value; similarly, for the second image value of the four image values (i.e., the value 5), the difference between the value 5 and the mean value (i.e., the value 4.5) can be calculated to be 0.5, so the ratio between 0.5 and the mean value (i.e., the value 4.5) (i.e., 0.11) can be used to update the second image value. Image value; for the third of the four image values (i.e., value 8), the difference between value 8 and the mean (i.e., value 4.5) is calculated to be 3.5. Therefore, the ratio between 3.5 and the mean (i.e., value 4.5) (i.e., 0.78) can be used to update the third image value; for the fourth of the four image values (i.e., value 2), the difference between value 2 and the mean (i.e., value 4.5) is calculated to be -2.5. Therefore, the ratio between -2.5 and the mean (i.e., value 4.5) (i.e., -0.46) can be used to update the fourth image value. This results in the following four regularized image values: -0.33, 0.11, 0.78, and -0.46.
[0091] Based on the above, it can be seen that by using the mean of the M image values at the i-th position in the current M frames of image data to regularize the corresponding M image values, the regularized M image values can be made to be within the interval of (-1,1) regardless of the value range of the image data. This can help the image generation model to perform subsequent processing, avoid the problem that the image generation model cannot recognize and process image values within the same value range due to the image generation model not learning the image values within the corresponding value range during the model training process, improve the processing success rate and generalization ability of the image generation model, and help the image generation model converge. In addition, since the data processed by the image generation model usually contains negative numbers, regularizing the M image values in this way can regularize each image value less than the mean to a negative number, and each image value greater than the mean to a positive number, so that the regularized M image values are more in line with the processing capabilities of the image generation model, thereby improving the processing accuracy of the image generation model.
[0092] In another embodiment, the statistical analysis mentioned above may include the processing of statistical maximum image values and minimum image values; in this case, the specific method of obtaining the image value regularization parameters of the i-th position point may be: read out M image values at the i-th position point from the current M frames of image data, arrange the read M image values in descending or ascending order, and determine the maximum and minimum values among the M image values based on the arrangement result.
[0093] Based on the above, when the statistical analysis includes mean calculation, the image value normalization parameter includes: the mean of the M image values at the i-th position in the current M frames of image data. Accordingly, in this case, based on the image value normalization parameter, a specific implementation method for normalizing the M image values at the i-th position in the current M frames of image data can be: calculating the difference between the maximum image value and the minimum image value among the M image values at the i-th position in the current M frames of image data to obtain a second difference; for the m-th image value among the M image values, calculating the difference between the m-th image value and the minimum image value, and updating the m-th image value using the ratio between the calculated difference and the second difference.
[0094] For example, take M=4 (i.e., a total of M frames of image data) and i=1 (i.e., the first position point); assuming that the four image values at the first position point in the current M frames of image data are 3, 5, 8, and 2, respectively, it can be determined that the maximum image value is 8 and the minimum image value is 2, and the difference between the maximum image value and the minimum image value (i.e., the second difference) can be calculated as 8-2=6. For the first image value (i.e., the value 3) of these four image values, the difference between the first image value and the minimum image value (i.e., the value 2) can be calculated to be 1, so the first image value can be updated using the ratio between 1 and the second difference (i.e., the value 6) (i.e., 0.17); similarly, for the second image value (i.e., the value 5) of the four image values, the difference between the value 5 and the minimum image value (i.e., the value 2) can be calculated to be 3, so the second image value can be updated using the ratio between 3 and the second difference (i.e., the value 6) (i.e., 0.5). 2 image values; for the third of the four image values (i.e., value 8), the difference between 8 and the minimum image value (i.e., value 2) is calculated to be 6. Therefore, the ratio of 6 to the second difference (i.e., value 6) (i.e., 1) is used to update the third image value; for the fourth of the four image values (i.e., value 2), the difference between 2 and the minimum image value (i.e., value 2) is calculated to be 0. Therefore, the ratio of 0 to the second difference (i.e., value 6) (i.e., 0) is used to update the fourth image value. Thus, the four regularized image values are: 0.17, 0.5, 1, and 0, respectively.
[0095] Based on the above, it can be seen that by using the maximum and minimum image values in the current M frames of image data to regularize the corresponding M image values, the regularized M image values can be made to be within the interval [0, 1] regardless of the value range of the image data. This can help the image generation model perform subsequent processing, avoiding the problem of the image generation model not being able to recognize and process image values within the same value range due to not learning image values within the corresponding value range during the model training process. This improves the processing success rate and generalization ability of the image generation model, and helps the image generation model converge. In addition, the entire regularization process does not require tedious operations, which can also improve the efficiency of regularization.
[0096] Optionally, considering that the M frames of image data generated by each round of image processing are intermediate results in the iterative execution of multiple rounds of image processing, their content is still in an incomplete state and there may be problems such as ambiguity and displacement in the content, that is, there may be displacement problems of the image values originally at the same position point in the M frames of image data, which will result in the same image semantics of these positions not being completely eliminated. Based on this, in order to improve the accuracy of eliminating image semantics and avoid the influence of the same image semantics around the same pixel position, the computer device can perform the following operations after traversing multiple position points based on the above operations so that the M image values at any position point in the current M frames of image data are regularized:
[0097] Obtain a second convolutional layer, which includes M channels of second convolution kernels. Each second convolution kernel has a size of P×P, where P is an integer greater than 1 and can be set based on actual needs or empirical values. The second convolution kernel of the mth channel is used to perform convolution mean calculation on the mth frame of image data. Call the second convolutional layer to perform convolution mean calculation on the current M frames of image data, obtaining M second convolution maps. Each second convolution map includes J second convolution values, where J is a positive integer. The jth second convolution value in the mth second convolution map is used to represent the mean of the image values within the jth region of size P×P in the mth frame of image data, where j∈[1,J]. Regularize the j-th second convolution value in the M second convolution maps to obtain M regularized second convolution values. The regularization method here can be mean regularization or regularization based on maximum and minimum values. The specific implementation method can refer to the relevant description of the step of "regularizing the M image values at the i-th position point in the current M frames of image data" mentioned above, which is not repeated here. In the m-th frame of image data included in the current M frames of image data, each image value in the j-th area of size P×P is updated to the m-th regularized second convolution value.
[0098] For example, take M=4 (i.e., a total of M frames of image data), P=2 (i.e., the size of the convolution kernel is 2×2), and j=1 (i.e., the first area of size 2×2) as an example; suppose that the image values in the first area of size 2×2 in the first frame of image data are (2, 4, 3, 1), the image values in the first area of size 2×2 in the second frame of image data are (3, 1, 5, 3), the image values in the first area of size 2×2 in the third frame of image data are (2, 4, 3, 1), and the image values in the first area of size 2×2 in the fourth frame of image data are (1, 4, 5, 2). Then, the first second convolution value in the first second convolution map is (2+4+3+1) / 4=2.5, the first second convolution value in the second second convolution map is (3+1+5+3) / 4=3, the first second convolution value in the third second convolution map is (2+4+3+1) / 4=2.5, and the first second convolution value in the fourth second convolution map is (1+4+5+2) / 4=3. If the four second convolution values (i.e., 2.5, 3, 2.5, 3) are regularized by using the mean regularization method, the four regularized second convolution values are -0.09, 0.09, -0.09, and 0.09, respectively; then, the regularized second convolution values can be used to update the image values in the first 2×2 area in the corresponding frame image, so that the image values in the first 2×2 area in the first frame image data are all updated to (-0.09, -0.09, -0.09). 9, -0.09), the image values in the first 2×2 area in the second frame image data are all updated to (0.09, 0.09, 0.09, 0.09), the image values in the first 2×2 area in the third frame image data are all updated to (-0.09, -0.09, -0.09, -0.09), and the image values in the first 2×2 area in the fourth frame image data are all updated to (0.09, 0.09, 0.09, 0.09).
[0099] It should be noted that the aforementioned implementation for eliminating the influence of identical image semantics around the same location is based on 3D convolution technology; in other embodiments, other technologies may also be used. For example, M sliding windows can be placed on the current M frames of image data and controlled to slide synchronously on the current M frames of image data. One data window is placed on each frame of image data, and the M data windows are aligned in a target direction, which refers to the direction in which the current M frames of image data are aligned. With each sliding operation, the area covered by each sliding window on the corresponding frame of image data is used as the target region, resulting in M target regions. Reference image values for the M target regions are obtained, where the reference image value for any target region is obtained by averaging the image values within the corresponding target region. The reference image values for the M target regions are normalized to obtain a normalized image value for each target region, which can be specifically normalized by average or based on maximum and minimum values. In the current M frames of image data, all image values within any target region are updated to the normalized image value of the corresponding target region.
[0100] (2) When the graph optimization process includes smoothing the image values of each position point in each frame of image data, the specific implementation of step s11 can be as follows: obtain the third convolution layer, the third convolution layer includes a third convolution kernel of 1 channel, the size of the third convolution kernel is R×R, R is an integer greater than 1, and the value of R can be set according to the empirical value or actual needs. Poll the current M frames of image data to determine the currently polled m-th frame of image data, and then call the third convolution layer to perform convolution mean calculation on the m-th frame of image data to obtain T third convolution values, T is a positive integer, and the t-th third convolution value is used to represent: the mean of each image value in the t-th area of size R×R in the m-th frame of image data, t∈[1,T]. In the m-th frame of image data, each image value in the t-th area of size R×R is updated to the t-th third convolution value. It can be seen that the embodiment of the present application can use convolution technology to achieve smoothing processing of each frame of image data, so that the smoothing efficiency can be improved with the help of convolution technology.
[0101] Alternatively, the computer device may implement smoothing of each frame of image data using other methods. For example, the computer device may place a sliding window on the mth frame of image data among the current M frames of image data and control the sliding window to slide over the mth frame of image data. Each time the sliding window is moved, the image values of each position within the sliding window at the mth frame of image data are used as target image values to be smoothed; the mean of each target image value is calculated to obtain a target mean value; and each target image value is updated to the target mean value.
[0102] S12, after the image optimization process, the difference processing is performed on the two adjacent frames of image data in the current M frames of image data in turn to obtain M-1 inter-frame offset matrices. The difference processing here includes: calculating the difference between the two image values at the same position point in the two adjacent frames of image data. For example, see Figure 5b As shown: the difference between two image values at the same position point in the first frame image data and the second frame image data can be calculated to obtain the first inter-frame offset matrix, which is used to indicate the difference between the first frame image data and the second frame image data.
[0103] S13: Using the M-1 inter-frame offset matrices, construct an inter-frame offset sequence corresponding to the currently generated M frames of image data. Specifically, the computer device can directly arrange the M-1 inter-frame offset matrices to obtain the inter-frame offset sequence corresponding to the currently generated M frames of image data.
[0104] Alternatively, the computer device may also construct an inter-frame offset matrix composed of zero values, and add the constructed inter-frame offset matrix in front of M-1 inter-frame offset matrices to obtain M inter-frame offset matrices arranged in sequence, that is, the inter-frame offset matrix composed of zero values is in the first place. The M values at the same position point in the M inter-frame offset matrices are statistically analyzed to obtain the numerical regularization parameters of each position point. The statistical analysis here can be mean calculation, or statistical maximum and minimum value processing. Based on the numerical regularization parameters of each position point, the M values at the corresponding position points in the M inter-frame offset matrices are regularized to obtain the inter-frame offset sequence corresponding to the currently generated M frame image data. The specific regularization method can refer to the relevant description of the step of "regularizing the M image values at the i-th position point in the current M frame image data" mentioned above, which will not be repeated here.
[0105] It can be seen that by constructing a pure zero inter-frame offset matrix to obtain M inter-frame offset matrices, and regularizing the M inter-frame offset matrices to obtain an inter-frame offset sequence, the number of inter-frame offset matrices in the final inter-frame offset sequence can be made the same as the number of image vectors in the image vector sequence (both are M). This facilitates the subsequent dynamic adjustment of the image vector sequence based on the inter-frame offset sequence, improves the convenience of dynamic adjustment, and thus improves the efficiency of dynamic adjustment; and, in this way, dynamic adjustment of the first image vector in the image vector sequence can also be achieved, avoiding the first image vector used to guide the generation of the first video frame from remaining unchanged during multiple rounds of image processing, thereby improving the diversity of the first video frame.
[0106] It should be noted that the above steps s11-s13 are merely exemplary embodiments of obtaining an inter-frame offset sequence and are not exhaustive. For example, in other embodiments, step s11 may be omitted (i.e., image optimization processing may not be performed on the current M frames of image data), and step 12 may be performed directly to perform difference processing on two adjacent frames of image data in the current M frames of image data to obtain M-1 inter-frame offset matrices. These M-1 inter-frame offset matrices are then used in step s13 to construct an inter-frame offset sequence corresponding to the currently generated M frames of image data.
[0107] Based on the above description, it can be seen that the inter-frame offset sequence includes at least one inter-frame offset matrix, one inter-frame offset matrix corresponds to one frame of image data, and any inter-frame offset matrix is used to indicate the difference between the corresponding frame image data and the adjacent frame image data. Based on this, a specific implementation method for dynamically adjusting at least one image vector in the image vector sequence based on the inter-frame offset sequence can be: traversing each inter-frame offset matrix in the inter-frame offset sequence, and taking the frame of image data corresponding to the currently traversed inter-frame offset matrix as the target frame image data; from the image vector sequence, extracting the image vector used to generate the target frame image data as the target image vector. For example, if the target frame image data is the 5th frame of image data, the target image vector is the 5th image vector. The target image vector is dynamically adjusted using the currently traversed inter-frame offset matrix to obtain the dynamically adjusted target image vector.
[0108] Specifically, the product operation can be performed on the currently traversed inter-frame offset matrix and the target image vector, and the result of the product operation is used as the dynamically adjusted target image vector. Alternatively, see Figure 5c As shown in the figure: based on the size of the currently traversed inter-frame offset matrix, the size of the target image vector can be transformed (MLP+Reshape) to obtain an intermediate image vector; the cross attention network is called to dynamically adjust the intermediate image vector according to the currently traversed inter-frame offset matrix to obtain a dynamically adjusted intermediate image vector. Specifically, the intermediate image vector can be used as the attention parameter Q, and the currently traversed inter-frame offset matrix can be used as the attention parameters K and attention parameters V, so as to achieve dynamic adjustment of the intermediate image vector, as shown in the figure: Figure 5d Furthermore, the size of the dynamically adjusted intermediate image vector can be transformed based on the original size of the target image vector to obtain the dynamically adjusted target image vector; wherein the original size refers to the size of the target image vector before the size of the target image vector is transformed.
[0109] S404 , generating a video frame sequence according to M frames of image data generated by the last round of image processing in the multiple rounds of image processing.
[0110] In an embodiment of the present application, an image vector sequence can be generated by encoding a reference image for indicating the video content in a video frame sequence to be generated, and multiple rounds of image processing are performed based on the image vector sequence and an image generation model, thereby generating a video frame sequence based on M frames of image data generated by the last round of image processing. Each time M frames of image data are generated through a round of image processing, an inter-frame offset sequence can be obtained to indicate the difference between any two adjacent frames of image data in the currently generated M frames of image data, and at least one image vector in the image vector sequence can be dynamically adjusted based on the inter-frame offset sequence. The dynamically adjusted image vector sequence is then input into the image generation model for the next round of image processing. This allows the image generation model to understand the current image generation situation based on the dynamically adjusted image vector sequence during the next round of image processing, thereby deeply perceiving the generation direction of the video frame sequence, improving the generation performance of the entire model, and accurately achieving the purpose of generating a video frame sequence. Moreover, by dynamically adjusting the image vector sequence, it is also possible to avoid the situation where the information referred to by the image generation model remains unchanged during the entire video generation process, thereby avoiding the image generation model being restricted in the content of the final generated video frame sequence due to the unchanged reference information. While maintaining the video content information indicated by the reference image, the image generation model can exert its own diverse generation capabilities, making the video content of the final generated video frame sequence more colorful, thereby improving the video quality of the video frame sequence.
[0111] Based on the above description, an embodiment of the present application proposes a method for dynamically adjusting an image vector sequence based on an inter-frame offset calculation mechanism to control video sequence generation. This method can support generating a video frame sequence with specified content after a reference image (ref image) is input into the system. In this video frame sequence, a richer content form is generated, which is not restricted by the content form in the ref image. Moreover, the generated video frame sequence maintains the content information consistent with the ref image to a great extent, greatly improving the performance of the entire generated video frame sequence.
[0112] Specifically, the method can be constructed by constructing a system for generating an image-to-video (image-to-video) image, calculating the inter-frame offset of the M-frame feature maps generated by each denoising, and dynamically adjusting the image vector sequence input into the model in real time based on the calculated inter-frame offset sequence. In this way, the model is not constrained by the content information of only one reference image provided in the process of generating a video frame sequence. At the same time, it can also give full play to the model's diversified generation capabilities, making the video content of the generated video frame sequence richer, greatly improving the effect of video generation. This method is constructed based on the Stable Diffusion model in the process of denoising the video frame sequence. During each denoising iteration of the denoising network in the Stable Diffusion model, the inter-frame offset can be calculated for the M-frame feature maps (i.e., latent variables) completed by the current iteration to understand the denoising situation in the entire model, and the content direction that the current model wants to generate the final video frame sequence can be obtained through the offset calculation.
[0113] In a specific implementation, the embodiment of the present application can use the image generation model as the basic architecture to build a Stable Diffusion model that integrates the motion module in advance. This model serves as a backbone generation network of the embodiment of the present application. Then, based on this backbone generation network, two innovative mechanisms are constructed, namely, an inter-frame offset calculation mechanism and a reference information dynamic adjustment mechanism, to dynamically adjust the reference guidance information used in each noise reduction iteration during the generation process in real time. Based on these two innovative mechanisms, the embodiment of the present application designs two modules, namely, an inter-frame offset calculation module and a reference information adjustment module. Figure 5e As shown. Among them, the inter-frame offset calculation module is mainly used to use the inter-frame offset calculation mechanism to calculate the inter-frame offset for the M frame feature maps calculated in the current noise reduction iteration process, and then after a series of fusion and regularization to obtain the inter-frame offset sequence, it is input into the reference information adjustment module together with the image vector sequence calculated in advance, so that the reference information adjustment module uses the reference information dynamic adjustment mechanism to calculate the image vector sequence used in the next noise reduction iteration process. The calculation principle is: use the inter-frame offset sequence to dynamically adjust the image vector sequence, and input the dynamically adjusted image vector sequence into the model for the next noise reduction iteration process.
[0114] Taking M=16 as an example, the inter-frame offset calculation mechanism and the reference information dynamic adjustment mechanism proposed in the embodiments of the present application are respectively described:
[0115] (1) Inter-frame offset calculation mechanism, which receives 16-frame feature maps (a latent sequence result) calculated by each denoising iteration. Each frame in the 16-frame feature map represents the intermediate result of the model at the current moment when calculating the video frame sequence. Figure 5f The calculation process of the inter-frame offset calculation mechanism proposed in the embodiment of the present application is shown:
[0116] The first step is to calculate the 16-frame feature map through a 16-channel 1x1 convolution layer. The main function of this convolution is to regularize the common features between the 16-frame feature maps. Specifically, the 16-channel 1×1 convolution layer can be called to perform the mean operation on the 16 eigenvalues at the same position in the 16-frame feature map. Then, the calculated mean can be used to adjust the 16 eigenvalues at the corresponding position in the 16-frame feature map. Specifically, any eigenvalue at the corresponding position is subtracted from the calculated mean, and the difference is divided by the mean. In this way, any eigenvalue can be regularized. The physical meaning of this step is to use the eigenvalues at the same position in all the generated 16-frame feature maps, and then perform the fusion and regularization of all the eigenvalues at the position to remove the common content and semantics. Because this mechanism is to calculate the inter-frame offset between feature maps, its main purpose is to calculate the difference between the feature values of each corresponding position point between frames. Therefore, this step is mainly to be able to eliminate the same image semantics implied by the feature values at the same position point in the 16-frame feature map.
[0117] The second step is to calculate through a 16-channel 2x2 convolution layer. The main function of this step is the same as the 16-channel 1x1 convolution in the first step, but a larger size 2x2 convolution is used here mainly because each feature map calculated here is an intermediate result in the denoising iterative process, so its content is still in an incomplete state, and there is a high probability of content blur and displacement. Therefore, a larger size convolution is used here to avoid the influence of the same image semantics around the same position point. After this convolution calculation, the influence of the same image semantics around the same position point can be eliminated (of course, a larger size convolution can be used as this step in practical applications). Specifically, the 16-channel 2×2 convolution layer is called to perform mean calculation in the 2×2 area in each frame to obtain 16 mean values a, calculate the mean between the 16 mean values a, and obtain the mean b. The mean b is used to regularize each mean a (in the same way as in the first step), and the feature value in the 2×2 area corresponding to any mean a is updated to the regularization result corresponding to the corresponding mean a.
[0118] The third step is a 1-channel 2x2 convolutional layer. Compared to the first two steps, this step mainly smoothes the feature values on each frame's feature map. This is mainly to eliminate the blur and displacement problems in each feature map during the noise reduction iteration. (Of course, this step is the same as the above. Subsequent experiments can use a larger convolution layer as this step.)
[0119] The fourth step is to calculate the inter-frame offset sequence. For the 16-frame feature maps extracted through the previous steps, the difference between the two eigenvalues of the corresponding position points between each two-frame feature map in the 16-frame feature map can be calculated to obtain the inter-frame offset matrix between the corresponding two-frame feature maps. After calculation, 15 inter-frame offset matrices can be obtained. A pure 0 inter-frame offset matrix is added in front of the 15 inter-frame offset matrices to obtain 16 inter-frame offset matrices arranged in sequence. The mean of these 16 inter-frame offset matrices is used to regularize these 16 offset matrices (the regularization method is the same as the first step), thereby obtaining an inter-frame offset sequence.
[0120] (2) Reference information dynamic adjustment mechanism, which receives an inter-frame offset sequence, which represents the situation of the feature map generated during the last noise reduction iteration. Therefore, this mechanism can adjust the image vector image required for the next noise reduction iteration according to the result of the last noise reduction iteration, so that the video frame sequence of the entire model continues to be generated and calculated in the current direction. It can be seen that, except for the first round of noise reduction iteration, the embodiment of the present application uses an inter-frame offset sequence to dynamically update the image vector sequence starting from the second round of noise reduction iteration, so that the dynamically updated image vector sequence is input into the model to participate in the next round of noise reduction. Specifically, as mentioned above Figure 5c As shown, the calculation process of the reference information dynamic adjustment mechanism proposed in the embodiment of the present application is as follows:
[0121] Given a received inter-frame offset sequence, each inter-frame offset matrix in the inter-frame offset sequence is traversed. The image vector associated with the currently traversed inter-frame offset matrix is extracted from the image vector sequence as the target image vector. The target image vector is then passed through multiple fully-connected layers and reshape operations to resize the target image vector to the same size as the currently traversed inter-frame offset matrix, yielding an intermediate image vector. The intermediate image vector and the currently traversed inter-frame offset matrix are then fed into the attention calculation. This calculation utilizes a criss-cross attention network with the same structure as the criss-cross attention layer in U-Net. Because the intermediate image vector is adjusted, it serves as the query in the criss-cross attention layer. Since the intermediate image vector is adjusted based on the inter-frame offset matrix, the inter-frame offset matrix serves as the key and value in the criss-cross attention network. After the attention calculation, the adjusted intermediate image vector is obtained. This dynamically adjusted intermediate image vector is then passed through a series of fully-connected layers and reshape operations to convert it back to its original size, yielding the dynamically adjusted target image vector. This is done to preserve the size of the image vectors used by the overall model. After a series of frame calculations, the dynamically adjusted image vector sequence can be calculated and input into the model network for the next denoising iterative calculation, just like the text vector sequence.
[0122] Based on the above description, this method designs a conv-based module to capture and calculate the inter-frame offset, and uses the attention mechanism to regularize the common characteristics of the inter-frame offset, so that the directional offset of the model in the current denoising iteration can be calculated more accurately. The calculated inter-frame offset sequence is then passed to the reference image encoding module, in which an innovative dynamic real-time image vector adjustment mechanism is constructed. By using an input reference image and the currently calculated inter-frame offset sequence, an image vector sequence with a dynamic offset can be generated, which is then passed to the next model generation denoising iteration, thereby improving the final generation effect.
[0123] In summary, the embodiments of the present application can have at least the following beneficial effects:
[0124] (1) A method for dynamically adjusting an image vector sequence based on an inter-frame offset calculation mechanism to control the generation of a video frame sequence is constructed. This method is based on providing a reference image information to generate a dynamic video frame sequence. By innovatively constructing a real-time dynamic adjustment method for the image vector sequence used in the model for generating a video frame sequence, the model can be improved to maintain a high degree of consistency with the image information of the provided reference image in the final generated video frame sequence. At the same time, it is not constrained by the limited image content of the reference image, and reduces the difficulty of preparing the reference image required for generating a video frame sequence (no need to provide multiple reference images), which greatly improves the generation performance and effect of the model.
[0125] (2) An innovative mechanism for calculating the offset between feature maps generated by the model is constructed. This mechanism can calculate the offset between frames in the current M-frame feature map for each round of denoising during each denoising iteration of the model, and obtain an inter-frame offset sequence. The inter-frame offset sequence is then passed to the next denoising process of the model after a series of standardization and regularization, so as to adjust at least one image vector of the image vector sequence to calculate the guidance content. This allows the model to feel the actual denoising situation of the current generation work in each denoising iteration in real time, and deeply perceive the target direction of the current model to generate the video frame sequence, so as to improve the denoising performance of the model and ultimately improve the video quality of the video frame sequence generated by the model.
[0126] (3) Innovatively based on the inter-frame offset sequence, a method for dynamically adjusting the image vector sequence to control the generation of the video sequence is calculated. This method can dynamically adjust the provided image vector sequence by calculating the inter-frame offset sequence in the previous round of noise reduction iteration, and input the dynamically adjusted image vector sequence into the current noise reduction iteration. After the method dynamically adjusts the information of the image vector sequence in real time, it will not cause the model to be restricted in the final video content due to the unchanging information of the image vector sequence (because in the process of generating the video frame sequence, there is often only one reference image, and when generating all video frames, this image is used as a reference, which can easily cause solidification). The dynamically adjusted information can, on the one hand, maintain the content information in the original reference image to a great extent, and on the other hand, give full play to the model's own diversity generation ability, so that the content of the final generated video frame sequence can be richer and more exciting, thereby improving the quality of the final generated video product.
[0127] Based on the description of the above method embodiment, the present application embodiment also discloses a video generation device; see Figure 6 As shown, the video generation device can run the following units:
[0128] An acquisition unit 601 is configured to acquire a reference image, where the reference image is used to indicate video content in a video frame sequence to be generated, where the video frame sequence includes M video frames, where M is an integer greater than 1.
[0129] A processing unit 602 is configured to generate an image vector sequence based on the reference image encoding, where the image vector sequence includes M image vectors, each of which is used to represent the reference image; the m-th image vector in the image vector sequence is used to guide generation of the m-th video frame in the video frame sequence, where m∈[1,M].
[0130] The processing unit 602 is further configured to perform multiple rounds of image processing based on the image vector sequence and the image generation model, each round of image processing generating M frames of image data, the m-th frame of image data corresponding to the m-th video frame; wherein, during the multiple rounds of image processing, each time M frames of image data are generated after a round of image processing, an inter-frame offset sequence corresponding to the current M frames of image data is obtained, at least one image vector in the image vector sequence is dynamically adjusted based on the inter-frame offset sequence, and the dynamically adjusted image vector sequence is input into the image generation model for the next round of image processing; the inter-frame offset sequence is used to indicate the difference between any two adjacent frames of image data in the corresponding M frames of image data;
[0131] The processing unit 602 is further configured to generate the video frame sequence according to the M frames of image data generated by the last round of image processing in the multiple rounds of image processing.
[0132] In a specific embodiment, each frame of image data includes image values of multiple position points; accordingly, when the processing unit is used to obtain the inter-frame offset sequence corresponding to the current M frames of image data, it can be specifically used to:
[0133] Performing graph optimization processing on the current M frames of image data, wherein the graph optimization processing includes at least one of the following: eliminating identical image semantics at the same position point in the M frames of image data, and smoothing image values of each position point in each frame of image data;
[0134] After the image optimization processing, performing difference processing on two adjacent frames of image data in the current M frames of image data in sequence to obtain M-1 inter-frame offset matrices; the difference processing includes: calculating the difference between two image values at the same position point in the two adjacent frames of image data;
[0135] The M-1 inter-frame offset matrices are used to construct an inter-frame offset sequence corresponding to the currently generated M frames of image data.
[0136] In another specific embodiment, before performing the image optimization process, the image value of any position point in each frame of image data is used to represent the image semantics of the corresponding position point;
[0137] Accordingly, when the image optimization processing includes eliminating identical image semantics at the same position in the M frames of image data, the processing unit 602, when performing the image optimization processing on the current M frames of image data, may be specifically configured to:
[0138] Traverse the plurality of position points, and take the currently traversed position point as the i-th position point, i∈[1,G], where G is the number of position points;
[0139] Obtaining an image value regularization parameter for the i-th position point, where the image value regularization parameter is obtained by statistically analyzing M image values at the i-th position point in the current M frames of image data;
[0140] Based on the image value regularization parameters, the M image values at the i-th position point in the current M frames of image data are regularized; wherein, the regularization refers to: a process of adjusting the values according to a preset rule.
[0141] In another specific implementation, the statistical analysis includes mean calculation; accordingly, when the processing unit 602 is used to obtain the image value normalization parameter of the i-th position point, it can be specifically used to:
[0142] Obtain a first convolutional layer, where the first convolutional layer includes M-channel first convolution kernels, each of which has a size of 1×1, and the M-channel first convolution kernels are used to perform convolution mean calculation on M image values at the same position in the current M frames of image data;
[0143] Calling the first convolution layer to perform convolution mean calculation on the current M frames of image data to obtain a first convolution map; wherein the first convolution map includes G first convolution values, and the i-th first convolution value is used to represent: the mean of the M image values at the i-th position in the current M frames of image data;
[0144] From the first convolution map, obtain the i-th first convolution value as the image value regularization parameter of the i-th position point.
[0145] In another specific implementation, the image value normalization parameter includes: the mean value of M image values at the i-th position point in the current M frames of image data;
[0146] Accordingly, when the processing unit 602 is used to regularize the M image values at the i-th position point in the current M frames of image data based on the image value regularization parameter, it can be specifically used to:
[0147] The mean of the M image values at the i-th position point in the current M frames of image data is used as the mean of the image values corresponding to the i-th position point;
[0148] Traverse the M image values at the i-th position in the current M frames of image data, and use the currently traversed image value as the m-th image value;
[0149] Calculating a difference between the m-th image value and a mean of the image values corresponding to the i-th position point to obtain a first difference;
[0150] The m-th image value is updated by using a ratio between the first difference and a mean of the image values corresponding to the i-th position point.
[0151] In another specific implementation, the image value normalization parameters include: a maximum image value and a minimum image value among the M image values at the i-th position point in the current M frames of image data;
[0152] Accordingly, when the processing unit 602 is used to regularize the M image values at the i-th position point in the current M frames of image data based on the image value regularization parameter, it can be specifically used to:
[0153] Calculating a difference between a maximum image value and a minimum image value among the M image values at the i-th position point in the current M frames of image data to obtain a second difference value;
[0154] For an mth image value among the M image values, a difference between the mth image value and the minimum image value is calculated, and the mth image value is updated by using a ratio between the calculated difference and the second difference.
[0155] In another specific embodiment, after all the multiple position points are traversed, the processing unit 602, when performing image optimization processing on the current M frames of image data, may further be used to:
[0156] Obtain a second convolutional layer, where the second convolutional layer includes second convolution kernels of M channels, each of which has a size of P×P, where P is an integer greater than 1; the second convolution kernel of the mth channel is used to perform convolution mean calculation on the mth frame of image data;
[0157] Call the second convolutional layer to perform convolution mean calculation on the current M frames of image data to obtain M second convolution maps; each second convolution map includes J second convolution values, where J is a positive integer, and the j-th second convolution value in the m-th second convolution map is used to represent: the mean of each image value in the j-th region of size P×P in the m-th frame of image data, j∈[1,J];
[0158] Regularizing the j-th second convolution value in the M second convolution images to obtain M regularized second convolution values;
[0159] In the m-th frame of image data included in the current M frames of image data, each image value in the j-th region with a size of P×P is updated to the m-th regularized second convolution value.
[0160] In another specific embodiment, when the graph optimization processing includes smoothing the image values of each position point in each frame of image data, the processing unit 602, when performing the graph optimization processing on the current M frames of image data, may be specifically configured to:
[0161] Obtain a third convolutional layer, where the third convolutional layer includes a third convolution kernel of one channel, and the size of the third convolution kernel is R×R, where R is an integer greater than 1;
[0162] Poll the current M frames of image data to determine the currently polled m-th frame of image data;
[0163] Calling the third convolutional layer to perform convolution mean calculation on the m-th frame image data to obtain T third convolution values, where T is a positive integer, and the t-th third convolution value is used to represent: the mean of each image value in the t-th region of size R×R in the m-th frame image data, t∈[1,T];
[0164] In the m-th frame image data, each image value in the t-th region with a size of R×R is updated to the t-th third convolution value.
[0165] In another specific implementation, when the processing unit 602 is configured to construct the inter-frame offset sequence corresponding to the currently generated M frames of image data using the M-1 inter-frame offset matrices, it may be specifically configured to:
[0166] Constructing an inter-frame offset matrix consisting of zero values, and adding the constructed inter-frame offset matrix in front of the M-1 inter-frame offset matrices to obtain M inter-frame offset matrices arranged in sequence;
[0167] Performing statistical analysis on M values at the same position point in the M inter-frame offset matrices to obtain a numerical regularization parameter for each position point;
[0168] Based on the numerical regularization parameters of each position point, the M numerical values at the corresponding position points in the M inter-frame offset matrices are regularized to obtain the inter-frame offset sequence corresponding to the currently generated M frames of image data; wherein, the regularization refers to: performing numerical adjustment processing according to preset rules.
[0169] In another specific embodiment, the inter-frame offset sequence includes at least one inter-frame offset matrix, one inter-frame offset matrix corresponds to one frame of image data, and any inter-frame offset matrix is used to indicate the difference between the corresponding frame of image data and an adjacent frame of image data;
[0170] Accordingly, when the processing unit 602 is configured to dynamically adjust at least one image vector in the image vector sequence based on the inter-frame offset sequence, it may be specifically configured to:
[0171] Traversing each inter-frame offset matrix in the inter-frame offset sequence, and taking a frame of image data corresponding to the currently traversed inter-frame offset matrix as target frame image data;
[0172] Extracting, from the image vector sequence, an image vector used to generate the target frame image data as a target image vector;
[0173] The target image vector is dynamically adjusted using the currently traversed inter-frame offset matrix to obtain a dynamically adjusted target image vector.
[0174] In another specific implementation, when the processing unit 602 is configured to dynamically adjust the target image vector using the currently traversed inter-frame offset matrix to obtain the dynamically adjusted target image vector, it may be specifically configured to:
[0175] Based on the size of the currently traversed inter-frame offset matrix, the size of the target image vector is transformed to obtain an intermediate image vector;
[0176] Calling a cross attention network to dynamically adjust the intermediate image vector according to the currently traversed inter-frame offset matrix to obtain a dynamically adjusted intermediate image vector;
[0177] Based on the original size of the target image vector, the size of the dynamically adjusted intermediate image vector is transformed to obtain the dynamically adjusted target image vector; wherein the original size refers to: the size of the target image vector before the size of the target image vector is transformed.
[0178] In another specific embodiment, the image generation model includes a denoising network; when the processing unit 602 is used for the nth round of image processing, it can be specifically used to:
[0179] Obtaining M frames of feature maps for performing the n-th round of image processing, where n is a positive integer; when n=1, the M frames of feature maps are obtained by encoding M noise images, and one frame of feature map corresponds to one noise data; when n>1, the M frames of feature maps are M frames of image data generated by the n-1th round of image processing;
[0180] Inputting the M frame feature maps and the current image vector sequence into a denoising network in the image generation model, causing the denoising network to predict the noise present in each frame feature map of the M frame feature maps based on the input data, thereby obtaining a predicted noise for each frame feature map;
[0181] The denoising network is controlled to perform denoising on the corresponding feature map based on the predicted noise of each frame feature map, and the denoised M frames feature map are used as the M frames image data generated by the n-th round of image processing.
[0182] In another specific implementation, when the processing unit 602 is used to perform multiple rounds of image processing based on the image vector sequence and the image generation model, it can be specifically used to:
[0183] Obtaining content description text, where the content description text is used to indicate a display form of the video content in the video frame sequence;
[0184] Generate a text vector sequence based on the content description text, the text vector sequence including M text vectors, each of which is used to represent the content description text;
[0185] Based on the text vector sequence, the image vector sequence and the image generation model, multiple rounds of image processing are performed; wherein any round of image processing is performed by the image generation model based on the text vector sequence and the currently input image vector sequence.
[0186] According to another embodiment of the present application, Figure 6 The various units in the video generating device shown can be respectively or entirely merged into one or several other units to constitute, or a certain (some) unit therein can also be split into multiple smaller units in function to constitute, which can achieve the same operation without affecting the realization of the technical effect of the embodiments of the application. The above-mentioned units are divided based on logical functions. In actual applications, the function of a unit can also be realized by multiple units, or the function of multiple units can be realized by one unit. In other embodiments of the application, other units can also be included based on the video generating device. In actual applications, these functions can also be assisted by other units to be realized, and can be realized by the collaboration of multiple units.
[0187] According to another embodiment of the present application, a computer program (including one or more instructions) capable of executing each step involved in any of the above method embodiments can be constructed by running the computer program (including one or more instructions) on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access memory (RAM), and a read-only memory (ROM). Figure 6The video generation device shown in the embodiment of the present application is used to implement the relevant methods proposed in the embodiment of the present application. The computer program can be recorded on a computer-readable storage medium, for example, and loaded into the above-mentioned computing device through the computer-readable storage medium and run therein.
[0188] It is worth noting that, in the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can include a part of the overall module or unit of the module or unit function.
[0189] In an embodiment of the present application, an image vector sequence can be generated by encoding a reference image for indicating the video content in a video frame sequence to be generated, and multiple rounds of image processing are performed based on the image vector sequence and an image generation model, thereby generating a video frame sequence based on M frames of image data generated by the last round of image processing. Each time M frames of image data are generated through a round of image processing, an inter-frame offset sequence can be obtained to indicate the difference between any two adjacent frames of image data in the currently generated M frames of image data, and at least one image vector in the image vector sequence can be dynamically adjusted based on the inter-frame offset sequence. The dynamically adjusted image vector sequence is then input into the image generation model for the next round of image processing. This allows the image generation model to understand the current image generation situation based on the dynamically adjusted image vector sequence during the next round of image processing, thereby deeply perceiving the generation direction of the video frame sequence, improving the generation performance of the entire model, and accurately achieving the purpose of generating a video frame sequence. Moreover, by dynamically adjusting the image vector sequence, it is also possible to avoid the situation where the information referred to by the image generation model remains unchanged during the entire video generation process, thereby avoiding the image generation model being restricted in the content of the final generated video frame sequence due to the unchanged reference information. While maintaining the video content information indicated by the reference image, the image generation model can exert its own diverse generation capabilities, making the video content of the final generated video frame sequence more colorful, thereby improving the video quality of the video frame sequence.
[0190] Based on the description of the above method embodiment and apparatus embodiment, the present application embodiment also provides a computer device. Figure 7, the computer device at least includes a processor 701, an input interface 702, an output interface 703 and a computer storage medium 704. Among them, the processor 701, input interface 702, output interface 703 and computer storage medium 704 in the computer device can be connected via a bus or other means. The computer storage medium 704 can be stored in the memory of the computer device, and the computer storage medium 704 is used to store a computer program, and the computer program includes one or more instructions, and the processor 701 is used to execute one or more instructions in the computer program stored in the computer storage medium 704. The processor 701 (or CPU (Central Processing Unit)) is the computing core and control core of the computer device, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to realize the various method processes or corresponding functions in the embodiments of the present application. For example, in one embodiment, the processor 701 described in the embodiment of the present application can be used to perform the aforementioned based on a reference image. Figure 2 or Figure 4 A series of video generation processes in the method flow shown.
[0191] The embodiment of the present application also provides a computer storage medium (Memory), which is a memory device in a computer device for storing computer programs and data. It is understandable that the computer storage medium here can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer storage medium provides a storage space that stores the operating system of the computer device. In addition, a computer program is also stored in the storage space, and the computer program includes one or more instructions suitable for being loaded and executed by the processor 701. It should be noted that the computer storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer storage medium located away from the aforementioned processor. In one embodiment, the processor can load and execute one or more instructions stored in the computer storage medium to implement the corresponding steps in the various method embodiments mentioned above.
[0192] In an embodiment of the present application, an image vector sequence can be generated by encoding a reference image for indicating the video content in a video frame sequence to be generated, and multiple rounds of image processing are performed based on the image vector sequence and an image generation model, thereby generating a video frame sequence based on M frames of image data generated by the last round of image processing. Each time M frames of image data are generated through a round of image processing, an inter-frame offset sequence can be obtained to indicate the difference between any two adjacent frames of image data in the currently generated M frames of image data, and at least one image vector in the image vector sequence can be dynamically adjusted based on the inter-frame offset sequence. The dynamically adjusted image vector sequence is then input into the image generation model for the next round of image processing. This allows the image generation model to understand the current image generation situation based on the dynamically adjusted image vector sequence during the next round of image processing, thereby deeply perceiving the generation direction of the video frame sequence, improving the generation performance of the entire model, and accurately achieving the purpose of generating a video frame sequence. Moreover, by dynamically adjusting the image vector sequence, it is also possible to avoid the situation where the information referred to by the image generation model remains unchanged during the entire video generation process, thereby avoiding the image generation model being restricted in the content of the final generated video frame sequence due to the unchanged reference information. While maintaining the video content information indicated by the reference image, the image generation model can exert its own diverse generation capabilities, making the video content of the final generated video frame sequence more colorful, thereby improving the video quality of the video frame sequence.
[0193] It should be noted that, according to one aspect of the present application, a computer program product or computer program is also provided, and the computer program product or computer program includes one or more instructions, and the one or more instructions are stored in a computer storage medium. The processor of the computer device reads the one or more instructions from the computer storage medium, and the processor executes the one or more instructions, so that the computer device performs the methods provided in various optional manners of the above-mentioned method embodiments. It should be understood that what is disclosed above is only a preferred embodiment of the present application, and it is certainly not used to limit the scope of the rights of the present application. Therefore, equivalent changes made in accordance with the claims of the present application are still within the scope covered by the present application.
Claims
1. A video generation method, characterized in that: include: Acquire a reference image, where the reference image is used to indicate video content in a video frame sequence to be generated, where the video frame sequence includes M video frames, where M is an integer greater than 1; Generate an image vector sequence based on the reference image encoding, the image vector sequence including M image vectors, each of which is used to represent the reference image; the m-th image vector in the image vector sequence is used to guide the generation of the m-th video frame in the video frame sequence, m∈[1,M]; Performing multiple rounds of image processing based on the image vector sequence and the image generation model, each round of image processing generating M frames of image data, the m-th frame of image data corresponding to the m-th video frame; wherein, during the multiple rounds of image processing, each time M frames of image data are generated after a round of image processing, obtaining an inter-frame offset sequence corresponding to the current M frames of image data, dynamically adjusting at least one image vector in the image vector sequence based on the inter-frame offset sequence, and inputting the dynamically adjusted image vector sequence into the image generation model for the next round of image processing; the inter-frame offset sequence is used to indicate: a difference between any two adjacent frames of image data in the corresponding M frames of image data; The video frame sequence is generated according to the M frames of image data generated by the last round of image processing in the multiple rounds of image processing.
2. The method according to claim 1, wherein Each frame of image data includes image values of multiple position points, and obtaining the inter-frame offset sequence corresponding to the current M frames of image data includes: Performing graph optimization processing on the current M frames of image data, wherein the graph optimization processing includes at least one of the following: eliminating identical image semantics at the same position point in the M frames of image data, and smoothing image values of each position point in each frame of image data; After the image optimization processing, performing difference processing on two adjacent frames of image data in the current M frames of image data in sequence to obtain M-1 inter-frame offset matrices; the difference processing includes: calculating the difference between two image values at the same position point in the two adjacent frames of image data; The M-1 inter-frame offset matrices are used to construct an inter-frame offset sequence corresponding to the currently generated M frames of image data.
3. The method according to claim 2, wherein Before performing the image optimization process, the image value of any position point in each frame of image data is used to represent the image semantics of the corresponding position point; When the image optimization processing includes eliminating identical image semantics at the same position in the M frames of image data, performing the image optimization processing on the current M frames of image data includes: Traverse the plurality of position points, and take the currently traversed position point as the i-th position point, i∈[1,G], where G is the number of position points; Obtaining an image value regularization parameter for the i-th position point, where the image value regularization parameter is obtained by statistically analyzing M image values at the i-th position point in the current M frames of image data; Based on the image value regularization parameters, the M image values at the i-th position point in the current M frames of image data are regularized; wherein, the regularization refers to: a process of adjusting the values according to a preset rule.
4. The method according to claim 3, wherein The statistical analysis includes mean value calculation; and obtaining the image value regularization parameter of the i-th position point includes: Obtain a first convolutional layer, where the first convolutional layer includes M-channel first convolution kernels, each of which has a size of 1×1, and the M-channel first convolution kernels are used to perform convolution mean calculation on M image values at the same position in the current M frames of image data; Calling the first convolution layer to perform convolution mean calculation on the current M frames of image data to obtain a first convolution map; wherein the first convolution map includes G first convolution values, and the i-th first convolution value is used to represent: the mean of the M image values at the i-th position in the current M frames of image data; From the first convolution map, obtain the i-th first convolution value as the image value regularization parameter of the i-th position point.
5. The method according to claim 3, wherein The image value regularization parameter includes: the mean value of the M image values at the i-th position point in the current M frames of image data; The step of regularizing the M image values at the i-th position point in the current M frames of image data based on the image value regularization parameter includes: The mean of the M image values at the i-th position point in the current M frames of image data is used as the mean of the image values corresponding to the i-th position point; Traverse the M image values at the i-th position in the current M frames of image data, and use the currently traversed image value as the m-th image value; Calculating a difference between the m-th image value and a mean of the image values corresponding to the i-th position point to obtain a first difference; The m-th image value is updated by using a ratio between the first difference and a mean of the image values corresponding to the i-th position point.
6. The method according to claim 3, wherein The image value regularization parameters include: the maximum image value and the minimum image value among the M image values at the i-th position point in the current M frames of image data; The step of regularizing the M image values at the i-th position point in the current M frames of image data based on the image value regularization parameter includes: Calculating a difference between a maximum image value and a minimum image value among the M image values at the i-th position point in the current M frames of image data to obtain a second difference value; For an mth image value among the M image values, a difference between the mth image value and the minimum image value is calculated, and the mth image value is updated by using a ratio between the calculated difference and the second difference.
7. The method according to claim 3, wherein After all the plurality of position points are traversed, the image optimization processing of the current M frames of image data further includes: Obtain a second convolutional layer, where the second convolutional layer includes second convolution kernels of M channels, each of which has a size of P×P, where P is an integer greater than 1; the second convolution kernel of the mth channel is used to perform convolution mean calculation on the mth frame of image data; Call the second convolutional layer to perform convolution mean calculation on the current M frames of image data to obtain M second convolution maps; each second convolution map includes J second convolution values, where J is a positive integer, and the j-th second convolution value in the m-th second convolution map is used to represent: the mean of each image value in the j-th region of size P×P in the m-th frame of image data, j∈[1,J]; Regularizing the j-th second convolution value in the M second convolution images to obtain M regularized second convolution values; In the m-th frame of image data included in the current M frames of image data, each image value in the j-th region with a size of P×P is updated to the m-th regularized second convolution value.
8. The method according to claim 2, wherein When the image optimization processing includes smoothing the image values of each position point in each frame of image data, performing the image optimization processing on the current M frames of image data includes: Obtain a third convolutional layer, where the third convolutional layer includes a third convolution kernel of one channel, and the size of the third convolution kernel is R×R, where R is an integer greater than 1; Poll the current M frames of image data to determine the currently polled m-th frame of image data; Calling the third convolutional layer to perform convolution mean calculation on the m-th frame image data to obtain T third convolution values, where T is a positive integer, and the t-th third convolution value is used to represent: the mean of each image value in the t-th region of size R×R in the m-th frame image data, t∈[1,T]; In the m-th frame image data, each image value in the t-th region with a size of R×R is updated to the t-th third convolution value.
9. The method according to claim 2, wherein The using the M-1 inter-frame offset matrices to construct an inter-frame offset sequence corresponding to the currently generated M frames of image data includes: Constructing an inter-frame offset matrix consisting of zero values, and adding the constructed inter-frame offset matrix in front of the M-1 inter-frame offset matrices to obtain M inter-frame offset matrices arranged in sequence; Performing statistical analysis on M values at the same position point in the M inter-frame offset matrices to obtain a numerical regularization parameter for each position point; Based on the numerical regularization parameters of each position point, the M numerical values at the corresponding position points in the M inter-frame offset matrices are regularized to obtain the inter-frame offset sequence corresponding to the currently generated M frames of image data; wherein, the regularization refers to: performing numerical adjustment processing according to preset rules.
10. The method according to claim 1, wherein The inter-frame offset sequence includes at least one inter-frame offset matrix, one inter-frame offset matrix corresponds to one frame of image data, and any inter-frame offset matrix is used to indicate the difference between the corresponding frame of image data and an adjacent frame of image data; The dynamically adjusting at least one image vector in the image vector sequence based on the inter-frame offset sequence comprises: Traversing each inter-frame offset matrix in the inter-frame offset sequence, and taking a frame of image data corresponding to the currently traversed inter-frame offset matrix as target frame image data; Extracting, from the image vector sequence, an image vector used to generate the target frame image data as a target image vector; The target image vector is dynamically adjusted using the currently traversed inter-frame offset matrix to obtain a dynamically adjusted target image vector.
11. The method according to claim 10, wherein The dynamically adjusting the target image vector by using the currently traversed inter-frame offset matrix to obtain the dynamically adjusted target image vector includes: Based on the size of the currently traversed inter-frame offset matrix, the size of the target image vector is transformed to obtain an intermediate image vector; Calling a cross attention network to dynamically adjust the intermediate image vector according to the currently traversed inter-frame offset matrix to obtain a dynamically adjusted intermediate image vector; Based on the original size of the target image vector, the size of the dynamically adjusted intermediate image vector is transformed to obtain the dynamically adjusted target image vector; wherein the original size refers to: the size of the target image vector before the size of the target image vector is transformed.
12. The method according to claim 1, wherein The image generation model includes a denoising network; the nth round of image processing includes: Obtaining M frames of feature maps for performing the n-th round of image processing, where n is a positive integer; when n=1, the M frames of feature maps are obtained by encoding M noise images, and one frame of feature map corresponds to one noise data; when n>1, the M frames of feature maps are M frames of image data generated by the n-1th round of image processing; Inputting the M frame feature maps and the current image vector sequence into a denoising network in the image generation model, causing the denoising network to predict the noise present in each frame feature map of the M frame feature maps based on the input data, thereby obtaining a predicted noise for each frame feature map; The denoising network is controlled to perform denoising on the corresponding feature map based on the predicted noise of each frame feature map, and the denoised M frames feature map are used as the M frames image data generated by the n-th round of image processing.
13. The method according to claim 1, wherein The performing multiple rounds of image processing based on the image vector sequence and the image generation model includes: Obtaining content description text, where the content description text is used to indicate a display form of the video content in the video frame sequence; Generate a text vector sequence based on the content description text, the text vector sequence including M text vectors, each of which is used to represent the content description text; Based on the text vector sequence, the image vector sequence and the image generation model, multiple rounds of image processing are performed; wherein any round of image processing is performed by the image generation model based on the text vector sequence and the currently input image vector sequence.
14. A video generating device, characterized in that: include: an acquisition unit, configured to acquire a reference image, wherein the reference image is used to indicate video content in a video frame sequence to be generated, wherein the video frame sequence includes M video frames, where M is an integer greater than 1; a processing unit, configured to generate an image vector sequence based on the reference image encoding, the image vector sequence comprising M image vectors, each of which is used to represent the reference image; the m-th image vector in the image vector sequence is used to guide the generation of the m-th video frame in the video frame sequence, m∈[1,M]; The processing unit is further configured to perform multiple rounds of image processing based on the image vector sequence and the image generation model, each round of image processing generating M frames of image data, the m-th frame of image data corresponding to the m-th video frame; wherein, during the multiple rounds of image processing, each time M frames of image data are generated after a round of image processing, an inter-frame offset sequence corresponding to the current M frames of image data is obtained, and at least one image vector in the image vector sequence is dynamically adjusted based on the inter-frame offset sequence, and the dynamically adjusted image vector sequence is input into the image generation model for the next round of image processing; the inter-frame offset sequence is used to indicate: a difference between any two adjacent frames of image data in the corresponding M frames of image data; The processing unit is further configured to generate the video frame sequence based on the M frames of image data generated by the last round of image processing in the multiple rounds of image processing.
15. A computer device comprising an input interface and an output interface, characterized in that: Also includes: processors and computer storage media; The processor is suitable for implementing one or more instructions, the computer storage medium stores one or more instructions, and the one or more instructions are suitable for being loaded by the processor and executing the video generation method according to any one of claims 1-13.
16. A computer storage medium, characterized in that The computer storage medium stores one or more instructions, and the one or more instructions are suitable for being loaded by a processor and executing the video generation method according to any one of claims 1 to 13.
17. A computer program product, characterized in that The computer program product includes one or more instructions; when the one or more instructions in the computer program are executed by a processor, the video generation method according to any one of claims 1 to 13 is implemented.
Citation Information
Patent Citations
Image-based video generation method and device, equipment and storage medium
CN117830483A
Image processing method and electronic equipment
CN118101856A