Video generation method, system, electronic device and storage medium
By adding noise to static images and predicting noise, and then combining this with a video generation model to generate a target video frame sequence, the problem of poor fidelity in video generation from static images is solved, achieving higher accuracy and versatility in video generation.
Patent Information
- Application Number
- CN202410176643.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-07
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2044-02-07
AI Technical Summary
Existing technologies suffer from poor fidelity when generating videos based on static images, especially for image generation from images outside of specific image domains, where the fidelity of the generated videos is insufficient, and existing methods lack universality.
By adding noise to the image to be processed, an initial video frame sequence is generated. Then, a pre-trained video generation model is used for noise prediction and denoising. Combined with video description information, a target video frame sequence is generated, and finally, video content matching the original image is generated.
It improves the fidelity and versatility of video generation, ensuring more accurate videos are generated without losing local image details.
Smart Images

Figure CN118714417B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and more specifically, to a method, system, electronic device, and storage medium for generating video. Background Technology
[0002] In today's digital media era, static images can be used to record users' lives. However, while static images can present necessary details and information about objects, they still lack interactivity and immersion. Based on this, video can be used as a more dynamic and intuitive medium. However, compared to creating static images, video creation is more expensive and technically challenging. Therefore, it is crucial to find a way to automatically generate corresponding videos from relatively low-cost static images—that is, how to make images "move" in a reasonable way.
[0003] In related technologies, the generation of videos based on static images mainly focuses on specific image domains, such as faces, handwritten letters, or simple doodles. Generative Adversarial Networks (GANs) or Variational Autoencoders (VAEs) can be used to process images in these domains. However, this approach is not only applicable to these specific image domains and lacks generality. Furthermore, the generated videos can only approximate the given image or retain a similar style, losing local details and exhibiting limited fidelity. Therefore, the technical problem of poor fidelity in image-generated videos persists.
[0004] There is currently no effective solution to the above problems. Summary of the Invention
[0005] This application provides a video generation method, system, electronic device, and storage medium to at least solve the technical problem of poor fidelity in video generated from images.
[0006] According to one aspect of the embodiments of this application, a video generation method is provided. The method may include: identifying an image to be processed; adding noise to the image to be processed using initial noise information to obtain an initial video frame sequence; inputting the initial video frame sequence into a video generation model, wherein the video generation model is trained based on video samples; using video description information to guide the video generation model to predict noise in the initial video frame sequence to obtain predicted noise information, and using the predicted noise information to denoise the initial video frame sequence to obtain a target video frame sequence, wherein the video description information is used to describe video content matching the image to be processed; and generating a target video based on the target video frame sequence, wherein the target video includes video content.
[0007] According to another aspect of the embodiments of this application, another video generation method is provided. This method can be applied to an e-commerce platform and may include: identifying an image to be processed, wherein the image content of the image to be processed includes a transaction object; adding noise to the image to be processed using initial noise information to obtain an initial video frame sequence; inputting the initial video frame sequence into a video generation model, wherein the video generation model is trained based on video samples; using video description information to guide the video generation model to predict noise in the initial video frame sequence to obtain predicted noise information, and using the predicted noise information to denoise the initial video frame sequence to obtain a target video frame sequence, wherein the video description information is used to describe video content matching the image to be processed; generating a motion-effect video based on the target video frame sequence, wherein the motion-effect video includes video content and is used to display the transaction object; and pushing the motion-effect video to a media platform for playback.
[0008] According to another aspect of the embodiments of this application, another method for generating video is provided. The method may include: displaying an image to be processed on the operation interface in response to an input command applied to the operation interface; and displaying a target video including video content matching the image to be processed on the operation interface in response to a video generation command applied to the operation interface, wherein the target video is generated based on a target video frame sequence, the target video frame sequence is obtained by a video generation model using predicted noise information to denoise an initial video frame sequence, the video generation model is trained based on video samples, the predicted noise information is obtained by using video description information to guide the video generation model to predict noise in the initial video frame sequence, the video description information is used to describe the video content, and the initial video frame sequence is obtained by adding noise to the image to be processed using initial noise information.
[0009] According to another aspect of the embodiments of this application, another method for generating video is provided. The method may include: displaying an image to be processed on the display screen of a virtual reality (VR) device or an augmented reality (AR) device; adding noise to the image to be processed using initial noise information to obtain an initial video frame sequence; inputting the initial video frame sequence into a video generation model in the VR or AR device, wherein the video generation model is trained based on video samples; using video description information to guide the video generation model to predict noise in the initial video frame sequence to obtain predicted noise information, and using the predicted noise information to denoise the initial video frame sequence to obtain a target video frame sequence, wherein the video description information is used to describe video content matching the image to be processed; generating a target video based on the target video frame sequence, wherein the target video includes video content; and driving the VR or AR device to display the target video.
[0010] According to another aspect of the embodiments of this application, a video generation system is provided. The system may include: an e-commerce platform for identifying an image to be processed; adding noise to the image to be processed using initial noise information to obtain an initial video frame sequence; inputting the initial video frame sequence into a video generation model, wherein the video generation model is trained based on video samples; using video description information to guide the video generation model to predict noise in the initial video frame sequence to obtain predicted noise information, and using the predicted noise information to denoise the initial video frame sequence to obtain a target video frame sequence, wherein the video description information is used to describe video content matching the image to be processed; generating a target video based on the target video frame sequence, wherein the target video includes video content; and a media platform for playing the target video.
[0011] According to another aspect of the embodiments of this application, an electronic device is also provided. The electronic device may include a memory and a processor: the memory stores computer-executable instructions, and the processor executes the computer-executable instructions, wherein when the processor executes the computer-executable instructions, it implements the video generation method of the embodiments of this application.
[0012] According to another aspect of the embodiments of this application, a processor is also provided. This processor is used to run a program, wherein the video generation method of the embodiments of this application is executed during program execution.
[0013] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided. This computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device where the storage medium is located to perform the video generation method described in the embodiments of this application.
[0014] According to another aspect of the embodiments of this application, a computer program product is also provided. This computer program product includes a computer program that, when executed by a processor, implements the video generation method described in the embodiments of this application.
[0015] In this embodiment, if it is necessary to generate a corresponding video based on a certain image, the image to be processed for video generation can be identified, and initial noise information can be used to add noise to the image to be processed, resulting in a noisy initial video frame sequence. The initial video frame sequence can be input into a video generation model, which is pre-trained based on video samples to process the initial video frame sequence. In the video generation model, matching video description information that describes the video content generated based on the image to be processed can be analyzed from the initial video frame sequence. Using the video description information, noise prediction is performed on the initial video frame sequence to obtain predicted noise information, which can then be used to denoise the initial video frame sequence, resulting in a target video frame sequence. Thus, a target video containing video content matching the current image can be generated based on the target video frame sequence. Since the embodiments of this application can add noise to the image to be processed to add more image details, thereby achieving the purpose of improving the accuracy of video generation by increasing image details, and can use a pre-trained video generation model to generate videos, it can use any image in the natural open domain, thereby also achieving the purpose of improving the versatility of video generation. Through the above method, without losing the local details of the image to be processed, video generation can be performed more accurately, thereby achieving the technical effect of improving the fidelity of image-based videos and solving the technical problem of poor fidelity of image-based videos.
[0016] It is worth noting that the general description above and the detailed description that follow are merely for illustrative purposes and do not constitute a limitation on this application. Attached Figure Description
[0017] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0018] Figure 1 This is a schematic diagram illustrating an application scenario of a video generation method according to an embodiment of this application;
[0019] Figure 2 This is a flowchart of a video generation method according to an embodiment of this application;
[0020] Figure 3This is a flowchart of another video generation method according to an embodiment of this application;
[0021] Figure 4 This is a flowchart of another video generation method according to an embodiment of this application;
[0022] Figure 5 This is a flowchart of another video generation method according to an embodiment of this application;
[0023] Figure 6 This is a flowchart of a video generation system according to an embodiment of this application;
[0024] Figure 7 This is a flowchart of a high-fidelity video motion effect generation method based on a diffusion model according to an embodiment of this application;
[0025] Figure 8 This is a schematic diagram illustrating a feature dimension conversion process using a video motion effect model according to an embodiment of this application;
[0026] Figure 9 This is a schematic diagram of a video motion effect generation process based on noise correction according to an embodiment of this application;
[0027] Figure 10(a) is a schematic diagram of an example of generating video from a graph according to an embodiment of this application;
[0028] Figure 10(b) is a schematic diagram of another example of generating video from a graph according to an embodiment of this application;
[0029] Figure 11 This is a schematic diagram of a video generation apparatus according to an embodiment of this application;
[0030] Figure 12 This is a schematic diagram of another video generation apparatus according to an embodiment of this application;
[0031] Figure 13 This is a schematic diagram of another video generation apparatus according to an embodiment of this application;
[0032] Figure 14 This is a schematic diagram of another video generation apparatus according to an embodiment of this application;
[0033] Figure 15 This is a structural block diagram of a computer terminal according to an embodiment of this application;
[0034] Figure 16 This is a block diagram of an electronic device according to an embodiment of the present application of a video generation method;
[0035] Figure 17This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a video generation method according to an embodiment of this application.
[0036] Figure 18 This is a structural block diagram of the computing environment for a video generation method according to an embodiment of this application. Detailed Implementation
[0037] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0038] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0039] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:
[0040] Artificial Intelligence Generated Content (AIGC) refers to content of various types, such as text, images, audio, or video, generated through artificial intelligence technology.
[0041] A diffusion model is a mathematical model used to simulate and predict the diffusion trend of data in space or events. It is a deep generative model that learns to eliminate noise through forward diffusion and backward denoising processes, and can generate high-quality data samples.
[0042] The Denoising Diffusion Implicit Model (DDIM) can accelerate the content generation process with fewer sampling steps by using non-Markovian ideas.
[0043] Variational autoencoders (VAEs) are generative models that learn the complex probability distribution of input data by encoding the input data into a latent representation space and by sampling to generate new data.
[0044] Example 1
[0045] According to an embodiment of this application, a method for generating video is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0046] The video generation method provided in this application embodiment can be applied to, for example, Figure 1 The application scenarios shown are not limited to these. In, for example... Figure 1 In the application scenario shown, the server 10 can be in the cloud. The server 10 can connect to one or more client devices 20 via a local area network (LAN), wide area network (WAN), internet connection, or other types of data network. These client devices 20 can include, but are not limited to, smartphones, tablets, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices. These client devices collectively constitute the client relative to the server. An interactive interface for obtaining file upload requests can be deployed on the graphical user interface of the client device. This interactive interface can be an operation interface for inputting images to be processed. The client device 20 can interact with the user through the graphical user interface to implement the video generation method provided in this embodiment.
[0047] In this embodiment, the system consisting of a client device and a server can perform the following steps: If the client needs to generate a video of a certain image, the client can input the image to be processed on the interactive interface of the client device and send it to the server via the network. For example, the image to be processed can be published on an e-commerce platform on the client device. After receiving the image to be processed, the e-commerce platform can perform the following steps on its server: Step S102, identify the image to be processed; Step S104, use initial noise information to add noise to the image to be processed to obtain an initial video frame sequence; Step S106, input the initial video frame sequence into the video generation model, wherein the video generation model is trained based on video samples; Step S108, use video description information to guide the video generation model to predict noise in the initial video frame sequence to obtain predicted noise information, and use the predicted noise information to denoise the initial video frame sequence to obtain a target video frame sequence, wherein the video description information is used to describe the video content that matches the image to be processed; Step S110, generate a target video based on the target video frame sequence, wherein the target video includes video content. During the above process, the target video generated on the server can be sent to the client device via the network. The target video can then be displayed on the client device's interface, for example, played on a media platform on the client device.
[0048] Because the embodiments of this application can add noise to the image to be processed, thereby increasing the image detail and improving the accuracy of video generation, and because a pre-trained video generation model can be used to generate videos, it can use any image in the natural open domain, thus improving the versatility of video generation. Through the above method, video generation can be performed more accurately without losing the local details of the image to be processed, thereby improving the fidelity of image-based videos and solving the technical problem of poor fidelity in image-based videos.
[0049] This application provides the following method from a technical implementation perspective. In the aforementioned application scenarios, this application provides the following... Figure 2 The method for generating the video shown. Figure 2 This is a flowchart of a video generation method according to an embodiment of this application, such as... Figure 2 As shown, the method may include the following steps:
[0050] Step S202: Identify the image to be processed.
[0051] In the technical solution provided in step S202 of this application, the image to be processed can be a given image used to generate the video, also referred to as a picture or input image. If the video generation scenario is an e-commerce scenario, the image to be processed can be a picture including products, and the user can be a merchant; no specific restrictions are made here.
[0052] In this embodiment, the image to be processed can be identified. Optionally, the user can take a picture according to their own needs. If a video needs to be generated based on the picture, the picture can be used as the image to be processed and uploaded to the server, where the server can identify the image to be processed.
[0053] For example, if a merchant needs to obtain a product video corresponding to a specific product image, they can upload that product image to an e-commerce platform. For instance, the merchant can input the product image into the e-commerce platform's interface on their client device. When the e-commerce platform detects the product image, it can upload it to a server associated with the platform. The server then performs image recognition on the product.
[0054] It should be noted that the above-described use cases for generating videos are merely illustrative examples. Other use cases could include specialized image editing or video editing applications. No specific limitations are imposed here. Any scenario that requires video animation effects falls within the protection scope of this application's embodiments.
[0055] Step S204: Using the initial noise information, noise is added to the image to be processed to obtain the initial video frame sequence.
[0056] In the technical solution provided in step S204 of this application, the initial noise information can be the video noise obtained by initial sampling of the image to be processed, that is, the real noise, which can also be called initial sampling noise or video noise. The initial video frame sequence can be all the frames of the video within a certain period of time obtained by noise addition processing, which can also be called video frames.
[0057] In this embodiment, after identifying the image to be processed, the initial noise information can be used to add noise to the image to be processed, thereby obtaining an initial video frame sequence.
[0058] Optionally, before video generation, initial sampling noise can be obtained by sampling random noise that follows a Gaussian distribution. That is, ground truth noise can be initially sampled from random noise. The large amount of ground truth noise is then summarized into initial noise information. The sampled noise and the size of the video frame after VAE encoding are consistent. In other words, the sampled initial noise information has the same length and width as the encoded video frame.
[0059] Optionally, the image to be processed can be input into a VAE encoder for encoding. The encoded image to be processed and the initial noise information can be combined to perform padding processing. That is, the encoded image to be processed is padded into the initial noise information to generate an initial video frame sequence that is consistent with the overall style and layout of the image to be processed.
[0060] For example, if the initial noise information consists of L real noises obtained from the video frames, then padding the initially encoded image to be processed—that is, adding L real noises to the encoded image to obtain a noisy image—can then be used to form the initial video frame sequence of the image to be processed. It should be noted that the number of real noises and the number of images included in the generated initial video frame sequence are merely illustrative examples and are not subject to specific limitations; they can be adjusted according to the actual video generation situation.
[0061] In this embodiment, since the initial noise information can be used to add noise to the image to be processed, more image details are added to the image to be processed. That is, noise is introduced into the encoded image information by padding, which increases the image details and improves the accuracy of the subsequently generated video. In this way, the technical effect of high fidelity of the video generated based on the image can be achieved.
[0062] Step S206: Input the initial video frame sequence into the video generation model.
[0063] In the technical solution provided in step S206 of this application, the video generation model is trained based on video samples and can be a diffusion model, also known as a video motion effect model. The diffusion model can incorporate a dilated 3D U-Net network structure (Dilated 3D U-Net, abbreviated as 3D U-Net), thus it can be called a 3D U-Net model. The 3D U-Net model can include a spatial module and a temporal module. The spatial module treats the input initial video frame sequence as multiple batches of image features, performing convolution and attention operations in the spatial dimension to spatially encode and extract features from the initial video frame sequence. The temporal module can transform the dimension of the initial video frame sequence, performing attention operations in the temporal dimension to learn the temporal correlation of the initial video frame sequence. The video samples can be videos input to the untrained video generation model.
[0064] In this embodiment, after adding noise to the image to be processed using the initial noise information to obtain an initial video frame sequence, the initial video frame sequence can be input into the video generation model. The video generation model can then be used to process the initial video frame sequence to obtain the corresponding target video.
[0065] Since the above method only processes specific images using CAN and VAE to generate corresponding videos, it lacks universality, and the generated videos can only approximate the given image or retain a similar image style, resulting in the loss of local details. Therefore, the technical problem of low fidelity in image-generated videos persists. However, in the embodiments of this application, noise can be added to the image to be processed to obtain an initial video frame sequence. In this way, more image details can be added to the image to be processed. A video generation model can be used to predict and denoise the initial video frame sequence, thereby achieving the technical effect of improving the fidelity of image-generated videos.
[0066] For example, during the training of the video generation model, the input video samples can be subjected to a forward noise addition process and then fed into a 3D U-Net model. The 3D U-Net model is then used to predict the noise in the video samples. If the predicted noise is accurate, the 3D U-Net model can be identified as the video generation model in this embodiment.
[0067] Step S208: Using video description information, guide the video generation model to predict noise in the initial video frame sequence to obtain predicted noise information, and use the predicted noise information to denoise the initial video frame sequence to obtain the target video frame sequence.
[0068] In the technical solution provided in step S208 of this application, the video description information can be used to describe the video content matching the image to be processed, also known as video description, which can provide more descriptive guidance for the image to be processed. The video content can be the motion effect video content corresponding to the image to be processed. The prediction noise information can be the prediction noise obtained by using a video generation model to predict the initial video frame sequence. The target video frame sequence can be the video obtained by using a video generation model to denoise the prediction noise information in the initial video frame sequence. The denoising process can be backward denoising.
[0069] In this embodiment, after inputting the initial video frame sequence into the video generation model, video description information can be used to guide the model to predict noise in the initial video frame sequence, thus obtaining predicted noise information. The initial video frame sequence can then be denoised to remove the predicted noise information, resulting in the target video frame sequence. This denoising process can be an iterative DDIM denoising process.
[0070] Optionally, the video description information can be input into a Contrastive Language-Image Pretraining (CLIP) encoder for encoding to obtain the encoded video description information. This encoded video description information can then be input into a video generation model.
[0071] Optionally, after the video generation model receives the encoded video description information and the initial video frame sequence, it can use the video generation model to perform noise prediction on the initial video frame sequence to obtain the predicted noise information.
[0072] Optionally, the denoising iterative process using DDIM can be achieved by predicting the noise at each step of the initial video frame sequence using a video generation model, and gradually eliminating the noise to generate the final target video frame sequence. It should be noted that the above process and method for obtaining predicted noise information through noise prediction are merely illustrative examples and are not intended to impose specific limitations.
[0073] Step S210: Generate the target video based on the target video frame sequence.
[0074] In the technical solution provided by step S210 of this application, the target video may include video content.
[0075] In this embodiment, after using video description information to guide the video generation model to predict noise in the initial video frame sequence, obtain predicted noise information, and then using the predicted noise information to denoise the initial video frame sequence to obtain the target video frame sequence, the target video corresponding to the target video frame sequence can be generated.
[0076] Optionally, after obtaining the target video frame sequence, the target video frame sequence can be input into the VAE decoder, and the VAE decoder can be used to decode the target video to obtain the target video formed by the image motion effect to be processed.
[0077] In this embodiment, to provide the capability of generating video motion effects, a video generation model can be employed. This model receives information from two parts. To improve the fidelity of the generated video, noise can be added to the encoded initial video frame sequence using a padding method. This initial video frame sequence is then input into the video generation model as part of the information to enhance image details in the image to be processed. The other part is encoded video description information. Both parts of information are input into the video generation model. Noise correction is performed during the backward denoising process in the video generation model. By reducing the accumulated error caused by the predicted noise information, the target video is formed. This method requires no additional training and can achieve the goal of generating video motion effects from the image to be processed through simple padding and noise correction, thereby achieving the technical effect of improving the fidelity of the generated video.
[0078] Through steps S202 to S210 of this application, if it is necessary to generate a corresponding video based on a certain image, the image to be processed for video generation can be identified, and initial noise information can be used to add noise to the image to be processed, resulting in a noisy initial video frame sequence. The initial video frame sequence can be input into a video generation model, which is pre-trained based on video samples to process the initial video frame sequence. In the video generation model, matching video description information that describes the video content generated based on the image to be processed can be analyzed from the initial video frame sequence. Using the video description information, noise prediction is performed on the initial video frame sequence to obtain predicted noise information, which can then be used to denoise the initial video frame sequence, resulting in a target video frame sequence. Therefore, a target video containing video content matching the current image can be generated based on the target video frame sequence. Since the embodiments of this application can add noise to the image to be processed to add more image details, thereby achieving the purpose of improving the accuracy of video generation by increasing image details, and can use a pre-trained video generation model to generate videos, it can use any image in the natural open domain, thereby also achieving the purpose of improving the versatility of video generation. Through the above method, without losing the local details of the image to be processed, video generation can be performed more accurately, thereby achieving the technical effect of improving the fidelity of image-based videos and solving the technical problem of poor fidelity of image-based videos.
[0079] The method described in this embodiment will be further described below.
[0080] As an optional implementation, step S204, using the initial noise information to add noise to the image to be processed to obtain an initial video frame sequence, includes: extracting image features from the image to be processed that have the same data dimension as the initial noise information; and using the initial noise information to add noise to the image features to obtain the features of the initial video frame sequence.
[0081] In this embodiment, after adding noise to the image to be processed using the initial noise information to obtain an initial video frame sequence, image features with the same data dimension as the initial noise information can be extracted from the image to be processed. Then, the initial noise information is used to add noise to these image features to obtain the features of the initial video frame sequence. These image features can be latent features in the image, also known as video latent space features. The features of the initial video frame sequence can be noisy latent space features.
[0082] In this embodiment, since a diffusion model is used as the baseline model for video motion effect generation, latent space features can be used to improve encoding efficiency and save model space. For the initial video frame sequence that needs to be input into the diffusion model, it can be encoded into the features of the initial video frame sequence, i.e., video latent space features, through a VAE encoder.
[0083] Optionally, for a given image I 0 And the initial sampled video noise, that is, the real noise n 0:L-1 A VAE encoder can be used to encode a given image to obtain its image features, i.e., its latent features z. 0 Dimensional broadcasting can maintain the same data dimension z as the initial sampling noise. 0:L-1 That is, the latent features z of the image can be... 0 The data dimension is adjusted to be the same as the data dimension of the initial noise information. The encoded image latent feature padding can be added to the video noise, that is, the initial noise information is used to add noise to the image features, where L can be used to represent the number of frames of the generated video.
[0084] As an optional implementation, the image features are noise-added using the initial noise information to obtain the features of the initial video frame sequence, including: using the initial noise information to perform noise-adding processing on the image features at multiple time steps to obtain the features of the initial video frame sequence corresponding to the multiple time steps respectively.
[0085] In this embodiment, during the process of adding noise to the image features using the initial noise information to obtain the features of the initial video frame sequence, the initial noise information can be used to add noise to the image features at multiple time steps to obtain the features of the initial video frame sequence corresponding to the multiple time steps. The multiple time steps can be T steps, which is only an example and is not a specific limitation.
[0086] Alternatively, during the padding process, a forward noise addition process using a diffusion model can be employed, using initial noise information n. 0:L-1 For image features z 0:L-1 Adding noise, for example, by adding noise for T steps, can yield noisy latent space features, that is, features of the initial video frame sequence. By using the above method, it can be ensured that the overall style and layout of the final generated video are consistent with the given image to be processed.
[0087] As an optional implementation, step S208, where multiple time steps include a first time step and a second time step, the second time step being the previous time step of the first time step, involves using video description information to guide the video generation model to predict noise in the initial video frame sequence to obtain predicted noise information, and using the predicted noise information to denoise the initial video frame sequence to obtain the target video frame sequence. This includes: a processing step, using video description information to guide the video generation model to predict noise in the features of the initial video frame sequence corresponding to the first time step to obtain predicted noise information corresponding to the first time step, and using the predicted noise information corresponding to the first time step to denoise the features of the initial video frame sequence corresponding to the first time step to obtain the features of the target video frame sequence corresponding to the second time step; and a determination step, where if the second time step has a previous time step among the multiple time steps, the previous time step among the multiple time steps is determined as the second time step, the features of the target video frame sequence corresponding to the second time step are determined as the features of the initial video frame sequence corresponding to the first time step, and the processing step is returned to be executed until the second time step is the first time step among the multiple time steps.
[0088] In this embodiment, during the process of using video description information to guide the video generation model to perform noise prediction and denoising on the initial video frame sequence, processing steps and determination steps can be executed. The multiple time steps may include a first time step and a second time step. The second time step can be the time step preceding the first time step; for example, if the first time step is t, then the second time step can be t-1.
[0089] Optionally, in the processing steps, video description information can be used to guide the video generation model to predict noise in the features of the initial video frame sequence corresponding to the first time step, thereby obtaining the predicted noise information corresponding to the first time step. The predicted noise information corresponding to the first time step can then be used to remove noise from the features of the initial video frame sequence corresponding to the first time step, resulting in the features of the target video frame sequence for the second time step.
[0090] Optionally, in the determination step, it can be determined whether a previous time step exists for the second time step. If it does, the previous time step can be determined as the second time step, and the features of the target video frame sequence corresponding to the second time step can be determined as the features of the initial video frame sequence corresponding to the first time step. The process can then return to the processing step. Alternatively, it can be detected in real-time whether the second time step has a previous time step. If not, it indicates that the second time step can be the first time step, and in this case, there is no need to return to the processing step.
[0091] Since perfect loss cannot be achieved during video generation model training, meaning there will always be some bias when using the video generation model for noise prediction, leading to a decrease in video fidelity, this embodiment addresses this issue. To mitigate error accumulation during denoising, the predicted noise information at each time step can be corrected using initial noise information. Specifically, the difference between the predicted noise information and the initial noise information can be calculated, and a weighted calculation method can be used to appropriately adjust the predicted noise information.
[0092] Optionally, during the DDIM noise removal process, if we take the input of the video generation model at time step t, that is, the noisy latent space features, it can be represented as: The noisy latent space features mentioned above can be input into the dilated 3D U-Net model to obtain the predicted noise information. The initial noise information at this point can be n 0:L-1 Based on the predicted noise information and the initial noise information mentioned above, noise correction can be performed.
[0093] Optionally, during the noise correction process, predicted noise information can be calculated. and initial noise information n 0:L-1 The differences between them.
[0094] Optionally, the difference between the two noise information determined above indicates that there is a certain deviation in the predicted noise information. The predicted noise information can be appropriately adjusted according to the difference to make the predicted noise information more accurate. That is, the predicted noise information can be corrected to obtain the corrected predicted noise information.
[0095] Optionally, the video generation model can use the corrected predicted noise information to perform the DDIM denoising process to obtain the latent space features at time t-1, that is, to obtain the features of the initial video frame sequence of the previous time step of the current time step.
[0096] In the embodiments of this application, the noise error of the first frame in the initial video frame sequence can be eliminated by the above method. That is, the prediction noise information of the first frame in the initial video frame sequence is corrected, so that the fidelity of the first frame can be accurately improved. At the same time, the other frames will also maintain the temporal content consistency with the first frame, thereby achieving the technical effect of improving the accuracy of the video generated from the image.
[0097] As an optional implementation, step S210, generating a target video based on the target video frame sequence, includes: if the second time step is the first time step of multiple time steps, then decoding the features of the target video frame sequence corresponding to the second time step to obtain the target video.
[0098] In this embodiment, during the process of generating a target video based on a target video frame sequence, it can be determined whether the second time step is the first time step among multiple time steps. If so, the features of the target video frame sequence corresponding to the second time step can be decoded to obtain the corresponding target video.
[0099] Optionally, after noise correction is completed for the predicted noise information at each time step, that is, after the second time step of the final denoising is the first time step, the features of the final target video frame obtained after denoising can be decoded, that is, the latent space features of the final video frame obtained after denoising can be decoded to generate a target video.
[0100] For example, the latent space features of the final video frame obtained after denoising are decoded using VAE to generate a motion-effect video. It should be noted that the decoding process described above is only an example and is not intended to impose any specific limitations.
[0101] As an optional implementation, step S208, which involves denoising the initial video frame sequence using the predicted noise information to obtain the target video frame sequence, includes: adjusting the predicted noise information using the initial noise information to obtain target noise information; and denoising the features of the initial video frame sequence using the target noise information to obtain the features of the target video frame sequence.
[0102] In this embodiment, during the process of denoising the initial video frame sequence using predicted noise information to obtain the target video frame sequence, the initial noise information can be used to adjust the predicted noise information to obtain the corresponding target noise information. The features of the initial video frame sequence can be denoised using the target noise information to obtain the features of the corresponding target video frame sequence. The target noise information can be the noise information obtained after adjusting the predicted noise information.
[0103] Optionally, after determining the initial noise information n 0:L-1 and predicted noise information Subsequently, the predicted noise information can be adjusted using the initial noise information. That is, when the accuracy of the predicted noise information predicted by the video generation model is low, the initial noise information can be used to correct the noise in the predicted noise information. Thus, the more accurate target noise information after correction can be used to perform a denoising process on the features of the initial video frame sequence to obtain the features in the final target video frame sequence.
[0104] As an optional implementation, adjusting the predicted noise information using the initial noise information to obtain the target noise information includes: acquiring the difference noise information between the predicted noise information and the initial noise information; and adjusting the predicted noise information using the difference noise information to obtain the target noise information.
[0105] In this embodiment, during the process of adjusting the predicted noise information using the initial noise information to obtain the target noise information, the difference noise information between the predicted noise information and the initial noise information can be determined. This difference noise information is then used to adjust the predicted noise information to obtain the target noise information. The difference noise information can represent the difference between the predicted noise information and the initial noise information. The target noise information can be the corrected predicted noise.
[0106] Optionally, after determining the initial noise information and predicted noise information The differential noise information can then be determined using the following formula:
[0107]
[0108] in, It can be used to represent the difference noise information at time step t; It can be used to represent the prediction noise information at time step t.
[0109] Optionally, the predicted noise can be appropriately adjusted using the aforementioned differential noise information. For example, a weighted approach can be used to adjust the predicted noise based on the differential noise information to obtain the adjusted target noise information.
[0110] As an optional implementation, the predicted noise information includes predicted noise information corresponding to different video frames in the initial video frame sequence, and the initial noise information includes initial noise information corresponding to different video frames. The different video frames are temporally correlated. The step of obtaining the difference noise information between the predicted noise information and the initial noise information includes: obtaining the first difference noise information between the predicted noise information corresponding to the first video frame in the different video frames and the corresponding initial noise information; and obtaining the second difference noise information between the predicted noise information corresponding to the target video frame in the different video frames and the corresponding initial noise information. The target video frame is any video frame other than the first video frame in the different video frames.
[0111] In this embodiment, during the process of obtaining the difference noise information between the predicted noise information and the initial noise information, the first difference noise information between the predicted noise information corresponding to the first video frame in different video frames and the corresponding initial noise information can be obtained, or the second difference noise information between the predicted noise information corresponding to the target video frame in different video frames and the corresponding initial noise information can be obtained.
[0112] Optionally, the predicted noise information may include predicted noise information corresponding to different video frames in the initial video frame sequence. The initial noise information may include initial noise information corresponding to different video frames, which are temporally correlated. The target video frame can be any video frame other than the first video frame. The first difference noise information can be the noise difference between the first video frame and the initial noise information, i.e., the difference of the first frame. The second difference noise can be the noise difference between the predicted noise information and the initial noise information of the remaining video frames (excluding the first video frame), i.e., the differences of the remaining frames.
[0113] Optionally, the difference between the predicted noise information and the initial noise information of the first video frame in the initial video frame sequence, i.e., the difference of the first frame, can be determined.
[0114] Optionally, the difference between the predicted noise information from other frames in the initial video frame sequence and the initial noise information, i.e., the difference in the target video frame.
[0115] As an optional implementation, the predicted noise information is adjusted using the differential noise information to obtain the target noise information, including: weighting the first differential noise information using a first weight and weighting the second differential noise information using a second weight, wherein the second weight is negatively correlated with the first weight; and adjusting the predicted noise information corresponding to the target video frame using the weighted first differential noise information and the weighted second differential noise information to obtain the target noise information corresponding to the target video frame.
[0116] In this embodiment, during the adjustment of predicted noise information using differential noise information, the first differential noise information can be weighted using a first weight, and the second differential noise information can be weighted using a second weight. The weighted first differential noise information can be used to adjust the predicted noise information corresponding to the target video frame, thereby obtaining the target noise information corresponding to the target video frame. The first weight can be used as a correction adjustment coefficient for the first differential noise information. The second weight is negatively correlated with the first weight and can be one minus the correction adjustment coefficient corresponding to the first differential noise information.
[0117] Optionally, the predicted noise information can be appropriately adjusted using a weighted approach.
[0118] For example, target noise information can be determined using the following formula:
[0119]
[0120] in, It can be used to represent target noise information; ω 0:L-1 It can contain L weight values, ranging from 0 to 1, which represent the correction adjustment coefficients for each frame, i.e., the first weight; It can be used to represent the difference of taking only the first frame. Then a copy operation is performed to expand the dimension to L frames.
[0121] Optionally, based on the target noise information determined above, a DDIM denoising process can be performed using a video generation model to obtain the latent space features at time t-1.
[0122] Optionally, the noise error in the first frame can be eliminated using the above method, ensuring the first frame maintains fidelity. Simultaneously, the remaining frames will also maintain temporal consistency with the first frame. The latent space features of the final video frame obtained after denoising are then analyzed. VAE decoding is performed to generate a motion-effect video.
[0123] As an optional implementation, the method further includes adjusting the first weight or the second weight.
[0124] In this embodiment, either the first weight or the second weight can be adjusted.
[0125] In this embodiment of the application, the correction adjustment coefficient ω of different frames during the noise correction process is adjusted. 0:L-1 This allows for the adjustment and trade-off between the motion effects and fidelity of the generated video, thereby improving the technical effect of the generated video's fidelity.
[0126] Optionally, if the correction adjustment coefficient ω of the i-th frame is... i A larger adjustment indicates that the correction process of the i-th frame tends to reference the noise difference of the first frame more, resulting in higher fidelity but relatively weaker motion effects. Conversely, if the correction adjustment coefficient ω of the i-th frame is increased... i A smaller adjustment indicates that the correction process for the i-th frame tends to reference its own noise differences more, resulting in higher motion fidelity but relatively lower fidelity. Therefore, based on the above analysis, you can determine whether you prioritize fidelity or motion fidelity based on your own needs, and adjust the corresponding correction adjustment coefficient accordingly.
[0127] As an optional implementation, the predicted noise information is adjusted using the differential noise information to obtain the target noise information, including: adjusting the predicted noise information corresponding to the first video frame using the first differential noise information to obtain the target noise information corresponding to the first video frame.
[0128] In this embodiment, during the process of adjusting the predicted noise information using the difference noise information to obtain the target noise information, the first difference noise information can be used to adjust the predicted noise information corresponding to the first video frame to obtain the target noise information corresponding to the first video.
[0129] Optionally, the prediction noise information corresponding to the first video frame can be corrected using the first difference noise information of the first video frame. During the correction process, the noise error of the first video frame can be eliminated, thereby ensuring the fidelity of the first frame and ensuring that the remaining frames are consistent with the first frame in terms of temporal content.
[0130] As an optional implementation, the method further includes: determining the data dimension of noise information based on the number of video frames corresponding to the target video to be generated.
[0131] In this embodiment, the data dimension of noise information can be determined based on the number of video frames corresponding to the target video to be generated.
[0132] Optionally, considering the characteristics of the noise information in the latent space at time step t of the DDIM denoising process. The dimension is Where B can be used to represent the number of videos in a batch; L can be used to represent the number of video frames; H′ and W′ can be used to represent the size of the latent space feature; and C can be used to represent the number of latent space feature channels.
[0133] Optionally, for a dilated 3D U-Net module, the batch dimensions and video frame dimensions can be merged before being input into the spatial module, treating the video frames as multiple batches of images, i.e., the dimensions are converted. When passing through the spatial module, the convolutional and attention modules operate on the spatial dimension H′×W′, and the dimension of the output features remains the same. Restore the dimension to
[0134] Optionally, before inputting into the time-series module, the dimensions of the batch processing and the spatial dimensions of the video are merged, i.e., the dimensions are converted. When passing through the temporal module, attention operations are performed on the temporal dimension L, and the dimension of the output features remains the same. Restore the dimension to Similarly, after processing by multiple downsampling and upsampling dilated 3D U-Net modules, the model obtains the noise features at time step t-1.
[0135] This application also provides a method for generating videos, which can be applied to e-commerce platforms. Figure 3 This is a flowchart of a video generation method according to an embodiment of this application, such as... Figure 3 As shown, the method may include the following steps:
[0136] Step S302: Identify the image to be processed, wherein the image content of the image to be processed includes the object to be traded.
[0137] In the technical solution provided by step S302 of this application, the image to be processed can be identified on the e-commerce platform, wherein the image content of the image to be processed may include the object to be traded.
[0138] Optionally, merchants can take relevant images of their potential customers based on their needs. If a video needs to be generated from this image, it can be used as the image to be processed and uploaded to the e-commerce platform. The e-commerce platform's server can then be used to perform recognition on the image to be processed.
[0139] Step S304: Using the initial noise information, noise is added to the image to be processed to obtain the initial video frame sequence.
[0140] In the technical solution provided by step S304 of this application, after identifying the image to be processed, the initial noise information can be used to add noise to the image to be processed to obtain an initial video frame sequence.
[0141] Optionally, before video generation, initial sampling noise can be obtained by sampling random noise that follows a Gaussian distribution. That is, real noise can be initially sampled from random noise, and the above-mentioned large amount of real noise can be summarized into initial noise information.
[0142] Optionally, the image to be processed can be input into a VAE encoder for encoding. The encoded image to be processed and the initial noise information can be combined to perform padding processing. That is, the encoded image to be processed is padded into the initial noise information to generate an initial video frame sequence that is consistent with the overall style and layout of the image to be processed.
[0143] Step S306: Input the initial video frame sequence into the video generation model, wherein the video generation model is trained based on video samples.
[0144] In the technical solution provided in step S306 of this application, after adding noise to the image to be processed using the initial noise information to obtain the initial video frame sequence, the initial video frame sequence can be input into the video generation model. The video generation model can then be used to process the initial video frame sequence to obtain the corresponding target video.
[0145] In this embodiment, the image to be processed can be noise-added to obtain an initial video frame sequence. In this way, more image details can be added to the image to be processed. A video generation model can be used to perform noise prediction and denoising on the initial video frame sequence, thereby achieving the technical effect of improving the fidelity of the image-generated video.
[0146] Step S308: Using video description information, guide the video generation model to predict noise in the initial video frame sequence to obtain predicted noise information, and use the predicted noise information to denoise the initial video frame sequence to obtain the target video frame sequence. The video description information is used to describe the video content that matches the image to be processed.
[0147] In the technical solution provided in step S308 of this application, after the initial video frame sequence is input into the video generation model, the video description information can be used to guide the video generation model to predict noise in the initial video frame sequence, thereby obtaining predicted noise information. The initial video frame sequence can then be denoised to remove the predicted noise information, thus obtaining the target video frame sequence.
[0148] Optionally, the video description information can be input into a CLIP encoder for encoding to obtain encoded video description information. The encoded video description information can then be input into a video generation model.
[0149] Optionally, after the video generation model receives the encoded video description information and the initial video frame sequence, it can use the video generation model to predict noise in the initial video frame sequence, thereby obtaining predicted noise information. The denoising iterative process using DDIM allows the video generation model to predict the noise at each step of the initial video frame sequence and gradually eliminate the noise to generate the final target video frame sequence.
[0150] Step S310: Generate a motion effect video based on the target video frame sequence, wherein the motion effect video includes video content and is used to display the transaction object.
[0151] In the technical solution provided by step S310 of this application, after using video description information to guide the video generation model to perform noise prediction on the initial video frame sequence, obtain predicted noise information, and use the predicted noise information to perform denoising processing on the initial video frame sequence to obtain the target video frame sequence, the motion effect video corresponding to the target video frame sequence can be generated.
[0152] Optionally, after obtaining the target video frame sequence, the target video frame sequence can be input into the VAE decoder, and the VAE decoder can be used to decode the target video frame sequence to obtain the motion effect video formed by the image motion effect to be processed, so as to display the transaction object through motion effects.
[0153] Step S312: Push the motion effect video to the media platform for playback.
[0154] In the technical solution provided in step S312 of this application, the generated motion effect video can be transmitted to a corresponding media platform. The motion effect video can then be played on the media platform.
[0155] Through steps S302 to S312 of this application, an image to be processed is identified, wherein the image content of the image to be processed includes the object to be traded; noise is added to the image to be processed using initial noise information to obtain an initial video frame sequence; the initial video frame sequence is input into a video generation model, wherein the video generation model is trained based on video samples; video description information is used to guide the video generation model to predict noise in the initial video frame sequence to obtain predicted noise information, and the predicted noise information is used to denoise the initial video frame sequence to obtain a target video frame sequence, wherein the video description information is used to describe the video content that matches the image to be processed; a motion effect video is generated based on the target video frame sequence, wherein the motion effect video includes video content and is used to display the object to be traded; the motion effect video is pushed to a media platform for playback, thereby achieving the technical effect of improving the fidelity of the video generated based on the image and solving the technical problem of poor fidelity of the video generated based on the image.
[0156] This application also provides a method for generating video. Figure 4 This is a flowchart of a video generation method according to an embodiment of this application, such as... Figure 4 As shown, the method may include the following steps:
[0157] Step S402: In response to the input command applied to the operation interface, display the image to be processed on the operation interface.
[0158] In the technical solution provided by step S402 of this application, according to the user's needs for video generation, the operation of inputting the image to be processed can be performed on the operation interface to form a corresponding input command to input the image to be processed into the operation interface. The image to be processed can also be displayed on the operation interface.
[0159] Step S404: In response to the video generation command applied to the operation interface, a target video including video content matching the image to be processed is displayed on the operation interface. The target video is generated based on a target video frame sequence. The target video frame sequence is obtained by the video generation model using predicted noise information to denoise the initial video frame sequence. The video generation model is trained based on video samples. The predicted noise information is obtained by using video description information to guide the video generation model to predict noise in the initial video frame sequence. The video description information is used to describe the video content. The initial video frame sequence is obtained by adding noise to the image to be processed using initial noise information.
[0160] In the technical solution provided in step S404 of this application, after obtaining the image to be processed through the operation interface, a corresponding video generation operation can be performed on the operation interface to form a corresponding video generation instruction. The generated target video can be displayed on the operation interface. The target video can be generated based on a target video frame sequence. The target video frame sequence can be obtained by the video generation model using predicted noise information to denoise the initial video frame sequence. The video generation model can be trained based on video samples. The predicted noise information can be obtained by using video description information to guide the video generation model to predict noise in the initial video frame sequence. The video description information can be used to describe the video content. The initial video frame sequence can be obtained by adding noise to the image to be processed using initial noise information.
[0161] Optionally, the image to be processed can be input into a VAE encoder for encoding. The encoded image to be processed and the initial noise information can be combined to perform padding processing. That is, the encoded image to be processed is padded into the initial noise information to generate an initial video frame sequence that is consistent with the overall style and layout of the image to be processed.
[0162] Optionally, the initial video frame sequence can be input into the video generation model, which can then be used to process the initial video frame sequence to obtain the corresponding target video.
[0163] Optionally, noise can be added to the image to be processed to obtain an initial video frame sequence. In this way, more image details can be added to the image to be processed. A video generation model can be used to perform noise prediction and denoising on the initial video frame sequence, thereby achieving the technical effect of improving the fidelity of the image-generated video.
[0164] Optionally, the video description information can be input into a CLIP encoder for encoding to obtain encoded video description information. The encoded video description information can then be input into a video generation model. After the video generation model receives the encoded video description information and the initial video frame sequence, it can use the video generation model to predict noise in the initial video frame sequence to obtain predicted noise information.
[0165] Optionally, the denoising iterative process using DDIM can predict the noise at each step in the initial video frame sequence using a video generation model, and gradually eliminate the noise to generate the final target video frame sequence.
[0166] Optionally, after obtaining the target video frame sequence, the target video frame sequence can be input into the VAE decoder, and the VAE decoder can be used to decode the target video to obtain the target video formed by the image motion effect to be processed.
[0167] Through steps S402 to S404 of this application, in response to an input command applied to the operation interface, an image to be processed is displayed on the operation interface; in response to a video generation command applied to the operation interface, a target video including video content matching the image to be processed is displayed on the operation interface. The target video is generated based on a target video frame sequence, which is obtained by a video generation model using predicted noise information to denoise an initial video frame sequence. The video generation model is trained based on video samples. The predicted noise information is obtained by using video description information to guide the video generation model to predict noise in the initial video frame sequence. The video description information is used to describe the video content. The initial video frame sequence is obtained by adding noise to the image to be processed using initial noise information. This achieves the technical effect of improving the fidelity of video generated from images and solves the technical problem of poor fidelity in video generated from images.
[0168] This application also provides a method for generating video. Figure 5 This is a flowchart of a video generation method according to an embodiment of this application, such as... Figure 5 As shown, the method may include the following steps:
[0169] Step S502: Display the image to be processed on the presentation screen of the virtual reality (VR) device or augmented reality (AR) device.
[0170] In the technical solution provided by step S502 of this application, the image to be processed can be displayed on the display screen of a virtual reality (VR) device or an augmented reality (AR) device.
[0171] Step S504: Using the initial noise information, noise is added to the image to be processed to obtain the initial video frame sequence.
[0172] In the technical solution provided by step S504 of this application, after displaying the image to be processed, the initial noise information can be used to add noise to the image to be processed to obtain an initial video frame sequence.
[0173] Optionally, before video generation, initial sampling noise can be obtained by sampling random noise that follows a Gaussian distribution. That is, real noise can be initially sampled from random noise, and the above-mentioned large amount of real noise can be summarized into initial noise information.
[0174] Optionally, the image to be processed can be input into a VAE encoder for encoding. The encoded image to be processed and the initial noise information can be combined to perform padding processing. That is, the encoded image to be processed is padded into the initial noise information to generate an initial video frame sequence that is consistent with the overall style and layout of the image to be processed.
[0175] Step S506: Input the initial video frame sequence into the video generation model in the VR device or AR device, wherein the video generation model is trained based on video samples.
[0176] In the technical solution provided by step S506 of this application, after adding noise to the image to be processed using the initial noise information to obtain the initial video frame sequence, the initial video frame sequence can be input into the video generation model in the VR device or AR device. The video generation model can then be used to process the initial video frame sequence to obtain the corresponding target video.
[0177] Optionally, during the training of the video generation model, the input video samples can be subjected to a forward noise addition process and then fed into the 3D U-Net model to predict the noise in the video samples. If the predicted noise is accurate, the 3D U-Net model can be identified as the video generation model in this embodiment.
[0178] Step S508: Using video description information, guide the video generation model to predict noise in the initial video frame sequence to obtain predicted noise information, and use the predicted noise information to denoise the initial video frame sequence to obtain the target video frame sequence. The video description information is used to describe the video content that matches the image to be processed.
[0179] In the technical solution provided in step S508 of this application, after the initial video frame sequence is input into the video generation model, the video description information can be used to guide the video generation model to predict noise in the initial video frame sequence, thereby obtaining predicted noise information. The initial video frame sequence can then be denoised to remove the predicted noise information, thus obtaining the target video frame sequence.
[0180] Optionally, the video description information can be input into a contrastive language-image pre-trained encoder for encoding to obtain encoded video description information. This encoded video description information can then be input into a video generation model.
[0181] Optionally, after the video generation model receives the encoded video description information and the initial video frame sequence, it can use the video generation model to perform noise prediction on the initial video frame sequence to obtain the predicted noise information.
[0182] For example, the denoising iterative process using DDIM can predict the noise at each step in the initial video frame sequence through a video generation model, and gradually eliminate the noise to generate the final target video frame sequence.
[0183] Step S510: Generate a target video based on the target video frame sequence, wherein the target video includes video content.
[0184] In the technical solution provided by step S510 of this application, after obtaining the target video frame sequence, the target video frame sequence can be input into the VAE decoder, and the VAE decoder can be used to decode the target video frame sequence to obtain the target video formed by the image motion effect to be processed.
[0185] Step S512: Drive the VR or AR device to display the target video.
[0186] In the technical solution provided in step S512 of this application, the target video can be transmitted to a VR device or an AR device. Controlling the VR device or AR device allows for the display of the target video.
[0187] Through steps S502 to S512 of this application, the image to be processed is displayed on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device; noise is added to the image to be processed using initial noise information to obtain an initial video frame sequence; the initial video frame sequence is input into a video generation model in the VR or AR device, wherein the video generation model is trained based on video samples; video description information is used to guide the video generation model to predict noise in the initial video frame sequence to obtain predicted noise information, and the predicted noise information is used to denoise the initial video frame sequence to obtain a target video frame sequence, wherein the video description information is used to describe the video content that matches the image to be processed; a target video is generated based on the target video frame sequence, wherein the target video includes video content; and the VR or AR device is driven to display the target video, thereby achieving the technical effect of improving the fidelity of the video generated based on the image and solving the technical problem of poor fidelity of the video generated based on the image.
[0188] Example 2
[0189] According to an embodiment of this application, an embodiment of a video generation system is also provided. Figure 6 This is a schematic diagram of a video generation system according to an embodiment of this application, such as... Figure 6 As shown, the video generation system 600 may include an e-commerce platform 601 and a media platform 602.
[0190] E-commerce platform 601 is used to identify the image to be processed; using initial noise information, the image to be processed is denoised to obtain an initial video frame sequence; the initial video frame sequence is input into a video generation model, which is trained based on video samples; using video description information, the video generation model is guided to predict noise in the initial video frame sequence to obtain predicted noise information, and the predicted noise information is used to denoise the initial video frame sequence to obtain a target video frame sequence, where the video description information is used to describe the video content that matches the image to be processed; based on the target video frame sequence, a target video is generated, where the target video includes video content.
[0191] In this embodiment, the e-commerce platform 601 can receive images to be processed, which need to be used to generate videos, transmitted by merchants. The e-commerce platform 601 can identify the images to be processed. Initial noise information is used to add noise to the images to be processed, resulting in an initial video frame sequence. This initial video frame sequence can be input into a video generation model, which can then process the initial video frame sequence to obtain the corresponding target video. Video description information can be used to guide the video generation model to predict noise in the initial video frame sequence, obtaining predicted noise information. The initial video frame sequence can then be denoised to remove the predicted noise information, resulting in the target video frame sequence, and the target video corresponding to the target video frame sequence is generated.
[0192] Optionally, if a merchant needs to obtain a product video corresponding to a specific product image, they can upload the product image to the server on the e-commerce platform. For example, the merchant can input a product image into the e-commerce platform's interface on their client device. When the e-commerce platform detects the product image, it can upload it to the server associated with the platform. The server then performs image recognition on the product.
[0193] Optionally, if the initial noise information is composed of L real noises obtained from L frames of video, then padding is performed on the image to be processed after initial encoding, that is, L real noises are added to the image to be processed after encoding, and the corresponding image after noise addition can be obtained. The L images after noise addition can be used to form the initial video frame sequence of the image to be processed.
[0194] Optionally, noise can be added to the image to be processed to obtain an initial video frame sequence. In this way, more image details can be added to the image to be processed. A video generation model can be used to perform noise prediction and denoising on the initial video frame sequence, thereby achieving the technical effect of improving the fidelity of the image-generated video.
[0195] Optionally, during the training of the video generation model, the input video samples can be subjected to a forward noise addition process and then fed into the 3D U-Net model to predict the noise in the video samples. If the predicted noise is accurate, the 3D U-Net model can be identified as the video generation model in this embodiment.
[0196] Optionally, after the video generation model receives the encoded video description information and the initial video frame sequence, it can use the video generation model to predict noise in the initial video frame sequence, thereby obtaining predicted noise information. The denoising iterative process using DDIM allows the video generation model to predict the noise at each step of the initial video frame sequence and gradually eliminate the noise to generate the final target video frame sequence.
[0197] Optionally, after obtaining the target video frame sequence, the target video frame sequence can be input into the VAE decoder, and the VAE decoder can be used to decode the target video to obtain the target video formed by the image motion effect to be processed.
[0198] Optionally, after the target video is generated on the e-commerce platform 601, it can be transmitted to the media platform 602 via the network.
[0199] Media platform 602 is used to play the target video.
[0200] In this embodiment, after the media platform 602 receives the target video from the e-commerce platform 601, the target video can be played on the media platform.
[0201] In this embodiment, a video generation system is provided. An image to be processed is identified through an e-commerce platform 601; initial noise information is used to add noise to the image to obtain an initial video frame sequence; the initial video frame sequence is input into a video generation model, which is trained based on video samples; video description information is used to guide the video generation model to predict noise in the initial video frame sequence, obtaining predicted noise information, and the predicted noise information is used to denoise the initial video frame sequence to obtain a target video frame sequence, wherein the video description information describes the video content matching the image to be processed; a target video is generated based on the target video frame sequence, wherein the target video includes video content; the target video is played through a media platform, thereby achieving the technical effect of improving the fidelity of image-based generated videos and solving the technical problem of poor fidelity in image-based generated videos.
[0202] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application, such as the data to be verified, are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0203] Example 3
[0204] Currently, with the content-driven development of e-commerce platforms, consumers have extremely high expectations for the visual experience of products, and video content, with its vividness and emotional connection, has become a new benchmark. Product presentation in short video format is more attractive to users and is gradually becoming a focus for merchants' promotions. However, the cost and technical barriers to video production are far higher than the cost of producing product images. Driven by the content trend, the demand for video creation is exploding, and merchants and advertisers urgently need a more efficient and cost-effective video production method.
[0205] In one related technology, image-to-video methods mainly focus on specific image domains, such as faces, human poses, or simple natural scenes or objects, such as clouds, flowing water, handwritten letters, simple doodles, etc., using traditional generative models (such as GANs and VAEs) to animate the images. The drawback of this type of method is that it can only be applied to specific images and cannot be applied to all images in the natural open domain, lacking universality. Therefore, the technical problem of poor fidelity in image-generated videos still exists.
[0206] Another related technique uses video diffusion models to generate videos from images. However, this type of method has the drawback that the generated video can only approximate a given image or retain a similar image style, losing the local details of the original image. It cannot generate videos from a given image, meaning its fidelity is limited. Furthermore, some methods require retraining the backbone network to adapt to additional image input, consuming computational resources. Therefore, the technical problem of poor fidelity in image-generated videos persists.
[0207] Furthermore, this application provides a high-fidelity video motion effect generation method based on a diffusion model. This method, based on a video diffusion model, uses a plug-and-play noise correction approach to enable open-domain images to "move," achieving motion effect generation from image to video. Since the embodiments of this application can add noise to the image to be processed, thereby increasing the image detail and improving the accuracy of video generation, and can use a pre-trained video generation model to generate videos, it can use any image in the natural open domain, thus improving the versatility of video generation. Through the above method, video generation can be performed more accurately without losing the local details of the image to be processed, thereby improving the fidelity of image-generated videos and solving the technical problem of poor fidelity in image-generated videos.
[0208] The method described in this embodiment will be further described below.
[0209] In this embodiment, with the rapid development of e-commerce, consumers increasingly rely on rich product displays and high-quality visual experiences in their purchasing decisions. While static images can present necessary product details and information, they lack interactivity and immersion. In the fiercely competitive consumer market, video, as a more dynamic and intuitive medium, can effectively enhance user engagement and purchase intention. Against the backdrop of AIGC technology and the content-driven nature of e-commerce platforms, the production and demand for video creatives are expected to experience exponential growth, bringing new opportunities to many businesses. However, compared to traditional text and image creation, video creatives are more costly and technically challenging. Businesses often possess a large amount of image material but limited video material. If image material could be made to "move" effectively, it could greatly enrich the video material library and provide advertisers with more convenient video editing tools. Therefore, a method for generating video motion effects is proposed, which can achieve image-to-video generation based on a pre-trained motion effect generation module. Utilizing AIGC's video generation technology to generate video creatives beyond the original materials of businesses enhances the richness of the creative content.
[0210] In this embodiment, a video diffusion model is employed to provide motion effect generation capabilities. To improve the fidelity of the generated video, two measures are taken: noise is introduced into the encoded image information through padding to enhance image details; and noise correction is performed during the backward denoising process to reduce the cumulative error caused by model prediction noise. Based on existing video motion effect models, no additional training is required; video motion effects can be generated from images simply through the "padding + noise correction" strategy.
[0211] In this embodiment, Figure 7This is a flowchart of a high-fidelity video motion effect generation method based on a diffusion model according to an embodiment of this application, such as... Figure 7 As shown, the method may include the following steps:
[0212] Step S702: Obtain the input image.
[0213] In this embodiment, a given image can be input into an e-commerce platform. A video overlay image is then created on the e-commerce platform.
[0214] Step S704: Perform video padding processing on the input image.
[0215] In this embodiment, the input image can be processed into a video pad image to obtain an initial video frame sequence.
[0216] Optionally, to improve the fidelity of the generated video, a "pad image + noise correction" generation strategy can be set. For a given image I 0 And the initial sampled video noise, i.e., the true noise n 0:L-1 Where L can represent the number of frames in the generated video.
[0217] Optionally, the input image is encoded using a VAE encoder to obtain the latent features z of the image. 0 Furthermore, by using dimension broadcasting, the data dimension z is maintained to be the same as the initial sampling noise. 0:L-1 Then, the encoded latent features, known as "pad maps," are added to the video noise. The specific operation of "pad maps" involves a forward noise addition process using a diffusion model, employing the initial sampled real noise n. 0:L-1 For image latent features z 0:L-1 After adding noise for T steps, the noisy latent space features are obtained. In this way, after DDIM denoising is completed, the overall style and layout of the generated video are consistent with the given image.
[0218] Step S706: The video after video pad image processing is transmitted to the video motion effect model for noise reduction processing to achieve noise correction.
[0219] In this embodiment, the video after video pad image processing can be transmitted to the video motion effect model for noise reduction processing to correct noise.
[0220] Figure 8 This is a schematic diagram illustrating a feature dimension transformation process using a video motion effect model according to an embodiment of this application, such as... Figure 8As shown, a diffusion model can be used as the baseline model for video motion effect generation. To improve encoding efficiency and save model space, latent space features are chosen. For the input video frames, they are first encoded into video latent space features by a VAE encoder. Considering the temporal consistency of video, the visual content between frames needs to be correlated. Therefore, an inflated 3D U-Net model structure is designed in the video diffusion model 800, which includes a spatial module 801 and a temporal module 802. The spatial module 801 can treat the input video frames as multiple batches of image features, perform convolution and attention operations in the spatial dimension, and encode and extract features from the video frames spatially. The temporal module 802 can transform the dimensions of the video frames, perform attention operations in the temporal dimension, and learn the temporal correlation between video frames.
[0221] Optionally, such as Figure 8 As shown, during training, the input video is first subjected to a forward noise addition process, and then fed into the dilated 3D U-Net model, where the model predicts the noise in the video. During inference, noise is sampled from a random Gaussian distribution, and then through the DDIM denoising iterative process, the model predicts the noise at each step, gradually eliminating the noise to generate the final video. The proposed solution uses an existing pre-trained motion effects generation module, processing L=16 frames at a time, ultimately generating a 2-second video with 8 fps.
[0222] During the inference process, consider the noise characteristics of the latent space at time step t of the DDIM denoising process. The dimension is Where B represents the number of videos in the batch, L represents the number of video frames, H′ and W′ represent the dimensions of the latent space features, and C represents the number of latent space feature channels. For a dilated 3D U-Net Block module, before inputting into the spatial module 801, the dimensions of the batch processing and the dimensions of the video frames are first merged, treating the video frames as multiple batches of images, i.e., the dimensions are converted to... When passing through the "spatial module," the convolutional module and attention module operate on the spatial dimension H′×W′, and the dimension of the output features remains the same. Restore the dimension to Before being input into the timing module 802, the dimensions of the batch processing and the spatial dimensions of the video are merged, i.e., the dimensions are converted. When passing through the "temporal module", attention operations are performed on the temporal dimension L, and the dimension of the output features remains the same. Restore the dimension to Similarly, after processing by multiple downsampling and upsampling dilated 3D U-Net modules, the model obtains the noise features at time step t-1.
[0223] Optionally, considering that model training cannot achieve perfect loss, meaning that the model's noise prediction will always have biases, thus reducing the fidelity of the video, in order to alleviate the accumulation of errors in the model during the DDIM denoising process, for the noise predicted by the model at each step, "noise correction" is performed using the noise from the initial sampling. That is, by calculating the difference between the predicted noise and the real noise, a weighted calculation method is used to appropriately adjust the predicted noise.
[0224] Optionally, at time step t of the DDIM denoising process, the model input (including noisy latent space features) is represented as: The input is given to the dilated 3D U-Net model, resulting in predicted noise. The previously recorded true noise n of the initial sample 0:L-1 Noise correction is performed.
[0225] Optionally, calculate the prediction noise. and real noise n 0:L-1 Differences between them:
[0226]
[0227] in, It can be used to represent the difference noise information at time step t; It can be used to represent the prediction noise information at time step t.
[0228] Optionally, a weighted approach can be used to appropriately adjust the prediction noise.
[0229] in, It can be used to represent target noise information; ω 0:L-1 It can contain L weight values, ranging from 0 to 1, which represent the correction adjustment coefficients for each frame, i.e., the first weight; This can be used to represent differences only in the first frame. Then a copy operation is performed to expand the dimension to L frames.
[0230] Optionally, based on the target noise information determined above, a DDIM denoising process can be performed using a video generation model to obtain the latent space features at time t-1.
[0231] Optionally, this setting can eliminate noise errors in the first frame. Because the noise in the first frame is corrected, the first frame achieves fidelity, while the remaining frames maintain temporal consistency with the first frame. The latent space features of the final video frame obtained after denoising are then analyzed. VAE decoding is performed to generate a motion-effect video. In addition, the correction adjustment coefficient ω for different frames in noise correction is adjusted.0:L-1 The solution also allows for adjustment and trade-off between the motion effects and fidelity of the generated video.
[0232] Optionally, if the correction adjustment coefficient ω of the i-th frame is... i A larger adjustment indicates that the correction process of the i-th frame tends to reference the noise difference of the first frame more, resulting in higher fidelity but relatively weaker motion effects. Conversely, if the correction adjustment coefficient ω of the i-th frame is increased... i A smaller adjustment indicates that the correction process for the i-th frame tends to reference its own noise differences more, resulting in higher motion fidelity but relatively lower fidelity. Therefore, based on the above analysis, you can determine whether you prioritize fidelity or motion fidelity based on your own needs, and adjust the corresponding correction adjustment coefficient accordingly.
[0233] In this embodiment, Figure 9 This is a schematic diagram of a video motion effect generation process based on noise correction according to an embodiment of this application, such as... Figure 9 As shown, this process can utilize a VAE encoder 901, a padding module 902, a CLIP encoder 903, an inflated 3D U-Net module 904, a noise correction module 905, and a VAE decoder 906. Based on these modules, only any single image needs to be input to generate a video (the first frame of the generated video is identical to the input image, while the remaining frames produce motion effects). The "padding + noise correction" strategy proposed in this application does not require additional training and can be directly applied to the inference stage of the motion effect model to improve the fidelity of the image-generated video.
[0234] Optionally, such as Figure 9 As shown, a given image can be input into the VAE encoder 901 for encoding to obtain the image's latent features. Real noise and the image's latent features can be input into the padding module 902 for padding, i.e., noise addition processing can be performed to obtain noisy latent space features. Video description information can be input into the CLIP encoder 903 for encoding. The encoded video description information and noisy latent space features are then input into the dilated 3D U-Net module 904 for noise prediction to obtain predicted noise. Predicted noise and real noise can be input into the noise correction module 905 to correct the predicted noise, obtaining corrected predicted noise. The latent space features at time t-1 can be obtained based on the corrected predicted noise. Thus, the final latent space features of the video frame are obtained. The final video frame latent space features can be input into the VAE decoder 906 to decode and obtain the generated video.
[0235] Step S708: Output the generated video.
[0236] In this embodiment, the video with motion effects generated based on a given image can be output to a corresponding media platform for playback.
[0237] For example, Figure 10(a) is a schematic diagram of an example of generating video from an image according to an embodiment of the present application. As shown in Figure 10(a), the input image A can be a picture of a girl's profile with her eyes open. The above method can be used to create motion effects to obtain the corresponding generated video A. The generated video A can be a video of a girl blinking, in which one frame shows the girl with her eyes closed.
[0238] For another example, Figure 10(b) is a schematic diagram of another example of generating video from an image according to an embodiment of this application. As shown in Figure 10(b), the input image B can be a picture of a flower. The animation effect can be applied using the above method to obtain the corresponding generated video B. The generated video B can be a video of flowers blooming from buds on a tree branch, in which one frame is a picture of several flowers when they are buds.
[0239] In this embodiment, if it is necessary to generate a corresponding video based on a certain image, the image to be processed for video generation can be identified, and initial noise information can be used to add noise to the image to be processed, resulting in a noisy initial video frame sequence. The initial video frame sequence can be input into a video generation model, which is pre-trained based on video samples to process the initial video frame sequence. In the video generation model, matching video description information that describes the video content generated based on the image to be processed can be analyzed from the initial video frame sequence. Using the video description information, noise prediction is performed on the initial video frame sequence to obtain predicted noise information, which can then be used to denoise the initial video frame sequence, resulting in a target video frame sequence. Thus, a target video containing video content matching the current image can be generated based on the target video frame sequence. Since the embodiments of this application can add noise to the image to be processed to add more image details, thereby achieving the purpose of improving the accuracy of video generation by increasing image details, and can use a pre-trained video generation model to generate videos, it can use any image in the natural open domain, thereby also achieving the purpose of improving the versatility of video generation. Through the above method, without losing the local details of the image to be processed, video generation can be performed more accurately, thereby achieving the technical effect of improving the fidelity of image-based videos and solving the technical problem of poor fidelity of image-based videos.
[0240] Example 4
[0241] According to embodiments of this application, a method for implementing the above is also provided. Figure 2 The video generation method shown is a video generation device.
[0242] Figure 11 This is a schematic diagram of a video generation apparatus according to an embodiment of this application, such as... Figure 11 As shown, the video generation device 1100 may include: a first recognition unit 1102, a first processing unit 1104, a first input unit 1106, a first prediction unit 1108, and a first generation unit 1110.
[0243] The first recognition unit 1102 is used to recognize the image to be processed.
[0244] The first processing unit 1104 is used to add noise to the image to be processed using the initial noise information to obtain an initial video frame sequence.
[0245] The first input unit 1106 is used to input the initial video frame sequence into the video generation model, wherein the video generation model is trained based on video samples.
[0246] The first prediction unit 1108 is used to guide the video generation model to predict noise in the initial video frame sequence using video description information, obtain predicted noise information, and use the predicted noise information to denoise the initial video frame sequence to obtain the target video frame sequence. The video description information is used to describe the video content that matches the image to be processed.
[0247] The first generation unit 1110 is used to generate a target video based on a target video frame sequence, wherein the target video includes video content.
[0248] Here, the first identification unit 1102, the first processing unit 1104, the first input unit 1106, the first prediction unit 1108, and the first generation unit 1110 correspond to steps S202 to S210 in Embodiment 1. The five units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units can be hardware or software components stored in memory (e.g., memory 1704) and processed by one or more processors (e.g., processors 1702a, 1702b, ..., 1702n). The above units can also be part of a device and run in the computer terminal 170 provided in Embodiment 8.
[0249] According to embodiments of this application, a method for implementing the above is also provided. Figure 3 The video generation method shown is a video generation device.
[0250] Figure 12 This is a schematic diagram of a video generation apparatus according to an embodiment of this application, such as... Figure 12As shown, the video generation device 1200 may include: a second recognition unit 1202, a second processing unit 1204, a second input unit 1206, a second prediction unit 1208, a second generation unit 1210, and a playback unit 1212.
[0251] The second recognition unit 1202 is used to identify the image to be processed, wherein the image content of the image to be processed includes the object to be traded.
[0252] The second processing unit 1204 is used to add noise to the image to be processed using the initial noise information to obtain an initial video frame sequence.
[0253] The second input unit 1206 is used to input the initial video frame sequence into the video generation model, wherein the video generation model is trained based on video samples.
[0254] The second prediction unit 1208 is used to guide the video generation model to predict noise in the initial video frame sequence using video description information, obtain predicted noise information, and use the predicted noise information to denoise the initial video frame sequence to obtain the target video frame sequence. The video description information is used to describe the video content that matches the image to be processed.
[0255] The second generation unit 1210 is used to generate a motion effect video based on the target video frame sequence, wherein the motion effect video includes video content and is used to display the transaction object.
[0256] Push unit 1212 is used to push motion effect videos to media platforms for playback.
[0257] It should be noted that the second recognition unit 1202, the second processing unit 1204, the second input unit 1206, the second prediction unit 1208, the second generation unit 1210, and the push unit 1212 correspond to steps S302 to S312 in Embodiment 1. The six units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units can be hardware or software components stored in memory (e.g., memory 1704) and processed by one or more processors (e.g., processors 1702a, 1702b, ..., 1702n). The above units can also be part of a device and run in the computer terminal 170 provided in Embodiment 8.
[0258] According to embodiments of this application, a method for implementing the above is also provided. Figure 4 The video generation method shown is a video generation device.
[0259] Figure 13This is a schematic diagram of a video generation apparatus according to an embodiment of this application, such as... Figure 13 As shown, the video generation device 1300 may include: a first display unit 1302 and a second display unit 1304.
[0260] The first display unit 1302 is used to respond to input commands applied to the operation interface and display the image to be processed on the operation interface.
[0261] The second display unit 1304 is used to respond to a video generation command applied to the operation interface and display a target video on the operation interface, which includes video content matching the image to be processed. The target video is generated based on a target video frame sequence. The target video frame sequence is obtained by the video generation model using predicted noise information to denoise the initial video frame sequence. The video generation model is trained based on video samples. The predicted noise information is obtained by using video description information to guide the video generation model to predict noise in the initial video frame sequence. The video description information is used to describe the video content. The initial video frame sequence is obtained by adding noise to the image to be processed using initial noise information.
[0262] It should be noted that the first display unit 1302 and the second display unit 1304 mentioned above correspond to steps S402 to S404 in Embodiment 1. The two units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units can be hardware components or software components stored in memory (e.g., memory 1704) and processed by one or more processors (e.g., processors 1702a, 1702b, ..., 1702n). The above units can also be part of a device and run in the computer terminal 170 provided in Embodiment 8.
[0263] According to embodiments of this application, a method for implementing the above is also provided. Figure 5 The video generation method shown is a video generation device.
[0264] Figure 14 This is a schematic diagram of a video generation apparatus according to an embodiment of this application, such as... Figure 14 As shown, the video generation device 1400 may include: a third display unit 1402, a third processing unit 1404, a third input unit 1406, a third prediction unit 1408, a third generation unit 1410, and a fourth display unit 1412.
[0265] The third display unit 1402 is used to display the image to be processed on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device.
[0266] The third processing unit 1404 is used to add noise to the image to be processed using the initial noise information to obtain an initial video frame sequence.
[0267] The third input unit 1406 is used to input the initial video frame sequence into the video generation model in the VR device or AR device, wherein the video generation model is trained based on video samples.
[0268] The third prediction unit 1408 is used to guide the video generation model to predict noise in the initial video frame sequence using video description information, obtain predicted noise information, and use the predicted noise information to denoise the initial video frame sequence to obtain the target video frame sequence. The video description information is used to describe the video content that matches the image to be processed.
[0269] The third generation unit 1410 is used to generate a target video based on the target video frame sequence, wherein the target video includes video content.
[0270] The fourth display unit 1412 is used to drive VR or AR devices to display target videos.
[0271] It should be noted that the third display unit 1402, the third processing unit 1404, the third input unit 1406, the third prediction unit 1408, the third generation unit 1410, and the fourth display unit 1412 correspond to steps S502 to S512 in Embodiment 1. The six units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units can be hardware or software components stored in memory (e.g., memory 1704) and processed by one or more processors (e.g., processors 1702a, 1702b, ..., 1702n). The above units can also be part of a device and run in the computer terminal 170 provided in Embodiment 8.
[0272] In this video generation device, if a video needs to be generated based on a certain image, the image to be processed can be identified, and initial noise information can be used to add noise to the image, resulting in a noisy initial video frame sequence. This initial video frame sequence can be input into a video generation model, which is pre-trained based on video samples to process it. Within the video generation model, matching video description information describing the video content generated from the image to be processed can be analyzed from the initial video frame sequence. Using this video description information, noise prediction is performed on the initial video frame sequence to obtain predicted noise information. This predicted noise information can then be used to denoise the initial video frame sequence, resulting in a target video frame sequence. Finally, a target video containing video content matching the current image can be generated based on the target video frame sequence. Since the embodiments of this application can add noise to the image to be processed to add more image details, thereby achieving the purpose of improving the accuracy of video generation by increasing image details, and can use a pre-trained video generation model to generate videos, it can use any image in the natural open domain, thereby also achieving the purpose of improving the versatility of video generation. Through the above method, without losing the local details of the image to be processed, video generation can be performed more accurately, thereby achieving the technical effect of improving the fidelity of image-based videos and solving the technical problem of poor fidelity of image-based videos.
[0273] Example 5
[0274] Embodiments of this application may provide a computer terminal, which may be any computer terminal device in a group of computer terminals. Optionally, in this embodiment, the aforementioned computer terminal may also be replaced by a mobile terminal or other terminal device.
[0275] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.
[0276] In this embodiment, the computer terminal described above can execute the program code for the following steps in the video generation method: identifying the image to be processed; using initial noise information to add noise to the image to be processed to obtain an initial video frame sequence; inputting the initial video frame sequence into a video generation model, wherein the video generation model is trained based on video samples; using video description information to guide the video generation model to predict noise in the initial video frame sequence to obtain predicted noise information, and using the predicted noise information to denoise the initial video frame sequence to obtain a target video frame sequence, wherein the video description information is used to describe the video content that matches the image to be processed; and generating a target video based on the target video frame sequence, wherein the target video includes video content.
[0277] Optionally, Figure 15 This is a structural block diagram of a computer terminal according to an embodiment of this application. Figure 15 As shown, the computer terminal A may include one or more (only one is shown in the figure) processors 1502, memory 1504, and transmission devices 1506.
[0278] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the video generation method and apparatus in this application embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned video generation method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0279] The processor can call the information and application program stored in the memory through the transmission device to perform the following steps: extracting image features from the image to be processed that have the same data dimension as the data dimension of the initial noise information; using the initial noise information to add noise to the image features to obtain the features of the initial video frame sequence.
[0280] Optionally, the processor may also execute program code that performs the following steps: using initial noise information to add noise to the image features at multiple time steps to obtain features of the initial video frame sequence corresponding to the multiple time steps.
[0281] Optionally, the processor may also execute program code with the following steps: multiple time steps including a first time step and a second time step, the second time step being the previous time step of the first time step, wherein, using video description information, the video generation model is guided to perform noise prediction on the initial video frame sequence to obtain predicted noise information, and the initial video frame sequence is denoised using the predicted noise information to obtain a target video frame sequence, including: a processing step, using video description information to guide the video generation model to perform noise prediction on the features of the initial video frame sequence corresponding to the first time step to obtain predicted noise information corresponding to the first time step, and using the predicted noise information corresponding to the first time step to denoise the features of the initial video frame sequence corresponding to the first time step to obtain the features of the target video frame sequence corresponding to the second time step; a determination step, if the second time step has a previous time step among the multiple time steps, then the previous time step among the multiple time steps is determined as the second time step, the features of the target video frame sequence corresponding to the second time step are determined as the features of the initial video frame sequence corresponding to the first time step, and the processing step is returned to be executed until the second time step is the first time step among the multiple time steps.
[0282] Optionally, the processor may also execute program code that performs the following steps: if the second time step is the first time step of a plurality of time steps, then the features of the target video frame sequence corresponding to the second time step are decoded to obtain the target video.
[0283] Optionally, the processor may also execute program code that performs the following steps: adjusting the predicted noise information using the initial noise information to obtain the target noise information; and denoising the features of the initial video frame sequence using the target noise information to obtain the features of the target video frame sequence.
[0284] Optionally, the processor may also execute program code that performs the following steps: obtaining the difference noise information between the predicted noise information and the initial noise information; and adjusting the predicted noise information using the difference noise information to obtain the target noise information.
[0285] Optionally, the processor may also execute program code that performs the following steps: obtaining the predicted noise information corresponding to the first video frame in different video frames, and the first difference noise information between the predicted noise information and the corresponding initial noise information; obtaining the predicted noise information corresponding to the target video frame in different video frames, and the second difference noise information between the predicted noise information and the corresponding initial noise information, wherein the target video frame is any video frame other than the first video frame in different video frames.
[0286] Optionally, the processor may also execute program code that performs the following steps: weighting the first difference noise information with a first weight and weighting the second difference noise information with a second weight, wherein the second weight is negatively correlated with the first weight; adjusting the prediction noise information corresponding to the target video frame using the weighted first difference noise information and the weighted second difference noise information to obtain the target noise information corresponding to the target video frame.
[0287] Optionally, the processor may also execute program code that adjusts the first weight or the second weight.
[0288] Optionally, the processor may also execute program code that performs the following steps: using the first difference noise information to adjust the prediction noise information corresponding to the first video frame to obtain the target noise information corresponding to the first video frame.
[0289] Optionally, the processor may also execute program code that performs the following steps: determining the data dimension of noise information based on the number of video frames corresponding to the target video to be generated.
[0290] Optionally, the processor can invoke information and application programs stored in the memory via a transmission device to perform the following steps: identifying the image to be processed, wherein the image content of the image to be processed includes the object to be traded; adding noise to the image to be processed using initial noise information to obtain an initial video frame sequence; inputting the initial video frame sequence into a video generation model, wherein the video generation model is trained based on video samples; using video description information to guide the video generation model to predict noise in the initial video frame sequence to obtain predicted noise information, and using the predicted noise information to denoise the initial video frame sequence to obtain a target video frame sequence, wherein the video description information is used to describe the video content matching the image to be processed; generating a motion effect video based on the target video frame sequence, wherein the motion effect video includes video content and is used to display the object to be traded; and pushing the motion effect video to a media platform for playback.
[0291] Optionally, the processor can invoke information and application programs stored in the memory via a transmission device to perform the following steps: in response to an input command applied to the operation interface, display the image to be processed on the operation interface; in response to a video generation command applied to the operation interface, display a target video including video content matching the image to be processed on the operation interface, wherein the target video is generated based on a target video frame sequence, the target video frame sequence is obtained by a video generation model using predicted noise information to denoise an initial video frame sequence, the video generation model is trained based on video samples, the predicted noise information is obtained by using video description information to guide the video generation model to predict noise in the initial video frame sequence, the video description information is used to describe the video content, and the initial video frame sequence is obtained by adding noise to the image to be processed using initial noise information.
[0292] Optionally, the processor can invoke information and applications stored in the memory via a transmission device to perform the following steps: displaying the image to be processed on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device; adding noise to the image to be processed using initial noise information to obtain an initial video frame sequence; inputting the initial video frame sequence into a video generation model in the VR or AR device, wherein the video generation model is trained based on video samples; using video description information to guide the video generation model to predict noise in the initial video frame sequence to obtain predicted noise information, and using the predicted noise information to denoise the initial video frame sequence to obtain a target video frame sequence, wherein the video description information is used to describe the video content matching the image to be processed; generating a target video based on the target video frame sequence, wherein the target video includes video content; and driving the VR or AR device to display the target video.
[0293] This application provides a method for generating a video. In this embodiment, if it is necessary to generate a corresponding video based on a certain image, the image to be processed for video generation can be identified, and initial noise information can be used to add noise to the image, resulting in a noisy initial video frame sequence. The initial video frame sequence can be input into a video generation model, which is pre-trained based on video samples to process the initial video frame sequence. In the video generation model, matching video description information describing the video content generated based on the image to be processed can be analyzed from the initial video frame sequence. Using the video description information, noise prediction is performed on the initial video frame sequence to obtain predicted noise information. This predicted noise information can then be used to denoise the initial video frame sequence, resulting in a target video frame sequence. Therefore, a target video containing video content matching the current image can be generated based on the target video frame sequence. Since the embodiments of this application can add noise to the image to be processed to add more image details, thereby achieving the purpose of improving the accuracy of video generation by increasing image details, and can use a pre-trained video generation model to generate videos, it can use any image in the natural open domain, thereby also achieving the purpose of improving the versatility of video generation. Through the above method, without losing the local details of the image to be processed, video generation can be performed more accurately, thereby achieving the technical effect of improving the fidelity of image-based videos and solving the technical problem of poor fidelity of image-based videos.
[0294] Those skilled in the art will understand that Figure 15 The structure shown is for illustrative purposes only. Computer terminal A can also be a smartphone (such as an Android phone, iOS phone, etc.), tablet computer, PDA, mobile Internet device (MID), PAD, and other terminal devices. Figure 15 This does not limit the structure of the aforementioned computer terminal A. For example, computer terminal A may also include components that are more complex than those described above. Figure 15 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 15 The different configurations shown.
[0295] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0296] Example 6
[0297] Embodiments of this application also provide a computer-readable storage medium. Optionally, in this embodiment, the aforementioned computer-readable storage medium can be used to store the program code executed by the video generation method provided in Embodiment 1.
[0298] Optionally, in this embodiment, the computer-readable storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0299] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: identifying the image to be processed; adding noise to the image to be processed using initial noise information to obtain an initial video frame sequence; inputting the initial video frame sequence into a video generation model, wherein the video generation model is trained based on video samples; using video description information to guide the video generation model to predict noise in the initial video frame sequence to obtain predicted noise information, and using the predicted noise information to denoise the initial video frame sequence to obtain a target video frame sequence, wherein the video description information is used to describe the video content that matches the image to be processed; generating a target video based on the target video frame sequence, wherein the target video includes video content.
[0300] Optionally, the computer-readable storage medium may also execute program code that performs the following steps: extracting image features from the image to be processed that have the same data dimension as the initial noise information; and using the initial noise information to add noise to the image features to obtain the features of the initial video frame sequence.
[0301] Optionally, the computer-readable storage medium may also execute program code that performs the following steps: using initial noise information, performing noise processing on image features at multiple time steps to obtain features of the initial video frame sequence corresponding to the multiple time steps respectively.
[0302] Optionally, the aforementioned computer-readable storage medium may also execute program code containing the following steps: multiple time steps including a first time step and a second time step, the second time step being the previous time step of the first time step, wherein, using video description information, the video generation model is guided to perform noise prediction on the initial video frame sequence to obtain predicted noise information, and the initial video frame sequence is denoised using the predicted noise information to obtain a target video frame sequence, including: a processing step, using video description information to guide the video generation model to perform noise prediction on the features of the initial video frame sequence corresponding to the first time step to obtain predicted noise information corresponding to the first time step, and using the predicted noise information corresponding to the first time step to denoise the features of the initial video frame sequence corresponding to the first time step to obtain the features of the target video frame sequence corresponding to the second time step; a determination step, if the second time step has a previous time step among the multiple time steps, then the previous time step among the multiple time steps is determined as the second time step, the features of the target video frame sequence corresponding to the second time step are determined as the features of the initial video frame sequence corresponding to the first time step, and the processing step is returned to be executed until the second time step is the first time step among the multiple time steps.
[0303] Optionally, the computer-readable storage medium may also execute program code that performs the following steps: if the second time step is the first time step of a plurality of time steps, then the features of the target video frame sequence corresponding to the second time step are decoded to obtain the target video.
[0304] Optionally, the computer-readable storage medium may also execute program code that performs the following steps: adjusting the predicted noise information using the initial noise information to obtain the target noise information; and using the target noise information to denoise the features of the initial video frame sequence to obtain the features of the target video frame sequence.
[0305] Optionally, the computer-readable storage medium may also execute program code that performs the following steps: acquiring the difference noise information between the predicted noise information and the initial noise information; and adjusting the predicted noise information using the difference noise information to obtain the target noise information.
[0306] Optionally, the computer-readable storage medium may also execute program code that performs the following steps: obtaining the predicted noise information corresponding to the first video frame in different video frames, and the first difference noise information between the predicted noise information and the corresponding initial noise information; obtaining the predicted noise information corresponding to the target video frame in different video frames, and the second difference noise information between the predicted noise information and the corresponding initial noise information, wherein the target video frame is any video frame other than the first video frame in different video frames.
[0307] Optionally, the computer-readable storage medium may also execute program code that performs the following steps: weighting the first difference noise information with a first weight and weighting the second difference noise information with a second weight, wherein the second weight is negatively correlated with the first weight; adjusting the prediction noise information corresponding to the target video frame using the weighted first difference noise information and the weighted second difference noise information to obtain the target noise information corresponding to the target video frame.
[0308] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: adjusting the first weight or the second weight.
[0309] Optionally, the computer-readable storage medium may also execute program code that performs the following steps: using the first difference noise information, adjusting the prediction noise information corresponding to the first video frame to obtain the target noise information corresponding to the first video frame.
[0310] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: determining the data dimension of noise information based on the number of video frames corresponding to the target video to be generated.
[0311] Optionally, the aforementioned computer-readable storage medium may also execute program code for the following steps: identifying an image to be processed, wherein the image content of the image to be processed includes the object to be traded; adding noise to the image to be processed using initial noise information to obtain an initial video frame sequence; inputting the initial video frame sequence into a video generation model, wherein the video generation model is trained based on video samples; using video description information to guide the video generation model to predict noise in the initial video frame sequence to obtain predicted noise information, and using the predicted noise information to denoise the initial video frame sequence to obtain a target video frame sequence, wherein the video description information is used to describe the video content matching the image to be processed; generating a motion effect video based on the target video frame sequence, wherein the motion effect video includes video content and is used to display the object to be traded; and pushing the motion effect video to a media platform for playback.
[0312] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: in response to an input command applied to the operation interface, displays the image to be processed on the operation interface; in response to a video generation command applied to the operation interface, displays a target video on the operation interface including video content matching the image to be processed, wherein the target video is generated based on a target video frame sequence, the target video frame sequence is obtained by a video generation model using predicted noise information to denoise an initial video frame sequence, the video generation model is trained based on video samples, the predicted noise information is obtained by using video description information to guide the video generation model to predict noise in the initial video frame sequence, the video description information is used to describe the video content, and the initial video frame sequence is obtained by adding noise to the image to be processed using initial noise information.
[0313] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: displaying the image to be processed on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device; adding noise to the image to be processed using initial noise information to obtain an initial video frame sequence; inputting the initial video frame sequence into a video generation model in the VR or AR device, wherein the video generation model is trained based on video samples; using video description information to guide the video generation model to predict noise in the initial video frame sequence to obtain predicted noise information, and using the predicted noise information to denoise the initial video frame sequence to obtain a target video frame sequence, wherein the video description information is used to describe the video content that matches the image to be processed; generating a target video based on the target video frame sequence, wherein the target video includes video content; and driving the VR or AR device to display the target video.
[0314] Example 7
[0315] Embodiments of this application may provide an electronic device that may include a memory and a processor.
[0316] Figure 16 This is a block diagram of an electronic device for a video generation method according to an embodiment of this application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present application described and / or claimed herein.
[0317] like Figure 16As shown, device 1600 includes a computing unit 1601, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1602 or a computer program loaded into random access memory (RAM) 1603 from storage unit 1608. RAM 1603 may also store various programs and data required for the operation of device 1600. The computing unit 1601, ROM 1602, and RAM 1603 are interconnected via bus 1604. Input / output (I / O) interface 1605 is also connected to bus 1604.
[0318] Multiple components in device 1600 are connected to I / O interface 1605, including: input unit 1606, such as keyboard, mouse, etc.; output unit 1604, such as various types of monitors, speakers, etc.; storage unit 1608, such as disk, optical disk, etc.; and communication unit 1609, such as network card, modem, wireless transceiver, etc. Communication unit 1609 allows device 1600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0319] The computing unit 1601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1601 performs the various methods and processes described above, such as video generation methods. For example, in some embodiments, the video generation method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1608. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1600 via ROM 1602 and / or communication unit 1609. When the computer program is loaded into RAM 1603 and executed by the computing unit 1601, one or more steps of the video generation method described above may be performed. Alternatively, in other embodiments, the computing unit 1601 may be configured to perform a video generation method by any other suitable means (e.g., by means of firmware).
[0320] Example 8
[0321] Embodiments of this application also provide a computer program product. Optionally, in this embodiment, the computer program product may include a computer program that, when executed by a processor, implements the video generation method of the embodiments of this application.
[0322] According to an embodiment of this application, a method for generating video is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0323] The method embodiment provided in Embodiment 8 of this application can be executed in a mobile terminal, computer terminal or similar computing device. Figure 17 This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a video generation method according to an embodiment of this application, such as... Figure 17 As shown, the computer terminal 170 (or mobile device) may include one or more processors 1702 (shown as 1702a, 1702b, ..., 1702n in the figure) 1702 (processor 1702 may include, but is not limited to, a microprocessor (MCU) or a programmable gate array (FPGA), etc.), a memory 1704 for storing data, and a transmission device 1706 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a Universal Serial Bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 17 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, the computer terminal 170 may also include... Figure 17 The more or fewer components shown, or having the same Figure 17 The different configurations shown.
[0324] Figure 17 The hardware structure block diagram shown can serve not only as an exemplary block diagram of the aforementioned computer terminal 170 (or mobile device), but also as an exemplary block diagram of the aforementioned server. In one optional embodiment, Figure 18 The use of the above is illustrated in a block diagram. Figure 17 The computer terminal 170 (or mobile device) shown is an embodiment of a computing node in computing environment 1801.
[0325] Figure 18 This is a structural block diagram of the computing environment for a video generation method according to an embodiment of this application, such as... Figure 18As shown, computing environment 1801 includes multiple compute nodes (such as servers) running on a distributed network (represented in the diagram as 1810-1, 1810-2, ...). Each compute node contains local processing and memory resources, and end user 1802 can remotely run applications or store data within computing environment 1801. Applications can be provided as multiple services 1820-1, 1820-2, 1820-3, and 1820-4 within computing environment 1801, representing services "F", "G", "I", and "H", respectively.
[0326] End user 1802 can provide and access services through a web browser or other software application on the client. In some embodiments, the provisioning and / or requests of end user 1802 can be provided to ingress gateway 1830. Ingress gateway 1830 may include a corresponding agent to handle the provisioning and / or requests for services (one or more services provided in computing environment 1801).
[0327] The service is provided or deployed based on various virtualization technologies supported by the Computing Environment 1801. In some embodiments, the service may be provided based on Virtual Machine (VM)-based virtualization, container-based virtualization, and / or similar methods. VM-based virtualization can simulate a real computer by initializing a virtual machine, executing programs and applications without directly accessing any actual hardware resources. While the machine is virtualized by a virtual machine, container-based virtualization can launch containers to virtualize an entire operating system so that multiple workloads can run on a single instance of the operating system.
[0328] In one embodiment based on container virtualization, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). For example, such as Figure 18 As shown, service 1820-2 can be equipped with one or more Pods 1840-1, 1840-2, ..., 1840-N (collectively referred to as Pods). A Pod can include a proxy 1845 and one or more containers 1842-1, 1842-2, ..., 1842-M (collectively referred to as containers). One or more containers within a Pod handle requests related to one or more corresponding functions of the service. Proxy 1845 typically controls service-related network functions such as routing and load balancing. Other services can also be equipped with Pods similar to Pods.
[0329] During operation, executing a user request from end user 1802 may require invoking one or more services in computing environment 1801, and executing one or more functions of one service may require invoking one or more functions of another service. For example... Figure 18As shown, service "F" 1820-1 receives user requests from terminal user 1802 from ingress gateway 1830. Service "F" 1820-1 can call service "G" 1820-2, and service "G" 1820-2 can request service "I" 1820-3 to perform one or more functions.
[0330] The aforementioned computing environment can be a cloud computing environment, where resource allocation is managed by cloud services, allowing functionality development without needing to consider implementation, adjustment, or server scaling. This computing environment allows developers to execute event-responsive code without building or maintaining complex infrastructure. Services can be partitioned into a set of functions that can automatically and independently scale, rather than scaling a single hardware device to handle potential loads.
[0331] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0332] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0333] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0334] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0335] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0336] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD)) for displaying information to the user; a monitor; and a keyboard and pointing device (e.g., a mouse or pathball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0337] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include Local Area Networks (LANs), Wide Area Networks (WANs), and the Internet.
[0338] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0339] It should be noted that the sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0340] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0341] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of units or modules may be electrical or other forms.
[0342] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0343] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0344] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory, random access memory, portable hard drive, magnetic disk, or optical disk.
[0345] The above are merely preferred embodiments of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for generating video, characterized in that, include: Identify the image to be processed; Using the initial noise information, the image to be processed is subjected to noise addition processing to obtain an initial video frame sequence; The initial video frame sequence is input into the video generation model, wherein the video generation model is trained based on video samples; Using video description information, the video generation model is guided to predict noise in the initial video frame sequence to obtain predicted noise information, and the predicted noise information is used to denoise the initial video frame sequence to obtain a target video frame sequence. The video description information is used to describe the video content that matches the image to be processed. A target video is generated based on the target video frame sequence, wherein the target video includes the video content; In the initial video frame sequence, different video frames correspond to the predicted noise information, and the different video frames correspond to the initial noise information. The different video frames are temporally correlated. The method further includes: correcting the predicted noise information based on first difference noise information and second difference noise information to obtain target noise information. The first difference noise information is used to represent the noise difference between the predicted noise information of the first video frame in the different video frames and the initial noise information. The second difference noise information is used to represent the noise difference between the predicted noise information of the video frames other than the first video frame in the different video frames and the initial noise information. The method of denoising the initial video frame sequence using the predicted noise information to obtain the target video frame sequence includes: using the target noise information to perform a denoising diffusion implicit operation on the features of the initial video frame sequence to obtain the features of the target video frame sequence.
2. The method according to claim 1, characterized in that, Using initial noise information, the image to be processed is subjected to noise addition processing to obtain an initial video frame sequence, including: Extract image features from the image to be processed that have the same data dimension as the initial noise information; Using the initial noise information, the image features are subjected to noise-adding processing to obtain the features of the initial video frame sequence.
3. The method according to claim 2, characterized in that, Using the initial noise information, the image features are subjected to noise addition processing to obtain the features of the initial video frame sequence, including: Using the initial noise information, the image features are subjected to noise addition processing at multiple time steps to obtain features of the initial video frame sequence corresponding to the multiple time steps.
4. The method according to claim 3, characterized in that, The plurality of time steps includes a first time step and a second time step, wherein the second time step is the time step preceding the first time step. The video generation model is guided to predict noise in the initial video frame sequence using video description information to obtain predicted noise information. The predicted noise information is then used to denoise the initial video frame sequence to obtain a target video frame sequence, including: The processing steps involve using the video description information to guide the video generation model to perform noise prediction on the features of the initial video frame sequence corresponding to the first time step, thereby obtaining the predicted noise information corresponding to the first time step, and using the predicted noise information corresponding to the first time step to perform denoising processing on the features of the initial video frame sequence corresponding to the first time step, thereby obtaining the features of the target video frame sequence corresponding to the second time step. The determination step involves determining that if the second time step has a previous time step among the plurality of time steps, then the previous time step among the plurality of time steps is determined as the second time step, and the features of the target video frame sequence corresponding to the second time step are determined as the features of the initial video frame sequence corresponding to the first time step. The process is then repeated until the second time step is the first time step among the plurality of time steps.
5. The method according to claim 4, characterized in that, Based on the target video frame sequence, the target video is generated, including: If the second time step is the first time step of the plurality of time steps, then the features of the target video frame sequence corresponding to the second time step are decoded to obtain the target video.
6. The method according to claim 1, characterized in that, The initial video frame sequence is denoised using the predicted noise information to obtain the target video frame sequence, including: The predicted noise information is adjusted using the initial noise information to obtain the target noise information; The features of the initial video frame sequence are denoised using the target noise information to obtain the features of the target video frame sequence.
7. The method according to claim 6, characterized in that, The target noise information is obtained by adjusting the predicted noise information using the initial noise information, including: Obtain the difference noise information between the predicted noise information and the initial noise information; The predicted noise information is adjusted using the differential noise information to obtain the target noise information.
8. The method according to claim 7, characterized in that, Obtaining the difference noise information between the predicted noise information and the initial noise information includes: Obtain the first difference noise information between the predicted noise information corresponding to the first video frame in the different video frames and the corresponding initial noise information; Obtain the second difference noise information between the predicted noise information corresponding to the target video frame in the different video frames and the corresponding initial noise information, wherein the target video frame is any video frame other than the first video frame in the different video frames.
9. The method according to claim 8, characterized in that, Based on the first and second difference noise information, the predicted noise information is corrected to obtain the target noise information, including: The first differential noise information is weighted using a first weight, and the second differential noise information is weighted using a second weight, wherein the second weight is negatively correlated with the first weight; The predicted noise information corresponding to the target video frame is adjusted using the weighted first difference noise information and the weighted second difference noise information to obtain the target noise information corresponding to the target video frame.
10. The method according to claim 9, characterized in that, The method further includes: Adjust either the first weight or the second weight.
11. The method according to claim 8, characterized in that, Using the differential noise information, the predicted noise information is adjusted to obtain the target noise information, including: Using the first difference noise information, the predicted noise information corresponding to the first video frame is adjusted to obtain the target noise information corresponding to the first video frame.
12. The method according to any one of claims 1 to 11, characterized in that, The method further includes: The data dimension of the noise information is determined based on the number of video frames corresponding to the target video to be generated.
13. A method for generating a video, characterized in that, Applied to e-commerce platforms, including: The image to be processed is identified, wherein the image content of the image to be processed includes the object to be traded; Using the initial noise information, the image to be processed is subjected to noise addition processing to obtain an initial video frame sequence; The initial video frame sequence is input into the video generation model, wherein the video generation model is trained based on video samples; Using video description information, the video generation model is guided to predict noise in the initial video frame sequence to obtain predicted noise information, and the predicted noise information is used to denoise the initial video frame sequence to obtain a target video frame sequence. The video description information is used to describe the video content that matches the image to be processed. Based on the target video frame sequence, a motion effect video is generated, wherein the motion effect video includes the video content and is used to display the transaction object; The animated video will be pushed to a media platform for playback; In the initial video frame sequence, different video frames correspond to the predicted noise information, and the different video frames correspond to the initial noise information. The different video frames are temporally correlated. The method further includes: correcting the predicted noise information based on first difference noise information and second difference noise information to obtain target noise information. The first difference noise information is used to represent the noise difference between the predicted noise information of the first video frame in the different video frames and the initial noise information. The second difference noise information is used to represent the noise difference between the predicted noise information of the video frames other than the first video frame in the different video frames and the initial noise information. The method of denoising the initial video frame sequence using the predicted noise information to obtain the target video frame sequence includes: using the target noise information to perform a denoising diffusion implicit operation on the features of the initial video frame sequence to obtain the features of the target video frame sequence.
14. A method for generating a video, characterized in that, include: In response to input commands applied to the user interface, the image to be processed is displayed on the user interface. In response to a video generation command applied to the operation interface, a target video including video content matching the image to be processed is displayed on the operation interface. The target video is generated based on a target video frame sequence, which is obtained by a video generation model using predicted noise information to denoise an initial video frame sequence. The video generation model is trained based on video samples. The predicted noise information is obtained by using video description information to guide the video generation model to predict noise in the initial video frame sequence. The video description information describes the video content. The initial video frame sequence is obtained by adding noise to the image to be processed using initial noise information. In this initial video frame sequence, different video frames correspond to the predicted noise information, and the different video frames correspond to the initial noise information. The different video frames are temporally correlated. The predicted noise information is used to correct the target noise information based on the first difference noise information and the second difference noise information. The first difference noise information represents the noise difference between the predicted noise information of the first video frame and the initial noise information. The second difference noise information represents the noise difference between the predicted noise information of the video frames other than the first video frame and the initial noise information. The feature of the target video frame is obtained by performing a denoising diffusion implicit operation on the feature of the initial video frame sequence using the target noise information.
15. A method for generating a video, characterized in that, include: Display the image to be processed on the screen of a virtual reality (VR) device or an augmented reality (AR) device; Using the initial noise information, the image to be processed is subjected to noise addition processing to obtain an initial video frame sequence; The initial video frame sequence is input into the video generation model in the VR device or the AR device, wherein the video generation model is trained based on video samples; Using video description information, the video generation model is guided to predict noise in the initial video frame sequence to obtain predicted noise information, and the predicted noise information is used to denoise the initial video frame sequence to obtain a target video frame sequence. The video description information is used to describe the video content that matches the image to be processed. A target video is generated based on the target video frame sequence, wherein the target video includes the video content; Drive the VR device or the AR device to display the target video; In the initial video frame sequence, different video frames correspond to the predicted noise information, and the different video frames correspond to the initial noise information. The different video frames are temporally correlated. The method further includes: correcting the predicted noise information based on first difference noise information and second difference noise information to obtain target noise information. The first difference noise information is used to represent the noise difference between the predicted noise information of the first video frame in the different video frames and the initial noise information. The second difference noise information is used to represent the noise difference between the predicted noise information of the video frames other than the first video frame in the different video frames and the initial noise information. The method of denoising the initial video frame sequence using the predicted noise information to obtain the target video frame sequence includes: using the target noise information to perform a denoising diffusion implicit operation on the features of the initial video frame sequence to obtain the features of the target video frame sequence.
16. A video generation system, characterized in that, include: E-commerce platforms are used to identify images to be processed; Using the initial noise information, the image to be processed is subjected to noise addition processing to obtain an initial video frame sequence; The initial video frame sequence is input into a video generation model, which is trained based on video samples. Using video description information, the video generation model is guided to predict noise in the initial video frame sequence to obtain predicted noise information. This predicted noise information is then used to denoise the initial video frame sequence to obtain a target video frame sequence. The video description information describes the video content that matches the image to be processed. Based on the target video frame sequence, a target video is generated, which includes the video content. A media platform for playing the target video; In the initial video frame sequence, different video frames correspond to the predicted noise information, and the different video frames correspond to the initial noise information. The different video frames are temporally related. The e-commerce platform is further configured to perform the following steps: based on the first difference noise information and the second difference noise information, the predicted noise information is corrected to obtain the target noise information. The first difference noise information is used to represent the noise difference between the predicted noise information of the first video frame in the different video frames and the initial noise information. The second difference noise information is used to represent the noise difference between the predicted noise information of the video frames other than the first video frame in the different video frames and the initial noise information. The method of denoising the initial video frame sequence using the predicted noise information to obtain the target video frame sequence includes: using the target noise information to perform a denoising diffusion implicit operation on the features of the initial video frame sequence to obtain the features of the target video frame sequence.
17. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method according to any one of claims 1 to 15.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the storage medium is located to perform the method according to any one of claims 1 to 15.
19. A computer program product, characterized in that, Includes computer instructions that, when executed by a processor, implement the method described in any one of claims 1 to 15.
Citation Information
Patent Citations
Video generation method, and method and device for training video generation model
CN116863003A
Video editing method and device, electronic equipment and storage medium
CN116980541A
Data processing method and device
CN117217284A