Method and apparatus for generating video content, device, and storage medium

By constructing 3D point cloud data and controlling the movement of a virtual camera, video content that matches preset camera movements is generated, solving the problem of insufficient camera movement control in traditional video generation solutions and improving the expressiveness and diversity of video content.

WO2025245877A1PCT designated stage Publication Date: 2025-12-04BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/096829
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-31
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Traditional video generation methods lack effective control over camera movement, which limits the expressiveness and diversity of video content.

Method used

Three-dimensional point cloud data is constructed based on target images and depth information. Images in multiple states are generated according to the motion trajectory of the virtual camera. Video content is generated using the target model, and the camera movement is controlled to match the preset camera movement.

Benefits of technology

It improves the quality of generated video content, achieves matching with preset camera movements, and enhances the expressiveness and diversity of video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024096829_04122025_PF_FP_ABST
    Figure CN2024096829_04122025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a method and apparatus for generating video content, a device, and a storage medium. The method provided herein comprises: constructing three-dimensional point cloud data on the basis of a target image and depth information associated with the target image; on the basis of the motion trajectory of a virtual camera and on the basis of the three-dimensional point cloud data, generating a first group of images corresponding to a plurality of states of the virtual camera; and on the basis of the first group of images and a prompt item, using a target model to generate video content. In this way, according to the embodiments of the present disclosure, video content that matches a preset camera movement can be generated, thereby improving the quality of the generated video content.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, device and storage medium for generating video content TECHNICAL FIELD

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to a method, apparatus, device and computer-readable storage medium for generating video content. BACKGROUND

[0002] In recent years, video generation technology has made significant progress, especially video synthesis technology based on text prompts or image inputs. These technologies can generate videos with rich dynamic content through machine learning models (e.g., diffusion models). However, traditional video generation schemes lack effective control over camera movements when generating videos, which limits the expressiveness and diversity of video content.

[0003] SUMMARY

[0004] In a first aspect of the present disclosure, a method for generating video content is provided. The method comprises: constructing three-dimensional point cloud data based on a target image and depth information associated with the target image; generating a first set of images corresponding to a plurality of states of a virtual camera based on the three-dimensional point cloud data according to a motion trajectory of the virtual camera; and generating video content using a target model based on the first set of images and a prompt item.

[0005] In a second aspect of the present disclosure, an apparatus for generating video content is provided. The apparatus comprises: a point cloud construction module configured to construct three-dimensional point cloud data based on a target image and depth information associated with the target image; an image generation module configured to generate a first set of images corresponding to a plurality of states of a virtual camera based on the three-dimensional point cloud data according to a motion trajectory of the virtual camera; and a video generation module configured to generate video content using a target model based on the first set of images and a prompt item.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The device comprises at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to perform the method of the first aspect.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium has stored thereon a computer program executable by a processor to implement the method of the first aspect.

[0008] It should be understood that the content described in this section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent through the following description. Attached Figure Description

[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0010] Figure 1 shows a schematic diagram of an example environment in which some embodiments of the present disclosure can be implemented;

[0011] Figure 2 shows a flowchart of a process for generating video content according to some embodiments of the present disclosure;

[0012] Figure 3 illustrates an example architecture of a video generation system according to some embodiments of the present disclosure;

[0013] Figure 4 shows a schematic structural block diagram of an example apparatus for generating video content according to some embodiments of the present disclosure; and

[0014] Figure 5 shows a block diagram of an electronic device capable of implementing several embodiments of the present disclosure. Detailed Implementation

[0015] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0016] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0017] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0018] The data of the user, the acquisition and / or use of the data, etc. can be involved in the embodiments of the present disclosure. These aspects all comply with the corresponding laws and regulations and relevant provisions. In the embodiments of the present disclosure, the collection, acquisition, processing, processing, forwarding, use, etc. of all data are performed on the premise that the user is aware of and confirms. Accordingly, when implementing the embodiments of the present disclosure, the type of data or information that can be involved, the use range, the use scenario, etc. should be notified to the user and the authorization of the user should be obtained through appropriate means according to the relevant laws and regulations. The specific notification and / or authorization manner can vary according to the actual situation and application scenario, and the scope of the present disclosure is not limited in this aspect.

[0019] In the present specification and embodiments, if the scheme involves processing of personal information, the processing is performed on the premise of having a legal basis (for example, obtaining the consent of the subject of personal information, or being necessary for the performance of a contract, etc.), and is performed only within the prescribed or agreed range. The user refuses to process personal information other than the necessary information required for the basic function, which does not affect the user's use of the basic function.

[0020] As briefly mentioned above, some video generation techniques can utilize machine learning models (e.g., diffusion models) to generate videos with rich dynamic content. However, traditional video generation schemes lack effective control over camera movements when generating videos, which limits the expressiveness and diversity of video content.

[0021] Embodiments of the present disclosure propose a scheme for generating video content. According to the scheme, three-dimensional point cloud data can be constructed based on a target image and depth information associated with the target image. Further, a first set of images corresponding to multiple states of a virtual camera can be generated based on the three-dimensional point cloud data according to a motion trajectory of the virtual camera. Accordingly, video content can be generated using a target model based on the first set of images and a prompt.

[0022] In this way, embodiments of the present disclosure can generate video content that matches a preset camera movement, thereby improving the quality of the generated video content.

[0023] Various example implementations of the scheme are described in further detail below in conjunction with the accompanying drawings.

[0024] Example Environment

[0025] FIG. 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in FIG. 1, the example environment 100 can include an electronic device 110.

[0026] In some embodiments, the electronic device 110 can obtain an input target image 120 and a prompt item 130. Further, the electronic device 110 can utilize a video generation system 140 to generate video content 150 based on the target image 120 and the prompt item 130. Such a prompt item 130 may, for example, include textual content for describing the video content 150 to be generated, also as a textual prompt item or prompt word, etc.

[0027] In some embodiments, such a video generation system 140 may, for example, be deployed locally at the electronic device 110, or can be deployed at a suitable remote device.

[0028] In some examples, the electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a tablet computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a palmtop computer, a portable gaming terminal, a VR / AR device, a Personal Communication System (PCS) device, a personal navigation device, a Personal Digital Assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. In some embodiments, the electronic device 110 can also support any type of interface to a user (such as “wearable” circuitry, etc.).

[0029] It should be understood that the structures and functions of the various elements in the environment 100 are described for illustrative purposes only, and do not imply any limitation on the scope of the present disclosure.

[0030] Some example embodiments of the present disclosure will be described below with continued reference to the drawings.

[0031] Example process

[0032] FIG. 2 shows a flowchart of an example process 200 for generating video content, according to some embodiments of the present disclosure. The process 200 may, for example, be implemented at the electronic device 110 as shown in FIG. 1. The process 200 will be described below with reference to FIG. 1.

[0033] As shown in FIG. 2, at block 210, the electronic device 110 constructs three-dimensional point cloud data based on a target image and depth information associated with the target image.

[0034] FIG. 3 illustrates a block diagram of an example video generation system 300, according to some embodiments of the present disclosure. Such a video generation system 300 can correspond to the video generation system 140 shown in FIG. 1.

[0035] In particular, as shown in FIG. 3, the electronic device 110 can obtain the target image 120, and can further determine depth information 305 associated with the target image 120. Such depth information 305 may, for example, include a depth map corresponding to the target image 120 to indicate a depth value for each pixel.

[0036] In some embodiments, the target image 120 may, for example, include a depth image, and its depth information can be determined directly based on the depth image. In some embodiments, the target image 120 may, for example, include a two-dimensional image, and the electronic device 110 may, for example, utilize a depth estimation module to determine the depth information 305. For example, the depth estimation module can utilize any appropriate depth estimation model to determine the depth information 305 corresponding to the target image 120.

[0037] Further, as shown in FIG. 3, the electronic device 110 can convert a plurality of pixels in the target image 120 to a three-dimensional space based on the target image 120 and the depth information 305 to generate three-dimensional point cloud data 315.

[0038] In particular, the construction process of the three-dimensional point cloud data 315 can be represented as:

[0039] wherein, represents the three-dimensional point cloud data 315, φ represents a mapping function from a two-dimensional image space to a three-dimensional space, represents the target image 120, D0represents the depth information 305, K represents internal parameters of a virtual camera corresponding to the target image 120, and P0represents external parameters of the virtual camera.

[0040] With reference back to FIG. 2, at block 220, the electronic device 110 generates, based on the three-dimensional point cloud data, a first set of images corresponding to a plurality of states of the virtual camera according to a motion trajectory of the virtual camera.

[0041] As shown in FIG. 3, the electronic device 110 can determine a motion trajectory 320 of the virtual camera, which may, for example, be represented as [K, P i ], where K is an internal parameter of the virtual camera, and P i represents external parameters of the virtual camera at a plurality of states. Such external parameters may, for example, include parameters such as a position, a rotation, a translation, and a pose of the virtual camera in the three-dimensional space.

[0042] Further, the electronic device 110 can project the three-dimensional point cloud data to corresponding two-dimensional planes based on the extrinsic parameters of the virtual camera in the plurality of states to generate a corresponding first set of images. This process can be represented as, for example:

[0043] where I i represents an image corresponding to the i-th state of the virtual camera, and ψ represents a projection function.

[0044] In this way, the electronic device 110 can obtain a set of images corresponding to different states of the camera, and the dynamic switching between such images can correspond to the dolly effect of the camera, e.g., zoom, pan, rotation, etc.

[0045] With reference to FIG. 2, at block 230, the electronic device 110 generates video content using the target model based on the first set of images and the prompt item.

[0046] In some embodiments, as shown in FIG. 3, the electronic device 110 can directly input the projected first set of images as input to a video generation model 335 for generating a corresponding second set of images as the plurality of video frames of the video content 150, for example.

[0047] Given that the three-dimensional point cloud data 315 has holes at certain location points, the first set of images generated by projecting the three-dimensional point cloud data 315 can have empty pixels. In some scenarios, to improve the generation quality of the video content, the electronic device 110 can further process one or more images in the first set of images using an image inpainting model to generate a repaired set of images (also referred to as a third set of images). For example, the image inpainting model can fill in the empty pixels in the images using appropriate image inpainting techniques.

[0048] In some embodiments, the electronic device 110 can input such third set of images as input to the video generation model 335 to generate the second set of images.

[0049] In addition, since the filled-in empty pixels can be misaligned in the three-dimensional point cloud space, the electronic device 110 can further project the repaired third set of images to the three-dimensional space corresponding to the three-dimensional point cloud data 315.

[0050] Further, the electronic device 110 can update the three-dimensional point cloud data 315 by aligning the plurality of locations corresponding to the third set of images in the three-dimensional space. This alignment process can be represented as, for example:

[0051] where and represent the modified third set of images and the corresponding depth information, respectively, di denotes the depth parameter to be optimized, M denotes and the overlapping region, ||·|| denotes the L1 loss.

[0052] Based on such a manner, embodiments of the present disclosure can find the optimal depth coefficient, so that the point cloud representations of the front and rear images in the overlapping region are as consistent as possible. Thus, embodiments of the present disclosure can correct the relative depth error caused by monocular depth estimation, thereby maintaining the coherence and consistency of objects and scenes in the generated video.

[0053] Additionally, the electronic device 110 can generate a fourth set of images corresponding to a plurality of states of the virtual camera based on the updated three-dimensional point cloud data. The process can be represented as:

[0054] Taking FIG. 3 as an example, the fourth set of images generated by the electronic device 110 may, for example, be a set of images 325 as shown in FIG. 3. Accordingly, the electronic device 110 can process the set of images 325 and the prompt item 130 using a video generation model 335 to generate a second set of images as a plurality of video frames.

[0055] In some embodiments, the video generation model 335 may, for example, include a diffusion model. Specifically, in the noise adding stage, the diffusion model can generate a set of hidden features corresponding to the fourth set of images by adding noise satisfying a preset distribution. Further, in the noise removing stage, the diffusion model can generate the corresponding second set of images based on the set of hidden features. The process can be represented as:

[0056] Specifically, formula (5) is used to obtain a noise latent representation from the rendered image sequence V0325 through a forward diffusion process 330. is a variance used in a DDIM (Denoising Diffusion Implicit Models) scheduler; ∈ is random noise sampled from a standard normal distribution, used to perturb the latent representation and increase the diversity of generated images; t0 is a time step in the diffusion process, which determines the strength of the noise.

[0057] Formula (6) describes the steps of generating a video using the noise latent representation through a reverse diffusion process. is the denoised latent representation at time step t-1. a t is a scheduling parameter used to control the denoising process. ∈ is random noise sampled at each step. is the noise prediction part of the diffusion model. s t Determining whether the denoising process is deterministic or probabilistic, typically set to 1 to encourage diversity in the generated results. t is the time step in the diffusion process.

[0058] In some embodiments, the electronic device 110 may, for example, also balance the realism and diversity of the video content by controlling the time step t0. In the generation process, the time step t0 used to generate the noise latent representation is a key factor that affects this trade-off. A larger t0 value will make the generated video closer to the original guided camera motion, but may sacrifice the dynamic and reasonableness of the video content. Conversely, a smaller t0 value can produce more reasonable videos, but may not fully comply with the desired camera motion.

[0059] In this way, embodiments of the present disclosure are able to control the camera motion in the video generation process by operating noise in latent space, achieving video generation that controls camera motion without additional training.

[0060] Further, the electronic device 110 may, for example, generate the video content 150 by sequentially combining the second set of images such that the plurality of video frames of the video content 150 are able to correspond to the plurality of states [K, P i ] of the virtual camera. It should be understood that such correspondence is intended to indicate that the video frames have similar panning control, and is not intended to indicate that the video frames correspond completely to the extrinsic parameters of the virtual camera in the corresponding state.

[0061] In some embodiments, the electronic device 110 may, for example, also generate the final video content 150 by adding other appropriate elements such as audio content, subtitles, etc.

[0062] Thus, embodiments of the present disclosure are able to generate videos with rich dynamic content and high fidelity using explicit rearrangement of image layouts in 3D point cloud space and layout priors of noise latent representations, while maintaining the efficiency of the processing process and the robustness of the model.

[0063] In addition, embodiments of the present disclosure are able to support complex hybrid camera motion and can be widely applied in the fields of 3D video generation, virtual reality content creation, etc., thereby improving the flexibility and innovation of video content creation.

[0064] Example apparatus and devices

[0065] Embodiments of the present disclosure also provide a corresponding apparatus for implementing the above method or process. FIG. 4 shows a schematic structural block diagram of an example apparatus 400 for generating video content according to certain embodiments of the present disclosure. The apparatus 400 can be implemented as or included in an electronic device. Various modules / components in the apparatus 400 can be implemented by hardware, software, firmware, or any combination thereof.

[0066] As shown in FIG. 4, the apparatus 400 includes a point cloud construction module 410 configured to construct three-dimensional point cloud data based on a target image and depth information associated with the target image; an image generation module 420 configured to generate a first set of images corresponding to a plurality of states of a virtual camera based on the three-dimensional point cloud data according to a motion trajectory of the virtual camera; and a video generation module 430 configured to generate video content using a target model based on the first set of images and a prompt.

[0067] In some embodiments, the image generation module 420 is further configured to determine a plurality of sets of extrinsic parameters of the virtual camera at the plurality of states based on the motion trajectory of the virtual camera, and generate the first set of images corresponding to the plurality of sets of extrinsic parameters by projecting the three-dimensional point cloud data.

[0068] In some embodiments, the video generation module 430 is further configured to process at least one image in the first set of images using an image inpainting model to generate a third set of images, and generate a second set of images using the target model based on the third set of images and the prompt.

[0069] In some embodiments, the video generation module 430 is further configured to project the third set of images to a three-dimensional space corresponding to the three-dimensional point cloud data, update the three-dimensional point cloud data by aligning a plurality of positions corresponding to the third set of images in the three-dimensional space, generate a fourth set of images corresponding to the plurality of states of the virtual camera based on the updated three-dimensional point cloud data, and provide the fourth set of images and the prompt to the target model to generate the second set of images.

[0070] In some embodiments, the target model is a diffusion model configured to generate a set of hidden features corresponding to the fourth set of images by adding noise satisfying a preset distribution, and generate the corresponding second set of images based on the set of hidden features.

[0071] In some embodiments, the apparatus 400 further includes a depth estimation module configured to process the target image using a depth estimation model to determine the depth information.

[0072] In some embodiments, the video content includes a plurality of video frames corresponding to the plurality of states of the virtual camera.

[0073] FIG. 5 illustrates a block diagram of an electronic device 500 in which one or more embodiments of the disclosure can be implemented. It should be understood that the electronic device 500 illustrated in FIG. 5 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 500 illustrated in FIG. 5 can be used for the electronic device 110 as illustrated in FIG. 1.

[0074] As illustrated in FIG. 5, the electronic device 500 is in the form of a general electronic device. Components of the electronic device 500 can include, but are not limited to, one or more processors or processing units 510, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processing unit 510 can be a real or virtual processor and capable of executing various processing according to programs stored in the memory 520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve parallel processing capability of the electronic device 500.

[0075] The electronic device 500 typically includes a number of computer storage media. Such media can be any available media that is accessible by the electronic device 500 and includes both volatile and non-volatile media, removable and non-removable media. The memory 520 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be a removable or non-removable media and can include a machine-readable medium, such as a flash drive, a magnetic disk, or any other medium that is capable of storing information and / or data and that can be accessed by the electronic device 500.

[0076] The electronic device 500 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5, a disk drive and a disk drive interface can be provided for reading from or writing to a removable, non- volatile magnetic disk (e.g., a "hard disk that is not removable and for which access by other computer systems is typically not desired). In these cases, the disk drive can be connected to the bus by one or more data media interfaces. The memory 520 can include a computer program product 525 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.

[0077] The communication unit 540 enables communication through the communication medium with other electronic devices. Additionally, the functionality of the components of the electronic device 500 can be implemented in a single computing cluster or a plurality of computer machines capable of communicating over a communication connection. As such, the electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0078] The input device 550 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 560 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 500 can also communicate with one or more external devices (not shown), such as a storage device, a display device, etc., through the communication unit 540, as necessary, with one or more devices that enable a user to interact with the electronic device 500, or with any device (e.g., a network card, a modem, etc.) that enables the electronic device 500 to communicate with one or more other electronic devices. Such communication can be carried out via an input / output (I / O) interface (not shown).

[0079] According to an example implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, where the computer-executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, there is also provided a computer program product tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, where the computer-executable instructions are executed by a processor to implement the method described above.

[0080] Various aspects of the disclosure are now described with reference to the drawings. In general, the drawings described below are diagrammatic and schematic representations of actual or conceptual structures and processes, and are not limiting of the scope of the present disclosure. In the drawings, the same reference numerals are used to represent similar components. The embodiments of the present disclosure will be described with reference to the drawings, beginning with FIG. 1.

[0081] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including a processor executable program of instructions which implement the function / act specified in the flowchart and / or block diagram block or blocks.

[0082] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0083] The flow diagrams and block diagrams in the attached figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various implementations of the present disclosure. In this regard, each block in the flow diagrams and block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logic functions (s). It should also be noted that in some alternative implementations, the functions noted in the block can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or acts or combinations of

[0084] implementations have been described above, the description is intended to be illustrative, and not restrictive, and is not intended to exclude other implementations from the scope of the implementations disclosed herein. Many modifications and variations to the implementations described herein are possible and can be apparent to those of ordinary skill in the art. The specific implementations described herein are shown by way of example, and any person skilled in the art can realize changes of or substitutions of equivalents of certain features between the various implementations described herein, without departing from the scope and spirit of the implementations described. The choice of words in this document is intended to express the best of the concepts, practical applications or improvements to the art in the various implementations, or to enable other ordinary skilled persons in the art to understand the various implementations disclosed herein.

Claims

1. A method for generating video content, comprising: constructing three-dimensional point cloud data based on target images and depth information associated with the target images; generating, according to a motion trajectory of a virtual camera, a first set of images corresponding to a plurality of states of the virtual camera based on the three-dimensional point cloud data; and generating, based on the first set of images and a prompt, video content using a target model.

2. The method of claim 1, wherein generating, according to a motion trajectory of a virtual camera, a first set of images corresponding to a plurality of states of the virtual camera based on the three-dimensional point cloud data comprises: determining, based on the motion trajectory of the virtual camera, a plurality of extrinsic parameters of the virtual camera at the plurality of states; and generating the first set of images corresponding to the plurality of extrinsic parameters by projecting the three-dimensional point cloud data.

3. The method of claim 1, wherein generating, based on the first set of images and a prompt, video content using a target model comprises: processing at least one image in the first set of images using an image inpainting model to generate a third set of images; and generating, based on the third set of images and the prompt, a second set of images in the video content using the target model.

4. The method of claim 3, wherein generating, based on the third set of images and the prompt, a second set of images in the video content using the target model comprises: projecting the third set of images to a three-dimensional space corresponding to the three-dimensional point cloud data; updating the three-dimensional point cloud data by aligning a plurality of positions corresponding to the third set of images in the three-dimensional space; generating a fourth set of images corresponding to the plurality of states of the virtual camera based on the updated three-dimensional point cloud data; and providing the fourth set of images and the prompt to the target model to generate the second set of images.

5. The method of claim 4, wherein the target model is a diffusion model configured to: generate a set of hidden features corresponding to the fourth set of images by adding noise satisfying a preset distribution; and generate the corresponding second set of images based on the set of hidden features.

6. The method of claim 1, further comprising: processing the target images using a depth estimation model to determine the depth information.

7. The method of claim 1, wherein the video content comprises a plurality of video frames corresponding to the plurality of states of the virtual camera.

8. An apparatus for generating video content, comprising: a point cloud construction module configured to construct three-dimensional point cloud data based on target images and depth information associated with the target images; an image generation module configured to generate, according to a motion trajectory of a virtual camera, a first set of images corresponding to a plurality of states of the virtual camera based on the three-dimensional point cloud data; and a video generation module configured to generate, based on the first set of images and a prompt, video content using a target model.

9. An electronic device, comprising: at least one processing unit; and ​ ​ ​ ​ ​ ​ ​ at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the electronic device to perform the method according to any one of claims 1 to 7.

10. A computer readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video generation method and device, equipment and storage medium

    CN114387326A

  • Video generation method and device

    CN117014651A

  • Arbitrary track three-dimensional scene construction and roaming video generation method and system guided by plain text

    CN117853686A

  • Display control device, display control method, and display control program

    WO2023195301A1