Method and system for generating three-dimensional video
By introducing three-dimensional data generation methods and multi-layer four-dimensional voxel representation, combined with staged motion enhancement training, the problem of insufficient depth information acquisition in two-dimensional image generation methods is solved, precise control and natural smoothness of object movements in three-dimensional videos are achieved, and dynamic three-dimensional scenes are generated in line with the creator's intentions.
Patent Information
- Application Number
- CN202510777431.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-05
AI Technical Summary
Existing video generation methods based on two-dimensional image generation cannot accurately obtain depth information, resulting in inaccurate representation of the distance and position between objects, lack of scene diversity, blurred objects, and inability to effectively capture the precise shape and structure of objects, especially when the objects are occluded or show details on the back.
A three-dimensional data generation method is introduced. By receiving text and motion control vectors containing objects and their movements, static three-dimensional videos are generated, and the object movement is driven according to the motion control vectors. Multi-layer four-dimensional voxel representation and staged motion enhancement training are used to ensure the natural smoothness and precise control of movements.
It achieves precise control of the object movements in three-dimensional videos, and the generated video movements are smooth and natural, avoiding artifacts, meeting the creator's personalized needs and the smooth movement requirements of dynamic scenes.
Smart Images

Figure CN120602736A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to video processing, and in particular, to a method, system and computer program product for generating three-dimensional video. Background Art
[0002] Video generation technology can automate the video content creation process, reducing the time and cost of manual video production. This is particularly important for industries that require large amounts of video content, such as news, education, and advertising. By analyzing user preferences through algorithms, video generation technology can provide customized video content for each user, enhancing the user experience and meeting individual needs. In the rapidly changing information age, video generation technology can quickly respond to news events or market changes, generating relevant video content in real time to ensure the timeliness and relevance of information. For individuals or small teams without professional video production skills, video generation technology provides a simple and easy way to create high-quality video content.
[0003] Currently, the technology for generating images based on a piece of text is relatively mature, and it is possible to generate videos based on the generated two-dimensional images. This allows videos to be generated directly from a piece of text. However, video generation methods based on two-dimensional image generation technology still have many shortcomings. First, the use of two-dimensional data limits the acquisition of depth information, which directly affects the accuracy of the distance between objects in the video and the precise representation of their positions. Second, due to the limitation of viewing angle, the scenes presented in the video lack diversity, making it difficult to express complex three-dimensional spatial relationships. In addition, two-dimensional images cannot effectively capture the precise shape and structure of objects, resulting in blurred objects in the video. Moreover, when the object is occluded or the video needs to show the details of the back of the object, the method based on two-dimensional data shows obvious shortcomings.
[0004] To overcome these shortcomings, researchers have begun exploring the introduction of 3D data into the video generation process. By incorporating 3D information, the depth, shape, and structure of objects can be more accurately simulated, greatly improving the quality and realism of video generation. Summary of the Invention
[0005] The following description includes exemplary methods, systems, and computer program products that embody the present invention. However, it should be understood that, in one or more aspects, the described invention can be practiced without these specific details. In other cases, well-known structures and techniques are not shown in detail to avoid obscuring the present invention.
[0006] According to one aspect of the present invention, a method for generating a three-dimensional video is disclosed. The method includes: receiving a text describing an object and its action, as well as a motion control vector for the object; generating a static three-dimensional video of the object and its action based on the content of the text; and, based on the motion control vector, causing the object in each picture in the static three-dimensional video to move according to the motion control vector, thereby obtaining a three-dimensional video in which the object's action is controllable.
[0007] According to another aspect of the present invention, a computer system is disclosed, comprising: a memory; and at least one processor operatively coupled to the memory and configured to execute the method described above.
[0008] According to yet another aspect of the present invention, a computer program product is disclosed. The computer program product includes program instructions, and the program instructions can be executed by a computing device to enable the computing device to perform the method described above.
[0009] The innovative motion-driven dynamic 3D scene generation method proposed in this paper not only uses a local deformation module to match the motion path domain for domain-adaptive alignment, thereby ensuring natural and smooth motion, but also introduces phased motion enhancement training, which gradually increases the intensity of motion control to ensure the quality of generated content and solve the problem of image artifacts caused by large motion amplitudes. Furthermore, the present invention also achieves motion-controlled dynamic 3D scene generation, which can adapt to drastic changes in motion and meet the requirements of smooth motion, enabling the creation of dynamic 3D scenes that meet the creator's intentions. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The invention itself, as well as the mode of use, objects, features and advantages of its preferred embodiments, may be better understood by reading the following detailed description of illustrative embodiments with reference to the accompanying drawings in which:
[0011] Figure 1 An example of a three-dimensional video generated for a text "a tiger is dancing" containing a description of an object and its actions is shown;
[0012] Figure 2 A structural block diagram of a system for generating three-dimensional video according to an embodiment of the present invention is shown;
[0013] Figure 3 A flowchart of a method for generating a static 3D video of an object and its actions, used by a static 3D video generation module according to an embodiment of the present invention, is shown;
[0014] Figure 4 A schematic diagram of a four-dimensional scene is schematically shown;
[0015] Figure 5 A schematic diagram of a normalized multi-layer four-dimensional voxel model is shown schematically;
[0016] Figure 6 A flowchart of a method for generating a 3D video with controllable object motion using a second neural network model by a 3D video generation module according to an embodiment of the present invention is shown;
[0017] Figure 7 Shows the Figure 1 An example of adding motion control vectors to a three-dimensional video generated by a text describing an object and its movements, "A tiger is dancing," to obtain a dynamic three-dimensional video; and
[0018] Figure 8 A flow chart of a method for generating three-dimensional video according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0019] The following description includes exemplary methods, systems, and storage media that embody the present invention. However, it should be understood that, in one or more aspects, the described invention can be practiced without these specific details. In other cases, well-known protocols, structures, and techniques are not shown in detail in order to avoid obscuring the present invention.
[0020] Embodiments of the present invention are described below with reference to the accompanying drawings. In the following description, many specific details are set forth in order to more fully understand the present invention. However, it will be apparent to those skilled in the art that the present invention may be implemented without some of these specific details. Furthermore, it should be understood that the present invention is not limited to the specific embodiments described. On the contrary, any combination of the following features and elements may be considered to implement the present invention, regardless of whether they relate to different embodiments. Therefore, the following aspects, features, embodiments, and advantages are provided for illustrative purposes only and should not be considered as elements or limitations of the appended claims unless expressly set forth in the claims.
[0021] As mentioned earlier, because video generation methods based on 2D image generation technology still have many shortcomings, researchers have begun exploring the integration of 3D data into the video generation process. Some methods have been developed, such as the Make-A-Video3D (MAV3D) method, that generate 3D videos from text describing an object and its actions.
[0022] Figure 1The example of a 3D video generated for a text describing an object and its actions, "A tiger is dancing," is shown. The 3D video generated by this method includes a perspective 1 video and a perspective 2 video. The perspective 1 video contains three pictures showing the front view of the object "tiger" performing the action "dancing"; the perspective 2 video contains three pictures showing the back view of the object "tiger" performing the action "dancing." For the sake of simplicity, Figure 1 Only some exemplary images of the generated video are shown. Those skilled in the art will appreciate that Perspective 1 and Perspective 2 videos can contain more than three images, depending on the length of the generated video and the number of video frames per unit time. Perspective 1 and Perspective 2 videos utilize two perspectives to create a 3D visual effect.
[0023] While existing technologies can successfully create impressive three-dimensional visual effects, these videos often lack precise control over motion during the generation process. This limitation means that the generated videos may not fully meet the user's specific needs or expectations. Users hope to be able to customize the motion elements in the three-dimensional video according to their own creativity and needs, thereby creating more personalized video content that meets the requirements of specific scenarios. This user-friendly video generation method not only improves the flexibility of video generation, but also provides content creators with greater creative space, allowing them to express their creativity and ideas more freely. Therefore, a method and system for generating three-dimensional videos that can control the motion of objects in the video are needed to meet the above-mentioned user needs.
[0024] The present invention proposes a method and system for generating three-dimensional videos, designed to allow users to drive the generation of video content according to a specified motion trajectory during the generation of a three-dimensional video. The method includes receiving a text description of an object and its motion, as well as a motion control vector for the object; generating a static three-dimensional video of the object and its motion based on the text; and, based on the motion control vector, causing each image in the static three-dimensional video to move according to the motion control vector, thereby generating a three-dimensional video with controllable object motion. This method enhances the controllability of the generated results by adding user-input object motion control vectors, allowing users to more precisely control the dynamic elements of the video. The present invention not only ensures the natural and smooth movement of objects in the video, but also achieves the generation of dynamic three-dimensional scenes with controlled motion. This can adapt to drastic changes in motion and meet the requirements of smooth motion, enabling the creation of dynamic three-dimensional scenes that meet the creator's intentions.
[0025] In addition, the present invention introduces a multi-layer four-dimensional voxel representation method for depicting dynamic three-dimensional environments. The local deformation module proposed in the present invention ensures smooth and natural transitions along the motion path, improving the consistency of dynamic scenes. The present invention proposes a training strategy through staged motion enhancement, which effectively avoids potential artifacts caused by changes in motion amplitude and ensures the quality of the generated results. The present invention introduces a multi-layer four-dimensional voxel representation method and motion loss to achieve precise control over the generation of dynamic three-dimensional scenes.
[0026] Figure 2 FIG. 2 shows a structural block diagram of a system 200 for generating three-dimensional video according to an embodiment of the present invention. Figure 2 As shown, system 200 includes a static 3D video generation module 210 and a 3D video generation module 220. Static 3D video generation module 210 is configured to receive a text 201 describing an object and its motion, and based on the content of text 201, generate a static 3D video 203 depicting the object and its motion. 3D video generation module 220 is configured to receive a motion control vector 205 for the object and the static 3D video 203, and based on the motion control vector 205, for each image in the static 3D video 203, cause the object in the image to move according to the motion control vector 205, thereby generating a 3D video 207 in which the object's motion is controllable.
[0027] First, consider how the static 3D video generation module 210 generates a static 3D video 203 about an object and its movements based on the content of the text 201 . Figure 3 FIG. 3 is a flow chart showing a method 300 for generating a static 3D video of an object and its actions by the static 3D video generation module 210 according to an embodiment of the present invention. Figure 3 The method 300 includes: in step 310, the static three-dimensional video generation module 210 generates an initial three-dimensional video from the static multi-layer four-dimensional image, and the initial three-dimensional video is represented by an initial multi-layer four-dimensional voxel model; in step 320, the static three-dimensional video generation module 210 optimizes the first frame scene of the initial three-dimensional video by using a first neural network model to obtain a static three-dimensional video 201, wherein the static three-dimensional video generation module 210 uses the loss function of the controlled fractional distillation sampling of the Vincent graph diffusion model as the loss function of the first neural network model, and the static three-dimensional video corresponds to a static multi-layer four-dimensional voxel model.
[0028] Regarding the above step 310 , in existing 3D video generation technologies, a 4D scene is generally used. Figure 4 A schematic diagram of a four-dimensional scene is shown schematically. Figure 4, the four-dimensional scene contains three-dimensional space (x, y, z) and one-dimensional time (t). Each dimension of the four-dimensional scene usually uses six planes with a coordinate range of -1 to 1 and the center of the plane as the origin, namely, six planes composed of xy, xz, yz, xt, yt, and zt, to capture the key features of the scene. Using the sampling method of the existing volume rendering process to sample continuous points in the four-dimensional scene, discrete sampling points and their coordinates can be obtained. In the four-dimensional space, a normalized sampling point q is obtained: q = (x, y, z, t), which represents the position at a specific time t and three-dimensional space coordinates (x, y, z). The existing technology uses the above-mentioned six-plane feature representation method to compress the information of four dimensions into two dimensions, that is, the four-dimensional space feature is compressed into the feature of one plane in the six planes, which will cause the information of the two dimensions to be strongly coupled, resulting in large artifacts in the generated video result. In order to overcome this problem, the present invention discloses a new method that can be used in step 310 to generate an initial three-dimensional video, which is represented by an initial multi-layer four-dimensional voxel model. When representing a multi-layer 4D voxel model, it is necessary to normalize the time and space dimensions of the 4D scene, ensuring that the time dimension is normalized to t∈[0,1] and the space dimension is normalized to (x,y,z)∈[0,1] 3 This ensures the consistency of the sampling points in the four-dimensional space, thus providing a solid foundation for subsequent feature calculation and video generation. Figure 5 Schematic diagram of a normalized multi-layer 4D voxel model is shown, where the coordinates of each voxel range from 0 to 1. Figure 5 , the multi-layer four-dimensional voxel model contains multiple layers H1, H2, ..., H i , each layer contains a 4D voxel, each 4D voxel Si has a different resolution r Si , for example, r of H1 S1 is 128, H2's r S2 is 64, r Si The larger the size, the higher the resolution. Each four-dimensional voxel contains three-dimensional space voxels and one-dimensional time voxels. The four-dimensional voxels of each layer are composed of r Si ×r Si ×r Si ×r Si feature vectors, and all feature vectors of the four-dimensional voxels of each layer are initially random feature vectors. The sampling method of the existing volume rendering process can still be used to sample the continuous points in the four-dimensional scene to obtain the coordinates of the discrete sampling points. Then, by using the existing technology, by calculating which element of each layer of four-dimensional voxels the sampling points (the number of sampling points is not less than the number of pixels of the rendered video) are located in, the eigenvalue si(q) of the sampling point q on each layer of four-dimensional voxel feature Si can be effectively calculated, that is,
[0029]
[0030] For the sampling point q, its multi-layer features can be obtained by splicing four-dimensional voxel features of different resolutions: where N C is the number of 4D voxels of different resolutions, and the resolution of each 4D voxel is r Si ,and (i=2,...,N C In one embodiment, it can be set to r S1 =8, and N C =5 By inputting the multi-layer feature vectors of all sample points into an existing renderer, the initial 3D video can be rendered. The 4D scene features represented by the multi-layer 4D voxel model described above avoid the strong coupling of dimensional information caused by dimensional information compression, ensuring the independence of the feature space across different dimensions, thereby reducing artifacts in the generated initial video.
[0031] Regarding step 320, in the prior art, the loss function for controlling fractional distillation sampling of the existing Vincent graph diffusion model can be used, and an existing first neural network model can be used to optimize a common 3D video, thereby obtaining a static 3D video corresponding to the common 3D video. For example, the first neural network model can sample existing MAV3D, 4d-fy, and other neural network models. In step 320, the same method as the prior art can be used, except that the common 3D video is replaced with the initial 3D video output from step 310, thereby obtaining a static 3D video. The static 3D video corresponds to a static multi-layer 4D voxel model. All feature vectors of the 4D voxels in each layer are initially random feature vectors. The sampling method of the volume rendering process is used to sample continuous points in the 4D scene to obtain discrete sampling points and their coordinates. The multi-layer feature vectors of all sampling points can be rendered by a renderer to obtain the initial 3D video.
[0032] Next, consider how the 3D video generation module 220, based on the motion control vector 205, can generate a 3D video 207 with controllable object motion by causing the object in each picture in the static 3D video 203 to move according to the motion control vector 205. Specifically, the 3D video with controllable object motion can be generated by optimizing the static 3D video using a second neural network model, wherein a combination of a loss function using the control fractional distillation sampling of the Vincent video diffusion model and a motion loss function is used as the loss function of the second neural network model, and a staged motion enhancement training strategy is employed in the optimization of the static 3D video.
[0033] Figure 6 FIG. 6 is a flow chart showing a method 600 for generating a 3D video with controllable object motion by using a second neural network model in the 3D video generation module 220 according to an embodiment of the present invention. Figure 6 In step 610, it is determined whether the second neural network model has converged. Convergence conditions include, for example, the number of iterations reaching a predetermined requirement or the loss function reaching a predetermined requirement. If it is determined that the second neural network model has not converged, the process proceeds to step 620, where a combination of the loss function for controlling fractional distillation sampling and the motion loss function of the Vincent video diffusion model is used to control the convergence of the second neural network. The process then proceeds to step 630, where a staged motion enhancement training strategy is employed.
[0034] In one embodiment, step 620 uses the loss function of the Vincent video diffusion model to control the convergence of the second neural network, and the motion loss function. The loss function of the Vincent video diffusion model is widely used in the existing Vincent video diffusion model and will not be described in detail here. The motion loss function can be defined as
[0035]
[0036] Among them, M represents the action of the object, N N is the number of random sampling points, qi is the random sampling points around the origin, N T is the number of frames of the output video, and R(q,t) is the rendered color and density at time t and spatial position q. Furthermore, those skilled in the art will know how to use the loss function to control the convergence of the second neural network, which will not be elaborated here.
[0037] In one embodiment, the training strategy of staged motion enhancement includes: during the training process of the second neural network model, initially using a smaller motion control vector, and gradually increasing it to the received motion control vector. Since the received motion control vector is directly used, when the motion amplitude is significant, directly driving the motion may cause drastic changes in content, thereby causing visual artifacts. To solve this problem, the motion intensity can be gradually increased to promote the smooth transition of the dynamic three-dimensional scene to a state that satisfies the motion constraints. During the training process, a staged strategy can be used to adjust the motion intensity. Specifically, for a given motion M, a staged motion enhancement function P(M) can be defined: Where W is the total number of stages, iter is the number of iterations of the current training, stage(iter) is the current stage, k is the number of training iterations of each stage, and U is the total number of iterations of motion enhancement training in all stages, and Based on this, the local deformation loss L corresponding to the staged motion enhancement drive can be derived am(P(M), and
[0038]
[0039] where N N is the number of random sampling points, q i are random sampling points around the origin, N T is the number of frames in the output video, and R(q,r) is the rendered color and density at spatial position q at time t. By gradually reducing the local deformation loss function, the naturalness and effectiveness of the motion can be ensured when it converges to a predetermined range. Current methods such as SGD and ADAM can be used to gradually reduce the local loss function. Based on this, those skilled in the art can devise other training strategies for motion enhancement at this stage, such as using different local deformation loss functions.
[0040] If it is determined in step 610 that the second neural network model has converged, then in step 640 , a three-dimensional video of the object's motion being controllable is output, and then the method 600 ends.
[0041] In one embodiment, the second neural network model includes: a local deformation part, which is used to locally deform the static multi-layer four-dimensional voxel model to obtain a deformed static multi-layer four-dimensional voxel model; a renderer, which is used to sample continuous points in the four-dimensional scene by using the sampling method of the volume rendering process on the deformed static multi-layer four-dimensional voxel model to obtain discrete sampling points and their coordinates. The multi-layer feature vectors of all sampling points can be rendered by the renderer to obtain an intermediate three-dimensional video; a control part, which is used to optimize the intermediate three-dimensional video by using a combination of the loss function of the control fraction distillation sampling of the Vincent video diffusion model and the motion loss function, and control the movement of the object to conform to the received motion control vector; wherein a staged motion enhancement training strategy is used in the training process of the second neural network model.
[0042] In one embodiment, the local deformation component automatically adjusts the local displacement of key points during training, thereby achieving natural alignment of motion trajectories at different time points. This functionality is achieved by calculating the offset of the sampled points: (dx, dy, dz) = ADM(x, y, z, t), where ADM represents the local deformation component, using an existing MLP neural network, and (x', y', z') represents the local deformation of the original coordinates (x, y, z) at a specific time t.
[0043] The renderer can adopt any existing renderer. The control part can adopt the above Figure 6 The method used in step 620 of the method 600 is shown.
[0044] Figure 7Shows the Figure 1 The three-dimensional video generated by the text "a tiger dancing" containing a description of an object and its action is added with the motion control vector (in Figure 1 An example of a dynamic 3D video (indicated by the white arrow in the figure). This demonstrates that the video generated by this method exhibits smoother and more natural motion, with clear, artifact-free images. The generated result can generate the corresponding object motion effect based on the motion control input, enabling the creation of dynamic 3D scenes and videos that meet the creator's intent.
[0045] In one embodiment, Figure 8 FIG. 8 is a flow chart showing a method 800 for generating a three-dimensional video according to an embodiment of the present invention. Figure 8 In step 810, a text describing an object and its motion, as well as a motion control vector for the object, is received. In step 820, a static 3D video depicting the object and its motion is generated based on the text. In step 830, the object in each picture of the static 3D video is moved according to the motion control vector, thereby generating a 3D video in which the object's motion is controllable.
[0046] In one embodiment, step 820 includes: generating an initial three-dimensional video, wherein the initial three-dimensional video is represented using an initial multi-layer four-dimensional voxel model; and optimizing the first frame scene of the initial three-dimensional video by using a first neural network model to obtain the static three-dimensional video, wherein the loss function of the controlled fractional distillation sampling of the Vincent graph diffusion model is used as the loss function of the first neural network model, and the static three-dimensional video corresponds to a static multi-layer four-dimensional voxel model.
[0047] In one embodiment, the multi-layer four-dimensional voxel model includes multiple layers, each layer includes a four-dimensional voxel, each four-dimensional voxel has a different resolution, and each four-dimensional voxel includes a three-dimensional spatial voxel and a one-dimensional temporal voxel.
[0048] In one embodiment, all feature vectors of the four-dimensional voxels of each layer are initially random feature vectors. The continuous points in the four-dimensional scene are sampled using the sampling method of the volume rendering process to obtain discrete sampling points and their coordinates. The multi-layer feature vectors of all sampling points are rendered by a renderer to obtain the initial three-dimensional video.
[0049] In one embodiment, step 830 includes: optimizing the static three-dimensional video by using a second neural network model to obtain a three-dimensional video in which the object motion is controllable, wherein a combination of a loss function of a controlled fractional distillation sampling of a Vincent video diffusion model and a motion loss function is used as the loss function of the second neural network model, and wherein a staged motion enhancement training strategy is used in the process of optimizing the static three-dimensional video.
[0050] In one embodiment, the staged motion enhancement training strategy includes: during the second neural network model training process, initially using a smaller motion control vector and gradually increasing it to the received motion control vector.
[0051] In one embodiment, the motion loss function is:
[0052]
[0053] Among them, M represents the action of the object, N N is the number of random sampling points, qi is the random sampling points around the origin, N T is the number of frames of the output video, and R(q,t) is the rendered color and density at time t and spatial position q.
[0054] In one embodiment, the second neural network model includes: a local deformation part, which is used to obtain a deformed static multi-layer four-dimensional voxel model by locally deforming the static multi-layer four-dimensional voxel model; a renderer, which is used to sample continuous points in the four-dimensional scene using the sampling method of the volume rendering process on the deformed static multi-layer four-dimensional voxel model to obtain discrete sampling points and their coordinates. The multi-layer feature vectors of all sampling points can be rendered by the renderer to obtain an intermediate three-dimensional video; a control part, which is used to optimize the intermediate three-dimensional video by using a combination of the loss function of the control fraction distillation sampling of the Vincent video diffusion model and the motion loss function, and control the movement of the object to conform to the motion control vector.
[0055] In one embodiment, an embodiment of the present invention also discloses a computer system, comprising: one or more processing units; a memory, wherein the memory stores computer program instructions, and the computer program instructions can be executed by the one or more processing units to enable the one or more processing units to perform the steps of the method described above.
[0056] In one embodiment, an embodiment of the present invention further discloses a computer program product, comprising a computer-readable storage medium, wherein the computer-readable storage medium stores computer program instructions, and the program instructions can be executed by one or more processing units to enable the one or more processing units to perform the steps of the method described above.
[0057] The present invention may be a system, method, and / or computer program product. The computer program product includes a computer-readable storage medium. The computer-readable storage medium carries computer-readable program instructions for causing a processor to implement various aspects of the present invention. The method of the present invention may be executed on a standalone computer system, a distributed computing system, or a cloud platform.
[0058] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, and any suitable combination of the foregoing.
[0059] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in a computer-readable storage medium within the corresponding computing / processing device.
[0060] The computer-readable program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages and conventional procedural programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a local area network or a wide area network, or may be connected to an external computer.
[0061] The flow chart and block diagram in the accompanying drawings have shown the possible architecture, function and operation of the system, method and computer-readable storage medium according to multiple embodiments of the present invention.In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and the part for this module, program segment or instruction comprises one or more executable instructions for realizing the logical function of regulation.In some as replacement implementations, the function marked in the box also can occur in a sequence different from that marked in the accompanying drawings.For example, two continuous boxes can actually be performed substantially in parallel, and they also can be performed in reverse order sometimes, and this depends on the function involved.
[0062] The description of the present invention is presented for the purpose of illustration and description and is not intended to be exhaustive or to limit the invention to the disclosed form. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The embodiments are selected and described in order to best explain the principles of the invention, practical applications, and to enable others of ordinary skill in the art to understand the various embodiments of the invention with various modifications that are suitable for the specific purposes contemplated. The terms used herein are selected to best explain the principles of the embodiments, practical applications, or technical improvements existing on the market, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A method for generating a three-dimensional video, comprising: Receive a text containing a description of an object and its motion and a motion control vector for the object; generating a static three-dimensional video of the object and its movements according to the content of the text; as well as According to the motion control vector, for each picture in the static three-dimensional video, an object in the picture is moved according to the motion control vector, thereby obtaining a three-dimensional video with controllable object motion.
2. The method according to claim 1, wherein generating a static 3D video of the object and its movement based on the content of the text comprises: generating an initial three-dimensional video, wherein the initial three-dimensional video is represented by an initial multi-layer four-dimensional voxel model; as well as The static three-dimensional video is obtained by using a first neural network model to optimize the first frame scene of the initial three-dimensional video, wherein the loss function of the controlled fractional distillation sampling of the Vincent graph diffusion model is used as the loss function of the first neural network model, and the static three-dimensional video corresponds to a static multi-layer four-dimensional voxel model.
3. The method according to claim 2, wherein the multi-layer 4D voxel model comprises multiple layers, each layer comprises a 4D voxel, each 4D voxel has a different resolution, and each 4D voxel comprises a 3D spatial voxel and a 1D temporal voxel.
4. The method according to claim 3, wherein all feature vectors of the four-dimensional voxels of each layer are initially random feature vectors, continuous points in the four-dimensional scene are sampled using the sampling method of the volume rendering process to obtain discrete sampling points and their coordinates, and the multi-layer feature vectors of all sampling points are rendered by a renderer to obtain the initial three-dimensional video.
5. The method according to claim 4, wherein, for each picture in the static 3D video, causing an object in the picture to move according to the motion control vector, thereby obtaining a 3D video with controllable object motion, comprises: The static three-dimensional video is optimized using a second neural network model to obtain a three-dimensional video in which the object motion is controllable, wherein a combination of a loss function of a controlled fractional distillation sampling of a Vincent video diffusion model and a motion loss function is used as the loss function of the second neural network model, and a staged motion enhancement training strategy is used in the process of optimizing the static three-dimensional video.
6. The method according to claim 5, wherein the phased motion enhancement training strategy comprises: During the training of the second neural network model, a smaller motion control vector is initially used and then gradually increased to the motion control vector.
7. The method according to claim 5 or 6, wherein the motion loss function is: in, M represents the action of the object, N N is the number of random sampling points, q i are random sampling points around the origin, N T is the number of frames of the output video, and R(q,t) is the rendered color and density at time t and spatial position q.
8. The method according to any one of claims 5 to 7, wherein the second neural network model comprises: A local deformation part, configured to obtain a deformed static multi-layer four-dimensional voxel model by locally deforming the static multi-layer four-dimensional voxel model; The renderer is configured to sample continuous points in the four-dimensional scene using a sampling method of a volume rendering process on the deformed static multi-layer four-dimensional voxel model to obtain discrete sampling points and their coordinates. The multi-layer feature vectors of all the sampling points can be rendered by the renderer to obtain an intermediate three-dimensional video. The control part is used to optimize the intermediate three-dimensional video by using a combination of a loss function of a control fraction distillation sampling of a Vincent video diffusion model and a motion loss function, and control the movement of the object to conform to the motion control vector.
9. A computer system comprising: one or more processing units; A memory storing computer program instructions, wherein the computer program instructions are executable by one or more processing units to cause the one or more processing units to perform the steps of the method according to any one of claims 1 to 8.
10. A computer program product comprising a computer-readable storage medium storing computer program instructions, wherein the program instructions are executable by one or more processing units to cause the one or more processing units to perform the steps of the method according to any one of claims 1 to 8.