Video generation method and device, equipment and medium
Through point cloud rendering and multimodal encoding feature processing, the quality problem of 4D view synthesis in the prior art in complex scenarios is solved, and a higher quality 4D new perspective video generation is achieved.
Patent Information
- Application Number
- CN202510257323.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-03
AI Technical Summary
When existing 4D view synthesis technology has occlusion, large lighting changes or complex object movement in the scene, the synthetic viewing angle may be distorted, blurred or incoherent, resulting in the generated 4D new viewing angle video quality.
By obtaining the source video and camera track parameters, converting them into point cloud data and performing point cloud rendering processing, view encoding features are generated. Combining the video description text and reference encoding features, self-attention processing and cross-attention processing are performed to generate multimodal encoding features, and finally the target video indicated by the camera track parameters is generated.
It improves the generation quality of the target video, enhances the accuracy and temporal and spatial consistency of the view, and reduces the distortion of the occlusion area and the inconsistency of dynamic scenes.
Smart Images

Figure CN120091196A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a video generation method, apparatus, device, and medium. Background Art
[0002] A 4D (four-dimensional) new perspective video refers to a video form that adds a time dimension on the basis of a traditional three-dimensional space (length, width, and height). The 4D new perspective video allows an object to view video content from different perspectives, bringing a more immersive visual experience to the viewer.
[0003] In the current 4D view synthesis technology, deep learning technologies such as convolutional neural networks (CNNs) and generative adversarial networks (GANs) can be used to learn the three-dimensional structure and appearance information of the same object or scene from multiple monocular videos or existing multi-view video data. When synthesizing a 4D new perspective video, according to the input monocular video or partial perspective video, the images of other perspectives can be predicted and coherence can be maintained in the time dimension. However, when there are occlusions, large lighting changes, or complex object movements in the scene, the synthesized perspectives may be distorted, blurred, or incoherent, resulting in too low quality of the synthesized 4D new perspective video. Summary of the Invention
[0004] Embodiments of the present application provide a video generation method, apparatus, device, and medium, which can improve the generation quality of the target video.
[0005] On the one hand, an embodiment of the present application provides a video generation method, including:
[0006] Obtain a source video and camera trajectory parameters, and obtain a video description text of the source video;
[0007] Convert the source video into point cloud data, perform point cloud rendering processing on the point cloud data according to the camera trajectory parameters to obtain a point cloud rendering result, and perform encoding processing on the point cloud rendering result to obtain view encoding features;
[0008] Perform encoding processing on the source video to obtain reference encoding features, and perform encoding processing on the video description text to obtain text encoding features;
[0009] Convert the view encoding features into first self-attention features, and convert the text encoding features into second self-attention features;
[0010] Perform cross-attention processing on the first self-attention features and the reference encoding features to obtain video interaction features, and obtain multi-modal encoding features according to the video interaction features and the second self-attention features;
[0011] Decode the multi-modal encoded features to generate the target video indicated by the camera trajectory parameters.
[0012] One aspect of the embodiments of the present application provides a video generation device, including:
[0013] A data acquisition module, configured to acquire a source video and camera trajectory parameters, and acquire the video description text of the source video;
[0014] A point cloud rendering module, configured to convert the source video into point cloud data, perform point cloud rendering processing on the point cloud data according to the camera trajectory parameters to obtain a point cloud rendering result, and perform encoding processing on the point cloud rendering result to obtain view encoded features;
[0015] An encoding module, configured to perform encoding processing on the source video to obtain reference encoded features, and perform encoding processing on the video description text to obtain text encoded features;
[0016] A first attention processing module, configured to convert the view encoded features into first self-attention features, and convert the text encoded features into second self-attention features;
[0017] A second attention processing module, configured to perform cross-attention processing on the first self-attention features and the reference encoded features to obtain video interaction features, and obtain multi-modal encoded features according to the video interaction features and the second self-attention features;
[0018] A decoding module, configured to decode the multi-modal encoded features to generate the target video indicated by the camera trajectory parameters.
[0019] Wherein, when the data acquisition module acquires the video description text of the source video, it is specifically configured to perform the following steps:
[0020] Perform frame division on the source video to obtain a video frame sequence of the source video, perform feature extraction on each video frame in the video frame sequence to obtain the visual features of each video frame;
[0021] Perform encoding processing on the visual features of each video frame to obtain the image semantic features of the video frame sequence;
[0022] Perform decoding processing on the image semantic features to generate the video description text of the source video.
[0023] Wherein, when the point cloud rendering module converts the source video into point cloud data, it is specifically configured to perform the following steps:
[0024] Divide the source video into multiple video segments, generate the initial noise distribution of each video segment in the multiple video segments, and the initial noise distribution of each video segment is used to characterize the initial depth distribution of a video segment;
[0025] Feature extraction is performed on each video segment to obtain the segment content features of each video segment. According to the segment content features, the initial noise distribution of each video segment is adjusted to obtain the depth distribution estimation of each video segment.
[0026] The depth distribution estimations of each video segment are stitched together to obtain the video depth sequence of the source video.
[0027] According to the video depth sequence, inverse perspective projection is performed on each video frame in the source video to obtain point cloud data.
[0028] Among them, when the point cloud rendering module stitches together the depth distribution estimations of each video segment to obtain the video depth sequence of the source video, it is specifically used to perform the following steps:
[0029] M video segment pairs are determined among multiple video segments. The two video segments in each video segment pair are two adjacent video segments among the multiple video segments, and M is a positive integer.
[0030] According to the overlapping data between the two video segments in each video segment pair and the depth distribution estimations of the two video segments in each video segment pair, the interpolation parameters between the two video segments in each video segment pair are determined.
[0031] According to the interpolation parameters, the depth distribution estimations of the two video segments in each video segment pair are fused to obtain the video depth sequence of the source video.
[0032] Among them, when the point cloud rendering module performs inverse perspective projection on each video frame in the source video according to the video depth sequence to obtain point cloud data, it is specifically used to perform the following steps:
[0033] Determine the original camera parameters of the source video, and construct an inverse perspective projection matrix according to the original camera parameters.
[0034] Obtain the image coordinates of each video frame in the source video, and convert the image coordinates of each video frame into three-dimensional space coordinates according to the inverse perspective projection matrix and the video depth sequence to obtain the first set of spatial coordinates.
[0035] Filter the invalid points in the first set of spatial coordinates to obtain the second set of spatial coordinates, and generate point cloud data according to the second set of spatial coordinates.
[0036] Among them, when the point cloud rendering module performs point cloud rendering processing on the point cloud data according to the camera trajectory parameters to obtain the point cloud rendering result, it is specifically used to perform the following steps:
[0037] Create a camera object according to the camera trajectory parameters, configure the renderer parameters for the point cloud renderer, and perform rendering processing on the point cloud data and the camera object according to the renderer parameters to obtain a sequence of point cloud rendered images;
[0038] Generate a sequence of occlusion mask images for the sequence of point cloud rendered images based on the depth information of each point cloud rendered image in the sequence of point cloud rendered images, and determine the sequence of point cloud rendered images and the sequence of occlusion mask images as the point cloud rendering result.
[0039] Among them, when the point cloud rendering module encodes the point cloud rendering result to obtain the view encoding feature, it is specifically used to perform the following steps:
[0040] Input the point cloud rendering result into the video encoder in the video diffusion model, and perform video compression processing on the point cloud rendering result through the video encoder to obtain a video compression feature;
[0041] Convert the video compression feature into a view latent feature, generate initial noise data with the same dimension as the view latent feature, and combine the view latent feature and the initial noise data to obtain a view-noisy feature;
[0042] Perform block processing on the view-noisy feature to obtain at least two feature blocks, and construct a view feature vector based on the at least two feature blocks;
[0043] Perform rotational position encoding on the at least two feature blocks to obtain position embedding features corresponding to the at least two feature blocks respectively;
[0044] Combine the view feature vector and the position embedding features corresponding to the at least two feature blocks respectively to obtain the view encoding feature.
[0045] Among them, when the encoding module encodes the video description text to obtain the text encoding feature, it is specifically used to perform the following steps:
[0046] Input the video description text into the text encoder in the video diffusion model, and divide the video description text through the text encoder to obtain the first token sequence of the source video;
[0047] Add a flag character to the first token sequence to obtain a second token sequence, perform vector transformation on the second token sequence to obtain the text embedding vector of the second token sequence;
[0048] Perform self-attention processing on the text embedding vector to obtain the third self-attention feature of the text embedding vector, and perform feature transformation on the third self-attention feature to obtain the text encoding feature.
[0049] Among them, when the first attention processing module converts the view encoding features into the first self-attention features and the text encoding features into the second self-attention features, it is specifically used to perform the following steps:
[0050] Input the view encoding features and the text encoding features into the video-specific network in the video diffusion model. The video-specific network adopts a cascaded structure of a first diffusion transformation component and a second diffusion transformation component. The video-specific network includes N first diffusion transformation components and N second diffusion transformation components, where N is a positive integer;
[0051] Concatenate the view encoding features and the text encoding features to obtain the first fusion feature;
[0052] According to the full attention layer in the first first diffusion transformation component, perform self-attention processing on the first fusion feature to obtain the second fusion feature;
[0053] Split the second fusion feature into the first self-attention features of the view encoding features and the second self-attention features of the text encoding features.
[0054] Among them, when the second attention processing module performs cross-attention processing on the first self-attention features and the reference encoding features to obtain the video interaction features, and obtains the multi-modal encoding features according to the video interaction features and the second self-attention features, it is specifically used to perform the following steps:
[0055] Input the reference encoding features into N first diffusion transformation components. In the first first diffusion transformation component, perform a linear transformation on the reference encoding features to obtain a key matrix and a value matrix, and perform a linear transformation on the first self-attention features to obtain a query matrix;
[0056] According to the cross-attention layer in the first first diffusion transformation component, perform a dot product operation on the query matrix and the transposed matrix of the key matrix to obtain a candidate weight matrix, and obtain the number of columns of the query matrix;
[0057] Normalize the ratio between the candidate weight matrix and the square root of the number of columns to obtain an attention weight matrix, and perform a dot product operation on the attention weight matrix and the value matrix to obtain the video interaction features;
[0058] In the N - 1 first diffusion models except the first first diffusion model and the N second diffusion transformation components, perform multi-modal fusion on the video interaction features, the second self-attention features, and the reference encoding features to obtain the multi-modal encoding features.
[0059] Among them, when the decoding module decodes the multi-modal encoding features to generate the target video indicated by the camera trajectory parameters, it is specifically used to perform the following steps:
[0060] Input the multi-modal encoded features into the video decoder in the video diffusion model, and perform upsampling processing on the multi-modal encoded features through the video decoder to obtain the reconstructed video data of the source video;
[0061] Perform activation processing on the reconstructed video data to obtain the target video with the same format as the source video.
[0062] Wherein, the video generation device further includes:
[0063] A training set acquisition module, configured to acquire a first sample training set and a second sample training set. The first sample training set includes multiple video training pairs, and each video training pair includes a sample monocular video and a first point cloud rendering video. The second sample training set includes multiple sample triples, and each sample triple includes a first perspective video, a second perspective video, and a second point cloud rendering video. The first perspective video and the second perspective video have different perspectives;
[0064] A model acquisition module, configured to acquire a diffusion model to be trained, and the diffusion model to be trained includes a video encoder to be trained, a text encoder to be trained, a dedicated network to be trained, and a video decoder to be trained;
[0065] A model training module, configured to fix the network parameters of the cross-attention layer in the dedicated network to be trained, and train the diffusion model to be trained according to multiple video training pairs to obtain a pre-trained diffusion model;
[0066] The model training module is further configured to correct the network parameters of the cross-attention layer in the pre-trained diffusion model according to multiple sample triples to obtain a video diffusion model.
[0067] Wherein, when the training set acquisition module acquires the first sample training set and the second sample training set, it is specifically configured to perform the following steps:
[0068] Acquire a sample monocular video, sample the camera transformation matrix of the sample monocular video, and generate a candidate view sequence of the sample monocular video according to the camera transformation matrix;
[0069] Perform inverse projection on the candidate view sequence to obtain a first point cloud rendering video, construct a video training pair from the sample monocular video and the first point cloud rendering video, and add the video training pair to the first sample training set;
[0070] Acquire multi-view static data, construct global point cloud data of the multi-view static data, and sample a first perspective video and a second perspective video from the multi-view static data;
[0071] Obtain the target camera parameters of the first-person view video. According to the target camera parameters, perform point cloud rendering processing on the global point cloud data to obtain the second point cloud rendering video. Construct a sample triple from the first-person view video, the second-person view video, and the second point cloud rendering video, and add the sample triple to the second sample training set.
[0072] In one aspect, an embodiment of the present application provides a computer device, including a memory and a processor. The memory is connected to the processor. The memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method provided in the above aspect of the embodiment of the present application.
[0073] In one aspect, an embodiment of the present application provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. The computer program is adapted to be loaded and executed by a processor so that a computer device with a processor executes the method provided in the above aspect of the embodiment of the present application.
[0074] According to one aspect of the present application, a computer program product is provided. The computer program product may include a computer program, and the computer program is stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program so that the computer device executes the method provided in the above aspect.
[0075] In the embodiment of the present application, the source video, the camera trajectory parameters, and the video description text of the source video can be obtained. The source video is converted into point cloud data. According to the camera trajectory parameters, point cloud rendering processing is performed on the point cloud data to obtain a point cloud rendering result. The point cloud rendering result is encoded into a view encoding feature. By performing self-attention processing on the view encoding feature and the text encoding feature of the video description text, geometric alignment can be achieved and the accuracy of the view can be maintained. By performing cross-attention processing on the view encoding feature and the reference encoding feature of the source video, detailed information in the source video can be injected, the spatio-temporal consistency between the target video and the source video can be improved, and thus the generation quality of the target video can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0077] Figure 1 is a schematic structural diagram of a network architecture provided by an embodiment of the present application;
[0078] Figure 2 It is a schematic diagram of a video generation scenario provided by an embodiment of the present application;
[0079] Figure 3 It is a schematic flowchart of a video generation method provided by an embodiment of the present application;
[0080] Figure 4 It is a schematic structural diagram of a video diffusion model provided by an embodiment of the present application;
[0081] Figure 5 It is a schematic structural diagram of a first diffusion transformation component in a video diffusion model provided by an embodiment of the present application;
[0082] Figure 6 It is a schematic diagram of the training of a video diffusion model provided by an embodiment of the present application;
[0083] Figure 7 It is a schematic structural diagram of a video generation device provided by an embodiment of the present application;
[0084] Figure 8 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0085] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0086] For ease of understanding, the basic concepts related to the embodiments of the present application are described below.
[0087] 4D (Four-Dimensional) View Synthesis: A technology in the fields of computer graphics and computer vision that can generate views containing the time dimension from a given series of two-dimensional or three-dimensional data. Here, "4D" can refer to the three-dimensional space and the time dimension. Through the 4D view synthesis technology, new perspective videos with temporal (the 4th dimension) and spatial (3D) consistency can be generated in dynamic scenes, supporting user-defined camera trajectories.
[0088] Camera Trajectory Parameters: The movement path of a camera in three-dimensional space. It can describe the changes in the position and orientation (pose) of the camera during the shooting process over time or the shooting sequence. This camera trajectory can include position parameters, rotation parameters, and focal length parameters, etc.
[0089] Among them, the position parameter can represent the coordinate position of the camera in three-dimensional space, and this position parameter can be represented by three-dimensional space coordinates (x, y, z), where x, y, and z respectively represent the coordinate values of the camera on the three axes in the space coordinate system. The focal length parameter can be used to represent the viewing angle range of the camera and the magnification of imaging.
[0090] The rotation parameter can represent the pose of the camera, and this rotation parameter can include yaw angle (Yaw), pitch angle (Pitch), and roll angle (Roll). The yaw angle can represent the rotation angle of the camera around the vertical axis (such as the Z axis), and can be used to control the direction of the camera turning left and right; the pitch angle can represent the rotation angle of the camera around the transverse axis (such as the Z axis), and can be used to represent the degree of the camera looking up or down; the roll angle can represent the rotation angle of the camera around the longitudinal axis (such as the Y axis), and can be used to reflect the tilt state of the camera itself.
[0091] Dynamic Point Cloud: Spatiotemporally consistent point cloud data constructed by the depth estimation method of monocular video, which can be used to render the geometric structure of the target view. The point cloud data can refer to a set of discrete data points in three-dimensional space. Each point in the point cloud data can contain coordinate information in three-dimensional space, and these points together depict the surface shape of a three-dimensional object or scene. In addition to carrying coordinate information, each point in the point cloud data can also carry other attribute information, such as color, normal vector, etc.
[0092] Monocular Video: A video obtained by shooting with a single camera. All video frames in the monocular video are shot from one angle.
[0093] Diffusion Model: A generative model based on the thermodynamic random diffusion process. This thermodynamic random diffusion process can include a forward process of gradually adding noise to the samples from the data distribution, and a backward process of training a neural network to reverse the forward process by gradually removing the noise. Diffusion models can be widely applied in various tasks such as video generation tasks, image generation tasks, and natural language generation tasks.
[0094] Double-reprojection Strategy: Training data can be generated through random view transformation and inverse projection, which can solve the problems of occlusion and geometric distortion.
[0095] Ref-DiT Block (Reference Conditional Diffusion Transformation Block): By means of the cross-attention mechanism, injecting the detailed information of the source video into the target view generation process can improve content consistency.
[0096] Three-dimensional Variational Autoencoder (3D VAE): The 3D VAE maps 3D data to a latent space through an encoder. This latent space is usually a continuous, low-dimensional representation that can capture the key features and variations of the data. This latent space can better represent the intrinsic features and structure of the data, achieving data compression and simplification.
[0097] Please refer to Figure 1 , Figure 1 FIG. is a schematic structural diagram of a network architecture provided by an embodiment of the present application; the network architecture may include a server 10d and a terminal cluster. The terminal cluster may include one or more terminal devices, and the number of terminal devices included in the terminal cluster and the number of servers included in the network architecture are not limited herein. As Figure 1 shown, the terminal cluster may specifically include terminal devices 10a, 10b, and 10c, etc.; all terminal devices in the terminal cluster (for example, terminal devices 10a, 10b, and 10c, etc.) can be network-connected to the server 10d, so that each terminal device can perform data interaction with the server 10d through this network connection.
[0098] Among them, Figure 1 the terminal devices in the terminal cluster shown may include, but are not limited to: electronic devices such as smart phones, tablet computers, laptop computers, palmtop computers, desktop computers, wearable devices (such as smart watches, smart bracelets, etc.), intelligent voice interaction devices, smart home appliances (such as smart TVs, etc.), vehicle-mounted devices, and aircraft. The type of terminal device in the present application is not limited.
[0099] Figure 1 The server 10d shown may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The type of server in the present application is not limited.
[0100] In the embodiments of the present application, Figure 1Each terminal device in the shown terminal cluster can run a browser or a client. Among them, the browser can be used to access web pages and interact with server 10d through protocols such as HTTP (HyperText Transfer Protocol) to obtain web page content for presentation to users; the browser generally does not need to be installed or configured for a specific service or application. The client is usually developed for a specific application or service and generally needs to be downloaded, installed, and updated to adapt to the needs and changes of a specific service; the client involved in this application can be any client with video synthesis capabilities, and this application does not limit the type of the client. When the client runs on each terminal device, it can perform data interaction with Figure 1 the shown server 10d. At this time, server 10d can be the background server corresponding to the client. Among them, the client running on each terminal device can be an independent client or an embedded sub-client integrated in a certain independent client, and this application does not limit this.
[0101] It should be understood that Figure 1 the browser or client running on each of the shown terminal devices can call a pre-trained video diffusion model. For example, the trained video diffusion model can be launched in the browser or client. If object A wants to convert the monocular video it shoots into a new perspective video, then object A can upload the monocular video and input any camera trajectory parameters in the client installed on the terminal device (for example, terminal device 10a). Terminal device 10a can obtain the monocular video and camera trajectory parameters uploaded by object A and call the video diffusion model through the client or browser running on this terminal device 10a. Using the monocular video and camera trajectory parameters as the input data of this video diffusion model, the video diffusion model can predict a new perspective video, and the application process of this video diffusion model (such as the process of generating a new perspective video) will be described in the subsequent content.
[0102] Among them, the monocular video can refer to the video captured by object A using devices such as mobile phones and cameras; or, the monocular video can refer to the video generated by object A using AI (Artificial Intelligence) technology, and so on. The camera trajectory parameters can be used to describe the changes in the position and pose of the camera during the shooting process over time, and can include the motion information of the camera in three-dimensional space, such as position coordinates, rotation angles, focal lengths, and other parameters. The camera trajectory parameters can be any camera trajectory parameters defined by object A; the camera trajectory parameters can be used to indicate the new perspective that object A needs to generate, such as the "new perspective" for indicating the new perspective video. The new perspective video can refer to a video with different perspectives generated on the basis of the monocular video through 4D view synthesis technology.
[0103] It can be understood that the video diffusion model involved in the embodiments of this application can be a dual-stream conditional video diffusion model, which combines the dual conditional inputs of point cloud rendering and the source video to generate a high-fidelity and consistent 4D new perspective video. Among them, the point cloud rendering condition can refer to the point cloud rendering result obtained by performing point cloud rendering processing according to the camera trajectory parameters and the point cloud data constructed from the monocular video, which can ensure the geometric alignment of the generated new perspective with the camera trajectory parameters and ensure view accuracy. The source video condition can refer to the monocular video. By injecting the detailed information of the monocular video into the video diffusion model through the cross-attention mechanism, the spatio-temporal consistency between the generated new perspective video and the monocular video, as well as the fidelity of the new perspective video, can be improved. The point cloud rendering processing can refer to the process of converting the point cloud data in three-dimensional space into a two-dimensional image or a visualization result. The point cloud rendering result can refer to the final output obtained after the point cloud rendering processing, usually presented in the form of a two-dimensional image or a visualization scene.
[0104] In the embodiments of this application, the video diffusion model can refer to the trained dual-stream conditional video diffusion model; the training process of the dual-stream conditional video diffusion model can include two stages, respectively denoted as the first training stage and the second training stage. The dual-stream conditional video diffusion model in the first training stage can be called the diffusion model to be trained, and the dual-stream conditional video diffusion model after the first training stage can be called the pre-trained diffusion model. For the sake of understanding, the dual-stream conditional video diffusion model in the second training stage can also be called the pre-trained diffusion model, and the dual-stream conditional video diffusion model after the second training stage can be called the video diffusion model. Among them, the video diffusion model, the pre-trained diffusion model, and the diffusion model to be trained are dual-stream conditional video diffusion models in different stages. That is to say, the video diffusion model, the pre-trained diffusion model, and the diffusion model to be trained have the same network structure but different network parameters.
[0105] It should be noted that the training process and application process of the dual-stream conditional video diffusion model can be executed by a computer device, which can be the server 10d in the network architecture shown in Figure 1 (for example, the server 10d can be the background server of the client), or can be any terminal device in the terminal cluster, or can be a computer program (including program code, for example, the client installed in the terminal device), etc. This application does not make any limitations in this regard.
[0106] As a 4D view synthesis technology, the trained video diffusion model can be applied in the scenarios of film and television and game content creation, or can be applied in virtual reality (VR) and metaverse platforms, or can be applied in e-commerce and advertising interaction systems, or can be applied in other scenarios. This application does not make any limitations in this regard.
[0107] In the scenarios of film and television and game content creation, by using the video diffusion model, it can provide the dynamic scene perspective expansion function for post-production of film and television and game engines, support directors of film and television works and developers of game engines to freely design the camera movement trajectory, and generate multi-perspective shots or game cutscenes. It can reduce the cost of multi-camera shooting and improve the efficiency of virtual scene construction.
[0108] In scenarios such as virtual reality social interaction and virtual exhibitions, the video diffusion model can be used to generate 3D dynamic content with user-defined perspectives in real time, enhancing the user's immersive experience; it can break through the fixed perspective limitation of traditional 360° panoramic videos and support dynamic interactive exploration.
[0109] In the e-commerce and advertising interaction system, by using the video diffusion model, it can generate 3D dynamic previews with arbitrary rotation perspectives for product display videos (such as clothing, electronic products, etc.). Using the video diffusion model instead of traditional 3D modeling, only monocular shooting is required to achieve high-interactivity display.
[0110] Next, taking the e-commerce and advertising interaction system as an example, the application process of the video diffusion model will be described. Please refer to Figure 2 , Figure 2 which is a schematic diagram of a video generation scenario provided by an embodiment of this application; as Figure 2 shown, when object A wants to create a 3D dynamic preview for the product (such as clothing) to be displayed, it can start the e-commerce client in the terminal device 20a and perform a trigger operation on the "video production" control in the e-commerce client. At this time, the terminal device 20a can respond to the trigger operation for the "video production" control and display the video production page 20b in the e-commerce client.
[0111] On the video production page 20b, a monocular video 20c of the object A shot for the service can be uploaded, and custom camera trajectory parameters 20d can be input. After the monocular video 20c and the camera trajectory parameters 20d are successfully uploaded, a trigger operation can be performed on the "Generate Video" control 20e on the video production page 20b. At this time, the terminal device 20a can obtain the monocular video 20c and the camera trajectory parameters 20d, and call the video diffusion model 20n in the e-commerce client. The network structure of the video diffusion model 20n will be described in the subsequent steps.
[0112] The terminal device 20a can construct point cloud data 20f based on the monocular video 20c, and perform point cloud rendering processing on the point cloud data 20f according to the camera trajectory parameters 20d input by the object A to obtain a point cloud rendering result 20g; image encoding processing can be performed on the point cloud rendering result 20g to obtain view latent features. The terminal device 20a can randomly initialize a Gaussian noise 20h (which can also be called initial noise data), and the Gaussian noise 20h has the same dimension as the video compression feature. The Gaussian noise 20h is combined with the view latent features, and the combined noisy features are block-processed to obtain view encoded features 20i.
[0113] Among them, the view latent features are a compressed representation of the point cloud rendering result 20g, removing redundant information and noise in the point cloud rendering result 20g. The view latent features can be used to represent the key information and essential information in the point cloud rendering result 20g. The view encoded features 20i can refer to the output result obtained by block-processing the combined noisy features. The view encoded features 20i can be represented in the form of view tokens. Here, the tokens represent multiple small patches after the combined noisy features are block-processed.
[0114] The terminal device 20a calls any publicly available video description generator to perform video content understanding on the monocular video 20c and generate a video description text 20j of the monocular video 20c; text encoding processing is performed on the video description text 20j to obtain text encoded features 20k. Among them, the video description generator refers to a video content understanding tool based on a large language model (LLM), which can generate appropriate description texts for video data. The text encoded features 20k can be represented in the form of text tokens. Here, the text tokens can refer to the encoded features of each token in the video description text 20j.
[0115] The terminal device 20a can perform image encoding processing on the monocular video 20c to obtain the reference encoding feature 20m of the monocular video 20c, and the reference encoding feature 20m can be represented in the same form as the view encoding feature 20i.
[0116] The terminal device 20a can input the text encoding feature 20k, the reference encoding feature 20m, and the view encoding feature 20i into the video diffusion model 20n. Through the video diffusion model 20n, the target video indicated by the camera trajectory parameter 20d can be generated. The target video can be the dynamic preview effect 20p of the clothing in the monocular video 20c. The dynamic preview effect 20p can include preview images of the clothing from different perspectives, and the different perspectives here are specified by the camera trajectory parameter 20d. Among them, the generation process of the target video will be described in the subsequent content. The dynamic preview effect 20p generated by the video diffusion model 20n can be displayed on the video production page 20b.
[0117] Optionally, the video production page 20b may further include an "upload" control. The object A can perform a triggering operation on the "upload" control on the video production page 20b to upload the dynamic preview effect 20p to the product display page of the e-commerce client, so that the users in the e-commerce client can preview the display images of the clothing in all directions, which can improve the interactivity of the products in the e-commerce client.
[0118] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of a video generation method provided by an embodiment of the present application; it can be understood that the video generation method can be executed by a computer device, and the computer device can be a terminal device, such as Figure 1 any terminal device in the terminal cluster shown, or can be a server, such as Figure 1 the server 10d shown, and the present application does not make any limitation in this regard. As Figure 3 shown, the video generation method may include the following steps S101 to step S107:
[0119] Step S101, obtain a source video and a camera trajectory parameter, and obtain a video description text of the source video.
[0120] Among them, the source video may refer to a monocular video captured by a single camera; or, the source video may refer to a monocular video generated using AI technology. This application does not limit the acquisition method of the source video. The source video can be any type of video data captured from one angle. For example, the source video can be a video of a film and television scene, a person, an object, etc. captured from one angle, or it can be a video of a scene, a player character, a non-player character, etc. in a game engine, or it can be a monocular video of any commodity on an e-commerce platform, or it can be a video of any scene, object, building, etc. in the real world, or a video of any exhibit in a virtual exhibition scene, etc. This application does not limit this.
[0121] The source video obtained in the embodiments of this application can be expressed as where I s represents the source video, represents any video frame in the source video, a is a positive integer less than or equal to B; B represents the number of video frames in the source video, h represents the height of the source video, w represents the width of the source video, and the value 3 represents three channels; for example, when each video frame in the source video is an RGB (color system, R represents red, G represents green, B represents blue) image, the channel of the source video is 3.
[0122] The camera trajectory parameters may refer to the perspective specified for an object to be generated, and the camera trajectory parameters may include parameters such as position coordinates, rotation angles, and focal lengths. The camera trajectory parameters can be expressed as where, T r represents the camera trajectory parameters, and the camera trajectory parameters can be a sequence of length B; represents the camera parameters at any time, and the camera parameters at each time can be represented by a 4×4 matrix, which can be used to represent the perspective specified for the a-th video frame in the source video.
[0123] The video description text may refer to the prompt text of the natural language description generated for the source video by understanding the video content of the source video, and the video description text may be related to the scene, object, etc. in the source video. For example, in the embodiments of this application, a Captioner (which can be called a video description generator) can be used to generate the video description text of the source video; or, this application can also use other methods other than Captioner to generate the video description text of the source video. This application does not limit the method of generating the video description text.
[0124] For ease of understanding, the following description takes the Captioner as an example. The Captioner can be trained using a large amount of image-text pair data and can adopt an encoder-decoder architecture. The encoder in the Captioner can be used to perform image encoding on each video frame in the input source video, and the decoder is used to gradually generate video description text according to the output result of the encoder. For example, the source video can be frame-divided to obtain a video frame sequence of the source video, feature extraction is performed on each video frame in the video frame sequence to obtain the visual features of each video frame; the visual features of each video frame are encoded to obtain the image semantic features of the video frame sequence; the image semantic features are decoded to generate the video description text of the source video.
[0125] Among them, frame division can refer to the process of decomposing a continuous source video into a series of individual video frames. For example, a video frame can be extracted from the source video at a preset interval time, and finally a video frame sequence can be obtained; among them, the interval time can refer to the interval duration for extracting video frames from the source video, and the video frame sequence can include all the video frames extracted from the source video.
[0126] Visual features can refer to the information that can reflect the essence and characteristics of video frames extracted from each video frame in the video frame sequence. The visual features can include at least one of features such as color features, texture features, shape features, and spatial relationship features of each video frame in the video frame sequence. Among them, the color feature can be used to reflect the overall or local color distribution of the video frame, and the color feature is a global feature. The texture feature can be used to describe the spatial distribution pattern of pixel gray values in the video frame and can reflect the surface structure and texture information of the video frame. The shape feature can be used to describe the shape and contour information of the objects in the video frame. The spatial relationship feature can be used to describe the spatial positions and mutual relationships between different objects or regions in the video frame.
[0127] Image semantic features can refer to the deep and meaningful information expressed by the content of the source video, which can go beyond the pixel values and surface features of the video frames; the image semantic features can be understood as features deeper than visual features and can be used to reflect the semantics of objects, scenes, space-time, events, etc. between each video frame in the video frame sequence.
[0128] In one or more embodiments, the visual features and image semantic features can be obtained through the encoder in the Captioner. The visual features can be the feature points and their descriptors extracted by a feature extraction algorithm, and the feature extraction algorithm can include but is not limited to: SIFT (Scale Invariant Feature Transform), SURF (Speeded Up Robust Features), ORB (Oriented FAST and Rotated BRIEF), convolutional neural network (CNN), etc. The image semantic features can be the encoded features obtained by sequence modeling of the visual features through any network such as a recurrent neural network or a long short-term memory network. The decoder in the Captioner can gradually generate the video description text according to the image semantic features output by the encoder; during the generation process, greedy search or beam search can be used to select the next word or phrase, and finally the complete video description text is generated.
[0129] Step S102: Convert the source video into point cloud data, and perform point cloud rendering processing on the point cloud data according to the camera trajectory parameters to obtain a point cloud rendering result.
[0130] Specifically, through a monocular depth estimation method, a temporally consistent video depth sequence can be extracted from the source video, and then each video frame in the source video can be converted into point cloud data according to the video depth sequence and inverse perspective projection.
[0131] Among them, the monocular depth estimation method can predict the distance from each pixel in the image (video frame) to the camera through a single image (video frame), so as to restore the three-dimensional structure information of the scene; the embodiments of the present application can adopt DepthCrafter (a depth estimation method) or other publicly available depth estimation methods, and the present application does not make any limitations in this regard. The video depth sequence can include the depth information of each video frame in the source video. The video depth sequence refers to a set of a series of depth information obtained after depth estimation of each video frame in the source video. The depth information can refer to the distance from each pixel in a single video frame to the camera. The video depth sequence can be expressed as where D s represents the video depth sequence, represents the depth information of any video frame (such as the a-th video frame) in the source video.
[0132] Inverse perspective projection may refer to the process of back-projecting points or objects in a two-dimensional image from the image plane into three-dimensional space. The point cloud data involved in the embodiments of this application may be dynamic point cloud, and the point cloud data may be represented as where P a represents the dynamic point cloud of the a-th video frame in the source video.
[0133] In one or more embodiments, taking the monocular depth estimation method DepthCrafter as an example, the depth estimation process of the source video is described. Specifically, the source video can be divided into multiple video segments. Here, the multiple video segments can be multiple overlapping segments, and the length of these video segments can be determined according to the processing ability of DepthCrafter and the specific situation of the source video. Each video segment can contain a certain number of consecutive video frames, and the number of video frames contained in each video segment can be flexibly set according to specific requirements. This application does not make any limitations in this regard.
[0134] For each of the multiple video segments, the depth sequence can be estimated by using noise initialization. In the process of estimating the depth sequence by using noise initialization, the initial noise distribution of each of the multiple video segments can be generated, and the initial noise distribution of each video segment can be used to characterize the initial depth distribution of a video segment; that is to say, this initial noise distribution can be regarded as a preliminary guess of the segment depth distribution.
[0135] Feature extraction can be performed on each video segment to obtain the segment content features of each video segment. According to the segment content features of each video segment, the initial noise distribution of each video segment is adjusted and transformed, so that it gradually converges to a reasonable depth distribution estimate, and the depth distribution estimate of each video segment is obtained; this depth distribution estimate can be used to characterize the approximate scale and offset of the depth distribution of each video segment. The depth distribution estimates of each video segment can be stitched together to obtain the video depth sequence of the source video According to the video depth sequence, inverse perspective projection is performed on each video frame in the source video, and point cloud data can be obtained where the point cloud data in which P a can be calculated as shown in formula (1):
[0136]
[0137] where represents inverse perspective projection, and K represents the camera intrinsic matrix.
[0138] In one or more embodiments, the stitching method between the depth distribution estimates of multiple video segments may include, but is not limited to: M pairs of video segments may be determined among the multiple video segments, where the two video segments in each pair of video segments are two adjacent video segments among the multiple video segments, and M is a positive integer; that is, two adjacent video segments among the multiple video segments may be combined to obtain M pairs of video segments.
[0139] According to the overlap data between the two video segments in each pair of video segments, and the depth distribution estimates of the two video segments in each pair of video segments, interpolation parameters between the two video segments in each pair of video segments are determined; according to the interpolation parameters, the depth distribution estimates of the two video segments in each pair of video segments are fused and adjusted to obtain the video depth sequence of the source video. Through these interpolation parameters, the stitched video depth sequence can be made coherent and consistent in both time and space.
[0140] In one or more embodiments, the process of obtaining point cloud data using inverse perspective projection may include, but is not limited to: the original camera parameters of the source video may be determined, and an inverse perspective projection matrix may be constructed according to the original camera parameters; the image coordinates of each video frame in the source video are obtained, and according to the inverse perspective projection matrix and the video depth sequence, the image coordinates of each video frame are converted into three-dimensional space coordinates to obtain a first set of space coordinates; the invalid points in the first set of space coordinates are filtered to obtain a second set of space coordinates, and point cloud data is generated according to the second set of space coordinates.
[0141] Among them, the original camera parameters may refer to the camera parameters of each video frame in the source video, such as those that can be used to indicate the viewing angle of each video frame in the source video. The original camera parameters may include the camera intrinsic matrix and the camera extrinsic matrix. The inverse perspective projection matrix is the inverse matrix of the perspective projection matrix. Perspective projection is the process of projecting points in three-dimensional space onto a two-dimensional image plane, and this transformation can be represented by a projection matrix. The inverse perspective projection matrix, on the other hand, performs the opposite operation, that is, it back-projects points on the two-dimensional image plane back into three-dimensional space, and the inverse perspective projection matrix can be calculated from the camera intrinsic matrix and the camera extrinsic matrix, such as the product between the inverse matrix of the camera intrinsic matrix and the inverse matrix of the camera extrinsic matrix.
[0142] The image coordinates may refer to the coordinates of each pixel of each video frame in the source video on the two-dimensional image plane. The first set of space coordinates may refer to the set of points in three-dimensional space obtained using inverse perspective projection, and the second set of space coordinates may refer to the result obtained by filtering the points in the first set of space coordinates. Among them, for the calculation method of any point in three-dimensional space within the first set of space coordinates, it may be as shown in formula (2):
[0143]
[0144] Among them, the three-dimensional space coordinates (X, Y, Z) can be expressed as the coordinates of any point in the first set of space coordinates. X represents the horizontal axis coordinate in the three-dimensional space coordinate system, Y represents the vertical axis coordinate in the three-dimensional space coordinate system, and Z represents the vertical axis coordinate in the three-dimensional space coordinate system. The two-dimensional coordinates (u, v) represent the coordinates of each pixel in each video frame of the source video on the two-dimensional image. u represents the coordinate in the horizontal direction in the two-dimensional image plane, and v represents the coordinate in the vertical direction in the two-dimensional image plane. Their value ranges can both be expressed as [0, 1]. d represents the depth information (depth value) of each pixel in each video frame of the source video, and P -1 represents the inverse perspective projection matrix.
[0145] Among them, the invalid points in the first set of space coordinates may include but are not limited to noise points, occluded points, and invisible points, etc. For example, they can be removed by methods such as setting thresholds. The thresholds here can include a first threshold and a second threshold, and the first threshold is greater than the second threshold. Points with a depth value greater than the first threshold or points with a depth value less than the second threshold can be determined as noise points (erroneously estimated points), and these noise points can be removed. According to the geometric relationship of the scene in the source video and the camera view, it can be judged which points are occluded or invisible points outside the camera's field of view, and they can be removed. For example, by comparing the magnitudes of different depth values, if a point has a depth value greater than its adjacent points and there are no other objects in front of it, it may be an occluded point and can be removed.
[0146] After removing the invalid points in the first set of space coordinates, a second set of space coordinates can be obtained. The points in this second set of space coordinates can be organized into point cloud data in a certain format and structure (for example, it can be dynamic point cloud). The format of the point cloud data can include but is not limited to PLY (Polygon File Format, a point cloud storage format), PCD (a format consisting of a file header and a data part), etc. In some formats, in addition to recording the three-dimensional space coordinates (X, Y, Z), attribute information such as color and normal vector can also be included.
[0147] In one or more embodiments, according to the camera trajectory parameters, a point cloud renderer can be used to perform point cloud rendering processing on the point cloud data to obtain a point cloud rendering result. Among them, the point cloud renderer can refer to a software tool or algorithm module for processing point cloud data and converting it into a visual image or video. The type of the point cloud renderer is not limited in this application. For ease of understanding, the embodiments of this application can use a differentiable point cloud renderer (for example, PyTorch3D, a deep learning library based on 3D data) to generate a point cloud rendering result.
[0148] Among them, the process of obtaining the point cloud rendering result may include but is not limited to: a camera object can be created according to the camera trajectory parameters. The camera object is an abstract representation of a virtual camera and is used to simulate the behavior and attributes of a real-world camera in a three-dimensional scene. When creating the camera object, the camera trajectory parameters can be passed into the camera object. Through the camera trajectory parameters, the position and viewing angle of the camera in the three-dimensional scene can be determined.
[0149] Renderer parameters can be configured for the point cloud renderer. According to the renderer parameters, the point cloud data and the camera object are rendered to obtain a sequence of point cloud rendering images; according to the depth information of each point cloud rendering image in the sequence of point cloud rendering images, a sequence of occlusion mask images of the sequence of point cloud rendering images is generated, and the sequence of point cloud rendering images and the sequence of occlusion mask images are determined as the point cloud rendering result.
[0150] Among them, the renderer parameters can be used to control the visualization effect of the point cloud data and may include but are not limited to the rasterization resolution, the number of samples, etc. A point cloud renderer object can be created according to the camera object and the renderer parameters. The renderer parameters and the camera object can be passed into the point cloud renderer object. The point cloud renderer object can refer to a tool for visually presenting the point cloud data on a screen or other display device.
[0151] The point cloud data can be converted into a format suitable for input to the point cloud renderer. For example, the three-dimensional spatial coordinates and attribute information of the points in the point cloud data can be organized into a tensor, and operations such as dimension expansion can be performed as needed. Through the point cloud renderer object, the input point cloud data is rendered to generate a point cloud rendering image of the viewing angle indicated by the camera trajectory parameters. Through the camera trajectory parameters, multiple point cloud rendering images can be generated, and the multiple point cloud rendering images can be constructed into a sequence of point cloud rendering images in chronological order.
[0152] During the point cloud rendering process, an occlusion mask image corresponding to each point cloud rendering image can be generated through the depth information of each pixel in each point cloud rendering image. The occlusion mask images corresponding to each point cloud rendering image are combined in chronological order to obtain a sequence of occlusion mask images. The sequence of occlusion mask images and the sequence of point cloud rendering images can be determined as the point cloud rendering result. That is to say, the point cloud rendering result can be two videos, one is the sequence of point cloud rendering images, and the other is the sequence of occlusion mask images.
[0153] It can be understood that each point cloud rendering image may include an occluded area and a non-occluded area. The mask value of the occluded area is 1 and can be marked as the area to be repaired. The mask value of the non-occluded area is 0, and the point cloud rendering result of the points in the point cloud data is retained. The occluded area and the non-occluded area in the point cloud rendering image can be determined by the occlusion mask image corresponding to the point cloud rendering image, and the occluded area needs to be repaired by the video diffusion model.
[0154] Step S103: Encode the point cloud rendering result to obtain view encoding features.
[0155] Specifically, the point cloud rendering result can be input into the video encoder in the video diffusion model. The video encoder performs video compression processing on the point cloud rendering result to obtain video compression features; the video compression features are converted into view latent features, and initial noise data with the same dimension as the view latent features is generated. The view latent features and the initial noise data are combined to obtain view-noisy features; the view-noisy features are divided into blocks to obtain at least two feature blocks, and a view feature vector is constructed according to the at least two feature blocks; the at least two feature blocks are subjected to rotational position encoding to obtain position embedding features corresponding to the at least two feature blocks respectively; the view feature vectors corresponding to the at least two feature blocks are combined with the position embedding features corresponding to the at least two feature blocks respectively to obtain view encoding features.
[0156] Among them, the video diffusion model can be a two-stream conditional video diffusion model based on 3D VAE (3D Variational Autoencoder) and DiT (Diffusion Transformer, a diffusion model based on Transformer, and Transformer represents a deep learning model based on the attention mechanism). That is to say, the basic architecture of the video diffusion model is 3D VAE and DiT. The video diffusion model can be any network structure based on 3D VAE and DiT and introducing a two-stream conditional mechanism. The specific network structure of the video diffusion model in this application is not limited.
[0157] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a video diffusion model provided by an embodiment of the present application. As Figure 4As shown, the video diffusion model may include a video encoder 30f, a text encoder 30h, a video-specific network 30n, and a video decoder 30p. Among them, the video encoder 30f can be used to encode the point cloud rendering result or the original video. The video encoder 30f can be a 3D VAE, or a three-dimensional causal variational autoencoder (3D Causal Variational Autoencoder, 3D Causal VAE), or other network structures. The present application does not limit the network structure of the video encoder 30f. For ease of understanding, the embodiments of the present application are described by taking the video encoder 30f as a 3D Causal VAE as an example. Compared with the 3D VAE, the 3D (three-dimensional) convolution in the 3D VAE can be replaced with a 3D causal convolution to compress the video in space and time, ensuring a higher compression ratio and greatly improving the continuity of video reconstruction. Among them, the 3D causal convolution requires that the output of each time step only depends on the data of the current time step and the previous time steps, and does not depend on the future time steps.
[0158] The text encoder 30h can be used to encode the video description text. The text encoder 30h can be any one of encoders such as BERT (Bidirectional Encoder Representations from Transformers, a bidirectional encoder based on Transformers), RoBERTa (a text encoder improved on the basis of BERT), T5 (Text-to-Text Transfer Transformer, a pre-trained language model based on the Transformer architecture), etc. The present application does not limit the network structure of the text encoder 30h. For ease of understanding, the embodiments of the present application are described by taking the text encoder 30h as a T5 encoder as an example.
[0159] The video decoder 30p can be used to decode the output of the video dedicated network 30n. When the video encoder 30f is a 3D Causal VAE, the video decoder 30p can be the decoder of the 3D Causal VAE. The video dedicated network 30n can adopt a cascaded structure of a first diffusion transformation component and a second diffusion transformation component. The video dedicated network can include N first diffusion transformation components and N second diffusion transformation components, where N is a positive integer. Among the N first diffusion transformation components and N second diffusion transformation components, a second diffusion transformation component can be connected behind each first diffusion transformation component. In the embodiments of the present application, the first diffusion transformation component is a Ref-DiT component and the second diffusion transformation component is a DiT component as an example for description. Among them, the network structure of the DiT component (the second diffusion transformation component) can be the Expert Transformer block in CogVideoX (a video generation model based on a diffusion model), which can be called an expert Transformer component and can combine the advantages of the diffusion model and the Transformer structure. The network structure of the DiT component will not be described in detail in the embodiments of the present application. Compared with the DiT component, the first diffusion transformation component can add a cross-attention mechanism on the basis of the DiT component, and its network structure will be described in detail in the following content.
[0160] As Figure 4 shown, after the computer device obtains the source video 30a, it can convert the source video 30a into point cloud data 30b, obtain the camera trajectory parameters 30c custom-input by the object, and input the camera trajectory parameters 30a and the point cloud data 30b into the point cloud renderer 30d. Using the point cloud renderer 30d, perform point cloud rendering processing on the camera trajectory parameters 30a and the point cloud data 30b to obtain a point cloud rendering result 30e. The point cloud rendering result 30e can include a point cloud rendering image sequence and an occlusion mask image sequence. Among them, the process of obtaining the point cloud rendering result 30e can refer to the relevant description in step S102 and will not be elaborated here.
[0161] When the video encoder 30f is a 3D Causal VAE, the point cloud rendering result 30e can be input into the video encoder 30f of the video diffusion model. Through the video encoder 30f, the point cloud rendering result 30e can be compressed from the original high resolution and high frame rate to a lower dimension. For example, the video encoder 30f can compress the spatial dimension of the point cloud rendering result 30e from h×w to h / 8×w / 8, and the time dimension from 4f + 1 to f. The compression ratio can be 4×8×8 = 256 times, which can effectively reduce the data volume and reduce the computational cost for subsequent processing.
[0162] It can be understood that the spatial dimension of the point cloud rendering image sequence and the occlusion mask image sequence in the point cloud rendering result 30e is h×w, and the temporal dimension is 4f + 1. By performing video compression processing on the point cloud rendering image sequence and the occlusion mask image sequence through the video encoder 30f, video compression features with a spatial dimension of h / 8×w / 8 and a temporal dimension of f can be obtained. It should be noted that since the 3D Causal VAE is migrated from the image VAE, and the image VAE is not suitable for directly compressing video data, the 1 in 4f + 1 can be encoded separately. It can be understood that the first video frame (the first point cloud rendering image in the point cloud rendering image sequence, or the first occlusion mask image in the occlusion mask image sequence) is repeated as 4 frames of images, which can simulate the temporal continuity of the point cloud rendering image sequence and the occlusion mask image sequence, retain the pre-training ability of the original image VAE as much as possible, and also retain the video content in the first video frame, which can improve the video compression effect of the 3D Causal VAE.
[0163] The video compression features can be transformed into a latent space through the video encoder 30f, and this latent space can capture the latent features of the point cloud rendering result 30e, which can be called video latent features. In this latent space, the point cloud rendering result 30e is encoded into a latent representation, and this latent representation can be used to control the generation process of the target video. The 3D Causal VAE can obtain the causal structure in the video (for example, the point cloud rendering image sequence and the occlusion mask image sequence in the point cloud rendering result 30e). By introducing causal constraints, the causal relationships between different video frames and different features in the video can be mined. The video encoder 30f can incorporate the learned causal relationships into the video latent features, so that the video latent features can not only contain the semantic information of the video, but also contain causal information. Through the learned causal relationships, it is helpful to understand the dynamic changes and logical relationships in the video, and then improve the encoding and reasoning abilities for video content.
[0164] The computer device can obtain the initial noise data 30k, which can be a randomly generated Gaussian noise. The initial noise data 30k and the video latent features output by the video encoder 30f can be added together to obtain the view-noisy features. The initial noise data 30k and the video latent features have the same dimension. Furthermore, patchify (a chunking tool) can be used to split the view-noisy features into at least two feature chunks (patches), and at least two feature chunks (patches) can be constructed into a view feature vector. For example, the size of the patch can be set according to actual needs, and the view-noisy features can be split according to the size of the patch to obtain at least two patches, which can reduce the computational complexity and help capture the local features of the video frames.
[0165] For example, when the dimension of the view-noisy feature is B×c×h×w, by performing a chunking process on the view-noisy feature, a sequence with a length of B×h / p×w / p can be obtained. At this time, the sequence can be called a view feature vector; where p×p can represent the size of the set patch, and c represents the number of channels of the view-noisy feature.
[0166] After the view-noisy feature undergoes patchify, each position can be represented by a three-dimensional spatial coordinate. Therefore, rotational position encoding can be performed on each position separately, and then the rotational position encodings of each position can be concatenated along the channels to obtain the position embedding features corresponding to at least two feature chunks. Adding the position embedding features and the view feature vector together can obtain the view encoding feature 30m.
[0167] Step S104: Perform encoding processing on the source video to obtain a reference encoding feature, and perform encoding processing on the video description text to obtain a text encoding feature.
[0168] Specifically, as Figure 4 shown, the source video 30a can be input into the video encoder 30f (for example, 3D Causal VAE) in the video diffusion model. Through the 3D Causal VAE, encoding processing is performed on the source video 30a to obtain the reference latent feature of the source video 30a; among them, the process of obtaining the reference latent feature can refer to the relevant description of the view latent feature above, and will not be elaborated here. Furthermore, patchify can be used to perform a chunking process on the reference latent feature, and rotational position encoding is added to the result of the chunking process to obtain the reference encoding feature 30j. Among them, the chunking process and rotational position encoding process of the view latent feature can refer to the chunking process and rotational position encoding process of the view-noisy feature above, and will not be elaborated here. It can be understood that initial noise data 30k needs to be added during the process from the point cloud rendering result 30e to the view encoding feature 30m, while no noise data needs to be added during the process from the source video 30a to the reference encoding feature 30j.
[0169] As Figure 4 shown, after generating the video description text through the video description generator 30g, the video description text can be input into the text encoder 30h in the video diffusion model. Here, the text encoder 30h is taken as an example of the T5 encoder for description. By dividing the video description text through the text encoder 30h, the first token sequence of the source video 30a is obtained. Flag characters are added to the first token sequence, for example, a start marker (which can be the character “ <s>”), add an end marker (which can be the character "< / s> ”) is added at the beginning of the first token sequence to obtain the second token sequence.
[0170] Perform vector transformation on the second token sequence to obtain the text embedding vector of the second token sequence. For example, each token in the second token sequence can be converted into its corresponding integer index, and these integer indexes can correspond to specific tokens in the vocabulary of the T5 encoder. In the above manner, the text data can be converted into a numerical form that can be processed by the text encoder 30h. The integer index of each token in the second token sequence can be converted into a word vector to obtain the word vector sequence of the second token sequence; relative position encoding can be performed on each token in the second token sequence to obtain the position encoding feature of the second token sequence, and the position encoding feature is added to the word vector sequence to obtain the text embedding vector of the second token sequence, which helps the T5 encoder to obtain the position information of each token in the second token sequence. Among them, the T5 encoder can include a word embedding matrix, and each row in the word embedding matrix can correspond to the vector representation of a token in the vocabulary of the T5 encoder. By looking up this word embedding matrix, the word vector of each token in the second token sequence can be obtained.
[0171] The text encoder 30h can include one or more Transformer encoder blocks. A Transformer encoder block can include a multi-head self-attention mechanism and a feed-forward neural network. For each Transformer encoder block, self-attention processing can be performed on the text embedding vector through the multi-head self-attention mechanism to obtain the third self-attention feature of the text embedding vector; the third self-attention feature can be subjected to feature transformation through the feed-forward neural network to obtain the output result of a Transformer encoder block, and this output result can be used as the input data of the next Transformer encoder block until the output result of the last Transformer encoder block in the text encoder 30h is obtained. The output result of the last Transformer encoder block can be referred to as the text encoding feature 30i. Among them, the processing process of the text embedding vector of the second token sequence in one or more Transformer encoder blocks included in the text encoder 30h is the same as the processing process of the existing Transformer encoder blocks, and the embodiments of the present application will not be described one by one here.
[0172] Step S105, convert the view encoding feature into the first self-attention feature, and convert the text encoding feature into the second self-attention feature.
[0173] Specifically, such as at the end of the first meta-sequenceAs shown, the computer device can use the view encoding feature 30m and the text encoding feature 30i as the input data of the video-specific network 30n in the video diffusion model and input them into the first first diffusion transformation component in the video-specific network 30n (such as Figure 4 the first diffusion transformation component 11 shown). Similarly, the reference encoding feature 30j can also be used as the input data of the video-specific network 30n and input into each first diffusion transformation component in the video-specific network 30n. For example, it can be respectively input into the first diffusion transformation component 11, the first diffusion transformation component 12,..., the first diffusion transformation component 1N.
[0174] Among them, all the first diffusion transformation components and the second diffusion transformation components in the video-specific network 30n can be dual-path DiT components. Compared with the second diffusion transformation component, a cross-attention mechanism is added to the first diffusion transformation component. Therefore, the first diffusion transformation component can be called a reference conditional diffusion transformation component (Ref-DiT Block). Through this cross-attention mechanism, the detailed information of the source video 30a can be injected into the video encoding feature 30m.
[0175] It can be understood that the view encoding feature 30m and the text encoding feature 30i can respectively pass through one branch of the dual-path DiT component and interact in the full attention layer of the dual-path DiT component. Among them, the full attention layer in the dual-path DiT component can be a 3D full attention mechanism (3D Full Attention). The 3D full attention mechanism can operate in three-dimensional space or three-dimensional data. For example, in scenarios such as processing video data and point cloud data, it is used to obtain the long-term dependencies of data in the spatial and temporal dimensions. This 3D full attention mechanism can consider the mutual relationships between all positions instead of focusing on local or specific parts, which can improve spatio-temporal consistency.
[0176] In the first diffusion transformation component 11, the view encoding feature 30m and the text encoding feature 30i can be concatenated to obtain a first fusion feature. According to the full attention layer in the first diffusion transformation component 11, self-attention processing is performed on the first fusion feature to obtain a second fusion feature. After passing through the full attention layer in the first diffusion transformation component 11, the second fusion feature can be split into the first self-attention feature of the view encoding feature and the second self-attention feature of the text encoding feature. Among them, the data calculation process of the full attention layer can refer to the data calculation process of the multi-head self-attention mechanism in the existing Transformer encoder block, which will not be described one by one here.
[0177] It can be understood that each first diffusion transformation component and each second diffusion transformation component in the video dedicated network 30n may include a full attention layer. The input data of each full attention layer is the concatenated data of the two branches of the DiT component, and its output data needs to be split and input into the two branches respectively. The operations of concatenation and splitting are opposite. For example, when the view encoding feature 30m and the text encoding feature 30i are concatenated into the first fusion feature, assuming the dimension of the first fusion feature is 400, where the first 200-dimensional feature represents the view encoding feature 30m and the last 200-dimensional feature represents the text encoding feature 30i. Then when splitting the second fusion feature, it can also be split in the same way, taking the first half-dimension feature of the second fusion feature as the first self-attention feature and the second half-dimension feature of the second fusion feature as the second self-attention feature.
[0178] Step S106, perform cross-attention processing on the first self-attention feature and the reference encoding feature to obtain a video interaction feature, and obtain a multimodal encoding feature according to the video interaction feature and the second self-attention feature.
[0179] Specifically, in the first first diffusion transformation component of the video dedicated network (such as Figure 4 the first diffusion transformation component 11 in the video dedicated network 30n shown), a linear transformation can be performed on the reference encoding feature 30j to obtain a key matrix K 1 and a value matrix V 1 , perform a linear transformation on the first self-attention feature to obtain a query matrix Q 1 ; according to the cross-attention layer in the first first diffusion transformation component (the first diffusion transformation component 11), perform a dot product operation on the query matrix Q 1 and the transposed matrix of the key matrix K 1 to obtain a candidate weight matrix, and obtain the number of columns of the query matrix Q 1 ; normalize the ratio between the candidate weight matrix and the square root of the number of columns to obtain an attention weight matrix, and perform a dot product operation on the attention weight matrix and the value matrix V 1 to obtain a video interaction feature; in the N - 1 first diffusion models except the first first diffusion model and the N second diffusion transformation components, perform multimodal fusion on the video interaction feature, the second self-attention feature, and the reference encoding feature to obtain a multimodal encoding feature.
[0180] Please refer to Figure 4 , Figure 5 which is a schematic structural diagram of a first diffusion transformation component in a video diffusion model provided by an embodiment of the present application. Assume Figure 5 the network structure of the first diffusion transformation component 11 shown is as Figure 4As shown, the full attention layer 1 can be the full attention layer in the first diffusion transformation component 11, and the full attention layer 2 can be the full attention layer in the first second diffusion transformation component within the video-specific network, such as Figure 5 the full attention layer in the second diffusion transformation component 21 shown. Figure 4 The structure in the region 30t shown is the cross-attention mechanism in the first diffusion transformation component 11, and the structure in the region 30t is the difference between the first diffusion transformation component and the second diffusion transformation component.
[0181] The computer device can splice the view encoding feature 30m and the text encoding feature 30i to obtain the first fusion feature; through Figure 5 the full attention layer 1 shown, perform self-attention processing on the first fusion feature to obtain the second fusion feature; and then the second fusion feature can be split to obtain the first self-attention feature 30q and the second self-attention feature 30r. Among them, the first self-attention feature 30q can refer to the result obtained after the view encoding feature 30m passes through the 3D full attention mechanism, and the second self-attention feature 30r can refer to the result obtained after the text encoding feature 30i passes through the 3D full attention mechanism.
[0182] According to the linear layer 1 in the first diffusion transformation component 11, a linear transformation can be performed on the reference encoding feature 30j to obtain the key matrix K 1 and the value matrix V 1 . Among them, the linear layer 1 can include two parameter matrices, such as the parameter matrix W k and the parameter matrix W v . The two parameter matrices of the linear layer 1 can be learned during the training process. Multiplying the reference encoding feature 30j with the parameter matrix W k can obtain the key matrix K 1 ; multiplying the reference encoding feature 30j with the parameter matrix W v can obtain the value matrix V 1 .
[0183] According to the linear layer 2 in the first diffusion transformation component 11, a linear transformation can be performed on the first self-attention feature 30q to obtain the query matrix Q 1 . Among them, the linear layer 2 can include a parameter matrix, such as the parameter matrix W q , and the parameter matrix W q can be learned during the training process. Multiplying the first self-attention feature 30q with the parameter matrix W q can obtain the query matrix Q 1 .
[0184] Such as Figure 5As shown, according to the cross-attention layer in the first diffusion transformation component 11, the query matrix Q 1 can be dot-multiplied with the transposed matrix of the key matrix K 1 to obtain a candidate weight matrix (which can be denoted as ). This candidate weight matrix can be the inner product (which can also be called dot product) of each row vector in the query matrix Q 1 and the key matrix K 1 . To prevent the inner product from being too large, the number of columns corresponding to the query matrix Q 1 can be obtained. (The query matrix Q 1 has the same number of columns as the key matrix K 1 , which can also be called the vector dimension). The ratio between the candidate weight matrix and the square root of the number of columns (which can be denoted as ) is normalized to obtain an attention weight matrix. According to the dot product between the attention weight matrix and the value matrix V 1 , the video interaction feature 30s is obtained. Among them, the attention weight matrix can be expressed as . The softmax function is a function used for normalization. The softmax function can be used to calculate the attention coefficient of a single feature for other features. Through the softmax function, can be softmaxed for each row. The dot product between the attention weight matrix and the value matrix V 1 is determined as the output feature of the cross-attention layer (which can be expressed as ). At this time, the output feature of the cross-attention layer can be used as the video interaction feature 30s.
[0185] The video interaction feature 30s and the second self-attention feature 30r can be input into the second diffusion transformation component 21 in the video-specific network 30n shown in Figure 5 . For example, the video interaction feature 30s and the second self-attention feature 30r can be concatenated to obtain a third fusion feature. According to the full-attention layer in the second diffusion transformation component 21 (such as the full-attention layer 2 shown in Figure 4 ), self-attention processing is performed on the third fusion feature. Among them, the calculation process of the third fusion feature in the full-attention layer 2 can refer to the calculation process of the first fusion feature in the full-attention layer 1, which will not be elaborated here. The output result of the full-attention layer 2 still needs to be split into two parts and input into the first diffusion transformation component 22 of the video-specific network 30n shown in Figure 5 .
[0186] It should be noted that, as shown in Figure 4As shown, the data processing procedures of each first diffusion transformation component in the video-specific network 30n are similar. Except for the same reference coding features, the input data of each first diffusion transformation component are different. The data processing procedures of each second diffusion transformation component in the video-specific network 30n are also similar. The input data of each second diffusion transformation component are different, and the input data are the output results of the previous first diffusion transformation component. Here, the data processing procedures of all first diffusion transformation components and second diffusion transformation components will not be described one by one. In the embodiments of the present application, the output data of the second diffusion transformation component 2N in the video-specific network 30n can be referred to as multi-modal coding features.
[0187] It can be understood that in actual application scenarios, the video diffusion model can perform multiple predictions on the view coding feature 30m. For example, the output data of the second diffusion transformation component 2N can be re-input into the first diffusion transformation component 11 for secondary prediction until the number of predictions reaches the preset maximum number of predictions T. th During the th T-th prediction process, the output data of the second diffusion transformation component 2N are determined as multi-modal coding features. The multi-modal coding features can be used as the input data of the video decoder 30p. Using the video diffusion model for multiple predictions can improve the generation quality of the target video.
[0188] In one or more embodiments, when the computer device obtains the view coding feature, initial noise data is added, and the time step t of the initial noise data can be determined. The time step t can be understood as the number of predictions in the video generation process. The first diffusion transformation component can include a view normalization layer and a text normalization layer; in the case of the time step t, the view normalization layer can generate a first visual scale parameter and a first visual shift parameter for the view coding feature; the text normalization layer can generate a first text scale parameter and a first text shift parameter for the text coding feature.
[0189] Through the first visual scale parameter and the first visual shift parameter, the view coding feature is normalized to obtain a first view normalized feature. Through the first text scale parameter and the first text shift parameter, the text coding feature is normalized to obtain a first text normalized feature. The first view normalized feature and the first text normalized feature are concatenated to obtain a first fusion feature. Through the full attention layer included in the first diffusion transformation component in the video diffusion model, self-attention processing is performed on the first fusion feature to obtain a second fusion feature, and the second fusion feature is split into a first self-attention feature and a second self-attention feature.
[0190] Furthermore, according to the cross-attention mechanism in the first first diffusion transformation component, cross-attention processing can be performed on the first self-attention features to obtain video interaction features. The view normalization layer can generate first visual gate parameters for the video interaction features, and the text normalization layer can generate first text gate parameters for the second self-attention features. According to the first visual gate parameters, the video interaction features are normalized to obtain second view-normalized features; according to the first text gate parameters, the second self-attention features are normalized to obtain second text-normalized features.
[0191] Add the view-encoded features and the second view-normalized features to obtain first view-combined features; add the text-encoded features and the second text-normalized features to obtain first text-combined features. In the case of the time step t, the view normalization layer can generate second visual scale parameters and second visual shift parameters for the first view-combined features; the text normalization layer can generate second text scale parameters and second text shift parameters for the second text-normalized features.
[0192] Normalize the first view-combined features through the second visual scale parameters and the second visual shift parameters to obtain third view-normalized features. Normalize the first text-combined features through the second text scale parameters and the second text shift parameters to obtain third text-normalized features. Concatenate the third view-normalized features and the third text-normalized features to obtain third fused features. Through the feed-forward network layer included in the first diffusion transformation component in the video diffusion model, feature transformation is performed on the third fused features to obtain fourth fused features, and then the fourth fused features can be split into view transformation features and text transformation features.
[0193] The view normalization layer can generate second visual gate parameters for the view transformation features, and the text normalization layer can generate second text gate parameters for the text transformation features. According to the second visual gate parameters, the view transformation features are normalized to obtain fourth view-normalized features; according to the second text gate parameters, the text transformation features are normalized to obtain fourth text-normalized features.
[0194] Add the first view combination feature and the fourth view normalization feature to obtain the second view combination feature; add the first text combination feature and the fourth text normalization feature to obtain the second text combination feature. The second view combination feature and the second text combination feature at this time can be used as the output results of the two branches of the first first diffusion transformation component, and are respectively input into the first second diffusion transformation component of the video dedicated network. The first second diffusion transformation component can perform the remaining processing other than cross-attention processing on the input data to obtain the output data of the first second diffusion transformation component, and so on, until the output data of the Nth second diffusion transformation component is obtained, and the output data of the Nth second diffusion transformation component is used as the multi-modal feature.
[0195] Step S107: Decode the multi-modal encoded feature to generate the target video indicated by the camera trajectory parameter.
[0196] Specifically, as Figure 4 shown, the multi-modal encoded feature can be input into the video decoder 30p in the video diffusion model, and the video decoder 30p performs upsampling processing on the multi-modal encoded feature to obtain the reconstructed video data of the source video; furthermore, the reconstructed video data can be activated to obtain the target video 30q with the same format as the source video. Among them, the source video 30a is a monocular video, and the target video 30q can refer to a 4D new view video predicted based on the camera trajectory parameter on the basis of the source video 30a. The target video 30q can be a multi-view video.
[0197] In the embodiment of the present application, the source video, the camera trajectory parameter, and the video description text of the source video can be obtained, the source video is converted into point cloud data, and the point cloud data is subjected to point cloud rendering processing according to the camera trajectory parameter to obtain a point cloud rendering result, and the point cloud rendering result is encoded into a view encoded feature. By performing self-attention processing on the view encoded feature and the text encoded feature of the video description text, geometric alignment can be achieved and the accuracy of the view can be maintained; by separating the deterministic view transformation (point cloud rendering processing) from the random content generation (diffusion model), precise camera control can be achieved. By performing cross-attention processing on the view encoded feature and the reference encoded feature of the source video, the detailed information in the source video can be injected, the spatio-temporal consistency between the target video and the source video can be improved, the repair effectiveness of the occlusion area in the point cloud rendering image and the coherence of the dynamic scene can be improved, and thus the generation quality of the target video can be improved.
[0198] The training process of the video diffusion model will be described below. Please refer to Figure 4 , Figure 6It is a schematic diagram of the training of a video diffusion model provided by an embodiment of the present application; it can be understood that the training process of the video diffusion model can be executed by a computer device, and the computer device can be a terminal device, such as Figure 6 any one of the terminal devices in the terminal cluster shown, or it can be a server, such as Figure 1 the server 10d shown, and the present application does not make any limitations in this regard. As Figure 1 shown, this can include the following steps S201 to step S207:
[0199] Step S201, obtain a sample monocular video, sample the camera transformation matrix of the sample monocular video, and generate a candidate view sequence of the sample monocular video according to the camera transformation matrix.
[0200] Specifically, the computer device can obtain a sample monocular video, and the number of the sample monocular videos can be multiple. For example, a single camera can be used to shoot the same scene or object from different perspectives to obtain multiple sample monocular videos; or a single camera can be used to shoot different objects or scenes from the same angle to obtain multiple sample monocular videos; or a single camera can be used to shoot different scenes or objects from different angles to obtain multiple sample monocular videos; or AI technology can be used to synthesize sample monocular videos, etc. Through the above methods, a large number of sample monocular videos can be collected.
[0201] For any sample monocular video, the camera transformation matrix of the sample monocular video can be sampled, and according to the camera transformation matrix, the sample monocular video can be transformed to generate a candidate view sequence of the sample monocular video. Among them, the camera transformation matrix can be a 4×4 matrix describing the position, attitude and projection relationship of the camera in the three-dimensional space, which can be used to convert points in the three-dimensional space to the coordinate system of the camera, and then map them to the two-dimensional image plane through projection transformation. In order to generate views from different perspectives, the camera transformation matrix can be randomly sampled within a certain range to increase the diversity of the candidate view sequence. The candidate view sequence can refer to the view sequence generated by mapping the image information in the sample monocular video to a new perspective, and the candidate view sequence can include one or more candidate views.
[0202] Step S202, perform inverse projection on the candidate view sequence to obtain a first point cloud rendering video, construct a video training pair from the sample monocular video and the first point cloud rendering video, and add the video training pair to the first sample training set.
[0203] Specifically, the computer device can use the inverse matrix of the camera transformation matrix to restore the candidate views in the candidate view sequence to the original views; that is, use the inverse matrix of the camera transformation matrix to recover the candidate view sequence. For example, it can map two-dimensional image points back to the three-dimensional space coordinates corresponding to the original views to obtain sample point cloud data. By performing point cloud rendering processing on the sample point cloud data, a first point cloud rendering video can be obtained, and the first point cloud rendering video can include multiple point cloud rendering diagrams.
[0204] Among them, in the actual scenario, there may be occlusion relationships between different objects in the sample monocular video. Therefore, when generating the first point cloud rendering video, the point cloud rendering diagrams in the first point cloud rendering video may include occluded areas. The sample monocular video and the first point cloud rendering video can form a video training pair, and a large number of video training pairs can constitute a first sample training set. The first sample training set can include multiple video training pairs, and each video training pair can include a sample monocular video and the first point cloud rendering video corresponding to the sample monocular video. The sample monocular video in a video training pair can be used as the real video (expected result) during the training process of the video diffusion model, and the first point cloud rendering video can be used as the input video during the training process of the video diffusion model.
[0205] Step S203: Obtain multi-view static data, construct the global point cloud data of the multi-view static data, and sample a first-view video and a second-view video from the multi-view static data.
[0206] Among them, the multi-view static data can refer to a data set obtained by collecting and recording a static scene, object, or phenomenon from multiple different angles or viewpoints. The number of the multi-view static data can be one or more. A reconstruction method, such as MASt3R (a method for 3D reconstruction), can be used to construct the global point cloud data of the multi-view static data. Two segments with different viewpoints can be sampled from the multi-view static data of the same scene, and each segment can include one or more images with different viewpoints. These two segments can be called the first-view video and the second-view video. The first-view video and the second-view video have different viewpoints. The different images in the first-view video can have the same or different viewpoints, and the different images in the second-view video can also have the same or different viewpoints.
[0207] Step S204: Obtain the target camera parameters of the first-view video, perform point cloud rendering processing on the global point cloud data according to the target camera parameters to obtain a second point cloud rendering video, construct the sample triple with the first-view video, the second-view video, and the second point cloud rendering video, and add the sample triple to the second sample training set.
[0208] Specifically, the computer device can obtain the target camera parameters of the first - perspective video, which can refer to the camera parameters calculated based on the depth information of the first - perspective video. According to the target camera parameters, point - cloud rendering processing can be performed on the point - cloud data of the second - perspective video to obtain the second point - cloud rendering video. The first - perspective video, the second - perspective video, and the second point - cloud rendering video can be constructed into a sample triple, and a large number of sample triples can form the second sample training set.
[0209] Among them, the second sample training set can include multiple sample triples, and each sample triple can include the first - perspective video, the second - perspective video, and the second point - cloud rendering video. The first - perspective video can be used as the target video during the training process of the video diffusion model, the second - perspective video can be used as the source video during the training process of the video diffusion model, and the second point - cloud rendering video can be used as the input data during the training process of the video diffusion model.
[0210] Step S205: Obtain the diffusion model to be trained, where the diffusion model to be trained includes a video encoder to be trained, a text encoder to be trained, a dedicated network to be trained, and a video decoder to be trained.
[0211] Among them, the diffusion model to be trained can refer to a two - stream conditional video diffusion model that needs to be trained. The network structure of the diffusion model to be trained is the same as that of the video diffusion model, such as Figure 6 the network structure of the video diffusion model shown, which will not be elaborated here. Among them, the video encoder in the training stage can be called the video encoder to be trained, the text encoder in the training stage can be called the text encoder to be trained, the video dedicated network in the training stage can be called the dedicated network to be trained, and the video decoder in the training stage can be called the video decoder to be trained.
[0212] Step S206: Fix the network parameters of the cross - attention layer in the dedicated network to be trained, and train the diffusion model to be trained according to multiple video training pairs to obtain a pre - trained diffusion model.
[0213] Specifically, the training process involved in the embodiments of the present application can include two stages, which are respectively denoted as the first training stage and the second training stage. In the first training stage, the network parameters of the cross - attention mechanism in the first diffusion transformation component included in the dedicated network to be trained can be fixed. That is to say, the network parameters of the cross - attention mechanism in the first diffusion transformation component are fixed and unchanged in the first training stage, such as Figure 4 the network parameters of the linear layer 1, the linear layer 2, and the cross - attention layer shown, which remain unchanged in the first training stage.
[0214] In the first training stage, a first sample description text of the sample monocular video in the video training pair can be generated by a video description generator. The generation process of the first sample description text can refer to the generation process of the aforementioned video description text, which will not be elaborated here. The first point cloud rendering video in the video training pair can be input into the to-be-trained video encoder in the to-be-trained diffusion model. The to-be-trained video encoder encodes the first point cloud rendering video, and diffusion noise data is added to the output data of the to-be-trained video encoder to obtain a sample noisy view feature. Furthermore, the sample noisy view feature can be block-processed, rotationally position-encoded, etc. to obtain a first sample view encoding feature.
[0215] The first sample description text is input into the to-be-trained text encoder in the to-be-trained diffusion model. The to-be-trained text encoder encodes the first sample description text, and relative position encoding is performed on the output data of the to-be-trained text encoder to obtain a first sample text encoding feature. The process of obtaining the first sample text encoding feature can refer to the relevant description of the process of obtaining the text encoding feature in step S104, which will not be elaborated here.
[0216] The first sample text encoding feature and the first sample view encoding feature can be input into the to-be-trained dedicated network in the to-be-trained diffusion model. Through N first diffusion transformation components and N second diffusion transformation components in the to-be-trained dedicated network, the first sample text encoding feature and the first sample view encoding feature are calculated to obtain a first sample multimodal feature. The first sample multimodal feature is input into the to-be-trained video decoder, and the to-be-trained video decoder decodes the first sample multimodal feature to generate a first predicted video. According to the error between the first predicted video and the sample monocular video in the video training pair, other network parameters in the to-be-trained diffusion model except for the network parameters of the cross-attention mechanism are corrected to obtain a pre-trained diffusion model. The loss function of the to-be-trained diffusion model can be the loss function of the current diffusion model, which will not be described one by one here. Through the first training stage, geometric distortion can be repaired, and the video generation quality of the pre-trained diffusion model can be improved.
[0217] Step S207: Correct the network parameters of the cross-attention layer in the pre-trained diffusion model according to multiple sample triples to obtain a video diffusion model.
[0218] Specifically, multiple sample triples in the second sample training set can be used to fine-tune the network parameters of the cross-attention mechanism in the pre-trained diffusion model to enhance 4D consistency. Among them, the second perspective video in the sample triple can be regarded as Figure 5 the source video in the corresponding embodiment, and the second point cloud rendering video can be regarded as Figure 3For the point cloud rendering results in the corresponding embodiments, the calculation process of the second - perspective video and the second point cloud rendering video in the pre - trained diffusion model can refer to the calculation process of the source video and the point cloud rendering results in the video diffusion model, which will not be elaborated here.
[0219] For example, in the second training stage, the video description generator can generate the second sample description text of the second - perspective video in the sample triple. Through the text encoder in the pre - trained diffusion model, the second sample text encoding features of the second sample description text can be obtained. The second point cloud rendering video in the sample triple can be input into the video encoder in the pre - trained diffusion model. Through the video encoder in the pre - trained diffusion model, the second sample view encoding features can be obtained. The second - perspective video in the sample triple can be input into the video encoder in the pre - trained diffusion model. Through the video encoder in the pre - trained diffusion model, the sample reference encoding features can be obtained, and the sample reference encoding features can be input into the N first diffusion transformation components in the pre - trained diffusion model.
[0220] The second sample text encoding features, the second sample view encoding features, and the sample reference encoding features can be input into the video - specific network in the pre - trained diffusion model. Through the N first diffusion transformation components and N second diffusion transformation components in the video - specific network, the second sample text encoding features, the second sample view encoding features, and the sample reference encoding features are calculated, and the second sample multimodal features can be obtained. The second sample multimodal features are input into the video decoder in the pre - trained diffusion model. Through this video decoder, the second predicted video can be generated. According to the error between the second predicted video and the first - perspective video in the sample triple, the network parameters of the pre - trained diffusion model can be corrected, and the video diffusion model can be obtained.
[0221] In the embodiments of this application, by performing double reprojection on the monocular video and using multi - perspective static data to expand the sample training set, the dependence on pair - perspective data can be reduced. Through the first training stage and the second training stage, by simultaneously using dynamic monocular videos and static multi - perspective data for training, the generalization of the video diffusion model can be improved.
[0222] It can be understood that in the specific implementation of this application, it may involve relevant information such as the user's login information in the client or browser, video upload records, etc. When the above embodiments of this application are applied to specific products or technologies, the permission or consent of relevant institutions or departments, or the user himself / herself is required, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards in the relevant regions.
[0223] Please refer to Figure 3 , Figure 7It is a schematic structural diagram of a video generation device provided by an embodiment of the present application. As Figure 7 shown, the video generation device 1 may include: a data acquisition module 101, a point cloud rendering module 102, an encoding module 103, a first attention processing module 104, a second attention processing module 105, and a decoding module 106;
[0224] The data acquisition module 101 is configured to acquire a source video and camera trajectory parameters, and acquire a video description text of the source video;
[0225] The point cloud rendering module 102 is configured to convert the source video into point cloud data, perform point cloud rendering processing on the point cloud data according to the camera trajectory parameters to obtain a point cloud rendering result, and perform encoding processing on the point cloud rendering result to obtain a view encoding feature;
[0226] The encoding module 103 is configured to perform encoding processing on the source video to obtain a reference encoding feature, and perform encoding processing on the video description text to obtain a text encoding feature;
[0227] The first attention processing module 104 is configured to convert the view encoding feature into a first self-attention feature, and convert the text encoding feature into a second self-attention feature;
[0228] The second attention processing module 105 is configured to perform cross-attention processing on the first self-attention feature and the reference encoding feature to obtain a video interaction feature, and obtain a multi-modal encoding feature according to the video interaction feature and the second self-attention feature;
[0229] The decoding module 106 is configured to perform decoding processing on the multi-modal encoding feature to generate a target video indicated by the camera trajectory parameters.
[0230] In one or more embodiments, when the data acquisition module 101 acquires the video description text of the source video, it is specifically configured to perform the following steps:
[0231] Perform frame splitting on the source video to obtain a video frame sequence of the source video, perform feature extraction on each video frame in the video frame sequence to obtain visual features of each video frame;
[0232] Perform encoding processing on the visual features of each video frame to obtain image semantic features of the video frame sequence;
[0233] Perform decoding processing on the image semantic features to generate a video description text of the source video.
[0234] In one or more embodiments, when the point cloud rendering module 102 converts the source video into point cloud data, it is specifically configured to perform the following steps:
[0235] Divide the source video into multiple video segments, generate the initial noise distribution for each of the multiple video segments, and the initial noise distribution for each video segment is used to characterize the initial depth distribution of a video segment;
[0236] Extract features from each video segment to obtain the segment content features of each video segment, and adjust the initial noise distribution of each video segment according to the segment content features to obtain the depth distribution estimate of each video segment;
[0237] Stitch the depth distribution estimates of each video segment to obtain the video depth sequence of the source video;
[0238] According to the video depth sequence, perform inverse perspective projection on each video frame in the source video to obtain point cloud data.
[0239] In one or more embodiments, when the point cloud rendering module 102 stitches the depth distribution estimates of each video segment to obtain the video depth sequence of the source video, it is specifically used to perform the following steps:
[0240] Determine M video segment pairs among the multiple video segments, where the two video segments in each video segment pair are two adjacent video segments among the multiple video segments, and M is a positive integer;
[0241] Determine the interpolation parameters between the two video segments in each video segment pair according to the overlapping data between the two video segments in each video segment pair and the depth distribution estimates of the two video segments in each video segment pair;
[0242] Fuse the depth distribution estimates of the two video segments in each video segment pair according to the interpolation parameters to obtain the video depth sequence of the source video.
[0243] In one or more embodiments, when the point cloud rendering module 102 performs inverse perspective projection on each video frame in the source video according to the video depth sequence to obtain point cloud data, it is specifically used to perform the following steps:
[0244] Determine the original camera parameters of the source video, and construct an inverse perspective projection matrix according to the original camera parameters;
[0245] Obtain the image coordinates of each video frame in the source video, and convert the image coordinates of each video frame into three-dimensional space coordinates according to the inverse perspective projection matrix and the video depth sequence to obtain the first set of spatial coordinates;
[0246] Filter the invalid points in the first set of spatial coordinates to obtain the second set of spatial coordinates, and generate point cloud data according to the second set of spatial coordinates.
[0247] In one or more embodiments, when the point cloud rendering module 102 performs point cloud rendering processing on point cloud data according to camera trajectory parameters to obtain a point cloud rendering result, it is specifically used to perform the following steps:
[0248] Create a camera object according to the camera trajectory parameters, configure renderer parameters for the point cloud renderer, and perform rendering processing on the point cloud data and the camera object according to the renderer parameters to obtain a sequence of point cloud rendering images;
[0249] Generate a sequence of occlusion mask images for the sequence of point cloud rendering images according to the depth information of each point cloud rendering image in the sequence of point cloud rendering images, and determine the sequence of point cloud rendering images and the sequence of occlusion mask images as the point cloud rendering result.
[0250] In one or more embodiments, when the point cloud rendering module 102 performs encoding processing on the point cloud rendering result to obtain view encoding features, it is specifically used to perform the following steps:
[0251] Input the point cloud rendering result into the video encoder in the video diffusion model, and perform video compression processing on the point cloud rendering result through the video encoder to obtain video compression features;
[0252] Convert the video compression features into view latent features, generate initial noise data with the same dimension as the view latent features, and combine the view latent features and the initial noise data to obtain view-noisy features;
[0253] Perform block processing on the view-noisy features to obtain at least two feature blocks, and construct a view feature vector according to the at least two feature blocks;
[0254] Perform rotational position encoding on the at least two feature blocks to obtain position embedding features corresponding to the at least two feature blocks respectively;
[0255] Combine the view feature vector and the position embedding features corresponding to the at least two feature blocks respectively to obtain view encoding features.
[0256] In one or more embodiments, when the encoding module 103 performs encoding processing on the video description text to obtain text encoding features, it is specifically used to perform the following steps:
[0257] Input the video description text into the text encoder in the video diffusion model, and divide the video description text through the text encoder to obtain a first token sequence of the source video;
[0258] Add flag characters to the first token sequence to obtain a second token sequence, and perform vector conversion on the second token sequence to obtain text embedding vectors of the second token sequence;
[0259] Perform self-attention processing on the text embedding vector to obtain the third self-attention feature of the text embedding vector, and perform feature transformation on the third self-attention feature to obtain the text encoding feature.
[0260] In one or more embodiments, when the first attention processing module 104 converts the view encoding feature into the first self-attention feature and converts the text encoding feature into the second self-attention feature, it is specifically used to perform the following steps:
[0261] Input the view encoding feature and the text encoding feature into the video-specific network in the video diffusion model. The video-specific network adopts a cascaded structure of the first diffusion transformation component and the second diffusion transformation component. The video-specific network includes N first diffusion transformation components and N second diffusion transformation components, where N is a positive integer;
[0262] Concatenate the view encoding feature and the text encoding feature to obtain the first fusion feature;
[0263] According to the full attention layer in the first first diffusion transformation component, perform self-attention processing on the first fusion feature to obtain the second fusion feature;
[0264] Split the second fusion feature into the first self-attention feature of the view encoding feature and the second self-attention feature of the text encoding feature.
[0265] In one or more embodiments, when the second attention processing module 105 performs cross-attention processing on the first self-attention feature and the reference encoding feature to obtain the video interaction feature, and obtains the multi-modal encoding feature according to the video interaction feature and the second self-attention feature, it is specifically used to perform the following steps:
[0266] Input the reference encoding feature into N first diffusion transformation components. In the first first diffusion transformation component, perform linear transformation on the reference encoding feature to obtain the key matrix and the value matrix, and perform linear transformation on the first self-attention feature to obtain the query matrix;
[0267] According to the cross-attention layer in the first first diffusion transformation component, perform dot product operation on the query matrix and the transposed matrix of the key matrix to obtain the candidate weight matrix, and obtain the number of columns of the query matrix;
[0268] Normalize the ratio between the candidate weight matrix and the square root of the number of columns to obtain the attention weight matrix, and perform dot product operation on the attention weight matrix and the value matrix to obtain the video interaction feature;
[0269] In the N - 1 first diffusion models except the first first diffusion model and the N second diffusion transformation components, perform multi-modal fusion on the video interaction feature, the second self-attention feature, and the reference encoding feature to obtain the multi-modal encoding feature.
[0270] In one or more embodiments, when the decoding module 106 decodes the multi-modal encoded features to generate the target video indicated by the camera trajectory parameters, it is specifically configured to perform the following steps:
[0271] Input the multi-modal encoded features into the video decoder in the video diffusion model, and perform upsampling processing on the multi-modal encoded features through the video decoder to obtain the reconstructed video data of the source video;
[0272] Perform activation processing on the reconstructed video data to obtain the target video with the same format as the source video.
[0273] In one or more embodiments, the video generation device 1 further includes: a training set acquisition module 107, a model acquisition module 108, and a model training module 109;
[0274] The training set acquisition module 107 is configured to acquire a first sample training set and a second sample training set. The first sample training set includes multiple video training pairs, and each video training pair includes a sample monocular video and a first point cloud rendering video. The second sample training set includes multiple sample triples, and each sample triple includes a first perspective video, a second perspective video, and a second point cloud rendering video, and the first perspective video and the second perspective video have different perspectives;
[0275] The model acquisition module 108 is configured to acquire a diffusion model to be trained, and the diffusion model to be trained includes a video encoder to be trained, a text encoder to be trained, a dedicated network to be trained, and a video decoder to be trained;
[0276] The model training module 109 is configured to fix the network parameters of the cross-attention layer in the dedicated network to be trained, and train the diffusion model to be trained according to multiple video training pairs to obtain a pre-trained diffusion model;
[0277] The model training module 109 is further configured to correct the network parameters of the cross-attention layer in the pre-trained diffusion model according to multiple sample triples to obtain a video diffusion model.
[0278] In one or more embodiments, when the training set acquisition module 107 acquires the first sample training set and the second sample training set, it is specifically configured to perform the following steps:
[0279] Acquire a sample monocular video, sample the camera transformation matrix of the sample monocular video, and generate a candidate view sequence of the sample monocular video according to the camera transformation matrix;
[0280] Perform inverse projection on the candidate view sequence to obtain a first point cloud rendering video, construct a video training pair from the sample monocular video and the first point cloud rendering video, and add the video training pair to the first sample training set;
[0281] Obtain multi-view static data, construct global point cloud data of the multi-view static data, and sample a first-view video and a second-view video from the multi-view static data;
[0282] Obtain the target camera parameters of the first-view video. According to the target camera parameters, perform point cloud rendering processing on the global point cloud data to obtain a second point cloud rendering video. Construct a sample triple from the first-view video, the second-view video, and the second point cloud rendering video, and add the sample triple to the second sample training set.
[0283] According to an embodiment of the present application, the relevant steps involved in the video generation method described above Figure 7 can be executed by each module in the video generation device 1 shown in Figure 3 For example, Figure 7 the step S101 shown in Figure 3 can be executed by the data acquisition module 101 shown in Figure 7 the steps S102 and S103 shown in Figure 3 can be executed by the point cloud rendering module 102 shown in Figure 7 the step S104 shown in Figure 3 can be executed by the encoding module 103 shown in Figure 7 the step S105 shown in Figure 3 can be executed by the first attention processing module 104 shown in Figure 7 the step S106 shown in Figure 3 can be executed by the second attention processing module 105 shown in Figure 7 the step S107 shown in Figure 3 can be executed by the decoding module 106 shown in, etc.
[0284] According to an embodiment of the present application, Figure 7 each module in the video generation device 1 shown in can be separately or all combined into one or several modules to form, or a certain one (or some) of the modules can be further split into at least two smaller functional units, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above modules are divided based on logical functions. In practical applications, the function of one module can also be realized by at least two units, or the functions of at least two modules are realized by one module. In other embodiments of the present application, the video generation device 1 may also include other modules or units. In practical applications, these functions can also be assisted by other modules and can be realized by the cooperation of at least two modules.
[0285] In an embodiment of the present application, a source video, camera trajectory parameters, and a video description text of the source video can be obtained, the source video can be converted into point cloud data, and point cloud rendering processing is performed on the point cloud data according to the camera trajectory parameters to obtain a point cloud rendering result, and the point cloud rendering result is encoded as a view coding feature. By performing self-attention processing on the view coding features and the text coding features of the video description text, geometric alignment can be achieved to maintain the accuracy of the view. Cross-attention processing is performed on the view coding features and the reference coding features of the source video, and the detailed information in the source video can be injected, which improves the spatiotemporal consistency of the target video and the source video, thereby improving the generation quality of the target video.
[0286] See also Figure 7 , Figure 8 Schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 8 As shown, the computer device 1000 may be a terminal device, for example, Figure 8 The terminal device 10a in the corresponding embodiment may also be a server, for example, Figure 1 The server 10d in the corresponding embodiment will not be limited here. For ease of understanding, this application takes a computer device as an example of a terminal device. The computer device 1000 may include: a processor 1001, a network interface 1004 and a memory 1005. In addition, the computer device 1000 may also include: a user interface 1003, and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, the user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or it may be a non-volatile memory (non-volatile memory), such as at least one disk storage. The memory 1005 may optionally also be at least one storage device located away from the aforementioned processor 1001. As Figure 1 As shown, the memory 1005 as a computer-readable storage medium may include an operating system, a network communication module, a user interface module, and a device control application program.
[0287] The network interface 1004 in the computer device 1000 can also provide a network communication function, and the optional user interface 1003 can also include a display screen (Display) and a keyboard (Keyboard). Figure 8In the computer device 1000 shown, the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an interface for users to input; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement:
[0288] Obtain the source video and camera trajectory parameters, and obtain the video description text of the source video;
[0289] Convert the source video into point cloud data, perform point cloud rendering processing on the point cloud data according to the camera trajectory parameters to obtain a point cloud rendering result, and perform encoding processing on the point cloud rendering result to obtain view encoding features;
[0290] Perform encoding processing on the source video to obtain reference encoding features, and perform encoding processing on the video description text to obtain text encoding features;
[0291] Convert the view encoding features into first self-attention features, and convert the text encoding features into second self-attention features;
[0292] Perform cross-attention processing on the first self-attention features and the reference encoding features to obtain video interaction features, and obtain multi-modal encoding features according to the video interaction features and the second self-attention features;
[0293] Perform decoding processing on the multi-modal encoding features to generate the target video indicated by the camera trajectory parameters.
[0294] It should be understood that the computer device 1000 described in the embodiments of the present application can execute the description of the video generation method in any one of the foregoing Figure 8 、 Figure 3 embodiments, and can also execute the description of the video generation device 1 in the corresponding embodiment of the foregoing Figure 6 which will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either.
[0295] In addition, it should be pointed out here that: The embodiments of the present application also provide a computer-readable storage medium, and the computer-readable storage medium stores the computer program executed by the foregoing video generation device 1, and the computer program includes computer instructions. When the processor executes the computer instructions, it can execute the foregoing Figure 7 、 Figure 3Any description of the video generation method in the embodiments will not be repeated here. In addition, the beneficial effects of using the same method will not be described again. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), or the like. For technical details not disclosed in the embodiments of the computer-readable storage medium involved in this application, please refer to the description of the method embodiments of this application. As an example, the program instructions can be deployed to be executed on a computer device, or on multiple computer devices located at one location, or on multiple computer devices distributed at multiple locations and interconnected by a communication network. The multiple computer devices distributed at multiple locations and interconnected by a communication network can form a blockchain system.
[0296] In addition, it should be noted that: The embodiments of this application also provide a computer program product, which may include a computer program that can be stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor can execute the computer program, so that the computer device executes the Figure 6 , Figure 3 Any description of the video generation method in the embodiments will not be repeated here. In addition, the beneficial effects of using the same method will not be described again. For technical details not disclosed in the embodiments of the computer program product or computer program involved in this application, please refer to the description of the method embodiments of this application.
[0297] The terms "first", "second", etc. in the specification, claims, and drawings of the embodiments of this application are used to distinguish different media contents, rather than to describe a specific order. In addition, the term "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product, or equipment that includes a series of steps or units is not limited to the listed steps or modules, but may optionally further include steps or modules not listed, or may optionally further include other step units inherent to these processes, methods, devices, products, or equipment.
[0298] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0299] The methods and related devices provided in the embodiments of this application are described with reference to the method flowcharts and / or structural schematic diagrams provided in the embodiments of this application. Specifically, each process and / or block of the method flowchart and / or structural schematic diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 6 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 Figure 1 one process or multiple processes and / or structural schematic one block or multiple blocks.
[0300] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of the module or unit.
[0301] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of the rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A video generation method, characterized in that: include: Obtain source video and camera trajectory parameters, and obtain video description text of the source video; Convert the source video into point cloud data, perform point cloud rendering processing on the point cloud data according to the camera trajectory parameters to obtain a point cloud rendering result, and perform encoding processing on the point cloud rendering result to obtain a view encoding feature; Encoding the source video to obtain reference encoding features, and encoding the video description text to obtain text encoding features; Converting the view encoding feature into a first self-attention feature, and converting the text encoding feature into a second self-attention feature; Performing cross-attention processing on the first self-attention feature and the reference coding feature to obtain a video interaction feature, and obtaining a multimodal coding feature according to the video interaction feature and the second self-attention feature; The multimodal coding features are decoded to generate a target video indicated by the camera trajectory parameters.
2. The method according to claim 1, characterized in that The step of obtaining the video description text of the source video includes: The source video is processed by frame division to obtain a video frame sequence of the source video, and features are extracted from each video frame in the video frame sequence to obtain visual features of each video frame; Encoding the visual features of each video frame to obtain image semantic features of the video frame sequence; The image semantic features are decoded to generate a video description text of the source video.
3. The method according to claim 1, characterized in that The converting the source video into point cloud data comprises: Dividing the source video into a plurality of video segments, generating an initial noise distribution of each of the plurality of video segments, wherein the initial noise distribution of each video segment is used to characterize an initial depth distribution of a video segment; Extracting features from each of the video clips to obtain a clip content feature of each of the video clips, and adjusting an initial noise distribution of each of the video clips according to the clip content feature to obtain a depth distribution estimate of each of the video clips; splicing the depth distribution estimation of each video clip to obtain a video depth sequence of the source video; According to the video depth sequence, each video frame in the source video is subjected to inverse perspective projection to obtain point cloud data.
4. The method according to claim 3, characterized in that: The step of splicing the depth distribution estimation of each video clip to obtain the video depth sequence of the source video includes: Determine M video segment pairs from the multiple video segments, where two video segments in each video segment pair are two adjacent video segments from the multiple video segments, and M is a positive integer; Determining an interpolation parameter between the two video segments in each video segment pair according to overlapping data between the two video segments in each video segment pair and depth distribution estimates of the two video segments in each video segment pair; According to the interpolation parameters, the depth distribution estimates of the two video segments in each video segment pair are fused to obtain a video depth sequence of the source video.
5. The method according to claim 3, characterized in that: The step of performing inverse perspective projection on each video frame in the source video according to the video depth sequence to obtain point cloud data includes: Determine original camera parameters of the source video, and construct an inverse perspective projection matrix according to the original camera parameters; Obtaining image coordinates of each video frame in the source video, and converting the image coordinates of each video frame into three-dimensional space coordinates according to the inverse perspective projection matrix and the video depth sequence to obtain a first space coordinate set; Invalid points in the first spatial coordinate set are filtered to obtain a second spatial coordinate set, and point cloud data is generated according to the second spatial coordinate set.
6. The method according to claim 1, characterized in that The step of performing point cloud rendering processing on the point cloud data according to the camera trajectory parameters to obtain a point cloud rendering result includes: Creating a camera object according to the camera trajectory parameters, configuring renderer parameters for a point cloud renderer, and rendering the point cloud data and the camera object according to the renderer parameters to obtain a point cloud rendering image sequence; An occlusion mask image sequence of the point cloud rendering image sequence is generated according to the depth information of each point cloud rendering image in the point cloud rendering image sequence, and the point cloud rendering image sequence and the occlusion mask image sequence are determined as the point cloud rendering result.
7. The method according to claim 1, characterized in that The encoding process of the point cloud rendering result to obtain a view encoding feature includes: Inputting the point cloud rendering result into a video encoder in a video diffusion model, and performing video compression processing on the point cloud rendering result by the video encoder to obtain a video compression feature; Converting the video compression feature into a view latent feature, generating initial noise data having the same dimension as the view latent feature, and combining the view latent feature with the initial noise data to obtain a view noise-added feature; Performing block processing on the view noise feature to obtain at least two feature blocks, and constructing a view feature vector according to the at least two feature blocks; Performing rotation position encoding on the at least two feature blocks to obtain position embedding features corresponding to the at least two feature blocks respectively; The view feature vector and the position embedding features corresponding to the at least two feature blocks are combined to obtain the view coding feature.
8. The method according to claim 1, characterized in that The encoding process of the video description text to obtain text encoding features includes: Inputting the video description text into a text encoder in a video diffusion model, dividing the video description text by the text encoder, and obtaining a first word sequence of the source video; Adding a flag character to the first word-gram sequence to obtain a second word-gram sequence, performing vector conversion on the second word-gram sequence to obtain a text embedding vector of the second word-gram sequence; Perform self-attention processing on the text embedding vector to obtain a third self-attention feature of the text embedding vector, and perform feature transformation on the third self-attention feature to obtain the text encoding feature.
9. The method according to claim 1, characterized in that: The converting the view encoding feature into a first self-attention feature and converting the text encoding feature into a second self-attention feature comprises: Inputting the view encoding feature and the text encoding feature into a video-specific network in a video diffusion model, wherein the video-specific network adopts a cascade structure of a first diffusion transformation component and a second diffusion transformation component, and the video-specific network includes N first diffusion transformation components and N second diffusion transformation components, where N is a positive integer; Concatenate the view encoding feature and the text encoding feature to obtain a first fusion feature; According to the full attention layer in the first diffusion transform component, the first fused feature is self-attention processed to obtain a second fused feature; The second fused feature is split into a first self-attention feature of the view encoding feature and a second self-attention feature of the text encoding feature.
10. The method according to claim 9, characterized in that The step of performing cross-attention processing on the first self-attention feature and the reference coding feature to obtain a video interaction feature, and obtaining a multimodal coding feature according to the video interaction feature and the second self-attention feature, comprises: Input the reference coding feature into the N first diffusion transformation components, in the first first diffusion transformation component, perform a linear transformation on the reference coding feature to obtain a key matrix and a value matrix, and perform a linear transformation on the first self-attention feature to obtain a query matrix; According to the cross attention layer in the first diffusion transformation component, performing a dot product operation on the query matrix and the transposed matrix of the key matrix to obtain a candidate weight matrix, and obtaining the number of columns of the query matrix; Normalizing the ratio between the candidate weight matrix and the square root of the number of columns to obtain an attention weight matrix, and performing a dot multiplication operation on the attention weight matrix and the value matrix to obtain the video interaction feature; In the N-1 first diffusion models other than the first first diffusion model, and the N second diffusion transformation components, the video interaction feature, the second self-attention feature and the reference coding feature are multimodally fused to obtain a multimodal coding feature.
11. The method according to claim 1, characterized in that: The decoding process of the multimodal coding feature to generate a target video indicated by the camera trajectory parameter includes: Inputting the multimodal coding features into a video decoder in a video diffusion model, and performing upsampling processing on the multimodal coding features through the video decoder to obtain reconstructed video data of the source video; The reconstructed video data is activated to obtain a target video having the same format as the source video.
12. The method according to any one of claims 1 to 10, characterized in that: The method further comprises: Obtain a first sample training set and a second sample training set, wherein the first sample training set includes a plurality of video training pairs, each of which includes a sample monocular video and a first point cloud rendering video, and the second sample training set includes a plurality of sample triplets, each of which includes a first-perspective video, a second-perspective video, and a second point cloud rendering video, wherein the first-perspective video and the second-perspective video have different perspectives; Acquire a diffusion model to be trained, wherein the diffusion model to be trained includes a video encoder to be trained, a text encoder to be trained, a dedicated network to be trained, and a video decoder to be trained; Fixing the network parameters of the cross attention layer in the dedicated network to be trained, and training the diffusion model to be trained according to the multiple video training pairs to obtain a pre-trained diffusion model; According to the multiple sample triplets, the network parameters of the cross-attention layer in the pre-trained diffusion model are modified to obtain a video diffusion model.
13. The method according to claim 12, characterized in that The step of obtaining the first sample training set and the second sample training set comprises: Acquire a sample monocular video, sample a camera transformation matrix for the sample monocular video, and generate a candidate view sequence for the sample monocular video according to the camera transformation matrix; Performing reverse projection on the candidate view sequence to obtain the first point cloud rendering video, constructing the sample monocular video and the first point cloud rendering video into a video training pair, and adding the video training pair to a first sample training set; Acquire multi-view static data, construct global point cloud data of the multi-view static data, and sample a first-view video and a second-view video from the multi-view static data; Obtain target camera parameters of the first perspective video, perform point cloud rendering processing on the global point cloud data according to the target camera parameters to obtain the second point cloud rendering video, construct the first perspective video, the second perspective video and the second point cloud rendering video into a sample triplet, and add the sample triplet to the second sample training set.
14. A video generating device, characterized in that: include: A data acquisition module, used to acquire source video and camera trajectory parameters, and acquire video description text of the source video; A point cloud rendering module is used to convert the source video into point cloud data, perform point cloud rendering processing on the point cloud data according to the camera trajectory parameters to obtain a point cloud rendering result, and encode the point cloud rendering result to obtain a view encoding feature; An encoding module, used for encoding the source video to obtain reference encoding features, and encoding the video description text to obtain text encoding features; A first attention processing module, configured to convert the view encoding feature into a first self-attention feature, and convert the text encoding feature into a second self-attention feature; A second attention processing module, configured to perform cross attention processing on the first self-attention feature and the reference coding feature to obtain a video interaction feature, and obtain a multimodal coding feature according to the video interaction feature and the second self-attention feature; A decoding module is used to decode the multimodal coding features to generate a target video indicated by the camera trajectory parameters.
15. A computer device, characterized in that: including memory and processor; The memory is connected to the processor, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method according to any one of claims 1 to 13.
16. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1 to 13.
17. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 13 is implemented.
Citation Information
Cited By
Article propaganda video generation method and device, medium and program product
CN121218002A
Multi-view video compression method
CN121357323A
Video processing method and device, computer readable storage medium and computer program product
CN121509749A
Space-time consistency data generation method for visual target tracking
CN121527140A
A spatiotemporal consistency data generation method for visual target tracking
CN121527140B