Video generation model training method, video generation method, device and equipment

By adjusting the video generation model through decoding and reward feedback mechanisms, the problems of low camera control precision and inconsistency in 3D geometry were solved, achieving efficient and high-quality 3D scene reconstruction and improving training efficiency and camera control precision.

CN121531179APending Publication Date: 2026-02-13BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511666186.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing technologies suffer from low camera control precision, inconsistent 3D geometry, and high reward computational resource consumption. Traditional video decoders ignore the 3D geometric structure behind camera motion, resulting in poor 3D reconstruction quality and low training efficiency.

Method used

By acquiring latent variables and camera parameters from sample videos, 3D scene information is decoded, the video is rendered, and reward information is generated based on the differences. The parameters of the video generation model are adjusted to improve camera control accuracy, and a lightweight decoding process is introduced to reduce computational resource consumption.

Benefits of technology

It achieves efficient generation of high-quality 3D scene information, reduces computing resources and video memory consumption, improves camera control accuracy and consistency of 3D reconstruction, and avoids time-consuming scene-by-scene optimization process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121531179A_ABST
    Figure CN121531179A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multimedia, in particular to a video generation model training method, a video generation method, a video generation device and video generation equipment, and the video generation model training method comprises the steps: obtaining a first sample video latent variable and a first sample camera parameter; decoding the first sample video latent variable and the first sample camera parameter to obtain first sample three-dimensional scene information; rendering the first sample three-dimensional scene information to obtain a first rendered video; and according to the reward information generated by the difference between the first real video tag and the first rendered video, adjusting the initial video generation model to obtain a target video generation model, thereby realizing direct generation from the video latent variable to the three-dimensional scene information, effectively improving the training efficiency, and improving the user experience. And the reward information can be generated through the difference between the first rendered video and the first real video tag, so that the initial video generation model is stimulated to generate the first sample video latent variable which is more aligned with the first sample camera parameter, and the model processing precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of multimedia technology, and in particular to a video generation model training method, a video generation method, an apparatus, and a device. Background Technology

[0002] Camera-controlled video generation techniques in related technologies typically employ datasets containing paired videos and camera motion trajectories to fine-tune a pre-trained video diffusion model. This enables the model to learn to generate video sequences that conform to the given motion based on an initial image and camera motion trajectory. For applications requiring 3D, these techniques utilize a separate, time-consuming scene optimization process to reconstruct the 3D scene after generating multi-view videos.

[0003] However, while related technologies incorporate camera conditions, they often struggle to precisely follow the given camera motion trajectory, leading to deviations between the generated video and the expected camera motion. This deviation affects the quality of subsequent 3D reconstruction, resulting in inconsistent geometry and poor convergence. Furthermore, reward calculations in these technologies typically require decoding the latent variables generated by the model into RGB pixel video, a process that consumes significant computational resources and GPU memory, resulting in low training efficiency. In addition, traditional video decoders focus only on two-dimensional (2D) pixel information, ignoring the three-dimensional (3D) geometry underlying camera motion, which limits the effectiveness of 3D models in the task. Summary of the Invention

[0004] This disclosure provides a video generation model training method, a video generation method, an apparatus, and a device to at least solve problems in related technologies such as low camera control precision, inconsistent 3D geometry, high computational cost, and lack of 3D perception in existing reward feedback learning methods. The technical solution of this disclosure is as follows: According to a first aspect of the present disclosure, a video generation model training method is provided, comprising: Obtain the first sample video latent variable and the first sample camera parameter corresponding to the first sample video latent variable; the first sample video latent variable is obtained by video generation processing of the sample image and the initial sample camera parameter corresponding to the sample image based on the initial video generation model, and the first sample video latent variable is labeled with the first real video label. The first sample video latent variables and the first sample camera parameters are decoded to align the first sample video latent variables and the first sample camera parameters to obtain the first sample three-dimensional scene information; Render the three-dimensional scene information of the first sample to obtain a first rendered video that conforms to the camera parameters of the first sample; Reward information is generated based on the difference between the first real video tag and the first rendered video; The model parameters of the initial video generation model are adjusted according to the reward information so that the latent variables of the first sample video generated by the initial video generation model meet the preset conditions, thereby obtaining the target video generation model.

[0005] In an optional embodiment, decoding the first sample video latent variables and the first sample camera parameters to align the first sample video latent variables and the first sample camera parameters to obtain the first sample 3D scene information includes: Input the latent variables of the first sample video and the camera parameters of the first sample into the target decoder; Based on the target decoder, the first sample video latent variables and the first sample camera parameters are decoded to align the first sample video latent variables and the first sample camera parameters, thereby obtaining the first sample three-dimensional scene information; The target decoder is obtained by training a preset decoder based on the second real video label corresponding to the second sample video latent variable and the second rendered video; the second rendered video is obtained by rendering the second sample three-dimensional scene information, and the second sample three-dimensional scene information is obtained by decoding the second sample video latent variable and the second sample camera parameters corresponding to the second sample video latent variable. In an optional embodiment, the target decoder includes a first camera encoding module, a first video encoding module, and a first Transformer module. The step of decoding the first sample video latent variables and the first sample camera parameters based on the target decoder to align the first sample video latent variables and the first sample camera parameters to obtain the first sample 3D scene information includes: The camera features of the first sample camera are obtained by extracting camera features from the first sample camera parameters based on the first camera encoding module. Based on the first video encoding module, video features are extracted from the latent variables of the first sample video, and the extracted features are segmented to obtain the first sample video features. The first sample camera features and the first sample video features are spliced ​​together to obtain sample spliced ​​features; Attention processing is performed on the sample stitching features based on the first Transformer module to obtain the three-dimensional scene information of the first sample.

[0006] In an optional embodiment, the training process of the target decoder includes: Obtain the latent variables of the second sample video and the camera parameters of the second sample; The second sample video latent variable and the second sample camera parameter are input into the preset decoder for decoding, so as to align the second sample video latent variable and the second sample camera parameter to obtain the second sample three-dimensional scene information corresponding to the second sample video latent variable; Render the second sample's 3D scene information to obtain the second rendered video; Loss data is generated based on the difference between the second real video label and the second rendered video; The model parameters of the preset decoder are adjusted based on the loss data to obtain the target decoder.

[0007] In an optional embodiment, generating loss data based on the difference between the second real video tag and the second rendered video includes: Determine the first mean square error and the first learned perceptual image patch similarity between the second real video label and the second rendered video; The loss data is generated based on the first mean square error and the first learned perceptual image patch similarity.

[0008] In an optional embodiment, the first sample camera parameters include camera motion trajectory, the first sample 3D scene information includes first sample 3D scene information, the first sample 3D scene information includes position information, and the process of generating the position information includes: A depth map is obtained from the three-dimensional scene information of the first sample; the depth map includes the depth value of each pixel; The camera motion trajectory is converted into the origin and direction of the ray corresponding to each pixel; The location information is obtained by combining the depth value, the origin of the ray, and its direction.

[0009] In an optional embodiment, generating reward information based on the difference between the first real video tag and the first rendered video includes: Obtain a depth map from the three-dimensional scene information of the first sample; Based on the depth map and the intrinsic and extrinsic parameters of the camera corresponding to the first rendered video, determine the target video frame that is visible in the first real video tag among the candidate video frames included in the first rendered video. The difference between the target video frame and the target video frame label is determined to obtain the reward information; The target video frame tag is the video tag corresponding to the target video frame in the first real video tag.

[0010] In an optional embodiment, the depth map includes the depth value of each pixel in the candidate video frames, and the step of determining the target video frame visible in the first real video tag among the candidate video frames included in the first rendered video, based on the depth map and the intrinsic and extrinsic parameters of the camera corresponding to the first rendered video, includes: Based on the intrinsic and extrinsic parameters of the camera corresponding to the first rendered video and the depth value of each pixel, each pixel in the candidate video frame is projected onto world coordinates to obtain the world coordinates corresponding to each pixel. Based on the intrinsic and extrinsic parameters of the camera corresponding to the first rendered video and the depth value of each pixel, the world coordinates corresponding to each pixel are projected onto the first real video label to obtain the pixel coordinates of each pixel in the first real view label. Determine the pixel coordinates of each pixel in the first real video label, and the target depth in the first real video label; Determine the world coordinates of each pixel and its initial depth in the first real video label; If the difference between the initial depth and the target depth is less than a preset threshold, and the target depth is valid, the candidate video frame is determined to be the target video frame.

[0011] In an optional embodiment, determining the difference between the target video frame and the target video frame label to obtain the reward information includes: Determine the second mean square error and the second learned perceptual image patch similarity between the target video frame and the target video frame label; The reward information is generated based on the second mean square error and the second learned perceptual image patch similarity.

[0012] In an optional embodiment, the initial video generation model includes a second camera encoding module, a second video encoding module, and a second Transformer module, and the generation process of the first sample video latent variables includes: Obtain the sample image, the initial sample camera parameters, and the random noise latent variable; The sample image, the initial sample camera parameters, and the random noise latent variable are input into the initial video generation model. Based on the second camera encoding module, the camera features in the initial sample camera parameters are extracted to obtain the second sample camera features. Based on the second video encoding module, the sample image is filled, the filled image and the random noise latent variable are spliced ​​together, and the spliced ​​features are segmented to obtain the second sample video features. The second sample camera features and the second sample video features are fused to obtain fused features; Attention processing is performed on the fused features based on the second Transformer module to obtain the latent variables of the first sample video.

[0013] According to a second aspect of the present disclosure, a video generation method is provided, comprising: Obtain the image to be processed and the target camera parameters corresponding to the image to be processed; Input the image to be processed and the target camera parameters into the target video generation model to obtain the latent variables of the target video; The latent variables of the target video and the parameters of the target camera are decoded to align the image to be processed and the parameters of the target camera, thereby obtaining the target three-dimensional scene information; Render the target 3D scene information to obtain a target rendered video that conforms to the target camera parameters; The target video generation model is obtained based on the video generation model training method described in any of the above embodiments.

[0014] According to a third aspect of the present disclosure, a video generation model training apparatus is provided, the apparatus comprising: The first sample data acquisition module is configured to acquire the first sample video latent variable and the first sample camera parameter corresponding to the first sample video latent variable; the first sample video latent variable is obtained by video generation processing of the sample image and the initial sample camera parameter corresponding to the sample image based on the initial video generation model, and the first sample video latent variable is labeled with a first real video label. The first sample 3D scene information generation module is configured to decode the first sample video latent variables and the first sample camera parameters to align the first sample video latent variables and the first sample camera parameters to obtain the first sample 3D scene information. The first rendering video generation module is configured to render the three-dimensional scene information of the first sample to obtain a first rendering video that conforms to the camera parameters of the first sample. The reward information generation module is configured to generate reward information based on the difference between the first real video tag and the first rendered video. The target video generation model generation module is configured to adjust the model parameters of the initial video generation model according to the reward information, so that the latent variables of the first sample video generated by the initial video generation model meet the preset conditions, thereby obtaining the target video generation model.

[0015] In an optional embodiment, the first sample 3D scene information generation module is configured to perform: Input the latent variables of the first sample video and the camera parameters of the first sample into the target decoder; Based on the target decoder, the first sample video latent variables and the first sample camera parameters are decoded to align the first sample video latent variables and the first sample camera parameters, thereby obtaining the first sample three-dimensional scene information; The target decoder is obtained by training a preset decoder based on the second real video label corresponding to the second sample video latent variable and the second rendered video; the second rendered video is obtained by rendering the second sample three-dimensional scene information, and the second sample three-dimensional scene information is obtained by decoding the second sample video latent variable and the second sample camera parameters corresponding to the second sample video latent variable. In an optional embodiment, the target decoder includes a first camera encoding module, a first video encoding module, and a first Transformer module, and the first sample 3D scene information generation module is configured to perform: The camera features of the first sample camera are obtained by extracting camera features from the first sample camera parameters based on the first camera encoding module. Based on the first video encoding module, video features are extracted from the latent variables of the first sample video, and the extracted features are segmented to obtain the first sample video features. The first sample camera features and the first sample video features are spliced ​​together to obtain sample spliced ​​features; Attention processing is performed on the sample stitching features based on the first Transformer module to obtain the three-dimensional scene information of the first sample.

[0016] In an optional embodiment, the video generation model training apparatus further includes: The second sample data acquisition module is configured to acquire the second sample video latent variables and the second sample camera parameters. The second sample 3D scene information generation module is configured to input the second sample video latent variable and the second sample camera parameter into the preset decoder for decoding, so as to align the second sample video latent variable and the second sample camera parameter to obtain the second sample 3D scene information corresponding to the second sample video latent variable; The second rendering video generation module is configured to render the second sample 3D scene information to obtain the second rendering video. The loss data generation module is configured to generate loss data based on the difference between the second real video label and the second rendered video. The target decoder generation module is configured to adjust the model parameters of the preset decoder based on the loss data to obtain the target decoder.

[0017] In an optional embodiment, the loss data generation module is configured to perform: Determine the first mean square error and the first learned perceptual image patch similarity between the second real video label and the second rendered video; generate the loss data based on the first mean square error and the first learned perceptual image patch similarity.

[0018] In an optional embodiment, the first sample camera parameters include camera motion trajectory, the first sample 3D scene information includes position information, and the video generation model training device further includes: The depth map acquisition module is configured to acquire a depth map from the first sample 3D scene information; the depth map includes the depth value of each pixel; The origin and orientation determination module is configured to convert the camera motion trajectory into the origin and orientation of the ray corresponding to each pixel; The location information generation module is configured to combine the depth value, the origin and direction of the ray to obtain the location information.

[0019] In an optional embodiment, the reward information generation module is configured to perform: Obtain a depth map from the three-dimensional scene information of the first sample; Based on the depth map and the intrinsic and extrinsic parameters of the camera corresponding to the first rendered video, determine the target video frame that is visible in the first real video tag among the candidate video frames included in the first rendered video. The difference between the target video frame and the target video frame label is determined to obtain the reward information; The target video frame tag is the video tag corresponding to the target video frame in the first real video tag.

[0020] In an optional embodiment, the depth map includes a depth value for each pixel, and the reward information generation module is configured to perform: Based on the intrinsic and extrinsic parameters of the camera corresponding to the first rendered video and the depth value of each pixel, each pixel in the candidate video frame is projected onto world coordinates to obtain the world coordinates corresponding to each pixel. Based on the intrinsic and extrinsic parameters of the camera corresponding to the first rendered video and the depth value of each pixel, the world coordinates corresponding to each pixel are projected onto the first real video label to obtain the pixel coordinates of each pixel in the first real view label. Determine the pixel coordinates of each pixel in the first real video label, and the target depth in the first real video label; Determine the world coordinates of each pixel and its initial depth in the first real video label; If the difference between the initial depth and the target depth is less than a preset threshold, and the target depth is valid, the candidate video frame is determined to be the target video frame.

[0021] In an optional embodiment, the reward information generation module is configured to perform: Determine the second mean square error and the second learned perceptual image patch similarity between the target video frame and the target video frame label; The reward information is generated based on the second mean square error and the second learned perceptual image patch similarity.

[0022] In an optional embodiment, the initial video generation model includes a second camera encoding module, a second video encoding module, and a second Transformer module, and the video generation model training device further includes: The initial data acquisition module is configured to acquire the sample image, the initial sample camera parameters, and random noise latent variables; The second sample video feature generation module is configured to input the sample image, the initial sample camera parameters, and the random noise latent variable into the initial video generation model; extract camera features from the initial sample camera parameters based on the second camera encoding module to obtain second sample camera features; perform padding processing on the sample image based on the second video encoding module; concatenate the padded image and the random noise latent variable; and perform block processing on the concatenated features to obtain second sample video features. The fusion feature generation module is configured to perform fusion processing on the second sample camera features and the second sample video features to obtain fusion features; The first sample video latent variable generation module is configured to perform attention processing on the fused features based on the second Transformer module to obtain the first sample video latent variables.

[0023] According to a second aspect of the present disclosure, a video generation apparatus is provided, comprising: The data acquisition module is configured to acquire the image to be processed and the target camera parameters corresponding to the image to be processed. The target video latent variable generation module is configured to input the image to be processed and the target camera parameters into the target video generation model to obtain the target video latent variables; The target 3D scene information generation module is configured to decode the target video latent variables and the target camera parameters to align the image to be processed and the target camera parameters to obtain target 3D scene information; The target rendering video generation module is configured to render the target 3D scene information to obtain a target rendering video that conforms to the target camera parameters; The target video generation model is obtained based on the video generation model training method described above.

[0024] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the video generation model training method or video generation method as described above.

[0025] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein when instructions in the computer-readable storage medium are executed by a processor of a server, the server is enabled to perform the video generation model training method or the video generation method as described above.

[0026] According to a fifth aspect of the present disclosure, a computer program product is provided, the computer program product including a computer program stored in a readable storage medium, wherein at least one processor of a computer device reads from the readable storage medium and executes the computer program, causing the device to perform the video generation model training method or the video generation method described above.

[0027] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects: This disclosure first obtains the first sample video latent variables and the corresponding first sample camera parameters. The first sample video latent variables are obtained by performing video generation processing on sample images and the corresponding initial sample camera parameters based on the initial video generation model. The first sample video latent variables are labeled with a first real video tag. Next, the first sample video latent variables and the first sample camera parameters are decoded to align them, thereby obtaining the first sample 3D scene information, i.e., high-quality first sample 3D scene information is directly generated from the video latent variables. Then, the first sample 3D scene information is rendered to obtain a first rendered video that conforms to the first sample camera parameters. Next, reward information is generated based on the difference between the first real video tag and the first rendered video. Then, the model parameters of the initial video generation model are adjusted based on the reward information so that the first sample video latent variables generated by the initial video generation model meet preset conditions, thereby obtaining the target video generation model. As can be seen, this disclosure provides a method for directly generating high-quality first-sample 3D scene information from video latent variables using a feedforward approach. This achieves direct, feedforward generation of first-sample 3D scene information from video latent variables, avoiding the time-consuming scene-by-scene optimization process of related technologies, reducing the consumption of computing resources and GPU memory during the decoding process, and effectively improving training efficiency. Furthermore, this disclosure does not require decoding the latent variables into high-resolution RGB video to calculate the reward, but instead operates directly through a lightweight decoding process, greatly reducing GPU memory consumption and computation time. In addition, this disclosure introduces a novel reward feedback mechanism, namely, generating reward information based on the difference between the first rendered video obtained from the first-sample 3D scene information and the first real video label, adjusting the model parameters of the initial video generation model that generates the first-sample video latent variables, thereby "incentivizing" the initial video generation model to generate first-sample video latent variables that are more aligned with the first-sample camera parameters. Because such latent variables can obtain higher reward information (i.e., lower loss) after decoding, this significantly enhances the initial video generation model's ability to follow camera motion conditions and further improves camera control accuracy. Furthermore, since the reward information originates from a first sample of 3D scene information with inherent 3D consistency, the initial video generation model can generate more geometrically coherent and static content through reward feedback learning, effectively suppressing unnecessary dynamic elements or artifacts and enhancing the 3D consistency of the scene.

[0028] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0029] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0030] Figure 1 This is a schematic diagram illustrating an application environment of a video generation model training method or a video generation method according to an exemplary embodiment.

[0031] Figure 2 This is a flowchart illustrating a video generation model training method according to an exemplary embodiment. Figure 1 .

[0032] Figure 3 This is a schematic diagram illustrating the structure of an initial video generation model according to an exemplary embodiment of this disclosure.

[0033] Figure 4 This is a flowchart illustrating a video generation model training method according to an exemplary embodiment. Figure 2 .

[0034] Figure 5 This is a schematic diagram illustrating the structure of a preset decoder according to an exemplary embodiment of the present disclosure.

[0035] Figure 6 This is a schematic diagram illustrating a location information generation process according to an exemplary embodiment of this disclosure.

[0036] Figure 7 This is a schematic diagram illustrating a video generation method according to an exemplary embodiment of this disclosure.

[0037] Figure 8 This is a block diagram of a video generation model training device according to an exemplary embodiment.

[0038] Figure 9 This is a block diagram of a video generation apparatus according to an exemplary embodiment.

[0039] Figure 10 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar first objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0042] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0043] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an application environment of a video generation model training method or a video generation method according to an exemplary embodiment, such as... Figure 1 As shown, the application environment may include terminal 01 and server 02.

[0044] In an optional embodiment, server 02 can be used to generate a target video generation model and generate a video based on the target video generation model. Exemplarily, server 02 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0045] In an optional embodiment, terminal 01 can be used to display the generated video. Specifically, terminal 01 can be, but is not limited to, electronic devices such as smartphones, desktop computers, tablets, laptops, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices. Terminal 01 has a resource processing platform terminal installed. Optionally, the operating system running on the electronic device can be, but is not limited to, Android, iOS, Linux, and Windows.

[0046] In addition, it should be noted that, Figure 1The example shown is merely one application environment for the video generation model training or video generation method provided in this disclosure. Other application environments may exist in other scenarios, and this disclosure does not limit these.

[0047] Figure 2 This is a flowchart illustrating a video generation model training method according to an exemplary embodiment. Figure 1 ,like Figure 2 As shown, the execution target of this video generation model training or video generation method is a server, and the method includes at least the following steps: In step S11, the first sample video latent variable and the first sample camera parameter corresponding to the first sample video latent variable are obtained; the first sample video latent variable is obtained by performing video generation processing on the sample image and the initial sample camera parameter corresponding to the sample image based on the initial video generation model, and the first sample video latent variable is labeled with the first real video label.

[0048] In some embodiments, the server can construct an initial video generation model that accepts camera motion as a condition. This initial video generation model has the ability to generate a predicted denoised video latent variable based on the input image and camera parameters. Then, the server inputs the latent variable of the sample image to be processed and the corresponding initial sample camera parameters into the initial video generation model for video latent variable prediction processing to obtain a first sample video latent variable. For example, this initial video generation model can be an I2V model, where the I2V model is an image-to-video generation model.

[0049] Optionally, the sample image can be from various fields, such as game development, navigation, virtual reality, film and television production, etc.

[0050] Optionally, "initial sample camera parameters corresponding to the sample image" refers to the parameters of the camera corresponding to the sample image, where "camera corresponding to the sample image" refers to a pre-defined camera for the sample image, so as to generate a video sequence that conforms to the motion of the given camera based on the sample image.

[0051] Optionally, the initial sample camera parameters can be a representation of various forms of camera conditions. For example, the initial sample camera parameters may include: camera motion trajectory (including rotation, translation, etc.), point cloud rendering conditions, etc. The camera motion trajectory can be further parameterized into a format that can be processed by a neural network, such as camera pose, and even further, it can be Plücker embeddings.

[0052] Optionally, the latent variables of the sample image refer to abstracted, compressed hidden feature representations extracted from the sample image, used to capture the core semantic, structural, and contextual information of the image (rather than directly storing pixel values). They typically include the following key features: Semantic features: Encode high-level semantic information of an image, that is, "what the image is expressing". Specifically, this includes object category, attributes and states, scene type, etc.

[0053] Structural features: capturing the spatial relationships and layout of objects in an image, i.e., "where the objects are and how they are arranged." Specifically, this includes: location and accessibility, hierarchical structure, and set shape.

[0054] Style characteristics: Extracting the visual style and artistic features of an image, i.e., "what style the image looks like". Specifically, this includes: color distribution, texture details, artistic style, etc.

[0055] Context and global features: Encoding the global contextual information of an image, i.e., "implicit associations beyond the image". Specifically, this includes: temporal / sequence associations, cross-image associations, and attempted inferences.

[0056] Cross-modal alignment features capture the consistency of motion between frames, which can be used for video prediction or action recognition. These include semantic alignment, sentiment and intent, etc.

[0057] Optionally, the first sample video latent variable refers to a low-dimensional, highly abstract feature representation extracted from the sample images. It is used to capture spatiotemporal dynamic information, semantic content, and contextual relationships in the video. The core of the first sample video latent variable is to simultaneously model the complex relationship between the temporal dimension (inter-frame motion) and the spatial dimension (intra-frame objects), thereby supporting tasks such as video understanding, generation, and retrieval. The first sample video latent variable mainly includes: Spatiotemporal dynamic characteristics: The essence of video is a "sequence of moving images". The latent variables of the first sample video need to capture the joint information of time and space: motion trajectory, action pattern, optical flow and changes.

[0058] Dynamic semantic features: The latent variables of the first sample video need to be transformed from pixel-level motion information into high-level semantic understanding.

[0059] Global context features: The latent variables of the first sample video need to capture the global correlation over a long time, rather than just the information of local frames.

[0060] Compression and generation features: The latent variables of the first sample video are compressed representations of the video data, used to generate new videos.

[0061] Optionally, the first sample camera parameter corresponding to the first sample video latent variable can be the parameters of the camera corresponding to the first sample video latent variable, wherein "the camera corresponding to the first sample video latent variable" refers to a camera pre-given for the first sample video latent variable, so as to generate a video sequence that conforms to the motion of the given camera based on the first sample video latent variable.

[0062] Optionally, similar to the initial sample camera parameters, the first sample camera parameters can be a representation of various forms of camera conditions. For example, the first sample camera parameters may include: camera motion trajectory (including rotation, translation, etc.), point cloud rendering conditions, etc. The camera motion trajectory can be further parameterized into a format that can be processed by a neural network, such as camera pose, and even further, it can be Plücker embeddings.

[0063] Optionally, the first real video label is a real video sequence generated based on the latent variables of the first sample video.

[0064] In step S12, the latent variables of the first sample video and the parameters of the first sample camera are decoded to align the latent variables of the first sample video and the parameters of the first sample camera, thereby obtaining the three-dimensional scene information of the first sample.

[0065] In some embodiments, the server can use various methods to decode the first sample video latent variables and the first sample camera parameters to align the first sample video latent variables and the first sample camera parameters to obtain the first sample 3D scene information. For example, it can be decoded by a decoder, or a model can be pre-trained and decoded using the model.

[0066] Optionally, the first sample 3D scene information may include, but is not limited to, 3D Gaussian parameters (3DGS) and other explicit or implicit 3D representations. For example, when the first sample camera parameters are camera motion trajectories, the 3D Gaussian parameters may include, but are not limited to, position, covariance matrix (rotation, scaling), spherical harmonic coefficients, opacity, etc.

[0067] Here, "position" refers to a 3D vector representing the coordinates of Gauss's center point in world space, which defines Gauss's position in space.

[0068] The covariance matrix is ​​a 3×3 symmetric matrix that controls the shape, size, and orientation of the Gaussian sphere. It describes the distribution of the Gaussian sphere in space (i.e., how it is stretched and rotated). For optimization convenience, the covariance matrix is ​​decomposed into scaling and rotation. Scaling refers to a 3D vector representing the scaling factor of the Gaussian sphere along each axis, which determines its size and degree of anisotropy. Rotation refers to a quaternion that defines the direction of rotation of the Gaussian sphere. The quaternion can be converted into a rotation matrix for orienting the Gaussian sphere.

[0069] Opacity refers to a scalar value that controls the transparency of a Gaussian.

[0070] Spherical harmonics refer to a set of coefficients used to represent view-dependent color (appearance). Spherical harmonic functions (for example, using 4 bands, i.e., 16 coefficients) encode directional lighting effects, allowing Gaussians to display different colors at different viewpoints.

[0071] For example, other explicit or implicit 3D representations may include, but are not limited to, variations of meshes, voxels, or neural radiation fields (NeRF).

[0072] In other embodiments, for the generation and reconstruction of dynamic scenes, the server can also extend the output of the 3D decoder from 3DGS to 4DGS (spacetime Gaussian sputtering) to enable it to model the dynamic changes of the scene.

[0073] In step S13, the first sample 3D scene information is rendered to obtain the first rendered video that conforms to the first sample camera parameters.

[0074] In some embodiments, the server can render a video sequence from the first sample 3D scene information to obtain a first rendered video of the first sample camera parameters.

[0075] In step S14, reward information is generated based on the difference between the first real video label and the first rendered video.

[0076] In step S15, the model parameters of the initial video generation model are adjusted according to the reward information so that the latent variables of the first sample video generated by the initial video generation model meet the preset conditions, thereby obtaining the target video generation model.

[0077] In some embodiments, the server can calculate the difference between the first real video label and the first rendered video to generate reward information, and adjust the model parameters of the initial video generation model according to the reward information through reward feedback learning (ReFL) so that the latent variables of the first sample video generated by the initial video generation model meet the preset conditions, thereby obtaining the target video generation model.

[0078] Optionally, the phrase "the first sample video latent variable generated by the initial video generation model satisfies the preset condition" can refer to the following: the alignment between the first sample video latent variable generated by the initial video generation model and the first sample camera parameters is greater than a threshold. This allows reward information to be generated based on the difference between the first rendered video (obtained by rendering the first sample 3D scene information obtained from decoding the first sample video latent variable) and the first real video label. The model parameters of the initial video generation model that generates the first sample video latent variable are then adjusted based on this reward information to "incentivize" the initial video generation model to generate first sample video latent variables that are more aligned with the first sample camera parameters. Since such latent variables, after decoding, yield higher reward information (i.e., lower loss), this significantly enhances the initial video generation model's ability to follow camera motion conditions and further improves camera control accuracy.

[0079] Optionally, the difference between the first real video label and the first rendered video may include, but is not limited to, mean squared error loss (MSE loss), learning-aware patch similarity loss (LPIPS loss), etc.

[0080] Therefore, this disclosure provides a method for directly generating high-quality first-sample 3D scene information from video latent variables using a feedforward approach. This achieves direct, feedforward generation of first-sample 3D scene information from video latent variables, avoiding the time-consuming scene-by-scene optimization process of related technologies, reducing the consumption of computing resources and GPU memory during the decoding process, and effectively improving training efficiency. Furthermore, this disclosure does not require decoding latent variables into high-resolution RGB video to calculate rewards, but instead operates directly through a lightweight decoding process, greatly reducing GPU memory consumption and computation time. In addition, this disclosure introduces a novel reward feedback mechanism, namely, generating reward information based on the difference between the first rendered video obtained from the first-sample 3D scene information and the first real video label, adjusting the model parameters of the initial video generation model that generates the first-sample video latent variables, thereby "incentivizing" the initial video generation model to generate first-sample video latent variables that are more aligned with the first-sample camera parameters. Because such latent variables can obtain higher reward information (i.e., lower loss) after decoding, this significantly enhances the initial video generation model's ability to follow camera motion conditions and further improves camera control accuracy. Furthermore, since the reward information originates from a first sample of 3D scene information with inherent 3D consistency, the initial video generation model can generate more geometrically coherent and static content through reward feedback learning, effectively suppressing unnecessary dynamic elements or artifacts and enhancing the 3D consistency of the scene.

[0081] The following section explains the training of the initial video generation model and the process of generating the first sample video latent variables based on the trained initial video generation model.

[0082] First, this disclosure pre-constructs an initial video generation model (e.g., an I2V diffusion model) that is capable of accepting camera motion as a condition. Figure 3 This is a schematic diagram illustrating the structure of an initial video generation model according to an exemplary embodiment of this disclosure, such as... Figure 3 As shown, this initial video generation model, based on a pre-trained I2V model, injects camera conditions through a ControlNet architecture, which is an extension of the diffusion model. Specifically, the camera embedding sequence is processed by a trainable camera encoder, and its output is added to the input of the backbone Transformer module of the I2V model. Optionally, this initial video generation model includes a second camera encoding module, a second video encoding module, and a second Transformer module. The second video encoding model further includes a zero-padding module and a patchify module. Figure 3 In Indicates splicing, It indicates fusion.

[0083] During training, the latent variables of an image, camera parameters (e.g., camera motion trajectory, represented as a Plück coordinate embedding sequence), and a random noise latent variable are input into the initial video generation model. Random noise refers to random fluctuations in the data that cannot be explained by the model; it typically originates from measurement errors, unobserved minor perturbations, or random factors outside the system, and is inherently unpredictable. The random noise latent variable is a special type of latent variable that is inherently random and is primarily used to model random noise components in the data that are not explained by other structured latent variables.

[0084] For the camera branch, the server extracts camera features from the camera parameters based on the second camera encoding module and outputs camera features (Camera Tokens). These Camera Tokens are discretized representations of camera parameters / constraints and are used to guide the viewpoint and motion logic of subsequent video generation.

[0085] For the video branch, the server performs zero-padding on the latent variables of the image based on the second video encoding module, resulting in a padded image. This is to adjust the dimensions / sequence length to fit subsequent block operations. Next, the server concatenates the padded image with random noise latent variables and divides the concatenated features into video features (VisualTokens). Each block corresponds to a local spatiotemporal unit of the video, serving as a carrier of fine-grained features of the video content. Then, the server fuses the camera features and video features to obtain fused features (the final token sequence). This fusion method ensures the decoupling and coordination between video content and camera control; the visual branch guarantees the richness of the generated video content, while the camera branch ensures the controllability of viewpoint and motion.

[0086] Finally, the server performs attention processing on the final token sequence based on the second Transformer module to obtain the predicted video latent variables; then, it calculates the difference between the predicted video latent variables and the video latent variable labels, determines the loss data based on the difference, and adjusts the model parameters of the initial video generation model based on the loss data to obtain the trained initial video generation model.

[0087] In some embodiments, continue as follows Figure 3 As shown, this disclosure allows the addition of a trainable Transformer module to the pre-trained DiT structure. During training, the trainable Transformer module can be fine-tuned or expanded based on the loss data to adapt to a specific task without changing the pre-trained weights. This method of freezing the pre-trained model and adding a trainable module can fine-tune the added parts for downstream tasks, saving computational resources and training time. Furthermore, this design combines the powerful feature extraction capabilities of the pre-trained model with the flexibility of the trainable module.

[0088] In some embodiments, the second Transformer module may further include: Multi-head self-attention mechanism: Through self-attention, each token at a given position can be directly associated with tokens at any other position in the sequence, thus efficiently modeling the global context. Furthermore, "multi-head" decomposes self-attention into multiple independent "heads," each learning different attention patterns (such as syntactic dependencies, semantic associations, positional relationships, etc.). Finally, the outputs of all heads are concatenated and linearly transformed to obtain a richer feature representation. In addition, attention weights are calculated by the dot product of query, key, and value, implicitly encoding relative or absolute positional information in the sequence (in conjunction with subsequent positional encoding). The process of multi-head self-attention mechanism includes: for each head, the input fused features are projected into query, key, and value through three linear layers respectively; attention weights are calculated based on clicks on the query and key; and the value is weighted and summed according to these attention weights to obtain the attention calculation results for each head. Finally, the attention calculation results of all heads are concatenated to obtain the latent variables of the first sample video.

[0089] Feedforward neural networks: The output of self-attention is further processed non-linearly (usually two fully connected layers + ReLU / GELU activation function) to transform the "global correlation information" extracted by attention into higher-level semantic features.

[0090] Residual connections and normalization layers: The output of each sub-layer (self-attention, FFN) is added to the input. The purpose is to preserve the low-level features of the original input, alleviate the degradation problem of deep networks, and accelerate training. Layer normalization (normalizing the mean and variance of the feature dimensions of each sample) constrains the distribution of features, prevents activation values ​​from being too large or too small, and improves training efficiency and model generalization ability.

[0091] In some embodiments, taking the initial video generation model as a diffusion model as an example, the core of the diffusion model is to gradually "de-noise" pure noise to generate target samples (such as 3D scenes and images). The time step is the "progress indicator" of the denoising process. The closer to pure noise, the earlier the denoising stage (requiring the establishment of a basic structure from chaos); the closer to a clean sample, the later the denoising stage (only requiring optimization of details). The essence of 3D generation is "reconstructing structure + perspective from noise," while camera control directly determines the "angle of observation of the world"—if the camera pose / motion is not correctly modeled in the early stages, the subsequently generated geometry and textures will "grow" based on an incorrect perspective, ultimately leading to problems such as inconsistent perspectives, structural distortion, and incorrect object occlusion. That is, early camera control is "determining the direction," while later details are "filling in the content"—if the direction is wrong, the content will also be incorrect. Therefore, to allow the model to focus on learning the early (large t) camera motion modeling ability and improve the effectiveness of camera control, instead of conventional uniform sampling (covering all t), a truncated normal distribution sampling time step is used during training: Normal distribution: The mean is set to a large t value to ensure that the sampling center is biased towards the "early stage".

[0092] Truncation: Limit the range of values ​​for t to completely exclude sampling of smaller t values ​​(late steps).

[0093] The effect of this is that the model encounters "large values" of t almost exclusively during training, meaning the model processes the input from the "early denoising stage" more frequently, and gradient updates prioritize optimizing the parameters of this step. The ultimate goal is to teach the model to "determine the camera early and fill in details later." By truncating normal sampling, the model will specifically optimize its "early camera motion modeling ability" during training.

[0094] After the above training, the initial video generation model can generate a predicted denoised video latent variable based on the input image and camera motion trajectory. Accordingly, the generation process of the first sample video latent variable includes: Acquire sample images, initial sample camera parameters, and random noise latent variables.

[0095] The sample image, initial sample camera parameters, and random noise latent variables are input into the initial video generation model. The camera features in the initial sample camera parameters are extracted based on the second camera encoding module to obtain the second sample camera features. The sample image is filled based on the second video encoding module. The filled image and random noise latent variables are spliced ​​together, and the spliced ​​features are segmented to obtain the second sample video features.

[0096] The camera features of the second sample and the video features of the second sample are fused to obtain the fused features.

[0097] Attention processing is applied to the fused features using the second Transformer module to obtain the latent variables of the first sample video.

[0098] In this embodiment, during the process of generating the first sample video latent variable using the trained initial video model, the sample image, initial sample camera parameters, and random noise latent variable can be input into the initial video generation model.

[0099] For the camera branch, the server extracts camera features from the initial sample camera parameters based on the second camera encoding module and outputs second sample camera features. These second sample camera features are discretized representations of the initial sample camera parameters / constraints and are used to guide the viewpoint and motion logic of subsequent video generation.

[0100] For the video branch, the server performs zero-padding on the latent variables of the image based on the second video encoding module, obtaining a padded image. This is to adjust the dimensions / sequence length to fit subsequent block operations. Next, the server concatenates the padded image with random noise latent variables and divides the concatenated features into blocks for second-sample video features. Each block corresponds to a local spatiotemporal unit of the video, serving as a carrier of fine-grained features of the video content. Then, the server fuses the second-sample camera features and the second-sample video features to obtain fused features. This fusion method ensures the decoupling and coordination between video content and camera control. The visual branch guarantees the richness of the generated video content, while the camera branch ensures the controllability of viewpoint and motion.

[0101] Finally, the server performs attention processing on the fused features based on the Transformer module. Through self-attention, the token at each position can be directly associated with the token at any other position in the sequence, thereby efficiently modeling the global context and obtaining the latent variables of the first sample video.

[0102] It should be noted that when the initial sample camera parameters are rendered as point cloud conditions, the initial sample camera parameters are not directly encoded as input to the initial video model. Instead, some incomplete images are rendered on the point cloud, and the initial video model completes them.

[0103] Because the camera branch accurately extracts camera features from the initial sample camera parameters, which are then used to guide the viewpoint and motion logic of subsequent video generation. The video branch accurately extracts fine-grained features of the video content. By fusing the second sample camera features and the second sample video features, the decoupling and coordination between video content and camera control are ensured. The visual branch guarantees the richness of the generated video content, while the camera branch ensures the controllability of viewpoint and motion. Attention processing of the fused features allows each token to be directly associated with any other token in the sequence, thus efficiently modeling the global context and improving the accuracy and efficiency of extracting latent variables from the first sample video.

[0104] Figure 4 This is a flowchart illustrating a video generation model training method according to an exemplary embodiment. Figure 2 ,like Figure 4 As shown, in step S12 above, decoding the first sample video latent variables and the first sample camera parameters to align the first sample video latent variables and the first sample camera parameters to obtain the first sample 3D scene information includes: In step S121, the first sample video latent variables and the first sample camera parameters are input into the target decoder.

[0105] In step S122, the first sample video latent variables and the first sample camera parameters are decoded based on the target decoder to align the first sample video latent variables and the first sample camera parameters, thereby obtaining the first sample 3D scene information; wherein, the target decoder is obtained by training a preset decoder based on the second real video label corresponding to the second sample video latent variables and the second rendered video; the second rendered video is obtained by rendering the second sample 3D scene information, and the second sample 3D scene information is obtained by decoding the second sample video latent variables and the second sample camera parameters corresponding to the second sample video latent variables.

[0106] In this embodiment, a preset decoder can be trained in advance using the second real video label corresponding to the second sample video latent variable and the second rendered video to obtain the target decoder. The second rendered video is obtained by rendering the second sample 3D scene information, and the second sample 3D scene information is obtained by decoding the second sample video latent variable and the second sample camera parameters corresponding to the second sample video latent variable. Optionally, the preset decoder and the target decoder can be 3D decoders.

[0107] Optionally, second-sample video latent variables refer to low-dimensional, highly abstract feature representations extracted from certain sample images. They are used to capture spatiotemporal dynamic information, semantic content, and contextual relationships in the video. Unlike image latent variables (which only process single-frame static information), the core of second-sample video latent variables is to simultaneously model the complex relationships between the temporal dimension (inter-frame motion) and the spatial dimension (intra-frame objects), thereby supporting tasks such as video understanding, generation, and retrieval. Second-sample video latent variables mainly include: spatiotemporal dynamic features, dynamic semantic features, global contextual features, and compression and generation features.

[0108] Optionally, the second true label refers to the true video sequence generated based on the latent variables of the second sample video.

[0109] Optionally, the second sample camera parameters corresponding to the second sample video latent variable can be the parameters of the camera corresponding to the second sample video latent variable, wherein "the camera corresponding to the second sample video latent variable" refers to a camera pre-given for the second sample video latent variable, so as to generate a video sequence that conforms to the motion of the given camera based on the second sample video latent variable.

[0110] For example, the second sample camera parameters refer to, similar to the initial sample camera parameters, representations of various forms of camera conditions. For example, these second sample camera parameters may include: camera motion trajectories (including rotation, translation, etc.), point cloud-based rendering conditions, etc. The camera motion trajectory can be further parameterized into a format suitable for neural network processing, such as camera pose, and even further, Plücker embeddings.

[0111] It should be noted that the first sample video latent variable is the sample data used to train the initial video generation model, and the second sample video latent variable is the sample data used to train the target decoder. The first sample video latent variable and the second sample video latent variable can be the same, different, or partially overlap, and there are no specific restrictions on this.

[0112] In this embodiment, after the target decoder is trained, the server can directly input the first sample video latent variables and the first sample camera parameters into the target decoder for decoding. The target decoder aligns the first sample video latent variables and the first sample camera parameters and directly outputs the first sample 3D scene information.

[0113] It should be noted that when the first sample camera parameters are used as point cloud rendering conditions, the first sample camera parameters will not be directly encoded as input to the target decoder. Instead, some incomplete images will be rendered on the point cloud, and the target decoder will fill in the missing parts.

[0114] Because the preset decoder can directly decode the second sample video latent variables and the corresponding second sample camera parameters during training to obtain the second sample 3D scene information, that is, directly decode the second sample video latent variables into a 3D scene representation, and its decoding quality is highly sensitive to the alignment between the input video latent variables and the camera pose, the trained target decoder has a method to directly generate high-quality sample 3D scene information from video latent variables. It realizes the direct, feedforward generation of sample 3D scene information from video latent variables, avoids the scene-by-scene optimization process that takes several minutes to several hours in related technologies, reduces the consumption of computing resources and GPU memory in the decoding process, and effectively improves the generation efficiency of the first sample 3D scene information.

[0115] In some embodiments, the training process of the target decoder described above is as follows: Obtain the latent variables of the second sample video and the camera parameters of the second sample.

[0116] The second sample video latent variables and the second sample camera parameters are input into a preset decoder for decoding, so as to align the second sample video latent variables and the second sample camera parameters, and obtain the second sample three-dimensional scene information corresponding to the second sample video latent variables.

[0117] Render the 3D scene information of the second sample to obtain the second rendered video.

[0118] Loss data is generated based on the difference between the second real video label and the second rendered video.

[0119] The target decoder is obtained by adjusting the model parameters of the preset decoder based on the loss data.

[0120] Figure 5 This is a schematic diagram illustrating the structure of a preset decoder according to an exemplary embodiment of the present disclosure, such as... Figure 5 As shown, the preset decoder includes a first camera encoding module, a first video encoding module, a first Transformer module, and a 3D deconvolution module (Deconv 3D). Among them, Figure 5 In Indicates splicing.

[0121] Optionally, the first camera encoding module is used to extract camera-related features from the input second sample camera parameters and output camera features. Optionally, the first video encoding module includes an encoding module and a segmentation module: the encoding module extracts features from the input first sample video latent variable sequence, and the segmentation module segments the extracted video features to obtain visual features. Optionally, the first Transformer module performs attention processing on the stitched camera features and visual features to obtain initial second sample 3D scene information. Optionally, the 3D deconvolution module is an upsampling operation for 3D data, mainly used to recover high-resolution spatial details from low-resolution, initial second sample 3D scene information to obtain the final second sample 3D scene information.

[0122] In some embodiments, the first Transformer module may further include a multi-head self-attention mechanism. Optionally, the process of the multi-head self-attention mechanism includes: for each head, projecting the input concatenated features through three linear layers as Query, Key, and Value respectively; calculating attention weights based on clicks on Query and Key; and weighted summing of Values ​​according to these attention weights to obtain the attention calculation results for each head. Finally, concatenating the attention calculation results for each head yields the initial second sample 3D scene information.

[0123] Feedforward neural networks: The output of self-attention is further processed non-linearly (usually two fully connected layers + ReLU / GELU activation function) to transform the "global correlation information" extracted by attention into higher-level semantic features.

[0124] Residual connections and normalization layers: The output of each sub-layer (self-attention, FFN) is added to the input. The purpose is to preserve the low-level features of the original input, alleviate the degradation problem of deep networks, and accelerate training. Layer normalization (normalizing the mean and variance of the feature dimensions of each sample) constrains the distribution of features, prevents activation values ​​from being too large or too small, and improves training efficiency and model generalization ability.

[0125] Optionally, after the server decodes the second sample 3D scene information based on the preset decoder, the second sample 3D scene information can be rendered to obtain a second rendered video. Then, the difference between the second rendered video and the second real video label is calculated to obtain loss data. Finally, the model parameters of the preset decoder are adjusted according to the loss data until the loss data meets the preset conditions or the number of iterations reaches the preset number, and the trained target decoder is obtained.

[0126] Because the pre-defined decoder can directly decode the second sample video latent variables and the corresponding second sample camera parameters during training to obtain the second sample 3D scene information, that is, directly decode the second sample video latent variables into a 3D scene representation, and its decoding quality is highly sensitive to the alignment between the input video latent variables and camera parameters, the trained target decoder has a method to directly generate high-quality sample 3D scene information from video latent variables. This achieves direct, feedforward generation from video latent variables to sample 3D scene information, avoiding the time-consuming scene-by-scene optimization process in related technologies, reducing the consumption of computing resources and GPU memory during the decoding process, and effectively improving the generation efficiency of second sample 3D scene information. Furthermore, it generates loss data based on the difference between the second real video label and the second rendered video, so that if the motion contained in the input second sample video latent variables does not match the input second sample camera parameters, the 3D scene information decoded and reconstructed by the pre-defined decoder will become inconsistent, resulting in a blurry and degraded final rendered video. Thus, by continuously adjusting the model parameters of the pre-defined decoder, it can effectively align the second sample video latent variables and the second sample camera parameters, improving the training accuracy and decoding accuracy of the target decoder.

[0127] In an optional embodiment, generating loss data based on the difference between the second real video label and the second rendered video may include: Determine the first mean square error and the first learned perceptual image patch similarity between the second real video label and the second rendered video.

[0128] Loss data is generated based on the first mean square error and the first learned perceptual image patch similarity.

[0129] In this embodiment, the loss function can be determined based on the pixel-level and / or perceptual-level differences between the rendered second video and the second real video label. Optionally, the pixel-level difference can be represented by a first mean squared error, and the perceptual difference can be represented by a first learned perceptual patch similarity (LPIPS similarity).

[0130] Optionally, the calculation process for the first mean square error (MSE) can be as follows: ; in, For the second rendered video, The second real video is tagged with N, where N is the number of videos.

[0131] Optionally, the calculation process for the first learned perceptual image patch similarity can be as follows: input the second real video label and the second rendered video into the pre-trained network, extract activation features from multiple layers, normalize the features of each layer to eliminate scale effects, calculate the cosine similarity of the features of each layer, and sum them layer by layer to obtain the final LPIPS similarity.

[0132] In one approach, the server can directly use the first mean squared error or the first learned perceptual image patch similarity as the loss data. In another approach, the server can use the sum of the first mean squared error and the first learned perceptual image patch similarity as the loss data. In a third approach, the process continues as follows... Figure 5 As shown, the server can assign a weight to the similarity of the first learned perceptual image patch based on its contribution to the loss data. The loss data is calculated using the following formula: L= + ; Where L refers to the loss data. This refers to the first mean square error. This refers to the similarity of the first learned perceptual image patches.

[0133] Therefore, by calculating the loss data through the pixel-level and perceptual-level differences between the second rendered video and the second real video label, the calculated loss data can contain both pixel-level and perceptual differences. This allows for the adjustment of the model parameters of the preset decoder from both pixel-level and perceptual-level dimensions to align the second real video label and the second rendered video. This minimizes the difference between the second rendered video obtained based on the 3D scene information of the second sample and the real video, thereby improving the training accuracy of the target decoder. The trained target decoder can accurately decode the input video latent variables and camera parameters to obtain the corresponding 3D scene information.

[0134] In some embodiments, in step S122 above, the decoding of the first sample video latent variables and the first sample camera parameters based on the target decoder to align the first sample video latent variables and the first sample camera parameters to obtain the first sample 3D scene information may include: The camera features of the first sample camera are obtained by extracting camera features from the parameters of the first sample camera based on the first camera encoding module.

[0135] Based on the first video encoding module, video features are extracted from the latent variables of the first sample video, and the extracted features are segmented to obtain the first sample video features.

[0136] The camera features and video features of the first sample are spliced ​​together to obtain the spliced ​​sample features.

[0137] Attention processing is applied to the sample stitching features based on the first Transformer module to obtain the three-dimensional scene information of the first sample.

[0138] In this embodiment, since the target encoder is obtained by training a preset encoder, the target encoder also includes a first camera encoding module, a first video encoding module, a first Transformer module, and Deconv 3D.

[0139] Optionally, after inputting the first sample video latent variables and the first sample camera parameters into the target decoder, camera-related features in the input second sample camera parameters can be extracted based on the first camera encoding module in the target decoder, outputting the first sample camera features. Features are then extracted from the input first sample video latent variable sequence based on the encoding module, and the extracted video features are segmented using a segmentation module to obtain the first sample video features. Next, the first sample camera features and the first sample video features are concatenated to obtain sample concatenated features. Then, attention processing is applied to the sample concatenated features based on the first Transformer module to obtain initial first sample 3D scene information. Finally, high-resolution spatial details are recovered from the low-resolution initial first sample 3D scene information using a 3D deconvolution module to obtain the final first sample 3D scene information.

[0140] In some embodiments, the first Transformer module may further include: Multi-head self-attention mechanism: The process of the multi-head self-attention mechanism includes: For each head, the input sample concatenated features are projected into Query, Key, and Value through three linear layers respectively. Attention weights are calculated based on clicks on Query and Key. The Value is then weighted and summed according to these attention weights to obtain the attention calculation results for each head. Finally, the attention calculation results of each head are concatenated to obtain the first sample's 3D scene information.

[0141] Feedforward neural networks: further nonlinear processing (usually two fully connected layers + ReLU / GELU activation function) is performed on the output of self-attention to transform the "global correlation information" extracted by attention into higher-level semantic features.

[0142] Residual connections and normalization layers: The output of each sub-layer (self-attention, FFN) is added to the input. The purpose is to preserve the low-level features of the original input, alleviate the degradation problem of deep networks, and accelerate training. Layer normalization (normalizing the mean and variance of the feature dimensions of each sample) constrains the distribution of features, prevents activation values ​​from being too large or too small, and improves training efficiency and model generalization ability.

[0143] As can be seen, this disclosure provides a method for directly generating high-quality first-sample 3D scene information from the latent variables of the first-sample video using a feedforward approach. This achieves direct, feedforward generation of first-sample 3D scene information from video latent variables, avoiding the scene-by-scene optimization process that takes several minutes to several hours in related technologies. It also reduces the consumption of computing resources and GPU memory during the decoding process, effectively improving training efficiency. Furthermore, this disclosure does not require decoding the latent variables into high-resolution RGB video to calculate the reward, but instead directly operates through the decoding process of a lightweight target decoder, greatly reducing GPU memory consumption and computation time.

[0144] In some embodiments, as shown above, when the first sample camera parameters are the camera motion trajectory, the 3DGS may include, but is not limited to: position, covariance matrix (rotation, scaling), spherical harmonic coefficients, opacity, etc. The covariance matrix (rotation, scaling), spherical harmonic coefficients, and opacity are directly output by the Deconv 3D of the target decoder, but the position information is not directly output by Deconv 3D, but is obtained through depth + projection.

[0145] Figure 6 This is a schematic diagram illustrating a location information generation process according to an exemplary embodiment of this disclosure, such as... Figure 6 As shown, in an optional embodiment, the process of generating the above location information may include: In step S21, a depth map is obtained from the first sample 3D scene information; the depth map includes the depth value of each pixel.

[0146] Optionally, the depth map refers to a 3D representation of the first sample's 3D scene information itself, from which the server can obtain the depth map, which includes the depth value of each pixel.

[0147] The depth value of each pixel refers to the distance from the 3D point corresponding to the pixel to the optical center of the camera along the ray direction.

[0148] In step S22, the camera motion trajectory is converted into the origin and direction of the ray corresponding to each pixel.

[0149] Optionally, the camera motion trajectory is a continuous or discrete sequence of camera pose in the time dimension. The camera pose is a mathematical representation of the camera's position (origin) and orientation (direction) in the three-dimensional world. Each pixel captured by the camera corresponds to a 3D ray that starts from the camera's optical center (origin) and extends along a specific direction.

[0150] Optionally, the method of converting the camera motion trajectory to the origin and direction of the ray corresponding to each pixel may include: Camera pose is typically represented by extrinsic parameters (rotation matrix and translation vector). Transforming the camera coordinate system to the world coordinate system involves: Origin of the ray (world coordinate system): The coordinates of the camera's optical center in the world coordinate system are... (Since the translation vector is the coordinate of the origin of the world coordinate system in the camera coordinate system, it needs to be inverted and rotated to transform back to world coordinates.) Here, O refers to the camera optical center (i.e., the origin of the ray), R refers to the rotation matrix, and t refers to the average vector.

[0151] The direction of the ray (world coordinate system): for a pixel in the image ( u , v First, the camera intrinsic parameter matrix K is transformed into a normalized direction vector (x,y,1) in the camera coordinate system through the inverse transformation. Then through the rotation matrix. R Transform it to the world coordinate system to obtain the direction vector of the ray. (Usually normalized to a unit vector).

[0152] In step S23, the depth value, the origin of the ray, and the direction are combined to obtain the position information.

[0153] In this embodiment, the position information can be obtained by combining the depth value, the origin of the ray, and the direction using the following formula: ; Where P refers to the position information, d refers to the direction of the ray, and depth refers to the depth value.

[0154] It is evident that camera parameters (e.g., camera motion trajectory) play a dual role in the decoder. Firstly, they serve as input conditions: the camera embedding, as part of the model's input, influences the model's prediction of 3D parameters. Secondly, they act as projection parameters: the 3D spatial location information of the 3D Gaussian points is calculated by combining the model's predicted depth values ​​with the origin and direction of the rays corresponding to each pixel, thus improving the accuracy of location determination and consequently enhancing the accuracy of generating the first sample of 3D scene information.

[0155] In some embodiments, in step S14 above, the server can calculate the reward information based on the pixel-level difference and / or perceptual-level difference between the first real video tag and the first rendered video. The pixel-level difference can be represented by the mean squared error, and the perceptual difference can be represented by the learned perceptual image patch similarity. Exemplarily, the server can calculate the sum of the mean squared error and the learned perceptual image patch similarity to obtain the reward information. Alternatively, the mean squared error can be calculated separately to obtain the reward information, and the learned perceptual image patch similarity can be calculated separately to obtain the reward information; no specific limitation is made in this regard.

[0156] In other embodiments, in step S14 above, generating reward information based on the difference between the first real video tag and the first rendered video includes: Depth maps are obtained from the three-dimensional scene information of the first sample.

[0157] Based on the depth map and the intrinsic and extrinsic parameters of the camera corresponding to the first rendered video, the target video frames visible in the first real video label are determined among the candidate video frames included in the first rendered video.

[0158] The difference between the target video frame and the target video frame label is determined to obtain the reward information.

[0159] Among them, the target video frame label is the video label corresponding to the target video frame in the first real video label.

[0160] In some embodiments, considering the randomness of video generation, newly generated content (such as new areas seen after camera rotation) may not have a corresponding counterpart in the real video. Directly calculating the difference might penalize the model's creativity. Therefore, this disclosure introduces a visibility mask. For each candidate video frame of the first rendered video, the server can obtain a depth map from the first sample 3D scene information. Combined with the intrinsic and extrinsic parameters of the camera corresponding to the first rendered video, and through geometric transformation (back projection followed by forward projection), determine the target video frames visible in the first real video label among the candidate video frames included in the first rendered video, thereby calculating whether each pixel is visible in the first real video. The final reward information is only calculated in the deterministic pixel regions marked as visible by the mask (i.e., the target video frames). This allows the video generation model to accurately align camera movement in areas with existing content while also allowing reasonable content generation in unknown areas, improving the rationality and accuracy of the video generation model training.

[0161] Optionally, the depth map refers to a 3D representation of the first sample's 3D scene information itself, from which the server can obtain the depth map. The depth map includes the depth value for each pixel.

[0162] Optionally, the "camera corresponding to the first rendered video" may refer to the camera corresponding to the first sample camera parameters, that is, the camera corresponding to the first sample video latent variables. Here, "camera corresponding to the first sample video latent variables" refers to a camera pre-given for the first sample video latent variables, so as to generate a video sequence that conforms to the motion of the given camera based on the first sample video latent variables.

[0163] Optionally, the first real video tag includes the video frame tag corresponding to each candidate video frame. Therefore, the video tag corresponding to the target video frame is determined as the target video frame tag.

[0164] In some embodiments, determining the target video frame visible in the first real video tag among the candidate video frames included in the first rendered video, based on the depth map and the intrinsic and extrinsic parameters of the camera corresponding to the first rendered video, may include: Based on the intrinsic and extrinsic parameters of the camera corresponding to the first rendered video and the depth of each pixel, each pixel in the candidate video frame is projected onto world coordinates to obtain the world coordinates corresponding to each pixel.

[0165] Based on the intrinsic and extrinsic parameters of the camera corresponding to the first real video tag, the world coordinates corresponding to each pixel are projected onto the first real video tag to obtain the pixel coordinates of each pixel in the first real video tag.

[0166] Determine the pixel coordinates of each pixel in the first real video label, and the target depth in the first real video label; Determine the world coordinates corresponding to each pixel and the initial depth in the first real video label.

[0167] If the difference between the initial depth and the target depth is less than a preset threshold, and the target depth is valid, the candidate video frame is determined as the target video frame.

[0168] Optionally, based on the intrinsic and extrinsic parameters of the camera corresponding to the first rendered video and the depth of each pixel, each pixel in the candidate video frame is projected onto world coordinates, and the formula for calculating the world coordinates corresponding to each pixel can be as follows: ; Where D(u, v) is the depth of each pixel, K refers to the intrinsic parameters of the camera corresponding to the first rendered video, and E is the extrinsic parameters of the camera corresponding to the first rendered video. This refers to the world coordinates corresponding to each pixel. This refers to converting each pixel coordinate into camera-normalized coordinates. This refers to multiplying the normalized coordinates by the depth to obtain the 3D point in the target camera coordinate system.

[0169] Optionally, based on the intrinsic and extrinsic parameters of the camera corresponding to the first real video label, the world coordinates corresponding to each pixel are projected onto the first real video label, and the calculation formula for obtaining the pixel coordinates of each pixel in the first real video label can be as follows: ; in, This refers to the pixel coordinates of each pixel in the first real-view label. K0 refers to the intrinsic parameters of the camera corresponding to the first real-view label, and E0 refers to the extrinsic parameters of the camera corresponding to the first real-view label. "The camera corresponding to the first real-view label" can refer to the camera used to generate the first real-view.

[0170] Optionally, the server can calculate the world coordinates corresponding to each pixel using the following formula ( The initial depth in the initial conditional coordinate system (i.e., the first real video label): ; in, This refers to the initial depth, and z refers to the z-component of the camera's coordinate system corresponding to the first real video label.

[0171] Optionally, the server can calculate the target depth in the first real video label by using the following formula: (i.e., the projection point in the first real video label) the pixel coordinates of each pixel in the first real video label. ; in, This refers to the target depth. This refers to the projection point of each pixel within the pixels of the first real video tag.

[0172] Alternatively, the server can perform depth comparison and determine visibility using the following formula: ; in, This refers to the visibility judgment result. A value of 1 indicates that the candidate video frame is a target video frame that is visible in the first real video tag. A value of 0 indicates that the candidate video frame is not visible in the first real video tag. This refers to the preset threshold. This refers to the difference between the initial depth and the target depth. This refers to the target depth being effective.

[0173] Optionally, if the difference between the initial depth and the target depth is less than a preset threshold, it indicates that the depth of the world coordinates corresponding to each pixel in the camera corresponding to the first real video label is almost the same as the depth of the projection point in the camera corresponding to the first real video label. This means that the projection of this point in the first real video label is not occluded by other closer objects (if occluded, ...). It will be the depth of the occluded object, and (The differences are significant). In this case, it indicates that the projection point is within the valid range of the first real video tag (the depth is positive, and it is not the background or invalid area). Therefore, if the difference between the initial depth and the target depth is less than a preset threshold, and the target depth is valid, the candidate video frame is considered a visible target video frame; otherwise, it is an invisible video frame.

[0174] Therefore, by utilizing the depth map and camera intrinsic and extrinsic parameters, and through a geometric transformation of back projection-forward projection-depth comparison, the visibility of pixels in the first real video tag can be accurately and quantitatively determined. Furthermore, if the difference between the initial depth and the target depth is less than a preset threshold, it indicates that the depth of the world coordinates corresponding to each pixel in the camera corresponding to the first real video tag is almost identical to the depth of the projected point in the camera corresponding to the first real video tag. This means that the projection of that point in the first real video tag is not occluded by other closer objects. In this case, it indicates that the projection point is within the effective range of the first real video label (the depth is positive, and it is not the background or invalid area). Therefore, if the difference between the initial depth and the target depth is less than the preset threshold, and the target depth is valid, the candidate video frame is considered a visible target video frame. This can further improve the accuracy of pixel visibility judgment, thereby improving the accuracy of reward information determination, so that the initial video generation model can accurately generate the first sample video latent variables that meet the preset conditions.

[0175] In some embodiments, determining the difference between the target video frame and the target video frame label to obtain reward information includes: Determine the second mean square error and the second learned perceptual image patch similarity between the target video frame and the target video frame label.

[0176] Reward information is generated based on the second mean square error and the similarity of the second learned perceptual image patches.

[0177] In one approach, the server can directly calculate the second mean square error or the second learned perceptual image patch similarity between the target video frame and the target video frame label, and use the second mean square error or the second learned perceptual image patch similarity as reward information.

[0178] In one approach, the server can assign a weight to the similarity of the second learned perceptual image patches based on their contribution to the loss data, and obtain this reward information using the following formula: ; in, This refers to reward information. This refers to the second mean square error. This refers to the second-learned perceptual image patch similarity. This refers to the target video frame. This refers to the target video frame label. It refers to the mask. This refers to weight.

[0179] Therefore, by calculating loss data through pixel-level and perceptual-level differences between target video frames and their labels, the calculated loss data can reflect both pixel-level and perceptual differences. This allows for adjustment of the initial video generation model parameters from both pixel-level and perceptual-level perspectives, aligning the latent variables of the first sample video with the first sample camera parameters. This minimizes the difference between the first rendered video obtained based on the first sample 3D scene information and the real video, thereby improving the training accuracy of the initial video generation model. The trained initial video generation model can then accurately decode the input video latent variables and camera parameters to obtain the corresponding 3D scene information.

[0180] The video generation model training method of this application has the following beneficial effects: 1) This disclosure designs the decoder to be both an efficient decoder capable of converting latent variables into 3D scene information and a differentiable reward function sensitive to camera-video alignment. It utilizes a dual mechanism of prototype parameters (e.g., camera trajectory, and more specifically, camera pose) as both input and projection parameters to achieve efficient "one-time" 3D reconstruction, significantly improving training efficiency and resource utilization, enhancing scene 3D consistency, and substantially improving camera control precision and video quality.

[0181] Regarding "significantly improving camera control precision and video quality": This disclosure optimizes the camera reward, enabling the generated video to far surpass related technologies in following the camera trajectory. It achieves optimal results in camera control metrics such as rotation error (Rerr) and displacement error (Terr), while also significantly outperforming in video fidelity metrics (FID, FVD).

[0182] Regarding "achieving efficient 'one-time' 3D reconstruction": the camera-aware 3D decoder disclosed herein enables direct, feedforward generation of 3D scenes from video latent variables, completely avoiding the scene-by-scene optimization process that takes several minutes to several hours in the two-stage methods of related technologies. Regarding "significantly improving training efficiency and resource utilization": Compared to traditional ReFL methods, this disclosure eliminates the need to decode latent variables into high-resolution RGB video for reward calculation. Instead, it directly operates through a lightweight 3D decoder, greatly reducing memory consumption and computation time. The memory overhead of the 3D decoder in this disclosure is only about 1 / 5 that of a standard video VAE decoder, and its time overhead is only about 1 / 10.

[0183] Regarding "enhanced 3D consistency of the scene": Since the reward signal originates from a 3DGS representation with inherent 3D consistency, through ReFL's "knowledge distillation", the video generation model itself also learns to generate geometrically more coherent and static content, effectively suppressing unnecessary dynamic elements or artifacts.

[0184] 2) The camera reward in this disclosure cleverly distinguishes between "deterministic regions" and "random regions" in the generated content through geometric calculations, and applies pixel-level alignment rewards only to deterministic regions. This ensures both the accuracy of camera motion and the diversity of model generation, effectively avoiding the "reward hacking problem".

[0185] 3) Reward feedback learning loop in latent space: This disclosure constructs an efficient ReFL framework in which the entire "generation-evaluation-feedback" optimization loop is completed in the computationally inexpensive latent space, bypassing the expensive pixel space decoding steps in related technologies, and providing a feasible path for fine-tuning alignment on large video models.

[0186] Figure 7 This is a schematic diagram illustrating a video generation method according to an exemplary embodiment of this disclosure, such as... Figure 7 As shown, the video generation method includes: In step S31, the image to be processed and the target camera parameters corresponding to the image to be processed are obtained.

[0187] In step S32, the image to be processed and the target camera parameters are input into the target video generation model to obtain the target video latent variables.

[0188] In step S33, the latent variables of the target video are decoded to align the image to be processed and the target camera parameters to obtain the target three-dimensional scene information.

[0189] In step S34, the target 3D scene information is rendered to obtain a target rendered video that conforms to the target camera parameters; wherein, the target video generation model is obtained based on the video generation model training method of any of the above embodiments.

[0190] In this embodiment, during the inference application stage, only one image to be processed and the target camera parameters corresponding to the custom image to be processed are required to perform instant 3D reconstruction and high-quality video generation operations.

[0191] Optionally, for the real-time 3D scene reconstruction process: the server can input the image to be processed and the target camera parameters into the target video generation model for processing to obtain the target video latent variables. Then, the target video latent variables are decoded to align the image to be processed and the target camera parameters to obtain the target 3D scene information. Furthermore, the server can input the target video latent variables into the aforementioned trained target decoder, which generates a 3DGS representation of the scene in a single feedforward, allowing for rendering and exploration from any new perspective without any additional optimization. For example, the process of the target decoder decoding the target video latent variables to align the image to be processed and the target camera parameters to obtain the target 3D scene information may include: extracting camera features from the target camera parameters based on the first camera encoding module to obtain target camera features; extracting video features from the target video latent variables based on the first video encoding module, and performing block processing on the extracted features to obtain target video features; concatenating the target camera features and target video features to obtain target concatenated features; and performing attention processing on the target concatenated features based on the first Transformer module to obtain the target 3D scene information.

[0192] Optionally, for high-quality video generation scenarios: the server can render the target 3D scene information to obtain a target rendered video that conforms to the target camera parameters.

[0193] Optionally, the target camera parameters may include, but are not limited to: camera motion trajectory (including rotation, translation, etc.), point cloud rendering conditions, etc. Among them, the camera motion trajectory can be further parameterized into a format that can be processed by neural networks, such as camera pose, and even further, it can be Plücker embeddings.

[0194] Optionally, the target 3D scene information may include, but is not limited to, 3D Gaussian parameters (3DGS) and other explicit or implicit 3D representations. For example, when the first sample camera parameters are the camera motion trajectory, the 3D Gaussian parameters may include, but are not limited to, position, covariance matrix (rotation, scaling), spherical harmonic coefficients, opacity, etc.

[0195] Therefore, in the video generation stage, only one image to be processed and the corresponding target camera parameters of the image need to be provided to perform real-time 3D reconstruction and high-quality video generation, which reduces the cost of video generation and improves the quality and efficiency of video generation.

[0196] The following is a general explanation of the training methods for the video generation model and the video generation methods described above: 1) Build and train the basic initial video generation model.

[0197] This disclosure pre-constructs an initial video generation model (e.g., an I2V diffusion model) that accepts camera motion as a condition. Figure 3 As shown, this initial video generation model is based on a pre-trained I2V model, with camera conditions injected through the ControlNet architecture. Specifically, the camera embedding sequence is processed by a trainable camera encoder, and its output is added to the input of the backbone Transformer module of the I2V model. To improve the effectiveness of camera control, this disclosure employs a truncated normal distribution to sample time steps during training, biasing it towards larger values ​​(i.e., early denoising), thereby focusing on optimizing the model's ability to model camera motion in the early stages.

[0198] After training, the initial video generation model is able to generate a predicted latent variable for the denoised video based on the input image and camera trajectory.

[0199] 2) Design and train a camera-aware target decoder.

[0200] This target decoder replaces the traditional high-overhead video decoder and acts as the reward model. The target decoder includes a first camera encoding module, a first video encoding module, and a first Transformer module. The training process of this target decoder includes: Obtain the latent variables of the second sample video and the parameters of the second sample camera; input the latent variables of the second sample video and the parameters of the second sample camera into a preset decoder for decoding to align the latent variables of the second sample video and the parameters of the second sample camera, and obtain the second sample 3D scene information corresponding to the latent variables of the second sample video; render the second sample 3D scene information to obtain the second rendered video; generate loss data based on the difference between the second real video label and the second rendered video; adjust the model parameters of the preset decoder based on the loss data to obtain the target decoder.

[0201] 3) Implement camera reward optimization This step utilizes the target decoder trained in step two as the reward function, and uses ReFL to fine-tune the initial video generation model in step one, thereby further improving camera control accuracy. The specific process includes: Generation and Decoding: Obtain the sample image and the corresponding initial sample camera parameters; input the sample image and the corresponding initial sample camera parameters into the initial video generation model trained in step one to obtain a first sample video latent variable; input the first sample video latent variable and the first sample camera parameters into the target decoder trained in step two for decoding to obtain the first sample 3D scene information.

[0202] Rendering and reward calculation: Render the first rendered video from the first sample 3D scene information, and generate reward information based on the difference between the first real video label and the first rendered video.

[0203] In some embodiments, this disclosure introduces a visibility-aware reward strategy: considering the randomness of video generation, newly generated content (such as new areas seen after the camera rotates) has no corresponding counterpart in the real video, and directly calculating pixel differences would penalize the model's creativity. Therefore, we introduce visibility masks, Mask Calculation: For each candidate video frame of the first rendered video, this disclosure utilizes the depth map provided by the 3D decoder and camera intrinsic and extrinsic parameters to calculate whether each pixel is visible in the initial conditional image through geometric transformation (back projection followed by forward projection). The final reward loss function is calculated only in the deterministic pixel regions marked as visible by the mask. This forces the model to accurately align camera motion in areas with existing content while allowing reasonable content generation in unknown areas.

[0204] Gradient backpropagation: Adjust the model parameters of the initial video generation model in step one based on the reward information. This process will incentivize the initial video generation model to produce the first sample video latent variables that are in camera condition alignment, thus obtaining the target video generation model. This is because such latent variables can obtain higher rewards (i.e. lower losses) after passing through the target decoder.

[0205] 4) Model Inference Once training is complete, this disclosure can efficiently accomplish the following tasks by providing only an image to be processed and a set of custom target camera parameters (e.g., camera trajectory): High-quality video generation: Video latent variables are generated through an initial video generation model, and then a video with precise camera control can be obtained by using a standard video variational autoencoder (VAE).

[0206] Real-time 3D scene reconstruction: The generated video latent variables are input into the trained target decoder, which can feed forward to generate the 3DGS representation of the scene in one go, and render and explore any new perspective without any additional optimization.

[0207] This disclosure also provides a video generation model training device. Figure 8 This is a block diagram illustrating a video generation model training apparatus according to an exemplary embodiment. (Refer to...) Figure 8 As shown, the training device for the video generation model includes: The first sample data acquisition module 41 is configured to acquire the first sample video latent variable and the first sample camera parameter corresponding to the first sample video latent variable; the first sample video latent variable is obtained by video generation processing of the sample image and the initial sample camera parameter corresponding to the sample image based on the initial video generation model, and the first sample video latent variable is labeled with a first real video label. The first sample 3D scene information generation module 42 is configured to decode the first sample video latent variables and the first sample camera parameters to align the first sample video latent variables and the first sample camera parameters to obtain the first sample 3D scene information. The first rendering video generation module 43 is configured to render the three-dimensional scene information of the first sample to obtain a first rendering video that conforms to the camera parameters of the first sample. The reward information generation module 44 is configured to generate reward information based on the difference between the first real video tag and the first rendered video. The target video generation model generation module 45 is configured to adjust the model parameters of the initial video generation model according to the reward information, so that the latent variables of the first sample video generated by the initial video generation model meet the preset conditions, thereby obtaining the target video generation model.

[0208] In an optional embodiment, the first sample 3D scene information generation module is configured to perform: Input the latent variables of the first sample video and the camera parameters of the first sample into the target decoder; Based on the target decoder, the first sample video latent variables and the first sample camera parameters are decoded to align the first sample video latent variables and the first sample camera parameters, thereby obtaining the first sample three-dimensional scene information; The target decoder is obtained by training a preset decoder based on the second real video label corresponding to the second sample video latent variable and the second rendered video; the second rendered video is obtained by rendering the second sample three-dimensional scene information, and the second sample three-dimensional scene information is obtained by decoding the second sample video latent variable and the second sample camera parameters corresponding to the second sample video latent variable. In an optional embodiment, the target decoder includes a first camera encoding module, a first video encoding module, and a first Transformer module, and the first sample 3D scene information generation module is configured to perform: The camera features of the first sample camera are obtained by extracting camera features from the first sample camera parameters based on the first camera encoding module. Based on the first video encoding module, video features are extracted from the latent variables of the first sample video, and the extracted features are segmented to obtain the first sample video features. The first sample camera features and the first sample video features are spliced ​​together to obtain sample spliced ​​features; Attention processing is performed on the sample stitching features based on the first Transformer module to obtain the three-dimensional scene information of the first sample.

[0209] In an optional embodiment, the video generation model training apparatus further includes: The second sample data acquisition module is configured to acquire the second sample video latent variables and the second sample camera parameters. The second sample 3D scene information generation module is configured to input the second sample video latent variable and the second sample camera parameter into the preset decoder for decoding, so as to align the second sample video latent variable and the second sample camera parameter to obtain the second sample 3D scene information corresponding to the second sample video latent variable; The second rendering video generation module is configured to render the second sample 3D scene information to obtain the second rendering video. The loss data generation module is configured to generate loss data based on the difference between the second real video label and the second rendered video. The target decoder generation module is configured to adjust the model parameters of the preset decoder based on the loss data to obtain the target decoder.

[0210] In an optional embodiment, the loss data generation module is configured to perform: Determine the first mean square error and the first learned perceptual image patch similarity between the second real video label and the second rendered video; generate the loss data based on the first mean square error and the first learned perceptual image patch similarity.

[0211] In an optional embodiment, the first sample camera parameters include camera motion trajectory, the first sample 3D scene information includes position information, and the video generation model training device further includes: The depth map acquisition module is configured to acquire a depth map from the first sample 3D scene information; the depth map includes the depth value of each pixel; The origin and orientation determination module is configured to convert the camera motion trajectory into the origin and orientation of the ray corresponding to each pixel; The location information generation module is configured to combine the depth value, the origin and direction of the ray to obtain the location information.

[0212] In an optional embodiment, the reward information generation module is configured to perform: Obtain a depth map from the three-dimensional scene information of the first sample; Based on the depth map and the intrinsic and extrinsic parameters of the camera corresponding to the first rendered video, determine the target video frame that is visible in the first real video tag among the candidate video frames included in the first rendered video. The difference between the target video frame and the target video frame label is determined to obtain the reward information; The target video frame tag is the video tag corresponding to the target video frame in the first real video tag.

[0213] In an optional embodiment, the depth map includes a depth value for each pixel, and the reward information generation module is configured to perform: Based on the intrinsic and extrinsic parameters of the camera corresponding to the first rendered video and the depth value of each pixel, each pixel in the candidate video frame is projected onto world coordinates to obtain the world coordinates corresponding to each pixel. Based on the intrinsic and extrinsic parameters of the camera corresponding to the first rendered video and the depth value of each pixel, the world coordinates corresponding to each pixel are projected onto the first real video label to obtain the pixel coordinates of each pixel in the first real view label. Determine the pixel coordinates of each pixel in the first real video label, and the target depth in the first real video label; Determine the world coordinates of each pixel and its initial depth in the first real video label; If the difference between the initial depth and the target depth is less than a preset threshold, and the target depth is valid, the candidate video frame is determined to be the target video frame.

[0214] In an optional embodiment, the reward information generation module is configured to perform: Determine the second mean square error and the second learned perceptual image patch similarity between the target video frame and the target video frame label; The reward information is generated based on the second mean square error and the second learned perceptual image patch similarity.

[0215] In an optional embodiment, the initial video generation model includes a second camera encoding module, a second video encoding module, and a second Transformer module, and the video generation model training device further includes: The initial data acquisition module is configured to acquire the sample image, the initial sample camera parameters, and random noise latent variables; The second sample video feature generation module is configured to input the sample image, the initial sample camera parameters, and the random noise latent variable into the initial video generation model; extract camera features from the initial sample camera parameters based on the second camera encoding module to obtain second sample camera features; perform padding processing on the sample image based on the second video encoding module; concatenate the padded image and the random noise latent variable; and perform block processing on the concatenated features to obtain second sample video features. The fusion feature generation module is configured to perform fusion processing on the second sample camera features and the second sample video features to obtain fusion features; The first sample video latent variable generation module is configured to perform attention processing on the fused features based on the second Transformer module to obtain the first sample video latent variables.

[0216] This disclosure also provides a video generation apparatus. Figure 9 This is a block diagram illustrating a video generation apparatus according to an exemplary embodiment. (Refer to...) Figure 9 As shown, the video generation device includes: The data acquisition module 51 is configured to acquire the image to be processed and the target camera parameters corresponding to the image to be processed. The target video latent variable generation module 52 is configured to input the image to be processed and the target camera parameters into the target video generation model to obtain the target video latent variables; The target 3D scene information generation module 53 is configured to decode the target video latent variables and the target camera parameters to align the image to be processed and the target camera parameters to obtain target 3D scene information; The target rendering video generation module 54 is configured to render the target 3D scene information to obtain a target rendering video that conforms to the target camera parameters. The target video generation model is obtained based on the video generation model training method described above.

[0217] Regarding the apparatus in the above embodiments, the specific manner in which each set of modules performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0218] In an exemplary embodiment, an electronic device is also provided, including a processor; a memory for storing processor-executable instructions; wherein, when the processor is configured to execute the instructions stored in the memory, it implements the steps of any of the video generation model training or video generation methods in the above embodiments.

[0219] The electronic device can be a terminal, a server, or a similar computing device. Taking a server as an example... Figure 10 This is a block diagram illustrating an electronic device 60 according to an exemplary embodiment. The electronic device 60 can vary significantly due to different configurations or performance characteristics. It may include one or more Central Processing Units (CPUs) 61 (CPUs 61 may include, but are not limited to, microprocessors (MCUs) or programmable logic devices (FPGAs), a memory 63 for storing data, and one or more storage media 62 (e.g., one or more mass storage devices) for storing application programs 623 or data 622. The memory 63 and storage media 62 may be temporary or persistent storage. The program stored in the storage media 62 may include one or more modules, each module including a series of instruction operations on the electronic device. Furthermore, the CPU 61 may be configured to communicate with the storage media 62 and execute the series of instruction operations stored in the storage media 62 on the electronic device 60. Electronic device 60 may also include one or more power supplies 66, one or more wired or wireless network interfaces 65, one or more input / output interfaces 64, and / or one or more operating systems 621, such as Windows Server™, Mac OSX™, Unix™, Linux™, FreeBSD™, etc.

[0220] The input / output interface 64 can be used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the electronic device 60. In one example, the input / output interface 64 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In an exemplary embodiment, the input / output interface 64 may be a radio frequency (RF) module for wireless communication with the Internet.

[0221] Those skilled in the art will understand that Figure 10 The structure shown is for illustrative purposes only and does not limit the structure of the electronic device described above. For example, electronic device 60 may also include... Figure 10 The more or fewer components shown, or having the same Figure 10 The different configurations shown.

[0222] In an exemplary embodiment, a computer-readable storage medium is also provided, which, when executed by a processor of an electronic device, enables the electronic device to perform the steps of any of the video generation model training or video generation methods described above.

[0223] In an exemplary embodiment, a computer program product is also provided, including a computer program that, when executed by a processor, implements the video generation model training or video generation method provided in any of the above embodiments.

[0224] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this disclosure can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0225] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0226] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A method for training a video generation model, characterized in that, The method includes: Obtain the first sample video latent variable and the first sample camera parameter corresponding to the first sample video latent variable; the first sample video latent variable is obtained by video generation processing of the sample image and the initial sample camera parameter corresponding to the sample image based on the initial video generation model, and the first sample video latent variable is labeled with the first real video label. The first sample video latent variables and the first sample camera parameters are decoded to align the first sample video latent variables and the first sample camera parameters to obtain the first sample three-dimensional scene information; Render the three-dimensional scene information of the first sample to obtain a first rendered video that conforms to the camera parameters of the first sample; Reward information is generated based on the difference between the first real video tag and the first rendered video; The model parameters of the initial video generation model are adjusted according to the reward information so that the latent variables of the first sample video generated by the initial video generation model meet the preset conditions, thereby obtaining the target video generation model.

2. The video generation model training method according to claim 1, characterized in that, Decoding the first sample video latent variables and the first sample camera parameters to align the first sample video latent variables and the first sample camera parameters to obtain the first sample 3D scene information includes: Input the latent variables of the first sample video and the camera parameters of the first sample into the target decoder; Based on the target decoder, the first sample video latent variables and the first sample camera parameters are decoded to align the first sample video latent variables and the first sample camera parameters, thereby obtaining the first sample three-dimensional scene information; The target decoder is obtained by training a preset decoder based on the second real video label corresponding to the second sample video latent variable and the second rendered video; the second rendered video is obtained by rendering the second sample three-dimensional scene information, and the second sample three-dimensional scene information is obtained by decoding the second sample video latent variable and the second sample camera parameters corresponding to the second sample video latent variable.

3. The video generation model training method according to claim 2, characterized in that, The target decoder includes a first camera encoding module, a first video encoding module, and a first Transformer module. The step of decoding the first sample video latent variables and the first sample camera parameters based on the target decoder to align the first sample video latent variables and the first sample camera parameters, thereby obtaining the first sample 3D scene information, includes: The camera features of the first sample camera are obtained by extracting camera features from the first sample camera parameters based on the first camera encoding module. Based on the first video encoding module, video features are extracted from the latent variables of the first sample video, and the extracted features are segmented to obtain the first sample video features. The first sample camera features and the first sample video features are spliced ​​together to obtain sample spliced ​​features; Attention processing is performed on the sample stitching features based on the first Transformer module to obtain the three-dimensional scene information of the first sample.

4. The video generation model training method according to claim 2, characterized in that, The training process of the target decoder includes: Obtain the latent variables of the second sample video and the camera parameters of the second sample; The second sample video latent variable and the second sample camera parameter are input into the preset decoder for decoding, so as to align the second sample video latent variable and the second sample camera parameter to obtain the second sample three-dimensional scene information corresponding to the second sample video latent variable; Render the second sample's 3D scene information to obtain the second rendered video; Loss data is generated based on the difference between the second real video label and the second rendered video; The model parameters of the preset decoder are adjusted based on the loss data to obtain the target decoder.

5. The video generation model training method according to claim 4, characterized in that, The step of generating loss data based on the difference between the second real video tag and the second rendered video includes: Determine the first mean square error and the first learned perceptual image patch similarity between the second real video label and the second rendered video; The loss data is generated based on the first mean square error and the first learned perceptual image patch similarity.

6. The video generation model training method according to any one of claims 1 to 5, characterized in that, The first sample camera parameters include camera motion trajectory, and the first sample 3D scene information includes position information. The process of generating the position information includes: A depth map is obtained from the three-dimensional scene information of the first sample; the depth map includes the depth value of each pixel; The camera motion trajectory is converted into the origin and direction of the ray corresponding to each pixel; The location information is obtained by combining the depth value, the origin of the ray, and its direction.

7. The video generation model training method according to any one of claims 1 to 5, characterized in that, The step of generating reward information based on the difference between the first real video tag and the first rendered video includes: Obtain a depth map from the three-dimensional scene information of the first sample; Based on the depth map and the intrinsic and extrinsic parameters of the camera corresponding to the first rendered video, determine the target video frame that is visible in the first real video tag among the candidate video frames included in the first rendered video. The difference between the target video frame and the target video frame label is determined to obtain the reward information; The target video frame tag is the video tag corresponding to the target video frame in the first real video tag.

8. The video generation model training method according to claim 7, characterized in that, The depth map includes the depth value of each pixel in the candidate video frames. The step of determining the target video frame visible in the first real video tag among the candidate video frames included in the first rendered video, based on the depth map and the intrinsic and extrinsic parameters of the camera corresponding to the first rendered video, includes: Based on the intrinsic and extrinsic parameters of the camera corresponding to the first rendered video and the depth value of each pixel, each pixel in the candidate video frame is projected onto world coordinates to obtain the world coordinates corresponding to each pixel. Based on the intrinsic and extrinsic parameters of the camera corresponding to the first rendered video and the depth value of each pixel, the world coordinates corresponding to each pixel are projected onto the first real video label to obtain the pixel coordinates of each pixel in the first real view label. Determine the pixel coordinates of each pixel in the first real video label, and the target depth in the first real video label; Determine the world coordinates of each pixel and its initial depth in the first real video label; If the difference between the initial depth and the target depth is less than a preset threshold, and the target depth is valid, the candidate video frame is determined to be the target video frame.

9. The video generation model training method according to claim 7, characterized in that, The step of determining the difference between the target video frame and the target video frame label to obtain the reward information includes: Determine the second mean square error and the second learned perceptual image patch similarity between the target video frame and the target video frame label; The reward information is generated based on the second mean square error and the second learned perceptual image patch similarity.

10. The video generation model training method according to any one of claims 1 to 5, characterized in that, The initial video generation model includes a second camera encoding module, a second video encoding module, and a second Transformer module. The generation process of the latent variables of the first sample video includes: Obtain the sample image, the initial sample camera parameters, and the random noise latent variable; The sample image, the initial sample camera parameters, and the random noise latent variable are input into the initial video generation model. Based on the second camera encoding module, the camera features in the initial sample camera parameters are extracted to obtain the second sample camera features. Based on the second video encoding module, the sample image is filled, the filled image and the random noise latent variable are spliced ​​together, and the spliced ​​features are segmented to obtain the second sample video features. The second sample camera features and the second sample video features are fused to obtain fused features; Attention processing is performed on the fused features based on the second Transformer module to obtain the latent variables of the first sample video.

11. A video generation method, characterized in that, The video generation method includes: Obtain the image to be processed and the target camera parameters corresponding to the image to be processed; Input the image to be processed and the target camera parameters into the target video generation model to obtain the latent variables of the target video; The latent variables of the target video and the parameters of the target camera are decoded to align the image to be processed and the parameters of the target camera, thereby obtaining the target three-dimensional scene information; Render the target 3D scene information to obtain a target rendered video that conforms to the target camera parameters; The target video generation model is obtained based on the video generation model training method according to any one of claims 1 to 10.

12. A video generation model training device, characterized in that, The device includes: The first sample data acquisition module is configured to acquire the first sample video latent variable and the first sample camera parameter corresponding to the first sample video latent variable; the first sample video latent variable is obtained by video generation processing of the sample image and the initial sample camera parameter corresponding to the sample image based on the initial video generation model, and the first sample video latent variable is labeled with a first real video label. The first sample 3D scene information generation module is configured to decode the first sample video latent variables and the first sample camera parameters to align the first sample video latent variables and the first sample camera parameters to obtain the first sample 3D scene information. The first rendering video generation module is configured to render the three-dimensional scene information of the first sample to obtain a first rendering video that conforms to the camera parameters of the first sample. The reward information generation module is configured to generate reward information based on the difference between the first real video tag and the first rendered video. The target video generation model generation module is configured to adjust the model parameters of the initial video generation model according to the reward information, so that the latent variables of the first sample video generated by the initial video generation model meet the preset conditions, thereby obtaining the target video generation model.

13. A video generation apparatus, characterized in that, The video generation device includes: The data acquisition module is configured to acquire the image to be processed and the target camera parameters corresponding to the image to be processed. The target video latent variable generation module is configured to input the image to be processed and the target camera parameters into the target video generation model to obtain the target video latent variables; The target 3D scene information generation module is configured to decode the target video latent variables and the target camera parameters to align the image to be processed and the target camera parameters to obtain target 3D scene information; The target rendering video generation module is configured to render the target 3D scene information to obtain a target rendering video that conforms to the target camera parameters; The target video generation model is obtained based on the video generation model training method according to any one of claims 1 to 10.

14. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing a computer program that is loaded and executed by the processor to implement the video generation model training method as described in any one of claims 1 to 10, or the video generation method as described in claim 11.