Video generation method with camera control without training and display repair

By rearrangeing and resampling potential variables during the denoising process of video generation, and combining the noise heavy injection mechanism, the high training cost and cumbersome inference process of camera-controlled video generation in the prior art is solved, and a lightweight and concise camera-controlled video generation method is realized.

CN119996853APending Publication Date: 2025-05-13XIAMEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510105755.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When realizing camera-controlled video generation, the prior art faces high training costs and cumbersome inference processes, and ordinary users are unable to withstand the needs of high computing resources and complex models.

Method used

By performing rearrangement operations on potential variables during the denoising process, simulate camera actions such as translation, scaling, and rotation, and applying resampling strategies and cross-frame fusion alignment strategies in the latent space, combined with the noise reinjection mechanism, the view angle conversion of video generated is realized.

Benefits of technology

The camera-controlled video generation without training and display repair is realized, which simplifies the model structure, reduces the computing resource requirements, and improves the quality and stability of video generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996853A_ABST
    Figure CN119996853A_ABST
Patent Text Reader

Abstract

The invention discloses a video generation method with camera control, which does not need training and display repair, so that a common base model can also have camera control capability, and the method is operated in a potential space, does not need an additional repair model and a depth estimation model, and realizes conciseness and light weight. According to the video generation method, a specific time step # imgabs0 # in a denoising process executes rearrangement operation on potential variables of each frame; simulating a specific camera action by changing the arrangement sequence of the potential variables; afterwards, a resampling strategy is applied in the potential space to fill the new view angle area, and meanwhile, a cross-frame fusion alignment strategy is combined to ensure that the sampling process is kept consistent between frames; a noise re-injection mechanism is introduced, and the noise is re-injected into a potential variable in the later stage of denoising, so that the denoising time is prolonged, the distribution offset phenomenon caused by rearrangement and re-sampling is relieved, and the video generation quality is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of controllable video generation, and is a subtask of controllable video generation in deep learning, and in particular refers to a video generation method with camera control that does not require training or display repair, and involves using additional input parameters to more accurately control the viewing angle of video generation. Background Art

[0002] With the emergence of Sora, many open source video generation frameworks have emerged, making video generation a research hotspot. Inspired by the growing needs of the film and virtual reality industries, researchers have recently begun to explore the use of additional control conditions to achieve complex and sophisticated camera movements, aiming to promote the application of camera-controlled video generation technology in real life.

[0003] Current research can be divided into two categories: training-based camera control and training-free camera control. Training-based camera control video generation methods mainly use displayed camera information to accurately control the viewing angle of video generation. For example, MotionCtrl and CameraCtrl train an additional encoder to process camera parameters and integrate the encoded features into the temporal attention layer of the video base for viewing angle control. However, the training of additional encoders and fine-tuning of the base require a lot of overhead, and the lack of high-quality camera motion data limits the development of this type of method. Some recent studies have begun to try to enable the basic model to generate videos with camera control in a training-free way. For example, CamTrol uses the depth estimation model ZoeDepth to convert the input image into a three-dimensional point cloud and render different perspectives through the camera parameters input by the user. Although these methods avoid resource overhead during training. However, they all introduce additional models for point cloud modeling and vacancy repair in the inference stage, making the model very bloated.

[0004] In summary, although existing methods have made some progress, they still face many challenges: 1. High training costs CameraCtrl uses 16 A100 GPUs to train the encoder and fine-tune the model, which is too costly and computationally demanding for ordinary users; in addition, the lack of a specific camera parameter dataset also limits the generalization ability of the model.

[0005] 2. Complicated reasoning process Although CamTrol uses a training-free approach to achieve camera control, it requires an additional depth estimation model ZoeDepth and a restoration model StableDiffusion, which significantly increases the computational overhead and complexity. Summary of the invention

[0006] The main purpose of the present invention is to provide a video generation method with camera control that does not require training or display repair, so as to solve the problems existing in the prior art and enable ordinary base models to have camera control capabilities. The method operates in the latent space and does not require additional repair models and depth estimation models, thus achieving simplicity and lightweight.

[0007] In order to achieve the above object, the solution of the present invention is: A training-free, display-free video generation method with camera control at specific time steps during denoising A rearrangement operation is performed on the latent variables of each frame; a specific camera action is simulated by changing the arrangement order of the latent variables; then, a resampling strategy is applied in the latent space to fill the new view area, and a cross-frame fusion alignment strategy is combined to ensure that the sampling process remains consistent between frames; a noise re-injection mechanism is introduced to extend the denoising time and alleviate the distribution shift caused by rearrangement and resampling by reinjecting noise into the latent variables in the later stage of denoising.

[0008] The rearrangement operation defines the positive direction as right, downward, inward and counterclockwise rotation, and the user can input the movement parameters. , , and rotation angle , the basic perspective transformation of translation, scaling and rotation is realized through the combination of motion parameters, and the original latent variables Mapping to new latent variables To achieve perspective conversion, Indicates the number of frames; The translation operation is performed by the motion parameters input by the user. and Update Latent variables of frames For the Frame, the camera is translated along the X axis , translate along the Y axis units, and then the translated latent variables are clipped back to the original position, where represents the width of the latent space, represents the height of the latent space; The scaling operation is performed by applying the latent variables of each frame to Interpolation is applied to simulate the effect of perspective zoom, first using a Gaussian kernel It is processed, interpolated to obtain an intermediate latent variable, and finally its center is cropped to a specific size to obtain the final latent variable. The above process is expressed as: ; ; ; ; in, Represents an interpolation operation; represents the Gaussian blur kernel; Indicates The scaling factor of the frame, whose value is ; Indicates the operation of taking the minimum value; Indicates cropping operation according to the preset size; The rotation operation introduces the point cloud projection theory into the latent space, which is the latent variable Each position in Generate its 3D point cloud representation, where and Represent the coordinates of any point respectively; then project them to different perspectives according to different parameters: ; ; Among them, the superscript Represents the transpose of a matrix; Represents the intrinsic parameter matrix of the camera; Indicates The rotation matrix of the frame; Indicates the depth information of the corresponding position; parameter and The calculation formula is as follows: ; For each frame, the parameters The latent variables are rotated, where Evenly distributed in arrive within the range.

[0009] The resampling strategy specifically introduces a subject mask during the sampling process to limit the sampling range, and the new viewing area The token is sampled from a column of a specific row or a row of a specific column, and the orientation of these rows or columns is determined by the mask The background area in is defined as follows: ; ; in, , and Respectively represent the index of any position in the latent variable; Represents the new viewing area for each frame. Row sampling and column sampling will be performed separately to ensure that all new viewing areas are fully filled. At the initial time step In , the cross attention map of the last upsampling module is stored, and the response map corresponding to the target token is extracted and reshaped into , Represents the real number space; then, the attention map is binarized by comparing it with the average value frame by frame, and the discrete information is filtered out using corrosion and expansion operations to generate the mask of the final target token. .

[0010] Preferably, based on the assumption that the motion change between adjacent frames is small, the masks of adjacent frames are combined, expressed as: ; in, Indicates the combined frame mask; , Respectively represent the first Frame mask, Frame mask.

[0011] Preferably, the cross-frame fusion alignment strategy is When sampling a new perspective, the first The sampling results of the frames fill the common new view area, and the remaining areas are sampled individually.

[0012] The noise re-injection mechanism adds noise back into the latent variables at the end of denoising. Convert back , this process can be expressed as: ; in, , Respectively represent Step and The latent variables of the step; represents a set of preset variance schedules, consistent with the diffusion process; represents a set of sampled noise; Indicates that the noise follows a standard Gaussian distribution; Subsequently, the denoising process continues until the latent variable is reached in the last step .

[0013] After adopting the above technical solution, the present invention has the following technical effects: The present invention introduces a latent variable rearrangement strategy in the denoising process, greatly simulates camera actions such as translation, scaling, and rotation, and proposes a mask-based sampling strategy for filling new viewing areas; then, through the attention map in the cross-attention process, the sampling range is automatically limited to avoid repeated sampling of the subject; in addition, in order to solve the time inconsistency phenomenon that may be caused in the sampling process, the present invention introduces a cross-frame fusion alignment mechanism, so that the sampling process between frames can interact and ensure the timing consistency. Therefore, the present invention enables ordinary base models to have camera control capabilities, and does not require additional repair models and depth estimation models, achieving simplicity and lightweight. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is an overall flow chart of a specific embodiment of the present invention.

[0015] Figure 2 This is a diagram explaining the distribution deviation phenomenon of a specific embodiment of the present invention.

[0016] Figure 3 It is a quantitative comparison between the specific embodiments of the present invention and other prior arts.

[0017] Figure 4 It is a qualitative comparison between the specific embodiments of the present invention and other prior arts. DETAILED DESCRIPTION

[0018] In order to further explain the technical solution of the present invention, the present invention is described in detail below through specific embodiments.

[0019] refer to Figure 1 The overall flow chart shown in FIG. 1 discloses a video generation method with camera control that does not require training or display repair, and uses the prior knowledge of the video diffusion model to achieve perspective conversion in the video, avoiding dependence on specific data fine-tuning or complex camera parameters. Figure 1 As shown at the top of the figure, the standard denoising process generates a video with a fixed viewing angle. Therefore, the present invention uses a specific time step in the denoising process. Performs a permutation operation on the latent variables for each frame. Figure 1 As shown in the middle part of , by changing the order of the latent variables, the present invention simulates specific camera actions such as translation, scaling, and rotation; then, the present invention applies a resampling strategy in the latent space to fill the new viewing area, and combines the cross-frame fusion alignment strategy to ensure that the sampling process remains consistent between frames, thereby improving the overall coherence and stability of the generated video. In addition, through detailed experiments, it is found that the rearrangement and resampling operations in the denoising process will cause the distribution of the latent variables to deviate from the receptive field of the original diffusion model, thereby affecting the quality of the generated video. In order to further alleviate this distribution shift phenomenon, refer to Figure 1 As shown at the bottom of the figure, the present invention introduces a noise re-injection mechanism, which re-injects noise into the latent variable in the later stage of denoising, prolongs the denoising time, and gives the model more opportunities to correct the distribution shift, thereby generating a more stable and natural video.

[0020] Through the above scheme, the present invention introduces a latent variable rearrangement strategy in the denoising process, greatly simulates camera actions such as translation, scaling, and rotation, and proposes a mask-based sampling strategy for filling new viewing areas; then, through the attention map in the cross-attention process, the sampling range is automatically limited to avoid the problem of repeated sampling of the subject; in addition, in order to solve the time inconsistency phenomenon that may be caused in the sampling process, the present invention introduces a cross-frame fusion alignment mechanism, so that the sampling process between frames can interact and ensure the timing consistency. Therefore, the present invention enables ordinary base models to have camera control capabilities, and does not require additional repair models and depth estimation models, achieving simplicity and lightweight.

[0021] Specific embodiments of the present invention are shown below.

[0022] The following describes the rearrangement operation of the present invention: The positive direction is defined as right, downward, inward, and counterclockwise, and the user can enter the motion parameters , , (corresponding to the X-axis, Y-axis and Z-axis respectively) and the rotation angle , by combining motion parameters to achieve basic perspective transformations such as translation, scaling, and rotation without complex camera adjustments, thus converting the original latent variables Mapping to new latent variables To achieve perspective conversion (ref. Figure 1 ), where Indicates the number of frames.

[0023] The translation operation is performed by the motion parameters entered by the user. and Update Latent variables of frames For the Frame, the camera is translated along the X axis , translate along the Y axis units, and then the translated latent variables are clipped back to the original position, where represents the width of the latent space, represents the height of the latent space.

[0024] The scaling operation is performed by applying the latent variable Interpolation is applied to simulate the effect of perspective zoom, first using a Gaussian kernel It is processed, interpolated to obtain an intermediate latent variable, and finally its center is cropped to a specific size to obtain the final latent variable. The above process is expressed as: ; ; ; ; in, Represents an interpolation operation; represents the Gaussian blur kernel; Indicates The scaling factor of the frame, whose value is ; Indicates the operation of taking the minimum value; Indicates that the cropping operation is performed according to the preset size.

[0025] The rotation operation introduces the point cloud projection theory into the latent space, which is the latent variable Each position in Generate its 3D point cloud representation, where and Represent the coordinates of any point respectively; then project them to different perspectives according to different parameters: ; ; Among them, the superscript Represents the transpose of a matrix; Represents the intrinsic parameter matrix of the camera; Indicates The rotation matrix of the frame; Indicates the depth information of the corresponding position.

[0026] Above, it is not difficult to prove through simple derivation that, under the assumption of ignoring the translation matrix, the projection results of the 3D point cloud under different viewing angles have nothing to do with the depth information, so it can be ignored. , and the discussion will focus on the following parameters and Top (with Y axis as the rotation axis, the same applies to other cases): .

[0027] Similar to pixel space, the optical center of the camera Usually set in the central region of the latent space However, the focal length in the latent space , It is not the same as the pixel space and has no actual physical meaning. To simplify the operation, the present invention defines In order to perform point cloud projection in latent space. For each frame, the parameters The latent variables are rotated, where Evenly distributed in arrive within the range.

[0028] The resampling strategy of the present invention is introduced below: Since there is semantic similarity between adjacent tokens in the latent space, the tokens in the new perspective area should maintain semantic consistency with the adjacent tokens in the old perspective. The existing resampling strategy is to sample the corresponding tokens for the new perspective area from the random columns of the corresponding rows or the random rows of the corresponding columns. However, this strategy may lead to sampling errors: for example, when generating a video containing the prompt word "a cat in the garden", the new perspective area may mistakenly sample the subject "cat", resulting in problems such as object duplication and semantic confusion. In order to solve this problem, the present invention proposes a sampling method based on fine-grained mask guidance, and combines it with cross-frame alignment technology to improve sampling accuracy and video generation quality. The resampling strategy of the present invention is specifically: Introduce a subject mask during sampling to limit the sampling range and the new viewing area The tokens are no longer sampled from random columns of the same row or random rows of the same column, but from columns of specific rows or rows of specific columns, whose orientations are determined by the mask The background area in is defined as follows: ; ; in, , and Respectively represent the index of any position in the latent variable; Represents the new viewing area for each frame. Row sampling and column sampling will be performed separately to ensure that all new viewing areas are fully filled. At the initial time step In this paper, we store the cross attention map of the last upsampling module, extract the response map corresponding to the target token and reshape it into , Represents the real number space; then, the attention map is binarized by comparing it with the average value frame by frame, and the discrete information is filtered out using corrosion and expansion operations to generate the mask of the final target token. .

[0029] However, the temporal attention mechanism in the video generation framework will cause the cross-attention map to be blurred between frames, which cannot accurately reflect the response relationship between frames. Based on the assumption that the motion change between adjacent frames is small, the present invention combines the masks of adjacent frames and expresses them as: ; in, Indicates the combined frame mask; , Respectively represent the first Frame mask, In this way, the present invention can achieve more accurate mask extraction without the need for the user to manually input a mask for a specific foreground.

[0030] The following introduces the cross-frame fusion alignment strategy of the present invention: In actual videos, there must be a common new viewing angle area between adjacent frames, which contradicts the independent sampling of each frame mentioned above. If the consistency of sampling between frames is not considered, the time inconsistency of the final generated video will be caused, resulting in flickering, artifacts and other phenomena. In order to solve this problem, the present invention needs to ensure consistent sampling in the common new viewing angle area: see Figure 1 , in the When sampling a new viewing angle, the present invention preferentially uses the first The sampling results of the frames fill the common new view area, and the remaining area is sampled separately. This cross-frame fusion sampling alignment method enables the sampling information of the previous frame to be transferred to the common area of ​​the subsequent frames, thereby ensuring consistency between frames.

[0031] The noise re-injection mechanism of the present invention is introduced below: See also Figure 2 , the training-free update operation (i.e., the above-mentioned rearrangement and resampling) will change the overall joint probability distribution of the latent variables, causing it to deviate from the expected distribution range of the diffusion model, ultimately resulting in poor quality of the generated results. Although the subsequent denoising process can correct this deviation to a certain extent, for some complex camera motions involving a large number of new perspectives, this distribution shift phenomenon will be very significant, thus affecting the quality of the final generated video. In order to solve this problem, the present invention introduces a noise re-injection mechanism, referring to Figure 1 As shown at the bottom of , in the later stage of denoising, the noise is added back into the latent variable, Convert back , this process can be expressed as: ; in, , Respectively represent Step and The latent variables of the step; represents a set of preset variance schedules, consistent with the diffusion process; represents a set of sampled noise; Indicates that the noise follows a standard Gaussian distribution; Subsequently, the denoising process continues until the latent variable is reached in the last step Therefore, by appropriately extending the denoising step, the present invention can effectively alleviate the distribution shift caused by the latent space update.

[0032] In order to verify the effectiveness of the method proposed in this invention, a large number of experiments were carried out, and the results are as follows: Figure 3 , 4 As shown in the figure, it can be seen that the method of the present invention is superior to the existing video generation method with camera control in both quantitative and qualitative evaluations. Specifically, the method of the present invention surpasses all other methods in the FVD index and also surpasses all existing methods in the text alignment Clip-Score; in addition, the translation, scaling, and rotation errors of the present invention are smaller than those of all other methods.

[0033] The above embodiments and drawings do not limit the product form and style of the present invention. Any appropriate changes or modifications made by ordinary technicians in the relevant technical field should be deemed to be within the patent scope of the present invention.

Claims

1. A video generation method with camera control without training or display repair, characterized in that: At a specific time step in the denoising process A rearrangement operation is performed on the latent variables of each frame; a specific camera action is simulated by changing the arrangement order of the latent variables; then, a resampling strategy is applied in the latent space to fill the new view area, and a cross-frame fusion alignment strategy is combined to ensure that the sampling process remains consistent between frames; a noise re-injection mechanism is introduced to extend the denoising time and alleviate the distribution shift caused by rearrangement and resampling by reinjecting noise into the latent variables in the later stage of denoising.

2. The video generation method with camera control without training or display repair as claimed in claim 1, characterized in that: The rearrangement operation defines the positive direction as right, downward, inward and counterclockwise rotation, and the user can input the movement parameters. , , and rotation angle , the basic perspective transformation of translation, scaling and rotation is realized through the combination of motion parameters, and the original latent variables Mapping to new latent variables To achieve perspective conversion, Indicates the number of frames; The translation operation is performed by the motion parameters input by the user. and Update Latent variables of frames For the Frame, the camera is translated along the X axis , translate along the Y axis units, and then the translated latent variables are clipped back to the original position, where represents the width of the latent space, represents the height of the latent space; The scaling operation is performed by applying the latent variables of each frame to Interpolation is applied to simulate the effect of perspective zoom, first using a Gaussian kernel It is processed, interpolated to obtain an intermediate latent variable, and finally its center is cropped to a specific size to obtain the final latent variable. The above process is expressed as: ; ; ; ; in, Represents an interpolation operation; represents the Gaussian blur kernel; Indicates The scaling factor of the frame, whose value is ; Indicates the operation of taking the minimum value; Indicates cropping operation according to the preset size; The rotation operation introduces the point cloud projection theory into the latent space, which is the latent variable Each position in Generate its 3D point cloud representation, where and Represent the coordinates of any point respectively; then project them to different perspectives according to different parameters: ; ; Among them, the superscript Represents the transpose of a matrix; Represents the intrinsic parameter matrix of the camera; Indicates The rotation matrix of the frame; Indicates the depth information of the corresponding position; parameter and The calculation formula is as follows: ; For each frame, the parameters The latent variables are rotated, where Evenly distributed in arrive within the range.

3. The video generation method with camera control without training or display repair as claimed in claim 1, characterized in that: The resampling strategy specifically introduces a subject mask during the sampling process to limit the sampling range, and the new viewing area The token is sampled from a column of a specific row or a row of a specific column, and the orientation of these rows or columns is determined by the mask The background area in is defined as follows: ; ; in, , and Respectively represent the index of any position in the latent variable; Represents the new viewing area for each frame. Row sampling and column sampling will be performed separately to ensure that all new viewing areas are fully filled. At the initial time step In , the cross attention map of the last upsampling module is stored, and the response map corresponding to the target token is extracted and reshaped into , Represents the real number space; then, the attention map is binarized by comparing it with the average value frame by frame, and the discrete information is filtered out using corrosion and expansion operations to generate the mask of the final target token. .

4. The video generation method with camera control without training or display repair as claimed in claim 3, characterized in that: Based on the assumption that the motion changes between adjacent frames are small, the masks of adjacent frames are combined and expressed as: ; in, Indicates the combined frame mask; , Respectively represent the first Frame mask, Frame mask.

5. The video generation method with camera control without training or display repair as claimed in claim 4, characterized in that: The cross-frame fusion alignment strategy is When sampling a new perspective, the first The sampling results of the frames fill the common new view area, and the remaining areas are sampled individually.

6. The video generation method with camera control without training or display repair as claimed in claim 1, characterized in that: The noise re-injection mechanism adds noise back into the latent variables at the end of denoising. Convert back , this process can be expressed as: ; in, , Respectively represent Step and The latent variables of the step; represents a set of preset variance schedules, consistent with the diffusion process; represents a set of sampled noise; Indicates that the noise follows a standard Gaussian distribution; Subsequently, the denoising process continues until the latent variable is reached in the last step .