Video generation method and device with controllable mirror operation mode, and electronic equipment

By constructing a joint training mechanism between a camera movement information encoder and a video generation model, the problems of inaccurate camera movement control and high computational overhead in existing video generation technologies are solved, achieving high-quality and flexible video generation.

CN121099153AActive Publication Date: 2025-12-09CHENGDU SOBEY DIGITAL TECH CO LTD

Patent Information

Application Number
CN202511136446.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-12-09
Estimated Expiration
2045-08-14

AI Technical Summary

Technical Problem

Existing video generation technologies lack precise camera control and high-quality methods when generating video content. Traditional methods are difficult to adapt to diverse video generation needs and have high computational costs.

Method used

A camera movement information encoder and a video generation model are jointly trained. Camera movement videos are generated by generating time-series camera pose parameters and text descriptions, and precise camera movement control is achieved using deep neural networks.

Benefits of technology

It enables video generation with controllable camera movement, improves the accuracy and quality of video generation, meets diverse video generation needs, and reduces computational overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121099153A_ABST
    Figure CN121099153A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video generation method and device with a controllable mirror operation mode and electronic equipment, and relates to the technical field of video generation and deep learning, and the method comprises the steps: generating a camera pose parameter of a time sequence according to a selected mirror operation mode; the camera pose parameters are preprocessed and input into a mirror moving information encoder to obtain mirror moving information encoding features, and finally the mirror moving information encoding features, the pictures and the text description are input into a video generation model to generate a corresponding mirror moving video. According to the method, a mirror moving information encoder is constructed, and the mirror moving information encoder and a video generation model are jointly trained, so that the original video generation model can accept mirror moving information, video generation with a controllable mirror moving mode is realized, creativity of the video generation model is fully utilized, and accurate mirror moving control is provided.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video generation and deep learning, in particular to a video generation method with controllable camera movement, a video generation device and an electronic device. BACKGROUND

[0002] With the rapid development of video production technology, video generation technology has become an important research direction in the field of artificial intelligence and computer vision. Traditional video production methods mainly rely on high-quality pre-recorded materials and professional video editing techniques. In recent years, deep learning-based generation technology has made it possible to automate and intelligently generate videos. This technology has broad application prospects in fields such as advertising production, virtual reality, game development, and film special effects.

[0003] However, current video generation technology faces many challenges when generating video content. Among them, the controllability of camera movement is particularly prominent. Camera movement refers to adjusting the camera's angle, position, or movement trajectory during video shooting or generation to achieve a specific visual effect. Reasonable camera movement design can enhance the narrative and aesthetic qualities of a video, but achieving precise and flexible camera movement control is a technical challenge.

[0004] Traditional methods of producing camera movement videos mainly include: rule-based camera control, which involves pre-setting fixed trajectories or angle change rules to crop, pan, or translate videos or images, but lacks flexibility and is difficult to adapt to diverse video generation needs; camera movement generation based on physical simulation, which uses physical engines to simulate camera movement to generate natural movement trajectories, but has high computational overhead and requires high scene complexity and real-time performance.

[0005] With the rapid development of diffusion models and video generation methods, some methods attempt to use deep neural networks to learn the features of camera movement in videos, thereby controlling the camera movement of generated videos. Current methods achieve this capability by simply influencing the attention features of video generation models or fine-tuning, but their accuracy of camera movement trajectories and video quality do not meet the needs of practical use.

[0006] In summary, there is currently a lack of a video generation method that can simultaneously satisfy precise camera movement control and high generation quality in the field of video generation technology. Therefore, developing a video generation method with controllable camera movement can not only make up for the shortcomings of existing technology, but also promote the development of related application fields. SUMMARY

[0007] Embodiments of the present application provide a video generation method with controllable camera movement, a video generation device, and an electronic device to solve the technical problems existing in the prior art.

[0008] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.

[0009] According to a first aspect of the embodiments of this application, a video generation method with controllable camera movement is provided, including: Generate timing-series camera pose parameters based on the selected camera movement method; After preprocessing the camera pose parameters, they are input into the camera motion information encoder for encoding to obtain the camera motion information encoded features; The camera movement information encoding features, images, and text descriptions are input into the video generation model to generate a camera movement video. The camera movement information encoder includes a feature compression module, a feature serialization module, and a feature extraction module. The video generation model includes a text feature extractor, a video encoder, a video decoder, and a video denoising model.

[0010] In some embodiments of this application, based on the foregoing scheme, the step of generating temporal camera pose parameters according to the selected camera movement method includes: Select the preset camera movement method, video frame rate, video resolution, and reference camera movement video; New camera pose parameters are sampled from the camera pose parameters of the corresponding camera movement method based on the video frame rate and video resolution. The SFM method is used to process the reference camera movement video to estimate the camera pose parameters for each frame of the video.

[0011] In some embodiments of this application, based on the foregoing scheme, the step of using the SFM method to process the reference camera movement video and estimate the camera pose parameters for each frame of the video includes: Key points are extracted from each frame of the reference camera movement video using a feature detection algorithm; Find the correspondence between key points in different images and perform matching; The matching keypoint pairs are used to compute the fundamental matrix F, which describes the geometric constraints between the two images; The fundamental matrix F is transformed into an essential matrix E, which contains the relative rotation and translation information between the cameras in the two frames of images; The camera's rotation matrix R and translation vector t are decomposed from the essential matrix E.

[0012] In some embodiments of this application, based on the foregoing scheme, the step of preprocessing the camera pose parameters and then encoding them to obtain camera movement information encoding features includes: The obtained camera pose parameters are preprocessed to obtain the Planck coordinates of each pixel in the image; Input the Planck coordinates into the camera motion information encoder to obtain the camera motion information encoded features.

[0013] In some embodiments of this application, based on the foregoing scheme, the preprocessing of the obtained camera pose parameters includes: Based on the camera pose parameters, calculate the ray origin and direction from the camera optical center to the pixel in each frame of the image; The Planck coordinates of each pixel in each frame are calculated based on the origin and direction of the ray from the camera's optical center to the image pixel.

[0014] In some embodiments of this application, based on the foregoing scheme, generating a camera movement video based on the camera movement information encoding features, images, and text descriptions includes: Randomly initialize a Gaussian noise as a latent space feature; The input image is processed using a video encoder to obtain image features, and then zero-padding is performed to the same length as the video to obtain the first frame image features. The first frame image features and latent space features are concatenated in the channel dimension to obtain the final latent space features; Text features are obtained by processing text descriptions using a text feature extractor. The text features and the final latent space features are directly input into the video denoising model, while the camera movement information encoding features are added to the corresponding intermediate features to predict the noise. The predicted noise is subtracted from the final latent space features by the forward sampler to obtain new latent space features. The new latent space features are then iterated and denoised repeatedly until the final denoised result is obtained. The video decoder is used to process the final denoising result to obtain the video corresponding to the camera movement.

[0015] In some embodiments of this application, based on the foregoing scheme, before encoding camera movement information, the method further includes: constructing a camera movement information encoder and a video generation model, and performing joint training, including: A large amount of video data is acquired through camera shooting and network collection, and is divided into two categories: static shots and moving shots. Use the SFM method to estimate camera parameters for videos without camera parameters; The video description model is used to process video data to obtain a text description for each video; Construct a camera motion information encoder and a video generation model, and train the camera motion information encoder using video data of camera motion, text descriptions, and camera parameters; The video generation model is fine-tuned using still video data, text descriptions, and camera parameters.

[0016] In some embodiments of this application, based on the foregoing scheme, the training of the camera movement information encoder using video data of camera motion, text description, and camera parameters includes: Randomly initialize the camera movement information encoder; Configure the optimizer, learning rate, decay rate, training batches, and number of training iterations; The camera parameters are processed into Planck coordinates and input into the camera motion information encoder. The features obtained from each block in the feature extraction module are saved as camera motion information encoded features. The video data is processed using a video encoder to obtain latent space features. The current denoising time step is randomized, and Gaussian noise of corresponding intensity is added to the latent space features to obtain the noisy latent space features. The first frame features of the noisy latent space features are zero-padding to the same length as the video to obtain the first frame image features. The first frame image features and the noisy latent space features are then concatenated along the channel dimension to obtain the final latent space features. Text features are obtained by processing text descriptions using a text feature extractor. Text features and final latent space features are directly input into the video denoising model, while the camera movement information encoding features and corresponding intermediate features are added together to predict the noise. The denoised latent space features are obtained by subtracting the predicted noise from the noisy latent space features using a forward sampler. The L2 loss and temporal loss between the denoised latent space features and the latent space features are calculated, and backpropagation is performed to optimize the camera movement information encoder.

[0017] According to a second aspect of the embodiments of this application, a video generation apparatus with controllable camera movement is provided, comprising: The first generation unit is used to generate the timing camera pose parameters according to the selected camera movement method; The encoding unit is used to preprocess the camera pose parameters and then input them into the camera motion information encoder for encoding to obtain camera motion information encoding features; The second generation unit is used to input the camera movement information encoding features, images and text descriptions into the video generation model to generate a camera movement video. The camera movement information encoder includes a feature compression module, a feature serialization module, and a feature extraction module. The video generation model includes a text feature extractor, a video encoder, a video decoder, and a video denoising model.

[0018] According to a third aspect of the embodiments of this application, an electronic device is provided, including: a memory and a processor; The memory is used to store computer instructions; The processor is configured to invoke computer instructions stored in the memory, causing the electronic device to execute the method described in the first aspect.

[0019] According to a fourth aspect of the embodiments of this application, a computer-readable storage medium is provided, the storage medium storing computer instructions that, when executed on a computer, cause the computer to perform the method as described in the first aspect.

[0020] The technical solution of this application first generates temporal camera pose parameters based on the selected camera movement method. Then, the camera pose parameters are preprocessed and input into a camera movement information encoder to obtain camera movement information encoded features. Finally, the camera movement information encoded features, images, and text descriptions are input into a video generation model to generate the corresponding camera movement video. This method constructs a camera movement information encoder and jointly trains it with the video generation model, enabling the original video generation model to accept camera movement images. This achieves video generation with controllable camera movement, fully utilizing the creativity of the video generation model and providing precise camera movement control.

[0021] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0022] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings: Figure 1 A schematic flowchart of a video generation method with controllable camera movement according to an embodiment of this application is shown; Figure 2 A schematic diagram of an SFM camera parameter estimation process according to an embodiment of this application is shown; Figure 3 A schematic diagram of the Planck coordinate calculation process according to an embodiment of this application is shown; Figure 4 A schematic diagram of a specific training method for a camera movement information encoder according to an embodiment of this application is shown; Figure 5 A schematic diagram of the process for generating camera movement video according to an embodiment of this application is shown; Figure 6 This illustration shows a schematic diagram of the effect generated by camera movement video according to an embodiment of this application; Figure 7A block diagram of a camera movement controllable video generation apparatus according to an embodiment of this application is shown; Figure 8 A block diagram of an electronic device according to one embodiment of this application is shown; Figure 9 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown. Detailed Implementation

[0023] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.

[0024] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0025] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0026] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such uses of these terms can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described.

[0028] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0029] The following detailed description of some embodiments of this application will be provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0030] In existing technologies, rule-based camera control or physics-based camera generation are commonly used. These methods lack flexibility, struggle to adapt to diverse video generation needs, and incur significant computational overhead. With the rapid development of diffusion models and video generation methods, some approaches attempt to utilize deep neural networks to learn the characteristics of camera motion in videos, thereby controlling the camera movement of the generated videos. However, the accuracy of the camera trajectory and the video quality still fall short of practical application requirements.

[0031] Therefore, in order to solve the above problems, this application proposes a video generation method with controllable camera movement. This method constructs a camera movement information encoder and performs joint training with a video generation model, so that the original video generation model can accept camera movement information, thereby realizing video generation with controllable camera movement. It makes full use of the creativity of the video generation model and provides powerful camera movement control, which has higher practical value. The following is a detailed description of the method.

[0032] See Figure 1 The diagram shows a flow chart of a video generation method with controllable camera movement according to an embodiment of this application.

[0033] like Figure 1 As shown, a video generation method with controllable camera movement is demonstrated, specifically including steps S100 to S300.

[0034] See Figure 1 Step S100: Generate timing camera pose parameters according to the selected camera movement method.

[0035] In some feasible embodiments, based on the foregoing scheme, step S100 includes: Step S110: Select the preset camera movement method, video frame rate, video resolution, and reference camera movement video; Step S120: Sample new camera pose parameters from the camera pose parameters of the corresponding camera movement method based on the video frame number and video resolution; Step S130: Use the SFM method to process the reference camera movement video and estimate the camera pose parameters for each frame of the video.

[0036] It should be noted that SFM stands for Structure from Motion, a core technology in the fields of image processing, computer vision, and photogrammetry. Its goal is to automatically recover the three-dimensional structure (point cloud) of a scene and the camera's motion trajectory from a series of two-dimensional images with overlapping areas.

[0037] In some feasible embodiments, based on the foregoing scheme, the process of step S130 is as follows: Figure 2 ,include: Step S131: Use a feature detection algorithm to extract key points from each frame of the reference camera movement video.

[0038] Step S132: Find the correspondence between key points in different images and perform matching.

[0039] Step S133: Calculate the fundamental matrix F using the matched keypoint pairs, whereby the fundamental matrix F describes the geometric constraint relationship between the two images; The calculation formula is as follows: ; in, and This represents the pixel coordinates of the matching keypoints in both views. It is the fundamental matrix that needs to be solved.

[0040] Step S134: Convert the basic matrix F into the essential matrix E, which contains the relative rotation and translation information between the cameras in the two frames of images. The calculation formula is as follows: ; in, This represents the known camera intrinsic parameter matrix. It is the obtained fundamental matrix. It is an essential matrix.

[0041] Step S135: Decompose the camera's rotation matrix R and translation vector t from the essential matrix E.

[0042] See also Figure 1 In step S200, the camera pose parameters are preprocessed and then input into the camera motion information encoder for encoding to obtain camera motion information encoded features; wherein, the camera motion information encoder includes a feature compression module, a feature serialization module and a feature extraction module.

[0043] It should be noted that in this embodiment, the feature compression module uses a spatial and temporal rearrangement method to stack spatial and temporal information into the channel, compressing the data size without losing information; the feature serialization module processes camera movement information into a token sequence through convolution and linear layers; the feature extraction module consists of 10 Transformers blocks, each containing a 3D full attention layer.

[0044] In some feasible embodiments, based on the foregoing scheme, step S200 includes: Step S210: Preprocess the obtained camera pose parameters to obtain the Planck coordinates of each pixel in the image; Step S220: Input the Planck coordinates into the camera motion information encoder to obtain the camera motion information encoded features.

[0045] In some feasible embodiments, based on the foregoing scheme, the process of step S210 is as follows: Figure 3 ,include: Step S211: Calculate the ray origin and direction from the camera optical center to the pixel in each frame of the image based on the camera pose parameters. The calculation formula is as follows: ; ; in, , , Let these represent the camera rotation matrix, intrinsic parameter matrix, and translation vector for the f-th frame, respectively. Indicates the first On the frame ( The direction of the ray at the coordinates, Indicates the first The starting point of the ray on the frame.

[0046] Step S212: Calculate the Planck coordinates of each pixel in each frame of the image based on the starting point and direction of the ray from the camera's optical center to the image pixel. The calculation formula is as follows: ; ; in, Indicates the first On the frame ( The direction of the normal vector of the plane containing the ray at coordinate ) Indicates the first On the frame ( Planck coordinates at the coordinates.

[0047] In some feasible embodiments, based on the aforementioned scheme, before encoding camera movement information, the method further includes: constructing a camera movement information encoder and a video generation model, and performing joint training, including: Step A10 involves acquiring a large amount of video data through camera shooting and network collection, and dividing it into two categories: still shots and moving shots. Step A20: Use the SFM method to estimate camera parameters for video without camera parameters; Step A30: Process the video data using the video description model to obtain a text description for each video; Step A40: Construct a camera motion information encoder and a video generation model. Train the camera motion information encoder using video data of camera motion, text descriptions, and camera parameters. Step A50: Using still video data, text descriptions, and camera parameters, fine-tune the video generation model to improve the dynamism of the generated video.

[0048] In some feasible embodiments, based on the foregoing scheme, the process of training the camera movement information encoder using video data of camera motion, text descriptions, and camera parameters is described in [link to documentation]. Figure 4 Specifically, it includes: Step A41: Randomly initialize the camera motion information encoder, initialize the video generation model from the existing model weights and freeze its weights, and optimize only the camera motion information encoder during training.

[0049] Step A42: Set hyperparameters such as optimizer, learning rate, decay rate, training batches, and number of training iterations.

[0050] Step A43: Process the camera parameters into Planck coordinates, input them into the camera motion information encoder, and save the features obtained from each block in the feature extraction module as camera motion information encoded features.

[0051] Step A44: Use a video encoder to process video data to obtain latent space features, randomize the current denoising time step, and add Gaussian noise of corresponding intensity to the latent space features to obtain the noisy latent space features.

[0052] Step A45: Zero-padding the first frame features of the latent space features to the same length as the video to obtain the first frame image features, and then concatenating the channel dimension and the noisy latent space features to obtain the final latent space features.

[0053] Step A46: Use a text feature extractor to process the text description to obtain text features.

[0054] Step A47 involves directly inputting the text features and the final latent space features into the video denoising model, while the camera movement information encoding features and the corresponding intermediate features are added together to predict the noise.

[0055] Step A48: Subtract the predicted noise from the noisy latent space features using a forward sampler to obtain the denoised latent space features.

[0056] Step A49: Calculate the L2 loss and temporal loss between the denoised latent space features and perform backpropagation to optimize the camera movement information encoder.

[0057] See also Figure 1 In step S300, the camera movement information encoding features, images, and text descriptions are input into the video generation model to generate a camera movement video; wherein, the video generation model includes a text feature extractor, a video encoder, a video decoder, and a video denoising model.

[0058] In some feasible embodiments, based on the foregoing scheme, the process of step S300 is as follows: Figure 5 Specifically, it includes: Step S310: Randomly initialize a Gaussian noise as a latent space feature.

[0059] Step S320: The input image is processed using a video encoder to obtain image features, and then zero-padding is performed to the same length as the video to obtain the first frame image features.

[0060] Step S330: The features of the first frame image and the latent space features are concatenated in the channel dimension to obtain the final latent space features.

[0061] Step S340: Use a text feature extractor to process the text description to obtain text features.

[0062] In step S350, the text features and the final latent space features are directly input into the video denoising model, while the camera movement information encoding features are added to the corresponding intermediate features to predict the noise.

[0063] In step S360, the predicted noise is subtracted from the latent space features by the forward sampler to obtain new latent space features. The new latent space features are then used as latent space features to repeat the processing of steps S320 to S350 above. The process is iterated multiple times to obtain the final denoising result.

[0064] Step S370: Use a video decoder to process the final denoising result to obtain a video with the corresponding camera movement. The video generation effect is as follows: Figure 6 As shown.

[0065] The following describes an embodiment of the apparatus described in this application, which can be used to execute a video generation method with controllable camera movement as described in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0066] Reference Figure 7 As shown, a video generation apparatus 700 with controllable camera movement according to an embodiment of this application includes: The first generation unit 701 is used to generate the timing camera pose parameters according to the selected camera movement method; The encoding unit 702 is used to preprocess the camera pose parameters and then input them to the camera motion information encoder for encoding to obtain camera motion information encoding features. The second generation unit 703 is used to input the camera movement information encoding features, images and text descriptions into the video generation model to generate a camera movement video. The camera movement information encoder includes a feature compression module, a feature serialization module, and a feature extraction module. The video generation model includes a text feature extractor, a video encoder, a video decoder, and a video denoising model.

[0067] like Figure 8 As shown, this application embodiment also provides an electronic device 800, including a memory 810, a processor 820, and a computer program 811 stored in the memory 810 and executable on the processor. When the processor 820 executes the computer program 811, it implements the steps of the above-mentioned video generation method with controllable camera movement.

[0068] Since the electronic device described in this embodiment is the device used to implement a video generation device with controllable camera movement in the embodiments of this application, those skilled in the art can understand the specific implementation method and various variations of the electronic device in this embodiment based on the method described in the embodiments of this application. Therefore, how the electronic device implements the method in the embodiments of this application will not be described in detail here. Any device used by those skilled in the art to implement the method in the embodiments of this application falls within the scope of protection of this application.

[0069] In practice, when the computer program 811 is executed by the processor, it can implement any of the embodiments corresponding to the first aspect.

[0070] Figure 9 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown.

[0071] It should be noted that, Figure 9 The computer system 900 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0072] like Figure 9As shown, the computer system 900 includes a Central Processing Unit (CPU) 901, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 902 or programs loaded from storage portion 908 into Random Access Memory (RAM) 903, such as performing the methods described in the above embodiments. The RAM 903 also stores various programs and data required for system operation. The CPU 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0073] The following components are connected to I / O interface 905: an input section 906 including a keyboard, mouse, etc.; an output section 907 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to I / O interface 905 as needed. Removable media 911, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 910 as needed so that computer programs read from them can be installed into storage section 908 as needed.

[0074] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 909, and / or installed from removable medium 911. When the computer program is executed by central processing unit (CPU) 901, it performs various functions defined in the system of this application.

[0075] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0076] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0077] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0078] In another aspect, this application also provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform a video generation method with controllable camera movement as described in the above embodiments.

[0079] In another aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to implement the video generation method with controllable camera movement described in the above embodiments.

[0080] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0081] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the methods according to the embodiments of this application.

[0082] Other embodiments of this application will readily conceive of by those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. It should be understood that this application is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A video generation method with controllable camera movement, characterized in that, include: Generate timing-series camera pose parameters based on the selected camera movement method; After preprocessing the camera pose parameters, they are input into the camera motion information encoder for encoding to obtain the camera motion information encoded features; The camera movement information encoding features, images, and text descriptions are input into the video generation model to generate a camera movement video. The camera movement information encoder includes a feature compression module, a feature serialization module, and a feature extraction module. The video generation model includes a text feature extractor, a video encoder, a video decoder, and a video denoising model.

2. The method according to claim 1, characterized in that, The generation of time-series camera pose parameters based on the selected camera movement method includes: Select the preset camera movement method, video frame rate, video resolution, and reference camera movement video; New camera pose parameters are sampled from the camera pose parameters of the corresponding camera movement method based on the video frame rate and video resolution. The SFM method is used to process the reference camera movement video to estimate the camera pose parameters for each frame of the video.

3. The method according to claim 2, characterized in that, The process of using the SFM method to process the reference camera movement video and estimate the camera pose parameters for each frame of the video includes: Key points are extracted from each frame of the reference camera movement video using a feature detection algorithm; Find the correspondence between key points in different images and perform matching; The matching keypoint pairs are used to compute the fundamental matrix F, which describes the geometric constraints between the two images; The fundamental matrix F is transformed into an essential matrix E, which contains the relative rotation and translation information between the cameras in the two frames of images; The camera's rotation matrix R and translation vector t are decomposed from the essential matrix E.

4. The method according to claim 1, characterized in that, The step of preprocessing and encoding the camera pose parameters to obtain camera movement information encoding features includes: The obtained camera pose parameters are preprocessed to obtain the Planck coordinates of each pixel in the image; Input the Planck coordinates into the camera motion information encoder to obtain the camera motion information encoded features.

5. The method according to claim 4, characterized in that, The preprocessing of the obtained camera pose parameters includes: Based on the camera pose parameters, calculate the ray origin and direction from the camera optical center to the pixel in each frame of the image; The Planck coordinates of each pixel in each frame are calculated based on the origin and direction of the ray from the camera's optical center to the image pixel.

6. The method according to claim 1, characterized in that, The process of generating a camera movement video based on the encoded features of the camera movement information, images, and text descriptions includes: Randomly initialize a Gaussian noise as a latent space feature; The input image is processed using a video encoder to obtain image features, and then zero-padding is performed to the same length as the video to obtain the first frame image features. The first frame image features and latent space features are concatenated in the channel dimension to obtain the final latent space features; Text features are obtained by processing text descriptions using a text feature extractor. The text features and the final latent space features are directly input into the video denoising model, while the camera movement information encoding features are added to the corresponding intermediate features to predict the noise. The predicted noise is subtracted from the final latent space features by the forward sampler to obtain new latent space features. The new latent space features are then iterated and denoised repeatedly until the final denoised result is obtained. The video decoder is used to process the final denoising result to obtain the video corresponding to the camera movement.

7. The method according to claim 1, characterized in that, Before encoding camera movement information, the process also includes: constructing a camera movement information encoder and a video generation model, and jointly training them, including: A large amount of video data is acquired through camera shooting and network collection, and is divided into two categories: static shots and moving shots. Use the SFM method to estimate camera parameters for videos without camera parameters; The video description model is used to process video data to obtain a text description for each video; Construct a camera motion information encoder and a video generation model, and train the camera motion information encoder using video data of camera motion, text descriptions, and camera parameters; The video generation model is fine-tuned using still video data, text descriptions, and camera parameters.

8. The method according to claim 7, characterized in that, The training of the camera movement information encoder using video data of camera motion, text descriptions, and camera parameters includes: Randomly initialize the camera movement information encoder; Configure the optimizer, learning rate, decay rate, training batches, and number of training iterations; The camera parameters are processed into Planck coordinates and input into the camera motion information encoder. The features obtained from each block in the feature extraction module are saved as camera motion information encoded features. The video data is processed using a video encoder to obtain latent space features. The current denoising time step is randomized, and Gaussian noise of corresponding intensity is added to the latent space features to obtain the noisy latent space features. The first frame features of the noisy latent space features are zero-padding to the same length as the video to obtain the first frame image features. The first frame image features and the noisy latent space features are then concatenated along the channel dimension to obtain the final latent space features. Text features are obtained by processing text descriptions using a text feature extractor. Text features and final latent space features are directly input into the video denoising model, while the camera movement information encoding features and corresponding intermediate features are added together to predict the noise. The denoised latent space features are obtained by subtracting the predicted noise from the noisy latent space features using a forward sampler. The L2 loss and temporal loss between the denoised latent space features and the latent space features are calculated, and backpropagation is performed to optimize the camera movement information encoder.

9. A video generation device with controllable camera movement, characterized in that, include: The first generation unit is used to generate the timing camera pose parameters according to the selected camera movement method; The encoding unit is used to preprocess the camera pose parameters and then input them into the camera motion information encoder for encoding to obtain camera motion information encoding features; The second generation unit is used to input the camera movement information encoding features, images and text descriptions into the video generation model to generate a camera movement video. The camera movement information encoder includes a feature compression module, a feature serialization module, and a feature extraction module. The video generation model includes a text feature extractor, a video encoder, a video decoder, and a video denoising model.

10. An electronic device, characterized in that, include: Memory and processor; The memory is used to store computer instructions; The processor is configured to invoke computer instructions stored in the memory, causing the electronic device to perform the method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Video processing method and device, electronic equipment and storage medium

    CN116233534A

  • Video abstract generation method based on motion information assistance

    CN116233569A

  • Mirror operation control method, device and equipment and storage medium

    CN118158340A

  • Method for optimizing robot positioning and map building based on YOLOv5s and key frame selection

    CN118762055A

  • Video generation method and device, medium, electronic equipment and program product

    CN119364142A

Cited By

  • Motion decoupling control method and system based on motion mirror big data

    CN121888097A