Video generation method and apparatus, medium, electronic device and program product

By acquiring camera parameters from video frames and using a temporal attention mechanism encoder to generate target camera parameter features, the problem of users shooting high-quality motion videos is solved, and automated control of camera movement is achieved to generate high-quality motion videos.

WO2026092380A1PCT designated stage Publication Date: 2026-05-07BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2025-10-27
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Users lack professional shooting skills and find it difficult to shoot high-quality video movements.

Method used

By acquiring camera parameters from video frames, a temporal attention mechanism encoder is used to generate target camera parameter features, thereby enabling automated control of camera movement to generate high-quality motion video.

Benefits of technology

It enables the generation of high-quality camera movement videos without requiring professional skills, improving the automation and accuracy of video generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025130222_07052026_PF_FP_ABST
    Figure CN2025130222_07052026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a video generation method and apparatus, a medium, an electronic device and a program product. The method comprises: acquiring a first video and a target camera movement type, the target camera movement type carrying camera parameters that correspond to video frames in the first video; on the basis of the camera parameters corresponding to the video frames in the first video, determining a target camera parameter feature, the target camera parameter feature being used for representing the camera parameters that respectively correspond to the video frames in the first video and the time correlation of the camera parameters of different video frames; and on the basis of the first video and the target camera parameter feature, generating a second video involving a camera movement that corresponds to the target camera movement type. The present application not only considers the camera parameters corresponding to video frames in the first video, but also considers the time correlation of the camera parameters of different video frames in the first video, thus accurately controlling the viewing angle of a camera during video generation, and achieving automatic generation of camera movement videos having high-quality camera movement effects.
Need to check novelty before this filing date? Find Prior Art

Description

Video generation methods, apparatus, media, electronic devices and software products

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese Patent Application No. 202411525491.1, filed on October 29, 2024, the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0003] This disclosure relates to a video generation method, apparatus, medium, electronic device, and program product. Background Technology

[0004] In films, television shows, and short videos, camera movements such as rotation and zoom-in are common. These techniques require users to handhold the camera and control its movement to track the subject, creating the effect. This shooting method, often simply called camera movement, demands professional shooting skills to effectively control the speed and stability of the camera movement, resulting in high-quality video with smooth camera movements. For users without these skills, capturing high-quality video with smooth camera movements is extremely difficult. Summary of the Invention

[0005] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0006] In a first aspect, this disclosure provides a video generation method, including:

[0007] Obtain a first video and a target camera movement type, wherein the target camera movement type carries camera parameters corresponding to video frames in the first video;

[0008] Based on the camera parameters corresponding to the video frames in the first video, target camera parameter features are determined. The target camera parameter features are used to characterize the camera parameters corresponding to the video frames in the first video and the temporal correlation of camera parameters between different video frames.

[0009] Based on the first video and the target camera parameter features, a second video with a camera movement corresponding to the target camera movement type is generated.

[0010] Secondly, this disclosure provides a video generation apparatus, comprising:

[0011] The acquisition module is configured to acquire a first video and a target camera movement type, wherein the target camera movement type carries camera parameters corresponding to video frames in the first video;

[0012] The determining module is configured to determine target camera parameter features based on the camera parameters corresponding to the video frames in the first video. The target camera parameter features are used to characterize the camera parameters corresponding to the video frames in the first video and the temporal correlation of camera parameters between different video frames.

[0013] The generation module is configured to generate a second video with a camera movement corresponding to the target camera movement type, based on the first video and the target camera parameter features.

[0014] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect.

[0015] Fourthly, this disclosure provides an electronic device, comprising:

[0016] A storage device on which computer programs are stored;

[0017] A processing device for executing the computer program in the storage device to implement the steps of the method in the first aspect.

[0018] Fifthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.

[0019] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description

[0020] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:

[0021] Figure 1 is a flowchart illustrating a video generation method according to an exemplary embodiment of the present disclosure.

[0022] Figure 2 is a schematic diagram illustrating the training process of a video generation model according to an exemplary embodiment of the present disclosure.

[0023] Figure 3 is a block diagram illustrating a video generation apparatus according to an exemplary embodiment of the present disclosure.

[0024] Figure 4 is a schematic diagram of the structure of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation

[0025] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0026] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0027] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0028] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0029] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0030] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0031] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0032] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0033] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0034] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0035] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0036] In related technologies, professional shooting skills are required to effectively control the speed and stability of camera movement, thereby capturing high-quality video footage. For users without professional shooting skills, it is difficult to produce high-quality video footage with impressive camera movements.

[0037] In view of this, the present disclosure provides a video generation method, apparatus, medium, electronic device and program product, which precisely controls the camera's angle during the video generation process and realizes the automated generation of videos with camera movement effects.

[0038] The embodiments of this disclosure will be further explained and described below with reference to the accompanying drawings.

[0039] Figure 1 is a flowchart illustrating a video generation method according to an exemplary embodiment of the present disclosure. This video generation method can be applied to an electronic device and can be executed by a video generation apparatus, which can be implemented in software and / or hardware and is configurable in the electronic device. Referring to Figure 1, the video generation method may include the following steps:

[0040] Step 110: Obtain the first video and the target camera movement type. The target camera movement type carries the camera parameters corresponding to the video frames in the first video.

[0041] The first video can be a video containing a preset number of frames, which is obtained by preprocessing a video or image uploaded by the user.

[0042] The target camera movement type can be specified by the user. After the user specifies the target camera movement type, the electronic device can generate camera parameters corresponding to the video frames in the first video according to the target camera movement type. The camera parameter sequence consisting of the camera parameters corresponding to all video frames in the first video is used to describe the attitude information of the camera during movement.

[0043] The target camera movement type can be specified from the following candidate camera movement types. Incorporating candidate camera movement types during video generation can achieve camera movement effects such as rotation, scaling, and panning of the image within the shot. The following describes the playback effects of videos with corresponding camera movements for different candidate camera movement types.

[0044] 1) Push-in shot type: The camera movement corresponding to the push-in shot type can be understood as the camera moving away during the shooting process. The camera movement corresponding to the push-in shot type can present the process of gradually changing from a large scene to a close-up of a local scene. That is, the subject of the video can present a dynamic change process from small to large.

[0045] 2) Pull-up shot type, which is the opposite of the push-up shot type, involves the camera zooming in during the shooting process. The pull-up shot type can present the process of gradually changing from a close-up to a wide shot, and the subject being photographed can present a dynamic change process from large to small.

[0046] 3) Rotation type: This refers to the process in which the camera rotates around a certain axis during the shooting process, and the subject being photographed rotates and changes in the frame.

[0047] Camera parameters can include at least one of the camera's extrinsic and intrinsic parameters. The extrinsic parameters describe the camera's position and orientation in the world coordinate system and are used to control the camera's attitude. The intrinsic parameters describe the camera's inherent characteristics and can be parameters such as focal length and principal point. It is understandable that when multiple camera parameters are included, it is possible to superimpose camera movements, for example, the subject being photographed may dynamically change from large to small while simultaneously undergoing rotation.

[0048] Step 120: Determine the target camera parameter features based on the camera parameters corresponding to the video frames in the first video. The target camera parameter features are used to characterize the camera parameters corresponding to the video frames in the first video and the temporal correlation of camera parameters between different video frames.

[0049] It should be understood that target camera parameter features are coded representations of the camera parameters corresponding to video frames.

[0050] Step 130: Generate a second video with a camera movement corresponding to the target camera movement type based on the first video and the target camera parameter features.

[0051] It should be understood that the second video can be obtained by adding a camera movement that corresponds to the target camera movement type to the first video.

[0052] Through the above technical solution, in the video generation process, not only the camera parameters corresponding to the video frames in the first video are considered, but also the temporal correlation of camera parameters between different video frames in the second video are considered. This allows for precise control of the camera's perspective during video generation, achieving automated generation of high-quality camera movement videos.

[0053] In some embodiments, the aforementioned video frame may be each video frame in the first video; in other embodiments, the aforementioned video frame may be a portion of the video frames in the first video. This embodiment does not impose any limitations on this.

[0054] In some embodiments, the step of determining the target camera parameter features based on the camera parameters corresponding to the video frames in the first video can be implemented in the following manner: determining the camera parameter embedding representation based on the camera parameters corresponding to the video frames in the first video; encoding the camera parameter embedding representation by an encoder based on a temporal attention mechanism to obtain the target camera parameter features.

[0055] The camera parameter embedding representation can be characterized by providing a geometric interpretation of each pixel in the first video in three-dimensional space. For example, the camera parameter embedding representation can be represented by a vector B*F*6*H*W, where B represents the number of the first video (i.e., B is 1 in this embodiment), F represents the number of video frames in the second video, 6 represents six dimensions, and H*W represents the size of each video frame in the first video. The camera parameter embedding representation provides a six-dimensional vector for each pixel, which contains the direction and position information of the line segment from the camera center to the pixel, thereby ensuring that the camera's pose information is fully expressed in each frame of the first video.

[0056] The encoder is an encoder that incorporates a temporal attention mechanism. For example, the encoder may include convolutional layers and a temporal attention layer. The convolutional layers encode the camera parameters corresponding to the video frames in the first video, as represented in the camera parameter embedding representation. The temporal attention layer captures the temporal relationships between the camera parameters corresponding to the video frames in the first video, as represented in the camera parameter embedding representation.

[0057] In the above manner, the encoder based on the temporal attention mechanism encodes the camera parameter embedding representation to obtain target camera parameter features that characterize the camera parameters corresponding to the video frames in the first video and the temporal correlation of camera parameters between different video frames in the first video. This provides a basis for precise camera control based on camera parameters when generating the second video.

[0058] In some embodiments, the step of encoding the camera parameter embedding representation using an encoder based on a temporal attention mechanism to obtain the target camera parameter features can be implemented as follows: the camera parameter embedding representation is encoded using convolutional layers of different scales in the encoder to obtain a first feature at the corresponding scale output by the convolutional layers, the first feature being used to characterize the camera parameters corresponding to the video frames in the first video; the corresponding first features are encoded using temporal attention layers corresponding to the convolutional layers in the encoder to obtain a second feature at the corresponding scale output by the temporal attention layers, the second feature being used to characterize the temporal correlation of camera parameters between different video frames in the first video; and the target camera parameter features are obtained based on all the first features and all the second features.

[0059] It should be understood that the target camera parameter features are multi-scale features. The encoder comprises multiple convolutional layers and a corresponding temporal attention layer for each convolutional layer. Different convolutional layers and temporal attention layers can extract features at different scales. As an example, the encoder includes three convolutional layers and a corresponding temporal attention layer for each convolutional layer.

[0060] By using the above method, the encoder obtains features at different scales that describe the camera parameters of video frames in the first video and the temporal correlation of camera parameters between different video frames in the first video. This can improve the feature representation capability of camera parameters and provide a foundation for precise camera control based on camera parameters when generating the second video.

[0061] In some embodiments, the step of generating a second video with a camera movement corresponding to the target camera movement type based on the first video and the target camera parameter features may include: determining a frame embedding vector corresponding to the frame features of the first video and a camera parameter embedding vector corresponding to the target camera parameter features; performing a diffusion process based on the frame embedding vector and the camera parameter embedding vector using a pre-generated video generation model to generate a second video carrying a camera movement corresponding to the target camera movement type; wherein the diffusion process includes a noise addition process and a noise reduction process, the camera parameter embedding vector is injected into the temporal attention layer involved in the noise reduction process of the video generation model in a manner based on a temporal attention mechanism, the noise addition process is used to add noise to the frame embedding vector, and the noise reduction process is used to generate a denoised second video based on the camera parameter embedding vector and the noise-added frame embedding vector.

[0062] It should be understood that the image embedding vector is a vectorized representation of the image features of the video frame in the first video; the camera parameter embedding vector is a vectorized representation of the target camera parameter features. These vectorized representations can be obtained by processing the input data through a network layer used for output vector representation. Here, the network layer is referred to as the embedding module. The embedding module can be integrated into the video generation model.

[0063] The video generation model can be based on the U-Net architecture. Similar to the embedding module mentioned above, the encoder can be integrated into the video generation model so that the encoder, embedding module, and U-Net architecture model can be trained jointly during model training, improving model performance. Furthermore, the training method for the video generation model can be referred to in the following embodiment, which will not be elaborated upon here.

[0064] Since camera movement usually causes changes in the image between frames, we can consider injecting the camera parameter embedding vector into the temporal attention layer involved in the denoising process of the video generation model in a way based on the temporal attention mechanism, thereby improving the control over camera movement throughout the entire video generation process.

[0065] In some embodiments, the video generation model described above is a model generated in the following manner: obtaining a training sample set, which includes multiple sample video data, the sample video data including sample camera movement videos and sample camera parameters corresponding to sample video frames in the sample camera movement videos, the sample camera movement videos being camera movement videos including foreground motion images or camera movement videos not including foreground motion images; training an initial model based on the training sample set to obtain a video generation model.

[0066] It should be understood that foreground motion images refer to images of moving targets in the foreground portion of a video (i.e., the area closest to the camera). These targets can be people, animals, etc., and the movement can be a movement within the target frame. Videos with foreground motion images include not only the moving target but also the movement of the camera to capture the scene. Conversely, videos without foreground motion images generally only include the background image, with no moving targets in the frame. The movement of the camera (i.e., camera movement) captures the scene, resulting in a video without foreground motion images.

[0067] The sample camera parameters include at least one of the camera's extrinsic parameters and the camera's intrinsic parameters.

[0068] As described above, both the embedding module and the encoder can be integrated into the video generation model. Therefore, corresponding to the model training process, the initial model can include an initial embedding module, an initial encoder, and an initial U-Net architecture model. The initial embedding module, initial encoder, and initial U-Net architecture model are jointly trained based on the training sample set to generate the video generation model. It is understandable that the initial embedding module, initial encoder, and initial U-Net architecture are similar in function and structure to the aforementioned embedding module, initial encoder, and initial U-Net architecture.

[0069] Figure 2 is a schematic diagram illustrating the training process of a video generation model according to an exemplary embodiment of the present disclosure. Referring to Figure 2, the initial model may include an initial embedding module, an initial encoder, and an initial U-Net architecture model. Taking the initial embedding module as an example, which includes a first embedding module, a second embedding module, and a third embedding module, the training process of the model involved in the embodiment of the present disclosure will be illustrated by way of example. Referring to Figure 2, the training process of the video generation model may be as follows:

[0070] First, the sample camera parameters are processed by the first embedding module to obtain the sample camera parameter embedding representation. Then, the sample camera parameter embedding representation is encoded by the initial encoder to determine the corresponding sample camera parameter features. The sample camera parameter embedding representation and the camera parameter embedding representation have the same physical meaning, and the sample camera parameter features and the target camera parameter features have the same physical meaning. Please refer to the above explanation and description of the camera parameter embedding representation and the target camera parameter features.

[0071] Next, the second embedding module is used to process the sample camera parameter features to obtain the sample camera parameter embedding vector. The third embedding module is used to process the sample camera movement video to obtain the sample image embedding vector. The sample camera parameter embedding vector and the camera parameter embedding vector have the same physical meaning, and the sample image embedding vector and the image embedding vector have the same physical meaning. For related explanations, please refer to the explanations and descriptions of the camera parameter embedding vector and the image embedding vector above.

[0072] Next, the initial U-Net architecture performs noise addition and denoising processes based on the sample camera parameter embedding vector and the sample image embedding vector to generate a predicted camera movement video with camera movement corresponding to the sample camera parameters. The loss value is determined based on the difference between the predicted camera movement video and the sample camera movement video. The parameters of the initial embedding module, the initial encoder and the initial U-Net architecture are iteratively updated based on the loss value until the stopping iteration condition is met. The initial embedding module, the initial encoder and the initial U-Net architecture that meet the stopping iteration condition are used as the generated video generation model.

[0073] The stopping iteration condition can be that the loss value is less than a preset threshold, or that the number of iterations of the parameters of the initial embedding module, the initial encoder, and the initial U-Net architecture reaches a preset number. This embodiment will not elaborate on this.

[0074] In other embodiments, the loss value can be determined based on the difference between the noise used in the noise addition process and the noise predicted in the noise removal process. The parameters of the initial embedding module, the initial encoder, and the initial U-Net architecture are iteratively updated based on this loss value until the stopping iteration condition is met.

[0075] By constructing sample camera movement videos that include foreground motion images and sample camera movement videos that do not include foreground motion images together for model training, the model avoids the bias of learning that some camera movement videos do not have foreground motion. In addition, the sample camera parameters include at least one of the camera's extrinsic and intrinsic parameters, specifically, the sample camera parameters include both the camera's extrinsic and intrinsic parameters, thereby improving the richness of the sample video data and thus improving the generalization of the model trained on the sample video data.

[0076] In some embodiments, before training the initial model based on the training sample set to obtain the video generation model, the video generation method may further include: deleting target sample video data from the training sample set to obtain an updated training sample set, wherein the sample camera movement videos corresponding to the target sample video data are camera movement videos with abrupt changes in screen content. Based on this, the initial model is trained using the updated training sample set to obtain the video generation model.

[0077] Specifically, features can be extracted from adjacent frames in a sample video movement, and the similarity between adjacent frames can be calculated based on the extracted features. If the similarity between any two adjacent frames is less than or equal to a preset similarity threshold, the sample video movement is determined to have a sudden change in the scene content, meaning its corresponding sample video data is the target sample video data. If the similarity between all adjacent frames is greater than the preset similarity threshold, the sample video movement is determined to have no sudden changes in the scene content, meaning its corresponding sample video data is not the target sample video data.

[0078] It should be understood that for video shots with sudden changes in content, the changes are abrupt and cannot be represented by camera parameters. Therefore, such sample videos are considered noisy data. Thus, by using the above method, video shots with sudden changes in content are detected and deleted, reducing noise during model training and thereby improving the performance of the trained model.

[0079] Based on the same inventive concept, this disclosure also provides a video generation apparatus. FIG3 is a block diagram of a video generation apparatus according to an exemplary embodiment of this disclosure. Referring to FIG3, the video generation apparatus 300 includes:

[0080] The acquisition module 301 is configured to acquire a first video and a target camera movement type, wherein the target camera movement type carries camera parameters corresponding to video frames in the first video;

[0081] The determining module 302 is configured to determine target camera parameter features based on the camera parameters corresponding to the video frames in the first video. The target camera parameter features are used to characterize the camera parameters corresponding to the video frames in the first video and the temporal correlation of camera parameters between different video frames.

[0082] The generation module 303 is configured to generate a second video with a camera movement corresponding to the target camera movement type based on the first video and the target camera parameter features.

[0083] Optionally, the determining module 302 includes:

[0084] The first determining submodule is configured to determine the camera parameter embedding representation based on the camera parameters corresponding to the video frames in the first video;

[0085] The encoding submodule is configured to encode the camera parameter embedding representation using an encoder based on a temporal attention mechanism to obtain the target camera parameter features.

[0086] Optionally, the encoding submodule is further configured as follows:

[0087] The camera parameter embedding representation is encoded by convolutional layers of different scales in the encoder to obtain a first feature of the corresponding scale output by the convolutional layer. The first feature is used to characterize the camera parameters corresponding to the video frame in the first video.

[0088] The first feature is encoded by the temporal attention layer corresponding to the convolutional layer in the encoder to obtain the second feature of the corresponding scale output by the temporal attention layer. The second feature is used to characterize the temporal correlation of camera parameters between different video frames in the first video.

[0089] Based on all the first features and all the second features, the target camera parameter features are obtained.

[0090] Optionally, the generation module 303 includes:

[0091] The second determining submodule is configured to determine the image embedding vector corresponding to the image features of the first video, and the camera parameter embedding vector corresponding to the target camera parameter features.

[0092] The generation submodule is configured to perform a diffusion process based on the image embedding vector and the camera parameter embedding vector using a pre-generated video generation model to generate a second video carrying a camera movement corresponding to the target camera movement type.

[0093] The diffusion process includes a noise-adding process and a noise-reducing process. The camera parameter embedding vector is injected into the temporal attention layer involved in the noise reduction process of the video generation model in a manner based on a temporal attention mechanism. The noise-adding process is used to add noise to the image embedding vector. The noise reduction process is used to generate the denoised second video based on the camera parameter embedding vector and the image embedding vector after the noise-adding process.

[0094] Optionally, the video generation model is a model generated based on the following method:

[0095] Obtain a training sample set, which includes multiple sample video data. The sample video data includes sample camera movement videos and sample camera parameters corresponding to sample video frames in the sample camera movement videos. The sample camera movement videos are either camera movement videos that include foreground motion images or camera movement videos that do not include foreground motion images.

[0096] The initial model is trained based on the training sample set to obtain the video generation model.

[0097] Optionally, the video generation device 300 further includes:

[0098] The deletion module is configured to delete the target sample video data in the training sample set to obtain an updated training sample set. The sample camera movement video corresponding to the target sample video data is a camera movement video with sudden changes in the content of the scene. The initial model is trained based on the updated training sample set to obtain the video generation model.

[0099] Optionally, the camera parameters include at least one of the following:

[0100] Camera external parameters;

[0101] The camera's internal parameters.

[0102] The implementation methods of each module in the video generation device 300 can be referred to the above method embodiments, and will not be repeated here.

[0103] Based on the same inventive concept, embodiments of this disclosure also provide a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the above-described method.

[0104] Based on the same inventive concept, this disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.

[0105] Based on the same inventive concept, this disclosure also provides an electronic device, including:

[0106] A storage device on which computer programs are stored;

[0107] A processing device for executing the computer program in the storage device to implement the steps of the above method.

[0108] Referring now to FIG4, a schematic diagram of the structure of an electronic device 400 suitable for implementing embodiments of the present disclosure is shown. The terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG4 is merely an example and should not impose any limitation on the functionality and scope of use of embodiments of the present disclosure.

[0109] As shown in Figure 4, the electronic device 400 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. The RAM 403 also stores various programs and data required for the operation of the electronic device 400. The processing unit 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0110] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 shows electronic device 400 with various devices, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0111] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 409, or installed from storage device 408, or installed from ROM 402. When the computer program is executed by processing device 401, it performs the functions defined in the methods of embodiments of this disclosure.

[0112] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0113] In some implementations, electronic devices can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0114] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0115] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire a first video and a target camera movement type, the target camera movement type carrying camera parameters corresponding to video frames in the first video; determine target camera parameter features based on the camera parameters corresponding to the video frames in the first video, the target camera parameter features being used to characterize the camera parameters corresponding to the video frames in the first video and the temporal correlation of camera parameters between different video frames; and generate a second video with a camera movement corresponding to the target camera movement type based on the first video and the target camera parameter features.

[0116] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0117] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0118] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules are not, in some cases, intended to limit the functionality of the module itself.

[0119] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0120] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0121] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0122] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0123] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.

Claims

1. A video generation method, comprising: Obtain a first video and a target camera movement type, wherein the target camera movement type carries camera parameters corresponding to video frames in the first video; Based on the camera parameters corresponding to the video frames in the first video, target camera parameter features are determined. The target camera parameter features are used to characterize the camera parameters corresponding to the video frames in the first video and the temporal correlation of camera parameters between different video frames. Based on the first video and the target camera parameter features, a second video with a camera movement corresponding to the target camera movement type is generated.

2. The method according to claim 1, wherein, The step of determining the target camera parameter features based on the camera parameters corresponding to the video frames in the first video includes: Based on the camera parameters corresponding to the video frames in the first video, determine the camera parameter embedding representation; The target camera parameter features are obtained by encoding the camera parameter embedding representation using an encoder based on a temporal attention mechanism.

3. The method according to claim 2, wherein, The process of encoding the camera parameter embedding representation using an encoder based on a temporal attention mechanism to obtain the target camera parameter features includes: The camera parameter embedding representation is encoded by convolutional layers of different scales in the encoder to obtain a first feature of the corresponding scale output by the convolutional layer. The first feature is used to characterize the camera parameters corresponding to the video frame in the first video. The first feature is encoded by the temporal attention layer corresponding to the convolutional layer in the encoder to obtain the second feature of the corresponding scale output by the temporal attention layer. The second feature is used to characterize the temporal correlation of camera parameters between different video frames in the first video. The target camera parameter features are obtained based on all the first features and all the second features.

4. The method according to any one of claims 1-3, wherein, The step of generating a second video with a camera movement corresponding to the target camera movement type based on the first video and the target camera parameter features includes: Determine the image embedding vector corresponding to the image features of the first video, and the camera parameter embedding vector corresponding to the target camera parameter features; A second video carrying a camera movement corresponding to the target camera movement type is generated by performing a diffusion process based on the image embedding vector and the camera parameter embedding vector using a pre-generated video generation model. The diffusion process includes a noise-adding process and a noise-reducing process. The camera parameter embedding vector is injected into the temporal attention layer involved in the noise reduction process of the video generation model in a manner based on a temporal attention mechanism. The noise-adding process is used to add noise to the image embedding vector. The noise reduction process is used to generate the denoised second video based on the camera parameter embedding vector and the image embedding vector after the noise-adding process.

5. The method according to claim 4, wherein, The video generation model is generated based on the following method: Obtain a training sample set, which includes multiple sample video data. The sample video data includes sample camera movement videos and sample camera parameters corresponding to sample video frames in the sample camera movement videos. The sample camera movement videos are either camera movement videos that include foreground motion images or camera movement videos that do not include foreground motion images. The initial model is trained based on the training sample set to obtain the video generation model.

6. The method according to claim 5, further comprising: The target sample video data in the training sample set is deleted to obtain an updated training sample set. The sample camera movement video corresponding to the target sample video data is a camera movement video with sudden changes in the content of the scene. The initial model is trained based on the updated training sample set to obtain the video generation model.

7. The method according to any one of claims 1-6, wherein the camera parameters include at least one of the following: Camera external parameters; The camera's internal parameters.

8. A video generation apparatus, comprising: The acquisition module is configured to acquire a first video and a target camera movement type, wherein the target camera movement type carries camera parameters corresponding to video frames in the first video; The determining module is configured to determine target camera parameter features based on the camera parameters corresponding to the video frames in the first video. The target camera parameter features are used to characterize the camera parameters corresponding to the video frames in the first video and the temporal correlation of camera parameters between different video frames. The generation module is configured to generate a second video with a camera movement corresponding to the target camera movement type, based on the first video and the target camera parameter features.

9. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by the processing device, it implements the steps of the method according to any one of claims 1-7.

10. An electronic device, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-7.

11. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Video processing method and related device

    CN114979785A

  • Video generation method

    CN116939325A

  • Video generation method and device, medium, electronic equipment and program product

    CN119364142A

  • Composition generation method, composition operation method, mirror display method, apparatus, electronic device and storage medium

    WO2024131010A1