Video generation method and device

CN120017928APending Publication Date: 2025-05-16VIVO MOBILE COMM CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510161194.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-16

Smart Images

  • Figure CN120017928A_ABST
    Figure CN120017928A_ABST
Patent Text Reader

Abstract

The invention discloses a video generation method and device, and belongs to the technical field of electronics. The method comprises the following steps: determining a motion feature of a camera based on an acquired first motion parameter and a first image, the first motion parameter being used for representing a rotation angle and a translation distance of the camera; and inputting the motion feature and the first image into a first video generation model, and generating a first video based on the motion feature and the image feature of the first image through the first video generation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of electronic technology, and specifically relates to a video generation method and device thereof. Background Art

[0002] With the continuous advancement of video generation technology, video generation models have been widely used due to their high-quality generation effects, such as the Diffusion Transformer (DiT). DiT can compress the duration and spatial information of the video through the Variational Autoencoder (VAE) to obtain a compact representation of the video features, and then decompose these compact representations into multiple patches. Each patch contains a part of the spatiotemporal information of the video. DiT can then extract the temporal and spatial features of these patches through the temporal self-attention layer and the spatial self-attention layer respectively, and then calculate the features between patches through the cross-attention layer. Finally, the generated video is obtained through the decoder.

[0003] However, while DiT can generate high-quality videos, it lacks fine-grained control over the camera's trajectory. Because controlling the camera's trajectory requires a complex motion matrix, it's difficult to incorporate it into the DiT framework, making it difficult to achieve precise control of the camera's trajectory. Consequently, when electronic devices use video generation models to generate videos, the accuracy of camera trajectory control is poor. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a video generation method and device thereof, which can improve the accuracy of camera motion trajectory control when an electronic device uses a video generation model to generate a video.

[0005] In a first aspect, an embodiment of the present application provides a video generation method, the method comprising: determining the motion characteristics of a camera based on an acquired first motion parameter and a first image, the first motion parameter being used to represent the rotation angle and translation distance of the camera; inputting the motion characteristics and the first image into a first video generation model, and generating a first video based on the motion characteristics and image characteristics of the first image through the first video generation model.

[0006] In a second aspect, embodiments of the present application provide a video generation device, comprising: a determination module and an execution module. The determination module is configured to determine a camera's motion characteristics based on acquired first motion parameters and a first image, wherein the first motion parameters represent the camera's rotation angle and translation distance. The execution module is configured to input the motion characteristics determined by the determination module and the first image into a first video generation model, and generate a first video using the first video generation model based on the motion characteristics and image features of the first image.

[0007] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.

[0008] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0009] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the method described in the first aspect.

[0010] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.

[0011] In an embodiment of the present application, the electronic device can first determine the motion characteristics of the camera based on the acquired first motion parameters and the first image, where the first motion parameters are used to represent the rotation angle and translation distance of the camera, and then input the motion characteristics and the first image into the first video generation model, and generate the first video based on the motion characteristics and the image characteristics of the first image through the first video generation model. In this solution, since the electronic device can determine the motion characteristics of the camera based on the first motion parameters and the first image used to represent the rotation angle and translation distance of the camera, after inputting the motion characteristics and the first image into the first video generation model, the first video generation model can generate the first video based on the motion characteristics and the image characteristics of the first image, that is, based on the content of the first image, the trajectory of the camera can be controlled in combination with the motion of the camera to generate a first video in which the camera moves according to the trajectory. In this way, the accuracy of the electronic device in controlling the camera motion trajectory when generating a video using the first video generation model is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 This is one of the flow charts of the video generation method provided in the embodiment of the present application;

[0013] Figure 2 This is the second flowchart of the video generation method provided in the embodiment of the present application;

[0014] Figure 3 is a schematic diagram of the conversion process of a sparse motion field provided in an embodiment of the present application;

[0015] Figure 4is a schematic diagram of a sparse sports field provided in an embodiment of the present application;

[0016] Figure 5 This is the third flow chart of the video generation method provided in the embodiment of the present application;

[0017] Figure 6 This is the fourth flow chart of the video generation method provided in the embodiment of the present application;

[0018] Figure 7 is a schematic diagram of the feature fusion process provided in an embodiment of the present application;

[0019] Figure 8 This is the fifth flow chart of the video generation method provided in the embodiment of the present application;

[0020] Figure 9 This is the sixth flowchart of the video generation method provided in the embodiment of the present application;

[0021] Figure 10 is a schematic diagram of the execution process of the video generation method provided in an embodiment of the present application;

[0022] Figure 11 This is one of the effect diagrams of the video generation method provided in the embodiment of the present application;

[0023] Figure 12 This is the second schematic diagram of the effect of the video generation method provided in the embodiment of the present application;

[0024] Figure 13 This is one of the schematic diagrams of the video generating device provided in the embodiment of the present application;

[0025] Figure 14 This is the second schematic diagram of the video generating device provided in an embodiment of the present application;

[0026] Figure 15 is a structural diagram of an electronic device provided in an embodiment of the present application;

[0027] Figure 16 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0028] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0029] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0030] The terms "at least one" and "at least one of" in this application refer to any one, any two, or a combination of more than two of the objects included. For example, at least one of a, b, and c can be represented by: "a", "b", "c", "a and b", "a and c", "b and c", and "a, b, and c", where a, b, and c can be single or multiple. Similarly, "at least two" means two or more, and its meaning is similar to "at least one".

[0031] The video generation method and apparatus provided by the embodiments of the present application are described in detail below with reference to specific embodiments and their application scenarios in conjunction with the accompanying drawings.

[0032] The embodiments of the present application can be applied to scenarios where video is generated through image and camera motion. Specifically, the electronic device generates high-quality video with rich dynamic effects and perspective changes by combining motion features and image features.

[0033] The following uses some specific scenarios of the embodiments of the present application as examples to exemplify the video generation method provided in the embodiments of the present application.

[0034] Scenario 1: Suppose a user takes a photo of a distant mountain while traveling and wishes to generate a video with the viewpoint gradually shifting to the right. The user inputs the photo and the corresponding motion parameters, including translation distance and rotation angle. The electronic device then samples the motion matrix through the sparse motion coding module to generate a sparse motion field, which is then compressed through the VAE to obtain motion features. The electronic device then inputs these motion features into the motion latent space layer of the DiT. Simultaneously, the photo is input into the temporal attention layer of the DiT. The motion features are aligned and embedded with the image features through the motion feature alignment and embedding module to obtain a comprehensive feature. Finally, the electronic device generates a video with translation and rotation camera motion angles based on the comprehensive features through the cross-attention layer. The resulting video not only retains the aesthetic of the original photo but also adds dynamic effects, making the scenery captured by the user more vivid.

[0035] Scenario 2: Suppose a user captures a video showing city streets and buildings. The last frame of the video is a panoramic view of a landmark building. The user wishes to generate subsequent videos with the camera's perspective gradually moving upward, showing the top of the building and the sky. The user inputs the video and corresponding motion parameters, including rotation angle and translation distance. The electronic device then uses the sparse motion coding module to sample the motion matrix, generating a sparse motion field. This is then compressed using a VAE to obtain motion features. The electronic device then inputs these motion features into the motion latent space layer of the DiT and the last frame into the temporal attention layer of the DiT. The motion features are aligned and embedded with the image features through the motion feature alignment embedding module to obtain a comprehensive feature. Finally, the electronic device uses the cross-attention layer to generate a video with an upward camera motion angle based on the comprehensive features. The generated video continues the last frame of the original video, adding new perspective content and providing a more complete presentation of the cityscape.

[0036] It should be noted that the above-mentioned scenarios 1 and 2 are merely examples of some possible scenarios in which the embodiments of the present application may be applied. In actual implementation, the embodiments of the present application can also be applied to more scenarios requiring video generation, and the embodiments of the present application are not limited here.

[0037] The embodiment of the present application provides a video generation method and apparatus thereof. Since the electronic device can determine the motion characteristics of the camera based on the first motion parameter and the first image used to represent the rotation angle and translation distance of the camera, after the motion characteristics and the first image are input into the first video generation model, the first video generation model can generate a first video based on the motion characteristics and the image characteristics of the first image. That is, the electronic device can control the trajectory of the camera based on the content of the first image and the motion of the camera, and generate a first video in which the camera moves according to the trajectory. In this way, the accuracy of the electronic device in controlling the camera motion trajectory when generating a video using the first video generation model is improved.

[0038] The video generation method provided in the embodiment of the present application may be executed by a video generation device, which may be an electronic device, or a functional module or functional entity in the electronic device. The technical solution provided in the embodiment of the present application is described below using an electronic device as an example.

[0039] Figure 1 A flow chart of a video generation method provided by an embodiment of the present application is shown. Figure 1 As shown, the video generation method provided in the embodiment of the present application may include the following steps 201 and 202.

[0040] Step 201: The electronic device determines a motion feature of a camera based on an acquired first motion parameter and a first image.

[0041] In the embodiment of the present application, the first motion parameter is used to represent the rotation angle and translation distance of the camera.

[0042] In the embodiment of the present application, the above-mentioned camera can be a virtual camera or a physical camera with AI capabilities, such as a SLR camera, a micro single-lens camera, etc. Among them, the virtual camera is a camera simulated in computer graphics and the first video generation model. The virtual camera has no physical form, but can simulate the behavior of an actual camera through mathematical models and parameters. The physical camera can process image data in real time through AI capabilities, extract motion parameters, such as rotation angle and translation distance, to provide support for image processing tasks. In the embodiment of the present application, the electronic device can simulate the movement and perspective change of the camera through the rotation matrix and translation vector.

[0043] Optionally, in an embodiment of the present application, the above-mentioned first motion parameter may be a motion matrix sequence, which includes at least one motion matrix. The motion matrix may be in the form of a combination of a rotation matrix R and a translation vector t.

[0044] Among them, R is a 3×3 rotation matrix used to describe the rotation angle of the camera. For example, if the camera rotates 45 degrees around the Z axis in the spatial coordinate system, the rotation matrix R can be expressed as:

[0045]

[0046] in,

[0047] t is a 3×1 translation vector that describes the camera's translation distance. For example, if the camera translates 1 meter to the right, the translation vector t can be expressed as:

[0048]

[0049] Alternatively, the first motion parameter may be other forms of motion description, such as Euler angles, quaternions, or rotation vectors, etc. The specific form may be determined based on actual usage requirements and is not limited in the embodiments of the present application.

[0050] Optionally, in the embodiment of the present application, the first motion parameter may be a default value of the electronic device or a preset value by the user, and may be determined based on actual use requirements, and is not limited in the embodiment of the present application.

[0051] Optionally, in an embodiment of the present application, the above-mentioned first image can be an image captured by a user through a camera of an electronic device, or an image of an interface in an electronic device captured by a user using an electronic device, or an image imported into the electronic device by a user from other storage devices, or an image downloaded by the electronic device from the Internet. The specific image can be determined according to actual usage requirements and is not limited in the embodiment of the present application.

[0052] Optionally, in an embodiment of the present application, the first image may be a static image, such as a photo, or the first image may be a dynamic image, such as a video.

[0053] Optionally, in the embodiment of the present application, Figure 1 ,like Figure 2 As shown, the above step 201 can be specifically implemented through the following steps 201a to 201c.

[0054] Step 201a: The electronic device determines a pixel points from the first image based on the resolution of the first image through a sparse motion coding module, and determines a pixel coordinates that correspond one-to-one to the a pixel points.

[0055] In the embodiment of the present application, a is a positive integer.

[0056] In the embodiment of the present application, the sparse motion coding module is a module in the electronic device for processing the first motion parameter. Taking the first motion parameter as a motion matrix sequence as an example, the sparse motion coding module can convert the complex motion matrix sequence into a sparse motion field through sampling and encoding.

[0057] In the embodiment of the present application, the sparse motion field is a simplified representation of the motion of the virtual camera. Compared with the motion matrix sequence, the sparse motion field retains the key information used to describe the motion of the virtual camera while reducing the amount of data, thereby facilitating subsequent calculations.

[0058] Optionally, in an embodiment of the present application, the sparse motion field includes Plücker coordinates of a sampling pixels.

[0059] It should be noted that the above-mentioned Plücker coordinates are also called Planck coordinate vectors, which are used to represent the direction and position of a straight line in three-dimensional space through six parameters.

[0060] It should be noted that the pixel point mentioned above is the basic element of a digital image. It is a single point or pixel. For example, a pixel point can refer to a specific point in a digital image. Pixel is the smallest unit of a digital image and is used to describe the resolution of a digital image. For example, the width of a digital image can be 1080 pixels.

[0061] Optionally, in an embodiment of the present application, the electronic device may determine the Plücker coordinates of each pixel in the first image, and then make these Plücker coordinates form a motion field, which may describe the motion information of each pixel in the first image.

[0062] Optionally, in an embodiment of the present application, the electronic device may determine all pixel points from the first image.

[0063] For example, assuming that the resolution of the first image is 1920×1080, the electronic device may determine 1920×1080=2073600 pixels from the first image.

[0064] Optionally, in an embodiment of the present application, the electronic device may determine a sampling pixel every preset number of pixels in the first image, and calculate the Plücker coordinates of these sampling pixels, so that the Plücker coordinates of these sampling pixels form a sparse motion field to simplify the calculation.

[0065] Optionally, in the embodiment of the present application, the preset number may be a default number of the electronic device or a preset number of the user. For example, the preset number may be 10, 50, 100, etc. The specific number may be determined according to actual use requirements and is not limited in the embodiment of the present application.

[0066] For example, assuming that the electronic device determines one pixel point for every 50 pixels, and assuming that the resolution of the first image is 1920×1080, the electronic device can determine 1920×1080÷50=41472 pixel points from the first image.

[0067] It should be noted that the above pixel coordinates are used to represent the position of a pixel point in the plane where the image is located. For example, the pixel coordinates may be in the form of (u, v), where u is the column coordinate and v is the row coordinate.

[0068] In an embodiment of the present application, since the representation of the camera imaging plane needs to be projected and transformed based on the real three-dimensional world coordinates, the electronic device can determine the spatial coordinates of a pixel points in the first image, and then map these a pixel points to the two-dimensional plane where the first image is located.

[0069] Optionally, in an embodiment of the present application, the electronic device may determine the spatial coordinates of a pixel point in the first image, and then convert the a spatial coordinates into a camera coordinates. The spatial coordinates are used to represent the spatial position of the pixel point in the first image, and the camera coordinates are used to represent the position of the pixel point in the first image relative to the camera. Specifically, the spatial coordinate X of the pixel point A in homogeneous coordinate form is X = [x, y, z, 1] T , the electronic device can use the rotation matrix R and the translation term t to convert the space coordinate X into the camera coordinate X c, the specific conversion method is shown in the following formula (1):

[0070] X c =PX=[R|t]X Formula (1)

[0071] Among them, X c is the camera coordinate of pixel A, R is a 3×3 rotation matrix used to describe the camera's rotation angle, and t is a 3×1 translation vector used to describe the camera's translation distance. The specific forms of R and T are described above and will not be repeated here. P is a 4×4 homogeneous transformation matrix used to represent the transformation of spatial coordinates. The specific form is as follows:

[0072]

[0073] The electronic device can convert the spatial coordinate X of pixel point A into the camera coordinate X by formula (1): c .

[0074] For example, assume that the camera rotates 45 degrees around the Z axis in the spatial coordinate system and translates 1 meter to the right. Assume that the spatial coordinate X of pixel point A in the spatial coordinate system is [0, 0, 0, 1] T , then the rotation matrix R is:

[0075]

[0076] in, Therefore, the specific rotation matrix R is:

[0077]

[0078] The translation vector t for a 1 meter shift to the right can be expressed as:

[0079]

[0080] Then the electronic device can determine the homogeneous transformation matrix P as:

[0081]

[0082] The electronic device can convert the spatial coordinates of pixel point A into [0,0,0,1] T Substituting the homogeneous transformation matrix P into formula (1), we get:

[0083]

[0084] The X coordinates are [0,0,0,1] T Convert to camera coordinates [1,0,0,1] T .

[0085] Optionally, in an embodiment of the present application, the electronic device can convert a camera coordinates into a pixel coordinates through a camera intrinsic parameter matrix. The camera intrinsic parameter matrix is ​​a matrix that describes the internal parameters of the camera, including information such as the focal length, principal point, and image size of the camera. The camera intrinsic parameter matrix is ​​used to map points in three-dimensional space to a two-dimensional image plane. Specifically, for the camera coordinate X c =PX=[R|t]X=RX+t. The camera coordinates can be converted to pixel coordinates through the camera intrinsic parameter matrix K. The specific conversion method is shown in the following formula (2):

[0086] x=KX c =K(RX+t) Formula (2)

[0087] Among them, x is the pixel coordinate of pixel point A, K is the camera intrinsic parameter matrix, and the specific form is as follows:

[0088]

[0089] Among them, f x and f y is the focal length of the camera, C x and C y are the coordinates of the camera's optical center.

[0090] It should be noted that the X obtained in formula (1) c In homogeneous coordinate form, substitute X into formula (2) c You need to convert it into Cartesian coordinates first. For example, for homogeneous coordinates [X, Y, Z, W] T , we can divide X, Y, and Z by W respectively to get These are Cartesian coordinates.

[0091] The following describes formula (2) with reference to specific examples.

[0092] For example, assume that the camera coordinate X of pixel A is c is [1,0,0,1] T , assuming the focal length of the camera is f x =1000 pixels, f y =1000 pixels, optical center coordinate C x =960 pixels, C y =540 pixels, assuming the image resolution is 1920×1080, the electronic device can now c Homogeneous coordinates are converted to Cartesian coordinates, that is, [1,0,0,1] T Divide the 1, 0, and 0 in the image by the 1 at the end to get the camera coordinate X of pixel A. c The Cartesian coordinate form is [1,0,0] T, then the electronic device can [1,0,0] T Substituting into formula (2), we get the pixel coordinate x:

[0093]

[0094] The final pixel coordinate x is [1000,0,1] T , the electronic device can take the first two elements as pixel coordinates, that is, pixel coordinate x = [1000,0].

[0095] In this way, the electronic device can determine the corresponding pixel coordinates for each of the a pixel points, thereby determining a pixel coordinates.

[0096] Step 202b: The electronic device converts a pixel coordinates into a Plücker coordinates through a sparse motion coding module to obtain a sparse motion field.

[0097] Optionally, in the embodiment of the present application, the electronic device can convert a pixel point into a 2D pixel coordinate by combining the camera's intrinsic parameter matrix K to 3D camera coordinate X. img =K -1 [x,y,1] T , and in camera coordinate X img Based on this, combined with the camera's external parameters, namely the rotation matrix R and the translation vector t, the spatial coordinates Q of a pixel point are determined. x,y , the specific calculation method is shown in the following formula (3):

[0098] Q x,y =RK -1 [x,y,1] T +t Formula (3)

[0099] Among them, the above Q x,y is the spatial coordinate of a pixel point.

[0100] For example, assuming that the pixel coordinates of pixel point A are x=[1000,0], and the camera intrinsic parameter matrix K is:

[0101]

[0102] Then the electronic device can calculate the camera coordinate X img for:

[0103]

[0104] Assume that the camera rotates 45 degrees around the Z axis and translates 1 meter to the right, that is, the rotation matrix R is:

[0105]

[0106] The translation vector t is:

[0107]

[0108] Then the electronic device can take the camera coordinate X img , the rotation matrix R and the translation vector t are substituted into formula (3) to calculate the spatial coordinate Q x,y for:

[0109]

[0110] That is, the spatial coordinate Q of pixel point A x,y for

[0111] Optionally, in the embodiment of the present application, the electronic device can use the spatial coordinates [o c ,1], the spatial coordinates of the pixel point are converted into the Plücker coordinate vector, where o c is the Cartesian coordinate of the optical center o. The specific conversion method is shown in the following formula (4):

[0112] P x,y =[o c ,1](RK -1 [x,y,1] T +t) Formula (4)

[0113] Combined with formula (3), P x,y =[o c ,1]Q x,y .

[0114] For example, it is assumed that the spatial coordinates of the camera optical center o in homogeneous coordinate form are [o c ,1], where o c =[0,0,0], that is, the optical center is at the origin. Assume that the spatial coordinates of pixel A are Q x,y for:

[0115]

[0116] Then the electronic device can [o c ,1] and Q x,y Substituting into formula (4), we get:

[0117]

[0118] Calculation yields:

[0119]

[0120] Since the Planck coordinate vector contains six parameters, Px,y for:

[0121]

[0122] That is, the Planck coordinate of pixel point A is P x,y =[1,0,0,0,0,0] T .

[0123] Optionally, in an embodiment of the present application, the electronic device may sample the first image at fixed intervals to obtain a sparse pixel points, and calculate the Plücker coordinates of these pixel points to form a sparse motion vector field, i.e., a sparse motion field. Assuming that the resolution of the first image is W×H, the electronic device may sample the first image every s in the column coordinate direction. x pixels, every s in the horizontal direction y Pixels are sampled to obtain the sparse point sequence {(x i ,y j )}, then the electronic device can calculate the Plücker coordinates of these pixel points using the following formula (5), which is as follows:

[0124]

[0125] Among them, x i =i·s x , yj=j·s y , i and j are sampling sequences. In this way, we get a sparse motion field M is the number of pixels sampled in the horizontal direction of the first image, and the calculation formula is M=W / s x , s x is the sampling interval in the horizontal direction, N is the number of pixels sampled in the vertical direction of the first image, and the calculation formula is N = H / s y , s y is the sampling interval in the vertical direction, L represents the length of the time series, Represents the field of real numbers.

[0126] In the embodiment of the present application, the electronic device can convert the first motion parameter into a sparse motion field. The conversion process is as follows: Figure 3 As shown, the schematic diagram of the sparse motion field is as follows Figure 4 shown.

[0127] In this way, the electronic device can uniformly sample the first image according to the resolution of the first image and calculate the Plücker coordinates of these sampled pixels, thereby significantly reducing the amount of calculation, alleviating the calculation burden, and improving the efficiency of subsequently obtaining the sparse motion field.

[0128] Step 201c: The electronic device compresses the sparse motion field through VAE to obtain motion features.

[0129] In the embodiment of the present application, the above-mentioned VAE is also called a variational autoencoder, which is a deep learning model that can generate new samples similar to real data.

[0130] In an embodiment of the present application, the above-mentioned motion features are low-dimensional latent space features obtained by compressing the sparse motion field through VAE. These features capture the key motion information in the sparse motion field and can be used for subsequent video generation.

[0131] Optionally, in an embodiment of the present application, the electronic device can compress the high-dimensional sparse motion field into a low-dimensional latent space through VAE to obtain the motion feature z p , z p The dimension is l×m×n×4, that is, Wherein l=L / 4, m=M / 8, n=N / 8.

[0132] In an embodiment of the present application, the electronic device needs to use Dit to generate video later. Dit, also known as diffusion transformer, is a new generation model that combines the transformer (Transformer) architecture and the diffusion model for image and video generation. When training DiT, the electronic device can map the video clip into a low-dimensional latent space through video VAE, compress the video duration information of the video clip by 4 times, and compress the spatial information by 8 times in the width and height dimensions respectively, so that a compact representation of the video clip can be obtained. Therefore, the electronic device here can compress the sparse motion field with the same compression ratio as VAE sampling to obtain motion features, so as to ensure the dimensional consistency of the sparse motion field and the temporal attention layer in DiT when the subsequent motion features are fused with the image features of the first image.

[0133] In the embodiments of the present application, the electronic device can embed the camera's intrinsic and extrinsic parameters into a motion field based on Plücker coordinates through mathematical derivation, thereby reducing the burden of small perturbations during camera motion capture. Furthermore, the electronic device guides the training of the camera trajectory VAE through a sparse motion coding module, achieving efficient sparse coding of the camera trajectory.

[0134] In this way, when compressing the sparse motion field, the electronic device can choose to compress the sparse motion field at the same compression ratio as when using DiT to generate the video, so as to maintain the dimensional consistency of the sparse motion field and the temporal attention layer of DiT.

[0135] Step 202: The electronic device inputs the motion feature and the first image into a first video generation model, and generates a first video based on the motion feature and the image feature of the first image through the first video generation model.

[0136] In the embodiment of the present application, the first video generation model is DiT.

[0137] In an embodiment of the present application, the above-mentioned image features are also called temporal latent features, which are feature vectors extracted from the first image through a feature extraction network, and are used to represent the relationship and dependency between different time steps in the video.

[0138] Optionally, in an embodiment of the present application, the electronic device may input the motion feature into the motion latent space layer in the first video generation model, and input the first image into the time attention layer (Time Attention Layer) in the first video generation model.

[0139] In the embodiment of the present application, the motion latent space layer is a low-dimensional latent space obtained by compressing the sparse motion field through VAE. The temporal attention layer is a neural network layer for processing temporal dependencies in sequence data.

[0140] Optionally, in the embodiment of the present application, Figure 1 ,like Figure 5 As shown, the above step 202 can be specifically implemented through the following steps 202a and 202b.

[0141] Step 202a: The electronic device aligns and embeds the motion feature and the image feature through the motion feature alignment and embedding module in the first video generation model to obtain a first feature.

[0142] In the embodiment of the present application, the motion feature alignment and embedding module is used to align and embed the motion feature and the image feature to generate the first feature.

[0143] In the embodiments of the present application, the first feature described above, also referred to as a comprehensive feature, refers to the feature obtained by aligning and embedding motion features and image features using the motion feature alignment and embedding module in the first video generation model. The comprehensive feature combines motion and image information to more comprehensively describe the video content, providing a feature representation for subsequent video generation.

[0144] In an embodiment of the present application, the electronic device can introduce a motion feature alignment embedding module to align the temporal attention module and the camera motion module to the same scale, thereby realizing the injection of long video motion information.

[0145] Optionally, in the embodiment of the present application, Figure 5 ,like Figure 6 As shown, the above step 202a can be specifically implemented through the following steps 202a1 to 202a3.

[0146] Step 202a1: The electronic device splits the first image through a motion feature alignment and embedding module to obtain p image patches.

[0147] In the embodiment of the present application, p is a positive integer.

[0148] In the embodiment of the present application, the above-mentioned image patch is a part of the image feature of the first image.

[0149] Optionally, in an embodiment of the present application, the electronic device may extract image features from the first image and split the image features into multiple patches, each of which is referred to as an image patch. For example, the electronic device may perform feature extraction on the first image. Assuming that a 27×24×13×4 feature representation is obtained after feature extraction, the electronic device may split the feature representation into 4×4 feature patches.

[0150] Optionally, in an embodiment of the present application, the electronic device may first split the first image into p patches, and then perform feature extraction on each patch to obtain multiple image patches. The size of each patch may be a default size of the electronic device or a preset size by the user. For example, the length and width of the patch may both be 16 pixels, i.e., 16×16. For another example, the length and width of the patch may both be 32 pixels, i.e., 32×32. The specific size may be determined based on actual usage requirements and is not limited by the embodiment of the present application.

[0151] For example, assuming that the size of the first image is 1080×1920, the electronic device splits the first image into 16×16 patches, and the size of each patch is 16×16, then the electronic device can split the first image into p=(1080 / 16)×(1920 / 16)=675 patches, and then perform feature extraction on each patch to obtain multiple image patches.

[0152] Step 202a2: The electronic device performs sparse spatial encoding on the motion features through the motion feature alignment and embedding module to obtain q feature patches.

[0153] In the embodiment of the present application, q is a positive integer.

[0154] In the embodiment of the present application, the above-mentioned feature patch is obtained by splitting the motion feature into multiple small blocks by the electronic device, and each small block is called a feature patch.

[0155] In an embodiment of the present application, the electronic device can split the motion features into q feature patches. The specific splitting method is the same as the above-mentioned method of splitting the first image into p image patches, which will not be repeated here.

[0156] Step 202a3: The electronic device performs feature fusion on the p image patches and the q feature patches to obtain a first feature.

[0157] Optionally, in an embodiment of the present application, the electronic device may input p image patches and q feature patches into the cross-attention layer in DiT, and combine the image features and motion features through the cross-attention mechanism to generate a richer feature representation.

[0158] Optionally, in the embodiment of the present application, the electronic device may first perform the image feature z of the kth layer. (k) Perform layer normalization (LN) processing, which is also called normalization. The specific processing formula is shown in the following formula (6):

[0159]

[0160] Among them, μ k is the mean value of the k-th layer image feature, σ k is the variance of the k-th layer image features, and ∈ is a small constant used to avoid division by zero. Formula (6) subtracts the mean and divides by the standard deviation to convert z (k) Convert to standard normal distribution to achieve normalization of image features.

[0161] Optionally, in the embodiment of the present application, the electronic device may convert the normalized time potential feature into and motion characteristics Perform feature fusion processing to obtain comprehensive features The specific process is shown in the following formula (7):

[0162]

[0163] The above ⊕ represents the feature fusion operation, which can be a simple concatenation or weighted summation.

[0164] Optionally, in the embodiment of the present application, the electronic device can Perform linear projection and adjustment to obtain the time potential feature z of the next layer (k+1) The specific process is shown in the following formula (8):

[0165]

[0166] Among them, γ k is the linear projection matrix, used to transform Projected into the new feature space, β k is the bias vector used to translate the projected features. Represents matrix multiplication.

[0167] In the embodiment of the present application, the electronic device can perform feature fusion of p image patches and q feature patches through the motion feature alignment embedding module in the first video generation model. The specific process is as follows: Figure 7 shown.

[0168] In this way, the electronic device can utilize image features and motion features, and generate comprehensive features that combine image information and motion information through the motion feature alignment embedding module in the first video generation model, so as to more comprehensively represent the video content and provide rich feature representation for subsequent video generation.

[0169] Step 202b: The electronic device generates a first video based on the first feature using a first video generation model.

[0170] Optionally, in an embodiment of the present application, the electronic device may use DiT to process the first feature through its internal cross-attention layer and self-attention layer. Specifically, DiT may start from generating noisy data and then gradually remove the noise until a clear video frame is generated to obtain the first video.

[0171] An embodiment of the present application provides a video generation method. Since the electronic device can determine the motion characteristics of the camera based on the first motion parameter and the first image used to represent the rotation angle and translation distance of the camera, after the motion characteristics and the first image are input into the first video generation model, the first video generation model can generate a first video based on the motion characteristics and the image characteristics of the first image. That is, the electronic device can control the trajectory of the camera based on the content of the first image and the motion of the camera to generate a first video in which the camera moves according to the trajectory. In this way, the accuracy of the electronic device in controlling the camera motion trajectory when generating a video using the first video generation model is improved.

[0172] Optionally, in the embodiment of the present application, the motion feature alignment embedding module includes a multi-layer perceptron. Figure 6 ,like Figure 8 As shown, before the above step 202a1, the video generation method provided in the embodiment of the present application further includes the following step 301.

[0173] Step 301: The electronic device adjusts the number and dimension of the q feature patches to be the same as the p image patches through a multi-layer perceptron.

[0174] It should be noted that the Multi-Layer Perceptron (MLP) described above is a feedforward neural network with a multi-layer structure consisting of multiple neurons, typically including an input layer, one or more hidden layers, and an output layer. Each neuron generates an output by performing a weighted summation of its input values ​​and applying an activation function.

[0175] Optionally, in an embodiment of the present application, the electronic device can input q feature patches into the input layer of the MLP, perform nonlinear transformation on these features through one or more hidden layers of the MLP, adjust the dimension of the features, and output them through the output layer. The number and dimension of neurons in the output layer are the same as the p image patches, ensuring that the output feature patches are consistent with the image patches in number and dimension.

[0176] In this way, the electronic device can use the multi-layer perceptron to adjust the number and dimension of the q feature patches to the same as the p image patches, providing consistent feature representation for subsequent feature fusion and video generation.

[0177] Optionally, in an embodiment of the present application, before the above step 202, the video generation method provided in the embodiment of the present application further includes the following steps 401 and 402, and the above step 202 can be implemented by the following step 202c.

[0178] Step 401: The electronic device obtains a first prompt word.

[0179] In the embodiment of the present application, the above-mentioned prompt word (prompt) is also called a text prompt, which is text information input into the generation model to guide the generation model to generate specific content. The prompt word can be a simple text description or a complex sentence or paragraph, depending on the requirements of the generation task.

[0180] Optionally, in an embodiment of the present application, the electronic device may obtain a prompt word (prompt) input by the user, and perform subsequent video generation based on the prompt word input by the user.

[0181] Step 402: The electronic device extracts features from the first prompt word to obtain text features.

[0182] In the embodiment of the present application, the above text features are feature vectors extracted from the text, which are used to represent the semantic information of the text. These features can be word embeddings, sentence embeddings, and other feature representations.

[0183] Optionally, in an embodiment of the present application, the electronic device may perform text feature extraction on the first prompt word to obtain text features, thereby providing DiT with semantic information of the text, enabling DiT to understand the meaning of the prompt word and generate content that meets user expectations.

[0184] Step 202c: The electronic device inputs the motion features, the first image, and the text features into a first video generation model, and generates a first video based on the motion features, the image features of the first image, and the text features through the first video generation model.

[0185] Optionally, in an embodiment of the present application, after receiving the first prompt word input by the user, the electronic device can extract the text features of the first prompt word, and then the electronic device can input the motion features, image features and text features into the first video generation model, i.e., DiT. DiT can perform normalization, alignment and feature fusion operations on these input features through its internal cross-attention layer and self-attention layer to obtain the fused first feature, that is, the first feature that meets the first prompt word input by the user, and generate the first video that meets the first prompt word based on the first feature. Please refer to the above description for the specific feature fusion operation, which will not be repeated here.

[0186] For example, assuming that the prompt word input by the user is "a person running in the park", the electronic device can generate a video of a person running in the park based on the text features of the prompt word, combined with the motion features and image features.

[0187] For another example, assuming that the prompt word input by the user is "a sunset scene at the seaside", the electronic device can generate a video of a sunset scene at the seaside based on the text features of the prompt word, combined with the motion features and image features.

[0188] In this way, the electronic device can receive the prompt words input by the user and extract their text features, so that the electronic device can better understand the user's intentions, and thus combine these text features when generating videos subsequently to generate higher quality videos that are more in line with user expectations, thereby improving the user experience.

[0189] Optionally, in the embodiment of the present application, Figure 1 ,like Figure 9 As shown, before the above step 201, the video generation method provided in the embodiment of the present application further includes the following steps 501 to 503.

[0190] Step 501: The electronic device determines a sample motion feature of a camera based on acquired sample motion parameters and sample images.

[0191] In the embodiment of the present application, the above-mentioned sample motion parameters are used to represent the rotation angle and translation distance of the camera.

[0192] In the embodiment of the present application, the sample motion parameters may be a sufficiently large number of sample motion parameters, and the sample images may be the same number of sample images as the sample motion parameters, for example, 5000 sample motion parameters and 5000 sample images, with a one-to-one correspondence between the sample motion parameters and the sample images.

[0193] It should be noted that the specific sample motion parameters and the number of sample image answers can be determined according to the actual needs during the model training process, and this application does not limit them here.

[0194] Step 502: The electronic device inputs the sample motion features and the sample image into a preset model, and generates a sample video based on the sample motion features and the image features of the sample image through the preset model.

[0195] In the embodiment of the present application, the composition of the above-mentioned preset model is roughly the same as the composition of the above-mentioned first video generation model. For details, please refer to the description of the above-mentioned embodiment, and the embodiment of the present application will not be repeated here.

[0196] In the embodiment of the present application, the above-mentioned sample motion features are motion features extracted by the electronic device from the sample motion parameters. The extraction process is roughly the same as that in the above-mentioned embodiment. For details, please refer to the description of the above-mentioned embodiment, which will not be repeated here.

[0197] In the embodiment of the present application, the above-mentioned sample video is a video generated by the electronic device through a preset model according to the sample motion characteristics and image characteristics of the sample image.

[0198] Step 503: The electronic device trains a preset model based on the sample video and the target video to obtain a first video generation model.

[0199] In the embodiment of the present application, the target video is a video that the electronic device is expected to generate based on the sample motion parameters and the sample image answers, and the target video has a one-to-one correspondence with the sample motion parameters and the sample image.

[0200] Optionally, in an embodiment of the present application, the electronic device may calculate a loss value based on the sample video and the target video, where the loss value is used to represent the difference between the sample video and the target video.

[0201] Optionally, in an embodiment of the present application, the electronic device may adjust the model parameters in the preset model based on the loss value to obtain a first video generation model.

[0202] It should be noted that the smaller the difference between the sample video and the target video, the closer the sample video is to the video required by the user, and the smaller the loss value.

[0203] Optionally, in an embodiment of the present application, the electronic device may use a back propagation algorithm to iteratively update the model parameters in the preset model based on the loss value until convergence, so as to train and obtain the first video generation model.

[0204] In this way, the electronic device can continuously calculate the loss value through the sample video and the target video, and adjust the model parameters of the preset model according to the loss value, thereby training the preset model into the first video generation model.

[0205] The embodiment of the present application provides a video generation method, Figure 10FIG. 1 shows a flow chart of a video generation method provided by an embodiment of the present application, which can be applied to electronic devices. Figure 10 As shown, the video generation method provided in the embodiment of the present application may include the following steps 10 to 15.

[0206] Step 10: The electronic device performs sparse motion coding on the motion matrix sequence and inputs it into VAE for compression to obtain a motion matrix.

[0207] Step 11: The electronic device extracts features from the noise latent space to obtain an image matrix.

[0208] Step 12: The electronic device extracts features from the text prompt to obtain text features.

[0209] Step 13: The electronic device inputs the image features into the temporal attention layer and the spatial attention layer in DiT, inputs the motion matrix into the temporal attention layer, and inputs the text features into the cross attention layer.

[0210] Step 14: The electronic device performs feature fusion processing on the image features, motion matrix and text features in the cross attention layer to obtain comprehensive features.

[0211] Step 15: The electronic device generates a video based on the comprehensive features.

[0212] Optionally, in an embodiment of the present application, the electronic device may obtain three inputs: a camera motion matrix sequence, a noise latent space, and a text prompt.

[0213] Optionally, in the embodiment of the present application, the electronic device may perform processing based on three main steps:

[0214] (1) Sparse motion coding based on Plücker coordinate representation. The motion matrix sequence is a series of motion matrices of the camera matrix. Each key frame has a 3*3 rotation matrix and a 3*1 displacement matrix, a total of 12 dimensions.

[0215] (2) Motion feature alignment and embedding based on adaptive layer normalization. The sparse motion matrix is ​​aligned with the backbone network (DiT) generated by the video, and the noise latent space is respectively sent to the time and space layers through DiT for attention calculation. After weighted calculation, the two are cross-attention calculated with the features obtained after text prompt processing to obtain the latent space distribution after conditional injection.

[0216] (3) Overall training process based on reconstruction loss and KL loss (Kullback-Leibler Divergence). The output latent space is decoded to obtain the output video sequence. Under the supervision of reconstruction loss and KL loss, the video generation with accurate motion information and the fine-tuning of the overall framework are completed.

[0217] Optionally, in an embodiment of the present application, the electronic device may train a VAE for predicting camera motion sequences. The training strategy includes a reconstruction process, including a reconstruction loss and a KL loss. The reconstruction loss aims to minimize the difference between the predicted results and the true results, while the KL loss minimizes the difference between the output distribution of the VAE and the standard normal distribution.

[0218] Optionally, in the embodiment of the present application, when using LN and MLP to fine-tune the motion latent space layer, the electronic device can freeze the network parameters except the temporal attention layer, and introduce low-rank adaptation (LoRA) when updating the self-attention to reduce the use of video memory. At the same time, the electronic device can introduce a loss function for fine-tuning, which converts p c As the camera pose motion condition input, the form of the function is shown in the following formula (9):

[0219]

[0220] in, is the loss function when the model parameter is θ, For random variables z0,c,t,∈,p m The expected value of z0 represents the initial data point or initial state, c represents the conditional variable, t represents the time step, ∈ represents the actual random noise, and p m are the parameters or hyperparameters of the model, ∈ θ (z t ,c,t,p c ) represents the model θ under given z t ,c,t,p c The output generated under the condition of , represents the model's estimate of random noise. represents the squared dinorm between the model output and the actual random noise, that is, the squared difference between them. Formula (9) trains the model by minimizing the squared difference between the model output and the actual random noise, enabling the model to better fit the data.

[0221] In the embodiment of the present application, the electronic device can effectively inject the camera motion sequence into the temporal attention layer based on the sparse motion coding module and motion feature alignment embedding module represented by Plücker coordinates, and realize the effective embedding of the motion latent space layer by converting the motion into sampling points in RGB space and combining layer normalization and multi-layer perceptron to achieve precise control of camera motion in long video generation. Figure 11 and Figure 12 , which is a schematic diagram of the effect of the video generation method provided in an embodiment of the present application.

[0222] Each of the above-mentioned method embodiments, or various possible implementation methods in each method embodiment, can be executed separately, or any two or more of them can be executed in combination with each other. The specific implementation can be determined according to actual usage requirements, and the embodiments of this application do not limit this.

[0223] The video generation method provided in the embodiment of the present application can be executed by a video generation device. In the embodiment of the present application, the video generation device provided in the embodiment of the present application is described by taking the video generation method executed by the video generation device as an example.

[0224] Figure 13 FIG. 1 shows a possible structural diagram of a video generating device involved in some embodiments of the present application. Figure 13 As shown, the video generating device 70 may include: a determination module 71 and an execution module 72 .

[0225] The determining module 71 is configured to determine the motion characteristics of the camera based on the acquired first motion parameters and the first image, wherein the first motion parameters are used to represent the rotation angle and translation distance of the camera;

[0226] The execution module 72 is configured to input the motion features and the first image determined by the determination module 71 into a first video generation model, and generate a first video based on the motion features and the image features of the first image through the first video generation model.

[0227] In one possible implementation, the above-mentioned determination module 71 is specifically used to: determine a pixel points from the first image based on the resolution of the first image through the sparse motion coding module, and determine a pixel coordinates corresponding to the a pixel points one by one; convert the a pixel coordinates into a Plücker coordinates through the sparse motion coding module to obtain a sparse motion field; compress the sparse motion field through VAE to obtain motion features.

[0228] In one possible implementation, the above-mentioned execution module 72 is specifically used to: align and embed the motion features and image features through the motion feature alignment and embedding module in the first video generation model to obtain the first feature; and generate the first video based on the first feature through the first video generation model.

[0229] In one possible implementation, the above-mentioned execution module 72 is specifically used to: split the first image through the motion feature alignment embedding module to obtain p image patches, where p is a positive integer; sparsely spatially encode the motion features through the motion feature alignment embedding module to obtain q feature patches, where q is a positive integer; and perform feature fusion on the p image patches and the q feature patches to obtain the first feature.

[0230] In a possible implementation, for example, combining Figure 13 ,like Figure 14 As shown, the video generation device provided in the embodiment of the present application further includes: an acquisition module 73. The acquisition module 73 is configured to acquire a first prompt word before inputting the motion feature and the first image into the first video generation model; the execution module 72 is further configured to perform feature extraction on the first prompt word acquired by the acquisition module 73 to obtain text features; the execution module 72 is specifically configured to input the motion feature, the first image, and the text features into the first video generation model, and generate the first video through the first video generation model based on the motion feature, the image features of the first image, and the text features.

[0231] In an embodiment of the present application, a video generation device is provided. Since the video generation device can determine the motion characteristics of the camera based on a first motion parameter representing the rotation angle and translation distance of the camera and a first image, after the motion characteristics and the first image are input into a first video generation model, the first video generation model can generate a first video based on the motion characteristics and the image characteristics of the first image. That is, the video generation device can control the trajectory of the camera based on the content of the first image and the camera's motion, and generate a first video in which the camera moves according to the trajectory. In this way, the accuracy of the camera motion trajectory control when the video generation device uses the first video generation model to generate a video is improved.

[0232] The video generation device in the embodiment of the present application can be an electronic device or a component in the electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a mobile Internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc. It can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine or a self-service machine, etc., and the embodiment of the present application does not specifically limit it.

[0233] The video generation device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0234] The video generation device provided in the embodiment of the present application can implement each process implemented in the above method embodiment. To avoid repetition, it will not be described here.

[0235] Alternatively, as Figure 15 As shown, an embodiment of the present application further provides an electronic device 1000, including a processor 1001 and a memory 1002, wherein the memory 1002 stores a program or instruction that can be run on the processor 1001, and when the program or instruction is executed by the processor 1001, the various steps of the above-mentioned video generation method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0236] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.

[0237] Figure 16 A schematic diagram of the hardware structure of an electronic device implementing an embodiment of the present application.

[0238] The electronic device 100 includes but is not limited to components such as a radio frequency unit 101 , a network module 102 , an audio output unit 103 , an input unit 104 , a sensor 105 , a display unit 106 , a user input unit 107 , an interface unit 108 , a memory 109 , and a processor 110 .

[0239] Those skilled in the art will understand that the electronic device 100 may also include a power source (such as a battery) to power each component, and the power source may be logically connected to the processor 110 through a power management system, thereby implementing functions such as charging, discharging, and power consumption management through the power management system. Figure 16 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be repeated here.

[0240] The processor 110 is configured to determine a motion characteristic of the camera based on the acquired first motion parameter and the first image, where the first motion parameter is used to represent a rotation angle and a translation distance of the camera;

[0241] The processor 110 is configured to input the motion features and the first image into a first video generation model, and generate a first video based on the motion features and the image features of the first image through the first video generation model.

[0242] An embodiment of the present application provides an electronic device that can determine the motion characteristics of a camera based on a first motion parameter and a first image used to represent the rotation angle and translation distance of the camera. After the motion characteristics and the first image are input into a first video generation model, the first video generation model can generate a first video based on the motion characteristics and the image characteristics of the first image. That is, the electronic device can control the trajectory of the camera based on the content of the first image and the motion of the camera to generate a first video in which the camera moves along the trajectory. In this way, the accuracy of the electronic device in controlling the camera motion trajectory when generating a video using the first video generation model is improved.

[0243] Optionally, the above-mentioned processor 110 is specifically used to: determine a pixel points from the first image based on the resolution of the first image through a sparse motion coding module, and determine a pixel coordinates corresponding to the a pixel points one by one; convert the a pixel coordinates into a Plücker coordinates to obtain a sparse motion field; compress the sparse motion field through a variational autoencoder VAE to obtain motion features.

[0244] Optionally, the processor 110 is specifically configured to: align and embed the motion features and the image features through a motion feature alignment and embedding module in the first video generation model to obtain a first feature; and generate a first video based on the first feature through the first video generation model.

[0245] Optionally, the above-mentioned processor 110 is specifically used to: split the first image through a motion feature alignment embedding module to obtain p image patches, where p is a positive integer; perform sparse spatial encoding on the motion features through a motion feature alignment embedding module to obtain q feature patches, where q is a positive integer; and perform feature fusion on the p image patches and the q feature patches to obtain the first feature.

[0246] Optionally, the above-mentioned processor 110 is used to obtain a first prompt word before inputting the motion features and the first image into the first video generation model; the above-mentioned processor 110 is also used to perform feature extraction on the first prompt word to obtain text features; the above-mentioned processor 110 is specifically used to input the motion features, the first image and the text features into the first video generation model, and generate the first video through the first video generation model based on the motion features, the image features of the first image and the text features.

[0247] The electronic device provided in the embodiment of the present application can implement each process implemented in the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here. The beneficial effects of various implementations in this embodiment can be specifically referred to the beneficial effects of the corresponding implementations in the above method embodiment. To avoid repetition, it will not be repeated here.

[0248] It should be understood that in an embodiment of the present application, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042, and the graphics processor 1041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 106 may include a display panel 1061, and the display panel 1061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 107 includes a touch panel 1071 and at least one of other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include two parts: a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and an operating stick, which will not be repeated here.

[0249] The memory 109 can be used to store software programs and various data. The memory 109 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 109 may include a volatile memory or a non-volatile memory, or the memory 109 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct memory bus random access memory (DRRAM). The memory 109 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.

[0250] Processor 110 may include one or more processing units. Optionally, processor 110 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 110.

[0251] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned video generation method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0252] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0253] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned video generation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0254] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0255] An embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned video generation method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0256] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0257] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0258] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. A video generation method, characterized in that: include: Determining a motion feature of the electronic device based on the acquired first motion parameter and the first image, wherein the first motion parameter is used to represent a rotation angle and a translation distance of the electronic device; The motion features and the first image are input into a first video generation model, and the first video generation model generates a first video based on the motion features and image features of the first image.

2. The method according to claim 1, characterized in that The step of determining the motion characteristics of the camera based on the acquired first motion parameter and the first image includes: Determine a pixel points from the first image based on the resolution of the first image through a sparse motion coding module, and determine a pixel coordinates corresponding to the a pixel points one by one, where a is a positive integer; The a pixel coordinates are converted into a Plücker coordinates by the sparse motion coding module to obtain a sparse motion field; The sparse motion field is compressed by a variational autoencoder (VAE) to obtain the motion feature.

3. The method according to claim 1, characterized in that The generating a first video based on the motion feature and the image feature of the first image by using the first video generation model includes: The motion feature and the image feature are aligned and embedded by a motion feature alignment and embedding module in the first video generation model to obtain a first feature; The first video is generated based on the first feature by using the first video generation model.

4. The method according to claim 3, characterized in that The step of aligning and embedding the motion feature and the image feature through the motion feature alignment embedding module in the first video generation model to obtain the first feature includes: Splitting the first image through the motion feature alignment embedding module to obtain p image patches, where p is a positive integer; The motion feature is sparsely spatially encoded through the motion feature alignment embedding module to obtain q feature patches, where q is a positive integer; The p image patches are fused with the q feature patches to obtain a first feature.

5. The method according to claim 1, characterized in that Before inputting the motion feature and the first image into a first video generation model, the method further includes: Get the first prompt word; Performing feature extraction on the first prompt word to obtain text features; The step of inputting the motion feature and the first image into a first video generation model, and generating a first video based on the motion feature and the image feature of the first image by the first video generation model, comprises: The motion features, the first image and the text features are input into a first video generation model, and the first video is generated by the first video generation model based on the motion features, the image features of the first image and the text features.

6. A video generating device, characterized in that: include: Determine modules and execute modules; The determination module is used to determine the motion characteristics of the camera based on the acquired first motion parameters and the first image, wherein the first motion parameters are used to represent the rotation angle and translation distance of the camera; The execution module is used to input the motion features determined by the determination module and the first image into a first video generation model, and generate a first video based on the motion features and image features of the first image through the first video generation model.

7. The device according to claim 6, characterized in that The determining module is specifically used for: Determine a pixel points from the first image based on the resolution of the first image through a sparse motion coding module, and determine a pixel coordinates corresponding to the a pixel points one by one, where a is a positive integer; The a pixel coordinates are converted into a Plücker coordinates by the sparse motion coding module to obtain a sparse motion field; The sparse motion field is compressed by a variational autoencoder (VAE) to obtain the motion feature.

8. The device according to claim 6, characterized in that The execution module is specifically used for: The motion feature and the image feature are aligned and embedded by a motion feature alignment and embedding module in the first video generation model to obtain a first feature; and The first video is generated based on the first feature by using the first video generation model.

9. The device according to claim 8, characterized in that The execution module is specifically used for: Splitting the first image through the motion feature alignment embedding module to obtain p image patches, where p is a positive integer; The motion feature is sparsely spatially encoded through the motion feature alignment embedding module to obtain q feature patches, where q is a positive integer; The p image patches are fused with the q feature patches to obtain a first feature.

10. The device according to claim 6, characterized in that The device also includes: an acquisition module; The acquisition module is used to acquire a first prompt word before inputting the motion feature and the first image into the first video generation model; The execution module is further used to extract features from the first prompt word obtained by the obtaining module to obtain text features; The execution module is specifically used to input the motion features, the first image and the text features into a first video generation model, and generate the first video based on the motion features, the image features of the first image and the text features through the first video generation model.

Citation Information

Cited By

  • Video generation method and device with controllable mirror operation mode, and electronic equipment

    CN121099153A

  • Video processing method and device, storage medium and program product

    CN122372769A