Video generation method and apparatus, device and medium

By acquiring the three-dimensional spatial trajectory information of the target area in image-to-video technology and generating motion videos of the target object using a target network model, the problem of inaccurate motion control in local image regions in existing technologies is solved, and the realism and dynamic viewing experience of the video are improved.

WO2026092335A1PCT designated stage Publication Date: 2026-05-07BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2025-10-24
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately control the movement of local image regions when generating videos, resulting in videos with poor dynamic viewing and low realism, failing to meet diverse user needs.

Method used

By responding to the region selection operation of the target reference image, the trajectory information of the target region is obtained as a three-dimensional spatial representation. The motion video of the target object is generated using a preset target network model. Users can select any target region that needs to move and guide the target network model to generate the motion video of the target object.

Benefits of technology

It achieves controllable effects for moving objects and motion trajectories, enhances the three-dimensional spatial motion effects and realism of videos, meets user needs, and enriches the effects of image-to-video generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025129968_07052026_PF_FP_ABST
    Figure CN2025129968_07052026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a video generation method and apparatus, a device and a medium. The method comprises: in response to a region selection operation for a target reference image and on the basis of the region selection operation, determining a target region; acquiring target trajectory information of the motion of the target region in the target reference image, the target trajectory information being trajectory information represented by means of three-dimensional spatial information; and, on the basis of the target region, the target reference image and the target trajectory information, using a preset target network model to generate a motion video of a target object, the target object being the target region or an object contained in the target region.
Need to check novelty before this filing date? Find Prior Art

Description

Video generation methods, apparatus, equipment and media

[0001] Cross-reference to related applications

[0002] This application claims priority to Chinese Patent Application No. 202411554890.0, filed on November 1, 2024, entitled "Video Generation Method, Apparatus, Device and Medium", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure relates to the field of computer technology, and in particular to a video generation method, apparatus, device, and medium. Background Technology

[0004] In scenarios requiring image processing, such as computer vision and multimedia editing, image-based video generation technology has gradually become a research hotspot. For example, by inputting an image, a video showing changes in a specific object within that image can be generated. Summary of the Invention

[0005] This disclosure provides a video generation method, apparatus, device, and medium.

[0006] This disclosure provides a video generation method, the method comprising: responding to a region selection operation on a target reference image, determining a target region based on the region selection operation; acquiring target trajectory information of the target region moving in the target reference image; wherein the target trajectory information is trajectory information represented by three-dimensional spatial information; generating a motion video of a target object using a preset target network model based on the target region, the target reference image, and the target trajectory information; wherein the target object is the target region or an object contained in the target region.

[0007] Optionally, the three-dimensional spatial information includes three-dimensional optical flow information, and the step of obtaining the target trajectory information of the target region moving in the target reference image includes: when receiving expected trajectory information set for the target region, generating three-dimensional optical flow information corresponding to the expected trajectory information, and obtaining the target trajectory information of the target region moving in the target reference image based on the three-dimensional optical flow information corresponding to the expected trajectory information; when not receiving expected trajectory information set for the target region, using preset three-dimensional optical flow information as the target trajectory information of the target region moving in the target reference image; wherein, the preset three-dimensional optical flow information is characterized by a specific value.

[0008] Optionally, the target network model is obtained through the following steps: acquiring target video samples; wherein the target video samples are videos with moving objects in the foreground and a static background; determining motion regions in the target video samples and acquiring the three-dimensional optical flow information corresponding to the motion regions; wherein the motion regions are regions containing moving objects; and adjusting the parameters of a preset generation model based on the target video samples, the motion regions, and the three-dimensional optical flow information corresponding to the motion regions, so as to obtain the target network model based on the preset generation model with adjusted parameters.

[0009] Optionally, obtaining the target video sample includes: obtaining a first video sample; wherein the number of the first video samples is multiple; obtaining optical flow information corresponding to the first video sample, and based on the optical flow information corresponding to the first video sample, filtering from the multiple first video samples to obtain a second video sample; wherein the second video sample is a non-still video; determining target mapping information corresponding to a video frame group in the second video sample based on the optical flow information corresponding to the second video sample; wherein the video frame group includes two video frame images with a preset interval; filtering from the second video samples to obtain a third video sample based on the target mapping information, and obtaining a target video sample based on the third video sample; wherein the target video sample is the third video sample, or the target video sample is a sample obtained by segmenting the third video sample.

[0010] Optionally, the optical flow information corresponding to the second video sample includes dense optical flow information corresponding to the video frame group in the second video sample; determining the target mapping information corresponding to the video frame group in the second video sample based on the optical flow information corresponding to the second video sample includes: obtaining multiple sparse optical flow points using a preset uniform sampling strategy based on the dense optical flow information corresponding to the video frame group in the second video sample; determining successfully matched keypoint pairs corresponding to the video frame group in the second video sample based on the multiple sparse optical flow points; filtering out keypoint pairs corresponding to the background of the video frames in the video frame group from the successfully matched keypoint pairs, and obtaining the target mapping information corresponding to the video frame group in the second video sample based on the mapping information of the keypoint pairs corresponding to the background of the video frames.

[0011] Optionally, determining the motion region in the target video sample includes: performing object segmentation processing on the video frame image of the target video sample to obtain an object region; determining the two-dimensional optical flow information corresponding to the video frame image of the target video sample based on the video frame image of the target video sample and the subsequent frame image adjacent to the video frame image of the target video sample; and determining the motion region in the video frame image of the target video sample based on the object region and the two-dimensional optical flow information.

[0012] Optionally, determining the motion region in the video frame image of the target video sample based on the object region and the two-dimensional optical flow information includes: filtering out a first region from the object region based on the two-dimensional optical flow information; wherein the optical flow value of the two-dimensional optical flow information corresponding to the first region is greater than a preset threshold; determining a second region from the video frame image of the target video sample based on the two-dimensional optical flow information; wherein the second region is a region other than the first region, and the optical flow value of the two-dimensional optical flow information corresponding to the second region is greater than the preset threshold; and obtaining the motion region in the video frame image of the target video sample based on the first region and the second region.

[0013] Optionally, obtaining the three-dimensional optical flow information corresponding to the motion region includes: obtaining the distance gradient field information corresponding to the motion region; obtaining the two-dimensional optical flow information corresponding to the motion region, and obtaining the three-dimensional optical flow information corresponding to the motion region based on the two-dimensional optical flow information and the distance gradient field information.

[0014] Optionally, obtaining the distance gradient field information corresponding to the motion region includes: obtaining the distance between the pixels of the motion region and the edge of the motion region; and determining the distance gradient field information corresponding to the motion region based on the distances between the pixels of the motion region and the distances between the adjacent pixels of the pixels.

[0015] Optionally, the preset generation model includes a denoising network; wherein the first frame image information of the target video sample, the mask image corresponding to the motion region, and the three-dimensional optical flow information corresponding to the motion region are introduced into the denoising network by splicing input channels.

[0016] Optionally, the method further includes: when initializing the preset generation model, setting the weights of the first newly added channel and the second newly added channel of the denoising network to zero; wherein, the first newly added channel is used to input the mask map corresponding to the motion region, and the second newly added channel is used to input the three-dimensional optical flow information corresponding to the motion region.

[0017] This disclosure also provides a video generation apparatus, comprising: a region determination module, configured to determine a target region based on a region selection operation on a target reference image in response to such operation; a trajectory acquisition module, configured to acquire target trajectory information of the target region moving in the target reference image; wherein the target trajectory information is trajectory information characterized by three-dimensional spatial information; and a video generation module, configured to generate a motion video of a target object using a preset target network model based on the target region, the target reference image, and the target trajectory information; wherein the target object is the target region or an object contained within the target region.

[0018] This disclosure also provides an electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement the video generation method provided in this disclosure.

[0019] This disclosure also provides a computer-readable storage medium storing a computer program for performing the video generation method provided in this disclosure.

[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0021] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0022] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 is a flowchart illustrating a video generation method provided in an embodiment of this disclosure;

[0024] Figure 2 is a schematic diagram of an image interaction provided in an embodiment of this disclosure;

[0025] Figure 3 is a schematic diagram of determining a motion area according to an embodiment of this disclosure;

[0026] Figure 4 is a schematic diagram of model training provided in an embodiment of this disclosure;

[0027] Figure 5 is a schematic diagram of the structure of a preset generation model provided in an embodiment of this disclosure;

[0028] Figure 6 is a schematic diagram of the structure of a video generation device provided in an embodiment of this disclosure;

[0029] Figure 7 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0030] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0031] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0032] In scenarios requiring image processing, such as computer vision and multimedia editing, image-based video generation technology has gradually become a research hotspot. For example, by inputting an image, a video showing changes in a specific object within that image can be generated. However, the inventors have found that the image-to-video generation methods used in related technologies are inadequate. For instance, they require global control of image content through text or images to generate motion, making it difficult to accurately control the motion of local areas within the image, nor can they directionally control the motion trajectory of local areas. Furthermore, the generated videos lack a good sense of dynamic movement and realism. Therefore, related technologies cannot meet the diverse needs of users, thus necessitating new image-to-video technologies. To improve upon these problems, this disclosure provides a video generation method, apparatus, device, and medium, which are described in detail below.

[0033] The technical solution provided in this disclosure can respond to a region selection operation on a target reference image, determine a target region based on the region selection operation, and further obtain target trajectory information represented by three-dimensional spatial information. Based on the target region, the target reference image, and the target trajectory information, a motion video of the target object is generated using a preset target network model, where the target object is the target region or an object contained within the target region. Through this method, users can select any target region in the image that requires motion according to their needs. Based on this, the target trajectory information guides the target network model to generate a motion video of the target object, achieving the effect of controllable motion objects and controllable motion trajectories. Furthermore, since the target trajectory information is represented by three-dimensional spatial information, it helps to present the motion effect of the target object in three-dimensional space, better meeting user needs, enriching the image-generated video effect, and enhancing the video's appeal.

[0034] Figure 1 is a flowchart illustrating a video generation method according to an embodiment of this disclosure. This method can be executed by a video generation device, which can be implemented using software and / or hardware and is generally integrated into an electronic device. As shown in Figure 1, the method mainly includes the following steps S102 to S106:

[0035] Step S102: In response to the region selection operation for the target reference image, the target region is determined based on the region selection operation.

[0036] The target reference image can be an image input by the user, and the content contained in the target reference image is not limited here. For example, the region selection operation includes a smearing operation or a bounding box operation. In practical applications, a region selection control such as a brush can be provided to the user, allowing the user to perform a region selection operation through this control. For example, the user can perform a smearing or bounding box operation on the target reference image using a brush, or the user can perform a smearing or bounding box operation by touching the screen with their finger, drawing lines, etc., thereby determining the target region. In some embodiments, the smeared area corresponding to the smearing operation is the target region, and the bounding box area corresponding to the bounding box operation is the target region. It should be noted that the target region can be any region, such as the entire object, a part of an object, or a non-object region; there are no limitations here.

[0037] Step S104: Obtain the target trajectory information of the target region moving in the target reference image; wherein, the target trajectory information is trajectory information represented by three-dimensional spatial information.

[0038] In practical applications, target trajectory information can be determined based on a user-specified trajectory. If the user does not specify a trajectory, the target trajectory information can be default information or information automatically generated based on the target region in the target reference image; there are no restrictions here. Target trajectory information represented by three-dimensional spatial information is more helpful in controlling the model to generate motion trajectories with three-dimensional visual effects, effectively improving the realism of the final generated video and making it easier to present users with a good dynamic viewing experience.

[0039] In some implementation examples, the three-dimensional spatial information includes three-dimensional optical flow information. That is, the target trajectory information can be represented by the three-dimensional optical flow information. Based on this, the target trajectory information of the target region moving in the target reference image can be obtained by referring to (1) and (2) below:

[0040] (1) Upon receiving the expected trajectory information set for the target region, generate the three-dimensional optical flow information corresponding to the expected trajectory information, and obtain the target trajectory information of the target region moving in the target reference image based on the three-dimensional optical flow information corresponding to the expected trajectory information. For example, the three-dimensional optical flow information corresponding to the expected trajectory information can be directly used as the target trajectory information.

[0041] For example, users can draw a desired trajectory in the form of lines on a target area using trajectory drawing controls or by tapping with their fingers to obtain two-dimensional trajectory information. Alternatively, users can be provided with trajectory input controls for the X, Y, or Z directions, allowing them to input information in each direction numerically to obtain three-dimensional trajectory information. In practical applications, any one of the above methods can be chosen, or a combination of both. For example, the user can first draw a two-dimensional trajectory and then input the depth information in the Z direction numerically to obtain a three-dimensional trajectory. The specific method for obtaining the desired trajectory information can be flexibly set and is not limited here.

[0042] For ease of understanding, please refer to Figure 2, which illustrates an image interaction diagram. In Figure 2, the left side (a) corresponds to the target reference image, and the right side (b) corresponds to the marked smeared area and the expected trajectory corresponding to that area. That is, the arm area of ​​the cartoon object in Figure 2 is the smeared area, also known as the target area. Arrows pointing upwards are drawn on both sides of the arms; these trajectories can also be drawn by the user, indicating the desired upward movement of the cartoon object's arms. It is understood that the expected trajectory drawn by the user on the image is usually a two-dimensional trajectory. Based on this, X-dimensional and Y-dimensional information can be obtained. For example, Z-dimensional information can be obtained by automatically filling preset values ​​or by directly acquiring the user's input information for the Z-dimensional dimension; this is not limited here. Through the above methods, three-dimensional expected trajectory information can be obtained, thereby corresponding to the three-dimensional optical flow information. The method of converting trajectory information into optical flow information can refer to related technologies, such as calculating the corresponding optical flow information based on key pixels in the trajectory information. For example, the starting point, ending point, and trajectory of a key pixel can be determined based on the trajectory information, thereby obtaining the optical flow information corresponding to that key pixel. The above is only an illustrative example and is not limited here.

[0043] (2) In the absence of receiving the expected trajectory information set for the target area, the preset three-dimensional optical flow information is used as the target trajectory information for the target area to move in the target reference image; wherein, the preset three-dimensional optical flow information can be characterized by specific values, such as setting preset optical flow values ​​based on X, Y and Z dimensions respectively, which can be set flexibly and is not restricted here.

[0044] Step S106: Based on the target region, target reference image and target trajectory information, generate a motion video of the target object using a preset target network model; wherein, the target object is the target region or an object contained in the target region.

[0045] The target network model can be implemented using a generative model, or in some specific implementation examples, a diffusion model or a network model containing a denoising network can be used; there are no restrictions here. For example, a motion video of the target object can be generated using a preset target network model based on the target region, target reference image, target trajectory information, and randomly generated noise data. In practical applications, the target object can be the complete target region or the main object contained within the target region; the specific settings can be flexible.

[0046] In this way, users can select any target area in the image that they want to move according to their needs. Based on this, the target trajectory information guides the target network model to generate a motion video of the target object, which can achieve the effect of controllable motion object and controllable motion trajectory. Furthermore, since the target trajectory information is represented by three-dimensional spatial information, it helps to present the motion effect of the target object in three-dimensional space, better meet the user's needs, enrich the image-generated video effect, and enhance the video's appeal.

[0047] For ease of understanding, this disclosure provides a method for training the target network model. Exemplarily, the target network model is obtained through the following steps A to C:

[0048] Step A: Obtain the target video sample; the target video sample is a video in which a moving object exists in the foreground and the background remains stationary. This video is also known as a video with a stable shot.

[0049] In some implementation examples, step A can be performed as follows: Steps A1 to A4

[0050] Step A1: Obtain the first video sample; there are multiple first video samples. The first video sample is the initially obtained video.

[0051] Step A2: Obtain the optical flow information corresponding to the first video sample; based on the optical flow information corresponding to the first video sample, select the second video sample from multiple first video samples; wherein the second video sample is a non-still video.

[0052] To improve processing efficiency and reduce computational resource consumption, in some specific implementation examples, multiple video frame groups can be extracted from the first video sample. Each video frame group contains a first video frame image and a second video frame image, which are two frames separated by a preset number of frames. Then, the optical flow information corresponding to each of the multiple video frame groups is obtained and used as the optical flow information corresponding to the first video sample. In practical applications, the number of video frame groups and the preset number of frames between the first and second video frame images can be flexibly set according to requirements. For example, setting the number of video frame groups to 4, with a 4-frame interval between the first and second video frame images, allows for the uniform extraction of 4 video frame groups from the first video sample. Each video frame group contains two frames separated by 4 frames, meaning a total of 8 frames can be extracted from the first video sample for analysis to obtain the optical flow information corresponding to each video frame group. Subsequently, based on the optical flow information of the multiple video frame groups corresponding to the first video sample, videos based on still images can be excluded, thus obtaining the second video sample. For example, for each first video sample, it is determined whether the maximum absolute value of the optical flow of the multiple video frame groups corresponding to the first video sample is greater than a preset value. If it is, the first video sample is used as the selected second video sample; otherwise, the first video sample is filtered out. In this way, non-still videos can be filtered out efficiently and reliably.

[0053] Step A3: Based on the optical flow information corresponding to the second video sample, determine the target mapping information corresponding to the video frame group in the second video sample; wherein, the video frame group includes two video frame images with a preset interval, and for example, the target mapping information can be represented by a homography matrix.

[0054] In some specific implementation examples, the optical flow information corresponding to the second video sample includes the dense optical flow information corresponding to the video frame group in the second video sample. It should be noted that the video frame group in the first video sample (which can be simply referred to as the first video frame group) and the video frame group in the second video sample (which can be simply referred to as the second video frame group) can be different. Specifically, the number of the first video frame group and the second video frame group and / or the interval between the two video frame images contained therein can be different. For example, each pair of adjacent frame images in the second video sample can be considered as a second video frame group. Assuming the second video sample has a total of N frames, the number of second video frame groups is N-1, and the preset interval between the two video frame images in the second video frame group is zero. The above is one implementation example and should not be considered a limitation. In practical applications, it can be flexibly set, and no limitation is imposed here. Based on the above, step A3 can be performed with reference to steps A3.1 to A3.3 as follows:

[0055] Step A3.1: Based on the dense optical flow information corresponding to the video frame group in the second video sample, multiple sparse optical flow points are obtained using a preset uniform sampling strategy. In this way, dense optical flow can be converted into sparse optical flow on grid points. For example, the uniform sampling strategy can be equidistant sampling, performing uniform sampling on both the long and short sides with a specified number of points. For instance, performing uniform sampling on both the long and short sides with a total of 30 points yields a total of 900 sparse optical flow points.

[0056] Step A3.2: Based on multiple sparse optical flow points, determine the key point pairs that are successfully matched for the video frame groups in the second video sample.

[0057] This disclosure embodiment can transform sparse optical flow into the form of successfully matched keypoints. Specifically, for each sparse optical flow point obtained from the above steps, its current coordinates are denoted as (x, y), and its corresponding optical flow value is (u, v). Then, a point pair of the form {(x, y), (x+u, y+v)} is defined as a successfully matched keypoint pair. In other words, assuming the pixel coordinates in the previous frame of a video frame group are (x, y), then the coordinates of the pixel that successfully matches it in the subsequent frame of the video frame group are (x+u, y+v).

[0058] Step A3.3 involves filtering out keypoint pairs corresponding to the background of video frames within the video frame group from the successfully matched keypoint pairs, and obtaining the target mapping information corresponding to the video frame group in the second video sample based on the mapping information of the keypoint pairs corresponding to the background of the video frames. For example, the mapping information of the keypoint pairs corresponding to the background of the video frames can be directly used as the target mapping information.

[0059] In some specific implementation examples, based on successfully matched keypoint pairs, the RANSAC algorithm can be used to implement step A3.3 above. That is, the RANSAC algorithm filters out keypoint pairs whose positions are within moving foreground objects from the successfully matched keypoint pairs, obtaining the mapping information of the keypoint pairs corresponding to the background of the video frame, represented in the form of a homography matrix. This method effectively ensures that the obtained target mapping information accurately reflects the background movement of the second video sample.

[0060] Using the above method, the target mapping information (such as homography matrix) of each of the multiple video frame groups corresponding to the second video sample can be obtained and stored for further processing.

[0061] Step A4: Based on the target mapping information, a third video sample is obtained from the second video sample, and a target video sample is obtained based on the third video sample; wherein, the target video sample is the third video sample, or the target video sample is a sample obtained by segmenting the third video sample.

[0062] The second video sample is a non-static video. Based on the target mapping information of each of the multiple video frame groups corresponding to the second video sample, the background of the video can be kept stationary by using the principle of affine transformation and a preset threshold. The third video sample obtained by the screening has a moving object in the foreground and the background remains stationary. For example, the target mapping information can be represented by a homography matrix, referred to here as the target homography matrix, assuming a size of M*M, such as a 3*3 matrix. By comparing the target homography matrix with a preset standard homography matrix, for example, the standard homography matrix is ​​a matrix where the diagonal elements are all 1 and the remaining elements are 0. By comparing the numerical difference between each element of the target homography matrix and the element at the same position in the standard homography matrix, the third video sample can be screened from the second video sample based on the difference corresponding to each of the nine elements of the homography matrix and the target threshold. For example, if the fused value (such as the sum, average, etc.) of the absolute values ​​of the differences corresponding to the nine elements is less than the target threshold, it indicates that the target homography matrix indicates that the background is roughly stationary. Using the above method, it is possible to reasonably and reliably filter out videos with static backgrounds from non-static videos, thus obtaining a third video sample.

[0063] The third video sample can then be used directly as the target video sample, or it can be segmented according to requirements. For example, if the third video sample is M frames, but the model only needs to output N frames of video, then the third video sample can be segmented to obtain the target video sample.

[0064] Step B involves identifying the motion regions within the target video sample and acquiring the corresponding 3D optical flow information; the motion region is a region containing moving objects. In some specific implementation examples, steps 1 through 3 can be used to determine the motion regions in the target video sample:

[0065] Step 1: Perform object segmentation on the video frame images of the target video sample to obtain the object regions. Specific methods for object segmentation can be found in relevant technologies and will not be elaborated here. In practical applications, object regions can be represented using object masks. In practice, object segmentation can be performed on each video frame image of the target video sample. For ease of processing, object regions in subsequent frames can be quickly determined based on the object segmentation results of previous frames using methods such as object tracking.

[0066] Step 2: Based on the video frame images of the target video sample and the subsequent frame images adjacent to the video frame images of the target video sample, determine the two-dimensional optical flow information corresponding to the video frame images of the target video sample. In other words, for each video frame image (excluding the last frame image) in the target video sample, the two-dimensional optical flow information between two adjacent frame images can be determined based on that video frame image and the subsequent frame images adjacent to that video frame image.

[0067] Step 3: Based on the object region and two-dimensional optical flow information, determine the motion region in the video frame image of the target video sample. It can be understood that object segmentation primarily identifies salient objects, and then, combined with two-dimensional optical flow information, motion objects can be filtered out from the salient objects. Furthermore, the two-dimensional optical flow information can be used to further determine regions with optical flow values ​​greater than a threshold, which can then be added as motion regions.

[0068] In some specific implementation examples, step 3 can be performed as follows: Steps 3.1 to 3.3:

[0069] Step 3.1: Based on the two-dimensional optical flow information, a first region is selected from the object region; wherein the optical flow value of the two-dimensional optical flow information corresponding to the first region is greater than a preset threshold. In this way, a moving first region can be selected from the object region.

[0070] Step 3.2: Based on two-dimensional optical flow information, determine a second region from the video frame image of the target video sample; wherein, the second region is the region other than the first region, and the optical flow value of the two-dimensional optical flow information corresponding to the second region is greater than a preset threshold. The second region is a motion region supplemented on the basis of the moving object; although this region does not belong to the object, it also exhibits motion.

[0071] Step 3.3: Based on the first and second regions, obtain the motion regions in the video frame image of the target video sample. In practical applications, both the first and second regions can be directly used as motion regions. Alternatively, post-processing operations such as erosion can be performed on the first and second regions, and the post-processed first and second regions can be used as motion regions. Specifically, the first and second regions each correspond to a label value to distinguish different regions. Then, erosion operations are performed on the first and second regions, which can be directly based on the masks corresponding to each region. Through the above methods, it can be effectively ensured that different regions are not connected.

[0072] For ease of understanding, refer to Figure 3 for a schematic diagram of motion region determination. Regions labeled 1, 2, 3, and 4 are object regions obtained through object segmentation. Region 5 is a region where the optical flow value determined based on 2D optical flow information is greater than the optical flow threshold. Region 5 also indicates region 6, which is also an object region, but one that was missed during object segmentation. That is, the second region obtained based on 2D optical flow information will also cover the inaccurately segmented object region. Based on the aforementioned steps, the target region is the region corresponding to 2, 4, 5, and 6. Region 5 can be considered as excluding regions 2, 4, and 6. In other words, even for cases of inaccurate object segmentation, 2D optical flow information can provide reasonable supplementation. Alternatively, the target region can be described as the region corresponding to 2, 4, and 5, where region 5 includes the unsegmented region 6. Both of these classification methods are acceptable, and the appropriate method can be chosen flexibly. Furthermore, it should be noted that regions with optical flow values ​​greater than the optical flow threshold determined based on two-dimensional optical flow information do not necessarily include the object detection region. For example, the small, independent region corresponding to label 5 in Figure 3 does not include the object detection region, but it can still be considered as a motion region. Using the above method, motion regions can be determined relatively comprehensively and accurately.

[0073] In some implementations after determining the motion region, the three-dimensional optical flow information corresponding to the motion region can be obtained by referring to the following steps 1) and 2):

[0074] Step 1) Obtain the distance gradient field information corresponding to the moving region. Specifically, the corresponding distance field can be calculated for each moving region to obtain the distance between the pixels of the moving region and the edge of the moving region; based on the distances corresponding to the pixels of the moving region and the distances corresponding to the adjacent pixels of the pixels, the distance gradient field information corresponding to the moving region can be determined, and the direction pointing to the center of the object can be determined based on the distance gradient field information.

[0075] Step 2) Obtain the two-dimensional optical flow information corresponding to the moving region, and based on the two-dimensional optical flow information and the distance gradient field information, obtain the three-dimensional optical flow information corresponding to the moving region. The X and Y directions can be determined based on the two-dimensional optical flow information, and then the Z direction information can be determined based on the distance gradient field information. The Z direction is the direction pointing towards the center of the object. For the X and Y directions, the average optical flow of the entire moving region can be directly calculated as the component of planar motion. Specifically, the average value of the optical flow values ​​of all pixels in the moving region in the X direction can be used as the X motion component value corresponding to that moving region, and the average value of the optical flow values ​​of all pixels in the moving region in the Y direction can be used as the Y motion component value corresponding to that moving region. When an object moves away from the camera, the pixels in the area occupied by the object will shrink towards the center of the object; similarly, when the object approaches the camera, the pixels in the area occupied by the object will spread towards the edge of the object. Based on this principle, for the Z direction (i.e., the direction of the distance gradient field), the average value of the optical flow values ​​of all pixels in the moving region in the Z direction can be used as the Z motion component value corresponding to that moving region.

[0076] In practical applications, three-dimensional optical flow information of each motion region in the target video sample can be obtained. This three-dimensional optical flow information can accurately present the motion direction and intensity of the motion region in three-dimensional space.

[0077] Step C: Based on the target video sample, the motion region, and the corresponding 3D optical flow information of the motion region, adjust the parameters of the preset generation model to obtain the target network model based on the preset generation model with adjusted parameters.

[0078] In practical applications, the structure of the preset generative model is the same as that of the target network model, with the main difference being the parameters. When training the preset generative model, the parameters of the preset generative model can be adjusted based on the differences between the video output by the preset generative model and the target video samples, thereby obtaining a preset generative model that meets the requirements (i.e., the preset generative model with adjusted parameters). Then, the preset generative model with adjusted parameters can be directly used as the target network model.

[0079] For ease of understanding, a model training diagram as shown in Figure 4 can be used as an example. This diagram focuses on illustrating the methods for acquiring target video samples, motion regions, and 3D optical flow information. Ultimately, by training a preset generation model and adjusting its parameters based on the differences between the model's output video and the target video samples, a target network model whose output results meet expectations can be obtained. The specific processes for acquiring target video samples, motion regions, and 3D optical flow information can be found in the aforementioned related content and will not be repeated here.

[0080] Furthermore, the preset generation model provided in this embodiment includes a denoising network; wherein, the first frame image information of the target video sample, the mask map corresponding to the motion region, and the three-dimensional optical flow information corresponding to the motion region are introduced into the denoising network by concatenating input channels. The first frame image information of the target video sample can be the encoded result of the first frame image of the target video sample. For ease of understanding, a schematic diagram of the structure of a preset generation model shown in Figure 5 illustrates the encoder, denoising network, and decoder. The denoising network can be implemented using a Unet network, and this is not limited thereto. The target video sample can be input to the encoder to obtain the encoder's encoding result for the target video sample. The encoded result of the first frame image of the target video sample, the scale-processed mask map corresponding to the motion region, and the three-dimensional optical flow information corresponding to the motion region can be input into the denoising network by concatenating input channels to obtain the denoising result of the denoising network. This result is then decoded based on the decoder to obtain the output video. The model parameters can be adjusted based on the information difference between the output video and the target video sample to make the model's output video as close as possible to the target video sample, so that the final model can output a video result that meets expectations. It should be noted that Figure 5 mainly illustrates the differences between the input of the denoising network and the existing network. In practical applications, the denoising network may also have noise data input, and the generative model may have other network modules, which will not be shown one by one in Figure 5.

[0081] By using the above-mentioned input channel splicing method, the mask image corresponding to the motion region and the newly added control information such as the three-dimensional optical flow information corresponding to the motion region can be introduced into the original model framework. It is understood that even if the existing image-generated video model may also have a mask image (such as one used to indicate the frame to be generated) as input, considering the subsequent zero-impact initialization, this embodiment of the disclosure will not input the mask image corresponding to the motion region on the original basis, but will add an input channel for the mask image corresponding to the motion region.

[0082] For example, when initializing the preset generated model, the weights of the first newly added channel and the second newly added channel of the denoising network are set to zero; wherein, the first newly added channel is used to input the mask map corresponding to the motion region, and the second newly added channel is used to input the three-dimensional optical flow information corresponding to the motion region.

[0083] In other words, the weights corresponding to the first and second newly added channels are not randomly initialized. Instead, they are set to zero during initialization, such as by multiplying control information like the mask image and 3D optical flow information corresponding to the motion region by zero. This achieves zero impact, maximizing the preservation of the model's original image-to-video conversion capability. The weights are then gradually adjusted during subsequent model training. Through this method, the ability to add region interaction control can be added without significantly affecting the original model's image-to-video conversion capability. After zero-impact initialization, for the same input and arbitrary control information (the mask image and 3D optical flow information corresponding to the motion region), the output format of the original image-to-video model can be consistent with the method provided in this embodiment. In summary, this embodiment can achieve new image-to-video effects without significant adjustments to the existing image-to-video model and is easy to implement.

[0084] In summary, the target network model used in the video generation method provided in this embodiment is easy to implement, and users can select any target region in the image that needs to move according to their needs. Based on this, the target network model is guided by the target trajectory information to generate a motion video of the target object, which can achieve the effect of controllable motion object and controllable motion trajectory. Furthermore, since the target trajectory information is represented by three-dimensional spatial information, it helps to present the motion effect of the target object in three-dimensional space, better meet the user's needs, enrich the image-generated video effect, and enhance the video's appeal.

[0085] Corresponding to the aforementioned video generation method, this disclosure further provides a video generation apparatus. Figure 6 is a schematic diagram of the structure of a video generation apparatus provided in this disclosure. This apparatus can be implemented by software and / or hardware, and is generally integrated into an electronic device. As shown in Figure 6, the video generation apparatus includes:

[0086] The region determination module 602 is used to determine the target region based on the region selection operation in response to a region selection operation on the target reference image;

[0087] The trajectory acquisition module 604 is used to acquire the target trajectory information of the target region moving in the target reference image; wherein, the target trajectory information is trajectory information represented by three-dimensional spatial information;

[0088] The video generation module 606 is used to generate a motion video of the target object based on the target region, the target reference image, and the target trajectory information using a preset target network model; wherein the target object is the target region or an object contained in the target region.

[0089] With the above-mentioned device, users can select any target area in the image that they want to move according to their needs. Based on this, the target trajectory information guides the target network model to generate a motion video of the target object, which can achieve the effect of controllable motion object and controllable motion trajectory. Furthermore, since the target trajectory information is represented by three-dimensional spatial information, it helps to present the motion effect of the target object in three-dimensional space, better meet the user's needs, enrich the image-generated video effect, and enhance the video's appeal.

[0090] In some embodiments, the three-dimensional spatial information includes three-dimensional optical flow information, and the trajectory acquisition module 604 is specifically used to: generate three-dimensional spatial information corresponding to the expected trajectory information when receiving expected trajectory information set for the target region, and obtain target trajectory information of the target region moving in the target reference image based on the three-dimensional optical flow information corresponding to the expected trajectory information; and use preset three-dimensional optical flow information as the target trajectory information of the target region moving in the target reference image when not receiving expected trajectory information set for the target region; wherein the preset three-dimensional optical flow information is characterized by a specific value.

[0091] In some embodiments, the apparatus further includes a model acquisition module for obtaining a target network model through the following steps: acquiring a target video sample; wherein the target video sample is a video in which a moving object exists in the foreground and the background remains stationary; determining a motion region in the target video sample and acquiring the three-dimensional optical flow information corresponding to the motion region; wherein the motion region is a region containing a moving object; and adjusting the parameters of a preset generation model based on the target video sample, the motion region, and the three-dimensional optical flow information corresponding to the motion region, so as to obtain a target network model based on the preset generation model with adjusted parameters.

[0092] In some implementations, the model acquisition module is specifically used for: acquiring a first video sample; wherein the number of the first video samples is multiple; acquiring optical flow information corresponding to the first video sample, and based on the optical flow information corresponding to the first video sample, selecting a second video sample from the multiple first video samples; wherein the second video sample is a non-still video; determining target mapping information corresponding to a video frame group in the second video sample based on the optical flow information corresponding to the second video sample; wherein the video frame group includes two video frame images with a preset interval; selecting a third video sample from the second video samples based on the target mapping information, and obtaining a target video sample based on the third video sample; wherein the target video sample is the third video sample, or the target video sample is a sample obtained by segmenting the third video sample.

[0093] In some implementations, the optical flow information corresponding to the second video sample includes dense optical flow information corresponding to the video frame group in the second video sample; the model acquisition module is specifically used to: obtain multiple sparse optical flow points based on the dense optical flow information corresponding to the video frame group in the second video sample using a preset uniform sampling strategy; determine successfully matched key point pairs corresponding to the video frame group in the second video sample based on the multiple sparse optical flow points; filter out key point pairs corresponding to the background of the video frame in the video frame group from the successfully matched key point pairs, and obtain target mapping information corresponding to the video frame group in the second video sample based on the mapping information of the key point pairs corresponding to the background of the video frame.

[0094] In some implementations, the model acquisition module is specifically used to: perform object segmentation processing on the video frame image of the target video sample to obtain an object region; determine the two-dimensional optical flow information corresponding to the video frame image of the target video sample based on the video frame image of the target video sample and the subsequent frame image adjacent to the video frame image of the target video sample; and determine the motion region in the video frame image of the target video sample based on the object region and the two-dimensional optical flow information.

[0095] In some implementations, the model acquisition module is specifically used to: filter out a first region from the object region based on the two-dimensional optical flow information; wherein the optical flow value of the two-dimensional optical flow information corresponding to the first region is greater than a preset threshold; determine a second region from the video frame image of the target video sample based on the two-dimensional optical flow information; wherein the second region is a region other than the first region, and the optical flow value of the two-dimensional optical flow information corresponding to the second region is greater than the preset threshold; and obtain the motion region in the video frame image of the target video sample based on the first region and the second region.

[0096] In some implementations, the model acquisition module is specifically used to: acquire distance gradient field information corresponding to the motion region; acquire two-dimensional optical flow information corresponding to the motion region, and obtain three-dimensional optical flow information corresponding to the motion region based on the two-dimensional optical flow information and the distance gradient field information.

[0097] In some implementations, the model acquisition module is specifically used to: obtain the distance between the pixels of the motion region and the edge of the motion region; and determine the distance gradient field information corresponding to the motion region based on the distances corresponding to the pixels of the motion region and the distances corresponding to the adjacent pixels of the pixels.

[0098] In some implementations, the preset generation model includes a denoising network; wherein the first frame image information of the target video sample, the mask image corresponding to the motion region, and the three-dimensional optical flow information corresponding to the motion region are introduced into the denoising network by concatenating input channels.

[0099] In some embodiments, the device further includes a weight setting module, used to set the weights corresponding to the first newly added channel and the second newly added channel of the denoising network to zero when initializing the preset generation model; wherein the first newly added channel is used to input the mask map corresponding to the motion region, and the second newly added channel is used to input the three-dimensional optical flow information corresponding to the motion region.

[0100] The video generation apparatus provided in this disclosure can execute the video generation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.

[0101] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device embodiments can be referred to the corresponding process in the method embodiments, and will not be repeated here.

[0102] This disclosure provides an electronic device, which includes: a storage device storing a computer program thereon; and a processing device for executing the computer program in the storage device to implement the steps of any method of this disclosure.

[0103] Referring now to FIG7, a schematic diagram of the structure of an electronic device 700 suitable for implementing embodiments of the present disclosure is shown. The terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG7 is merely an example and should not impose any limitation on the functionality and scope of use of embodiments of the present disclosure.

[0104] As shown in Figure 7, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0105] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 shows electronic device 700 with various devices, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0106] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the methods of embodiments of this disclosure.

[0107] In addition to the methods and devices described above, embodiments of this disclosure can also be computer program products, comprising computer program instructions that, when executed by a processor, cause the processor to perform the methods provided in the embodiments of this disclosure. The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this disclosure. These programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user computing device, partially on a user device, as a standalone software package, partially on a user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0108] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the video generation method provided in embodiments of this disclosure.

[0109] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0110] This disclosure also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the video generation method of this disclosure.

[0111] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0112] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0113] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0114] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0115] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0116] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A video generation method, comprising: In response to a region selection operation on a target reference image, a target region is determined based on the region selection operation; Obtain target trajectory information of the target region moving in the target reference image; wherein, the target trajectory information is trajectory information represented by three-dimensional spatial information; Based on the target region, the target reference image, and the target trajectory information, a motion video of the target object is generated using a preset target network model; wherein, the target object is the target region or an object contained within the target region.

2. The method according to claim 1, wherein, The three-dimensional spatial information includes three-dimensional optical flow information, and the step of obtaining the target trajectory information of the target region moving in the target reference image includes: Upon receiving expected trajectory information set for the target region, three-dimensional optical flow information corresponding to the expected trajectory information is generated, and target trajectory information of the target region moving in the target reference image is obtained based on the three-dimensional optical flow information corresponding to the expected trajectory information. In the absence of receiving the expected trajectory information set for the target region, the preset three-dimensional optical flow information is used as the target trajectory information for the movement of the target region in the target reference image; wherein, the preset three-dimensional optical flow information is characterized by a specific value.

3. The method according to claim 1, wherein, The target network model is obtained through the following steps: Obtain a target video sample; wherein the target video sample is a video in which a moving object exists in the foreground and the background remains stationary; The motion region in the target video sample is determined, and the three-dimensional optical flow information corresponding to the motion region is obtained; wherein, the motion region is a region containing a moving object; Based on the target video sample, the motion region, and the corresponding three-dimensional optical flow information of the motion region, the parameters of the preset generation model are adjusted to obtain the target network model based on the preset generation model with adjusted parameters.

4. The method according to claim 3, wherein, The acquisition of target video samples includes: Obtain a first video sample; wherein, the number of the first video samples is multiple; Obtain optical flow information corresponding to the first video sample, and based on the optical flow information corresponding to the first video sample, filter out a second video sample from multiple first video samples; wherein, the second video sample is a non-still video. Based on the optical flow information corresponding to the second video sample, the target mapping information corresponding to the video frame group in the second video sample is determined; wherein, the video frame group includes two video frame images with a preset interval; Based on the target mapping information, a third video sample is obtained by filtering from the second video sample, and a target video sample is obtained based on the third video sample; wherein, the target video sample is the third video sample, or the target video sample is a sample obtained by segmenting the third video sample.

5. The method according to claim 4, wherein, The optical flow information corresponding to the second video sample includes dense optical flow information corresponding to the video frame group in the second video sample; the step of determining the target mapping information corresponding to the video frame group in the second video sample based on the optical flow information corresponding to the second video sample includes: Based on the dense optical flow information corresponding to the video frame group in the second video sample, multiple sparse optical flow points are obtained using a preset uniform sampling strategy. Based on the multiple sparse optical flow points, the key point pairs that are successfully matched are determined for the video frame groups in the second video sample; From the successfully matched keypoint pairs, keypoint pairs corresponding to the background of the video frames in the video frame group are selected, and based on the mapping information of the keypoint pairs corresponding to the background of the video frames, the target mapping information corresponding to the video frame group in the second video sample is obtained.

6. The method according to claim 3, wherein, Determining the motion region in the target video sample includes: The target video sample's video frame image is subjected to object segmentation processing to obtain the object region; Based on the video frame image of the target video sample and the subsequent frame image adjacent to the video frame image of the target video sample, determine the two-dimensional optical flow information corresponding to the video frame image of the target video sample; Based on the object region and the two-dimensional optical flow information, the motion region in the video frame image of the target video sample is determined.

7. The method according to claim 6, wherein, The step of determining the motion region in the video frame image of the target video sample based on the object region and the two-dimensional optical flow information includes: Based on the two-dimensional optical flow information, a first region is selected from the object region; wherein the optical flow value of the two-dimensional optical flow information corresponding to the first region is greater than a preset threshold. Based on the two-dimensional optical flow information, a second region is determined from the video frame image of the target video sample; wherein, the second region is a region other than the first region, and the optical flow value of the two-dimensional optical flow information corresponding to the second region is greater than the preset threshold; Based on the first region and the second region, the motion region in the video frame image of the target video sample is obtained.

8. The method according to claim 3, wherein, The step of obtaining the three-dimensional optical flow information corresponding to the motion region includes: Obtain the distance gradient field information corresponding to the motion region; Two-dimensional optical flow information corresponding to the motion region is obtained, and three-dimensional optical flow information corresponding to the motion region is obtained based on the two-dimensional optical flow information and the distance gradient field information.

9. The method according to claim 8, wherein, The step of obtaining the distance gradient field information corresponding to the motion region includes: Obtain the distance between the pixels of the motion region and the edge of the motion region; Based on the distances corresponding to pixels in the motion region and the distances corresponding to adjacent pixels of the pixel, the distance gradient field information corresponding to the motion region is determined.

10. The method according to claim 3, wherein, The preset generation model includes a denoising network; wherein, the first frame image information of the target video sample, the mask image corresponding to the motion region, and the three-dimensional optical flow information corresponding to the motion region are introduced into the denoising network by splicing the input channels.

11. The method according to claim 10, wherein, The method further includes: When initializing the preset generation model, the weights corresponding to the first newly added channel and the second newly added channel of the denoising network are set to zero; wherein, the first newly added channel is used to input the mask map corresponding to the motion region, and the second newly added channel is used to input the three-dimensional optical flow information corresponding to the motion region.

12. A video generation apparatus, comprising: A region determination module is used to determine a target region based on a region selection operation on a target reference image in response to the region selection operation. The trajectory acquisition module is used to acquire target trajectory information of the target region moving in the target reference image; wherein, the target trajectory information is trajectory information represented by three-dimensional spatial information; The video generation module is used to generate a motion video of the target object based on the target region, the target reference image, and the target trajectory information using a preset target network model; wherein the target object is the target region or an object contained in the target region.

13. An electronic device, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the video generation method according to any one of claims 1-11.

14. A computer-readable storage medium storing a computer program for performing the video generation method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Audio and video generation method and device, computer equipment and storage medium

    CN116934638A

  • Image processing method and electronic equipment

    CN118101856A

  • Video generation method and device, electronic equipment and storage medium

    CN118644411A

  • Video generation method and device, equipment and medium

    CN119520923A

  • Machine learning based controllable animation of still images

    US20240005587A1