Method, apparatus, device and computer program for generating video
By allowing users to input content and mark objects with bounding boxes in start and end frames, the method addresses the challenge of accurately controlling object movement in videos, ensuring the generated content meets user expectations and enhances video creation flexibility.
Patent Information
- Application Number
- JP2024114013
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-30
- Filing Date
- 2024-07-17
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2044-07-17
AI Technical Summary
Existing video generation technologies struggle to accurately interpret user requirements for object movement patterns in videos, especially when precise motion schemes are described in natural language, leading to inaccuracies in generating desired video content.
A method and apparatus that allow users to input content information, such as text or images, and mark objects with bounding boxes in start and end frames to provide precise movement control, enabling the application to generate videos that meet user expectations by using a video generation model that incorporates bounding box-guided motion control.
Enables accurate and precise control of object movement in generated videos, enhancing user experience and flexibility in video creation by allowing users to clearly define object positions and trajectories, resulting in videos that align with their intended scenarios and effects.
Smart Images

Figure 2025117515000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to the field of artificial intelligence, and more particularly to methods, apparatus, electronic devices and computer programs for generating video. [Background technology]
[0002] Text-guided video generation is a technology that uses natural language text to guide the creation of video content. Deep learning and artificial intelligence technologies enable the system to understand input text descriptions, convert them into specific visual content, and generate the corresponding video. This method can be applied to fields such as filmmaking, virtual reality, and video production, providing creators with a more intuitive and efficient way to express their ideas.
[0003] Motion control refers to the precise control of object or camera movement to achieve various effects and dynamics in video. This technology can be implemented through programming or automated systems, making video production more creative and dynamic. Motion control is widely applied in fields such as film and virtual reality, providing viewers with a more immersive and engaging visual experience.
[0004] Combining text-guided video creation with motion control can enable smarter and more personalized video creation. Text guidance allows creators to express their desired scenarios and effects in a natural language manner, and motion control ensures that these ideas appear in the video in an accurate and smooth manner, providing greater flexibility and creativity for the creation process. Summary of the Invention [Problem to be solved by the invention]
[0005] To solve the problems of the prior art. [Means for solving the problem]
[0006] In a first aspect of an embodiment of the present disclosure, there is provided a method for generating a video, the method including: obtaining content information, the content information including at least one of text or images, related to content of a video to be generated; obtaining position information indicating a position of an object in the video at a start frame; obtaining control information constraining a position of the object at an end frame; and generating the video based on the content information, the position information, and the control information.
[0007] In a second aspect of an embodiment of the present disclosure, there is provided an apparatus for generating a video, the apparatus including: a content information acquisition module configured to acquire content information relating to content of a video to be generated, the content information including at least one of text or images; a position information acquisition module configured to acquire position information indicating a position of an object in the video at a start frame; a control information acquisition module configured to acquire control information constraining a position of the object in an end frame; and a video generation module configured to generate the video based on the content information, the position information, and the control information.
[0008] In a third aspect of an embodiment of the present disclosure, there is provided an electronic device, the electronic device including one or more processors and a storage device for storing one or more programs, the one or more programs, when executed by the one or more processors, causing the one or more processors to implement a method for generating a video, the method including obtaining content information, the content information including at least one of text or images, related to content of the video to be generated, the method further including obtaining position information indicating a position of an object in a start frame of the video, and the method further including obtaining control information constraining a position of the object in an end frame, the method further including generating the video based on the content information, the position information, and the control information.
[0009] In a fourth aspect of an embodiment of the present disclosure, there is provided a computer program product, the computer program product being tangibly stored on a non-transitory computer-readable medium and including machine-executable instructions that, when executed, cause a machine to perform a method for generating a video. The method includes obtaining content information, the content information including at least one of text or images, related to content of a video to be generated. The method further includes obtaining position information indicating a position of an object in a start frame of the video. The method further includes obtaining control information that constrains a position of the object in an end frame. The method further includes generating the video based on the content information, the position information, and the control information.
[0010] This Summary is provided to introduce in a simplified form a selection of concepts, which are further described in the specific embodiments below. This Summary is not intended to identify key features or primary characteristics of the subject matter for which protection is sought, nor is it intended to limit the scope of the subject matter for which protection is sought. [Brief explanation of the drawings]
[0011] These and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent from the following detailed description taken in conjunction with the drawings, in which like or similar reference numerals represent like or similar elements, and in which: [Figure 1] 1 shows a schematic diagram of an exemplary environment in which several embodiments of the present disclosure may be implemented. [Figure 2] 1 illustrates a flowchart of a method for generating video according to some embodiments of the present disclosure. [Figure 3A] 1 illustrates a schematic diagram of an example in which a user inputs multiple bounding boxes at a start frame and an end frame, respectively, according to some embodiments of the present disclosure. [Figure 3B] 1 illustrates a schematic diagram of an example in which a user inputs multiple bounding boxes at a start frame and an end frame, respectively, according to some embodiments of the present disclosure. [Figure 4A] 1 shows a schematic diagram of an example of a user inputting a bounding box and motion trajectory of an object at a starting frame according to some embodiments of the present disclosure. [Figure 4B] 1 shows a schematic diagram of an example of a user inputting a bounding box and motion trajectory of an object at a starting frame according to some embodiments of the present disclosure. [Figure 5A] 10 shows a schematic diagram of an example in which a user inputs multiple soft bounding boxes at an end frame and does not input a text description, according to some embodiments of the present disclosure. [Figure 5B] 10 shows a schematic diagram of an example in which a user inputs multiple soft bounding boxes at an end frame and does not input a text description, according to some embodiments of the present disclosure. [Figure 6A] 10A-10C show schematic diagrams of an example of a user inputting multiple bounding boxes on a starting frame for which no image is provided, according to some embodiments of the present disclosure. [Figure 6B] 10A-10C show schematic diagrams of an example of a user inputting multiple bounding boxes on a starting frame for which no image is provided, according to some embodiments of the present disclosure. [Figure 7] 10A-10C illustrate schematic diagrams of an example of a user creating an intermediate frame and inputting multiple soft bounding boxes in the intermediate frame according to some embodiments of the present disclosure. [Figure 8] 1 shows a block diagram of an apparatus for generating video according to some embodiments of the present disclosure. [Figure 9] 1 shows a block diagram of a device capable of implementing several embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0012] It should be understood that all user-related data involved in this technical solution should be collected and used only after obtaining the user's permission. This means that if the technical solution needs to use the user's personal information, the user's explicit consent and permission must be obtained before collecting such data, otherwise the relevant data will not be collected or used. It should also be understood that when implementing this technical solution, relevant laws and regulations must be strictly observed during the collection, use, and storage of data, and necessary technologies and measures must be taken to ensure the security of user data and ensure the safe use of data.
[0013] The following describes in more detail the embodiments of the present disclosure with reference to the drawings. Although the drawings show some embodiments of the present disclosure, it should be understood that the present disclosure can be realized in various forms and should not be construed as being limited to the embodiments described herein, but rather these embodiments are provided for a better understanding and complete comprehension of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are merely illustrative and are not intended to limit the protection scope of the present disclosure.
[0014] In describing embodiments of the present disclosure, the term "comprises" and similar terms should be understood as an open-ended "includes," i.e., "including, but not limited to." The term "based on" should be understood as "based at least in part on." The term "one embodiment" or "this embodiment" should be understood as "at least one embodiment." Unless explicitly stated otherwise, terms such as "first," "second," etc. may refer to different or the same object. Other explicit and implicit definitions may be included below.
[0015] In video generation scenarios guided by some text or reference images, a user may wish to provide information about the movement pattern of objects in the generated video by inputting a text description. For example, a user may provide a reference image of a building and input the text description, "Tilt the camera up to expose the top of the building." In this case, the user wishes the camera to gradually raise the lens from a viewing angle that captures the ground, eventually capturing the top of the building. However, in related art, although a video with relatively high image quality and slow lens movement can be generated based on the reference image and text description provided by the user, the model cannot adequately understand the user's requirements for the movement pattern of objects in the video, and therefore cannot accurately expose the top of the building in the generated video.
[0016] In addition, in some usage scenarios, when a user's requirements for the motion scheme are very precise, it is difficult to accurately describe the image in mind in written language. For example, if a user wants two puppies to run toward the camera in a video, one of the puppies, a white puppy, will approach the camera and run to the center of the screen, filling one-third of the screen. At the same time, the other black puppy will also approach the camera, but will run toward a toy next to the camera, moving away from the center of the screen and eventually disappearing to the right side of the screen. For an average user, it is very difficult to accurately describe such motion requirements, which makes it impossible to generate the desired video.
[0017] Therefore, an embodiment of the present disclosure provides a method for generating a video. In this method, a user can input content information related to the content of the video to be generated in a user interface provided by an application. This content information can be a text description, video keyframes, or both. The user can then mark a motion control object with a bounding box in a start frame and input control information in the user interface for how the object should move, the control information including at least the object's position in an end frame. The application can then generate a video based on the content information, the bounding box marking the object, and the control information.
[0018] In this manner, a user can accurately mark an object to be controlled using a bounding box in the start frame, and can accurately express the movement position of the marked object from the start frame to the end frame, allowing the application to receive precise movement control demands and generate a video that meets the user's expectations.
[0019] FIG. 1 illustrates a schematic diagram of an exemplary environment 100 in which embodiments of the present disclosure may be implemented. As illustrated in FIG. 1, the environment 100 includes a user 102 and a computing device 104, which may be a user terminal, a mobile device, a computer, a computing system, a single server, a distributed server, or a cloud-based server. The computing device 104 includes an application 106 capable of providing video generation functionality. The application 106 may be any application, such as a local application, a remote application, a browser / server architecture application, a client / server architecture application, or the like.
[0020] As shown in FIG. 1 , in environment 100, user 102 can interact with application 106 via user interface 108. On user interface 108, user 102 can input content information 110 related to the content of the video to be generated. The content information 110 can be a text description 112, a reference image 114, or both. For example, text description 112 can be "two puppies running toward the camera," and reference image 114 can be an image of two puppies running across a lawn. It should be noted that in some implementations, text description 112 can be provided indirectly through user interaction methods, such as audio, and thus text description 112 can also include text provided indirectly through methods such as audio. User 102 can mark an object whose motion he or she wishes to control by drawing bounding boxes 118-1, 118-2, ..., 118-N (collectively referred to as bounding boxes 118) in an area displaying starting frame 116. 1, a user may draw only one bounding box 118 to control the movement of only one object. If the content information 110 includes a reference image 114, the reference image 114 may be displayed in the area of the starting frame 116, thereby facilitating the user to use the bounding box 118 to mark the object whose movement they wish to control.
[0021] 1, in environment 100, user 102 may input control information 120 via user interface 108 to describe at least the location of the controlled object in the ending frame of the video. In some embodiments, control information 120 may be a bounding box drawn by user 102 in the ending frame region to represent the desired location of the controlled object. In some embodiments, control information 120 may be a motion trajectory of the controlled object drawn by user 102 in the starting frame 116 region.
[0022] In the environment 100, after the user 102 inputs the content information 110, the bounding box 118 in the start frame 116, and the control information 120, the application 106 can generate a video 122 based on these user inputs and provide the video 122 to the user 102 via the user interface 108. For example, the computing device 104 can send these user inputs to a server, receive an address for the video 122 from the server, or generate the video 122 locally by the computing device 104. In the user interface 108, the video 122 can be displayed to the user 102 via, for example, video playback controls or provided to the user 102 in the form of download controls. The content of the video 122 relates to the content information 110 and is marked in the video 122 by the bounding box 118. The control object moves from a position in the start frame 116 to a specified position in the end frame according to the constraints of the control information 120.
[0023] It should be understood that while the environment 100 includes the content information 110, start frame 116, and control information 120 in a single user interface 108, in some embodiments the user 102 may enter each of these pieces of information in different user interfaces, and the video 122 may be provided to the user 102 in a separate user interface.
[0024] In this manner, the user 102 can accurately mark an object that they want to control using a bounding box 118 in the start frame 116. In addition, the user 102 can accurately express the movement position of the marked object from the start frame 116 to the end frame, and the application 106 can receive precise movement control demands and thereby generate a video that meets the user's expectations.
[0025] 2 illustrates a flowchart of a method 200 for generating a video according to some embodiments of the present disclosure. As shown in FIG. 2, in box 202, the method 200 can obtain content information including at least one of text or images related to the content of the video to be generated. For example, in the environment 100 illustrated in FIG. 1, the computing device 104 can obtain content information 110 related to the video content to be generated input by the user 102, the content information 110 including at least one of a text description 112 and a reference image 114.
[0026] In box 204, method 200 can obtain position information indicating the position of an object in the video in the starting frame. This position information can be position-related information such as a bounding box, a contour, a coordinate value, a coordinate range, etc. For example, in the environment 100 shown in FIG. 1 , computing device 104 can obtain a bounding box 118 input by user 102 in starting frame 116, and the bounding box 118 can be used to mark a motion-control object in starting frame 116. If user 102 provides a reference image 114, the content of starting frame 116 can be the reference image 114, thereby facilitating the user to directly mark a motion-control object on the reference image 114. If user 102 does not provide a reference image 114, the content of starting frame 116 can be blank, and the user can mark the position and size of the motion-control object in the blank area with a bounding box 118.
[0027] In box 206, method 200 can obtain control information that constrains the position of the object in the end frame. For example, in environment 100 shown in FIG. 1 , computing device 104 can obtain control information 120 entered by user 102, where control information 120 can constrain the position of a controlled object, labeled with bounding box 118, in the end frame. In some embodiments, user 102 can provide control information 120 in the form of drawing a bounding box in the end frame. In some embodiments, user 102 can provide control information 120 in the form of drawing a motion trajectory in start frame 116.
[0028] In box 208, method 200 can generate a video based on the content information, the position information, and the control information. For example, in environment 100 shown in FIG. 1, computing device 104 can generate video 122 based on content information 110, bounding box 118, and control information 120. The content of video 122 relates to content information 110, and a control object, identified by bounding box 118 in video 122, moves from a position in starting frame 116 to a specified position according to the constraints of control information 120.
[0029] In this manner, the user can accurately mark the object to be controlled using the position information in the start frame. In addition, the user can accurately express the movement position of the marked object from the start frame to the end frame, so that the application can receive precise movement control demands and generate a video that meets the user's expectations.
[0030] In some embodiments, the object is a first object, the position information is a first bounding box having a first color, the control information is the first control information having the first color, and generating the video may include obtaining a second bounding box indicating a position of the second object in the video at a start frame, the second bounding box having a second color different from the first color. Additionally, second control information may be obtained that constrains a position of the second object at an end frame, the second control information having the second color. The video may then be generated based on the content information, the first bounding box, the first control information, the second bounding box, and the second control information.
[0031] In some embodiments, the position information is a start bounding box, and the control information is an end bounding box in the end frame of the object, and the start bounding box and the end bounding box are rectangular boxes. In some embodiments, in response to a user selecting a first type as the type of the end bounding box, a video may be generated by moving the object from a position indicated by the start bounding box to a specific position indicated by the end bounding box, where the size of the object in the end frame corresponds to the end bounding box. In some embodiments, image content marked by the bounding box in the start frame is determined as the object, and a video may be generated based on the image, object, and control information of the start frame, where the content of the video is related to the image.
[0032] In some embodiments, in response to the ending bounding box being close to a left or right boundary of the ending frame and the width of the ending bounding box being less than a threshold width, the video can be generated by moving an object from a position in the starting frame to outside the left or right boundary of the ending frame, the magnitude of the object as it moves outside the left or right boundary being related to the height of the ending bounding box. In some embodiments, in response to the ending bounding box being close to a top or bottom boundary of the ending frame and the height of the ending bounding box being less than a threshold height, the video can be generated by moving an object from a position in the starting frame to outside the top or bottom boundary of the ending frame, the magnitude of the object as it moves outside the top or bottom boundary being related to the width of the ending bounding box, and the content of the video is associated with the content information.
[0033] In some embodiments, in response to a starting bounding box being close to a left or right boundary of the starting frame and the width of the starting bounding box being less than a threshold width, the video is generated by moving an object from outside the left or right boundary to a position constrained by the control information, wherein the size of the object as it enters the left or right boundary is related to the height of the starting bounding box. In some embodiments, in response to a starting bounding box being close to a top or bottom boundary of the starting frame and the height of the starting bounding box being less than a threshold height, the video is generated by moving an object from outside the top or bottom boundary of the starting frame to a position constrained by the control information, wherein the size of the object as it enters the top or bottom boundary is related to the width of the ending bounding box, and wherein the content of the video is related to the content information.
[0034] In some embodiments, the object moves gradually from a first position to a second position in the generated video. In some embodiments, the generated video is adapted to move the object relative to the camera by changing the camera viewing angle.
[0035] 3A-3B illustrate schematic diagrams of an example 300 in which a user inputs multiple bounding boxes at a start frame and an end frame, respectively, according to some embodiments of the present disclosure. As shown in FIG. 3A, the example 300 includes an input user interface 302, which provides a control 304 for inputting a text description. A user can input a text description in the text input control 304 to describe the content of the video to be generated. For example, in the example 300, if the user inputs the text description "two puppies running toward the camera," the content of the generated video should include two puppies running toward the camera. The user interface 302 further provides a select picture control 306 for inputting a reference image, and the user can select the reference image by interacting with the control 306. If the user selects a reference image, the reference image can guide the content of the generated video. For example, in example 300, if the user selects an image of two puppies (including one white puppy and one black puppy) running on the grass, the video generated should include these two puppies running on the grass in the background.
[0036] As shown in FIG. 3A , the user interface 302 further provides a start frame control 308, and if the user selects a reference image, the content of the start frame control 308 can display this reference image. If the user has not selected a reference image, the content of the start frame control 308 can be blank (similar to the control 318 for representing the end frame in FIG. 3 ). The user interface 302 further provides a color indicator control 310 corresponding to object 1 and a color indicator control 312 corresponding to object 2, associated with the start frame control 308, where the colors corresponding to each object are different. When the user interacts with the color indicator controls 310 or 312, the user can select which object is currently indicated on the start frame control 308. For example, in the example 300, object 1 is selected, and the next operation performed on the start frame control 308 indicates that the operation is for object 1. The user interface 302 further provides an operation type selection control 314, where the operation type includes a bounding box and a motion trajectory, and when the value of the operation type selection control 314 is a bounding box, a start bounding box marking the object 1 can be input in a drag manner in the start frame control 308. When the value of the operation type selection control 314 is a motion trajectory, a motion trajectory for the object 1 can be input in a brush manner in the start frame control 308.
[0037] 3A , the user interface 302 further provides an end frame control 318, the content of which may be blank, allowing the user to draw an end bounding box in the end frame control 318, which constrains the location and size of the object to which the object will move. The user interface 302 further provides a color indicator control 320 corresponding to object 1 and a color indicator control 312 corresponding to object 2, associated with the end frame control 318, where the value of the color indicator control 320 is the same as the value of the color indicator control 310, and the value of the color indicator control 322 is the same as the value of the color indicator control 312. In other words, the same object is represented by the same color in the start frame control 308 and the end frame control 318. It should be understood that, in the embodiment of the present disclosure, each object is assigned a specific color, such as black and white, but this is not intended to limit the color corresponding to each object, and the color may be other colors, such as yellow, purple, green, etc.
[0038] As shown in FIG. 3A , the user interface 302 further provides a bounding box type selection control 324. Bounding boxes come in two types: hard bounding boxes and soft bounding boxes. A hard bounding box is used to specify a specific position and a specific size of an object. This indicates that the object will be generated at the coordinates (e.g., the center coordinates of the bounding box) specified by the hard bounding box in the generated image frame, and the size of the object corresponds to the size of the hard bounding box. A soft bounding box is used to specify a position range and a size range of an object. This indicates that the object will be generated within the range defined by the soft bounding box in the generated image frame, and the size of the object will not exceed this range. If the user selects an object 1 to be manipulated in the end frame and the value of the bounding box type control is a hard bounding box, then the user can use the end frame control 318 to draw a hard bounding box that constrains a specific position of the object 1 after it moves, for example, by dragging. If the user selects an object 1 to be manipulated in the end frame and the value of the bounding box type control is a soft bounding box, then the user can use the end frame control 318 to draw a soft bounding box that constrains a position range of the object 1 after it moves.
[0039] 3A, in example 300, the user marks the white puppy as Subject 1 with a black bounding box 330 in the start frame control 308, and marks a particular location and size that Subject 1 should move to with a black hard bounding box 334 in the end frame control 318. Note that the user marks the black puppy as Subject 2 with a white bounding box 332 in the start frame control 308, and marks a range of locations and sizes that Subject 2 should move to with a white hard bounding box 336 in the end frame control 318.
[0040] 3A, in the end frame control 318 of example 300, the hard bounding box 336 of Subject 2 is close to the right boundary of the end frame and its width is relatively narrow (e.g., smaller than the threshold width), which indicates that the black puppy corresponding to Subject 2 runs out of the right boundary of the video, and its size is related to the height of the hard bounding box 336 when the black puppy runs out of the right boundary. In other words, the height of the hard bounding box 336 is the projection of Subject 2's height onto the right boundary when it moves out of the right boundary.
[0041] The user can then click on the video generation control 328 to generate the video, and the generated video will be as shown in Figure 3B. In the video 340 shown in Figure 3B, the white puppy (Subject 1) runs from a position in the start frame to a position specified by the hard bounding box 334 in the end frame, and the size of the white puppy in the end frame corresponds to the hard bounding box 334. At the same time, the black puppy (Subject 2) runs from a position in the start frame to off-screen, and when the black puppy starts running off-screen in the third frame, its height may be relative to the hard border frame 324.
[0042] In this way, the user can use different colors to mark multiple control objects, thereby making the content of the generated video more dynamic and rich. Furthermore, by marking a bounding box in the end frame, the position and size of the object's destination can be precisely controlled, allowing the user to accurately express and control the object's movement, improving the motion effect of the generated video and enhancing the user experience. In addition, the effect of moving an object off-screen can be achieved, and the position and size of the object when it moves off-screen can be specified, thereby providing rich motion control instructions and meeting the needs of users in different scenarios.
[0043] In some embodiments, the position information is a bounding box and the control information is a motion trajectory that is drawn on a starting frame. In some embodiments, a video can be generated by moving an object along the motion trajectory from a position indicated by the bounding box, where the content of the video is associated with content information.
[0044] 4A-4B illustrate schematic diagrams of an example 400 in which a user inputs a bounding box and motion trajectory of an object at a start frame, according to some embodiments of the present disclosure. As shown in FIG. 4A, in example 400, a user inputs a text description, "A person is throwing a Frisbee," in a text input control 404 of a user interface 402. As shown in start frame control 408, the user selects a reference image, which includes a person holding a Frisbee. In example 400, the user labels object 1 (i.e., the Frisbee) with a black bounding box 432 in start frame control 408. The user then selects the value of operation type selection control 414 as motion trajectory and draws a motion trajectory 434 for the Frisbee in start frame control 408. As shown in FIG. 4A, the Frisbee's motion trajectory 434 resembles a "U" shape, with a start position and an end position close together, representing the Frisbee flying out from the start position and eventually flying back to the position specified by the trajectory arrow. Note that the user has marked the person as Subject 2 with a white bounding box 430 in the start frame control 408, and marked a hard bounding box 436 for Subject 2 in the end frame control 418, indicating that Subject 2 will move from its position in the start frame to the position specified by the hard bounding box 436, with its size after the movement corresponding to the size of the hard bounding box 436.
[0045] 4B shows the resulting video 440, in which, as shown in the drawing, a Frisbee (subject 1) flies from a starting position along a motion trajectory 434, and finally flies to an end position of the motion trajectory 434. At the same time, a person (subject 2) is located at a position specified by a hard bounding box 436 after undergoing a series of motions, and the size of the person corresponds to the hard bounding box 436.
[0046] In this manner, the user can control the movement of the object by drawing the movement trajectory of the object, and since the movement trajectory contains more information about the movement process, the movement of the object can be controlled more precisely. In addition, since it is difficult for users to express a somewhat complicated movement trajectory in text description, by drawing the movement trajectory, the user can easily express how the object moves, thereby improving the user experience.
[0047] In some embodiments, in response to a user selecting the second type as the type of the ending bounding box, a video is generated by moving an object from a position indicated by the starting bounding box to within a range of positions indicated by the ending bounding box, where the size of the object in the ending frame does not exceed the ending bounding box, and the content of the video is associated with the content information.
[0048] 5A-5B illustrate schematic diagrams of an example 500 in which a user inputs multiple soft bounding boxes but does not input a text description at an end frame, according to some embodiments of the present disclosure. As shown in FIG. 5A, in example 500, the user has not input any content in a text input control 504 of a user interface 502, but has selected a reference image. The reference image includes multiple eggs in a basket, as shown in a start frame control 508. In example 500, the user marks object 1 (i.e., the egg located to the right of the basket) with a black bounding box 528, object 2 (i.e., the egg located behind the basket) with a white bounding box 530, and object 3 (i.e., the egg located to the left of the basket) with a gray bounding box 532 in the start frame control 508. In the end frame, the user selects the values of the bounding box type selection controls 522, 524, and 526 for object 1, object 2, and object 3 as soft bounding boxes, and in the end frame control 518, specifies the position range and size range of the destination of the egg on the right side of the basket with a black soft bounding box 534, i.e., the position of the egg on the right side of the basket in the end frame does not exceed the range specified by the soft bounding box 534, and its size does not exceed the soft bounding box 534. Note that the user also specifies the position range and size range of the destination of the egg on the left side of the basket with a white soft bounding box 536, and specifies the position range and size range of the destination of the egg behind the basket with a gray soft bounding box 538 in the end frame control 518.
[0049] 5B shows the resulting video 540, in which, as shown in the figure, the egg on the right of the basket (subject 1) moves from a starting position to within an area defined by soft bounding box 534, and its size does not exceed the range defined by soft bounding box 534. The egg at the back of the basket (subject 2) moves from a starting position to within an area defined by soft bounding box 536, and its size does not exceed the range defined by soft bounding box 536. The egg on the left of the basket (subject 3) moves from a starting position to within an area defined by soft bounding box 538, and its size does not exceed the range defined by soft bounding box 538.
[0050] In this way, the user can use the soft bounding box to expand the constraints on the movement of the controlled object, thereby increasing the variety of the generated video when the constraints are met. In addition, the requirements on the user can be reduced, i.e., only a certain range of constraints are required, thereby making the operation easier and improving the user experience.
[0051] In some embodiments, noun phrases in the text can be identified and the noun phrases can be determined to be targets, and a video can be generated based on the text, targets, and control information, where the content of the video is related to the text.
[0052] 6A-6B illustrate a schematic diagram of an example 600 in which a user inputs bounding boxes on a start frame for which no image is provided, according to some embodiments of the present disclosure. As shown in FIG. 6, in example 600, a user inputs a text description describing the video content, "Four pigs running in the snow," in a text input control 604 of a user interface 602. However, because the user has not selected a reference image, the background of the start frame control 608 is blank. In example 600, the user marks the start position and size of the four pigs in the start frame control 608 with a black bounding box 630, a white bounding box 632, a gray bounding box 634, and a yellow bounding box 636, respectively. In the end frame control 618, the user marks the position and size of the four pigs in the end frame with a black hard bounding box 640, a white hard bounding box 642, a gray hard bounding box 646, and a yellow hard bounding box 646, respectively.
[0053] When generating the video, the video generation model can identify noun phrases in these descriptions. Because many of these phrases are abstract nouns rather than concrete object names, these noun phrases can be filtered, leaving only phrases that represent concrete object names. These filtered noun phrases can then be processed to identify control objects and associate these objects with bounding boxes.
[0054] 6B shows a generated video 650, in which the video generation model generates four pigs based on the text description and at locations and sizes specified by bounding boxes 630, 632, 634, and 636, and the four generated pigs are associated with the four bounding boxes. The four pigs run from their initial positions to locations specified by hard bounding boxes 640, 642, 644, and 646, respectively, and the sizes of the four pigs correspond to the hard bounding boxes.
[0055] In this way, when a reference image is not provided, it is possible to realize an indication for an object in the text description, so that the user can generate a desired video even when a reference image cannot be provided; in this way, the prerequisites for the user to generate a video using the application can be reduced, so that more users will generate videos using this application.
[0056] In some embodiments, third control information that constrains the position of the object in the intermediate frames can be obtained, and the video can be generated based on the content information, the bounding box, the first control information, and the third control information.
[0057] FIG. 7 illustrates a schematic diagram of an example 700 in which a user creates an intermediate frame and inputs multiple soft bounding boxes in the intermediate frame, according to some embodiments of the present disclosure. As shown in FIG. 7 , in example 700, the user inputs a text description describing the video content, "Two puppies running toward the camera," in a text input control 704 of a user interface 702 and selects a reference image (as shown in a start frame control 708). In the start frame control 708, the user labels the white puppy as subject 1 in a black bounding box 730 and the black puppy as subject 2 in a white bounding box 732. In the end frame control 718, the user specifies the destination position and size of the white puppy in a black hard bounding box 740 and the destination position and size of the black puppy in a white hard bounding box 742. However, the bounding boxes in the end frame cannot directly participate in the movement process of the white puppy moving from the position of bounding box 730 to the position of bounding box 740. Thus, in example 700, the user inserts one intermediate frame between the start frame and the end frame, and in response, user interface 702 displays an intermediate frame control 728 and a set of controls associated with the intermediate frame control 728 (e.g., color display controls for each object, bounding box type selection controls, etc.), the functionality of which is the same as the functionality of the set of controls associated with the end frame control 718.
[0058] 7 , in example 700, the user uses intermediate frame control 728 to constrain the destination position and size of the white puppy with black soft bounding box 750 and the destination position and size of the black puppy with white soft bounding box 752. Thus, in the generated video, as the white puppy runs from the position of bounding box 730 to the position specified by hard bounding box 740, it first passes through a position in the area specified by soft bounding box 750 and then reaches the position specified by hard bounding box 740. As the black puppy runs from the position of bounding box 732 to the position specified by hard bounding box 742, it first passes through a position in the area specified by soft bounding box 752. Note that in user interface 702, the user can further insert more intermediate frames by interacting with image frame insertion controls 760 and 762, thereby achieving more precise control over the movement process of the two puppies.
[0059] In this way, by inserting an intermediate frame between the start frame and the end frame, the movement process of the controlled object can be controlled precisely. Compared with the movement trajectory, the method of inserting an intermediate frame can further control the size of the object in the movement process, thereby making the movement control function of the application more complete.
[0060] To realize bounding box-guided video generation, a motion control module may be inserted in a traditional video generation model, which processes bounding boxes as control tokens and utilizes a self-attention layer to fuse the control tokens with visual tokens to generate image frames, thereby generating fused visual tokens, which include the motion control information provided by the bounding boxes.
[0061] One exemplary architecture includes a spatial self-attention layer, a multilayer perceptron, a motion control module, and a spatial cross-attention layer. The spatial self-attention layer and the spatial cross-attention layer may be modules in a video diffusion model based on a three-dimensional U-network (3D U-Net) architecture, for example. The video diffusion model can iteratively predict noise vectors in noisy video inputs, thereby gradually transforming pure Gaussian noise into high-quality video frames. The 3D U-Net consists of alternating convolutional and attention blocks. Each block includes two components: a spatial component that processes each image frame as a single image, and a temporal component that facilitates information exchange between image frames. In each attention block, the spatial component typically includes a self-attention layer followed by a cross-attention layer, which is used to adjust video generation based on text presentation. A motion control module is inserted between these two attention layers, allowing the model to manage motion control in video generation.
[0062] This exemplary architecture inserts a motion control module between the spatial self-attention layer and the cross-spatial attention layer of the original video diffusion model. The spatial self-attention layer receives frame-level visual tokens and generates visual tokens based on the frame-level visual tokens. The motion control module receives visual tokens and control tokens as input and outputs fused visual tokens, where each control token corresponds to a corresponding object (or bounding box). Because the control tokens include motion control information provided by the bounding box, the fused visual token also includes motion control information provided by the bounding box. The visual tokens are then input to the cross-spatial attention layer, which can generate updated frame-level visual tokens based on the visual tokens and text tokens. The video diffusion model can then generate image frames based on the updated frame-level visual tokens. To avoid changing the original structure of the cross-spatial attention layer, the number of visual tokens and the number of visual tokens may be kept the same. In this way, by fixing the parameters of the original video diffusion model (including the spatial self-attention layer and spatial cross-attention layer) during the training phase and only adjusting the parameters of the motion control module, retraining due to modification of the structure of the video diffusion model can be avoided, saving costs and avoiding the deterioration of the accuracy of the original video diffusion model due to retraining.
[0063] In this exemplary architecture, the number of control tokens is determined by the number of bounding boxes that the video generation model supports simultaneously existing in an image frame, and the control tokens and bounding boxes have a one-to-one correspondence. For example, if the video generation model only supports an image frame containing a bounding box for one object, the number of control tokens is 1. If the video generation model supports an image frame containing five bounding boxes for five objects simultaneously, the number of control tokens is 5. If the video generation model supports providing five bounding boxes simultaneously in an image frame but only needs to control the movement of two objects in the video to be generated (i.e., only two bounding boxes are provided), the three empty control tokens can be filled with specific learnable tokens. In this exemplary architecture, text tokens are not required; i.e., if the user does not provide a text description for the video to be generated, the empty text tokens can be filled with learnable tokens.
[0064] To generate control tokens, the coordinates of the bounding boxes, unique object identifiers for labeling the bounding boxes, and bounding box types may be determined. Then, the control tokens are generated based on the coordinates, object identifiers, and bounding box types. For example, the object identifiers may be represented in a color RGB space, where each object corresponds to a bounding box with a unique color, and the object identifiers are vectors with normalized 3D RGB values between 0 and 1. The coordinates, object identifiers, and bounding box types are concatenated into a vector, and a corresponding embedding is generated using a Fourier embedding operation. This embedding is then input to a multilayer perceptron to generate control tokens. By generating object identifiers using RGB values, corresponding bounding boxes can be generated in image frames based on the object identifiers during the training phase, thereby facilitating the alignment of the generated bounding boxes with ground truth bounding boxes and improving the effectiveness of model training.
[0065] It should be understood that while this exemplary architecture illustrates generating control tokens based on bounding box coordinates, target indicators, and bounding box type, in some embodiments the target indicator and bounding box type are not required. For example, in some embodiments, if only one particular type of bounding box (e.g., hard bounding box) is supported, control tokens may be generated based on coordinates alone. In some embodiments, if only multiple particular types of bounding boxes are supported, control tokens may be generated based on coordinates and target indicators alone.
[0066] In this manner, the motion control module can provide accurate motion control information to the original video diffusion model, thereby improving the effect of the generated image frame and allowing the object to move in a manner desired by the user. Note that, because the inserted motion control module does not change the structure and parameters of the original video diffusion model, this exemplary architecture can reuse the capabilities of the trained video diffusion model, thereby ensuring the visual quality of the generated video and improving the motion control for the object in the video.
[0067] 8 illustrates a block diagram of an apparatus 800 for generating a video according to some embodiments of the present disclosure. As illustrated in FIG. 8 , the apparatus 800 includes a content information acquisition module 802 configured to acquire content information, including at least one of text and images, related to the content of the video to be generated. The apparatus 800 further includes a position information acquisition module 804 configured to acquire position information indicating a position of an object in the video at a start frame. The apparatus 800 further includes a control information acquisition module 806 configured to acquire control information constraining a position of the object in an end frame. The apparatus 800 further includes a video generation module 808 configured to generate the video based on the content information, the position information, and the control information.
[0068] As can be seen, the device 800 of the present disclosure can be used to achieve at least one of the many advantages that the above-described method or process can achieve. For example, the device 800 allows a user to accurately mark an object to be controlled using a bounding box in a start frame. In addition, the user can accurately express the movement position of the marked object from the start frame to the end frame, allowing the application to receive precise movement control demands and thereby generate a video that meets the user's expectations.
[0069] FIG. 8 shows a block diagram of a device 800 capable of implementing multiple embodiments of the present disclosure. The device 800 may be a device or apparatus described in the embodiments of the present disclosure. As shown in FIG. 8, the device 800 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 801, which can perform various appropriate operations and processes based on computer program instructions stored in a read-only memory (ROM) 802 or loaded from a storage unit 808 into a random access memory (RAM) 803. The RAM 803 may further store various programs and data necessary for the operation of the device 800. The CPU / GPU 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804. Although not shown in FIG. 8, the device 800 may further include a coprocessor.
[0070] Several components of device 800 are connected to I / O interface 805, such as input unit 806, e.g., keyboard, mouse, etc., output unit 807, e.g., various types of displays, speakers, etc., storage unit 808, e.g., magnetic disk, optical disk, etc., and communication unit 809, e.g., network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks.
[0071] Each of the methods or processes described above may be executed by the CPU / GPU 801. For example, in some embodiments, the methods may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed into the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the CPU / GPU 801, it may perform one or more steps or operations in the methods or processes described above.
[0072] In some embodiments, the methods and processes described above may be implemented as a computer program product, which may include a computer-readable storage medium having computer-readable program instructions thereon for carrying out aspects of the present disclosure.
[0073] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile optical disk (DVD), memory stick, floppy disk, mechanical coding devices such as punch cards or groove-in-groove structures having instructions stored thereon, and any suitable combination of the above. As used herein, computer-readable storage medium is not to be construed as a momentary signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated by a waveguide or other transmission medium (e.g., light pulses through a fiber optic cable), or an electrical signal transmitted over an electrical wire.
[0074] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device, or may be downloaded to an external computer or external storage device over a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, fiber optic transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0075] Computer program instructions for carrying out the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source or target code written in any combination of one or more programming languages, including object-oriented programming languages and conventional procedural programming languages. The computer-readable program instructions may execute entirely on the user computer, partially on the user computer, as a separate software package, partially on the user computer and partially on a remote computer, or entirely on a remote computer or server. When a remote computer is involved, the remote computer may be connected to the user computer by any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., connected via the Internet using an Internet service provider). In some embodiments, state information from the computer-readable program instructions is used to customize electronic circuitry capable of executing the computer-readable program instructions, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), to implement aspects of the present disclosure.
[0076] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, or other programmable data processing device to produce an apparatus that, when executed by the processing unit of the computer or other programmable data processing device, implements the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may be stored on a computer-readable storage medium that causes the computer, programmable data processing device, and / or other device to operate in a particular manner, such that the computer-readable medium having the instructions stored thereon includes an article of manufacture containing instructions that implement various aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0077] These computer program instructions may be loaded into a computer, other programmable data processing device, or other apparatus, and cause the computer, other programmable data processing device, or other apparatus to perform a series of operational steps to produce a computer-implemented process, whereby the instructions executing on the computer, other programmable data processing device, or other apparatus implement the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0078] The flowcharts and block diagrams in the figures illustrate possible system architectures, functions, and operations of apparatuses, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams may represent a module, program segment, or part of an instruction set, which includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than the order marked in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, or they may be executed in reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or operation, or by a combination of dedicated hardware and computer instructions.
[0079] Below are listed some exemplary implementations of the present disclosure.
[0080] (Example 1) 1. A method for generating a video, comprising: obtaining content information, including at least one of text or images, related to the content of the video to be generated; obtaining position information indicating a position of an object in the video at a starting frame; obtaining control information that constrains the position of the object in an end frame; generating the video based on the content information, the location information, and the control information.
[0081] (Example 2) The object is a first object, the position information is a first bounding box having a first color, the control information is first control information having the first color, and generating the video includes: obtaining a second bounding box indicating a location in the starting frame of a second object in the video, the second bounding box having a second color different from the first color; obtaining second control information that constrains a position of the second object in the end frame, the second control information having the second color; generating the video based on the content information, the first bounding box, the first control information, the second bounding box, and the second control information.
[0082] (Example 3) The method of Examples 1-2, wherein the position information is a start bounding box and the control information is an end bounding box in the end frame of the object, and the start bounding box and the end bounding box are rectangular boxes.
[0083] (Example 4) generating the video based on the content information, the location information, and the control information; 4. The method of any one of Examples 1 to 3, further comprising: in response to a user selecting a first type as the type of the ending bounding box, generating the video by moving the object from a position indicated by the starting bounding box to a particular position indicated by the ending bounding box, wherein a size of the object in the ending frame corresponds to the ending bounding box.
[0084] (Example 5) generating the video based on the content information, the location information, and the control information; and generating the video by moving the object from a position indicated by the start bounding box to within a range of positions indicated by the end bounding box in response to a user selecting a second type as the type of the end bounding box, wherein a size of the object in the end frame does not exceed the end bounding box, and content of the video is related to the content information.
[0085] (Example 6) generating the video based on the content information, the location information, and the control information; in response to the ending bounding box being close to a left or right boundary of the ending frame and the width of the ending bounding box being less than a threshold width, generating the video by moving the object from its position in the starting frame to outside the left or right boundary of the ending frame, wherein a magnitude of the object as it moves outside the left or right boundary is related to a height of the ending bounding box; or generating the video by moving the object from its position in the start frame to outside the upper or lower boundary of the end frame in response to the ending bounding box being close to an upper or lower boundary of the end frame and the height of the ending bounding box being less than a threshold height, wherein a magnitude of the object as it moves outside the upper or lower boundary is related to a width of the ending bounding box; The method of any one of Examples 1 to 5, wherein the content of the video is related to the content information.
[0086] (Example 7) generating the video based on the content information, the location information, and the control information; in response to the starting bounding box being close to a left or right boundary of a starting frame and the width of the starting bounding box being less than a threshold width, generating the video by moving the object from outside the left or right boundary to a position constrained by the control information, wherein a size of the object as it enters the left or right boundary is related to a height of the starting bounding box; or generating the video by moving the object from outside the upper or lower boundary of the starting frame to a position constrained by the control information in response to the starting bounding box being close to an upper or lower boundary of the starting frame and the height of the starting bounding box being less than a threshold height, wherein a size of the object when it enters the upper or lower boundary is related to a width of the ending bounding box; The method according to any one of Examples 1 to 6, wherein the content of the video is related to the content information.
[0087] (Example 8) The method according to any one of Examples 1 to 7, wherein the position information is a bounding box and the control information is a motion trajectory to be drawn on the starting frame.
[0088] (Example 9) generating the video based on the content information, the location information, and the control information; 9. The method of any one of Examples 1 to 8, further comprising generating the video by moving the object along the motion trajectory from a position indicated by the bounding box, wherein the content of the video is related to the content information.
[0089] (Example 10) the position information is a bounding box, and the content information includes a user-selected image for the starting frame, and generating the video based on the content information, the position information, and the control information includes: determining image content in the starting frame marked by the bounding box as the target; and generating the video based on the image of the starting frame, the object, and the control information, wherein the content of the video is related to the image.
[0090] (Example 11) The content information includes user-entered text that describes the content of the video, examples of which include: identifying noun phrases in the text; determining said noun phrase as said object; and generating the video based on the text, the subject, and the control information, wherein content of the video is related to the text.
[0091] (Example 12) The control information is first control information, and the example is obtaining third control information that constrains a position of the object in the intermediate frame; 12. The method of any one of Examples 1 to 11, further comprising: generating the video based on the content information, the location information, the first control information, and the third control information.
[0092] (Example 13) The method of claim 1 , wherein the object moves gradually from a first position to a second position in the video.
[0093] (Example 14) The method of claim 1 , wherein the video is captured by varying the camera viewing angle to allow the subject to move relative to the camera.
[0094] (Example 15) 1. An apparatus for generating video, comprising: a content information acquisition module configured to acquire content information, including at least one of text or images, related to the content of the video to be generated; a position information acquisition module configured to acquire position information indicating a position of an object in the video at a starting frame; a control information acquisition module configured to acquire control information that constrains a position of the object in an end frame; a video generation module configured to generate the video based on the content information, the location information, and the control information.
[0095] (Example 16) The object is a first object, the position information is a first bounding box having a first color, the control information is first control information having the first color, and generating the video includes: a second bounding box acquisition module configured to acquire a second bounding box indicating a position of a second object in the video at the starting frame, the second bounding box having a second color different from the first color; a second control information acquisition module configured to acquire second control information that constrains a position of the second object in the end frame, the second control information having the second color; and a second bounding box usage module configured to generate the video based on the content information, the first bounding box, the first control information, the second bounding box, and the second control information.
[0096] (Example 17) The device of Examples 15-16, wherein the position information is a start bounding box, the control information is an end bounding box in the end frame of the target, and the start bounding box and the end bounding box are rectangular boxes.
[0097] (Example 18) generating the video based on the content information, the location information, and the control information; The apparatus of Examples 15-17, further comprising: a first type video generation module configured to, in response to a user selecting a first type as the type of the ending bounding box, generate the video by moving the object from a position indicated by the starting bounding box to a specific position indicated by the ending bounding box, wherein a size of the object in the ending frame corresponds to the size of the ending bounding box.
[0098] (Example 19) generating the video based on the content information, the location information, and the control information; The device described in Examples 15 to 18, wherein a second type video generation module generates the video by moving the object from a position indicated by the start bounding box to within a range of positions indicated by the end bounding box in response to a user selecting a second type as the type of the end bounding box, wherein the size of the object in the end frame does not exceed the end bounding box, and the content of the video is related to the content information.
[0099] (Example 20) generating the video based on the content information, the location information, and the control information; a first boundary determination module configured to, in response to the ending bounding box being close to a left or right boundary of the ending frame and the width of the ending bounding box being less than a threshold width, generate the video by moving the object from a position in the starting frame to outside the left or right boundary of the ending frame, wherein a size of the object as it moves outside the left or right boundary is related to a height of the ending bounding box; or a second boundary determination module configured to, in response to the ending bounding box being close to an upper boundary or a lower boundary of the ending frame and a height of the ending bounding box being less than a threshold height, generate the video by moving the object from a position in the starting frame to outside the upper boundary or the lower boundary of the ending frame, wherein a size of the object when it moves outside the upper boundary or the lower boundary is related to a width of the ending bounding box; 20. The apparatus of any one of Examples 15 to 19, wherein the content of the video is associated with the content information.
[0100] (Example 21) generating the video based on the content information, the location information, and the control information; a third boundary determination module configured to, in response to the starting bounding box being close to a left or right boundary of a starting frame and the width of the starting bounding box being less than a threshold width, generate the video by moving the object from outside the left or right boundary to a position constrained by the control information, wherein a size of the object when it enters the left or right boundary is related to a height of the starting bounding box; or a fourth boundary determination module configured to, in response to the starting bounding box being close to an upper boundary or a lower boundary of the starting frame and a height of the starting bounding box being less than a threshold height, generate the video by moving the object from outside the upper boundary or the lower boundary of the starting frame to a position constrained by the control information, wherein a size of the object when it enters the upper boundary or the lower boundary is related to a width of the ending bounding box; 21. The apparatus of any one of Examples 15 to 20, wherein the content of the video is related to the content information.
[0101] (Example 22) 22. The device of Examples 15-21, wherein the position information is a bounding box and the control information is a motion trajectory to be drawn on the starting frame.
[0102] (Example 23) generating the video based on the content information, the location information, and the control information; The device of Examples 15-22, wherein the motion trajectory usage module is configured to generate the video by moving the object along the motion trajectory from a position indicated by the bounding box, wherein the content of the video is related to the content information.
[0103] (Example 24) The content information includes an image for the starting frame selected by a user, and generating the video based on the content information, the position information, and the control information includes: an object determination module configured to determine image content in the starting frame marked by the bounding box as the object; and generating the video based on the image, the object, and the control information of the starting frame, wherein the content of the video is related to the image.
[0104] (Example 25) The content information includes user-entered text that describes the content of the video, examples of which include: a noun identification module configured to identify noun phrases in the text; a noun usage module configured to determine the noun phrase as the target; and a subject usage module configured to generate the video based on the text, the subject, and the control information, wherein content of the video is related to the text.
[0105] (Example 26) The control information is first control information, and the example is a third control information acquisition module configured to acquire third control information that constrains a position of the object in the intermediate frame; and a third control information using module configured to generate the video based on the content information, the location information, the first control information, and the third control information.
[0106] (Example 27) 27. The apparatus of Examples 15-26, wherein the object moves gradually from a first position to a second position in the video.
[0107] (Example 28) The apparatus of Examples 15 to 27, wherein the video is captured by changing the camera viewing angle to move the subject relative to the camera.
[0108] (Example 29) a processor; 1. An electronic device including a processor and a memory coupled to the processor, the memory having instructions stored therein that, when executed by the processor, cause the electronic device to perform a method for generating video, the method comprising: obtaining content information, including at least one of text or images, related to the content of the video to be generated; obtaining position information indicating a position of an object in the video at a starting frame; obtaining control information that constrains the position of the object in an end frame; generating the video based on the content information, the location information, and the control information.
[0109] (Example 30) The object is a first object, the position information is a first bounding box having a first color, the control information is first control information having the first color, and generating the video includes: obtaining a second bounding box indicating a location in the starting frame of a second object in the video, the second bounding box having a second color different from the first color; obtaining second control information that constrains a position of the second object in the end frame, the second control information having the second color; and generating the video based on the content information, the first bounding box, the first control information, the second bounding box, and the second control information.
[0110] (Example 31) The device of Examples 29 to 30, wherein the position information is a start bounding box, the control information is an end bounding box in the end frame of the target, and the start bounding box and the end bounding box are rectangular boxes.
[0111] (Example 32) generating the video based on the content information, the location information, and the control information; 32. The device of Examples 29-31, wherein in response to a user selecting a first type as the type of the ending bounding box, generating the video by moving the object from a position indicated by the starting bounding box to a specific position indicated by the ending bounding box, wherein a size of the object in the ending frame corresponds to the size of the ending bounding box.
[0112] (Example 33) generating the video based on the content information, the location information, and the control information; The device described in Examples 29 to 32 includes, in response to a user selecting a second type as the type of the ending bounding box, generating the video by moving the object from a position indicated by the starting bounding box to within a range of positions indicated by the ending bounding box, wherein the size of the object in the ending frame does not exceed the ending bounding box, and the content of the video is related to the content information.
[0113] (Example 34) generating the video based on the content information, the location information, and the control information; in response to the ending bounding box being close to a left or right boundary of the ending frame and the width of the ending bounding box being less than a threshold width, generating the video by moving the object from its position in the starting frame to outside the left or right boundary of the ending frame, wherein a magnitude of the object as it moves outside the left or right boundary is related to a height of the ending bounding box; or generating the video by moving the object from its position in the start frame to outside the upper or lower boundary of the end frame in response to the ending bounding box being close to an upper or lower boundary of the end frame and the height of the ending bounding box being less than a threshold height, wherein a magnitude of the object as it moves outside the upper or lower boundary is related to a width of the ending bounding box; 34. The device of any one of Examples 29 to 33, wherein the video content is related to the content information.
[0114] (Example 35) generating the video based on the content information, the location information, and the control information; in response to the starting bounding box being close to a left or right boundary of a starting frame and the width of the starting bounding box being less than a threshold width, generating the video by moving the object from outside the left or right boundary to a position constrained by the control information, wherein a size of the object as it enters the left or right boundary is related to a height of the starting bounding box; or generating the video by moving the object from outside the upper or lower boundary of the starting frame to a position constrained by the control information in response to the starting bounding box being close to an upper or lower boundary of the starting frame and the height of the starting bounding box being less than a threshold height, wherein a size of the object when it enters the upper or lower boundary is related to a width of the ending bounding box; The device of Examples 29 to 34, wherein the video content is related to the content information.
[0115] (Example 36) The device of Examples 29 to 35, wherein the position information is a bounding box and the control information is a motion trajectory to be drawn on the starting frame.
[0116] (Example 37) generating the video based on the content information, the location information, and the control information; The device of Examples 29-36, further comprising generating the video by moving the object along the motion trajectory from a position indicated by the bounding box, wherein the content of the video is related to the content information.
[0117] (Example 38) the content information includes an image for the starting frame selected by a user, and generating the video based on the content information, the position information, and the control information includes: determining image content in the starting frame marked by the bounding box as the target; and generating the video based on the image, the object, and the control information of the starting frame, wherein the content of the video is related to the image.
[0118] (Example 39) The content information includes user-entered text that describes the content of the video, examples of which include: identifying noun phrases in the text; determining said noun phrase as said object; and generating the video based on the text, the subject, and the control information, wherein the content of the video is related to the text.
[0119] (Example 40) The control information is first control information, and the example is obtaining third control information that constrains a position of the object in the intermediate frame; 40. The device of Examples 29-39, further comprising: generating the video based on the content information, the location information, the first control information, and the third control information.
[0120] (Example 41) The apparatus of Examples 29-40, wherein the object moves gradually from a first position to a second position in the video.
[0121] (Example 42) The device of Examples 29 to 41, wherein the video can be moved relative to the object and the camera by changing the camera viewing angle.
[0122] Although the embodiments of the present disclosure have been described above, the above description is illustrative and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used in this specification is intended to best interpret the principles, practical applications, or technical improvements of the technology in the marketplace of the embodiments, or to enable those skilled in the art to understand the embodiments disclosed herein.
Claims
1. 1. A method for generating a video, comprising: obtaining content information, including at least one of text or images, related to the content of the video to be generated; obtaining position information indicating a position of an object in the video at a starting frame; obtaining control information that constrains the position of the object in an end frame; generating the video based on the content information, the location information, and the control information; A method comprising:
2. the object is a first object, the position information is a first bounding box having a first color, and the control information is first control information having the first color; generating the video obtaining a second bounding box indicating a position of a second object in the video at the starting frame, the second bounding box having a second color different from the first color; obtaining second control information that constrains a position of the second object in the end frame, the second control information having the second color; generating the video based on the content information, the first bounding box, the first control information, the second bounding box, and the second control information; The method of claim 1 , comprising:
3. The method of claim 1 , wherein the position information is a start bounding box and the control information is an end bounding box of the object in the end frame, the start bounding box and the end bounding box being rectangular boxes.
4. generating the video based on the content information, the location information, and the control information; generating the video by moving the object from a position indicated by the start bounding box to a particular position indicated by the end bounding box in response to a user selecting a first type as the type of the ending bounding box; The size of the object in the ending frame corresponds to the ending bounding box. The method of claim 3.
5. generating the video based on the content information, the location information, and the control information; in response to a user selecting a second type as the type of the ending bounding box, generating the video by moving the object from a position indicated by the starting bounding box to within a range of positions indicated by the ending bounding box; The method of claim 3 or 4, wherein the size of the object in the ending frame does not exceed the ending bounding box, and the content of the video is related to the content information.
6. generating the video based on the content information, the location information, and the control information; in response to the ending bounding box being close to a left or right boundary of the ending frame and the width of the ending bounding box being less than a threshold width, generating the video by moving the object from its position in the starting frame to outside the left or right boundary of the ending frame, wherein a magnitude of the object as it moves outside the left or right boundary is related to a height of the ending bounding box; or generating the video by moving the object from its position in the start frame to outside the upper or lower boundary of the end frame in response to the ending bounding box being close to an upper or lower boundary of the end frame and the height of the ending bounding box being less than a threshold height, wherein a magnitude of the object as it moves outside the upper or lower boundary is related to a width of the ending bounding box; The method of claim 3 , wherein the content of the video is related to the content information.
7. generating the video based on the content information, the location information, and the control information; in response to the starting bounding box being close to a left or right boundary of a starting frame and the width of the starting bounding box being less than a threshold width, generating the video by moving the object from outside the left or right boundary to a position constrained by the control information, wherein a size of the object as it enters the left or right boundary is related to a height of the starting bounding box; or generating the video by moving the object from outside the upper or lower boundary of the starting frame to a position constrained by the control information in response to the starting bounding box being close to an upper or lower boundary of the starting frame and the height of the starting bounding box being less than a threshold height, wherein a size of the object when it enters the upper or lower boundary is related to a width of the ending bounding box; The method of claim 3 , wherein the content of the video is related to the content information.
8. The method of claim 1 , wherein the position information is a bounding box and the control information is a motion trajectory that is drawn on the starting frame.
9. generating the video based on the content information, the location information, and the control information; generating the video by moving the object along the motion trajectory from a position indicated by the bounding box; The method of claim 8 , wherein the content of the video is related to the content information.
10. the location information is a bounding box, and the content information includes a user-selected image for the starting frame; generating the video based on the content information, the location information, and the control information; determining image content in the starting frame marked by the bounding box as the target; generating the video based on the image of the starting frame, the object, and the control information; the content of the video is related to the image; The method of claim 1.
11. the content information includes user-entered text that describes the content of the video; The method comprises: identifying noun phrases in the text; determining said noun phrase as said object; generating the video based on the text, the object, and the control information; further comprising the content of the video is related to the text; The method of claim 10.
12. the control information is first control information, The method comprises: obtaining third control information that constrains a position of the object in the intermediate frame; generating the video based on the content information, the location information, the first control information, and the third control information; The method of claim 1 further comprising:
13. The method of claim 1 , wherein the object moves gradually from a first position to a second position in the video.
14. The method of claim 1 , wherein the video is captured by varying the camera angle to allow relative movement between the subject and the camera.
15. 1. An apparatus for generating video, comprising: a content information acquisition module configured to acquire content information, including at least one of text or images, related to the content of the video to be generated; a position information acquisition module configured to acquire position information indicating a position of an object in the video at a starting frame; a control information acquisition module configured to acquire control information that constrains a position of the object in an end frame; a video generation module configured to generate the video based on the content information, the location information, and the control information; 1. An apparatus comprising:
16. a processor; an electronic device including a memory coupled to the processor, 15. An electronic device, wherein the memory has instructions stored therein that, when executed by a processor, cause the electronic device to perform a method according to any one of claims 1 to 14.
17. 15. A computer program product tangibly stored on a non-transitory computer-readable medium and comprising machine-executable instructions that, when executed, cause a machine to perform the method of any one of claims 1 to 14.
Citation Information
Patent Citations
Interactive video generation
JP2018503279A
Interactive video generation
JP2019154045A
Video generation program, video generation device, and video generation method
JP2021033961A
Video processing method, device, electronic apparatus, and storage medium
JP2021193559A