Methods, apparatus, devices, and computer programs for generating video.
By using bounding boxes to input content and position constraints, the method addresses the challenge of accurately representing complex object movements in videos, resulting in videos that meet user expectations and enhance creativity.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-07-17
- Publication Date
- 2026-03-31
AI Technical Summary
Existing video generation technologies struggle to accurately interpret user requirements for object movement in videos, particularly when precise motion control is needed, as users find it difficult to describe complex movements in natural language.
A method and apparatus that allows users to input content information and position constraints using bounding boxes in start and end frames, enabling precise motion control by generating videos based on these inputs.
Enables the generation of videos that accurately meet user expectations by allowing users to mark objects and control their movement using bounding boxes, enhancing the flexibility and creativity of video creation.
Smart Images

Figure 0007838031000001 
Figure 0007838031000002 
Figure 0007838031000003
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of artificial intelligence, and more specifically, to a method, an apparatus, an electronic device, and a computer program for generating a video.
Background Art
[0002] Video generation guided by text is a technology that uses natural language text to guide video content generation. With deep learning and artificial intelligence technologies, a system can understand the input text description, convert it into specific visual content, and generate the corresponding video. Such a method is applicable in fields such as movie production, virtual reality, and video production, providing a more intuitive and efficient way for creators to express their ideas.
[0003] Motion control refers to realizing various effects and dynamic effects in a video by precisely controlling the movement of an object or a camera. Such technology can be realized by a programming or automation system, making video production more creative and dynamic. Motion control is widely applied in fields such as movies and virtual reality, providing a more immersive and attractive visual experience for viewers.
[0004] By combining video generation guided by text and motion control, it is possible to realize more intelligent and personalized video creation. By being guided by text, creators can express the scenarios and effects they desire in a natural language manner, and motion control ensures that these ideas are presented in the video accurately and smoothly, providing greater flexibility and creativity for the creation process.
Summary of the Invention
Problems to be Solved by the Invention
[0005] Solve the problems of the prior art. [Means for solving the problem]
[0006] A first embodiment of the embodiments of this disclosure provides a method for generating a video. This method includes obtaining content information, which includes at least one of text or an image, relating to the content of the video to be generated. This method further includes obtaining location information indicating the position of a subject in the video at the beginning of the video. This method further includes obtaining control information that constrains the position of the subject at the end of the video. This method further includes generating the video based on the content information, location information, and control information.
[0007] A second embodiment of the embodiments of this disclosure provides an apparatus for generating a video. This apparatus includes a content information acquisition module configured to acquire content information, which includes at least one of text or an image relating to the content of the video to be generated. This apparatus further includes a location information acquisition module configured to acquire location information indicating the position of a subject in the starting frame of the video. This apparatus further includes a control information acquisition module configured to acquire control information that constrains the position of the subject in the ending frame. This apparatus further includes a video generation module configured to generate the video based on the content information, location information, and control information.
[0008] A third embodiment of the embodiments of the present disclosure provides an electronic device comprising one or more processors and a memory device for storing one or more programs, wherein when one or more programs are executed by one or more processors, the one or more processors implement a method for generating a video. This method includes obtaining content information, which includes at least one of text or images, relating to the content of the video to be generated. This method further includes obtaining location information indicating the position of a subject in the video at the beginning frame. This method further includes obtaining control information that constrains the position of the subject at the end frame. This method further includes generating the video based on the content information, location information, and control information.
[0009] A fourth aspect of the embodiments of this disclosure provides a computer program which includes instructions that are physically stored in a non-temporary computer-readable medium and are device-executable when these instructions are executed, causing a device to implement a method for generating a video. This method includes obtaining content information which includes at least one of text or images relating to the content of the video to be generated. This method further includes obtaining position information which indicates the position of a subject in the video at the beginning of the video. This method further includes obtaining control information which constrains the position of the subject at the end of the video. This method further includes generating the video based on the content information, position information and control information.
[0010] The summary section of the invention is provided to briefly introduce the selection of concepts and is further described in the specific embodiments below. The summary section of the invention is not intended to mark any important or main features of the subject matter for which protection is sought, nor is it intended to limit the scope of the subject matter for which protection is sought. [Brief explanation of the drawing]
[0011] The above and other features, advantages and aspects of each embodiment of this disclosure will become more apparent by referring to the following detailed description in conjunction with the drawings. In the drawings, the same or similar reference numerals represent the same or similar elements, where, [Figure 1] A schematic diagram of an exemplary environment in which several embodiments of this disclosure may be realized is shown. [Figure 2] A flowchart of a method for generating video according to some embodiments of this disclosure is shown. [Figure 3A] The following are schematic diagrams illustrating examples of how a user inputs multiple bounding boxes in the start and end frames, respectively, according to some embodiments of this disclosure. [Figure 3B] The following are schematic diagrams illustrating examples of how a user inputs multiple bounding boxes in the start and end frames, respectively, according to some embodiments of this disclosure. [Figure 4A] The following are schematic diagrams illustrating examples of how a user inputs the target bounding box and motion trajectory in the starting frame, according to some embodiments of this disclosure. [Figure 4B] The following are schematic diagrams illustrating examples of how a user inputs the target bounding box and motion trajectory in the starting frame, according to some embodiments of this disclosure. [Figure 5A] The following are schematic diagrams illustrating examples of the present disclosure in which a user enters multiple soft bounding boxes in the end frame and does not enter any text description. [Figure 5B] The following are schematic diagrams illustrating examples of the present disclosure in which a user enters multiple soft bounding boxes in the end frame and does not enter any text description. [Figure 6A] The following is a schematic diagram illustrating an example in which a user, according to some embodiments of this disclosure, inputs multiple bounding boxes on a starting frame where no image is provided. [Figure 6B] The following is a schematic diagram illustrating an example in which a user, according to some embodiments of this disclosure, inputs multiple bounding boxes on a starting frame where no image is provided. [Figure 7] The following are schematic diagrams illustrating examples of how a user creates an intermediate frame and inputs multiple soft bounding boxes within that intermediate frame, according to some embodiments of this disclosure. [Figure 8] A block diagram of an apparatus for generating video according to some embodiments of this disclosure is shown. [Figure 9] A block diagram of equipment capable of realizing multiple embodiments of this disclosure is shown. [Modes for carrying out the invention]
[0012] To ensure clarity, all user-related data related to this proposed technology should be acquired and used only after obtaining the user's permission. This means that if it is necessary to use a user's personal information in this proposed technology, the user's explicit consent and permission are required before acquiring this data; otherwise, the relevant data will not be collected or used. Furthermore, it should be understood that when implementing this proposed technology, all applicable laws and regulations regarding data collection, use, and storage must be strictly observed, and necessary technologies and measures must be taken to guarantee the security of user data and ensure its safe use.
[0013] The following describes embodiments of the Disclosure in more detail with reference to the drawings. While the drawings illustrate several embodiments of the Disclosure, it should be understood that the Disclosure can be implemented in various forms and should not be interpreted as being limited to the embodiments described herein. Rather, these embodiments are provided for a better and more complete understanding of the Disclosure. It should also be understood that the drawings and embodiments of the Disclosure are illustrative and not intended to limit the scope of protection of the Disclosure.
[0014] In the descriptions of the embodiments of this disclosure, the term “including” and similar terms should be understood as non-restrictive “including,” i.e., “including, but not limited to.” The term “based on” should be understood as “based at least in part.” The term “one embodiment” or “this embodiment” should be understood as “at least one embodiment.” Unless explicitly stated otherwise, terms such as “first,” “second,” etc., may refer to different or the same subject. The following may include other explicit and implicit definitions.
[0015] In a video generation scenario guided by some text or reference images, the user desires to provide information regarding the movement method for an object in the generated video by inputting a text description. For example, the user may provide a reference image in which a building is photographed and input a text description such as "tilt the camera upward to expose the top of the building". At this time, the user hopes that in the generated video, the camera gradually raises the lens from the perspective of shooting the ground and finally shoots up to the top of the building. However, in the related art, based on the reference image and text description provided by the user, a video with relatively high screen quality and a lens that moves slowly can be generated, but the model cannot well understand the user's requirements for the movement method of the object in the video, so the top of the building cannot be accurately exposed in the generated video.
[0016] In addition to this, in some usage scenarios, when the user's requirements for the movement method are very precise, it is difficult to accurately describe the envisioned screen in language characters. For example, when the user hopes that in the video, two puppies run towards the camera, one of the white puppies approaches the camera and runs to the center position of the screen. At this time, this white puppy fills one-third of the screen. At the same time, the other black puppy also approaches the camera, but the running direction is towards a toy next to the camera, so it deviates from the center of the screen and finally disappears from the right side of the screen. For the average user, it is very difficult to accurately describe such movement requirements, and thus the desired video cannot be generated.
[0017] Therefore, embodiments of the present disclosure provide a solution for generating a video. In this solution, a user can input content information related to the content of the video to be generated in a user interface provided by an application. Such content information may be a single text description, a key frame of the video, or both may be provided. And the user can label the motion control target with a bounding box at the start frame, and can input control information in the user interface on how this target moves. The control information includes at least the position of this target at the end frame. And the application can generate a video based on the content information, the bounding box labeling the target, and the control information.
[0018] In such a manner, the user can accurately label the target to be controlled using the bounding box at the start frame. In addition to this, the user can accurately represent the movement position of the labeled target from the start frame to the end frame, and the application can receive an accurate motion control requirement, thereby generating a video that meets the user's expectations.
[0019] FIG. 1 shows a schematic diagram of an exemplary environment 100 in which multiple embodiments of the present disclosure can be realized. As shown in FIG. 1, the environment 100 includes a user 102 and a computing device 104. The computing device 104 may be a device such as a user terminal, a mobile device, a computer, etc., or may be a device such as a computing system, a single server, a distributed server, or a cloud-based server. The computing device 104 includes an application 106 that can provide a function of generating a video. The application 106 may be any application such as a local application, a remote application, an application of a browser / server architecture, or an application of a client / server architecture.
[0020] As shown in Figure 1, in environment 100, user 102 can interact with application 106 via user interface 108. User 102 may input content information 110 related to the content of the video to be generated on user interface 108. Content information 110 may be a text description 112, a reference image 114, or both simultaneously. For example, the text description 112 may be "Two puppies are running towards the camera," and the reference image 114 may be an image of two puppies running on the grass. It should be noted that in some implementations, the text description 112 may be provided indirectly through a user interaction method such as audio, and therefore the text description 112 also includes text provided indirectly through methods such as audio. User 102 can mark the object whose movement they want to control by drawing bounding boxes 118-1, 118-2, ..., 118-N (collectively referred to as bounding boxes 118) in the area where the start frame 116 is displayed. It should be understood that although multiple bounding boxes 118 are shown in Figure 1, the user may draw only one bounding box 118 to control the movement of only one object. If the content information 110 includes a reference image 114, the reference image 114 may be displayed in the area of the starting frame 116, thereby making it easier for the user to use the bounding box 118 to mark the object whose movement they want to control.
[0021] As shown in Figure 1, in environment 100, user 102 may input control information 120 via user interface 108 so as to describe at least the position of the video to be controlled at the end frame. In some embodiments, the control information 120 may be a bounding box drawn by user 102 in the end frame region to represent the position to which the controlled object can be moved. In some embodiments, the control information 120 may be the movement trajectory of the controlled object drawn by user 102 in the region of the start frame 116.
[0022] In environment 100, after user 102 inputs content information 110, a bounding box 118 at the start frame 116, and control information 120, application 106 can generate a video 122 based on these user inputs and provide the video 122 to user 102 via user interface 108. For example, computing device 104 may send these user inputs to a server and receive the address of video 122 from the server, or computing device 104 may generate video 122 locally. In user interface 108, video 122 may be displayed to user 102, for example, via a video playback control, or provided to user 102 in the form of a download control. The content of video 122 relates to content information 110 and is marked by a bounding box 118 in video 122. The controlled object moves from its position at the start frame 116 to a specified position at the end frame according to the constraints of control information 120.
[0023] It should be understood that in environment 100, content information 110, start frame 116, and control information 120 are included in a single user interface 108. However, in some embodiments, user 102 may input this information through different user interfaces. Furthermore, video 122 may be provided to user 102 through a separate user interface.
[0024] This method allows user 102 to accurately mark the object they want to control using the bounding box 118 in the starting frame 116. In addition, user 102 can accurately represent the position of the marked object's movement from the starting frame 116 to the ending frame, and application 106 can receive precise motion control requests, thereby generating a video that meets the user's expectations.
[0025] Figure 2 shows a flowchart of a method 200 for generating a video according to some embodiments of the present disclosure. As shown in Figure 2, in box 202, the method 200 can obtain content information which includes at least one of text or an image relating to the content of the video to be generated. For example, in the environment 100 shown in Figure 1, a computing device 104 can obtain content information 110 relating to the video content to be generated, which is input by user 102, and the content information 110 includes at least one of a text description 112 and a reference image 114.
[0026] In box 204, method 200 can acquire positional information indicating the position of an object in the video at the start frame. This positional information may be position-related information such as a bounding box, contour, coordinate values, or coordinate range. For example, in the environment 100 shown in Figure 1, computing device 104 can acquire a bounding box 118 input by user 102 at the start frame 116, and the bounding box 118 is used to mark the motion-controlled object at the start frame 116. If user 102 provides a reference image 114, the content of the start frame 116 may be the reference image 114, thereby facilitating the user to directly mark the motion-controlled object on the reference image 114. If user 102 does not provide a reference image 114, the content of the start frame 116 may be blank, and the user can mark the position and size of the motion-controlled object in the blank area with the bounding box 118.
[0027] In box 206, method 200 can acquire control information that constrains the position of the target at the end frame. For example, in the environment 100 shown in Figure 1, computing device 104 can acquire control information 120 input by user 102, and the control information 120 can constrain the position of the controlled object at the end frame, which is marked by a bounding box 118. In some embodiments, user 102 may provide the control information 120 by drawing a bounding box at the end frame. In some embodiments, the user may provide the control information 120 by drawing a motion trajectory at the start frame 116.
[0028] In box 208, method 200 can generate a video based on content information, position information, and control information. For example, in the environment 100 shown in Figure 1, computing device 104 can generate a video 122 based on content information 110, bounding box 118, and control information 120. The content of video 122 relates to content information 110, and the controlled object marked by the bounding box 118 in video 122 moves from its position in the starting frame 116 to a specified position according to the constraints of control information 120.
[0029] This method allows the user to accurately mark the object they want to control using positional information from the starting frame. In addition, the user can accurately represent the position of the marked object's movement from the starting frame to the ending frame, allowing the application to receive precise motion control requests and thereby generate videos that meet the user's expectations.
[0030] In some embodiments, the above object is a first object, the above position information is a first bounding box having a first color, the above control information is first control information having a first color, and generating a video may include obtaining a second bounding box indicating the position of a second object in the video at the start frame, the second bounding box having a second color different from the first color. It is also possible to obtain second control information that constrains the position of the second object at the end frame, the second control information having a second color. Then, a video can be generated based on the content information, the first bounding box, the first control information, the second bounding box, and the second control information.
[0031] In some embodiments, the position information is a start bounding box, the control information is an end bounding box in the target's end frame, and both the start and end bounding boxes are rectangular boxes. In some embodiments, in response to a user selecting a first type as the type of end bounding box, a video may be generated by moving the target from a position indicated by the start bounding box to a specific position indicated by the end bounding box, where the size of the target in the end frame corresponds to the end bounding box. In some embodiments, the image content marked by the bounding box in the start frame can be determined as the target, and a video can be generated based on the image, target, and control information of the start frame, where the video content is related to the image.
[0032] In some embodiments, the video can be generated by moving an object from its position in the start frame to outside the left or right boundary of the end frame in response to the end bounding box being close to the left or right boundary of the end frame and the width of the end bounding box being smaller than a threshold width, the size of the object when it moves outside the left or right boundary being related to the height of the end bounding box. In some embodiments, the video can be generated by moving an object from its position in the start frame to outside the upper or lower boundary of the end frame in response to the end bounding box being close to the upper or lower boundary of the end frame and the height of the end bounding box being smaller than a threshold height, the size of the object when it moves outside the upper or lower boundary being related to the width of the end bounding box, where the content of the video is related to content information.
[0033] In some embodiments, a video is generated by moving an object from outside the left or right boundary to a position constrained by control information in response to the start boundary box being close to the left or right boundary of the start frame and the width of the start boundary box being smaller than a threshold width, and the size of the object when it enters the left or right boundary is related to the height of the start boundary box. In some embodiments, a video is generated by moving an object from outside the upper or lower boundary of the start frame to a position constrained by control information in response to the start boundary box being close to the upper or lower boundary of the start frame and the height of the start boundary box being smaller than a threshold height, and the size of the object when it enters the upper or lower boundary is related to the width of the end boundary box, where the content of the video is related to content information.
[0034] In some embodiments, the object moves gradually from a first position to a second position in the generated video. In some embodiments, the generated video can move the object and the camera relative to each other by changing the camera's viewing angle.
[0035] Figures 3A-3B show schematic diagrams of Example 300 in which the user inputs multiple bounding boxes in the start and end frames, respectively, according to some embodiments of the present disclosure. As shown in Figure 3A, Example 300 includes an input user interface 302, which provides a control 304 for inputting a text description. The user can input a text description in the text input control 304 to describe the content of the video to be generated. For example, in Example 300, if the user inputs the text description "Two puppies running towards the camera", the content of the generated video should include two puppies running towards the camera. The user interface 302 further provides a selection picture control 306 for inputting a reference image, which the user can select by interacting with the control 306. If the user selects a reference image, the reference image can guide the content of the generated video. For example, in Example 300, if the user selects an image of two puppies (including one white puppy and one black puppy) running on a lawn, the generated video should include these two puppies running on the lawn in the background.
[0036] As shown in Figure 3A, the user interface 302 further provides a start frame control 308, and if the user selects a reference image, the content of the start frame control 308 can display this reference image. If the user has not selected a reference image, the content of the start frame control 308 may be blank (similar to the control 318 for representing the end frame in Figure 3). The user interface 302 further provides a color display control 310 corresponding to object 1 and a color display control 312 corresponding to object 2, associated with the start frame control 308, where the colors corresponding to each object are all different. When the user interacts with the color display control 310 or 312, they can select which object to currently mark on the start frame control 308. For example, in Example 300, object 1 is selected, and the operation performed next on the start frame control 308 is to indicate that the operation will be on object 1. The user interface 302 further provides an operation type selection control 314, where the operation types include bounding box and motion trajectory. If the value of the operation type selection control 314 is bounding box, then a starting bounding box marking target 1 can be input by dragging in the start frame control 308. If the value of the operation type selection control 314 is motion trajectory, then a motion trajectory for target 1 can be input by brushing in the start frame control 308.
[0037] As shown in Figure 3A, the user interface 302 further provides an end frame control 318, the content of which may be blank, and the user can draw an end bounding box in the end frame control 318, which constrains the position and size of the target's destination. The user interface 302 further provides a color display control 320 corresponding to target 1 and a color display control 312 corresponding to target 2, associated with the end frame control 318, where the value of color display control 320 is the same as the value of color display control 310, and the value of color display control 322 is the same as the value of color display control 312. In other words, the same target is represented by the same color in the start frame control 308 and the end frame control 318. It should be understood that in the embodiments of this disclosure, a specific color such as black and white is assigned to each target, but this is not intended to limit the color corresponding to each target, and this color may be other colors such as yellow, purple, or green.
[0038] As shown in Figure 3A, the user interface 302 further provides a bounding box type selection control 324. There are two types of bounding boxes: hard bounding boxes and soft bounding boxes. A hard bounding box is used to specify a specific position and size of an object, and it indicates that in the generated image frame, the object will be generated at the coordinates specified by the hard bounding box (e.g., the center coordinates of the bounding box), and the size of the object will correspond to the size of the hard bounding box. A soft bounding box is used to specify a position range and size range of an object, and it indicates that in the generated image frame, the object will be generated within the range defined by the soft bounding box, and the size of the object will not exceed this range. If the user selects object 1 to be manipulated in the end frame and the value of the bounding box type control is hard bounding box, then the end frame control 318 can draw a hard bounding box that constrains the specific position of object 1 after it has moved, for example, by dragging. If the user selects object 1 to be manipulated in the end frame and the value of the bounding box type control is soft bounding box, then the end frame control 318 can draw a soft bounding box that constrains the position range of object 1 after it has moved.
[0039] As shown in Figure 3A, in Example 300, the user marks the white puppy as Object 1 with a black bounding box 330 in the start frame control 308, and marks the specific position and size to which Object 1 should move with a black hard bounding box 334 in the end frame control 318. The user also marks the black puppy as Object 2 with a white bounding box 332 in the start frame control 308, and marks the position range and size range to which Object 2 should move with a white hard bounding box 336 in the end frame control 318.
[0040] As shown in Figure 3A, in the end frame control 318 of Example 300, the hard bounding box 336 for subject 2 is close to the right boundary of the end frame and is relatively narrow in width (e.g., smaller than the threshold width), which means that when the black puppy corresponding to subject 2 runs out of the right boundary of the video, its size is related to the height of the hard bounding box 336. In other words, the height of the hard bounding box 336 is the projection of the height onto the right boundary when subject 2 moves outside the right boundary.
[0041] The user can then generate a video by clicking the video generation control 328, and the generated video is as shown in Figure 3B. In the video 340 as shown in Figure 3B, the white puppy (object 1) runs from its position in the starting frame to the position specified by the hard bounding box 334 in the ending frame, and the size of the white puppy in the ending frame corresponds to the hard bounding box 334. At the same time, the black puppy (object 2) runs from its position in the starting frame to off-screen, and when the black puppy in the third frame begins to run off-screen, its height may be related to the hard border frame 324.
[0042] In this manner, users can mark multiple control objects using different colors, thereby making the generated video content more dynamic and richer. Furthermore, by marking the bounding box in the final frame, the position and size of the object's movement can be precisely controlled, allowing users to accurately represent and control the object's movement, improving the motion effects of the generated video and enhancing the user experience. In addition, it is possible to implement effects that allow objects to move off-screen, and to specify their position and size when they move off-screen, thereby providing rich motion control commands and meeting the needs of users in different scenarios.
[0043] In some embodiments, the position information is a bounding box, and the control information is a motion trajectory drawn on the starting frame. In some embodiments, a video can be generated by moving an object from a position indicated by the bounding box along the motion trajectory, where the content of the video is related to the content information.
[0044] Figures 4A-4B show schematic diagrams of Example 400, in which the user inputs a bounding box and motion trajectory of an object in a starting frame, according to some embodiments of the present disclosure. As shown in Figure 4A, in Example 400, the user inputs the text description "One person is throwing a frisbee" in the text input control 404 of the user interface 402, and as shown in the starting frame control 408, the user has selected a reference image, which includes one person holding a frisbee. In Example 400, the user marks object 1 (i.e., the frisbee) with a black bounding box 432 in the starting frame control 408. The user then selects the value of the operation type selection control 414 as the motion trajectory and draws the motion trajectory 434 of the frisbee in the starting frame control 408. As shown in Figure 4A, the motion trajectory 434 of the frisbee is similar in shape to a "U", with the starting and ending positions close together, indicating that the frisbee flies out from the starting position and eventually flies back to the position indicated by the arrows in the trajectory. Furthermore, in the start frame control 408, the user marks a person as object 2 with a white bounding box 430, and in the end frame control 418, marks the hard bounding box 436 of object 2. This indicates that object 2 moves from its position in the start frame to the position specified by the hard bounding box 436, and that its size after the move corresponds to the size of the hard bounding box 436.
[0045] Figure 4B shows the generated video 440, in which, as shown in the figure, the frisbee (object 1) flies from the starting position along the motion trajectory 434 and finally flies to the endpoint of the motion trajectory 434. At the same time, the person (object 2), after going through a series of movements, is positioned at the location specified by the hard bounding box 436, and the size of the person corresponds to the hard bounding box 436.
[0046] In this method, users can control the movement of an object by drawing its movement trajectory, and because the movement trajectory contains richer information about the movement process, they can control the object's movement more precisely. In addition, since it is difficult for users to describe slightly complex movement trajectories in text, drawing the movement trajectory makes it easier for users to represent how an object moves, thereby enhancing the user experience.
[0047] In some embodiments, in response to a user selecting a second type as the type of end bounding box, a video is generated by moving the object from the position indicated by the start bounding box to the position range indicated by the end bounding box, where the size of the object in the end frame does not exceed the end bounding box, and the content of the video is related to content information.
[0048] Figures 5A-5B show schematic diagrams of Example 500 in which the user enters multiple soft bounding boxes in the end frame and does not enter any text description, according to some embodiments of the present disclosure. As shown in Figure 5A, in Example 500, the user has selected a reference image but has not entered any content in the text input control 504 of the user interface 502. As shown in the start frame control 508, the reference image shows multiple eggs in a basket. In Example 500, the user marks object 1 (i.e., the egg located to the right of the basket) with a black bounding box 528, object 2 (i.e., the egg located behind the basket) with a white bounding box 530, and object 3 (i.e., the egg located to the left of the basket) with a gray bounding box 532 in the start frame control 508. In the final frame, the user selects the values of the bounding box type selection controls 522, 524, and 526 for target 1, target 2, and target 3 as soft bounding boxes. In the final frame control 518, the user specifies the position range and size range for the destination of the right egg in the basket using the black soft bounding box 534. That is, the position of the right egg in the basket in the final frame does not exceed the range specified by the soft bounding box 534, and its size does not exceed the soft bounding box 534. Furthermore, in the final frame control 518, the user specifies the position range and size range for the destination of the left egg in the basket using the white soft bounding box 536, and the position range and size range for the destination of the back egg in the basket using the gray soft bounding box 538.
[0049] Figure 5B shows the generated video 540, and as shown in the figure, the egg on the right of the basket (object 1) moves from the starting position to within the area defined by the soft bounding box 534, and its size does not exceed the range defined by the soft bounding box 534. The egg at the back of the basket (object 2) moves from the starting position to within the area defined by the soft bounding box 536, and its size does not exceed the range defined by the soft bounding box 536. The egg on the left of the basket (object 3) moves from the starting position to within the area defined by the soft bounding box 538, and its size does not exceed the range defined by the soft bounding box 538.
[0050] In this manner, users can use soft bounding boxes to expand constraints on the movement of the controlled object, thereby increasing the diversity of the resulting video if the constraints are met. In addition, it can reduce the demands on the user; that is, only a certain range of constraints are required, thereby simplifying operation and improving the user experience.
[0051] In some embodiments, nominal phrases in text can be identified and determined as targets. Then, based on the text, targets, and control information, a video can be generated, where the video content is related to the text.
[0052] Figures 6A-6B show schematic diagrams of Example 600 in which a user enters multiple bounding boxes on a start frame where no image is provided, according to some embodiments of the present disclosure. As shown in Figure 6, in Example 600, the user enters the text description "Four pigs running on the snow," which describes video content, in the text input control 604 of the user interface 602. However, since the user has not selected a reference image, the background of the start frame control 608 is blank. In Example 600, the user marks the start position and size of the four pigs in the start frame control 608 with a black bounding box 630, a white bounding box 632, a gray bounding box 634, and a yellow bounding box 636, respectively. In the end frame control 618, the user marks the position and size of the four pigs in the end frame with a black hard bounding box 640, a white hard bounding box 642, a gray hard bounding box 640, and a yellow hard bounding box 646, respectively.
[0053] When generating a video, the video generation model can identify nominal phrases in these descriptions. Since many of these phrases are abstract nouns rather than specific object names, these nominal phrases can be filtered out, leaving only phrases that represent specific object names. These filtered nominal phrases can then be processed to identify the objects to be controlled and associate these objects with bounding boxes.
[0054] Figure 6B shows the generated video 650, and as shown in the figure, the video generation model generates four pigs based on a text description and at positions and sizes specified by bounding boxes 630, 632, 634, and 636, with the four generated pigs associated with the four bounding boxes. These four pigs each run from their initial position to positions specified by hard bounding boxes 640, 642, 644, and 646, and the sizes of the four pigs correspond to these hard bounding boxes.
[0055] In this manner, when a reference image is not provided, it is possible to implement markers for the target in the text description, thereby allowing the user to generate the desired video even if they cannot provide a reference image. In this way, the prerequisites for users to generate videos using the application are reduced, and therefore more users will use this application to generate videos.
[0056] In some embodiments, a third control information that constrains the position in the target intermediate frame can be obtained, and a video can be generated based on the content information, bounding box, first control information, and third control information.
[0057] Figure 7 shows a schematic diagram of Example 700 in which a user creates an intermediate frame and inputs multiple soft bounding boxes in the intermediate frame, according to some embodiments of the present disclosure. As shown in Figure 7, in Example 700, the user inputs a text description describing video content, "Two puppies are running towards the camera," in the text input control 704 of the user interface 702, and selects a reference image (as shown in the start frame control 708). In the start frame control 708, the user labels the white puppy as target 1 with a black bounding box 730 and the black puppy as target 2 with a white bounding box 732. In the end frame control 718, the user specifies the destination position and size of the white puppy with a black hard bounding box 740 and the destination position and size of the black puppy with a white hard bounding box 742. However, the bounding boxes in the end frame cannot directly participate in the motion process of the white puppy moving from the position of bounding box 730 to the position of bounding box 740. Therefore, in Example 700, the user inserts an intermediate frame between the start frame and the end frame, and accordingly, the user interface 702 displays an intermediate frame control 728 and a set of controls associated with the intermediate frame control 728 (e.g., a color display control for each object, a bounding box type selection control, etc.), the functions of which are the same as those of the set of controls associated with the end frame control 718.
[0058] As shown in Figure 7, in Example 700, the user constrains the position and size of the white puppy's destination using a black soft bounding box 750 and a white soft bounding box 752 in the intermediate frame control 728. Thus, in the generated video, in the process of the white puppy running from the position of bounding box 730 to the position specified by the hard bounding box 740, it first passes through a certain position in the region specified by the soft bounding box 750 before going to the position specified by the hard bounding box 740. In the process of the black puppy running from the position of bounding box 732 to the position specified by the hard bounding box 742, it first passes through a certain position in the region specified by the soft bounding box 752. Furthermore, in the user interface 702, the user can insert more intermediate frames by interacting with image frame insertion controls 760 and 762, thereby enabling more precise control over the movement processes of the two puppies.
[0059] In this way, by inserting intermediate frames between the start and end frames, the motion process of the controlled object can be controlled with greater precision. Compared to the motion trajectory, the method of inserting intermediate frames allows for even greater control over the size of the object in the motion process, thereby making the motion control function of the application more complete.
[0060] To achieve bounding box-guided video generation, a motion control module may be inserted into a conventional video generation model. The motion control module processes the bounding box as a control token and, using a self-attention layer, can fuse the control token with a visual token for generating image frames, thereby generating a fused visual token that contains motion control information provided by the bounding box.
[0061] One exemplary architecture includes a spatial self-attention layer, a multilayer perceptron, a motion control module, and a spatial cross-attention layer. The spatial self-attention layer and spatial cross-attention layer may also be modules in a video spread model based on, for example, a three-dimensional U-net (3D U-Net) architecture. The video spread model can iteratively predict noise vectors in a noisy video input, thereby gradually transforming pure Gaussian noise into high-quality video frames. A 3D U-Net consists of alternating convolutional blocks and attention blocks. Each block includes two components: a spatial component that processes each image frame as a standalone image, and a temporal component that facilitates information exchange between image frames. In each attention block, the spatial component typically includes a self-attention layer, followed by a cross-attention layer, which is used to regulate video generation based on text presentation. A motion control module is inserted between these two attention layers, allowing the model to manage motion control in video generation.
[0062] This exemplary architecture inserts a motion control module between the spatial self-attention layer and the spatial cross-attention layer of the original video diffusion model. The spatial self-attention layer receives frame-level visual tokens and generates visual tokens based on these frame-level visual tokens. The motion control module receives visual tokens and control tokens as input and outputs a fused visual token, where each control token corresponds to a corresponding object (or bounding box). Since the control tokens contain motion control information provided by the bounding box, the fused visual token also contains motion control information provided by the bounding box. The visual token is then input to the spatial cross-attention layer, which can generate updated frame-level visual tokens based on the visual tokens and text tokens. The video diffusion model can then generate image frames based on the updated frame-level visual tokens. To avoid altering the original structure of the spatial cross-attention layer, the number of visual tokens may be kept the same. In this method, by fixing the parameters of the original video diffusion model (including the spatial self-attention layer and spatial cross-attention layer) during the training phase and adjusting only the parameters of the motion control module, it is possible to avoid retraining due to modifications to the structure of the video diffusion model, thereby saving costs and avoiding a decrease in the accuracy of the original video diffusion model due to retraining.
[0063] In this exemplary architecture, the number of control tokens is determined by the number of bounding boxes simultaneously present in a single image frame supported by the video generation model, and there is a one-to-one correspondence between control tokens and bounding boxes. For example, if the video generation model supports an image frame containing only one bounding box for a single object, the number of control tokens is 1; if the video generation model supports an image frame containing five bounding boxes for five objects simultaneously, the number of control tokens is 5. If the video generation model supports providing five bounding boxes simultaneously in a single image frame, but the video to be generated only needs to control the movement of two objects (i.e., only two bounding boxes are provided), the three remaining control tokens can be filled with specific learnable tokens. In this exemplary architecture, text tokens are not required; that is, if the user does not provide a text description for the video to be generated, the remaining text tokens can be filled with learnable tokens.
[0064] To generate control tokens, the coordinates of the bounding box, a unique target marker to label the bounding box, and the bounding box type may be determined. Then, control tokens are generated based on the coordinates, target marker, and bounding box type. For example, the target marker may be represented in the color RGB space, where each target corresponds to a bounding box with a unique color, and thus the target marker is a vector with 3D RGB values standardized between 0 and 1. The coordinates, target marker, and bounding box type are concatenated into a single vector, and the corresponding embedding is generated by a Fourier embedding operation. This embedding is then input into a multilayer perceptron, which generates control tokens. By generating target markers using RGB values, during the training phase, the corresponding bounding box can be generated in the image frame based on the target marker, thereby facilitating the alignment of the generated bounding box with the true bounding box of the ground and improving the effectiveness of model training.
[0065] It should be understood that, while this exemplary architecture demonstrates generating control tokens based on bounding box coordinates, target identifiers, and bounding box type, in some embodiments, the target identifiers and bounding box type are not mandatory. For example, in some embodiments, if only one specific type of bounding box (e.g., hard bounding box) is supported, control tokens can be generated based solely on coordinates. In some embodiments, if only multiple specific types of bounding boxes are supported, control tokens can be generated based solely on coordinates and target identifiers.
[0066] In this manner, the motion control module can provide precise motion control information to the original video diffusion model, thereby improving the effect of the generated image frames and allowing the subject to move in the manner desired by the user. Furthermore, since the inserted motion control module does not alter the structure and parameters of the original video diffusion model, this exemplary architecture can reuse the capabilities of the trained video diffusion model, thereby ensuring the screen quality of the generated video and improving motion control of the subject in the video.
[0067] Figure 8 shows a block diagram of an apparatus 800 for generating a video according to some embodiments of the present disclosure. As shown in Figure 8, the apparatus 800 includes a content information acquisition module 802 configured to acquire content information, which includes at least one of text or images, relating to the content of the video to be generated. The apparatus 800 further includes a location information acquisition module 804 configured to acquire location information indicating the position of a subject in the starting frame of the video, and a control information acquisition module 806 configured to acquire control information that constrains the position of the subject in the ending frame. The apparatus 800 further includes a video generation module 808 configured to generate the video based on the content information, location information, and control information.
[0068] To make it clear, the apparatus 800 of this disclosure can be used to realize at least one of the many advantages that can be achieved by the methods or processes described above. For example, the apparatus 800 allows a user to precisely mark an object to be controlled using a bounding box in the starting frame. In addition, the user can accurately represent the position of the movement of the marked object from the starting frame to the ending frame, and the application can receive precise motion control requirements and thereby generate a video that meets the user's expectations.
[0069] Figure 8 shows a block diagram of a device 800 that can implement several embodiments of the present disclosure. The device 800 may be any device or apparatus described in the embodiments of the present disclosure. As shown in Figure 8, the device 800 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 801 which can perform various appropriate operations and processes based on computer program instructions stored in read-only memory (ROM) 802 or computer program instructions loaded from storage unit 808 into random access memory (RAM) 803. The RAM 803 may further store various programs and data necessary for the operation of the device 800. The CPU / GPU 801, ROM 802 and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804. Although not shown in Figure 8, the device 800 may further include a coprocessor.
[0070] Multiple components of the device 800, such as input units 806 including a keyboard and mouse, output units 807 including various types of displays and speakers, storage units 808 including magnetic disks and optical disks, and communication units 809 including network cards, modems, and wireless communication transceivers, are connected to the I / O interface 805. The communication unit 809 allows the device 800 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks.
[0071] Each of the methods or processes described above may be executed by the CPU / GPU 801. For example, in some embodiments, the method may be implemented as a computer software program, which is tangibly contained in a device-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed into the device 800 via the ROM 802 and / or communication unit 809. When the computer program is loaded into the RAM 803 and executed by the CPU / GPU 801, one or more steps or operations in the methods or processes described above can be performed.
[0072] In some embodiments, the methods and processes described above may be implemented as a computer program product. The computer program product may include a computer-readable storage medium on which computer-readable program instructions for performing each aspect of the present disclosure are contained.
[0073] A computer-readable storage medium may be a tangible device capable of holding and storing instructions used by an instruction execution device. A computer-readable storage medium may be, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (a non-exhaustive list) of computer-readable storage media include portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital multifunction optical disks (DVDs), memory sticks, floppy disks, mechanical coding devices such as punch cards or grooved projection structures on which instructions are stored, and any suitable combination of the above. The computer-readable storage medium as used herein is not interpreted as an instantaneous signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated by waveguides or other transmission media (e.g., optical pulses via optical fiber cables), or electrical signals transmitted via wires.
[0074] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer-readable program instructions from the network, transfers these computer-readable program instructions, and stores them in a computer-readable storage medium in each computing / processing device.
[0075] Computer program instructions for performing the operations of the Disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or target code written in any combination of one or more programming languages, including object-oriented programming languages and conventional procedural programming languages. Computer-readable program instructions may be executed entirely on a user computer, partially on a user computer, as a standalone software package, partially on a user computer and partially on a remote computer, or entirely on a remote computer or server. If a remote computer is involved, the remote computer may be connected to the user computer by any type of network, including a local area network (LAN) or wide area network (WAN), or it may be connected to an external computer (e.g., connected via the Internet using an Internet service provider). In some embodiments, the embodiments of the Disclosure are realized by using the state information of the computer-readable program instructions to customize an electronic circuit capable of executing the computer-readable program instructions, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA).
[0076] By providing these computer-readable program instructions to the processing unit of a general-purpose computer, a dedicated computer, or other programmable data processing device, a device can be generated that, when executed by the processing unit of the computer or other programmable data processing device, realizes the functions / operations defined in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may be stored in a computer-readable storage medium, and by operating the computer, programmable data processing device, and / or other device in a specific manner using these instructions, the computer-readable medium storing the instructions will contain a product containing instructions that realize various modes of the functions / operations defined in one or more blocks of a flowchart and / or block diagram.
[0077] These computer program instructions may be loaded into a computer, other programmable data processing device, or other equipment, thereby generating a process implemented by the computer by executing a series of operational steps on the computer, other programmable data processing device, or other equipment, and thereby the instructions executed on the computer, other programmable data processing device, or other equipment implement the functions / operations defined in one or more blocks of a flowchart and / or block diagram.
[0078] The flowcharts and block diagrams in the drawings illustrate the implementable system architectures, functions, and operations of the devices, methods, and computer program products according to several embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, program segment, or portion of an instruction, which includes one or more executable instructions for implementing a defined logical function. In some alternative implementations, the functions marked in the blocks may occur in an order different from the order marked in the drawings. For example, two consecutive blocks may actually be executed almost in parallel, or they may be executed in reverse order, depending on the functions involved. Furthermore, each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a system based on dedicated hardware that performs the defined function or operation, or by a combination of dedicated hardware and computer instructions.
[0079] The following lists some exemplary implementations of this disclosure.
[0080] (Example 1) A method for generating a video, Obtaining content information, which includes at least one of text or an image, relating to the content of the video to be generated, To obtain positional information indicating the position of the target in the starting frame of the aforementioned video, To obtain control information that constrains the position in the end frame of the target, A method comprising generating the video based on the content information, the location information, and the control information.
[0081] (Example 2) The aforementioned object is a first object, the position information is a first bounding box having a first color, the control information is first control information having the first color, and the generation of the video is Obtaining a second bounding box indicating the position of a second object in the video at the start frame, wherein the second bounding box has a second color different from the first color. Obtaining second control information that constrains the position of the second object in the end frame, wherein the second control information has the second color, The method according to Example 1, comprising generating the video based on the content information, the first bounding box, the first control information, the second bounding box, and the second control information.
[0082] (Example 3) The method according to Examples 1-2, wherein the position information is a start boundary box, the control information is an end boundary box in the target end frame, and the start boundary box and the end boundary box are rectangular boxes.
[0083] (Example 4) The video is generated based on the content information, the location information, and the control information. The method according to Examples 1-3, which includes generating the video by moving the object from a position indicated by the start bounding box to a specific position indicated by the end bounding box, in response to the user selecting a first type as the type of the end bounding box, wherein the size of the object in the end frame corresponds to the end bounding box.
[0084] (Example 5) The video is generated based on the content information, the location information, and the control information. The method according to Examples 1-4, which includes generating the video by moving the object from the position indicated by the start bounding box to the position range indicated by the end bounding box in response to the user selecting a second type as the type of the end bounding box, wherein the size of the object in the end frame does not exceed the end bounding box, and the content of the video is related to the content information.
[0085] (Example 6) The video is generated based on the content information, the location information, and the control information. The process involves generating the video by moving the object from its position in the start frame to outside the left or right boundary of the end frame in response to the end boundary box being close to the left or right boundary of the end frame and the width of the end boundary box being less than a threshold width, wherein the size of the object's movement outside the left or right boundary is related to the height of the end boundary box, or The video is generated by moving the object from its position in the start frame to outside the upper or lower boundary of the end frame in response to the end boundary box being close to the upper or lower boundary of the end frame and the height of the end boundary box being less than a threshold height, wherein the magnitude of the object's movement outside the upper or lower boundary is related to the width of the end boundary box. Here, the content of the video is the method described in Examples 1 to 5, relating to the content information.
[0086] (Example 7) The video is generated based on the content information, the location information, and the control information. The process involves generating the video by moving the object from outside the left or right boundary to a position constrained by the control information, in response to the start boundary box being close to the left or right boundary of the start frame and the width of the start boundary box being smaller than a threshold width, wherein the size of the object when it enters the left or right boundary is related to the height of the start boundary box, or The video is generated by moving the object from outside the upper or lower boundary of the start frame to a position constrained by the control information, in response to the start boundary box being close to the upper or lower boundary of the start frame and the height of the start boundary box being less than a threshold height, wherein the size of the object when it enters the upper or lower boundary is related to the width of the end boundary box. Here, the content of the video is the method described in Examples 1 to 6, relating to the content information.
[0087] (Example 8) The method according to Examples 1 to 7, wherein the position information is a bounding box and the control information is a motion trajectory drawn on the start frame.
[0088] (Example 9) The video is generated based on the content information, the location information, and the control information. The method comprising generating the video by moving the object from the position indicated by the bounding box along the motion trajectory, wherein the content of the video is related to the content information, as described in Examples 1 to 8.
[0089] (Example 10) The position information is a bounding box, the content information includes an image for the start frame selected by the user, and the video is generated based on the content information, the position information, and the control information. The image content marked by the bounding box in the starting frame is determined to be the target, The method according to Examples 1 to 9, comprising generating the video based on the image, object, and control information of the start frame, wherein the content of the video is related to the image.
[0090] (Example 11) The content information includes user-entered text describing the content of the video, and the example is: Identifying noun phrases in the aforementioned text, The aforementioned noun phrase is determined to be the aforementioned object, The method according to Examples 1 to 10, further comprising generating the video based on the text, the subject, and the control information, wherein the content of the video is related to the text.
[0091] (Example 12) The control information is the first control information, and the example above is To obtain a third control information that constrains the position in the intermediate frame of the aforementioned target, The method according to Examples 1 to 11, further comprising generating the video based on the content information, the location information, the first control information, and the third control information.
[0092] (Example 13) The method according to claim 1, wherein the object moves gradually from a first position to a second position in the video.
[0093] (Example 14) The method according to claim 1, wherein the video can be moved relative to the object and the camera by changing the camera's viewing angle.
[0094] (Example 15) A device for generating video, A content information acquisition module configured to acquire content information including at least one of text or an image related to the content of the video to be generated, A location information acquisition module configured to acquire location information indicating the position of the target in the starting frame of the video, A control information acquisition module configured to acquire control information that constrains the position in the target's end frame, An apparatus comprising: a video generation module configured to generate the video based on the content information, the location information, and the control information.
[0095] (Example 16) The aforementioned object is a first object, the position information is a first bounding box having a first color, the control information is first control information having the first color, and the generation of the video is The second bounding box acquisition module is configured to acquire a second bounding box indicating the position of a second object in the video at the start frame, wherein the second bounding box has a second color different from the first color. The second control information acquisition module is configured to acquire second control information that constrains the position of the second target in the end frame, wherein the second control information has the second color. The apparatus according to Example 15, wherein a second bounding box-using module is configured to generate the video based on the content information, the first bounding box, the first control information, the second bounding box, and the second control information.
[0096] (Example 17) The apparatus according to Examples 15-16, wherein the position information is a start boundary box, the control information is an end boundary box in the target end frame, and the start boundary box and the end boundary box are rectangular boxes.
[0097] (Example 18) The video is generated based on the content information, the location information, and the control information. The apparatus according to Examples 15-17, wherein the first type video generation module is configured to generate the video in response to a user selecting the first type as the type of the end bounding box, by moving the object from a position indicated by the start bounding box to a specific position indicated by the end bounding box, wherein the size of the object in the end frame corresponds to the end bounding box.
[0098] (Example 19) The video is generated based on the content information, the location information, and the control information. The second type video generation module, in response to a user selecting the second type as the type of the end bounding box, generates the video by moving the object from the position indicated by the start bounding box to the position range indicated by the end bounding box, wherein the size of the object in the end frame does not exceed the end bounding box, and the content of the video is related to the content information, as described in Examples 15-18.
[0099] (Example 20) The video is generated based on the content information, the location information, and the control information. The first boundary determination module is configured to generate the video by moving the object from its position in the start frame to outside the left or right boundary of the end frame in response to the end boundary box being close to the left or right boundary of the end frame and the width of the end boundary box being less than a threshold width, wherein the magnitude of the object's movement outside the left or right boundary is related to the height of the end boundary box, or The second boundary determination module is configured to generate the video by moving the object from its position in the start frame to outside the upper or lower boundary of the end frame in response to the end boundary box being close to the upper or lower boundary of the end frame and the height of the end boundary box being less than a threshold height, wherein the magnitude of the object's movement outside the upper or lower boundary is related to the width of the end boundary box. Here, the video content is the apparatus described in Examples 15-19, relating to the content information.
[0100] (Example 21) The video is generated based on the content information, the location information, and the control information. The third boundary determination module is configured to generate the video by moving the object from outside the left or right boundary to a position constrained by the control information, in response to the start boundary box being close to the left or right boundary of the start frame and the width of the start boundary box being less than a threshold width, wherein the size of the object when it enters the left or right boundary is related to the height of the start boundary box, or The fourth boundary determination module is configured to generate the video by moving the object from outside the upper or lower boundary of the starting frame to a position constrained by the control information in response to the starting boundary box being close to the upper or lower boundary of the starting frame and the height of the starting boundary box being less than a threshold height, wherein the size of the object when it enters the upper or lower boundary is related to the width of the ending boundary box. Here, the video content is the apparatus described in Examples 15-20, relating to the content information.
[0101] (Example 22) The apparatus according to Examples 15-21, wherein the position information is a bounding box and the control information is a motion trajectory drawn on the start frame.
[0102] (Example 23) The video is generated based on the content information, the location information, and the control information. The motion trajectory module is configured to generate the video by moving the object from a position indicated by the bounding box along the motion trajectory, wherein the content of the video is related to the content information, as described in Examples 15-22.
[0103] (Example 24) The content information includes an image for the start frame selected by the user, and the video is generated based on the content information, the position information, and the control information. The target determination module is configured to determine the image content marked by the bounding box in the starting frame as the target, The apparatus according to Examples 15-23, comprising generating the video based on the image, object, and control information of the starting frame, wherein the content of the video is related to the image.
[0104] (Example 25) The content information includes user-entered text describing the content of the video, and the example is: A noun identification module configured to identify noun phrases in the aforementioned text, A noun usage module configured to determine the aforementioned noun phrase as the aforementioned object, The apparatus further includes an object-using module configured to generate the video based on the text, the object, and the control information, wherein the content of the video is related to the text, as described in Examples 15-24.
[0105] (Example 26) The control information is the first control information, and the example above is A third control information acquisition module configured to acquire third control information that constrains the position in the intermediate frame of the target, The method according to Examples 15-25, further comprising a third control information user module configured to generate the video based on the content information, the location information, the first control information, and the third control information.
[0106] (Example 27) The apparatus described in Examples 15-26, wherein the object moves gradually from a first position to a second position in the video.
[0107] (Example 28) The aforementioned video is an apparatus according to Examples 15-27, which allows the object and the camera to move relative to each other by changing the camera's viewing angle.
[0108] (Example 29) Processor and An electronic device including a processor and a memory coupled thereto, wherein the memory has instructions stored therein, and when the instructions are executed by the processor, the electronic device is made to execute a method for generating video, the method is Obtaining content information, which includes at least one of text or an image, relating to the content of the video to be generated, To obtain positional information indicating the position of the target in the starting frame of the aforementioned video, To obtain control information that constrains the position in the end frame of the target, An electronic device that generates the video based on the content information, the location information, and the control information.
[0109] (Example 30) The aforementioned object is a first object, the position information is a first bounding box having a first color, the control information is first control information having the first color, and the generation of the video is Obtaining a second bounding box indicating the position of a second object in the video at the start frame, wherein the second bounding box has a second color different from the first color. Obtaining second control information that constrains the position of the second object in the end frame, wherein the second control information has the second color, The apparatus according to Example 29, comprising generating the video based on the content information, the first bounding box, the first control information, the second bounding box, and the second control information.
[0110] (Example 31) The apparatus according to Examples 29-30, wherein the position information is a start boundary box, the control information is an end boundary box in the target end frame, and the start boundary box and the end boundary box are rectangular boxes.
[0111] (Example 32) The video is generated based on the content information, the location information, and the control information. The apparatus, as described in Examples 29-31, includes generating the video by moving the object from a position indicated by the start bounding box to a specific position indicated by the end bounding box, in response to the user selecting a first type as the type of the end bounding box, wherein the size of the object in the end frame corresponds to the end bounding box.
[0112] (Example 33) The video is generated based on the content information, the location information, and the control information. The apparatus, as described in Examples 29-32, includes generating the video by moving the object from the position indicated by the start bounding box to the position range indicated by the end bounding box, in response to the user selecting a second type as the type of the end bounding box, wherein the size of the object in the end frame does not exceed the end bounding box, and the content of the video is related to the content information.
[0113] (Example 34) The video is generated based on the content information, the location information, and the control information. The process involves generating the video by moving the object from its position in the start frame to outside the left or right boundary of the end frame in response to the end boundary box being close to the left or right boundary of the end frame and the width of the end boundary box being less than a threshold width, wherein the size of the object's movement outside the left or right boundary is related to the height of the end boundary box, or The video is generated by moving the object from its position in the start frame to outside the upper or lower boundary of the end frame in response to the end boundary box being close to the upper or lower boundary of the end frame and the height of the end boundary box being less than a threshold height, wherein the magnitude of the object's movement outside the upper or lower boundary is related to the width of the end boundary box. Here, the video content is the equipment described in Examples 29-33, relating to the content information.
[0114] (Example 35) The video is generated based on the content information, the location information, and the control information. The process involves generating the video by moving the object from outside the left or right boundary to a position constrained by the control information, in response to the start boundary box being close to the left or right boundary of the start frame and the width of the start boundary box being smaller than a threshold width, wherein the size of the object when it enters the left or right boundary is related to the height of the start boundary box, or The video is generated by moving the object from outside the upper or lower boundary of the start frame to a position constrained by the control information, in response to the start boundary box being close to the upper or lower boundary of the start frame and the height of the start boundary box being less than a threshold height, wherein the size of the object when it enters the upper or lower boundary is related to the width of the end boundary box. Here, the video content is the equipment described in Examples 29-34, relating to the content information.
[0115] (Example 36) The apparatus according to Examples 29-35, wherein the position information is a bounding box and the control information is a motion trajectory drawn on the start frame.
[0116] (Example 37) The video is generated based on the content information, the location information, and the control information. The process includes generating the video by moving the object from the position indicated by the bounding box along the motion trajectory, wherein the content of the video is the device described in Examples 29-36, relating to the content information.
[0117] (Example 38) The content information includes an image for the start frame selected by the user, and the video is generated based on the content information, the position information, and the control information. The image content marked by the bounding box in the starting frame is determined to be the target, The device comprises generating the video based on the image, the object, and the control information of the start frame, wherein the content of the video is related to the image, as described in Examples 29-37.
[0118] (Example 39) The content information includes user-entered text describing the content of the video, and the example is: Identifying noun phrases in the aforementioned text, The aforementioned noun phrase is determined to be the aforementioned object, The device further comprises generating the video based on the text, the object, and the control information, wherein the content of the video is related to the text, as described in Examples 29-38.
[0119] (Example 40) The control information is the first control information, and the example above is To obtain a third control information that constrains the position in the intermediate frame of the aforementioned target, The apparatus according to Examples 29-39, further comprising generating the video based on the content information, the location information, the first control information, and the third control information.
[0120] (Example 41) The object is the device described in Examples 29-40, which moves gradually from a first position to a second position in the video.
[0121] (Example 42) The aforementioned video is an apparatus as described in Examples 29-41, which allows the camera to move relative to the subject by changing the camera's viewing angle.
[0122] While the embodiments of this disclosure have been described above, the above descriptions are illustrative, not exhaustive, and are not limited to the embodiments disclosed. Many modifications and changes will be apparent to those skilled in the art without departing from the scope and spirit of the embodiments described. The choice of terms used herein is intended to best interpret the principles, practical applications, or technical improvements in the technology in the market of each embodiment, or to enable those skilled in the art to understand each embodiment disclosed herein.
Claims
1. A method for generating a video, Obtaining content information, including images for creating the video, related to the content of the video to be generated, The acquisition of position information, which is input by the user and marks the position of the controlled object in the starting frame of the image, via a user interface, wherein the controlled object is included in the image and is marked by the image by the user via the user interface, and the acquisition is performed accordingly. The acquisition of control information that constrains the position of the controlled object at the end frame, via a user interface, wherein at least one of a drawn bounding box and a drawn motion trajectory is included in the control information. The process includes generating the video based on the content information, the location information, and the control information, method.
2. The control target is a first target, the position information is a first bounding box having a first color, and the control information is first control information having the first color. To generate the aforementioned video, Obtaining a second bounding box indicating the position of a second object in the video at the start frame, wherein the second bounding box has a second color different from the first color. Obtaining second control information that constrains the position of the second object in the end frame, wherein the second control information has the second color, Based on the content information, the first bounding box, the first control information, the second bounding box, and the second control information, the video is generated. The method according to claim 1, including the method described in claim 1.
3. The method according to claim 1, wherein the position information is a start boundary box, the control information is an end boundary box in the end frame of the controlled object, and the start boundary box and the end boundary box are rectangular boxes.
4. The video is generated based on the content information, the location information, and the control information. The process includes generating the video by moving the controlled object from a position indicated by the start bounding box to a specific position indicated by the end bounding box, in response to the user selecting a first type as the type of the end bounding box. The size of the termination frame of the controlled object corresponds to the termination bounding box, The method according to claim 3.
5. The video is generated based on the content information, the location information, and the control information. The process includes generating the video by moving the controlled object from the position indicated by the start bounding box to the position range indicated by the end bounding box, in response to the user selecting a second type as the type of the end bounding box. The method according to claim 3, wherein the size of the controlled object in the termination frame does not exceed the termination bounding box, and the content of the video is related to the content information.
6. The video is generated based on the content information, the location information, and the control information. The video is generated by moving the controlled object from its position in the start frame to outside the left or right boundary of the end frame in response to the end boundary box being close to the left or right boundary of the end frame and the width of the end boundary box being smaller than a threshold width, wherein the magnitude of the movement of the controlled object outside the left or right boundary is related to the height of the end boundary box, or The video is generated by moving the controlled object from its position in the start frame to outside the upper or lower boundary of the end frame in response to the end boundary box being close to the upper or lower boundary of the end frame and the height of the end boundary box being less than a threshold height, wherein the magnitude of the movement of the controlled object outside the upper or lower boundary is related to the width of the end boundary box. The video content relates to the content information, according to the method of claim 3.
7. The video is generated based on the content information, the location information, and the control information. The video is generated by moving the controlled object from outside the left or right boundary to a position constrained by the control information, in response to the start boundary box being close to the left or right boundary of the start frame and the width of the start boundary box being smaller than a threshold width, wherein the size of the controlled object when it enters the left or right boundary is related to the height of the start boundary box, or The video is generated by moving the controlled object from outside the upper or lower boundary of the starting frame to a position constrained by the control information, in response to the start boundary box being close to the upper or lower boundary of the starting frame and the height of the start boundary box being less than a threshold height, wherein the size of the controlled object when it enters the upper or lower boundary is related to the width of the end boundary box. The video content relates to the content information, according to the method of claim 3.
8. The method according to claim 1, wherein the position information is a bounding box and the control information is a motion trajectory drawn on the start frame.
9. The video is generated based on the content information, the location information, and the control information. The process includes generating the video by moving the controlled object from the position indicated by the bounding box along the motion trajectory, The method according to claim 8, wherein the content of the video relates to the content information.
10. The aforementioned position information is a bounding box, The video is generated based on the content information, the location information, and the control information. The image content labeled by the bounding box in the starting frame is determined to be the control target, This includes generating the video based on the image of the starting frame, the controlled object, and the control information, The content of the aforementioned video is related to the aforementioned image, The method according to claim 1.
11. The content information includes user-entered text describing the content of the video. The aforementioned method, Identifying noun phrases in the aforementioned text, The aforementioned noun phrase is determined to be the object of control, Based on the aforementioned text, the controlled object, and the control information, the video is generated. It further includes, The content of the aforementioned video is related to the aforementioned text, The method according to claim 10.
12. The aforementioned control information is the first control information, The aforementioned method, To acquire third control information that constrains the position of the controlled object in the intermediate frame, Based on the content information, the location information, the first control information, and the third control information, the video is generated. The method according to claim 1, further comprising:
13. The method according to claim 1, wherein the controlled object moves gradually from a first position to a second position in the video.
14. The method according to claim 1, wherein the video can move the controlled object and the camera relative to each other by changing the camera's viewing angle.
15. A device for generating video, A content information acquisition module configured to acquire content information including images for generating the video, relating to the content of the video to be generated, A location information acquisition module that acquires location information, input by the user, which marks the position of a controlled object in the starting frame of the image, via a user interface, wherein the controlled object is included in the image and is marked by the image via the user interface by the user, A control information acquisition module that acquires control information constraining the position of the controlled object at the end frame via a user interface, wherein the control information is configured such that at least one of a drawn bounding box and a drawn motion trajectory is included in the control information. A video generation module configured to generate the video based on the content information, the location information, and the control information, is included. Device.
16. Processor and An electronic device including a memory coupled to the aforementioned processor, The memory has instructions stored therein, and when executed by the processor, the instructions cause the electronic device to perform the method according to any one of claims 1 to 14. electronic equipment.
17. A computer program that, when an executable instruction is performed, causes the method according to any one of claims 1 to 14 to be implemented.
Citation Information
Patent Citations
interactive video generation
JP2018503279A
Interactive video generation
JP2019154045A
Video generation program, video generation device, and video generation method
JP2021033961A
Video processing method, device, electronic apparatus, and storage medium
JP2021193559A