Method, apparatus, electronic device and computer program product for generating video

By employing bounding boxes as control tokens to constrain object positions and sizes, the video generation models achieve precise motion control, addressing the challenge of accurately representing user-defined motion patterns.

JP7789134B2Active Publication Date: 2025-12-19BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024114360
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2024-01-30
Filing Date
2024-07-17
Publication Date
2025-12-19
Estimated Expiration
2044-07-17

AI Technical Summary

Technical Problem

Existing video generation models struggle to accurately understand and implement precise motion patterns described in text, leading to a mismatch between user requirements and generated video content.

Method used

The use of bounding boxes as control tokens to constrain object positions and sizes, combined with visual tokens, allows for precise motion control in video generation, enhancing the alignment between user intent and generated video.

Benefits of technology

Improves the accuracy and consistency of motion effects in generated videos, providing a better user experience by accurately expressing desired object movements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007789134000002
    Figure 0007789134000002
  • Figure 0007789134000003
    Figure 0007789134000003
  • Figure 0007789134000004
    Figure 0007789134000004
Patent Text Reader

Abstract

To provide an improved method of generating a video.SOLUTION: A method according to the present invention has a step of acquiring a visual sense token for use in generation of an image frame in a video. The method further has a step of acquiring a control token for use in restricting positional information of an object in the image frame. The method further has a step of generating an image frame in the video based on the visual sense token and the control token, and an object in the image frame satisfies the positional information. Thus, content of the image frame with use of the control token is restricted, which makes it possible to improve a motion effect of the generated video and improves a matching degree between the generated video and a user request, thereby improving user experience.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates generally to the field of artificial intelligence, and more particularly to methods, apparatus, electronic devices and computer program products for generating video. [Background technology]

[0002] Text-guided video generation is a technique that uses text descriptions to guide the generation of video content. In such video generation tasks, a model receives text descriptions in natural language format, generates image frames corresponding to the text based on these descriptions, and then synthesizes these image frames into a video. One of the key challenges of this task involves making associations between the text descriptions and the video content, understanding the objects, movements, spatiotemporal relationships, etc. in the text descriptions, and converting this information into a sequence of image frames.

[0003] Motion control refers to controlling the motion of objects, scenes, and cameras in a generated video using text descriptions. For example, the text descriptions may contain information about the motion of objects or people, so it is necessary to control the motion of objects or people in the generated video according to the text descriptions. In related technologies, machine learning models are usually used to achieve motion control in video generation tasks. Summary of the Invention

[0004] According to a first aspect of an embodiment of the present disclosure, there is provided a method for generating a video, the method including obtaining visual tokens for generating image frames in the video, the method further including obtaining control tokens for constraining position information of objects in the image frames, and generating image frames in the video based on the visual tokens and the control tokens, wherein the objects in the image frames satisfy the position information.

[0005] According to a second aspect of an embodiment of the present disclosure, there is provided an apparatus for generating a video, the apparatus comprising: a visual token acquisition module configured to acquire visual tokens for generating image frames in the video; a control token acquisition module configured to acquire control tokens for constraining position information of objects in the image frames; and a video image generation module configured to generate image frames in the video based on the visual tokens and the control tokens, wherein the objects in the image frames satisfy the position information.

[0006] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device comprising: one or more processors; and a storage device for storing one or more programs, the one or more programs, when executed by the one or more processors, causing the one or more processors to implement a method for generating a video. The method includes obtaining visual tokens for generating image frames in the video. The method further includes obtaining control tokens for constraining position information of objects in the image frames. The method further includes generating the image frames in the video based on the visual tokens and the control tokens, wherein the objects in the image frames satisfy the position information.

[0007] According to a fourth aspect of an embodiment of the present disclosure, there is provided a computer program product, the computer program product including machine-executable instructions tangibly stored on a non-transitory computer-readable medium, the machine-executable instructions, when executed, causing a machine to perform a method for generating a video, the method including obtaining visual tokens for generating image frames in the video, the method further including obtaining control tokens for constraining position information of objects in the image frames, and generating the image frames in the video based on the visual tokens and the control tokens, wherein the objects in the image frames satisfy the position information.

[0008] This Summary is provided to present a selection of concepts in a simplified form that are further described in the specific embodiments below. This Summary is not intended to identify key features or primary features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. [Brief explanation of the drawings]

[0009] These and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent with reference to the accompanying drawings and the following detailed description, in which the same or similar reference numerals indicate the same or similar elements. [Figure 1] 1 shows a schematic diagram of an exemplary environment in which several embodiments of the present disclosure may be implemented. [Figure 2] 1 illustrates a flowchart of a method for generating a video according to some embodiments of the present disclosure. [Figure 3] 1 illustrates a schematic diagram of an example architecture for generating video according to some embodiments of the present disclosure. [Figure 4] FIG. 1 illustrates a schematic diagram of an example of generating training data from an existing video dataset according to some embodiments of the present disclosure. [Figure 5] 1A-1C show schematic diagrams of an exemplary process of a self-alignment operation according to some embodiments of the present disclosure. [Figure 6] 1 illustrates an example process flow chart of a multi-stage training process according to some embodiments of the present disclosure. [Figure 7] 1 illustrates a schematic diagram of an exemplary process for generating a video by generating a hard bounding box and expanding it into a soft bounding box when a hard bounding box at an end frame is provided, according to some embodiments of the present disclosure. [Figure 8]1 illustrates a schematic diagram of an exemplary process for generating a video by generating a hard bounding box and expanding it to a soft bounding box when a motion trajectory of an object is provided, according to some embodiments of the present disclosure. [Figure 9] 10 illustrates a schematic diagram of an example in which multiple bounding boxes are provided in a start frame and a right-bounding bounding box is provided in an end frame according to some embodiments of the present disclosure. [Figure 10] 1 shows a schematic diagram of an example in which an object bounding box, a motion trajectory, and a bounding box of another object are provided in a start frame according to some embodiments of the present disclosure, and the bounding box of the other object is provided in an end frame. [Figure 11] 1 shows a block diagram of an apparatus for generating video according to some embodiments of the present disclosure. [Figure 12] 1 shows a block diagram of a possible implementation of a device according to several embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0010] It should be understood that all user-related data in this technical solution must be obtained and used only after authorization from the user. This means that if the technical solution needs to use the user's personal information, explicit authorization and authentication from the user is required before obtaining such data, otherwise the relevant data will not be collected or used. Furthermore, when implementing this technical solution, relevant laws and regulations will be strictly observed in the data collection, use, and storage processes, and necessary technologies and measures will be adopted to protect the safety of user data and ensure the safe use of data.

[0011] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the drawings. Although specific embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be realized in various embodiments and should not be construed as being limited to the embodiments described herein, but rather these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are merely illustrative and are not intended to limit the scope of protection of the present disclosure.

[0012] In describing embodiments of the present disclosure, the terms "comprising" and similar terms should be understood as an open-ended inclusion, i.e., "including, but not limited to." The term "based on" should be understood as "based at least in part on." The terms "one embodiment" or "the embodiment" should be understood as "at least one embodiment." Terms such as "first," "second," and the like may refer to different objects or the same object unless expressly stated otherwise. The following may include other explicit and implicit definitions.

[0013] As video generation technology advances, some video generation models may generate videos based on text presentations or key image frames. As an example of a video generation model, the video diffusion model is an extension of the image diffusion model, incorporating the architecture of a U-Net model in the image model and adding a temporal layer to support the generation of multiple image frames. The text-to-video (T2V) diffusion model is typically the basis for various forms of video generation models with constraints. In the T2V diffusion model, image frames are created based on text descriptions, and then video can be generated based on the text descriptions and pre-generated image frames. This method allows the video generation model to use still images as references to focus on the dynamic aspects of video generation, thereby improving the quality of the generated video.

[0014] For convenience of explanation, some embodiments in this specification will be described using a video diffusion model with a U-Net architecture as an example, but it should be understood that this is not intended to limit the specific architecture of the video generation model. The technical solution of the present disclosure can be applied to a video generation model that arbitrarily generates visual tokens and generates image frames based on the visual tokens.

[0015] In some usage scenarios, a user may wish to provide information about the movement patterns of objects in the generated video by inputting a text description. For example, a user may provide a reference image of a building and input the text description, "Tilt the camera upward to expose the top of the building." In this case, the user may wish to have the camera gradually raise its lens from a ground-level perspective in the generated video, so that the top of the building is captured in the final generated video. However, while related technologies can generate high-quality video with gentle camera movement based on the reference image and the text description provided by the user, the model cannot fully understand the user's requirements for the movement patterns of objects in the video, and therefore cannot accurately expose the top of the building in the generated video.

[0016] Furthermore, in some usage scenarios, when a user's requirements for a specific motion pattern are precise, it is difficult to accurately describe them in words. For example, a user may want to capture in a video two puppies running toward the camera, one of which, a white puppy, approaches the camera and runs to the center of the screen, taking up one-third of the screen, while the other, a black puppy, also approaches the camera but runs toward a toy near the camera, gradually moving away from the center of the screen and eventually disappearing to the right side of the screen. For an average user, it is very difficult to accurately describe such motion requirements, and the desired video cannot be generated.

[0017] To this end, the embodiments of the present disclosure provide a technical solution for generating a video, in which a user provides a bounding box to constrain the position and size of an object in the video, and a video generation model can obtain a set of visual tokens and control tokens for generating an image frame based on the provided bounding box, and generate an image frame based on the visual tokens and control tokens, in which the object in the video is displayed within the bounding box in the generated image frame.

[0018] In this way, the video generation model can use the bounding box control tokens to understand the location and size of the target object's destination, and can use the control tokens to constrain the content of the generated image frames. In this way, the user can easily and accurately express the desired object movement form. In addition, compared to providing only a text description or reference images, this approach can improve the match between the generated video and the user's requirements, thereby improving the user experience.

[0019] FIG. 1 illustrates a schematic diagram of an exemplary environment 100 in which several embodiments of the present disclosure can be implemented. As shown in FIG. 1 , the environment 100 includes a video generation model 102, also referred to herein as a base model. The video generation model 102 may be a video diffusion model that generates video based on text descriptions or key image frames, as described above, or may be a machine learning model for generating video using other neural network techniques. The environment 100 also includes a motion control module 104 that combines with the video generation model 102 as a plug-in to enhance the motion control capabilities of the video generation model 102. As shown in FIG. 1 , the video generation model 102 can generate visual tokens 112-1, 112-2, ..., 112-N (collectively referred to as visual tokens 112). The visual tokens 112 are a set of vectors generated based on information, such as text descriptions or reference images, that contain information about the image to be generated.

[0020] As shown in FIG. 1 , environment 100 includes bounding boxes 106 and 108 for constraining the position and size of an object in a video in an image frame to be generated. In embodiments of the present disclosure, the term “object” may refer to an independent object (e.g., a puppy), a part of an independent object (e.g., a person's hand), or a combination of multiple objects (e.g., a person riding a horse). Furthermore, the term “motion” may refer to the motion of an object relative to a camera (or lens) or the motion of a camera relative to an object. For example, if bounding box 106 or 108 indicates the position and size of an object after it has moved, an object capable of autonomous motion (e.g., a puppy, a car, etc.) can move to a specified position and appear at a specified size, while an object that cannot move autonomously (e.g., a rock, a building, etc.) can appear on the screen at a specified position and size by moving the camera. In some embodiments, bounding boxes such as bounding boxes 106 and 108 are rectangular boxes, and two types of bounding boxes can be implemented: hard bounding boxes and soft bounding boxes. A hard bounding box is used to specify a specific position and a specific size of an object, and in the generated image frame, the object is generated at the coordinates specified by the hard bounding box (e.g., the center coordinates of the bounding box), and the size of the object corresponds to the size of the hard bounding box. A soft bounding box is used to specify a range of positions and a range of sizes of an object, and in the generated image frame, the object is generated within the range defined by the soft bounding box, and the size of the object does not exceed the range.

[0021] In environment 100, bounding box 106 is treated as control token 116, and bounding box 108 is treated as control token 118, so that control tokens 116 and 118 contain motion control information for the objects corresponding to bounding boxes 106 and 108, respectively. Motion control module 104 can then generate new visual tokens 114-1, 114-2, ..., 114-N (collectively referred to as visual tokens 114) based on visual token 112, control token 116, and control token 118. In this manner, visual token 114 may contain the motion control information provided by bounding boxes 106 and 108. Video generation model 102 can then generate image frames 120 based on visual tokens 114. In image frame 120, the two puppies move from their positions in the previous image frame 122 to positions specified by bounding boxes 106 and 108, with the size of the white puppy corresponding to bounding box 106 and the size of the black puppy corresponding to bounding box 108. Multiple image frames, such as image frames 120 and 122, can then form video 124.

[0022] In this way, the motion control module 104 can utilize the control tokens 116 and 118 of the bounding boxes 106 and 108 to provide motion control information to the video generation model 102, thereby improving the motion effects of the generated video, increasing the match between the video and the user's requirements, and improving the user experience. 2 illustrates a flowchart of a method 200 for generating a video according to some embodiments of the present disclosure. As shown in FIG. 2, at block 202, the method 200 obtains visual tokens for generating image frames in the video. For example, in the environment 100 illustrated in FIG. 1, the motion control module 104 may obtain visual tokens 112 generated by the video generation model 102 based on information such as a text description or a reference image, where the visual tokens 112 contain visual information for the image frames to be generated and are used to generate the image frames 120.

[0023] In block 204, the method 200 obtains control tokens for constraining the position information of the object in the image frame. The position information may be position-related information such as a bounding box, a contour, a coordinate value, a coordinate range, etc. For example, in the environment 100 shown in FIG. 1 , the motion control module 104 obtains the control tokens 116 and 118. The control tokens 116 and 118 are generated based on the bounding boxes 106 and 108, and therefore contain the motion control information of the bounding boxes 106 and 108. The bounding boxes 106 and 108 function to constrain the position and size of the object in the generated image frame, thereby controlling the object motion.

[0024] In block 206, method 200 generates an image frame in the video based on the visual tokens and the control tokens, where the objects in the image frame satisfy the position information. For example, in the environment 100 shown in FIG. 1 , motion control module 104 generates visual token 114 based on visual token 112, control token 116, and control token 118, and video generation model 102 generates image frame 120 based on visual token 114. In image frame 120, two puppies appear within the range defined by bounding boxes 106 and 108. In different embodiments, the positions of the two puppies may exactly correspond to the positions of bounding boxes 106 and 108, or the positions of the two puppies may be located within the range defined by bounding boxes 106 and 108.

[0025] In this way, method 200 can use the control tokens generated based on the bounding box to help the user understand the desired position of the target object, and then use the control tokens to constrain the content of the generated image frames. In this way, the user can easily and accurately express the desired object movement form. Compared with providing only a text description or reference image, this method can improve the movement effect of the generated video, increase the consistency between the generated video and the user's requirements, and improve the user experience.

[0026] In some embodiments, the position information is a bounding box, and to obtain a control token for the bounding box, coordinates of the bounding box in the image frame may be determined and a control token may be generated based on the coordinates. In some embodiments, using multiple bounding boxes to constrain the movement patterns of multiple objects is supported. In these embodiments, an object identifier for the bounding box may be generated based on the color of the bounding box, and a control token may be generated based on the coordinates and the object identifier. In some embodiments, hard bounding boxes (also referred to herein as first type bounding boxes) and soft bounding boxes (also referred to herein as second type bounding boxes) may be supported simultaneously. In these embodiments, a type of bounding box may be determined, and the type of bounding box may include a hard bounding box that constrains a specific position and a specific size of the object to be generated and a soft bounding box that constrains a range of positions and a range of sizes of the object to be generated, and a control token may be generated based on the coordinates, the object identifier, and the type.

[0027] In some embodiments, in response to the bounding box type being a hard bounding box, the center position of the object coincides with the center position of the bounding box, and the size of the object corresponds to the size of the bounding box. In some embodiments, in response to the bounding box type being a soft bounding box, the center position of the object is located within the bounding box, and the object has a size that does not exceed the bounding box. In some embodiments, multiple embeddings may be generated based on the coordinates, object identifiers, and types, and control tokens may be generated based on the multiple embeddings using a multilayer perceptron. In some embodiments, a second set of visual tokens may be generated based on the first set of visual tokens and the control tokens, and the number of visual tokens in the first set of visual tokens is the same as the number of visual tokens in the second set of visual tokens.

[0028] FIG. 3 illustrates a schematic diagram of an exemplary architecture 300 for generating videos according to some embodiments of the present disclosure. As shown in FIG. 3, the architecture 300 includes a spatial self-attention layer 302, a multi-layer perceptron 308, a motion control module 318, and a cross-spatial attention layer 322. The spatial self-attention layer 302 and the cross-spatial attention layer 322 may be modules of a video diffusion model (e.g., the video generation model 102 in FIG. 1 may be a video diffusion model) based on a three-dimensional (3D) U-Net architecture. The video diffusion model can iteratively predict noise vectors in a noisy video input to gradually transform pure Gaussian noise into high-quality video frames. The 3D U-Net consists of alternating convolutional and attention blocks. Each block includes two components: a spatial component that processes each image frame as an individual image and a temporal component that facilitates information exchange between image frames. In each attention block, the spatial component typically includes a self-attention layer followed by a mutual attention layer, which adjusts video generation based on text presentation. A motion control module is inserted between these two attention layers, allowing the modes to manage motion control during video generation.

[0029] As shown in Figure 3, the architecture 300 inserts a motion control module 318 between the spatial self-attention layer 302 and the spatial cross-attention layer 322 of the original video diffusion model. The spatial self-attention layer 302 receives frame-level visual tokens 304 and generates visual tokens 306-1, 306-2, ..., 306-N (collectively referred to as visual tokens 306) based on the frame-level visual tokens 304. The motion control module 318 receives as input the visual tokens 306 and control tokens 316-1, 316-2, ..., 316-N (collectively referred to as control tokens 316) and outputs visual tokens 320-1, 320-2, ..., 320-N (collectively referred to as visual tokens 320), where each control token in the control tokens 316 corresponds to a corresponding object (or bounding box). Because the control tokens 316 include the motion control information provided by the bounding boxes, the newly generated visual tokens 320 also include the motion control information provided by the bounding boxes. The visual tokens 320 are then input to the cross-spatial attention layer 322, which may generate frame-level visual tokens 326 based on the visual tokens 320 and text tokens 324-1, 324-2, ..., 324-N (collectively referred to as text tokens 324). The video diffusion model can then generate image frames based on the frame-level visual tokens 326. To keep the original structure of the cross-spatial attention layer 322 unchanged, the number of visual tokens 306 can be kept the same as the visual tokens 320. In this way, the parameters of the original video diffusion model (including the spatial self-attention layer 302 and the cross-spatial attention layer 322) are fixed during the training phase, and only the parameters of the motion control module 318 are adjusted, thereby avoiding retraining due to changing the structure of the video diffusion model, thereby saving costs and avoiding a loss in accuracy of the original video diffusion model due to retraining.

[0030] In the architecture 300, if the frame-level visual tokens 304 of the image frame to be generated are denoted by v, the sequence of text tokens 324 are denoted by htext, and the sequence of control tokens 316 are denoted by hbox, the enhanced spatial attention block can be expressed as the following equations (1)-(3): v = v + SelfAttn(v) (1) v = v + TS(SelfAttn([v,hbox])) (2) v = v + CrossAttn(v,htext) (3) where TS(·) represents a token selection operation that takes visual tokens into particular consideration, SelfAttn represents the spatial self-attention layer 302, and CrossAttn represents the spatial cross-attention layer 322.

[0031] In the architecture 300, the number of control tokens 316 depends on the number of bounding boxes that the video generation model can support simultaneously in an image frame, and the control tokens 316 correspond one-to-one to the bounding boxes. For example, if the video generation model only supports including a bounding box for one object in an image frame, the number of control tokens 316 is 1. If the video generation model supports including five bounding boxes for five objects simultaneously in an image frame, the number of control tokens 316 is 5. If the video generation model supports providing five bounding boxes simultaneously in an image frame but only needs to control the movement of two objects in the video to be generated (i.e., provide only two bounding boxes), the three empty control tokens can be filled with specific learnable tokens. In the architecture 300, the text tokens 324 are not required; in other words, if the user does not provide a text description of the video to be generated, the empty text tokens can be filled with learnable tokens.

[0032] 3, to generate a control token 316, the coordinates 310 of a bounding box 328, a unique object identifier 312 for identifying the bounding box 328 (or the object corresponding to the bounding box 328), and a bounding box type 314 can be specified. Each control token 316 can be defined by the following equation (4): tb=MLP(Fourier([bloc, bid, bflag])) (4) where bloc represents a four-dimensional vector (i.e., coordinate 310) including the top-left and bottom-right coordinates of the bounding box, and is normalized between 0 and 1. bid represents an object identifier 312, which is used to identify and link bounding boxes between each image frame. bflag represents a bounding box type 314, e.g., 1 represents a hard bounding box and 0 represents a soft bounding box. Furthermore, Fourier represents a Fourier embedding operation, and MLP represents a multi-layer perceptron. In this way, the multi-layer perceptron can be utilized to improve the performance of the motion control module 318 and enhance the motion effect of the generated image frames, since the control tokens contain higher-level and more abstract semantic features.

[0033] In some embodiments, bid may be expressed in a color RGB space, where each object corresponds to a bounding box with a unique color, and bid is a vector with three-dimensional RGB values ​​normalized to values ​​between 0 and 1. bloc, bid, and bflag are concatenated into a vector to generate a corresponding embedding through a Fourier embedding operation. The embedding is then input to the multi-layer perceptron 308 to generate control tokens 316. By generating object identifiers using RGB values, corresponding bounding boxes can be generated in image frames based on the object identifiers during the training phase, which can facilitate alignment between the generated bounding boxes and ground truth bounding boxes and improve the effectiveness of model training.

[0034] When encoding bloc, bid, and bflag using Fourier embedding, it is possible to ensure that all of the dimensions of the input are scaled between 0 and 1. Any given input x within this range has its Fourier embedding defined by equation (5) below:

[0035]

number

[0036] In some embodiments, the Fourier embeddings of each input can be combined to generate an overall embedding with dimension 128. These embeddings can then be processed by a multi-layer perceptron with three hidden layers, each with dimension 512. The output control tokens can then be scaled to match the dimension of the visual tokens (i.e., 1024).

[0037] Although architecture 300 shows generating control token 316 based on coordinates 310, object identifier 312, and bounding box type 314 of bounding box 328, it should be understood that in some embodiments, object identifier 312 and bounding box type 314 are not required. For example, in some embodiments, if only one particular type of bounding box (e.g., a hard bounding box) is supported, control token 316 may be generated based only on coordinates 310. In some embodiments, if only multiple particular types of bounding boxes are supported, control token 316 may be generated based only on coordinates 310 and object identifier 312.

[0038] In this way, the motion control module 318 can provide accurate motion control information for the original video diffusion model to improve the effect of the generated image frame, allowing the object to move according to the user's desired motion pattern. Furthermore, because the inserted motion control module 318 does not change the structure and parameters of the original video diffusion model, the architecture 300 can reuse the function of the trained video diffusion model, ensuring the visual quality of the generated video and improving the motion control of the object in the video.

[0039] In the training phase, to obtain a training dataset, suitable training data can be obtained from an existing publicly accessible video dataset. For example, each video in the existing video dataset is evaluated and the embeddings of its start and end frames are compared. If the cosine similarity between the embeddings of the start and end frames is lower than a predetermined threshold, it means that the video displays more significant object motion or camera motion, and the data can be put into the dataset to form a selected dataset.

[0040] For a video in a selected dataset, the starting frame of the video can be taken and a description of the content of the starting frame can be generated using an existing model. Noun phrases (e.g., young man, white shirt, etc.) can then be extracted from these descriptions as object representations. These object representations can then be used to identify rectangular bounding boxes that surround the object in the starting frame. These bounding boxes can then be tracked and propagated across all image frames of the video to obtain a number of objects enclosed by the bounding boxes.

[0041] In the training process, a video may be randomly cropped according to a particular aspect, and all of the bounding boxes may be projected into the cropping region. If a bounding box lies entirely outside the cropping region, it may be projected as a line segment (or a bounding box approximating a line segment) along the boundary of the cropping region. FIG. 4 illustrates a schematic diagram of an example 400 of generating training data from an existing video dataset according to some embodiments of the present disclosure. As shown in FIG. 4, the example 400 includes an image 402 containing identified bounding boxes 404, 406, and 408. Because the dimensions of the image 402 are larger than those required for the video generation model, the image 402 is cropped to the dimensions required for the video generation model to obtain the cropping region 412. As shown in FIG. 4, the bounding box 404 lies entirely within the cropping region 412, so no additional operations are required. Because a portion of bounding box 406 lies outside crop region 412, bounding box 406 can be cropped to leave only the portion of bounding box 406 that lies within crop region 412, i.e., bounding box 416. Furthermore, because bounding box 408 lies entirely outside crop region 412, it can be projected onto bounding box 418 located on the boundary of crop region 412, where bounding box 418 can be considered a line segment or a very small rectangular box with a width that is related to the height of bounding box 418. In the training dataset, bounding box 418 can represent the entry of an object from outside the image frame or the movement of an object from within the image frame to outside the image frame.

[0042] In this way, training samples that can be used for training the motion control module of the present disclosure can be generated from the existing training dataset, which can solve the problem of lack of training data corresponding to the video generation method provided by the embodiment of the present disclosure. The training data generated in this way has good diversity and can improve the training effect.

[0043] In some embodiments, objects in a video can be annotated in three steps. In step 1, dynamic video clips can be selected by comparing the start and end frames of each 4-second video clip in the dataset. In some embodiments, these image frames are processed to calculate the cosine similarity of their feature embeddings in an average pooling layer. Video clips with a similarity score below 0.65 are retained for further processing. In step 2, for each selected video clip, a three-sentence description of the video content is created and noun phrases in these descriptions are identified. Because most of these phrases are abstract nouns rather than concrete object names, these noun phrases are filtered, leaving only phrases that represent concrete object names. These filtered noun phrases are then processed to identify initial bounding boxes in the start frame of the video clip. These bounding boxes are then tracked in subsequent frames. Several bounding boxes may be provided for each detected object, with each image frame in the video clip having one bounding box. Then, objects whose bounding boxes cannot be detected in a particular image frame or whose detection confidence is lower than a threshold are filtered out, and the successfully tracked objects can be used as ground truth data for training.

[0044] During training, in some embodiments, the motion control module can be trained by fixing parameters of the base model and adjusting parameters of the motion control module. In some embodiments, the motion control module can be trained by applying a self-alignment operation, which includes generating, based on a target bounding box in the training dataset, an identification image frame including an identification bounding box that identifies an object bounded by the target bounding box, and aligning the identification bounding box with the target bounding box to train the motion control module. In some embodiments, the motion control module can be trained by identifying a loss between the identification bounding box and the target bounding box, and the loss satisfying a predetermined condition.

[0045] FIG. 5 shows a schematic diagram of an example process 500 of a self-alignment operation according to some embodiments of the present disclosure. Process 500 generates bounding boxes of different colors for each coded object in each image frame and trains a model to specify the color in the object's control token. Such a method separates the problem of matching between bounding boxes and objects and maintaining temporal consistency across multiple image frames into two more manageable tasks: generating bounding boxes with the correct color for each object, and aligning these blocks with bounding boxes that provide motion control information in each image frame. In this way, it can be ensured that bounding boxes of the same color always enclose the same object in different image frames. For hard bounding boxes, the model only needs to generate bounding boxes at specified coordinates; for soft bounding boxes, the model may generate bounding boxes within a specified region. The self-aligned bounding boxes can be used as an intermediate representation, and the model can guide the generation of these self-aligned bounding boxes according to the constraints provided by the target bounding box, thereby guiding the generation of the visual object. After the training phase that performs the self-alignment operation is complete, the same dataset can be continued to train the model further to remove bounding boxes in the generated image frames.

[0046] 5, when generating image frames 502, 512, 522, 532, and 542, process 500 may predict bounding boxes 504, 506, 514, 516, 524, 526, 534, and 536 surrounding each of the controlled objects and render these bounding boxes as part of the corresponding image frames. Here, bounding boxes identifying the same object are represented by the same color, and bounding boxes identifying different objects are represented by different colors. For example, bounding boxes 504, 514, 524, and 534 identifying the same object are all represented by white, bounding boxes 506, 516, 526, and 536 identifying the same object are all represented by black, and bounding boxes 504 and 506 identifying different objects are represented by different colors.

[0047] In this way, the self-alignment operation can effectively map bounding boxes to objects while maintaining temporal consistency across multiple frames. Furthermore, the model can quickly learn to stop generating visible bounding boxes, while still retaining the ability to align these bounding boxes. In this way, the self-alignment operation can help the model build an appropriate internal representation.

[0048] In some embodiments, a first training data set including hard bounding boxes may be obtained, and a self-alignment operation may be applied to train the motion control module based on the first training data set. In some embodiments, a second training data set may be generated by converting some of the bounding boxes of the first training data set into soft bounding boxes, and a self-alignment operation may be applied to train the motion control module based on the second training data set. In some embodiments, a motion control module may be trained based on the second training data set by utilizing a loss relative to a base model without applying a self-alignment operation.

[0049] 6 illustrates a flowchart of an example process 600 of a multi-stage training process according to some embodiments of the present disclosure. As shown in FIG. 6, the first stage is described in block 602, where process 600 can train a model using all of the provided ground truth bounding boxes as hard bounding boxes. Because the motion control of hard bounding boxes is easier to learn than that of soft bounding boxes, this stage can be an initial stage to establish the model's initial understanding of coordinates and object identifiers.

[0050] In the second stage described in block 604, process 600 replaces a portion of the hard bounding box with a soft bounding box to train the model. For example, 80% of the hard bounding box is replaced with a soft bounding box. Process 600 randomly extends the hard bounding box along four directions (up, down, left, and right) respectively (without going beyond the boundaries of the image frame), and the extended bounding box becomes the soft bounding box. In the first and second stages of blocks 602 and 604, training can be performed by applying a self-alignment operation.

[0051] In the third stage, described at block 606, process 600 trains a model using the training data from the second stage, but does not perform a self-alignment operation. In this way, the first and second stages effectively equip the model with the ability to handle hard and soft bounding boxes, and the third stage generates image frames that do not contain self-aligned bounding boxes. This improves the training efficiency of the model and also allows the self-alignment operation to be utilized to improve the training effectiveness of the model.

[0052] FIG. 7 illustrates a schematic diagram of an example process 700 for generating a video by generating a hard bounding box and expanding it into a soft bounding box when a hard bounding box in an end frame is provided, according to some embodiments of the present disclosure. During the inference stage, a user may identify a bounding box in only a few image frames (e.g., a start frame and an end frame). As shown in FIG. 7 , the example process 700 identifies a bounding box 704 in a start frame 702, identifies a bounding box 734 in an end frame 732, and indicates that the object identified by the bounding box 704 should move from the location of the bounding box 704 to the position specified by the bounding box 734, and that the size of the object in the end frame 732 should correspond to the size of the bounding box 734. The model may insert soft bounding boxes in intermediate frames 712 and 722 to provide more stable motion control. In some embodiments, hard bounding boxes 714 and 724 can be generated by applying linear interpolation between user-specified bounding boxes 704 and 734 to intermediate frames 712 and 722. Process 700 then generates soft bounding boxes 716 and 726 by appropriately expanding hard bounding boxes 714 and 724. In this manner, process 700 generates intermediate frame 712 with soft bounding box 716 and intermediate frame 722 with soft bounding box 726.

[0053] In this way, we can ensure that objects essentially follow expected trajectories during the motion process, while the soft bounding boxes allow for variations in the model, improving the diversity of the generated videos.

[0054] 8 illustrates a schematic diagram of an exemplary process for generating a video by generating hard bounding boxes and expanding them into soft bounding boxes when a motion trajectory of an object is provided, according to some embodiments of the present disclosure. As shown in FIG. 8 , a user provides a bounding box 804 and a motion trajectory 808 in a starting frame 802. The process 800 interpolates in subsequent image frames 812, 822, and 832 and generates hard bounding boxes 814, 824, and 834 along the motion trajectory 808. For example, the hard bounding boxes 814, 824, and 834 are generated by generating hard bounding boxes whose centers are located on the motion trajectory. The process 800 then generates soft bounding boxes 816, 826, and 836 by expanding the hard bounding boxes 814, 824, and 834. In this way, the object identified by bounding box 804 moves along motion trajectory 808 within the range enclosed by soft bounding box 836 in image frame 832, and the object in image frames 812 and 822 can be kept within the range enclosed by soft bounding boxes 816 and 826.

[0055] In this way, the user can specify the motion trajectory of the object, allowing for more precise motion control and achieving effects that cannot be expressed by text description.

[0056] 9 illustrates an example 900 in which multiple bounding boxes are provided in a start frame and a bounding box near the right boundary is provided in an end frame, according to some embodiments of the present disclosure. As shown in FIG. 9 , in the example 900, a start frame 902 and a text description 908 (i.e., "A puppy is running toward the camera") are provided, and the start frame 902 includes a white puppy and a black puppy. In the example 900, the start frame 902 includes a black bounding box 904 identifying the white puppy and a white bounding box 906 identifying the black puppy. Furthermore, the end frame 912 includes a black hard bounding box 914 and a white hard bounding box 916, which approaches the right boundary of the end frame 912 and has a very narrow width (e.g., less than a threshold width), indicating that the black puppy needs to jump out of the screen from the right side before the end frame.

[0057] In example 900, the model generates a sequence of image frames 920 in which two puppies gradually run towards the camera, with the white puppy running to a position specified by a hard bounding box 914 in the end frame, and the white puppy having a size corresponding to the hard bounding box 914. Furthermore, the black puppy is seen to jump out of the right boundary in the end frame and have a size corresponding to the hard bounding box 916 in the image frame in which the black puppy jumped out of the right boundary.

[0058] 10 shows a schematic diagram of an example 1000 in which an object bounding box, a motion trajectory, and a bounding box of another object are provided in a start frame and a bounding box of the other object is provided in an end frame according to some embodiments of the present disclosure. As shown in FIG. 10, in the example 1000, a start frame 1002 including a person and a Frisbee and a text description 1010 (i.e., "A person is throwing a Frisbee") are provided. In the example 1000, the start frame 1002 includes a white hard bounding box 1004 identifying the person, a black hard bounding box 1006 identifying the Frisbee, and a motion trajectory 1008 identifying the Frisbee motion trajectory.

[0059] In this example 1000, the model generates an image frame sequence 1020 in which the person is located at a position specified by a hard bounding box 1014 in the end frame, and the size of the person corresponds to the hard bounding box 1014. Furthermore, a Frisbee is thrown by the person in the image frame sequence 1020, flies along a motion trajectory 1008, and eventually flies back to the end position indicated by the motion trajectory 1008. This allows the motion control module to achieve precise motion control according to the bounding boxes, managing the movement of foreground and background objects and correcting the pose of relatively large objects by adjusting their smaller components. Furthermore, in image-based video generation scenarios, a user simply selects an object by drawing a hard bounding box around it. This visually based method is easier to operate than text- or language-based control. Additionally, for intermediate frames lacking a user-provided bounding box, the motion control module algorithmically generates soft bounding boxes to approximate the motion trajectory. These soft bounding boxes may be constructed based on the bounding boxes at the start and end frames specified by the user, or may be constructed based on the bounding box and motion trajectory specified by the user, thereby achieving precise motion control and improving the user experience.

[0060] 11 shows a block diagram of an apparatus 1100 for generating a video according to some embodiments of the present disclosure. As shown in FIG. 11, the apparatus 1100 includes a visual token acquisition module 1102 configured to acquire visual tokens for generating image frames in the video. The apparatus 1100 further includes a control token acquisition module 1104 configured to acquire control tokens for constraining position information of objects in the image frames. The apparatus 1100 also includes a video image generation module 1106 configured to generate image frames in the video based on the visual tokens and the control tokens, where the objects in the image frames satisfy the position information.

[0061] It should be appreciated that the device 1100 of the present disclosure can achieve at least one of many advantages that can be achieved by the above-described method or process. For example, the device 1100 can use the control tokens generated based on the bounding box to understand the user's expected destination location of the target object, and then use the control tokens to constrain the content of the generated image frame. In this way, the user can easily and accurately express the desired object movement form. Compared with providing only a text description or a reference image, this method can improve the movement effect of the generated video, increase the consistency between the generated video and the user's request, and improve the user experience.

[0062] FIG. 12 shows a block diagram of a device 1200 according to several embodiments of the present disclosure. The device 1200 is a device or apparatus described in the embodiments of the present disclosure. As shown in FIG. 12, the device 1200 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 1201, and can perform various appropriate operations and processes based on computer program instructions stored in a read-only memory (ROM) 1202 or loaded from a storage device 1208 into a random access memory (RAM) 1203. The RAM 1203 may also store various programs and data necessary for the operation of the device 1200. The CPU / GPU 1201, the ROM 1202, and the RAM 1203 are connected to one another via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204. Although not shown in FIG. 12, the device 1200 may include a coprocessor.

[0063] Several elements of device 1200 are connected to an I / O interface 1205, which comprises input units 1206 such as a keyboard, a mouse, etc., output units 1207 such as various types of displays, speakers, etc., storage devices 1208 such as a disk, an optical disk, etc., and communication units 1209 such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1209 enables device 1200 to exchange information / data with other devices via computer networks, for example the Internet, and / or various telecommunication networks.

[0064] Each of the methods or processes described above may be executed by the CPU / GPU 1201. For example, in some embodiments, the methods may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage device 1208. In some embodiments, some or all of the computer program may be loaded and / or installed into the device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program is loaded into the RAM 1203 and executed by the CPU / GPU 1201, it may perform one or more steps or operations of the methods or processes described above.

[0065] In some embodiments, the methods and processes described above may be implemented as a computer program product, which may include a computer-readable storage medium containing computer-readable program instructions for carrying out aspects of the present disclosure.

[0066] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions for use by an instruction execution device. Computer-readable storage media include, but are not limited to, electrical, magnetic, optical, electromagnetic, and semiconductor storage devices, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanical encoding devices such as punch cards or grooved projections on which instructions are stored, and any suitable combination of the above. As used herein, a computer-readable storage medium is not to be construed as a transitory signal per se, such as an electric wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., light pulses through a fiber optic cable), or an electrical signal transmitted through an electrical wire.

[0067] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or downloaded to an external computer or external storage device over a network, such as, for example, a local area network (LAN), a wide area network (WAN), and / or a wireless network. The network can include copper transmission cables, fiber optic transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface of each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium of each computing / processing device.

[0068] Computer program instructions for carrying out the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source or object code programmed in one or more programming languages, including object-oriented programming languages ​​and traditional procedural programming languages. The computer-readable program instructions may execute entirely on the user computer, partially on the user computer, as a standalone software package, partially on the user computer, partially on a remote computer, or entirely on a remote computer or server. A remote computer may be connected to the user computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet via an Internet service provider). In some embodiments, aspects of the present disclosure are achieved by utilizing state information from the computer-readable program instructions to personalize an electronic circuit capable of executing computer-readable program instructions, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA).

[0069] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, or other programmable data processing device, and when executed by the processing unit of the computer or other programmable data processing device, the instructions can produce an apparatus that implements the functions / acts specified in one or more boxes in the flowcharts and / or block diagrams. These computer-readable program instructions can also be stored on computer-readable storage media that cause a computer, programmable data processing device, and / or other device to operate in a particular manner, and computer-readable media having instructions stored thereon can include articles of manufacture containing various instructions that implement the functions / acts defined in one or more boxes in the flowcharts and / or block diagrams.

[0070] The computer-readable program instructions are loaded into a computer, other programmable data processing device, or other device, which then executes a series of operational steps on the computer, other programmable data processing device, or other device to produce a computer-implemented process, such that the instructions executed by the computer, other programmable data processing device, or other device implement the functions / acts specified in one or more boxes of the flowcharts and / or block diagrams.

[0071] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams may represent a module, program segment, or part of an instruction, which includes one or more executable instructions for implementing a given logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order from the order noted in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or may be executed in the reverse order depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented in a dedicated hardware-based system that performs a given function or operation, or in a combination of dedicated hardware and computer instructions.

[0072] The following presents some exemplary aspects of the present disclosure.

[0073] Example 1. A method for generating a video, comprising: obtaining visual tokens for generating image frames in the video; obtaining a control token for constraining position information of an object in the image frame; generating the image frames in the video based on the visual tokens and the control tokens; The object in the image frame satisfies the position information.

[0074] Example 2. The position information is a bounding box, Obtaining the control token for constraining the position information of an object in the image frame includes: determining the coordinates of the bounding box in the image frame; generating the control token based on the coordinates.

[0075] Example 3. Generating the control token based on the coordinates comprises: generating an object identifier for the bounding box based on a color of the bounding box; and generating the control token based on the coordinates and the object identifier.

[0076] Example 4. Generating the control token based on the coordinates and the object identifier comprises: determining a type of the bounding box, the type including a first type that constrains a specific position and a specific size of the object to be generated, and a second type that constrains a range of positions and a range of sizes of the object to be generated; and generating the control token based on the coordinates, the object identifier, and the type.

[0077] Example 5. Generating the image frames in the video based on the visual tokens and the control tokens comprises: In response to the type of the bounding box being the first type, a center position of the object coincides with a center position of the bounding box, and a size of the object corresponds to a size of the bounding box; In response to the type of the bounding box being the second type, the center position of the object is within the bounding box and the size of the object does not exceed the bounding box.

[0078] Example 6. Generating the control token based on the coordinates, the object identifier, and the type includes: generating a plurality of embeddings based on the coordinates, the object identifier, and the type; and generating the control tokens based on the plurality of embeddings using a multi-layer perceptron.

[0079] Example 7. The visual tokens are a first set of visual tokens; generating the image frames in the video based on the visual tokens and the control tokens, generating a second set of visual tokens based on the first set of visual tokens and the control tokens; The method of Examples 1-6, wherein the number of visual tokens in the first set of visual tokens and the number of visual tokens in the second set of visual tokens are the same.

[0080] Example 8. The first set of visual tokens and the image frames are generated by a base model, and the second set of visual tokens are generated by a motion control module; The method comprises: The method of any one of Examples 1 to 7, further comprising fixing parameters of the base model and adjusting parameters of the motion control module to trail the motion control module.

[0081] Example 9. Trailing the motion control module by applying a self-alignment operation, the self-alignment operation comprising: generating an identification image frame based on the target bounding box in the training dataset, the identification image frame including an identification bounding box that identifies the object constrained by the target bounding box; and trailing the motion control module by aligning the identification bounding box with the target bounding box.

[0082] Example 10. Trailing the motion control module by aligning the identification bounding box with the target bounding box comprises: determining a loss between the discriminant bounding box and the target bounding box; and trailing the motion control module such that the loss satisfies a predetermined condition.

[0083] Example 11. Fixing the parameters of the base model and adjusting the parameters of the motion control module to trail the motion control module includes: obtaining a first training data set including bounding boxes having the first type; 11. The method of any one of Examples 1 to 10, further comprising: applying the self-alignment operation based on the first training data set to train the motion control module.

[0084] Example 12. Fixing the parameters of the base model and adjusting the parameters of the motion control module to trail the motion control module includes: generating a second training data set by converting a portion of the bounding boxes of the first training data set into bounding boxes having the second type; 12. The method of any one of Examples 1 to 11, further comprising: applying the self-alignment operation based on the second training data set to train the motion control module.

[0085] Example 13. Fixing the parameters of the base model and adjusting the parameters of the motion control module to trail the motion control module includes: 13. The method of any one of Examples 1 to 12, further comprising: training the motion control module based on the second training dataset without using the self-alignment operation and using a loss relative to the base model.

[0086] Example 14. A device for generating video a visual token acquisition module configured to acquire visual tokens for generating image frames in the video; a control token obtaining module configured to obtain a control token for constraining position information of an object in the image frame; a video image generation module arranged to generate the image frames in the video based on the visual tokens and the control tokens; The object in the image frame satisfies the position information.

[0087] Example 15. The position information is a bounding box, Obtaining the control token for constraining the position information of an object in the image frame includes: a coordinate determination module arranged to determine coordinates of the bounding box in the image frame; and a coordinate utilization module configured to generate the control token based on the coordinates.

[0088] Example 16. Generating the control token based on the coordinates comprises: an identifier generation module configured to generate an object identifier for the bounding box based on a color of the bounding box; and an identifier utilization module configured to generate the control token based on the coordinates and the object identifier.

[0089] Example 17. Generating the control token based on the coordinates and the object identifier comprises: a type determination module arranged to determine a type of the bounding box, the type including a first type that constrains a specific position and a specific size of the object to be generated, and a second type that constrains a range of positions and a range of sizes of the object to be generated; and a type utilization module configured to generate the control token based on the coordinates, the object identifier, and the type.

[0090] Example 18. Generating the image frames in the video based on the visual tokens and the control tokens comprises: In response to the type of the bounding box being the first type, a center position of the object coincides with a center position of the bounding box, and a size of the object corresponds to a size of the bounding box; In response to the type of the bounding box being the second type, the center position of the object is within the bounding box and the size of the object does not exceed the bounding box.

[0091] Example 19. Generating the control token based on the coordinates, the object identifier, and the type includes: an embedding generation module configured to generate a plurality of embeddings based on the coordinates, the object identifier, and the type; and an embedding utilization module configured to utilize a multi-layer perceptron to generate the control tokens based on the plurality of embeddings.

[0092] Example 20. The visual tokens are a first set of visual tokens; generating the image frames in the video based on the visual tokens and the control tokens, a visual token generation module configured to generate a second set of visual tokens based on the first set of visual tokens and the control tokens; 20. The apparatus of Examples 14-19, wherein the number of visual tokens in the first set of visual tokens and the number of visual tokens in the second set of visual tokens are the same.

[0093] Example 21. The first set of visual tokens and the image frames are generated by a base model, and the second set of visual tokens are generated by a motion control module; The device comprises: 21. The apparatus of Examples 14-20, further comprising a motion control training module arranged to fix parameters of the base model and adjust parameters of the motion control module to trail the motion control module.

[0094] Example 22. Trailing the motion control module by applying a self-alignment operation; The self-alignment operation includes: an identification frame generation module configured to generate, based on a target bounding box in the training dataset, an identification image frame including an identification bounding box that identifies an object constrained by the target bounding box; and a bounding box alignment module positioned to trail the motion control module by aligning the identification bounding box with the target bounding box.

[0095] Example 23. Trailing the motion control module by aligning the identification bounding box with the target bounding box comprises: a loss determination module arranged to determine a loss between the discriminant bounding box and the target bounding box; and a loss utilization module arranged to trail the motion control module when the loss satisfies a predetermined condition.

[0096] Example 24. Fixing the parameters of the base model and adjusting the parameters of the motion control module to trail the motion control module includes: a first training set acquisition module configured to acquire a first training data set including bounding boxes having the first type; and a first training set utilization module arranged to trail the motion control module by applying the self-alignment operation based on the first training data set.

[0097] Example 25. Fixing the parameters of the base model and adjusting the parameters of the motion control module to trail the motion control module includes: a second training set acquisition module configured to generate a second training data set by converting a portion of the bounding boxes of the first training data set into bounding boxes having the second type; and and a second training set utilization module arranged to trail the motion control module by applying the self-alignment operation based on the second training data set.

[0098] Example 26. Fixing the parameters of the base model and adjusting the parameters of the motion control module to trail the motion control module includes: 26. The apparatus of Examples 14-25, further comprising a third training set utilization module configured to trail the motion control module by not utilizing the self-alignment operation and utilizing a loss relative to the base model based on the second training data set.

[0099] Example 27. An electronic device, a processor; a memory coupled to the processor; The memory may include instructions that, when executed by the processor, cause the electronic device to perform actions, including: obtaining visual tokens for generating image frames in the video; obtaining a control token for constraining position information of an object in the image frame; generating the image frames in the video based on the visual tokens and the control tokens, wherein the objects in the image frames satisfy the position information.

[0100] Example 28. Obtaining the control token for constraining a bounding box of an object in the image frame comprises: determining the coordinates of the bounding box in the image frame; and generating the control token based on the coordinates.

[0101] Example 29. Generating the control token based on the coordinates includes: generating an object identifier for the bounding box based on a color of the bounding box; and generating the control token based on the coordinates and the object identifier.

[0102] Example 30. Generating the control token based on the coordinates and the object identifier comprises: determining a type of the bounding box, the type including a first type that constrains a specific position and a specific size of the object to be generated, and a second type that constrains a range of positions and a range of sizes of the object to be generated; and generating the control token based on the coordinates, the object identifier, and the type.

[0103] Example 31. Generating the image frames in the video based on the visual tokens and the control tokens comprises: In response to the type of the bounding box being the first type, a center position of the object coincides with a center position of the bounding box, and a size of the object corresponds to a size of the bounding box; In response to the type of the bounding box being the second type, the center position of the object is within the bounding box and the size of the object does not exceed the bounding box.

[0104] Example 32. Generating the control token based on the coordinates, the object identifier, and the type includes: generating a plurality of embeddings based on the coordinates, the object identifier, and the type; and generating the control token based on the plurality of embeddings using a multi-layer perceptron.

[0105] Example 33. The visual tokens are a first set of visual tokens, and generating the image frames in the video based on the visual tokens and the control tokens comprises: The device of Examples 27 to 32, further comprising generating a second set of visual tokens based on the first set of visual tokens and the control tokens, wherein the number of visual tokens in the first set of visual tokens is the same as the number of visual tokens in the second set of visual tokens.

[0106] Example 34. The first set of visual tokens and the image frames are generated by a base model, and the second set of visual tokens are generated by a motion control module, and the device: 34. The apparatus of any one of Examples 27 to 33, further comprising fixing parameters of the base model and adjusting parameters of the motion control module to trail the motion control module.

[0107] Example 35. Trailing the motion control module by applying a self-alignment operation, the self-alignment operation comprising: generating an identification image frame based on the target bounding box in the training dataset, the identification image frame including an identification bounding box that identifies the object constrained by the target bounding box; and trailing the motion control module by aligning the identification bounding box with the target bounding box.

[0108] Example 36. Trailing the motion control module by aligning the identification bounding box with the target bounding box includes: determining a loss between the discriminant bounding box and the target bounding box; and trailing the motion control module when the loss satisfies a predetermined condition.

[0109] Example 37. Fixing the parameters of the base model and adjusting the parameters of the motion control module to trail the motion control module includes: obtaining a first training data set including bounding boxes having the first type; 37. The apparatus of Examples 27-36, further comprising: applying the self-alignment operation based on the first training data set to train the motion control module.

[0110] Example 38. Fixing the parameters of the base model and adjusting the parameters of the motion control module to trail the motion control module includes: generating a second training data set by converting a portion of the bounding boxes of the first training data set into bounding boxes having the second type; 38. The apparatus of Examples 27-37, further comprising: applying the self-alignment operation based on the second training data set to train the motion control module.

[0111] Example 39. Fixing the parameters of the base model and adjusting the parameters of the motion control module to trail the motion control module includes: 39. The apparatus of Examples 27-38, further comprising: training the motion control module based on the second training data set without using the self-alignment operation and using a loss relative to the base model.

[0112] While various embodiments of the present disclosure have been described above, the above description is illustrative and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the illustrated embodiments. The terms used in this specification are selected to best explain the principles, applications, or improvements to commercially available technologies of each embodiment, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. 1. A method for generating a video, comprising: obtaining, by an apparatus, visual tokens for generating image frames in the video, the visual tokens indicating visual information of the image frames to be generated; obtaining, by the device, a control token indicating motion control information for the object for constraining position information of the object in the image frame; generating, by the device, the image frames in the video based on the visual tokens and the control tokens; the object in the image frame satisfies the position information; The control token is generated based on a bounding box provided by a user; the visual tokens are a first set of visual tokens; generating, by the device, the image frames in the video based on the visual tokens and the control tokens, generating, by the device, a second set of visual tokens based on the first set of visual tokens and the control tokens; generating, by the device, the image frames in the video based on the second set of visual tokens; A method wherein the number of visual tokens in the first set of visual tokens and the number of visual tokens in the second set of visual tokens are the same.

2. the position information is a bounding box, Obtaining the control token for constraining the position information of an object in the image frame includes: determining the coordinates of the bounding box in the image frame; and generating the control token based on the coordinates.

3. generating the control token based on the coordinates, generating an object identifier for the bounding box based on a color of the bounding box; generating the control token based on the coordinates and the object identifier.

4. generating the control token based on the coordinates and the object identifier, determining a type of the bounding box, the type including a first type that constrains a specific position and a specific size of the object to be generated, and a second type that constrains a range of positions and a range of sizes of the object to be generated; generating the control token based on the coordinates, the object identifier, and the type.

5. generating the image frames in the video based on the visual tokens and the control tokens, In response to the type of the bounding box being the first type, a center position of the object coincides with a center position of the bounding box, and a size of the object corresponds to a size of the bounding box; 5. The method of claim 4, further comprising: in response to the type of the bounding box being the second type, determining that a center position of the object is within the bounding box and that a size of the object does not exceed the bounding box.

6. generating the control token based on the coordinates, the object identifier, and the type, generating a plurality of embeddings based on the coordinates, the object identifier, and the type; and generating the control tokens based on the plurality of embeddings using a multi-layer perceptron.

7. the first set of visual tokens and the image frames are generated by a base model, and the second set of visual tokens are generated by a motion control module; The method comprises: The method of claim 6 , further comprising fixing parameters of the base model and adjusting parameters of the motion control module to trail the motion control module.

8. applying a self-alignment operation to trail the motion control module, the self-alignment operation comprising: generating an identification image frame based on the target bounding box in the training dataset, the identification image frame including an identification bounding box that identifies the object constrained by the target bounding box; and trailing the motion control module by aligning the identified bounding box with the target bounding box.

9. Trailing the motion control module by aligning the identification bounding box with the target bounding box includes: determining a loss between the discriminant bounding box and the target bounding box; and trailing the motion control module with the loss satisfying a predetermined condition.

10. Fixing parameters of the base model and adjusting parameters of the motion control module to trail the motion control module includes: obtaining a first training data set including bounding boxes having the first type; The method of claim 8 , further comprising: trailing the motion control module by applying the self-alignment operation based on the first training data set.

11. Fixing parameters of the base model and adjusting parameters of the motion control module to trail the motion control module includes: generating a second training data set by converting a portion of the bounding boxes of the first training data set into bounding boxes having the second type; The method of claim 10 , further comprising: trailing the motion control module by applying the self-alignment operation based on the second training data set.

12. Fixing parameters of the base model and adjusting parameters of the motion control module to trail the motion control module includes: The method of claim 11 , further comprising: training the motion control module based on the second training data set by not using the self-alignment operation and by using a loss relative to the base model.

13. 1. An apparatus for generating video, comprising: a visual token acquisition module for generating image frames in the video, the visual token acquisition module being configured to acquire visual tokens indicative of visual information of the image frames to be generated; a control token obtaining module configured to obtain a control token indicating motion control information for the object to constrain position information of the object in the image frame; a video image generation module arranged to generate the image frames in the video based on the visual tokens and the control tokens; the object in the image frame satisfies the position information; The control token is generated based on a bounding box provided by a user; the visual tokens are a first set of visual tokens; The video image generation module: generating a second set of visual tokens based on the first set of visual tokens and the control tokens; arranged to generate the image frames in the video based on the second set of visual tokens; The number of visual tokens in the first set of visual tokens and the number of visual tokens in the second set of visual tokens are the same.

14. An electronic device, a processor; a memory coupled to the processor; An electronic device, wherein the memory stores instructions that, when executed by a processor, cause the electronic device to perform a method according to any one of claims 1 to 12.

15. A computer-readable storage medium having machine-executable instructions stored thereon, A computer readable storage medium, the machine executable instructions which, when executed, cause a machine to perform the method of any one of claims 1 to 12.

16. A computer program which, when executed by a computing device, causes the computing device to implement a method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Animation data generation method and device based on neural network, and related product

    CN115222847A

  • Machine learning systems and methods for augmenting images

    US10529137B1

  • Controllable video characters with natural motions extracted from real-world videos

    US11017560B1

  • Generating images of object motion using one or more neural networks

    US20230147641A1