Video Generation With Bounding Boxes for Precise Object Motion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video generation models struggle to accurately understand and implement user-defined motion requirements for objects in generated videos, especially when precise descriptions are difficult to provide, leading to mismatches between user expectations and generated content.

Innovation Solution

A method and apparatus that utilize bounding boxes to constrain object positions and sizes, generating control tokens to guide the video generation process, ensuring objects move as intended by incorporating a motion control module into the video diffusion model architecture.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If text descriptions are used to guide video generation, then the system can understand user requirements, but the accuracy of object motion control deteriorates when precise descriptions are difficult to provide

Engineering Contradiction:
Improvetext description understandingVSAvoidobject motion control accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces bounding boxes as an intermediary element between text descriptions and video generation. The bounding boxes serve as a mediator that translates user intent into precise spatial constraints, allowing the system to achieve accurate motion control without requiring perfectly precise natural language descriptions. The bounding boxes act as a bridge that converts abstract text requirements into concrete positional parameters that the generation model can directly utilize.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameter representation from purely text-based to a combination of text and spatial parameters (bounding boxes). By introducing explicit spatial parameters (x, y, width, height coordinates) alongside text descriptions, the system gains finer control over object motion. This parameter enrichment allows users to specify both what should happen (text) and where/how precisely (bounding boxes), resolving the contradiction between adaptability and precision.

Inventive Principle:
Principle #35Parameter changes

2Extent of automation

If machine learning models are used for motion control, then the system can process text descriptions, but the precision of object position control deteriorates

Engineering Contradiction:
Improveautomatic motion controlVSAvoidobject position precision
Core Design Contradiction:
Extent of automationVSManufacturing precision

Solution Approach 1:

The bounding boxes serve as an intermediary that bridges the gap between automated machine learning processing and precise position control. Instead of relying solely on the model to interpret text and infer positions, the bounding boxes provide explicit positional guidance that the model can directly follow, significantly improving position precision while maintaining automation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent performs preliminary action by pre-defining bounding boxes that specify object positions before video generation occurs. These pre-established spatial constraints are then used to guide the generation process, ensuring that objects appear at the correct positions from the outset rather than relying on post-processing or iterative adjustment to achieve precise positioning.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If bounding boxes are used to constrain object positions, then the precision of object position control improves, but the device complexity increases

Engineering Contradiction:
Improveobject position precisionVSAvoidsystem architecture complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent makes the bounding box module universal and multi-functional. The same bounding box mechanism handles multiple tasks: it defines object positions, constrains motion ranges, guides generation processes, and provides spatial context for multiple objects simultaneously. This multi-functionality reduces the need for separate specialized components, thereby limiting the increase in overall system complexity despite the gain in position precision.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Reliability

If control tokens are generated to guide video generation, then the matching degree between user requirements and generated videos improves, but the generation time increases

Engineering Contradiction:
Improveuser requirement matchingVSAvoidvideo generation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The control tokens are generated in advance as preliminary guidance before the main video generation process. By pre-computing the control tokens from bounding boxes and text descriptions, the system establishes a clear roadmap for generation, which actually accelerates the overall process by reducing iterative adjustments and rejections during generation. The preliminary token generation phase is computationally efficient compared to full video generation.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP4597449A1Method and apparatus for generating video, electronic device, and computer program product
Publication Date: 2025.08.06 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • EP4597449A1 patent drawingFigure 1
  • EP4597449A1 patent drawingFigure 2
  • EP4597449A1 patent drawingFigure 3

AI summary

The present disclosure relates to a method and apparatus for generating a video, an electronic device, and a computer program product. The method includes obtaining (202) a visual token for generating an image frame in the video. The method further includes obtaining (204) a control token for constraining position information of an object in the image frame. In addition, the method also includes generating (206) the image frame in the video based on the visual token and the control token, where the object in the image frame satisfies the position information.