Video Generation With Bounding-Box Motion Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video generation models struggle to accurately understand and implement user-defined motion modes for objects in generated videos, especially when precise descriptions are difficult to articulate, leading to a mismatch between user expectations and generated content.

Innovation Solution

A method and apparatus that utilize bounding boxes to constrain object positions and sizes, generating control tokens to guide the video generation process, ensuring objects move as intended by incorporating a motion control module into the video diffusion model architecture.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If text descriptions are used to guide video generation, then the system can understand user intent, but the precision of motion control deteriorates when precise descriptions are difficult to articulate

Engineering Contradiction:
Improvetext understanding capabilityVSAvoidmotion control precision
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces bounding boxes as an intermediary tool between text descriptions and video generation. The bounding box module converts imprecise text motion descriptions into precise spatial constraints by defining rectangular regions that objects must follow, thereby mediating the gap between user intent and precise motion control

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent adds a spatial dimension to text-based video generation by introducing bounding box coordinates. Instead of relying solely on natural language motion descriptions, the system incorporates explicit spatial positioning information through bounding box parameters, enabling precise control in the spatial dimension

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If bounding boxes are used to constrain object positions, then motion control precision is improved, but the device complexity increases due to additional modules

Engineering Contradiction:
Improveobject position precisionVSAvoidsystem structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges the bounding box module with the existing video diffusion model, integrating spatial constraint functionality into the generation process. The control token generation is embedded within the diffusion model's architecture, allowing bounding box constraints to be applied without requiring completely separate processing systems

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The bounding box module serves multiple functions: it constrains object positions, defines motion trajectories, and generates control tokens for the diffusion model. This multi-functionality reduces the need for separate specialized modules for each control aspect

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12412319B2Method and apparatus for generating video, electronic device, and computer program product
Publication Date: 2025.09.09 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • US12412319B2 patent drawing
  • US12412319B2 patent drawing
  • US12412319B2 patent drawing

AI summary

The present disclosure relates to a method and apparatus for generating a video, an electronic device, and a computer program product. The method includes obtaining a visual token for generating an image frame in the video. The method further includes obtaining a control token for constraining position information of an object in the image frame. In addition, the method also includes generating the image frame in the video based on the visual token and the control token, where the object in the image frame satisfies the position information.