Video Generation with Bounding-Box Object Motion Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video generation technologies struggle to accurately understand user requirements for object motion in videos, especially when precise motion control is needed, making it difficult to generate desired video effects.

Innovation Solution

A method and apparatus that allows users to input content information, such as text or images, and use bounding boxes to identify and control object positions in starting and ending frames, enabling precise motion control through video generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If text-guided video generation is used, then creativity and ease of operation are improved, but manufacturing precision and measurement precision of object motion are insufficient

Engineering Contradiction:
Improveease of video creationVSAvoidprecision of object motion control
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent segments the video generation process into distinct controllable components: content information (text/image), position information (starting frame object locations), and control information (ending frame constraints). This segmentation allows users to independently control each aspect, achieving both ease of operation through modular input and precision through specific parameter control for object motion.

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If motion control is added to text-guided video generation, then manufacturing precision of object motion is improved, but device complexity increases

Engineering Contradiction:
Improveprecision of object motion controlVSAvoidcomplexity of video generation system
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent implements a universal video generation system that handles multiple input types (text and images) through a single integrated process. The system universally processes content information, position information, and control information regardless of the specific input modality, reducing complexity by avoiding separate specialized systems for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces position information as an intermediary element that bridges content information and control information. This intermediary structure organizes the complex data flow by first identifying object positions in the starting frame, then applying ending frame constraints, making the overall system more manageable and less complex.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Manufacturing precision

If precise motion control parameters are required, then manufacturing precision is improved, but ease of operation deteriorates

Engineering Contradiction:
Improveprecision of object position controlVSAvoidease of video generation
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The patent performs preliminary action by first obtaining position information that identifies object locations in the starting frame before applying control constraints. This preliminary step establishes a clear reference framework that simplifies subsequent control operations, making precision control more intuitive and easier to operate.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies local quality by allowing different levels of control precision for different aspects of video generation. Users can provide detailed position and control information specifically for objects requiring precise motion control, while other elements can be generated with standard text-guided controls, optimizing ease of operation where precision is not critical.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12608851B2Method, apparatus, device and computer program product for generating video
Publication Date: 2026.04.21 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • US12608851B2 patent drawing
  • US12608851B2 patent drawing
  • US12608851B2 patent drawing

AI summary

The present disclosure relates to a method and apparatus for generating a video, a device, and a computer program product. The method includes obtaining content information related to content of the video to be generated, where the content information includes at least one of a text or an image. The method further includes obtaining position information indicating a position of an object in the video in a starting frame. The method also includes obtaining control information that constrains a position of the object in an ending frame. In addition, the method further includes generating the video based on the content information, the position information, and the control information.