Video Generation Using Bounding Boxes for Precise Motion Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video generation technologies struggle to accurately capture user-defined motion requirements, particularly when precise control over object movement is desired, as users find it difficult to express these requirements through text descriptions alone.

Innovation Solution

A method and apparatus that allow users to input content information, such as text or images, and use bounding boxes to identify and control the position of objects in both starting and ending frames, enabling precise motion control during video generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If text descriptions are used to guide video generation, then the ease of operation is improved, but the manufacturing precision of motion control deteriorates

Engineering Contradiction:
Improveease of operationVSAvoidmotion control precision
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent segments the motion control process into two distinct input modes: text descriptions for general guidance and bounding box annotations for precise object localization. This segmentation allows users to choose the appropriate level of detail for different aspects of video generation, maintaining ease of operation while improving motion control precision when needed.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces bounding boxes as an intermediary tool between the user's intent and the video generation system. These visual markers serve as a mediator that translates abstract motion requirements into concrete spatial constraints, enabling precise object tracking and movement control without requiring complex text descriptions.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If bounding boxes are used to control object positions, then the manufacturing precision of motion control is improved, but the device complexity increases

Engineering Contradiction:
Improvemotion control precisionVSAvoiddevice complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent uses simplified 2D bounding box representations as copies of the actual 3D object positions and movements. These bounding box copies capture the essential spatial information needed for motion control without requiring complex 3D modeling or tracking systems, thus improving precision while maintaining system simplicity.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent changes the control parameters from abstract text descriptions to concrete spatial coordinates defined by bounding boxes. This parameter transformation enables precise mathematical computation of object movements and camera trajectories, improving motion control precision through straightforward coordinate geometry rather than complex interpretation algorithms.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4597494A1Method, apparatus, device and computer program product for generating video
Publication Date: 2025.08.06 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • EP4597494A1 patent drawingFigure 1
  • EP4597494A1 patent drawingFigure 2
  • EP4597494A1 patent drawingFigure 3A

AI summary

The present disclosure relates to a method and apparatus for generating a video, a device, and a computer program product. The method (200) includes obtaining (202) content information related to content of the video to be generated, where the content information includes at least one of a text or an image. The method (200) further includes obtaining (204) position information indicating a position of an object in the video in a starting frame. The method (200) also includes obtaining (206) control information that constrains a position of the object in an ending frame. In addition, the method (200) further includes generating (208) the video based on the content information, the position information, and the control information.