3D Model Placement in Video Using Plane Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for placing three-dimensional models in videos, such as video-in advertisements, are inefficient and require complex calculations involving feature points and large parallax, limiting their applicability and speed.

Innovation Solution

A method and apparatus that place a target three-dimensional model on a video frame, track a target plane to calculate poses in subsequent frames, and replace it with a displaylink model without relying on image feature points or large parallax, using a neural network for plane detection and homography matrices to achieve accurate and efficient placement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional methods use feature point calculation and large parallax for placing three-dimensional models in videos, then placement accuracy can be achieved, but computational complexity increases and processing speed decreases

Engineering Contradiction:
Improveplacement accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and removes the dependency on feature point calculation and large parallax from the conventional video-in advertisement process. By using plane tracking instead of feature point matching, the method eliminates complex computational steps while maintaining placement accuracy, directly resolving the contradiction between measurement precision and device complexity

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the fundamental parameters of the approach by shifting from feature point coordinates to plane homography parameters. This parameter transformation allows the system to achieve accurate three-dimensional model placement through simpler homography matrix calculations rather than complex feature point correspondence, reducing computational complexity while preserving placement precision

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If conventional methods rely on feature point calculation for video-in advertisement, then placement can be achieved, but processing speed decreases

Engineering Contradiction:
Improveplacement accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSSpeed

Solution Approach 1:

The patent extracts the time-consuming feature point calculation step from the processing pipeline and replaces it with efficient plane tracking. This removal of computational bottleneck directly increases processing speed while the extracted plane homography method maintains placement accuracy through alternative mathematical formulation

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent performs preliminary plane detection and homography matrix calculation on the first frame before processing subsequent frames. This preliminary action establishes the transformation relationship in advance, allowing rapid tracking and placement in following frames without repeating complex feature point calculations, thereby significantly improving overall processing speed

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If conventional methods use large parallax for three-dimensional model placement, then accuracy can be improved, but the method becomes less versatile and applicable to fewer videos

Engineering Contradiction:
Improveplacement accuracyVSAvoidvideo applicability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent implements a universal plane tracking approach that can be applied to any video containing detectable planes, regardless of camera motion characteristics. This multi-functional method replaces the specialized large parallax requirement with a general-purpose homography-based solution, enhancing video applicability while maintaining placement accuracy through robust plane detection and tracking

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

Instead of requiring large parallax to achieve accuracy, the patent inverts the approach by using plane homography transformation that works with various camera motions. This inversion of the fundamental assumption allows the method to achieve placement accuracy without being constrained by specific video characteristics, thereby improving versatility across different video types

Inventive Principle:
Principle #13The other way round (Inversion)

4Ease of manufacture

If conventional video advertisement methods are used (adding ads to beginning/end or floating layers), then implementation is simple, but user experience and traffic coverage are inferior

Engineering Contradiction:
Improveimplementation simplicityVSAvoiduser experience and traffic coverage
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent replaces the simple but ineffective mechanical approaches (floating layers, static ads) with an intelligent computer vision-based system. By substituting the mechanical placement method with automated plane detection, tracking, and three-dimensional model rendering, the system achieves both improved user experience through immersive ads and maintains operational efficiency through automated processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentEP3923248B1Image processing method and apparatus, electronic device and computer-readable storage medium
Publication Date: 2025.10.22 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • EP3923248B1 patent drawingFigure 1~2
  • EP3923248B1 patent drawingFigure 3
  • EP3923248B1 patent drawingFigure 4~5

AI summary

The present disclosure provides an image processing method and apparatus, an electronic device, and a computer-readable storage medium. The method includes: obtaining a to-be-processed video; placing a target three-dimensional model on a target plane of a first frame of image of the to-be-processed video; obtaining three-dimensional coordinates of a plurality of feature points of a target surface in a world coordinate system and pixel coordinates of the plurality of feature points of the target surface on the first frame of image; obtaining a pose of a camera coordinate system of the first frame of image relative to the world coordinate system; obtaining, according to the target plane, the three-dimensional coordinates of the plurality of feature points of the target surface in the world coordinate system, and the pixel coordinates of the plurality of feature points of the target surface on the first frame of image, poses of camera coordinate systems of a second frame of image to an mth frame of image of the to-be-processed video relative to the world coordinate system; and replacing the target three-dimensional model with a target displaylink model and placing the target displaylink model on the world coordinate system to generate, according to the pose of the camera coordinate system of each frame of image of the to-be-processed video relative to the world coordinate system, a target video including the target displaylink model.