3D Visual Perception Model Position Encoding for Camera Adaptability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current neural network models for three-dimensional visual perception tasks suffer from poor generalization, leading to inaccurate and unreliable results when applied to images captured by cameras with significantly different parameters.

Innovation Solution

A method that involves obtaining images from cameras mounted on movable devices, determining position information based on camera parameters, generating position encoding and fusion feature maps, and using these maps to produce accurate three-dimensional visual perception results.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a neural network model is trained on images from a specific camera, then the model can achieve good performance on that camera, but the model shows poor generalization when applied to images from cameras with significantly different parameters

Engineering Contradiction:
Improvethree-dimensional visual perception accuracyVSAvoidmodel generalization capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies parameter changes by explicitly incorporating camera parameters (focal length, optical center, pixel size) into the neural network model through position encoding feature maps. This allows the model to adapt to different camera configurations by adjusting how spatial positions are encoded, thereby maintaining accurate three-dimensional visual perception across cameras with significantly different parameters without retraining the entire model

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments the camera parameter integration into a separate position encoding module that generates position encoding feature maps. This modular approach allows the camera parameter information to be independently computed and then fused with image features, enabling the model to handle different camera parameters without affecting the core three-dimensional visual perception architecture

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If camera parameter information is integrated into the model calculation process, then the model can handle different camera parameters, but the model complexity increases

Engineering Contradiction:
Improvecamera parameter adaptabilityVSAvoidmodel structure complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces position encoding feature maps as an intermediary that carries camera parameter information into the neural network model. Instead of directly integrating complex camera parameter transformations into the main model architecture, the position encoding maps serve as a mediator that encodes spatial positions according to camera parameters, which are then seamlessly fused with image features in existing fusion layers, thereby adding adaptability without significantly increasing overall model complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250131635A1Three-dimensional visual perception method, model training method and apparatus, medium, and device
Publication Date: 2025.04.24 BEIJING HORIZON INFORMATION TECH CO LTD
  • US20250131635A1 patent drawing
  • US20250131635A1 patent drawing
  • US20250131635A1 patent drawing

AI summary

Disclosed are a three-dimensional visual perception method, a model training method and a device. The three-dimensional visual perception method includes: obtaining an image captured by a camera mounted on a movable device; determining, based on a camera parameter corresponding to the image, position information respectively corresponding to at least partial pixels in the image within a camera coordinate system; generating a position encoding feature map based on the position information respectively corresponding to the at least partial pixels; generating a fusion feature map based on the image and the position encoding feature map; and generating, based on the fusion feature map, a three-dimensional visual perception result corresponding to the image by using a three-dimensional visual perception model. According to the embodiments of this disclosure, accuracy and reliability of the three-dimensional visual perception result are well ensured.