3D Understanding Model Pretraining With Image, Text, And Point Clouds

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

3D visual understanding models are limited by small datasets and high data collection and annotation costs, hindering their generalization and real-world applications.

Innovation Solution

A 3D visual understanding framework (ULIP-2) generates well-aligned, holistic multimodal data for 3D understanding by learning unified representations of image, text, and point cloud using an efficient multimodal pre-training architecture, aligning features across these modalities without requiring manual annotations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional 3D visual understanding models are used with small datasets, then model training is feasible with limited data, but generalization ability and real-world application performance deteriorate

Engineering Contradiction:
Improvegeneralization abilityVSAvoiddataset size
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent creates synthetic 3D training data by rendering images from 3D models and generating corresponding point cloud representations. This copying approach allows the model to learn from artificially generated data that mimics real-world scenarios, thereby improving generalization ability without requiring large amounts of expensive real annotated 3D data

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary preprocessing to convert 3D models into point cloud representations and generate associated 2D images before training. This preliminary action creates a structured training dataset that enables the model to learn the relationship between 3D structures and 2D observations in advance, improving downstream task performance

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If manual annotation is used for 3D data, then data quality and accuracy are improved, but data collection and annotation costs increase

Engineering Contradiction:
Improvedata annotation accuracyVSAvoiddata collection cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The system generates its own training data by automatically rendering images from 3D models and creating point cloud representations without requiring manual annotation. This self-service approach eliminates the need for expensive human annotators while maintaining data quality, as the synthetic data inherently contains accurate ground truth information

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Instead of manually annotating real 3D data, the patent creates synthetic copies of 3D scenes with automatically generated point clouds and images. These copied representations serve as high-quality training data without the cost and time requirements of manual annotation processes

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If multimodal pretraining is implemented, then recognition ability and cross-domain task performance are improved, but system complexity increases

Engineering Contradiction:
Improvecross-domain task capabilityVSAvoidmodel architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges multiple modalities (3D point clouds, 2D images, and text descriptions) into a unified training framework. By combining these different data types during pretraining, the model learns complementary representations that improve its ability to perform diverse downstream tasks while managing complexity through integrated processing

Inventive Principle:
Principle #5Merging (Combining)

4Reliability

If large-scale 3D datasets are collected, then model generalization is improved, but data collection time and resources increase

Engineering Contradiction:
Improvemodel generalizationVSAvoiddata collection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent generates large-scale training datasets by copying and transforming existing 3D models into various viewpoints and conditions through rendering. This approach creates extensive synthetic data quickly without the time-consuming process of collecting and annotating real-world 3D data at scale

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary generation of diverse training samples by rendering images from multiple angles and creating varied point cloud representations before training begins. This preliminary action creates a comprehensive training dataset that improves generalization without requiring time-consuming real-world data collection

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12430849B2Systems and methods for multimodal pretraining for three-dimensional understanding models
Publication Date: 2025.09.30 SALESFORCE INC
  • US12430849B2 patent drawing
  • US12430849B2 patent drawing
  • US12430849B2 patent drawing

AI summary

A method of training a neural network based three-dimensional (3D) encoder is provided. A first plurality of samples of a training dataset are generated using a first 3D model. An image generator with multi-view rendering is used to generate a plurality of two-dimensional (2D) images having different viewpoints of the first 3D model. A first language model is used to generate a plurality of texts corresponding to the plurality of 2D images respectively. A first text for a first image is generated by using one or more text descriptions generated by the first language model. A point cloud is generated by randomly sampling points in the 3D model. The first plurality of samples are generated using the plurality of 2D images, the corresponding plurality of texts, and the point cloud. The neural network based 3D encoder is trained using the training dataset including the first plurality of samples.