3D Point Cloud Graph Generation Without VLM Inference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for predicting 3-dimensional scene graphs from limited training data face challenges in complexity and abstraction, requiring computationally expensive vision-language models (VLMs) for inference, and lack efficient methods to construct 3D maps directly from 3-dimensional point cloud data.

Innovation Solution

A method involving preprocessing networks and a graph neural network to refine feature vectors from 3D point clouds, combined with vision-language models for 2D image data, to directly generate graph representations of instances and their relationships without relying on VLMs at inference time, allowing for explicit semantic relationship prediction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If vision-language models (VLMs) are used for predicting 3-dimensional scene graphs, then the ability to handle open-vocabulary instances and relationships is improved, but the computational expense and complexity increase

Engineering Contradiction:
Improveopen-vocabulary capabilityVSAvoidmodel complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments the scene graph prediction task into distinct components: a point cloud processing network that extracts geometric features from 3D data, and a language processing component that handles open-vocabulary relationships. This segmentation allows each component to be optimized independently, reducing overall complexity while maintaining versatility.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediate representation layer that bridges 3D geometric data and language descriptions. This intermediary structure enables the system to handle open-vocabulary relationships without requiring the full complexity of end-to-end VLM processing, thus reducing computational expense while preserving adaptability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If vision-language models are used for inference, then semantic relationship prediction capability is improved, but the computational expense increases

Engineering Contradiction:
Improvesemantic relationship prediction accuracyVSAvoidcomputational expense
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary processing of 3D point cloud data to extract geometric features and potential relationships before semantic analysis. This pre-processing step organizes the data in a way that reduces the computational burden during the semantic relationship prediction phase, maintaining accuracy while lowering overall computational expense.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the computationally intensive mechanism of using VLMs for 3D scene graph prediction with a hybrid approach that substitutes geometric reasoning based on point cloud analysis for parts of the language model processing. This substitution maintains semantic prediction capability while significantly reducing computational requirements.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Loss of information

If 2-dimensional VLMs are used to process input data, then semantic understanding is improved, but the ability to directly process 3-dimensional point cloud data is lost

Engineering Contradiction:
Improvesemantic understanding capabilityVSAvoiddata processing pipeline complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent extends 2D language processing capabilities into the 3D domain by processing point cloud data directly along the depth dimension. This allows the system to maintain semantic understanding capabilities while operating natively in 3D space, avoiding the need to project 3D data to 2D and back, thus simplifying the data processing pipeline.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentEP4603988A1Device and method for generating a graph representation from a 3-dimensional point cloud
Publication Date: 2025.08.20 ROBERT BOSCH GMBH
  • EP4603988A1 patent drawingFigure 1
  • EP4603988A1 patent drawingFigure 2
  • EP4603988A1 patent drawingFigure 3

AI summary

Method for training a first machine learning system (10) for generating a graph representation of objects and their relationships in a 3D environment scene from 3D point cloud input data (P). For each object i and each pair of objects i and j in the scene initial node feature vectors ϕi and initial edge feature vectors ϕij are determined from the point cloud input data (P) and are arranged in an initial graph structure (ϕi, ϕij, ϕi). A refined graph structure (11) is determined by a graph neural network (10). From 2-dimensional image sensor data (2, 3) of the environment scene, feature vectors (21) of objects i are determined by a second machine learning system (20) and feature vectors (31) of object pairs i and j are determined by a third machine learning system (30). Parameters of the first machine learning system (10) are adjusted with respect to a training objective defined by an optimization of a difference between the refined node feature vector (12) of object i and the corresponding feature vector (21) of object i and/or an optimization of a difference between refined edge feature vector (13) of object i and j and the corresponding feature vector (31) of object pair i and j.