Multi-View Scene Graph Fusion for Real-Time 3D Relation Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating multi-dimensional scene graphs face challenges in real-time processing and computational overhead, particularly in handling incomplete and dynamic 3D data, limiting their application in complex visual tasks.

Innovation Solution

A method involving object detection, entity re-identification, and multi-view-direction information fusion using neural networks and geometric constraints to establish a multi-dimensional scene graph with complex light field, leveraging Faster RCNN and Neural-Motifs algorithms for object detection and relation prediction, and Dempster Shafer evidence theory for feature fusion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional methods are used to generate multi-dimensional scene graphs, then processing accuracy can be maintained, but computational overhead increases and real-time processing becomes difficult

Engineering Contradiction:
Improvereal-time processing capabilityVSAvoidcomputational overhead
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent segments the scene graph generation process into multiple independent modules: object detection module, relation prediction module, re-identification module, and scene graph generation module. Each module processes specific tasks separately, allowing for optimized computation and reduced overall computational overhead while maintaining real-time processing capability through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by pre-processing 2D images to extract entity features and semantic relations before generating the final 3D scene graph. The object detection and relation prediction are performed in advance on 2D images, which then guide the 3D scene graph generation, reducing the computational burden during real-time processing.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If comprehensive multi-view-direction information is processed, then scene graph accuracy improves, but processing time and computational complexity increase

Engineering Contradiction:
Improvescene graph accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies local quality by processing different view directions with varying levels of detail based on their importance and availability. The system prioritizes processing information from view directions that provide the most significant constraints for accurate 3D scene graph generation, while reducing processing intensity for less critical view directions, thereby balancing accuracy with processing time.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent employs partial action by selectively processing only the necessary view directions and information components required for accurate scene graph generation. Instead of uniformly processing all multi-view-direction information, the system identifies and processes the essential constraints and relationships, reducing overall processing time while maintaining sufficient accuracy.

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If complex light field information is integrated, then scene understanding capability improves, but device complexity and algorithm difficulty increase

Engineering Contradiction:
Improvescene understanding capabilityVSAvoidalgorithm complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary approach by using 2D images as a mediator to bridge the gap between simple image processing and complex 3D scene graph generation. The 2D images serve as an intermediate representation that captures essential scene information, which is then transformed into 3D scene graphs through standardized conversion algorithms, reducing the direct complexity of processing complex light field information.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies universality by designing a multi-functional processing framework that can handle various types of input data (2D images from different view directions) and generate unified 3D scene graphs. The same core algorithmic framework processes different view directions and light field information, reducing overall system complexity through reusable components and standardized procedures.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12469213B2Method, system and apparatus for generating multi-dimensional scene graph with complex light field
Publication Date: 2025.11.11 TSINGHUA UNIVERSITY
  • US12469213B2 patent drawing
  • US12469213B2 patent drawing
  • US12469213B2 patent drawing

AI summary

In a method for generating a multi-dimensional scene graph with a complex light field, entity features of entities are obtained by inputting respective 2-Dimensional (2D) images captured in multiple view directions into an object detection model, and features of a respective single-view-direction scene graph are obtained by predicting a semantic relation among the entities contained in a corresponding 2D image captured in each view direction. An entity correlation for an entity among the multiple view directions is determined as an entity re-identification result. A multi-dimensional bounding box for the entity is established based on the entity re-identification result and a geometric constraint of camera parameters. A feature fusion result is obtained by fusing features of respective single-view-direction scene graphs in the multiple view directions. The multi-dimensional scene graph with the complex light field is established.