3D Scene Representation With Semantic Latent Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for constructing a representation of a three-dimensional space for robotic navigation and interaction lack efficient incorporation of semantic information, leading to limited functionality and accuracy in real-time applications.

Innovation Solution

A system that processes image data to generate latent representations associated with semantic segmentations and depth maps, using an optimisation engine to jointly optimise these representations in a latent space, thereby improving semantic consistency and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If sparse techniques are used to generate three-dimensional representations, then real-time processing capability is improved, but semantic information accuracy deteriorates

Engineering Contradiction:
Improvereal-time processing capabilityVSAvoidsemantic information accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the three-dimensional representation into multiple views (front view, back view, side views) and processes each view separately through neural networks to generate semantic segmentations. This segmentation allows real-time processing of individual views while maintaining comprehensive semantic coverage of the entire object across multiple perspectives.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from processing two-dimensional images to generating and processing three-dimensional representations composed of multiple views. By representing objects in 3D space with multiple viewpoint images, the system achieves both real-time processing efficiency and accurate semantic information capture that single-view 2D processing cannot provide.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If dense mapping techniques are used to generate three-dimensional representations, then semantic information accuracy is improved, but computational requirements and processing time increase

Engineering Contradiction:
Improvesemantic information accuracyVSAvoidcomputational requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the complex dense mapping task into separate processing streams for different views (front, back, sides). Each view is processed independently through dedicated neural networks, reducing the computational burden on any single processing unit while collectively achieving comprehensive semantic coverage of the three-dimensional object.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent processes multiple views of the object, which is more than a single view would provide. This excessive action of capturing and processing multiple perspectives ensures comprehensive semantic information is obtained, while each individual view remains computationally manageable for real-time processing.

Inventive Principle:
Principle #16Partial or excessive action

3Ease of operation

If traditional computer vision systems are used, then geometric representation is achieved, but semantic understanding capability deteriorates

Engineering Contradiction:
Improvegeometric representation capabilityVSAvoidsemantic understanding capability
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent merges geometric representation with semantic understanding by integrating three-dimensional geometric modeling with neural network-based semantic segmentation. The system simultaneously produces both the geometric structure (multiple view images representing 3D space) and semantic information (class labels for each pixel region), eliminating the information loss present in traditional systems that handle only geometry.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a composite representation system that combines traditional geometric computer vision techniques with modern deep learning semantic segmentation. This composite approach integrates the strengths of both methodologies: accurate geometric modeling from computer vision and rich semantic understanding from neural networks, producing a unified representation that contains both geometric and semantic information.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS12205297B2Scene representation using image processing
Publication Date: 2025.01.21 IMPERIAL COLLEGE INNVOATIONS LTD
  • US12205297B2 patent drawing
  • US12205297B2 patent drawing
  • US12205297B2 patent drawing

AI summary

Certain examples described herein relate to a system for processing image data. In such examples, the system includes an input interface to receive the image data, which is representative of at least one view of a scene. The system also includes an initialisation engine to generate a first latent representation associated with a first segmentation of at least a first view of the scene, wherein the first segmentation is a semantic segmentation. The initialisation engine is also arranged to generate a second latent representation associated with at least a second view of the scene. The system additionally includes an optimisation engine to jointly optimise the first latent representation and the second latent representation, in a latent space, to obtain an optimised first latent representation and an optimised second latent representation.