GAN Point Cloud Repair for Immersive Viewpoint Navigation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems for providing immersive experiences with captured video content, such as VR and AR, face limitations in allowing users to freely navigate viewpoints due to pre-defined or restricted camera positions, leading to discomfort and reduced immersion, especially when occlusions and image resolution issues arise, making it difficult to capture sufficient image information and derive necessary lighting and depth information.

Innovation Solution

A digital model repair system that captures and processes content using multiple cameras to generate a three-dimensional representation of the environment, allowing users to freely navigate viewpoints through image fusion and depth estimation, and utilizes a generative adversarial network to improve the accuracy of point clouds by reducing noise and errors, enabling more comprehensive and immersive experiences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple cameras are used to capture content from different positions, then viewpoint navigation freedom is improved, but device complexity increases

Engineering Contradiction:
Improveviewpoint navigation freedomVSAvoidcamera system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments the camera system into multiple independent camera units distributed throughout the environment. Each camera captures content from its specific position, and the processing system segments the captured content into discrete viewpoint data that can be independently selected and rendered, enabling free viewpoint navigation without requiring a single complex omnidirectional camera system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The multiple camera system is designed with universal functionality where each camera can capture content that serves multiple purposes: creating 2D video content, generating 3D immersive content, and providing various viewpoints for different user preferences. The same hardware infrastructure supports multiple content types and delivery formats, reducing overall system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If image processing is performed to improve quality and reduce noise, then image quality is improved, but processing time increases

Engineering Contradiction:
Improveimage qualityVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary processing of captured content during the content creation phase, pre-computing multiple viewpoints and storing them in an optimized format. This preliminary action reduces the need for intensive real-time processing during content delivery, as the heavy computational work of viewpoint generation and noise reduction is completed beforehand during content production.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies selective processing based on the specific content and delivery requirements. Not all captured content undergoes full processing pipelines; instead, processing is applied partially or excessively only where needed to achieve the desired quality level, balancing processing time against quality improvement for different content types and delivery scenarios.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11501118B2Digital model repair system and method
Publication Date: 2022.11.15 SONY INTERACTIVE ENTERTAINMENT LLC
  • US11501118B2 patent drawing
  • US11501118B2 patent drawing
  • US11501118B2 patent drawing

AI summary

A digital model repair method includes: providing a point cloud digital model of a target object as input to a generative network of a trained generative adversarial network ‘GAN’, the input point cloud comprising a plurality of points erroneously perturbed by one or more causes, and generating, by the generative network of the GAN, an output point cloud in which the erroneous perturbation of some or all of the plurality of points has been reduced; where the generative network of the GAN was trained using input point clouds comprising a plurality of points erroneously perturbed by said one or more causes, and a discriminator of the GAN was trained to distinguish point clouds comprising a plurality of points erroneously perturbed by said one or more causes and point clouds substantially without such perturbations.