Multi-view Neural Human Rendering via Sparse View Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating high-quality 3D models of humans in motion are limited by reliance on dense camera setups, manual processing, and inability to handle occlusions and complex clothing, leading to poor reconstruction quality and artifacts.

Innovation Solution

The proposed method employs a multi-view neural human rendering technique using PointNet++ for feature extraction, U-Net for image synthesis, and anti-aliased convolutional neural networks to decode feature maps into images and foreground masks, eliminating the need for dense view samples and improving translation invariance, while also using visual hull reconstruction to patch holes and refine geometry.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional modeling and rendering pipelines are used with dense camera systems, then high fidelity reconstruction can be achieved, but device complexity and cost increase significantly

Engineering Contradiction:
Improvereconstruction fidelityVSAvoidcamera system complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical/optical system of dense camera arrays with a neural network-based rendering system. The neural renderer learns to generate photorealistic images from sparse input views, substituting the need for complex physical camera systems with an intelligent software solution that achieves similar or superior reconstruction fidelity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates virtual copies of the scene from sparse input views by training a neural network to synthesize new viewpoints. Instead of capturing all views physically, the system learns to copy and generate missing views computationally, reducing the need for dense physical camera systems.

Inventive Principle:
Principle #26Copying

2Device complexity

If image-based modeling and rendering methods are used, then fewer cameras are needed, but occlusions and fine details cannot be preserved

Engineering Contradiction:
Improvecamera system complexityVSAvoiddetail preservation
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent replaces traditional image-based rendering methods with a neural network-based approach that can preserve fine details and handle occlusions. The neural renderer learns to infer hidden details and resolve occlusions by analyzing patterns across multiple input views, achieving detail preservation that traditional IBMR methods cannot accomplish.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces a neural network as an intermediary between the sparse input views and the final rendered output. This neural intermediary learns to fill in missing information, resolve occlusions, and preserve fine details by synthesizing intermediate representations that capture the essential features of the scene.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Manufacturing precision

If SMPL model fitting is used to improve proxy geometry, then shape accuracy improves, but the model cannot handle clothing or strong shape variations

Engineering Contradiction:
Improveshape accuracyVSAvoidhandling of clothing and shape variations
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal rendering system that can handle multiple types of objects and appearances, including humans with various clothing styles and body shapes. The neural network learns to generalize across different appearances and clothing types, making the system versatile rather than limited to bare-skin models like SMPL.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent moves from fixed parametric models like SMPL to a flexible neural network approach that can adapt to various shape variations and clothing types. The neural renderer learns continuous representations that can capture the full range of human appearance variations without being constrained by predefined model parameters.

Inventive Principle:
Principle #35Parameter changes

4Manufacturing precision

If manual work is performed to produce commercial quality results, then output quality improves, but productivity decreases

Engineering Contradiction:
Improveoutput qualityVSAvoidprocessing speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent implements a self-service system where the neural network automatically processes the rendering task without requiring manual intervention. The trained neural renderer can generate photorealistic images autonomously from sparse input views, eliminating the need for manual post-processing while maintaining commercial quality output.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual human work with an automated neural network system. The neural renderer learns to perform the complex task of high-quality rendering automatically, substituting human operators with an intelligent algorithm that can process images at much higher speeds while maintaining or improving output quality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11880935B2Multi-view neural human rendering
Publication Date: 2024.01.23 SHANGHAI TECH UNIV
  • US11880935B2 patent drawing
  • US11880935B2 patent drawing
  • US11880935B2 patent drawing

AI summary

An image-based method of modeling and rendering a three-dimensional model of an object is provided. The method comprises: obtaining a three-dimensional point cloud at each frame of a synchronized, multi-view video of an object, wherein the video comprises a plurality of frames; extracting a feature descriptor for each point in the point cloud for the plurality of frames without storing the feature descriptor for each frame; producing a two-dimensional feature map for a target camera; and using an anti-aliased convolutional neural network to decode the feature map into an image and a foreground mask.