3D Perception Network Adaptation Across Camera Rig Viewpoints

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Autonomous vehicles face challenges in developing generalized deep neural networks for multi-view or 3D perception due to the diversity of camera rig configurations across different vehicle models, leading to degraded performance when processing data from different camera rigs, especially in nonplanar road environments.

Innovation Solution

The system adapts a 3D perception network by training layers using real source rig data and simulated source and target rig data, employing feature statistics to transform features and minimize differences across training channels, enabling viewpoint-adaptive perception across various camera configurations without the need for extensive data collection for each vehicle model.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a dedicated DNN is trained for each vehicle model using real source rig data, then detection accuracy is improved, but resource intensity and development complexity increase significantly

Engineering Contradiction:
Improvedetection accuracyVSAvoiddevelopment complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent develops a universal 3D perception network that can process data from multiple camera rig configurations across different vehicle models. Instead of creating separate dedicated networks for each vehicle model, a single generalized network is trained to handle diverse camera arrangements, making the system multi-functional and adaptable to various platforms without requiring model-specific retraining

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent employs domain adaptation techniques that transform features from source domain (one vehicle model's camera rig) to target domain (another vehicle model's camera rig) by adjusting statistical parameters. This involves modifying feature distributions through parameter changes rather than retraining the entire network, thereby maintaining detection accuracy across different configurations while reducing development resources

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If data collection is performed for each vehicle model, then detection accuracy is improved, but time consumption and resource intensity increase

Engineering Contradiction:
Improvedetection accuracyVSAvoiddata collection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary domain adaptation training using simulated target rig data before deploying the network to actual target vehicles. This preliminary action with synthetic data prepares the network to handle target domain characteristics, reducing the need for extensive real-world data collection from each new vehicle model and accelerating the deployment timeline

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses simulated data as a copy or proxy for real target rig data. By training on synthesized images that mimic the target vehicle's camera configuration, the system avoids the time-consuming process of collecting, annotating, and processing real-world data from each new vehicle model, while still achieving effective domain adaptation

Inventive Principle:
Principle #26Copying

3Ease of manufacture

If a generalized DNN is trained without model-specific data, then development resources are reduced, but performance degrades on specific vehicle models

Engineering Contradiction:
Improvedevelopment efficiencyVSAvoiddetection accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent introduces domain adaptation as an intermediary process between source domain training and target domain deployment. The adaptation module acts as a mediator that bridges the gap between generalized pretraining and model-specific performance requirements, allowing the system to maintain both development efficiency and detection accuracy by translating features across domains without requiring model-specific retraining

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240317263A1Viewpoint-adaptive perception for autonomous machines and applications using real and simulated sensor data
Publication Date: 2024.09.26 NVIDIA CORP
  • US20240317263A1 patent drawing
  • US20240317263A1 patent drawing
  • US20240317263A1 patent drawing

AI summary

Systems and methods are disclosed relating to viewpoint adapted perception for autonomous machines and applications. A 3D perception network may be adapted to handle unavailable target rig data by training one or more layers of the 3D perception network as part of a training network using real source rig data and simulated source and target rig data. Feature statistics extracted from the real source data may be used to transform the features extracted from the simulated data during training. The paths for real and simulated data through the resulting network may be alternately trained on real and simulated data to update shared weights for the different paths. As such, one or more of the paths through the training network(s) may be designated as the 3D perception network, and target rig data may be applied to the 3D perception network to perform one or more perception tasks.