ROV Vision Training With Synthetic Data for Real-World Generalization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Training machine learning models for remotely operated vehicles (ROVs) using synthetic data is challenging due to differences in textures and lighting between simulated and real-world data, which can lead to poor generalization from synthetic to real data.
Innovation Solution
The system uses a synthetic training engine that replays real examples in a virtual world, constraining features extracted from both real and virtual images to be equal, and includes a simulator module that generates synthetic data from real mission telemetry and 3D model data, allowing for automatic annotation and training of machine learning models that can perform well on real data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If synthetic data is used to train machine learning models, then the dataset size increases and annotation cost decreases, but the model generalization to real data deteriorates due to differences in textures and lighting
Solution Approach 1:
The patent creates synthetic copies of real-world scenes by rendering 3D models with captured camera poses and sensor data. These synthetic images serve as training data while preserving the structural and geometric properties of real scenes, enabling model training without direct use of expensive annotated real images.
Solution Approach 2:
The system varies multiple parameters in synthetic data generation including lighting conditions, camera poses, sensor noise levels, and material properties. By adjusting these parameters to match real-world variations, the synthetic data becomes more representative of actual operating conditions, improving model generalization.
2Measurement precision
If human annotators are used to label real images, then the annotation accuracy improves, but the time consumption and cost increase significantly
Solution Approach 1:
Instead of manually annotating real images, the system generates synthetic images with automatically known ground truth annotations. The 3D models provide precise geometric information, and the rendering process inherently provides accurate labels for objects, surfaces, and spatial relationships without human intervention.
Solution Approach 2:
The synthetic data generation process is self-annotating. The rendering engine automatically produces images with embedded ground truth information from the 3D models, eliminating the need for external human annotators. The system serves its own annotation needs through the physical-based rendering process.
Data Source
AI summary
The present invention provides systems and methods for leveraging synthetic data to train machine learning models. A synthetic training engine may be used to train machine learning models. The synthetic training engine can automatically annotate real images for valuable tasks, such as object segmentation, depth map estimation, and classifying whether a structure is in an image. The synthetic training engine can also train the machine learning model with synthetic images in such a way that the machine learning model will work on real images. The output of the machine learning model may perform valuable tasks, such as the detection of integrity threats in underwater structures.


