Sub-Pixel Ray Tracing for Scalable Ground Truth Image Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing real capture data for machine learning and computer vision tasks is less scalable and less accurate due to manual image capture and post-processing, which can introduce inaccuracies in ground truth data.
Innovation Solution
A computer device simulates virtual environments using ray tracing to generate sub-pixel data for each ray, storing it in image files for scalable and accurate training of machine learning models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If real capture data is used for machine learning training, then the data reflects real-world conditions, but the scalability is limited and ground truth accuracy is reduced due to manual capture and post-processing requirements
Solution Approach 1:
The patent creates virtual copies of real-world scenes through 3D environment simulation. Instead of capturing real scenes manually, the system generates synthetic images by rendering virtual representations of environments, objects, and lighting conditions. This copying approach allows unlimited data generation while maintaining ground truth accuracy through programmatic control of scene parameters.
Solution Approach 2:
The system varies multiple parameters of the virtual environment including lighting conditions, camera positions, object positions, and scene geometry to generate diverse training data. By programmatically changing these parameters, the system can produce unlimited variations of images with known ground truth, achieving both scalability and accuracy simultaneously.
2Measurement precision
If manual image capture is performed for real data collection, then the data represents actual conditions, but the process is time-consuming and less scalable
Solution Approach 1:
The patent replaces the mechanical process of physical camera capture with a computational rendering system. Instead of using actual cameras to capture scenes, the system uses computer-generated images from virtual environments. This substitution eliminates time-consuming manual capture operations while maintaining measurement precision through controlled virtual scene parameters.
Solution Approach 2:
The system performs preliminary actions by pre-defining the virtual environment, objects, and lighting conditions before generating images. The ground truth labels are determined in advance through programmatic control of the virtual scene, eliminating the need for time-consuming post-processing annotation steps required by real capture data.
3Adaptability or versatility
If ground truth data is generated through post-processing steps, then real capture data can be utilized, but the process reduces scalability and may introduce inaccuracies
Solution Approach 1:
The patent determines ground truth data in advance during the virtual environment setup phase rather than through subsequent post-processing. The scene graph and rendering parameters are configured beforehand, automatically providing accurate labels for objects, positions, and properties. This preliminary determination eliminates the need for time-consuming and error-prone post-processing annotation.
Solution Approach 2:
The virtual environment system serves itself by automatically generating both the images and their corresponding ground truth labels through the same rendering process. The system does not require separate manual annotation processes, as the virtual scene configuration inherently provides the accurate labels needed for training data.
Data Source
AI summary
A computer device includes a processor configured to simulate a virtual environment based on a set of virtual environment parameters, and perform ray tracing to render a view of the simulated virtual environment. The ray tracing includes generating a plurality of rays for one or more pixels of the rendered view of the simulated virtual environment. The processor is further configured to determine sub-pixel data for each of the plurality of rays based on intersections between the plurality of rays and the simulated virtual environment, and store the determined sub-pixel data for each of the plurality of rays in an image file.


