Synthetic Labeled Training Data Generation from CAD Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Creating large, balanced labeled data sets for AI algorithms like neural networks is time-consuming and expensive, especially in industrial environments, and manual labeling is prone to errors due to lack of concentration and monotony, while reflecting environmental influences is challenging.
Innovation Solution
A computer-implemented method that automatically generates labeled training data sets using CAD models, where the CAD model's information about electronic components' coordinates and relationships is used to create labeled render images, which are then used to train image recognition functions, allowing for efficient and error-free data set creation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual labeling is performed, then labeling accuracy can be maintained, but time consumption and cost increase significantly
Solution Approach 1:
The patent creates synthetic renderings of electronic components and assemblies that replicate real-world visual characteristics. These synthetic images serve as copies of real images but with known ground truth labels from the CAD model, eliminating the need for manual labeling while maintaining accuracy through consistent rendering pipelines that preserve visual fidelity.
Solution Approach 2:
The system uses the CAD model itself to generate both the structural information and the visual representations. The CAD model's coordinate data and component definitions are automatically transformed into labeled training data through rendering, making the system self-sufficient without requiring external manual annotation efforts.
2Loss of information
If manual labeling is performed, then detailed component information can be captured, but cost increases due to specialist knowledge requirements
Solution Approach 1:
The patent extracts component information directly from the CAD model's bill of materials and coordinate data, creating digital copies of component metadata that are automatically associated with rendered images. This eliminates the need for human experts to annotate component information while preserving complete technical details through structured data extraction.
Solution Approach 2:
The patent replaces the mechanical process of human expert annotation with automated computational processes. The system uses computer vision algorithms and rendering engines to automatically generate labels and associate component information with images, substituting human cognitive work with automated software systems.
3Stability of the object's composition
If intentionally introducing errors is performed to create balanced data, then dataset balance can be achieved, but additional time is required
Solution Approach 1:
The patent generates diverse training data by systematically varying rendering parameters such as camera angles, lighting conditions, component positions, and visual characteristics. This automated parameter variation creates balanced datasets representing different scenarios and error conditions without requiring manual intervention, achieving dataset balance through computational diversity generation.
4Reliability
If large labeled datasets are created, then AI algorithm performance improves, but time and resource consumption increase
Solution Approach 1:
The patent generates synthetic training data that replicates the visual characteristics and complexity of real electronic assemblies. These synthetic images provide sufficient training material for high-performance AI algorithms without requiring the time and resources needed to capture, process, and label large volumes of real images, thereby improving productivity while maintaining algorithm performance.
Data Source
Figure 1
Figure 2
Figure 3~6
AI summary
The invention relates to a computer-implemented method for providing a labelled training dataset, wherein - at least one sub-object (21, 22, 23, 24, 25, 26) is selected in a CAD model (1) of an object (2) comprising a plurality of sub-objects, - a plurality of different render images (45, 46, 47, 48) is generated, wherein the different render images (45, 46, 47, 48) contain the at least one selected sub-object (21, 22, 23, 24, 25, 26), - the different render images are labelled on the basis of the CAD model to provide a training dataset based on the labelled render images (49).