Synthetic Data Generation via Parametric Node Relationships

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data capture and processing techniques face challenges in efficiently generating parametric representations of objects and their relationships for applications like computer vision, as they require manual specification of objects and complex queries, which can be time-consuming and inefficient.

Innovation Solution

A method and system for generating parametric representations by querying a database of nodes, evaluating relationships, and assigning semantic categories, allowing for the extraction of parametric representations and the creation of synthetic datasets that include elements of specific semantic categories, thereby simplifying the process of object selection and dataset generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual specification of objects and complex queries are used for data capture, then object selection accuracy can be maintained, but time consumption and processing efficiency deteriorate

Engineering Contradiction:
Improveobject selection accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent creates synthetic data copies that replicate the statistical properties and relationships of real data without requiring manual specification. By generating synthetic datasets that mirror the structure and semantics of source data, the system achieves accurate object selection automatically, eliminating time-consuming manual queries while maintaining measurement precision through faithful reproduction of data characteristics

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system enables self-service data generation where the synthetic data creation process automatically queries and processes source data without human intervention. The framework autonomously performs data capture, relationship extraction, and synthetic instance generation, allowing the system to serve itself in producing training datasets without requiring manual object specification or complex user queries

Inventive Principle:
Principle #25Self-service

2Reliability

If photorealistic image rendering is used to create realistic 3D representations, then visual realism is improved, but computational complexity and processing time deteriorate

Engineering Contradiction:
Improvevisual realismVSAvoidcomputational complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts only the essential parametric properties and semantic relationships needed for AI training from complex photorealistic images, separating these critical features from the computationally intensive visual rendering components. By taking out and preserving only the necessary data elements (object parameters, relationships, semantics) while discarding redundant visual complexity, the system achieves reliable training data generation without requiring full photorealistic rendering

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of investing heavy computational resources in creating permanent photorealistic 3D models, the system generates disposable synthetic data instances that are sufficient for training purposes. These simplified synthetic representations serve their purpose effectively for AI model training without the need for expensive, long-term 3D asset creation and maintenance

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Quantity of substance

If comprehensive data capture of all objects and relationships is performed, then dataset completeness is improved, but data processing complexity and resource requirements deteriorate

Engineering Contradiction:
Improvedataset completenessVSAvoiddata processing complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent segments the data capture and processing task into distinct modular components: data source identification, relationship extraction, parametric property extraction, and synthetic instance generation. By dividing the comprehensive data capture process into these manageable segments, the system can process complete datasets through a structured pipeline that reduces overall complexity while maintaining data completeness through systematic coverage of all necessary elements

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from processing data in traditional dimensional space to parametric space, where objects are represented by their key properties and relationships. This dimensional transformation allows comprehensive data capture to be achieved more efficiently by working with condensed parametric representations rather than full-dimensional raw data, reducing processing complexity while preserving completeness through faithful parametric modeling

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20240428497A1Systems and methods for generating synthetic data from a database of nodes and relationships
Publication Date: 2024.12.26 LEXSET AI LLC
  • US20240428497A1 patent drawing
  • US20240428497A1 patent drawing
  • US20240428497A1 patent drawing

AI summary

Embodiments of the present disclosure include methods and systems for generating a parametric representation and a synthetic dataset. In some embodiments, a database of nodes may be constructed. The database of nodes may include a plurality of nodes, wherein each node may represent an object associated with one or more parametric properties. Embodiments include evaluating a set of relationships for at least one subset of nodes. A method may include assigning a semantic category to at least one node and extracting the parametric representation. In some embodiments, a first node of the plurality of nodes is selected based at least in part on the semantic category, and a second node is selected based at least in part on a relationship of the second node with the first node. Embodiments may include generating a synthetic dataset based on the parametric representation or a semantic category that links nodes.