Synthetic Data Generation Using Category Attribute Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for generating synthetic data with category attributes are inefficient as they require generating synthetic data for each set of categories, leading to decreased calculation efficiency with increasing sets of categories.
Innovation Solution
A synthetic data generation apparatus and method that codes category attributes into numerical attributes using a coding rule, generates synthetic data using numerical attribute methods, converts out-of-range numerical values, and decodes the numerical values back to category attributes, maintaining attribute relationships with improved efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If synthetic data is generated for each set of categories using conventional methods, then the relationships among attributes are maintained, but the calculation efficiency decreases as the number of category sets increases
Solution Approach 1:
The patent introduces a coding unit as an intermediary that converts category attributes into numerical attributes before synthetic data generation. This mediator enables the use of efficient numerical attribute generation methods while preserving category attribute relationships through the coding/decoding process, thereby maintaining reliability without sacrificing productivity
Solution Approach 2:
The patent changes the parameter representation of category attributes from categorical values to numerical codes through the coding unit. This parameter transformation allows the synthetic data generation to operate on numerical attributes with high efficiency, while the decoding unit restores the original category values, thus resolving the contradiction between efficiency and relationship maintenance
2Productivity
If category attributes are coded into numerical attributes, then the synthetic data generation efficiency improves, but the complexity of the generation process increases due to additional coding and decoding steps
Solution Approach 1:
The patent segments the synthetic data generation process into distinct functional modules: a coding unit for converting category attributes to numerical attributes, a data formatting unit for generating synthetic data from the coded numerical attributes, and a decoding unit for converting back to category attributes. This segmentation improves efficiency by allowing specialized optimization of each module while managing overall complexity through modular architecture
Data Source
AI summary
A synthetic data generation apparatus codes a value of each of category attributes contained in original data into a value of a numerical attribute in accordance with a coding rule; generates first synthetic data from the original data after coding using a synthetic data generation method for numerical attributes; if the value of the numerical attribute which is contained in the first synthetic data and corresponds to the value of one of the category attributes exceeds a range of values that can be assumed by the value of that numerical attribute, converts the value of that numerical attribute to a value included in the range of values that can be assumed by the value of that numerical attribute; and decodes the value of the numerical attribute which is contained in the first synthetic data after conversion and corresponds to the value of one of the category attributes to the value of that category attribute in accordance with the coding rule to obtain synthetic data.


