Synthetic Dataset Generation for Network Equipment Configuration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The availability of relevant data for configuring, training, and monitoring systems is limited due to privacy, contractual, and regulatory restrictions, making it difficult to use actual data from production environments in development or test environments.
Innovation Solution
The use of synthetically generated datasets using AI/ML techniques to mirror the properties of actual data, allowing for the configuration, training, and modification of systems without directly using the actual data, thus overcoming usage restrictions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If actual data from production environments is used in development and test environments, then system configuration, training, and performance determination can be performed with real-world data, but privacy, contractual, and regulatory restrictions prevent the availability and usage of such data
Solution Approach 1:
The patent generates synthetic data that copies the statistical properties, patterns, and characteristics of actual production data without containing the actual sensitive information. This synthetic copy enables system training and testing while avoiding privacy and regulatory issues associated with using real customer data.
Solution Approach 2:
Synthetic data acts as an intermediary between the need for real-world data in system training and the restrictions on using actual production data. It mediates by providing a proxy that preserves the essential characteristics needed for system development while eliminating the legal and privacy concerns of using real data.
2Loss of information
If synthetic datasets are generated to overcome data usage restrictions, then data availability for system training and configuration is improved, but the accuracy and representativeness of the data may be compromised
Solution Approach 1:
The system adjusts parameters of the synthetic data generation process to optimize the balance between data availability and accuracy. By controlling parameters such as the degree of synthetic transformation and the preservation of statistical properties, the system maintains data representativeness while ensuring compliance with privacy requirements.
Data Source
AI summary
A configuration system may generate synthetic datasets based on machine learning, and may provide the synthetic datasets in order to configure, train, test, and/or otherwise modify operation of network equipment and/or other devices. The configuration system may receive a sampling of data used by the particular network equipment, and may determine parameters that define values for the variables represented in the sampling. The configuration system may generate different datasets based on the parameters, and may determine an accuracy of each dataset to the sampling based on the values of each dataset conforming by different amounts to the parameters that define the values for the variables represented in the sampling by different amounts. The configuration system may modify the network equipment operation using the values from a particular dataset in response to determining that the accuracy of the particular dataset is greatest to the sampling.


