LLM Synthetic Test Data Generation for Privacy and Edge Cases
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for generating synthetic test data are inefficient, time-consuming, and lack representation of real-world conditions, often exposing sensitive information and failing to cover edge cases, leading to gaps in coverage and security vulnerabilities.
Innovation Solution
Utilizing generative AI models, particularly large language models (LLMs), to generate tailored synthetic test data that mimics real-world data structures and patterns while ensuring privacy, coverage, and control, by receiving user input, processing task properties, forming training datasets, and configuring data generation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional manual methods are used to generate synthetic test data, then the process is simple to implement, but it is time-consuming and inefficient
Solution Approach 1:
The patent replaces manual mechanical data generation processes with an automated generative AI system. The system uses machine learning models to automatically generate synthetic test data based on input parameters and templates, eliminating the need for manual creation of test datasets while significantly improving efficiency and reducing time requirements.
2Reliability
If real customer data is used for testing, then the data is representative of real-world conditions, but it exposes sensitive information and creates security vulnerabilities
Solution Approach 1:
The patent creates synthetic copies of real-world data that preserve the statistical properties, patterns, and relationships of actual customer data without containing any real sensitive information. The generative AI system learns from real data distributions and generates artificial datasets that mimic real-world conditions while being completely anonymized, thus maintaining reliability without security risks.
3Adaptability or versatility
If manual simulations using available tools are used, then the implementation is straightforward, but they fail to cover edge cases and lack comprehensive representation
Solution Approach 1:
The patent implements a dynamic data generation system that can adapt to different testing requirements, data types, and scenario complexities. The system uses configurable templates and parameters that allow it to generate diverse test data including edge cases, rare events, and various attack scenarios. This dynamic approach enables comprehensive coverage without requiring complex manual configuration for each test case.
4Adaptability or versatility
If academic datasets are used for testing, then external data sources are available, but they become obsolete over time and may not align with specific user needs
Solution Approach 1:
The patent implements a self-updating generative AI system that can continuously learn from new data sources and adapt to changing requirements. The system allows users to configure specific parameters, data distributions, and scenario requirements, and it automatically generates current, relevant test data without requiring external updates. This self-service capability ensures the test data remains current and aligned with specific user needs indefinitely.
Data Source
AI summary
Systems and methods for generating synthetic test data for testing a software solution. Systems and methods include receiving a testing task from a user, identifying test properties of the testing task, gathering initial information based on the test properties of the testing task, forming a training dataset based on the initial information and the test properties, pretraining a generative AI model based on a large language model (LLM) using the training dataset, configuring synthetic test data based on the test properties, and generating synthetic test data according to the testing task using the generative AI model.


