Synthetic Document Generation for API Testing Compliance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data privacy regulations, such as GDPR, restrict the processing of personal data, making it challenging for service providers to analyze and test application programming interfaces (APIs) without violating user privacy, especially when dealing with large volumes of similar documents, which can be inefficient and resource-intensive.
Innovation Solution
Generating encoded documents that retain structural information but remove personal data, allowing for analysis and regression testing using synthetic documents that comply with data regulations, enabling machine learning analysis and resource-efficient processing while maintaining compliance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If personal data is processed for API analysis and testing, then testing completeness and accuracy are improved, but data privacy compliance deteriorates
Solution Approach 1:
The patent creates synthetic documents that are copies of real user documents but with personally identifiable information replaced by synthetic data. These synthetic documents maintain the same structural characteristics, field relationships, and data patterns as real documents, enabling comprehensive API testing without using actual personal data. This resolves the contradiction by providing testing completeness through realistic document structures while ensuring privacy compliance through synthetic data generation.
2Measurement precision
If real user documents are used for regression testing, then testing accuracy is improved, but resource consumption increases
Solution Approach 1:
The system generates synthetic documents that replicate the structural and relational characteristics of real user documents without copying the actual large-volume data. By creating simplified synthetic versions with the same schema, field types, and relationships, the system achieves accurate regression testing while significantly reducing computational resources, storage requirements, and processing time compared to using actual user documents.
3Object-affected harmful factors
If personal data is removed from documents, then data privacy compliance is improved, but information utility deteriorates
Solution Approach 1:
The patent removes personally identifiable information from real documents and replaces it with synthetic data that preserves the document structure, field relationships, and data patterns. The synthetic documents maintain all necessary information for API testing including document schemas, field types, validation rules, and relational structures, while ensuring complete privacy compliance. This approach retains full information utility for testing purposes without compromising user privacy.
Data Source
AI summary
The present disclosure involves systems, software, and computer-implemented methods for generating data regulation-compliant data from application interface data. One example method includes receiving a request for creation of document data. The request includes personal data of a user. Document data, including at least some of the personal data, is created based on the request. The document data is encoded into an encoded document that does not include any personal data of the user and includes structural information that describes the structure of the document data. A request to use the encoded document is received and the encoded document is decoded. A synthetic document is generated using the structural information included in the encoded document. Generation of the synthetic document includes insertion of synthetic user data into the synthetic document at positions in the synthetic document that correspond to positions of personal data within the document data.


