Synthetic Document Generation for API Testing Compliance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data privacy regulations, such as GDPR, restrict the processing of personal data, making it challenging for service providers to analyze and test application programming interfaces (APIs) without violating user privacy, especially when dealing with large volumes of similar documents, which can be inefficient and resource-intensive.

Innovation Solution

Generating encoded documents that retain structural information but remove personal data, allowing for analysis and regression testing using synthetic documents that comply with data regulations, enabling machine learning analysis and resource-efficient processing while maintaining compliance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If personal data is processed for API analysis and testing, then testing completeness and accuracy are improved, but data privacy compliance deteriorates

Engineering Contradiction:
Improvetesting completenessVSAvoidprivacy violation
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic documents that are copies of real user documents but with personally identifiable information replaced by synthetic data. These synthetic documents maintain the same structural characteristics, field relationships, and data patterns as real documents, enabling comprehensive API testing without using actual personal data. This resolves the contradiction by providing testing completeness through realistic document structures while ensuring privacy compliance through synthetic data generation.

Inventive Principle:
Principle #26Copying

2Measurement precision

If real user documents are used for regression testing, then testing accuracy is improved, but resource consumption increases

Engineering Contradiction:
Improvetesting accuracyVSAvoidresource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system generates synthetic documents that replicate the structural and relational characteristics of real user documents without copying the actual large-volume data. By creating simplified synthetic versions with the same schema, field types, and relationships, the system achieves accurate regression testing while significantly reducing computational resources, storage requirements, and processing time compared to using actual user documents.

Inventive Principle:
Principle #26Copying

3Object-affected harmful factors

If personal data is removed from documents, then data privacy compliance is improved, but information utility deteriorates

Engineering Contradiction:
Improveprivacy complianceVSAvoiddata utility
Core Design Contradiction:
Object-affected harmful factorsVSLoss of information

Solution Approach 1:

The patent removes personally identifiable information from real documents and replaces it with synthetic data that preserves the document structure, field relationships, and data patterns. The synthetic documents maintain all necessary information for API testing including document schemas, field types, validation rules, and relational structures, while ensuring complete privacy compliance. This approach retains full information utility for testing purposes without compromising user privacy.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12079284B2Generating data regulation compliant data from application interface data
Publication Date: 2024.09.03 SAP SE
  • US12079284B2 patent drawing
  • US12079284B2 patent drawing
  • US12079284B2 patent drawing

AI summary

The present disclosure involves systems, software, and computer-implemented methods for generating data regulation-compliant data from application interface data. One example method includes receiving a request for creation of document data. The request includes personal data of a user. Document data, including at least some of the personal data, is created based on the request. The document data is encoded into an encoded document that does not include any personal data of the user and includes structural information that describes the structure of the document data. A request to use the encoded document is received and the encoded document is decoded. A synthetic document is generated using the structural information included in the encoded document. Generation of the synthetic document includes insertion of synthetic user data into the synthetic document at positions in the synthetic document that correspond to positions of personal data within the document data.