Synthetic Data Mirroring for Privacy-Safe ML Testing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data sharing, particularly cross-border data sharing for research and development, is hindered by privacy and security concerns, making it difficult to obtain high-quality testing data for machine learning applications.

Innovation Solution

A method and system that generate synthetic data using machine learning models to replicate the features and structural characteristics of real data, allowing secure and compliant data sharing without exposing sensitive information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If real data is shared for research and development purposes, then the quality of testing data improves, but privacy and security concerns worsen

Engineering Contradiction:
Improvequality of testing dataVSAvoidprivacy and security concerns
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic data that copies the statistical properties, distributions, and relationships of real data without containing actual sensitive information. Machine learning models generate artificial datasets that replicate the structural characteristics and feature correlations of source data, enabling research and development while eliminating privacy and security risks associated with sharing real data.

Inventive Principle:
Principle #26Copying

2Reliability

If data localization regulations are implemented, then data security improves, but the ability to share data with external partners deteriorates

Engineering Contradiction:
Improvedata securityVSAvoidability to share data
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent introduces synthetic data as an intermediary that mediates between data security requirements and data sharing needs. Instead of sharing real data directly, organizations can share synthetic data with external partners, researchers, and developers. This intermediary form maintains security compliance while enabling collaboration and development activities that would otherwise be restricted by data localization regulations.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If synthetic data is generated using machine learning models, then data security improves, but the complexity of the data generation process worsens

Engineering Contradiction:
Improvedata securityVSAvoidcomplexity of data generation process
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements self-service capabilities where the synthetic data generation system automatically analyzes source data characteristics, selects appropriate generation methods, and produces synthetic datasets without requiring extensive manual configuration or expert intervention. The system autonomously handles data profiling, model selection, parameter optimization, and quality validation, reducing operational complexity while maintaining security benefits.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250355898A1Data mirror
Publication Date: 2025.11.20 HSBC SOFTWARE DEV (GUANGDONG) LTD
  • US20250355898A1 patent drawing
  • US20250355898A1 patent drawing
  • US20250355898A1 patent drawing

AI summary

At least one processor may receive a sample data set, determine at least one feature of data in the sample data set, and determine at least one structural characteristic of the sample data set. The at least one processor may determine that at least a portion of the data is categorical data from the at least one feature and the at least one structural characteristic. By operating a machine learning (ML) model, the at least one processor may generate synthetic data having the same at least one feature as the categorical data. The at least one processor may package the synthetic data into a synthetic data set having the same at least one feature and at least one structural characteristic as the sample data set.