Synthetic Data Generation via Nearest Neighbor Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methodologies for generating synthetic data, such as multiple imputation, fail to create optimal datasets that accurately preserve the statistical properties of the original data while ensuring privacy, leading to inaccurate relations among variables and potential risks of tracing back original records.
Innovation Solution
The use of the nearest neighbor searching algorithm combined with propensity score methods to identify and replace data items in a way that maintains the statistical characteristics of the original dataset, ensuring computational efficiency and preventing record tracing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If multiple imputation methodology is used to generate synthetic data, then data synthesis is achieved, but statistical properties of original data are not accurately preserved and variable relations become inaccurate
Solution Approach 1:
The patent creates synthetic data records by copying and combining elements from existing real records in the database. Instead of generating entirely artificial data through statistical imputation, the system identifies similar real records and synthesizes new records by selecting and combining actual data elements from these source records, thereby preserving the authentic statistical properties and variable relationships of the original data.
2Object-affected harmful factors
If synthetic data is generated to protect privacy, then original records are protected, but risk of tracing back original records remains
Solution Approach 1:
The patent applies local quality by selectively combining different elements from multiple source records to create each synthetic record. Each field or attribute of the synthetic record may be drawn from different source records, creating a mosaic structure where no single source record can be easily identified. This local differentiation across multiple dimensions makes tracing back to original records significantly more difficult while preserving overall data utility.
3Productivity
If traditional imputation methods are used, then data synthesis is achieved, but computational efficiency is reduced and optimal datasets are not created
Solution Approach 1:
The patent performs preliminary actions by pre-identifying and categorizing source records based on their similarity characteristics before the actual synthesis process. The system pre-processes the database to organize records into groups or clusters that share similar properties, so that when synthetic records need to be generated, the selection and combination process can proceed efficiently using these pre-organized structures rather than searching through the entire database each time.
Data Source
AI summary
Systems and methods for constructing sets of synthetic data. A single data record is identified from a first set of data. The first set of data comprises a first plurality of data records, each of the data records including multiple items of data describing an entity. Using pattern recognition, the single data record is processed to identify a group of records from within the first set that have corresponding characteristics equivalent to the single data record. The identified group of records comprises a target set of variables and the group of records from the first set that are not identified comprises a control set of variables. The target set of variables and the control set of variables are processed, using probability estimation and optimization constraints, to determine a score for each of the records in the first set. The score describes how similar each of the records in the first set is to the single data record. The records associated with a percentage of the highest scores are identified. The data associated with the single data record is replaced with data associated with the identified records identified, item-by-item.


