Anonymizing Statistical Data for Secure Multi-Hospital Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data analysis methods fail to effectively preserve patient privacy while analyzing aggregated statistical data from multiple hospitals, as raw data disclosure is often unacceptable and current approaches do not adequately address privacy concerns.
Innovation Solution
A computer-implemented method that calculates statistical information, aggregates data to determine valid ranges, removes outlier data, creates pair lists with target data, replaces target data with random numbers within defined bins, and swaps pairs in random order for secure transfer and analysis, ensuring privacy preservation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If raw patient data is disclosed for analysis, then data analysis accuracy is improved, but patient privacy is compromised
Solution Approach 1:
The patent extracts and removes sensitive information from patient data through multiple processing steps including outlier removal based on valid ranges, anonymization by replacing identifiers with random numbers, and selective exclusion of sensitive attributes. This allows the essential statistical data to be retained for analysis while the harmful sensitive components are extracted and eliminated.
Solution Approach 2:
The patent introduces an intermediary processing system that acts as a mediator between raw patient data and analysis requirements. This system performs statistical validation, outlier detection, anonymization, and data transformation to create processed data that maintains analytical value while eliminating privacy risks. The intermediary processing layer enables both data utility and privacy protection simultaneously.
2Object-affected harmful factors
If statistical data is anonymized and processed, then patient privacy is preserved, but data analysis capability is reduced
Solution Approach 1:
The patent applies preliminary actions by pre-calculating statistical information, determining valid ranges, and identifying outliers before the actual data analysis. This preliminary processing preserves the statistical properties and relationships in the data while removing sensitive information, ensuring that the anonymized data retains sufficient capability for meaningful analysis.
Solution Approach 2:
The patent changes parameters by transforming raw data into processed data with modified characteristics - replacing sensitive identifiers with random numbers, adjusting data formats through statistical aggregation, and transforming individual records into validated statistical sets. These parameter changes maintain the essential statistical relationships needed for analysis while eliminating privacy risks.
3Object-affected harmful factors
If multiple data processing steps are applied, then privacy preservation is improved, but processing complexity increases
Solution Approach 1:
The patent segments the privacy preservation process into distinct modular steps: statistical information calculation, valid range determination, outlier detection and removal, anonymization processing, and final validation. Each segment performs a specific function and can be independently implemented or optimized. This segmentation makes the complex privacy preservation process more manageable and systematically applicable.
Data Source
AI summary
A method is provided for anonymizing statistical data for a secure transfer. The method calculates statistical information for each of the statistical data. The method aggregates the statistical information to calculate a valid range for each of the statistical information. The method removes outlier data based on the valid range for each of the statistical data. The method creates pair lists from each of the statistical data and target data, the pair lists having a respective member from both the statistical data and the target data. The method replaces each respective member of the target data by a random number existing in a range of a corresponding one of a plurality of target data bins. The method swaps each pair in each pair list in a random order using the randomized number, wherein the random number used for swapping is different for different ones of the pair lists.


