Data Processing Device Noise Removal Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Differential privacy technologies face challenges in achieving high accuracy with small sample sizes due to increased errors in statistical results, particularly when data types are numerous, leading to inefficiencies in data utilization.
Innovation Solution
A data processing system and method that includes a noise removal unit, measurement unit, and data set updating unit to optimize the dictionary size based on data distribution, combining or dividing data types to reduce errors and improve accuracy, utilizing local type differential privacy to enhance reliability and accuracy of statistical results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a large quantity of data is collected to improve accuracy of statistical results, then measurement precision improves, but loss of information increases due to noise addition requirements
Solution Approach 1:
The patent segments data types into multiple categories and collects data separately for each category. This allows the system to reduce the amount of noise needed per category while maintaining overall statistical accuracy, as noise requirements are proportional to the number of data types. By dividing the data collection into manageable segments, the system achieves better precision without excessive noise addition.
Solution Approach 2:
The patent introduces a new dimension of data type classification to organize and manage collected data. By categorizing data into multiple types and handling each type separately in the collection process, the system optimizes noise addition requirements and improves statistical accuracy without requiring uniformly large sample sizes across all data.
2Adaptability or versatility
If data types are increased to improve data utilization, then adaptability improves, but device complexity increases due to dictionary management
Solution Approach 1:
The patent implements a dynamic dictionary management system that automatically adjusts data type categories based on collected data characteristics. The system can add, remove, or merge data types as needed, optimizing adaptability while managing complexity through automated procedures rather than static, manually-configured categories.
Solution Approach 2:
The system performs self-service dictionary management by automatically analyzing collected data and determining optimal data type categorizations. This reduces manual intervention and simplifies complexity management while maintaining high adaptability to different data collection scenarios.
3Reliability
If noise is added to protect privacy, then reliability of privacy protection improves, but measurement precision deteriorates due to increased error in statistical results
Solution Approach 1:
By segmenting data into multiple types and collecting them separately, the patent reduces the noise addition burden for each individual category. Since noise requirements scale with the number of data types, segmentation allows each category to receive appropriate noise levels that protect privacy while maintaining statistical accuracy for that specific category.
Solution Approach 2:
The patent changes the parameter of noise addition by applying different noise levels to different data types based on their specific characteristics and sensitivity requirements. This allows optimization of the balance between privacy protection and measurement precision for each data type rather than applying uniform noise across all data.
Data Source
AI summary
Provided is a data processing device including: a noise removal unit that removes noise from data to which noise has been added, the data having been received from a terminal device; a measurement unit that measures the data for each data type constituting a data set and indicating a classification of the data; and a data set updating unit that updates the data set on the basis of a measurement result of the measurement unit.


