Image Segmentation for Privacy-Safe Data Labeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Neural network models require large amounts of training data labeled by workers, which may expose sensitive personal information during the labeling process.
Innovation Solution
A method for processing data that involves segmenting original images into regions, generating processing target images by combining these regions, and providing them to a data processor terminal while masking sensitive information to prevent exposure.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If original data is used for labeling to improve training data quality, then neural network model accuracy is improved, but personal information exposure risk increases
Solution Approach 1:
The patent segments original images into multiple regions (e.g., first region, second region) and generates processing target images by combining partial regions from different original images. This segmentation prevents workers from viewing complete original images containing personal information while still providing sufficient visual data for accurate labeling, thus resolving the contradiction between model accuracy and privacy protection.
Solution Approach 2:
The patent introduces processing target images as an intermediary between original images and workers. These intermediary images are generated by combining partial regions from multiple original images, allowing workers to perform labeling tasks without directly accessing original images containing personal information, thereby protecting privacy while maintaining labeling quality.
2Manufacturing precision
If original images are provided to workers for labeling, then labeling accuracy is improved, but security risk of personal information disclosure increases
Solution Approach 1:
The patent divides original images into multiple segmented regions and recombines them to create processing target images. This segmentation ensures that workers receive images with sufficient detail for accurate labeling while personal information in any single region cannot be traced back to the original image, thus maintaining both labeling accuracy and information security.
Solution Approach 2:
The patent creates processed copies of original images by combining partial regions from multiple sources. These copied and recombined images provide workers with adequate visual information for accurate labeling while being fundamentally different from any single original image, preventing reverse engineering to obtain personal information.
3Reliability
If data is processed to protect personal information, then privacy security is improved, but data utility for training may be reduced
Solution Approach 1:
The patent segments original images into multiple regions and recombines them to create processing target images that maintain sufficient visual information for effective model training. The segmentation and recombination process preserves essential features and patterns needed for learning while removing personal information, thus balancing privacy security with data utility.
Solution Approach 2:
The patent merges partial regions from multiple original images to create processing target images. This merging process maintains rich visual information and diversity needed for effective training while ensuring that no single original image's personal information is preserved, thereby maintaining data utility while protecting privacy.
Data Source
AI summary
Provided are a method of processing data for protecting personal information and an apparatus using the same. The method of processing data, which is a method for providing a data processor terminal with processing target data, includes acquiring a first original image and a second original image, segmenting the first original image into a plurality of regions including a first region, segmenting the second original image into a plurality of regions including a second region, generating a first processing target image including at least a partial region of the first region and at least a partial region of the second region, and providing the generated first processing target image to a data processor terminal.


