Gradient Transfer Noise Masking for Imbalanced Data Security
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data protection methods for machine learning models in imbalanced binary classification tasks face challenges in ensuring data security and consistency of gradient transfer information, leading to potential data security risks due to differentiation in gradient-related information from positive and negative samples.
Innovation Solution
A method that acquires gradient correlation information for target and reference samples within the same batch, generates data noise based on the comparison of this information, and corrects the initial gradient transfer value to ensure consistency of gradient transfer information across different categories, thereby ensuring data security by adjusting the joint training model parameters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If gradient correlation information is directly transferred without correction, then data processing efficiency is improved, but data security deteriorates due to differentiation in gradient-related information from positive and negative samples
Solution Approach 1:
The patent introduces noise as an intermediary element to mask the gradient correlation information. By adding noise to the gradient transfer values, the system prevents direct differentiation between positive and negative samples while still allowing the passive party to receive useful training signals. This mediator approach resolves the contradiction by enabling efficient data processing without compromising data security.
Solution Approach 2:
The patent modifies the gradient transfer values by changing their parameters through noise addition. Instead of transferring raw gradient information directly, the system transforms the parameters by adding controlled noise, which obscures the original gradient correlations while preserving the essential training information needed for model improvement.
2Reliability
If noise is added to gradient transfer information to protect privacy, then data security is improved, but training effect deteriorates due to inconsistency in gradient transfer information across different categories
Solution Approach 1:
The patent applies different noise characteristics to different gradient transfer values based on their local properties. By analyzing the gradient correlation information and applying appropriate noise levels or types to specific samples or features, the system maintains data security while preserving the essential training signals needed for effective model training.
3Manufacturing precision
If gradient transfer information is corrected for consistency, then training effect is improved, but data security deteriorates due to potential exposure of sample category information
Solution Approach 1:
The patent converts the potentially harmful differentiation in gradient information into a beneficial training signal. By adding noise that masks category information while preserving gradient direction and magnitude patterns, the system transforms what would be a security risk into a useful feature for robust model training that generalizes better across different data distributions.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Disclosed are a data protection method and apparatus, and a server and a medium. A particular embodiment of the method comprises: acquiring gradient associated information, which respectively corresponds to a target sample that belongs to a binary classification sample set with unbalanced distribution and a reference sample that belongs to the same batch as the target sample; generating information of data noise to be added; according to the information of said data noise, correcting an initial gradient transfer value corresponding to the target sample, such that corrected gradient transfer information corresponding to samples in the sample set that belong to different types is consistent; and sending the gradient transfer information to a passive party of a joint training model. By means of the embodiment, there is no significant difference between corrected gradient transfer information corresponding to positive and negative samples, thereby effectively protecting the security of data.