Multipath Mixing Learning Data Acquisition for Privacy and Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In distributed networks, training artificial neural networks is limited by the small amount of learning data generated by individual terminals, and existing methods for data exchange face challenges such as information leakage and reduced learning accuracy due to transmission capacity constraints.
Innovation Solution
A learning data acquisition apparatus and method that mixes data from multiple terminals based on specific ratios, classifies and re-mixes the data according to the number of terminals, and adjusts re-mixing ratios to generate re-mixed learning data for improved accuracy while preventing personal information leakage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If learning data is directly exchanged between terminals, then the amount of learning data increases, but personal information may be leaked
Solution Approach 1:
The patent introduces a server as an intermediary that receives mixed data from multiple terminals, performs re-mixing operations, and distributes the re-mixed data back to terminals. This intermediary structure allows data aggregation and processing without direct terminal-to-terminal data sharing, thereby preventing personal information leakage while still enabling increased learning data volume through centralized collection and redistribution.
2Object-affected harmful factors
If a learning model is exchanged between terminals, then information leakage is prevented, but the data transmission size becomes very large
Solution Approach 1:
The patent extracts and transmits only the essential mixed data features rather than complete learning models. By separating the necessary data components from unnecessary model structures, the system prevents information leakage while significantly reducing transmission size compared to exchanging full learning models between terminals.
3Object-affected harmful factors
If an output distribution of a learning model is exchanged, then information leakage is prevented and transmission size is reduced, but learning accuracy is not improved to required level
Solution Approach 1:
The patent merges data from multiple terminals through mixing and re-mixing operations performed by the server. This combining approach aggregates information from diverse sources to create enriched learning datasets that improve accuracy, while the server-mediated process maintains privacy protection. The merging of multiple data sources compensates for the limitations of exchanging only output distributions.
4Object-affected harmful factors
If data mixing method is applied to prevent information leakage, then personal information protection is improved, but the amount of data increases or learning accuracy is lowered
Solution Approach 1:
The patent implements dynamic mixing ratios where the server adjusts the proportion of data from different terminals based on data quality, diversity, and learning progress. This dynamic adjustment optimizes the balance between privacy protection and data efficiency, preventing unnecessary data accumulation while maintaining adequate learning accuracy through adaptive re-mixing strategies.
Data Source
AI summary
The present disclosure provides a learning data acquisition apparatus and method for receiving, from each of a plurality of terminals, mixed data in which a plurality of pieces of learning data are mixed according to a mixing ratio, identifying the mixed data transmitted from each of the plurality of terminals according to an included label, and acquire re-mixed learning data for training a pre-stored learning model by re-mixing each identified label according to a re-mixing ratio configured in correspondence to the number of terminals having transmitted the mixed data, thereby enabling learning performance and security to be improved by re-mixing the mixed data transmitted from each of the plurality of terminals in a data mixing manner.


