Multipath Mixing Learning Data Acquisition for Privacy and Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In distributed networks, training artificial neural networks is limited by the small amount of learning data generated by individual terminals, and existing methods for data exchange face challenges such as information leakage and reduced learning accuracy due to transmission capacity constraints.

Innovation Solution

A learning data acquisition apparatus and method that mixes data from multiple terminals based on specific ratios, classifies and re-mixes the data according to the number of terminals, and adjusts re-mixing ratios to generate re-mixed learning data for improved accuracy while preventing personal information leakage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If learning data is directly exchanged between terminals, then the amount of learning data increases, but personal information may be leaked

Engineering Contradiction:
Improveamount of learning dataVSAvoidpersonal information leakage
Core Design Contradiction:
Quantity of substanceVSObject-affected harmful factors

Solution Approach 1:

The patent introduces a server as an intermediary that receives mixed data from multiple terminals, performs re-mixing operations, and distributes the re-mixed data back to terminals. This intermediary structure allows data aggregation and processing without direct terminal-to-terminal data sharing, thereby preventing personal information leakage while still enabling increased learning data volume through centralized collection and redistribution.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If a learning model is exchanged between terminals, then information leakage is prevented, but the data transmission size becomes very large

Engineering Contradiction:
Improveinformation leakageVSAvoiddata transmission size
Core Design Contradiction:
Object-affected harmful factorsVSQuantity of substance

Solution Approach 1:

The patent extracts and transmits only the essential mixed data features rather than complete learning models. By separating the necessary data components from unnecessary model structures, the system prevents information leakage while significantly reducing transmission size compared to exchanging full learning models between terminals.

Inventive Principle:
Principle #2Taking out (Extraction)

3Object-affected harmful factors

If an output distribution of a learning model is exchanged, then information leakage is prevented and transmission size is reduced, but learning accuracy is not improved to required level

Engineering Contradiction:
Improveinformation leakageVSAvoidlearning accuracy
Core Design Contradiction:
Object-affected harmful factorsVSMeasurement precision

Solution Approach 1:

The patent merges data from multiple terminals through mixing and re-mixing operations performed by the server. This combining approach aggregates information from diverse sources to create enriched learning datasets that improve accuracy, while the server-mediated process maintains privacy protection. The merging of multiple data sources compensates for the limitations of exchanging only output distributions.

Inventive Principle:
Principle #5Merging (Combining)

4Object-affected harmful factors

If data mixing method is applied to prevent information leakage, then personal information protection is improved, but the amount of data increases or learning accuracy is lowered

Engineering Contradiction:
Improvepersonal information protectionVSAvoidamount of data
Core Design Contradiction:
Object-affected harmful factorsVSQuantity of substance

Solution Approach 1:

The patent implements dynamic mixing ratios where the server adjusts the proportion of data from different terminals based on data quality, diversity, and learning progress. This dynamic adjustment optimizes the balance between privacy protection and data efficiency, preventing unnecessary data accumulation while maintaining adequate learning accuracy through adaptive re-mixing strategies.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20220327426A1Multipath mixing-based learning data acquisition apparatus and method
Publication Date: 2022.10.13 IND ACADEMIC COOP FOUND YONSEI UNIV
  • US20220327426A1 patent drawing
  • US20220327426A1 patent drawing
  • US20220327426A1 patent drawing

AI summary

The present disclosure provides a learning data acquisition apparatus and method for receiving, from each of a plurality of terminals, mixed data in which a plurality of pieces of learning data are mixed according to a mixing ratio, identifying the mixed data transmitted from each of the plurality of terminals according to an included label, and acquire re-mixed learning data for training a pre-stored learning model by re-mixing each identified label according to a re-mixing ratio configured in correspondence to the number of terminals having transmitted the mixed data, thereby enabling learning performance and security to be improved by re-mixing the mixed data transmitted from each of the plurality of terminals in a data mixing manner.