Federated Learning Server Weighting Data Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing federated machine learning approaches suffer from inaccuracies and unreliability due to variations in data quality across multiple data sources, which are not adequately accounted for in current decentralized training methods.
Innovation Solution
A method of federated machine learning that transmits a global machine learning model to multiple data sources, receives training updates, and updates the model based on data quality parameters associated with each source, weighting the updates to account for differences in feature and label quality, thereby improving accuracy and reliability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If federated learning is used to train across multiple data sources without centralizing data, then data privacy and security are improved, but model accuracy and reliability deteriorate due to unaccounted data quality variations
Solution Approach 1:
The patent applies local quality by assigning different weight values to different data sources based on their individual data quality characteristics. Each data source is evaluated independently for features like label quality, feature quality, and data completeness, and weights are assigned accordingly. This allows the system to maintain data privacy while accounting for local quality variations, resolving the contradiction between privacy preservation and model accuracy.
Solution Approach 2:
The patent changes the parameter of data source weighting from uniform (equal weights) to variable (quality-based weights). By introducing data quality parameters such as label accuracy, feature quality metrics, and data completeness scores, the system dynamically adjusts the contribution of each data source to the federated learning process, thereby improving model reliability without compromising privacy.
2Device complexity
If uniform weighting is applied to all data sources in federated learning, then system complexity is reduced, but model reliability deteriorates due to unaccounted data quality differences
Solution Approach 1:
The patent applies preliminary action by performing data quality assessment and weight assignment before the federated learning training process begins. Data sources are evaluated for their quality metrics (label quality, feature quality, completeness) in advance, and weight values are predetermined based on these assessments. This preliminary characterization simplifies the ongoing training process while ensuring reliable model convergence, addressing the contradiction between system complexity and model reliability.
3Measurement precision
If data quality parameters are incorporated into federated learning updates, then model accuracy is improved, but communication overhead and processing complexity increase
Solution Approach 1:
The patent extracts data quality assessment from the main federated learning training loop and performs it separately as a preliminary step. Quality metrics such as label accuracy, feature quality, and data completeness are evaluated and weights are assigned outside the iterative training process. This extraction reduces the communication overhead and processing time during actual training, while still incorporating quality information to improve model accuracy.
Data Source
AI summary
There is provided a method of federated machine learning using at least one processor, the method including: transmitting a current global machine learning model to each of a plurality of data sources; receiving a plurality of training updates from the plurality of data sources, respectively, each of the plurality of training updates being generated by the respective data source in response to the global machine learning model received; and updating the current global machine learning model based on the plurality of training updates received and a plurality of data quality parameters associated with the plurality of data sources, respectively, to generate an updated global machine learning model. There is also provided a corresponding server for federated machine learning.


