Analysis Device for Site-Specific Prediction Models Without Data Transfer
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning technologies face challenges in generating prediction models that are site-specific without the need for data transfer, particularly in scenarios where data confidentiality and resource constraints are concerns, and existing federated learning methods often require redundant relearning and may not achieve optimal prediction accuracy due to non-iid data distributions.
Innovation Solution
An analysis device that communicates with multiple learning devices to transform and analyze features locally, allowing for the generation of site-specific prediction models without data transfer, using a distribution analysis unit to determine appropriate learning methods based on similarity analysis of transformed features, thereby facilitating personalized or non-personalized federated learning.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If data is transferred from multiple sites to generate a prediction model, then prediction accuracy is improved, but data confidentiality is compromised
Solution Approach 1:
The patent introduces a server as an intermediary that collects transformed features from multiple learning devices without requiring transfer of raw confidential data. The server performs distribution analysis and model generation based on these transformed features, thereby achieving high prediction accuracy while preserving data confidentiality at each site.
Solution Approach 2:
The patent extracts only the necessary transformed features from the raw data at each learning device, rather than transferring the complete raw datasets. This extraction approach allows the server to generate accurate prediction models while minimizing the exposure of confidential information, as only processed feature representations are transmitted.
2Loss of information
If federated learning is performed without data transfer, then data confidentiality is maintained, but prediction accuracy deteriorates due to non-iid data distributions
Solution Approach 1:
The patent transforms the raw features into transformed features using predetermined transformation rules before transmission to the server. This parameter transformation enables the server to perform effective distribution analysis and generate accurate prediction models even when the underlying data distributions are non-iid, while maintaining data confidentiality through the transformation process.
Solution Approach 2:
The patent performs distribution analysis and model generation tailored to each learning device's specific data characteristics. The server determines appropriate learning methods based on the distribution analysis results, allowing each site to receive customized models that account for local data properties, thereby improving prediction accuracy for non-iid distributions while maintaining data privacy.
3Measurement precision
If learning models are generated for each site individually, then site-specific prediction accuracy is improved, but resource consumption increases
Solution Approach 1:
The patent merges the learning processes by having multiple learning devices transmit their transformed features to a central server that performs distribution analysis and generates prediction models. This centralized approach consolidates computational resources, reducing the overall resource consumption compared to each device performing independent learning, while still generating site-specific accurate models through the distribution analysis.
Solution Approach 2:
The server performs multiple functions including collecting transformed features, performing distribution analysis, determining learning methods, and generating prediction models for multiple learning devices. This multi-functional approach eliminates the need for each device to independently perform all these functions, thereby reducing redundant resource consumption while achieving site-specific model generation.
4Measurement precision
If transformed features are collected from multiple learning devices, then distribution analysis accuracy is improved, but communication overhead increases
Solution Approach 1:
The patent extracts only the essential transformed features needed for distribution analysis from the raw data at each learning device, rather than transmitting complete datasets. This selective extraction reduces communication overhead while providing sufficient information for accurate distribution analysis and model generation at the server.
Data Source
AI summary
An object of the present invention is to achieve generation of a prediction model appropriate for each site without a necessity of transfer of data located at a plurality of sites to the outside of the sites.An analysis device capable of communicating with a plurality of learning devices includes a reception unit (301, 401, 1501) that receives transformed features obtained by transforming, in accordance with a predetermined rule, features contained in pieces of learning data individually retained in the plurality of learning devices, a distribution analysis unit (302) that analyzes distributions of a plurality of the features of the plurality of learning devices on the basis of the transformed features received by the reception unit (301, 401, 1501) for each of the learning devices, and an output unit (304, 1504) that outputs a distribution analysis result analyzed by the distribution analysis unit (302).


