Federated Learning Model Update for Small Medical Facilities

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Small medical facilities face challenges in generating accurate models due to insufficient training data, which hampers their ability to effectively learn and improve their local models in distributed learning systems.

Innovation Solution

A learning system that includes a central server and multiple sites, where sites with larger data sets, like candidate sites, assist smaller sites by updating their models based on data distributions, selecting the most similar data sets from larger cohorts to enhance model accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If distributed learning is used to protect patient data confidentiality, then data security is improved, but model accuracy deteriorates due to insufficient training data at individual sites

Engineering Contradiction:
Improvedata securityVSAvoidmodel accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent combines data from multiple sites through federated learning to improve model accuracy while maintaining data security. The server aggregates models from multiple client sites, effectively merging the computational benefits of distributed data without physically combining the sensitive data itself, thus resolving the contradiction between data security and model accuracy.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The learned model is designed to be universally applicable across different medical sites. By training a generalizable model that can function effectively at multiple sites with varying data volumes, the system achieves both data security (through distributed learning) and acceptable model accuracy (through universal applicability).

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If local models are trained at each client site, then data confidentiality is maintained, but model performance deteriorates when training data is insufficient

Engineering Contradiction:
Improvedata confidentialityVSAvoidmodel performance
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The server acts as an intermediary that facilitates model improvement without direct access to client data. It coordinates the federated learning process, aggregates models from multiple clients, and redistributes improved models, enabling performance enhancement while maintaining the confidentiality barrier between client sites and the server.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

Multiple local models from different client sites are merged through aggregation at the server level. This combining of models from sites with different data characteristics compensates for individual data insufficiencies and improves overall model performance while each site maintains its data confidentiality.

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If more training data is collected at small medical facilities, then model accuracy would improve, but data collection capability worsens due to limited patient volume

Engineering Contradiction:
Improvemodel accuracyVSAvoidtraining data volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent transitions from a single-site data collection approach to a multi-site collaborative approach. Instead of small facilities trying to collect more data locally (one dimension), the system leverages data from multiple sites through federated learning (adding another dimension), thereby achieving sufficient training data volume without requiring individual sites to increase their patient volume.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The overall training data requirement is segmented across multiple client sites rather than requiring one site to provide all data. Each site contributes a portion of the training process through its local model updates, and the aggregation of these segmented contributions achieves the equivalent of having large-volume training data without concentrating it at a single location.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20230020543A1Learning system, learning device, learning method, and storage medium
Publication Date: 2023.01.19 CANON KK
  • US20230020543A1 patent drawing
  • US20230020543A1 patent drawing
  • US20230020543A1 patent drawing

AI summary

A learning system includes processing circuitry. The processing circuitry is configured to acquire a first data distribution for a first data set out of data sets based on a first cohort, to select a second cohort that is used to update a first model out of a plurality of second cohorts on the basis of the acquired first data distribution, and to update the first model on the basis of at least part of a second data set out of data sets based on the selected second cohort.