Privacy-Preserving Federated Data Normalization via Ontology Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In federated learning systems, clients face challenges in forming a common data encoding without a prior global data schema, leading to potential data privacy breaches and difficulties in training a global machine learning model.

Innovation Solution

A method and system for automatically generating a local data ontology, creating synthetic data, and computing risk and utility scores to transform client data to a common normalization schema using ontology matching algorithms, ensuring data privacy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If clients share raw data for training global machine learning models, then model training accuracy is improved, but data privacy is compromised

Engineering Contradiction:
Improvemodel training accuracyVSAvoiddata privacy breach
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent generates synthetic data that copies the statistical properties and patterns of raw client data without replicating actual sensitive information. This synthetic data is then shared across the federated learning system, enabling model training while preserving the privacy of original data through the use of generated replicas rather than actual data copies

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces synthetic data as an intermediary between raw client data and the global model training process. This intermediary element allows information to be transmitted for training purposes while blocking direct access to sensitive raw data, thus mediating between the need for accurate training and data privacy protection

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If a global data schema is established before data collection, then data integration efficiency is improved, but adaptability to new data sources is reduced

Engineering Contradiction:
Improvedata integration efficiencyVSAvoidadaptability to new data sources
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent implements a dynamic schema evolution mechanism where the global data schema is not fixed but can adapt and evolve as new data sources are introduced. The system automatically adjusts the schema to accommodate new data types and formats while maintaining compatibility with existing data, allowing the system to be both efficient and adaptable

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent segments the global data schema into modular components that can be independently configured and adapted. This segmentation allows different parts of the schema to be optimized for specific data sources while maintaining overall integration, enabling both efficient processing and flexibility for new data types

Inventive Principle:
Principle #1Segmentation

3Manufacturing precision

If data transformations are manually configured for each client, then data quality control is improved, but system complexity and implementation time increase

Engineering Contradiction:
Improvedata quality controlVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent implements self-service data transformation where the system automatically generates and applies transformation rules based on the global data schema and client data characteristics. Rather than requiring manual configuration, the system autonomously performs data quality control transformations, reducing complexity while maintaining precision through automated schema-based validation and transformation generation

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250284727A1Data quality assessment and transformation in a privacy preserving federated system
Publication Date: 2025.09.11 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20250284727A1 patent drawing
  • US20250284727A1 patent drawing
  • US20250284727A1 patent drawing

AI summary

A computer-implemented method for automatically transforming client data to a common data normalization schema associated with a collaborative multi-client federated learning system while preserving data privacy. The method may include automatically generating a local data ontology based on the client data associated with a client, and automatically generating synthetic data based on the client data and the local data ontology. The method may also include automatically computing an inference risk score comprising determining a privacy risk associated with sharing the synthetic data, and automatically computing a task utility score comprising determining a utility of the synthetic data. The method may further include generating a global data ontology using ontology matching algorithms on the synthetic data associated with each local data ontology. The method may also include automatically recommending and implementing data transformations to the client data.