Semantic Data Linking for Privacy Risk Assessment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data privacy management techniques struggle to effectively protect individual privacy preferences and personally identifiable information while relying on data aggregation, which often involves linking information across multiple datasets.

Innovation Solution

A method that receives a target dataset, determines semantic representations of attributes, and iteratively augments the dataset by identifying and linking with auxiliary datasets to fill gaps in information, ultimately enhancing data privacy risk assessment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data aggregation techniques are used to protect individual privacy preferences and personally identifiable information, then data privacy protection is improved, but data linking across multiple datasets becomes more complex and time-consuming

Engineering Contradiction:
Improvedata privacy protectionVSAvoiddata linking complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces semantic representations as an intermediary layer between raw data and privacy protection mechanisms. These semantic representations serve as mediators that enable data linking across multiple datasets through meaningful associations rather than direct data aggregation, reducing the complexity of privacy-preserving data linkage operations

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms data from its original form into semantic representations with different properties. By changing the parameter space from raw data values to semantic concepts and relationships, the system enables more efficient and less complex data linking while maintaining privacy protection capabilities

Inventive Principle:
Principle #35Parameter changes

2Reliability

If data aggregation is used to protect privacy, then privacy preferences are better protected, but the time required for data processing and analysis increases

Engineering Contradiction:
Improveprivacy preference protectionVSAvoiddata processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary computation of semantic representations and their relationships before actual data linking operations. By pre-computing semantic associations and indexing them, the system reduces the time required for subsequent data aggregation and privacy analysis operations

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Semantic representations act as intermediaries that enable faster data linking by pre-establishing meaningful relationships between data elements. This intermediary layer eliminates the need for time-consuming direct comparison of raw data across multiple datasets while maintaining privacy protection

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If semantic representations are used to link data across datasets, then data privacy risk assessment accuracy is improved, but computational resources and processing complexity increase

Engineering Contradiction:
Improveprivacy risk assessment accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies partial computation by selecting and processing only the semantic representations and data relationships necessary for privacy risk assessment, rather than computing all possible semantic associations across entire datasets. This partial action approach maintains assessment accuracy while reducing computational resource consumption

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12321483B2Augmented privacy datasets using semantic based data linking
Publication Date: 2025.06.03 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12321483B2 patent drawing
  • US12321483B2 patent drawing
  • US12321483B2 patent drawing

AI summary

Disclosed are techniques for linking information about individual entities across multiple datasets. A target dataset with some information corresponding to at least one attribute of an entity is received. Semantic processing is performed on the target dataset to extract semantic representations of the information and corresponding attributes, which is utilized to search at least one other dataset for additional information that is absent from the target dataset, corresponding to at least one attribute of the entity, which are used to augment the target dataset with additional information corresponding to the entity. This is repeated iteratively, with each subsequent iteration including semantic representations of information found in the searches of previous iterations until no additional information about the entity is found when searching the multiple datasets with semantic representations of the now augmented target dataset. In some embodiments the augmented target dataset is used for determining privacy risks for the entity.