Avatar Data Generation for Privacy-Preserving Analytics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current anonymization techniques fail to effectively protect sensitive data from re-identification while maintaining data utility, particularly in large datasets, as they often rely on reversible methods or distort data, leading to security vulnerabilities and loss of information precision.

Innovation Solution

A method that creates avatars by generating new attribute values from the k nearest neighbors of an individual, weighted by coefficients, to produce synthetic, non-identifiable data that preserves the original data's structure and utility, using multivariate analysis and stochastic weighting to ensure anonymity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If substitution or pseudonymization is used to anonymize data, then data can be accessed and traced, but the anonymization is reversible and security is compromised

Engineering Contradiction:
Improveanonymization securityVSAvoiddata traceability
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent creates synthetic copies (avatars) of real data that replicate statistical properties and relationships without containing actual personal information. These avatar records are generated through stochastic processes that mimic the distribution and correlations of the original data, allowing analysis while preventing re-identification of individuals.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms real data parameters into synthetic parameters through mathematical transformations. Continuous variables are transformed using cumulative distribution functions and stochastic sampling, while categorical variables are transformed through probability-based mapping to avatar categories, fundamentally changing the data representation while preserving statistical structure.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If hashing is used for anonymization, then the process becomes irreversible, but correlation tables can still be reconstituted through reiteration attacks

Engineering Contradiction:
Improveanonymization irreversibilityVSAvoidcorrelation table security
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces synthetic avatar records as an intermediary between the original sensitive data and any analysis performed. These avatars act as a buffer that preserves statistical properties for analysis while breaking the direct link to individual identities, making it impossible to reconstruct original data or create meaningful correlation tables.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If noise addition or masking is applied to sensitive data, then anonymization is improved, but data becomes distorted and less relevant for analysis

Engineering Contradiction:
Improveanonymization strengthVSAvoiddata precision
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent employs dynamic, data-driven transformation processes that adapt to the specific characteristics of each dataset. The stochastic transformation parameters are determined by the actual distribution and relationships in the data, allowing the anonymization process to preserve relevant statistical properties while ensuring privacy protection.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies targeted parameter transformations that change only the necessary aspects of data for anonymization while preserving analytical properties. Continuous variables undergo transformation through cumulative distribution functions that maintain relative positions and relationships, and categorical variables are transformed through probability-based mapping that preserves class distributions and associations.

Inventive Principle:
Principle #35Parameter changes

4Reliability

If aggregation is used to combine values into classes, then re-identification risk decreases, but data precision and information content are lost

Engineering Contradiction:
Improvere-identification protectionVSAvoiddata precision
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent creates synthetic copies that replicate the precision and detail of original data without copying actual personal information. The avatar records maintain the same level of measurement precision as the original data through stochastic transformation processes that preserve distribution characteristics and relationship structures.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12026281B2Method for creating avatars for protecting sensitive data
Publication Date: 2024.07.02 BIG DATA SANTE
  • US12026281B2 patent drawing
  • US12026281B2 patent drawing
  • US12026281B2 patent drawing

AI summary

The present invention relates to a method for creating avatars from an initial sensitive data set stored in a database of a computer system, the initial data comprising attributes relating to a plurality of individuals, the method comprising: a) choosing a number {k} of nearest neighbors to be used from all the individuals in the initial data set, b) identifying, for attributes relating to a given individual, the k nearest neighbors from among the other individuals in the data set, c) generating, for at least one attribute relating to said individual, a new attribute value from quantities which are characteristic of the attribute in the identified k nearest neighbors and weighted by a coefficient, and d) creating avatar data comprising the new attribute value(s), so as to ensure the sensitive data relating to the individual are non-identifiable.