K-Anonymity via Word Embedding Vector Substitution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional k-anonymization techniques are inaccurate and struggle with handling conversational data, such as transcripts of phone calls, due to their reliance on simple rules and the scarcity of training datasets, leading to inadequate anonymization of personal information.

Innovation Solution

The method employs linguistic similarity and embeddings distances between words to identify and replace word pairs, progressively anonymizing documents until k-anonymity is achieved, using a system comprising client devices and servers interconnected through a network to perform k-anonymization operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional k-anonymization techniques use simple rules and machine learning models for named entity recognition, then the anonymization process can be performed, but the accuracy of identifying personal information is poor

Engineering Contradiction:
Improveaccuracy of PII identificationVSAvoidreliability of anonymization
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent transforms the anonymization approach by changing from rule-based and traditional ML parameters to word embedding vector space parameters. By representing words as vectors and using cosine similarity to measure semantic proximity, the system achieves more accurate identification of personal information while maintaining reliability through mathematical rigor in vector space operations.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces mechanical rule-based systems and traditional machine learning models with a vector space model using word embeddings. This substitution enables the system to capture semantic relationships between words, significantly improving the accuracy of personal information identification in diverse contexts including conversational data.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If conventional approaches use clean labeling and structured data, then the anonymization can be performed on structured data, but the system cannot handle unstructured conversational data effectively

Engineering Contradiction:
Improvecapability to handle diverse data typesVSAvoidaccuracy of PII detection in unstructured data
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent creates a universal anonymization system based on word embeddings that can process both structured and unstructured data types. The vector space model provides a unified framework that captures semantic relationships across different data formats, enabling accurate PII detection in emails, transcripts, and other unstructured conversational data while maintaining effectiveness on structured data.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If simple rule-based methods are used for anonymization, then the process is computationally simple, but the results are inaccurate and cannot guarantee formal k-anonymity

Engineering Contradiction:
Improveguarantee of k-anonymityVSAvoidcomplexity of anonymization system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces word embedding vectors as an intermediary representation between raw text and anonymization decisions. This intermediary layer captures semantic meaning in a computationally tractable vector space, enabling the system to guarantee formal k-anonymity through mathematically rigorous similarity calculations while avoiding the need for complex rule-based systems.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11704481B1K-anonymity guarantee in text anonymization using word embeddings
Publication Date: 2023.07.18 INTUIT INC
  • US11704481B1 patent drawing
  • US11704481B1 patent drawing
  • US11704481B1 patent drawing

AI summary

Systems and methods for k-anonymizing a corpus of documents using linguistic similarities and embeddings distances between words. For instance, a word pair is selected based on linguistic similarity (e.g., belonging to the same part of speech) and small embeddings distance. For the selected word pair, a plurality of words is retrieved, also based on linguistic similarity to, and embeddings distances from, the selected word pair. Out of the plurality of words, a third word is identified that has a closer linguistic similarity to the word pair and also has smaller embeddings distances from the word pair. Each word in the word pair is then replaced by the third word. The process is repeated until k-anonymity is achieved.