K-Anonymity via Word Embedding Vector Substitution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional k-anonymization techniques are inaccurate and struggle with handling conversational data, such as transcripts of phone calls, due to their reliance on simple rules and the scarcity of training datasets, leading to inadequate anonymization of personal information.
Innovation Solution
The method employs linguistic similarity and embeddings distances between words to identify and replace word pairs, progressively anonymizing documents until k-anonymity is achieved, using a system comprising client devices and servers interconnected through a network to perform k-anonymization operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional k-anonymization techniques use simple rules and machine learning models for named entity recognition, then the anonymization process can be performed, but the accuracy of identifying personal information is poor
Solution Approach 1:
The patent transforms the anonymization approach by changing from rule-based and traditional ML parameters to word embedding vector space parameters. By representing words as vectors and using cosine similarity to measure semantic proximity, the system achieves more accurate identification of personal information while maintaining reliability through mathematical rigor in vector space operations.
Solution Approach 2:
The patent replaces mechanical rule-based systems and traditional machine learning models with a vector space model using word embeddings. This substitution enables the system to capture semantic relationships between words, significantly improving the accuracy of personal information identification in diverse contexts including conversational data.
2Adaptability or versatility
If conventional approaches use clean labeling and structured data, then the anonymization can be performed on structured data, but the system cannot handle unstructured conversational data effectively
Solution Approach 1:
The patent creates a universal anonymization system based on word embeddings that can process both structured and unstructured data types. The vector space model provides a unified framework that captures semantic relationships across different data formats, enabling accurate PII detection in emails, transcripts, and other unstructured conversational data while maintaining effectiveness on structured data.
3Reliability
If simple rule-based methods are used for anonymization, then the process is computationally simple, but the results are inaccurate and cannot guarantee formal k-anonymity
Solution Approach 1:
The patent introduces word embedding vectors as an intermediary representation between raw text and anonymization decisions. This intermediary layer captures semantic meaning in a computationally tractable vector space, enabling the system to guarantee formal k-anonymity through mathematically rigorous similarity calculations while avoiding the need for complex rule-based systems.
Data Source
AI summary
Systems and methods for k-anonymizing a corpus of documents using linguistic similarities and embeddings distances between words. For instance, a word pair is selected based on linguistic similarity (e.g., belonging to the same part of speech) and small embeddings distance. For the selected word pair, a plurality of words is retrieved, also based on linguistic similarity to, and embeddings distances from, the selected word pair. Out of the plurality of words, a third word is identified that has a closer linguistic similarity to the word pair and also has smaller embeddings distances from the word pair. Each word in the word pair is then replaced by the third word. The process is repeated until k-anonymity is achieved.


