Automated Reference Data Structure Generation for Text Anonymization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for generating reference data structures for K-anonymity models, particularly for text data, are manual, time-consuming, and require extensive domain knowledge, making them inefficient and language-specific, with challenges in handling context-dependent meanings and large data sizes.
Innovation Solution
An automated method using machine learning techniques to generate a vector space for text data, allowing for the conversion of input text into numerical values, clustering similar records, and creating a reference data structure with meaningful identifiers, which can be updated and applied across multiple languages without the need for extensive human effort or additional files.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual methods are used to generate reference data structures for text data, then domain knowledge and semantic understanding can be applied, but the process becomes time-consuming and labor-intensive
Solution Approach 1:
The patent replaces manual mechanical processes with an automated computational system. A processing system automatically generates reference data structures by converting text data into numerical vectors, clustering similar vectors, and extracting generalized terms without requiring manual domain expertise or semantic analysis by humans.
Solution Approach 2:
The system performs self-service by automatically generating reference data structures from raw text data without external human intervention. The processing system independently completes all steps including vector conversion, clustering, and generalized term extraction, making the system autonomous and eliminating dependency on manual domain knowledge.
2Reliability
If pre-made reference data structures are delivered with datasets, then K-anonymity can be implemented, but the data size increases
Solution Approach 1:
The patent extracts only the essential generalized terms needed for K-anonymity from the text data, rather than delivering complete pre-made reference data structures. This extraction approach provides the necessary anonymization capability while minimizing the amount of additional data that needs to be stored and transmitted.
3Measurement precision
If extensive knowledge of word taxonomy and semantic meaning is used, then accurate clustering can be achieved, but the complexity of the system increases
Solution Approach 1:
The patent replaces complex manual semantic analysis and word taxonomy knowledge with an automated vector-based system. By converting text to numerical vectors and using computational clustering algorithms, the system achieves accurate clustering without requiring explicit domain knowledge or complex manual rule systems.
Data Source
AI summary
A method and a system of using machine learning to automatically generate a reference data structure for a K-anonymity model. A vector space is generated from reference text data, where the vector space is defined by numerical vectors representative of semantic meanings of the reference text words. Input text words are converted into numerical vectors using the vector space. Word clusters are formed according to semantic similarity between the input text words, where the semantic similarity between pairs of input text words is represented by metric values determined from pairs of numerical vectors. The word clusters define nodes of the reference data structure. A text label is applied to each node of the reference data structure, where the text label is representative of the semantic meaning shared by elements of the word cluster.


