Automated Reference Data Structure Generation for Text Anonymization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for generating reference data structures for K-anonymity models, particularly for text data, are manual, time-consuming, and require extensive domain knowledge, making them inefficient and language-specific, with challenges in handling context-dependent meanings and large data sizes.

Innovation Solution

An automated method using machine learning techniques to generate a vector space for text data, allowing for the conversion of input text into numerical values, clustering similar records, and creating a reference data structure with meaningful identifiers, which can be updated and applied across multiple languages without the need for extensive human effort or additional files.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual methods are used to generate reference data structures for text data, then domain knowledge and semantic understanding can be applied, but the process becomes time-consuming and labor-intensive

Engineering Contradiction:
Improvesemantic accuracyVSAvoidgeneration time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical processes with an automated computational system. A processing system automatically generates reference data structures by converting text data into numerical vectors, clustering similar vectors, and extracting generalized terms without requiring manual domain expertise or semantic analysis by humans.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-service by automatically generating reference data structures from raw text data without external human intervention. The processing system independently completes all steps including vector conversion, clustering, and generalized term extraction, making the system autonomous and eliminating dependency on manual domain knowledge.

Inventive Principle:
Principle #25Self-service

2Reliability

If pre-made reference data structures are delivered with datasets, then K-anonymity can be implemented, but the data size increases

Engineering Contradiction:
Improveanonymity implementationVSAvoiddata size
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential generalized terms needed for K-anonymity from the text data, rather than delivering complete pre-made reference data structures. This extraction approach provides the necessary anonymization capability while minimizing the amount of additional data that needs to be stored and transmitted.

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If extensive knowledge of word taxonomy and semantic meaning is used, then accurate clustering can be achieved, but the complexity of the system increases

Engineering Contradiction:
Improveclustering accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces complex manual semantic analysis and word taxonomy knowledge with an automated vector-based system. By converting text to numerical vectors and using computational clustering algorithms, the system achieves accurate clustering without requiring explicit domain knowledge or complex manual rule systems.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11301639B2Methods and systems for generating a reference data structure for anonymization of text data
Publication Date: 2022.04.12 HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
  • US11301639B2 patent drawing
  • US11301639B2 patent drawing
  • US11301639B2 patent drawing

AI summary

A method and a system of using machine learning to automatically generate a reference data structure for a K-anonymity model. A vector space is generated from reference text data, where the vector space is defined by numerical vectors representative of semantic meanings of the reference text words. Input text words are converted into numerical vectors using the vector space. Word clusters are formed according to semantic similarity between the input text words, where the semantic similarity between pairs of input text words is represented by metric values determined from pairs of numerical vectors. The word clusters define nodes of the reference data structure. A text label is applied to each node of the reference data structure, where the text label is representative of the semantic meaning shared by elements of the word cluster.