Document Anonymization via Selective Token Modification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current techniques for anonymizing writing styles do not account for an author's personality vector score or profile, making it difficult to maintain author anonymity as personality characteristics can still be identified through writing style.

Innovation Solution

An artificial intelligence platform with a natural language manager, document manager, and director that processes documents using natural language processing to identify and modify personality vector scores, selectively amending tokens in new documents to create a new version that deviates from the original author's characteristics, thereby obscuring identity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If current anonymization techniques are used, then basic text processing is achieved, but author anonymity is not maintained because personality characteristics remain identifiable

Engineering Contradiction:
Improveauthor anonymityVSAvoidpersonality vector score
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The system modifies specific linguistic parameters (tokens) within the document to change the personality vector score. By selectively amending tokens based on their contribution to personality characteristics, the system transforms the document's stylistic parameters to deviate from the author's original profile, thereby achieving effective anonymization while preserving meaningful content.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system employs a feedback loop where the personality vector score is calculated before and after token modification. The difference in vector scores guides the selection of tokens to be amended, ensuring that modifications effectively change the authorial personality profile while maintaining document coherence and meaning.

Inventive Principle:
Principle #23Feedback

2Reliability

If tokens are selectively amended to change personality vector score, then author anonymity is improved, but document meaning may be altered

Engineering Contradiction:
Improveauthor anonymityVSAvoiddocument meaning preservation
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The system applies local quality by selectively modifying only specific tokens that contribute to personality characteristics rather than uniformly altering the entire document. This targeted approach allows the system to change authorial style while preserving the overall meaning and structure of the document, as only localized linguistic patterns are modified.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system replaces manual, rule-based anonymization with an AI-driven approach that uses natural language processing and machine learning models. This substitution enables more nuanced and context-aware token selection, allowing the system to distinguish between tokens that define personality characteristics and those that are essential for maintaining document meaning.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11334716B2Document anonymization including selective token modification
Publication Date: 2022.05.17 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11334716B2 patent drawing
  • US11334716B2 patent drawing
  • US11334716B2 patent drawing

AI summary

Embodiments relate to an intelligent computer platform to selectively amend one or more tokens in a document. A first document set is subjected to natural language processing (NLP) and a vector score is identified for two or more documents of the first document set. Upon receipt of a new document, the new document is subjected to NLP and a new document vector score is identified. The new document is analyzed against the first document set, and the identified vector score of the first document set is compared to the vector score of the new document. One or more tokens of the new document are amended responsive to the comparison, and a new document version is created from the selective amendment.