Document Anonymization via Selective Token Modification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current techniques for anonymizing writing styles do not account for an author's personality vector score or profile, making it difficult to maintain author anonymity as personality characteristics can still be identified through writing style.
Innovation Solution
An artificial intelligence platform with a natural language manager, document manager, and director that processes documents using natural language processing to identify and modify personality vector scores, selectively amending tokens in new documents to create a new version that deviates from the original author's characteristics, thereby obscuring identity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If current anonymization techniques are used, then basic text processing is achieved, but author anonymity is not maintained because personality characteristics remain identifiable
Solution Approach 1:
The system modifies specific linguistic parameters (tokens) within the document to change the personality vector score. By selectively amending tokens based on their contribution to personality characteristics, the system transforms the document's stylistic parameters to deviate from the author's original profile, thereby achieving effective anonymization while preserving meaningful content.
Solution Approach 2:
The system employs a feedback loop where the personality vector score is calculated before and after token modification. The difference in vector scores guides the selection of tokens to be amended, ensuring that modifications effectively change the authorial personality profile while maintaining document coherence and meaning.
2Reliability
If tokens are selectively amended to change personality vector score, then author anonymity is improved, but document meaning may be altered
Solution Approach 1:
The system applies local quality by selectively modifying only specific tokens that contribute to personality characteristics rather than uniformly altering the entire document. This targeted approach allows the system to change authorial style while preserving the overall meaning and structure of the document, as only localized linguistic patterns are modified.
Solution Approach 2:
The system replaces manual, rule-based anonymization with an AI-driven approach that uses natural language processing and machine learning models. This substitution enables more nuanced and context-aware token selection, allowing the system to distinguish between tokens that define personality characteristics and those that are essential for maintaining document meaning.
Data Source
AI summary
Embodiments relate to an intelligent computer platform to selectively amend one or more tokens in a document. A first document set is subjected to natural language processing (NLP) and a vector score is identified for two or more documents of the first document set. Upon receipt of a new document, the new document is subjected to NLP and a new document vector score is identified. The new document is analyzed against the first document set, and the identified vector score of the first document set is compared to the vector score of the new document. One or more tokens of the new document are amended responsive to the comparison, and a new document version is created from the selective amendment.


