Document Masking Device Using NLP Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing masking processes for documents are inefficient and labor-intensive, particularly for long documents, as they rely on manual blackening of personal information, leading to potential omissions and increased device load when adapting to changes in masking targets based on document content and recipients.
Innovation Solution
A document masking device and method using natural language processing to extract and present masking candidates, allowing users to select and output masked words, thereby reducing manual effort and device load by associating text data with image data and using AI to identify concealment target attributes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual blackening method is used to mask personal information in long documents, then the masking process can be performed with simple equipment, but the processing time increases significantly and masking omissions are likely to occur
Solution Approach 1:
The patent replaces the manual mechanical blackening process with an automated computer-based system that uses optical character recognition (OCR) to identify text and automatically applies masking. The system captures document images, converts them to text data, identifies personal information through keyword matching, and automatically generates masked images, eliminating the need for manual visual inspection and blackening operations.
Solution Approach 2:
The system enables self-service masking by automatically performing the entire masking workflow without human intervention. The computer device autonomously captures documents, recognizes text, identifies masking targets through predefined keywords, and generates masked outputs, allowing users to simply input masking requirements and receive processed results without manual operation.
2Productivity
If search function is used to extract masking target words from text data, then the masking process efficiency is improved, but the device load increases when adapting to changes in masking targets based on document content and recipients
Solution Approach 1:
The patent changes the approach from dynamic keyword extraction based on document analysis to using predefined keyword lists stored in the system. The computer device maintains a database of masking keywords that can be selected and applied based on document type and recipient, avoiding the computational complexity of real-time keyword extraction while maintaining adaptability through parameter selection rather than complex processing.
Solution Approach 2:
The system performs preliminary preparation by storing predefined masking keywords and categories in advance. Before actual masking operations, users can select appropriate keyword sets based on document type and recipient, and the system pre-loads these keywords for efficient matching. This eliminates the need for complex real-time analysis during the masking process, reducing device load while maintaining productivity.
3Ease of manufacture
If manual masking operation is performed on paper documents, then the equipment requirements are minimal, but the processing speed is slow and human error is likely
Solution Approach 1:
The patent replaces manual paper-based masking operations with an automated digital system that captures document images, performs optical character recognition, and automatically applies masking. The system uses computer vision and pattern recognition algorithms to identify and mask personal information, achieving high processing speeds while maintaining ease of use through a user-friendly interface that requires minimal input from operators.
Data Source
AI summary
This document masking device includes an extraction unit, a presentation unit, and an output unit so as to be able to flexibly respond to changing a word to be masked, and make a masking process performed on a document efficient while suppressing an increase in the load on a device. The extraction unit extracts, from text data of a document, words which belong to the attribute to be concealed that indicates the kinds of words to be masked by using natural language processing techniques. The presentation unit presents the extracted words as masking candidates. The output unit outputs a document in which words to be masked, designated as a masking target among the masking candidates, are masked.


