Document Image Anonymization via Named Entity Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional natural language processing techniques face difficulties in handling business form documents due to their lack of sentence structures and frequent use of conjunctions and particles, making them harder to comprehend and requiring burdensome registration of secret words for security, which complicates the generation of disclosable document images from confidential documents.
Innovation Solution
An information processing apparatus that extracts named entities from document images, replaces them with attribute tags, and generates attribute tag document data, allowing for the creation of disclosable learning data without revealing confidential information, using a system that includes an obtaining unit for character blocks, an extraction unit for named entities, and a generation unit for attribute tag document data, along with a database management system for secure storage and processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If secrete words are replaced by candidate words with the same attribute to convert confidential documents to disclosable documents, then security is improved, but the registration and management of secrete words becomes burdensome
Solution Approach 1:
The patent extracts and removes confidential information (named entities) from documents entirely, rather than replacing them with other words. This eliminates the need for registering and managing replacement word lists, while still achieving security goals through redaction or anonymization techniques.
Solution Approach 2:
The patent introduces an intermediary processing step that identifies named entities and replaces them with anonymized representations (such as masked text or generic placeholders) rather than requiring a pre-defined dictionary of replacement words. This intermediary approach simplifies management while maintaining security.
2Adaptability or versatility
If business form documents are used for natural language processing, then data utility is improved, but comprehension difficulty increases due to lack of sentence structures
Solution Approach 1:
The patent segments business form documents into structured components by identifying and extracting named entities with specific attributes (person names, organization names, locations, dates, amounts). This segmentation transforms unstructured or semi-structured form data into organized, machine-readable formats that improve comprehensibility while retaining data utility for NLP tasks.
Solution Approach 2:
The patent changes the parameter representation of document data by extracting named entities and representing them with standardized attribute tags (e.g., PERSON, ORGANIZATION, LOCATION, DATE, MONEY). This parameter transformation converts diverse business form formats into a unified structured representation that enhances both utility and comprehendibility.
Data Source
AI summary
The present disclosure relates to a technique of generating a disclosable document image based on a document image including confidential information, without using the confidential information. A document input unit obtains a document image scanned with a scanner, separates the document image into character information and background information, and then outputs them to an extraction unit. The extraction unit performs named entity extraction processing on the obtained character information and background information to extract named entities in the document and attributes thereof, and output an extraction result to a generation unit. The generation unit replaces the named entities in the document image with attribute tags and obtains superimposable ranges to generate attribute tag document data. A management unit registers the received extraction result of the named entities and the attribute tag document data in a database.


