Keyword Extraction Target Determination in Document Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In offices and companies, extracting keywords from a large volume of documents, including digitized and non-digitized materials, often results in irrelevant words being selected as keywords due to the lack of targeted selection processes.
Innovation Solution
A document processing apparatus comprising an image processing unit, a target determination unit, and a keyword extraction unit that identifies and processes images to extract relevant keywords from documents, using methods such as OCR and character recognition, and determines the relevance based on content, format, and specific character strings, to differentiate between keyword extraction targets and non-targets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If keywords are extracted from all documents without selection, then the quantity of extracted keywords increases, but the quality and relevance of keywords deteriorates
Solution Approach 1:
The target determination unit performs preliminary filtering of documents before keyword extraction by analyzing document images to identify those containing keywords worth extracting. This preliminary action eliminates irrelevant documents from the extraction process, ensuring that only documents with potential valuable keywords are processed, thus maintaining high keyword relevance while managing extraction volume effectively.
2Loss of information
If keyword extraction is performed on all documents, then comprehensive keyword coverage is achieved, but processing time and computational resources increase
Solution Approach 1:
The system extracts and analyzes specific features from document images (such as text content, layout patterns, and visual characteristics) to determine whether the document contains keywords worth extracting. By taking out only the essential visual features for analysis rather than processing entire documents uniformly, the system achieves comprehensive keyword coverage from relevant documents while significantly reducing processing time and computational resource consumption.
3Measurement precision
If document images are analyzed in detail to determine extraction targets, then keyword extraction accuracy improves, but image processing complexity increases
Solution Approach 1:
The target determination unit applies different levels of image analysis to different documents based on their characteristics. For documents that appear to contain valuable keywords based on initial analysis, more detailed local examination is performed. For clearly irrelevant documents, minimal processing is applied. This local quality approach ensures high keyword extraction accuracy for relevant documents while avoiding unnecessary complex processing of irrelevant ones.
Data Source
AI summary
A document processing apparatus includes an image processing unit, a target determination unit, a text reading unit, and a keyword extraction unit. The image processing unit processes an image. The target determination unit determines whether an image that is a target of a process executed by the image processing unit is to be used as a keyword extraction target. The text reading unit reads a text from the image processed by the image processing unit. The keyword extraction unit extracts a keyword from the text read by the text reading unit from an image determined to be a keyword extraction target by the target determination unit.


