Keyword Extraction Model Using Multi-Modal Visual and Text Features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional keyword extraction methods from images rely solely on OCR text, ignoring text visual information, leading to lower accuracy due to lost visual information and high OCR errors, which results in inappropriate keyword extraction.
Innovation Solution
A deep learning-based keyword extraction model that utilizes multi-modal information, including text content, text visual information, and image visual information, to enhance keyword extraction by incorporating visual features and reducing OCR errors through mode selection and parallel processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional OCR-based keyword extraction is used, then the process is simple, but the extraction accuracy is low due to lost visual information and high OCR errors
Solution Approach 1:
The patent merges text-based OCR information with visual-based text detection information into a unified keyword extraction model. This combination allows the system to leverage both the semantic understanding from OCR and the visual accuracy from text detection, resolving the contradiction between extraction accuracy and processing complexity by integrating multiple information sources into a single comprehensive model.
Solution Approach 2:
The extraction model uses composite feature representation by combining text features from OCR with visual features from text detection algorithms. This composite approach creates a more robust feature set that maintains high accuracy while managing complexity through structured integration of multiple feature types.
2Reliability
If only OCR text is used for keyword extraction, then the processing is fast, but the reliability is low due to OCR errors
Solution Approach 1:
The patent introduces visual text detection as an intermediary component that bridges the gap between OCR text and keyword extraction. This intermediary provides visual verification and supplementation of OCR results, improving reliability by filtering and validating text information before it reaches the keyword extraction stage, thereby maintaining reasonable processing speed.
Solution Approach 2:
The patent partially replaces the purely mechanical OCR-based text extraction with a visual-based text detection mechanism. This substitution reduces reliance on error-prone OCR by using visual pattern recognition to identify and validate text regions, thereby improving reliability while keeping processing time acceptable through efficient visual algorithms.
3Adaptability or versatility
If visual information is ignored in keyword extraction, then the method is simple, but the extracted keywords are inappropriate due to lost visual context
Solution Approach 1:
The patent adds the visual dimension to the traditional text-only keyword extraction process. By incorporating visual features such as text layout, positioning, and appearance characteristics, the model gains another dimension of information that enhances keyword extraction quality. This dimensional expansion is managed through efficient feature fusion techniques that control model complexity.
Data Source
AI summary
A method for keyword extraction, an apparatus, an electronic device, and a computer-readable storage medium, which relate to the field of artificial intelligence are provided. The method includes collecting feature information corresponding to an image to be processed, the feature information including text representation information and image visual information and then extracting keywords from the image to be processed based on the feature information. The text representation information includes text content and text visual information corresponding to each text line in the image to be processed. The method for keyword extraction, apparatus, electronic device, and computer-readable storage medium provided in the embodiments of the disclosure may extract the keywords from an image to be processed.


