Keyword Extraction Model Using Multi-Modal Visual and Text Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional keyword extraction methods from images rely solely on OCR text, ignoring text visual information, leading to lower accuracy due to lost visual information and high OCR errors, which results in inappropriate keyword extraction.

Innovation Solution

A deep learning-based keyword extraction model that utilizes multi-modal information, including text content, text visual information, and image visual information, to enhance keyword extraction by incorporating visual features and reducing OCR errors through mode selection and parallel processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional OCR-based keyword extraction is used, then the process is simple, but the extraction accuracy is low due to lost visual information and high OCR errors

Engineering Contradiction:
Improvekeyword extraction accuracyVSAvoidextraction model complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges text-based OCR information with visual-based text detection information into a unified keyword extraction model. This combination allows the system to leverage both the semantic understanding from OCR and the visual accuracy from text detection, resolving the contradiction between extraction accuracy and processing complexity by integrating multiple information sources into a single comprehensive model.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The extraction model uses composite feature representation by combining text features from OCR with visual features from text detection algorithms. This composite approach creates a more robust feature set that maintains high accuracy while managing complexity through structured integration of multiple feature types.

Inventive Principle:
Principle #40Composite materials

2Reliability

If only OCR text is used for keyword extraction, then the processing is fast, but the reliability is low due to OCR errors

Engineering Contradiction:
Improvekeyword extraction reliabilityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent introduces visual text detection as an intermediary component that bridges the gap between OCR text and keyword extraction. This intermediary provides visual verification and supplementation of OCR results, improving reliability by filtering and validating text information before it reaches the keyword extraction stage, thereby maintaining reasonable processing speed.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent partially replaces the purely mechanical OCR-based text extraction with a visual-based text detection mechanism. This substitution reduces reliance on error-prone OCR by using visual pattern recognition to identify and validate text regions, thereby improving reliability while keeping processing time acceptable through efficient visual algorithms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If visual information is ignored in keyword extraction, then the method is simple, but the extracted keywords are inappropriate due to lost visual context

Engineering Contradiction:
Improvekeyword extraction qualityVSAvoidmodel structure complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent adds the visual dimension to the traditional text-only keyword extraction process. By incorporating visual features such as text layout, positioning, and appearance characteristics, the model gains another dimension of information that enhances keyword extraction quality. This dimensional expansion is managed through efficient feature fusion techniques that control model complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12135940B2Method for keyword extraction and electronic device implementing the same
Publication Date: 2024.11.05 SAMSUNG ELECTRONICS CO LTD
  • US12135940B2 patent drawing
  • US12135940B2 patent drawing
  • US12135940B2 patent drawing

AI summary

A method for keyword extraction, an apparatus, an electronic device, and a computer-readable storage medium, which relate to the field of artificial intelligence are provided. The method includes collecting feature information corresponding to an image to be processed, the feature information including text representation information and image visual information and then extracting keywords from the image to be processed based on the feature information. The text representation information includes text content and text visual information corresponding to each text line in the image to be processed. The method for keyword extraction, apparatus, electronic device, and computer-readable storage medium provided in the embodiments of the disclosure may extract the keywords from an image to be processed.