Document Area Extraction Using Linguistic Characteristics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for extracting an area of interest in documents require manual construction of new rules or patterns for each type of document, which is time-consuming and inefficient.

Innovation Solution

A method and apparatus that use linguistic characteristics, such as part-of-speech distribution ratios and image characteristics, to automatically extract areas of interest from documents without the need for new rules or patterns, employing classification models to identify target pages and sentences based on normalized frequency values.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If rule-based methods are used to separate specific areas in documents, then the specific area can be appropriately separated in documents for which analysis has been completed, but a new rule or pattern must be added each time when targeting a new type of document

Engineering Contradiction:
Improveaccuracy of area extractionVSAvoidcomplexity of rule construction
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs self-learning by automatically analyzing document characteristics and generating extraction rules without manual intervention. The model learns from the document itself, identifying patterns in text distribution, formatting, and structure to create extraction rules tailored to each document type, thereby eliminating the need for manual rule construction for new document types.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes the approach from fixed manual rules to dynamic parameter-based extraction. By analyzing various parameters such as text density, font characteristics, spacing patterns, and structural elements, the system automatically adjusts extraction parameters to suit different document types, maintaining accuracy without requiring new manual rules.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If manual rule construction is performed for each new document type, then extraction accuracy can be maintained, but the time required to construct new rules or patterns increases

Engineering Contradiction:
Improveaccuracy of area extractionVSAvoidtime for rule construction
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary analysis of document characteristics before extraction. By pre-processing the document to identify its type, structure, and key features, the system prepares extraction parameters in advance, enabling rapid and accurate extraction without time-consuming manual rule construction for each new document type.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The automated model performs self-learning and self-adjustment by analyzing new document types and generating appropriate extraction rules autonomously. This eliminates the time required for manual rule construction while maintaining extraction accuracy, as the system adapts to new document types automatically through machine learning.

Inventive Principle:
Principle #25Self-service

3Productivity

If rule-based methods are used, then extraction can be performed for documents with known patterns, but the system lacks adaptability to new types of documents without additional rules

Engineering Contradiction:
Improveextraction efficiencyVSAvoidadaptability to new document types
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The extraction system is designed with universal adaptability to handle multiple document types through a single unified model. By using machine learning to identify and adapt to various document structures and characteristics, the system maintains high extraction efficiency across different document types without requiring separate rule sets, thereby achieving both productivity and versatility.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system transitions from static manual rules to dynamic adaptive extraction. The model continuously learns from new document types and adjusts its extraction parameters dynamically, enabling it to maintain high productivity while adapting to diverse and evolving document formats without requiring manual rule updates.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20240078827A1Method and apparatus for extracting area of interest in a document
Publication Date: 2024.03.07 SAMSUNG SDS CO LTD
  • US20240078827A1 patent drawing
  • US20240078827A1 patent drawing
  • US20240078827A1 patent drawing

AI summary

A method for extracting an area of interest in a document is provided. The method may comprise extracting one or more target pages from a document composed of a plurality of pages and extracting an area of interest including a plurality of sentences from a target page based on a first part-of-speech characteristic and a sentence characteristic of the target page.