Document Content Extraction Using Learning Model

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face challenges in extracting relevant content from large documents, such as PDFs, to generate a concise version, as existing methods are inefficient and require significant human effort for searching and editing.

Innovation Solution

A method involving text extraction from documents using a predetermined learning model, where keywords are defined by users to identify and integrate relevant pages, with the model trained on user-defined samples to improve accuracy and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If manual searching and editing methods are used to extract relevant content from large documents, then users can obtain customized document versions, but the process requires significant human effort and time

Engineering Contradiction:
Improvedocument content extraction automationVSAvoidtime for searching and editing
Core Design Contradiction:
Extent of automationVSLoss of time

Solution Approach 1:

The system performs self-service by automatically analyzing document content, identifying relevant sections based on user-defined keywords, and generating condensed document versions without requiring manual searching or editing. The learning model autonomously completes the extraction task that would otherwise require human effort.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual process of searching and editing documents with an automated learning model. The model uses machine learning algorithms to identify and extract relevant content, substituting human cognitive and manual labor with computational processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If existing manual methods are used for document content extraction, then relevant content can be identified, but the process is inefficient and requires significant human effort

Engineering Contradiction:
Improvedocument processing efficiencyVSAvoidhuman effort for searching and editing
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The system autonomously performs document analysis and content extraction without requiring users to manually search through documents or edit content. The learning model independently identifies relevant sections and generates condensed versions, making the process self-service oriented.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces inefficient manual document processing with an automated learning model that efficiently analyzes and extracts relevant content. This substitution dramatically improves productivity by eliminating the need for human users to manually search and edit large documents.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If a condensed version of the document is generated to focus on specific content, then document specificity is improved, but the extraction process becomes problematic with existing methods

Engineering Contradiction:
Improvedocument parsing accuracyVSAvoidextraction process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces complex manual extraction processes with a learning model that automatically parses documents and identifies relevant content. The model uses machine learning to accurately distinguish between relevant and irrelevant sections, achieving high parsing accuracy without manual intervention.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The learning model performs self-service by autonomously analyzing document structure, identifying relevant content based on keywords, and generating condensed versions. This self-service capability simplifies the extraction process while maintaining high accuracy in content selection.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20230342385A1Method for analyzing document for desired content and exracting same, electronic device employing method, and non-transitory storage medium
Publication Date: 2023.10.26 HONG FU JIN PRECISION IND (WUHAN) CO LTD
  • US20230342385A1 patent drawing
  • US20230342385A1 patent drawing
  • US20230342385A1 patent drawing

AI summary

A method for extracting text relevant to a particular topic from a document to generate a narrower, condensed, more specific, and smaller version obtains information as to the text of a large document and searches the text information of the document based on first keywords to extract first pages. The first pages are inputted into a predetermined learning model to extract second keywords from the first pages, and the first keywords and the second keywords are integrated to obtain third keywords. The method further searches the text of the first pages based on the third keywords to extract second pages and a determination is made as to whether the second pages meet a predetermined page standard. If yes, the second pages are integrated and output. An electronic device and a non-transitory storage medium are also disclosed.