Database Query Generation for Accurate Patent Text Corpus Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automated text processing methods, particularly for patent documents, fail to accurately extract statements necessary for pre-training, training, or retraining classification and clustering models due to insufficient completeness and specificity.

Innovation Solution

A method for generating a request to a database that involves identifying and parsing natural language texts into segments, marking up parts for analysis, and extracting main and associative entities to form a text corpus, which can be used for pre-training, training, or retraining classification and clustering models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated text processing methods are used to extract statements from patent documents, then processing efficiency is improved, but extraction accuracy and completeness deteriorate

Engineering Contradiction:
Improvetext processing efficiencyVSAvoidstatement extraction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies segmentation by dividing the text processing into distinct stages: identifying natural language texts with at least three segments, selecting specific segments (first, second, and third segments), and parsing marked-up parts within those segments. This multi-level segmentation enables automated processing while maintaining extraction accuracy by focusing on relevant portions of the text.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary markup step where parts to be parsed are marked up before semantic and syntactic analysis. This intermediary representation serves as a bridge between raw text and extracted statements, enabling accurate extraction while maintaining automated processing efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If comprehensive text processing is performed on all patent document parts, then extraction completeness is improved, but processing complexity increases

Engineering Contradiction:
Improveextraction completenessVSAvoidprocessing system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent applies local quality by treating different segments of the patent document with different processing approaches. The first segment is marked up with only one part to be parsed, while the second and third segments have at least one part marked up for parsing. This differentiated local processing ensures comprehensive extraction without requiring uniform complex processing across the entire document.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements partial action by selectively processing only the necessary parts of each segment rather than analyzing the entire patent document uniformly. This approach achieves sufficient extraction completeness for training purposes while reducing overall processing complexity.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If multiple segments are processed and analyzed, then statement extraction accuracy is improved, but processing time increases

Engineering Contradiction:
Improveextraction accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by performing markup of parts to be parsed before semantic and syntactic analysis. This preparation step is done in advance for the selected segments, enabling more efficient subsequent processing and reducing overall processing time while maintaining extraction accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20260087251A1Method for generating a request to a database
Publication Date: 2026.03.26 KRAVCHENKO ARTEM ALEKSANDROVICH
  • US20260087251A1 patent drawing
  • US20260087251A1 patent drawing

AI summary

The proposed technical solution relates to methods of automated text processing and can be used in the generating of text corpuses. The technical problem solved by the claimed invention is the creation of a method and/or a computer device and/or a system and/or a machine-readable data carrier that do not have the disadvantages of analogs and thus ensure accurate automated generation of a text corpus, which can subsequently be used for pre-training, or training, or additional training of classification models and/or clustering models.