Unstructured Document Migrator Using Verb Noun Ratio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The migration of unstructured documents to Darwin Information Typing Architecture (DITA) is hindered by the complexity of determining the correct topic type, requiring manual metadata addition and resulting in poor partial conversion with excessive post-conversion manual efforts, as existing conversion scripts cannot automatically identify source and target topic types.

Innovation Solution

A method that calculates the verb to noun ratio of an unstructured document to assign a weight, which is used to migrate the document to a specific DITA topic type, utilizing a processor to execute these calculations and migrations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual metadata addition is used to enable accurate topic type identification, then topic typing precision is improved, but conversion time and labor cost increase significantly

Engineering Contradiction:
Improvetopic typing precisionVSAvoidconversion time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs self-service by automatically analyzing the unstructured document content and determining the appropriate DITA topic type without requiring manual metadata addition. The processor examines the document structure, content patterns, and linguistic features to autonomously classify the content into the correct topic type (task, concept, reference, etc.), thereby eliminating the need for manual intervention while maintaining high classification accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The manual mechanical process of adding metadata is replaced with an automated computational system. The processor uses algorithms to analyze document characteristics and automatically assign topic types, substituting human manual work with machine-based content analysis and classification mechanisms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If conversion scripts are used without manual metadata, then conversion speed is improved, but topic type identification accuracy deteriorates

Engineering Contradiction:
Improveconversion speedVSAvoidtopic type identification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The conversion script performs self-service by automatically determining topic types through content analysis rather than requiring pre-added metadata. The system autonomously evaluates the unstructured document and classifies it into the appropriate DITA topic type, maintaining both high conversion speed and accurate topic identification.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes the parameters of topic type identification from relying on pre-added metadata to analyzing intrinsic document characteristics such as content structure, linguistic patterns, and semantic features. This parameter change enables accurate classification while maintaining automated processing speed.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If generic DITA topic type is used for all documents, then conversion complexity is reduced, but content reuse capability is diminished

Engineering Contradiction:
Improveconversion script complexityVSAvoidcontent reuse capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The system segments the document classification process into distinct analysis stages, examining different document characteristics (structure, content patterns, linguistic features) to determine the appropriate topic type. This segmentation enables accurate classification without requiring overly complex monolithic conversion scripts.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The conversion system performs self-service by automatically determining the most appropriate specific topic type for each document based on its content characteristics. This eliminates the need to use generic topic types for all documents, thereby preserving content reuse capability while keeping the conversion process automated and reasonably complex.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10592538B2Unstructured document migrator
Publication Date: 2020.03.17 EDISON VAULT LLC
  • US10592538B2 patent drawing
  • US10592538B2 patent drawing
  • US10592538B2 patent drawing

AI summary

Aspects migrate an unstructured document to a specific document type definition Darwin Information Typing architecture wherein processors are configured to calculate a verb to noun ratio of an unstructured document by dividing a of plurality verbs of the unstructured document by a plurality of nouns of the unstructured document, assign a first weight to the unstructured document based on the calculated verb to noun ratio, and migrate the unstructured document to a specific document type definition Darwin Information Typing Architecture based on the first weight.