Artificial Document Generation with Metadata Ratio Override

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current artificial data generation methods struggle to effectively mimic the structure and content of source documents across various domains, such as finance and medicine, while maintaining the ratios and dependencies between terms, which limits their applicability in generating realistic and context-specific synthetic data.

Innovation Solution

The system employs a processor-based method that derives metadata from source documents, trains a model to output artificial documents, and overrides generated words to retain the ratios and dependencies, using techniques like recurrent neural networks and n-gram applications to ensure the artificial documents mimic the source documents' characteristics, including sentence structures and grammatical guidelines.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If current artificial data generation methods are used, then data generation can be performed, but the structure and content of source documents cannot be effectively mimicked while maintaining ratios and dependencies between terms

Engineering Contradiction:
Improvemimicking accuracyVSAvoiddomain applicability
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The system segments the document generation process into distinct components: metadata extraction from source documents, dictionary and n-gram generation for domain-specific terminology, neural network model training, and controlled text generation. This segmentation allows each component to be optimized independently, improving both mimicking accuracy and domain adaptability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes parameters by deriving metadata parameters (such as term ratios, sentence structures, and formatting characteristics) from source documents and using these as controlled parameters during artificial document generation. This ensures the generated documents maintain the same statistical properties and dependencies as the source documents across different domains.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If a neural network model generates artificial documents, then document generation speed increases, but the ratios and dependencies between terms are not maintained

Engineering Contradiction:
Improvegeneration speedVSAvoidratio preservation
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The system implements feedback mechanisms where the neural network model is trained on metadata extracted from source documents, including term ratios and n-grams. During generation, the model receives feedback through the loss function that measures deviation from target ratios and dependencies, continuously adjusting predictions to maintain accuracy while preserving statistical properties.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary actions by extracting metadata, generating dictionaries and n-grams, and training the neural network model before actual document generation. This pre-processing establishes the statistical framework and term relationships that the model will maintain during high-speed generation, ensuring both speed and precision.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If metadata parameters are derived from source documents and used to control generation, then document realism improves, but system complexity increases

Engineering Contradiction:
Improvedocument realismVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system extracts only the essential metadata parameters needed for realistic document generation, such as term frequency ratios, n-gram sequences, sentence structure patterns, and formatting characteristics. By taking out only these critical features rather than copying entire source documents, the system achieves high document realism while managing complexity through selective feature extraction.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11586815B2Method, system and computer program product for generating artificial documents
Publication Date: 2023.02.21 PROOV SYST LTD
  • US11586815B2 patent drawing
  • US11586815B2 patent drawing
  • US11586815B2 patent drawing

AI summary

An artificial document generation method generating a larger set of artificial documents which mimics a smaller set of source documents. The method may include deriving metadata parameters from source documents, each parameter characterizing the source documents and having more than one possible value. The deriving may include: determining which ratio of the source documents has each of the values, defining metadata ratios characteristic of the source documents; training a model, on the source documents, to output artificial documents; and running the model as trained thereby to output artificial documents and overriding at least some words generated by the model in the draft artificial documents. Overriding is configured to ensure that at least some of the ratios characteristic of the smaller set of source documents are retained in the larger set.