Long-Range Event Relation Extraction via Synthetic Data Augmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document analysis systems face challenges with computational accuracy and operational flexibility in determining long-range event relations within digital documents, primarily due to limited training data consisting of short-range event relations, leading to inaccuracies and inflexibility in handling event pairs separated by greater distances.

Innovation Solution

The system generates a synthetically augmented long-range event relation dataset by inserting contextually coherent synthetic sentences between event pairs in digital documents, using a generative language model to expand the range of event relations beyond the initial short-range threshold, thereby training an event relation extraction model capable of identifying long-range relations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If existing document analysis systems are trained only on short-range event relation data, then the training data requirement is reduced, but the accuracy and flexibility in determining long-range event relations deteriorates

Engineering Contradiction:
Improvetraining data quantityVSAvoidevent relation extraction accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent uses a generative language model to create synthetic event relation datasets that replicate the structure and patterns of real long-range event relations. These synthetic datasets copy the essential characteristics of actual event relations while providing the necessary training data that would otherwise be unavailable or insufficient, thereby enabling accurate long-range event relation extraction without requiring extensive real-world data

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary data generation and augmentation before the actual event relation extraction task. By pre-generating synthetic long-range event relation datasets and pre-training the model on this augmented data, the system prepares the necessary learning material in advance, allowing the model to accurately handle long-range event relations when they actually need to be extracted from documents

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If existing document analysis systems are trained only on short-range event relations, then the system complexity is reduced, but the adaptability to various event pair distances deteriorates

Engineering Contradiction:
Improvesystem complexityVSAvoidevent pair distance adaptability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent modifies the training data parameters by creating synthetic datasets with varying event relation distances, contexts, and structures. This parameter transformation allows the model to learn patterns across different event pair distances without fundamentally changing the system architecture, thereby achieving adaptability to various event distances while maintaining manageable system complexity through data-level transformation rather than system-level complexity

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If synthetic sentences are inserted to create long-range event relation datasets, then the event relation extraction accuracy for long-range relations is improved, but the data generation process complexity increases

Engineering Contradiction:
Improvelong-range event relation accuracyVSAvoiddata generation process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The generative language model serves as an intermediary between the existing short-range event relation data and the desired long-range event relation training data. This intermediary component automatically transforms and augments the data, generating synthetic long-range event relations by inserting contextual sentences between event pairs, thereby achieving high accuracy in long-range event relation extraction while managing data generation complexity through automated intelligent transformation rather than manual complex processing

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240378370A1Generating and utilizing models for long-range event relation extraction
Publication Date: 2024.11.14 ADOBE INC
  • US20240378370A1 patent drawing
  • US20240378370A1 patent drawing
  • US20240378370A1 patent drawing

AI summary

The present disclosure relates to systems, methods, and non-transitory computer-readable media that generates a long-range event relation dataset by augmenting a digital document with a set of synthetic sentences. For example, the disclosed systems access a digital document from a short-range event relation dataset that includes an event pair. In some embodiments, the disclosed systems generate a set of synthetic sentences utilizing a generative language model for inserting within the digital document between the event pair to satisfy a long-range event relation threshold. In these or other embodiments, the disclosed systems generate a long-range event relation dataset by augmenting the digital document within the short-range event relation dataset to include the set of synthetic sentences. In certain cases, the disclosed systems generate an event relation extraction model to determine long-range event relations by learning model parameters for the event relation extraction model from the long-range event relation dataset.