Long-Document Prediction With Cross-Domain Counterfactual Augmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The analysis of earnings call transcripts is hindered by their lengthiness, scarcity, and noisiness, making it challenging for machine learning models to effectively predict and explain market volatility.
Innovation Solution
A system utilizing a transformer-based encoder with hierarchical sparse self-attention and a multi-source counterfactual augmentation module to generate perturbed earnings call transcripts, leveraging unlabeled sentences from external sources to enhance model training and prediction accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If deep learning methods are used for automatic earnings call analysis, then prediction capability is improved, but data scarcity becomes a limiting factor
Solution Approach 1:
The patent creates synthetic training data by copying and adapting sentences from external financial news sources into earnings call transcript formats. This copying approach generates additional training examples without requiring more actual earnings call data, directly addressing the data scarcity problem while maintaining prediction accuracy through the use of realistic synthetic samples
Solution Approach 2:
The system uses a universal sentence encoder that can process both real earnings call transcripts and synthetic sentences from multiple sources. This multi-functional encoder handles diverse input types (actual transcripts, synthetic augmentations, cross-domain sentences) uniformly, allowing the model to leverage limited labeled data effectively while benefiting from abundant unlabeled external sources
2Quantity of substance
If traditional analysis methods are used for earnings call transcripts, then data scarcity is tolerated, but analysis efficiency deteriorates
Solution Approach 1:
The system implements self-service through automated synthetic data generation and unsupervised training mechanisms. The model automatically creates its own training data from external sources and performs self-supervised learning on unlabeled transcripts, eliminating the need for manual data annotation while improving both data utilization and analysis efficiency through automated processing pipelines
Solution Approach 2:
The patent performs preliminary action by pre-processing and augmenting earnings call data with synthetic sentences before the main training process. This preliminary data preparation includes generating synthetic augmentations, filtering external sentences, and creating enhanced training sets in advance, which accelerates the subsequent training and analysis efficiency while maximizing data utilization
3Measurement precision
If earnings call transcripts are analyzed in detail, then prediction accuracy is improved, but noise interference increases
Solution Approach 1:
The system extracts and isolates salient information from noisy earnings call transcripts by identifying and focusing on key sentences while filtering out irrelevant content. The attention mechanism extracts important features from the input transcripts, separating signal from noise and improving prediction accuracy by concentrating on relevant information rather than processing all content equally
Solution Approach 2:
The patent applies local quality by using sentence-level and token-level attention mechanisms that assign different importance weights to different parts of the transcript. Rather than treating all content uniformly, the model identifies and emphasizes locally important sentences and phrases while downweighting noisy or irrelevant sections, thereby improving prediction accuracy through selective focus on high-quality information
4Measurement precision
If more earnings call data is collected for training, then model performance is improved, but data acquisition time increases
Solution Approach 1:
The system creates synthetic training data by copying and transforming sentences from external financial news sources into earnings call transcript formats. This copying approach generates additional training examples immediately without requiring time-consuming collection of actual earnings call data, improving model performance while avoiding data acquisition delays
Solution Approach 2:
The patent performs preliminary data augmentation and preprocessing steps before main training, including generating synthetic sentences, filtering external sources, and preparing enhanced training sets in advance. This preliminary action reduces the need for extensive data collection during the training process, improving model performance while minimizing data acquisition time through upfront preparation
Data Source
AI summary
Unsupervised cross-domain data augmentation techniques for long-text document based prediction and explanation are provided. In one aspect, a system for long-document based prediction includes: an encoder for creating embeddings of long-document texts with hierarchical sparse self-attention, and making predictions using the embeddings of the long-document texts; and a multi-source counterfactual augmentation module for generating perturbed long-document texts using unlabeled sentences from at least one external source to train the encoder. A method for long-document based prediction is also provided.


