Unstructured Financial Data Extraction Using Parallel Summarization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for extracting financial information from unstructured data sources, such as news articles and press releases, rely heavily on manual effort and are inefficient, lacking automation and accuracy.
Innovation Solution
A method involving scraping, summarization, and postprocessing of unstructured data using a pre-trained summarizer to convert unstructured financial information into structured data, utilizing techniques like abstractive summarization and neural networks for efficient parallel processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual crowdsourcing is used to extract financial information from news articles and press releases, then information can be gathered from multiple sources, but the process requires significant manual effort and time
Solution Approach 1:
The patent replaces manual mechanical extraction processes with an automated neural network system. The model automatically scrapes text from news articles and press releases, extracts funding and revenue information, and structures the data without human intervention, thereby eliminating the time-consuming manual crowdsourcing process while maintaining extraction accuracy
Solution Approach 2:
The system enables self-service information extraction where the neural network model independently performs data scraping, extraction, and structuring. The automated pipeline processes multiple news sources simultaneously and generates structured financial information without requiring manual oversight for each extraction task
2Productivity
If traditional automated extraction methods are used, then data acquisition speed increases, but the extraction accuracy and reliability decrease
Solution Approach 1:
The patent transforms the extraction approach by changing key parameters: using pre-trained language models with fine-tuning on financial data, adjusting attention mechanisms to focus on funding-specific patterns, and optimizing the neural network architecture for both speed and accuracy. This enables the system to process data rapidly while maintaining high extraction precision
Solution Approach 2:
The system performs preliminary training of the neural network model on annotated financial data before deployment. This pre-training phase establishes accurate extraction patterns that enable the model to quickly and accurately process new news articles without requiring manual validation, achieving both high speed and high accuracy simultaneously
3Quantity of substance
If comprehensive text scraping is performed from multiple news sources, then more financial information becomes available, but the computational burden increases
Solution Approach 1:
The patent extracts only the essential funding and revenue information from news articles using targeted neural network extraction. Instead of processing and analyzing entire documents, the model identifies and extracts specific financial entities and relationships, significantly reducing computational requirements while maintaining comprehensive data coverage from multiple sources
Solution Approach 2:
The system segments the information extraction process into distinct neural network components: text scraping, information extraction, and structured output generation. This segmentation allows parallel processing of multiple news sources and efficient resource utilization, enabling the system to handle large volumes of data from numerous sources without overwhelming computational resources
Data Source
AI summary
A method includes extracting information from an unstructured data source, the method including: scraping, by at least one processor, a plurality of texts from the unstructured data source, extracting, by the at least one processor, from the plurality of texts a chunk of relevant text, summarizing, by the at least one processor, using a pre-trained summarizer, the chunk of relevant text to obtain semi-structured information comprising a set of sentences that summarize the chunk of relevant texts, and postprocessing, by the at least one processor, the semi-structured information to obtain structured information. The method can be executed highly efficiently, in particular on massively parallel hardware.


