Automated Data Curation for Corporate Action Prospectuses
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data curation techniques for generating corporate action prospectus narratives are resource-intensive and not timely due to the complexity and uncommon vocabulary in corporate action prospectuses, making real-time curation infeasible with reasonable resource allocations.
Innovation Solution
Implementing a method that uses natural language processing, expression patterns, and scoring algorithms to automatically curate prospectus and narrative data in real-time by retrieving electronic documents, converting them into machine-readable formats, preprocessing to identify linguistic units, extracting key attributes, and generating messages based on these attributes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional data curation techniques are used to generate message narratives from corporate action prospectuses, then the narratives can be produced with sufficient accuracy, but the process requires large resource allocations and cannot be completed in real-time
Solution Approach 1:
The patent segments the corporate action prospectus into multiple linguistic units (paragraphs and sentences), and further segments the extraction process into multiple stages including preprocessing, attribute identification using expression patterns, scoring, and selection. This segmentation allows the system to process complex documents systematically without requiring excessive resources while maintaining accuracy.
Solution Approach 2:
The patent changes the parameter of resource allocation by implementing a scoring algorithm that ranks extracted attributes based on multiple criteria (keyword hit criterion, context criterion, word clustering criterion, word embedding criterion). This parameter change enables the system to prioritize processing of most relevant information, reducing overall resource requirements while maintaining curation quality.
2Loss of energy
If conventional data curation techniques are used with reasonable resource allocations, then resource usage is optimized, but timely curation of message narratives becomes infeasible
Solution Approach 1:
The patent performs preliminary actions by pre-processing the corporate action prospectus to identify linguistic units and pre-defining expression patterns for attribute extraction before the actual curation process. This preliminary preparation enables faster processing during real-time message narrative generation without increasing resource consumption during the critical curation phase.
Solution Approach 2:
The patent replaces manual mechanical curation processes with automated computational methods including natural language processing, expression pattern matching, and scoring algorithms. This substitution eliminates the need for human reviewers while maintaining curation quality, thereby reducing both resource consumption and time requirements.
3Measurement precision
If manual curation methods are used to handle complex corporate action prospectuses with uncommon verbiage, then accurate extraction of key information is achieved, but large resource allocations are required
Solution Approach 1:
The patent introduces expression patterns as intermediaries between the raw text of corporate action prospectuses and the extraction process. These pre-defined patterns serve as mediators that capture uncommon verbiage and specialized terminology, enabling accurate information extraction without requiring human expertise. The scoring algorithm acts as another intermediary that objectively evaluates extracted attributes based on multiple criteria.
Solution Approach 2:
The system enables self-service by allowing the corporate action prospectus to extract its own key information through automated processing. The document's own linguistic structures and expression patterns are utilized by the system to identify and extract relevant attributes without external human intervention, thereby eliminating the need for manual curation resources.
Data Source
AI summary
A method for facilitating automated data curation in real-time is disclosed. The method includes retrieving electronic documents from a source; converting the electronic documents into data sets, the data sets corresponding to a predetermined format; preprocessing the data sets to identify linguistic units, the linguistic units relating to paragraphs and sentences; extracting, by using a model, attributes based on the linguistic units, the attributes relating to a key detail in the electronic documents; and generating, in real-time, messages based on the extracted attributes. Additionally, the electronic documents include a corporate action prospectus that provides information for a corresponding corporate event, the information including term and condition information, date information, and restriction information.


