Automated Data Curation for Corporate Action Prospectuses

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data curation techniques for generating corporate action prospectus narratives are resource-intensive and not timely due to the complexity and uncommon vocabulary in corporate action prospectuses, making real-time curation infeasible with reasonable resource allocations.

Innovation Solution

Implementing a method that uses natural language processing, expression patterns, and scoring algorithms to automatically curate prospectus and narrative data in real-time by retrieving electronic documents, converting them into machine-readable formats, preprocessing to identify linguistic units, extracting key attributes, and generating messages based on these attributes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional data curation techniques are used to generate message narratives from corporate action prospectuses, then the narratives can be produced with sufficient accuracy, but the process requires large resource allocations and cannot be completed in real-time

Engineering Contradiction:
Improvecuration accuracyVSAvoidcuration speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the corporate action prospectus into multiple linguistic units (paragraphs and sentences), and further segments the extraction process into multiple stages including preprocessing, attribute identification using expression patterns, scoring, and selection. This segmentation allows the system to process complex documents systematically without requiring excessive resources while maintaining accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter of resource allocation by implementing a scoring algorithm that ranks extracted attributes based on multiple criteria (keyword hit criterion, context criterion, word clustering criterion, word embedding criterion). This parameter change enables the system to prioritize processing of most relevant information, reducing overall resource requirements while maintaining curation quality.

Inventive Principle:
Principle #35Parameter changes

2Loss of energy

If conventional data curation techniques are used with reasonable resource allocations, then resource usage is optimized, but timely curation of message narratives becomes infeasible

Engineering Contradiction:
Improveresource consumptionVSAvoidcuration time
Core Design Contradiction:
Loss of energyVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-processing the corporate action prospectus to identify linguistic units and pre-defining expression patterns for attribute extraction before the actual curation process. This preliminary preparation enables faster processing during real-time message narrative generation without increasing resource consumption during the critical curation phase.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual mechanical curation processes with automated computational methods including natural language processing, expression pattern matching, and scoring algorithms. This substitution eliminates the need for human reviewers while maintaining curation quality, thereby reducing both resource consumption and time requirements.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If manual curation methods are used to handle complex corporate action prospectuses with uncommon verbiage, then accurate extraction of key information is achieved, but large resource allocations are required

Engineering Contradiction:
Improveinformation extraction accuracyVSAvoidresource allocation
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent introduces expression patterns as intermediaries between the raw text of corporate action prospectuses and the extraction process. These pre-defined patterns serve as mediators that capture uncommon verbiage and specialized terminology, enabling accurate information extraction without requiring human expertise. The scoring algorithm acts as another intermediary that objectively evaluates extracted attributes based on multiple criteria.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system enables self-service by allowing the corporate action prospectus to extract its own key information through automated processing. The document's own linguistic structures and expression patterns are utilized by the system to identify and extract relevant attributes without external human intervention, thereby eliminating the need for manual curation resources.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11947901B2Method and system for automated data curation
Publication Date: 2024.04.02 JPMORGAN CHASE BANK NA
  • US11947901B2 patent drawing
  • US11947901B2 patent drawing
  • US11947901B2 patent drawing

AI summary

A method for facilitating automated data curation in real-time is disclosed. The method includes retrieving electronic documents from a source; converting the electronic documents into data sets, the data sets corresponding to a predetermined format; preprocessing the data sets to identify linguistic units, the linguistic units relating to paragraphs and sentences; extracting, by using a model, attributes based on the linguistic units, the attributes relating to a key detail in the electronic documents; and generating, in real-time, messages based on the extracted attributes. Additionally, the electronic documents include a corporate action prospectus that provides information for a corresponding corporate event, the information including term and condition information, date information, and restriction information.