Automated Text Novelty Detection for Financial Documents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for analyzing financial documents, such as SEC filings, are inefficient due to the lack of automated tools for processing unstructured textual information, requiring manual expertise and consuming significant time, especially in identifying market standard language and potential inaccuracies.

Innovation Solution

A system and method for analyzing clusters of conceptually-related text documents to develop a model document, calculate novelty measurements, and merge documents based on common neighbors similarity, providing direct access to a large corpus of documents scored for conformity and novelty to market standard language.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual analysis methods are used to process unstructured text in financial documents, then analysis accuracy can be maintained through expert judgment, but time consumption increases significantly and productivity decreases

Engineering Contradiction:
Improveanalysis accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent creates a model document that replicates the structure and language patterns of standard financial documents. This model serves as a template against which actual documents are compared, enabling automated identification of deviations from market standard language without requiring manual expert analysis of each document.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system transforms unstructured text analysis into a structured comparison process by defining specific parameters such as novelty measurements and common neighbors similarity metrics. These parameters enable quantitative automated comparison while maintaining the ability to detect subtle linguistic deviations that would indicate atypical language usage.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If automated processing is implemented to increase productivity, then processing speed improves, but the ability to detect subtle linguistic deviations and maintain analysis precision deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoiddetection accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent replaces manual mechanical analysis with an automated computational system that uses algorithmic comparisons. The system substitutes human expert judgment with automated novelty measurements and similarity calculations, achieving both speed and precision through mathematical rather than mechanical processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The model document serves as an intermediary between the actual financial documents and the analysis process. It mediates the comparison by providing a standardized reference that captures market standard language, allowing the automated system to detect deviations without needing to understand the full context of each document.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If comprehensive manual review is performed to ensure reliability, then detection accuracy improves, but time consumption increases and the system becomes less adaptable to large volumes of documents

Engineering Contradiction:
Improvedetection reliabilityVSAvoidscalability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent segments the analysis process into distinct components: creating a model document, calculating novelty measurements for individual documents, and computing common neighbors similarity. This segmentation allows the system to process documents independently and parallelize operations, enabling scalable handling of large document volumes while maintaining reliable detection through consistent application of the same analytical segments.

Inventive Principle:
Principle #1Segmentation

4Measurement precision

If expert manual analysis is used to identify market standard language, then measurement precision is maintained, but device complexity and operational complexity increase

Engineering Contradiction:
Improvelanguage identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system enables self-service by allowing automated processing without requiring expert manual intervention. The model document and automated comparison algorithms serve themselves to identify market standard language and detect deviations, eliminating the need for external expert analysis while maintaining precision through consistent algorithmic application.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS9690849B2Systems and methods for determining atypical language
Publication Date: 2017.06.27 THOMSON REUTERS ENTERPRISE CENTRE GMBH
  • US9690849B2 patent drawing
  • US9690849B2 patent drawing
  • US9690849B2 patent drawing

AI summary

A method includes analyzing a cluster of conceptually-related portions of text to develop a model and calculating a novelty measurement between a first identified conceptually-related portion of text and the model. The method further includes transmitting a second identified conceptually-related portion of text and a score associated with the novelty measurement from a server to an access device via a signal. Another method includes determining at least two corpora of conceptually-related portions of text. The method also includes calculating a common neighbors similarity measurement between the at least two corpora of conceptually-related portions of text and if the common neighbors similarity measurement exceeds a threshold, merging the at least two corpora of conceptually-related portions of text into a cluster or if the common neighbors similarity measurement does not exceed a threshold, maintaining a non-merge of the at least two corpora of conceptually-related portions of text.