Machine Learning Data Identification System for Financial Alerts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The manual search and review of unstructured data from various sources, such as news articles and social media, is time-consuming and delays the identification of data of interest, particularly in financial institutions where timely decision-making is crucial.
Innovation Solution
A system that retrieves unstructured data from internet sources and inputs it into multiple machine learning models, including Naïve Bayes, LSTM, NER, SRL, and GBRT, to calculate sentiment scores and identify data of interest, generating data alerts for financial decisioning systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual search and review of unstructured data is used, then information can be searched, but the process is time-consuming and delays identification of data of interest
Solution Approach 1:
The patent replaces the manual mechanical search and review process with automated machine learning models. Multiple ML models (including NLP, classification, and information extraction models) automatically analyze unstructured data from internet sources, eliminating the need for human reviewers to manually scan and evaluate information, thus dramatically reducing time loss while maintaining or improving identification accuracy.
Solution Approach 2:
The system enables self-service by allowing the machine learning models to autonomously retrieve, analyze, and identify data of interest without human intervention. The models automatically process unstructured data, generate relevance scores, and present findings, making the information identification process self-sufficient and eliminating dependency on manual human effort.
2Measurement precision
If multiple machine learning models are used to identify data of interest, then accuracy is improved, but system complexity increases
Solution Approach 1:
The patent segments the identification task into multiple specialized machine learning models, each handling specific aspects of data analysis (e.g., NLP for text processing, classification for categorization, information extraction for entity recognition). This segmentation allows each model to focus on its strength, improving overall accuracy while organizing system complexity into manageable, modular components that can be independently trained and maintained.
Solution Approach 2:
The patent merges the outputs of multiple machine learning models into a unified identification system. By combining the results from different models (such as integrating NLP analysis, classification scores, and entity extraction data), the system achieves enhanced identification accuracy that leverages the complementary strengths of each model while presenting a unified interface to users.
3Productivity
If automated machine learning systems are implemented, then manual effort is reduced, but the initial setup and training requirements increase complexity
Solution Approach 1:
The patent applies preliminary action by pre-training the machine learning models with extensive datasets before deployment. The models are trained in advance on curated training data to learn patterns and relationships, so that when deployed for actual data identification, they can immediately begin automated analysis without requiring real-time human intervention for training or setup, thus achieving high productivity from the start.
Data Source
AI summary
Systems and methods for identifying data of interest are disclosed. The system may retrieve unstructured data from an internet data source via an alert system or RSS feed. The system may input the unstructured data into various models and scoring systems to determine whether the data is of interest. The models and scoring systems may be executed in order or in parallel. For example, the system may input the unstructured data into a Naïve Bayes machine learning model, a long short-term memory (LSTM) machine learning model, a named entity recognition (NER) model, a semantic role labeling (SRL) model, a sentiment scoring algorithm, and/or a gradient boosted regression tree (GBRT) machine learning model. Based on determining that the unstructured data is of interest, a data alert may be generated and transmitted for manual review or as part of an automated decisioning process.


