Watchlist Entity Resolution Using NLP Identity Profiles
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional watchlist identification methods rely heavily on manual review of structured data sources, leading to errors due to name misspellings, incorrect PII, and lack of consideration for aliases, while external unstructured data sources like social media and news articles are underutilized due to difficulty in normalization and relevance determination.
Innovation Solution
A system and method utilizing natural language processing and machine learning to process and normalize external data sources, constructing holistic identity profiles through unsupervised semantic identity modeling and graph-based clustering, enabling accurate watchlist matching and transaction risk assessment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual review of structured data sources is used for watchlist identification, then simplicity of implementation is maintained, but accuracy deteriorates due to name misspellings, incorrect PII, and lack of consideration for aliases
Solution Approach 1:
The patent replaces manual mechanical review processes with automated natural language processing and machine learning systems. The NLP model automatically extracts and compares identity characteristics from unstructured text sources, while the machine learning system performs semantic matching that accounts for misspellings, aliases, and variations in personally identifiable information, thereby maintaining high accuracy without manual intervention
Solution Approach 2:
The system transforms traditional exact-match parameters into semantic similarity parameters. Instead of requiring exact string matching for names and identifiers, the system uses semantic equivalence detection that evaluates meaning-based similarity, allowing matches even when names are misspelled or presented in different formats, thus improving accuracy while managing complexity through parameter transformation
2Loss of information
If external unstructured data sources like social media and news articles are utilized, then data completeness is improved, but difficulty in normalization and relevance determination increases
Solution Approach 1:
The patent employs natural language processing to automatically extract and normalize information from unstructured external data sources. The NLP system parses social media posts, news articles, and other unstructured text to identify and standardize identity characteristics, replacing manual data normalization processes and enabling efficient utilization of external data sources
Solution Approach 2:
The system introduces an intermediary NLP processing layer between raw external data sources and the watchlist matching system. This intermediary automatically extracts relevant identity information from unstructured sources, normalizes it to consistent formats, and filters for relevance before feeding it into the matching algorithm, thereby reducing the complexity of direct data integration
3Loss of information
If traditional structured data sources only are used, then data consistency is maintained, but data richness deteriorates
Solution Approach 1:
The patent creates a universal data processing framework that can handle multiple data types and sources through a common NLP interface. The system processes both structured traditional data and unstructured external data using the same extraction and normalization pipeline, enabling rich data integration while maintaining consistency through unified processing rules
Solution Approach 2:
The system transforms diverse data formats from various sources into a standardized parameter set. By converting different data types (structured databases, unstructured text, social media posts) into common identity characteristic parameters, the system enriches the available information while maintaining consistency through parameter standardization
4Measurement precision
If human intervention is used to assess external data content, then relevance determination accuracy is improved, but processing speed deteriorates
Solution Approach 1:
The patent replaces manual human assessment of external data relevance with automated machine learning models. The system uses trained algorithms to evaluate the relevance of extracted information from social media and news articles, achieving both high accuracy through sophisticated pattern recognition and high speed through automated processing, eliminating the speed-accuracy trade-off of manual review
Data Source
AI summary
Provided are a method and system for identity correlation between a transaction applicant (TA) and a watchlist entity (WE). Preexisting watchlist data and other aggregated identity data (AID) are processed to provide for comparison to a collective identity of at least the TA. Using various categorizations for the AID and the collective identity, watchlist tags are generated that can then be matched to the collective identity. As a result of the matching, a watchlist candidacy demonstrating a probability that the identity of the TA does or does not correspond to that of the WE can be generated. Natural language processing may be used in connection with sourced textual information such as from news articles, social media posts and the like to supplement the AID available to the system and to further enhance matching functionality.


