Unstructured Watchlist Data Enrichment for Screening Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data watchlist screening tools face challenges in accurately screening customers, vendors, and transactions due to poor data quality and lack of differentiation between data types, resulting in high false positive rates, making it impractical to check publicly available data without expensive subscriptions and manual intervention.

Innovation Solution

A system and method for labeling, categorizing, and enriching unstructured data by processing raw data through checksum calculation, sanity checks, stability checks, and data enrichment, transforming it into a fully searchable format without human intervention, using machine learning techniques to reduce false positives and improve data integrity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If unstructured watchlist data is processed without enrichment and structuring, then processing speed is maintained, but data quality and screening accuracy deteriorate, resulting in high false positive rates

Engineering Contradiction:
Improvescreening accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary data enrichment and structuring operations on watchlist data before screening operations commence. This includes parsing unstructured data, extracting relevant entities, standardizing formats, and organizing into structured databases in advance, so that when screening occurs, high-quality structured data is already available without adding complexity during the critical screening phase

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary data enrichment layer between raw unstructured watchlist data and the screening system. This intermediary component parses, validates, and structures the data into standardized formats, acting as a mediator that transforms low-quality unstructured data into high-quality structured data suitable for accurate screening without requiring the screening system itself to handle the complexity of unstructured data processing

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If expensive subscription services are used to access watchlist data, then data availability is improved, but cost increases and false positive rates remain high due to poor data quality

Engineering Contradiction:
Improvedata qualityVSAvoidcost
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system implements self-service data enrichment capabilities that automatically parse, validate, and structure watchlist data without requiring external expensive subscription services. The system autonomously accesses publicly available data sources, extracts relevant information, and structures it into usable formats, eliminating the need to pay premium prices for data while maintaining high data quality through automated enrichment processes

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical system of manually purchasing and managing expensive subscription-based data services with an automated computational system that freely accesses, parses, and enriches publicly available watchlist data. This substitution eliminates ongoing subscription costs while maintaining data quality through automated data processing and enrichment algorithms

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If manual intervention is used to verify watchlist data, then data accuracy is improved, but processing time and labor costs increase significantly

Engineering Contradiction:
Improvedata accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system replaces manual human verification processes with automated computational algorithms that parse, validate, and verify watchlist data accuracy. Machine learning models and validation rules automatically assess data quality, verify entities, and flag suspicious entries, achieving high data accuracy without the time-consuming and labor-intensive manual review processes that previously required human analysts to examine each entry

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an automated data validation intermediary layer that sits between data ingestion and final screening results. This intermediary component performs automated accuracy checks, cross-references data against multiple sources, and validates extracted entities before presenting results, providing the accuracy function previously performed by manual review but at automated speeds without human intervention

Inventive Principle:
Principle #24Intermediary (Mediator)

4Adaptability or versatility

If all watchlist data is screened without differentiation between data types, then comprehensive coverage is achieved, but false positive rates increase due to inability to differentiate between data quality

Engineering Contradiction:
Improvedata type differentiationVSAvoidfalse positive rate
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system applies local quality by treating different data types and sources with different processing rules and validation criteria. Instead of applying a uniform screening approach to all watchlist data, the system identifies specific data types (individuals, entities, vessels, aircraft) and applies type-specific parsing, validation, and screening rules, allowing each data category to be processed with the appropriate level of scrutiny and differentiation, thereby reducing false positives while maintaining comprehensive coverage

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12130815B2System and method for processing data for electronic searching
Publication Date: 2024.10.29 CASTELLUM AI
  • US12130815B2 patent drawing
  • US12130815B2 patent drawing
  • US12130815B2 patent drawing

AI summary

A system and method for labelling, categorizing, structuring, and enriching unstructured data is disclosed. First, raw data is received, and a checksum is calculated and compared with an existing checksum. The raw data is downloaded to an object storage, wherein it is then parsed and transmitted to a sanity check system. The sanity checked data is then loaded into a staging database and a stability check is performed. The stability checked data is then loaded into a master database and transmitted to a search platform. Post checks are performed, and a report is generated, followed by the system updating the watchlist checksum.