Master Ingestion Automation for Rapid Corporate Entity Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data matching processes are incapable of processing name changes within an adequate time frame, leading to undesirable delays in corporate identity number creation, duplication of records, and incorrectly assigned trade data, missing critical events like bankruptcies and merger and acquisition activities.

Innovation Solution

A master ingestion and data automation system (MIDAS) that includes a collection services unit and an evaluation and decisioning unit, utilizing serverless architecture with Google Compute Engine and GKE clusters, to identify and match trade styles immediately, reducing latency in corporate identity number creation to a few hours.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If conventional data matching processes are used, then existing systems can process data, but they cannot process name changes within an adequate time frame, leading to delays in corporate identity number creation

Engineering Contradiction:
Improveprocessing speed of name changesVSAvoidtime delay in corporate identity number creation
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

The system performs preliminary actions by immediately capturing and storing trade style data at the moment it becomes available, before any matching or processing occurs. This preliminary capture enables subsequent rapid processing and matching against existing entities, reducing the overall time delay in corporate identity number creation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system maintains continuous operation through automated ingestion and processing of trade style data. The collection services unit continuously monitors for new data, and the evaluation and decisioning unit continuously processes matches, ensuring no gaps in processing and enabling timely response to name changes.

Inventive Principle:
Principle #20Continuity of useful action

2Reliability

If conventional intelligence engines are used, then data can be processed, but they cause duplication of records and incorrectly assigned trade data due to inability to handle name changes timely

Engineering Contradiction:
Improveaccuracy of corporate entity matchingVSAvoidduplicate records and incorrectly assigned trade data
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The system segments the data processing function into distinct specialized units: a collection services unit that captures trade style data, an evaluation and decisioning unit that processes matches, and a publisher unit that outputs results. This segmentation allows each unit to focus on specific tasks, improving overall reliability and reducing errors in matching.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system implements feedback mechanisms where the evaluation and decisioning unit continuously compares incoming trade style data against existing entity records and adjusts matching decisions based on the comparisons. This feedback loop ensures accurate identification of name changes and prevents duplication of records.

Inventive Principle:
Principle #23Feedback

3Productivity

If manual intervention is used in corporate identity number creation, then flexibility can be maintained, but it increases processing time and reduces productivity

Engineering Contradiction:
Improvespeed of corporate identity number creationVSAvoidmanual intervention in data processing
Core Design Contradiction:
ProductivityVSExtent of automation

Solution Approach 1:

The system performs self-service through automated processes where the collection services unit automatically captures trade style data, the evaluation and decisioning unit automatically processes matches against existing entities, and the publisher unit automatically outputs results. This eliminates the need for manual intervention while maintaining high productivity in corporate identity number creation.

Inventive Principle:
Principle #25Self-service

4Reliability

If trade style data is not captured immediately, then system complexity is reduced, but it leads to missed critical events such as bankruptcies and merger and acquisition activities

Engineering Contradiction:
Improvedetection of critical eventsVSAvoidcomplexity of data capture and processing system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The collection services unit performs preliminary action by immediately capturing and storing trade style data at the moment it becomes available. This preliminary capture ensures that no critical events are missed, as the data is ready for immediate processing and matching against existing entities.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system achieves multi-functionality by using a unified collection services unit that handles multiple data types and a centralized evaluation and decisioning unit that processes various matching scenarios. This universal approach detects multiple critical events (bankruptcies, mergers, acquisitions) through a single integrated process, reducing the need for separate specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250217322A1Master ingestion and data automation process and system
Publication Date: 2025.07.03 THE DUN & BRADSTREET CORP
  • US20250217322A1 patent drawing
  • US20250217322A1 patent drawing
  • US20250217322A1 patent drawing

AI summary

A master ingestion and data automation system which comprises: a source ingestion module which receives incoming files of meta data driven framework (MDD) which standardizes and normalizes data using source-specific rules; a recognizer engine which receives the incoming files to poll and recognize new files into the master ingestion and data automation system as recognized files; a loader which loads the recognized files into staging tables; a collection services engine, wherein the collection services engine processes the recognized files by: (i) collecting and/or creating at least one cluster of legal entities, names, and/or addresses; (ii) collecting data and capturing insights other than the legal entities, names, and/or addresses; and (iii) cluster data based on rules without storing the data redundantly; an evaluation and decisioning module that retrieves all of the data from the staging tables and/or the cluster data to determine which data to publish; and a publisher module which publishes output decisions from the evaluation and decisioning module.