Firmographic Database Aggregation via Clustering and Voting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data curation methods fail to effectively aggregate noisy, large, and overlapping datasets into a reliable master firmographic database, leading to inaccuracies and inconsistencies in business entity profiles.

Innovation Solution

A system and method that normalize, filter, clean, and deduplicate firmographic records by clustering and voting on values, generating master identifiers, and merging them into a unified database, while accounting for source reliability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If noisy datasets from multiple sources are aggregated directly, then the quantity of data increases, but data quality and accuracy deteriorate

Engineering Contradiction:
Improvequantity of dataVSAvoiddata quality
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent segments the data aggregation process into distinct stages: normalization, cleaning, clustering, voting, and master record generation. Each stage processes specific aspects of the data separately, allowing quality control at each step while maintaining overall data quantity growth.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediary processing steps (normalization layer, cleaning rules, clustering algorithms) between the raw noisy datasets and the final master database. These intermediaries filter and transform data to improve quality while preserving quantity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If data from multiple sources is merged without processing, then data completeness improves, but inconsistencies and noise increase

Engineering Contradiction:
Improvedata completenessVSAvoidinconsistencies and noise
Core Design Contradiction:
Loss of informationVSObject-affected harmful factors

Solution Approach 1:

The patent applies preliminary actions to data before final merging: normalization standardizes formats, cleaning removes obvious errors, and clustering groups related records. These preliminary processing steps ensure data completeness is achieved without introducing or amplifying inconsistencies.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent converts the harmful effect of noisy data into a benefit by using voting algorithms where multiple noisy sources collectively determine the most likely correct value. The noise is transformed into probabilistic information that, when aggregated, reveals the true signal.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

3Measurement precision

If extensive data cleaning and validation processes are applied, then data accuracy improves, but processing time and complexity increase

Engineering Contradiction:
Improvedata accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial cleaning actions focused on the most critical quality issues rather than exhaustive processing of all data aspects. The voting mechanism applies excessive action selectively only to clustered records needing resolution, balancing accuracy improvement with time constraints.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent changes parameters of the cleaning process dynamically: voting thresholds, clustering parameters, and validation strictness are adjusted based on data characteristics and processing requirements, optimizing the balance between accuracy and processing time.

Inventive Principle:
Principle #35Parameter changes

4Loss of information

If multiple data sources are integrated, then information coverage improves, but difficulty in managing overlapping records increases

Engineering Contradiction:
Improveinformation coverageVSAvoidrecord management complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent merges overlapping records from multiple sources by clustering them together and applying voting algorithms to determine the master record. This combining approach manages complexity by treating overlapping records as a single unit rather than separate entities requiring individual management.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The voting mechanism allows the data itself to resolve overlaps automatically based on consistency and frequency patterns, reducing the need for manual intervention and complex management systems. The system self-organizes overlapping records through algorithmic decision-making.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12111852B2Aggregation of noisy datasets into master firmographic database
Publication Date: 2024.10.08 6SENSE INSIGHTS INC
  • US12111852B2 patent drawing
  • US12111852B2 patent drawing
  • US12111852B2 patent drawing

AI summary

Aggregation of noisy datasets into a master firmographic database. In an embodiment, firmographic records are received from a plurality of sources, and normalized into a common schema. One or more firmographic records may be cleaned by replacing a value of one or more fields in those firmographic record(s) with a value of those field(s) in another firmographic record. The firmographic records may then be clustered, and each of the clusters may be collapsed into a single conflated firmographic record based on a voting process. A master identifier may be generated for each conflated firmographic record, and the conflated firmographic records may be merged into a master firmographic database that is indexed by master identifiers.