Firmographic Database Aggregation via Clustering and Voting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data curation methods fail to effectively aggregate noisy, large, and overlapping datasets into a reliable master firmographic database, leading to inaccuracies and inconsistencies in business entity profiles.
Innovation Solution
A system and method that normalize, filter, clean, and deduplicate firmographic records by clustering and voting on values, generating master identifiers, and merging them into a unified database, while accounting for source reliability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If noisy datasets from multiple sources are aggregated directly, then the quantity of data increases, but data quality and accuracy deteriorate
Solution Approach 1:
The patent segments the data aggregation process into distinct stages: normalization, cleaning, clustering, voting, and master record generation. Each stage processes specific aspects of the data separately, allowing quality control at each step while maintaining overall data quantity growth.
Solution Approach 2:
The patent introduces intermediary processing steps (normalization layer, cleaning rules, clustering algorithms) between the raw noisy datasets and the final master database. These intermediaries filter and transform data to improve quality while preserving quantity.
2Loss of information
If data from multiple sources is merged without processing, then data completeness improves, but inconsistencies and noise increase
Solution Approach 1:
The patent applies preliminary actions to data before final merging: normalization standardizes formats, cleaning removes obvious errors, and clustering groups related records. These preliminary processing steps ensure data completeness is achieved without introducing or amplifying inconsistencies.
Solution Approach 2:
The patent converts the harmful effect of noisy data into a benefit by using voting algorithms where multiple noisy sources collectively determine the most likely correct value. The noise is transformed into probabilistic information that, when aggregated, reveals the true signal.
3Measurement precision
If extensive data cleaning and validation processes are applied, then data accuracy improves, but processing time and complexity increase
Solution Approach 1:
The patent applies partial cleaning actions focused on the most critical quality issues rather than exhaustive processing of all data aspects. The voting mechanism applies excessive action selectively only to clustered records needing resolution, balancing accuracy improvement with time constraints.
Solution Approach 2:
The patent changes parameters of the cleaning process dynamically: voting thresholds, clustering parameters, and validation strictness are adjusted based on data characteristics and processing requirements, optimizing the balance between accuracy and processing time.
4Loss of information
If multiple data sources are integrated, then information coverage improves, but difficulty in managing overlapping records increases
Solution Approach 1:
The patent merges overlapping records from multiple sources by clustering them together and applying voting algorithms to determine the master record. This combining approach manages complexity by treating overlapping records as a single unit rather than separate entities requiring individual management.
Solution Approach 2:
The voting mechanism allows the data itself to resolve overlaps automatically based on consistency and frequency patterns, reducing the need for manual intervention and complex management systems. The system self-organizes overlapping records through algorithmic decision-making.
Data Source
AI summary
Aggregation of noisy datasets into a master firmographic database. In an embodiment, firmographic records are received from a plurality of sources, and normalized into a common schema. One or more firmographic records may be cleaned by replacing a value of one or more fields in those firmographic record(s) with a value of those field(s) in another firmographic record. The firmographic records may then be clustered, and each of the clusters may be collapsed into a single conflated firmographic record based on a voting process. A master identifier may be generated for each conflated firmographic record, and the conflated firmographic records may be merged into a master firmographic database that is indexed by master identifiers.


