Trade Line De-duplication via Supervised Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data de-duplication methods in credit reporting, such as 'Pick and Choose' and 'List and Stack,' are inefficient and prone to human error, leading to skewed results and inconsistencies in credit evaluations due to reliance on manual interpretation and emphasis on account number matches, which can slow down the credit evaluation process and provide an incomplete or inaccurate representation of a consumer's credit history.
Innovation Solution
A supervised machine learning method that compares attributes of trade lines across multiple industry reports, using a classifier trained with user feedback to identify and remove duplicates, thereby providing a more accurate and balanced credit picture by automating the de-duplication process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual interpretation methods (Pick and Choose, List and Stack) are used for de-duplication, then human specialists can review data, but the credit evaluation process slows down and produces inconsistent results
Solution Approach 1:
The patent replaces manual mechanical review processes with an automated machine learning system that uses supervised learning algorithms to identify and remove duplicate trade lines. The system substitutes human specialists' manual interpretation with computer-based attribute comparison and classification, thereby maintaining de-duplication accuracy while dramatically increasing credit evaluation speed.
Solution Approach 2:
The system enables self-service de-duplication by automatically comparing trade line attributes, identifying duplicates, and removing them without requiring human specialist intervention. The machine learning model independently performs the de-duplication task that previously required manual review, improving both efficiency and consistency.
2Ease of manufacture
If account number matching is emphasized for de-duplication, then duplicate identification becomes simpler, but false positives increase and accurate representation of credit history is compromised
Solution Approach 1:
The patent changes the parameters used for de-duplication from relying primarily on account number matching to using a comprehensive set of trade line attributes including account type, balance, payment amount, credit limit, and other financial characteristics. This multi-parameter approach reduces false positives while maintaining de-duplication simplicity through automated attribute comparison.
Solution Approach 2:
The system employs a universal de-duplication approach that can handle multiple types of trade lines (credit cards, mortgages, auto loans, etc.) by comparing multiple attributes simultaneously. This multi-functional attribute comparison system accurately identifies duplicates across diverse credit products without relying solely on account number matching, thereby improving credit history accuracy.
3Measurement precision
If blend methodology is used to merge duplicate data, then the most accurate data for each element is provided, but the process becomes complex and requires many different implementations
Solution Approach 1:
The patent extracts and removes duplicate trade lines before blending, rather than attempting to blend all duplicate data. By first identifying and eliminating duplicates through attribute comparison, the system simplifies the process while maintaining data accuracy. This extraction approach reduces implementation complexity compared to comprehensive blending methodologies.
Solution Approach 2:
The system segments the de-duplication process into distinct phases: attribute comparison, duplicate identification, and removal. This segmentation simplifies the overall process by breaking down the complex blending operation into manageable steps, reducing system complexity while preserving data accuracy through systematic attribute-based comparison.
Data Source
AI summary
A system, method, and computer program includes a communications interface configured to receive a set of industry reports from multiple industry sources, and circuitry to compare one or more attributes of at least two trade lines to identify whether the at least two trade lines are duplicates. The circuitry characterizes as a binary indication whether the comparing indicates the one or more attributes are a match, and display a representation of the binary indication and receive a user-identified indication whether the at least two trade lines are duplicates. The circuitry trains a classifier, records the indication whether the at least two trade lines are duplicates and removes at least one of the at least two trade lines from the set of industry reports, and runs the classifier. Subsequently, a supervised machine learning classifier is trained in fit on the training data and is evaluated for accuracy of the testing data.


