Hierarchical Data Cleansing Tree for Multi-Record De-duplication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The 'side-by-side' approach to data de-duplication becomes inefficient and time-consuming as data set complexity and size increase, limiting users to comparing and processing only two data sets at a time.
Innovation Solution
A graphical user interface utilizing a hierarchical data cleansing tree allows users to modify, merge, and consolidate data records by dragging and dropping nodes, with features like 'Keep' functionality for copying data and automatic comparison of attributes, enabling efficient data management and validation within a unified user interface.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If the side-by-side approach is used for data de-duplication, then the user interface is simple to understand, but the efficiency and time consumption worsen as data set complexity and size increase
Solution Approach 1:
The data comparison process is segmented into multiple levels: the data cleansing tree provides a hierarchical overview of all data sets, while detailed comparison can be performed by selecting specific nodes. This segmentation allows users to manage complex data sets without being overwhelmed by displaying all data simultaneously, thus maintaining ease of operation while improving productivity.
Solution Approach 2:
The patent transitions from a two-dimensional side-by-side view to a three-dimensional hierarchical tree structure. This dimensional change enables users to navigate through multiple data sets simultaneously in a hierarchical manner, allowing comparison of many data sets without increasing interface complexity, thereby resolving the contradiction between simplicity and efficiency.
2Ease of operation
If the side-by-side approach is used, then the interface is easy to use, but users are limited to comparing and processing only two data sets at a time
Solution Approach 1:
The data cleansing tree serves multiple functions: it displays hierarchical structure of data sets, enables selection of multiple data sets for comparison, allows drilling down to detailed views, and supports various comparison operations. This multi-functionality enables comparison of any number of data sets while maintaining a consistent, easy-to-use interface paradigm.
Solution Approach 2:
The patent implements a nested structure where the data cleansing tree contains nested data sets, which in turn contain nested data elements. This nesting allows users to work with multiple data sets simultaneously at different hierarchical levels, effectively increasing adaptability while keeping the interface simple through consistent nested navigation patterns.
3Productivity
If multiple data sets are processed simultaneously, then productivity improves, but the user interface complexity increases
Solution Approach 1:
The data cleansing tree is a dynamic structure that adapts to the number and complexity of data sets. Users can dynamically expand or collapse tree nodes, select different levels of detail, and adjust the view based on current needs. This dynamic behavior allows the interface to handle multiple data sets efficiently without presenting unnecessary complexity, as the interface adapts to the actual work requirements.
Data Source
AI summary
A system and method of de-duplicating data using a graphical user interface application. The graphical user interface application represents a model of the selected data records in a data tree. The graphical user interface application processes a selected target data record and potential duplicates data records. Nodes representing the potential duplicate data records can be added to the target data record. Nodes representing the potential duplicate data records can also be dragged and dropped into a node of the target data record. Nodes from the target data record can also be removed from the target data record. Differences between data associated with multiple nodes can be graphically presented with the graphical user interface application when multiple nodes are selected. Changes made to the data tree in the graphical user interface are applied to data records stored in a database.


