Hierarchical Data Cleansing Tree for Multi-Record De-duplication

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The 'side-by-side' approach to data de-duplication becomes inefficient and time-consuming as data set complexity and size increase, limiting users to comparing and processing only two data sets at a time.

Innovation Solution

A graphical user interface utilizing a hierarchical data cleansing tree allows users to modify, merge, and consolidate data records by dragging and dropping nodes, with features like 'Keep' functionality for copying data and automatic comparison of attributes, enabling efficient data management and validation within a unified user interface.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If the side-by-side approach is used for data de-duplication, then the user interface is simple to understand, but the efficiency and time consumption worsen as data set complexity and size increase

Engineering Contradiction:
Improveuser interface simplicityVSAvoiddata processing efficiency
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The data comparison process is segmented into multiple levels: the data cleansing tree provides a hierarchical overview of all data sets, while detailed comparison can be performed by selecting specific nodes. This segmentation allows users to manage complex data sets without being overwhelmed by displaying all data simultaneously, thus maintaining ease of operation while improving productivity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from a two-dimensional side-by-side view to a three-dimensional hierarchical tree structure. This dimensional change enables users to navigate through multiple data sets simultaneously in a hierarchical manner, allowing comparison of many data sets without increasing interface complexity, thereby resolving the contradiction between simplicity and efficiency.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Ease of operation

If the side-by-side approach is used, then the interface is easy to use, but users are limited to comparing and processing only two data sets at a time

Engineering Contradiction:
Improveinterface usabilityVSAvoiddata set comparison capacity
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The data cleansing tree serves multiple functions: it displays hierarchical structure of data sets, enables selection of multiple data sets for comparison, allows drilling down to detailed views, and supports various comparison operations. This multi-functionality enables comparison of any number of data sets while maintaining a consistent, easy-to-use interface paradigm.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements a nested structure where the data cleansing tree contains nested data sets, which in turn contain nested data elements. This nesting allows users to work with multiple data sets simultaneously at different hierarchical levels, effectively increasing adaptability while keeping the interface simple through consistent nested navigation patterns.

Inventive Principle:
Principle #7Nested doll (Nesting)

3Productivity

If multiple data sets are processed simultaneously, then productivity improves, but the user interface complexity increases

Engineering Contradiction:
Improvedata processing throughputVSAvoiduser interface complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The data cleansing tree is a dynamic structure that adapts to the number and complexity of data sets. Users can dynamically expand or collapse tree nodes, select different levels of detail, and adjust the view based on current needs. This dynamic behavior allows the interface to handle multiple data sets efficiently without presenting unnecessary complexity, as the interface adapts to the actual work requirements.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS9471609B2Data cleansing tool with new cleansing tree
Publication Date: 2016.10.18 SAP SE
  • US9471609B2 patent drawing
  • US9471609B2 patent drawing
  • US9471609B2 patent drawing

AI summary

A system and method of de-duplicating data using a graphical user interface application. The graphical user interface application represents a model of the selected data records in a data tree. The graphical user interface application processes a selected target data record and potential duplicates data records. Nodes representing the potential duplicate data records can be added to the target data record. Nodes representing the potential duplicate data records can also be dragged and dropped into a node of the target data record. Nodes from the target data record can also be removed from the target data record. Differences between data associated with multiple nodes can be graphically presented with the graphical user interface application when multiple nodes are selected. Changes made to the data tree in the graphical user interface are applied to data records stored in a database.