Heterogeneous Dataset Merge via Correlation Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Merging multi-dimensional heterogeneous datasets is challenging due to the difficulty in maintaining data cohesiveness and relativeness, as existing methods often introduce new dimensions that deteriorate the cohesiveness of the merged dataset.
Innovation Solution
A system and method utilizing a knowledge engine with an attribute manager, correlation manager, and merge manager to identify seed attributes, compute correlations, and iteratively amend mergeable attributes, leveraging cohesiveness characteristics to form a merged dataset that represents non-trivial similarities between datasets while minimizing the introduction of new dimensions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing methods merge heterogeneous datasets by introducing new dimensions, then the datasets are combined, but data cohesiveness deteriorates
Solution Approach 1:
The patent changes the parameter of dimension selection by using correlation-based filtering to identify and retain only highly correlated dimensions while removing redundant ones. This parameter change approach maintains data cohesiveness by ensuring that only dimensions with strong relationships are preserved in the merged dataset, rather than introducing all possible dimensions.
Solution Approach 2:
The patent replaces traditional mechanical merging approaches with a knowledge engine that uses correlation analysis and statistical measures. Instead of simply combining datasets with all dimensions, the system substitutes a sophisticated selection mechanism that evaluates dimension relationships and selectively merges only those dimensions that maintain cohesiveness.
2Measurement precision
If correlation analysis is performed on all attributes, then mergeable attributes are identified, but computational complexity increases
Solution Approach 1:
The patent segments the attribute correlation analysis into distinct phases: initial correlation computation, threshold-based filtering, and iterative refinement. By dividing the analysis into manageable segments with clear stopping criteria, the system achieves accurate correlation measurement without requiring exhaustive computation of all possible attribute combinations.
Solution Approach 2:
The patent applies partial action by computing correlations only for attributes that meet certain criteria or up to a specified threshold, rather than exhaustively analyzing all possible attribute pairs. This approach provides sufficient correlation accuracy for practical purposes while significantly reducing computational complexity.
3Loss of information
If multiple dimensions are introduced to represent dataset similarities, then more information is captured, but data cohesiveness deteriorates
Solution Approach 1:
The patent introduces correlation threshold as an intermediary mechanism that mediates between information retention and cohesiveness preservation. By using the threshold as a filtering criterion, the system selectively admits dimensions that meet the cohesiveness requirement while blocking those that would deteriorate it, thus resolving the contradiction through an intermediary selection layer.
Data Source
AI summary
Embodiments relate to a system, computer program product, and method to merge two or more heterogeneous datasets. Seed attributes of each dataset that is the subject of the merge are identified. The seed attributes are derived from candidate attributes of the respective datasets. A correlation is assessed to create a set of mergeable attributes and a set of non-mergeable attributes. A cohesiveness characteristic is leveraged to iteratively identify one or more attributes from the set of non-mergeable attributes, and to amend the set of mergeable attributes with the one or more attributes identified in the set of non-mergeable attributes. A merged dataset based on the amended set of mergeable attributes and representing non-trivial similarities between the first and second dataset is formed as output.


