Dataset Interpretation via Rule Cover Clustering and Exception Handling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data interpretation techniques struggle to provide a comprehensive view of customer data across multiple disparate sources, often resulting in fragmented views and low coverage of relationship data, failing to leverage all customer data holistically and leading to redundant results that are difficult to interpret.

Innovation Solution

A system and method for interpreting datasets that compute a rule set with pre-determined consequents based on antecedents, generate a rule cover, calculate distances between rule pairs, cluster overlapping rules, select representative rules, and determine exceptions to provide comprehensive interpretations of datasets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If data interpretation techniques analyze relationships among stored data sets, then the meaning of data is provided for decision-making, but the results are fragmented and redundant making them difficult to interpret

Engineering Contradiction:
Improvecomprehensive coverage of relationship dataVSAvoidinterpretability of results
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent merges multiple rules that cover the same transactions by calculating overlap between rules and combining them into unified rule covers. This consolidation eliminates redundant results while preserving comprehensive coverage of relationships in the data, directly addressing the contradiction between information completeness and interpretability

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent segments the rule set into multiple rule covers based on transaction overlap thresholds. By dividing the comprehensive rule set into manageable segments with controlled overlap, the system maintains thorough relationship coverage while organizing results in an interpretable structure that avoids redundancy

Inventive Principle:
Principle #1Segmentation

2Loss of information

If existing techniques provide data interpretation across multiple disparate sources, then customer data analysis is performed, but the view remains fragmented with low coverage of relationship data

Engineering Contradiction:
Improvecoverage of customer relationship dataVSAvoidcomplexity of integrating disparate data sources
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent implements a universal rule cover framework that works across multiple disparate data sources by defining rules and overlaps in a source-agnostic manner. The system calculates transaction coverage and rule overlaps uniformly regardless of data source, enabling comprehensive customer relationship analysis without being constrained by source-specific complexities

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces rule covers as an intermediary layer between disparate data sources and interpretation outputs. This intermediary structure standardizes the representation of relationships across different sources, managing integration complexity while achieving comprehensive data coverage through the unified rule cover mechanism

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10579931B2Interpretation of a dataset for co-occurring itemsets using a cover rule and clustering
Publication Date: 2020.03.03 TATA CONSULTANCY SERVICES LTD
  • US10579931B2 patent drawing
  • US10579931B2 patent drawing

AI summary

A method and system for interpreting a dataset is described herein. The method include computing a rule set pertaining to the dataset, followed by generating a rule cover pertinent to a subset of the rule set. Further, a plurality of distances between the plurality of rule pairs in the rule cover is calculated and a distance matrix based on the calculated plurality of distances is generated. Consequently, the overlapping rules within the rule cover are clustered using the distance matrix and a representative rule from each cluster is selected. Further, at least one exception for each representative rule is determined and the dataset is interpreted using the representative rules and the at least one exception. Thereby, the method provides succinct results in terms of rules and exceptions along with multiple interpretations of the same set of transactions from the dataset, thereby providing a holistic view about the dataset.