Web Mining Clustering via Minimum Description Length Cost Partitioning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional web mining techniques require manual intervention and are subjective, leading to erroneous patterns and a lack of scalability, as they fail to adapt to dynamic changes in online user behavior and content.

Innovation Solution

A web mining and clustering method based on Minimum Description Length (MDL) principles, which divides input data into primitive datasets, generates models using a grammar generator, computes costs, and partitions datasets into clusters to identify objective patterns with minimal assumptions, allowing for dynamic adaptation and scalability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual intervention by domain experts is used for establishing context in web mining, then patterns can be identified with domain knowledge, but the process becomes subjective and requires significant human time and effort

Engineering Contradiction:
Improvepattern identification accuracyVSAvoidmanual intervention time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables automatic web mining and pattern discovery without requiring manual domain expert intervention. The MDL-based algorithm autonomously processes web usage data, generates patterns, and identifies contexts, allowing the system to serve itself rather than relying on external human expertise for each mining task.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of manual domain expert analysis with an automated computational system based on Minimum Description Length principles. The algorithm systematically evaluates data compressibility and pattern significance, substituting human cognitive processes with objective mathematical calculations that eliminate subjectivity and reduce time requirements.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If conventional web mining techniques are used, then patterns can be identified, but they do not adapt to dynamic changes in user behavior and content

Engineering Contradiction:
Improvepattern validityVSAvoidadaptability to dynamic changes
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system implements dynamic web mining that continuously adapts to changing user behavior and web content. The MDL-based algorithm can process evolving data streams and automatically adjust to new patterns, making the system flexible and responsive to dynamic changes in the web environment rather than relying on static, pre-defined contexts.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system incorporates feedback mechanisms where mining results are continuously evaluated and used to refine future mining operations. The objective MDL-based evaluation provides feedback on pattern quality and data compressibility, allowing the system to iteratively improve its pattern discovery and adapt to changing conditions in web usage data.

Inventive Principle:
Principle #23Feedback

3Productivity

If conventional web mining techniques are used, then patterns can be discovered, but they generate erroneous patterns due to inaccurate assumptions

Engineering Contradiction:
Improvepattern discovery capabilityVSAvoidpattern accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent fundamentally changes the evaluation parameter from subjective domain expert judgment to objective Minimum Description Length measurement. By using data compressibility as the primary criterion for pattern evaluation, the system eliminates inaccurate assumptions inherent in conventional approaches and provides reliable, objective pattern discovery without compromising productivity.

Inventive Principle:
Principle #35Parameter changes

4Quantity of substance

If web mining systems process large amounts of web data, then comprehensive patterns can be identified, but the system complexity increases

Engineering Contradiction:
Improvedata processing volumeVSAvoidsystem complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system extracts and focuses on the essential characteristic of web usage data - patterns that can be described with minimal information. By applying MDL principles, the system identifies only the most significant patterns that capture the essence of user behavior, thereby processing large data volumes without proportionally increasing system complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS8521773B2System and method for web mining and clustering
Publication Date: 2013.08.27 NBCUNIVERSAL MEDIA LLC
  • US8521773B2 patent drawing
  • US8521773B2 patent drawing
  • US8521773B2 patent drawing

AI summary

A method and system for web mining and clustering is described. The method includes receiving and dividing input data into a plurality of primitive datasets. Additionally, one or more combinations of the plurality of primitive datasets may be created. Further, a model for each primitive dataset in the plurality of primitive datasets and each of the one or more combinations of the plurality of primitive datasets may be generated. Subsequently, a cost associated with a model corresponding to each primitive dataset in the plurality of primitive datasets, and each of the one or more combinations of the plurality of primitive datasets may be computed. Further, a sum of the costs associated with the models corresponding to each primitive dataset in the plurality of primitive datasets may be compared with the cost associated with each model corresponding to each of the one or more combinations of the plurality of primitive datasets. Finally, the plurality of primitive datasets may be partitioned into one or more clusters based on the comparison of the costs such that each primitive dataset is a part of a cluster in the one or more clusters or a stand-alone primitive dataset.