Web Mining Clustering via Minimum Description Length Cost Partitioning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional web mining techniques require manual intervention and are subjective, leading to erroneous patterns and a lack of scalability, as they fail to adapt to dynamic changes in online user behavior and content.
Innovation Solution
A web mining and clustering method based on Minimum Description Length (MDL) principles, which divides input data into primitive datasets, generates models using a grammar generator, computes costs, and partitions datasets into clusters to identify objective patterns with minimal assumptions, allowing for dynamic adaptation and scalability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual intervention by domain experts is used for establishing context in web mining, then patterns can be identified with domain knowledge, but the process becomes subjective and requires significant human time and effort
Solution Approach 1:
The system enables automatic web mining and pattern discovery without requiring manual domain expert intervention. The MDL-based algorithm autonomously processes web usage data, generates patterns, and identifies contexts, allowing the system to serve itself rather than relying on external human expertise for each mining task.
Solution Approach 2:
The patent replaces the mechanical process of manual domain expert analysis with an automated computational system based on Minimum Description Length principles. The algorithm systematically evaluates data compressibility and pattern significance, substituting human cognitive processes with objective mathematical calculations that eliminate subjectivity and reduce time requirements.
2Reliability
If conventional web mining techniques are used, then patterns can be identified, but they do not adapt to dynamic changes in user behavior and content
Solution Approach 1:
The system implements dynamic web mining that continuously adapts to changing user behavior and web content. The MDL-based algorithm can process evolving data streams and automatically adjust to new patterns, making the system flexible and responsive to dynamic changes in the web environment rather than relying on static, pre-defined contexts.
Solution Approach 2:
The system incorporates feedback mechanisms where mining results are continuously evaluated and used to refine future mining operations. The objective MDL-based evaluation provides feedback on pattern quality and data compressibility, allowing the system to iteratively improve its pattern discovery and adapt to changing conditions in web usage data.
3Productivity
If conventional web mining techniques are used, then patterns can be discovered, but they generate erroneous patterns due to inaccurate assumptions
Solution Approach 1:
The patent fundamentally changes the evaluation parameter from subjective domain expert judgment to objective Minimum Description Length measurement. By using data compressibility as the primary criterion for pattern evaluation, the system eliminates inaccurate assumptions inherent in conventional approaches and provides reliable, objective pattern discovery without compromising productivity.
4Quantity of substance
If web mining systems process large amounts of web data, then comprehensive patterns can be identified, but the system complexity increases
Solution Approach 1:
The system extracts and focuses on the essential characteristic of web usage data - patterns that can be described with minimal information. By applying MDL principles, the system identifies only the most significant patterns that capture the essence of user behavior, thereby processing large data volumes without proportionally increasing system complexity.
Data Source
AI summary
A method and system for web mining and clustering is described. The method includes receiving and dividing input data into a plurality of primitive datasets. Additionally, one or more combinations of the plurality of primitive datasets may be created. Further, a model for each primitive dataset in the plurality of primitive datasets and each of the one or more combinations of the plurality of primitive datasets may be generated. Subsequently, a cost associated with a model corresponding to each primitive dataset in the plurality of primitive datasets, and each of the one or more combinations of the plurality of primitive datasets may be computed. Further, a sum of the costs associated with the models corresponding to each primitive dataset in the plurality of primitive datasets may be compared with the cost associated with each model corresponding to each of the one or more combinations of the plurality of primitive datasets. Finally, the plurality of primitive datasets may be partitioned into one or more clusters based on the comparison of the costs such that each primitive dataset is a part of a cluster in the one or more clusters or a stand-alone primitive dataset.


