Website Classification via Clickstream Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Standard text classification approaches for websites are inefficient and struggle with dynamically changing topics, and cannot effectively analyze web pages with video news and flash animations, requiring extensive labeled examples and being unable to classify content without reviewing the actual content of the websites.
Innovation Solution
The method uses clickstream data to classify websites by tracking user behavior and identifying semantic tags, allowing for classification without examining the content, using a system that generates categories, seed URLs, and analyzes clickstream data to determine user interest and select URLs for classification, enabling dynamic and efficient categorization of websites based on user navigation patterns.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If standard text classification approaches are used to classify websites, then classification accuracy can be achieved with representative examples, but the system becomes too intensive to execute efficiently and cannot handle dynamically changing topics
Solution Approach 1:
The patent extracts the core classification task from content analysis by removing the need to examine actual website content. Instead, it uses clickstream data to infer website categories, thereby eliminating the intensive text processing while maintaining classification capability through user behavior patterns
Solution Approach 2:
The patent introduces clickstream data as an intermediary between user behavior and website classification. This mediator allows the system to classify websites indirectly through navigation patterns rather than directly analyzing website content, improving efficiency while maintaining accuracy
2Adaptability or versatility
If traditional text classification techniques are used to analyze web pages, then text-based content can be categorized, but the system cannot effectively analyze web pages with video news and flash animations
Solution Approach 1:
The patent replaces the mechanical text analysis system with a behavioral data-based classification system. By substituting content examination with clickstream analysis, the system becomes adaptable to various content types including video and flash animations that cannot be effectively processed by traditional text classification
3Measurement precision
If extensive labeled examples are used for website classification, then classification accuracy improves, but the system requires too many labeled examples to be practical
Solution Approach 1:
The patent enables the classification system to serve itself by automatically generating classifications through clickstream data analysis without requiring extensive manual labeling. The system uses inherent user navigation patterns to self-determine website categories, eliminating the need for large labeled datasets while maintaining accuracy
4Measurement precision
If website content is reviewed for classification, then accurate content-based categorization is achieved, but the system cannot classify websites without examining their actual content
Solution Approach 1:
The patent inverts the traditional classification approach by not examining website content directly. Instead, it analyzes user clickstream data to infer website categories, achieving accurate content-based categorization through indirect behavioral evidence rather than direct content review
Data Source
AI summary
One embodiment is a method that receives a seed Uniform Resource Locator (URL) that represents a category for website classification. Clickstream data generated from the seed URL and additional URLs are analyzed to determine whether the additional URLs belong to the category. The method selects one or more of the additional URLs to represent the category.


