Website Classification via Clickstream Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Standard text classification approaches for websites are inefficient and struggle with dynamically changing topics, and cannot effectively analyze web pages with video news and flash animations, requiring extensive labeled examples and being unable to classify content without reviewing the actual content of the websites.

Innovation Solution

The method uses clickstream data to classify websites by tracking user behavior and identifying semantic tags, allowing for classification without examining the content, using a system that generates categories, seed URLs, and analyzes clickstream data to determine user interest and select URLs for classification, enabling dynamic and efficient categorization of websites based on user navigation patterns.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If standard text classification approaches are used to classify websites, then classification accuracy can be achieved with representative examples, but the system becomes too intensive to execute efficiently and cannot handle dynamically changing topics

Engineering Contradiction:
Improveclassification accuracyVSAvoidexecution efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent extracts the core classification task from content analysis by removing the need to examine actual website content. Instead, it uses clickstream data to infer website categories, thereby eliminating the intensive text processing while maintaining classification capability through user behavior patterns

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces clickstream data as an intermediary between user behavior and website classification. This mediator allows the system to classify websites indirectly through navigation patterns rather than directly analyzing website content, improving efficiency while maintaining accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If traditional text classification techniques are used to analyze web pages, then text-based content can be categorized, but the system cannot effectively analyze web pages with video news and flash animations

Engineering Contradiction:
Improvecontent type coverageVSAvoidanalysis effectiveness
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent replaces the mechanical text analysis system with a behavioral data-based classification system. By substituting content examination with clickstream analysis, the system becomes adaptable to various content types including video and flash animations that cannot be effectively processed by traditional text classification

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If extensive labeled examples are used for website classification, then classification accuracy improves, but the system requires too many labeled examples to be practical

Engineering Contradiction:
Improveclassification accuracyVSAvoidnumber of labeled examples
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent enables the classification system to serve itself by automatically generating classifications through clickstream data analysis without requiring extensive manual labeling. The system uses inherent user navigation patterns to self-determine website categories, eliminating the need for large labeled datasets while maintaining accuracy

Inventive Principle:
Principle #25Self-service

4Measurement precision

If website content is reviewed for classification, then accurate content-based categorization is achieved, but the system cannot classify websites without examining their actual content

Engineering Contradiction:
Improvecontent categorization accuracyVSAvoidcontent examination requirement
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent inverts the traditional classification approach by not examining website content directly. Instead, it analyzes user clickstream data to infer website categories, achieving accurate content-based categorization through indirect behavioral evidence rather than direct content review

Inventive Principle:
Principle #13The other way round (Inversion)

Data Source

PatentUS9256692B2Clickstreams and website classification
Publication Date: 2016.02.09 MICRO FOCUS LLC
  • US9256692B2 patent drawing
  • US9256692B2 patent drawing
  • US9256692B2 patent drawing

AI summary

One embodiment is a method that receives a seed Uniform Resource Locator (URL) that represents a category for website classification. Clickstream data generated from the seed URL and additional URLs are analyzed to determine whether the additional URLs belong to the category. The method selects one or more of the additional URLs to represent the category.