Web Browsing History Classification via Keyword Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing web browsing history classification systems are inefficient, lacking user-friendly management and fast access, often requiring user interaction and consuming excessive memory due to reliance on cookies, which limits the classification and judicial usage of memory space.

Innovation Solution

A method and system that utilize machine learning and natural language processing to extract keywords from web pages, generate relevancy matrices, and store classifications in a non-volatile storage unit, allowing for automatic classification and efficient retrieval of web browsing history without user intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If cookies are used to store browsing history classification, then classification can be maintained, but storage space is minimal and cookies run out of space as indexes build up

Engineering Contradiction:
Improveclassification maintenanceVSAvoidstorage space
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts the classification data from cookies and stores it separately in a dedicated data structure on the user's device. This allows the browsing history to be classified without being constrained by cookie storage limitations, while still maintaining the classification information reliably.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs preliminary classification of browsing history into categories (e.g., news, social media, shopping) before storing the data. This preliminary organization enables efficient retrieval and management of browsing history without requiring extensive cookie storage space.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If manual user interaction is used for classification, then proper classification can be obtained, but it requires user time and effort

Engineering Contradiction:
Improveclassification accuracyVSAvoiduser interaction time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system automatically classifies browsing history without requiring user intervention. It uses the extracted keywords and pre-defined categories to autonomously organize browsing data, eliminating the need for users to manually categorize each visited webpage while maintaining classification accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual user classification actions with an automated computational system that uses keyword extraction and pattern matching. This substitution eliminates the mechanical user interaction required for classification while achieving comparable or superior classification precision through algorithmic processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Extent of automation

If frequency, label and metatag classification is used, then classification is performed, but there is no guarantee that important web pages can be easily accessible

Engineering Contradiction:
Improveautomatic classificationVSAvoidretrieval ease
Core Design Contradiction:
Extent of automationVSEase of operation

Solution Approach 1:

The patent segments browsing history into distinct categorical groups (e.g., news, entertainment, shopping, social media) based on extracted keywords. This segmentation allows users to easily access important web pages by navigating to specific category sections rather than searching through an unorganized list, significantly improving retrieval ease.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system adds a categorical dimension to the browsing history organization, transforming a single-dimensional chronological list into a multi-dimensional structure where pages can be accessed both by time and by category. This dimensional enhancement provides multiple access paths, making important pages more easily retrievable.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Adaptability or versatility

If existing classification techniques are used, then web classification is provided, but performance is slow and manageability is poor

Engineering Contradiction:
Improveclassification capabilityVSAvoidprocessing speed
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system performs preliminary keyword extraction and classification categorization as browsing history is generated, rather than processing everything at retrieval time. This preliminary action prepares the data in advance, enabling fast retrieval and improving overall productivity while maintaining versatile classification capabilities.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts only the essential classification information (keywords and category assignments) from browsing history and stores this condensed data structure separately. This extraction approach reduces the amount of data that needs to be processed during retrieval, significantly improving processing speed while preserving full classification functionality.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10482149B2Method and system for classification of web browsing history
Publication Date: 2019.11.19 WIPRO LTD
  • US10482149B2 patent drawing
  • US10482149B2 patent drawing
  • US10482149B2 patent drawing

AI summary

The present disclosure relates to method and computing device for classification of web browsing history by classification system. The classification system receives web browsing history from web browser associated with user, where web browsing history comprises details about one or more web pages browsed by user, extracts one or more keywords from each of one or more web pages browsed by user based on trained keyword dataset, determines a plurality of classifications for each of the one or more web pages based on one or more keywords, generates relevancy matrix between one or more keywords of web pages and corresponding plurality of classifications and identifies a classification from plurality of classifications for each of one or more webpages based on relevancy matrix, where snapshot of classification is stored in non-volatile storage unit of web browser. The use of non-volatile storage unit in present disclosure provides no restriction on storage space.