AI Domain Categorization Using Webpage Feature Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual categorization of web domains is time-consuming, labor-intensive, and prone to errors, making it impractical to keep pace with the large number of new domains created daily.

Innovation Solution

A classifier is trained using labeled training data to extract features from webpages, allowing for the automatic categorization of domains, reducing the need for human reviewers and enabling efficient categorization of new domains.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual human review is used to categorize webpages, then categorization accuracy can be maintained, but the process becomes extremely time consuming and labor intensive

Engineering Contradiction:
Improvecategorization accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

Human reviewers perform preliminary categorization of training data to create labeled datasets. This preliminary action enables the subsequent automated classifier to learn from pre-categorized examples, combining human accuracy with automated speed for processing new domains.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the mechanical system of manual human review with an automated machine learning classifier. The classifier uses training data labeled by humans to automatically categorize new webpages and domains, eliminating the need for continuous manual review while maintaining categorization accuracy through learned patterns.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If manual human review is used to categorize each new domain, then categorization quality can be ensured, but the process becomes impractical given the huge number of new domains created every day

Engineering Contradiction:
Improvecategorization qualityVSAvoidcategorization throughput
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent substitutes manual human review with an automated classifier system that can process numerous domains simultaneously. This mechanical substitution enables the system to handle the huge volume of new domains created daily while maintaining categorization quality through consistent application of learned classification rules.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The classifier system performs self-service by automatically categorizing new domains without requiring human intervention for each domain. The system uses its trained model to independently evaluate and categorize incoming domains, enabling high productivity while maintaining quality through automated consistency.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If human reviewers review each webpage of a domain, then comprehensive categorization can be achieved, but the process becomes extremely labor intensive

Engineering Contradiction:
Improvecategorization completenessVSAvoidhuman resource requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces the complex human resource system with a simplified automated classifier. The classifier processes webpage features and content automatically, eliminating the need for multiple human reviewers while achieving comprehensive categorization through systematic analysis of relevant features.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent extracts key features and content from webpages that are relevant for categorization, separating the essential classification information from the full webpage content. This extraction enables the classifier to achieve comprehensive categorization by focusing on the most relevant features without requiring review of entire webpages by human reviewers.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12039442B2Systems and method for categorizing domains using artificial intelligence
Publication Date: 2024.07.16 720 IT UAB
  • US12039442B2 patent drawing
  • US12039442B2 patent drawing
  • US12039442B2 patent drawing

AI summary

In an embodiment, a set of labeled training data that includes indicators of webpages is received. Each indicated webpage is labeled with one or more categories that were determined for the webpage by a human reviewer. Features, such as text and scripts, are extracted from each indicated webpage, and are used along with the labels to train a classifier to predict one or more categories for a webpage based on the features of the webpage. The trained classifier may be used to associate one or more categories with each domain of a plurality of domains given the categories predicted for some or all of the webpages associated with the domain. A list of domains and associated categories may be used for a variety of purposes including search engine optimization and content filtering.