Website Categorization Using Domain Token Probabilities and Vector Distance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for categorizing websites are inefficient and require manual entry of NAICS codes, making it impractical to classify the large number of existing websites, especially since they only classify to a 3-digit code level, lacking specificity and requiring extensive manual effort.
Innovation Solution
A system and method that analyzes website domain names and content to automatically categorize websites using token probabilities, TF-IDF scores, and reference vectors, allowing for a more granular classification based on keywords and a standardized category structure, such as NAICS 4-6 digit codes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If manual entry of NAICS codes is used for website categorization, then categorization can be performed, but the process requires extensive manual effort and time
Solution Approach 1:
The system enables websites to be automatically categorized through self-service mechanisms. The categorization system autonomously analyzes website content, domain names, and metadata to assign NAICS codes without requiring manual intervention. This self-categorizing capability eliminates the need for manual data entry while maintaining accurate classification.
Solution Approach 2:
The patent replaces the mechanical manual process of entering NAICS codes with an automated computational system. The system uses natural language processing, machine learning algorithms, and data analysis to automatically determine appropriate categorizations, substituting human manual labor with automated technological processes that analyze website characteristics and assign codes accordingly.
2Measurement precision
If manual categorization is performed, then websites can be classified, but the classification level is limited to 3-digit codes lacking specificity
Solution Approach 1:
The categorization system segments the classification process into multiple hierarchical levels. Instead of a single 3-digit code assignment, the system breaks down categorization into progressive stages, analyzing website characteristics and assigning increasingly specific NAICS codes (4-digit, 5-digit, 6-digit levels) based on the depth of analysis and available information about the website's content and purpose.
Solution Approach 2:
The system implements dynamic categorization that adapts to the specific characteristics of each website. Rather than applying a static classification rule, the system dynamically adjusts the level of specificity achieved based on the website's content, domain name, metadata, and other analyzed features, enabling more precise 4-6 digit code assignments when sufficient information is available.
3Productivity
If the number of websites to be categorized increases, then more data is available for analysis, but the manual effort required increases proportionally
Solution Approach 1:
The system enables continuous automated categorization of websites without interruption or manual intervention. As new websites are added to the database or existing websites are updated, the system continuously and automatically re-analyzes their content and assigns appropriate NAICS codes, maintaining constant productivity without the time losses associated with manual processing of each new entry.
Solution Approach 2:
The system changes the fundamental parameter of processing capacity by transitioning from manual to automated analysis. This parameter change enables the system to handle large volumes of websites simultaneously or in rapid succession, processing thousands of categorizations in the time it would take a human to process a handful, thereby dramatically increasing productivity while reducing total time investment.
Data Source
AI summary
Systems and methods for the categorization of websites are presented. A website is categorized using one or a combination of its domain name and its web page content. The domain name is tokenized, and the tokens compared to categories in a category structure to determine probabilities that the token belongs to each category. Combinations of tokens are similarly compared to the categories. A category may be determined with reference to a vector space in which a training set of websites having known categories is converted according to a methodology into reference vectors containing keyword frequencies. A target website is converted to a target vector using the same methodology, and a distance score of the target vector to each reference vector is calculated. The website represented by the target vector is assigned the category of the reference vector having the lowest distance score.


