Automated Entity Classification via Webpage Mining
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computer security systems rely on manual and often general industry classifications for determining security measures, which can be inaccurate and difficult for new or small companies to obtain, leading to inefficiencies in risk assessment and security implementation.
Innovation Solution
An automated system that mines webpages of entities with known classifications to create a classification structure, assigns class labels based on webpage content, and applies these labels to new entities with unknown classifications to determine their industry vertical and risk levels, enabling targeted security actions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If manual self-classification is used for industry classification, then companies can obtain their own classification labels, but the classifications become overly general and rigid, and new or small companies cannot obtain classifications
Solution Approach 1:
The system enables automatic self-classification by having entities classify themselves through webpage content analysis. The classification structure is applied automatically to entity webpages without requiring manual intervention, allowing new and small companies to obtain classifications independently and accurately based on their own digital footprint.
Solution Approach 2:
The patent replaces manual classification processes with automated computer-based analysis. The system uses computational algorithms to analyze webpage content, extract features, and assign classifications automatically, eliminating the need for manual self-classification while improving both accuracy and speed.
2Measurement precision
If manual classification processes are used, then classifications can be assigned to entities, but the process is time-consuming and difficult to obtain for new companies
Solution Approach 1:
The system performs preliminary classification by analyzing entity webpage content as soon as the entity establishes its online presence. The classification structure is pre-built and ready to be applied automatically, eliminating delays associated with manual classification processes and enabling new companies to be classified immediately upon webpage creation.
Solution Approach 2:
Manual classification tasks are replaced with automated computational processes that analyze webpage content, extract relevant features, and assign classifications instantly. This substitution eliminates time-consuming manual review while maintaining or improving classification accuracy through consistent algorithmic application.
3Adaptability or versatility
If general industry classifications are used, then security measures can be implemented, but the security measures may not be appropriately targeted to specific entities
Solution Approach 1:
The system applies local quality by tailoring security measures to each entity's specific classification and characteristics. Instead of applying uniform security measures based on broad industry categories, the system analyzes individual entity webpage content to determine specific risk factors and applies customized security measures appropriate to each entity's actual business activities and risk profile.
Solution Approach 2:
The system changes the parameter of classification granularity from general industry categories to specific entity-level classifications. By analyzing detailed webpage content and extracting specific features, the system transforms broad classification parameters into precise entity-specific parameters, enabling more accurate risk assessment and targeted security measure application.
Data Source
AI summary
The disclosed computer-implemented method for creating automatic computer-generated classifications may include (i) mining webpages of entities with a known classification, (ii) using information mined from the webpages to create a classification structure that assigns class labels to entities based on entity webpage content, (iii) applying, to the classification structure, one or more webpages of a new entity with an unknown classification, (iv) receiving, from the classification structure, a class label for the new entity, and (v) performing a security action based on the new entity's class label. Various other methods, systems, and computer-readable media are also disclosed.


