Web Data Scraping and Tokenization for Industrial Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for determining industrial classifications of entities in the insurance industry are error-prone, leading to inaccurate risk assessments and economic consequences, as they often fail to account for varied business operations and lack precise data, especially for new or small companies.
Innovation Solution
A computerized system that leverages electronic resources such as websites, social media, and third-party data to retrieve and analyze content using predictive models to determine and verify industrial classifications, providing a more accurate and reliable classification for insurance evaluations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If current methods for determining industrial classifications are used, then the process is simple and quick, but the accuracy and reliability of classification are poor
Solution Approach 1:
The system performs multiple functions: retrieving data from electronic resources, tokenizing text, generating token counts, applying predictive models, and determining industrial classifications. This multi-functional approach consolidates classification accuracy improvement without proportionally increasing system complexity.
Solution Approach 2:
The patent replaces manual classification methods with automated computerized predictive models that process tokenized data from electronic resources. This substitution of mechanical/manual processes with automated computational systems improves classification accuracy while managing system complexity through algorithmic efficiency.
2Reliability
If manual verification of industrial classifications is performed, then accuracy may improve, but the time and resources required increase significantly
Solution Approach 1:
The system performs self-verification by automatically retrieving data from electronic resources, processing it through predictive models, and generating classifications without requiring external manual verification. This self-service mechanism improves reliability while eliminating the time loss associated with manual verification processes.
Solution Approach 2:
The automated system enables continuous classification and verification operations without interruption. The predictive model continuously processes tokenized data from electronic resources, providing ongoing reliable classifications without the intermittent delays inherent in manual verification workflows.
3Measurement precision
If comprehensive data from multiple electronic resources is collected and analyzed, then classification precision improves, but data processing complexity and computational requirements increase
Solution Approach 1:
The system segments the complex data processing task into distinct stages: retrieving data from electronic resources, tokenizing the text data, generating token counts, and applying predictive models. This segmentation reduces processing complexity by breaking down the comprehensive data analysis into manageable, sequential operations that maintain classification precision.
4Measurement precision
If automated predictive models are used to determine industrial classifications, then accuracy improves, but the complexity of the system increases
Solution Approach 1:
The system introduces tokenized data and token counts as intermediary representations between the raw electronic resource data and the predictive model. This intermediary layer simplifies the input requirements for the predictive model, reducing system complexity while maintaining the accuracy benefits of automated classification.
Data Source
AI summary
A web server obtains URL data for an electronic resource about an entity, and scrapes content data about the entity from the resource. A content processor tokenizes the content data, generates token count data, and stores the token count data in one or more data storage devices. A predictive model processor applies the token count data to a trained predictive model trained to generate first data indicative of at least one industrial classification applicable to the entity and second data indicative of a likelihood the first data is applicable to the entity. The web server is configured to provide, by the communications device to a user device and responsive to application of the trained computerized predictive model to the token count data, a display including the first data and second data.


