Web Data Scraping and Tokenization for Industrial Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for determining industrial classifications of entities in the insurance industry are error-prone, leading to inaccurate risk assessments and economic consequences, as they often fail to account for varied business operations and lack precise data, especially for new or small companies.

Innovation Solution

A computerized system that leverages electronic resources such as websites, social media, and third-party data to retrieve and analyze content using predictive models to determine and verify industrial classifications, providing a more accurate and reliable classification for insurance evaluations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If current methods for determining industrial classifications are used, then the process is simple and quick, but the accuracy and reliability of classification are poor

Engineering Contradiction:
Improveclassification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs multiple functions: retrieving data from electronic resources, tokenizing text, generating token counts, applying predictive models, and determining industrial classifications. This multi-functional approach consolidates classification accuracy improvement without proportionally increasing system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent replaces manual classification methods with automated computerized predictive models that process tokenized data from electronic resources. This substitution of mechanical/manual processes with automated computational systems improves classification accuracy while managing system complexity through algorithmic efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If manual verification of industrial classifications is performed, then accuracy may improve, but the time and resources required increase significantly

Engineering Contradiction:
Improveclassification reliabilityVSAvoidverification time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs self-verification by automatically retrieving data from electronic resources, processing it through predictive models, and generating classifications without requiring external manual verification. This self-service mechanism improves reliability while eliminating the time loss associated with manual verification processes.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The automated system enables continuous classification and verification operations without interruption. The predictive model continuously processes tokenized data from electronic resources, providing ongoing reliable classifications without the intermittent delays inherent in manual verification workflows.

Inventive Principle:
Principle #20Continuity of useful action

3Measurement precision

If comprehensive data from multiple electronic resources is collected and analyzed, then classification precision improves, but data processing complexity and computational requirements increase

Engineering Contradiction:
Improveclassification precisionVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the complex data processing task into distinct stages: retrieving data from electronic resources, tokenizing the text data, generating token counts, and applying predictive models. This segmentation reduces processing complexity by breaking down the comprehensive data analysis into manageable, sequential operations that maintain classification precision.

Inventive Principle:
Principle #1Segmentation

4Measurement precision

If automated predictive models are used to determine industrial classifications, then accuracy improves, but the complexity of the system increases

Engineering Contradiction:
Improveclassification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system introduces tokenized data and token counts as intermediary representations between the raw electronic resource data and the predictive model. This intermediary layer simplifies the input requirements for the predictive model, reducing system complexity while maintaining the accuracy benefits of automated classification.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10109017B2Web data scraping, tokenization, and classification system and method
Publication Date: 2018.10.23 HARTFORD FIRE INSURANCE CO
  • US10109017B2 patent drawing
  • US10109017B2 patent drawing
  • US10109017B2 patent drawing

AI summary

A web server obtains URL data for an electronic resource about an entity, and scrapes content data about the entity from the resource. A content processor tokenizes the content data, generates token count data, and stores the token count data in one or more data storage devices. A predictive model processor applies the token count data to a trained predictive model trained to generate first data indicative of at least one industrial classification applicable to the entity and second data indicative of a likelihood the first data is applicable to the entity. The web server is configured to provide, by the communications device to a user device and responsive to application of the trained computerized predictive model to the token count data, a display including the first data and second data.