ISP Traffic Classification via Gradient Boosting and Firmographic Enrichment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems for distinguishing website traffic, such as those used in CRM and Web Analysis Systems, are inadequate in accurately identifying Internet Service Provider (ISP) traffic, leading to many false positives and negatives, and fail to leverage real web traffic data for accurate classification.

Innovation Solution

A method and system utilizing machine intelligence, including IP address mapping, attribute data collection, and a gradient boosting classifier, to identify ISPs based on website traffic data, providing a more robust and accurate classification by incorporating firmographic data and web traffic behavior analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If simple lists of known ISPs or high-profile businesses are used for identification, then the system is easy to implement, but the identification accuracy deteriorates with many false positives and false negatives

Engineering Contradiction:
Improveease of implementationVSAvoididentification accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent transforms the identification approach from static list-matching to dynamic machine learning classification. The system changes parameters by incorporating multiple features (IP address, user agent, browsing behavior, session duration, page views) and continuously trains models on new data, enabling adaptive parameter optimization that resolves the contradiction between implementation simplicity and identification accuracy

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the mechanical list-matching system with an intelligent machine learning system. Instead of manually maintaining static lists, the system uses automated gradient boosting classifiers that learn patterns from web traffic data, substituting manual mechanical processes with automated intelligent analysis to achieve higher accuracy without proportionally increasing implementation complexity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Device complexity

If conventional list-based methods are used, then the system complexity is low, but the system cannot identify Global Traffic with a native company name

Engineering Contradiction:
Improvesystem complexityVSAvoidability to identify global traffic
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal classification system that handles multiple traffic types through a single machine learning model. The gradient boosting classifier is designed to identify various entity types (ISPs, global companies, local businesses, bots) using the same infrastructure, enabling the system to adapt to global traffic identification needs without requiring separate specialized systems for each traffic category

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent adds dimensional depth to traffic identification by incorporating multiple analysis layers: device-level features (user agent, IP), session-level features (duration, page views), and entity-level features (company name matching, firmographic data). This multi-dimensional approach enables the system to identify global traffic across different dimensions simultaneously, overcoming the limitations of simple list-based methods

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If machine learning classifiers are implemented, then the identification accuracy improves, but the computational resources and system complexity increase

Engineering Contradiction:
Improveidentification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the machine learning system into distinct modular components: feature extraction module, model training module, classification module, and continuous learning module. Each component performs a specific function, allowing the system to achieve high identification accuracy through specialized processing while managing complexity through modular architecture that can be implemented and maintained independently

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11651253B2Machine learning classifier for identifying internet service providers from website tracking
Publication Date: 2023.05.16 THE DUN & BRADSTREET CORP
  • US11651253B2 patent drawing
  • US11651253B2 patent drawing
  • US11651253B2 patent drawing

AI summary

A method and system for identifying and classifying Visitor Information tracked on websites to identify Internet Service Providers (ISPs) and non-Internet Service Providers (non-ISPs). The technology employs machine intelligence to train a classifier on firmographically-enriched Visitor Intelligence from website tracking technology. The ISP classifier can distinguish ISPs from non-ISPs to identify website traffic for a given website that is attributable to ISPs.