Algorithmic Topic Clustering for User Behavior Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for topic clustering and classification of data are inefficient, leading to improper recording and classification of research data, making replication of studies questionable and increasing the likelihood of misconduct allegations. Additionally, the vast amount of internet content makes it difficult to discern true from fake information and track user behavior accurately.

Innovation Solution

The development of an algorithmic topic clustering system that classifies internet user behavior by tracking page views and recency, categorizing users into intender and nonintender groups, and serving targeted content based on advertising targeting parameters, using natural language processing and machine learning techniques to improve data classification and behavioral characterization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional topic clustering methods are used, then data classification can be performed, but the classification accuracy and reliability are insufficient leading to improper recording and classification of research data

Engineering Contradiction:
Improvedata classification accuracyVSAvoidresearch data reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent segments the data classification process into multiple specialized components: topic modeling module for identifying latent topics, cluster analysis module for grouping similar documents, and hierarchical classification module for organizing results. This segmentation allows each module to specialize in specific aspects of classification, improving both accuracy and reliability compared to conventional single-method approaches.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediary components including a feature extraction layer that transforms raw data into meaningful representations, and a validation layer that verifies classification results before final recording. These intermediaries ensure that data is properly transformed and validated throughout the classification process, preventing improper recording and enhancing reliability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If manual tracking of user behavior is performed, then user behavior can be characterized, but the process is inefficient and cannot handle the great volume of internet data

Engineering Contradiction:
Improveuser behavior analysis efficiencyVSAvoidvolume of internet data
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent replaces manual mechanical tracking methods with automated computational systems including web crawlers that automatically scrape content, machine learning models that automatically classify user behavior patterns, and algorithms that automatically analyze large datasets. This substitution enables the system to process vast volumes of internet data efficiently without human intervention.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms the analysis approach by changing parameters from tracking individual user actions manually to analyzing aggregated behavioral patterns through computational metrics. The system monitors multiple parameters simultaneously (clickstreams, dwell time, navigation paths) and uses algorithmic processing to derive insights, dramatically increasing productivity compared to manual methods.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If conventional classification methods are used, then data can be organized, but it becomes increasingly difficult to discern true from fake content and identify reliable trends

Engineering Contradiction:
Improveinformation reliabilityVSAvoidcontent verification difficulty
Core Design Contradiction:
Loss of informationVSDifficulty of detecting and measuring

Solution Approach 1:

The patent implements feedback mechanisms where classification results are continuously validated against ground truth data, and model predictions are refined based on performance metrics. The system incorporates feedback loops that adjust classification thresholds and parameters based on detected patterns of reliable versus unreliable content, progressively improving the ability to discern true from fake information.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent employs a composite classification approach that combines multiple classification methods (topic modeling, cluster analysis, sentiment analysis, fact-checking algorithms) into an integrated system. Each method contributes different strengths to the overall classification, creating a robust multi-layered verification process that enhances information reliability and makes it easier to identify credible content amidst vast amounts of data.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS20220335220A1Algorithmic topic clustering of data for real-time prediction and look-alike modeling
Publication Date: 2022.10.20 ZETA GLOBAL CORP
  • US20220335220A1 patent drawing
  • US20220335220A1 patent drawing
  • US20220335220A1 patent drawing

AI summary

The subject technology provides a user classification system comprising a communications network, a Front-End URL Handler (FEUH) establishing an entry point for content calls from a network user, and a Fast Retrieval (FR) store storing a set of behavioral segments. A classification engine performs operations comprising accessing a plurality of pages viewed by a communications network user, classifying the plurality of pages as pertaining to at least one topic of a plurality of topics, tracking a count of each of the pages viewed by the communications network user for each of the topics, tracking a recency or frequency with which each of the pages viewed by the communications network user was viewed for each of the topics, characterizing the communications network user as belonging to one or more of the behavioral segments based on the tracked count and tracked recency, and serving content to the communications network user based on a targeting parameter and the behavioral segment characterization.