Deep Similarity Modeling for User Clusters from Mixed Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies struggle to accurately determine consumer behavior patterns due to independently controlled data sources generating diverse and unstandardized data streams, making it difficult to predict retail potential and understand user behavior effectively.

Innovation Solution

A method involving deep similarity modeling using a trained model to derive behavioral attributes at an entity level by matching datasets, generating ground truth labels, and determining feature combinations to identify similar or contrasting entities based on their behavioral attributes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If multiple independently controlled data sources are used to collect consumer behavior data, then the quantity and diversity of data increase, but the data standardization and consistency deteriorate

Engineering Contradiction:
Improvedata quantityVSAvoiddata consistency
Core Design Contradiction:
Quantity of substanceVSStability of the object's composition

Solution Approach 1:

The patent introduces an intermediary processing layer that receives data from multiple independently controlled sources and transforms it into a standardized format. This intermediary system performs data cleaning, normalization, and feature extraction to convert diverse raw data into consistent structured data that can be processed by the deep learning model, thus resolving the contradiction between data diversity and data consistency.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies parameter changes by transforming various data formats and structures into a unified parameter schema. Through data normalization and feature engineering, different data sources with varying formats, granularities, and structures are converted into consistent parameters that the deep learning model can process, maintaining data consistency while preserving the diversity of sources.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If deep similarity modeling is performed on large-scale client data, then the accuracy of behavioral attribute derivation improves, but the computational complexity and processing time increase

Engineering Contradiction:
Improvebehavioral attribute accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the large-scale data processing task into multiple stages: data preprocessing, feature extraction, similarity computation, and behavioral attribute derivation. The deep learning model is divided into multiple layers that process data hierarchically, computing similarities at different levels of abstraction. This segmentation reduces the computational burden at any single stage while maintaining overall accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by pre-processing the data and pre-computing features before the main similarity modeling task. Data cleaning, normalization, and feature extraction are performed in advance, and the model is pre-trained on sample data. This preliminary processing reduces the complexity of the main computation and accelerates the overall processing time.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If ground truth labels are generated for training the deep learning model, then the model training accuracy improves, but the data preparation time and resource requirements increase

Engineering Contradiction:
Improvemodel training reliabilityVSAvoiddata preparation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent uses copying by creating synthetic ground truth labels through data augmentation and simulation. Instead of manually annotating all training data, the system generates additional training examples by copying and transforming existing labeled data, applying various transformations to create new valid training samples. This reduces the time required for data preparation while maintaining training reliability.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent implements self-service by enabling the system to automatically generate ground truth labels through self-supervised learning mechanisms. The model learns to predict certain attributes from the data itself, using the data's inherent structure and relationships to create its own training labels, thereby reducing manual intervention and time requirements for label generation.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12412094B1Server-based method of using a trained deep learning model and ground truth labels
Publication Date: 2025.09.09 AZIRA LLC
  • US12412094B1 patent drawing
  • US12412094B1 patent drawing
  • US12412094B1 patent drawing

AI summary

A server-implemented method for determining, at a server, using a deep learning model, a cluster of device identifiers associated with computing devices having characteristics that are related to a match set of computing devices based on locations data streams and ground truth labels.