Hash-Based Monitored Data Segmentation for Fraud Anomaly Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems struggle to effectively detect and isolate anomalous activities in multi-dimensional data streams, particularly in high-cardinality scenarios where fraud patterns change frequently, leading to operational expenses and reputational risks from identity theft attacks.

Innovation Solution

A computer-based system that dynamically calculates hash keys and anomaly scores for monitored segmentations in multi-dimensional data streams, using counting structures and chi-squared goodness of fit statistics to identify and mark devices with pre-generated labels based on predetermined thresholds.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional overall account volume monitoring is used to detect identity theft fraud, then the monitoring approach is simple, but it fails to effectively detect and isolate anomalous activities in high-cardinality multi-dimensional data streams where fraud patterns change frequently

Engineering Contradiction:
Improvefraud detection accuracyVSAvoidmonitoring system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent divides the multi-dimensional data stream into multiple segmentations using dynamically calculated hash keys. Each segmentation represents a different view of the data (e.g., by device, by account, by transaction type), allowing the system to detect anomalies at multiple granularities. This segmentation approach enables effective fraud detection in high-cardinality scenarios without requiring monitoring of every individual data point.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically calculates hash keys and anomaly scores in real-time as data flows through the system. The hash keys are regenerated based on current segmentation parameters, and anomaly scores are continuously updated using counting structures. This dynamic approach allows the system to adapt to changing fraud patterns while maintaining computational efficiency.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If multiple monitored segmentations are dynamically calculated and analyzed, then anomaly detection capability is improved, but processing complexity and computational resources increase

Engineering Contradiction:
Improveanomaly detection precisionVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces counting structures as intermediary data structures that efficiently aggregate information across multiple segmentations. These counting structures serve as a mediator between the raw multi-dimensional data and the anomaly detection logic, enabling precise measurement without requiring complex direct analysis of all data combinations.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system replaces traditional mechanical anomaly detection methods with hash-based counting structures. Instead of directly comparing and analyzing every data point across multiple dimensions, the patent uses hash functions to map data to counting structures, which efficiently aggregate frequencies. This substitution dramatically reduces computational complexity while maintaining detection precision.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If real-time anomaly detection is implemented with dynamically calculated hash keys and anomaly scores, then fraud identification speed is improved, but system resource consumption increases

Engineering Contradiction:
Improvefraud identification speedVSAvoidsystem resource consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The counting structures automatically update themselves as data flows through the system, without requiring external intervention or complex processing. Each incoming data point self-updates the relevant counting structures based on its hash key, enabling real-time detection with minimal system resource consumption. The system serves itself by leveraging the inherent structure of the data flow.

Inventive Principle:
Principle #25Self-service

Data Source

PatentEP4579503A1Computer-based systems configured to select a monitored data segmentation and methods of use thereof
Publication Date: 2025.07.02 CAPITAL ONE SERVICES LLC
  • EP4579503A1 patent drawingFigure 1
  • EP4579503A1 patent drawingFigure 2
  • EP4579503A1 patent drawingFigure 3A

AI summary

In some embodiments, the present disclosure provides an exemplary system and method that may include steps of identifying a device capable of processing a data stream; calculating a plurality of hash keys for a plurality of monitored segmentations associated with the device capable of the data stream; generating an increment data counter that corresponds to each hash key in a plurality of counting structures; calculating an anomaly score associated for the plurality of monitored segmentations; selecting a monitored segmentation based on the anomaly score; determining that a selected monitored segmentation meets a predetermined threshold associated with the anomaly score; and automatically marking the device capable of the data stream with a pre-generated label.