Major-Key-Shared Correlation for Fraud Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing correlation measures for categorical variables, such as the Chi-Squared test, are sensitive to sample size and inaccurate in large data sets, making them unsuitable for online fraud detection systems that handle millions or billions of user accounts.
Innovation Solution
The introduction of a Major-Key-Share-based correlation coefficient (MKS-based correlation) that measures the association between categorical variables by determining if one variable is generated by another, providing a precise correlation measure that is not biased by sample size, allowing for efficient fraud detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If Chi-Squared statistics are used to measure correlation between categorical variables, then the method is widely applicable and based on established statistical theory, but the measurement precision deteriorates when sample size is too large or too small, rendering it useless for big data fraud detection
Solution Approach 1:
The patent transforms the Chi-Squared statistic by taking its square root and normalizing it to create a new correlation coefficient that ranges from 0 to 1. This parameter transformation resolves the issue where Chi-Squared statistics become unreliable with large sample sizes, as the new coefficient maintains measurement precision across varying sample sizes including big data scenarios.
Solution Approach 2:
The patent replaces the traditional Chi-Squared statistical testing mechanism with a new correlation coefficient system that is specifically designed to handle categorical variables in big data contexts. This substitution eliminates the sample size sensitivity problem inherent in the original Chi-Squared approach while maintaining the ability to measure association between variables.
2Reliability
If Chi-Squared statistics are applied to large data sets with millions or billions of data points, then comprehensive coverage is achieved, but the reliability of correlation measurement deteriorates due to sensitivity to sample size
Solution Approach 1:
By transforming the Chi-Squared statistic through square root and normalization operations, the patent creates a reliable correlation coefficient that remains stable regardless of whether the data set contains millions or billions of points. This parameter change ensures that the measurement reliability does not deteriorate with increasing data volume.
3Loss of information
If traditional correlation measures are used for fraud detection, then existing statistical methods are applied, but erroneous association relationships are generated, polluting the fraud detection results
Solution Approach 1:
The patent substitutes the traditional Chi-Squared correlation measurement system with a new correlation coefficient system that is specifically calibrated for categorical variables in big data. This substitution eliminates the generation of erroneous association relationships that occur with traditional methods, thereby preventing pollution of fraud detection results.
Solution Approach 2:
The transformation of the Chi-Squared statistic into a normalized coefficient with a defined range (0 to 1) provides more accurate and interpretable correlation measurements. This parameter change prevents the generation of spurious associations that occur when using unnormalized statistics on large data sets, thus reducing information loss in fraud detection.
Data Source
AI summary
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for detecting fraudulent accounts. One of the methods includes obtaining raw data from network events associated with a collection of user accounts of an online service; processing the raw data including determining a feature set and applying the feature set to generate user groups each comprising one or more user account; evaluating each user group based on feature distributions including performing one or more Major-Key-Shared (MKS) correlation calculations on pairs of features from the feature set for the group; scoring each group based on the evaluation; and identifying groups having a score that exceeds a specified threshold as fraudulent.

