Survey Fraud Scoring Using Rules and Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional feedback systems fail to accurately identify fraudulent survey responses, leading to inaccurate data gathering and skewing of insights due to bad actors exploiting incentives, resulting in compromised and irrelevant information.
Innovation Solution
A fraudulent response determination system utilizing rule-based models and machine-learning models to generate fraud scores, identifying fraud indicators, and updating datasets by removing fraudulent survey responses.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional feedback systems gather information from target audiences, then data collection volume increases, but data accuracy deteriorates due to fraudulent responses
Solution Approach 1:
The system performs preliminary fraud detection and scoring on survey responses before they are fully integrated into the dataset. By applying rule-based models and machine learning models to identify fraud indicators in advance, the system prevents fraudulent data from contaminating the overall dataset, thus maintaining data accuracy while allowing continuous data collection from target audiences.
Solution Approach 2:
The fraud score acts as an intermediary mechanism between data collection and data utilization. The system generates fraud scores based on multiple indicators and uses these scores to determine whether to include or exclude specific responses. This intermediary layer allows the system to maintain high data collection volume while filtering out fraudulent responses to preserve data accuracy.
2Quantity of substance
If conventional systems include fraudulent data in datasets, then data completeness is maintained, but downstream analysis accuracy deteriorates
Solution Approach 1:
The system applies different quality standards to different data points based on their fraud scores. Rather than uniformly including or excluding all data, the system evaluates each response individually and applies local quality control measures. Responses with high fraud scores are excluded or flagged, while legitimate responses are included, thus maintaining overall data completeness while ensuring analysis accuracy.
3Measurement precision
If the system removes fraudulent responses from datasets, then data accuracy improves, but data processing complexity increases
Solution Approach 1:
The fraud detection system is segmented into multiple independent components: rule-based models for straightforward fraud detection, machine learning models for pattern recognition, and fraud indicator calculation modules. This segmentation allows the system to process data through specialized subsystems, improving accuracy while managing complexity through modular architecture where each component has a specific function.
Solution Approach 2:
The system changes parameters such as fraud score thresholds and weighting factors to optimize the balance between data accuracy and processing complexity. By adjusting these parameters, the system can adapt to different survey types and fraud patterns without requiring complete system redesign, thus maintaining accuracy while controlling processing complexity.
Data Source
AI summary
The present disclosure relates to systems, non-transitory computer-readable media, and methods for generating a fraud score for survey response data and updating a dataset of responses of a digital survey. In particular, in one or more embodiments, the disclosed systems utilize a fraud indicator identifying algorithm to determine fraud indicators and generate a fraud score for the survey response data. In addition, in one or more embodiments, the disclosed systems utilize a fraudulent response identifying machine-learning model to generate a fraud score. The disclosed systems then utilize the fraud score to generate a label for survey response data and update a dataset of responses to a digital survey based on the label. In one or more embodiments, based on the disclosed systems generating a fraudulent label for the survey response data, the disclosed systems remove survey response data from the dataset.


