Spam Detection via Contextual Frequency Differential Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current search engines face challenges in accurately identifying and filtering out spam business listings, which can manipulate search results to drive traffic to incorrect locations or websites, leading to user dissatisfaction and loss of trust in search services.
Innovation Solution
A system and method that utilize a processor to analyze business listing characteristics from trusted and untrusted sources, determining frequency differentials and assigning spam scores based on context-specific curves and thresholds, to identify and filter out spam listings by comparing characteristics such as title length, text terms, and phone numbers across different data sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If business listings are collected from multiple sources including untrusted sources, then the quantity of business listings increases, but the reliability of search results decreases due to spam listings
Solution Approach 1:
The patent segments business listings by context (e.g., type of business, geographic location) and analyzes characteristics within each segment. This allows the system to identify spam patterns specific to different contexts while preserving legitimate diverse listings, thus maintaining quantity while improving reliability through context-aware filtering
Solution Approach 2:
The patent introduces an intermediary spam detection system that acts as a mediator between multiple data sources and the final search results. This intermediary layer analyzes business listing characteristics, compares them across trusted and untrusted sources, and filters out spam before results are presented to users, thereby preserving reliability while allowing quantity to increase
2Measurement precision
If spam detection uses comprehensive analysis of business listing characteristics, then the measurement precision of spam identification improves, but the device complexity increases
Solution Approach 1:
The patent segments the spam detection process into distinct modules: characteristic extraction, frequency analysis, differential calculation, and spam scoring. Each module handles a specific aspect of the analysis, making the overall complex system more manageable and maintainable while achieving high measurement precision through comprehensive characteristic analysis
Solution Approach 2:
The patent transforms complex spam detection into a series of parameter comparisons. By extracting specific characteristics (title length, text terms, phone numbers, addresses) and analyzing their frequency differentials between trusted and untrusted sources, the system converts a complex judgment problem into measurable parameter comparisons, achieving high precision with manageable complexity
3Measurement precision
If the system compares business listing characteristics across multiple data sources, then the accuracy of spam detection improves, but the loss of time increases due to additional processing
Solution Approach 1:
The patent performs preliminary analysis of business listing characteristics from multiple sources in advance. By pre-calculating frequency distributions and identifying differential characteristics before actual spam detection queries, the system reduces processing time during runtime while maintaining high accuracy through the pre-established baseline data
Solution Approach 2:
The patent applies context-specific analysis where the system tailors its detection approach to different business listing contexts (e.g., different types of businesses, geographic regions). This allows the system to focus computational resources on the most relevant characteristics for each context, improving accuracy while reducing overall processing time by avoiding unnecessary analysis
Data Source
AI summary
Aspects of the disclosure provide for detection of spam business listings. Aspects operate to identify business listing characteristics in trusted sources and untrusted sources. As untrusted sources are likely to contain more spam, characteristics that are present in untrusted sources but not present in trusted sources are typically indicative of spam listings, and vice versa. Thus, statistical analysis of the frequency of characteristics within each source may be used to identify common characteristics of spam listings. These characteristics may further be analyzed in specific listing contexts, as different listing contexts (e.g., different types of businesses) typically use different terms and vocabularies, such that terms that are indicative of spam in one context may not be indicative of spam in another. Various methods for leveraging this context-specific statistical information to improve spam detection operations are disclosed.


