Spam Term Identification Using BTF-IDF Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Map spam, where businesses create fake listings on web mapping services to attract customers, leads to inaccurate information, harming users, web mapping services, and local businesses by diverting customers to distant businesses.
Innovation Solution
The implementation of a method to identify spam terms using a blacklist term frequency-inverse document frequency (BTF-IDF) score, which calculates the occurrence of terms in spam accounts versus non-spam accounts, allowing for the classification of accounts as spam or non-spam, thereby reducing map spam and maintaining accurate database information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If web mapping services allow businesses to create listings without verification, then account setup is simplified and businesses can be easily added, but map spam increases with fake listings diverting customers to distant businesses
Solution Approach 1:
The system performs preliminary analysis of account documents and business information before allowing listings to be created. By pre-screening for spam indicators and verifying business legitimacy during the account setup phase, the system prevents fake listings from entering the database while maintaining efficient onboarding processes
Solution Approach 2:
The system continuously monitors and analyzes business listings against established spam detection criteria, providing feedback loops that identify and remove fake listings. This feedback mechanism allows the system to maintain information accuracy while preserving ease of legitimate business registration
2Reliability
If manual verification of each business listing is performed, then information accuracy is improved, but processing time and system complexity increase
Solution Approach 1:
The system enables automatic self-verification through automated document analysis, business information validation, and spam detection algorithms. By implementing self-service verification mechanisms, the system achieves high accuracy without requiring manual intervention for each listing, thereby reducing processing complexity and time
Solution Approach 2:
The patent replaces manual verification processes with automated computational systems that use pattern recognition, document analysis, and machine learning algorithms. This substitution eliminates the need for human reviewers while maintaining or improving verification accuracy and reducing system complexity
3Reliability
If all account documents are stored for analysis, then comprehensive spam detection is possible, but storage space requirements increase
Solution Approach 1:
The system extracts and analyzes only the critical information and key features from account documents rather than storing and processing entire documents. By extracting essential business information, contact details, and location data for analysis, the system achieves comprehensive spam detection while significantly reducing storage requirements
Solution Approach 2:
The system discards unnecessary document data after extraction and analysis, retaining only essential information needed for spam detection and business operations. This approach reduces storage footprint while maintaining the ability to detect spam through preserved key features and extracted information
Data Source
AI summary
Methods, systems, and apparatus, including computer programs encoded on computer storage media, are described for identifying target terms, e.g., spam terms within a collection of documents. In one aspect, methods can include identifying spam terms by calculating a blacklist term frequency-inverse document frequency (BTF-IDF) score for multiple terms, and by selecting, as the spam terms, the terms that have scores above or below a threshold score. The multiple terms may be derived from documents that are associated with accounts that have been designated as spam accounts.


