Language Model Clustering for Suspicious Account Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for detecting suspicious account groups in social networks are inadequate as they fail to effectively analyze language fashion similarity and near-synonyms, leading to inefficiencies in identifying high speech homogeneity and interaction connections among accounts.
Innovation Solution
A method and system that establish a language model for each account group based on post content, compare similarities to cluster accounts, and update near-synonyms for newly added data, integrating and re-clustering accounts to continuously discover suspicious groups.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional tracking down technologies are used to identify account groups, then basic feature matching can be performed, but language fashion similarity and near-synonym analysis cannot be effectively conducted
Solution Approach 1:
The patent transforms the detection parameters from basic feature matching (name, phone number, address) to language model-based parameters including speech homogeneity, near-synonym frequency, and linguistic fashion similarity. This allows the system to detect accounts with identical or near-identical language patterns even when basic features differ.
Solution Approach 2:
The patent introduces an intermediary language model that analyzes the semantic and syntactic patterns of text content. This intermediary layer processes the raw text data to extract linguistic features, enabling the system to detect suspicious account groups through language fashion analysis rather than direct feature comparison.
2Productivity
If static account groups are analyzed once, then initial clustering can be completed, but continuous discovery of newly formed suspicious groups is not achieved
Solution Approach 1:
The patent implements a continuous monitoring mechanism that periodically collects new text content from social network accounts and updates the language models accordingly. The system continuously compares new accounts against existing suspicious groups and detects newly formed suspicious groups in real-time, rather than performing one-time analysis.
Solution Approach 2:
The patent incorporates feedback loops where the results of language model analysis are used to refine the detection thresholds and update the suspicious account group database. The system learns from detected patterns and adjusts its detection sensitivity, improving its ability to identify new suspicious groups while reducing false positives.
3Measurement precision
If detailed language model analysis is performed for each account, then speech homogeneity can be detected, but computational complexity increases significantly
Solution Approach 1:
The patent segments the language analysis process into distinct modules: text preprocessing, feature extraction, language model construction, and similarity comparison. Each module handles a specific aspect of the analysis, allowing for optimized processing at each stage and enabling parallel computation where applicable.
Solution Approach 2:
The patent implements a two-stage detection approach: first performing a quick filter using basic features and simple text metrics to eliminate obviously non-suspicious accounts, then applying the computationally intensive language model analysis only to accounts that pass the initial filter. This reduces the overall computational burden while maintaining detection precision.
Data Source
AI summary
In one exemplary embodiment, a system for discovering suspicious account groups establishes a language model according to the post contents from each account of a first group of accounts during a first time interval, to describe the speech of the account, and compares the similarity among a plurality of language models of the first group of accounts to cluster the first group of accounts; and for a plurality of newly added data during a second time interval, discovers near-synonyms of at least a monitored vocabulary set, and updates the near-synonyms to a plurality of language models of a second group of accounts. The system further integrates the first and the second groups of accounts, and re-clusters an integrated group of accounts.


