Outlier Detection in Social Networks via Seed-Based Sampling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Identifying outliers in large online social networks is challenging due to the vast amount of publicly available information, making it difficult for law enforcement and other agencies to detect anomalies or unusual behavior within these networks.
Innovation Solution
A method and system that utilize seed information to reduce the sampling size of OSN users, compare them to a social graph or generalized profile, and identify outliers by analyzing follower and following relationships, status updates, and geographical data, employing data collection techniques like constrained crawls and Metropolis-Hastings algorithms to characterize user behavior and network structures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the entire population of OSN users is analyzed to identify outliers, then the detection accuracy is improved, but the computational resources and time required increase significantly
Solution Approach 1:
The patent segments the large OSN user population into smaller, manageable subsets using sampling techniques. Instead of analyzing all users, the system divides the population into representative samples that can be processed efficiently while maintaining outlier detection accuracy. This segmentation allows law enforcement agencies to identify anomalies without being overwhelmed by the sheer volume of data in networks with hundreds of millions of users.
2Measurement precision
If the entire population of OSN users is analyzed to identify outliers, then the detection accuracy is improved, but the computational resources required increase significantly
Solution Approach 1:
The patent applies partial action by analyzing only a representative sample of the OSN user population rather than the entire population. The sampling methodology is designed to capture sufficient statistical information to detect outliers effectively, while consuming significantly fewer computational resources. This approach allows the system to achieve adequate detection accuracy without the prohibitive computational cost of processing all users in large social networks.
3Productivity
If sampling size is reduced to improve processing efficiency, then the analysis time and computational resources are reduced, but the detection accuracy may deteriorate
Solution Approach 1:
The patent employs parameter changes by systematically adjusting sampling parameters such as sample size, sampling rate, and selection criteria to optimize the balance between processing efficiency and detection accuracy. The system modifies these parameters based on the specific requirements of the analysis task, the size of the OSN population, and the desired level of confidence in outlier detection. This allows the system to maintain adequate accuracy while achieving significant improvements in processing efficiency.
Data Source
AI summary
A system that incorporates teachings of the present disclosure may include, for example, a process that reduces a sampling size of a total population of on-line social network users based on a comparison of seed information to a population of on-line social network users. The reduced sampling of on-line social network users is compared to a social graph of the on-line social network users, wherein the social graph is obtained from an algorithm applied to the reduced sampling of the on-line social network users. An outlier is determined in the reduced sampling of on-line social network users based on a characterizing of a cluster of social network users. Additional embodiments are disclosed.