Spammer Group Extraction Using NLP and Behavioral Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing malicious behavior detection and blocking algorithms on social network services are inadequate in identifying and addressing collective and intentional malicious activities, such as spreading false information, which interfere with fair trade and unbiased decision-making.

Innovation Solution

A spammer group extraction method and apparatus that utilizes natural language processing based on big data to collect and preprocess social network data, detect abnormal behavior, and extract spammer groups by analyzing user IDs and their connections, including identifying distributed phrases, keywords, Retweet activities, and temporal patterns.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional spam detection algorithms based on URL analysis and simple message characteristics are used, then individual spam messages can be detected, but collective malicious behavior patterns and spammer groups cannot be identified

Engineering Contradiction:
Improvespam detection accuracyVSAvoidability to detect collective malicious behavior
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the analysis into multiple dimensions: individual message characteristics (URLs, keywords), user behavior patterns (posting frequency, time distribution), and group-level metrics (follower ratios, message similarity). This segmentation allows detection at both individual and collective levels simultaneously.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from two-dimensional analysis (message content only) to multi-dimensional analysis by incorporating temporal dimensions (time span, posting frequency), social network dimensions (follower ratios, connection patterns), and linguistic dimensions (message similarity, keyword distribution). This enables identification of coordinated spammer groups.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If comprehensive big data analysis with natural language processing is applied, then spammer groups can be effectively identified, but system complexity and computational resources increase

Engineering Contradiction:
Improveability to identify spammer groupsVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary actions by pre-processing social network data to extract key features (user profiles, message histories, connection patterns) before analysis. This pre-processing organizes raw data into structured formats, reducing computational complexity during the actual spammer group detection phase.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces intermediary components including a data collection unit that aggregates social network data, a natural language processing unit that extracts linguistic features, and an abnormal behavior detection unit that identifies patterns. These intermediaries break down the complex analysis into manageable stages.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If analysis focuses on simple message characteristics and user ID ratios, then processing speed is maintained, but detection of intentional slander and false information is insufficient

Engineering Contradiction:
Improveprocessing speedVSAvoiddetection of intentional malicious behavior
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements continuous monitoring and analysis of social network data streams, maintaining processing speed through optimized data flow. The system continuously collects, pre-processes, and analyzes data without interruption, enabling real-time detection of malicious behavior patterns while preserving processing efficiency.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS9563770B2Spammer group extraction apparatus and method
Publication Date: 2017.02.07 ELECTRONICS & TELECOMM RES INST
  • US9563770B2 patent drawing
  • US9563770B2 patent drawing
  • US9563770B2 patent drawing

AI summary

The present invention relates to a spammer group extraction apparatus and method, which extract spammer groups that interfere with fair trade and unbiased decision making by sending messages aimed at intentionally slandering other companies (other persons, other products, etc.) on social network services. The spammer group extraction apparatus includes a data collection unit for collecting pieces of data corresponding to social network services. A natural language processing unit preprocesses the pieces of data using a natural language processing algorithm based on big data. An abnormal behavior detection unit detects abnormal behavior based on user identifications (IDs) respectively corresponding to pieces of data, preprocessing of which has been completed. A spammer extraction unit extracts a spammer group using a user ID causing the abnormal behavior and an ID of a user group including the user ID.