UGC Feature Extraction for Wide-Range Malicious Site Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies for detecting malicious sites are limited in detection range and lack the speed and accuracy needed for wide-area, real-time detection of malignant user-generated content.
Innovation Solution
An extraction device that processes user-generated content to extract features, trains on normal and malicious content, and determines malignancy using a trained model to identify and report threat information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If feature amount specialized for false live stream site fraud is used, then detection accuracy for specific fraud type is improved, but detection range is limited
Solution Approach 1:
The patent applies universality by designing a detection system that handles multiple types of fraudulent content (live streams, blogs, bulletin boards, videos) through a single unified platform. The system extracts feature quantities from various content types and uses a common trained model to detect malignancy across different services and content formats, enabling wide-range detection without sacrificing accuracy for specific fraud types
2Adaptability or versatility
If wide-range detection of malignant sites is performed, then detection coverage is improved, but detection speed and accuracy may deteriorate
Solution Approach 1:
The patent applies extraction by identifying and extracting key feature quantities from user-generated content that are indicative of malignancy. The system extracts specific features (such as URL patterns, content characteristics, metadata) from various content types and uses these extracted features for training and detection, enabling efficient wide-range detection without processing entire content sets
3Loss of time
If real-time detection of user-generated content is performed, then response time is improved, but processing load and complexity increase
Solution Approach 1:
The patent applies preliminary action by performing training in advance using feature quantities from both normal and malicious user-generated content. The trained model is prepared beforehand and can then perform rapid real-time detection without complex processing during actual detection. This pre-training approach reduces response time while managing processing complexity through advance preparation
Data Source
AI summary
An extraction device includes processing circuitry configured to access an entrance URL described in user-generated content generated by a user in a plurality of services in a predetermined period to extract a feature quantity of the user-generated content, perform training by using the extracted feature quantity of the user-generated content generated by a normal user and a feature quantity of content generated by a malicious user, and determine whether or not the user-generated content has been generated by the malicious user using a trained model.


