Automated Keyword Extraction for Malicious Web Page Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for detecting inducing elements on web pages that launch social engineering attacks rely on manual setting of keywords by analysts, which can be inaccurate and time-consuming, especially for inexperienced analysts, due to the large number of web pages and character strings involved.
Innovation Solution
An automated extraction apparatus that classifies web pages into clusters, extracts HTML elements from malicious and benign pages, and identifies character strings specific to malicious pages, allowing for the automatic extraction of keywords that characterize inducing elements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual keyword setting by analysts is used to detect inducing elements, then the detection process can be performed with simple tools, but the accuracy and efficiency deteriorate due to the large number of web pages and character strings involved
Solution Approach 1:
The system performs self-service by automatically extracting keywords from web page content and metadata without requiring manual analyst intervention. The keyword extraction unit autonomously analyzes web page data, identifies inducing elements, and generates keywords based on statistical frequency and relevance metrics, enabling the system to serve itself rather than relying on external human operators.
Solution Approach 2:
The patent replaces the mechanical manual process of keyword setting with an automated computational system. Instead of analysts manually reviewing web pages and setting keywords, the system uses processing circuits to automatically extract and analyze web page data, apply statistical algorithms, and generate keywords programmatically, substituting human mechanical work with automated information processing.
2Productivity
If manual keyword setting is used, then the system complexity remains low, but the productivity and comprehensiveness of analysis deteriorate
Solution Approach 1:
The system segments the keyword extraction process into distinct functional units: a web page data acquisition unit that collects data, a keyword extraction unit that processes the data, and a keyword determination unit that finalizes the keywords. This segmentation allows each unit to specialize in specific tasks, improving overall productivity while managing complexity through modular design.
Solution Approach 2:
The patent introduces an intermediary keyword extraction unit that acts as a mediator between raw web page data and the final keyword database. This intermediary layer processes and transforms unstructured web page content into structured keyword data, enabling efficient analysis without directly exposing the complexity of the underlying processing mechanisms to the rest of the system.
3Reliability
If comprehensive keyword coverage is attempted manually, then detection accuracy may improve, but the time and resources required increase significantly
Solution Approach 1:
The system changes parameters by using statistical frequency analysis and relevance scoring to dynamically determine keyword importance. Instead of manually reviewing all possible character strings, the system automatically adjusts keyword selection based on quantitative metrics such as occurrence frequency, positional importance, and relevance to inducing elements, achieving comprehensive coverage through parameter-driven automation.
Solution Approach 2:
The patent implements feedback mechanisms where the system continuously refines keyword extraction based on analysis results. The keyword determination unit uses feedback from the extraction process to adjust and optimize keyword selection, improving detection reliability over time while automatically managing the volume of data processed through iterative refinement rather than exhaustive manual review.
Data Source
AI summary
An extraction apparatus includes processing circuitry configured to receive an input of information about a plurality of web pages including a hypertext markup language (HTML) element that is known to reach a malicious web page through browser operation and an HTML element that is known to reach a benign web page through browser operation, classify the plurality of web pages whose input is received into clusters, extract an HTML element that reaches the malicious web page and an HTML element that reaches the benign web page from a web page of each cluster that is classified to extract a first character string included in HTML elements that are extracted, and extract, as a keyword, a second character string that characterizes the HTML element that reaches the malicious web page from the first character string.


