Harmful Site Detection via URL Link Circulation Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing harmful site blocking techniques rely heavily on manual maintenance of databases, which is time-consuming and inefficient due to the constant emergence of new harmful sites and changing website content and addresses.
Innovation Solution
A device and method that automatically determines harmful sites by analyzing URL connections, normalizing URLs, and calculating link circulations using a database to identify and rank potentially harmful sites, thereby reducing computational load and enhancing efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual maintenance of harmful site database is used, then the database can be kept updated, but it takes too much time and is difficult to maintain due to constant emergence of new harmful sites
Solution Approach 1:
The system automatically detects and updates harmful sites through web crawlers and link circulation analysis without requiring manual intervention. The harmful site collection device autonomously crawls websites, analyzes link patterns, identifies harmful sites based on connection patterns, and updates the database automatically, making the maintenance process self-service oriented.
Solution Approach 2:
The patent replaces manual mechanical maintenance with automated computational systems. Web crawlers automatically collect URL data, algorithms automatically analyze link circulations, and software automatically updates the harmful site database, substituting human manual work with automated mechanical/computational processes.
2Measurement precision
If real-time analysis method is used to determine harmful sites, then accurate detection can be achieved, but the process is complex and time-consuming compared to database blocking
Solution Approach 1:
The system performs preliminary analysis by pre-collecting URL data from web pages and pre-analyzing link circulations before actual harmful site detection is needed. The web crawler continuously collects URL information and the system pre-processes this data to identify potential harmful sites, so when detection is required, the work is already partially done, reducing real-time complexity.
Solution Approach 2:
The patent introduces link circulation analysis as an intermediary method between simple database blocking and complex real-time analysis. Instead of directly analyzing all website contents in real-time, the system uses link circulation patterns as an intermediate indicator to indirectly identify harmful sites, simplifying the detection process while maintaining accuracy.
3Reliability
If all extracted URLs are processed to check link circulation, then comprehensive harmful site identification can be achieved, but computational load increases significantly
Solution Approach 1:
The system extracts and removes duplicate URLs and non-harmful site URLs from the processing queue before conducting link circulation analysis. By taking out unnecessary URLs that would not contribute to harmful site detection, the system reduces the number of computations required while maintaining the completeness of harmful site identification.
Solution Approach 2:
The patent applies partial action by selectively processing only those URLs that are likely to be harmful based on initial filtering criteria, rather than processing all extracted URLs equally. The system performs link circulation analysis on a subset of URLs that pass initial filters, achieving sufficient detection coverage with reduced computational energy.
Data Source
AI summary
Provided are a harmful site collection device and method for determining a harmful site by analyzing a connection between harmful sites. The harmful site collection device extracts a URL linked to a web page of a harmful site; checks a link circulation on the basis of link information on a web page of the URL linked to the harmful site to determine whether the web page of the URL linked to the harmful site is a harmful site; and, when a URL of a prestored non-harmful site is extracted while the link circulation is checked, stops checking the link circulation that includes the URL of the non-harmful site. Accordingly, the harmful site collection device can more easily determine a harmful site merely with information on a URL linked to a web page and can reduce the amount of computation using information on a URL of a prestored non-harmful site.


