Web Page Structure Trait Analysis for Malicious Content Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods fail to effectively detect and prevent unwanted web contents, such as phishing, pornography, and malicious websites, especially those with changing domain names and mutating content, due to their sophisticated techniques and ease of implementation.
Innovation Solution
A system that generates and compares page structure traits by extracting markup language tags from web pages to identify and detect unwanted content, using a network of endpoint computers, support servers, and update servers to provide feedback and updates for filtering and blocking such content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If URL filtering is used to detect unwanted web contents, then detection capability is improved, but websites can evade detection by changing domain names
Solution Approach 1:
The patent extracts the domain name from the URL and separates it from the content analysis. By focusing on content structure traits rather than domain names, the system removes the evasive element (domain name changes) from the detection equation, allowing detection to proceed independently of URL filtering limitations
Solution Approach 2:
Instead of detecting unwanted content through URL/domain analysis (traditional approach), the patent inverts the approach by analyzing the structural traits of the web page content itself. This inversion makes detection independent of domain name changes, as the structural patterns remain consistent even when domains change
2Ease of manufacture
If traditional detection methods are used, then implementation is simple, but detection effectiveness against sophisticated techniques is poor
Solution Approach 1:
The patent introduces page structure traits as an intermediary between the web page content and the detection process. These traits serve as a mediator that captures essential structural characteristics without requiring complex analysis of the actual content, maintaining implementation simplicity while improving detection effectiveness against sophisticated techniques
Solution Approach 2:
The patent changes the detection parameter from domain names and surface-level URL characteristics to structural traits of the web page markup. This parameter change allows the system to detect sophisticated unwanted content that traditional methods miss, while the structural analysis remains computationally efficient
3Measurement precision
If content analysis is performed on all web pages, then detection accuracy is improved, but processing time and computational resources increase
Solution Approach 1:
The patent extracts only the structural traits from web pages, taking out the essential detection information while leaving out the time-consuming content analysis. By focusing solely on markup structure rather than full content processing, the system achieves good detection accuracy with significantly reduced processing time
Solution Approach 2:
The patent applies partial action by performing only structural analysis rather than complete content analysis. This partial approach focuses computational resources on the most discriminative features (structural traits) while avoiding unnecessary processing of content that would increase time consumption without proportionally improving detection accuracy
Data Source
AI summary
Unwanted web contents are detected in an endpoint computer. The endpoint computer receives a web page from a website. The reputation of the website is determined and the web page is scanned for malicious codes to protect the endpoint computer from web threats. To further protect the endpoint computer from web threats including mutating unwanted web contents, page structure traits of the web page are generated and compared to page structure traits of other web pages detected to contain unwanted web contents.


