Web Page Structure Trait Analysis for Malicious Content Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods fail to effectively detect and prevent unwanted web contents, such as phishing, pornography, and malicious websites, especially those with changing domain names and mutating content, due to their sophisticated techniques and ease of implementation.

Innovation Solution

A system that generates and compares page structure traits by extracting markup language tags from web pages to identify and detect unwanted content, using a network of endpoint computers, support servers, and update servers to provide feedback and updates for filtering and blocking such content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If URL filtering is used to detect unwanted web contents, then detection capability is improved, but websites can evade detection by changing domain names

Engineering Contradiction:
Improvedetection capabilityVSAvoidevasion capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent extracts the domain name from the URL and separates it from the content analysis. By focusing on content structure traits rather than domain names, the system removes the evasive element (domain name changes) from the detection equation, allowing detection to proceed independently of URL filtering limitations

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of detecting unwanted content through URL/domain analysis (traditional approach), the patent inverts the approach by analyzing the structural traits of the web page content itself. This inversion makes detection independent of domain name changes, as the structural patterns remain consistent even when domains change

Inventive Principle:
Principle #13The other way round (Inversion)

2Ease of manufacture

If traditional detection methods are used, then implementation is simple, but detection effectiveness against sophisticated techniques is poor

Engineering Contradiction:
Improveimplementation simplicityVSAvoiddetection effectiveness
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent introduces page structure traits as an intermediary between the web page content and the detection process. These traits serve as a mediator that captures essential structural characteristics without requiring complex analysis of the actual content, maintaining implementation simplicity while improving detection effectiveness against sophisticated techniques

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the detection parameter from domain names and surface-level URL characteristics to structural traits of the web page markup. This parameter change allows the system to detect sophisticated unwanted content that traditional methods miss, while the structural analysis remains computationally efficient

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If content analysis is performed on all web pages, then detection accuracy is improved, but processing time and computational resources increase

Engineering Contradiction:
Improvedetection accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts only the structural traits from web pages, taking out the essential detection information while leaving out the time-consuming content analysis. By focusing solely on markup structure rather than full content processing, the system achieves good detection accuracy with significantly reduced processing time

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by performing only structural analysis rather than complete content analysis. This partial approach focuses computational resources on the most discriminative features (structural traits) while avoiding unnecessary processing of content that would increase time consumption without proportionally improving detection accuracy

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9811664B1Methods and systems for detecting unwanted web contents
Publication Date: 2017.11.07 TREND MICRO INC
  • US9811664B1 patent drawing
  • US9811664B1 patent drawing
  • US9811664B1 patent drawing

AI summary

Unwanted web contents are detected in an endpoint computer. The endpoint computer receives a web page from a website. The reputation of the website is determined and the web page is scanned for malicious codes to protect the endpoint computer from web threats. To further protect the endpoint computer from web threats including mutating unwanted web contents, page structure traits of the web page are generated and compared to page structure traits of other web pages detected to contain unwanted web contents.