Webpage Malicious Content Detection via Structure and Language Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Credential stuffing attacks, where compromised credentials are reused across multiple accounts, pose a significant threat to organizations due to the ease of access and monetization of stolen information, leading to security compromises and financial losses.

Innovation Solution

A system is developed to detect and flag malicious webpages hosting web account checkers by calculating baseline page structure and language scores, identifying and flagging pages that exceed predetermined thresholds, thereby aiding in the identification and removal of credential stuffing websites.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If automated tools and scripts are made widely available for credential stuffing attacks, then the ease of operation for attackers is improved, but the security of target systems deteriorates due to increased attack volume

Engineering Contradiction:
Improveease of access to attack toolsVSAvoidsecurity compromises
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary detection and identification of malicious webpages before they can be used for credential stuffing attacks. By calculating baseline page structure scores and language scores, and comparing them against thresholds, the system proactively identifies and flags malicious sites, preventing their use in attacks before harm occurs.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If credential stuffing tools are hosted on websites and made widely available, then the productivity of attackers is improved, but the loss of information to organizations worsens due to increased security breaches

Engineering Contradiction:
Improveattack efficiencyVSAvoidmonetary loss to consumers and merchants
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The system converts the harmful presence of credential stuffing tools on the web into a benefit by using their characteristic page structures and language patterns as detection signatures. By analyzing and scoring these patterns, the system identifies malicious websites, thereby transforming the ubiquity of attack tools from a security vulnerability into an opportunity for automated detection and takedown.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

3Measurement precision

If the system analyzes content from webpages to detect malicious activity, then the measurement precision of malicious content is improved, but the use of energy for content collection and analysis worsens

Engineering Contradiction:
Improvedetection accuracyVSAvoidcomputational resources for analysis
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system applies partial action by focusing analysis only on specific linguistic and structural features of webpages that are indicative of credential stuffing tools. Rather than analyzing all content equally, it calculates targeted page structure scores and language scores based on predetermined thresholds, achieving effective detection with reduced computational overhead compared to comprehensive content analysis.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11720742B2Detecting webpages that share malicious content
Publication Date: 2023.08.08 PAYPAL INC
  • US11720742B2 patent drawing
  • US11720742B2 patent drawing
  • US11720742B2 patent drawing

AI summary

Methods and systems for detecting webpages that share malicious content are presented. A first set of webpages that hosts a web account checker is identified. A baseline page structure score and a baseline language score are calculated based on the identified first set of webpages. Content from a second set of webpages is collected and analyzed based on the calculated baseline page structure and the calculated baseline language scores. One or more of the second set of webpages is flagged as malicious based on the analyzing of the content collected from the second set of webpages.