ML Detection of Sensitive Resource Collection via Global Invariant Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for detecting sensitive resource collection, such as phishing attacks, face challenges with high training data requirements and computational burdens, especially when dealing with brand identification in webpages, leading to inefficiencies in browser security and endpoint protection.

Innovation Solution

A machine learning-based system utilizing few-shot learning techniques, which reduces data requirements by using entire screenshots and focusing on globally invariant features, rather than localized elements like logos, to detect sensitive resource collection with enhanced accuracy and scalability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional machine learning methods are used for brand identification in webpages, then detection accuracy can be achieved, but training data requirements and computational burden increase significantly

Engineering Contradiction:
Improvedetection accuracyVSAvoidtraining data requirements
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts and focuses on specific globally invariant features (such as layout structure, color schemes, and architectural patterns) from webpage screenshots, rather than using entire images or localized elements like logos. This extraction of essential features reduces the dimensionality of the data while maintaining detection accuracy, thereby reducing training data requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies local quality by treating different regions of the webpage screenshot differently - focusing computational attention on globally invariant features that remain consistent across brand websites, rather than uniformly processing all image data. This selective processing reduces computational burden while maintaining accuracy.

Inventive Principle:
Principle #3Local quality

2Reliability

If traditional machine learning methods are used for brand identification, then detection capability is achieved, but computational overhead and processing time increase

Engineering Contradiction:
Improvedetection capabilityVSAvoidcomputational overhead
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent extracts only the essential globally invariant features from webpage screenshots, such as layout structure, color palettes, and architectural patterns, rather than processing entire high-resolution images. This feature extraction significantly reduces computational overhead and processing time while maintaining detection capability.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameter representation from raw pixel data to extracted feature vectors that capture globally invariant properties. This parameter transformation reduces the computational complexity of subsequent machine learning operations while preserving the essential information needed for reliable detection.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If localized elements like logos are used for detection, then brand identification is achieved, but detection robustness decreases due to element absence or variation

Engineering Contradiction:
Improvebrand identification accuracyVSAvoiddetection robustness
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

Instead of focusing on localized elements like logos that may be absent or modified, the patent inverts the approach by focusing on globally invariant features that remain consistent across all brand websites. This inversion from local to global feature detection improves both accuracy and robustness.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent develops a detection approach that works universally across different brand websites by identifying globally invariant features that are common to all instances of a brand's web presence. This universal feature set enables reliable detection regardless of specific webpage content or localized element variations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12041085B2Machine learning-based sensitive resource collection agent detection
Publication Date: 2024.07.16 ZOHO OFFICE SUITE
  • US12041085B2 patent drawing
  • US12041085B2 patent drawing
  • US12041085B2 patent drawing

AI summary

Obtaining one or more metrics associated with a network location. Determining, based on the one or more metrics and one or more prefatory check conditions, a prefatory status of the network location, the prefatory status indicating a benign status, malicious status, or a suspicious status. If the prefatory status of the network location indicates the benign status or the malicious status, providing a notification of the prefatory status in response to the prefatory status being determined. If the prefatory status of the network location indicates a suspicious status, obtaining a document object model of the network location. Obtaining a screenshot of an entire page of content at the network location. Generating a null hypothesis based on the document object model, the null hypothesis including a potential brand list, the potential brand list including one or more potential brands. Obtaining a set of reference images for each of the one or more potential brands of the potential brand list. Extracting one or more globally invariant visual features from the screenshot of the entire page of the content. Generating, based on a machine learning model using the one or more globally invariant visual features, an alternate hypothesis, the alternate hypothesis indicating a list of potentially malicious content brands. Determining, based on the null hypothesis and the alternate hypothesis and the machine learning model, a classification result. Performing one or more responsive actions in response to determining the classification result.