ML Detection of Sensitive Resource Collection via Global Invariant Features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for detecting sensitive resource collection, such as phishing attacks, face challenges with high training data requirements and computational burdens, especially when dealing with brand identification in webpages, leading to inefficiencies in browser security and endpoint protection.
Innovation Solution
A machine learning-based system utilizing few-shot learning techniques, which reduces data requirements by using entire screenshots and focusing on globally invariant features, rather than localized elements like logos, to detect sensitive resource collection with enhanced accuracy and scalability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional machine learning methods are used for brand identification in webpages, then detection accuracy can be achieved, but training data requirements and computational burden increase significantly
Solution Approach 1:
The patent extracts and focuses on specific globally invariant features (such as layout structure, color schemes, and architectural patterns) from webpage screenshots, rather than using entire images or localized elements like logos. This extraction of essential features reduces the dimensionality of the data while maintaining detection accuracy, thereby reducing training data requirements.
Solution Approach 2:
The patent applies local quality by treating different regions of the webpage screenshot differently - focusing computational attention on globally invariant features that remain consistent across brand websites, rather than uniformly processing all image data. This selective processing reduces computational burden while maintaining accuracy.
2Reliability
If traditional machine learning methods are used for brand identification, then detection capability is achieved, but computational overhead and processing time increase
Solution Approach 1:
The patent extracts only the essential globally invariant features from webpage screenshots, such as layout structure, color palettes, and architectural patterns, rather than processing entire high-resolution images. This feature extraction significantly reduces computational overhead and processing time while maintaining detection capability.
Solution Approach 2:
The patent changes the parameter representation from raw pixel data to extracted feature vectors that capture globally invariant properties. This parameter transformation reduces the computational complexity of subsequent machine learning operations while preserving the essential information needed for reliable detection.
3Measurement precision
If localized elements like logos are used for detection, then brand identification is achieved, but detection robustness decreases due to element absence or variation
Solution Approach 1:
Instead of focusing on localized elements like logos that may be absent or modified, the patent inverts the approach by focusing on globally invariant features that remain consistent across all brand websites. This inversion from local to global feature detection improves both accuracy and robustness.
Solution Approach 2:
The patent develops a detection approach that works universally across different brand websites by identifying globally invariant features that are common to all instances of a brand's web presence. This universal feature set enables reliable detection regardless of specific webpage content or localized element variations.
Data Source
AI summary
Obtaining one or more metrics associated with a network location. Determining, based on the one or more metrics and one or more prefatory check conditions, a prefatory status of the network location, the prefatory status indicating a benign status, malicious status, or a suspicious status. If the prefatory status of the network location indicates the benign status or the malicious status, providing a notification of the prefatory status in response to the prefatory status being determined. If the prefatory status of the network location indicates a suspicious status, obtaining a document object model of the network location. Obtaining a screenshot of an entire page of content at the network location. Generating a null hypothesis based on the document object model, the null hypothesis including a potential brand list, the potential brand list including one or more potential brands. Obtaining a set of reference images for each of the one or more potential brands of the potential brand list. Extracting one or more globally invariant visual features from the screenshot of the entire page of the content. Generating, based on a machine learning model using the one or more globally invariant visual features, an alternate hypothesis, the alternate hypothesis indicating a list of potentially malicious content brands. Determining, based on the null hypothesis and the alternate hypothesis and the machine learning model, a classification result. Performing one or more responsive actions in response to determining the classification result.


