Machine learning based system and method using URL feature hashing, HTML encoding, and content page embedded images to detect phishing websites

A machine learning system using URL hashing, HTML encoding, and image embeddings with BERT and ResNet transfer learning addresses the challenge of real-time phishing detection, reducing false positives and adapting to new threats efficiently.

JP7766187B2Active Publication Date: 2025-11-07NETSKOPE INC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2024516418
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-09-14
Filing Date
2022-09-13
Publication Date
2025-11-07
Estimated Expiration
2042-09-13

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively detect phishing websites in real-time due to the transient nature of phishing campaigns and the high computational complexity of deep learning architectures, leading to high false positive rates and the inability to react quickly to new phishing threats.

Method used

A machine learning-based system utilizing URL feature hashing, HTML encoding, and content page embedded images, employing transfer learning techniques with deep learning architectures like BERT and ResNet, to classify URLs and content pages as phishing or non-phishing, enabling real-time detection and reducing false positives.

Benefits of technology

The system achieves low false positive rates and effective real-time detection of phishing websites, leveraging transfer learning to quickly adapt to new threats and neutralize suspicious URLs without requiring website visits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007766187000001
    Figure 0007766187000001
  • Figure 0007766187000002
    Figure 0007766187000002
  • Figure 0007766187000003
    Figure 0007766187000003
Patent Text Reader

Abstract

A phishing classifier is disclosed for classifying URLs and content pages as phishing or not, comprising a URL feature hasher that parses and hashes URLs into feature hashes, and a headless browser that visits and internally renders the pages of the URLs, extracts HTML tokens, and captures an image of the rendering. Also disclosed is a phishing classifier for classifying URLs and content pages accessed via the URLs as phishing or not, comprising a URL feature hasher that parses and hashes URLs into feature hashes, and a headless browser that visits and internally renders the pages of the URLs, extracts words from the rendering, and captures an image of the pages. Classifying URLs and content pages accessed via the URLs as phishing or not is further disclosed. In addition to one or more of the disclosures, there is a phishing classification layer, a URL embedder, and an HTML encoder.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and benefit of:

[0002] U.S. Application No. 17 / 475,236, entitled "A Machine Learning-Based System for Detecting Phishing Websites Using the URLs, Word Encodings and Images of Content Pages," filed September 14, 2021, which is now U.S. Patent No. 11,444,978 (Attorney Docket No. NSKO1052-1), issued September 13, 2022; and

[0003] U.S. Application No. 17 / 475,233, entitled "Detecting Phishing Websites via a Machine Learning-Based System Using URL Feature Hashes, HTML Encodings and Embedded Images of Content Pages," filed September 14, 2021, which is now U.S. Patent No. 11,336,689 (Attorney Docket No. NSKO1060-1), issued May 17, 2022; and

[0004] U.S. Application No. 17 / 475,230, filed September 14, 2021, entitled "Machine Learning-Based Systems and Methods of Using URLs and HTML Encodings for Detecting Phishing Websites," which is now U.S. Patent No. 11,438,377, issued September 6, 2022 (Attorney Docket No. NSKO1061-1).

[0005] Related Stories This application is also related to the following applications, which are incorporated by reference for all purposes as if fully set forth herein:

[0006] U.S. Application No. 17 / 390,803 (Attorney Docket No. 1037-2), entitled "Preventing Cloud-Based Phishing Attacks Using Shared Documents with Malicious Links," filed on July 30, 2021, is a continuation of U.S. Application No. 17 / 154,978, entitled "Preventing Phishing Attacks Via Document Sharing," filed on January 21, 2021, which is now U.S. Application No. 11,082,445 (Attorney Docket No. NSKO1037-1), issued on August 3, 2021.

[0007] Reference The following materials are incorporated by reference in this application: “KDE Hyper Parameter Determination,” Yi Zhang et al., Netskope, Inc. U.S. Non-Provisional Application No. 15 / 256,483 (Attorney Docket No. NSKO1004-2), entitled "Machine Learning Based Anomaly Detection," filed September 2, 2016 (now U.S. Patent No. 10,270,788, issued April 23, 2019); U.S. Non-Provisional Application No. 16 / 389,861 (Attorney Docket No. NSKO1004-3), entitled "Machine Learning Based Anomaly Detection," filed April 19, 2019 (now U.S. Patent No. 11,025,653, issued June 1, 2021); U.S. Non-Provisional Application No. 14 / 198,508 (Attorney Docket No. NSKO1000-3), entitled "Security For Network Delivered Services," filed March 5, 2014 (now U.S. Patent No. 9,270,765, issued February 23, 2016); U.S. Non-provisional Application No. 15 / 368,240 (Attorney Docket No. NSKO1003-2), entitled "Systems and Methods of Enforcing Multi-Part Policies on Data-Deficient Transactions of Cloud Computing Services," filed December 2, 2016 (now U.S. Patent No. 10,826,940, issued November 3, 2020), and U.S. Provisional Application No. 62 / 307,305 (Attorney Docket No. NSKO1003-1), entitled "Systems and Methods of Enforcing Multi-Part Policies on Data-Deficient Transactions of Cloud Computing Services," filed March 11, 2016; “Cloud Security for Dummies,Netskope Special Edition”by Cheng,Ithal,Narayanaswamy,and Malmskog,John Wiley&Sons,Inc.2015; “Netskope Introspection” by Netskope,Inc.; “Data Loss Prevention and Monitoring in the Cloud” by Netskope,Inc.; “The 5 Steps to Cloud Confidence” by Netskope, Inc.; “Netskope Active Cloud DLP” by Netskope,Inc.; "Repave the Cloud-Data Breach Collision Course" by Netskope, Inc., and “Netskope Cloud Confidence Index(TM)” by Netskope,Inc.

[0008] The disclosed technology relates generally to cloud-based security, and more specifically to systems and methods for detecting phishing websites using URLs, word encoding, and images of content pages. Also disclosed are methods and systems for using URL feature hashes, HTML encoding, and embedded images of content pages. The disclosed technology further relates to detecting phishing in real time via URL links and downloaded HTML through machine learning and statistical analysis. [Background technology]

[0009] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, it should not be assumed that the problems mentioned in this section, or problems related to the subject matter provided as background, have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which may themselves correspond to implementations of the claimed technology.

[0010] Phishing, sometimes called spearhead phishing, is on the rise. The exploitation of documents obtained using stolen passwords through phishing has punctuated national news. Typically, emails contain legitimate-looking links, leading to legitimate-looking pages where users enter their passwords, exposing them to phishing attacks. Cleaver phishing sites, such as credit card skimmers or gas pump or ATM shims, can forward entered passwords to real websites, taking them out of the loop so users don't detect the password theft when it occurs. Working from home has led to a significant increase in phishing attacks in recent years.

[0011] The term phishing refers to several methods for fraudulently obtaining confidential information from unsuspecting users over the web. Phishing stems, in part, from the use of increasingly sophisticated lures to extract confidential company information. These methods are commonly referred to as phishing attacks. Website users fall victim to phishing attacks when the rendered web page mimics the appearance of a legitimate login page. Victims of phishing attacks are lured to fraudulent websites, resulting in the disclosure of sensitive information such as bank account details, login passwords, and social security IDs.

[0012] Recent data breach investigation reports indicate that the trend of large-scale attacks based on social engineering is on the rise. This can be attributed in part to the increasing difficulty of exploits, and in part to the use of advances in machine learning (ML) algorithms to prevent and detect such exploits. Consequently, phishing attacks are becoming more frequent and sophisticated. New defensive solutions are needed.

[0013] Using ML / DL, an opportunity arises to classify URLs and the content pages accessed via these URLs as phishing or not phishing, and also to classify URLs and the HTML accessed and downloaded via these URL links as phishing or not in real time.

[0014] In the drawings, like reference characters generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead generally being placed upon illustrating the principles of the disclosed technology. In the following description, various implementations of the disclosed technology are described with reference to the following drawings: [Brief explanation of the drawings]

[0015] [Figure 1]According to an implementation of the disclosed technology, an architectural level diagram of a system is shown illustrating the classification of URLs and content pages accessed via the URLs as phishing or non-phishing. [Figure 2] 1 illustrates a high-level block diagram of the disclosed phishing detection engine, which utilizes ML / DL encoding of URL feature hashes, encoding of natural language (NL) words, and encoding of captured website images to detect phishing sites. [Figure 3] 1 illustrates an exemplary ResNet residual CNN block diagram for image classification for reference. [Figure 4] 1 illustrates a high-level block diagram of the disclosed phishing detection engine, where each example URL uses ML / DL with URL feature hashing, encoding of HTML extracted from the content page, and embedding of images captured from the example URL's content page, with a ground truth classification as phishing or not, to detect phishing sites. [Figure 5] 1 illustrates a block diagram of a reference residual neural network (ResNet) that is pre-trained for image classification before use in a phishing detection engine. [Figure 6] 6 illustrates a high-level block diagram of the disclosed phishing detection engine 602 utilizing ML / DL with a URL embedder and HTML encoder. [Figure 7] 1 shows precision-recall graphs for several of the disclosed phishing detection systems. [Figure 8] 1 illustrates an exemplary Receiver Operating Characteristic Curve (ROC) of the disclosed system for phishing website detection described herein. [Figure 9]1 illustrates an example receiver operating characteristic curve (ROC) for phishing website detection of the disclosed phishing detection engine utilizing ML / DL with URL embedder and HTML encoder. [Figure 10] 1 illustrates a block diagram of the functionality of a one-dimensional 1D convolutional neural network (Conv1D) URL embedder that generates URL embeddings using C++ code expressed in the Open Neural Network Exchange (ONNX) format. [Figure 11] 1 shows a block diagram of the functionality of the disclosed html encoder, which provides the html encoding that is input to the phishing classifier layer. [Figure 12A] 1 shows a schematic block diagram of the disclosed html encoder, which provides html encoding that is input to a phishing classifier layer. [Figure 12B] 6 illustrates, using C++ code expressed in Open Neural Network Exchange (ONNX) format, together with a computational dataflow graph of the functionality of the html encoder that results in the html encoding that is input to the phishing classifier layer 675. A section of the dataflow graph is shown in two columns separated by a dotted line, with the connector at the bottom of the left column flowing into the top of the right column, illustrating the input encoding and positional embedding. [Figure 12C] 6 illustrates a computational dataflow graph of the functionality of an html encoder using C++ code represented in Open Neural Network Exchange (ONNX) format, resulting in html encoding input to a phishing classifier layer 675. FIG. 7 illustrates a single iteration of multi-head attention, showing an example of a dataflow graph with computational nodes asynchronously transmitting data along data connections. [Figure 12D]6 illustrates, using C++ code expressed in Open Neural Network Exchange (ONNX) format, together with a computational dataflow graph of the functionality of the html encoder that results in the html encoding that is input to the phishing classifier layer 675. The summation, normalization, and feedforward functionality using ONNX operations is illustrated using three columns separated by dotted lines. [Figure 13] 1 illustrates a computational data flow graph of the disclosed phishing classifier layer functionality, using code represented in Open Neural Network Exchange (ONNX) format to generate likelihood score(s) signaling how likely a particular website is to be a phishing website. [Figure 14] FIG. 1 is a simplified block diagram of a computer system that may be used to classify a URL and a content page accessed via the URL as phishing or non-phishing, according to one implementation of the disclosed technology. DETAILED DESCRIPTION OF THE INVENTION

[0016] The following detailed description is made with reference to the drawings. Sample implementations are described to illustrate the disclosed technology, but not to limit its scope, as defined by the claims. This discussion is presented to enable any person skilled in the art to make and use the disclosed technology, and is provided in the context of a particular application and its requirements. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other implementations and applications without departing from the spirit and scope of the invention. Thus, the disclosed technology is not intended to be limited to the implementations shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0017] The problem addressed by the disclosed technology is the detection of phishing websites. Security teams attempt to catalog phishing campaigns as they occur. Security vendors rely on lists of phishing websites to power their security engines. Both proprietary and open-source sources for cataloging phishing links are available. Two open-source community examples of phishing universal resource locator (URL) lists are PhishTank and OpenPhish. Security teams use the lists to analyze malicious links and generate signatures from malicious URLs. Signatures are used to detect malicious links, typically by matching part or all of the URL or its compact hash. Generalization from signatures has become the primary approach for preventing zero-day phishing attacks that hackers can use to attack systems. A zero-day refers to a recently discovered security vulnerability that a vendor or developer has just learned about and has zero days to fix.

[0018] To avoid getting caught, phishing campaigns often end before the phishing link's website can be analyzed. Phishers can take down websites as soon as they are listed by security departments. Analysis of harvested URLs is more likely to persist than tracing malicious URLs to active phishing sites. Sites disappear as suddenly as they appeared. Due in part to disappearing sites, state-of-the-art technology has become URL analysis.

[0019] The disclosed technology applies machine learning / deep learning (ML / DL) to phishing detection with very low false positive rates and good recall. Three transfer learning techniques are presented: one based on text / image analysis and one based on HTML analysis.

[0020] In the first technique, we use transfer learning by leveraging novel deep learning architectures for multilingual natural language understanding and computer vision to embed the textual and visual content of web pages. The first generation of ML / DL applications to phishing detection uses concatenated embeddings of web page text and web page images. We leverage transfer learning from general training on text and image embeddings to train a detection classifier that uses the model's encoder function, such as the Bidirectional Encoder Representation from Transformer (BERT) and the decoder function of a Residual Neural Network (ResNet). Because they are trained on large amounts of data, the final layer of such a model serves as a reliable encoding for the visual and textual content of web pages. Because benign, non-phishing links are much more abundant than phishing sites, and blocking non-phishing links is a nuisance, care is taken to reduce false positives.

[0021] The second technique for applying ML / DL to phishing detection creates a new encoder-decoder pair that counterintuitively decodes HTML embeddings to replicate the browser's rendering on the display. Embeddings, of course, are lossy; decoding is much less accurate than what the browser achieves. The encoder-decoder approach to embedding HTML code facilitates transfer learning. Once the encoder is trained to embed HTML, a classifier replaces the decoder. Embedding-based transfer learning requires a relatively small training corpus and is practical. Currently, as few as 20k or 40k examples of phishing pages have proven sufficient to train a classifier with two fully connected layers that process embeddings. Second-generation embeddings of HTML can be enhanced by concatenating other embeddings, such as ResNet image embeddings, URL feature embeddings, or both ResNet image embeddings and URL feature embeddings.

[0022] However, the scale of new URLs may hinder real-time detection of web pages that use these contents due to the high computational complexity of deep learning architectures and the rendering and parsing times of web page content.

[0023] The third generation, which applies ML / DL to phishing detection, uses a URL embedder, HTML encoder, and phishing classifier layer to classify URLs and content pages accessed through these URLs as phishing or non-phishing, and can react in real time when malicious web pages are detected. This third technology effectively filters suspicious URLs using faster trained models without requiring website visits. Suspicious URLs can also be routed back to the first or second technology for final detection later.

[0024] Next, an exemplary system for detecting phishing via URL links and downloaded HTML in offline mode and in real time is described.

[0025] architecture FIG. 1 shows an architectural level schematic diagram of a system 100 for detecting phishing via URL links and downloaded HTML. The system 100 also includes functionality for detecting phishing via redirected or hidden URL links and real-time downloaded HTML. Because FIG. 1 is an architectural diagram, some details have been intentionally omitted for clarity. The description of FIG. 1 is structured as follows: first, the elements of the diagram are described, followed by a description of their interconnections. Next, the use of the elements in the system is described in more detail.

[0026] FIG. 1 includes a system 100 that includes an endpoint 166. The user endpoint 166 may include devices such as a computer 174, a smartphone 176, and a computer tablet 178 that provide access to and interaction with data stored on the cloud-based store 136 and the cloud-based services 138. In another organizational network, organizational users may utilize additional devices. An inline proxy 144 intervenes between the user endpoint 166 and the cloud-based services 138 through a network 155, particularly through a network security system 112 that includes a network administrator 122, a network policy 132, a rating engine 152, and a data store 164. The inline proxy 144 is accessible through the network 155 as part of the network security system 112. The inline proxy 144 provides monitoring and control of traffic between the user endpoint 166, the cloud-based store 136, and the other cloud-based services 138. The inline proxy 144 includes an active scanner 154 that collects HTML and web page snapshots and stores the data sets in a data store 164. Because features can be extracted from traffic in real time and snapshots are not collected from live traffic, an active scanner 154 is not required to crawl web page content for URLs, as in third-generation systems that apply ML / DL to phishing detection. Three ML / DL systems for detecting phishing websites are described in detail below. The inline proxy 144 monitors network traffic between user endpoints 166 and cloud-based services 138 to enforce network security policies, including, among other things, data loss prevention (DLP) policies and protocols. The reputation engine 152 checks database records for URLs deemed malicious through disclosed detection of phishing websites, and these phishing URLs are automatically and permanently blocked.

[0027] To detect phishing in real time via URL links and downloaded HTML, an inline proxy 144 positioned between the user endpoint 166 and the cloud-based storage platform inspects incoming traffic and forwards it to a phishing detection engine 202, 404, 602, described below. The inline proxy 144 may be configured to sandbox the content corresponding to the link and inspect / probe the link to ensure that the page pointed to by the URL is safe before allowing the user to access the page through the proxy. Links identified as malicious can then be quarantined and inspected for threats using known techniques, including secure sandboxing.

[0028] Continuing with FIG. 1 , cloud-based services 138 include cloud-based hosting services, web email services, video, messaging, and voice calling services, streaming services, file transfer services, and cloud-based storage services. Network security system 112 connects user endpoints 166 and cloud-based services 138 via public network 155. Data store 164 stores a list of malicious links and signatures from malicious URLs. The signatures are typically used to detect malicious links by matching part or all of the URL or its compact hash. Data store 164 stores information from one or more tenants in tables of a common database image to form an on-demand database service (ODDS), which can be implemented in many ways, such as a multi-tenant database system (MTDS). The database image can include one or more database objects. In other implementations, the database can be a relational database management system (RDBMS), an object-oriented database management system (OODBMS), a distributed file system (DFS), a no-schema database, or any other data storage system or computing device. In some implementations, the collected metadata is processed and / or normalized. In some cases, the metadata includes structured data, and functionality is targeted to a particular data structure provided by the cloud-based service 138. Unstructured data, such as free text, may also be provided by and targeted back to the cloud-based service 138. Both structured and unstructured data can be stored in semi-structured data formats, such as JSON (JavaScript Object Notation), BSON (Binary JSON), XML, Protobuf, Avro, or Thrift objects, where the semi-structured data format consists of string fields (or columns) and corresponding values ​​of potentially different types, such as numbers, strings, arrays, objects, etc.In other implementations, JSON objects can be nested and fields can be multi-valued, e.g., arrays, nested arrays, etc. These JSON objects are stored in a schemaless or NoSQL key-value metadata store 178, such as Apache Cassandra™, Google's Bigtable™, HBase™, Voldemort™, CouchDB™, MongoDB™, Redis™, Riak™, Neo4j™, etc., which uses keyspaces, similar to SQL databases, to store the parsed JSON objects. Each keyspace is divided into column families, which are similar to tables and consist of a set of rows and columns.

[0029] Continuing with the description of FIG. 1 , the system 100 can include any number of cloud-based services 138, namely, point-to-point streaming services, hosted services, cloud applications, cloud stores, cloud collaboration and messaging platforms, and cloud customer relationship management (CRM) platforms. Services can include peer-to-peer file sharing (P2P) via protocols for portal traffic such as BitTorrent (BT), User Datagram Protocol (UDP) streaming, and File Transfer Protocol (FTP), as well as voice, video, and messaging multimedia communication sessions such as instant messaging over Internet Protocol (IP) and mobile phone calling over LTE (VoLTE) via Session Initiation Protocol (SIP) and Skype. Services can handle Internet traffic, cloud application data, and Generic Routing Encapsulation (GRE) data. Network services or applications can be web-based (e.g., accessed via a Uniform Resource Locator (URL)) or native, such as a sync client. Examples include software-as-a-service (SaaS), platform-as-a-service (PaaS), and infrastructure-as-a-service (IaaS) offerings, as well as internal enterprise applications exposed via a URL. Examples of common cloud-based services today include Salesforce.com™, Box™, Dropbox™, Google Apps™, Amazon AWS™, Microsoft Office 365™, Workday™, Oracle on Demand™, Taleo™, Yammer™, Jive™, and Concur™.

[0030] In interconnecting the elements of system 100, network 155 communicatively couples computers, tablets, and mobile devices, cloud-based hosting services, web email services, video, messaging, and voice calling services, streaming services, file transfer services, cloud-based storage services 136, and network security system 112. Communication paths can be point-to-point over public and / or private networks. Communications can occur over a variety of networks, e.g., private networks, VPNs, MPLS circuits, or the Internet, and can use appropriate application program interfaces (APIs) and data exchange formats, e.g., REST, JSON, XML, SOAP, and / or JMS. All communications can be encrypted. This communication generally occurs over networks such as local area networks (LANs), wide area networks (WANs), telephone networks (public switched telephone networks (PSTNs)), session initiation protocol (SIP), wireless networks, point-to-point networks, star networks, token ring networks, hub networks, and the Internet, including mobile Internet, via protocols such as EDGE, 3G, 4G LTE, Wi-Fi, and WiMAX. Additionally, various authorization and authentication technologies such as username / password, OAuth, Kerberos, SecureID, digital certificates, etc. may be used to secure communications.

[0031] Continuing the description of the system architecture of FIG. 1 , the network security system 112 includes a data store 164, which can include one or more computers and computer systems communicatively coupled to each other. They can also be one or more virtual computing and / or storage resources. For example, the network security system 112 can be one or more Amazon EC2 instances, and the data store 164 can be Amazon S3™ storage. Rather than implementing the network security system 112 directly on a physical computer or traditional virtual machine, other computing-as-a-service platforms, such as Salesforce's Rackspace, Heroku, or Force.com, can be used. Additionally, one or more engines can be used and one or more points of presence (POPs) can be established to implement security functions. The engines or system components of FIG. 1 are implemented by software running on various types of computing devices. Exemplary devices are workstations, servers, computing clusters, blade servers, and server farms, or any other data processing system or computing device. The engines can be communicatively coupled to a database via different network connections.

[0032] While system 100 is described herein with reference to particular blocks, it should be understood that the blocks are defined for convenience of description and are not intended to require a particular physical arrangement of components. Furthermore, the blocks need not correspond to physically separate components. To the extent physically separate components are used, connections between components can be wired and / or wireless as appropriate. Different elements or components can be combined into a single software module, and multiple software modules can execute on the same processor.

[0033] Despite the best attempts of malicious actors, the content and appearance of phishing websites provide features that the disclosed deep learning models can exploit to reliably detect phishing websites. In the disclosed system described next, we use transfer learning by utilizing a novel deep learning architecture for multilingual natural language understanding and computer vision to embed the textual and visual content of web pages.

[0034] 2 illustrates a high-level block diagram 200 of the disclosed phishing detection engine 202, which employs ML / DL with URL feature hashes, natural language (NL) word encodings, and embeddings of captured website images to detect phishing sites. The disclosed phishing classifier layer 275 generates a likelihood score 285 that represents how likely a particular website is to be a phishing website. In one embodiment, the phishing detection engine 202 utilizes a multilingual bidirectional encoder representations from a transformer (BERT) model supporting over 100 languages ​​as the encoder 264 and a residual neural network for images (ResNet50) as the embedder 256. The URL feature hashes 242, word encodings 265, and image embeddings 257 are then passed to the neural network phishing classifier layer 275 for final training and inference, as described below.

[0035] An encoder can be trained by pairing an encoder and a decoder. The encoder and decoder can be trained to compress the input into an embedding space and then reconstruct the input from the embedding. Once the encoder is trained, it can be reused as described herein. The phishing classifier layer 275 utilizes URL feature hashes 242 of the URL n-grams, word encodings 265 of words extracted from the content page, and image embeddings 257 of images captured from the content page 216 at the URL 214 web address.

[0036] In one embodiment, the phishing detection engine 202 utilizes feature hashes of the web page content and security information present in response headers to supplement the features available on both benign and phishing web pages. The content is expressed in JavaScript in one implementation. In another embodiment, a different language, such as Python, can be used. The URL feature hasher 222 receives the URL 214, parses the URL into features, and hashes the features to generate URL feature hashes 242, providing dimensionality reduction for the URL n-gram. An example of domain features for a URL with headers + security information is listed below: “scanned_url”:[ “http: / / alfabeek.com / ” ], "header":{ “date”:”Tue, 02 Mar 2021 15:30:27 GMT”, “server”:”Apache”, “last-modified”:”Tue,08 Sep 2020 02:09:49 GMT”, “accept-ranges”:”bytes”, “vary”:”Accept-Encoding”, “content-encoding”:”gzip”, “content-length”:”23859”, “content-type”:”text / html”}, “security_info”:[ { “_subjectName”:”alfabeek.com”, “_issuer”:”Sectigo RSA Domain Validation Secure Server CA”, “_validFrom”:1583107200, “_validTo”:1614729599, “_protocol”:”TLS 1.3”, “_sanList”:[ “alfabeek.com”, “www.alfabeek.com” ]

[0037] Continuing with FIG. 2 , headless browser 226 is configured to access content from a URL, internally render the content page, extract words from the content page rendering, and capture an image of at least a portion of the content page rendering. Headless browser 226 receives URL 214, which is the web address of content page 216, and extracts words from content page 216. Headless browser 226 provides extracted words 246 to natural language encoder 264, which generates encodings from the extracted words: word encoding 265 in block diagram 200. Natural language (NL) encoder 264 is pre-trained on natural language and generates encodings of the words extracted from the content page. In an exemplary embodiment, encoder 264 utilizes a standard encoder, BERT for Natural Language. The encoder embeds inputs that the encoder processes in a relatively low-dimensional embedding space. BERT embeds natural language passages in an embedding space of 400-800 dimensions. The Transformer logic accepts natural language input and, in one example, generates a 768-dimensional vector that encodes and embeds the input. The dashed block outline of the pre-trained decoder 266 distinguishes it as pre-trained, i.e., BERT is trained before being used to detect phishing in URLs 214. The encoder 264 generates word encodings 265 of words extracted from the content page being screened for use by the phishing classifier layer 275 to detect phishing. Different implementations may utilize different ML / DL encoders, such as a universal sentence encoder. Different embodiments may utilize a long short-term memory (LSTM) model.

[0038] Continuing with FIG. 2 , a headless browser 226 receives a URL 214, which is the web address of a content page 216, and captures an image of the web page by mimicking a real user visiting the web page and taking a snapshot of the rendered web page. The headless browser 226 takes the snapshot and provides the captured image 248 to an image embedder 256, which is pre-trained on the image, to generate an embedding of the captured image from the content page. Image embedding can increase efficiency and improve phishing detection for obfuscated cases. The embedder 256 encodes the captured image 248 as an image embedding 257. In one embodiment, the embedder 256 utilizes a standard embedder, a residual neural network (ResNet50), with a pre-trained classifier 258 for the image. Different implementations can utilize different ML / DL pre-trained image embedders, such as Inception-v3, VGG-16, ResNet34, or ResNet-101. Continuing with the exemplary embodiment, ResNet50 embeds an image, such as an RGB 224x224 pixel image, and generates a 248-dimensional embedding vector that maps this image into an embedding space that is much more compact than the original input. The pre-trained ResNet50 embedder 256 generates image embeddings 257 of snapshots of the content pages being screened, which are used to detect phishing websites.

[0039] The phishing classifier layer 275 of the disclosed phishing detection engine 202 is trained on URL feature hashes, encodings of words extracted from content pages, and embeddings of image captures from content pages of example URLs, with each example URL accompanied by a ground truth classification as phishing or not. The phishing classifier layer 275 processes the URL feature hashes, word encodings, and image embeddings to generate at least one likelihood score representing the phishing risk of the URL and the content accessed via the URL. The likelihood score 285 represents how likely a particular website is to be a phishing website. In one embodiment, the input size to the phishing classifier layer 275 is 2048 + 768 + 1024, the output of BERT is 768, the ResNet50 embedding size is 2048, and the size of the feature hashes over the URL n-grams is 1024. The phishing detection engine 202 is well suited for semantically meaningful detection of phishing websites, regardless of the language they are in. The disclosed near-real-time crawling pipeline quickly captures the content of new suspicious web pages before they are neutralized, thus addressing the short lifecycle nature of phishing attacks, which helps accumulate larger training datasets for continuous retraining of a given deep learning architecture.

[0040] FIG. 3 illustrates a block diagram of a Reference Bidirectional Encoder Representations from Transformer (BERT) that may be utilized for natural language classification of words extracted from web content pages, as described in connection with the block diagram shown in FIG. 2 above.

[0041] The second system for applying ML / DL to phishing detection utilizes image transfer learning and generative pretraining (GPT) to learn HTML embeddings. This addresses the issue of having limited phishing datasets and provides a better representation of HTML content. Unlike the first approach, it eliminates the need for BERT text encoding. The HTML embedding network learns to represent the entire multimodal content of HTML content (text, JS, CSS, etc.) by a vector of 256 numbers. The theoretical foundation of this HTML embedding network is presented in "Open AI Generative Pretraining From Pixels," published in the Proceedings of the 37th International Conference on Machine Learning, PMLR 119:1691-1703, 2020.

[0042] 4 illustrates a high-level block diagram 400 of the disclosed phishing detection engine 402, where each example URL uses ML / DL with URL feature hashing, encoding of HTML extracted from the content page, and embedding of images captured from the example URL's content page, with a ground truth classification as phishing or not, to detect phishing sites. The disclosed phishing classifier layer 475 generates at least one likelihood score 485 that the URL and the content accessed via this URL present a phishing risk.

[0043] The phishing detection engine 402 uses a URL feature hasher 422 that parses the URL 414 into features and hashes the features to generate URL feature hashes 442, resulting in dimensionality reduction of the URL n-grams. An example of domain features for a URL with headers + security information is listed above.

[0044] The headless browser 426 extracts HTML tokens 446 and provides them to the HTML encoder 464. The phishing detection engine 202 utilizes the disclosed HTML encoder 464 and is trained on the HTML tokens 446 extracted from the content page of the example URL 416, encoded, and then decoded to recreate the image captured from a rendering of the content page. The dashed line distinguishes the dashed block outline of the generative training decoder 466, which generates the rendered image of the page from the encoder embedding from subsequent processing of the active URL, including the dashed line from the captured image 448. Learning the data distribution P(X) is highly beneficial for subsequent supervised modeling of P(X|Y), where Y is the binary class of phishing and non-phishing, and X is the HTML content. The HTML encoder 464 is pre-trained using generative pre-training (GPT), in which unsupervised pre-training over large amounts of unsupervised data is utilized to learn the data distribution P(X) for subsequent supervised decision-making using P(Y|X). Once the HTML encoder 464 is trained, it can be reused to generate an HTML encoding 465 of the HTML tokens 446 extracted from the content page 416.

[0045] The HTML is tokenized based on rules, and the HTML tokens 446 are passed to the HTML encoder 464. Community examples of phishing URL lists, open-source, collaborative clearinghouses of phishing data and information on the Internet, include PhishTank, OpenPhish, MalwarePatrol, and Kaspersky, which serve as sources of HTML files identified as phishing websites. Negative samples that do not contain phishing balance the dataset in proportions that represent current trends in phishing websites. The HTML encoder 464 is trained using a large, unlabeled dataset of HTML and page snapshots collected by the in-house active scanner 154. Because website users fall victim to attacks, especially when malicious rendered pages mimic the appearance of legitimate login pages, the training goal forces the HTML encoder to learn to represent HTML content in terms of rendered images of these pages.

[0046] For training, the HTML encoder 464 is initialized with random initial parameters and parameters for HTML. In one embodiment, 700K HTML files in the data store 164 were scanned, and the resulting extraction of the top 10K tokens representing content pages was used to configure the phishing detection engine 402 for classification without suffering from a significant number of false positive results. In one example content page, 800 valid tokens were extracted. In another example, 2K valid tokens were recognized, and in a third example, approximately 1K tokens were collected.

[0047] In another embodiment, an HTML parser can be used to extract HTML tokens from content pages accessed via URLs. Both the headless browser and the HTML parser can be configured to extract HTML tokens that belong to a predetermined token vocabulary and to ignore portions of the content that do not belong to the predetermined token vocabulary. In one embodiment, the phishing detection engine 402 includes a headless browser configured to extract and generate an HTML encoding of up to 64 HTML tokens. 64 is a specific, configurable system parameter. The extraction can take up to 10 milliseconds to render in some cases. In another embodiment, the headless browser can be configured to extract and generate an HTML encoding of up to 128, 256, 1024, or 4096 HTML tokens. Using more tokens slows training. Implementations of up to 2k tokens have been achieved. Training can be used to learn what ordering patterns of HTML tokens result in specific page views. Mathematical approximations can be used to learn what distribution of tokens should be used later for better classification.

[0048] Continuing with FIG. 4 , an image embedder pre-trained on images generates an image embedding for an image captured from a content page. The pre-trained embedder coefficients enable near-real-time, cost-effective embedding. A headless browser 426 is configured to access content from a URL and internally render the content page. The headless browser 226 receives a URL 414, which is the web address of the content page 416, and captures an image of the web page by mimicking a real user visiting the web page and taking a snapshot of the rendered web page. The headless browser 426 takes the snapshot and provides the captured image 448 to a rendered image embedder 456 pre-trained on images, which generates an embedding for the image captured from the content page. Image embedding can increase efficiency and improve phishing detection, which is particularly useful for obfuscated cases. The embedder 456 encodes the captured image 448 as an image embedding 457. In one embodiment, the rendered image embedder 456 utilizes a standard embedder, a residual neural network (ResNet50), along with a pre-trained classifier 458 for the image. Different implementations can utilize different ML / DL pre-trained image embedders, such as Inception-v3, VGG-16, ResNet34, or ResNet-101. Continuing with the exemplary embodiment, ResNet50 embeds an image, such as an RGB 224x224 pixel image, and generates a 2048-dimensional embedding vector that maps the image into an embedding space that is much more compact than the original input. The pre-trained ResNet50 embedder 456 generates image embeddings 457 for images captured from content pages to be used to detect phishing websites. The URL feature hashes 442, HTML encodings 465, and image embeddings 457 are passed to a neural network phishing classifier layer 475 for final training and inference, as described below.In one embodiment, the input size of the final classifier is 2048 (ResNet50 embedding size) + 256 (HTML encoder encoding size) + 1024 (size of feature hash over URL n-grams). New phishing websites are submitted hourly by the security team in one production system. In one exemplary phishing website, the HTML script starts, then a blank section is detected, and then the HTML script ends. The disclosed technology supports timely detection of new phishing websites.

[0049] The phishing classifier layer 475 of the disclosed phishing detection engine 402 is trained on URL feature hashes, HTML encodings of HTML tokens extracted from content pages, and captured image embeddings from the content pages of example URLs, with each example URL accompanied by a ground truth 472 classification as phishing or not. After training, the phishing classifier layer 275 processes the URL feature hashes 442, HTML encodings 465, and image embeddings 457 to generate at least one likelihood score 485 that the URL and content accessed via this URL 414 present a phishing risk. The likelihood score 485 represents how likely the URL and content accessed via this URL present a phishing risk. The following lists example pseudocode for training a model for classification loss, clf loss, a binary phishing / not phishing, and Gen_loss, the difference between what the classifier expects to see and what it sees. def training_step(self, batch, batch_idx): html_tokens, snapshot, label, resnet_embed, domain_features = batch # Tuning and classification if self.classify: Embedding, logits=self.gpt(x,classify=True) gen_loss = self.criterion(logits, y) clf_logits = self.concat_layer(torch.cat([embedding, resnet_embed, domain_features], dim=1)) clf_loss = self.clf_criterion(clf_logits, label) #Joint Loss for Classification loss = clf_loss + gen_loss #Generative pre-training else: generated_img = self.gpt(html_tokens) loss = self.criterion(generated_img, snapshot)

[0050] FIG. 5 illustrates a block diagram of a reference residual neural network (ResNet) that is pre-trained for image classification before use in the phishing detection engine 402.

[0051] In inline phishing, web pages are rendered user-side at the user endpoint 166, so a snapshot of the page is not available, so ResNet is not utilized and the phishing detection classifier does not have access to header information of the content page. Next, we describe a disclosed classifier system that utilizes URLs and HTML tokens extracted from the page to classify a URL and the content page accessed via that URL as phishing or not. This third system is particularly useful in production environments where snapshots of visited content and access to header information are not available, and can operate in real time in network security systems.

[0052] Another disclosed classifier system applies ML / DL to a snapshot of the visited content and when access to header information is not available to classify a URL and the content accessed via the URL as phishing or not. Figure 6 illustrates a high-level block diagram 600 of a disclosed phishing detection engine 602 that utilizes ML / DL with a URL embedder and HTML encoder. HTML encoding is extracted from the content page pointed to by the URL. The disclosed phishing classifier layer 675 generates at least one likelihood score 685 that the URL and the content accessed via the URL present a phishing risk.

[0053] The phishing detection engine 602 uses a URL link sequence extractor 622 that extracts characters within a predetermined character set from a URL 614 to generate a URL character sequence 642. A one-dimensional 1D convolutional neural network (Conv1D) URL embedder 652 generates a URL embedding 653. Prior to use for classification, the URL embedder 652 and URL classifier 654 are trained using example URLs with ground truth 632 that classify URLs as phishing or non-phishing. The dashed block outline of the trained URL classifier 654 distinguishes training from subsequent processing of active URLs. During training of the URL embedder 652, differences are backpropagated past the phishing classifier layer to the embedding layer used to generate the URL embedding.

[0054] Continuing with the description of the system 600 illustrated in FIG. 6 , the phishing detection engine 602 also utilizes the disclosed HTML encoder 664, which is trained using HTML tokens 646 parsed by the HTML parser 636 from the content page at the exemplary URL 616, encoded, and then decoded to recreate the image captured from the rendering of the content page. Parsing extracts meaning from available metadata. In one implementation, tokenization acts as the first step of parsing to identify HTML tokens within the stream of metadata, and parsing then proceeds to determine the meaning and / or type of information being referenced using the context in which the tokens are found. The HTML encoder 664 generates an HTML encoding 665 of the HTML tokens 646 extracted from the content page 616.

[0055] During training, the headless browser 628 captures an image of the URL's content page for use in pre-training. A dashed line from the captured image 648 to the dashed block outline of the generative training decoder 668, which generates a rendered image of the page from the encoder embedding, distinguishes training from subsequent processing of the active URL. For training, the HTML encoder 664 is initialized with random initial parameters and parameters for HTML. The HTML encoder 664 is pre-trained using generative pre-training, where unsupervised pre-training over a large amount of unsupervised data is utilized to learn a data distribution P(X) for subsequent supervised decision-making using P(Y|X). Once the HTML encoder 664 is trained, it is reused for production use. During encoder 664 training, the difference between the encoding layers used to generate the HTML encoding is backpropagated across the phishing classifier layers. The training data, in one embodiment, includes 206,224 benign pages and 69,808 phishing pages.

[0056] Continuing with the description of system 600, phishing classifier layer 675 is trained on URL embeddings and HTML encodings of example URLs, each example URL accompanied by a ground truth classification 632 as phishing or non-phishing. During training of URL embedder 652, the encoding layer differences used to generate URL embeddings 653 are backpropagated beyond the phishing classifier layer. That is, once HTML encoder 664 is pre-trained, the URL embedding 653 network, along with the rest of the network (classification layer 675 and the fine-tuning step of HTML encoder 654), is trained using a loss function and with the help of example URLs with ground truth 632 for the input information.

[0057] After training, phishing classifier layer 675 processes the concatenated input of URL embeddings 653 and HTML encodings 665 to generate at least one likelihood score that the URL and content accessed via the URL present a phishing risk. Phishing detection engine 602 applies phishing classifier layer 675 to the concatenated input of URL embeddings and HTML encodings to generate at least one likelihood score 685 that the URL and content accessed via the URL present a phishing risk.

[0058] The HTML parser 636 extracts HTML tokens from content pages accessed via URLs. In one example, the HTML parser 636 can be configured to extract HTML tokens from the content that belong to a predetermined token vocabulary and ignore parts of the content that do not belong to the predetermined token vocabulary, such as line breaks and line feeds. In one embodiment, in training to determine the number of HTML tokens to specify for training the HTML encoder 664, a scan of 700K HTML files in the data store 164 and the resulting extraction of the top 10K tokens representing content pages was used to configure the phishing detection engine 602 for classification that does not suffer from a significant number of false positives. In one example content page, 800 valid tokens were extracted, in another example, 2K valid tokens were recognized, and in a third example, approximately 1K tokens were collected. Training is used to learn which ordering patterns of HTML tokens result in specific content pages.

[0059] For real-time phishing detection systems utilizing inline implementations, care is taken to minimize the size of the vocabulary for speed considerations. For the phishing detection engine 602, in one embodiment, the HTML parser 636 is configured to extract for generation of an HTML encoding of up to 64 HTML tokens. Different implementations may utilize different numbers of HTML encodings. In other embodiments, the headless browser may be configured to extract for generation of an HTML encoding of up to 128, 256, 1024, or 4096 HTML tokens.

[0060] Phishing patterns are constantly evolving, and it is often difficult for detection methods to achieve a high true positive rate (TPR) while maintaining a low false positive rate (FPR). A precision-recall curve shows the relationship between precision (= positive predictive value) and recall (= sensitivity) for possible cutoffs.

[0061] FIG. 7 shows a precision-recall graph of several of the disclosed phishing detection systems. The accuracy of the results near the top of the graph, where precision is close to 1.000, is interesting due to the requirement for few false positive detections. The curve with HTML+URL+header content is represented by the dotted line 746. Snapshot(ResNet)+BERT, represented as the curve with long dashes 736, is more accurate than HTML+URL+header content, and Snapshot, represented as the solid curve 726, is the most accurate phishing technique. BERT is computationally expensive, and the graph shows that BERT is not required to obtain accurate phishing detection results.

[0062] FIG. 8 illustrates a receiver operating characteristic curve (ROC) for the phishing website detection described above. The ROC curve is a plot of the true positive rate (TPR) as a function of the false positive rate (FPR) at various threshold settings. The region of interest is the area under the curve with a very low FPR, since false positive identification of a content page as a phishing website is not sustainable. The ROC curve is useful for comparing systems for detecting phishing websites described above. For the phishing detection engine 202, the ROC curve illustrated, labeled +Snapshot+Bert836 and with long dashes, shows the results of a system that utilizes ML / DL with URL feature hashes, encoding of NL words, and embedding of captured website images. In a second system, the phishing detection engine 402 utilizes ML / DL with URL feature hashes, encoding of HTML tokens extracted from content pages, and embedding of images captured from content pages to detect phishing sites. The curve labeled +Snapshot 826 shows a higher system accuracy with fewer false positives than HTML-URL-Header 846, whose curve is illustrated with dots.

[0063] Continuing with the ROC curves in Figure 8 with different combinations of features, including multilingual Bert embeddings, the comparison illustrates how text embeddings such as BERT can sometimes impair the model's effectiveness due to the fact that the HTML encoder already takes text content into account. As a result, including more text encodings can lead to overfitting to the text of HTML pages. Snapshot 826 has higher accuracy and leads to the lowest number of FPs in production. Furthermore, a version of the model without image embeddings results in a system that can bypass the active scanner / headless browser functionality in environments where snapshots are not available, such as runtime environments. In one example, running a headless browser may be too expensive and unscalable for the vast number of URLs encountered in production environments. Furthermore, an attacker can avoid detection in such environments.

[0064] 9 illustrates a receiver operating characteristic curve (ROC) for phishing website detection of the phishing detection engine 602 utilizing ML / DL with a URL embedder and HTML encoder. The ROC curve 936 is a plot of the true positive rate (TPR) as a function of the false positive rate (FPR) at various threshold settings. The ROC curve 936 illustrates that the phishing detection engine 602 has a higher true positive rate than the phishing detection system whose ROC curve is shown in FIG. 8.

[0065] The URL embedder 652 and html encoder 664 of the phishing detection engine 602 comprise high-level programs with irregular memory access patterns or data-dependent flow control. The high-level programs are source code written in programming languages ​​such as C, C++, Java, Python, and Spatial. The high-level programs may implement the computational structures and algorithms of machine learning models such as AlexNet, VGG Net, GoogleNet, ResNet, ResNeXt, RCNN, YOLO, SqueezeNet, SegNet, GAN, BERT, ELMo, USE, Transformer, and Transformer-XL. In one example, the high-level programs may implement a convolutional neural network with several processing layers, such that each processing layer can include one or more nested loops. The high-level programs may perform irregular memory operations, involving accessing inputs and weights and performing matrix multiplications between the inputs and weights. A high-level program may include nested loops with high iteration counts and loop bodies that load and multiply input values ​​from a previous processing layer with the weight of the subsequent processing layer to generate the output of the subsequent processing layer. A high-level program may have loop-level parallelism in the outermost loop body, which may be exploited using coarse-grained pipelining. A high-level program may have instruction-level parallelism in the innermost loop body, which may be exploited using loop unrolling, single instruction, multiple data (SIMD) vectorization, and pipelining.

[0066] Figure 10 illustrates a computational dataflow graph of the functionality of a one-dimensional 1D convolutional neural network (Conv1D) URL embedder 652 that generates a URL embedding 653 using C++ code expressed in the Open Neural Network Exchange (ONNX) format. The URL input is the first 100 characters of the URL 1014, which, in one exemplary embodiment, uses one-hot encoding and results in a convolution block 1024 (sliding window-like, kernel size = 7) with weights of 256 x 56 x 7, generating a binary output of multidimensional features. The output 1064, with dimensions 1 x 32 x 8 (1 x 256), represents the final URL embedding 653 generated as input to the phishing classifier layer 675.

[0067] Figure 11 shows a diagram of the disclosed HTML encoder 664 blocks, which generate HTML encodings 665 as input to the phishing classifier layer 675. The HTML encoder architecture is pre-trained with the help of a convolutional decoder that reconstructs images seen in the training data. The decoder is typically a convolutional neural network (CNN). The training forces the HTML encoder to learn to represent HTML content in terms of their rendered images, thus skipping irrelevant parts of the HTML. This training is tailored to how phishing attacks are launched; as long as the rendered pages continue to mimic the appearance of legitimate pages, users will fall victim to the attack. Next, we provide an overview of the block functionality. Input Embedding 1112 takes the 64 HTML tokens extracted by the HTML Parser 636, as described above, and maps the HTML tokens to a vocabulary. Positional Encoding 1122 adds contextual information for the HTML token vectors. Multi-head Attention 1132 generates multiple self-attention vectors to identify which elements of the input to focus on. Abstract Vectors Q, K, and V are used to extract different components of the input and calculate attention vectors. The multiple attention vectors represent the relationships between HTML vectors. Multi-head attention 1132 is repeated four times in the example computational data flow graph described below with respect to Figures 12A-12D, representing second-order multi-head attention. In different embodiments, the number of heads can be doubled or even larger. Multi-head attention 1132 passes the attention vectors, one vector at a time, to feedforward network 1162, which transforms the vectors for the next block. Each block ends with an addition and normalization operation, indicated by addition and normalization 1172, to smooth and normalize the layers across each feature, compressing the HTML representation to 256 numbers that recreate the image seen in the training data.The output represents the final HTML encoding 665 produced as an input of 256 numbers to a phishing classifier layer 675 .

[0068] FIG. 12A shows a schematic block diagram of the disclosed html encoder 664, which results in an html encoding 665 that is input to the phishing classifier layer 675. The input encoding and positional embedding 1205 is described in connection with FIG. 11 above, and a detailed ONNX listing is illustrated in FIG. 12B. The multi-head attention 1225 is described in connection with FIG. 11 above, and FIG. 12C shows a detailed ONNX image with inputs from the input encoding and positional embedding shown as inputs to the operator schema for implementing the blocks from FIG. 12B. The summation and normalization and feedforward 1245 is also described in connection with FIG. 11, with a detailed ONNX listing in FIG. 12D. FIG. 12A also includes a Reduce Mean 1265 operator that reduces the dimensionality of the input tensor and calculates the mean value of the elements of the input tensor along a provided axis. The HTML encoding 665 output is a 1×256 vector 1285 that is generated and mapped as input to the phishing classifier layer 675. The details of the inputs, outputs, and operations performed for ONNX operators are well known to those skilled in the art.

[0069] Figures 12B, 12C, and 12D together illustrate a computational data flow graph of the functionality of html encoder 664, using C++ code expressed in Open Neural Network Exchange (ONNX) format, resulting in html encoding 665 that is input to phishing classifier layer 675. In alternative embodiments, the ONNX code can express a different programming language.

[0070] Figure 12B shows a section of the dataflow graph in two columns separated by a dotted line, with connectors at the bottom of the left column flowing into the top of the right column. The results at the bottom of the right column of Figure 12B flow into Figures 12C and 12D. Figure 12B illustrates input encoding and positional embedding, as indicated by the gather operator 1264. For input embedding 1112, the gather block gathers data 1264 of dimensions 64 x 256.

[0071] FIG. 12C illustrates a single iteration of multi-head attention 1225, showing an example dataflow graph with computational nodes asynchronously transmitting data along data connections. The dataflow graph represents the so-called multi-head attention module of the Transformer model. In one embodiment, the dataflow graph shows multiple loops running in parallel as separate processing pipelines to process input tensors across multiple processing pipelines, with loop nesting arranged in a hierarchy of levels, such that second-level loops are within first-level loops, and with gather, unpack, and concatenate operations. The gather operations (three in all) refer to the use of query, key, and value vectors in the multi-head attention layer. In this example disclosed model, two heads were utilized, resulting in a concatenation operation for each of these vectors. In the illustrated embodiment, the respective outputs of each of the processing pipelines are concatenated to produce concatenated outputs A2, B2, C2, and D2. As illustrated by A2, B2, C2, D2 at the bottom of FIG. 12C and at the top of FIG. 12D, the outputs from the multi-head attention functionality flow into summation and normalization and feedforward.

[0072] Figure 12D illustrates the summation and normalization and feedforward 1245 functionality using ONNX operations. Figure 12D is illustrated using three columns separated by dotted lines, with the connector at the bottom of the left column flowing into the top of the middle column, and the connector at the bottom of the middle column flowing into the operation in the right column. The output of the multi-head attention (shown in Figure 12C) flows into the summation and normalization and feedforward operation, and the softmax operation 1232 transforms and normalizes the input vector into a probability distribution that feeds into the matrix multiplier MatMul 1242. The output of the summation and normalization and feedforward 1245, Ax, Bx, Cx, Dx (shown in the lower right corner of Figure 12D), flows into the reduced mean 1265 operator (Figure 12A). The reduced mean 1265 operator reduces the dimensionality of the input tensors (Ax, Bx, Cx, Dx) and calculates the mean of the elements of the input tensors along the provided axes. The output of HTML encoding 665 is a 1x256 vector 1285.

[0073] Figure 13 illustrates a computational dataflow graph of the functionality of the phishing classifier layer 675, which generates likelihood scores 685 representing how likely a particular website is to be a phishing website. The C++ code is expressed in the Open Neural Network Exchange (ONNX) format. The input 1314 to the phishing classifier layer 675, which is 1x512 in size, is formed by concatenating a URL embedding and an HTML encoding. The two concatenated vectors are a 1x256 vector HTML encoding and a URL embedding with dimensions 1x32x8 (1x256), as previously described. Batch normalization 1324, as applied to the activations of the previous layer, standardizes the input and accelerates training. Operators GEneral Matrix Multiplication (GEMM) 1346, 1366 represent linear algebra routines that are fundamental operators in DL. The output 1374 of the final two-layer feedforward classifier, of size 1x2, is the likelihood that a web page is a phishing site and the likelihood that a web page is not a phishing site, for classifying the site as phishing or not phishing.

[0074] Computer Systems 14 is a simplified block diagram of a computer system 1000 that can be used to classify URLs and content pages accessed via the URLs as phishing or non-phishing. The computer system 1400 includes at least one central processing unit (CPU) 1472 that communicates with certain peripheral devices via a bus subsystem 1455 and a network security system 112 for providing the network security services described herein. These peripheral devices may include, for example, a storage subsystem 1410 including a memory device and a file storage subsystem 1436, a user interface input device 1438, a user interface output device 1476, and a network interface subsystem 1474. The input and output devices enable user interaction with the computer system 1400. The network interface subsystem 1474 provides an interface to external networks, including interfaces to corresponding interface devices in other computer systems.

[0075] In one implementation, the cloud-based security system 153 of FIG. 1 is communicatively linked to the storage subsystem 1410 and the user interface input device 1438.

[0076] User interface input devices 1438 can include pointing devices such as keyboards, mice, trackballs, touchpads, or graphics tablets, scanners, touchscreens integrated into displays, audio input devices such as voice recognition systems and microphones, and other types of input devices. In general, use of the term "input device" is intended to encompass all possible types of devices and methods for inputting information into computer system 1400.

[0077] The user interface output devices 1476 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as an LED display, a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem may also provide a non-visual display, such as an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computer system 1400 to a user or to another machine or computer system.

[0078] Storage subsystem 1410 stores programming and data structures that provide the functionality of some or all of the modules and methods described herein. Subsystem 1478 can be a graphics processing unit (GPU) or a field programmable gate array (FPGA).

[0079] The memory subsystem 1422 used within the storage subsystem 1410 may include some memory, including a main random access memory (RAM) 1432 for storing instructions and data during program execution, and a read-only memory (ROM) 1434 in which fixed instructions are stored. The file storage subsystem 1436 may provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of a particular implementation may be stored by the file storage subsystem 1436, within the storage subsystem 1410, or within another machine accessible by the processor.

[0080] Bus subsystem 1455 provides a mechanism for allowing the various components and subsystems of computer system 1400 to communicate with each other as intended. Although bus subsystem 1455 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0081] The computer system 1400 itself can be of a variety of types, including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a widely distributed set of loosely networked computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, the description of the computer system 1400 shown in Figure 14 is intended only as a specific example to illustrate a preferred embodiment of the present invention. Many other configurations of computer system 1400 are possible, having more or fewer components than the computer system shown in Figure 14.

[0082] Specific Implementations Some specific implementations and features for classifying a URL and the content accessed via that URL as phishing or non-phishing are described in the following discussion.

[0083] In one disclosed implementation, a phishing classifier that classifies a URL and content accessed via the URL as phishing or non-phishing includes a URL feature hasher that parses the URL into features and hashes the features to generate a URL feature hash, a headless browser configured to extract words from a rendered content page and capture an image of at least a portion of the rendered content page, a natural language encoder that generates word encodings of the extracted words, an image embedder that generates image embeddings of the captured image, and a phishing classifier layer that processes the concatenated input of the URL feature hash, word encodings, and image embeddings to generate at least one likelihood score that the URL and content accessed via the URL present a phishing risk.

[0084] In another disclosed implementation, a phishing classifier that classifies a URL and content accessed via the URL as phishing or not includes a URL feature hasher that parses the URL into features and hashes the features to generate a URL feature hash, and a headless browser configured to access the content of the URL and internally render a content page, extract words from the rendering of the content page, and capture images of at least a portion of the rendering of the content page. The disclosed implementation also includes a natural language encoder pre-trained on natural language that generates word encodings of the words extracted from the content page, and an image embedder pre-trained on images that generates image embeddings of images captured from the content page. The implementation further includes a phishing classifier layer trained on the URL feature hashes, word encodings, and image embeddings of the example URL, with each example URL having a ground truth classification as phishing or not phishing, that processes the concatenated input of the URL feature hashes, word encodings, and image embeddings of the URL to generate at least one likelihood score that the URL and content accessed via the URL present a phishing risk.

[0085] In some disclosed implementations of the phishing classifier, the natural language encoder is one of Bidirectional Encoder Representations from Transformer (abbreviated as BERT) and Universal Sentence Encoder, and the image embedder is one of Residual Neural Network (abbreviated as ResNet), Inception-v3, and VGG-16.

[0086] In one implementation, a disclosed computer-implemented method for classifying a URL and content accessed via the URL as phishing or non-phishing includes applying a URL feature hasher to extract features from the URL and hash the features to generate URL feature hashes. The disclosed method also includes applying a pre-trained natural language encoder for natural language to generate word encodings of words parsed from a rendering of the content, and applying a pre-trained image encoder for images to generate image embeddings of images captured from at least a portion of the rendering. The disclosed method further includes applying a phishing classifier layer trained on a concatenation of the URL feature hashes, word encodings, and image embeddings of example URLs with ground truth classifications as phishing or non-phishing, and processing the URL feature hashes, word encodings, and image embeddings to generate at least one likelihood score that the URL and content accessed via the URL present a phishing risk.

[0087] The methods described in this and other sections of the disclosed technology may include one or more of the following features and / or features described in connection with the additional methods disclosed. For brevity, combinations of features disclosed in this application are not individually listed or repeated with each basic set of features. The reader will understand how features identified in this method can be readily combined with the set of basic features identified as implementations.

[0088] One disclosed computer-implemented method further includes applying a headless browser to access content via a URL and internally render the content, parsing words from the rendered content, and capturing an image of at least a portion of the rendered content.

[0089] One embodiment of the disclosed computer-implemented method includes a natural language encoder as one of Bidirectional Encoder Representations from Transformer (BERT) and a Universal Sentence Encoder. Some embodiments of the disclosed computer-implemented method also include an image embedder as one of a Residual Neural Network (ResNet), Inception-v3, and VGG-16.

[0090] One disclosed computer-implemented method for training a phishing classifier layer to classify URLs and content accessed via the URLs as phishing or non-phishing includes receiving and processing URL feature hashes, word encodings of words extracted from content pages, and image embeddings of images captured from renderings of the content for example URLs to generate at least one likelihood score that each example URL and content accessed via the URL presents a phishing risk. The method also includes calculating a difference between the likelihood score for each example URL and a corresponding ground truth that the example URL and content page are phishing or non-phishing, and using the difference for the example URLs to train coefficients of the phishing classifier layer. The method further includes storing the trained coefficients for use in classifying production URLs and content pages accessed via the production URLs as phishing or non-phishing.

[0091] The disclosed computer-implemented method further includes not backpropagating the differences beyond the phishing classifier layer to an encoding layer used to generate word encodings, and not backpropagating the differences beyond the phishing classifier layer to an embedding layer used to generate image embeddings.

[0092] The disclosed computer-implemented method also includes generating a URL feature hash for each example URL, generating word encodings of words extracted from the rendering of the content page, and generating image embeddings of images captured from the rendering.

[0093] For many disclosed computer implementations, the disclosed method for training a phishing classifier layer to classify URLs and content accessed via the URLs as phishing or non-phishing includes generating word encodings using a Bidirectional Representation of Text from Transformers (BERT) encoder or a variant of the BERT encoder, and generating image embeddings using one of a residual neural network (ResNet), Inception-v3, and VGG-16.

[0094] In one disclosed implementation, a phishing classifier that classifies a URL and content accessed via the URL as phishing or non-phishing includes a URL feature hasher that parses the URL into features and hashes the features to generate a URL feature hash, and a headless browser configured to extract HTML tokens from a rendered content page and capture an image of at least a portion of the rendered content page. The disclosed classifier also includes an HTML encoder that generates an HTML encoding of the extracted HTML tokens, an image embedder that generates an image embedding of the captured image, and a phishing classifier layer that processes the URL feature hash, the HTML encoding, and the image embedding to generate at least one likelihood score that the URL and the content accessed via the URL present a phishing risk. In some implementations, the HTML tokens belong to a recognized vocabulary of HTML tokens.

[0095] In one implementation, a disclosed phishing classifier that classifies a URL and content accessed via the URL as phishing or non-phishing includes a URL feature hasher that parses the URL into features and hashes the features to generate a URL feature hash, and a headless browser configured to access the content of the URL and internally render a content page, extract HTML tokens from the content page, and capture an image of at least a portion of the rendering of the content page. The disclosed phishing classifier also includes an HTML encoder that is trained on HTML tokens extracted from the content page of an example URL, encoded, and then decoded to recreate the image captured from the rendering of the content page, and generates an HTML encoding of the HTML tokens extracted from the content page. Also included is an image embedder that is pre-trained on images and generates image embeddings for images captured from content pages; and a phishing classifier layer that is trained on URL feature hashes, HTML encodings, and image embeddings of example URLs, and that processes the URL feature hashes, HTML encodings, and image embeddings of example URLs, each example URL with a ground truth classification as phishing or not phishing, to generate at least one likelihood score that the URL and the content page accessed via the URL presents a phishing risk.

[0096] Some implementations of the disclosed method further include a headless browser configured to extract HTML tokens from the content that belong to a predetermined token vocabulary and ignore portions of the content that do not belong to the predetermined token vocabulary. Some disclosed implementations further include a headless browser configured to extract up to 64 HTML tokens for generation of an HTML encoding.

[0097] One implementation of the disclosed method for classifying a URL and a content page accessed via the URL as phishing or non-phishing includes applying a URL feature hasher to extract features from the URL and hash the features to generate URL feature hashes. The method also includes applying an HTML encoder trained on natural language and generating HTML encodings of HTML tokens extracted from the rendered content page. The method further includes applying an image embedder pre-trained on images and generating image embeddings of images captured from at least a portion of the rendered content page, and a phishing classifier layer trained on the URL feature hashes, HTML encodings, and image embeddings for example URLs classified with ground truth classifications as phishing or non-phishing, and processing the URL feature hashes, HTML encodings, and image embeddings of the URL to generate at least one likelihood score that the URL and content accessed via the URL presents a phishing risk. In one disclosed implementation, the method further includes applying a headless browser, accessing a content page via a URL and internally rendering the content page, parsing HTML tokens from the rendered content, and capturing an image of at least a portion of the rendered content.

[0098] Some disclosed implementations further include the headless browser parsing HTML tokens from the content that belong to a predetermined token vocabulary and ignoring portions of the content that do not belong to the predetermined token vocabulary. Some implementations also include the headless browser parsing up to 64 HTML tokens for generating the HTML encoding.

[0099] One implementation of a disclosed computer-implemented method for training a phishing classifier layer to classify URLs and content accessed via the URLs as phishing or non-phishing includes receiving and processing URL feature hashes, HTML encodings of HTML tokens extracted from content pages, and image embeddings of images captured from renderings of the content pages for example URLs to generate at least one likelihood score that each example URL and content accessed via the URL presents a phishing risk. The method includes calculating a difference between the likelihood score for each example URL and a corresponding ground truth about whether the example URL and content page are phishing or non-phishing, using the calculated difference for the example URLs to train coefficients of the phishing classifier layer, and storing the trained coefficients for use in classifying production URLs and content pages accessed via the production URLs as phishing or non-phishing.

[0100] Some implementations of the disclosed method include backpropagating the difference beyond the phishing classifier layer to an encoding layer used to generate the HTML encoding. Some implementations further include not backpropagating the difference beyond the phishing classifier layer to an embedding layer used to generate the image embedding.

[0101] Some implementations also include generating URL feature hashes for each of the example URLs, generating an HTML encoding of the HTML tokens extracted from the rendering of the content page, and generating an image embedding of the image captured from the rendering. Some implementations further include training an HTML encoder-decoder for a second example URL to generate an HTML encoding using the HTML tokens extracted from the content page of the second example URL, encoded, and then decoded to recreate the image captured from the content page of the second example URL. Some implementations of the disclosed method also include generating the image embedding using a ResNet embedder or a variant of a ResNet embedder pre-trained to embed images in an embedding space.

[0102] One implementation of the disclosed phishing classifier that classifies a URL and the content page accessed via the URL as phishing or not includes an input processor that accepts a URL for classification, a URL embedder that generates a URL embedding for the URL, an HTML parser that extracts HTML tokens from the content page accessed via the URL, an HTML encoder that generates an HTML encoding from the HTML tokens, and a phishing classifier layer that operates on the URL embedding and HTML encoding to classify the URL and the content accessed via the URL as phishing or not.

[0103] Some implementations of the disclosed phishing classifier also include a URL embedder that extracts characters in a predetermined character set from a URL to generate a string and is trained using the ground truth classification of the URL as phishing or not, and generates a URL embedding. The classifier further includes an HTML parser configured to access the content of the URL and extract HTML tokens from the content page. Also included are a disclosed HTML encoder that is trained on the HTML tokens extracted from the content page of each example URL along with a ground truth image captured from the content page accessed via the example URL, and generates an HTML encoding of the HTML tokens extracted from the content page; and a phishing classifier layer that is trained on the URL embedding and HTML encoding of each example URL along with the ground truth classification of the URL as phishing or not, and processes the concatenated input of the URL embedding and HTML encoding to generate at least one likelihood score that the URL and the content accessed via the URL present a phishing risk.

[0104] For some implementations of the disclosed phishing classifier, the input processor accepts URLs for classification in real time. In many implementations of the disclosed phishing classifier, the phishing classifier layer operates to classify URLs and content accessed via the URLs as phishing or non-phishing in real time. In some implementations, the disclosed phishing classifier can further include an HTML parser configured to extract HTML tokens from a content page that belong to a predetermined token vocabulary and ignore portions of the content page that do not belong to the predetermined token vocabulary, and to extract for generation of an HTML encoding of up to 64 HTML tokens.

[0105] One implementation of the disclosed computer-implemented method for classifying a URL and content accessed via the URL as phishing or non-phishing includes applying a URL embedder that extracts characters in a predetermined character set from the URL to generate a string to generate a URL embedding, and that generates the URL embedding trained and using ground truth classification of the URL as phishing or non-phishing. The method also includes applying an HTML parser to access the content of the URL, extracting HTML tokens from the content page, applying an HTML encoder to generate an HTML encoding of the extracted HTML tokens, and applying a phishing classifier layer to the concatenated input of the URL embedding and the HTML encoding to generate at least one likelihood score that the URL and content accessed via the URL present a phishing risk. Some implementations also include an HTML parser that extracts HTML tokens belonging to a predetermined token vocabulary from the content page and ignores portions of the content page that do not belong to the predetermined token vocabulary, and may further include an HTML parser that extracts up to 64 HTML tokens for generating an HTML encoding. The disclosed methods can also include applying, in real time, a URL embedder, an HTML parser, an HTML encoder, and a phishing classifier layer, in some cases, the phishing classifier layer operates to generate at least one likelihood score that a URL and content accessed via the URL presents a phishing risk in real time.

[0106] One implementation of the disclosed computer-implemented method for training a phishing classifier layer to classify URLs and content pages accessed via the URLs as phishing or non-phishing includes receiving and processing URL embeddings of characters extracted from the URLs and HTML encodings of HTML tokens extracted from the content pages for example URLs and the content pages accessed via the URLs to generate at least one likelihood score that each example URL and the content page accessed via the URL presents a phishing risk. The disclosed method also includes calculating a difference between the likelihood score for each example URL and a corresponding ground truth that the example URL and content page are phishing or non-phishing, using the difference for the example URLs to train coefficients of the phishing classifier layer, and storing the trained coefficients for use in classifying production URLs and content pages accessed via the production URLs as phishing or non-phishing.

[0107] Some implementations of the disclosed method further include applying a headless browser, accessing the content of the URL to internally render the content page, and capturing an image of at least a portion of the content page. The disclosed method can further include backpropagating the difference beyond the phishing classifier layer to an encoding layer used to generate HTML encodings, and can further include backpropagating the difference beyond the phishing classifier layer to an embedding layer used to generate URL embeddings.

[0108] Some implementations of the disclosed method further include generating a URL embedding of characters extracted from the exemplary URL, generating an HTML encoding of HTML tokens extracted from the content page accessed via the exemplary URL, and generating an image embedding of an image captured from the rendering.

[0109] Some implementations of the disclosed method further include training an HTML encoder-decoder for a second exemplary URL to generate an HTML encoding using HTML tokens extracted from the content page of the second exemplary URL, encoded, and then decoded to recreate an image captured from the content page of the second exemplary URL. The disclosed method may further include extracting HTML tokens from the content page that belong to a predetermined token vocabulary and ignoring portions of the content page that do not belong to the predetermined token vocabulary. The method may also include limiting the extraction to a predetermined number of HTML tokens and may further include generating an HTML encoding of up to 64 HTML tokens.

[0110] Other implementations of the methods described in this section may include a tangible, non-transitory computer-readable storage medium characterized by computer program instructions that, when executed on a processor, cause the processor to perform any of the methods described above. Yet another implementation of the methods described in this section may include a device including a memory and one or more processors operable to execute computer instructions stored in the memory to perform any of the methods described above.

[0111] Any data structures and code described or referenced above are, according to many implementations, stored on a computer-readable storage medium, which can be any device or medium capable of storing code and / or data for use by a computer system, including, but not limited to, volatile memory, non-volatile memory, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), magnetic and optical storage devices such as disk drives, magnetic tape, compact discs (CDs), digital versatile discs or digital video discs (DVDs), or other media capable of storing computer-readable media now known or later developed.

[0112] The foregoing description is presented to enable making and using the disclosed technology. Various modifications to the disclosed implementations will become apparent, and the general principles defined herein may be applied to other implementations and applications without departing from the spirit and scope of the disclosed technology. Thus, the disclosed technology is not intended to be limited to the implementations shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein. The scope of the disclosed technology is defined by the appended claims.

[0113] Terms The following clauses are disclosed:

[0114] Clause Set 1 1. A phishing classifier that classifies URLs and content pages accessed via the URLs as phishing or non-phishing, A URL feature hasher that parses a URL into features, hashes the features, and generates a URL feature hash. Access the content page of the URL, render the content page internally, Extract words from content page renderings, a headless browser configured to capture an image of at least a portion of a rendering of a content page; a natural language encoder that is pre-trained on a natural language and generates word encodings of words extracted from content pages; an image embedder that is pre-trained on images and generates image embeddings for images captured from content pages; Trained on URL feature hashes, word encodings, and image embeddings of example URLs, with each example URL accompanied by a ground truth classification as phishing or non-phishing; a phishing classifier layer that processes a concatenated input of URL feature hashes, word encodings, and image embeddings of the URL to generate at least one likelihood score that the URL and content accessed via the URL present a phishing risk. 2. The phishing classifier of clause 1, wherein the natural language encoder is one of: a Bidirectional Encoder Representations from Transformer (abbreviated as BERT) and a Universal Sentence Encoder. 3. The phishing classifier of clause 1, wherein the image embedder is one of a residual neural network (ResNet for short), Inception-v3, and VGG-16. 4. A computer-implemented method for classifying URLs and content pages accessed via the URLs as phishing or non-phishing, comprising: applying a URL feature hasher to extract features from the URL and hashing the features to generate a URL feature hash; applying a natural language encoder that is pre-trained on a natural language and that generates word encodings for words parsed from a rendering of a content page; applying an image encoder that is pre-trained on the image and that generates an image embedding of the image captured from at least a portion of the rendering; Applying a phishing classifier layer trained on a concatenation of URL feature hashes, word encodings, and image embeddings for example URLs with ground truth classifications as phishing or not phishing; and processing the URL feature hashes, word encodings, and image embeddings to generate at least one likelihood score that the URL and content accessed via the URL present a phishing risk. 5. Applying headless browsers and accessing the content page via a URL and internally rendering the content page; Parsing words from the rendered content page; and 5. The computer-implemented method of claim 4, further comprising: capturing an image of at least a portion of the rendered content page. 6. The computer-implemented method of clause 4, wherein the natural language encoder is one of: Bidirectional Encoder Representations from Transformer (abbreviated as BERT) and a Universal Sentence Encoder. 7. The computer-implemented method of clause 4, wherein the image embedder is one of a residual neural network (abbreviated as ResNet), Inception-v3, and VGG-16. 8. A non-transitory computer-readable storage medium characterized by computer program instructions for classifying URLs and content pages accessed via the URLs as phishing or non-phishing, the instructions, when executed on a processor, applying a URL feature hasher to extract features from the URL and hashing the features to generate a URL feature hash; applying a natural language encoder that is pre-trained on a natural language and that generates word encodings for words parsed from a rendering of a content page; applying an image encoder that is pre-trained on the image and that generates an image embedding of the image captured from at least a portion of the rendering; Applying a phishing classifier layer trained on a concatenation of URL feature hashes, word encodings, and image embeddings for example URLs with ground truth classifications as phishing or not phishing; 1. A non-transitory computer-readable storage medium implementing a method that includes processing URL feature hashes, word encodings, and image embeddings to generate at least one likelihood score that the URL and content accessed via the URL present a phishing risk. 9. Applying headless browsers and accessing the content page via a URL and internally rendering the content page; Parsing words from the rendered content page; and and capturing an image of at least a portion of the rendered content page. 10. The non-transitory computer-readable storage medium of clause 8, wherein the natural language encoder is one of: Bidirectional Encoder Representations from Transformer (abbreviated as BERT) and a Universal Sentence Encoder. 11. The non-transitory computer-readable storage medium of clause 8, wherein the image embedder is one of a residual neural network (abbreviated as ResNet), Inception-v3, and VGG-16. 12. A computer-implemented method for training a phishing classifier layer to classify URLs and content pages accessed via the URLs as phishing or non-phishing, comprising: For example URLs, receiving and processing URL feature hashes, word encodings of words extracted from the content page, and image embeddings of images captured from renderings of the content page; generating at least one likelihood score that each example URL and content page accessed via the URL presents a phishing risk; Calculating the difference between the likelihood score for each example URL and each corresponding ground truth that the example URL and content page is phishing or non-phishing; training coefficients of a phishing classifier layer using the differences for the example URLs; and storing the trained coefficients for use in classifying production URLs and content pages accessed via the production URLs as phishing or non-phishing. 13. The computer-implemented method of clause 12, further comprising not backpropagating the differences beyond the phishing classifier layer to the encoding layer used to generate the word encodings. 14. The computer-implemented method of clause 12, further comprising not backpropagating the difference beyond the phishing classifier layer to the embedding layer used to generate the image embeddings. 15. The computer-implemented method of clause 12, further comprising: generating a URL feature hash for each example URL; generating word encodings of words extracted from the rendering of the content page; and generating image embeddings of images captured from the rendering. 16. The computer-implemented method of clause 15, further comprising generating the word encodings using a Bidirectional Encoder Representation from Transformer (abbreviated BERT) encoder or a variant of the BERT encoder. 17. The computer-implemented method of clause 15, further comprising generating the image embeddings using one of a residual neural network (abbreviated as ResNet), Inception-v3, and VGG-16.

[0115] Clause Set 2 1. A phishing classifier that classifies URLs and content pages accessed via the URLs as phishing or non-phishing, A URL feature hasher that parses a URL into features, hashes the features, and generates a URL feature hash. Access the content page of the URL, render the content page internally, Extract HTML tokens from content pages, a headless browser configured to capture an image of at least a portion of a rendering of a content page; an HTML encoder that is trained on HTML tokens extracted from the content page of an exemplary URL, encoded, and then decoded to recreate an image captured from a rendering of the content page, and that generates an HTML encoding of the HTML tokens extracted from the content page; an image embedder that is pre-trained on images and generates image embeddings for images captured from content pages; Trained on URL feature hashes, HTML encodings, and image embeddings of example URLs, with each example URL accompanied by a ground truth classification as phishing or non-phishing; a phishing classifier layer that processes URL feature hashes, HTML encoding, and image embeddings of the URL to generate at least one likelihood score that the URL and a content page accessed via the URL presents a phishing risk. 2. The phishing classifier of clause 1, further comprising a headless browser configured to extract HTML tokens from a content page that belong to a predetermined token vocabulary, and to ignore portions of the content page that do not belong to the predetermined token vocabulary. 3. The phishing classifier of clause 1, further comprising a headless browser configured to extract for generation an HTML encoding of up to 64 HTML tokens. 4. A computer-implemented method for classifying URLs and content pages accessed via the URLs as phishing or non-phishing, comprising: applying a URL feature hasher to extract features from the URL and hashing the features to generate a URL feature hash; applying an HTML encoder trained on a natural language to generate an HTML encoding of HTML tokens extracted from the rendered content page; an image embedder that is pre-trained on images and generates image embeddings for images captured from at least a portion of a rendered content page; Applying a phishing classifier layer trained on URL feature hashes, HTML encodings, and image embeddings to example URLs classified with ground truth classifications as phishing or not phishing; 1. A computer-implemented method comprising: processing URL feature hashes, HTML encoding, and image embedding of a URL to generate at least one likelihood score that the URL and a content page accessed via the URL present a phishing risk. 5. Applying headless browsers and accessing the content page via a URL and internally rendering the content page; Parsing HTML tokens from the rendered content; and 5. The computer-implemented method of claim 4, further comprising: capturing an image of at least a portion of the rendered content. 6. The computer-implemented method of clause 5, further comprising the headless browser parsing HTML tokens from the content that belong to a predetermined token vocabulary and ignoring portions of the content that do not belong to the predetermined token vocabulary. 7. The computer-implemented method of clause 5, further comprising the headless browser parsing up to 64 HTML tokens to generate the HTML encoding. 8. A non-transitory computer-readable storage medium characterized by computer program instructions for classifying URLs and content pages accessed via the URLs as phishing or non-phishing, the instructions, when executed on a processor, applying a URL feature hasher to extract features from the URL and hashing the features to generate a URL feature hash; applying an HTML encoder trained on a natural language to generate an HTML encoding of HTML tokens extracted from the rendered content page; applying an image embedder to generate an image embedding of the captured image from at least a portion of the rendered content; and applying a phishing classifier layer that processes URL feature hashes, HTML encoding, and image embeddings of the URL to generate at least one likelihood score that the URL and a content page accessed via the URL present a phishing risk. 9. When an instruction is executed on a processor, 9. The non-transitory computer-readable storage medium of clause 8, further comprising training a phishing classifier layer on URL feature hashes, HTML encodings, and image embeddings for example URLs classified with ground truth classifications as phishing or not phishing. 10. When an instruction is executed on a processor, 9. The non-transitory computer-readable storage medium of clause 8, further comprising: training an HTML encoder to generate an HTML encoding of the HTML tokens extracted from the content page for the exemplary URL, encoded, and then decoded to recreate an image captured from a rendering of the content page. 11. When an instruction is executed on a processor, 9. The non-transitory computer-readable storage medium of clause 8, implementing the method, wherein the image embedder is a ResNet embedder or a variant of a ResNet embedder that is pre-trained to embed images in the embedding space. 12. When an instruction is executed on a processor, Applying headless browsers, Accessing the content via a URL and rendering the content internally; Parsing HTML tokens from the rendered content; and 9. The non-transitory computer-readable storage medium of claim 8, implementing a method further comprising: capturing an image of at least a portion of the rendered content. 13. The non-transitory computer-readable storage medium of clause 12, implementing a method wherein the instructions, when executed on a processor, further include causing the headless browser to parse HTML tokens from the content that belong to a predetermined token vocabulary and ignore portions of the content that do not belong to the predetermined token vocabulary. 14. The non-transitory computer-readable storage medium of clause 12, wherein the instructions, when executed on a processor, implement a method further comprising the headless browser parsing up to 64 HTML tokens to generate an HTML encoding. 15. A computer-implemented method for training a phishing classifier layer to classify URLs and content pages accessed via the URLs as phishing or non-phishing, comprising: For example URLs, receiving and processing URL feature hashes, HTML encodings of HTML tokens extracted from the content page, and image embeddings of images captured from renderings of the content page; generating at least one likelihood score that each example URL and content accessed via the URL presents a phishing risk; Calculating the difference between each example URL and each corresponding ground truth regarding whether the example URL and content page is phishing or non-phishing; training the coefficients of a phishing classifier layer using the calculated differences for the example URLs; and and storing the trained coefficients for use in classifying production URLs and content pages accessed via the production URLs as phishing or non-phishing. 16. The computer-implemented method of clause 15, further comprising backpropagating the differences beyond the phishing classifier layer to an encoding layer used to generate the HTML encoding. 17. The computer-implemented method of clause 15, further comprising not backpropagating the difference beyond the phishing classifier layer to the embedding layer used to generate the image embedding. 18. The computer-implemented method of clause 15, further comprising: generating a URL feature hash for each example URL; generating an HTML encoding of HTML tokens extracted from the rendering of the content page; and generating image embeddings of images captured from the rendering. 19. The computer-implemented method of clause 18, further comprising: for a second exemplary URL, training an HTML encoder-decoder to generate an HTML encoding using HTML tokens extracted from the content page of the second exemplary URL, encoded, and then decoded to generate an image captured from the content page of the second exemplary URL. 20. The computer-implemented method of clause 19, further comprising generating image embeddings using a ResNet embedder or a variant of a ResNet embedder.

[0116] Clause Set 3 1. A phishing classifier that classifies URLs and content pages accessed via the URLs as phishing or non-phishing, an input processor that accepts a URL for classification; a URL embedder that generates a URL embedding of a URL; an HTML parser that extracts HTML tokens from content pages accessed via URLs; an HTML encoder that generates an HTML encoding from the HTML tokens; a phishing classification layer that operates on URL embedding and HTML encoding to classify URLs and content accessed via URLs as phishing or not phishing. 2. A URL embedder that extracts characters in a predetermined character set from a URL to generate a string and is trained using ground truth classification of URLs as phishing or non-phishing, and generates a URL embedding; Access the content of the URL, an HTML parser configured to extract HTML tokens from the content page; each example URL is trained on HTML tokens extracted from the content page of the example URL accompanied by a ground truth image captured from the content page accessed via the example URL; an HTML encoder that generates an HTML encoding of the HTML tokens extracted from the content page; trained on URL embeddings and HTML encodings of example URLs with a ground truth classification of each example URL as phishing or non-phishing; 10. The phishing classifier of claim 1, further comprising: a phishing classifier layer that processes the concatenated input of URL embedding and HTML encoding to generate at least one likelihood score that the URL and content accessed via the URL present a phishing risk. 3. The phishing classifier of clause 1, wherein the input processor accepts URLs for classification in real time. 4. The phishing classifier of clause 1, wherein the phishing classifier layer operates to classify URLs and content accessed via the URLs as phishing or not phishing in real time. 5. The phishing classifier of clause 1, further comprising an HTML parser configured to extract HTML tokens from the content page that belong to a predetermined token vocabulary, and to ignore portions of the content page that do not belong to the predetermined token vocabulary. 6. The phishing classifier of clause 1, further comprising an HTML parser configured to extract for generation an HTML encoding of up to 64 HTML tokens. 7. A computer-implemented method for classifying URLs and content pages accessed via the URLs as phishing or non-phishing, comprising: applying a URL embedder that extracts characters in a predetermined character set from a URL to generate a string to generate the URL embedding, and that trains and uses ground truth classification of URLs as phishing or non-phishing; applying an HTML parser to access the content of the URL and extract HTML tokens from the content page; applying an HTML encoder to generate an HTML encoding of the extracted HTML tokens; and applying a phishing classifier layer to the concatenated input of the URL embedding and HTML encoding to generate at least one likelihood score that the URL and content accessed via the URL present a phishing risk. 8. The computer-implemented method of clause 7, further comprising an HTML parser that extracts HTML tokens from the content page that belong to a predetermined token vocabulary and ignores portions of the content page that do not belong to the predetermined token vocabulary. 9. The computer-implemented method of clause 7, further comprising the HTML parser extracting up to 64 HTML tokens for generation of the HTML encoding. 10. The computer-implemented method of clause 7, further comprising applying, in real time, a URL embedder, an HTML parser, an HTML encoder, and a phishing classifier layer. 11. The computer-implemented method of clause 7, wherein the phishing classifier layer operates to generate at least one likelihood score that a URL and content accessed via the URL presents a phishing risk in real time. 12. A non-transitory computer-readable storage medium characterized by computer program instructions for classifying URLs and content pages accessed via the URLs as phishing or non-phishing, the instructions, when executed on a processor, applying a URL embedder to extract characters in a predetermined character set from the URL to generate a string, thereby generating a URL embedding; applying an HTML parser to extract HTML tokens from the content page; applying an HTML encoder to generate an HTML encoding of the extracted HTML tokens; 1. A non-transitory computer-readable storage medium implementing a method that includes applying a phishing classifier layer to a concatenated input of URL embedding and HTML encoding to generate at least one likelihood score that the URL and content accessed via the URL present a phishing risk. 13. A non-transitory computer-readable storage medium as recited in clause 12, implementing a method wherein the instructions, when executed on a processor, further include an HTML parser that extracts HTML tokens from a content page that belong to a predetermined token vocabulary and ignores portions of the content page that do not belong to the predetermined token vocabulary. 14. The non-transitory computer-readable storage medium of clause 12, implementing a method wherein the instructions, when executed on a processor, further include an HTML parser extracting for generation of an HTML encoding of up to 64 HTML tokens. 15. The non-transitory computer-readable storage medium of clause 12, wherein the instructions, when executed on a processor, implement a method further comprising training an HTML encoder on HTML tokens extracted from content pages of exemplary URLs, each exemplary URL accompanied by ground truth images captured from the content pages accessed via the exemplary URL. 16. The non-transitory computer-readable storage medium of clause 12, implementing a method wherein the instructions, when executed on a processor, further include training a phishing classifier layer on URL embeddings and HTML encodings of example URLs, each example URL with a ground truth classification as phishing or not phishing. 17. A computer-implemented method for training a phishing classifier layer to classify URLs and content pages accessed via the URLs as phishing or non-phishing, comprising: For an example URL and a content page accessed via the URL: receiving and processing URL embedding of characters extracted from the URL and HTML encoding of HTML tokens extracted from the content page; generating at least one likelihood score that each example URL and content page accessed via the URL presents a phishing risk; Calculating the difference between the likelihood score for each example URL and each corresponding ground truth that the example URL and content page is phishing or non-phishing; training coefficients of a phishing classifier layer using the differences for the example URLs; and storing the trained coefficients for use in classifying production URLs and content pages accessed via the production URLs as phishing or non-phishing. 18. Applying headless browsers; Accessing the content of the URL and internally rendering the content page; and 20. The computer-implemented method of claim 17, further comprising: capturing an image of at least a portion of the content page. 19. The computer-implemented method of clause 17, further comprising backpropagating the differences beyond the phishing classifier layer to an encoding layer used to generate the HTML encoding. 20. The computer-implemented method of clause 17, further comprising backpropagating the difference beyond the phishing classifier layer to an embedding layer used to generate URL embeddings. 21. Generating a URL embedding of characters extracted from an example URL; 19. The computer-implemented method of claim 18, further comprising: generating an HTML encoding of the HTML tokens extracted from the content page accessed via the exemplary URL; and generating an image embedding of the image captured from the rendering. 22. The computer-implemented method of clause 18, further comprising: for a second exemplary URL, training an HTML encoder-decoder to generate an HTML encoding using HTML tokens extracted from the content page of the second exemplary URL, encoded, and then decoded to recreate an image captured from the content page of the second exemplary URL. 23. The computer-implemented method of clause 17, further comprising extracting HTML tokens from the content page that belong to a predetermined token vocabulary and ignoring portions of the content page that do not belong to the predetermined token vocabulary. 24. The computer-implemented method of clause 23, further comprising limiting the extraction to a predetermined number of HTML tokens. 25. The computer-implemented method of clause 23, further comprising generating an HTML encoding of up to 64 HTML tokens.

Claims

1. A phishing classifier that classifies universal resource locators (URLs) and content pages accessed via said URLs as phishing or non-phishing by executing computer program instructions on a processor, comprising: a URL feature hasher that parses the URL into features and hashes the features to generate a URL feature hash; Accessing the content of said URL and internally rendering the content page; Extracting HyperText Markup Language (HTML) tokens from the content page; a headless browser configured to capture an image of at least a portion of the rendering of the content page; and trained on HTML tokens extracted from content pages of example URLs, encoded into an embedding space, and then decoded to recreate images captured from renderings of said content pages; an HTML encoder that generates an HTML encoding of the HTML tokens extracted from the content page; an image embedder that is pre-trained on images and generates image embeddings of the images captured from the content pages; trained on URL feature hashes, HTML encodings, and image embeddings of the example URLs, each example URL accompanied by a ground truth classification as phishing or not phishing; a phishing classifier layer that processes the URL feature hashes, the HTML encoding, and the image embeddings of the URL to generate at least one likelihood score that represents how likely the URL and the content page accessed via the URL are to present a phishing risk; and a phishing classifier,

2. 2. The phishing classifier of claim 1, further comprising the headless browser configured to extract the HTML tokens from the content page that belong to a predetermined token vocabulary and to ignore portions of the content page that do not belong to the predetermined token vocabulary.

3. The phishing classifier of claim 1 , further comprising the headless browser configured to extract up to 64 of the HTML tokens for generation of an HTML encoding.

4. 1. A non-transitory computer-readable storage medium having computer program instructions recorded thereon for classifying universal resource locators (URLs) and content pages accessed via the URLs as phishing or non-phishing, the instructions, when executed on a processor, comprising: applying a URL feature hasher to parse the URL into features and hash the features to generate a URL feature hash; Applying a headless browser to access the content of the URL and internally rendering the content page; Extracting HyperText Markup Language (HTML) tokens from the content page; capturing an image of at least a portion of the rendering of the content page; trained on HTML tokens extracted from content pages of example URLs, encoded into an embedding space, and then decoded to reproduce images captured from renderings of said content pages; applying an HTML encoder to generate an HTML encoding of the HTML tokens extracted from the content page; applying an image embedder that is pre-trained on images and that generates an image embedding of the image captured from the content page; applying a phishing classifier layer trained on URL feature hashes, HTML encodings, and image embeddings of the example URLs, with each example URL being accompanied by a ground truth classification as either phishing or not phishing; applying a phishing classifier layer that processes the URL feature hashes, the HTML encoding, and the image embeddings of the URL to thereby generate at least one likelihood score that represents how likely the URL and the content page accessed via the URL are to present a phishing risk; 1. A non-transitory computer-readable storage medium that performs actions including:

5. 5. The non-transitory computer-readable storage medium of claim 4, wherein the actions further include training the HTML encoder to generate the HTML encoding for a second exemplary URL using HTML tokens extracted from a content page of the second exemplary URL, encoded, and then decoded to recreate an image captured from the content page of the second exemplary URL.

6. The non-transitory computer-readable storage medium of claim 4 , wherein the actions further comprise generating the image embedding using a ResNet embedder or a variant of a ResNet embedder.

7. A phishing classifier that classifies universal resource locators (URLs) and content pages accessed via said URLs as phishing or non-phishing by executing computer program instructions on a processor, comprising: a URL link sequence extractor that extracts characters in a predetermined character set from the URL to generate a URL character sequence; a URL embedder that generates a URL embedding of the URL character sequence; a headless browser configured to access the content page of the URL and extract HyperText Markup Language (HTML) tokens from the content page; an HTML encoder that is pre-trained on HTML tokens extracted from content pages of example URLs, encoded into an embedding space, and then decoded to recreate images captured from renderings of the content pages, and that generates HTML encodings of the HTML tokens extracted from the content pages; trained on URL embedding and HTML encoding of the example URLs, each example URL accompanied by a ground truth classification as phishing or not phishing; a phishing classifier layer that processes the URL embedding and the HTML encoding of the URL to generate at least one likelihood score that represents how likely the URL and the content page accessed via the URL are to present a phishing risk; and a phishing classifier,

8. 8. The phishing classifier of claim 7, further comprising the headless browser configured to extract the HTML tokens from the content page that belong to a predetermined token vocabulary and to ignore portions of the content page that do not belong to the predetermined token vocabulary.

9. The phishing classifier of claim 7 , further comprising the headless browser configured to extract up to 64 of the HTML tokens for generation of an HTML encoding.

10. 1. A non-transitory computer-readable storage medium having computer program instructions recorded thereon for classifying universal resource locators (URLs) and content pages accessed via the URLs as phishing or non-phishing, the instructions, when executed on a processor, comprising: applying a URL link sequence extractor to extract characters in a predetermined character set from the URL to generate a URL character sequence; applying a URL embedder to generate a URL embedding of the URL character sequence; applying a headless browser configured to access a content page of the URL and extract HyperText Markup Language (HTML) tokens from the content page; applying an HTML encoder that has been pre-trained on HTML tokens extracted from a content page of an example URL, encoded into an embedding space, and then decoded to recreate an image captured from a rendering of the content page, to generate an HTML encoding of the HTML tokens extracted from the content page; trained on URL embedding and HTML encoding of the example URLs, each example URL accompanied by a ground truth classification as phishing or not; applying a phishing classifier layer that processes the URL embedding and the HTML encoding of the URL to generate at least one likelihood score that represents how likely the URL and the content page accessed via the URL are to present a phishing risk; 1. A non-transitory computer-readable storage medium that performs actions including:

11. 11. The non-transitory computer-readable storage medium of claim 10, wherein the actions further comprise implementing the headless browser configured to extract the HTML tokens from the content page that belong to a predetermined token vocabulary and to ignore portions of the content page that do not belong to the predetermined token vocabulary.

12. 11. The non-transitory computer-readable storage medium of claim 10, wherein the actions further comprise implementing the headless browser configured to extract up to 64 of the HTML tokens for generating an HTML encoding.

Citation Information

Patent Citations

  • Identification method of phishing sites

    CN102708186A

  • People Engine Optimization

    US20130254214A1

  • System and method to provide automatic classification of phishing sites

    US20140033307A1

  • Discovering website phishing attacks

    US20190068638A1

  • Uniform Resource Locator Classifier and Visual Comparison Platform for Malicious Site Detection

    US20200314122A1