Application identification for phishing detection
Advanced application identification techniques in security platforms improve phishing detection by combining protocol and content analysis, addressing the limitations of existing methods and enhancing detection accuracy and efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- PALO ALTO NETWORKS INC
- Filing Date
- 2023-03-31
- Publication Date
- 2026-05-13
AI Technical Summary
Existing methods for detecting phishing attacks, such as signature detection and URL-based approaches, suffer from high false positive rates due to the similarity between legitimate and phishing sites, and are ineffective when cloaking techniques are used, necessitating improved techniques for accurate phishing site identification.
Implementing advanced application identification (App-ID) techniques in security platforms, combining protocol identification and website content analysis to enhance phishing detection, reducing false positives and improving detection efficiency.
The integration of advanced application identification techniques in firewalls and security platforms significantly enhances phishing detection accuracy and reduces false positive rates, enabling more effective and efficient inline blocking of phishing sites.
Smart Images

Figure 0007858074000001 
Figure 0007858074000002 
Figure 0007858074000003
Abstract
Description
Background Art
[0001] A firewall generally enables authorized communications to pass through the firewall while protecting the network from unauthorized access. A firewall is typically a device, such as a computer, or a set of devices, or software running on a device, that provides a firewall function for network access. For example, a firewall can be integrated into the operating system of a device (such as a computer, smartphone, or other type of network - communicable device). A firewall can also be integrated into or run as software on a computer server, gateway, network / routing device (such as a network router), or data appliance (such as a security appliance or other type of dedicated device).
[0002] A firewall typically rejects or permits network transmissions based on a set of rules. These sets of rules are often called policies. For example, a firewall can filter inbound traffic by applying a set of rules or policies. A firewall can also filter outbound traffic by applying a set of rules or policies. A firewall can also perform basic routing functions.
Brief Description of the Drawings
[0003] Various embodiments of the present invention are disclosed in the following detailed description and the accompanying drawings. [Figure 1] FIG. 1 shows one exemplary environment in which a malicious application is detected and prevented from causing harm. [Figure 2A]Figure 2A shows one embodiment of the data appliance. [Figure 2B] Figure 2B is a functional diagram of a logical component according to one embodiment of a data appliance. [Figure 3] Figure 3 shows an exemplary logical component that may be included in a system for analyzing a sample. [Figure 4] Figure 4 shows one exemplary embodiment of a threat detection engine that uses application identification, according to several embodiments. [Figure 5A] Figure 5A shows an exemplary phishing site that can be detected using application identification for phishing detection, according to several embodiments. [Figure 5B] Figure 5B shows an exemplary phishing site that can be detected using application identification for phishing detection, according to several embodiments. [Figure 5C] Figure 5C shows an exemplary phishing site that can be detected using application identification for phishing detection, according to several embodiments. [Figure 6] Figure 6 is a flowchart illustrating the process of application identification for phishing detection according to several embodiments. [Figure 7] Figure 7 is another flowchart relating to the process of application identification for phishing detection, according to several embodiments. [Modes for carrying out the invention]
[0004] The present invention can be implemented in numerous ways, including processes, apparatus, systems, compositions, computer program products embodied on computer-readable storage media, and / or instructions stored in memory, and / or processors configured to execute instructions stored and / or provided by memory coupled to the processor. In this specification, these implementations, or any other forms the present invention may take, may be referred to as techniques. Generally, the order of the steps of the disclosed process may be modified within the scope of the invention. Unless otherwise specified, components such as processors or memory described as configured to perform a task may be implemented as general-purpose components temporarily configured to perform a task at a given time, or as specific components manufactured to perform a task. As used herein, the term “processor” refers to one or more devices, circuits, and / or processing cores configured to process data, such as computer program instructions.
[0005] A detailed description of one or more embodiments of the present invention, along with accompanying drawings illustrating the principles of the present invention, is provided below. While the present invention is described in relation to such embodiments, it is not limited to any embodiment. The scope of the present invention is limited only by the claims, and the present invention encompasses numerous alternatives, modifications, and equivalents. To provide a complete understanding of the present invention, numerous specific details are provided below. These details are provided for illustrative purposes, and the present invention may be carried out in accordance with the claims without some or all of these specific details. For clarity, technical materials known in the art related to the present invention are not described in detail so as not to unnecessarily obscure the present invention.
[0006] A firewall generally allows authorized communications to pass through while protecting the network from unauthorized access. Typically, a firewall is a device, a set of devices, or software running on a device that provides firewall functionality for network access. For example, a firewall can be integrated into the operating system of a device (e.g., a computer, smartphone, or other type of network-enabled device). Firewalls can also be integrated or run as software applications on various types of devices or security devices, such as computer servers, gateways, network / routing devices (e.g., network routers), or data appliances (e.g., security equipment or other types of special-purpose devices), and in some implementations, specific operations can be implemented on special-purpose hardware such as ASICs or FPGAs. ru.
[0007] A firewall typically denies or allows network transmissions based on a set of rules. These sets of rules are often referred to as policies (e.g., network policies or network security policies). For example, a firewall can filter inbound traffic by applying a set of rules or policies to prevent unwanted external traffic from reaching the protected device. A firewall can also filter outbound traffic by applying a set of rules or policies (e.g., allow, block, monitor, notify, log, and / or other actions that may be specified in a firewall rule or firewall policy, which can be triggered based on various criteria, as described herein). A firewall can also filter local network (e.g., intranet) traffic by similarly applying a set of rules or policies.
[0008] Security devices (e.g., security equipment, security gateways, security services, and / or other security devices) can perform a variety of security operations (e.g., firewalls, anti-malware, intrusion prevention / detection, proxies, and / or other security functions), network functions (e.g., routing, quality of service (QoS), workload balancing of network-related resources, and / or other network functions), and / or other security and / or network-related functions. For example, routing functions can be based on source information (e.g., IP address and port), destination information (e.g., IP address and port), and protocol information.
[0009] A basic packet filtering firewall filters network communication traffic by inspecting individual packets transmitted over the network (e.g., a stateless packet filtering firewall, a packet filtering firewall, or a first-generation firewall). A stateless packet filtering firewall typically inspects the individual packets themselves and then applies rules based on the inspected packets (e.g., using a combination of source and destination address information, protocol information, and port number).
[0010] An application firewall can also perform application layer filtering (for example, using an application layer filtering firewall or a second-generation firewall that operates at the application level of the TCP / IP stack). An application layer filtering firewall or application firewall can generally identify a given application and protocol (e.g., web browsing using Hypertext Transfer Protocol (HTTP), Domain Name System (DNS) requests, file transfers using File Transfer Protocol (FTP), and various other types of applications and protocols such as Telnet, DHCP, TCP, UDP, and TFTP (GSS)). For example, an application firewall can block unauthorized protocols attempting to communicate on standard ports (for example, unauthorized / unauthorized policy protocols attempting to sneak through by using non-standard ports for that protocol can generally be identified using an application firewall).
[0011] A stateful firewall can also perform stateful-based packet inspection, where each packet is examined within the context of a set of packets associated with its network transmission packet flow. This firewall technique is commonly referred to as stateful packet inspection because it maintains a record of all connections passing through the firewall and can determine whether a packet is the start of a new connection, part of an existing connection, or an invalid packet. For example, the state of a connection can itself be one of the criteria that trigger rules in a policy.
[0012] Advanced or next-generation firewalls, as described above, can perform stateless and stateful packet filtering and application layer filtering. Next-generation firewalls can also perform additional firewall technologies. For example, a given new firewall, sometimes referred to as an advanced or next-generation firewall, can also identify users and content. In particular, a given next-generation firewall extends the list of applications these firewalls can automatically identify to thousands of applications. Examples of such next-generation firewalls are commercially available from Palo Alto Networks (e.g., Palo Alto Networks' PA Series Next-Generation Firewalls, Palo Alto Networks' VM Series Virtualization Next-Generation Firewalls, and CN Series Container Next-Generation Firewalls).
[0013] For example, Palo Alto Networks' next-generation firewalls use a variety of identification technologies to enable enterprises and service providers to identify and control applications, users, and content—not just ports, IP addresses, and packets. These identification technologies include Application ID (App-ID) for precise application identification, User ID (User-ID) for user identification (e.g., User ID), Content ID (Content-ID) for real-time content scanning (e.g., to control web surfing and restrict data and file transfers), and Device ID (Device-ID) (e.g., for identifying IoT device types). These identification technologies allow enterprises to securely enable application use using business-relevant concepts, instead of following the traditional approach provided by conventional port-blocking firewalls. Furthermore, purpose-specific hardware for next-generation firewalls generally offers a higher level of performance for application inspection than software running on general-purpose hardware (for example, security devices offered by Palo Alto Networks, which utilize dedicated, function-specific processing tightly integrated with a single-path software engine to minimize latency while maximizing network throughput for Palo Alto Networks' PA Series next-generation firewalls).
[0014] Overview of techniques for application identification for phishing detection
[0015] Phishing is a growing security threat, with approximately 1.5 million new phishing sites identified each month. A significant increase in phishing attacks has been observed, particularly since the start of the Covid-19 pandemic in late 2019 (resulting in a large number of companies allowing employees and contractors to work remotely from home, for example).
[0016] Existing approaches to detecting phishing attacks have several drawbacks. For example, signature detection approaches based on pattern matching are generally prone to false positives (FP) due to the high similarity between the original web page of a target site and the phishing page attempting to emulate / mimic that site (e.g., phishing pages are typically designed to mimic original web pages, such as login pages for banking sites, e-commerce sites, streaming sites, etc.). Another example is URL-based detection approaches, which generally require an offline analyzer to observe the exact same page as the customer (often not, given the various cloaking techniques that use geolocation targets, such as a customer's corporate firewall, where only a location target in Germany receives the phishing page, in contrast to customers elsewhere in Europe. And if the security service provider has servers located in a different geolocation than the target location, it will not result in a phishing site being presented for security analysis by the security service provider's cloud-based security platform, which can perform URL or other security analysis).
[0017] Therefore, new and improved techniques are needed to detect phishing.
[0018] Thus, various techniques for application identification (App-ID) for phishing detection are disclosed. For example, the disclosed techniques, as described herein, in combination with various other techniques, facilitate the integration of the use of advanced application identification (App-ID) (e.g., as used herein, advanced application identification (App-ID) includes (1) protocol identification (e.g., HTTP, HTTPS, SSL, TLS, etc.), and (2) website identification based on analysis of the content of web pages, generally referred to herein as site similarity) to provide more efficient and effective phishing detection (e.g., increased coverage and lower FP rate).
[0019] In some embodiments, a system / process / computer program product for application identification for phishing detection includes monitoring network activity associated with a session to detect requests to access a site, determining an advanced application identification associated with the site, and identifying the site as a phishing site based on the advanced application identification.
[0020] For example, the disclosed techniques for application identification for phishing detection can be performed to facilitate more effective and efficient in-line detection of phishing sites and blocking (e.g., in a security platform such as a perimeter firewall). The disclosed techniques also facilitate an improvement in the phishing detection rate as compared to existing stand-alone approaches. Further, the disclosed techniques result in a lower FP rate using such combinations of application identification and detection techniques (e.g., in contrast to existing approaches that simply perform only analysis of web page content) as further described below.
[0021] Therefore, a new and improved security solution that facilitates applying application identification for phishing detection is disclosed in some embodiments using a security platform (e.g., a firewall (FW) / next-generation firewall (NGFW), a network sensor operating in place of a firewall, or another (virtual) device / component that can implement security policies using the disclosed techniques, including, for example, Palo Alto Networks' PA Series next-generation firewalls, Palo Alto Networks' VM Series virtualized next-generation firewalls, and CN Series container next-generation firewalls, and / or other commercially available virtual-based or container-based firewalls that can be similarly implemented and configured to execute the disclosed techniques).
[0022] These and other embodiments and examples for applying application identification for phishing detection are further described below.
[0023] Exemplary system architecture for application identification for phishing detection
[0024] Accordingly, in some embodiments, the disclosed techniques include providing a security platform (for example, the security function / platform may be implemented using another (virtualized) device / component that runs on a virtual / physical NGFW solution commercially available from Palo Alto Networks, Inc., or another security platform / NFGW, such as a firewall (FW) / next-generation firewall (NGFW), a network sensor acting in place of a firewall, or PANOS running on a virtual / physical NGFW solution, including, for example, Palo Alto Networks' PA Series Next Generation Firewall, Palo Alto Networks' VM Series Virtualized Next Generation Firewall, and CN Series Container Next Generation Firewall, and / or other commercially available virtualization-based or container-based firewalls may similarly be implemented and configured to perform the disclosed techniques), the security platform being configured to provide DPI capabilities (e.g., stateful inspection) that apply App-ID for phishing detection based on policies (e.g., Layer 7 security and / or other security policy enforcement).
[0025] Figure 1 shows an exemplary environment in which a malicious application ("malware") is detected and prevented from causing harm. As will be described in more detail below, malware classifications (for example, as determined by security platform 122) can be shared and / or refined in various ways among the various entities included in the environment shown in Figure 1. Using the techniques described herein, devices such as endpoint client devices 104-110 can be protected from such malware (including, for example, phishing-related malware).
[0026] As used herein, “malware” refers to an application that engages in behavior that a user would not authorize or would not authorize if fully informed, whether or not it is kept secret (and illegal or not). Examples of malware include Trojans, viruses, rootkits, spyware, hacking tools, keyloggers, etc. One example of malware is a desktop application that encrypts a user’s stored data (e.g., ransomware). Another example of malware is a desktop application that collects end-user activity and / or various information associated with the user and reports it to a remote server (e.g., spyware). Other forms of malware can also be detected and blocked using improved phishing detection techniques (e.g., keyloggers), as further described herein.
[0027] The term “phishing” is used throughout this specification and collectively refers to email messages, text messages, and / or various other types of messages, including social or productivity-related messaging platforms, etc., through which an attacker sends malicious messages (e.g., spoofed, fake, or otherwise deceptive content and / or web links) designed to deceive a user / human into taking an action to facilitate malicious and / or unwanted activity (e.g., revealing sensitive information to the attacker, such as login credentials or other personal / confidential information or certificates, or facilitating the download of malware by visiting a spoofed website, deploying malware on the user's infrastructure, such as ransomware, etc.). As will be further discussed herein, detecting phishing activity is becoming an increasingly technical challenge. This is because phishing attacks are becoming more sophisticated and often transparently mirror the targeted site (e.g., a website), allowing attackers to observe everything while the victim navigates the site and cross any additional security boundaries with the victim.
[0028] The techniques described herein can be used with various platforms (e.g., desktops, mobile devices, gaming platforms, embedded systems, etc.) and / or with various forms of phishing attacks (e.g., via email messages, text messages, and / or various other types of messages, via social or productivity-related messaging platforms, etc.). In the exemplary environment shown in Figure 1, client devices 104-108 are a laptop computer, a desktop computer, and a tablet (each) located within the corporate network 140. Client device 110 is a laptop computer located outside the corporate network 140.
[0029] The data appliance 102 is configured to enforce policies regarding communication between client devices, such as client devices 104 and 106, and nodes outside the corporate network 140 (e.g., those reachable via the external network 118). Examples of such policies include those that manage traffic shaping, quality of service, and traffic routing. Other examples of policies include security policies that require scanning for threats in incoming (and / or outgoing) email attachments, website content, files exchanged via instant messaging programs, and / or other file transfers. In some embodiments, the data appliance 102 is also configured to enforce policies regarding traffic that remains within the corporate network 140.
[0030] One embodiment of the data appliance is shown in Figure 2A. The example shown is a representation of the physical components included in the data appliance 102 in various embodiments. Specifically, the data appliance 102 includes a high-performance multi-core central processing unit (CPU) 202 and random access memory (RAM) 204. The data appliance 102 also includes storage 210 (such as one or more hard disks or solid-state units). In various embodiments, the data appliance 102 stores information (in either RAM 204, storage 210, and / or other appropriate locations) used to monitor the enterprise network 110 and implement the disclosed technology. Examples of such information include application identifiers, content identifiers, user identifiers, requested URLs, IP address mappings, policy and other configuration information, signatures, hostname / URL classification information, malware profiles, machine learning models, IoT device classification information, etc. The data appliance 102 may also include one or more optional hardware accelerators. For example, the data appliance 102 may include a cryptographic engine 206 configured to perform encryption and decryption operations, and one or more field-programmable gate arrays 208 configured to perform matching, act as a network processor, and / or perform other tasks.
[0031] The functionality described herein as being performed by the data appliance 102 can be provided / implemented in a variety of ways. For example, the data appliance 102 may be a dedicated device or a set of devices. The functionality provided by the data appliance 102 may also be integrated or run as software on a general-purpose computer, computer server, gateway, and / or network / routing device. In some embodiments, at least some of the services described as being provided by the data appliance 102 are instead (or in addition to) provided to a client device (e.g., client device 104 or client device 110) by software running on the client device (e.g., endpoint protection application 132).
[0032] Whenever the data appliance 102 is described as performing a task, a single component, a subset of components, or all components of the data appliance 102 may work together to perform the task. Similarly, whenever a component of the data appliance 102 is described as performing a task, a subcomponent may perform the task, and / or a component may perform the task together with other components. In various embodiments, parts of the data appliance 102 are provided by one or more third parties. Depending on factors such as the amount of computing resources available to the data appliance 102, various logical components and / or features of the data appliance 102 may be omitted, and the techniques described herein will be adapted accordingly. Similarly, additional logical components / features may be included in embodiments of the data appliance 102 as applicable. One example of a component included in the data appliance 102 in various embodiments is an application identification engine configured to identify applications (e.g., using various application signatures to identify applications based on packet flow analysis). For example, the application identification engine may determine the type of traffic a session is involved in, such as web browsing-social networking, web browsing-news, SSH, etc. Furthermore, as described herein, the application identification engine disclosed herein can perform advanced application identification. Specifically, advanced application identification includes both identifying an application (e.g., using various application signatures to identify an application based on packet flow analysis) and identifying a target site based on site similarity analysis (e.g., site similarity analysis of a target site and a potential phishing site, as described further herein).
[0033] Figure 2B is a functional diagram of a logical component in one embodiment of a data appliance. The example shown is a representation of a logical component that may be included in the data appliance 102 in various embodiments. Unless otherwise specified, the various logical components of the data appliance 102 can generally be implemented in various ways, including a set of one or more scripts (e.g., written in Java®, Python, etc., where applicable).
[0034] As shown in the diagram, the data appliance 102 includes a firewall and contains a management plane 232 and a data plane 234. The management plane is responsible for managing user interaction by providing a user interface for setting policies and displaying log data. The data plane is responsible for data management by performing packet processing and session processing.
[0035] The network processor 236 is configured to receive packets from client devices, such as client device 108, and provide them to the data plane 234 for processing. Whenever the flow module 238 identifies a packet as part of a new session, it generates a new session flow. Subsequent packets are identified as belonging to the session based on the flow lookup. SSL decryption is applied by the SSL decryption engine 240 where applicable; otherwise, processing by the SSL decryption engine 240 is omitted. The decryption engine 240 helps the data appliance 102 inspect and control SSL / TLS and SSH encrypted traffic, and therefore helps stop threats that might otherwise remain hidden within encrypted traffic. The decryption engine 240 can also help prevent sensitive content from leaving the corporate network 140. Decryption can be selectively controlled (e.g., enabled or disabled) based on parameters such as URL category, traffic source, traffic destination, user, user group, and port. In addition to the decryption policy (for example, specifying the session to decrypt), a decryption profile may be assigned to control various options for the session controlled by the policy. For example, the use of a specific cipher suite and encryption protocol version may be required.
[0036] The Application Identification (APP-ID) engine 242 is configured to determine the type of traffic a session is involved in. For example, the application identification engine 242 can recognize a GET request in incoming data and conclude that the session requires an HTTP decoder. In some cases, the identified application, such as a web browsing session, can be modified, and such modifications are noted by the data appliance 102. For example, a user might first browse a company wiki (classified as "Web Browsing-Productivity" based on the visited URL) and then browse a social networking site (classified as "Web Browsing-Social Networking" based on the visited URL). Different types of protocols have corresponding decoders. In addition, the application identification engine disclosed herein can perform advanced application identification. Specifically, advanced application identification includes both identifying applications (for example, using various application signatures to identify applications based on packet flow analysis, as described above) and identifying target sites based on site similarity analysis (for example, site similarity analysis of target sites and potential phishing sites, as further described below).
[0037] Based on the decision made by the application identification engine 242, the packet is sent by the threat engine 244 to an appropriate decoder configured to assemble the packet (which may be received out of order), perform tokenization, and extract information. The threat engine 244 also performs signature matching to determine what should happen to the packet. If necessary, the SSL encryption engine 246 can re-encrypt the decrypted data. The packet is then forwarded using the forwarding module 248 for forwarding (e.g., to a destination).
[0038] Furthermore, as shown in Figure 2B, policy 252 is also received and stored in the management plane 232. A policy may include one or more rules, which can be specified using a domain name and / or host / server name, and the rules may apply one or more signatures or other matching criteria or discoverative methods, such as for security policy enforcement on subscriber / IP flows, based on various extracted parameters / information from the monitored session traffic flow. An exemplary policy may include a phishing detection policy that uses the disclosed advanced application identification techniques in combination with one or more other parameters (e.g., the IP address associated with the site, geolocation information associated with the site / IP address, and / or other information as further described below). An interface (I / F) communicator 250 is provided for management communications (e.g., via (REST) API, messages, or network protocol communications, or other communication mechanisms).
[0039] Security platform
[0040] Returning to Figure 1, let's assume a malicious individual has created malware 130 (using system 120) (for example, by delivering it to a user's endpoint device via a phishing site, where the phishing site's URL is sent to the target user within the content of an email). The malicious individual hopes that a client device, such as client device 104, will run a copy of the malware 130, thereby compromising the client device and, for example, turning it into a bot in a botnet. The compromised client device may then be instructed to perform tasks (for example, cryptocurrency mining or participating in a denial of service attack), report information to an external entity, such as a command and control (C&C) server 150, and, where applicable, receive instructions from the C&C server 150.
[0041] Suppose data appliance 102 intercepts an email sent (for example, by system 120) to a user "Alice" operating client device 104. In this example, Alice receives the email and clicks a link to a phishing site, which could result in an attempt to download malware 130 by Alice's client device 104. However, in this example, data appliance 102 can perform application identification disclosed in relation to the phishing attempt and block access to the phishing site from Alice's client device 104, thereby preempting and preventing any such download of malware 130 to Alice's client device 104. As further described below, data appliance 102 can perform advanced application identification and use additional information associated with the target site (e.g., the IP address associated with the site, geolocation information associated with the site / IP address, and / or other information as further described below) to detect and block such phishing attempts.
[0042] In various embodiments, the data appliance 102 is configured to work in cooperation with the security platform 122. As one example, the security platform 122 may provide the data appliance 102 with a set of signatures for known malicious files (for example, as part of a subscription). If the set includes a signature for malware 130 (e.g., the MD5 hash of malware 130), the data appliance 102 can accordingly prevent the transmission of malware 130 to the client device 104 (for example, by detecting that the MD5 hash of an email attachment sent to the client device 104 matches the MD5 hash of malware 130). The security platform 122 also provides the data appliance 102 with a list of known malicious domains and / or IP addresses, enabling the data appliance 102 to block traffic between the corporate network 140 and the C&C server 150 (for example, here the C&C server 150 is known to be malicious). A list of malicious domains (and / or IP addresses) can also help the data appliance 102 determine when one of its nodes was compromised. For example, if client device 104 attempts to contact C&C server 150, such attempts are a strong indicator that client 104 is being compromised by malware (and remedial action should be taken accordingly, such as preventing client device 104 from communicating with other nodes in the corporate network 140). As will be described in more detail below, the security platform 122 can also provide other types of information to the data appliance 102 (for example, as part of a subscription). This includes a set of information available to the data appliance 102 for performing inline analysis of files, such as a set of information for performing advanced application identification for phishing (e.g., target site information for performing site similarity, IP addresses associated with the site, geolocation information associated with the site / IP address, and / or other information as further described below).
[0043] In various embodiments, if the signature of an attachment cannot be found, the data appliance 102 may take various actions. As a first example, the data appliance 102 can fail-safe by blocking the transmission of any attachment that is not whitelisted as benign (e.g., does not match the signature of a known good file). The drawback of this approach is that many legitimate attachments may be unnecessarily blocked as potential malware when they are actually benign. As a second example, the data appliance 102 can fail-danger by allowing the transmission of any attachment that is not blacklisted as malicious (e.g., does not match the signature of a known bad file). The drawback of this approach is that newly created malware (not previously detected by platform 122) may cause harm. As a third example, the data appliance 102 may be configured to provide a file (e.g., malware 130) to the security platform 122 for static / dynamic analysis to determine whether it is malicious and / or classify it if it is not.
[0044] The security platform 122 stores a copy of the received sample in storage 142, and analysis is initiated (or scheduled, if applicable). One example of storage 142 is an Apache Hadoop Cluster (HDFS). The results of the analysis (and additional information about the application) are stored in database 146. If the application is determined to be malicious, the data appliance may be configured to automatically block file downloads based on the analysis results. Furthermore, a signature for the malware is generated and distributed (to data appliances such as data appliances 102, 136, and 148, for example) to automatically block future file transfer requests to download files determined to be malicious.
[0045] In various embodiments, the security platform 122 comprises one or more dedicated, commercially available hardware servers (e.g., having a multi-core processor, 32G+ RAM, a Gigabit network interface adapter, and a hard drive) running a typical server-class operating system (e.g., Linux®). The security platform 122 may be implemented across a scalable infrastructure including multiple such servers, solid-state drives, and / or other applicable high-performance hardware. The security platform 122 may comprise several distributed components, including components provided by one or more third parties. For example, some or all of the security platform 122 may be implemented using Amazon Elastic Compute Cloud (EC2) and / or Amazon Simple Storage Service (S3). Furthermore, as with the data appliance 102, whenever the security platform 122 is referred to as performing tasks such as storing or processing data, it should be understood that subcomponents or multiple subcomponents of the security platform 122 may cooperate (individually or in cooperation with third-party components) to perform those tasks. As one example, the security platform 122 can work with one or more virtual machine (VM) servers, such as VM server 124, to optionally perform static / dynamic analysis.
[0046] One example of a virtual machine server is a physical machine containing commercial server-class hardware (e.g., a multi-core processor, 32+ gigabytes of RAM, and one or more Gigabit network interface adapters) running commercially available virtualization software such as VMware ESXi, Citrix XenServer, or Microsoft Hyper-V. In some embodiments, the virtual machine server is omitted. Furthermore, the virtual machine server may be under the control of the same entity managing the security platform 122, but may also be provided by a third party. As one example, the virtual machine server may rely on EC2, and the rest of the security platform 122 is owned by the operator of the security platform 122 and provided by dedicated hardware under the control of that operator. VM server 124 is configured to provide one or more virtual machines 126-128 for emulating client devices. The virtual machines can run various operating systems and / or versions thereof. Observed behavior resulting from running applications within the virtual machines is logged and analyzed (e.g., for indicators that the application is malicious). In some embodiments, log analysis is performed by a VM server (e.g., VM server 124). In other embodiments, the analysis is performed at least partially by other components of the security platform 122, such as the coordinator 144.
[0047] In various embodiments, the security platform 122 makes the results of sample analysis available to the data appliance 102 as part of a subscription, via a list of signatures (and / or other identifiers). For example, the security platform 122 may periodically (e.g., daily, hourly, or at some other interval and / or based on events configured by one or more policies) send content packages that identify malware files, phishing sites, etc. One exemplary content package includes a list of identified phishing sites, along with information such as target site names, URLs, and site similarity information, as well as various other information about each target site (e.g., IP addresses associated with the target site, geolocation information associated with the site / IP address, and / or other information as further described below). The subscription can cover the analysis of those very files that have been intercepted by the data appliance 102 and sent to the security platform 122 by the data appliance 102, and can also cover the signatures of malware known to the security platform 122. As will be described in more detail below, platform 122 can also make other types of information available for phishing detection using site model 152, which is performed using disclosed advanced application identification techniques that can help data appliance 102 detect phishing sites and perform inline blocking of phishing sites.
[0048] In various embodiments, the security platform 122 is configured to provide security services to various entities in addition to (or, where applicable, instead of) the operators of the data appliance 102. For example, other companies having their respective corporate networks 114 and 116, and their respective data appliances 136 and 148, can contract with the operator of the security platform 122. Other types of entities can also utilize the services of the security platform 122. For example, an Internet service provider (ISP) providing internet services to a client device 110 can contract with the security platform 122 to analyze applications that the client device 110 attempts to download. As another example, the owner of the client device 110 can install software on the client device 110 that communicates with the security platform 122 (e.g., receiving content packages from the security platform 122, using the received content packages to check attachments according to the techniques described herein, and sending applications to the security platform 122 for analysis).
[0049] Sample analysis using static / dynamic analysis
[0050] Figure 3 shows an exemplary logical component that may be included in a system for analyzing samples. The analysis system 300 may be implemented using a single device. For example, the functionality of the analysis system 300 may be implemented in a malware analysis module 112 integrated into the data appliance 102. The analysis system 300 may also be implemented collectively across multiple separate devices. For example, the functionality of the analysis system 300 may be provided by a security platform 122.
[0051] In various embodiments, the analysis system 300 utilizes a list, database, or other collection (collectively shown as collection 314 in Figure 3) of known secure content and / or known bad content. Collection 314 may be obtained in various ways, including through a subscription service (e.g., provided by a third party) and / or as a result of other processing (e.g., performed by the data appliance 102 and / or security platform 122). Examples of information contained in collection 314 are as follows: That is, URLs, domain names, and / or IP addresses of known malicious servers; URLs, domain names, and / or IP addresses of known secure servers; URLs, domain names, and / or IP addresses of known command and control (C&C) domains; signatures, hashes, and / or other identifiers of known malicious applications; signatures, hashes, and / or other identifiers of known secure applications; signatures, hashes, and / or other identifiers of known malicious files (e.g., OS exploit files); signatures, hashes, and / or other identifiers of known secure libraries; and signatures, hashes, and / or other identifiers of known malicious libraries.
[0052] In various embodiments, when a new sample is received for analysis (for example, when no existing signature associated with the sample exists in the analysis system 300), it is added to queue 302. As shown in Figure 3, application 130 is received by system 300 and added to queue 302.
[0053] The coordinator 304 monitors the queue 302, and when a resource (e.g., a static analysis worker) becomes available, the coordinator 304 fetches a sample from the queue 302 for processing (e.g., fetching a copy of malware 130). In particular, the coordinator first provides the sample to the static analysis engine 306 for static analysis (305). In some embodiments, one or more static analysis engines are included within the analysis system 300, where the analysis system 300 is a single device. In other embodiments, static analysis is performed by a separate static analysis server containing multiple workers (i.e., multiple instances of the static analysis engine 306).
[0054] The static analysis engine obtains general information about the sample and includes it (along with heuristic information and other information, where applicable) in the static analysis report 308. The report may be generated by the static analysis engine or by a coordinator 304 (or another appropriate component) which may be configured to receive information from the static analysis engine 306. As one example, the static analysis of a target site may include site information for performing a site similarity analysis in order to perform the disclosed techniques for phishing detection. In some embodiments, the collected information is stored in a database record of the sample (e.g., in database 316) instead of, or in addition to, a separate static analysis report 308 being generated (i.e., a portion of the database record forms the report 308). In some embodiments, the static analysis engine also forms a verdict about the application (e.g., "safe", "suspicious", or "malicious"). For example, if an application contains even one “malicious” static feature (e.g., the application contains a hard link to a known malicious domain), the determination may be “malicious.” For another example, points may be assigned to each feature (e.g., based on severity if found, based on how reliable the feature is in predicting malice, etc.). Then, based on the number of points associated with the static analysis results, the static analysis engine 306 (or the coordinator 304, if applicable) may assign a determination.
[0055] Once the static analysis is complete, the coordinator 304 locates an available dynamic analysis engine 310 to perform dynamic analysis on the application. Similar to the static analysis engine 306, the analysis system 300 may directly include one or more dynamic analysis engines. In other embodiments, the dynamic analysis is performed by a separate dynamic analysis server, which includes multiple workers (i.e., multiple instances of the dynamic analysis engine 310).
[0056] Each dynamic analysis worker manages a virtual machine instance. In some embodiments, the results of a static analysis (e.g., performed by the static analysis engine 306) are provided as input to the dynamic analysis engine 310, whether in report format (308) and / or stored in the database 316 or otherwise. For example, static report information may be used to help the dynamic analysis engine 310 select / customize the virtual machine instances to be used (e.g., Microsoft Windows 7 SP 2 vs. Microsoft Windows 10 Enterprise, or iOS 11.0 vs. iOS 12.0). If multiple virtual machine instances are running concurrently, a single dynamic analysis engine may manage all instances, or multiple dynamic analysis engines may be used, where applicable (e.g., each managing its own virtual machine instance). During the dynamic part of the analysis, actions performed by the application (including network activity) are analyzed, as will be described in more detail below.
[0057] In various embodiments, static analysis of a sample is omitted or, if applicable, performed by a separate entity. As one example, conventional static and / or dynamic analysis may be performed on a file by a first entity. Once a given file is determined to be malicious (e.g., by the first entity), the file may be provided to a second entity (e.g., an operator of the security platform 122) for additional analysis, particularly regarding the use of malware in network activity (e.g., by the dynamic analysis engine 310).
[0058] The environment used by the analysis system 300 is instrumented / hooked so that any behavior observed while the application is running is logged as it occurs (e.g., using a customized kernel that supports hooking and logcat). Network traffic associated with the emulator is also captured (e.g., using pcap). Log / network data may be stored as temporary files on the analysis system 300. It may also be stored more permanently (e.g., using HDFS or another suitable storage technology, or a combination of technologies such as MongoDB). The dynamic analysis engine (or another suitable component) can compare connections made by the sample with a list of domains, IP addresses, etc. (314) and determine whether the sample communicated with (or attempted to communicate with) a malicious entity.
[0059] Similar to the static analysis engine, the dynamic analysis engine stores the results of its analysis in the database 316 in records associated with the application being tested (and / or, if applicable, includes the results in report 312). In some embodiments, the dynamic analysis engine also forms a determination about the application (e.g., “safe,” “suspicious,” or “malicious”). For example, the determination may be “malicious” even if only one “malicious” action is taken by the application (e.g., an attempt is made to contact a known malicious domain, or an attempt is observed to extract sensitive information). For another example, points may be assigned to actions performed (e.g., based on severity if found, based on how reliable the action is in predicting malice, etc.). The determination may then be assigned by the dynamic analysis engine 310 (or, if applicable, the coordinator 304) based on the number of points associated with the dynamic analysis results. In some embodiments, the final determination associated with the sample is made (e.g., by the coordinator 304) based on a combination of reports 308 and 312.
[0060] Advanced application identification for phishing detection
[0061] Figure 4 shows part of one exemplary embodiment relating to a threat detection engine using application identification, according to several embodiments. As similarly described above, in various embodiments, the data appliance 102 includes a threat engine 244. The threat engine includes an application identification (App-ID) engine 402 that performs advanced application identification. Thus, the App-ID engine 402 incorporates both protocol decoding for application identification and site similarity matching 406 during the respective decoder stage and a pattern matching stage performed inline in the data appliance 102. The results of the two stages are merged by a detector stage, as shown in 410. In some embodiments, the detector 410 also utilizes IP address information and / or other features, as shown in 408, in combination with advanced application identification to detect phishing sites (for example, combining advanced application identification with IP, URL, and / or domain anomaly detection features, as further described below).
[0062] When data appliance 102 receives a packet, it performs a session match to determine which session the packet belongs to (enabling data appliance 102 to support concurrent sessions). Each session has a session state involving a specific protocol decoder (e.g., web browsing decoder, FTP decoder, or SMTP decoder). When a file is sent as part of a session, the applicable protocol decoder can utilize the appropriate file-specific decoder (e.g., PE file decoder, JavaScript® decoder, or PDF decoder).
[0063] Site similarity is performed based on a periodic static analysis of the target site, which is analyzed to cache the expected visual representation of the target site for pattern matching comparative analysis with subsequent potential phishing site candidates. In one exemplary implementation, site similarity analysis is performed by abstracting the web page code, which may include, for example, the title, header, footer, form elements, copyright information, and / or other fields and information.
[0064] In some embodiments, the disclosed techniques for phishing detection, implemented using the threat detection engine shown in Figure 4, include combining application identification with IP, URL, and / or domain anomaly detection features.
[0065] As a first example, phishing detection may be performed using application identification in combination with pattern matching (e.g., site similarity pattern matching as described herein). For example, advanced application identification involves performing application identification in combination with site similarity to improve phishing detection, which can reduce false positives compared to phishing detection techniques that rely solely on site similarity pattern matching. Thus, the disclosed advanced application identification technique for phishing detection includes using application identification for identifying hosting servers as a pre-filter for pattern-based site similarity detection (e.g., the pre-filter may be used to reduce the load on performing pattern matching, which promotes better performance for implementing the disclosed technique).
[0066] As a second example, phishing detection can be performed using advanced application identification in combination with an IP address range associated with a target site (e.g., a legitimate / verified Amazon Web Services (AWS) site). Many web services have servers in a fixed range of IP addresses. Thus, if advanced application identification identifies, based on site similarity detection, that a potential phishing site's webpage is similar to a target site (e.g., a legitimate / verified Amazon Web Services (AWS) site, webpage sign-in), but these servers associated with the potential phishing site have IP addresses outside the known IP address range associated with the target site, the threat detection engine can identify the potential phishing site as likely to be a phishing site.
[0067] As a third example, phishing detection can be performed using advanced application identification in combination with URL categories associated with target sites (e.g., legitimate / verified Amazon Web Services (AWS) sites). Generally, many phishing sites are observed to frequently change their domains. Thus, for web page content that appears to be similar to well-known sites (e.g., various well-known sites such as the top 100, or top 1000, or top 10000, etc., which can be regularly monitored to generate site similarity information for performing such site similarity analyses), however, if the potential phishing site is associated with a newly registered domain category (e.g., various publicly and / or commercially available URL-related services can provide information on domain name registration date information), the threat detection engine can identify the potential phishing site as likely to be phishing.
[0068] Exemplary use case of application identification for phishing detection
[0069] Figures 5A–5C illustrate exemplary phishing sites that can be detected using application identification for phishing detection, according to several embodiments. For example, the techniques disclosed for application identification for phishing detection may be implemented to run using inline malware detection, for example, on data appliance 102 (and / or on endpoints such as client devices 104, 106, and 108, using a security agent running on protected endpoints, and also as a cloud-based security service, such as using security platform 122, as described herein).
[0070] Referring to Figure 5A, an exemplary Wells Fargo Bank phishing site webpage is shown. However, the disclosed technology can perform advanced application identification to identify such exemplary phishing sites. In this example, the candidate phishing site resides at hXXps: / / storage[.]googleapis[.]com / awells-putlogs-308643420 / index[.]html. Thus, traffic directed to this address (e.g., Uniform Resource Identifier (URI) / Uniform Resource Locator (URL)) is identified as "Google Cloud Storage ("google-cloud-storage")" using the disclosed advanced application identification technique (e.g., implemented by the protocol ID component 404 of the App-ID engine 402 as shown in Figure 4). In addition, using advanced application identification technologies disclosed for hosting App-IDs or storage App-IDs (for example, implemented by the site similarity component 406 of the App-ID engine 402 as shown in Figure 4), and using site similarity disclosed based on pattern matching to detect web pages that appear similar to the Wells Fargo Bank login page, can be used by the detector 410 of the threat detection engine 244 to efficiently and effectively detect such phishing sites (for example, and similar phishing sites), since the threat detection engine (for example, the threat detection engine 244 in Figure 4) is configured to identify such well-known sites that are not legally placed in the hosting App-ID or storage App-ID as phishing sites.
[0071] Referring to Figure 5B, an exemplary AWS phishing site webpage is shown. However, the techniques disclosed can perform advanced application identification in combination with IP address range validation to detect examples of such phishing sites. Thus, when advanced application identification in combination with IP range validation is applied to a phishing site such as https: / / howitfix[.]com / app / aws / , the IP associated with this exemplary phishing site will not resolve to known AWS-related IP ranges (for example, since this phishing site is not even owned by Amazon, AWS has a given set of publicly available IP ranges at https: / / ip-ranges.ip-ranges.com / ip-ranges.json, which are periodically updated and can be stored in the IP / other features component 408 of the threat detection engine 244). Therefore, identifying potential phishing sites as appearing similar to AWS sites, and then checking their IP address ranges, can be used by the detector 410 of the threat detection engine 244 to efficiently and effectively detect these phishing sites (e.g., and similar such phishing sites).
[0072] Referring to Figure 5C, an exemplary Netflix phishing site webpage is shown. However, the disclosed technology can be combined with other features to perform advanced application identification to determine whether a potential phishing site is a newly registered domain. In this example, the candidate phishing site domain, Canada-neflxt[.]com, is a recently registered domain (e.g., less than one day, less than one week, less than one month, or any other threshold for recently registered domains may be used for this feature). In this case, the threat engine identifies the potential phishing site as similar to the well-known site, Netflix, but if, as in this example, the URL category indicates that this site is a newly registered domain (NRD) (e.g., NRD information may be updated regularly and stored in the IP / other features component 408 of the threat detection engine 244), then the detector 410 of the threat detection engine 244 can use this additional NRD attribute to efficiently and effectively detect this phishing site (e.g., and similar such phishing sites).
[0073] Additional illustrative processes for the techniques disclosed for application identification for phishing detection are described below.
[0074] An exemplary process for application identification for phishing detection.
[0075] Figure 6 is a flowchart relating to a process for application identification for phishing detection, according to several embodiments. In some embodiments, the process 600 shown in Figure 6 is performed by security platforms and techniques similarly described above, including embodiments described with respect to Figures 1-5C. In one embodiment, the process 600 is performed by the data appliance 102 described with respect to Figure 1, the security platform 122 described with respect to Figure 1 (e.g., as a cloud-based security service), virtual appliances (e.g., Palo Alto Networks' VM Series virtualized next-generation firewalls, CN Series container next-generation firewalls, and / or other commercially available virtual-based or container-based firewalls may be similarly implemented and configured to perform the disclosed techniques), SDN security solutions, cloud security services, and / or combinations or hybrid implementations of the foregoing described herein.
[0076] In 602, monitoring of network activity associated with sessions is performed to detect requests to access the site. For example, data appliance 102 may be configured to monitor sessions and detect requests to access the site, as described above with respect to Figures 1-2B.
[0077] In step 604, the determination of the advanced application identification associated with the site is performed. For example, advanced application identification may be performed using the protocol ID component 404 and site similarity component 406 of the threat detection engine 244, as similarly described above with respect to Figures 4 and 5A.
[0078] In 606, the site is identified as a phishing site based on advanced application identification. For example, policy enforcement may include blocking a session from accessing a phishing site, logging attempts to access a phishing site, monitoring and logging access to a phishing site, and warning the user before allowing access to a phishing site. And / or other actions / responses, or a combination thereof, may be performed based on the policy (e.g., a phishing / security policy which may be stored in policy 252 as shown in Figure 2B).
[0079] Figure 7 is another flowchart relating to the process of application identification for phishing detection according to several embodiments. In some embodiments, the process 700 shown in Figure 7 is performed by the security platform and techniques similarly described above, including the embodiments described above with respect to Figures 1-5C. In one embodiment, the process 700 is performed by the data appliance 102 described above with respect to Figure 1, the security platform 122 described above with respect to Figure 1 (e.g., as a cloud-based security service), virtual appliances (e.g., Palo Alto Networks' VM Series virtualized next-generation firewall, CN Series container next-generation firewall, and other commercially available virtual-based or container-based firewalls may similarly be implemented and configured to perform the disclosed techniques), SDN security solutions, cloud security services, and / or a combination or hybrid implementation of the foregoing described herein.
[0080] In 702, monitoring of network activity associated with sessions is performed to detect requests to access the site. For example, data appliance 102 may be configured to monitor sessions and detect requests to access the site, as described above with respect to Figures 1-2B.
[0081] In step 704, the determination of the advanced application identification associated with the site is performed. For example, advanced application identification may be performed using the protocol ID component 404 and site similarity component 406 of the threat detection engine 244, as similarly described above with respect to Figures 4 and 5A.
[0082] In 706, the IP / other features associated with the request to access the site are determined. For example, the determination of the IP / other features associated with the request to access the site (e.g., IP address / range, NRD information, URL category information, etc.) may be performed using the IP / other features component 408 of the threat detection engine 244, as similarly described above with respect to Figures 4, 5B, and 5C. For example, advanced application identification may include protocol identification that identifies phishing sites based at least in part on the use of URL categories, where the URL category is selected using URL categorization based on extracted domains associated with the request to access the site (for example, URL categorization can be obtained from a firewall using a URL categorization solution / service such as PanDB, which is commercially available from Palo Alto Networks, and / or another commercially available or publicly available URL categorization solution / service).
[0083] In 708, the identification of a site as a phishing site is performed based on advanced application identification and other IP / other characteristics. For example, policy enforcement may include blocking a session from accessing a phishing site, logging attempts to access a phishing site, monitoring and logging access to a phishing site, and warning the user before allowing access to a phishing site. And / or other actions / responses, or a combination thereof, may be performed based on the policy (e.g., a phishing / security policy which may be stored in policy 252 as shown in Figure 2B).
[0084] While the embodiments described above have been explained in some detail for the purpose of clarifying understanding, the present invention is not limited to the details provided. Many alternative methods exist for carrying out the present invention. The disclosed embodiments are illustrative and not limiting.
Claims
1. A system including a processor and memory, The aforementioned processor, By monitoring network activity associated with sessions, requests to access the site are detected. Determine the advanced application identification associated with the aforementioned site, and, Based on the aforementioned advanced application identification, the site is identified as a phishing site. It is configured in such a way, The aforementioned memory is It is coupled to the aforementioned processor and configured to provide instructions to the aforementioned processor, The aforementioned processor further, Based on site similarity analysis, determine the IP address associated with the request to the site in order to compare it with the expected IP address range associated with a well-known similar site, or Based on site similarity analysis, determine the URL category associated with the request to the site in order to compare it with the expected URL category associated with a well-known similar site. It is structured in such a way. system.
2. The aforementioned advanced application identification includes protocol identification, The system according to claim 1.
3. The advanced application identification includes a protocol identification that identifies the phishing site based at least partially on the URL category, The aforementioned URL categories are selected using URL categorization based on the extracted domains associated with the requests to access the aforementioned sites. The system according to claim 1.
4. The aforementioned advanced application identification includes site similarity identification. The system according to claim 1.
5. The advanced application identification includes site similarity identification, which determines whether the request for the site is visually similar to a well-known site. The system according to claim 1.
6. The aforementioned advanced application identification includes protocol identification and site similarity identification. The system according to claim 1.
7. Detecting the aforementioned site as a phishing site is performed inline using a data appliance. The system according to claim 1.
8. The aforementioned processor further, Another characteristic for detecting the aforementioned site as a phishing site is determining that the site is a newly registered domain (NRD). The system according to claim 1, configured as described above.
9. It is a method, The system's processor performs a step of monitoring network activity associated with a session and detecting requests to access a site. The processor performs the steps of determining an advanced application identification associated with the site, The processor identifies the site as a phishing site based on the advanced application identification, Includes, The above method further, Based on site similarity analysis, determine the IP address associated with the request to the site in order to compare it with the expected IP address range associated with a well-known similar site, or Based on site similarity analysis, determine the URL category associated with the request to the site in order to compare it with the expected URL category associated with a well-known similar site. including, method.
10. The aforementioned advanced application identification includes protocol identification, The method according to claim 9.
11. The advanced application identification includes a protocol identification that identifies the phishing site based at least partially on the URL category, The aforementioned URL categories are selected using URL categorization based on the extracted domains associated with the requests to access the aforementioned sites. The method according to claim 9.
12. The aforementioned advanced application identification includes site similarity identification. The method according to claim 9.
13. The advanced application identification includes site similarity identification, which determines whether the request for the site is visually similar to a well-known site. The method according to claim 9.
14. The aforementioned advanced application identification includes protocol identification and site similarity identification. The method according to claim 9.
15. Detecting the aforementioned site as a phishing site is performed inline using a data appliance. The method according to claim 9.
16. A computer program stored on a non-temporary computer-readable storage medium, which includes multiple computer instructions, When the computer instruction is executed, the computer will, This step involves monitoring network activity associated with a session, detecting requests to access the site, and The steps include determining the advanced application identification associated with the aforementioned site, Based on the aforementioned advanced application identification, the steps include identifying the site as a phishing site, They will implement this, and furthermore, Based on site similarity analysis, determine the IP address associated with the request to the site in order to compare it with the expected IP address range associated with a well-known similar site, or Based on site similarity analysis, determine the URL category associated with the request to the site in order to compare it with the expected URL category associated with a well-known similar site. To have it implemented Computer program.
17. The aforementioned advanced application identification includes protocol identification, The computer program according to claim 16.
18. The aforementioned advanced application identification includes site similarity identification. The computer program according to claim 16.