Application Identification for Phishing Detection
Advanced application identification techniques, incorporating protocol and site similarity analysis, enhance phishing detection efficiency and accuracy, addressing the limitations of existing methods by reducing false positives and improving in-line detection capabilities.
Patent Information
- Application Number
- JP2024563038
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-04-26
- Filing Date
- 2023-03-31
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-03-31
AI Technical Summary
Existing approaches to detecting phishing attacks are inefficient and prone to false positives, particularly due to the high similarity between original web pages and phishing pages, and the requirement for offline analysis that may not account for geolocation-based cloaking techniques.
The implementation of advanced application identification techniques for phishing detection, which include protocol identification and website identification based on content analysis, allowing for more efficient and effective in-line detection and blocking of phishing sites.
This approach significantly improves the phishing detection rate while reducing false positives, enabling more effective protection against sophisticated phishing attacks.
Smart Images

Figure 2025516170000001_ABST
Abstract
Description
Background Art
[0001] A firewall generally enables authorized communications to pass through the firewall while protecting the network from unauthorized access. A firewall is typically a device, such as a computer, or a set of devices, or software running on a device, that provides a firewall function for network access. For example, a firewall can be integrated into the operating system of a device (such as a computer, smartphone, or other type of network - communicable device). A firewall can also be integrated into or run as software on a computer server, gateway, network / routing device (such as a network router), or data appliance (such as a security appliance or other type of dedicated device).
[0002] A firewall typically rejects or permits network transmissions based on a set of rules. These sets of rules are often called policies. For example, a firewall can filter inbound traffic by applying a set of rules or policies. A firewall can also filter outbound traffic by applying a set of rules or policies. A firewall can also perform basic routing functions.
Brief Description of the Drawings
[0003] Various embodiments of the present invention are disclosed in the following detailed description and the accompanying drawings.
Figure 1
Figure 2A
Figure 2B
Figure 3
Figure 4
Figure 5A
Figure 5B
Figure 5C
Figure 6
Figure 7
[0004] The present invention can be implemented in a number of ways, including a process, an apparatus, a system, a composition, a computer program product embodied on a computer-readable storage medium, and / or instructions stored in a memory and / or executed by a processor coupled to the memory and / or provided by a processor configured to execute the instructions. As used herein, these implementations, or any other form that the present invention may take, may be referred to as a technique. Generally, the order of steps of the disclosed process may be changed within the scope of the present invention. Unless otherwise specified, components such as a processor or a memory described as being configured to perform a task are implemented as general-purpose components temporarily configured to perform the task at a given time or as specific components manufactured to perform the task. As used herein, the term "processor" refers to one or more devices, circuits, and / or processing cores configured to process data, such as computer program instructions.
[0005] A detailed description of one or more embodiments of the present invention is provided below along with the accompanying drawings that illustrate the principles of the present invention. The present invention is described in relation to such embodiments, but the present invention is not limited to any embodiment. The scope of the present invention is limited only by the claims and the present invention includes numerous alternatives, modifications, and equivalents. In order to provide a complete understanding of the present invention, numerous specific details are set forth in the following description. These details are provided for illustrative purposes only and the present invention may be practiced according to the claims without some or all of these specific details. Technical material known in the technical field associated with the present invention is not described in detail so as not to unnecessarily obscure the present invention.
[0006] A firewall generally permits authorized communications to pass through the firewall while protecting the network from unauthorized access. A firewall is typically a device, a set of devices, or software executed on a device that provides firewall functionality for network access. For example, a firewall can be integrated into the operating system of a device (such as a computer, smartphone, or other type of network - communicable device). A firewall can also be integrated or executed as a software application on various types of devices or security devices, such as a computer server, gateway, network / routing device (such as a network router), or data appliance (such as a security appliance or other type of special - purpose device), and in some implementations, certain operations can be implemented in special - purpose hardware such as an ASIC or FPGA. It does.
[0007] A firewall typically rejects or permits network transmissions based on a set of rules. This set of rules is often referred to as a policy (e.g., a network policy or a network security policy). For example, a firewall can filter inbound traffic by applying a set of rules or policies to prevent unwanted external traffic from reaching a protected device. A firewall can also filter outbound traffic by applying a set of rules or policies (e.g., allow, block, monitor, notify, or log, and / or other actions that can be specified in a firewall rule or a firewall policy, which can be triggered based on various criteria as described herein). A firewall can also filter local network (e.g., intranet) traffic by applying a set of rules or policies in a similar manner.
[0008] A security device (e.g., a security appliance, a security gateway, a security service, and / or other security devices) can perform various security operations (e.g., firewall, anti-malware, intrusion prevention / detection, proxy, and / or other security functions), network functions (e.g., routing, quality of service (QoS), workload balancing of network-related resources, and / or other network functions), and / or other security and / or network-related functions. For example, a routing function can be based on source information (e.g., IP address and port), destination information (e.g., IP address and port), and protocol information.
[0009] A basic packet filtering firewall filters network communication traffic by inspecting individual packets transmitted over the network (e.g., a packet filtering firewall or first-generation firewall that is a stateless packet filtering firewall). A stateless packet filtering firewall typically inspects the individual packets themselves and applies rules based on the inspected packets (e.g., using a combination of source and destination address information, protocol information, and port numbers of the packets).
[0010] An application firewall can also perform application layer filtering (e.g., using an application layer filtering firewall or a second-generation firewall that functions at the application level of the TCP / IP stack). An application layer filtering firewall or application firewall can generally identify a given application and protocol (e.g., web browsing using the Hypertext Transfer Protocol (HTTP), Domain Name System (DNS) requests, file transfers using the File Transfer Protocol (FTP), and various other types of applications and other protocols such as Telnet, DHCP, TCP, UDP, and TFTP (GSS)). For example, an application firewall can block unauthorized protocols that attempt to communicate on standard ports (e.g., unauthorized / rogue protocols that attempt to sneak through by using non-standard ports for that protocol can generally be identified using an application firewall).
[0011] A stateful firewall can also perform stateful - based packet inspection, where each packet is inspected within a set of packet contexts associated with the packet flow of its network transmission. This firewall technology is generally referred to as stateful packet inspection. This is because it keeps a record of all connections passing through the firewall and can determine whether a packet is the start of a new connection, part of an existing connection, or an invalid packet. For example, the state of a connection can itself be one of the criteria that trigger rules in the policy.
[0012] Advanced or next - generation firewalls, as described above, can perform stateless and stateful packet filtering and application - layer filtering. Next - generation firewalls can also perform additional firewall technologies. For example, a certain new firewall, sometimes referred to as an advanced or next - generation firewall, can also identify users and content. In particular, certain next - generation firewalls have expanded the list of applications that these firewalls can automatically identify to thousands of applications. Examples of such next - generation firewalls are commercially available from Palo Alto Networks (e.g., Palo Alto Networks' PA series next - generation firewalls, Palo Alto Networks' VM series virtualized next - generation firewalls, and CN series container next - generation firewalls).
[0013] For example, Palo Alto Networks' next-generation firewall uses various identification technologies to enable enterprises and service providers to identify and control applications, users, and content - not just ports, IP addresses, and packets. The various identification technologies include Application ID (App-ID) for accurate application identification, User ID (User-ID) for user identification (e.g., User ID), Content ID (Content-ID) for real-time content scanning (e.g., to control web surfing and restrict data and file transfers), and Device ID (Device-ID) (e.g., for identifying IoT device types). With these identification technologies, enterprises can use business-related concepts to safely enable the use of applications instead of following the traditional approach provided by conventional port-blocking firewalls. Also, the dedicated hardware for next-generation firewalls generally provides a higher performance level for application inspection than software running on general-purpose hardware (e.g., security devices provided by Palo Alto Networks, which utilize dedicated, function-specific processing tightly integrated with a single-pass software engine to minimize latency and maximize network throughput for Palo Alto Networks' PA series next-generation firewalls).
[0014] Overview of Techniques for Application Identification for Phishing Detection
[0015] Phishing is a growing security threat, with approximately 1.5 million new phishing sites identified each month. In particular, since the start of the Covid-19 pandemic in late 2019 (resulting in a large number of companies enabling employees and contractors to work from home (remotely)), a significant increase in phishing attacks has been witnessed.
[0016] Existing approaches to detecting phishing attacks have various drawbacks. For example, signature detection approaches based on pattern matching generally tend to produce false positives (FP) due to the high similarity between the original web page of the target site and the phishing page that attempts to emulate / mimic that target site (e.g., phishing pages are typically created to mimic the original web page, such as the login pages of banking sites, e-commerce sites, streaming sites, etc.). As another example, URL-based detection approaches generally require an offline analyzer to observe exactly the same page as the customer (e.g., in the case of using various cloaking techniques based on geolocation targets such as when a customer's corporate firewall, but only for location targets in Germany, receives phishing pages as opposed to customers in other parts of Europe, and often not so. And when the security service provider has servers located in a different geolocation from the target location, it does not result in phishing sites being presented for security analysis by a cloud-based security platform of the security service provider that can perform URL or other security analysis).
[0017] Therefore, new and improved techniques for detecting phishing are needed.
[0018] Thus, various techniques for application identification (App-ID) for phishing detection are disclosed. For example, the disclosed techniques, as described herein, in combination with various other techniques, promote the use of advanced application identification (App-ID) (e.g., as used herein, advanced application identification (App-ID) includes (1) protocol identification (e.g., HTTP, HTTPS, SSL, TLS, etc.), and (2) website identification based on analysis of the content of web pages, generally referred to herein as site similarity) to provide more efficient and effective phishing detection (e.g., increased coverage and lower FP rate).
[0019] In some embodiments, a system / process / computer program product for application identification for phishing detection includes monitoring network activity associated with a session to detect requests to access a site, determining an advanced application identification associated with the site, and identifying the site as a phishing site based on the advanced application identification.
[0020] For example, the disclosed techniques for application identification for phishing detection can be performed to facilitate more effective and efficient in-line detection of phishing sites and blocking (e.g., in a security platform such as a perimeter firewall). The disclosed techniques also promote an improvement in the phishing detection rate compared to existing stand-alone approaches. Further, the disclosed techniques result in a lower FP rate using such combinations of application identification and detection techniques (e.g., as opposed to existing approaches that simply perform only analysis of web page content), as further described below.
[0021] Accordingly, a new and improved security solution that facilitates applying application identification for phishing detection is disclosed in some embodiments using a security platform (e.g., a firewall (FW) / next-generation firewall (NGFW), a network sensor operating in place of a firewall, or another (virtual) device / component that can implement security policies using the disclosed techniques, including, for example, Palo Alto Networks' PA series next-generation firewalls, Palo Alto Networks' VM series virtualized next-generation firewalls, and CN series container next-generation firewalls, and / or other commercially available virtual-based or container-based firewalls that can be similarly implemented and configured to execute the disclosed techniques).
[0022] These and other embodiments and examples for applying application identification for phishing detection are further described below.
[0023] Exemplary System Architecture for Application Identification for Phishing Detection
[0024] Accordingly, in some embodiments, the disclosed techniques can be implemented using another (virtualized) device / component that can be implemented using the disclosed techniques on a security platform (e.g., the security function / platform can be a firewall (FW) / next-generation firewall (NGFW), a network sensor operating in place of a firewall, or a virtual / physical NGFW solution commercially available from Palo Alto Networks or another security platform / NGFW, such as the PA series next-generation firewall from Palo Alto Networks, the VM series virtualized next-generation firewall from Palo Alto Networks, and the CN series container next-generation firewall, and PANOS running on it), and / or other commercially available virtualization-based or container-based firewalls can be similarly implemented and configured to execute the disclosed techniques), and the security platform is configured to provide a DPI capability (including, for example, stateful inspection) that applies App-ID for phishing detection based on policies (e.g., layer 7 security and / or other security policy enforcement), as further described below.
[0025] FIG. 1 shows one exemplary environment in which malicious applications (“malware”) are detected and prevented from causing harm. As described in more detail below, malware classification can be shared and / or refined among various entities included in the environment shown in FIG. 1 (e.g., as determined by security platform 122), and devices such as endpoint client devices 104-110 can be protected from such malware (including, for example, phishing-related malware) using the techniques described herein.
[0026] As used herein, "malware" refers to an application that engages in behavior that the user would not approve of / approve of if fully informed, whether or not it is secret (and whether or not it is illegal). Examples of malware include Trojan horses, viruses, rootkits, spyware, hacking tools, keyloggers, etc. One example of malware is a desktop application that encrypts the user's stored data (e.g., ransomware). Another example of malware is a desktop application that collects the end user's activities and / or various information associated with the user and reports it to a remote server (e.g., spyware). Other forms of malware can also be detected / blocked using improved phishing detection techniques (e.g., keyloggers) as further described herein.
[0027] The term “phishing” is used throughout this specification and collectively refers to things such as electronic mail (e-mail) messages, text messages, and / or various other types of messages, social or work productivity related messaging platforms, etc. Here, the attacker sends a fraudulent message (e.g., a spoofed, fake, or otherwise deceptive content and / or web link) designed to trick the user / human into performing an action (e.g., revealing confidential information such as login information, or other personal / confidential information or certificates, or visiting a spoofed website, or facilitating the download of malware such as ransomware by deploying malware on the user's infrastructure, etc.) that promotes malicious and / or unwanted activities. As further explained herein, the detection of phishing activities is becoming an increasingly technical challenge. Phishing attacks are becoming more sophisticated and often involve transparently mirroring the targeted site (e.g., a website) so that the attacker can observe everything while the victim is navigating the site and cross any additional security boundaries with the victim.
[0028] The techniques described in this specification can be used with various platforms (e.g., desktops, mobile devices, game platforms, embedded systems, etc.) and / or various forms of phishing attacks (e.g., via social or work productivity related messaging platforms, etc. through email (e-mail) messages, text messages, and / or various other types of messages). In the exemplary environment shown in FIG. 1, client devices 104 - 108 are laptop computers, desktop computers, and tablets (respectively) residing within enterprise network 140. Client device 110 is a laptop computer residing outside enterprise network 140.
[0029] Data appliance 102 is configured to enforce policies regarding communications between client devices, such as client devices 104 and 106, and nodes external to enterprise network 140 (e.g., those reachable via external network 118). Examples of such policies include those that manage traffic shaping, quality of service, and traffic routing. Other examples of policies include security policies that require scanning for threats in incoming (and / or outgoing) email attachment files, website content, files exchanged via instant messaging programs, and / or other file transfers. In some embodiments, data appliance 102 is also configured to enforce policies regarding traffic remaining within enterprise network 140.
[0030] One embodiment of a data appliance is shown in FIG. 2A. The example shown is a representation of the physical components included in data appliance 102 in various embodiments. Specifically, data appliance 102 includes a high-performance multi-core central processing unit (CPU) 202 and random access memory (RAM) 204. Data appliance 102 also includes storage 210 (such as one or more hard disks or solid state units). In various embodiments, data appliance 102 is used to monitor enterprise network 110 and store information (in either RAM 204, storage 210, and / or other appropriate locations) used to implement the disclosed techniques. Examples of such information include application identifiers, content identifiers, user identifiers, requested URLs, IP address mappings, policies and other configuration information, signatures, host name / URL classification information, malware profiles, machine learning models, IoT device classification information, etc. Data appliance 102 may also include one or more optional hardware accelerators. For example, data appliance 102 may include an encryption engine 206 configured to perform encryption and decryption operations, and one or more field programmable gate arrays 208 configured to perform matching, operate as a network processor, and / or perform other tasks.
[0031] The functionality described herein as being performed by data appliance 102 can be provided / implemented in a variety of ways. For example, data appliance 102 can be a dedicated device or set of devices. The functionality provided by data appliance 102 can also be integrated or executed as software on a general-purpose computer, computer server, gateway, and / or network / routing device. In some embodiments, at least some of the services described as being provided by data appliance 102 are instead (or in addition to) provided to client devices (e.g., client device 104 or client device 110) by software running on a client device (e.g., endpoint protection application 132).
[0032] Whenever the data appliance 102 is described as performing a task, a single component, a subset of components, or all components of the data appliance 102 can cooperate to perform the task. Similarly, whenever a component of the data appliance 102 is described as performing a task, a sub-component can perform the task and / or a component can perform the task with other components. In various embodiments, a portion of the data appliance 102 is provided by one or more third parties. Depending on factors such as the amount of computing resources available to the data appliance 102, various logical components and / or features of the data appliance 102 may be omitted and the techniques described herein adapted accordingly. Similarly, additional logical components / features can be included in embodiments of the data appliance 102 as applicable. One example of a component included in the data appliance 102 in various embodiments is an application identification engine configured to identify applications (e.g., using various application signatures to identify applications based on packet flow analysis). For example, the application identification engine can determine the type of traffic a session is involved in. Such as web browsing - social networking, web browsing - news, SSH, etc. Further, as described herein, the application identification engine disclosed herein can perform advanced application identification. Specifically, advanced application identification includes both identifying an application (e.g., using various application signatures to identify an application based on packet flow analysis) and identifying a target site based on site similarity analysis (e.g., site similarity analysis of a target site and a potential phishing site as further described herein).
[0033] Figure 2B is a functional diagram of logical components according to one embodiment of the data appliance. The example shown is a representation of logical components that may be included in the data appliance 102 in various embodiments. Unless otherwise specified, the various logical components of the data appliance 102 can generally be implemented in various ways, including a set of one or more scripts (e.g., written in Java®, python, etc. if applicable).
[0034] As shown, the data appliance 102 includes a firewall and includes a management plane 232 and a data plane 234. The management plane is responsible for managing user interaction, such as by providing a user interface for policy setting and viewing log data. The data plane is responsible for data management, such as by performing packet processing and session processing.
[0035] The network processor 236 is configured to receive packets from a client device, such as the client device 108, and provide them to the data plane 234 for processing. The flow module 238 generates a new session flow whenever it identifies a packet as part of a new session. Subsequent packets are identified as belonging to the session based on a flow lookup. When applicable, SSL decryption is applied by the SSL decryption engine 240. Otherwise, the processing by the SSL decryption engine 240 is omitted. The decryption engine 240 helps the data appliance 102 inspect and control SSL / TLS and SSH encrypted traffic and thus helps stop threats that might otherwise remain hidden within the encrypted traffic. The decryption engine 240 can also help prevent highly confidential content from leaving the enterprise network 140. Decryption can be selectively controlled (e.g., enabled or disabled) based on parameters such as URL category, traffic source, traffic destination, user, user group, and port. In addition to a decryption policy (e.g., one that specifies the sessions to decrypt), a decryption profile can be assigned to control various options for the sessions controlled by the policy. For example, the use of a particular cipher suite and encryption protocol version can be required.
[0036] The Application Identification (APP-ID) engine 242 is configured to determine the type of traffic involved in a session. As one example, the Application Identification engine 242 can recognize a GET request within received data and conclude that the session requires an HTTP decoder. In some cases, the identified application, such as a web browsing session, can change, and such changes are noted by the data appliance 102. For example, a user can first browse a corporate Wiki (classified as "Web Browsing-Productivity" based on the visited URL) and then browse a social networking site (classified as "Web Browsing-Social Networking" based on the visited URL). Different types of protocols have corresponding decoders. In addition, the Application Identification engine disclosed herein can perform advanced application identification. Specifically, advanced application identification includes both identifying an application (e.g., using various application signatures to identify an application based on packet flow analysis as described above) and identifying a target site based on site similarity analysis (e.g., site similarity analysis of a target site and potential phishing sites as further described below).
[0037] Based on the decision made by the application identification engine 242, the packet is sent to an appropriate decoder configured by the threat engine 244 to assemble the packets (which may be received out of order) into the correct order, perform tokenization, and extract information. The threat engine 244 also performs signature matching to determine what should happen to the packet. Optionally, the SSL encryption engine 246 can re-encrypt the decrypted data. The packet is forwarded for transfer (e.g., to the destination) using the forwarding module 248.
[0038] Also, as shown in Figure 2B, the policy 252 is also received and stored in the management plane 232. The policy can include one or more rules that can be specified using domain names and / or host / server names, and the rules can apply one or more signatures or other matching criteria or discovery methods, such as for the implementation of security policies for subscribers / IP flows, based on various extracted parameters / information from the monitored session traffic flow. Exemplary policies can include phishing detection policies that use the disclosed advanced application identification techniques in combination with one or more other parameters (e.g., IP addresses associated with the site, geolocation information associated with the site / IP address, and / or other information as further described below). An interface (I / F) communicator 250 is provided for management communication (e.g., via (REST) API, messages, or network protocol communication, or other communication mechanisms).
[0039] Security Platform
[0040] Returning to FIG. 1, assume that a malicious individual created malware 130 (using system 120) (e.g., distributed to a user's endpoint device via a phishing site, where the URL of the phishing site is sent to the target user within the content of an email). The malicious individual desires that a client device, such as client device 104, execute a copy of the malware 130, exposing the client device to risk and, for example, turning the client device into a bot within a botnet. The exposed client device is then instructed to perform a task (e.g., participate in cryptocurrency mining or a denial of service attack) and report information to an external entity, such as command and control (C&C) server 150, and, if applicable, receive instructions from the C&C server 150.
[0041] Assume that data appliance 102 intercepts an email sent to user “Alice” who operates client device 104 (e.g., by system 120). In this example, Alice receives the email and clicks on a link to a phishing site, which may result in an attempt to download malware 130 by Alice's client device 104. However, in this example, data appliance 102 performs application identification disclosed for phishing and blocks access to the phishing site from Alice's client device 104, thereby preempting and preventing any such download of malware 130 to Alice's client device 104. As further explained below, data appliance 102 performs advanced application identification and uses additional information associated with the target site (e.g., an IP address associated with the site, geolocation information associated with the site / IP address, and / or other information as further explained below) to detect and block such phishing attempts.
[0042] In various embodiments, data appliance 102 is configured to operate in cooperation with security platform 122. As one example, security platform 122 can provide a set of signatures of known malicious files to data appliance 102 (e.g., as part of a subscription). If the signature of malware 130 is included in the set (e.g., the MD5 hash of malware 130), data appliance 102 can accordingly prevent the transmission of malware 130 to client device 104 (e.g., by detecting that the MD5 hash of an email attachment file sent to client device 104 matches the MD5 hash of malware 130). Security platform 122 can also provide a list of known malicious domains and / or IP addresses to data appliance 102, enabling data appliance 102 to block traffic between enterprise network 140 and C&C server 150 (e.g., where C&C server 150 is known to be malicious). The list of malicious domains (and / or IP addresses) can also help data appliance 102 determine when one of its nodes has been exposed to danger. For example, if client device 104 attempts to contact C&C server 150, such an attempt is a strong indicator that client 104 has been exposed to malware (and remedial action, such as refraining from client device 104 communicating with other nodes within enterprise network 140, should be taken accordingly). As will be described in more detail below, security platform 122 can also provide other types of information to data appliance 102 (e.g., as part of a subscription). A set of information (e.g., target site information for performing site similarity, IP addresses associated with the site, geolocation information associated with the site / IP address, and / or other information as further described below) for performing advanced application identification for phishing, which is available for use by data appliance 102 for performing in-line analysis of files.
[0043] In various embodiments, when a signature of an attached file is not found, various actions can be taken by data appliance 102. As a first example, data appliance 102 can be fail-safe by blocking the transmission of any attached file that is not listed on the whitelist as benign (e.g., does not match the signature of a known good file). The drawback of this approach is that many legitimate attached files may be unnecessarily blocked as potential malware when they are actually benign. As a second example, data appliance 102 can be fail-danger by allowing the transmission of any attached file that is not listed on the blacklist as malicious (e.g., does not match the signature of a known bad file). The drawback of this approach is that newly created malware (previously not seen by platform 122) may cause harm and not be prevented. As a third example, data appliance 102 can be configured to provide a file (e.g., malware 130) to security platform 122 for static / dynamic analysis to determine whether it is malicious and / or, if not, classify it.
[0044] The security platform 122 stores a copy of the received sample in the storage 142, and then analysis is initiated (or scheduled if applicable). One example of the storage 142 is an Apache Hadoop Cluster (HDFS). The results of the analysis (and additional information about the application) are stored in the database 146. If the application is determined to be malicious, the data appliance can be configured to automatically block file downloads based on the analysis results. Further, a signature for the malware is generated and distributed (for example, to data appliances such as data appliances 102, 136, and 148), and future file transfer requests to download files determined to be malicious can be automatically blocked.
[0045] In various embodiments, the security platform 122 comprises one or more dedicated, commercially available hardware servers (e.g., having a multi-core processor, 32G+ of RAM, a gigabit network interface adapter, and a hard drive) running a typical server-class operating system (e.g., Linux®). The security platform 122 can be implemented across a scalable infrastructure that includes multiple such servers, solid state drives, and / or other applicable high-performance hardware. The security platform 122 can comprise several distributed components, including components provided by one or more third parties. For example, some or all of the security platform 122 can be implemented using Amazon Elastic Compute Cloud (EC2) and / or Amazon Simple Storage Service (S3). Further, similar to the data appliance 102, whenever the security platform 122 is referred to as performing tasks such as storing data or processing data, it should be understood that sub-components or multiple sub-components of the security platform 122 can cooperate (either individually or in cooperation with third-party components) to perform that task. As one example, the security platform 122 can optionally perform static / dynamic analysis in cooperation with one or more virtual machine (VM) servers, such as the VM server 124.
[0046] One example of a virtual machine server is a commercial server-class hardware (e.g., a multi-core processor, 32+ gigabytes of RAM, and one or more Gigabit network interface adapters) that runs commercial virtualization software such as VMware ESXi, Citrix XenServer, or Microsoft Hyper-V. In some embodiments, the virtual machine server is omitted. Further, the virtual machine server may be under the control of the same entity that manages the security platform 122, or may be provided by a third party. As one example, the virtual machine server can rely on EC2, and the rest of the security platform 122 is owned by the operator of the security platform 122 and provided by dedicated hardware under the control of that operator. The VM server 124 is configured to provide one or more virtual machines 126-128 for emulating client devices. The virtual machines can run various operating systems and / or versions thereof. The observed behavior resulting from running an application within a virtual machine is logged and analyzed (e.g., for indicators that the application is malicious). In some embodiments, the log analysis is performed by the VM server (e.g., VM server 124). In other embodiments, the analysis is performed at least in part by other components of the security platform 122, such as the coordinator 144.
[0047] In various embodiments, the security platform 122 makes the analysis results of samples available to the data appliance 102 as part of a subscription via a list of signatures (and / or other identifiers). For example, the security platform 122 can periodically (e.g., daily, hourly, or at some other interval and / or based on events configured by one or more policies) send a content package that identifies malware files, phishing sites, etc. One exemplary content package includes a list of identified phishing sites with information such as target site names, URLs, and site similarity information, as well as various other information about each target site (e.g., IP addresses associated with the target site, geolocation information associated with the site / IP address, and / or other information as further described below). The subscription can be intercepted by the data appliance 102 and cover the analysis of exactly those files that were sent by the data appliance 102 to the security platform 122, and can also cover malware signatures known to the security platform 122. As will be described in more detail below, the platform 122 can also use the site model 152 to make available other types of information for phishing detection, which is performed using the disclosed advanced application identification techniques that can help the data appliance 102 detect phishing sites and perform in-line blocking of phishing sites.
[0048] In various embodiments, security platform 122 is configured to provide security services to various entities in addition to (or, where applicable, instead of) the operator of data appliance 102. For example, other enterprises having respective enterprise networks 114 and 116, and respective data appliances 136 and 148, can contract with the operator of security platform 122. Other types of entities can also utilize the services of security platform 122. For example, an Internet service provider (ISP) providing Internet services to client device 110 can contract with security platform 122 to analyze an application that client device 110 attempts to download. As another example, the owner of client device 110 can install software that communicates with security platform 122 on client device 110 (e.g., receive a content package from security platform 122 and use the received content package to check attached files according to the techniques described herein, and send the application to security platform 122 for analysis).
[0049] Analysis of Samples Using Static / Dynamic Analysis
[0050] Figure 3 shows one exemplary logical component that may be included in a system for analyzing samples. Analysis system 300 can be implemented using a single device. For example, the functionality of analysis system 300 can be implemented in malware analysis module 112 incorporated in data appliance 102. Analysis system 300 can also be implemented collectively across a plurality of separate devices. For example, the functionality of analysis system 300 can be provided by security platform 122.
[0051] In various embodiments, the analysis system 300 utilizes a list, database, or other collection of known safe content and / or known bad content (collectively shown as collection 314 in FIG. 3). The collection 314 can be obtained in various ways, including via a subscription service (e.g., provided by a third party) and / or as a result of other processes (e.g., performed by data appliance 102 and / or security platform 122). Examples of information included in collection 314 are as follows. That is, URLs, domain names, and / or IP addresses of known malicious servers, URLs, domain names, and / or IP addresses of known safe servers, URLs, domain names, and / or IP addresses of known command and control (C&C) domains, signatures, hashes, and / or other identifiers of known malicious applications, signatures, hashes, and / or other identifiers of known safe applications, signatures, hashes, and / or other identifiers of known malicious files (e.g., OS exploit files), signatures, hashes, and / or other identifiers of known safe libraries, and signatures, hashes, and / or other identifiers of known malicious libraries.
[0052] In various embodiments, when a new sample is received for analysis (e.g., an existing signature associated with the sample does not exist in the analysis system 300), it is added to queue 302. As shown in FIG. 3, application 130 is received by system 300 and added to queue 302.
[0053] Coordinator 304 monitors queue 302, and when a resource (e.g., a static analysis worker) becomes available, Coordinator 304 fetches a sample from queue 302 for processing (e.g., fetches a copy of malware 130). In particular, the coordinator first provides a sample to static analysis engine 306 for static analysis (305). In some embodiments, one or more static analysis engines are included within analysis system 300, where analysis system 300 is a single device. In other embodiments, static analysis is performed by a separate static analysis server that includes multiple workers (i.e., multiple instances of static analysis engine 306).
[0054] The static analysis engine obtains general information about the sample and includes it in the static analysis report 308 (along with heuristic information and other information where applicable). The report can be created by the static analysis engine or by a coordinator 304 configured to receive information from the static analysis engine 306 (or by another appropriate component). As one example, the static analysis of a target site can include site information for performing a site similarity analysis to perform disclosed techniques for phishing detection. In some embodiments, the information collected is stored in a database record of the sample (e.g., within database 316) instead of or in addition to the separate static analysis report 308 being created (i.e., portions of the database record form the report 308). In some embodiments, the static analysis engine also forms a verdict regarding the application (e.g., "safe", "suspicious", or "malicious"). As one example, if there is even one "malicious" static characteristic within the application (e.g., the application includes a hard link to a known malicious domain), the verdict can be "malicious". As another example, points can be assigned to each of the characteristics (e.g., based on severity if found, based on how reliable the characteristic is for predicting malice, etc.). And based on the number of points associated with the static analysis results, a verdict can be assigned by the static analysis engine 306 (or by the coordinator 304 where applicable).
[0055] Once the static analysis is complete, the coordinator 304 locates an available dynamic analysis engine 310 to perform dynamic analysis on the application. Similar to the static analysis engine 306, the analysis system 300 can directly include one or more dynamic analysis engines. In other embodiments, the dynamic analysis is performed by a separate dynamic analysis server that includes a plurality of workers (i.e., multiple instances of the dynamic analysis engine 310).
[0056] Each dynamic analysis worker manages a virtual machine instance. In some embodiments, the results of the static analysis (e.g., performed by the static analysis engine 306), whether in report form (308) and / or stored in the database 316 or otherwise stored, are provided as input to the dynamic analysis engine 310. For example, the static report information can be used to help select / customize the virtual machine instance (e.g., Microsoft Windows 7 SP 2 vs. Microsoft Windows 10 Enterprise, or iOS 11.0 vs. iOS 12.0) used by the dynamic analysis engine 310. If multiple virtual machine instances are run simultaneously, a single dynamic analysis engine can manage all of the instances, or multiple dynamic analysis engines can be used, if applicable (e.g., each managing its own virtual machine instance). As described in more detail below, during the dynamic part of the analysis, actions (including network activities) performed by the application are analyzed.
[0057] In various embodiments, static analysis of the sample is either omitted or, if applicable, performed by a separate entity. As one example, conventional static and / or dynamic analysis may be performed on the file by a first entity. Once a given file is determined to be malicious (e.g., by the first entity), the file may be provided to a second entity (e.g., the operator of the security platform 122) for additional analysis regarding the use of malware for network activity (e.g., by the dynamic analysis engine 310).
[0058] The environment used by the analysis system 300 is instrumented / hooked such that the behavior observed while the applications are running is logged (e.g., using a customized kernel that supports hooking and logcat) as they occur. Network traffic associated with the emulator is also captured (e.g., using pcap). The log / network data can be stored as temporary files on the analysis system 300 and also, more persistently (e.g., using HDFS or another suitable storage technology, or a combination of technologies such as MongoDB). The dynamic analysis engine (or another suitable component) can compare the connections made by the sample against a list of domains, IP addresses, etc. (314) and determine whether the sample communicated (or attempted to communicate) with a malicious entity.
[0059] Similar to the static analysis engine, the dynamic analysis engine stores the results of its analysis in a record associated with the application being tested in database 316 (and / or, if applicable, includes the results in report 312). In some embodiments, the dynamic analysis engine also forms a determination regarding the application (e.g., "safe", "suspicious", or "malicious"). As one example, even if one "malicious" action is taken by the application (e.g., an attempt is made to contact a known malicious domain, or an attempt to extract sensitive information is observed), the determination can still be "malicious". As another example, points can be assigned to the actions taken (e.g., based on the severity if found, based on how reliable the action is for predicting malice, etc.). And based on the number of points associated with the dynamic analysis results, a determination can be assigned by the dynamic analysis engine 310 (or, if applicable, the coordinator 304). In some embodiments, the final determination associated with the sample is made based on a combination of reports 308 and 312 (e.g., by the coordinator 304).
[0060] Advanced Application Identification for Phishing Detection
[0061] FIG. 4 shows a portion of one exemplary embodiment of a threat detection engine that uses application identification, according to some embodiments. As described above, in various embodiments, data appliance 102 includes a threat engine 244. The threat engine includes an Application Identification (App-ID) engine 402 that performs advanced application identification. Thus, App-ID engine 402 incorporates both protocol decoding for application identification and site similarity matching 406 during each decoder stage and the pattern matching stage that is performed inline at data appliance 102. The results of the two stages are merged by a detector stage, as shown at 410. In some embodiments, detector 410 also utilizes IP address information and / or other characteristics, as shown at 408, in combination with advanced application identification to detect phishing sites (e.g., combining advanced application identification with IP, URL, and / or domain anomaly detection characteristics, as further described below).
[0062] When data appliance 102 receives a packet, data appliance 102 performs a session match to determine to which session the packet belongs (enabling data appliance 102 to support concurrent sessions). Each session has a session state associated with a particular protocol decoder (e.g., a web browsing decoder, an FTP decoder, or an SMTP decoder). When a file is sent as part of a session, the applicable protocol decoder can utilize an appropriate file-specific decoder (e.g., a PE file decoder, a JavaScript® decoder, or a PDF decoder).
[0063] Site similarity is analyzed based on periodic static analysis of the target site to cache the visual representation of the expected target site for subsequent pattern matching comparison analysis with potential phishing site candidates. In one exemplary implementation, site similarity analysis is performed by abstracting web page code that can include, for example, title, header, footer, form elements, copyright information, and / or other fields and information.
[0064] In some embodiments, the disclosed techniques for phishing detection implemented using the threat detection engine shown in FIG. 4 include combining application identification with IP, URL, and / or domain anomaly detection features.
[0065] As a first example, phishing detection can be performed using application identification in combination with pattern matching (e.g., site similarity pattern matching as described herein). For example, advanced application identification includes performing application identification in combination with site similarity to improve phishing detection, which can reduce false positives compared to phishing detection techniques that rely only on pattern matching of site similarity. Thus, the disclosed advanced application identification techniques for phishing detection include using application identification for identifying the hosting server as a prefilter for pattern-based site similarity detection (e.g., the prefilter can be used to reduce the load for performing pattern matching. It promotes better performance for implementing the disclosed techniques).
[0066] As a second example, phishing detection can be performed using advanced application identification in combination with an IP address range associated with a target site (e.g., a legitimate / verified Amazon Web Services (AWS) site). Many web services have servers at a fixed range of IP addresses. Thus, when advanced application identification identifies that a web page of a potential phishing site is similar to a target site (e.g., a legitimate / verified Amazon Web Services (AWS) site, the sign-in of the web page), based on site similarity detection, but these servers associated with the phishing site candidate have IP addresses outside the known IP address range associated with the target site, the threat detection engine can identify the potential phishing site as highly likely to be phishing.
[0067] As a third example, phishing detection can be performed using advanced application identification in combination with a URL category associated with a target site (e.g., a legitimate / verified Amazon Web Services (AWS) site). Generally, it has been observed that many phishing sites frequently change their domains. Thus, for web page content that appears to be similar to well-known sites (e.g., various well-known sites such as the top 100, or top 1000, or top 10000, etc., and such AWS sites that can be regularly monitored to generate site similarity information for performing such site similarity analysis), however, when a potential phishing site is associated with a newly registered domain category (e.g., various URL-related services that can be publicly and / or commercially available can provide information regarding domain name registration date information), the threat detection engine can identify the potential phishing site as highly likely to be phishing.
[0068] Exemplary Use Cases of Application Identification for Phishing Detection
[0069] Figures 5A-5C illustrate exemplary phishing sites that can be detected using application identification for phishing detection, according to some embodiments. For example, the techniques disclosed for application identification for phishing detection may be implemented to be executed using inline malware detection, such as on data appliance 102 (e.g., and / or, using a security agent executed on a protected endpoint, on endpoints such as client devices 104, 106, and 108, and also, as described herein, may be executed as a cloud-based security service, such as using security platform 122).
[0070] Referring to FIG. 5A, a web page of an exemplary Wells Fargo Bank phishing site is shown. However, the disclosed technology can perform advanced application identification to identify such exemplary phishing sites. In this case, the phishing site candidate exists at hXXps: / / storage[.]googleapis[.]com / awells-putlogs-308643420 / index[.]html. Thus, traffic directed to this address (e.g., Uniform Resource Identifier (URI) / Uniform Resource Locator (URL)) is identified as "google-cloud-storage" using the disclosed advanced application identification technique (e.g., implemented by the protocol ID component 404 of the App-ID engine 402 as shown in FIG. 4). Additionally, the disclosed advanced application identification technology for the hosting App-ID or storage App-ID (e.g., implemented by the site similarity component 406 of the App-ID engine 402 as shown in FIG. 4) is used, and also the disclosed site similarity based on pattern matching to detect a web page that appears to be similar to the Wells Fargo Bank login page is used by the threat detection engine (e.g., threat detection engine 244 of FIG. 4) to identify such well-known sites that are not legally placed on the hosting App-ID or storage App-ID as phishing sites. Therefore, it can be used by the detector 410 of the threat detection engine 244 to efficiently and effectively detect this phishing site (e.g., and similar such phishing sites).
[0071] Referring to FIG. 5B, a web page of one exemplary AWS phishing site is shown. However, the disclosed techniques can perform advanced application identification in combination with an IP address range verification function to detect examples of such phishing sites. Thus, when applying advanced application identification in combination with IP range verification to a phishing site such as https: / / howitfix[.]com / app / aws / , the IP associated with this exemplary phishing site does not resolve to a known AWS-related IP range (e.g., this phishing site is not even owned by Amazon, and AWS has a given set of IP ranges publicly available at https: / / ip-ranges.ip-ranges.com / ip-ranges.json, which are updated periodically and can be stored in the IP / other feature component 408 of the threat detection engine 244). Therefore, identifying a potential phishing site as something that appears similar to an AWS site and then checking its IP address range can be used by the detector 410 of the threat detection engine 244 to efficiently and effectively detect this phishing site (e.g., and similar such phishing sites).
[0072] Referring to FIG. 5C, a web page of one exemplary Netflix phishing site is shown. However, the disclosed technology can perform advanced application identification in combination with other features to determine whether a potential phishing site is a newly registered domain. In this example, the domain of the phishing site candidate, Canada-neflxt[.]com, is a recently registered domain (e.g., less than 1 day, less than 1 week, less than 1 month, or some other threshold for recently registered domains can be used similarly for this feature). In this case, the threat engine identifies the potential phishing site as being similar to a well-known site, Netflix. However, as in this example, if the URL category indicates this site as a newly registered domain (NRD) (e.g., NRD information is updated periodically and can be stored in the IP / other feature component 408 of the threat detection engine 244), then the detector 410 of the threat detection engine 244 can use this additional NRD attribute to efficiently and effectively detect this phishing site (e.g., and similar such phishing sites).
[0073] Additional exemplary processes for the application identification for phishing detection disclosed herein will now be described.
[0074] Exemplary Process for Application Identification for Phishing Detection
[0075] FIG. 6 is a flowchart relating to a process for application identification for phishing detection according to some embodiments. In some embodiments, process 600 shown in FIG. 6 is executed by the security platform and techniques described above, including the embodiments described above with respect to FIGS. 1-5C. In one embodiment, process 600 includes data appliance 102 described above with respect to FIG. 1, security platform 122 described above with respect to FIG. 1 (e.g., as a cloud-based security service), virtual appliances (e.g., Palo Alto Networks' VM series virtual next-generation firewall, CN series container next-generation firewall, and / or other commercially available virtual-based or container-based firewalls that can be similarly implemented and configured to execute the disclosed techniques), SDN security solutions, cloud security services, and / or combinations or hybrid implementations of the foregoing described herein.
[0076] At 602, monitoring network activity associated with a session is performed to detect requests to access a site. For example, data appliance 102 can be configured to monitor sessions and detect requests to access a site, similar to that described above with respect to FIGS. 1-2B.
[0077] At 604, determining an advanced application identification associated with the site is performed. For example, the advanced application identification can be performed using protocol ID component 404 and site similarity component 406 of threat detection engine 244, as similarly described above with respect to FIGS. 4 and 5A.
[0078] At 606, based on advanced application identification, the site is identified as a phishing site. For example, the implementation of the policy may include blocking a session from accessing a phishing site, logging an attempt to access a phishing site, monitoring and logging access to a phishing site, and warning the user before allowing access to a phishing site. And / or, other actions / responses, or combinations thereof, may be performed based on a policy (e.g., a phishing / security policy that may be stored in policy 252 as shown in FIG. 2B).
[0079] FIG. 7 is another flowchart relating to a process for application identification for phishing detection according to some embodiments. In some embodiments, the process 700 shown in FIG. 7 is performed by the security platform and techniques described above with respect to FIGS. 1-5C, including the embodiments described above. In one embodiment, the process 700 is performed by the data appliance 102 described above with respect to FIG. 1, the security platform 122 described above with respect to FIG. 1 (e.g., as a cloud-based security service), a virtual appliance (e.g., Palo Alto Networks' VM series virtualized next-generation firewall, CN series container next-generation firewall, other commercially available virtual-based or container-based firewalls may be similarly implemented and configured to perform the disclosed technology), an SDN security solution, a cloud security service, and / or a combination or hybrid implementation of the foregoing described herein.
[0080] At 702, monitoring of network activities associated with a session is performed to detect requests for accessing a site. For example, data appliance 102 may be configured to monitor a session and detect requests for accessing a site, as described above with respect to FIGS. 1-2B.
[0081] At 704, determining advanced application identification associated with a site is performed. For example, advanced application identification may be performed using protocol ID component 404 and site similarity component 406 of threat detection engine 244, as described above similarly with respect to FIGS. 4 and 5A.
[0082] At 706, IP / other characteristics associated with a request for accessing a site are determined. For example, determination of IP / other characteristics (e.g., IP address / range, NRD information, URL category information, etc.) associated with a request for accessing a site may be performed using IP / other characteristics component 408 of threat detection engine 244, as described above similarly with respect to FIGS. 4, 5B, and 5C. For example, advanced application identification can include protocol identification for identifying a phishing site based at least in part on using a URL category, where the URL category is selected using URL categorization based on the extracted domain associated with a request for accessing a site (e.g., URL categorization can be obtained from a firewall using a URL categorization solution / service such as PanDB commercially available from Palo Alto Networks and / or using another commercially available or publicly available URL categorization solution / service).
[0083] At 708, based on advanced application identification and other IP / other characteristics, identifying a site as a phishing site is performed. For example, enforcement of the policy may include blocking a session from accessing a phishing site, logging an attempt to access a phishing site, monitoring and logging access to a phishing site, warning a user before permitting access to a phishing site. And / or, other actions / responses, or combinations thereof, may be performed based on a policy (e.g., a phishing / security policy that may be stored in policy 252 as shown in FIG. 2B).
[0084] The foregoing embodiments have been described in some detail for purposes of clarity of understanding, but the present invention is not limited to the details provided. There are many alternative ways to implement the present invention. The disclosed embodiments are exemplary and not limiting.
Claims
1. A system including a processor and a memory, wherein the processor is configured to monitor network activities associated with a session to detect requests for accessing a site, determine a high - level application identification associated with the site, and identify the site as a phishing site based on the high - level application identification, and the memory is coupled to the processor and configured to provide instructions to the processor, a system.
2. The high - level application identification includes protocol identification, The system according to claim 1.
3. The high - level application identification includes protocol identification for identifying the phishing site based at least in part on a URL category, The URL category is selected using URL categorization based on an extracted domain associated with a request for accessing the site, The system according to claim 1.
4. The high - level application identification includes site similarity identification, The system according to claim 1.
5. The high - level application identification includes site similarity identification and determines whether the request for the site is visually similar to a well - known site, The system according to claim 1.
6. The high - level application identification includes protocol identification and site similarity identification, The system according to claim 1.
7. Detecting the site as a phishing site is performed in - line using a data appliance, The system according to claim 1.
8. The processor is further configured to determine an IP address associated with the request for the site for comparison with an expected IP address range associated with well - known similar sites based on site similarity analysis, The system according to claim 1.
9. The processor is further configured to determine a URL category associated with the request for the site for comparison with an expected URL category associated with well - known similar sites based on site similarity analysis, The system according to claim 1.
10. The processor is further configured to As another feature for detecting the site as a phishing site, determining that the site is a newly registered domain (NRD), The system according to claim 1, configured as such.
11. A method comprising: Monitoring network activities associated with a session, and detecting a request to access a site; Determining an advanced application identification associated with the site; Identifying the site as a phishing site based on the advanced application identification; A method including the above.
12. The advanced application identification includes protocol identification. The method according to claim 11.
13. The advanced application identification includes protocol identification for identifying the phishing site based at least in part on a URL category, The URL category is selected using URL categorization based on an extracted domain associated with a request to access the site. The method according to claim 11.
14. The advanced application identification includes site similarity identification. The method according to claim 11.
15. The advanced application identification includes site similarity identification, and determining whether the request for the site is visually similar to a well-known site. The method according to claim 11.
16. The advanced application identification includes protocol identification and site similarity identification. The method according to claim 11.
17. Detecting the site as a phishing site is performed in-line using a data appliance. The method according to claim 11.
18. A computer program stored in a non-transitory computer-readable storage medium, including a plurality of computer instructions, When the computer instructions are executed, causing the computer to: Monitor network activities associated with a session, and detect a request to access a site; Determine an advanced application identification associated with the site; Identify the site as a phishing site based on the advanced application identification; To perform the above. A computer program.
19. The high-level application identification includes protocol identification, The computer program according to claim 18. [
20. ] The high-level application identification includes site similarity identification, The computer program according to claim 18.
Citation Information
Patent Citations
Using DNS communication to filter domain names
JP2014519751A
Program for warning access to web page, method, and system
JP2016045754A
Credentials enforcement using a firewall
US20180309721A1
Visual Detection of Phishing Websites Via Headless Browser
US20210144174A1
Communication control apparatus
WO2008062542A1