Presumption of Innocence (IUPG), Adversarial-Resistant, and False-Positive-Resistant Deep Learning Models

The Presumption of Innocence (IUPG) framework addresses the vulnerabilities of existing malware detection systems by training deep neural networks to resist adversarial attacks and reduce false positives, enhancing the accuracy of malware classification.

JP7763301B2Active Publication Date: 2025-10-31PALO ALTO NETWORKS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024120874
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-26
Filing Date
2024-07-26
Publication Date
2025-10-31
Estimated Expiration
2041-06-03

AI Technical Summary

Technical Problem

Existing malware detection systems are vulnerable to adversarial attacks and produce high false-positive rates, failing to effectively distinguish between benign and malicious applications.

Method used

The development of adversary-resistant and false-positive-resistant deep learning models, known as the Presumption of Innocence (IUPG) framework, which utilizes a hybrid discriminative and generative loss function to train deep neural networks, learning unique input space prototypes and maximizing the distance between target and off-target classes.

Benefits of technology

The IUPG framework enhances malware detection by reducing false positives and improving the ability to identify adversarial attacks, ensuring accurate classification of applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007763301000075
    Figure 0007763301000075
  • Figure 0007763301000076
    Figure 0007763301000076
  • Figure 0007763301000077
    Figure 0007763301000077
Patent Text Reader

Abstract

To provide a system and method including innocent until proven guilty (IUPG) model for building and using adversary resistant and false positive resistant deep learning models.SOLUTION: A method of performing static analysis of a sample includes: storing a set comprising one or more IUPG models for static analysis of a sample; performing static analysis of a content associated with the sample, performing the static analysis of the model including using at least one stored IUPG model. The method further includes: determining that the sample is malicious on the basis of at least partially, the static analysis of the content associated with the sample; and in response to determining that the sample is malicious, performing an action on the basis of a security policy.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to deep learning models that are adversary resistant and false positive resistant.

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 034,843, filed June 4, 2020, entitled "INNOCENT UNTIL PROVEN GUILTY (IUPG): BILDING FALSE POSITIVE RESISTANT DEEP LEARNING MODELS," which is incorporated herein by reference for all purposes. [Background technology]

[0003] Malware is a general term commonly used to refer to malicious software (e.g., including a variety of hostile, intrusive, and / or other unwanted software). Malware can be in the form of code, scripts, active content, or other software. Examples of malware include disrupting the operation of computers or networks and / or stealing proprietary information (e.g., sensitive information such as personal, financial, and / or intellectual property-related information) and / or accessing private / proprietary computer systems and / or computer networks. Unfortunately, as technology is developed to help detect and mitigate malware, malicious authors find ways to circumvent these efforts. Thus, there is a continuing need for improved techniques to identify and mitigate malware. [Brief explanation of the drawings]

[0004] Various embodiments of the present invention are disclosed in the following detailed description and accompanying drawings. [Figure 1]FIG. 1 shows an example environment in which malicious applications ("malware") are detected and prevented from causing harm. [Figure 2A] FIG. 2A shows one embodiment of a data appliance. [Figure 2B] FIG. 2B is a functional diagram of the logical components of one embodiment of a data device. [Figure 3] FIG. 3 illustrates an example of logical components that may be included in a system for analyzing a sample. [Figure 4] FIG. 4 shows portions of one embodiment of a threat thread engine. [Figure 5A] FIG. 5A is a block diagram of an extended IUPG component in an abstract network N, according to some embodiments. [Figure 5B] FIG. 5B illustrates an example of using an IUPG network, according to some embodiments. [Figure 5C] FIG. 5C shows the desired force for on-target and off-target samples. [Figure 6] FIG. 6 illustrates the function N used in (A) MNIST and Fashion MNIST, (B) JS, and (C) URL, according to some embodiments. [Figure 7A] Figure 7A shows Table 1, including the malicious JS classification test set μ±SEM FNR across all non-benign classes. [Figure 7B] FIG. 7B shows Table 2, which contains the error rates for the non-noisy test set μ±SEM for image classification. [Figure 7C] FIG. 7C shows Table 3, which contains the test set μ±SEM error rates for image classification when the test set contains Gaussian noise images and the model is trained without noise. [Figure 7D] Figure 7D shows Table 4, which contains the test set FNR for malicious URL classification. [Figure 7E] Figure 7E shows Table 5, which contains the detections with a test set FPR threshold set at 0.005%, organized by VTS. [Figure 8] Figure 8 shows the results of the OOD attack simulation. [Figure 9] Figure 9 shows Table 6, which contains the results of the append attack simulation. [Figure 10A] Figure 10A shows the accuracy on correctly classified test images versus the scaling factor of the FGSM perturbations. [Figure 10B] Figure 10B shows the accuracy on correctly classified test images versus the scaling factor of the FGSM perturbations. [Figure 11] FIG. 11 is an example of a t-SNE visualization of U-vector space, according to some embodiments. [Figure 12A] FIG. 12A shows various examples of append attacks that can evade detection by existing malware detection solutions. [Figure 12B] FIG. 12B shows various examples of append attacks that can evade detection by existing malware detection solutions. [Figure 12C] FIG. 12C shows various examples of append attacks that can evade detection by existing malware detection solutions. [Figure 13] FIG. 13 illustrates the IPUG framework for malware JavaScript classification, according to some embodiments. [Figure 14] FIG. 14 is an example process for performing static analysis of a sample using the Presumption of Innocent (IUPG) model for malware classification, according to some embodiments. [Figure 15] FIG. 15 is an example process for generating a presumption of innocence (IUPG) model for malware classification according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0005] The present invention can be implemented in various ways, including as an apparatus, a system, an article of manufacture, a computer program product embodied in a computer-readable storage medium, and / or a processor configured to execute instructions stored in and / or provided by a memory coupled to the processor. These implementations, or other forms the present invention may take, are referred to in this specification as techniques. In general, the order of steps in disclosed processes may be varied within the scope of the present invention. Unless otherwise specified, components, such as a processor or memory, described as being configured to perform a task may be implemented as general components temporarily configured to perform the task at a given time, or as specific components manufactured to perform the task. As used herein, the term “processor” refers to one or more devices, circuits, and / or processing cores configured to process data, such as computer program instructions.

[0006] A detailed description of one or more embodiments of the present invention is provided below along with accompanying figures that illustrate the principles of the invention. While the present invention will be described in connection with such embodiments, the invention is not limited to any embodiment. The scope of the present invention is limited only by the claims, and the present invention encompasses numerous alternatives, modifications, and equivalents. In the following description, numerous specific details are set forth to provide a thorough understanding of the present invention. These details are provided for the purpose of example, and the present invention may be practiced according to the claims without some or all of these specific details. For the purposes of clarity, technical material known in the art related to the present invention has not been described in detail so as not to unnecessarily obscure the present invention.

[0007] Firewall Technology Overview

[0008] A firewall generally protects a network from unauthorized access while allowing authorized communications to pass through the firewall. A firewall is typically a device or set of devices, or software running on a device, that provides firewall functionality for network access. For example, a firewall can be integrated into the operating system of a device (e.g., a computer, smartphone, or other type of network-enabled device). Firewalls can also be integrated into or run as software on computer servers, gateways, network / routing devices (e.g., network routers), and data appliances (e.g., security appliances or other types of special-purpose devices).

[0009] Firewalls typically deny or allow network transmissions based on a set of rules. These sets of rules are often referred to as policies. For example, a firewall can filter inbound traffic by applying a set of rules or policies to prevent unwanted external traffic from reaching a protected device. A firewall can also filter outbound traffic by applying a set of rules or policies (e.g., allow, block, monitor, notify or log, and / or other actions can be specified in a firewall rule or firewall policy, which can be triggered based on various criteria, as described herein). A firewall can also filter local network (e.g., intranet) traffic by similarly applying a set of rules or policies.

[0010] A security device (e.g., a security appliance, a security gateway, a security service, or other security device) can include various security functions (e.g., firewalls, anti-malware, intrusion prevention / detection, data loss prevention (DLP), and / or other security functions), network functions (e.g., routing, quality of service (QoS), workload balancing of network-related resources, and other network functions), and / or other functions. For example, a routing function can be based on source information (e.g., IP addresses and ports), destination information (e.g., IP addresses and ports), and protocol information.

[0011] Basic packet filtering firewalls filter network communication traffic by inspecting individual packets sent over a network (e.g., packet filtering firewalls or first-generation firewalls, which are stateless packet filtering firewalls). Stateless packet filtering firewalls typically inspect the individual packets themselves and then apply rules based on the inspected packets (e.g., using a combination of the packet's source and destination address information, protocol information, and port numbers).

[0012] Application firewalls can also perform application layer filtering (e.g., application layer filtering firewalls or second-generation firewalls that operate at the application level of the TCP / IP stack). Application layer filtering firewalls or application firewalls can generally identify certain applications and protocols (e.g., web browsing using the Hypertext Transfer Protocol (HTTP), Domain Name System (DNS) requests, file transfers using the File Transfer Protocol (FTP), and various other types of applications and other protocols, such as Telnet, DHCP, TCP, UDP, and TFTP (GSS)). For example, an application firewall can block unauthorized protocols that attempt to communicate over standard ports (e.g., unauthorized / out-of-policy protocols trying to sneak in using non-standard ports for that protocol can generally be identified using an application firewall).

[0013] Stateful firewalls can also perform state-based packet inspection, in which each packet is inspected within the context of the set of packets associated with that network transmission's packet flow. This firewall technique is commonly referred to as stateful packet inspection because it keeps a record of all connections passing through the firewall and can determine whether a packet is the start of a new connection, part of an existing connection, or an invalid packet. For example, the state of a connection itself can be one of the criteria that triggers a rule in a policy.

[0014] Advanced or next-generation firewalls can perform stateless and stateful packet filtering and application layer filtering, as described above. Next-generation firewalls can also perform additional firewall techniques. For example, certain newer firewalls, sometimes referred to as advanced or next-generation firewalls, can also identify users and content (e.g., next-generation firewalls). In particular, certain next-generation firewalls have expanded the list of applications they can automatically identify to thousands of applications. Examples of such next-generation firewalls are commercially available from Palo Alto Networks (e.g., the Palo Alto Networks PA Series Firewall). For example, Palo Alto Networks next-generation firewalls enable enterprises to identify and control applications, users, and content—not just ports, IP addresses, and packets—using a variety of identification technologies, such as APP-ID for precise application identification, User-ID for user identification (e.g., user or user group), and Content-ID for real-time content scanning (e.g., controlling web surfing and restricting data and file transfers). These identification technologies allow enterprises to securely enable application usage using business-relevant concepts instead of following the traditional approach offered by traditional port-blocking firewalls. Additionally, special-purpose hardware for next-generation firewalls (such as those implemented as dedicated appliances) generally offers higher performance levels for application inspection than software running on general-purpose hardware (e.g., security appliances from Palo Alto Networks use dedicated, function-specific processing tightly integrated with single-pass software engines to maximize network throughput while minimizing latency).

[0015] Advanced or next-generation firewalls can also be implemented using virtualized firewalls. Examples of such next-generation firewalls are commercially available from Palo Alto Networks (e.g., the Palo Alto Networks VM Series Firewalls, e.g., VMware® ESXi). TM and NSX TM , Citrix® Netscaler SDX TM It supports a variety of commercial virtualization environments, including KVM / OpenStack (Centos / RHEL, Ubuntu®), and Amazon Web Services (AWS). For example, the virtualized firewall can support similar or identical next-generation firewall and advanced threat prevention capabilities available in physical form factor appliances, enabling enterprises to securely run applications in and across their private, public, and hybrid cloud computing environments. Automation capabilities such as VM monitoring, dynamic address groups, and a REST-based API enable enterprises to proactively monitor VM changes and dynamically incorporate that context into security policies, thereby eliminating policy lag that can occur when VMs change.

[0016] Example Environment

[0017] Figure 1 illustrates an example environment in which malicious applications ("malware") are detected and prevented from causing harm. As described in more detail below, malware classifications (e.g., as produced by security platform 122) may be shared and / or refined in various ways among various entities included in the environment illustrated in Figure 1. Techniques described herein can then be used to protect devices, such as endpoint client devices 104-110, from such malware.

[0018] The term "application" is used throughout this specification to collectively refer to programs, program bundles, manifests, packages, etc., regardless of format / platform. An "application" (also referred to herein as a "sample") may be a standalone file (e.g., a calculator application with the file name "calculator.apk" or "calculator.exe"), and may also be an independent component of another application (e.g., a mobile advertising SDK or library embedded within a calculator app).

[0019] As used herein, "malware" refers to an application, whether covert or not (and illegal or not), that performs actions that a user would not / would not approve of if fully informed. Examples of malware include Trojan horses, viruses, rootkits, spyware, hacking tools, keyloggers, etc. One example of malware is a desktop application that collects an end user's location and reports it to a remote server (but does not provide the user with location-based services, such as mapping services). Another example of malware is a malicious Android Application Package (APK) file that appears to the end user to be a free game but secretly sends premium SMS messages (e.g., costing $10 each) and drains the end user's phone bill. Another example of malware is the Apple iOS flashlight application that secretly collects a user's contacts and sends them to spammers. Other forms of malware (e.g., ransomware) can also be detected / thwarted using the techniques described herein. Additionally, although feature vectors are described herein as being generated to detect malicious JavaScript source code, the techniques described herein may also be used in various embodiments to generate feature vectors for other types of source code (e.g., HTML and / or other programming / scripting languages).

[0020] The techniques described herein can be used in conjunction with various platforms (e.g., desktops, mobile devices, gaming platforms, embedded systems, etc.) and / or various types of applications (e.g., Android .apk files, iOS applications, Windows PE files, Adobe Acrobat PDF files, etc.). In the example environment shown in Figure 1, client devices 104-108 are laptop computers, desktop computers, and tablets (respectively) that reside within enterprise network 140. Client device 110 is a laptop computer that resides outside enterprise network 140.

[0021] Data appliance 102 is configured to enforce policies regarding communications between client devices, such as client devices 104 and 106, and nodes outside enterprise network 140 (e.g., reachable via external network 118). Examples of such policies include policies controlling traffic shaping, quality of service, and traffic routing. Other example policies include security policies, such as those requiring scanning for threats in incoming (and / or outgoing) email attachments, website content, files exchanged via instant messaging programs, and other file transfers. In some embodiments, data appliance 102 is also configured to enforce policies regarding traffic remaining within enterprise network 140.

[0022] An example of a data device is shown in FIG. 2A. The illustrated example is a representation of the physical components included in a data device 102, in various embodiments. Specifically, the data device 102 includes a high-performance multi-core central processing unit (CPU) 202 and random access memory (RAM) 204. The data device 102 also includes storage 210 (such as one or more hard disk units or solid-state storage units). In various embodiments, the data device 102 monitors the enterprise network 110 and stores information (either in RAM 204, storage 210, and / or other suitable locations) used in implementing the disclosed techniques. Examples of such information include application identifiers, content identifiers, user identifiers, requested URLs, IP address mappings, policy and other configuration information, signatures, hostname / URL categorization information, malware profiles, machine learning models, IoT device classification information, etc. The data device 102 may also include one or more optional hardware accelerators. For example, the data device 102 may include an encryption engine 206 configured to perform encryption and decryption operations, and one or more field programmable gate arrays (FPGAs) 208 configured to perform matching, function as a network processor, and / or perform other tasks.

[0023] The functionality described herein as being performed by data device 102 may be provided / implemented in a variety of ways. For example, data device 102 may be a dedicated device or a suite of devices. The functionality provided by data device 102 may also be integrated with or performed as software on a general-purpose computer, computer server, gateway, and / or network / routing device. In some embodiments, at least some functionality described as being provided by data device 102 is instead (or additionally) provided to a client device (e.g., client device 104 or client device 106) by software executing on the client device.

[0024] Whenever a data device 102 is described as performing a task, a single component, a subset of components, or all components of the data device 102 may collaborate to perform the task. Similarly, whenever a component of the data device 102 is described as performing a task, a subcomponent may perform the task and / or the component may collaborate with other components to perform the task. In various embodiments, portions of the data device 102 are provided by one or more third parties. Depending on factors such as the amount of computing resources available to the data device 102, various logical components and / or functions of the data device 102 may be omitted and the techniques described herein may be adapted accordingly. Similarly, additional logical components / functions may be included in embodiments of the data device 102, where applicable. One example of a component included in the data device 102 in various embodiments is an application identification engine configured to identify applications (e.g., using various application signatures to identify applications based on packet flow analysis). For example, the application identification engine may determine what type of traffic a session includes, such as web browsing-social networking, web browsing-news, SSH, etc.

[0025] 2B is a functional diagram of the logical components of one embodiment of a data device. The illustrated example is a representation of the logical components that may be included in the data device 102 in various implementations. Unless otherwise specified, the various logical components of the data device 102 may generally be implemented in a variety of ways, including as a set of one or more scripts (e.g., written in Java, Python, etc., as applicable).

[0026] As shown, the data appliance 102 comprises a firewall and includes a management plane 232 and a data plane 234. The management plane is responsible for managing user interactions, such as by setting policies and providing a user interface for displaying log data. The data plane is responsible for managing data, such as by performing packet processing and session handling.

[0027] The network processor 236 is configured to receive packets from client devices, such as the client device 108, and provide them to the data plane 234 for processing. The flow module 238 creates a new session flow whenever it identifies a packet as being part of a new session. Subsequent packets are identified as belonging to the session based on the flow lookup. If applicable, SSL decryption is applied by the SSL decryption engine 240. Otherwise, processing by the SSL decryption engine 240 is skipped. The decryption engine 240 can help the data device 102 inspect and control SSL / TLS and SSH-encrypted traffic, and thus help prevent threats that might otherwise remain hidden in encrypted traffic. The decryption engine 240 can also help prevent sensitive content from leaving the enterprise network 110. Decryption can be selectively controlled (e.g., enabled or disabled) based on parameters such as URL category, traffic source, traffic destination, user, user group, and port. In addition to decryption policies (e.g., specifying which sessions to decrypt), decryption profiles can be assigned to control various options for the sessions controlled by the policy, for example, requiring the use of specific cipher suites and encryption protocol versions.

[0028] The application identification (APP-ID) engine 242 is configured to determine what type of traffic a session includes. As an example, the application identification engine 242 can recognize a GET request in the received data and conclude that the session requires an HTTP decoder. In some cases, such as during a web browsing session, the identified application may change, and such changes are recorded by the data device 102. For example, a user may first browse a corporate wiki (categorized based on the visited URL as "web browsing-productivity") and then subsequently browse a social networking site (categorized based on the visited URL as "web browsing-social networking"). Different types of protocols have corresponding decoders.

[0029] Based on the determinations made by the application identification engine 242, the packets are sent by the threat engine 244 to the appropriate decoder, which is configured to assemble the packets (which may be received out of order) into the correct order, perform tokenization, and extract information. The threat engine 244 also performs signature matching to determine what happens to the packets. If necessary, the SSL encryption engine 246 can re-encrypt the decrypted data. The packets are forwarded using the forwarding module 248 for transmission (e.g., to a destination).

[0030] 2B, a policy 252 is received and stored in the management plane 232. The policy can include one or more rules, which can be specified using domains and / or host / server names, and the rules can apply one or more signatures or other matching criteria or heuristics, to enforce security policies for subscriber / IP flows based on various parameters / information extracted from the monitored session traffic flows. An interface (I / F) communicator 250 is provided for managing communications (e.g., via (REST) ​​APIs, messages, or network protocol communications or other communication mechanisms).

[0031] Exemplary Security Platform

[0032] Returning to FIG. 1 , assume that a malicious individual (using system 120) has created malware 130. The malicious individual hopes that something, such as a client device 104, will run a copy of malware 130 and compromise the client device, for example, making the client device a bot in a botnet. The compromised client device can then be instructed to perform tasks (e.g., mine cryptocurrency or participate in a denial-of-service attack) as needed, as well as receive instructions from a C&C server 150, and report information to an external entity, such as a command and control (C&C) server 150.

[0033] Assume that the data device 102 intercepts an email sent (e.g., by the system 120) to a user “Alice,” operating the client device 104. A copy of malware 130 is attached to the message by the system 120. Alternatively, but in a similar scenario, the data device 102 could intercept an attempt by the client device 104 to download the malware 130 (e.g., from a website). In either scenario, the data device 102 determines whether a signature of the file (e.g., the email attachment or website download of the malware 130) is present on the data device 102. If present, the signature can indicate that the file is known to be safe (e.g., whitelisted) and can also indicate that the file is known to be malicious (e.g., blacklisted).

[0034] In various embodiments, the data appliance 102 is configured to operate in conjunction with the security platform 122. As an example, the security platform 122 can provide the data appliance 102 with a set of signatures of known malicious files (e.g., as part of a subscription). If the signature of the malware 130 is included in the set (e.g., an MD5 hash of the malware 130), the data appliance 102 can accordingly prevent transmission of the malware 130 to the client device 104 (e.g., by detecting that the MD5 hash of an email attachment sent to the client device 104 matches the MD5 hash of the malware 130). The security platform 122 can also provide the data appliance 102 with a list of known malicious domains and / or IP addresses, allowing the data appliance 102 to block traffic between the enterprise network 140 and the C&C server 150 (where the C&C server 150 is known to be malicious). The list of malicious domains (and / or IP addresses) also helps the data appliance 102 determine when one of its nodes has been compromised. For example, if a client device 104 attempts to contact a C&C server 150, such an attempt is a strong indicator that the client device 104 has been compromised by malware (and remedial action should be taken, such as isolating the client device 104 from communicating with other nodes in the enterprise network 140). As described in more detail below, the security platform 122 can also provide other types of information to the data appliance 102 (e.g., as part of a subscription), such as a set of machine learning models available to the data appliance 102 to perform inline analysis of files.

[0035] In various embodiments, if an attachment signature is not found, the data device 102 can take various actions. As a first example, the data device 102 can be fail-safe by blocking the transmission of any attachments that are not whitelisted as benign (e.g., do not match the signature of a known good file). A drawback of this method is that there may be many legitimate attachments that are unnecessarily blocked as potential malware when they are actually benign. As a second implementation, the data device 102 can be fail-danger by allowing the transmission of attachments that are not blacklisted as malicious (e.g., do not match the signature of a known bad file). A drawback of this approach is that it does not prevent newly created malware (previously unrecognized by the platform 122) from causing harm.

[0036] As a third example, the data device 102 can be configured to provide a file (e.g., malware 130) to the security platform 122 for static / dynamic analysis to determine whether it is malicious and / or otherwise classify it. While the security platform 122 is analyzing the attachment (for which a signature does not yet exist), the data device 102 can take various actions. As a first example, the data device 102 can prevent the email (and any attachments) from being delivered to Alice until it receives a response from the security platform 122. Assuming that it takes approximately 15 minutes for the platform 122 to thoroughly analyze the sample, this means that the incoming message to Alice will be delayed by 15 minutes. In this example, the attachment is malicious, so such a delay will not negatively impact Alice. In another example, assume that someone sends Alice a time-sensitive message with a benign, unsigned attachment. A 15-minute delay in delivering the message to Alice may be deemed unacceptable (e.g., by Alice). As described in more detail below, an alternative approach is to perform at least some real-time analysis on the attachment at the data device 102 (e.g., while awaiting a verdict from the platform 122). If the data device 102 can independently determine whether the attachment is malicious or benign, it can take initial action (e.g., block or allow delivery to Alice), and, if necessary, adjust / take additional action once it receives a verdict from the security platform 122.

[0037] The security platform 122 stores a copy of the received sample in storage 142 and initiates (or schedules, as needed) an analysis. An example of storage 142 is an Apache Hadoop Cluster (HDFS). The analysis results (and additional information about the application) are stored in database 146. If the application is determined to be malicious, the data device can be configured to automatically block file downloads based on the analysis results. Additionally, a signature can be generated and distributed (e.g., to data devices 102, 136, 148) for the malware, automatically blocking future file transfer requests to download files determined to be malicious.

[0038] In various embodiments, security platform 122 includes one or more dedicated, commercially available hardware servers (e.g., having multi-core processors, 32G+ RAM, Gigabit network interface adapters, and hard drives) running a typical server-class operating system (e.g., Linux). Security platform 122 can be implemented across a scalable infrastructure including multiple such servers, solid-state drives, and / or other applicable high-performance hardware. Security platform 122 may include several distributed components, including components provided by one or more third parties. For example, part or all of security platform 122 can be implemented using Amazon Elastic Compute Cloud (EC2) and / or Amazon Simple Storage Service (S3). Additionally, similar to data devices 102, whenever security platform 122 is referred to as performing a task, such as storing or processing data, it should be understood that a subcomponent or subcomponents of security platform 122 can cooperate (individually or in cooperation with third-party components) to perform that task. As an example, security platform 122 may optionally cooperate with one or more virtual machine (VM) servers, such as VM server 124, to perform static and dynamic analysis.

[0039] An example of a virtual machine server is a physical machine including commercially available server-class hardware (e.g., a multi-core processor, 32+ gigabytes of RAM, and one or more gigabit network interface adapters) running commercially available virtualization software such as VMware ESXi, Citrix XenServer, or Microsoft Hyper-V. In some embodiments, the virtual machine server is omitted. Furthermore, the virtual machine server may be under the control of the same entity that manages security platform 122, but may also be provided by a third party. As an example, the virtual machine server may rely on EC2, with the remainder of security platform 122 being provided by dedicated hardware owned and under the control of the operator of security platform 122. VM server 124 is configured to provide one or more virtual machines 126-128 to emulate client devices. The virtual machines may run various operating systems and / or versions thereof. Observed behavior as a result of running applications on the virtual machines is logged and analyzed (e.g., to indicate that the application is malicious). In some embodiments, log analysis is performed by the VM server (e.g., VM server 124). In other embodiments, the analysis is performed at least in part by other components of the security platform 122 , such as the coordinator 144 .

[0040] In various embodiments, the security platform 122 provides the analysis results of the samples to the data device 102 as part of the subscription via a list of signatures (and / or other identifiers). For example, the security platform 122 can periodically (e.g., daily, hourly, or some other interval and / or based on events configured by one or more policies) send a content package identifying malware apps. An example content package includes a list of identified malware apps, including information such as the package name, a hash value to uniquely identify the app, and the malware name (and / or malware family name) of each identified malware app. The subscription can cover analysis of only files intercepted by the data device 102 and transmitted by the data device 102 to the security platform 122, and can also cover the signatures of all malware known to the security platform 122 (or a subset thereof, such as merely mobile malware and not other forms of malware (e.g., PDF malware)). As described in more detail below, the platform 122 can also make available other types of information. Such as machine learning models (e.g., based on feature vectors) that help the data device 102 detect malware (e.g., through techniques other than hash-based signature matching).

[0041] In various embodiments, security platform 122 is configured to provide security services to various entities in addition to (or instead of, as needed) the operator of data device 102. For example, other businesses with their own respective enterprise networks 114 and 116 and their own respective data devices 136 and 148 may contract with the operator of security platform 122. Other types of entities may also utilize the services of security platform 122. For example, an Internet Service Provider (ISP) providing Internet service to client device 110 may contract with security platform 122 to analyze applications that client device 110 attempts to download. As another example, the owner of client device 110 may install software on client device 110 that communicates with security platform 122 (e.g., receive a content package from security platform 122, use the received content package to verify attachments according to techniques described herein, and then send the application to security platform 122 for analysis).

[0042] Analyzing samples using static and dynamic analysis

[0043] 3 illustrates an example of logical components that may be included in a system for analyzing a sample. Analysis system 300 may be implemented using a single device. For example, the functionality of analysis system 300 may be implemented in malware analysis module 112 embedded in data appliance 102. Analysis system 300 may also be implemented collectively on multiple different devices. For example, the functionality of analysis system 300 may be provided by security platform 122.

[0044] In various embodiments, analysis system 300 utilizes a list, database, or other collection of known good content and / or known bad content (collectively shown in FIG. 3 as collection 314). Collection 314 can be obtained in a variety of ways, including via a subscription service (e.g., provided by a third party) and / or as a result of other processing (e.g., performed by data device 102 and / or security platform 122). Examples of information included in collection 314 include URLs, domain names, and / or IP addresses of known malicious servers, URLs, domain names, and / or IP addresses of known safe servers, URLs, domain names, and / or IP addresses of known command and control (C&C) domains, signatures, hashes, and / or other identifiers of known malicious applications, signatures, hashes, and / or other identifiers of known safe applications, signatures, hashes, and / or other identifiers of known malicious files (e.g., Android exploit files), signatures, hashes, and / or other identifiers of known safe libraries, and signatures, hashes, and / or other identifiers of known malicious libraries.

[0045] Import (ingestion)

[0046] In various embodiments, when a new sample is received for analysis (e.g., there is no existing signature associated with the sample in analysis system 300), it is added to queue 302. As shown in Figure 3, malware 130 is received by system 300 and added to queue 302.

[0047] static analysis

[0048] The coordinator 304 monitors the queue 302, and as resources (e.g., static analysis workers) become available, the coordinator 304 fetches samples from the queue 302 for processing (e.g., obtains copies of malware 130). In particular, the coordinator 304 initially provides samples to the static analysis engine 306 for static analysis. In some embodiments, one or more static analysis engines are included within the analysis system 300, where the analysis system 300 is a single device. In other embodiments, the static analysis is performed by a separate static analysis server that includes multiple workers (i.e., multiple instances of the static analysis engine 306).

[0049] The static analysis engine obtains general information about the sample and includes it in a static analysis report 308 (along with heuristics and other information, if necessary). The report may be generated by the static analysis engine or by the coordinator 304, which can be configured to receive information from the static analysis engine 306 (or by another suitable component). In some embodiments, the collected information is stored in a database record for the sample (e.g., in database 316) (i.e., portions of the database record form the report 308) instead of or in addition to generating a separate static analysis report 308. In some embodiments, the static analysis engine also generates a verdict about the application (e.g., “safe,” “suspicious,” or “malicious”). As one example, if an application has at least one “malicious” static feature (e.g., the application contains a hard link to a known malicious domain), the verdict may be “malicious.” As another example, each feature can be assigned points (e.g., based on severity if found, based on how reliable the feature is for predicting malicious intent, etc.), and a verdict can be assigned by the static analysis engine 306 (or coordinator 304, if applicable) based on the number of points associated with the static analysis result.

[0050] Dynamic Analysis

[0051] Once the static analysis is complete, the coordinator 304 locates an available dynamic analysis engine 310 and runs dynamic analysis on the application. Similar to static analysis engine 306, analysis system 300 can include one or more dynamic analysis engines directly. In other embodiments, dynamic analysis is performed by a separate dynamic analysis server that includes multiple workers (i.e., multiple instances of dynamic analysis engine 310).

[0052] Each dynamic analysis worker manages a virtual machine instance. In some embodiments, the results of the static analysis (e.g., performed by static analysis engine 306), whether in report form 308 and / or stored in database 316 or otherwise, are provided as input to dynamic analysis engine 310. For example, static report information can be used to help select / customize the virtual machine instance used by dynamic analysis engine 310 (e.g., Microsoft Windows 7 SP 2 vs. Microsoft Windows 10 Enterprise, or iOS 11.0 vs. iOS 12.0). If multiple virtual machine instances run simultaneously, a single dynamic analysis engine can manage all instances, or multiple dynamic analysis engines can be used (e.g., each managing its own virtual machine instance), as desired. As described in more detail below, during the dynamic portion of the analysis, actions performed by the application (including network activity) are analyzed.

[0053] In various embodiments, if desired, static analysis of the sample may be omitted or performed by another entity. By way of example, traditional static and / or dynamic analysis may be performed on the file by a first entity. Once a given file is determined to be malicious (e.g., by the first entity), the file may be provided to a second entity (e.g., an operator of security platform 122) for further analysis of network activity, particularly for malware use (e.g., by dynamic analysis engine 310).

[0054] The environment used by analysis system 300 is instrumented / hooked (e.g., using a customized kernel that supports hooks and logcat) so that observed behavior while the application is running is recorded as it occurs. Network traffic associated with the emulator is also captured (e.g., using pcap). The log / network data can be stored in analysis system 300 as temporary files, and can also be stored more persistently (e.g., using HDFS or other suitable storage technology, or a combination of technologies, such as MongoDB). The dynamic analysis engine (or another suitable component) can compare connections made by the sample to a list (314) of domains, IP addresses, etc., and determine whether the sample communicated (or attempted to communicate) with malicious entities.

[0055] Similar to the static analysis engine, the dynamic analysis engine stores the results of its analysis in database 316 in a record associated with the application being tested (and / or includes the results in report 312, if desired). In some embodiments, the dynamic analysis engine also forms a verdict for the application (e.g., “safe,” “suspicious,” or “malicious”). As one example, the verdict may be “malicious” if any “malicious” actions were performed by the application (e.g., an attempt was made to connect to a known malicious domain or an attempt to surreptitiously extract sensitive information was observed). As another example, points may be assigned to the actions performed (e.g., based on severity if found, based on how reliable the feature is for predicting maliciousness, etc.), and a verdict may be assigned by dynamic analysis engine 310 (or coordinator 304, if applicable) based on the number of points associated with the dynamic analysis results. In some embodiments, a final verdict associated with the sample is determined (e.g., by coordinator 304) based on a combination of report 308 and report 312.

[0056] Additional details for threat engines

[0057] In various embodiments, the data device 102 includes a threat engine 244. The threat engine incorporates both protocol decoding and threat signature matching during the respective decoder and pattern matching stages. The results of the two stages are merged by the detector stage.

[0058] When the data device 102 receives a packet, it performs a session match to determine which session the packet belongs to (allowing the data device 102 to support simultaneous sessions). Each session has a session state that implies a particular protocol decoder (e.g., a web browsing decoder, an FTP decoder, or an SMTP decoder). If a file is sent as part of the session, the applicable protocol decoder can utilize an appropriate file-specific decoder (e.g., a PE file decoder, a JavaScript decoder, or a PDF decoder).

[0059] Portions of one example of the threat engine 244 are shown in FIG. 4. In one embodiment, for a given session, the decoder 402 walks the traffic byte stream according to the corresponding protocol and marking context. One example of a context is an end-of-file context (e.g., encountered during processing of a JavaScript file). The decoder 402 can mark the end-of-file context in the packet and then use this to trigger the execution of the appropriate model using the observed features of the file. In some cases (e.g., FTP traffic), there may not be an explicit protocol-level tag for the decoder 402 to identify / mark the context. In another embodiment, the decoder component 402 is configured to determine the file type associated with each file in the sample 404 (e.g., a malware sample may include various source code content and / or other types of content for malware analysis, such as JS code, HTML code, and / or other programming / scripting languages, as well as other structured text such as URLs or unstructured content such as images, etc.), and can then decode the files to perform static analysis using the IUPG model, as described further below.Also, as described in more detail below, in various embodiments, decoder 402 uses other information (e.g., the file size reported in the header) to determine when to stop feature extraction for the file (e.g., when the overlay section begins) and when to begin execution using an appropriate model (e.g., as described further below, decoder 402 can determine the file type associated with sample 404 and then select an IUPG model appropriate for the type of source code associated with that file type, such as a JS IUPG model for a JS file, an HTML IUPG model for an HTML file, etc., and analyzer 406 can perform static analysis of the sample using the appropriate IUPG model).

[0060] The threat engine 244 also includes an analyzer component 406 for performing static analysis of the samples 404 using a selected IUPG model, as described further below. The analyzer component 408 (e.g., using the target feature vector of the selected IUPG model) determines whether to classify each analyzed sample 404 as malicious or benign (e.g., based on a threshold score), as described further below. By way of example, the analyzer 406 and detector 408 can be implemented by the data device 102 and / or by security agents / software running on the client 110 (e.g., by the analyzer and detector 154 of the security platform 122, as also shown in FIG. 1 ) using the disclosed techniques of IUPG models applied to malware classification based on static analysis of source code samples. The detector 408 processes the output provided by the decoder 402 and analyzer 406 and performs various response actions (e.g., based on security policies / rules).

[0061] Introducing the Presumption of Innocent (IUPG): Building Adversarial-Resistant and False-Positive-Resistant Deep Learning Models

[0062] Categorical Cross-Entropy (CCE) loss is a standard supervised loss function used to train various deep neural network (DNN) classifiers. CCE produces purely discriminative models without built-in means to infer out-of-distribution (OOD) content or effectively exploit classes that lack uniquely identifiable structure. The disclosed technology, providing the Innocent Until Proven Guilty (IUPG) framework, offers alternative architectural components and a hybrid discriminative and generative loss function for training DNNs to classify mutually exclusive classes. IUPG involves learning a library of inputs in the original input space that, together with the network, prototype uniquely identifiable subsets of the input space. The network learns to map the input space to an output vector space, where prototypes and members of related input subsets map exclusively to a common point in the output vector space. The distance between noise (or any class of data lacking a prototype description) and all prototypes in the output vector space is maximized in training. Such classes are called "off-target," while the target class has one or more prototypes assigned. The off-target data helps chisel down the extracted features of the target class to those that are truly class-exclusive, as opposed to by chance.

[0063] For example, machine learning techniques (MLT) applied to computer / network security have a significant challenge—i.e., they generally must not make mistakes. Mistakes in one direction can lead to the intrusion of dangerous malware (e.g., malware being allowed into an enterprise network or executed on a computing entity (such as a server, computing endpoint, etc.)). Mistakes in the other direction can cause security solutions to, for example, block harmless traffic or executable files on a computing entity, which is also expensive for cybersecurity companies and a major headache for users. Generally, the amount of good (benign) traffic far outweighs the amount of malicious traffic; therefore, minimizing the amount of good traffic that is called malicious (e.g., what are commonly referred to as false positives, or FPs) is typically desired for an effective and efficient security solution. Malware creators understand this and attempt to make their malicious code appear more benign. The simplest way to accomplish this is generally known as an append attack (also known, e.g., as injection, bundling, etc.), in which an attacker obtains a (e.g., typically large) amount of benign content and then injects malicious content into it without impairing the functionality of the malware, as described further below. Because machine learning (ML) classifiers built using standard techniques are sensitive to its presence, a significant portion of benign content can confuse classification verdicts from positive malware verdicts, and sometimes cause the classifier to miss it entirely.

[0064] In the context of malware classification using the IUPG technique, this corresponds to learning the inseparable features of malware clusters that define their maliciousness while ignoring benign content. In one implementation, during inference, each sample is scanned for these inseparable qualities while ignoring all out-of-class structure, and thus the technique assumes that each sample is "presumed innocent." Intuitively, increasing the specificity of the learned features increases the network's resistance to arbitrary OOD content (e.g., noise resistance, as described further below). As a central hypothesis of the disclosed IUPG technique, we propose that this increased resistance to arbitrary OOD content is primarily responsible for the desired effects we explore, as described further below.

[0065] The disclosed IUPG technique achieves this by learning a library of abstracted inputs within the original representation of the data, which—together with the network's layer operations—learns a prototype subset of the input space. The network learns to map the input space to an output vector space, in which prototypes and members of the associated input subsets map exclusively to a common point. The distance between noise in the output vector space (or classes of data lacking a prototype description, commonly referred to as "off-target") and all prototype inputs is maximized during training. The off-target examples thus help prune the extracted features of the target class to those that are truly class-exclusive, as opposed to chance.

[0066] Increasing the specificity of the learned representations of classes (while still balancing the generalizability of the loss) naturally increases the network's resistance to input noise. We hypothesize that it is the noise-tolerance property built into the IUPG network that is primarily responsible for the desirable qualities explored in this work. We use an equivalent network topology trained with a CCE loss to measure baseline performance. We refer to this control network configuration as the IUPG network's CCE counterpart. In our evaluation, we (1) explore the test set classification performance of IUPG and its CCE counterpart across various cybersecurity and computer vision experimental settings, including different usages of noise; (2) measure the tendency of both frameworks to generate false positive (FP) responses for OOD inputs; (3) measure the resistance of both frameworks to black-box append attacks using both standard training and our custom adversarial training procedure; and (4) demonstrate the applicability of existing adversarial learning techniques to IUPG.

[0067] By increasing the specificity of structured class models through prototype-based learning and unique handling of structureless classes, as described below with respect to various embodiments, IUPG-trained networks can offer significant advantages in common real-world problem settings, such as predetermined noise-based adversarial attacks, handling distribution shifts, and out-of-distribution classification in computer / network security contexts. Because IUPG is general enough to be applied to any architecture where categorical cross-entropy (CCE) can be used, various opportunities for combining IUPG with existing adversarial learning / OOD detectors exist on their own, outperforming either technique used in isolation, as described further below.

[0068] We demonstrate that the unique advantages of IUPG are particularly useful in malware classification efforts. In the context of malware classification, append attacks can lead to risky false negatives, while OOD failures can lead to costly false positives. In summary, various novel aspects disclosed herein include, without limitation: (1) presenting the IUPG framework; (2) demonstrating some of the aforementioned advantages of using IUPG over CCE loss; (3) showing a novel architecture and training procedure for building an append attack-resistant DNN malware classifier; and (4) applying the IUPG framework to effectively and efficiently detect various forms of malware (e.g., JavaScript-related malware, URL-related malware, and / or other forms of malware can similarly be classified and detected using the disclosed IUPG techniques and framework, as further described below).

[0069] As explained further below, experimental results reveal that append attacks can be highly successful even against highly accurate classifiers. As an example, for a deep-learning JavaScript (JS) malware classifier built using categorical cross-entropy (CCE) loss, we found that even though the classifier achieved over 99% accuracy on the test set, just 10,000 characters of random benign content appended to malicious samples successfully overturned the verdict over 50% of the time. This is particularly concerning given the extremely low cost of exploiting this attack: an adversary does not need to know any details about the victim's classifier, while at the same time, benign content is extremely abundant and easy to create. If the adversary has access to sensitive information about the victim model, such as the loss function, the appended content can generally be designed with model-specific techniques that further increase the success rate.

[0070] To solve this technically challenging problem, content that is not uniquely indicative of malware must generally have a sufficiently small impact on the classification mechanism so that the verdict does not flip to benign. At a high level, our approach encourages the network to exclusively learn and recognize uniquely identifiable patterns of malicious classes, while being explicitly robust to all other content. A key observation is that malware patterns are highly structured and uniquely recognizable compared to the infinite number of benign patterns that may be encountered in the data. Thus, an example of the innovation of the disclosed IUPG technology is its use for learning, distinguishing between classes that have uniquely identifiable structure (e.g., patterns) and those that do not. In the context of malware classification, malware classes generally have uniquely identifiable structure (e.g., referred to herein as target classes), and benign classes are essentially random (e.g., referred to herein as off-target classes). As explained further below, the disclosed IUPG technique is specifically designed to learn uniquely identifiable structures within a target class, while only leveraging off-target classes to prune the target class representation to one that is truly inseparable. This encourages the neural network to reduce sensitive data patterns to only those malicious patterns that are clear indicators of a correct positive verdict within the neural network's overall receptive field. Only when no such malicious patterns are found is a benign verdict generated; that is, unknown files are presumed innocent (IUPG). This contrasts with traditional, unconstrained learning, which can help minimize losses but ultimately provides no information about the overall safety of the file. Furthermore, we assume that any benign patterns learned by a classifier are likely overfitted features to facets of the situational training data. Due to the nearly infinite number of possible manifestations, attempting to capture benign patterns is inefficient at best.In the worst case, it leads to overfitted features, opening the classifier to additional susceptibility to attacks. Worse yet, if the training, validation, and test splits are drawn from the same distribution (e.g., this is common practice), standard test set classification metrics are unlikely to reveal the problem, as benign features of the classifier may still yield good performance on the test set. Only once the classifier is placed in the real world with data outside the training distribution (e.g., if that is important) will nuisance classification errors and vulnerability to attacks occur.

[0071] As explained in more detail below, in our evaluation, we measure baseline performance—herein referred to as the IUPG network's CCE counterpart—using an equivalent network trained with a CCE loss. We (1) investigate the test set classification performance of IUPG and its CCE counterpart across a variety of cybersecurity and computer vision settings, including various uses of noise, (2) compare the tendency to generate false positive (FP) responses with OOD inputs, (3) compare the impact of recency bias (e.g., performance degradation due to distribution shifts) on classification accuracy, (4) compare the resistance of both frameworks to black-box append attacks, and (5) demonstrate the applicability of existing adversarial training techniques to IUPG.

[0072] Related research background

[0073] We summarize three main themes of related research, including: (1) prototype-based learning, (2) append attacks, and (3) out-of-distribution (OOD) attacks.

[0074] Prototype-Based Learning Some of the earliest work on prototype-based learning is Learning Vector Quantization (LVQ) (see, for example, T. Kohonen, The self-organizing map. Neurocomputing, 21(1):1-6, 1998, ISSN 0925-2312, doi:https: / / doi.org / 10.1016 / S0925-2312(98)00030-7), which can be thought of as a prototype-based k-nearest neighbor algorithm. In the taxonomy of LVQ variants presented in D. Nova and P.A. Est´evez. A review of learning vector quantization classifiers. Neural Comput. Appl. c25(3-4):511-524, Sep. 2014. ISSN 0941-0643. doi:10.1007 / s00521-013-1535-3, this work corresponds to GLVQ (see, for example, A. Sato and K. Yamada, Generalized learning vector quantization. In Proceedings of the 8th International Conference on Neural Information Processing Systems, NIPS'95, pages 423-429, Cambridge, MA, USA, 1995, MIT Press), which is based on margin maximization in the data space using Euclidean distance. IUPG, however, combines prototype learning with DNNs and uniquely uses off-target examples.Common goals of prototype-based learning in DNNs include low-shot learning (see, e.g., X. Liu et al. Meta-learning based prototype-relation network for few-shot classification, Neurocomputing, 383:224-234, 2020, ISSN 0925-2312, doi:https: / / doi.org / 10.1016 / j.neucom.2019.12.034) and interpretability of decisions (see, e.g., O. Li, H. Liu, C. Chen, and C. Rudin, Deep learning for case-based reasoning through prototypes: A neural network that explains its predictions, 2018). H.-M. Yang, X.-Y. Zhang, F. Yin, and C.-L. Liu, "Robust classification with convolutional prototype learning," 2018 IEEE / CVF Conference on Computer Vision and Pattern Recognition, June 2018, doi:10.1109 / cvpr.2018.00366 (Yang et al.) shares some similarities with IUPG. However, Yang et al.'s prototypes are defined in the model's output vector space. These prototypes are not human-interpretable and do not have intuitive initialization. Importantly, when prototypes are defined in this way, we observed frequent convergence to a solution in which multiple prototypes are merged at a common point. This is supported by Yang et al.'s results, which report similar or worse performance with multiple prototypes per class. We also could not find similar studies that utilize off-target samples, as IUPG does.This capability allows us to find considerable advantages, especially for problems such as malware classification, where only one class has a uniquely identifiable structure.

[0075] Append Attacks: Append attacks involve concatenating adversarial content into the input with the goal of confusing classification results (see, e.g., R.R. Wiyatno, A. Xu, O. Dia, and A. de Berker, Adversarial examples in modern machine learning: A review, 2019 (Wiyatno et al.)). This is particularly relevant to malware classification, where benign noise can be added to malware to fool classifiers while malicious activity remains intact. Conversely, malicious content can be inserted into larger benign files to evade detection (e.g., benign library injection or adding more whitespace is a common form of append attack; using custom bundle files with various jQuery plugins and website dependencies can lead to false benign classifications by many classifiers, as the injection of benign libraries / other content into such files can confuse classification results). In what are known as white-box attacks (see, e.g., Wiyatno et al.), the added noise can be crafted while exploiting model details. Typically, black-box adversarial attacks assume no knowledge of the model and are often the only possible attack against proprietary defenses. This work provides evidence that even the simplest types of append attacks can pose a serious threat to highly accurate models. To the best of our knowledge, previous work in deep learning lacks a general solution for append attacks against malware.

[0076] Out-of-Distribution (OOD) Classification: It is well understood that DNNs trained with CCE (save for some specialized, highly nonlinear options, such as RBF networks) are prone to generating highly overconfident posterior distributions for OOD inputs (see, e.g., V. Sehwag et al., Analyzing the robustness of open world machine learning, pages 105–116, Nov. 2019, ISBN 978-1-4503-6833-9, doi:10.1145 / 3338501.3357372). Reliably processing OOD content is a key requirement for practical systems. Orthogonal to our work, open-world frameworks often equip models with external detectors aimed at identifying and discarding OOD inputs (see, e.g., J. Chen, Y. Li, X. Wu, Y. Liang, and S. Jha, Robust out-of-distribution detection for neural networks, 2020). Other work relies on learning external rejection functions either simultaneously with or after training the classification network (see, e.g., Y. Geifman and R. El-Yaniv, Selective classification for deep neural networks, in I. Guyon, U.V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems 30, pages 4878-4887, Curran Associates, Inc., 2017). This work demonstrates built-in adaptability for processing OOD content arising solely from IUPG.

[0077] Combined Advantages: In general, any of the techniques recently developed for CCE-trained DNNs to overcome white-box attacks, such as many varieties of adversarial training (see, e.g., Wiyatno et al.), are equally applicable to IUPG-trained networks. The IUPG loss can be used as a drop-in replacement for CCE in these specialized training procedures. An example is provided below where we observe a consistently high IUPG success rate. Similarly, we suggest that the combination of IUPG and an external OOD detector is likely to outperform either in isolation, as will be described below with respect to various embodiments.

[0078] Overview of Techniques for Presumption of Innocence (IUPG): Building and Using Adversarial-Resistant and False-Positive-Resistant Deep Learning Models

[0079] Presumption of Innocence (IUPG) techniques for building and using adversarial-resistant and false-positive-resistant deep learning models are disclosed. In some embodiments, a system, process, and / or computer program product includes storing a set including one or more presumption of innocence (IUPG) models for static analysis of a sample, performing a static analysis of content associated with the sample, where performing the static analysis includes using at least one stored IUPG model, determining that the sample is malicious based at least in part on the static analysis of the content associated with the sample, and performing an action based on a security policy in response to determining that the sample is malicious.

[0080] Deep neural network classifiers trained with traditional categorical cross-entropy losses face challenges in real-world environments, including a tendency to generate highly reliable posterior distributions for out-of-distribution inputs, sensitivity to adversarial noise, and performance loss due to distribution shifts. We hypothesize that a central drawback—the inability to effectively handle out-of-distribution content in the input—exacerbates each of these obstacles. In response, we propose a novel learning framework, called Innocent Until Proven Guilty, that prototypes training data clusters or classes in the input space while uniquely exploiting noise and inherently random classes to discover uniquely discriminatory features for the modeled classes, while being robust to noise. In our evaluation, we leverage both academic computer vision datasets and real-world JavaScript and URL datasets for malware classification.

[0081] Across these interdisciplinary settings, we observe favorable classification performance on test data, reduced performance loss due to recency bias, fewer false positive responses in noisy samples, and reduced vulnerability to several noise-based attack simulations when compared to baseline networks of equal topology trained with categorical cross entropy.

[0082] The published IUPG framework demonstrates a significant reduction in vulnerability to black-box append attacks in malware. For example, by applying the well-known Fast-Gradient Sign Method, we demonstrate the feasibility of combining our framework with existing adversarial learning techniques and finding favorable performance by a significant margin. Our framework is general enough to be used with any network topology that could otherwise be trained with categorical cross entropy (CCE).

[0083] Exemplary System Embodiments of the IUPG Framework

[0084] 5A is a block diagram of an extended IUPG component in an abstract network N in accordance with some embodiments. In general, an IUPG network is an encoder that maps inputs and prototypes from a common input space to an output vector space.

[0085] As a brief overview of the IUPG framework and technology, a library of input samples and prototypes is processed by the network in a Siamese fashion for each forward pass. These prototypes exist in the same input space as regular data points. For example, if classifying 28x28 images, each prototype exists as a 28x28 matrix of learnable weights. In one implementation, prototypes can be defined in two ways: (1) directly as learnable members of the input space; or (2) more generally, as weights of a linear combination of a biased set of training data points. The latter is particularly useful when the input space is very large. The samples and prototypes are mapped to a final output vector space paired with a specially trained distance metric. IUPG learns the prototypes, network weights, and distance metric so that the output vector space orients all inputs. This is shown in Figure 5B using the IUPG network 500 described below with reference to Figure 5A.

[0086] In an ideal mapping, structured class members and their assigned prototypes are uniquely mapped to a common point, with space to spare, such that possible inputs that are not members of the structured class are mapped elsewhere. Now, obviously, a verdict is drawn by measuring the distance of the mapped input to all mapped prototypes in this output vector space. If the mapped sample is measured to be sufficiently close to the prototype, the prototype is predicted to be a member of the assigned class. As shown in the space at 530, a background of noise, also referred to here as off-target data, helps illuminate (and capture in prototypes) what is truly inseparable with respect to the target class, such as Class 1 prototype 532 and Class 2 prototype 534, as shown in Figure 5B. Note that the IUPG network 500 can still be successfully trained without off-target data or classes. We report stable or improved classification performance using several public datasets of this type. However, certain problems, such as malware classification, are naturally suited to leveraging this capacity to recognize off-target data.

[0087] The IUPG loss facilitates this ideal mapping by adjusting the pushing and pulling forces between the sample and all prototypes in the output vector space. The force applied to each anchor sample is determined based on its label. Figure 5C shows the necessary forces for on-target and off-target samples. Note that the off-target sample 550 is pushed away from all prototypes, including the target 1 prototype 552, target 2 prototype 554, and target 3 prototype 556, as shown in Figure 5C. This is achieved by using a zero vector for its one hot label vector.

[0088] Returning to binary malware classification, as mentioned above, we specify several prototypes for malicious classes while defining benign classes as off-targets. Hopefully, it is now clear why we need to learn uniquely identifying patterns for malware while encoding robustness against benign content. In the ideal case, the prototype and network mapping exclusively captures the inseparable features of a malware family, so that their activation is the strongest possible indicator of malware, and other features do not lead to significant activation. This traps an adversary in a situation where the only path to breaking the mapping to malicious prototypes is to distort or remove the pieces of malware that actually do something malicious. Patterns that are not directly malicious will rarely activate, making leveraging extra benign content useless for an adversary to bypass the malware classifier. Importantly, note that forming tight, robust clusters of malware families around prototypes in the output vector space is simultaneously balanced with the generalizability of the loss, so as to reliably catch, for example, orphan malware. In our experiments, prototypes do not generally map to single malware families or malicious patterns, as if the model were reduced to simple pattern matching. Instead, the network and prototypes learn to recognize complex, sophisticated combinations of patterns that generalize across malware families while still retaining robustness against benign activations.

[0089] Referring now to FIG. 5A, novel components of the IUPG framework 500 include the input and output layers of the DNN, as well as a special loss function, as described below with respect to FIG. 5A. The details of all hidden layers—including the number of layers of different functional types organized into any topology—can be varied as needed based on the relevance of the problem.

number

number

number

number

number

[0090] Data Guidelines

[0091] For CCE training using c classes, it is generally necessary and sufficient to obtain labeled examples of all c classes. Although IUPG can be trained using these datasets, we often find it useful to include off-target samples.

number

number

[0092] prototype

[0093] The IUPG network, in a Siamese fashion, uses the input and library of ρ prototypes.

number

number

number

number

number

number

number

number

number

number

number

number

[0094] Distance Function (Distance Function)

[0095]

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

[0096] Loss function (Loss Function)

[0097] IUPG loss is measured using a sample and

number

number

number

number

number

[0098]

number

[0099] y i If =1, we

number

number

number

[0100] The assignment of ρ prototypes to c target classes is

number

number

number

number

number

number

number

number

number

number

[0101] Training and inference complexity

[0102] If all weights remain unchanged,

number

number

number

number

number

number

number

number

number

[0103] experiment

[0104] The experiments consider not only the classification of malicious JavaScript (JS) and URLs, but also the classification of MNIST (see, e.g., Yann LeCun and Corinna Cortes, MNIST handwritten digit database, 2010) and Fashion-MNIST (see, e.g., Han Xiao, Kashif Rasul, and Roland Vollgraf, Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms, 2017). For JS, we consider both the binary general-purpose malware classification problem and the multi-class malware family tagging problem. All models are implemented in TensorFlow (see, e.g., M. Abadi et al., TensorFlow: Large-scale machine learning on heterogeneous distributed systems, 2015; TensorFlow open source software is available at tensorflow.org) and trained with the Adam optimizer (see, e.g., Diederik P. Kingma and Jimmy Ba, Adam: A method for stochastic optimization, 2014). The training batch size is 32 and the learning rate is 5 × 10. -5 is used throughout. Across all convolutional and fully connected layers, we use ReLU and sigmoid activation, respectively. These hyperparameters allow both IUPG and CCE trained networks to converge after roughly the same number of batches. The shared hyperparameters used in this work are tuned while using the CCE loss - and are therefore biased towards the CCE counterpart. For IUPG, we set k=32 throughout. All

number

number

[0105] Figure 6 shows the function N used for (A) MNIST and Fashion MNIST, (B) JS, and (C) URL, according to some embodiments. In (A), the model includes parallel convolutional layers. C:128@5x5 is a convolutional layer of 128 5x5 filters. maxP@2x2 is 2x2 max pooling. FCC is a fully connected convolutional layer. FC:512 is a fully connected layer with 512 units. In (B), char-level and token-level input representations are processed independently. EVL refers to an embedded vector lookup operation. seqComp refers to a sequence compression operation. globalmaxP refers to a global max pooling operation. In (C), C:128@11,3x30 indicates two different heights used in the filter bank: 11 for character-level input and 3 for token-level input.

[0106] Image Classification

[0107] For simplicity, MNIST (see, e.g., Yann LeCun and Corinna Cortes, MNIST handwritten digit database, 2010) and Fashion-MNIST (see, e.g., Han Xiao, Kashif Rasul, and Roland Vollgraf, Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms, 2017) are treated roughly the same. Both datasets are split into random 50k-10k-10k Train-Test-Val (TTV) splits. Optionally, to filter out chance true positives, we generate Gaussian noise images as well as random stroke images using a random forest classifier. Images are preprocessed using max-min scaling and mean subtraction. When using IUPG, each

number

[0108] Malicious JS and Static URL Categorization

[0109] Classifying malicious JS and static URLs is a challenging task in web security (see, for example, Aurore Fass, Michael Backes, and Ben Stock, Jstap: A static pre-filter for malicious javascript detection, In Proceedings of the 35th Annual Computer Security Applications Conference, ACSAC 2019, pages 257-269, New York, NY, USA, 2019, Association for Computing Machinery; Yann LeCun and Corinna Cortes, MNIST handwritten digit database, 2010; Doyen Sahoo, Chenghao Liu, and Steven CHHoi, Malicious URL detection using machine learning: A survey, 2017). Append attacks are particularly popular among JS malware, such as malicious injections into benign scripts. A simple but critical observation in malware classification is that benignity can only be defined in terms of non-maliciousness. For IUPG, we define benign data as an off-target class. That is, no prototypes are used for modeling.

[0110] Benign JS samples were collected by crawling the top 1M domains from the Tranco list (see, e.g., Victor Le Pochat, Tom van Goethem, and Wouter Joosen, "Rigging research results by manipulating top websites rankings," CorRR, abs / 1806.01156, 2018). In addition to Tranco filtering, we ignored samples flagged by state-of-the-art commercial URL filtering services. We utilized VirusTotal (VT) as our primary source of malicious JS samples (see, e.g., Gaurav Sood, "virustotal: R Client for the virustotal API," 2017, R package version 0.2.1). We required a VT score (VTS) of at least 3, which has been empirically shown to be reasonably accurate. Malicious and benign URL data was collected via Internet traffic as well as external data sources (e.g., VT) from industry cybersecurity companies using static and dynamic URL filters and analyzers. For binary JS malware classification, we used 450k-600k-600k TTV splits with benign-to-malicious ratios of 70:30, 96:4, and 96:4, respectively. Substantially benign samples were included to accurately measure performance under strict false positive rate (FPR) requirements. Due to the prohibitive cost of FPR, FPRs of 0.1% or less are common in industrial cybersecurity settings (see, for example, this site produced by Byte Productions and maintained by www.byte-productions.com, The cost of malware containment, January 2015). To build multi-class malware family tagging classifiers, we separated nine different malware families using 10k-1k-1k TTV samples per family.An equal portion of benign data was added to form a multi-class training dataset. To generate OOD samples, we evenly scrambled the order of tokens in benign scripts. For URLs, we used a 14M-2M-2M TTV split with a 50:50 class ratio. We also collected another 2M, 50:50 test set one year after the initial collection to test for recency bias.

[0111] The JS and URL classifier architectures shown in Figure 6(B) and (C) are built on various prior works in NLP (e.g., Xiang Zhang, Junbo Zhao, and Yann LeCun, In Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 1, NIPS 2015, pages 649-657, Cambridge, MA, USA, 2015, MIT Press; Yoon Kim, Convolutional neural networks for sentence classification, In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, EMNLP 2014, October 25-29, Doha, Qatar, A meeting of SIGDAT, the Special Interest Group of the ACL, pages 1746-1751, 2014; Michele Tufano, Cody Watson, Gabriele Bavota, Massimiliano Di Penta, Martin White, and Denys Poshyvanyk, Deep learning similarities from different representations of source code, In Proceedings of the 15 thInternational Conference on Mining Software Repositories, MSR 2018, pages542-553, New York,NY, USA, 2018, Association for Computing Machinery; IEEE Military Communications Conference, MILCOM 2019, Norfolk, VA, USA, November 12-14,pages1-8, IEEE, 2019; M.Sugiyama,およand R.Garnett, editors, Advances in Neural Information Processing Systems 28, pages919-927, Curran Associates, Inc., 2015; Rie Johnson and Tong Zhang, Effective use of word order for text categorization with convolutional neural networks, In Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 103-112, Denver, Colorado, May-June 2015, Association for Computational Linguistics; Yao Wang, Wan-dong Cai, and Peng-cheng Wei, A deep learning approach for detecting malicious JavaScript code, Sec. and Commun. Netw., 9(11):1520-1534, July 2016; and Hung Le, Quang Pham, Doyen Sahoo, and Steven CHHoi, Urlnet: Learning URL representation with deep learning for malicious URL detection, 2018. All inputs are represented at two levels of abstraction. A stream of characters (chars) and tokens. All URLs are padded to a fixed maximum size, but JS files are padded dynamically for each batch. For token-level representation, (B) uses a single-channel vocabulary of trained token embedding vectors selected based on frequency. For (C), we use Hung Le, Quang Pham, Doyen Sahoo, and Steven CHIt includes a char-by-word channel similar to Hoi, Urlnet: Learning a URL representation with deep learning for malicious URL detection, 2018. We additionally use an independently trained Hidden Markov Model to generate a randomness score for each token, which scales the learned embedding vector to generate a third randomness channel. When using IUPG, for JS.

number

number

number

[0112] Classification performance

[0113] We investigate classification performance on various datasets with different combinations of training and testing with noise. When training CCE correspondences with synthetic noise, we extend the dedicated noise class. For multi-class models, we define FP as off-target samples classified as any target class. We used a single confidence threshold for all target classes. Above this threshold, the most confident target class is predicted. When testing without noise in Table 2, the most confident target class is always predicted. Note also that in all tables except Table 2 in Figure 7B, the decision threshold is set to follow the maximum FPR. Results are shown across five trials with various random seeds, where applicable. As further explained below, we find reliably stable or reduced false negative rates (FNRs), error rates, and variability across Tables 1, 2, 3, and 4.

[0114] Figure 7A shows Table 1, which includes the malicious JS classification test set μ ± SEM FNR across all non-benign classes. For the multi-class model, the low-shot training dataset consists of 10 randomly selected samples per non-benign class. For the binary model, it is 1000 randomly selected malware samples.

[0115] Figure 7B shows Table 2, which contains the error percentages for the non-noise test set μ ± SEM for image classification.

[0116] Figure 7C shows Table 3, which contains the error rate for the test set μ ± SEM for image classification when the test set contains Gaussian noise images and the model is trained without noise. The decision threshold is set to comply with the maximum FPR.

[0117] Figure 7D shows Table 4, which contains the test set FNR for malicious URL classification.

number

number

[0118] Figure 7E shows Table 5, which contains the detections with a test set FPR threshold set at 0.005%, organized by VTS.

[0119] Note that Table 3 can also be interpreted as investigating the susceptibility to OOD attacks, assuming the model is trained without noise and then tasked with classifying a noisy test set. In Table 4, we see that IUPG better retains its performance in the presence of distribution shifts over a one-year period compared to CCE. Fitting the central hypothesis of IUPG, the noise-tolerance capabilities of the IUPG network are naturally more robust to distribution shifts of benign classes. The prototype definition strategy appears to affect performance. Clusters of malicious URLs are numerous and diverse. Intuitively, we would see benefits in defining a large number of prototypes using a bias set.

[0120] As further investigation of our central hypothesis, we trained a single IUPG model (e.g., Figure 6(B)) and a larger stacked ensemble of CCE networks for JS classification, ensuring that the test set FNR of the ensemble was lower than that of the IUPG model at the same FPR. We assembled a new collection of >5M JS samples taken from top-ranked popular websites. Using a test set FPR threshold below 0.005% (≤) FPR, we cross-referenced all detections from both models on this dataset with VT (see, e.g., Gaurav Sood, virustotal: R Client for the virustotal API, 2017, R package version 0.2.1). High VTS indicates strong consensus of malicious intent among multiple industry cybersecurity service providers. The VTS for both models' detections are shown in Table 5. Importantly, we observe a significant shift in the VT consensus for IUPG detections, despite the opposite performance gap on the test split. This is important to emphasize because constructing TTV splits from similar distributions often results in deploying models in more complex environments.

[0121] Out-of-Distribution (OOD) Attack Simulation

[0122] In addition to the exploration in Table 3, we explored the tendency of IUPG and its CCE counterpart to generate false positive (FP) responses for OOD inputs at a decision threshold, which represents the confidence level for in-distribution data. In this way, we take a closer look at the different tendencies of models to output similar confidence levels for OOD samples as for in-distribution samples. The results of our analysis are shown in Figure 8.

[0123] Figure 8 shows the results of an OOD attack simulation. The FPR was measured over the OOD test set with a decision threshold set based on the 75th-95th percentile of all confidence scores generated in the test data for the target class. (A) An image classification model trained without noise over the Gaussian OOD test set. (B) An image classification model trained without noise over the random-stroke OOD test set. (C) An image classification model trained with Gaussian noise over the random-stroke OOD test set. (D) A binary JS classifier over the randomized benign JS OOD test set.

[0124] When imposing a decision threshold that represents a confidence level typical of the target class in test data, we find, by a large margin, a smaller false positive rate (FPR) with IUPG. The lower tendency to generate FPs in OOD content allows for the use of looser decision thresholds in real-world systems, leading to higher recall. This result helps support the widening classification performance gap at tighter FPR requirements, as discussed above.

[0125] Append attack simulation

[0126] We explored the vulnerability of the JS malware classifier to append attacks. Simulation results are shown in Table 6, as further explained below. At each epoch, we also tried to dynamically change all benign classes, so that 33% of all members are appended with a random benign fragment with the same TTV split. The fragments are given a random size between 1000 and 5000 characters.

[0127] Figures 12A-12C show various examples of append attacks that can evade detection by existing malware detection solutions. Figures 12A and 12B illustrate why signature or hash matching is generally not sufficient for effective malware detection, and why advanced ML and DL models should also be deployed against "patient zero" malware (e.g., the malicious script in these examples). Both are examples of the same malicious campaign, which is difficult to catch because it generates many unique scripts and uses different obfuscation techniques for the injected pieces. Indeed, the example hash / SHA 256 shown in Figure 12B (i.e., cf9ac8b038e4a6df1c827dc31420818ad5809fceb7b41ef96cedd956a761afcd) was already known to VirtusTotal at the time of this writing, while the example hash / SHA 256 shown in Figure 12A (i.e., a248259f353533b31c791f79580f5a98a763fee585657b15013d1bb459734ba8) is new, i.e., not previously detected. Similar examples can be shown for malware injections that add whitespace and padding to attempt to fool existing ML classifiers.

[0128] In addition to redirectors and droppers, the publicly available IUPG framework is also effective and efficient in detecting JavaScript (JS) malware, such as phishing kits, clickjacking campaigns, malvertising libraries, and surviving exploit kits. For example, a similar script in Figure 12C (showing a sample of an obfuscated phishing JavaScript malware script in HTML, e.g., generating a fake Facebook login page) was discovered on over 60 websites, such as regalosycconcurso2021.blogspot.{al, am, bg, jp, ... com, co.uk}. Note that while the script uses heavy obfuscation techniques, it can nevertheless be accurately detected by a model trained on IUPG using the disclosed techniques, as described herein.

[0129] Figure 9 shows Table 6, which contains the simulation results of append attacks. Each cell is the percentage of malware where the model generates a malicious verdict for the original but benign verdict when appending a fragment of benign data of a given size in chars. 20 random fragments are tested for each malware. The decision threshold is set to follow a maximum FPR of 0.1% for the test set. Adversarial training is

number

[0130] We find a significant margin between the vulnerability of IUPG and its CCE counterparts, with or without adversarial training. Crucially, we need a binary approach to protect against append attacks that exceed the fragment size used during training.

number

[0131] Fast-Gradient Sign Method (FGSM) attack

[0132] To demonstrate the feasibility of combining IUPG with existing adversarial training techniques, we combined an image classifier with the Fast Gradient Sign Method (FGSM) training procedure (see, e.g., Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy, Explaining and Harnessing Adversarial Examples, arXiv e-prints, page arXiv:1412.6572, December 2014). We found that IUPG, with or without FGSM adversarial training, provides significantly higher resistance to FGSM attacks compared to its CCE counterpart. This is visualized in Figures 10A-10B.

[0133] Figures 10A-10B show the accuracy for correctly classified test images versus the scaling factor of the FGSM perturbations. Figure 10A shows the results using standard training, while Figure 10B shows the results when these models are trained with the FGSM training procedure described above.

[0134] Specifically, as shown in Figure 10B, we use typical FGSM training parameters of α = 0.9, ε = 0.25 on MNIST and α = 0.5, ε = 0.05 on Fashion-MNIST. Conforming to our core hypothesis, IUPG networks are, by design, less sensitive to low-level perturbations. This is due in particular to IUPG's prototyping mechanism, which encourages exclusive sensitivity to high-level information shared between subsets of data. Thus, both Figure 10B and Table 6 demonstrate that using IUPG in combination with specialized adversarial training outperforms either technique alone. Thus, combining strengths is generally recommended for maximum success in a variety of real-world environments, such as malware classification and similar applications of the disclosed IUPG technology.

[0135] Therefore, we present an IUPG training framework and demonstrate its impact on classification networks compared to CCE. Our core hypothesis is the boosted capacity to properly process OOD content provided by IUPG's inherent noise tolerance and increased feature specificity. This feature logically unites all the supportive results presented in this study: (1) improved or stable classification performance, (2) reduced performance loss due to recency bias, (3) reduced FP in OOD noise, and (4) reduced vulnerability to some noise-based attacks. Properly processing OOD content is generally important for models of real-world environments, such as malware classification, where benign content cannot reasonably be captured using finite samples.

[0136] As discussed above, IUPG's unique advantages are particularly useful for malware classification efforts. In that context, append attacks can lead to risky false negatives, while OOD failures can lead to costly false positives that can be handled or mitigated using the IUPG techniques described above. For example, IUPG has been shown to increase resistance to some attack types, as well as reduce FP responses to OOD inputs, as discussed above. Thus, IUPG can improve the efficiency and security of machine learning (ML) systems used in such ML systems, such as malware classification and detection using security platforms, as described herein. Black-box attack defense is particularly relevant to attacks utilized in proprietary systems, where an attacker can obtain benign data but not model details. Security of deep learning and ML in general is important for adoption and effectiveness, especially in safety-critical environments. Of particular relevance are cybersecurity service providers, who may directly benefit from the adoption of IUPG. These include improving the success rate of malware detection, increasing robustness against adversaries, enhancing customer trust, and furthering our shared ethical mission of keeping the digital world safe.

[0137] Example JavaScript (JS) data collection used in the experiment

[0138] Benign JS was collected from popular websites using filters. In particular, we used the top 1M domains from the Tranco list (e.g., publicly available at https: / / tranco-list.eu / ), which aggregates and cleans several popular lists and has thus been shown by researchers to be a more accurate and cleaner option (see, e.g., VLPochat, T. van Goethem, and W. Joosen, Rigging research results by manipulating top websites rankings, CorRR, abs / 1806.01156, 2018, available at http: / / arxiv.org / abs / 1806.01156). In addition to Tranco filtering, we ignored five samples flagged by state-of-the-art commercial URL filtering services. We used VirusTotal (VT) as our primary source of malicious JS samples (see, for example, G. Sood, virustotal: R Client for the virustotal API, 2017, URL https: / / www.virustotal.com, R package version 0.2.1). The problem is that VT's malicious file feed mostly contains HTML files rather than JS scripts. To accurately identify malicious scripts in HTML files, we extracted inline snippets and externally referenced scripts from the VT feed and resubmitted them to be verified by VT. We required at least three VT vendor hits, which empirically have been shown to be reasonably accurate. Data collection took place between 2014 and 2020. The most popular tokens among tags were "ExpKit," "Trojan," "Virus," "JS.Agent," and "HTML / Phishing."To supplement this malware data, we added malicious exploit kits kindly provided by major network and enterprise security companies. To source multi-class data for the malware family tagging problem, we isolated nine subsets of the malicious data. These subsets were determined through clustering the malware data using the method described in the section below on Using K-Means++ for IUPG. These clusters generally contain malware families and several obfuscation techniques whose outputs have a high degree of visual similarity. Each malware cluster is then described below.

[0139] 1. Angler Exploit KitThese samples aim to deliver malicious payloads via web browsers without any interaction from the victim (e.g., I. Nikolaev, M. Grill, and V. Valeros. Exploit kit website detection using http proxy logs, In Proceedings of the Fifth International Conference on Network, Communication and Computing, ICNCC 2016, pages 120-125, New York, NY, USA, 2016, Association for Computing Machinery, ISBN 9781450347938, doi: 10.1145 / 3033288.3033354. Available at https: / / doi.org / 10.1145 / 1323033288.3033354; B. Duncan. Understanding angler exploit kit - part 1: Exploit kit fundamentals, June 2016, URL (See https: / / unit42.paloaltonetworks.com / unit42-understanding-anglerexploit-kit-part-1-exploit-kit-fundamentals / ; F. Howard. A closer look at the angler-exploit kit, July 2019, available at https: / / news.sophos.com / en-us / 2015 / 07 / 21 / a-closer-look-at-the-angler-exploit-kit / ; A. Zaharia. The ultimate guide to angler exploit kit for non-technical people [updated], February 2017, URL https: / / heimdalsecurity.com / blog / ultimate-guide-angler-exploit kit-non-technical-people / ).It has high levels of both token and character (char) level randomness.

[0140] 2. "Hea2p" style obfuscation is often associated with phishing kits (see, e.g., O. Starov, Y. Zhou, and J. Wang, Detecting malicious campaigns in obfuscated javascript with scalable behavioral analysis, pages 218-223, May 2019, doi:10.1109 / SPW.2019.00048). It has high token-level similarity but high character-level randomness.

[0141] 3. Clickjacker (Clickjackers) focuses on generating artificial "like" button presses on social media websites (see id). High token-level similarity, but high character-level randomness. High token-level and character-level randomness.

[0142] 4. "Lololo" style obfuscation produces output with low token structure similarity but containing recognizable character-level patterns.

[0143] 5. Nemucod is a family of threats that attempt to download and install other malware on the device, including ransomware (see, for example, Microsoft. Js / nemucod, Mar 2015. URL https: / / www.microsoft.com / en-us / wdsi / threats / malware-encyclopedia-description?Name=JS%2FNemucod). It has a high degree of both token-level and character-level randomness.

[0144] 6. Various unnamed JS packersproduces output with high levels of both token-level and character-level randomness, assuming the packed code can reside anywhere in the original script.

[0145] 7. Various unnamed JS Trojans has high token-level similarity but high character-level randomness (see, e.g., C.E. Landwehr, A.R. Bull, J.P. McDermott, and W.S. Choi, A taxonomy of computer program security flaws, ACM Comput. Surv., 26(3) :211-254, September 1994, ISSN 0360-0300.doi:10.1145 / 185403.185412, URL https: / / doi.org / 10.1145 / 185403.185412).

[0146] 8. Another variety of unnamed JS Trojan has high token-level similarity but high character-level randomness (see id).

[0147] 9. Various anonymous encryption techniques produces output with high token-level similarity but high character-level randomness.

[0148] FIG. 11 is an example of a t-SNE visualization of the U-vector space, according to some embodiments. Specifically, the visualization shown in FIG. 11 is a real-world example of the output vector space after training a multi-class JS malware family classifier. The network was trained to recognize nine different JS malware families, listed in the legend, with off-target benign classes. Each of the nine target malware family classes is tightly grouped around one assigned prototype, while benign data is more arbitrarily mapped toward the center. This visualization was created using t-SNE on the mapped representation of the validation data and prototypes in the output vector space (see, e.g., van der Maaten & Hinton, GE (2008), Visualizing High-Dimensional Data Using t-SNE, Journal of Machine Learning Research, 9 (November), pages 2579-2605).

[0149] How to use K-Means++ for IUPG

[0150] K-means++ was utilized for our IUPG experiments (see, e.g., D. Arthur and S. Vassilvitskii, K-means++: The advantages of careful seeding, In Proceedings of the Eighteenth Annual ACM-SIAM Symposium on Discrete Algorithms, SODA 2007, pages 1027-1035, USA, 2007, Society for Industrial and Applied Mathematics, ISBN 9780898716245). Recall that clustering is used to discover intelligent IUPG prototype initializations. We used K-means++ (see, e.g., Y. LeCun and C. Cortes, MNIST database of handwritten digits, 2010, URL http: / / yann.lecun.com / exdb / mnist / ) and malicious JS datasets in this work.

[0151] MNIST clustering

[0152] We first grouped each digit class together in the training data and then clustered each subgroup separately. For each digit subgroup, we used principal component analysis (PCA) to project each image into the top 75 principal component vectors (see, e.g., I. Jolliffe and Springer-Verlag, Principal Component Analysis, Springer Series in Statistics, Springer, 2002, ISBN 9780387954424, URL https: / / books.google.com / books?id=_olByCrhjwIC). We performed K-means++ clustering with K = 1 across these compressed representations. We used Euclidean distance for clustering. We then calculated the Euclidean distance from the center of the resulting cluster to all images in the digit subgroup. The training image closest to the cluster center was selected as the prototype initialization. The selected image pixels were slightly perturbed with Gaussian noise to avoid potential overfitting.

[0153] Malicious JS Clustering

[0154] Since all benign data were assigned off-target labels, we first grouped the 54 malicious samples together. For the multi-class model, we grouped the malware families together and then clustered each one separately. For the binary model, we clustered all malware samples simultaneously. We further isolated only the token sequence representation of each malicious sample,

number

number

[0155] An Exemplary IUPG Framework for Malware JavaScript Classification

[0156] 13 illustrates the IPUG framework for malware JavaScript classification, according to some specific examples. A JavaScript document 1302 is tokenized using an OpenNMT Tokenizer 1304 (e.g., an open source tokenizer available at https: / / github.com / OpenNMT / Tokenizer) to produce Chars (Character Tokens) 1306 encoded as Char Encoding 1310 and Tokens 1308 encoded as Token Encoding 1312. Tokenization is followed by an ensemble of CNN feature extractors, including a CCE CNN feature extractor 1314 and an IUPG CNN feature extractor 1316 (e.g., implemented similarly to above using the disclosed IUPG techniques applicable to the JS malware classification context), followed by an XGB classifier 1318, which generates a JS malware classification verdict, as shown at 1320, based on the ensemble of CNN feature extractors using a combination of the CCE CNN and IUPG CNN classification techniques described above to facilitate a more effective and efficient JS malware detection solution (e.g., more robust against potential adversarial evasion techniques such as append attacks similar to those described above).

[0157] The disclosed IUPG framework and techniques can similarly be applied to URL categorization, such as URL filtering security solutions, and / or other computer / network security classification / detection for various security solutions as will now be apparent to those skilled in the art.

[0158] An exemplary process embodiment for performing the disclosed IUPG technique is further described below.

[0159] Exemplary Process Embodiments for Using a Presumed Innocent Model for Malware Classification

[0160] FIG. 14 illustrates an example process for performing static analysis of a sample using a presumption of innocence (IUPG) model for malware classification, according to some embodiments. In some embodiments, process 1400 is performed by security platform 122, and in particular by analyzer and detector 154. For example, analyzer and detector 154 can be implemented using a script (or set of scripts) written in a suitable scripting language (e.g., Python). In some embodiments, process 1400 is performed by data appliance 102, and in particular by threat engine 244. For example, threat engine 244 can be implemented using a script (or set of scripts) written in a suitable scripting language (e.g., Python). In some embodiments, process 1400 can also be performed on an endpoint, such as client device 110 (e.g., by an endpoint protection application running on client device 110). In some embodiments, process 1400 can also be performed by a cloud-based security service, such as using security platform 122, as described further below.

[0161] Process 1400 begins at 1402, where a set including one or more IUPG models for security analysis is stored on a network device. For example, IUPG models, such as JS code, HTML code, and / or other programming / scripting languages, as well as other structured text such as URLs or unstructured content such as images, can be generated (e.g., and / or periodically updated / replaced) based on training and validation data using the techniques described above.

[0162] At 1404, a static analysis of content associated with the received sample is performed at the network device using at least one stored IUPG classification model. As an example of processing performed at 1404, such as on the data device 102 and / or the client device 110, for a given session, an associated protocol decoder may call or otherwise utilize an appropriate file-specific decoder when the start of a file is detected by the protocol decoder. As described above, the file type is determined (e.g., by decoder 402) and associated with the session. In another example implementation, the file may be sent to a cloud-based security service (e.g., WildFire, a commercially available cloud security service offered by Palo Alto Networks, Inc.). TM Commercially available cloud-based security services, such as cloud-based malware analysis environments, include automated security analysis of malware samples as well as analysis by security experts, or similar solutions offered by other vendors may be used.

[0163] At 1406, it is determined whether the sample is malicious based at least in part on a static analysis of the content associated with the received sample. In one implementation, an appropriate IUPG model is used (e.g., applying an IUPG model to the JS code of a JS sample, applying an IUPG model to the HTML code of an HTML sample, etc.) to determine a class verdict for the file as malicious or benign (i.e., the final value obtained using the IUPG model is combined and compared with one or other classification models, such as a CCE CNN model similar to those described above, trained on appropriate content).

[0164] At 1408, an action is taken based on a security policy in response to the determination that the sample is malicious. Specifically, the action is taken in response to the determination at 1406. One example of a response action, in the case of the data device 102 and / or client device 110, is to terminate the session. Another example of a response action, in the case of the data device 102 and / or client device 110, is to allow the session to continue but prevent the file from being accessed and / or transmitted (and instead be placed in a quarantined area). Yet another example of a response action, in the case of the security platform 122, is to send a determination that the sample is malicious to the subscriber (e.g., the data device 102 and / or client device 110) that submitted the sample for analysis, notifying the subscriber that the sample was determined to be malicious so that the subscriber can take a response based on their locally configured security policy. In various embodiments, security platform 122, device 102, and / or client device 110 are configured to share their verdicts (benign verdicts, malicious verdicts, or both) with one or more other devices / platforms (e.g., security platform 122, device 102, and / or client device 110, etc.). As an example, once security platform 122 completes its independent analysis of the sample, the verdict reported by device 102 can be used for various purposes, including assessing the performance of the model that formed the verdict.

[0165] In one exemplary embodiment, security platform 122 is configured to target a particular false positive rate (e.g., 0.01%) when generating models for use by devices such as data device 102. Thus, in some cases (e.g., 1 in 1000 files), when performing inline analysis using a model according to the techniques described herein, data device 102 may erroneously determine a benign file to be malicious. In such a scenario, if security platform 122 subsequently determines that the file is in fact benign, it can add it to a whitelist so that it is not subsequently flagged (e.g., by another device) as malicious.

[0166] An embodiment of an exemplary process for building adversarial-resistant and false-positive-resistant deep learning models for security solutions

[0167] Figure 15 is an example of a process for generating a presumption of innocence (IUPG) model for malware classification, according to some embodiments. Specifically, one exemplary process for generating a presumption of innocence (IUPG) model for malware classification is shown in Figure 15. In various embodiments, process 15 is performed by security platform 122 (e.g., using model builder 152).

[0168] Process 1500 begins at 1502 when training data is received for training an IUPG model to classify malicious and benign content based on static analysis (e.g., the training data includes a set of files in an appropriate training context, such as JS files, HTML files, URLs, etc.).

[0169] 1504 extracts a set of tokens from a set of input files and generates a character encoding and a token encoding. As mentioned above, various techniques have been published for tokenizing content based on a set of characters and other tokens extracted from a set of input files, such as JS files.

[0170] At 1506, an IUPG CNN feature extractor is generated. Similar to above, additional feature vectors based on different abstraction levels / layers can also be generated based on different representations extracted from the set of input files.

[0171] At 1508, ensembling the IUPG CNN feature extractor with one or more other CNN-based feature extractors (e.g., a CCE CNN feature extractor or another form of CNN-based feature extractor) is performed to classify malicious and benign content based on static analysis of the sample. In one embodiment, following ensembling of the IUPG-based CNN feature extractor and the CCE-based CNN feature extractor, an XGB classifier is generated, such as to classify malicious and benign JS content based on static analysis of the sample, as similarly described.

[0172] Similarly, various IUPG models for one or more programming / scripting languages ​​or other content can be constructed using open source or other tools, and hyperparameter tuning, as described above, can be performed as needed to tune these IUPG models to perform efficiently, for example, for static analysis-based classification of samples executed / performed in various computing environments that may have different computing resources (e.g., memory resources, processor / CPU resources, etc. available to process these IUPG models). IUPG models (e.g., generated by model builder 152 using process 1500) can also be transmitted (e.g., as part of a subscription service) to data device 102, client device 110, and / or other applicable recipients (e.g., data devices 136 and 148, etc.).

[0173] In various embodiments, model builder 152 generates IUPG models (e.g., IUPG models for one or more types of source code, i.e., different programming / scripting languages ​​as described above, and / or other content) on a daily or other applicable / periodic basis. By performing process 1500 or otherwise generating models periodically, security platform 122 and / or cloud-based security services can help ensure that the various security classification models detect the latest types of malware threats (e.g., those recently deployed by nefarious individuals).

[0174] Although the above embodiments have been described in some detail for purposes of clarity of understanding, the present invention is not limited to the details provided. There are many alternative ways of implementing the present invention. The disclosed embodiments are illustrative and not restrictive.

Claims

1. 1. A system including a processor and a memory, The processor: storing a set including one or more presumption of innocence (IUPG) models for static analysis of samples on a network-connected device; performing a static analysis of content associated with the sample, wherein performing the static analysis of the content includes using at least one stored IUPG model and another type of CNN-based classifier, wherein the at least one stored IUPG model is selected based at least in part on a file type associated with the sample; determining that the sample is malicious based at least in part on the static analysis of the content associated with the sample, and performing an action based on a security policy in response to determining that the sample is malicious. It is structured as follows: The IUPG model is an adversarial-resistant and false-positive-resistant deep learning model, The memory includes: a processor coupled to the processor and configured to provide instructions to the processor; system.

2. The processor is further configured to receive at least one updated classification model. The system of claim 1 .

3. the processor is further configured to receive another IUPG model for another programming language. The system of claim 1 .

4. A method implemented by a data device, comprising: storing, by the data device, a set including one or more presumption of innocence (IUPG) models for static analysis of samples; performing, by the data device, a static analysis of content associated with the sample; performing the static analysis of the content includes using at least one stored IUPG model and another type of CNN-based classifier, the at least one stored IUPG model being selected based at least in part on a file type associated with the sample; Steps and determining, by the data device, that the sample is malicious based at least in part on the static analysis of the content associated with the sample; performing, by the data device, an action based on a security policy in response to determining that the sample is malicious; Including, The IUPG model is an adversarial-resistant and false-positive-resistant deep learning model. method.

5. The method further comprises: receiving, by said data device, at least one updated IUPG model; The method of claim 4, comprising:

6. The method further comprises: receiving, by said data device, another IUPG model for another programming language; The method of claim 4, comprising:

7. A computer program stored on a tangible computer-readable storage medium and including computer instructions, which, when executed, perform the following steps: storing a set including one or more presumption of innocence (IUPG) models for static analysis of samples; performing a static analysis of content associated with the sample; performing the static analysis of the content includes using at least one stored IUPG model and another type of CNN-based classifier, the at least one stored IUPG model being selected based at least in part on a file type associated with the sample; Steps and determining that the sample is malicious based at least in part on the static analysis of the content associated with the sample; performing an action based on a security policy in response to determining that the sample is malicious; and The IUPG model is an adversarial-resistant and false-positive-resistant deep learning model. Computer program.

8. The computer program further comprises: receiving at least one updated IUPG model; including computer instructions for 8. A computer program according to claim 7.

9. The computer program further comprises: receiving another IUPG model for another programming language; including computer instructions for 8. A computer program according to claim 7.

Citation Information

Patent Citations

  • Systems, devices, and methods for separating malware and background events

    JP2016091549A

  • Information processing device, information processing method, program, and computer-readable recording medium with program recorded

    JP2017162244A

  • Malware determining method, malware determining apparatus, malware determining program

    JP2018160172A

  • Sandbox environment for document preview and analysis

    US20190213325A1