Identification of.NET Malware with "Unmanaged IMPHASH"

By parsing the .NET header and calculating the unmanaged Imphash of .NET files, the system effectively addresses the limitations of existing methods for detecting malicious files, achieving improved accuracy and detection rates.

JP7694928B2Active Publication Date: 2025-06-18PALO ALTO NETWORKS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024529898
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-12-21
Filing Date
2022-12-05
Publication Date
2025-06-18
Estimated Expiration
2042-12-05

AI Technical Summary

Technical Problem

Existing methods for detecting malicious files, particularly .NET files, are unreliable due to similarities in PE file structures between malicious and benign files, leading to high false positives and low detection rates.

Method used

The system detects malicious .NET files by parsing the .NET header, extracting unmanaged imports, and calculating a hash of these imports, known as the unmanaged Imphash, which is then compared against a blacklist of known malicious files.

Benefits of technology

This approach provides accurate and low-latency detection of malicious .NET files, reducing false positives and improving detection rates compared to traditional PE header-based methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007694928000001
    Figure 0007694928000001
  • Figure 0007694928000002
    Figure 0007694928000002
  • Figure 0007694928000003
    Figure 0007694928000003
Patent Text Reader

Abstract

The present application provides a method, system, and computer system for detecting malicious files, the method including receiving a sample including a .NET file, obtaining imported API function names based at least in part on a .NET header of the .NET file, determining a hash of a list of unmanaged imported API function names, and determining whether the sample is malware based at least in part on the hash of the list of unmanaged imported API function names.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Malicious individuals attempt to compromise computer systems in various ways. As one example, such individuals may embed malicious software (“malware”) in an email attachment or otherwise include and send or cause to be sent malware to unsuspecting users. When executed, the malware compromises the victim's computer. Some types of malware instruct the compromised computer to communicate with a remote host. For example, the malware can turn the compromised computer into a “bot” in a “botnet” and receive instructions from, and / or report data to, a command and control (C&C) server under the control of a malicious individual. One approach to reducing the damage caused by malware is for a security company (or other appropriate entity) to attempt to identify the malware and prevent it from reaching and executing on an end-user computer. Another approach is to attempt to prevent the compromised computer from communicating with the C&C server. Unfortunately, malware authors are using increasingly sophisticated techniques to obfuscate the workings of their software. As one example, some types of malware use Domain Name System (DNS) queries to secretly exfiltrate data. Accordingly, there is a continuing need for improved techniques for detecting and preventing the harm of malware.

Brief Description of the Drawings

[0002] Various embodiments of the present invention are disclosed in the following detailed description and the accompanying drawings.

Figure 1

Figure 2

Figure 3A

Figure 3B

Figure 3C

Figure 3D

Figure 3E

Figure 3F

Figure 3G

Figure 4

Figure 5

Figure 6

Figure 7A

Figure 7B

Figure 8

Figure 9

Figure 10

Figure 11

[0003] The present invention can be implemented in a number of ways, including a process, an apparatus, a system, a composition, a computer program product embodied on a computer-readable storage medium, and / or instructions stored in a memory and / or instructions provided and / or stored by a memory coupled to a processor and executed by a processor including the processor. As used herein, these implementations, or any other form that the present invention may take, may be referred to as a technique. Generally, the order of steps of the disclosed process may be changed within the scope of the present invention. Unless otherwise specified, components such as a processor or a memory described as being configured to perform a task are implemented as a general-purpose component temporarily configured to perform the task at a given time or as a specific component manufactured to perform the task. As used herein, the term "processor" refers to one or more devices, circuits, and / or processing cores configured to process data, such as computer program instructions.

[0004] A detailed description of one or more embodiments of the present invention is provided below along with the accompanying drawings that illustrate the principles of the present invention. The present invention is described in relation to such embodiments, but the present invention is not limited to any particular embodiment. The scope of the present invention is limited only by the claims and the present invention encompasses numerous alternatives, modifications, and equivalents. In order to provide a complete understanding of the present invention, numerous specific details are set forth in the following description. These details are provided for illustrative purposes only and the present invention may be practiced according to the claims without some or all of these specific details. To clarify, technical material known in the technical field related to the present invention has not been described in detail so as not to unnecessarily obscure the present invention.

[0005] As used herein, a security entity is a network node (e.g., a device) that enforces one or more security policies with respect to information such as network traffic, files, etc. As one example, a security entity can be a firewall. As another example, a security entity can be implemented as a router, switch, DNS resolver, computer, tablet, laptop, smartphone, etc. As yet another example, security can be implemented as an application running on a device, such as an anti-malware application.

[0006] As used herein, malware refers to an application that engages in behavior that a user would not approve of / would not approve if fully informed, whether or not it is secret (and whether or not it is illegal). Examples of malware include Trojan horses, viruses, rootkits, spyware, hacking tools, keyloggers, etc. One example of malware is a desktop application that collects the end user's location and reports it to a remote server (but does not provide a location-based service such as a mapping service to the user). Another example of malware is a malicious Android (registered trademark) application package (Android Application Package).apk file that appears to be a free game to the end user but secretly (stealthily) sends SMS premium messages (e.g., each costing $10), increasing the end user's phone bill. Another example of malware is an Apple iOS flashlight application that secretly collects the user's contacts and sends them to spammers. Other forms of malware can also be detected / blocked using the techniques described herein (e.g., ransomware). Further, while malware signatures are described herein as being generated for malicious applications, the techniques described herein can also be used in various embodiments to generate profiles for other types of applications (e.g., adware profiles, goodware profiles, etc.).

[0007] As used herein, unmanaged code or an unmanaged function refers to an imported win32 API function, as contrasted with normal.NET code, which is called "managed code". As an example, such unmanaged code or unmanaged function is generally not reflected in / included in the PE header of a.NET file; rather, such unmanaged code or unmanaged function is imported via the.NET header of the.NET file.

[0008] According to the prior art, malware is identified using a machine learning model. The machine learning model according to the related art is trained / developed using the portable executable (PE) structure based on features such as imports, headers, and sections. The machine learning model uses such imports, headers, and sections to distinguish between malware and benign files. However, the PE file structure of Microsoft Windows® PE installer-based files appears to be very similar between malicious files and benign files. Therefore, using the PE file structure for Microsoft Windows PE installers to detect malware is not very reliable. This is because it is extremely difficult to distinguish between malicious files and benign files based on such PE file structures. For example, using the PE file structure to detect malicious files for Microsoft Windows PE installer files leads to higher false positives and lower detection rates. One example of a Microsoft Windows PE installer file used for benign purposes is the Microsoft Windows Nullsoft Scriptable Install System (NSIS) installer, which is widely used by legitimate products and in enterprise environments. Each machine learning model trained to analyze the PE structure to distinguish between malicious Microsoft Windows PE installer files and benign Microsoft Windows PE installer files cannot accurately detect malicious files.

[0009] A system, method, and / or device for detecting malicious files are disclosed. The system includes one or more processors and a memory coupled to the one or more processors and configured to provide instructions to the one or more processors. The one or more processors are configured to receive a sample including a.NET file, obtain API function names imported at least partially based on the.NET header of the.NET file, determine a hash of a list of unmanaged imported API function names, and determine whether the sample is malware based at least partially on the hash of the list of unmanaged imported API function names.

[0010] According to various embodiments, a system for detecting malicious files is implemented by one or more servers. The one or more servers may provide services to one or more customers and / or security entities. For example, the one or more servers may detect a malicious file or determine / evaluate whether a file is malicious and provide an indication as to whether the file is malicious to one or more customers and / or security entities. In response to determining that a file is malicious and / or in connection with updating the mapping of a file in response to an indication as to whether the file is malicious (e.g., an update to a blacklist including an identifier associated with the malicious file), the one or more servers provide an indication that the file is malicious to a security entity. As another example, in response to a request from a customer or security for an evaluation as to whether a file is malicious, the one or more servers determine whether the file is malicious and the one or more servers provide the result of such determination.

[0011] According to various embodiments, a system for detecting malicious files is implemented by a security entity. For example, a system for detecting malicious files is implemented by a firewall. As another example, a system for detecting malicious files is implemented by an application such as an anti-malware application running on a device (e.g., a computer, laptop, mobile phone, etc.). According to various embodiments, a security entity receives a.NET file, obtains the.NET header of the.NET file, and determines whether the.NET file is malicious based at least in part on the.NET header of the.NET file. In response to determining that the.NET file is malicious, the security entity applies one or more security entities to the.NET file. In response to determining that the.NET file is not malicious (e.g., the.NET file is benign), the security entity treats the.NET file as non-malicious traffic. In some embodiments, the security entity determines (e.g., obtains) the imported API function names based at least in part on the.NET header of the.NET file, determines (e.g., calculates) the hash of the list of unmanaged imported API function names, and determines whether the hash of the list of unmanaged imported API function names corresponding to the.NET file matches the hash associated with files considered malicious, at least in part based on this, to determine whether the file is malicious. For example, the security entity performs a lookup regarding the mapping of the hash (e.g., the hash of the unmanaged imported API function names) to malicious files to determine whether the mapping includes a matching hash (e.g., the mapping has a record of a file with a hash of unmanaged imported API function names that matches the calculated hash of the.NET file).

[0012] Portable executable (PE) files are often coded to import functions from external libraries in order to interact with various OS components. Related art methods for detecting malware use the import sequence, hash the import sequence to obtain a hash value, and then compare the hash value against a known block list of "Import Table Hashes" (imphash). Related art methods for detecting malware imports obtain the imported API function names and corresponding library names from the PE header of the file being analyzed. However, determining the API function names and corresponding library names from the PE header and using such API function names and corresponding library names to detect malware is not ideal for.NET files. This is because almost all.NET PE files have similar import tables. As an example, most.NET assemblies have a single imported function named "_CorExeMain" (EXE) or "_CorDllMain" (DLL) in the PE header. Generally, only a small portion of.NET assemblies have more imports in the PE header. Such.NET files are generally created using Visual C++ and C++ / CLI extensions. The imported functions contained in the PE header are generally determined by the.NET compiler and are not affected by the code itself. This phenomenon occurs because.NET code is not compiled into a native assembly but rather into intermediate language or intermediate bytecode (MSIL), which is then executed by the.NET runtime.

[0013] Therefore, the use of imported functions extracted from the PE header of a.NET file does not provide accurate detection of malware. However, various.NET malware families still need to interact directly with the win32 API, for example, to inject code into other processes. Such code can be injected from.NET into other processes, but the win32 functions for doing so are not reflected in the import table of the PE header for the.NET file. Rather, code injection functions are generally included (or imported) via the.NET header of the.NET file. The.NET header is a header included in the.NET file (in addition to, for example, the PE header). For example, the.NET header is different from the PE header of the.NET file. The.NET file contains both a PE header and a.NET header. The.NET header generally contains data streams and tables that contain various information about the.NET assembly. One such data stream contained in the.NET header is named "#Strings" and contains a list of strings used within the file. The list contained within the #Strings stream also includes the names of any unmanaged win32 API functions used. Additionally, one of the tables contained within the.NET header is named "ImplMap" and contains various information about any imported unmanaged functions.

[0014] Various embodiments parse the.NET header of a.NET file, extract unmanaged imports (e.g., unmanaged functions, libraries, etc.) from one or more fields within the.NET header, and determine whether the.NET file is malicious based at least in part on the extracted unmanaged imports. In some embodiments, the system determines a list of unmanaged imports corresponding to the.NET file (e.g., extracted from fields within the.NET header), and determines (e.g., calculates) a hash of the list of unmanaged imports. The hash of the list of unmanaged imports can be determined based on a predefined hash function. Examples of hash functions include the SHA-256 hash function, the MD5 hash function, the SHA-1 hash function, etc. Various other hash functions can be implemented. As used herein, unmanaged imphash refers to the value obtained by determining a hash of a list of unmanaged imports (e.g., unmanaged imports extracted from fields within the.NET header).

[0015] In accordance with various embodiments, the information contained in the.NET header of a.NET file is used in connection with determining whether the file is malicious. In some embodiments, the system uses the information contained in the ImplMap table and the information contained in the strings of the "#Strings" data stream to determine a set of unmanaged function <-> library name pairs. The system can determine a list of unmanaged import functions (e.g., those imported via the.NET header) for the.NET file. In some embodiments, the system determines a hash of the list of unmanaged import functions for the.NET file. For example, the system determines an unmanaged Imphash corresponding to the.NET file. The unmanaged Imphash can be used to determine whether the file is malicious. For example, the system can query a list of files considered malicious (e.g., a blacklist) to determine whether the list contains records having an unmanaged Imphash that matches the unmanaged Imphash determined for the.NET file.

[0016] In accordance with various embodiments, the system analyzes a.NET file in a sandbox environment. For example, the system parses the.NET file and extracts information from the.NET header within the sandbox environment. The system can be implemented by a virtual machine (VM) operating in the sandbox environment.

[0017] In some embodiments, the system receives historical information regarding the maliciousness of files (e.g., historical dataset of malicious files and historical dataset of benign files) from a third-party service such as VirusTotal(R). The third-party service may provide a set of files considered malicious and a set of files considered benign. As one example, the third-party service can analyze a file and provide an indication of whether the file is malicious or benign and / or a score indicating the likelihood that the file is malicious. The third-party service can provide the unmanaged Imphashes (e.g., blacklist of files, whitelist of files, etc.) corresponding to the files included in the historical dataset, or the list can include an indication of whether the historical unmanaged Imphashes are malicious. The system can receive updates from the third-party service such as newly identified benign or malicious files, corrections to previous mis-classifications, etc. (e.g., at a predefined interval when updates are available, etc.). In some embodiments, the indication is whether the files in the historical dataset correspond to a social score such as a community-based score or rating (e.g., reputation score) indicating that the file is malicious or likely to be malicious.

[0018] According to various embodiments, a security entity and / or a network node (e.g., a client, a device, etc.) handle a file based at least in part on an indication that the file is malicious and / or an indication that the file matches a file indicated to be malicious. In response to receiving an indication that a file (e.g., a sample) is malicious, the security network and / or the network node can update a mapping for an indication as to whether the corresponding file is malicious and / or a blacklist of files. In some embodiments, the security entity and / or the network node receive a signature for a file (e.g., a sample deemed malicious), and the security entity and / or the network node store the signature of the file for use in detecting whether the retrieved file is malicious (e.g., based at least in part on comparing the signature generated for the file to the signatures for files included in the blacklist of files) via, for example, network traffic. As one example, the signature can be a hash. In some embodiments, the signature of the file is an unmanaged Imphash corresponding to such a file.

[0019] A firewall typically rejects or permits network transmissions based on a set of rules. This set of rules is often referred to as a policy (e.g., a network policy or a network security policy). For example, a firewall can filter inbound traffic by applying a set of rules or policies to prevent unwanted external traffic from reaching a protected device. A firewall can also filter outbound traffic by applying a set of rules or policies (e.g., allow, block, monitor, notify, or log, and / or other actions that can be specified in a firewall rule or a firewall policy, which can be triggered based on various criteria as described herein). A firewall can also filter local network (e.g., intranet) traffic by applying a set of rules or policies in a similar manner.

[0020] A security device (e.g., a security appliance, a security gateway, a security service, and / or other security devices) can perform various security operations (e.g., firewall, anti-malware, intrusion prevention / detection, proxy, and / or other security functions), network functions (e.g., routing, quality of service (QoS), workload balancing of network-related resources, and / or other network functions), and / or other security and / or network-related functions. For example, routing can be performed based on source information (e.g., IP address and port), destination information (e.g., IP address and port), and protocol information.

[0021] A basic packet filtering firewall filters network communication traffic by inspecting individual packets transmitted over the network (e.g., a packet filtering firewall or first-generation firewall, which is a stateless packet filtering firewall). A stateless packet filtering firewall typically inspects the individual packets themselves and applies rules based on the inspected packets (e.g., using a combination of source and destination address information, protocol information, and port numbers of the packets).

[0022] An application firewall can also perform application layer filtering (e.g., using an application layer filtering firewall or a second-generation firewall that functions at the application level of the TCP / IP stack). An application layer filtering firewall or application firewall can generally identify given applications and protocols (e.g., web browsing using the Hypertext Transfer Protocol (HTTP), Domain Name System (DNS) requests, file transfers using the File Transfer Protocol (FTP), and various other types of applications and other protocols such as Telnet, DHCP, TCP, UDP, and TFTP (GSS)). For example, an application firewall can block unauthorized protocols attempting to communicate on standard ports (e.g., unauthorized / rogue protocols that attempt to sneak through by using non-standard ports for that protocol can generally be identified using an application firewall).

[0023] A stateful firewall can also perform stateful - based packet inspection, where each packet is inspected within a set of packet contexts related to the packet flow of its network transmission. This firewall technology is generally referred to as stateful packet inspection. This is because it can maintain a record of all connections passing through the firewall and determine whether a packet is the start of a new connection, part of an existing connection, or an invalid packet. For example, the state of a connection can itself be one of the criteria that trigger rules in the policy.

[0024] As described above, advanced or next-generation firewalls can perform stateless and stateful packet filtering and application layer filtering. Next-generation firewalls can also perform additional firewall technologies. For example, a given new firewall, often referred to as an advanced or next-generation firewall, can also identify users and content. In particular, given next-generation firewalls have expanded the list of applications that these firewalls can automatically identify to thousands of applications. Examples of such next-generation firewalls are commercially available from Palo Alto Networks (e.g., the PA series of firewalls from Palo Alto Networks). For example, the next-generation firewalls from Palo Alto Networks use various identification technologies to enable enterprises and service providers to identify and control applications, users, and content - not just ports, IP addresses, and packets. The various identification technologies include Application ID (App-ID) for accurate application identification, User ID (User-ID) for user identification (e.g., User ID), Content ID (Content-ID) for real-time content scanning (e.g., to control web surfing and restrict data and file transfers), and Device ID (Device-ID) (e.g., for identifying IoT device types). With these identification technologies, enterprises can use business-related concepts to safely enable the use of applications, rather than following the traditional approach provided by conventional port-blocking firewalls.Also, specialized hardware for next-generation firewalls (e.g., implemented as a dedicated device) generally provides a higher performance level for application inspection than software executed on general-purpose hardware (e.g., security appliances such as those provided by Palo Alto Networks, which utilize dedicated, function-specific processing tightly integrated with a single-pass software engine to minimize latency and maximize network throughput for Palo Alto Networks' PA series next-generation firewalls).

[0025] Advanced or next-generation firewalls can also be implemented using virtualized firewalls. Examples of such next-generation firewalls are commercially available from Palo Alto Networks (Palo Alto Networks' firewalls support various commercial virtualization environments including VMware(R) ESXi TM and NSX TM , Citrix(R) Netscaler SDX TM , KVM / OpenStack (Centos / RHEL, Ubuntu(R)), and Amazon Web Services (AWS)). For example, virtualized firewalls can support the same or fully identical next-generation firewall and advanced threat prevention functions available on physical form factor devices, enabling enterprises to securely enable the flow of applications into private, public, and hybrid cloud computing environments. Through automation functions such as VM monitoring, dynamic address groups, and REST-based APIs, enterprises can dynamically monitor changes to VMs and reflect their context in security policies, thereby eliminating policy lag that can occur during VM changes.

[0026] The system improves the detection of malicious files. Further, the system further improves the handling of network traffic by preventing (or improving the prevention of) malicious files from traversing a network such as between nodes within the network, or by preventing malicious files from entering the network. The system determines.NET files that are considered malicious or likely to be malicious, such as based on the.NET headers of the.NET files. Detection techniques of related art that use the structure of the PE header for files can be insufficient / inaccurate with respect to files having similar structures / profiles among malicious or benign files. Further, since.NET files are compiled into intermediate language, it is difficult to classify files as malicious / benign using a machine learning classifier or manually written YARA rules. YARA is a tool aimed at (but not limited to) helping malware researchers identify and classify malware samples. YARA rules are used to classify and identify malware samples by creating a description of the malware family based on text or binary patterns. Further, the system can provide accurate and low-latency updates to security entities (such as endpoints, firewalls, etc.) to enforce one or more security policies (such as default and / or customer-specific security policies) with respect to traffic containing malicious files (such as malicious.NET files). Accordingly, the system prevents the spread of malicious traffic (such as files) to nodes within the network.

[0027] FIG. 1 is a block diagram of an environment in which malicious files are detected or suspected, according to various embodiments. In the example shown, client devices 104-108 are laptop computers, desktop computers, and tablets (respectively) that exist within enterprise network 110 (which belongs to “Acme Company”). Data appliance 102 is configured to enforce policies (e.g., security policies) regarding communications between client devices, such as client devices 104 and 106, and nodes external to enterprise network 110 (e.g., those reachable via external network 118). Examples of such policies include those that manage traffic shaping, quality of service, and routing of traffic. Other examples of policies include security policies that require scanning for threats in incoming (and / or outgoing) email attachment files, website content, files exchanged via instant messaging programs, and / or other file transfers. In some embodiments, data appliance 102 is also configured to enforce policies regarding traffic that stays within (or enters) enterprise network 110.

[0028] The techniques described herein can be used with a variety of platforms (e.g., desktops, mobile devices, game platforms, embedded systems, etc.) and / or a variety of types of applications (e.g., Android.apk files, iOS applications, Windows PE files, Adobe Acrobat PDF files, Microsoft Windows PE installers, etc.). In the exemplary environment shown in FIG. 1, client devices 104-108 are laptop computers, desktop computers, and tablets (respectively) that exist within enterprise network 140. Client device 120 is a laptop computer that exists outside of enterprise network 110.

[0029] Data appliance 102 may be configured to operate in cooperation with remote security platform 140. The security platform 140 can provide various services. Performing static and dynamic analysis on malware samples, providing a list of signatures of known malicious files to data appliances such as data appliance 102 as part of a subscription, detecting malicious files (e.g., on-demand detection, or regular basis updates to the mapping of files to an indication of whether the file is malicious or benign), providing the likelihood that a file is malicious or benign, providing / updating a whitelist of files considered benign, identifying malicious domains, detecting malicious files, predicting whether a file is malicious or not, and providing an indication that a file is malicious (or benign). In various embodiments, the results of the analysis (and additional information regarding applications, domains, etc.) are stored in database 160. In various embodiments, the security platform 140 comprises one or more dedicated off-the-shelf hardware servers (e.g., having a multi-core processor, 32G+ of RAM, gigabit network interface adapter, and hard drive) that execute a typical server-class operating system (e.g., Linux (registered trademark)). The security platform 140 may be implemented across a scalable infrastructure including multiple such servers, solid state drives, and / or other applicable high-performance hardware. The security platform 140 may comprise several distributed components including components provided by one or more third parties. For example, part or all of the security platform 140 may be implemented using Amazon Elastic Compute Cloud (EC2), and / or Amazon Simple Storage Service (S3).Furthermore, whenever the security platform 140 is referred to as performing tasks such as storing or processing data, similar to the data appliance 102, it should be understood that sub-components or multiple sub-components of the security platform 140 can cooperate (either individually or in cooperation with third-party components) to perform that task. As one example, the security platform 140 can optionally perform static / dynamic analysis in cooperation with one or more virtual machine (VM) servers. One example of a virtual machine server is a physical machine that runs commercially available virtualization software (such as VMware ESXi, Citrix XenServer, or Microsoft Hyper-V) and includes commercially available server-class hardware (e.g., a multi-core processor, 32+ gigabytes of RAM, and one or more gigabit network interface adapters). In some embodiments, the virtual machine server is omitted. Further, the virtual machine server can be under the control of the same entity that manages the security platform 140, or it can be provided by a third party. As one example, the virtual machine server can rely on EC2, and the remainder of the security platform 140 can be provided by dedicated hardware that is owned by and under the control of the operator of the security platform 140.

[0030] According to various embodiments, the security platform 140 includes a DNS tunneling detector 138 and / or a malicious file detector 170. The malicious file detector 170 is used in connection with determining whether a file (e.g., a.NET file) is malicious. In response to receiving a sample, the malicious file detector 170 analyzes the file and determines whether the file is malicious. For example, the malicious file detector 170 determines whether an unmanaged Imphash corresponding to the file being analyzed matches a file included in a historical dataset (e.g., a list of files considered malicious, a list of files considered benign, etc.). In some embodiments, the malicious file detector 170 receives a sample including a.NET file, obtains imported API function names based at least in part on the.NET header of the.NET file, determines a hash of a list of unmanaged imported API function names (e.g., unmanaged Imphash), and determines whether the sample is malware based at least in part on the hash of the list of unmanaged imported API function names. In some embodiments, the malicious file detector 170 comprises one or more of a.NET file parser 172, an unmanaged function extractor 174, a prediction engine 176, and / or a cache 178.

[0031] .NET file parser 172 is used in relation to obtaining information about samples such as.NET files. In some embodiments, the.NET file parser 172 obtains the.NET header of the.NET file and / or information from the.NET header. The.NET file parser 172 obtains one or more data streams included in the.NET header and / or one or more tables included in (or referenced by) the.NET header. For example, the.NET file parser 172 obtains the #Strings stream included within the.NET header. As another example, the.NET file parser 172 obtains an ImplMap from the.NET file (e.g., from the.NET header). In some embodiments, the.NET file parser 172 determines a set of imported functions (e.g., imported API function names) that are imported into the.NET file.

[0032] The unmanaged function extractor 174 is used in relation to determining (e.g., obtaining) a set of unmanaged import functions that are imported into the.NET file. For example, the unmanaged function extractor 174 determines a set of unmanaged import functions from a set of imported functions (e.g., imported API function names) that are imported into the.NET file. In accordance with various embodiments, the unmanaged function extractor 174 uses information included in (or referenced by) the.NET header to determine unmanaged code or unmanaged functions (e.g., unmanaged functions and corresponding libraries). In some embodiments, the unmanaged function extractor 174 determines a set of used unmanaged win32 API functions that are imported into the.NET file. The unmanaged function extractor 174 provides a set of unmanaged import functions (or names of unmanaged functions) that are imported into the.NET file and / or corresponding libraries to the prediction engine 176.

[0033] The prediction engine 176 is used to determine whether a file (e.g., a.NET file) is malicious. The prediction engine 176 uses the information contained in the.NET header in connection with determining whether the corresponding.NET file is malicious. For example, the prediction engine 176 obtains a set of unmanaged import functions (or the names of unmanaged functions), and / or the corresponding libraries imported into the.NET file from the unmanaged function extractor 174. In some embodiments, the prediction engine 176 determines a set of unmanaged import functions (or the names of unmanaged functions) imported into the.NET file, and / or a hash (e.g., a hash value) for the corresponding libraries. For example, the prediction engine 176 calculates the unmanaged Imphash corresponding to the.NET file. The prediction engine 176 determines a list of a set of unmanaged import functions (or the names of unmanaged functions) imported into the.NET file, and / or a set of corresponding libraries, and then determines the hash of such a list. The list of unmanaged import functions and / or the set of corresponding libraries is determined according to a default order. For example, the ordering of the imported unmanaged functions and / or the corresponding libraries corresponds to the order in which the unmanaged functions are included in the elements of the.NET header (e.g., the order in which the unmanaged functions are included in the.NET table or the #Strings stream and / or table referenced thereby). Various other orders in which the unmanaged functions are added to the list (or the list is constructed) may be implemented. The prediction engine 176 formats the list and / or the unmanaged functions (e.g., unmanaged function names and / or the corresponding libraries) according to a default format.Examples of the established format include: (i) a lower-case alphanumeric string, (ii) removal of the file extension, (iii) removal of the library extension (e.g., removal of.dll from the corresponding library name), (iv) addition of the unmanaged function name and the corresponding library, and (v) use of a predefined separator between the unmanaged function name and the corresponding library. In some embodiments, the prediction engine 176 adds the function name (e.g., the unmanaged function name) to the corresponding library and separates the function name from the corresponding library by a dot or period (e.g., "."). In response to determining the string by adding the function name and library (e.g., using a predefined separator such as "."), the prediction engine 176 adds such an entry to the list of unmanaged import functions and / or the set of corresponding libraries.

[0034] According to various embodiments, the prediction engine 176 determines a hash for the list of unmanaged import functions and / or the set of corresponding libraries. In connection with determining the hash, various hash functions can be used. Examples of hash functions include the SHA-256 hash function, the MD5 hash function, the SHA-1 hash function, etc. Various other hash functions can be implemented. The prediction engine 176 uses the hash function to determine the unmanaged Imphash corresponding to the.NET file.

[0035] According to various embodiments, the prediction engine 176 uses information obtained from the.NET header of the.NET file to determine whether the.NET file is malicious. In some embodiments, the prediction engine 176 uses the unmanaged Imphash corresponding to the.NET file in relation to determining whether the.NET file is malicious. As an example, in response to determining the unmanaged Imphash corresponding to the.NET file, the prediction engine 176 determines whether the unmanaged Imphash matches the unmanaged Imphash of a file considered to be malicious. As an example, in response to determining the unmanaged Imphash corresponding to the.NET file, the prediction engine 176 determines whether the unmanaged Imphash matches the unmanaged Imphash of a file considered to be benign. In some embodiments, the malicious file detector 170 (e.g., the prediction engine 176) determines whether information about a particular file (e.g., the unmanaged Imphash corresponding to the.NET file being analyzed) is in a dataset of history files and a history dataset indicating whether the particular file is malicious (e.g., VirusTotal TMDetermine whether it is included in the historical information related to (such as third-party services). In response to determining that the information about a specific file is not included in or not available in the historical file and the dataset of historical information, the malicious file detector 170 may consider the file to be benign (for example, consider the file not to be malicious). One example of the historical information associated with the historical file indicating whether a specific file is malicious corresponds to the VirusTotal (R) (VT) score. If the VT score for a specific file is greater than 0, that specific file is considered malicious by the third-party service. In some embodiments, the historical information associated with the historical file indicating whether a specific file is malicious corresponds to a social score such as a community-based score or a rating (for example, a reputation score) indicating that the file is malicious or likely to be malicious. The historical information (from, for example, third-party services, community-based scores, etc.) indicates whether other vendors or cybersecurity organizations consider a specific file to be malicious.

[0036] In some embodiments, a malicious file detector 170 (e.g., prediction engine 176) determines that a received file is newly analyzed (e.g., the file is not in the history information / dataset, not on the whitelist or blacklist, etc.). The malicious file detector 170 (e.g., script extraction module 172) can detect that a file is newly analyzed in response to the security platform 140 receiving the file from a security entity (e.g., firewall) or endpoint within the network. For example, the malicious file detector 170 determines that the file is newly analyzed at the same time as the security platform 140 or the malicious file detector 170 receives the file. As another example, the malicious file detector 170 (e.g., prediction engine 176) determines that a file is newly analyzed according to a predefined schedule (e.g., daily, weekly, monthly, etc.), such as being related to a batch process. In response to determining that such a file has not yet been analyzed as to whether it is malicious (e.g., the system does not have historical information about such a file), the malicious file detector 170 determines whether to use the.NET header related to the file in connection with determining whether the file is malicious (e.g., in response to determining that the file is a.NET file), and the malicious file detector 170 uses a.NET file parser to analyze and / or extract information about the.NET file from the.NET header of the.NET file. In some embodiments, the.NET file parser 172 extracts information from the.NET header within the system's sandbox environment.

[0037] In accordance with various embodiments, in response to the prediction engine 176 determining that a file is malicious, the system transmits an indication that the file is malicious to a security entity (or an endpoint such as a client). For example, the malicious file detector 170 transmits an indication that the file is malicious to a security entity (e.g., a firewall) or a network node (e.g., a client). The indication that the file is malicious may correspond to an update to a blacklist of files (e.g., corresponding to malicious files) if the file is considered malicious, or an update to a whitelist of files (e.g., corresponding to non-malicious files) if the file is considered benign. In some embodiments, the malicious file detector 170 transmits a hash or signature corresponding to the file in connection with an indication that the file is malicious or benign. The security entity or endpoint may calculate the hash or signature of the file and perform a lookup against a mapping of the hash / signature to an indication of whether the file is malicious / benign (e.g., query a whitelist and / or blacklist). In some embodiments, the hash or signature uniquely identifies the file.

[0038] Cache 178 stores information about files. In some embodiments, Cache 178 stores a mapping to a particular file of an indication of whether the file is malicious (or likely to be malicious), or a mapping to a hash or signature corresponding to the file of an indication of whether the file is malicious (or likely to be malicious). Cache 178 can store additional information about a set of files. Such as script information of files within the set of files, hashes or signatures corresponding to the files within the set of files, other unique identifiers corresponding to the files within the set of files, executable files called by the files, Bitcoin wallets called by the files, pointers included in the files, and the like.

[0039] Returning to FIG. 1, assume that a malicious individual (using System 120) created malware 130. The malicious individual wants a client device, such as client device 104, to execute a copy of malware 130, expose the client device to danger, and turn the client device into a bot within a botnet. The exposed client device can then be instructed to execute tasks (e.g., cryptocurrency mining, or participate in a service disruption attack), and / or report information to an external entity such as a command and control (C&C) server 150 (e.g., extract confidential corporate data associated with such tasks), and, where applicable, receive instructions from C&C server 150.

[0040] Malware 130 can attempt to directly communicate the compromised client device with the C&C server 150 (e.g., by causing the client to send an email to the C&C server 150), but such obvious communication attempts are flagged (e.g., by the data appliance 102) as suspicious / harmful and can be blocked. Instead of generating such direct communication, the malware author uses a technique called DNS tunneling herein. DNS is a protocol that converts human-friendly URLs such as paloaltonetworks.com into machine-friendly IP addresses such as 199.167.52.137. DNS tunneling exploits the DNS protocol to tunnel malware and other data through the client-server model. In an example of an attached file, a malicious file (e.g., malware) is sent as an attached file to a message such as an email, an instant message, etc. Once the attached file is selected, the malware program can be installed on the client device. In one example of an attack, the attacker registers a domain such as badsite.com. The domain name server points to the attacker's server on which the tunneling malware program is installed. The attacker infects a computer. Since DNS requests have conventionally been permitted to move in and out of security devices, the infected computer is permitted to send a query to a DNS resolver (e.g., kj32hkjqfeuo32ylhkjshdflu23.badsite.com, where the subdomain part of the query encodes information for consumption by the C&C server). A DNS resolver is a server that relays requests for IP addresses to the root and top-level domain servers. The DNS resolver routes the query to the attacker's C&C server on which the tunneling program is installed. A connection is now established between the victim and the attacker via the DNS resolver.This tunnel can be used to extract data or for other malicious purposes.

[0041] Detecting and preventing DNS tunneling attacks is difficult for various reasons. Many legitimate services (e.g., content delivery networks, web hosting companies, etc.) legitimately use the subdomain part of a domain name to encode information to help support the use of those legitimate services. The encoding patterns used by such legitimate services can vary widely between providers, and benign subdomains may appear visually indistinguishable from malicious subdomains. A second reason is that, unlike other areas (e.g., computer research) that have large corpora of training set data for both known benign and known malicious, the training set data for DNS queries is highly skewed (e.g., having millions of examples of benign root domains and a very small number of malicious examples). Despite such difficulties, and using the techniques described herein, malicious domains can be detected efficiently and proactively (e.g., immediately after domain registration), and security policies can be enforced regarding malicious files within or entering the network, and such malicious files can be blocked or, otherwise, the user or administrator of the malicious file can be warned (e.g., by sending a notification, providing a prompt to the user, etc.).

[0042] The environment shown in FIG. 1 includes three Domain Name System (DNS) servers (122-126). As shown, DNS server 122 is under the control of ACME (for use by computing assets located within network 110), while DNS server 124 is publicly accessible (and can also be used by other devices such as computing assets located within network 110 and those located within other networks (e.g., networks 114 and 116)). DNS server 126 is publicly accessible but under the control of the malicious operator of C&C server 150. Enterprise DNS server 122 is configured to resolve enterprise domain names to IP addresses and is further configured to communicate with one or more external DNS servers (e.g., DNS servers 124 and 126) to resolve domain names when applicable.

[0043] As described above, in order to connect to a legitimate domain (e.g., www.example.com shown as Site 128), a client device, such as client device 104, must resolve the domain to the corresponding Internet Protocol (IP) address. One way such resolution can be performed is for client device 104 to forward a request to DNS server 122 and / or 124 to resolve the domain. In response to receiving a valid IP address for the requested domain name, client device 104 can use the IP address to connect to website 128. Similarly, in order to connect to malicious C&C server 150, client device 104 must resolve the domain "kj32hkjqfeuo32ylhkjshdflu23.badsite.com" to the corresponding Internet Protocol (IP) address. In this example, malicious DNS server 126 has authority for *.badsite.com, and the request from client device 104 is forwarded to DNS server 126 for resolution, ultimately enabling C&C server 150 to receive data from client device 104.

[0044] Data appliance 102 is configured to enforce policies regarding communications between client devices, such as client devices 104 and 106, and nodes external to enterprise network 140 (e.g., those reachable via external network 118). Examples of such policies include those that manage traffic shaping, quality of service, and routing of traffic. Other examples of policies include security policies that require scanning for threats in incoming (and / or outgoing) email attachment files, website content, files exchanged via instant messaging programs, and / or other file transfers. In some embodiments, data appliance 102 is also configured to enforce policies regarding traffic remaining within enterprise network 140.

[0045] In various embodiments, data appliance 102 includes a DNS module 134 configured to facilitate determining whether client devices (e.g., client devices 104 - 108) are attempting to engage in malicious DNS tunneling and / or to prevent connections to malicious DNS servers (e.g., by client devices 104 - 108). The DNS module 134 can be integrated into the appliance 102 (as shown in FIG. 1) and, in various embodiments, can also operate as a stand-alone appliance. And, similar to the other components shown in FIG. 1, the DNS module 134 can be provided by the same entity that provides the appliance 102 (or security platform 140) and can also be provided by a third party (e.g., one different from the provider of the appliance 102 or security platform 140). Further, in addition to preventing connections to malicious DNS servers, the DNS module 134 can take other actions such as individualized logging of tunneling attempts made by clients (indicating that a given client is at risk and should be quarantined or otherwise investigated by an administrator).

[0046] In various embodiments, when a client device (e.g., client device 104) attempts to resolve a domain, the DNS module 134 uses the domain as a query to the security platform 140. This query can be executed concurrently with the resolution of the domain (e.g., concurrently with requests sent to DNS servers 122, 124, and / or 126, and the security platform 140). As one example, the DNS module 134 can send a query (e.g., in JSON format) to the front end 142 of the security platform 140 via a REST API. Using the processes described in more detail below, the security platform 140 determines (e.g., using the DNS tunneling detector 138) whether the queried domain exhibits an attempt at malicious DNS tunneling, and returns the result (e.g., "malicious DNS tunneling" or "non-tunneling") to the DNS module 134.

[0047] In various embodiments, when a client device (e.g., client device 104) attempts to open a received file, such as via an attachment to an email or an instant message, or when the client device receives such a file, the DNS module 134 uses the file (i.e., the calculated hash or signature, or other unique identifier, etc.) as a query to the security platform 140. This query can be executed simultaneously with the receipt of the file or in response to a user request to scan the file. As one example, the data appliance 102 can send a query (e.g., in JSON format) to the front end 142 of the security platform 140 via a REST API. Using the processes described in more detail below, the security platform 140 determines whether the queried file is a malicious file (or is likely to be a malicious file) (e.g., using the malicious file detector 170), and returns the result (e.g., "malicious DNS tunneling" or "non-tunneling") to the DNS module 134.

[0048] In various embodiments, the DNS tunneling detector 138 (implemented either on the security platform 140, on the data appliance 102, or at any suitable location / location combination) uses a two - pronged approach when identifying malicious DNS tunneling. The first approach uses an anomaly detector 146 (implemented, for example, using python) to build a set of real - time profiles (156) of DNS traffic for root domains. The second approach uses signature generation and matching (also referred to herein as similarity detection and implemented, for example, using Go). The two approaches are complementary. The anomaly detector functions as a general - purpose detector that can identify tunneling traffic that was previously unknown. However, the anomaly detector may need to observe multiple DNS queries before detection can occur. To block the first DNS tunneling packet, the similarity detector 144 complements the anomaly detector 146 and extracts a signature from the detected tunneling traffic. It can be used to identify situations where an attacker has registered a new malicious tunneling root domain using tools / malware similar to the detected root domain.

[0049] When the data appliance 102 receives DNS queries (from, for example, the DNS module 134), the data appliance 102 provides them to the security platform 140, which performs both anomaly detection and similarity detection, respectively. In various embodiments, a domain (such as provided in a query received by the security platform 140) is classified as a malicious DNS tunneling root domain if either detector flags the domain.

[0050] The DNS tunneling detector 138 maintains a set of fully qualified domain names (FQDNs) that are grouped by their root domains (collectively shown as domain profiles 156 in FIG. 1) for each device from which data is received. (Through root-based grouping, domains are generally described herein, but it should be understood that the techniques described herein can also be extended to any level of domain.) In various embodiments, information regarding received queries for a given domain is persisted within the profile for a fixed amount of time (e.g., a 10-minute sliding time window).

[0051] As an example, DNS query information received from the data appliance 102 for various foo.com sites is grouped as follows (into a domain profile for the root domain foo.com). G(foo.com) = [mail.foo.com, coolstuff.foo.com, domain1234.foo.com]. The second root domain will have a second profile with similar application information (e.g., G(baddomain.com) = [lskjdf23r.baddomain.com, kj235hdssd233.baddomain.com]). Each root domain (e.g., foo.com or baddomain.com) is modeled using a set of characteristics unique to malicious DNS tunneling, such that even if the benign DNS patterns are diverse (e.g., k2jh3i8y35.legitimatesite.com, xxx888222000444.otherlegitimatesite.com), the likelihood that they will be misclassified as malicious tunneling is very low. The following are exemplary characteristics that can be extracted as features for a given group of domains (i.e., sharing a root domain) (e.g., into a feature vector).

[0052] In some embodiments, the malicious file detector 170 provides an indication as to whether a file is malicious to a security entity such as the data appliance 102. For example, in response to determining that a file is malicious, the malicious file detector 170 transmits an indication that the file is malicious to the data appliance 102, and the data appliance may then implement one or more security policies at least in part based on the indication that the file is malicious. The one or more security policies may include isolating the file, deleting the file, warning or prompting the user of the maliciousness of the file before the user opens / executes the file, etc. As another example, in response to determining that a file is malicious, the malicious file detector 170 provides an update to the security entity of a mapping of the file (or hash, signature, unmanaged Imphashes, or other unique identifier corresponding to the file) to the indication as to whether the corresponding file is malicious, or an update to a blacklist for malicious files (e.g., identifying the file domain) or a whitelist for benign files (e.g., identifying files not considered malicious).

[0053] FIG. 2 is a block diagram of a system for detecting malicious files according to various embodiments. According to various embodiments, the system 200 is implemented in connection with the system 100 of FIG. 1, such as for the malicious file detector 170. In various embodiments, the system 200 is implemented in connection with the process 400 of FIG. 4, the process 500 of FIG. 5, the process 600 of FIG. 6, the process 700 of FIG. 7A, the process 750 of FIG. 7B, the process 800 of FIG. 8, the process 900 of FIG. 9, the process 1000 of FIG. 10, and / or the process 1100 of FIG. 11. The system 200 may be implemented on one or more servers, a security entity such as a firewall, and / or an endpoint.

[0054] System 200 can be implemented by one or more devices such as a server. System 200 can be implemented at various locations on a network. In some embodiments, System 200 implements the malicious file detector 170 of System 100 in FIG. 1. As an example, System 200 is deployed as a service, such as a web service (e.g., System 200 determines whether a file is malicious and provides such a determination as a service). The service can be provided by one or more servers (e.g., System 200 or the malicious file detector monitors or receives files transmitted within or inside / outside the network via attached files to emails, instant messages, etc., determines whether the file is malicious, and sends / pushes a notification or update regarding the file, such as an indication of whether the file is malicious, and is deployed on a remote server). As another example, the malicious file detector is deployed on a firewall.

[0055] In the example shown, System 200 implements one or more modules related to predicting whether a file (e.g., a newly received file) is malicious, determining the likelihood that the file is malicious, and / or providing a notification or indication of whether the file is malicious. System 200 includes a communication interface 205, one or more processors 210, a storage device 215, and / or a memory 220. The one or more processors 210 include one or more of a communication module 225, a.NET header extraction module 230, an unmanaged function extraction module 235, a list generation module 240, a prediction module 245, and / or a notification module 250.

[0056] In some embodiments, system 200 includes a communication module 225. System 200 uses communication module 225 to communicate with various nodes or endpoints (e.g., client terminals, firewalls, DNS resolvers, data appliances, other security entities, etc.), or user systems such as an administrator system. For example, communication module 225 provides the information to be communicated to communication interface 205. As another example, communication interface 205 provides the information received by system 200 to communication module 225. Communication module 225 is configured to receive files to be analyzed, such as from a network endpoint or node such as a security entity (e.g., firewall). Communication module 225 is configured to query third-party services about information regarding the file (e.g., third-party scores or evaluations of the maliciousness of the file, community-based scores, evaluations, or reputations regarding the file, blacklists of the file, and / or services that disclose file information of a whitelist of the file, etc.). For example, system 200 uses communication module 225 to query third-party services. Communication module 225 is configured to receive one or more settings or configurations from an administrator. Examples of one or more settings or configurations include the configuration of the process for determining whether a file is malicious, the format in which information of a.NET file is compiled / arranged to determine the unmanaged Imphash of the.NET file, the hash function used in relation to the determination of the unmanaged Imphash of the file, information regarding a whitelist of domains (e.g., domains that are not considered suspicious and traffic or connections to which are permitted), information regarding a blacklist of domains (e.g., domains that are considered suspicious and traffic or connections to which are restricted).

[0057] In some embodiments, system 200 includes a.NET header extraction module 230. System 200 determines whether to extract information related to the header of a file (e.g., from the header), and uses.NET header extraction module 230 in connection with extracting information from the file (e.g., for analysis of whether the file is malicious). In some embodiments, the.NET header extraction module 230 receives the file to be analyzed. Such as a file included as an attachment to an email, instant message, or otherwise communicated via a network or from within / out of a network. In response to determining that the file is a.NET file, the.NET header extraction module 230 determines to perform extraction of information regarding the header of the file. As one example, the.NET header extraction module 230 determines that the file is a.NET file based on receiving an indication that the file corresponds to a.NET file. As another example, the.NET header extraction module 230 determines that the file is a.NET file based at least in part on determining that the file includes a.NET header. As another example, the.NET header extraction module 230 determines that the file is a.NET file based at least in part on determining that a directory entry within the PE header has a non-zero value (e.g., system 200 examines the binary structure of the file and determines whether the values within the optional header values include non-zero values / locations, and if so, determines that the file is a.NET file). As another example, the.NET header extraction module 230 determines that the file is a.NET file based at least in part on a check to determine whether the file imports the "_CorExeMain" or "_CorDllMain" functions.

[0058] In some embodiments, the.NET header extraction module 230 obtains.NET headers and / or information from the.NET header of a.NET file. In response to determining that the file is a.NET file, the.NET header extraction module 230 obtains information about the header of the file (e.g., information from the header). In some embodiments, the.NET header extraction module 230 determines the.NET header and obtains the import functions imported (or referenced) by the.NET header. For example, the.NET header extraction module 230 obtains the names of imported API functions based at least in part on the.NET header of the.NET file. The.NET header extraction module 230 obtains one or more data streams included in the.NET header and / or one or more tables included in (or referenced by) the.NET header. For example, the.NET header extraction module 230 obtains the #Strings stream included in the.NET header. As another example, the.NET header extraction module 230 obtains an ImplMap from the.NET file (e.g., from the.NET header). The ImplMap may contain various information about the imported unmanaged functions of the.NET file. In some embodiments, the.NET header extraction module 230 determines a set of import functions (e.g., the names of imported API functions) imported into the.NET file.

[0059] In accordance with various embodiments, in response to receiving a file to be analyzed to determine whether the file is malicious, system 200 places the file within a sandbox in which the file is to be analyzed. In some embodiments, a.NET header extraction module 230 extracts information related to the header of the file (e.g., information from the header). As one example, the.NET header extraction module 230 extracts header information from the.NET header and / or PE header for a.NET file within the sandbox. For example, system 200 invokes a sandbox for the analysis of a particular file. As another example, system 200 uses a common sandbox for the analysis of various files.

[0060] In some embodiments, system 200 includes an unmanaged function extraction module 235. System 200 uses the unmanaged function extraction module 235 to determine (e.g., obtain) a set of unmanaged functions included in or referenced by a.NET file, such as a list of unmanaged functions imported via the.NET header of the.NET file. For example, the unmanaged function extraction module 235 determines a set of unmanaged import functions from a set of import functions (e.g., imported API function names) imported into the.NET file. In accordance with various embodiments, the unmanaged function extraction module 235 uses information included in (or referenced by) the.NET header to determine unmanaged code or unmanaged functions (e.g., unmanaged functions and corresponding libraries). In some embodiments, the unmanaged function extraction module 235 determines a set of used unmanaged win32 API functions imported into the.NET file. The unmanaged function extraction module 235 provides a set of unmanaged import functions (or names of unmanaged functions) imported into the.NET file, and / or corresponding libraries, to a list generation module 240 and / or a prediction module 245.

[0061] In some embodiments, system 200 includes a list generation module 240. The system 200 uses the list generation module 240 to generate a list of unmanaged functions and / or corresponding libraries. In some embodiments, the system 200 uses the list generation module 240 to format a set of unmanaged import functions (or names of unmanaged functions) imported into a.NET file and / or corresponding libraries. The list generation module 240 formats the list and / or unmanaged functions (e.g., unmanaged function names and / or corresponding libraries) according to a default format. Examples of the default format include (i) a lowercase alphanumeric string, (ii) removal of the file extension, (iii) removal of the library extension (e.g., removal of.dll from the corresponding library name), (iv) addition of the unmanaged function name and corresponding library, and (v) use of a predefined separator between the unmanaged function name and the corresponding library. In some embodiments, the list generation module 240 appends the function name (e.g., unmanaged function name) to the corresponding library and separates the function name from the corresponding library by a dot or period (e.g., "."). In response to determining a string by adding the function name and library (using a default separator such as "."), the list generation module 240 adds such an entry to the list of unmanaged import functions and / or set of corresponding libraries. In response to determining that there are no additional unmanaged functions and / or corresponding libraries to add to the list according to various embodiments, the list generation module 240 provides the list to the prediction module 245.

[0062] In some embodiments, system 200 includes a prediction module 245. System 200 uses prediction module 245 to predict whether a file is malicious or to predict the likelihood that a file is malicious. According to various embodiments, prediction module 245 determines whether a file is malicious based at least in part on information included in (or referenced by) the.NET header of the file. For example, prediction module 245 determines the unmanaged Imphash corresponding to the file and determines whether the file is malicious based at least in part on the unmanaged Imphash.

[0063] According to various embodiments, prediction module 245 determines a hash with respect to a list of unmanaged import functions and / or a set of corresponding libraries. Various hash functions may be used in connection with determining the hash. Examples of hash functions include the SHA-256 hash function, the MD5 hash function, the SHA-1 hash function, etc. Various other hash functions may be implemented. Prediction module 245 uses a hash function to determine the unmanaged Imphash corresponding to the.NET file.

[0064] According to various embodiments, the prediction module 245 uses information obtained from the.NET header of the.NET file to determine whether the.NET file is malicious. In some embodiments, the prediction module 245 uses the unmanaged Imphash corresponding to the.NET file in connection with determining whether the.NET file is malicious. As an example, in response to determining the unmanaged Imphash corresponding to the.NET file, the prediction module 245 determines whether the unmanaged Imphash matches the unmanaged Imphash of a file considered to be malicious. As an example, in response to determining the unmanaged Imphash corresponding to the.NET file, the prediction module 245 determines whether the unmanaged Imphash matches the unmanaged Imphash of a file considered to be benign. In some embodiments, the prediction module 245 determines whether information about a particular file (e.g., the unmanaged Imphash corresponding to the.NET file being analyzed) is in a dataset of historical files and a historical dataset indicating whether a particular file is malicious (e.g., VirusTotal TMDetermine whether it is included in the historical information related to (such as third-party services). In response to determining that the information about a specific file is not included in or not available in the historical file and the dataset of historical information, the prediction module 245 may consider the file to be benign (for example, consider it not malicious). An example of historical information related to a historical file indicating whether a specific file is malicious corresponds to the VirusTotal (R) (VT) score. When the VT score for a specific file is greater than 0, that specific file is considered malicious by the third-party service. In some embodiments, the historical information related to a historical file indicating whether a specific file is malicious corresponds to a social score such as a community-based score or rating (for example, a reputation score) indicating that the file is malicious or likely to be malicious. The historical information (for example, from third-party services, community-based scores, etc.) indicates whether other vendors or cybersecurity organizations consider a specific file to be malicious.

[0065] System 200 can determine (e.g., calculate) a hash or signature corresponding to a file (e.g., unmanaged Imphash), and perform a lookup against historical information (e.g., whitelists, blacklists, etc.). In some implementations, prediction module 245 corresponds to, or is similar to, prediction engine 176. System 200 (e.g., prediction module 245) can query a third party (e.g., a third-party service) for historical information about a file (or a set of files or hashes / signatures for files previously determined to be malicious or benign) via communication interface 205. System 200 (e.g., prediction module 245) can query the third party at a predefined interval (e.g., a customer-specified interval, etc.). As one example, prediction module 245 can query the third party daily (or daily during the business week) about registration information for newly analyzed files.

[0066] In some embodiments, system 200 includes a notification module 250. System 200 uses the notification module 250 to provide an indication as to whether a file is malicious. For example, the notification module 250 obtains an indication (or likelihood that a file is malicious) as to whether a file is malicious from the prediction module 245, and provides the indication as to whether a file is malicious to one or more security entities and / or one or more endpoints. As another example, the notification module 250 provides updates to a white list of files and / or a black list of files to one or more security entities (e.g., a firewall), nodes, or endpoints (e.g., a client terminal). According to various embodiments, the notification module 250 obtains a hash, signature, or other unique identifier associated with the file (e.g., an unmanaged Imphash corresponding to the file), and provides an indication as to whether the file is malicious in relation to the hash, signature, or other unique identifier associated with the file.

[0067] According to various embodiments, the hash of a file corresponds to a hash using a predefined hash function (e.g., an unmanaged Imphash using an MD5 hash function, an MD5 hash of the file, etc.). A security entity or endpoint may calculate the hash of a received file (e.g., an attached file, etc.). The security entity or endpoint may determine whether the calculated hash corresponding to the file is included in a set such as a white list of benign files and / or a black list of malicious files. If a signature for malware (e.g., the hash of a received file) is included in a set of signatures of malicious files (e.g., a black list of malicious files), the security entity or endpoint may prevent the transmission of malware to an endpoint (e.g., a client device) and / or, in response, prevent the opening or execution of the malware.

[0068] In accordance with various embodiments, storage 215 includes one or more of file system data 260, hash data 262, and / or cache data 264. Storage 215 includes shared storage (e.g., a network storage system), and / or database data, and / or user activity data.

[0069] In some embodiments, file system data 260 comprises a database of one or more data sets (e.g., one or more data sets regarding files and / or file attributes, mapping of indicators of maliciousness to files or hashes, unmanaged Imphashes, signatures or other unique identifiers of files, mapping of indicators of benign files to files or hashes, signatures or other unique identifiers of files, etc.). File system data 260 includes data such as historical information regarding files (e.g., maliciousness of files), a white list of files considered to be safe (e.g., not suspicious), a black list of files considered to be suspicious or malicious (e.g., files for which the likelihood of being considered malicious exceeds a predefined / preset likelihood threshold), information associated with suspicious or malicious files, etc.

[0070] Hash data 262 includes data regarding one or more files, such as hash values for one or more files. In some embodiments, hash data 262 includes unmanaged Imphashes for files, such as files that are analyzed by system 200 to determine whether a file is malicious, or historical data sets that have previously been evaluated as malicious, such as by a third party. Hash data 262 includes a mapping of hash values (unmanaged Imphashes) to indications of maliciousness (e.g., an indication that the corresponding item is malicious or benign, etc.). In some embodiments, hash data 262 includes a relationship and association between a file or information related to the file (e.g., attributes such as scripts, bytes, structure, etc.) and an indication or likelihood that the file is malicious or benign. For example, hash data 262 includes a mapping of hash values (unmanaged Imphashes) to an indication of maliciousness (e.g., an indication as to whether the corresponding item is malicious or benign, etc.).

[0071] Cache data 264 includes information regarding the prediction of whether a file is malicious. As one example, predictive cache data 264 stores an indication of whether one or more files are malicious.

[0072] According to various embodiments, memory 220 includes executing application data 270. Executing application data 270 includes data that is obtained or used in connection with the execution of an application, such as an application that executes a hash function or an application that extracts information from a file. In an embodiment, the application includes one or more applications that perform one or more of receiving and / or executing a query or task, generating a report and / or configuration of information responsive to the executed query or task, and / or providing information responsive to the query or task to a user. Other applications include any other suitable applications (e.g., index maintenance applications, communication applications, machine learning model applications, applications for detecting suspicious traffic, document creation applications, report creation applications, user interface applications, data analysis applications, anomaly detection applications, user authentication applications, security policy management / update applications, etc.).

[0073] FIG. 3A is a diagram of an ImplMap table related to a.NET header for an exemplary.NET file. Table 300 shown in FIG. 3A provides an ImplMap table for a 32-bit DLL sample. In some embodiments, the ImplMap table for the sample is obtained by inputting the 32-bit DLL sample into a debugger and / or a.NET assembly editor. An example of such a debugger and / or.NET assembly editor is dnSpy. Various other debuggers and / or.NET assembly editors may be implemented.

[0074] In accordance with various embodiments, the system determines an index into the #Strings stream, at least in part, based on the ImplMap table. In some embodiments, the index into the #Strings stream corresponds to the ImportName table. The #Strings stream generally corresponds to an array of null-terminated strings where most of the strings of the.NET file reside. The system uses the ImportName and labeled columns to determine the index into the #Strings stream. The name of a function can be determined using the index value from the ImportName column. In Table 300, the name of the function corresponding to the value of ImportName is provided in the information column.

[0075] In response to determining the index value from ImportName and / or the function name of the import function, the system determines the library name for the corresponding library (e.g., DLL). In some embodiments, the system determines the library (e.g., library name), at least in part, based on the ModuleRef table of the.NET file.

[0076] FIG. 3B is a diagram of the ModuleRef table for an exemplary.NET file. Table 310 shown in FIG. 3B provides the ModuleRef table for a 32-bit DLL sample analyzed with respect to Table 300 of FIG. 3A. In some embodiments, the sample ModuleRef table is obtained by inputting the 32-bit DLL sample into a debugger and / or a.NET assembly editor.

[0077] According to various embodiments, the system uses the index value from the ImportName field (e.g., column) of the ImplMap table as an index to determine the corresponding library (e.g., library name). The system uses the index value as a lookup in the name column of Table 310. For example, the name column is an index for the #Strings stream. The Info column of the ModuleRef table contains an indication of the field. For example, the library name corresponding to the index value 0x18F3 in the Name column is kernel32.dll.

[0078] In connection with determining the unmanaged Imphash corresponding to a.NET file, the system determines a list of unmanaged functions. The list of unmanaged functions is generated using a default format or syntax. For example, the system obtains the library name corresponding to the import function and removes the extension. The system determines whether the library has a ".dll" extension, and if so, the system removes the extension. As another example, the system formats the function name and library name to convert any uppercase letters to lowercase. In some embodiments, the system constructs a string corresponding to a combination of the function name of an import function (e.g., an unmanaged import function) and the library name corresponding to the import function. For example, the system constructs the string by appending the library name to the function name, and a default separator (e.g., ".") is included between the library name and the function name. According to various embodiments, the system <libraryname> . <functionname>Determine a character string according to the following format. The system then adds the character string to a list of unmanaged import functions (for example, the list in which the unmanaged Imphash is determined).

[0079] Using the exemplary samples corresponding to Table 300 of FIG. 3A and Table 310 of FIG. 3B, the list of library-function name pairs is as follows: kernel32.createprocess, kernel32.getthreadcontext, kernel32.wow64getthreadcontext, kernel32.setthreadcontext, kernel32.wow64setthreadcontext, kernel32.readprocessmemory, kernel32.writeprocessmemory, ntdll.ntunmapviewofsection, kernel32.virtualallocex, kernel32.resumethread, kernel32.loadlibrary, kernel32.getprocaddress. The entries in the list can be separated by default separators such as commas, semicolons, colons, etc.

[0080] In response to determining a list of unmanaged import functions for a.NET file according to various embodiments, the system executes a hash of the list using a default hash function (for example, MD5, SHA1, SHA-256, etc.). As an example, for the list of unmanaged import functions using the samples analyzed in Table 300 of FIG. 3A and Table 310 of FIG. 3B, the unmanaged Imphash is (using SHA-256 as the hash function) 4b386faf53783c4fd17de6c043fd31374302a4455456b4af3d78abd74c865ff8.

[0081] Figure 3C is a diagram of the ImplMap table related to the.NET header for an exemplary.NET file. In table 320 shown in Figure 3C, the ImplMap table of a 32-bit EXE sample is provided. The 32-bit EXE was created using the Visual C++ compiler and C++ / CLI extensions. In some embodiments, the sample ImplMap table is obtained by inputting a 32-bit DLL sample into a debugger and / or a.NET assembly editor. One example of such a debugger and / or.NET assembly editor is dnSpy. Various other debuggers and / or.NET assembly editors may be implemented.

[0082] As shown in Figure 3C, the first three rows of table 320 have values contained in the Info column. Thus, for the first three rows, the process described in relation to table 300 of Figure 3A and table 310 of Figure 3B can be executed to determine the function name and the corresponding library name (e.g., to determine the library-function name pair). However, the remaining rows of the ImplMap table in Figure 3C have null values or zeros. Null values or zero entries may occur in the remaining rows because the corresponding functions are not used by the sample creator and are instead part of the C++ / CLI created runtime methods used internally. In some embodiments, the system determines the function name and library name based on the MethodDef table of the file. For example, the system uses the value contained in the MethodForward column of the ImplMap table (e.g., table 320) as an index into the MethodDef table. As one example, the MemberForward value is referred to as a coded index into the MethodDef table (e.g., the coded index is defined in the.NET standard ECMA-335).

[0083] According to various embodiments, the system obtains values from the remaining rows of the MethodForward column of the ImplMap table in FIG. 3C, decodes the values obtained from the MethodForward column, and uses the decoded values as indexes for the MethodDef table. As an example, the value from the fourth row of the MemberForward column is 0x99, which is decoded to correspond to 76. Thus, 76 is used as an index value to determine information from the MethodDef table.

[0084] FIG. 3D is a diagram of the MethodDef table for an exemplary.NET file. Table 330 shown in FIG. 3D provides a sample MethodDef table analyzed with respect to Table 320 of FIG. 3C. In some embodiments, the MethodDef table for the sample is obtained by inputting the sample into a debugger and / or a.NET assembly editor.

[0085] Using the index value obtained from the MethodForward column of ImplMap, the system determines an index for the #Strings stream. For example, the system uses the index value obtained from the MethodForward column of ImplMap as a lookup in the Name column in front of the MethodDef table. The index value of the Name column corresponding to the index 76 from the MethodForward column is 0x1831. As shown in Table 330, the Info column of the MethodDef table indicates that the function name is _amsg_exit.

[0086] FIG. 3E is a diagram of the ModuleRef table for an exemplary.NET file. Table 340 shown in FIG. 3E provides a sample ModuleRef table analyzed with respect to Table 320 of FIG. 3C. In some embodiments, the sample ModuleRef table is obtained by inputting the sample into a debugger and / or a.NET assembly editor.

[0087] In accordance with various embodiments, the system determines a library name in response to determining a function name (e.g., a function name obtained from the MethodDef table). For example, the system performs a lookup within the ImplMap table (e.g., table 320 of FIG. 3C) and determines that the value in the ImportScope column is 2 for the remaining rows. The system uses the index value 2 from the ImportScope column as the index into the ModuleRef table. In response to performing a lookup within the ModuleRef table (e.g., table 340) using the index value 2 (e.g., from the ImportScope column of the ImplMap table), the system also determines that the value in the Name column is also 0. Thus, the system determines that the ModuleRef does not provide an indication of the library name.

[0088] In response to determining that ModuleRef does not provide an indication of the library name corresponding to the function, the system determines to use the function name retrieved from the MethodDef table and to parse the import table within the PE header. The import table within the PE header contains all the statically used functions of the file having the library name (DLL) corresponding to those pairs in which there are statically used functions. For example, the system uses a PE header parser to parse the import table within the PE header. One example of a PE header parser is pefile, a PE header parser coded in Python (e.g., an open source project called pefile). In some embodiments, the system performs a one-to-one (1:1) comparison between the function name obtained from the MethodDef table and the functions within the import table of the PE header. As an example, a file written in C++ (such as a mixed assembly) can have so-called mangled names as the function names within the import table. The process of creating these names is known as name mangling. Such mangled function names are automatically created for all C++ functions by the C++ compiler, except when the function is defined as extern "C".

[0089] Figure 3F is a diagram of the Import table for an exemplary.NET file. In the example shown in Figure 3F, table 350 is the import table included in the PE header of the.NET file.

[0090] Figure 3G is a diagram of the MethodDef table for an exemplary.NET file. In the example shown, table 360 is the MethodDef table of the.NET file. In some embodiments, in response to the system determining that a function name obtained from the MethodDef table matches a function name in an import table included in the PE header, the system obtains the corresponding library name from the import table. In accordance with various embodiments, the system performs a lookup of matching function names between the MethodDef and the import table included in the PE header to obtain the corresponding library name.

[0091] In connection with determining the unmanaged Imphash corresponding to a.NET file, the system determines a list of unmanaged functions. The list of unmanaged functions is generated using a default format or syntax. For example, the system obtains the library name corresponding to an imported function and removes the extension. The system determines whether the library has a ".dll" extension, and if so, the system removes the extension. As another example, the system formats the function name and library name to convert any uppercase letters to lowercase. In some embodiments, the system constructs a string corresponding to a combination of the function name of an imported function (e.g., an unmanaged imported function) and the library name corresponding to the imported function. For example, the system constructs the string by appending the library name to the function name, and a default separator (e.g., ".") is included between the library name and the function name. In accordance with various embodiments, the system <libraryname> . <functionname>Determine a string according to the following format. The system then adds the string to a list of unmanaged import functions (for example, the list for which the unmanaged Imphash has been determined).

[0092] Using the exemplary samples corresponding to Table 300 of FIG. 3A and Table 310 of FIG. 3B, the list of library-function name pairs is as follows: ntdll.zwallocatevirtualmemory, ntdll.zwfreevirtualmemory, ntdll.ldrgetprocedureaddress, msvcr80.amsg_exit, kernel32.sleep,. <crtimplementationdetails>.throwmoduleloadexception,. <crtimplementationdetails>.throwmoduleloadexception,. <crtimplementationdetails>.dodlllanguagesupportvalidation,. <crtimplementationdetails>.thrownestedmoduleloadexception,. <crtimplementationdetails>.registermoduleuninitializer,. <crtimplementationdetails>.docallbackindefaultdomain, msvcr80._cexit, msvcr80._encode_pointer, msvcr80._decode_pointer, msvcr80._encoded_null, msvcr80._frameunwindfilter. Entries in the list may be separated by default separators such as commas, semicolons, colons, etc. Library function pairs <crtimplementationdetails>.throwmoduleloadexception is included twice because the function is a so-called overloaded function. Such functions have the same name but differ in their parameter types, number, or return value. Therefore, these are different functions.

[0093] In accordance with various embodiments, in response to determining a list of unmanaged import functions for a.NET file, the system executes a hash of the list using a default hash function (e.g., MD5, SHA1, SHA-256, etc.). As an example, for the list of unmanaged import functions using the samples analyzed in Table 300 of FIG. 3A and Table 310 of FIG. 3B, the unmanaged Imphash (using SHA-256 as the hash function) is cd2f27b642a85c3f0c10db4e887d504d3b4c0882f9264994367c0c9b4ea7a537.

[0094] In some embodiments, the system converts or formats the list according to a default format. For example, a list of library <-> function (<-> mappingflags) name pairs is converted to a comma-separated string. The system creates a comma-separated string having a single list item (library <-> function pair). In response to converting / creating a comma-separated string having a single list item, the system determines (e.g., calculates a hash) with respect to such a string.

[0095] FIG. 4 is a flowchart of a method for detecting malicious files according to various embodiments. In some embodiments, process 400 is at least partially implemented by system 100 of FIG. 1 and / or system 200 of FIG. 2. In some implementations, process 400 may be implemented by one or more servers related to providing services to a network (e.g., network endpoints such as security entities and / or client devices). In some implementations, process 400 may be implemented by a security entity (e.g., a firewall) related to implementing security policies regarding files communicated over a network or within and outside the network. In some implementations, process 400 may be implemented by a client device such as a laptop, smartphone, personal computer, etc., related to executing or opening a file such as an email attachment file.

[0096] At 410, a sample is received. In some embodiments, the system receives a sample (e.g., a.NET file) from a security entity (e.g., a firewall), an endpoint (e.g., a client device), etc. For example, in response to determining that a file is attached to a communication such as an email or an instant message, the security entity or the endpoint provides (e.g., transmits) the file to the system. The sample may be received in relation to a request to determine whether the file is malicious.

[0097] When process 400 is implemented by a security entity, the sample can be received, for example, by routing traffic to an applicable network endpoint (e.g., a firewall can obtain a sample from an email attachment of an email directed to a client device). When process 400 is implemented by a client device, the sample can be received by an application or layer that monitors incoming / outgoing information. For example, a process (e.g., an application, an operating system process, etc.) can monitor email attachment files, files exchanged via an instant messaging program, etc., and can be executed in the background to obtain them.

[0098] At 420, the imported API function name is obtained using the.NET header of the sample. In some embodiments, in response to receiving a request to evaluate whether the sample, and / or the sample (e.g., the.NET file) is malicious, the system analyzes the sample to obtain information related to (e.g., included in) the.NET header of the sample.

[0099] According to various embodiments, the system determines a.NET header and obtains the import functions imported (or referenced) by the.NET header. For example, the system obtains the imported API function names based at least in part on the.NET header of the.NET file. The system obtains one or more data streams included in the.NET header and / or one or more tables included in (or referenced by) the.NET header. For example, the system obtains the #Strings stream included in the.NET header. As another example, the system obtains an ImplMap from the.NET file (e.g., from the.NET header). The ImplMap can include various information regarding any imported unmanaged functions of the.NET file. In some embodiments, the system determines (e.g., obtains) a set of import functions (e.g., imported API function names) imported into the.NET file.

[0100] In some embodiments, the system determines (e.g., obtains) a set of unmanaged functions included in or referenced by the.NET file, such as a list of unmanaged functions imported through the.NET header of the.NET file. For example, the system determines a set of unmanaged import functions from a set of import functions (e.g., imported API function names) imported into the.NET file. The system uses the information included in (or referenced by) the.NET header to determine unmanaged code or unmanaged functions (e.g., unmanaged functions and corresponding libraries). In some embodiments, the system determines a set of used managed win32 API functions imported into the.NET file.

[0101] In 430, a hash of the list of unmanaged imported API function names is determined. In some embodiments, in response to obtaining the imported API function names, the system determines (e.g., calculates) a hash of the list of unmanaged imported API function names. For example, the system calculates an unmanaged Imphash corresponding to a.NET file.

[0102] In connection with determining the hash of a list of unmanaged imported API function names, the system determines a list of unmanaged import functions (or names of unmanaged functions) imported into a.NET file, and / or a corresponding set of libraries, and then determines the hash of such a list. The list of unmanaged import functions and / or corresponding set of libraries is determined according to a default order. For example, the ordering of the imported unmanaged functions and / or corresponding libraries corresponds to the order in which the unmanaged functions are included in the elements of the.NET header (e.g., the order in which the unmanaged functions are included in the.NET table or in the #Strings stream and / or table referenced thereby). Various other orders in which unmanaged functions are added to the list (or the list is arranged) may be implemented. The system formats the list and / or unmanaged functions (e.g., unmanaged function names and / or corresponding libraries) according to a default format or syntax. Examples of the default format include (i) lower-case alphanumeric strings, (ii) removal of file extensions, (iii) removal of library extensions (e.g., removal of.dll from the corresponding library name), (iv) addition of the unmanaged function name and corresponding library, and (v) use of a predefined separator between the unmanaged function name and the corresponding library. In some embodiments, the system adds the function name (e.g., unmanaged function name) to the corresponding library and separates the function name from the corresponding library by a dot or period (e.g., "."). In response to determining the string by adding the function name and library (using a default separator such as "."), the system adds such an entry to the list of unmanaged import functions and / or corresponding set of libraries.

[0103] According to various embodiments, the system determines a hash for an unmanaged import function and / or a list of a corresponding set of libraries. Various hash functions can be used in connection with determining the hash. Examples of hash functions include the SHA-256 hash function, the MD5 hash function, the SHA-1 hash function, etc. Various other hash functions can be implemented. The system uses a hash function to determine the unmanaged Imphash corresponding to a.NET file.

[0104] At 440, a determination is made as to whether the sample is malicious. In some embodiments, in response to determining the names of unmanaged imported API functions, the system uses a hash (e.g., unmanaged Imphash) in connection with determining whether the sample is malicious.

[0105] In some embodiments, the system uses the unmanaged Imphash corresponding to the.NET file in connection with determining whether the.NET file is malicious. As one example, in response to determining the unmanaged Imphash corresponding to the.NET file, the system determines whether the unmanaged Imphash matches the unmanaged Imphash of a file that is considered malicious. When the unmanaged Imphash for a sample (e.g., the file being analyzed) matches the unmanaged blacklist of malicious files (e.g., records included in the Imphash list of malicious files) within the historical dataset, the system considers the sample to be malicious. As one example, in response to determining the unmanaged Imphash corresponding to the.NET file, the system determines whether the unmanaged Imphash matches the unmanaged Imphash of a file that is considered benign. When the unmanaged Imphash of the sample (e.g., the file being analyzed) matches the unmanaged Imphash of the benign files (e.g., records included in the whitelist of benign files) within the historical dataset, the system considers the sample to be benign. In some embodiments, the system determines whether information about a particular file (e.g., the unmanaged Imphash corresponding to the.NET file being analyzed) is included in the dataset of historical files, and historical information related to the historical dataset indicating whether a particular third-party service (e.g., VirusTotal TM is malicious. As one example, in response to determining that information about a particular file is not included in or not available in the historical file and the dataset of historical information, the system considers that file to be benign (e.g., considers that file not to be malicious). Examples of historical information related to the historical file indicating whether a particular file is malicious include VirusTotal TM (VT) score. If the VT score for a particular file is greater than 0, that particular file is considered malicious by third - party services. In some embodiments, the historical information related to the historical file indicating whether a particular file is malicious corresponds to a social score such as a community - based score or rating (e.g., reputation score) indicating that the file is malicious or likely to be malicious. The historical information (e.g., from third - party services, community - based scores, etc.) indicates whether other vendors or cyber - security organizations consider a particular file to be malicious.

[0106] In response to the determination at 440 that the sample is malicious, process 400 proceeds to 450, where an indication that the sample is malicious is provided.

[0107] In response to determining at 440 that the sample is malicious, process 400 proceeds to 450, where an indication that the sample is malicious is provided. For example, the indication that the sample is malicious can be provided to the component from which the sample is received. As one example, the system provides the indication that the sample is malicious to a security entity. As another example, the system provides the indication that the sample is malicious to a client device. As one example, security provides the indication that the sample is malicious to a client device. In some embodiments, the indication that the sample is malicious is provided to a user such as the user of the client device and / or a network administrator.

[0108] In accordance with various embodiments, in response to receiving an indication that a sample is malicious, active measures may be executed. The active measures may be executed in accordance with (e.g., at least in part based on) one or more security policies. As one example, one or more security policies may be pre-set by a network administrator, a customer (e.g., an organization / company) of a service that provides detection of malicious files, etc. Examples of active measures that may be executed include the following. Isolating the file (e.g., quarantining the file), deleting the file, prompting the user to warn the user that a malicious file has been detected, providing a prompt to the user when the device attempts to open or execute the file, blocking the transmission of the file, updating a blacklist of malicious files (e.g., mapping the hash of the file to an indication that the file is malicious, etc.).

[0109] In response to determining at 440 that the sample is not malicious, process 400 proceeds to 460. In some embodiments, in response to determining that the sample is not malicious, a mapping of the file (or the hash / signature of the file) to an indication that the file is not malicious is updated. For example, a whitelist of benign files is updated to include the sample, or the hash, signature, or other unique identifier associated with the sample.

[0110] At 460, a determination is made as to whether process 400 has been completed. In some embodiments, in response to determining that no further samples should be analyzed (e.g., no further prediction about the file is needed), process 400 is determined to be complete, and the administrator may indicate that process 400 should be paused or stopped, etc. In response to determining that process 400 has been completed, process 400 ends. In response to determining that process 400 has not been completed, process 400 returns to 410.

[0111] FIG. 5 is a flowchart of a method for determining whether a file is malicious according to various embodiments. In some embodiments, process 500 is at least partially implemented by system 100 of FIG. 1 and / or system 200 of FIG. 2. In some implementations, process 500 can be implemented by one or more servers, such as in connection with providing services to a network (e.g., network endpoints such as security entities and / or client devices). In some implementations, process 500 can be implemented by a security entity (e.g., a firewall), such as in connection with enforcing security policies regarding files communicated over a network or within and outside of a network. In some implementations, process 500 can be implemented by a client device, such as a laptop, smartphone, personal computer, etc., such as in connection with executing or opening a file such as an email attachment file.

[0112] According to various embodiments, process 500 is called in relation to 440 of process 400 of FIG. 4.

[0113] At 510, a hash of the list is obtained. In some embodiments, the system obtains (e.g., receives, determines, etc.) an unmanaged Imphash. For example, the system receives a hash of a list of unmanaged imported API function names.

[0114] At 520, the hash is used in connection with a query of the mapping. In response to receiving the hash of the list (e.g., unmanaged Imphash), the system performs a lookup against a historical dataset of malicious files and / or benign files. For example, the historical dataset includes an association between unmanaged Imphashes and an indication of whether the corresponding file is malicious or benign.

[0115] In some embodiments, the system uses the unmanaged Imphash corresponding to a.NET file in connection with determining whether the.NET file is malicious. As one example, in response to determining the unmanaged Imphash corresponding to a.NET file, the system determines whether the unmanaged Imphash matches the unmanaged Imphash of a file considered to be malicious. As one example, in response to determining the unmanaged Imphash corresponding to a.NET file, the system determines whether the unmanaged Imphash matches the unmanaged Imphash of a file considered to be benign. In some embodiments, the system determines whether information about a particular file (e.g., the unmanaged Imphash corresponding to the.NET file being analyzed) is present in a historical file dataset and a historical dataset indicating whether the particular file is malicious (e.g., VirusTotal TM determine whether it is included in historical information related to such third - party services). As an example, in response to determining that information about a particular file is not included in, or not available in, the historical file and the dataset of historical information, the system considers the file to be benign (e.g., considers the file not to be malicious). An example of historical information related to a historical file indicating whether a particular file is malicious corresponds to a VirusTotal(R) (VT) score. When the VT score for a particular file is greater than 0, that particular file is considered malicious by a third - party service. In some embodiments, the historical information related to a historical file indicating whether a particular file is malicious corresponds to a social score, such as a community - based score or rating (e.g., a reputation score) indicating that the file is malicious or likely to be malicious. The historical information (e.g., from third - party services, community - based scores, etc.) indicates whether other vendors or cyber - security organizations consider a particular file to be malicious.

[0116] At 530, a determination is made as to whether the mapping indicates that the hash corresponds to a malicious file.

[0117] In response to determining at 530 that the mapping indicates that the hash corresponds to a malicious file, process 500 proceeds to 540, where the sample is determined to be malicious.

[0118] In response to determining that the mapping at 530 indicates that the hash does not correspond to a malicious file, process 500 proceeds to 550, where the sample is determined not to be malicious. In some embodiments, the system determines that the sample is benign in response to determining that the mapping of the hash to the file does not contain an indication that the hash maps to a malicious file. As one example, the system determines that the hash is not included in the mapping of the hash to a malicious file. As another example, the system determines that the mapping does not contain a record mapped to a malicious file (or an indication of a malicious file).

[0119] If the unmanaged Imphash of a sample (e.g., the file being analyzed) matches the unmanaged blacklist of malicious files (e.g., records included in the Imphash list of malicious files) within the historical dataset, the system considers the sample to be malicious.

[0120] If the unmanaged Imphash of a sample (e.g., the file being analyzed) matches the unmanaged Imphash of benign files (e.g., records included in the whitelist of benign files) within the historical dataset, the system considers the sample to be benign. In some embodiments, in response to determining that information about a particular file is not included in or not available in the historical file and the dataset of historical information, the system considers the file to be benign (e.g., considers the file not to be malicious).

[0121] At 560, the result of maliciousness is provided. In some embodiments, the system provides an indication that the hash corresponds to a malicious file. For example, the system provides an indication that the file corresponding to the hash is malicious.

[0122] At 570, a determination is made as to whether process 500 has been completed. In some embodiments, process 500 is determined to be complete in response to determining that no further hashes are to be analyzed (e.g., no further predictions about the file are needed), and the administrator indicates that process 500 should be paused or stopped, etc. In response to determining that process 500 has been completed, process 500 ends. In response to determining that process 500 has not been completed, process 500 returns to 510.

[0123] FIG. 6 is a flowchart of a method for detecting malicious files according to various embodiments. In some embodiments, process 600 is at least partially implemented by system 100 of FIG. 1 and / or system 200 of FIG. 2. In some implementations, process 600 can be implemented by one or more servers, such as in connection with providing services to a network (e.g., network endpoints such as security entities and / or client devices). In some implementations, process 600 can be implemented by a security entity (e.g., a firewall), such as in connection with enforcing security policies regarding files communicated over a network or within and outside of a network. In some implementations, process 600 can be implemented by client devices such as laptops, smartphones, personal computers, etc., such as in connection with executing or opening files such as email attachment files.

[0124] At 602, a sample is received.

[0125] At 604, a.NET assembly corresponding to the sample is obtained. In some embodiments, the sample is compressed (e.g., in ZIP format, etc.), and the.NET assembly is extracted from the compressed file. In some embodiments, obtaining the.NET assembly includes determining that the sample is a.NET file.

[0126] At 606, a determination is made as to whether the.NET assembly includes a ModuleRef table. In response to determining that the.NET assembly does not include a ModuleRef table, process 600 ends. Conversely, in response to determining that the.NET assembly includes a ModuleRef table, process 600 proceeds to 608.

[0127] At 608, a determination is made as to whether the.NET assembly includes an ImplMap table. In response to determining that the.NET assembly does not include an ImplMap table, process 600 ends. Conversely, in response to determining that the.NET assembly includes an ImplMap table, process 600 proceeds to 610.

[0128] At 610, the ImportName value is obtained. In some embodiments, the system obtains the ImplMap value from the ImportName column of the ImportName table. For example, the system obtains the ImportName value from the corresponding row (e.g., the selected row) of the ImportName column. In some embodiments, the system iterates over the rows of ImportName to obtain the value of each corresponding row of the ImportName column.

[0129] At 612, a determination is made as to whether the ImportName value obtained at 610 is equal to 0. In response to determining that the ImportName value obtained at 610 is equal to 0, process 600 proceeds to 626. In response to determining that the ImportName value obtained at 610 is not equal to 0, process 600 proceeds to 614.

[0130] At 614, the function name is obtained from the #Strings stream. In some embodiments, the system uses the selected row to obtain the function name from the #Strings stream of the.NET assembly.

[0131] At 616, the value from the ImportScope column is obtained. In some embodiments, the system obtains the value from the ImportScope column of the ImplMap table. The system obtains the ImplMap table using the.NET header (e.g., the.NET header includes the ImplMap table). The value of the ImportScope column is obtained from the selected row (e.g., the row of the ImplMap table from which the ImportName value is obtained, and / or the row of the Info column from which the #Strings stream information is obtained). In some embodiments, the value from the ImportScope column is used as an index in connection with performing a lookup into the ModuleRef table (e.g., regarding the name of the corresponding library).

[0132] At 618, the ImportScope value is used as a row index in the ModuleRef table to obtain a value. In some embodiments, the system obtains the ModuleRef table using the.NET header (e.g., the.NET header includes the ModuleRef table). The system performs a lookup within the ModuleRef table using the value obtained from the ImportScope column as an index. For example, the system uses the value obtained from the ImportScope column to determine the row of the ModuleRef table from which the system should obtain a value from the name column. In some embodiments, the value obtained from the name column of the ModuleRef table is used as an index into the #Strings stream.

[0133] At 620, the library name is obtained from the #Strings stream. In some embodiments, the system uses the.NET header to obtain the #Strings stream (e.g., the.NET header includes the #Strings stream). The system uses the value obtained from the name list as an index to perform a lookup within the #Strings stream.

[0134] At 624, a string corresponding to the library name-function name pair is added to the list. For example, the library name-function name pair is added to the list of unmanaged import functions. In some embodiments, the system generates the string according to a default format or syntax.

[0135] At 626, a determination is made as to whether the.NET assembly contains a MethodDef table. In response to determining that the.NET assembly does not contain a MethodDef table, process 600 proceeds to 6, where the system considers the function name for which the.NET header is empty. Conversely, in response to determining that the.NET assembly contains a MethodDef table, process 600 proceeds to 628.

[0136] At 628, a value is obtained from the MemberForwarded column of the ImplMap table. The system uses the.NET header to obtain the ImplMap table (e.g., the.NET header includes the ImplMap table). The value of the MemberForwarded column is obtained from the selected row (e.g., the row of the ImplMap table from which the MemberForwarded value is obtained). In some embodiments, the value from the MemberForwarded column is used as an index in connection with performing a lookup into the ModuleRef table (e.g., for the name of the corresponding library).

[0137] In 630, to obtain a value, the MemberForwarded value is used as a row index in the ModuleRef table. In some embodiments, the system uses the.NET header to obtain the ModuleRef table (e.g., the.NET header includes the ModuleRef table). The system performs a lookup within the ModuleRef table using the value obtained from the MemberForwarded column as an index. For example, the system uses the value obtained from the name column to determine, from the #Strings stream, the row of the ModuleRef table from which the system should obtain the value.

[0138] In 632, the function name is obtained from the #Strings stream. In some embodiments, the system uses the.NET header to obtain the #Strings stream (e.g., the.NET header includes the #Strings stream). The system performs a lookup of the #Strings stream using the value obtained from the name column of the ModuleRef table as an index.

[0139] In 636, a determination is made as to whether the PE header has an import table. In some embodiments, the system determines whether the PE header includes an import table in connection with determining the library name corresponding to the function name (e.g., the function name obtained from the #Strings stream).

[0140] In 636, in response to determining that the PE header does not have an import table, process 600 proceeds to 638, where the system considers the library name to be an empty library name. Conversely, in 636, in response to determining that the PE header has an import table, process 600 proceeds to 640, where the library name corresponding to the function name is obtained. In some embodiments, the system parses the import table to obtain the library name corresponding to the function name. In response to obtaining the library name, process 600 proceeds to 624.

[0141] In 624, after the function name and the corresponding library name are added to the list, process 600 proceeds to 642.

[0142] In 642, a determination is made as to whether the population (e.g., generation) of the list is complete. In some embodiments, the population / generation of the list is determined to be complete in response to determining that no further import functions and / or corresponding libraries are added to the list (e.g., no further import functions are included in the.NET header or are not referenced), and the administrator indicates that process 600 should be suspended or stopped, etc. In response to determining in 642 that no further import functions and / or corresponding libraries are added to the list, process 600 proceeds to 644, where the system determines (e.g., calculates) the hash of the list (e.g., the list of unmanaged import functions). In response to determining that process 600 is not complete, process 600 returns to 602. In some embodiments, in response to calculating the hash (e.g., the unmanaged Imphash), the system determines whether the sample is malicious based at least in part on the unmanaged Imphash. For example, in response to the calculation of the hash, process 500 of FIG. 5 is called.

[0143] In some embodiments, the system determines the hash of the list by invoking process 700 of FIG. 7A or process 750 of FIG. 7B.

[0144] FIG. 7A is a flowchart of a method for detecting malicious files according to various embodiments. In some embodiments, process 700 is at least partially implemented by system 100 of FIG. 1 and / or system 200 of FIG. 2. In some implementations, process 600 can be implemented by one or more servers, such as in connection with providing services to a network (e.g., network endpoints such as security entities and / or client devices). In some implementations, process 700 can be implemented by a security entity (e.g., a firewall), such as in connection with enforcing security policies regarding files communicated over a network or within and outside of a network. In some implementations, process 700 can be implemented by a client device such as a laptop, smartphone, personal computer, etc., such as in connection with executing or opening a file such as an email attachment file.

[0145] At 702, a library name and / or a function name is obtained. In some embodiments, the function names and corresponding library names obtained using.NET headers are obtained, such as in connection with generating a list of unmanaged import functions.

[0146] At 704, a determination is made as to whether the library name has a file extension. In response to determining at 704 that the library name has a file extension, process 700 proceeds to 706, where the file extension is removed. In response to determining at 704 that the library name does not have a file extension, process 700 proceeds to 708.

[0147] At 708, function names and library names are formatted. In some embodiments, the system formats library name - function name pairs according to a default format or syntax. As another example, the system formats function names and library names to convert any uppercase letters to lowercase letters.

[0148] At 710, function names and library names are combined. In some embodiments, the system constructs a string corresponding to a combination of a function name of an imported function (e.g., an unmanaged imported function) and a library name corresponding to the imported function. For example, the system constructs the string to append the library name to the function name, and includes a default separator (e.g., ".") between the library name and the function name. According to various embodiments, the system formats, i.e., <libraryname> . <functionname>, determine the string accordingly. Next, the system adds the string to a list of unmanaged import functions (e.g., the list where the unmanaged Imphash is determined).

[0149] According to various embodiments, various formats or syntaxes can be implemented in connection with a system that combines function names and library names. Examples of a default format are: (i) a lowercase alphanumeric string, (ii) removal of the file extension, (iii) removal of the library extension (e.g., removal of.dll from the corresponding library name), (iv) addition of the unmanaged function name and the corresponding library, and (v) use of a predefined separator between the unmanaged function name and the corresponding library, (vi) the order of concatenation of the library name and the function name to determine the string (e.g., <libraryname> . <functionname>or <functionname> . <libraryname>It includes ). In some embodiments, the system adds the function name (e.g., the unmanaged function name) to the corresponding library, and separates the function name from the corresponding library by a dot or period (e.g., "."). In response to determining the string by adding the function name and the library (using a default separator such as "."), the system adds such an entry to the list of unmanaged import functions and / or the set of corresponding libraries. Various other default formats / syntaxes may be implemented.

[0150] In 712, the combination of the function name and the library name is added to the list of unmanaged import functions.

[0151] In 714, a determination is made as to whether there are more functions to be added to the list. For example, the system determines whether more unmanaged import functions should be added to the list of unmanaged import functions for that file. In response to determining in 714 that there are no additional functions to be added to the list, process 700 proceeds to 716. In response to determining that additional functions should be added to the list, process 700 returns to 702.

[0152] In 716, a hash is calculated for the list. According to various embodiments, the system determines a hash for the list of unmanaged import functions and / or the set of corresponding libraries. Various hash functions can be used in connection with determining the hash. Examples of hash functions include the SHA-256 hash function, the MD5 hash function, the SHA-1 hash function, etc. Various other hash functions may be implemented. The system uses a hash function to determine the unmanaged Imphash corresponding to the.NET file.

[0153] In some embodiments, the system converts or formats a list according to a predefined format. For example, a list of library <-> function (<-> mappingflags) pairs is converted into a comma-separated string. Following such an example, instead of calculating a hash for the following list, kernel32.createprocess kernel32.getthreadcontext kernel32.wow64getthreadcontext kernel32.setthreadcontext kernel32.wow64setthreadcontext kernel32.readprocessmemory The system creates a comma-separated string with a single list item (library <-> function pair). That is, "kernel32.createprocess, kernel32.getthreadcontext, kernel32.wow64getthreadcontext, kernel32.setthreadcontext, kernel32.wow64setthreadcontext, kernel32.readprocessmemory, …". In response to converting / creating a comma-separated string with a single list item, the system determines (e.g., calculates a hash) with respect to such a string.

[0154] In 718, a hash is provided. In some embodiments, the system provides the hash to another system or module, such as a system or module related to determining whether a file is malicious. For example, the hash is provided in response to an invocation of process 700.

[0155] At 720, a determination is made as to whether process 700 has been completed. In some embodiments, process 700 is determined to be complete in response to determining that no further hashes should be calculated (e.g., no further predictions about the file are needed), and the administrator instructs that process 700 should be paused or stopped, etc. In response to determining that process 700 has been completed, process 700 ends. In response to determining that process 700 has not been completed, process 700 returns to 702.

[0156] FIG. 7B is a flowchart of a method for detecting malicious files according to various embodiments.

[0157] In some embodiments, process 700 is at least partially implemented by system 100 of FIG. 1 and / or system 200 of FIG. 2. In some implementations, process 600 can be implemented by one or more servers, such as in connection with providing services to a network (e.g., network endpoints such as security entities and / or client devices). In some implementations, process 700 can be implemented by a security entity (e.g., a firewall), such as in connection with implementing security policies regarding files communicated over a network or inside and outside the network. In some implementations, process 700 can be implemented by a client device, such as a laptop, smartphone, personal computer, etc., such as in connection with executing or opening a file such as an email attachment file.

[0158] According to various embodiments, 750 includes additional information of an unmanaged function (e.g., information regarding the way the unmanaged function is defined in the code by the creator, etc.), and 750 is more precise compared to 700 of FIG. 7A. As one example, the final hash of 750 is also more precise, and potentially has fewer false positive tendencies. However, the drawback is that the final hash of 750 may potentially hit in malware samples with fewer final hashes than that of 700.

[0159] In 752, the library name and / or the function name is obtained. In some embodiments, the function name and the corresponding library name obtained using the.NET header are obtained, such as in relation to generating a list of unmanaged import functions.

[0160] In 754, a determination is made as to whether the library name has a file extension. In response to determining in 754 that the library name has a file extension, process 750 proceeds to 756, where the file extension is removed. In response to determining in 754 that the library name does not have a file extension, process 700 proceeds to 760.

[0161] In 758, the function name and the library name are formatted. In some embodiments, the system formats the library name - function name pair according to a default format or syntax. As another example, the system formats the function name and the library name to convert any uppercase letters to lowercase.

[0162] At 760, the MappingFlags value is obtained. In some embodiments, the system obtains the MappingFlags value from the ImplMap table. For example, the system obtains the MappingFlags value from the row in the ImplMap table corresponding to the function. The MappingFlags value includes P / Invoke attributes.

[0163] At 762, the MappingFlags value, the function name, and the library name are combined. In some embodiments, the system constructs a string corresponding to the combination of the MappingFlags value, the function name of an import function (e.g., an unmanaged import function), and the library name corresponding to the import function. For example, the system constructs a string by appending the MappingFlags value and the library name to the function name, and includes a default separator (e.g., ".") between the library name and the function name. According to various embodiments, the system formats, i.e., <libraryname> . <functionname> . <mappingflags>Determine a string accordingly. Next, the system adds the string to a list of unmanaged import functions (e.g., the list for which the unmanaged Imphash was determined).

[0164] According to various embodiments, various formats or syntaxes may be implemented in connection with a system that combines function names and library names. Examples of a default format are (i) an alphanumeric string in lowercase, (ii) removal of the file extension, (iii) removal of the library extension (e.g., removal of.dll from the corresponding library name), (iv) addition of the unmanaged function name and the corresponding library, and (v) use of a default separator between the unmanaged function name and the corresponding library, (vi) the order of combination of the library name and the function name to determine the string (e.g., <libraryname> . <functionname> . <mappingflags> 、 <functionname> . <libraryname> . <mappingflags>It includes (such as,...). In some embodiments, the system adds the function name (for example, the unmanaged function name) to the corresponding library, and separates the function name from the corresponding library by a dot or period (for example, "."). In response to determining the character string by adding the function name and the library (using a default separator such as "."), the system adds such an entry to the list of unmanaged import functions and / or the set of corresponding libraries. Various other default formats / syntaxes may be implemented.

[0165] In 764, the combination of the function name and the library name is added to the list of unmanaged import functions.

[0166] In 766, a determination is made as to whether there are more functions to be added to the list. For example, the system determines whether more unmanaged import functions should be added to the list of unmanaged import functions of the file. In 766, in response to determining that there are no additional functions to be added to the list, process 750 proceeds to 752. In response to determining that additional functions should be added to the list, process 750 returns to 752.

[0167] In 768, a hash is calculated for the list. According to various embodiments, the system determines a hash for the list of unmanaged import functions and / or the set of corresponding libraries. Various hash functions can be used in relation to the determination of the hash. Examples of hash functions include the SHA-256 hash function, the MD5 hash function, the SHA-1 hash function, etc. Various other hash functions may be implemented. The system performs a hash function to determine the unmanaged Imphash corresponding to the.NET file.

[0168] In some embodiments, the system converts or formats a list according to a predefined format. For example, a list of library <-> function (<-> mappingflags) pairs is converted into a comma-separated string. Following such an example, instead of calculating a hash for the following list, kernel32.createprocess kernel32.getthreadcontext kernel32.wow64getthreadcontext kernel32.setthreadcontext kernel32.wow64setthreadcontext kernel32.readprocessmemory the system creates a comma-separated string having a single list item (library <-> function pair). That is, "kernel32.createprocess, kernel32.getthreadcontext, kernel32.wow64getthreadcontext, kernel32.setthreadcontext, kernel32.wow64setthreadcontext, kernel32.readprocessmemory, …". In response to converting / creating a comma-separated string having a single list item, the system determines (e.g., calculates a hash) with respect to such a string.

[0169] At 770, a hash is provided. In some embodiments, the system provides the hash to another system or module, such as a system or module related to determining whether a file is malicious. For example, the hash is provided in response to a call of process 750.

[0170] At 772, a determination is made as to whether process 750 has been completed. In some embodiments, process 750 is determined to be complete in response to determining that no further hash should be calculated (e.g., no further prediction about the file is needed), and the administrator instructs that process 750 should be paused or stopped, etc. In response to determining that process 750 has been completed, process 750 ends. In response to determining that process 750 has not been completed, process 750 returns to 752.

[0171] FIG. 8 is a flowchart of a method for detecting malicious files according to various embodiments. In some embodiments, process 800 is at least partially implemented by system 100 of FIG. 1 and / or system 200 of FIG. 2. In some implementations, process 800 can be implemented by one or more servers, such as in connection with providing services to a network (e.g., network endpoints such as security entities and / or client devices). In some implementations, process 800 can be implemented by a security entity (e.g., a firewall), such as in connection with implementing a security policy regarding files communicated over a network or inside and outside the network. In some implementations, process 800 can be implemented by a client device, such as a laptop, smartphone, personal computer, etc., such as in connection with executing or opening a file such as an email attachment file.

[0172] At 810, an indication that the sample is malicious is received. In some embodiments, the system receives an indication that the sample is malicious, and a sample or hash, signature, or other unique identifier associated with the sample. For example, the system may receive an indication that the sample is malicious from a service such as a security or malware service. The system may receive an indication that the sample is malicious from one or more servers.

[0173] According to various embodiments, an indication that the sample is malicious is received in connection with an update to a previously identified set of malicious files. For example, the system receives an indication that the sample is malicious as an update to a blacklist of malicious files.

[0174] At 820, an association between the sample and the indication that the sample is malicious is stored. In response to receiving an indication that the sample is malicious, the system stores the indication that the sample is malicious in association with the sample or an identifier corresponding to the sample, facilitating a lookup (e.g., a local lookup) as to whether a subsequently received file is malicious. In some embodiments, the identifier corresponding to the sample stored in association with the indication that the sample is malicious includes a hash of the file (or a portion of the file), a signature of the file (or a portion of the file), or another unique identifier associated with the file. In some embodiments, storing the sample in association with an indication as to whether the sample is malicious includes storing an unmanaged Imphash for a.NET file in association with an indication as to whether the sample is malicious.

[0175] At 830, traffic is received. The system can obtain traffic by routing traffic in / across a network, or mediating traffic in / out of a network such as a firewall, or in connection with monitoring email traffic or instant message traffic, etc.

[0176] At 840, a determination is made as to whether the traffic contains malicious files. In some embodiments, the system obtains files from the received traffic. For example, the system identifies a file as an attachment to an email, or identifies a file as being exchanged between two client devices via an instant messaging program or other file exchange program, etc. In response to obtaining a file from the traffic, the system determines whether the file corresponds to a file included in a previously identified set of malicious files, such as a blacklist of malicious files. In response to determining that the file is included in the set of files in the blacklist of malicious files, the system determines that the file is malicious (e.g., the system may further determine that the traffic contains a malicious file).

[0177] In some embodiments, the system determines whether the file corresponds to a file included in a previously identified set of benign files, such as a whitelist of benign files. In response to determining that the file is included in the set of files in the whitelist of benign files, the system determines that the file is not malicious (e.g., the system may further determine that the traffic contains a malicious file).

[0178] In accordance with various embodiments, in response to determining that a file is not included in a set of previously identified malicious files (e.g., a blacklist of malicious files), or a set of previously identified benign files (e.g., a whitelist of benign files), the system considers the file to be non-malicious (e.g., benign).

[0179] In accordance with various embodiments, in response to determining that a file is not included in a set of previously identified malicious files (e.g., a blacklist of malicious files), or a set of previously identified benign files (e.g., a whitelist of benign files), the system queries a malicious file detector to determine whether the file is malicious. For example, the system can quarantine the file until it receives a response from the malicious file detector regarding whether the file is malicious. The malicious file detector can perform an evaluation of whether the file is malicious, such as concurrently with the system's processing of traffic (e.g., in real time in response to a query from the system). The malicious file detector can correspond to the malicious file detector 170 of the system 100 of FIG. 1 and / or the system 200 of FIG. 2.

[0180] In some embodiments, the system determines whether a file is included in a set of previously identified malicious files or a set of previously identified benign files. This is done by calculating a hash associated with the file, or determining a signature or other unique identifier, and performing a lookup in a set of previously identified malicious files or a set of previously identified benign files for files that match the hash, signature, or other unique identifier. Various hashing techniques may be implemented. According to various embodiments, determining whether a file is included in a set of previously identified malicious files or a set of previously identified benign files includes determining an unmanaged Imphash corresponding to the file and determining whether the unmanaged Imphash is included in a historical dataset (e.g., a dataset including the results of previous determinations regarding maliciousness).

[0181] At 840, in response to determining that the traffic does not include malicious files, process 800 proceeds to 850 where the file is treated as non-malicious traffic / information.

[0182] At 840, in response to determining that the traffic includes malicious files, process 800 proceeds to 860 where the file is treated as malicious traffic / information. The system can handle malicious traffic / information based at least in part on one or more policies such as one or more security policies.

[0183] In accordance with various embodiments, handling malicious traffic / information of a file may include performing proactive measures. The proactive measures may be performed in accordance with (e.g., at least partially based on) one or more security policies. As one example, one or more security policies may be pre-set by a network administrator, a customer (e.g., an organization / company), etc., for a service that provides detection of malicious files. Proactive measures that may be performed include the following. Separating the file (e.g., quarantining the file), deleting the file, prompting the user to warn the user that a malicious file has been detected, providing a prompt to the user when the device attempts to open or execute the file, blocking the transmission of the file, updating the blacklist of malicious files (e.g., mapping the hash of the file to an indication that the file is malicious), etc.

[0184] At 870, a determination is made as to whether process 800 has completed. In some embodiments, in response to determining that no further samples should be analyzed (e.g., no further prediction for the file is required), process 800 is determined to be complete, and the administrator indicates that process 800 should be paused or stopped, etc. In response to determining that process 800 has completed, process 800 ends. In response to determining that process 800 has not completed, process 800 returns to 810.

[0185] FIG. 9 is a flowchart of a method for detecting malicious files according to various embodiments. In some embodiments, process 900 is at least partially implemented by system 100 of FIG. 1 and / or system 200 of FIG. 2. In some implementations, process 900 may be implemented by a security entity (e.g., a firewall), and / or an anti-malware application running on a client system, such as in connection with implementing a security policy regarding files communicated over a network or within and outside of a network. In some implementations, process 900 may be implemented by a client device such as a laptop, smartphone, personal computer, etc., such as in connection with executing or opening a file such as an email attachment file.

[0186] At 910, a file is obtained from traffic. The system may obtain traffic by routing traffic within / across a network, or mediating traffic inside / outside of a network such as a firewall, or in connection with monitoring email traffic or instant message traffic. In some embodiments, the system obtains a file from the received traffic. For example, the system may identify the file as an attachment to an email, or as a file being exchanged between two client devices via an instant messaging program or other file exchange program, etc.

[0187] At 920, the signature corresponding to the file is determined. In some embodiments, the system calculates a hash or determines a signature or other unique identifier associated with the file. Various hash techniques may be implemented. For example, the hash technique may be to determine (e.g., calculate) the MD5 hash for the file. In some embodiments, determining the signature corresponding to the file includes calculating the unmanaged Imphash of the.NET file.

[0188] At 930, a dataset of signatures of malicious samples is queried to determine whether the signature corresponding to the file matches the signature from a malicious sample. In some embodiments, the system performs a lookup within the dataset for the signature of a malicious sample for files that match a hash, signature, or other unique identifier. The dataset of signatures of malicious samples may be stored locally in the system or remotely in a storage system accessible to the system.

[0189] According to various embodiments, determining whether a file is included in a set of previously identified malicious files or a set of previously identified benign files includes determining the unmanaged Imphash corresponding to the file and determining whether the unmanaged Imphash is included in a historical dataset (e.g., a dataset including the results of previous determinations regarding maliciousness).

[0190] At 940, a determination as to whether a file is malicious is made, at least in part, based on whether the signature of the file matches the signature of a malicious sample. In some embodiments, the system determines whether a dataset of malicious signatures includes a record that matches the signature of a file obtained from traffic. In response to determining that the historical dataset includes an indication that a file corresponding to an unmanaged Imphash is malicious (e.g., the unmanaged Imphash is included in a field blacklist), the system considers the file obtained from traffic at 910 to be malicious.

[0191] At 950, a file is handled according to whether the file is malicious. In some embodiments, in response to determining that a file is malicious, the system applies one or more security policies with respect to the file. In some embodiments, in response to determining that a file is not malicious, the system handles the file as benign (e.g., the file is handled as normal traffic).

[0192] At 960, a determination is made as to whether process 900 is complete. In some embodiments, process 900 is determined to be complete in response to determining that no further samples should be analyzed (e.g., no further predictions about the file are required), and the administrator indicates that process 900 should be paused or stopped, etc. In response to determining that process 900 is complete, process 900 ends. In response to determining that process 900 is not complete, process 900 returns to 910.

[0193] FIG. 10 is a flowchart of a method for detecting malicious files according to various embodiments. In some embodiments, process 1000 is at least partially implemented by system 100 of FIG. 1 and / or system 200 of FIG. 2. In some implementations, process 1000 can be implemented by a security entity (e.g., a firewall) related to implementing a security policy regarding files communicated over a network or inside and outside the network. In some implementations, process 1000 can be implemented by a client device such as a laptop, smartphone, personal computer, etc. related to executing or opening a file such as an email attachment file.

[0194] At 1010, traffic is received. The system can obtain traffic by routing traffic in / across the network, or mediating traffic to / from the network such as a firewall, or related to monitoring email traffic or instant message traffic.

[0195] At 1020, a file is obtained from the traffic. In some embodiments, the system obtains a file from the received traffic. For example, the system identifies the file as an attachment file to an email, or as a file exchanged between two client devices via an instant messaging program or other file exchange program, etc.

[0196] At 1030, imported API function names are obtained using the.NET header of the file. In some embodiments, 1030 corresponds to or is similar to 420 of process 400 in FIG. 4.

[0197] At 1040, an unmanaged function is determined. In some embodiments, the system determines that a set of unmanaged functions from the imported API function names is obtained using the.NET header of the file. For example, the system determines which of the imported API functions correspond to unmanaged functions. As one example, at least a portion of process 600 may be called in connection with determining unmanaged functions.

[0198] At 1050, a determination is made as to whether the file is malicious. In some embodiments, the system determines whether the file is malicious based at least in part on unmanaged functions (e.g., a set of unmanaged functions imported into the file via the.NET header). In some embodiments, 1050 corresponds to or is similar to 440 of process 400 in FIG. 4. In some embodiments, process 500 in FIG. 5 is performed in connection with 1050.

[0199] In response to determining at 1050 that the file is malicious, process 1000 proceeds to 1060, where one or more security policies are applied with respect to the file. In some embodiments, 1060 corresponds to or is similar to 860 of process 800 in FIG. 8. Thereafter, process 1000 proceeds to 1070.

[0200] In response to determining at 1050 that the file is not malicious, process 1000 proceeds to 1070, where the file is treated as non-malicious traffic. In some embodiments, 1070 corresponds to or is similar to 850 of process 800 in FIG. 8.

[0201] At 1080, a determination is made as to whether process 1000 has been completed. In some embodiments, process 1000 is determined to be complete in response to determining that no further samples should be analyzed (e.g., no further prediction of the file is required), no further traffic should be analyzed, etc., and the administrator indicates that process 1000 should be paused or stopped. In response to determining that process 1000 has been completed, process 1000 ends. In response to determining that process 1000 has not been completed, process 1000 returns to 1010.

[0202] FIG. 11 is a flowchart of a method for detecting malicious files according to various embodiments. In some embodiments, process 1100 is at least partially implemented by system 100 of FIG. 1 and / or system 200 of FIG. 2. In some implementations, process 1100 may be implemented by a security entity (e.g., a firewall), and / or an anti-malware application running on a client system, such as related to implementing a security policy regarding files communicated over a network or inside and outside the network. In some implementations, process 1100 may be implemented by a client device such as a laptop, smartphone, personal computer, etc., such as related to executing or opening a file such as an email attachment file.

[0203] At 1110, traffic is received. In some embodiments, 1110 corresponds to or is similar to 1010 of process 1000 of FIG. 10.

[0204] At 1120, a file is obtained from the traffic. In some embodiments, 1120 corresponds to or is similar to 1020 of process 1000 of FIG. 10.

[0205] At 1130, the imported API function names are obtained using the.NET header of the file. In some embodiments, 1130 corresponds to, or is similar to, 1030 of process 1000 in FIG. 10.

[0206] At 1140, a hash of the list of unmanaged imported API function names is determined. In some embodiments, the system determines the unmanaged functions based at least in part on the imported API function names (e.g., via the.NET header of the file). In some embodiments, the system determines that a set of unmanaged functions is obtained from the imported API function names using the.NET header of the file, the system determines a list of unmanaged functions, and determines a hash based at least in part on the list. As one example, the list includes unmanaged function names corresponding to the set of unmanaged functions.

[0207] At 1150, a mapping of the hash to the file is queried. In some embodiments, the system queries the mapping of the hash to the file based at least in part on the hash of the list of unmanaged imported API function names. For example, the system performs a lookup regarding the mapping of the hash to the file to determine whether the mapping includes the hash of the list of unmanaged imported API function names (e.g., whether the mapping includes a record corresponding to the determined / computed hash). In some embodiments, 1150 corresponds to, or is similar to, 840 of process 800 in FIG. 8 and / or 930 of process 900 in FIG. 9.

[0208] At 1160, a determination is made as to whether the file is malicious. In some embodiments, 1160 corresponds to, or is similar to, 1050 of process 1000 in FIG. 10.

[0209] In response to determining at 1160 that the file is malicious, process 1100 proceeds to 1170, where one or more security policies are applied to the file. In some embodiments, 1170 corresponds to or is similar to 860 of process 800 in FIG. 8. Thereafter, process 1100 proceeds to 1190.

[0210] In response to determining at 1160 that the file is not malicious, process 1100 proceeds to 1180, where the file is treated as non-malicious traffic. In some embodiments, 1180 corresponds to or is similar to 850 of process 800 in FIG. 8.

[0211] At 1190, a determination is made as to whether process 1100 has completed. In some embodiments, in response to determining that no further samples should be analyzed (e.g., no further prediction of the file is required), no further traffic should be analyzed, etc., process 1000 is determined to be complete, and the administrator indicates that process 1000 should be paused or stopped. In response to determining that process 1100 has completed, process 1100 ends. In response to determining that process 1100 has not completed, process 1100 returns to 1110.

[0212] The various examples of the embodiments described herein are described in relation to flowcharts. An example may include several steps that are performed in a particular order, but in accordance with various embodiments, the various steps may be performed in various orders and / or the various steps may be combined into a single step or in parallel.

[0213] The foregoing embodiments have been described in some detail for purposes of clarity of understanding, but the present invention is not limited to the details provided. There are many alternative ways of implementing the present invention. The disclosed embodiments are illustrative and not restrictive.< / mappingflags> < / libraryname> < / functionname> < / mappingflags> < / functionname> < / libraryname> < / mappingflags> < / functionname> < / libraryname> < / libraryname> < / functionname> < / functionname> < / libraryname> < / functionname> < / libraryname> < / crtimplementationdetails> < / crtimplementationdetails> < / crtimplementationdetails> < / crtimplementationdetails> < / crtimplementationdetails> < / crtimplementationdetails> < / crtimplementationdetails> < / functionname> < / libraryname> < / functionname> < / libraryname>

Claims

1. A system comprising one or more processors and a memory, wherein the one or more processors are configured to: receive a sample containing a.NET file, the.NET file including (i) a PE header containing a set of managed imported API functions and (ii) a PE header containing a set of unmanaged imported API functions; obtain unmanaged imported API function names for a second set of the unmanaged imported API functions based at least in part on a.NET header of the.NET file; determine a hash of a list of the unmanaged imported API function names; and determine whether the sample is malware based at least in part on the hash of the list of the unmanaged imported API function names; and the memory is coupled to the one or more processors and is configured to provide instructions to the one or more processors. A system.

2. Obtaining the imported API function names includes: parsing the.NET header of the.NET file; and extracting the imported API function names from the parsed.NET header. The system according to claim 1.

3. The imported API function names are extracted from the parsed.NET header based at least in part on an index table. The system according to claim 2.

4. The index table includes an ImplMap table that indicates a set of unmanaged methods imported in relation to the execution of the.NET file. The system according to claim 3. **Claim 5** The imported API function names are extracted from the parsed.NET header, at least in part, based on the MethodDef table. The system according to claim 2. **Claim 6** Determining the hash of the list of unmanaged imported API function names includes: Determining a set of unmanaged imported API functions based on the imported API function names obtained at least in part based on the.NET header, and Generating the hash of the list of unmanaged imported API function names based at least in part on a predefined hash function. The system according to claim 1. **Claim 7** The predefined hash function includes at least one of the SHA-256 hash algorithm, the MD5 hash algorithm, and the SHA-1 hash algorithm. The system according to claim 6. **Claim 8** The one or more processors are configured to: Send an indication to a security entity that the sample is malicious. The system according to claim 1. The system according to claim 1. **Claim 9** Sending an indication to a security entity that the sample is malicious includes: Updating a blacklist of files considered malicious, where the blacklist of files is updated to include an identifier corresponding to the sample. The system according to claim 8, comprising

10. wherein the security entity is compatible with a firewall The system according to claim 8.

11. Determining whether the sample is malware based at least in part on the hash of the list of unmanaged imported API function names is performed in a sandbox environment The system according to claim 1.

12. Determining whether the sample is malware based at least in part on the hash of the list of unmanaged imported API function names is performed in a security entity The system according to claim 1.

13. The imported API function names are obtained based at least in part on values corresponding to the ImportName field The system according to claim 1.

14. In response to determining that the value corresponding to the ImportName field is not equal to 0 Obtaining the imported API function names comprises Determining a set of one or more library names based at least in part on the #Strings stream of the.NET file The system according to claim 13, comprising

15. In response to determining that the value corresponding to the ImportName field is equal to 0 Obtaining the imported application programming interface function names comprises Obtaining a value from a row of the MethodDef table Obtaining a function name from the #Strings stream based at least in part on the value from the row of the MethodDef table Determining whether the PE header of the.NET file includes an import table, and In response to determining that the PE header of the.NET file does not include an import table, determining a library name corresponding to the function name based at least in part on the import table The system according to claim 13, comprising **Claim 16** Determining a library name corresponding to the function name based at least in part on the import table comprises Parsing the import table to obtain the library name from the corresponding function name The system according to claim 15 **Claim 17** The one or more processors are further configured to Ensure that the library name corresponding to the unmanaged function does not have a file extension Determine a string based at least in part on the library name and the unmanaged function, and Add the string to a list of unmanaged imported API function names As configured The system according to claim 1 **Claim 18** The string is Ensuring that the library name and the name of the unmanaged function do not contain uppercase letters, and Appending the name of the unmanaged function to the library name using a default separator included between the name of the unmanaged function and the library name Determined by The system according to claim 17 **Claim 19** A method comprising: Receiving, by a processor in a system, a sample including a.NET file, wherein the.NET file includes (i) a PE header including a set of managed imported API functions, and (ii) a PE header including a set of unmanaged imported API functions; Obtaining, by the processor, unmanaged imported API function names for a second set of unmanaged imported API functions based at least in part on the.NET header of the.NET file; Determining, by the processor, a hash of a list of unmanaged imported API function names; Determining, by the processor, whether the sample is malware based at least in part on the hash of the list of unmanaged imported API function names; A method comprising the above steps.

20. A computer program stored on a non-transitory computer-readable storage medium, the computer program including a plurality of computer instructions, When the computer instructions are executed by a processor, causing the computer to: Receive a sample including a.NET file, wherein the.NET file includes (i) a PE header including a set of managed imported API functions, and (ii) a PE header including a set of unmanaged imported API functions; Obtain, based at least in part on the.NET header of the.NET file, unmanaged imported API function names for a second set of unmanaged imported API functions; Obtain unmanaged imported API function names for a second set of unmanaged imported API functions; Determine a hash of a list of unmanaged imported API function names; Determining whether the sample is malware based at least in part on the hash of the list of unmanaged imported API function names; A computer program for causing the above to be executed.

21. The list of unmanaged imported API function names comprises pairs, each pair including, for each unmanaged imported API function, (a) an indication of the unmanaged API function name and (b) an indication of the corresponding library name; The step of hashing the list of unmanaged imported API function names comprises: Generating a string representing the list of unmanaged imported API function names based at least in part on the indication of the unmanaged API function name and the indication of the corresponding library name for a plurality of unmanaged API functions; Calculating the hash based on the string representing the list of unmanaged imported API function names; The system according to claim 1, comprising the above.

Citation Information

Patent Citations

  • System and method for categorization of .net applications

    US20190243976A1

  • Securing an application framework from shared library sideload vulnerabilities

    US20210073374A1