Identifying. NET malware with "non-hosted IMPHASH"

By parsing the .NET header of .NET files and extracting the hash values ​​(Imphash) of unmanaged imported API function names to detect malware, this technology solves the problems of low detection rate and high false positive rate in existing technologies, and achieves accurate identification and prevention of malicious .NET software.

CN120974490APending Publication Date: 2025-11-18PALO ALTO NETWORKS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511066860.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2021-12-21
Filing Date
2022-12-05
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately detect malicious .NET software, especially since the PE header structure of .NET files is highly similar to that of benign files, resulting in low detection rates and high false positive rates.

Method used

By parsing the .NET header of a .NET file, extracting a list of unmanaged imported API function names and calculating their hash values ​​(unmanaged Imphash), and matching them against a predefined list of malicious files, it can be determined whether a file is malware.

Benefits of technology

It improves the accuracy of detecting malicious .NET files, reduces the false positive rate, and achieves effective identification and prevention of malware.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974490A_ABST
    Figure CN120974490A_ABST
Patent Text Reader

Abstract

The invention discloses a method and system for detecting malicious files and a computer system. The method includes receiving a sample including a. NET file, obtaining an import API function name based at least in part on a. NET header of the. NET file, determining a hash of a list of non-hosted import API function names, and determining whether the sample is malware based at least in part on the hash of the list of non-hosted import API function names.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This is a divisional application. The parent application is entitled "Identifying .NET Malware with 'Unmanaged IMPASH'", filed on December 5, 2022, with application number 202280077225.X. Background Technology

[0002] Malicious individuals attempt to compromise computer systems in various ways. As an example, such individuals might embed or otherwise include malware (“malware”) in email attachments and deliver or cause the malware to be delivered to unsuspecting users. When executed, the malware harms the victim's computer. Some types of malware instruct the compromised computer to communicate with remote hosts. For example, malware can turn a compromised computer into a “bot” in a “botnet,” receiving instructions from and / or reporting data to a command and control (C&C) server under the control of a malicious individual. One way to mitigate the damage caused by malware is for security companies (or other suitable entities) to attempt to identify malware and prevent it from reaching / executing on end-user computers. Another approach is to try to prevent the compromised computer from communicating with a C&C server. Unfortunately, malware authors are using increasingly sophisticated techniques to obfuscate the work of their software. As an example, some types of malware use Domain Name System (DNS) queries to leak data. Therefore, there is an ongoing need to improve technologies to detect malware and prevent its harm. Attached Figure Description

[0003] Various embodiments of the invention are disclosed in the following detailed description and accompanying drawings.

[0004] Figure 1 It is a block diagram of an environment in which malicious files are detected or suspected, according to various embodiments.

[0005] Figure 2 This is a block diagram of a system for detecting malicious files according to various embodiments.

[0006] Figure 3A This is a diagram of the ImplMap table in the .NET header of a sample .NET file.

[0007] Figure 3B This is a diagram of the ModuleRef table in a sample .NET file.

[0008] Figure 3C This is a diagram of the ImplMap table in the .NET header of a sample .NET file.

[0009] Figure 3D This is a diagram of the MethodDef table in a sample .NET file.

[0010] Figure 3E This is a diagram of the ModuleRef table in a sample .NET file.

[0011] Figure 3F This is a diagram of the import table in a sample .NET file.

[0012] Figure 3G This is a diagram of the MethodDef table in a sample .NET file.

[0013] Figure 4 This is a flowchart of a method for detecting malicious files according to various embodiments.

[0014] Figure 5 This is a flowchart of a method for determining whether a file is malicious, according to various embodiments.

[0015] Figure 6 This is a flowchart of a method for detecting malicious files according to various embodiments.

[0016] Figure 7A This is a flowchart of a method for detecting malicious files according to various embodiments.

[0017] Figure 7B This is a flowchart of a method for detecting malicious files according to various embodiments.

[0018] Figure 8 This is a flowchart of a method for detecting malicious files according to various embodiments.

[0019] Figure 9 This is a flowchart of a method for detecting malicious files according to various embodiments.

[0020] Figure 10 This is a flowchart of a method for detecting malicious files according to various embodiments.

[0021] Figure 11 This is a flowchart of a method for detecting malicious files according to various embodiments. Detailed Implementation

[0022] This invention can be implemented in many ways, including as a process; an apparatus; a system; a composition of matter; a computer program product embodied on a computer-readable storage medium; and / or a processor, such as a processor configured to execute instructions stored on and / or provided by memory coupled to and / or provided thereto. In this specification, these embodiments or any other form in which the invention may take may be referred to as technology. Generally, the order of steps of the disclosed processes can be varied within the scope of this invention. Unless otherwise stated, components described as configured to perform a task (such as processors or memory) can be implemented as general components temporarily configured to perform a task at a given time or manufactured as specific components to perform a task. As used herein, the term "processor" refers to one or more devices, circuits, and / or processing cores configured to process data (such as computer program instructions).

[0023] The following provides a detailed description of one or more embodiments of the present invention, together with accompanying drawings illustrating the principles of the invention. The invention has been described in conjunction with such embodiments, but is not limited to any particular embodiment. The scope of the invention is limited only by the claims, and the invention covers many alternatives, modifications, and equivalents. Numerous specific details are set forth in the following description in order to provide a thorough understanding of the invention. These details are provided for illustrative purposes, and the invention can be practiced according to the claims without requiring some or all of these specific details. For clarity, technical materials known in the art related to the invention have not been described in detail so that the invention is not unnecessarily obscured.

[0024] As used in this document, a security entity is a network node (e.g., a device) that implements one or more security policies relative to information such as network traffic and files. As an example, a security entity can be a firewall. As another example, a security entity can be implemented as a router, switch, DNS resolver, computer, tablet, laptop, smartphone, etc. Various other devices can be implemented as security entities. As yet another example, a security entity can be implemented as an application running on the device, such as an anti-malware application.

[0025] As used herein, malware refers to an application that, whether covertly (and whether illegally), would not be approved by the user if fully aware of its actions. Examples of malware include Trojans, viruses, hacks, spyware, hacking tools, keyloggers, etc. One example of malware is a desktop application that collects the end-user's location and reports it to a remote server (but does not provide the user with location-based services such as map services). Another example of malware is a malicious Android application package (.apk, APK) file that appears to the end-user as a free game but secretly sends SMS messages with added charges (e.g., $10 each), increasing the end-user's phone bill. Another example of malware is the Apple iOS flashlight app that secretly collects the user's contacts and sends those contacts to spammers. The techniques described herein can also be used to detect / thwart other forms of malware (e.g., ransomware). Furthermore, while the malware signatures described herein are for generating malicious applications, the techniques described herein can also be used in various embodiments to generate profiles for other types of applications (e.g., adware profiles, product software profiles, etc.).

[0026] As used in this article, unmanaged code or unmanaged functions refer to imported Win32 API functions, not regular .NET code known as "managed code." As an example, such unmanaged code or unmanaged functions are generally not reflected or included in the PE header of a .NET file; instead, they are imported via the .NET header of the .NET file.

[0027] According to relevant technologies, machine learning models are used to identify malware. These models are trained / developed using Portable Executable (PE) structures based on features such as imports, headers, and sections. The machine learning models use these imports, headers, and sections to distinguish between malicious and benign files. However, the PE file structure of files based on Microsoft Windows PE installers looks extremely similar to malicious and benign files. Therefore, using the PE file structure of Microsoft Windows PE installers to detect malware is not very reliable because distinguishing between malicious and benign files based on such a structure is extremely difficult. For example, using the PE file structure to detect malicious files in Microsoft Windows PE installer files leads to a higher false positive rate and a poorer detection rate. An example of a Microsoft Windows PE installer file used for benign purposes is the Microsoft Windows Nullsoft Scriptable InstallSystem (NSIS) installer, which is commonly used in legitimate products and corporate environments. Each machine learning model trained to analyze PE structures to distinguish between malicious and benign Microsoft Windows PE installer files will not be able to accurately detect malicious files.

[0028] A system, method, and / or apparatus for detecting malicious files is disclosed. The system includes one or more processors and a memory coupled to and configured to provide instructions to the one or more processors. The one or more processors are configured to receive a sample comprising a .NET file, obtain import API function names based at least in part on the .NET header of the .NET file, determine a hash of a list of unmanaged import API function names, and determine whether the sample is malware based at least in part on the hash of the list of unmanaged import API function names.

[0029] According to various embodiments, a system for detecting malicious files is implemented by one or more servers. The one or more servers may provide services to one or more clients and / or security entities. For example, one or more servers detect malicious files or determine / evaluate whether a file is malicious, and provide an indication to one or more clients and / or security entities whether the file is malicious. In response to the determination that a file is malicious and / or in conjunction with updates to the mapping from file to indication of whether a file is malicious (e.g., updates to a blacklist including one or more identifiers associated with one or more malicious files), one or more servers provide an indication that the file is malicious to a security entity. As another example, one or more servers determine whether a file is malicious in response to a request from a client or security entity for the purpose of evaluating whether a file is malicious, and one or more servers provide the result of such determination.

[0030] According to various embodiments, the system for detecting malicious files is implemented by a security entity. For example, the system for detecting malicious files is implemented by a firewall. As another example, the system for detecting malicious files is implemented by an application such as an anti-malware application running on a device (e.g., a computer, laptop, mobile phone, etc.). According to various embodiments, the security entity receives a .NET file, obtains the .NET header of the .NET file, and determines whether the .NET file is malicious based at least in part on the .NET header of the .NET file. In response to determining that the .NET file is malicious, the security entity applies one or more security policies relative to the .NET file. In response to determining that the .NET file is not malicious (e.g., the .NET file is benign), the security entity processes the .NET file as non-malicious business. In some embodiments, the security entity determines whether a file is malicious by determining (e.g., obtaining) imported API function names based at least in part on the .NET header of the .NET file, determining (e.g., calculating) a hash of a list of unmanaged imported API function names, and determining whether the hash of the list of unmanaged imported API function names corresponding to the .NET file matches a hash associated with a file considered malicious. For example, a secure entity performs a lookup relative to a hash (such as the hash of an unmanaged import API function name) to a malicious file to determine whether the mapping includes a matching hash (such as a mapping that includes records of files whose hashes of unmanaged import API function names match a calculated hash of a .NET file).

[0031] Portable Executable (PE) files are typically encoded to import functions from external libraries to interact with various OS components. Relevant domain methods for detecting malware use import sequences, hash the import sequences to obtain hash values, and compare the hash values ​​to a list of known blocks in the "imphash" of the import table. Relevant domain methods for detecting malware imports obtain the imported API function names and their corresponding library names from the PE header of the analyzed file. However, determining API function names and corresponding library names from the PE header, and using such API function names and corresponding library names to detect malware, is not ideal for .NET files because almost all .NET PE files have similar import tables. As an example, most .NET assemblies have a single imported function in the PE header called "_CorExeMain" (EXE) or "_CorDllMain" (DLL). Generally, only a small fraction of .NET assemblies have more imports in the PE header. Such .NET files are generally created using Visual C++ and C++ / CLI extensions. The imported functions included in the PE header are generally determined by the .NET compiler and are not affected by the code itself. This happens because .NET code is not compiled into native assembly, but into intermediate language or intermediate bytecode (MSIL), which is then executed by the .NET runtime.

[0032] Therefore, using import functions extracted from the PE header of a .NET file does not provide accurate detection of malware. However, various .NET malware families still require direct interaction with the Win32 API, such as injecting code into other processes. Such code can be injected into other processes from .NET, but the Win32 functions doing so are not reflected in the import table of the PE header of the .NET file. Instead, the injected functions are generally composed (or imported) via the .NET header of the .NET file. The .NET header is a header included in the .NET file (e.g., in addition to the PE header). For example, the .NET header is different from / different from the PE header of the .NET file. A .NET file includes both the PE header and the .NET header. The .NET header generally includes data streams and tables containing various information related to .NET assembly. One such data stream included in the .NET header is called "#Strings" and includes a list of strings used in the file. The list included in the #Strings stream also includes the names of one or more unmanaged Win32 API functions used. In addition, one of the tables included in the .NET header is called "ImplMap" and contains various information about any imported unmanaged functions.

[0033] Various embodiments parse the .NET header of a .NET file, extract unmanaged imports (e.g., unmanaged functions, libraries, etc.) from one or more fields in the .NET header, and determine whether the .NET file is malicious, at least in part, based on the extracted unmanaged imports. In some embodiments, the system determines a list of unmanaged imports corresponding to the .NET file (e.g., extracted from one or more fields in the .NET header) and determines (e.g., calculates) a hash of the unmanaged import list. The hash of the unmanaged import list can be determined based on a predefined hash function. Examples of hash functions include the SHA-256 hash function, the MD5 hash function, the SHA-1 hash function, etc. Various other hash functions can be implemented. As used herein, unmanaged imphash refers to the value obtained by determining the hash of the unmanaged import list (e.g., unmanaged imports extracted from one or more fields in the .NET header).

[0034] According to various embodiments, information included in the .NET header of a .NET file is used in conjunction with determining whether a file is malicious. In some embodiments, the system uses information included in the ImplMap table and information included in the strings of the "#Strings" data stream to determine a set of unmanaged function <-> library name pairs. The system may determine a list associated with unmanaged import functions of the .NET file (e.g., imported via the .NET header). In some embodiments, the system determines a list hash associated with the unmanaged import functions of the .NET file. For example, the system determines an unmanaged Imphash corresponding to the .NET file. The unmanaged Imphash can be used to determine whether a file is malicious. For example, the system may query a list of files considered malicious (e.g., a blacklist) to determine whether the list includes a record whose unmanaged Imphash matches the unmanaged Imphash determined for the .NET file.

[0035] According to various embodiments, the system analyzes .NET files in a sandbox environment. For example, the system parses .NET files and extracts information from .NET headers within the sandbox environment. The system can be implemented by a virtual machine (VM) operating in the sandbox environment.

[0036] In some embodiments, the system obtains data from third-party services (such as...). The system receives historical information related to file malice (e.g., historical datasets of malicious files and historical datasets of benign files). Third-party services can provide sets of files considered malicious and sets of files considered benign. For example, a third-party service can analyze a file and provide an indication of whether the file is malicious or benign and / or a score indicating the likelihood that the file is malicious. The third-party service can provide unmanaged impashes corresponding to files included in the historical dataset (e.g., a blacklist of files, a whitelist of files, etc.), or the list can include indications of whether historical unmanaged impashes were malicious. The system can receive updates from the third-party service (e.g., at predefined intervals, when updates are available, etc.), such as benign or malicious files with new identifiers, corrections to previous misclassifications, etc. In some embodiments, whether a file in the historical dataset corresponds to a social score (such as a community-based score or rating, e.g., a reputation score) indicating whether a file is malicious or potentially malicious.

[0037] According to various embodiments, security entities and / or network nodes (e.g., clients, devices, etc.) process files at least in part based on indications that the file is malicious and / or that the file matches a file indicated as malicious. In response to receiving an indication that a file (e.g., a sample is malicious), the security network and / or network node may update the mapping of the file to the corresponding indication of whether the file is malicious and / or a blacklist of files. In some embodiments, the security entity and / or network node receives a signature associated with a file (e.g., a sample considered malicious), and the security entity and / or network node stores the signature of the file for use in conjunction with detecting whether a file obtained via network traffic is malicious (e.g., at least in part based on comparing the signature generated for the file with the signatures of files included in a file blacklist). As an example, the signature may be a hash. In some embodiments, the signature of the file is an unmanaged imphash corresponding to such a file.

[0038] Firewalls typically deny or allow network traffic based on sets of rules. These sets of rules are often referred to as policies (e.g., network policies, network security policies, security policies, etc.). For example, a firewall can filter inbound traffic by applying sets of rules or policies to prevent unwanted external traffic from reaching the protected device. A firewall can also filter outbound traffic by applying sets of rules or policies (e.g., allow, block, monitor, notify, or log, and / or specify other actions in firewall rules or policies that can be triggered based on various criteria, such as those described in this document). A firewall can also filter local network (e.g., intranet) traffic by similarly applying sets of rules or policies.

[0039] Security devices (such as security appliances, security gateways, security services, and / or other security devices) may include a variety of security functions (such as firewalls, anti-malware, intrusion prevention / detection, data loss prevention (DLP), and / or other security functions), networking functions (such as routing, quality of service (QoS), workload balancing of network-related resources, and / or other networking functions), and / or other functions. For example, routing functions may be based on source information (such as IP addresses and ports), destination information (such as IP addresses and ports), and protocol information.

[0040] Basic packet-filtering firewalls filter network traffic by inspecting individual packets transmitted over the network (e.g., packet-filtering firewalls or first-generation firewalls, which are stateless packet-filtering firewalls). Stateless packet-filtering firewalls typically inspect each packet itself and apply rules based on the inspected packets (e.g., using a combination of packet source and destination address information, protocol information, and port numbers).

[0041] Application firewalls can also perform application-layer filtering (e.g., application-layer filtering firewalls or second-generation firewalls, which operate at the application layer of the TCP / IP stack). Application-layer filtering firewalls or application firewalls can typically identify certain applications and protocols (e.g., web browsing using Hypertext Transfer Protocol (HTTP), Domain Name System (DNS) requests, file transfer using File Transfer Protocol (FTP), and various other types of applications and protocols such as Telnet, DHCP, TCP, UDP, and TFTP (GSS)). For example, application firewalls can block unauthorized protocols attempting to pass through standard ports (e.g., application firewalls can generally be used to identify unauthorized / out-of-policy protocols attempting to sneak through using non-standard ports for that protocol).

[0042] Stateful firewalls can also perform state-based packet inspection, where each packet is examined within the context of a series of packets associated with a stream of packets traveling through the network. This firewall technique is generally referred to as stateful packet inspection because it maintains a record of all connections traversing the firewall and is able to determine whether a packet is the start of a new connection, part of an existing connection, or an invalid packet. For example, the state of a connection itself can be one of the criteria that triggers rules within a policy.

[0043] Advanced or next-generation firewalls can perform stateless and stateful packet filtering, as well as application-layer filtering, as discussed above. Next-generation firewalls can also perform additional firewall technologies. For example, some newer firewalls, sometimes referred to as advanced or next-generation firewalls, can also identify users and content (e.g., next-generation firewalls). In particular, some next-generation firewalls are expanding the list of applications they can automatically identify to thousands. Examples of such next-generation firewalls are commercially available from Palo Alto Networks (e.g., Palo Alto Networks' PA Series firewalls). For example, Palo Alto Networks' next-generation firewalls enable enterprises to use a variety of identification technologies to identify and control applications, users, and content, not just ports, IP addresses, and packets, such as: APP-ID for accurate application identification, user ID for user identification (e.g., by user or user group), and content ID for real-time content scanning (e.g., controlling web browsing and restricting data and file transfers). These identification technologies allow enterprises to securely enable application usage using business-relevant concepts, rather than following the traditional methods provided by traditional port-blocking firewalls. Moreover, dedicated hardware for next-generation firewalls (e.g., implemented as dedicated appliances) generally provides a higher performance level for application inspection than software running on general-purpose hardware (e.g., security appliances provided by Palo Alto Networks that use dedicated function-specific processing tightly integrated with a single-pass software engine to maximize network throughput while minimizing latency).

[0044] Advanced or next-generation firewalls can also be implemented using virtualized firewalls. Examples of such next-generation firewalls are commercially available from Palo Alto Networks (e.g., Palo Alto Networks' VM series firewalls, which support a variety of commercial virtualization environments, including, for example...). ESXi TM and NSX TM , Netscaler SDX TM , KVM / OpenStack (Centos / RHEL, (and Amazon Web Services (AWS)). For example, virtualized firewalls can support similar or identical next-generation firewalls and advanced threat prevention features available in physical form factor appliances, allowing enterprises to securely enable applications flowing into and across their private, public, and hybrid cloud environments. Automation features such as VM monitoring, dynamic address groups, and REST-based APIs allow enterprises to proactively monitor VM changes to dynamically feed that context into security policies, eliminating policy lag that can occur when VMs change.

[0045] The system improves the detection of malicious files. Furthermore, the system further improves network traffic handling by preventing malicious files from crossing the network (e.g., between nodes within the network) (or improving prevention against them) or preventing malicious files from entering the network. The system identifies .NET files that are considered malicious or potentially malicious, such as those based on the .NET header of the .NET file. Domain-specific detection techniques using the structure of the file's PE header may be insufficient / inaccurate compared to files with similar structures / profiles within malicious or benign files. Furthermore, because .NET files are compiled into an intermediate language, it is difficult to classify files as malicious / benign using machine learning classifiers or manually written YARA rules. YARA is a tool designed (but not limited to) to help malware researchers identify and classify malware samples. YARA rules are used to classify and identify malware samples by creating descriptions of malware families based on text or binary patterns. Furthermore, the system can provide accurate and low-latency updates to security entities (e.g., endpoints, firewalls, etc.) to implement one or more security policies (e.g., pre-defined and / or customer-specific security policies) relative to traffic containing malicious files (e.g., malicious .NET files). Therefore, this system prevents malicious transactions (such as files) from spreading to nodes within the network.

[0046] Figure 1 This is a block diagram of an environment in which malicious files are detected or suspected, according to various embodiments. In the illustrated example, client devices 104 to 108 are (respectively) laptops, desktop computers, and tablets residing in corporate network 110 (belonging to "Acme Corporation"). Data appliance 102 is configured to implement policies (e.g., security policies) regarding communication between client devices (such as client devices 104 and 106) and nodes outside corporate network 110 (e.g., accessible via external network 118). Examples of such policies include those managing business shaping, quality of service, and business routing. Other examples of policies include security policies such as those requiring the scanning of incoming (and / or outgoing) email attachments, website content, files exchanged via instant messaging programs, and / or other file transfers for threats. In some embodiments, data appliance 102 is also configured to enforce policies relative to businesses residing within (or accessing) corporate network 110.

[0047] The techniques described in this article can be used with various platforms (such as desktop computers, mobile devices, gaming platforms, embedded systems, etc.) and / or various types of applications (such as Android .apk files, iOS applications, Windows PE files, Adobe Acrobat PDF files, Microsoft Windows PE installers, etc.). Figure 1In the example environment shown, client devices 104 to 108 are (respectively) laptops, desktop computers, and tablets residing in corporate network 140. Client device 120 is a laptop computer residing outside corporate network 110.

[0048] Data appliance 102 can be configured to work in conjunction with a remote security platform 140. Security platform 140 can provide various services, including performing static and dynamic analysis on malware samples, providing a list of known malicious file signatures as part of a subscription to data appliances (such as data appliance 102), detecting malicious files (e.g., on-demand detection, or updates based on periodic mappings of file-to-file malicious or benign indications), providing the probability that a file is malicious or benign, providing / updating a whitelist of files considered benign, providing / updating files considered malicious, identifying malicious domains, detecting malicious files, predicting whether a file is malicious, and providing an indication that a file is malicious (or benign). In various embodiments, the analysis results (and additional information related to applications, domains, etc.) are stored in a database 160. In various embodiments, security platform 140 includes one or more dedicated commercially available hardware servers (e.g., having one or more multi-core processors, 32GB+ of RAM, one or more gigabit network interface adapters, and one or more hard drives) running a typical server-class operating system (e.g., Linux). Security platform 140 can be implemented across a scalable infrastructure comprising multiple such servers, solid-state drives, and / or other suitable high-performance hardware. Security platform 140 may include several distributed components, including components provided by one or more third parties. For example, some or all of security platform 140 may be implemented using Amazon Elastic Compute Cloud (EC2) and / or Amazon Simple Storage Service (S3). Further, similar to data appliance 102, whenever security platform 140 is referred to perform a task such as storing or processing data, it should be understood that sub-components or multiple sub-components of security platform 140 (whether individually or in collaboration with third-party components) may collaborate to perform that task. As an example, security platform 140 may optionally collaborate with one or more virtual machine (VM) servers to perform static / dynamic analytics. Examples of VM servers are physical machines comprising commercially available server-class hardware (e.g., multi-core processors, 32+ gigabytes of RAM, and one or more gigabit network interface adapters) running commercially available virtualization software such as VMware ESXi, Citrix XenServer, or Microsoft Hyper-V. In some embodiments, the VM server is omitted. Furthermore, the virtual machine server can be under the control of the same entity managing the security platform 140, but it can also be provided by a third party. As an example, the virtual machine server can rely on EC2, where the rest of the security platform 140 is provided by dedicated hardware owned and controlled by the operator of the security platform 140.

[0049] According to various embodiments, security platform 140 includes a DNS tunneling detector 138 and / or a malicious file detector 170. The malicious file detector 170 is used in conjunction with determining whether a file (e.g., a .NET file) is malicious. In response to receiving a sample, the malicious file detector 170 analyzes the file and determines whether the file is malicious. For example, the malicious file detector 170 determines whether the unmanaged imphash corresponding to the analyzed file matches files included in a historical dataset (e.g., a list of files considered malicious, a list of files considered benign, etc.). In some embodiments, the malicious file detector 170 receives a sample including a .NET file, obtains the import API function names based at least in part on the .NET header of the .NET file, determines the hash (unmanaged imphash) of the list of unmanaged import API function names, and determines whether the sample is malware based at least in part on the hash of the list of unmanaged import API function names. In some embodiments, the malicious file detector 170 includes one or more of a .NET file parser 172, an unmanaged function extractor 174, a prediction engine 176, and / or a cache 178.

[0050] The .NET file parser 172 is used in conjunction with obtaining information related to samples (such as .NET files). In some embodiments, the .NET file parser 172 obtains .NET headers and / or information from .NET headers in the .NET file. The .NET file parser 172 obtains one or more data streams included in the .NET header and / or one or more tables included in (or referenced by) the .NET header. For example, the .NET file parser 172 obtains the #Strings stream included in the .NET header. As another example, the .NET file parser 172 obtains an ImplMap from the .NET file (e.g., from the .NET header). In some embodiments, the .NET file parser 172 determines a set of imported functions (e.g., imported API function names) that are imported into the .NET file.

[0051] Unmanaged function extractor 174 is used in conjunction with determining (e.g., obtaining) a set of unmanaged imported functions imported into .NET files. For example, unmanaged function extractor 174 determines the set of unmanaged imported functions based on the set of imported functions (e.g., imported API function names) imported into .NET files. According to various embodiments, unmanaged function extractor 174 uses information included in (or referenced by) the .NET header to determine unmanaged code or unmanaged functions (e.g., unmanaged functions and corresponding libraries). In some embodiments, unmanaged function extractor 174 determines a set of used unmanaged Win32 API functions imported into .NET files. Unmanaged function extractor 174 provides prediction engine 176 with the set of unmanaged imported functions (or names of unmanaged functions) and / or corresponding libraries imported into .NET files.

[0052] Prediction engine 176 is used to determine whether a file (e.g., a .NET file) is malicious. Prediction engine 176 uses information included in the .NET header in conjunction with determining whether the corresponding .NET file is malicious. For example, prediction engine 176 obtains a set of unmanaged import functions (or names of unmanaged functions) and / or corresponding libraries imported into the .NET file from unmanaged function extractor 174. In some embodiments, prediction engine 176 determines a hash (e.g., a hash value) of the set of unmanaged import functions (or names of unmanaged functions) and / or corresponding libraries imported into the .NET file. For example, prediction engine 176 computes an unmanaged Imphash corresponding to the .NET file. Prediction engine 176 determines a list of sets of unmanaged import functions (or names of unmanaged functions) and / or corresponding libraries imported into the .NET file and determines the hash of such a list. The list of sets of unmanaged import functions and / or corresponding libraries is determined according to a predetermined order. For example, the order of unmanaged imported functions and / or their corresponding libraries corresponds to the order in which the unmanaged functions are included in the elements of the .NET header (e.g., the order in which unmanaged functions are included in the #Strings stream and / or tables included in or referenced by .NET tables). Various other orders in which unmanaged functions are added to the list (or the list is arranged) can be implemented. Prediction engine 176 formats the list and / or unmanaged functions (e.g., unmanaged function names and / or their corresponding libraries) according to a predetermined format. Examples of predetermined formats include (i) lowercase alphanumeric strings, (ii) removal of file extensions, (iii) removal of library extensions (e.g., removing .dll from the corresponding library name), (iv) appending unmanaged function names and their corresponding libraries, and (v) the use of a predefined separator between the unmanaged function names and their corresponding libraries. In some embodiments, prediction engine 176 appends the function name (e.g., the unmanaged function name) to the corresponding library and separates the function name from the corresponding library by a dot or period (e.g., "."). In response to identifying the string by appending the function name and library (e.g., using a predefined delimiter such as "."), prediction engine 176 adds such entries to a list of the collection of unmanaged imported functions and / or corresponding libraries.

[0053] According to various embodiments, prediction engine 176 determines a hash relative to a list of unmanaged import functions and / or corresponding libraries. Various hash functions can be used in conjunction with the determined hash. Examples of hash functions include SHA-256, MD5, SHA-1, etc. Various other hash functions can be implemented. Prediction engine 176 uses hash functions to determine the unmanaged imphash corresponding to a .NET file.

[0054] According to various embodiments, prediction engine 176 uses information obtained from the .NET header of a .NET file to determine whether the .NET file is malicious. In some embodiments, prediction engine 176 uses an unmanaged Imphash corresponding to the .NET file in conjunction with determining whether a .NET file is malicious. As an example, in response to determining an unmanaged Imphash corresponding to a .NET file, prediction engine 176 determines whether the unmanaged Imphash matches the unmanaged Imphash of a file considered malicious. As an example, in response to determining an unmanaged Imphash corresponding to a .NET file, prediction engine 176 determines whether the unmanaged Imphash matches the unmanaged Imphash of a file considered benign. In some embodiments, malicious file detector 170 (e.g., prediction engine 176) determines whether information associated with a particular file (e.g., an unmanaged Imphash corresponding to the analyzed .NET file) is included in a dataset of historical files and historical information (e.g., a third-party service, such as VirusTotal) associated with a historical dataset indicating whether a particular file is malicious. TM In response to determining that information related to a specific file is not included in or is unavailable in the dataset of historical files and historical information, the malicious file detector 170 may consider the file benign (e.g., not malicious). Examples of historical information associated with historical files indicating whether a specific file is malicious correspond to... (VT) rating. A file is considered malicious by a third-party service if its VT rating is greater than 0. In some embodiments, historical information associated with historical files indicating whether a file is malicious corresponds to social ratings that indicate a file is malicious or potentially malicious, such as community-based ratings or scores (e.g., reputation ratings). Historical information (e.g., from third-party services, community-based ratings, etc.) indicates whether other vendors or cybersecurity organizations consider a file malicious.

[0055] In some embodiments, a malicious file detector 170 (e.g., prediction engine 176) determines that a received file is newly analyzed (e.g., the file is not in historical information / datasets, not on a whitelist or blacklist, etc.). In response to security platform 140 receiving a file from a security entity (e.g., a firewall) or endpoint within the network, malicious file detector 170 (e.g., script extractor module 172) may detect that the file is newly analyzed. For example, malicious file detector 170 determines that the file is newly analyzed simultaneously with security platform 140 or malicious file detector 170 receiving the file. As another example, malicious file detector 170 (e.g., prediction engine 176) determines that a file is newly analyzed based on a predefined schedule (e.g., daily, weekly, monthly, etc.), such as in conjunction with a batch process. In response to determining that a file has not yet been analyzed for maliciousness (e.g., the system does not include historical information for such files), the malicious file detector 170 determines whether to combine the determination of whether a file is malicious (e.g., in response to determining that the file is a .NET file) with the use of the .NET header associated with the file, and the malicious file detector 170 uses a .NET file parser to parse and / or extract .NET file-related information from the .NET header of the .NET file, etc. In some embodiments, the .NET file parser 172 extracts information from the .NET header in a sandboxed environment of the system.

[0056] According to various embodiments, in response to prediction engine 176 determining that a file is malicious, the system sends an indication that the file is malicious to a security entity (or endpoint, such as a client). For example, malicious file detector 170 sends an indication that the file is malicious to a security entity (e.g., a firewall) or network node (e.g., a client). The indication that the file is malicious may correspond to an update to a blacklist of files (e.g., for malicious files) if the file is considered malicious, or an update to a whitelist of files (e.g., for non-malicious files) if the file is considered benign. In some embodiments, in conjunction with the indication that the file is malicious or benign, malicious file detector 170 sends a hash or signature corresponding to the file. The security entity or endpoint may compute the hash or signature of the file and perform a lookup (e.g., querying whitelists and / or blacklists) against the mapping of the hash / signature to whether the file is malicious / benign. In some embodiments, the hash or signature uniquely identifies the file.

[0057] Cache 178 stores information related to files. In some embodiments, cache 178 stores a mapping indicating whether a file is malicious (or potentially malicious) to a specific file, or a mapping indicating whether a file is malicious (or potentially malicious) to a hash or signature corresponding to the file. Cache 178 may store additional information related to a set of files, such as script information for the files in the set, hashes or signatures corresponding to the files in the set, other unique identifiers corresponding to the files in the set, executable files invoked by the files, pointers included in the files, etc.

[0058] Return to Figure 1 Suppose a malicious individual (using system 120) has created malware 130. The malicious individual wants client devices (such as client device 104) to execute a copy of malware 130, thereby compromising the client devices and turning them into botnets. The compromised client devices can then be instructed to perform tasks (such as cryptocurrency mining or denial of service attacks) and / or report information to external entities (such as those associated with such tasks, such as leaking sensitive corporate data, etc.), such as command and control (C&C) server 150, and receive instructions from C&C server 150 (if applicable).

[0059] While malware 130 might attempt to make an compromised client device communicate directly with C&C server 150 (e.g., by having the client send an email to C&C server 150), such overt communication attempts might be flagged as suspicious / harmful and blocked (e.g., by data appliance 102). Instead of enabling such direct communication, malware authors increasingly use a technique referred to herein as DNS tunneling. DNS is a protocol that translates human-friendly URLs (such as paloaltonetworks.com) into machine-friendly IP addresses (such as 199.167.52.137). DNS tunneling utilizes the DNS protocol to tunnel malware and other data through a client-server model. In the example attachment, a malicious file (such as malware) is sent as an attachment to a message such as an email or instant message. After selecting the attachment, the malware program can be installed on the client device. In the example attack, the attacker registers a domain such as badsite.com. The domain name server points to the attacker's server, on which the tunneling malware program is installed. The attacker infects the computer. Because DNS requests are traditionally allowed to move in and out of security devices, this allows an infected computer to send queries to a DNS resolver (e.g., to kj32hkjqfeuo32ylhkjshdflu23.badsite.com, where the subdomain portion of the query is encoded for consumption by a C&C server). A DNS resolver is a server that relays requests for IP addresses to the root domain server and top-level domain server. The DNS resolver routes the query to the attacker's C&C server, which has a tunneling program installed. A connection is now established between the victim and the attacker via the DNS resolver. This tunnel can be used to leak data or for other malicious purposes.

[0060] Detecting and preventing DNS tunneling attacks is difficult for a variety of reasons. Many legitimate services (such as content delivery networks, web hosting companies, etc.) legitimately use subdomain portions of domain names to encode information to help support the use of those legitimate services. The encoding patterns used by such legitimate services can vary significantly between providers, and benign subdomains may be visually indistinguishable from malicious subdomains. A second reason is that, unlike other fields (such as computer science) which have large amounts of known benign training data and known malicious training data, the training data for DNS queries heavily favors one side (e.g., millions of benign root domain examples and very few malicious examples). Despite these difficulties, and using the techniques described in this paper, malicious domains can be detected efficiently and proactively (e.g., shortly after domain registration), and security policies can be implemented relative to malicious files within or entering the network to block such malicious files, or otherwise warn users or administrators of the malicious file (e.g., sending notifications, providing alerts to users, etc.).

[0061] Figure 1 The environment shown includes three Domain Name System (DNS) servers (122-126). As shown, DNS server 122 is under the control of ACME (for use by computing assets located within network 110), while DNS server 124 is publicly accessible (and can also be used by computing assets located within network 110 as well as other devices located in other networks such as networks 114 and 116). DNS server 126 is publicly accessible but under the control of a malicious operator of C&C server 150. Enterprise DNS server 122 is configured to resolve enterprise domain names into IP addresses and is further configured to communicate with one or more external DNS servers (such as DNS servers 124 and 126) to resolve domain names (if applicable).

[0062] As mentioned above, in order to connect to a legitimate domain (e.g., www.example.com, described as site 128), a client device (such as client device 104) will need to resolve the domain into its corresponding Internet Protocol (IP) address. One way this resolution can occur is that client device 104 forwards a request to DNS servers 122 and / or 124 to resolve the domain. In response to receiving a valid IP address for the requested domain name, client device 104 can use that IP address to connect to website 128. Similarly, in order to connect to a malicious C&C server 150, client device 104 will need to resolve the domain “kj32hkjqfeuo32ylhkjshdflu23.badsite.com” into its corresponding Internet Protocol (IP) address. In this example, malicious DNS server 126 is authoritative on *.badsite.com, and client device 104's request will be forwarded (e.g.) to DNS server 126 for resolution, ultimately allowing C&C server 150 to receive data from client device 104.

[0063] Data appliance 102 is configured to implement policies regarding communication between client devices (such as client devices 104 and 106) and nodes outside the corporate network 140 (e.g., accessible via external network 118). Examples of such policies include those managing service shaping, quality of service, and service routing. Other examples of policies include security policies, such as those requiring the scanning of incoming (and / or outgoing) email attachments, website content, files exchanged via instant messaging programs, and / or other file transfers for potential threats. In some embodiments, data appliance 102 is also configured to enforce policies relative to services residing within the corporate network 140.

[0064] In various embodiments, data appliance 102 includes a DNS module 134 configured to facilitate the determination of whether client devices (e.g., client devices 104-108) attempt to participate in malicious DNS tunneling, and / or prevent (e.g., by client devices 104-108) from connecting to malicious DNS servers. The DNS module 134 may be integrated into appliance 102 (e.g., ...). Figure 1 As shown in the illustration, and in various embodiments can also be operated as a stand-alone device. Furthermore, with Figure 1Like the other components shown, DNS module 134 can be provided by the same entity that provides appliance 102 (or security platform 140), and it can also be provided by a third party (e.g., a third party different from the provider of appliance 102 or security platform 140). Furthermore, in addition to preventing connections to malicious DNS servers, DNS module 134 can also take other actions, such as personalizing tunneling attempts made by clients (indicating that a given client is compromised and should be isolated, or otherwise investigated by an administrator).

[0065] In various embodiments, when a client device (e.g., client device 104) attempts to resolve a domain, the DNS module 134 uses that domain as a query to the security platform 140. This query can be performed concurrently with the domain resolution (e.g., concurrently with requests sent to DNS servers 122, 124, and / or 126 and the security platform 140). As an example, the DNS module 134 can send the query (e.g., in JSON format) to the front end 142 of the security platform 140 via a REST API. Using the processing described in more detail below, the security platform 140 will determine (e.g., using a DNS tunneling detector 138) whether the queried domain indicates a malicious DNS tunneling attempt and will provide the result back to the DNS module 134 (e.g., "malicious DNS tunneling" or "non-tunneling").

[0066] In various embodiments, when a client device (e.g., client device 104) attempts to open a file, such as an email attachment, an instant message received, or otherwise exchanged over a network, or when the client device receives such a file, the DNS module 134 uses that file (or a calculated hash or signature, or other unique identifier, etc.) as a query to the security platform 140. This query can be performed concurrently with the receipt of the file or in response to a request from a user to scan for files. As an example, data appliance 102 can send a query (e.g., in JSON format) to the front end 142 of the security platform 140 via a REST API. Using the processing described in more detail below, the security platform 140 will (e.g., using a malicious file detector 170) determine whether the queried file is malicious (or likely to be malicious) and provide the result back to the DNS module 134 (e.g., "malicious DNS tunneling" or "non-tunneling").

[0067] In various embodiments, the DNS tunneling detector 138 (whether implemented on security platform 140, data appliance 102, or other suitable location / combination) employs a two-pronged approach in identifying malicious DNS tunneling traffic. The first approach uses an anomaly detector 146 (e.g., implemented in Python) to construct a live profile set (156) of DNS traffic for the root domain. The second approach uses signature generation and matching (also referred to herein as similarity detection, and implemented, for example, in Go). These two approaches are complementary. The anomaly detector serves as a general detector capable of identifying previously unknown tunneling traffic. However, the anomaly detector may need to observe multiple DNS queries before performing detection. To block the first DNS tunneling packet, a similarity detector 144 complements the anomaly detector 146 and extracts a signature from the detected tunneling traffic, which can be used to identify cases where an attacker has registered a new malicious tunneling root domain, but has already registered using tools / malware similar to the detected root domain.

[0068] When data appliance 102 receives DNS queries (e.g., from DNS module 134), it provides them to security platform 140, which performs both anomaly detection and similarity detection. In various embodiments, if either detector flags a domain, that domain (e.g., provided in a query received by security platform 140) is classified as a malicious DNS tunneling root domain.

[0069] DNS tunneling detector 138 maintains a set of fully qualified domain names (FQDNs) for each device (from which it receives data), based on its root domain grouping (in... Figure 1 The diagram is shown as domain profile 156. (Although grouping by root domain is generally described in the specification, it should be understood that the techniques described herein can be extended to domains at any level.) In various embodiments, information about received queries for a given domain is kept in the profile for a fixed amount of time (e.g., a sliding time window of ten minutes).

[0070] As an example, DNS query information received from data appliance 102 for various foo.com sites is grouped (in the domain profile of the root domain foo.com) as: G(foo.com) = [mail.foo.com, coolstuff.foo.com, domain1234.foo.com]. A second root domain will have a second profile with similar applicable information (e.g., G(baddomain.com) = [lskjdf23r.baddomain.com, kj235hdssd233.baddomain.com]). Each root domain (e.g., foo.com or baddomain.com) is modeled using a set of characteristics unique to malicious DNS tunneling, making it highly unlikely that benign DNS patterns (e.g., k2jh3i8y35.legitimatesite.com, xxx888222000444.otherlegitimatesite.com) will be misclassified as malicious tunneling. The following are example properties that can be used as features for a given group of domains (i.e., shared root domains) for feature extraction (e.g., extraction into feature vectors).

[0071] In some embodiments, the malicious file detector 170 provides an indication to a security entity, such as data device 102, of whether a file is malicious. For example, in response to determining that a file is malicious, the malicious file detector 170 sends an indication that the file is malicious to data device 102, and the data device can then implement one or more security policies based at least in part on the indication that the file is malicious. One or more security policies may include orphaning files, deleting files, warning or prompting the user about the maliciousness of a file before the user opens / executes the file, etc. As another example, in response to determining that a file is malicious, the malicious file detector 170 provides the security entity with an update to the mapping of a file (or a hash, signature, unmanaged imphalash, or other unique identifier corresponding to the file) to an indication of whether the file is malicious, or an update to a blacklist of malicious files (e.g., identifying file domains) or a whitelist of benign files (e.g., identifying files not considered malicious).

[0072] Figure 2 This is a block diagram of a system for detecting malicious files according to various embodiments. According to various embodiments, system 200 is combined with... Figure 1 The system 100 is implemented as, for example, for a malicious file detector 170. In various embodiments, the system 200 combines... Figure 4 Process 400 Figure 5 Process 500 Figure 6 Process 600 Figure 7A The process 700 Figure 7B Process 750 Figure 8 The process 800 Figure 9 The process 900 Figure 10 Process 1000 and / or Figure 11 The process 1100 is implemented. System 200 can be implemented in one or more servers, security entities such as firewalls, and / or endpoints.

[0073] System 200 can be implemented by one or more devices (such as a server). System 200 can be implemented at various locations on a network. In some embodiments, system 200 is implemented... Figure 1 System 100 includes a malicious file detector 170. As an example, system 200 is deployed as a service, such as a network service (e.g., system 200 determines whether a file is malicious and provides such determination as a service). This service may be provided by one or more servers (e.g., system 200 or the malicious file detector is deployed on a remote server that monitors or receives files transmitted within or outside the network, such as via email attachments, instant messages, etc., and determines whether the file is malicious, and sends / issues file-related notifications or updates, such as indications that the file is malicious). As another example, the malicious file detector is deployed on a firewall.

[0074] In the example shown, system 200 implements one or more modules in conjunction with predicting whether a file (e.g., a newly received file) is malicious, determining the probability that the file is malicious, and / or providing notification or indication that a file is malicious. System 200 includes a communication interface 205, one or more processors 210, a storage device 215, and / or a memory 220. The one or more processors 210 include one or more of a communication module 225, a .NET header extraction module 230, an unmanaged function extraction module 235, a list generation module 240, a prediction module 245, and / or a notification module 250.

[0075] In some embodiments, system 200 includes a communication module 225. System 200 uses communication module 225 to communicate with various nodes or endpoints (e.g., client terminals, firewalls, DNS resolvers, data appliances, other security entities, etc.) or user systems such as administrator systems. For example, communication module 225 provides information to be transmitted to communication interface 205. As another example, communication interface 205 provides information received by system 200 to communication module 225. Communication module 225 is configured to receive files to be analyzed, such as from network endpoints or nodes such as security entities (e.g., firewalls). Communication module 225 is configured to query one or more third-party services (e.g., services that expose information about files, such as third-party ratings or file malice assessments, community-based ratings, assessments or reputations related to files, blacklists and / or whitelists of files, etc.) for information related to the files. For example, system 200 uses communication module 225 to query one or more third-party services. Communication module 225 is configured to receive one or more settings or configurations from an administrator. Examples of one or more settings or configurations include configurations for determining whether a file is malicious, the format of the unmanaged Imphash for determining the .NET file based on which information is being organized / arranged for the .NET file, the hash function used in conjunction with determining the unmanaged Imphash for the file, information related to a domain whitelist (e.g., domains not considered suspicious and for which business or attachments are permitted), and information related to a domain blacklist (e.g., domains considered suspicious and for which business or attachments are restricted).

[0076] In some embodiments, system 200 includes a .NET header extraction module 230. System 200 uses the .NET header extraction module 230 in conjunction with determining whether to extract information related to the header of a file (e.g., from the header) and extracting information about the file (e.g., for analysis to determine if the file is malicious). In some embodiments, the .NET header extraction module 230 receives a file to be analyzed, such as a file included as an email attachment, an instant message, or otherwise transmitted across or within / outside a network. In response to determining that the file is a .NET file, the .NET header extraction module 230 determines to perform the extraction of information related to the file's header. As an example, based on receiving an indication that the file corresponds to a .NET file, the .NET header extraction module 230 determines that the file is a .NET file. As another example, the .NET header extraction module 230 determines that the file is a .NET file based at least in part on the determination that the file includes a .NET header. As another example, the .NET header extraction module 230 determines that a file is a .NET file, at least in part based on the determination that the directory entries in the PE header have non-zero values ​​(e.g., system 200 examines the binary structure of the file and determines whether the values ​​in the optional header values ​​include non-zero values / positions, and if so, then determines that the file is a .NET file). As yet another example, the .NET header extraction module 230 determines that a file is a .NET file, at least in part based on checking whether the file imports the "_CorExeMain" or "_CorDllMain" function.

[0077] In some embodiments, the .NET header extraction module 230 obtains information about .NET headers and / or .NET headers from a .NET file. In response to determining that the file is a .NET file, the .NET header extraction module 230 obtains information related to the file's headers (e.g., from the headers). In some embodiments, the .NET header extraction module 230 determines the .NET headers and obtains imported functions imported (or referenced) by the .NET headers. For example, the .NET header extraction module 230 obtains the names of imported API functions based at least in part on the .NET headers of the .NET file. The .NET header extraction module 230 obtains one or more data streams included in the .NET headers and / or one or more tables included (or referenced) by the .NET headers. For example, the .NET header extraction module 230 obtains the #Strings stream included in the .NET headers. As another example, the .NET header extraction module 230 obtains an ImplMap from the .NET file (e.g., from the .NET headers). The ImplMap may include various information about any unmanaged functions imported by the .NET file. In some embodiments, the .NET header extraction module 230 determines a set of imported functions (e.g., imported API function names) that are imported into a .NET file.

[0078] According to various embodiments, in response to receiving a file to be analyzed to determine whether the file is malicious, system 200 places the file in a sandbox for analysis. In some embodiments, .NET header extraction module 230 extracts information related to the file's header (e.g., from the header itself). As an example, .NET header extraction module 230 extracts header information from the .NET header and / or PE header of the .NET file in the sandbox. For example, system 200 invokes the sandbox for the analysis of a specific file. As another example, system 200 uses a public sandbox for the analysis of various files.

[0079] In some embodiments, system 200 includes an unmanaged function extraction module 235. System 200 uses the unmanaged function extraction module 235 to determine (e.g., obtain) a set of unmanaged functions included in or referenced by a .NET file, such as a list of unmanaged functions imported via the .NET header of the .NET file. For example, the unmanaged function extraction module 235 determines the set of unmanaged imported functions based on the set of imported functions (e.g., imported API function names) imported into the .NET file. According to various embodiments, the unmanaged function extraction module 235 uses information included in (or referenced by) the .NET header to determine unmanaged code or unmanaged functions (e.g., unmanaged functions and corresponding libraries). In some embodiments, the unmanaged function extraction module 235 determines the set of used unmanaged Win32 API functions imported into the .NET file. The unmanaged function extraction module 235 provides the list generation module 240 and / or prediction module 245 with the set of unmanaged imported functions (or names of unmanaged functions) and / or corresponding libraries imported into the .NET file.

[0080] In some embodiments, system 200 includes a list generation module 240. System 200 uses list generation module 240 to generate a list of unmanaged functions and / or corresponding libraries. In some embodiments, system 200 uses list generation module 240 to format a collection of unmanaged imported functions (or names of unmanaged functions) and / or corresponding libraries imported into a .NET file. List generation module 240 formats the list and / or unmanaged functions (e.g., unmanaged function names and / or corresponding libraries) according to a predetermined format. Examples of predetermined formats include (i) lowercase alphanumeric strings, (ii) removal of file extensions, (iii) removal of library extensions (e.g., removing .dll from the corresponding library name), (iv) appending unmanaged function names and corresponding libraries, and (v) the use of a predefined separator between unmanaged function names and corresponding libraries. In some embodiments, list generation module 240 appends function names (e.g., unmanaged function names) to corresponding libraries and separates function names from corresponding libraries by a dot or period (e.g., "."). In response to determining a string by appending function names and libraries (e.g., using predefined delimiters such as "."), list generation module 240 adds such entries to a list of the set of unmanaged imported functions and / or corresponding libraries. According to various embodiments, in response to determining that no other unmanaged functions and / or corresponding libraries need to be added to the list, list generation module 240 provides a list to prediction module 245.

[0081] In some embodiments, system 200 includes a prediction module 245. System 200 uses prediction module 245 to predict whether a file is malicious, or the likelihood that a file is malicious. According to various embodiments, prediction module 245 determines whether a file is malicious based at least in part on information included in (or referenced by) the file's .NET header. For example, prediction module 245 determines an unmanaged imphash corresponding to the file and determines whether the file is malicious based at least in part on the unmanaged imphash.

[0082] According to various embodiments, prediction module 245 determines a hash relative to a list of unmanaged import functions and / or corresponding libraries. Various hash functions can be used in conjunction with the determined hash. Examples of hash functions include SHA-256, MD5, SHA-1, etc. Various other hash functions can be implemented. Prediction module 245 uses the hash function to determine the unmanaged imphash corresponding to a .NET file.

[0083] According to various embodiments, the prediction module 245 uses information obtained from the .NET header of the .NET file to determine whether the .NET file is malicious. In some embodiments, the prediction module 245 uses an unmanaged Imphash corresponding to the .NET file in conjunction with determining whether the .NET file is malicious. As an example, in response to determining the unmanaged Imphash corresponding to the .NET file, the prediction module 245 determines whether the unmanaged Imphash matches the unmanaged Imphash of a file considered malicious. As an example, in response to determining the unmanaged Imphash corresponding to the .NET file, the prediction module 245 determines whether the unmanaged Imphash matches the unmanaged Imphash of a file considered benign. In some embodiments, the prediction module 245 determines whether information associated with a particular file (e.g., the unmanaged Imphash corresponding to the analyzed .NET file) is included in a dataset of historical files and historical information (e.g., a third-party service, such as VirusTotal) associated with a historical dataset indicating whether a particular file is malicious. TM In response to determining that information related to a particular file is not included in or is unavailable in the dataset of historical files and historical information, prediction module 245 may consider the file to be benign (e.g., not malicious). Examples of historical information associated with historical files indicating whether a particular file is malicious correspond to... (VT) rating. A file is considered malicious by a third-party service if its VT rating is greater than 0. In some embodiments, historical information associated with historical files indicating whether a file is malicious corresponds to social ratings that indicate a file is malicious or potentially malicious, such as community-based ratings or scores (e.g., reputation ratings). Historical information (e.g., from third-party services, community-based ratings, etc.) indicates whether other vendors or cybersecurity organizations consider a file malicious.

[0084] System 200 can determine (e.g., calculate) a hash or signature (e.g., an unmanaged imphalite) corresponding to a file and perform lookups against historical information (e.g., whitelists, blacklists, etc.). In some implementations, prediction module 245 corresponds to or is similar to prediction engine 176. System 200 (e.g., prediction module 245) can query a third party (e.g., a third-party service) via communication interface 205 for historical information related to a file (or a set of files or a hash / signature of a file previously considered malicious or benign). System 200 (e.g., prediction module 245) can query the third party at predetermined intervals (e.g., customer-specified intervals, etc.). As an example, prediction module 245 can query a third party daily (or daily during a work week) for registration information of newly analyzed files.

[0085] In some embodiments, system 200 includes a notification module 250. System 200 uses notification module 250 to provide an indication of whether a file is malicious. For example, notification module 250 obtains an indication of whether a file is malicious (or the probability that a file is malicious) from prediction module 245 and provides this indication to one or more security entities and / or one or more endpoints. As another example, notification module 250 provides updates to a whitelist and / or blacklist of files to one or more security entities (e.g., firewalls), nodes, or endpoints (e.g., client terminals). According to various embodiments, notification module 250 obtains a hash, signature, or other unique identifier associated with a file (e.g., an unmanaged imphash corresponding to the file) and combines it with the hash, signature, or other unique identifier associated with the file to provide an indication of whether the file is malicious.

[0086] According to various embodiments, the hash of a file corresponds to a hash using a predetermined hash function (e.g., an unmanaged Imphash using the MD5 hash function, the MD5 hash of the file, etc.). A secure entity or endpoint can compute the hash of a received file (e.g., a file attachment). The secure entity or endpoint can determine whether the computed hash corresponding to the file is included in a set, such as a whitelist of benign files and / or a blacklist of malicious files. If the signature of malware (e.g., the hash of a received file) is included in a set of signatures of malicious files (e.g., a blacklist of malicious files), the secure entity or endpoint can prevent malware from being transmitted to the endpoint (e.g., a client device) and / or accordingly prevent the opening or execution of malware.

[0087] According to various embodiments, storage device 215 includes one or more of file system data 260, hash data 262, and / or cache data 264. Storage device 215 includes shared storage devices (e.g., network storage systems) and / or database data and / or user activity data.

[0088] In some embodiments, file system data 260 includes a database, such as one or more datasets (e.g., one or more datasets of files and / or file attributes, malicious indicators to files or file hashes, unmanaged imphalos, mappings of signatures or other unique identifiers, benign file indicators to files or file hashes, mappings of signatures or other unique identifiers, etc.). File system data 260 includes data such as historical information related to files (e.g., the maliciousness of files), a whitelist of files considered safe (e.g., not suspicious), a blacklist of files considered suspicious or malicious (e.g., files whose probability of being malicious exceeds a predetermined / preset probability threshold), and information associated with suspicious or malicious files.

[0089] Hash data 262 includes data associated with one or more files, such as hash values ​​associated with one or more files. In some embodiments, hash data 262 includes an unmanaged impash of a file (such as a file analyzed by system 200 to determine whether such a file is malicious) or a historical dataset previously assessed for malice by a third party. Hash data 262 includes a mapping of hash values ​​(unmanaged impash) to indications of malice (e.g., indications of whether it is malicious or benign). In some embodiments, hash data 262 includes relationships and associations between files or file-related information (e.g., scripts, attributes such as bytes, structure, etc.) and indications or probabilities that a file is malicious or benign. For example, hash data 262 includes a mapping of hash values ​​(unmanaged impash) to indications of malice (e.g., indications of whether it is malicious or benign).

[0090] Cache data 264 includes information related to predictions of whether a file is malicious. As an example, prediction cache data 264 stores indications of whether one or more files are malicious.

[0091] According to various embodiments, memory 220 includes application execution data 270. Application execution data 270 includes data acquired or used in conjunction with an application (such as an application executing a hash function or an application extracting information from a file). In embodiments, applications include one or more applications that perform one or more of receiving and / or executing queries or tasks, generating reports in response to executed queries or tasks, and / or configuring and / or providing information to a user in response to queries or tasks. Other applications include any other suitable applications (e.g., index maintenance applications, communication applications, machine learning model applications, applications for detecting suspicious transactions, document preparation applications, report preparation applications, user interface applications, data analysis applications, anomaly detection applications, user authentication applications, security policy management / update applications, etc.).

[0092] Figure 3A This is a diagram of the ImplMap table in the .NET header of a sample .NET file. Figure 3A Table 300, illustrated in the figure, provides an ImplMap table for a 32-bit DLL sample. In some embodiments, the ImplMap table for the sample is obtained by inputting the 32-bit DLL sample into a debugger and / or a .NET assembly editor. An example of such a debugger and / or .NET assembly editor is dnSpy. Various other debuggers and / or .NET assembly editors can be implemented.

[0093] According to various embodiments, the system determines the index of the #Strings stream based at least in part on the ImplMap table. In some embodiments, the index of the #Strings stream corresponds to the ImportName table. The #Strings stream generally corresponds to an array of null-terminated strings where most strings of a .NET file reside. The system uses a column labeled ImportName to determine the index of the #Strings stream. The name of a function can be determined using the index value from the ImportName column. In Table 300, the name of the function corresponding to the value of ImportName is provided in an information column.

[0094] In response to determining an index value based on the ImportName and / or function name of the imported function, the system determines the library name (e.g., a DLL) of the corresponding library. In some embodiments, the system determines the library (e.g., the library name) based at least in part on the ModuleRef table of the .NET file.

[0095] Figure 3B This is a diagram of the ModuleRef table in a sample .NET file. Figure 3B Table 310, as shown in the figure, provides information relative to... Figure 3A Table 300 analyzes the ModuleRef table of the 32-bit DLL sample. In some embodiments, the ModuleRef table of the sample is obtained by inputting the 32-bit DLL sample into a debugger and / or a .NET assembly editor.

[0096] According to various embodiments, the system uses the index value of the ImportName field (e.g., a column) from the ImplMap table as an index to determine the corresponding library (e.g., the library name). The system uses the index value for lookup in the Name column of table 310. For example, the Name column is the index for the #Strings stream. The Information columns of the ModuleRef table include indications of the fields. For example, the library name corresponding to the index value 0x18F3 in the Name column is kernel32.dll.

[0097] In conjunction with determining the unmanaged Imphash corresponding to the .NET file, the system determines a list of unmanaged functions. The list of unmanaged functions is generated using a predefined format or syntax. For example, the system obtains the library name corresponding to the imported function and removes any extensions. The system determines whether the library has a ".dll" extension, and if so, removes that extension. As another example, the system formats the function name and library name to convert any uppercase letters to lowercase. In some embodiments, the system constructs a string corresponding to the combination of the function name of the imported function (e.g., an unmanaged imported function) and the library name corresponding to the imported function. For example, the system constructs the string by appending the library name to the function name, and a predefined separator (e.g., ".") is included between the library name and the function name. According to various embodiments, the system determines the string based on the following format: <libraryname> . <functionname>The system then adds the string to the list of unmanaged import functions (e.g., to determine the list of unmanaged Imphash).

[0098] Use and Figure 3A Table 300 and Figure 3B The example sample corresponding to Table 310 lists the following library-function name pairs: kernel32.createprocess, kernel32.getthreadcontext, kernel32.wow64getthreadcontext, kernel32.setthreadcontext, kernel32.wow64setthreadcontext, kernel32.readprocessmemory, kernel32.writeprocessmemory, ntdll.ntunmapviewofsection, kernel32.virtualallocex, kernel32.resumethread, kernel32.loadlibrary, and kernel32.getprocaddress. Entries in the list can be separated by predefined delimiters, such as commas, semicolons, and colons.

[0099] According to various embodiments, in response to determining a list of unmanaged import functions for a .NET file, the system performs a hash of the list using a predefined hash function (e.g., MD5, SHA1, SHA-256, etc.). As an example, using... Figure 3A Table 300 and Figure 3B The unmanaged Imphash for the unmanaged import function list of the samples analyzed in Table 310 (using SHA-256 as the hash function) is: 4b386faf53783c4fd17de6c043fd31374302a4455456b4af3d78abd74c865ff8.

[0100] Figure 3C This is a diagram of the ImplMap table in the .NET header of a sample .NET file. Figure 3C Table 320, illustrated in the figure, provides an ImplMap table for a 32-bit EXE sample. The 32-bit EXE is created using the Visual C++ compiler and C++ / CLI extensions. In some embodiments, the ImplMap table for the sample is obtained by inputting the 32-bit DLL sample into a debugger and / or a .NET assembly editor. An example of such a debugger and / or .NET assembly editor is dnSpy. Various other debuggers and / or .NET assembly editors can be implemented.

[0101] like Figure 3C As illustrated in the diagram, the first three rows of Table 320 contain the values ​​included in the information columns. Therefore, for the first three rows, a combination can be performed. Figure 3A Table 300 and Figure 3B Table 310 describes the process of determining function names and corresponding library names (e.g., determining library-function name pairs). However, Figure 3C The remaining rows of the ImplMap table contain null values ​​or 0. Null or zero entries may occur in the remaining rows because the corresponding functions are not used by the sample authors, but are part of internal runtime methods created by C++ / CLI. In some embodiments, the system determines function names and library names based on the file's MethodDef table. For example, the system uses values ​​included in the MethodForward column of the ImplMap table (e.g., Table 320) as indices to the MethodDef table. As an example, the MemberForward value is referred to as the encoded index in the MethodDef table (e.g., encoded indexes are defined in the .NET specification ECMA-335).

[0102] According to various embodiments, the system from Figure 3C The remaining rows of the MethodForward column in the ImplMap table are used to obtain values. These values ​​are then decoded, and the decoded values ​​are used as indexes in the MethodDef table. For example, the value from row 4 of the MemberForward column is 0x99, which is decoded to correspond to 76. Therefore, 76 is used as the index value to determine information based on the MethodDef table.

[0103] Figure 3D This is a diagram of the MethodDef table in a sample .NET file. Figure 3D Table 330 shown in the figure provides information relative to... Figure 3C Table 320 analyzes the MethodDef table of the sample. In some embodiments, the MethodDef table of the sample is obtained by inputting the sample into a debugger and / or a .NET assembly editor.

[0104] Using the index value obtained from the MethodForward column of ImplMap, the system determines the index of the #Strings stream. For example, the system uses the index value obtained from the MethodForward column of ImplMap as the lookup in the Name column of the MethodDef table. The index value in the Name column corresponding to index 76 from the MethodForward column is 0x1831. As illustrated in Table 330, the Information column of the MethodDef table indicates that the function name is _amsg_exit.

[0105] Figure 3E This is a diagram of the ModuleRef table in a sample .NET file. Figure 3E Table 340, as shown in the figure, provides information relative to... Figure 3C Table 320 analyzes the ModuleRef table of the sample. In some embodiments, the ModuleRef table of the sample is obtained by inputting the sample into a debugger and / or a .NET assembly editor.

[0106] According to various embodiments, the system determines the library name in response to determining the function name (e.g., a function name obtained from the MethodDef table). For example, the system determines the library name in the ImplMap table (e.g., Figure 3C A lookup is performed in Table 320, and the value in the ImportScope column is determined to be 2 for the remaining row. The system uses the index value 2 from the ImportScope column as the index for the ModuleRef table. In response to performing a lookup in the ModuleRef table (e.g., Table 340) using the index value 2 (e.g., from the ImportScope column of the ImplMap table), the system determines that the value in the Name column is also 0. Therefore, the system determines that ModuleRef does not provide an indication of the library name.

[0107] According to various embodiments, in response to determining that ModuleRef does not provide an indication of a library name corresponding to a function, the system determines to use the function name retrieved from the MethodDef table and parses the import table in the PE header. The import table in the PE header includes all statically used functions in the file and the corresponding library names (DLLs) where these statically used functions reside. For example, the system uses a PE header parser to parse the import table in the PE header. An example of a PE header parser is pefile (e.g., an open-source project called pefile), which is a PE header parser coded in Python. In some embodiments, the system performs a 1:1 comparison between the function names obtained from the MethodDef table and the functions in the import table of the PE header. As an example, files written in C++ (such as mixed assembly) can use so-called mangling names as function names in the import table. The process of creating these names is called mangling. Such mangling function names are automatically created by the C++ compiler for each C++ function, except when the function is defined as an external "C".

[0108] Figure 3F This is a diagram of the import table in a sample .NET file. Figure 3F In the example shown, Table 350 is an import table included in the PE header of a .NET file.

[0109] Figure 3G This is a diagram of the MethodDef table in an example .NET file. In the example shown, table 360 ​​is the MethodDef table of the .NET file. In some embodiments, in response to the system determining that a function name obtained from the MethodDef table matches a function name in the import table included in the PE header, the system obtains the corresponding library name from the import table. According to various embodiments, the system performs a lookup on the matching function names between the MethodDef and the import table included in the PE header to obtain the corresponding library name.

[0110] In conjunction with determining the unmanaged Imphash corresponding to the .NET file, the system determines a list of unmanaged functions. The list of unmanaged functions is generated using a predefined format or syntax. For example, the system obtains the library name corresponding to the imported function and removes any extensions. The system determines whether the library has a ".dll" extension, and if so, removes that extension. As another example, the system formats the function name and library name to convert any uppercase letters to lowercase. In some embodiments, the system constructs a string corresponding to the combination of the function name of the imported function (e.g., an unmanaged imported function) and the library name corresponding to the imported function. For example, the system constructs the string by appending the library name to the function name, and a predefined separator (e.g., ".") is included between the library name and the function name. According to various embodiments, the system determines the string based on the following format: <libraryname> . <functionname>The system then adds the string to the list of unmanaged import functions (e.g., to determine the list of unmanaged Imphash).

[0111] Use and Figure 3A Table 300 and Figure 3B The example samples corresponding to Table 310, the list of library-function name pairs are: ntdll.zwallocatevirtualmemory, ntdll.zwfreevirtualmemory, ntdll.ldrgetprocedureaddress, msvcr80._amsg_exit, kernel32.sleep, etc. <crtimplementationdetails>.throwmoduleloadexception、. <crtimplementationdetails>.throwmoduleload exception、. <crtimplementationdetails>.dodlllanguagesupportvalidation、. <crtimplementationdetails>.thrownestedmoduleloadexception、.<crtimple mentationdetails>.registermoduleuninitializer、. <crtimplementationdetails>`.docallbackindefaultdomain`, `msvcr80._cexit`, `msvcr80._encode_pointer`, `msvcr80._decode_pointer`, `msvcr80._encoded_null`, `msvcr80.__frameunwindfilter`. Entries in the list can be separated by predefined delimiters, such as commas, semicolons, colons, etc. Library-function pairs. <crtimplementationdetails>The `.throwModuleLoadException` is included twice because the function is a so-called overloaded function. Such functions have the same name but differ in parameter types, numeric values, or return values. Therefore, they are different functions.

[0112] According to various embodiments, in response to determining a list of unmanaged import functions for a .NET file, the system performs a hash of the list using a predefined hash function (e.g., MD5, SHA1, SHA-256, etc.). As an example, using... Figure 3A Table 300 and Figure 3B The unmanaged Imphash for the unmanaged import function list of the samples analyzed in Table 310 (using SHA-256 as the hash function) is: cd2f27b642a85c3f0c10db4e887d504d3b4c0882f9264994367c0c9b4ea7a537.

[0113] In some embodiments, the system transforms or formats a list according to a predefined format. For example, a list of library <-> function (<-> mappingflags) name pairs is transformed into a comma-separated string. The system creates a comma-separated string using a single list item (library <-> function pair). In response to transforming / creating a comma-separated string using a single list item, the system determines (e.g., calculates a hash) relative to such a string.

[0114] Figure 4 This is a flowchart of a method for detecting malicious files according to various embodiments. In some embodiments, process 400 is at least partially performed by... Figure 1 System 100 and / or Figure 2 System 200 is implemented. In some implementations, process 400 may be implemented by one or more servers, such as in conjunction with providing services to a network (e.g., security entities and / or network endpoints, such as client devices). In some implementations, process 400 may be implemented by a security entity (e.g., a firewall), such as in conjunction with implementing security policies relative to files transferred across or within / outside the network. In some implementations, process 400 may be implemented by a client device (e.g., a laptop computer, smartphone, personal computer, etc.), such as in conjunction with executing or opening files (e.g., email attachments).

[0115] At 410, a sample is received. In some embodiments, the system receives a sample (e.g., a .NET file) from a security entity (e.g., a firewall), an endpoint (e.g., a client device), etc. For example, in response to determining that a file has been attached to a communication such as an email or instant message, the security entity or endpoint provides (e.g., sends) the file to the system. Receiving the sample may be combined with a request to determine whether the file is malicious.

[0116] When process 400 is implemented by a security entity, samples can be received, such as in conjunction with routing traffic to applicable network endpoints (e.g., a firewall obtaining samples from email attachments in emails directed to a client device). When process 400 is implemented by a client device, samples can be received by an application or layer that monitors incoming / outgoing information. For example, processes (e.g., applications, operating system processes, etc.) can run in the background to monitor and obtain email attachments, files exchanged via instant messaging programs, etc.

[0117] At 420, the imported API function name is obtained using the sample's .NET header. In some embodiments, in response to receiving a sample and / or a request to evaluate whether a sample (e.g., a .NET file) is malicious, the system parses the sample to obtain information related to (e.g., included therein) the sample's .NET header.

[0118] According to various embodiments, the system determines the .NET header and obtains the imported functions imported (or referenced) by the .NET header. For example, the system obtains the imported API function names based at least in part on the .NET header of the .NET file. The system obtains one or more data streams included in the .NET header and / or one or more tables included (or referenced) by the .NET header. For example, the system obtains the #Strings stream included in the .NET header. As another example, the system obtains an ImplMap from the .NET file (e.g., from the .NET header). The ImplMap may include various information about any unmanaged functions imported into the .NET file. In some embodiments, the system determines a set of imported functions (e.g., imported API function names) imported into the .NET file.

[0119] In some embodiments, the system determines (e.g., obtains) a set of unmanaged functions included in or referenced by a .NET file, such as a list of unmanaged functions imported via the .NET header of the .NET file. For example, the system determines the set of unmanaged imported functions based on the set of imported functions (e.g., imported API function names) imported into the .NET file. The system uses information included in (or referenced by) the .NET header to determine unmanaged code or unmanaged functions (e.g., unmanaged functions and their corresponding libraries). In some embodiments, the system determines the set of used unmanaged Win32 API functions imported into the .NET file.

[0120] At 430, a hash of the list of unmanaged imported API function names is determined. In some embodiments, in response to obtaining the imported API function names, the system determines (e.g., computes) a hash of the list of unmanaged imported API function names. For example, the system computes an unmanaged Imphash corresponding to a .NET file.

[0121] Combined with a hash of the list that determines the names of unmanaged imported API functions, the system determines a list of collections of unmanaged imported functions (or the names of unmanaged functions) and / or corresponding libraries imported into .NET files, and determines the hash of such a list. The list of collections of unmanaged imported functions and / or corresponding libraries is determined according to a predetermined order. For example, the order of unmanaged imported functions and / or corresponding libraries corresponds to the order in which unmanaged functions are included in elements of the .NET header (e.g., the order in which unmanaged functions are included in #Strings streams and / or tables included in or referenced by .NET tables). Various other orders in which unmanaged functions are added to the list (or the list is arranged) can be implemented. The system formats the list and / or unmanaged functions (e.g., unmanaged function names and / or corresponding libraries) according to a predetermined format or syntax. Examples of predefined formats include (i) lowercase alphanumeric strings, (ii) removal of file extensions, (iii) removal of library extensions (e.g., removing .dll from the corresponding library name), (iv) appending unmanaged function names and corresponding libraries, and (v) the use of predefined separators between unmanaged function names and corresponding libraries. In some embodiments, the system appends the function name (e.g., an unmanaged function name) to the corresponding library and separates the function name from the corresponding library by a dot or period (e.g., "."). In response to determining the string by appending the function name and library (e.g., using predefined separators such as "."), the system adds such entries to a list of the set of unmanaged imported functions and / or corresponding libraries.

[0122] According to various embodiments, the system determines a hash relative to a list of unmanaged import functions and / or corresponding libraries. Various hash functions can be used in conjunction with the determined hash. Examples of hash functions include SHA-256, MD5, SHA-1, and others. Various other hash functions can be implemented. The system uses a hash function to determine the unmanaged imphash corresponding to a .NET file.

[0123] At 440, a determination is made as to whether the sample is malicious. In some embodiments, in response to determining a hash of a list of unmanaged import API function names, the system uses a hash (e.g., unmanaged Imphash) in conjunction with the determination of whether the sample is malicious.

[0124] In some embodiments, the system uses an unmanaged impash corresponding to the .NET file in conjunction with determining whether a .NET file is malicious. As an example, in response to determining the unmanaged impash corresponding to a .NET file, the system determines whether the unmanaged impash matches the unmanaged impash of a file considered malicious. If the unmanaged impash of a sample (e.g., the file being analyzed) matches the unmanaged impash of a malicious file in a historical dataset (e.g., records included in a malicious file blacklist), the system considers the sample malicious. As an example, in response to determining the unmanaged impash corresponding to a .NET file, the system determines whether the unmanaged impash matches the unmanaged impash of a file considered benign. If the unmanaged impash of a sample (e.g., the file being analyzed) matches the unmanaged impash of a benign file in a historical dataset (e.g., records included in a benign file whitelist), the system considers the sample benign. In some embodiments, the system determines whether information associated with a particular file (e.g., the unmanaged impash corresponding to the analyzed .NET file) is included in a dataset of historical files and historical information (e.g., a third-party service, such as VirusTotal) associated with a historical dataset indicating whether a particular file is malicious. TM In this context, as an example, in response to determining that information related to a particular file is not included in or is unavailable in the dataset of historical files and historical information, the system assumes the file is benign (e.g., not malicious). An example of historical information associated with historical files indicating whether a particular file is malicious corresponds to... (VT) rating. A file is considered malicious by a third-party service if its VT rating is greater than 0. In some embodiments, historical information associated with historical files indicating whether a file is malicious corresponds to social ratings that indicate a file is malicious or potentially malicious, such as community-based ratings or scores (e.g., reputation ratings). Historical information (e.g., from third-party services, community-based ratings, etc.) indicates whether other vendors or cybersecurity organizations consider a file malicious.

[0125] In response to the determination that the sample is malicious at 440, process 400 continues to 450, where an indication that the sample is malicious is provided.

[0126] In response to determining at 440 that the sample is malicious, process 400 continues to 450, where an indication that the sample is malicious is provided. For example, the indication that the sample is malicious may be provided to a component that received the sample. As an example, the system provides an indication that the sample is malicious to a security entity. As another example, the system provides an indication that the sample is malicious to a client device. As an example, a security entity provides an indication that the sample is malicious to a client device. In some embodiments, the indication that the sample is malicious is provided to a user, such as a user of the client device and / or a network administrator.

[0127] According to various embodiments, proactive measures can be taken in response to receiving an indication that a sample is malicious. These proactive measures can be performed based on (e.g., at least in part on) one or more security policies. As an example, one or more security policies may be preset by a network administrator, a customer (e.g., an organization / company) providing a malicious file detection service, etc. Examples of proactive measures that can be performed include: orphaning files (e.g., isolating files), deleting files, alerting the user to the detection of a malicious file, providing a prompt to the user when a device attempts to open or execute a file, blocking file transfer, updating a malicious file blacklist (e.g., a mapping from a file hash to an indication that the file is malicious), etc.

[0128] In response to determining at 440 that the sample is not malicious, process 400 continues to 460. In some embodiments, in response to determining that the sample is not malicious, the mapping of the file (or the hash / signature of the file) to the indication that the file is not malicious is updated. For example, the benign file whitelist is updated to include the sample or the hash, signature, or other unique identifier associated with the sample.

[0129] At 460, a determination is made regarding whether process 400 is complete. In some embodiments, in response to determining that no further sample analysis is needed (e.g., further prediction of the document is not required), the administrator instructs to pause or stop process 400, etc., to determine whether process 400 should complete. In response to determining that process 400 is complete, process 400 ends. In response to determining that process 400 is not complete, process 400 returns to 410.

[0130] Figure 5 This is a flowchart of a method for determining whether a file is malicious, according to various embodiments. In some embodiments, process 500 is at least partially performed by... Figure 1 System 100 and / or Figure 2 System 200 is implemented. In some implementations, process 500 may be implemented by one or more servers, such as in conjunction with providing services to a network (e.g., security entities and / or network endpoints, such as client devices). In some implementations, process 500 may be implemented by a security entity (e.g., a firewall), such as in conjunction with implementing security policies relative to files transferred across or within / outside the network. In some implementations, process 500 may be implemented by a client device (e.g., a laptop computer, smartphone, personal computer, etc.), such as in conjunction with executing or opening files (e.g., email attachments).

[0131] According to various embodiments, with Figure 4 The process 400 is combined with the call process 500.

[0132] At 510, the hash of the list is obtained. In some embodiments, the system obtains (e.g., receives, determines, etc.) an unmanaged Imphash. For example, the system receives the hash of a list of unmanaged imported API function names.

[0133] At position 520, hashes are used in conjunction with query mappings. In response to a received list of hashes (e.g., unmanaged Imphashes), the system performs a lookup against a historical dataset of malicious and / or benign files. For example, the historical dataset might include associations between unmanaged Imphashes and indications of whether the corresponding file is malicious or benign.

[0134] In some embodiments, the system uses an unmanaged Imphash corresponding to the .NET file in conjunction with determining whether a .NET file is malicious. As an example, in response to determining an unmanaged Imphash corresponding to a .NET file, the system determines whether the unmanaged Imphash matches an unmanaged Imphash of a file considered malicious. As an example, in response to determining an unmanaged Imphash corresponding to a .NET file, the system determines whether the unmanaged Imphash matches an unmanaged Imphash of a file considered benign. In some embodiments, the system determines whether information associated with a particular file (e.g., an unmanaged Imphash corresponding to the analyzed .NET file) is included in a dataset of historical files and historical information (e.g., third-party services such as VirusTotal) associated with a historical dataset indicating whether a particular file is malicious. TM In this context, as an example, in response to determining that information related to a particular file is not included in or is unavailable in the dataset of historical files and historical information, the system assumes the file is benign (e.g., not malicious). An example of historical information associated with historical files indicating whether a particular file is malicious corresponds to... (VT) rating. A file is considered malicious by a third-party service if its VT rating is greater than 0. In some embodiments, historical information associated with historical files indicating whether a file is malicious corresponds to social ratings that indicate a file is malicious or potentially malicious, such as community-based ratings or scores (e.g., reputation ratings). Historical information (e.g., from third-party services, community-based ratings, etc.) indicates whether other vendors or cybersecurity organizations consider a file malicious.

[0135] At 530, a determination is made as to whether the mapping indicates that the hash corresponds to a malicious file.

[0136] In response to the determination at 530 that the mapping indicator hash corresponds to a malicious file, process 500 continues to 530, where the sample is determined to be malicious.

[0137] In response to determining at 530 that the mapping indication hash does not correspond to a malicious file, process 500 continues to 550, where the sample is determined to be non-malicious. In some embodiments, in response to determining that the mapping from hash to file does not include an indication that the hash is mapped to a malicious file, the system determines that the sample is benign. As an example, the system determines that the hash is not included in the mapping from hash to malicious file. As another example, the system determines that the mapping does not include a record (or indication of a malicious file) mapped to a malicious file.

[0138] If the unmanaged Imphash of a sample (e.g., the file being analyzed) matches the unmanaged Imphash of a malicious file in a historical dataset (e.g., a record included in a malicious file blacklist), the system considers the sample to be malicious.

[0139] The system considers the sample to be benign if its unmanaged impash matches the unmanaged impash of benign files in a historical dataset (e.g., records included in a whitelist of benign files). In some embodiments, the system considers a file to be benign (e.g., not malicious) in response to determining that information associated with a particular file is not included in or is unavailable in a dataset of historical files and historical information.

[0140] At position 560, a malicious result is provided. In some embodiments, the system provides an indication that the hash corresponds to a malicious file. For example, the system provides an indication that the file corresponding to the hash is malicious.

[0141] At 570, a determination is made regarding whether process 500 is complete. In some embodiments, in response to determining that no further hash analysis is needed (e.g., further prediction of the file is not required), the administrator instructs to pause or stop process 500, etc., to determine that process 500 is complete. In response to determining that process 500 is complete, process 500 ends. In response to determining that process 500 is not complete, process 500 returns to 510.

[0142] Figure 6 This is a flowchart of a method for detecting malicious files according to various embodiments. In some embodiments, process 600 is at least partially performed by... Figure 1 System 100 and / or Figure 2 The system 200 is implemented as follows. In some implementations, process 600 may be implemented by one or more servers, such as in conjunction with providing services to a network (e.g., security entities and / or network endpoints, such as client devices). In some implementations, process 600 may be implemented by a security entity (e.g., a firewall), such as in conjunction with implementing security policies relative to files transferred across or within / outside the network. In some implementations, process 600 may be implemented by a client device (e.g., a laptop computer, smartphone, personal computer, etc.), such as in conjunction with executing or opening files (e.g., email attachments).

[0143] At position 602, the sample is received.

[0144] At 604, the .NET assembly corresponding to the sample is obtained. In some embodiments, the sample is compressed (e.g., in ZIP format, etc.), and the .NET assembly is extracted from the compressed file. In some embodiments, obtaining the .NET assembly includes determining that the sample is a .NET file.

[0145] At 606, a determination is made as to whether the .NET assembly includes the ModuleRef table. If it is determined that the .NET assembly does not include the ModuleRef table, process 600 ends. Conversely, if it is determined that the .NET assembly includes the ModuleRef table, process 600 continues to 608.

[0146] At point 608, a determination is made as to whether the .NET assembly includes the ImplMap table. If it is determined that the .NET assembly does not include the ImplMap table, process 600 ends. Conversely, if it is determined that the .NET assembly includes the ImplMap table, process 600 continues to 610.

[0147] At position 610, the ImportName value is obtained. In some embodiments, the system obtains the ImportName value from the ImportName column of the ImplMap table. For example, the system obtains the ImportName value from the applicable rows (e.g., selected rows) of the ImportName column. In some embodiments, the system iterates over the rows of ImportName to obtain the value for each applicable row of the ImportName column.

[0148] At 612, it is determined whether the ImportName value obtained at 610 is equal to 0. In response to determining that the ImportName value obtained at 610 is equal to 0, process 600 continues to 626. In response to determining that the ImportName value obtained at 610 is not equal to 0, process 600 continues to 614.

[0149] At 614, the function name is obtained from the #Strings stream. In some embodiments, the system uses the selected line to obtain the function name from the #Strings stream of .NET assembly.

[0150] At position 616, the value is obtained from the ImportScope column. In some embodiments, the system obtains the value from the ImportScope column of the ImplMap table. The system obtains the ImplMap table using a .NET header (e.g., the .NET header includes the ImplMap table). The value in the ImportScope column is obtained from the selected row (e.g., the row in the ImplMap table from which the ImportName value is obtained and / or the row in the Information column from which the #Strings flow information is obtained). In some embodiments, the value from the ImportScope column is used as an index in conjunction with performing a lookup of the ModuleRef table (e.g., for the name of the corresponding library).

[0151] At 618, the ImportScope value is used as a row index in the ModuleRef table to retrieve the value. In some embodiments, the system obtains the ModuleRef table using a .NET header (e.g., the .NET header includes the ModuleRef table). The system uses the value obtained from the ImportScope column as an index to perform a lookup in the ModuleRef table. For example, the system uses the value obtained from the ImportScope column to determine the row in the ModuleRef table from which the system will retrieve the value from the Name column. In some embodiments, the value obtained from the Name column of the ModuleRef table is used as an index for the #Strings stream.

[0152] At position 620, the library name is obtained from the #Strings stream. In some embodiments, the system obtains the #Strings stream using the .NET header (e.g., the .NET header includes the #Strings stream). The system uses the value obtained from the name column as an index to perform a lookup in the #Strings stream.

[0153] At position 624, the string corresponding to the library name-function name pair is added to the list. For example, the library name-function name pair is added to the list of unmanaged imported functions. In some embodiments, the system generates the string according to a predetermined format or syntax.

[0154] At 626, a determination is made as to whether the .NET assembly includes the MethodDef table. In response to determining that the .NET assembly does not include the MethodDef table, procedure 600 continues to 634, where the system assumes that the .NET header has empty function names. Conversely, in response to determining that the .NET assembly includes the MethodDef table, procedure 600 continues to 628.

[0155] At position 628, the value is obtained from the MemberForwarded column of the ImplMap table. The system obtains the ImplMap table using a .NET header (e.g., the .NET header includes the ImplMap table). The value in the MemberForwarded column is obtained from the selected row (e.g., the row in the ImplMap table from which the MemberForwarded value is obtained). In some embodiments, the value from the MemberForwarded column is used as an index in conjunction with performing a lookup of the ModuleRef table (e.g., against the name of the corresponding library).

[0156] At 630, the MemberForwarded value is used as a row index in the ModuleRef table to retrieve the value. In some embodiments, the system obtains the ModuleRef table using a .NET header (e.g., the .NET header includes the ModuleRef table). The system uses the value obtained from the MemberForwarded column as an index to perform a lookup in the ModuleRef table. For example, the system uses the value obtained from the Name column to determine the row in the ModuleRef table from which the system will retrieve the value from the #Strings stream.

[0157] At position 632, the function name is obtained from the #Strings stream. In some embodiments, the system obtains the #Strings stream using the .NET header (e.g., the .NET header includes the #Strings stream). The system uses the value obtained from the name column of the ModuleRef table as an index to perform a lookup in the #Strings stream.

[0158] At 636, the system determines whether the PE header contains an import table. In some embodiments, this is done in conjunction with determining the library name corresponding to a function name (e.g., a function name obtained from the #Strings stream).

[0159] In response to determining at 636 that the PE header does not have an import table, process 600 continues to 638, where the system considers the library name to be an empty library name. Conversely, in response to determining at 636 that the PE header has an import table, process 600 continues to 640, where the library name corresponding to the function name is obtained. In some embodiments, the system resolves the import table to obtain the library name corresponding to the function name. In response to obtaining the library name, process 600 continues to 624.

[0160] After adding the function name and corresponding library name to the list at 624, process 600 continues to 642.

[0161] At 642, a determination is made regarding whether the overall population (e.g., generation) of the list is complete. In some embodiments, in response to determining that no further imported functions and / or corresponding libraries are to be added to the list (e.g., no further imported functions are included in or referenced in the .NET header), the administrator instructs to pause or stop process 600, etc., and determines that the overall / generation of the list is complete. In response to determining at 642 that no further imported functions and / or corresponding libraries are to be added to the list, process 600 continues to 644, where the system determines (e.g., calculates) the hash of the list (e.g., the list of unmanaged imported functions). In response to determining that process 600 is not complete, process 600 returns to 602. In some embodiments, in response to calculating the hash (e.g., an unmanaged Imphash), the system determines whether the sample is malicious, at least in part, based on the unmanaged Imphash. For example, in response to calculating the hash, a call is made... Figure 5 The process is 500.

[0162] In some embodiments, the system calls Figure 7A Process 700 or Figure 7B The process 750 determines the hash of the list.

[0163] Figure 7A This is a flowchart of a method for detecting malicious files according to various embodiments. In some embodiments, process 700 is at least partially performed by... Figure 1 System 100 and / or Figure 2 System 200 is implemented. In some implementations, process 600 may be implemented by one or more servers, such as in conjunction with providing services to a network (e.g., security entities and / or network endpoints, such as client devices). In some implementations, process 700 may be implemented by a security entity (e.g., a firewall), such as in conjunction with implementing security policies relative to files transferred across or within / outside the network. In some implementations, process 700 may be implemented by a client device (e.g., a laptop computer, smartphone, personal computer, etc.), such as in conjunction with executing or opening files (e.g., email attachments).

[0164] At 702, obtain the library name and / or function name. In some embodiments, obtain the function name and corresponding library name obtained using the .NET header, such as in combination with a list of generated unmanaged import functions.

[0165] At step 704, it is determined whether the library name has a file extension. In response to the determination at step 704 that the library name has a file extension, process 700 continues to step 706, where the file extension is removed. In response to the determination at step 704 that the library name does not have a file extension, process 700 continues to step 708.

[0166] At 708, the function name and library name are formatted. In some embodiments, the system formats the library name-function name pair according to a predetermined format or syntax. As another example, the system formats the function name and library name to convert any uppercase letters to lowercase.

[0167] At 710, the function name and library name are combined. In some embodiments, the system constructs a string corresponding to the combination of the function name of the imported function (e.g., an unmanaged imported function) and the library name corresponding to the imported function. For example, the system constructs the string by appending the library name to the function name, and a predefined separator (e.g., ".") is included between the library name and the function name. According to various embodiments, the system determines the string based on a format: <libraryname> . <functionname>The system then adds the string to the list of unmanaged import functions (e.g., to determine the list of unmanaged Imphash).

[0168] According to various embodiments, various formats or syntaxes can be implemented by combining system-combined function names and library names. Examples of predefined formats include (i) lowercase alphanumeric strings, (ii) removal of file extensions, (iii) removal of library extensions (e.g., removing .dll from the corresponding library name), (iv) appending unmanaged function names and corresponding libraries, and (v) the use of predefined separators between unmanaged function names and corresponding libraries, and (vi) combining library names and function names to determine the order of strings (e.g., ...). <libraryname> . <functionname>or <functionname> . <libraryname>In some embodiments, the system appends the function name (e.g., an unmanaged function name) to the corresponding library and separates the function name from the corresponding library by a dot or period (e.g., "."). In response to determining the string by appending the function name and library (e.g., using a predefined separator such as "."), the system adds such entries to a list of the collection of unmanaged imported functions and / or corresponding libraries. Various other predefined formats / syntax can be implemented.

[0169] At 712, the combination of function name and library name is added to the list of unmanaged imported functions.

[0170] At 714, a determination is made as to whether more functions should be added to the list. For example, the system determines whether more unmanaged import functions should be added to the file's list of unmanaged import functions. In response to determining at 714 that no further functions should be added to the list, process 700 proceeds to 716. In response to determining that additional functions should be added to the list, process 700 returns to 702.

[0171] At position 716, the hash is calculated relative to a list. According to various embodiments, the system determines the hash relative to a list of unmanaged import functions and / or corresponding libraries. Various hash functions can be used in conjunction with the hash determination. Examples of hash functions include SHA-256, MD5, SHA-1, etc. Various other hash functions can be implemented. The system uses the hash function to determine the unmanaged imphash corresponding to a .NET file.

[0172] In some embodiments, the system transforms or formats the list according to a predefined format. For example, a list of library <-> function (<-> mappingflags) name pairs is transformed into a comma-separated string. Instead of calculating the hash on the table below, this is done according to such an example.

[0173] kernel32.createprocess

[0174] kernel32.getthreadcontext

[0175] kernel32.wow64getthreadcontext

[0176] kernel32.setthreadcontext

[0177] kernel32.wow64setthreadcontext

[0178] kernel32.readprocessmemory

[0179] The system creates a comma-separated string using a single list item (library <-> function pair): "kernel32.createprocess,kernel32.getthreadcontext,kernel32.wow64getthreadcontext,kernel32.setthreadcontext,kernel32.wow64setthreadcontext,kernel32.readprocessmemory,..." In response to transforming / creating a comma-separated string using a single list item, the system determines (e.g., calculates a hash) relative to such a string.

[0180] At 718, a hash is provided. In some embodiments, the system provides a hash to another system or module, such as to determine whether a file is malicious. For example, a hash is provided in response to a call to procedure 700.

[0181] At 720, a determination is made regarding whether process 700 is complete. In some embodiments, in response to determining that no further hashing is required (e.g., further prediction of the file is not needed), the administrator instructs to pause or stop process 700, etc., to determine that process 700 is complete. In response to determining that process 700 is complete, process 700 ends. In response to determining that process 700 is not complete, process 700 returns to 702.

[0182] Figure 7B This is a flowchart of a method for detecting malicious files according to various embodiments.

[0183] In some embodiments, process 700 is at least partially constituted by Figure 1 System 100 and / or Figure 2 System 200 is implemented. In some implementations, process 600 may be implemented by one or more servers, such as in conjunction with providing services to a network (e.g., security entities and / or network endpoints, such as client devices). In some implementations, process 700 may be implemented by a security entity (e.g., a firewall), such as in conjunction with implementing security policies relative to files transferred across or within / outside the network. In some implementations, process 700 may be implemented by a client device (e.g., a laptop computer, smartphone, personal computer, etc.), such as in conjunction with executing or opening files (e.g., email attachments).

[0184] According to various embodiments, with Figure 7A Compared to 700, 750 is more stringent because it includes additional information about unmanaged functions (such as information related to how the author defined unmanaged functions in the code). As an example, the final hash of 750 may also be more stringent and less prone to false positives. However, the downside is that the final hash of 750 may hit fewer malware samples than the final hash of 700.

[0185] At 752, obtain the library name and / or function name. In some embodiments, obtain the function name and corresponding library name obtained using the .NET header, such as in combination with a list of generated unmanaged import functions.

[0186] At 754, it is determined whether the library name has a file extension. In response to the determination at 754 that the library name has a file extension, process 750 continues to 756, where the file extension is removed. In response to the determination at 754 that the library name does not have a file extension, process 700 continues to 760.

[0187] At 758, the function name and library name are formatted. In some embodiments, the system formats the library name-function name pair according to a predetermined format or syntax. As another example, the system formats the function name and library name to convert any uppercase letters to lowercase.

[0188] At position 760, the MappingFlags value is obtained. In some embodiments, the system obtains the MappingFlags value from the ImplMap table. For example, the system obtains the MappingFlags value from the row in the ImplMap table corresponding to the function. The MappingFlags value includes the P / Invoke attribute.

[0189] At 762, the MappingFlags value, function name, and library name are combined. In some embodiments, the system constructs a string corresponding to the combination of the MappingFlags value, the function name of the imported function (e.g., an unmanaged imported function), and the library name corresponding to the imported function. For example, the system constructs the string by appending the MappingFlags value and the library name to the function name, and a predefined separator (e.g., ".") is included between the library name and the function name. According to various embodiments, the system determines the string based on a format: <libraryname> . <functionname> . <mappingflags>The system then adds the string to the list of unmanaged import functions (e.g., to determine the list of unmanaged Imphash).

[0190] According to various embodiments, various formats or syntaxes can be implemented by combining system-combined function names and library names. Examples of predefined formats include (i) lowercase alphanumeric strings, (ii) removal of file extensions, (iii) removal of library extensions (e.g., removing .dll from the corresponding library name), (iv) appending unmanaged function names and corresponding libraries, and (v) the use of predefined separators between unmanaged function names and corresponding libraries, and (vi) combining library names and function names to determine the order of strings (e.g., ...). <libraryname> . <functionname> . <mappingflags> 、 <functionname> . <libraryname> . <mappingflags>(etc.). In some embodiments, the system appends the function name (e.g., an unmanaged function name) to the corresponding library and separates the function name from the corresponding library by a dot or period (e.g., "."). In response to determining the string by appending the function name and library (e.g., using a predefined separator, such as "."), the system adds such entries to a list of the collection of unmanaged imported functions and / or corresponding libraries. Various other predefined formats / syntax can be implemented.

[0191] At 764, the combination of function name and library name is added to the list of unmanaged imported functions.

[0192] At point 766, a determination is made as to whether more functions should be added to the list. For example, the system determines whether more unmanaged import functions should be added to the file's list of unmanaged import functions. In response to determining at 766 that no further functions should be added to the list, process 750 proceeds to 752. In response to determining that additional functions should be added to the list, process 750 returns to 752.

[0193] At position 768, the hash is calculated relative to a list. According to various embodiments, the system determines the hash relative to a list of unmanaged import functions and / or corresponding libraries. Various hash functions can be used in conjunction with the hash determination. Examples of hash functions include SHA-256, MD5, SHA-1, etc. Various other hash functions can be implemented. The system uses the hash function to determine the unmanaged imphash corresponding to a .NET file.

[0194] In some embodiments, the system transforms or formats the list according to a predefined format. For example, a list of library <-> function (<-> mappingflags) name pairs is transformed into a comma-separated string. Instead of calculating the hash on the following list, this is done according to such an example.

[0195] kernel32.createprocess

[0196] kernel32.getthreadcontext

[0197] kernel32.wow64getthreadcontext

[0198] kernel32.setthreadcontext

[0199] kernel32.wow64setthreadcontext

[0200] kernel32.readprocessmemory

[0201] The system creates a comma-separated string using a single list item (library <-> function pair): "kernel32.createprocess,kernel32.getthreadcontext,kernel32.wow64getthreadcontext,kernel32.setthreadcontext,kernel32.wow64setthreadcontext,kernel32.readprocessmemory,..." In response to transforming / creating a comma-separated string using a single list item, the system determines (e.g., calculates a hash) relative to such a string.

[0202] At 770, a hash is provided. In some embodiments, the system provides a hash to another system or module, such as to determine whether a file is malicious. For example, a hash is provided in response to a call to procedure 750.

[0203] At 772, a determination is made regarding whether process 750 is complete. In some embodiments, in response to determining that no further hashing is required (e.g., further prediction of the file is not needed), the administrator instructs to pause or stop process 750, etc., to determine that process 750 is complete. In response to determining that process 750 is complete, process 750 ends. In response to determining that process 750 is not complete, process 750 returns to 752.

[0204] Figure 8 This is a flowchart of a method for detecting malicious files according to various embodiments. In some embodiments, process 800 is at least partially constituted by Figure 1 System 100 and / or Figure 2 The system 200 is implemented. In some implementations, process 800 may be implemented by one or more servers, such as in conjunction with providing services to a network (e.g., security entities and / or network endpoints, such as client devices). In some implementations, process 800 may be implemented by a security entity (e.g., a firewall), such as in conjunction with implementing security policies relative to files transferred across or within / outside the network. In some implementations, process 800 may be implemented by a client device (e.g., a laptop computer, smartphone, personal computer, etc.), such as in conjunction with executing or opening files (e.g., email attachments).

[0205] At 810, an indication that the received sample is malicious is received. In some embodiments, the system receives an indication that the sample is malicious, along with the sample or a hash, signature, or other unique identifier associated with the sample. For example, the system may receive an indication that the sample is malicious from a service, such as a security or malware service. The system may receive an indication that the sample is malicious from one or more servers.

[0206] According to various embodiments, receiving a sample is an indication of maliciousness in conjunction with an update to a previously identified set of malicious files. For example, the system receives an indication that a sample is malicious as an update to a blacklist of malicious files.

[0207] At 820, the system stores the association between the sample and an indication that the sample is malicious. In response to receiving an indication that the sample is malicious, the system stores the indication that the sample is malicious in association with the sample or an identifier corresponding to the sample to facilitate a subsequent lookup (e.g., a local lookup) to determine whether a received file is malicious. In some embodiments, the identifier corresponding to the sample stored in association with the indication that the sample is malicious includes a hash of the file (or a portion of the file), a signature of the file (or a portion of the file), or another unique identifier associated with the file. In some embodiments, storing the sample in association with the indication that the sample is malicious includes storing an unmanaged imphalite of a .NET file in association with the indication that the sample is malicious.

[0208] At 830, services are received. The system can acquire services, such as those routed within / across the network, those mediating traffic entering or leaving the network (e.g., through a firewall), or those monitoring email or instant messaging services.

[0209] At point 840, the system determines whether the service includes a malicious file. In some embodiments, the system obtains a file from a received service. For example, the system identifies the file as an email attachment, or as being exchanged between two client devices via an instant messaging program or another file exchange program. In response to obtaining the file from the service, the system determines whether the file corresponds to a file included in a previously identified set of malicious files (such as a malicious file blacklist). In response to determining that the file is included in the set of files on the malicious file blacklist, the system determines that the file is malicious (e.g., the system may further determine that the service includes a malicious file).

[0210] In some embodiments, the system determines whether a file corresponds to a file included in a previously identified set of benign files (such as a benign file whitelist). In response to determining that a file is included in a set of files on the benign file whitelist, the system determines that the file is not malicious (e.g., the system may further determine that the business includes malicious files).

[0211] According to various embodiments, in response to determining that a file is not included in a previously identified set of malicious files (e.g., a malicious file blacklist) or a previously identified set of benign files (e.g., a benign file whitelist), the system considers the file to be non-malicious (e.g., benign).

[0212] According to various embodiments, in response to determining that a file is not included in a previously identified set of malicious files (e.g., a malicious file blacklist) or a previously identified set of benign files (e.g., a benign file whitelist), the system queries a malicious file detector to determine whether the file is malicious. For example, the system may isolate the file until it receives a response from the malicious file detector regarding whether the file is malicious. The malicious file detector may perform the assessment of whether a file is malicious, such as simultaneously with the system's processing of business operations (e.g., in real-time with queries from the system). The malicious file detector may correspond to... Figure 1 System 100 and / or Figure 2 The system has a malicious file detector 170.

[0213] In some embodiments, by calculating a hash or determining a signature or other unique identifier associated with the file, and performing a lookup in a previously identified set of malicious files or a previously identified set of benign files that matches the hash, signature, or other unique identifier, the system determines whether a file is included in a previously identified set of malicious files or a previously identified set of benign files. Various hashing techniques can be implemented. According to various embodiments, determining whether a file is included in a previously identified set of malicious files or a previously identified set of benign files includes determining an unmanaged imphash corresponding to the file and determining whether the unmanaged imphash is included in a historical dataset (e.g., a dataset that includes previously determined results of malicious activity).

[0214] In response to the determination at 840 that the business does not contain malicious files, process 800 continues to 850, where the file is processed as a non-malicious business / information.

[0215] In response to determining at 840 that the transaction does not include malicious files, process 800 continues to 860, where the file is processed as malicious transaction / information. The system may process malicious transactions / information based at least in part on one or more policies (such as one or more security entities).

[0216] According to various embodiments, the handling of malicious file transactions / information may include performing proactive measures. These proactive measures can be performed based on (e.g., at least in part on) one or more security policies. As an example, one or more security policies may be preset by a network administrator, a customer (e.g., an organization / company) providing a malicious file detection service, etc. Examples of proactive measures that can be performed include: orphaning files (e.g., isolating files), deleting files, alerting the user to the detection of a malicious file, providing a prompt to the user when a device attempts to open or execute a file, blocking file transfer, updating a malicious file blacklist (e.g., a mapping from a file hash to an indication that the file is malicious), etc.

[0217] At 870, a determination is made as to whether process 800 is complete. In some embodiments, in response to determining that no further sample analysis is needed (e.g., further prediction of the document is not required), the administrator instructs to pause or stop process 800, etc., to determine that process 800 is complete. In response to determining that process 800 is complete, process 800 ends. In response to determining that process 800 is not complete, process 800 returns to 810.

[0218] Figure 9 This is a flowchart of a method for detecting malicious files according to various embodiments. In some embodiments, process 900 is at least partially performed by... Figure 1 System 100 and / or Figure 2 The system 200 is implemented. In some implementations, process 900 may be implemented by a security entity (such as a firewall) and / or an anti-malware application running on the client system, such as in conjunction with implementing security policies relative to files transferred across or within / outside the network. In some implementations, process 900 may be implemented by a client device (such as a laptop computer, smartphone, personal computer, etc.), such as in conjunction with executing or opening files (such as email attachments).

[0219] At 910, a file is obtained from a service. The system may obtain services, such as those routing services within / across the network, those mediating in / out of the network (e.g., firewalls), or those monitoring email or instant messaging services. In some embodiments, the system obtains a file from a received service. For example, the system may identify the file as an email attachment, or as an exchange between two client devices via an instant messaging program or other file exchange program.

[0220] At position 920, a signature corresponding to the file is determined. In some embodiments, the system calculates a hash or determines a signature or other unique identifier associated with the file. Various hashing techniques can be implemented. For example, a hashing technique could be determining (e.g., calculating) the MD5 hash of the file. In some embodiments, determining a signature corresponding to a file includes calculating an unmanaged imphalite of the .NET file.

[0221] At point 930, the system queries a dataset of signatures from malicious samples to determine if the signature corresponding to a file matches a signature from the malicious sample. In some embodiments, the system performs a lookup in the dataset of signatures from malicious samples for files that match a hash, signature, or other unique identifier. The dataset of signatures from malicious samples may be stored locally on the system or remotely on a storage system accessible to the system.

[0222] According to various embodiments, determining whether a file is included in a previously identified set of malicious files or a previously identified set of benign files includes determining the unmanaged Imphash corresponding to the file and determining whether the unmanaged Imphash is included in a historical dataset (e.g., a dataset that includes previously identified results of maliciousness).

[0223] At 940, the determination of whether a file is malicious is made at least in part based on whether the file's signature matches the signature of a malicious sample. In some embodiments, the system determines whether the dataset of malicious signatures includes records that match the signature of a file obtained from the business. In response to determining that the historical dataset includes an indication that a file corresponding to an unmanaged Imphash is malicious (e.g., the unmanaged Imphash is included in a field blacklist), the system considers the file obtained from the business to be malicious at 910.

[0224] At 950, the file is processed based on whether it is malicious. In some embodiments, in response to determining that a file is malicious, the system applies one or more security policies to the file. In some embodiments, in response to determining that a file is not malicious, the system processes the file as benign (e.g., processes the file as normal business).

[0225] At 960, a determination is made as to whether process 900 is complete. In some embodiments, in response to determining that no further sample analysis is needed (e.g., further prediction of the document is not required), the administrator instructs to pause or stop process 900, etc., to determine that process 900 is complete. In response to determining that process 900 is complete, process 900 ends. In response to determining that process 900 is not complete, process 900 returns to 910.

[0226] Figure 10 This is a flowchart of a method for detecting malicious files according to various embodiments. In some embodiments, process 1000 is at least partially constituted by Figure 1 System 100 and / or Figure 2 The system 200 is implemented. In some implementations, process 1000 may be implemented by a security entity (e.g., a firewall) and / or an anti-malware application running on the client system, such as in conjunction with implementing security policies relative to files transferred across or within / outside the network. In some implementations, process 1000 may be implemented by a client device (e.g., a laptop computer, smartphone, personal computer, etc.), such as in conjunction with executing or opening files (e.g., email attachments).

[0227] At 1010, services are received. The system can acquire services, such as those routed within / across the network, those mediating traffic entering or leaving the network (e.g., through a firewall), or those monitoring email or instant messaging services.

[0228] At 1020, a file is obtained from a service. In some embodiments, the system obtains a file from a received service. For example, the system may identify the file as an email attachment, or as an exchange between two client devices via an instant messaging program or other file exchange program.

[0229] At 1030, the imported API function name is obtained using the file's .NET header. In some embodiments, 1030 corresponds to or is similar to... Figure 4 The process is 400 out of 420.

[0230] At 1040, the unmanaged function is determined. In some embodiments, the system determines the set of unmanaged functions based on the imported API function names obtained using the .NET header of the file. For example, the system determines which of the imported API functions corresponds to the unmanaged function. As an example, in conjunction with determining the unmanaged function, at least a portion of procedure 600 may be invoked.

[0231] At point 1050, a determination is made regarding whether the file is malicious. In some embodiments, the system determines whether a file is malicious based at least in part on unmanaged managed functions (e.g., a set of unmanaged functions imported into the file via .NET headers). In some embodiments, 1050 corresponds to or is similar to... Figure 4 The process 440 of 400. In some embodiments, Figure 5 The process is executed by combining 500 and 1050.

[0232] In response to determining at 1050 that the file is malicious, process 1000 continues to 1060, where one or more security policies are applied relative to the file. In some embodiments, 1060 corresponds to or is similar to Figure 8 The process is 860 of 800. After that, process 1000 continues to 1070.

[0233] In response to determining at 1050 that the file is not malicious, process 1000 continues to 1070, where the file is processed as non-malicious. In some embodiments, 1070 corresponds to or is similar to... Figure 8 The process is 800 out of 850.

[0234] At 1080, a determination is made regarding whether process 1000 is complete. In some embodiments, in response to determining that no further sample analysis is needed (e.g., no further prediction of documents is required), no further business analysis is performed, the administrator instructs to pause or stop process 1000, etc., to determine whether process 1000 is complete. In response to determining that process 1000 is complete, process 1000 ends. In response to determining that process 1000 is not complete, process 1000 returns to 1010.

[0235] Figure 11 This is a flowchart of a method for detecting malicious files according to various embodiments. In some embodiments, process 1100 is at least partially performed by... Figure 1 System 100 and / or Figure 2 The system 200 is implemented. In some implementations, process 1100 may be implemented by a security entity (e.g., a firewall) and / or an anti-malware application running on the client system, such as in conjunction with implementing security policies relative to files transferred across or within / outside the network. In some implementations, process 1100 may be implemented by a client device (e.g., a laptop computer, smartphone, personal computer, etc.), such as in conjunction with executing or opening files (e.g., email attachments).

[0236] At 1110, a service is received. In some embodiments, 1110 corresponds to or is similar to... Figure 10 The process of 1000 to 1010.

[0237] At 1120, a document is obtained from the business. In some embodiments, 1120 corresponds to or is similar to... Figure 10 The process is 1000 to 1020.

[0238] At 1130, the imported API function name is obtained using the file's .NET header. In some embodiments, 1130 corresponds to or is similar to... Figure 10 The process is 1000 out of 1030.

[0239] At 1140, a hash of the list of unmanaged imported API function names is determined. In some embodiments, the system determines the unmanaged functions at least in part based on the imported API function names (e.g., via the .NET header of a file). In some embodiments, the system determines a set of unmanaged functions based on the imported API function names obtained using the .NET header of a file, determines a list of unmanaged functions, and determines the hash at least in part based on the list. As an example, the list includes unmanaged function names corresponding to the set of unmanaged functions.

[0240] At 1150, a hash-to-file mapping is queried. In some embodiments, the system queries the hash-to-file mapping based at least in part on the hash of a list of unmanaged import API function names. For example, the system performs a query relative to the hash-to-file mapping to determine whether the mapping includes the hash of a list of unmanaged import API function names (e.g., determining whether the mapping includes a record corresponding to the determined / computed hash). In some embodiments, 1150 corresponds to or is similar to... Figure 8 Process 800 of 840 and / or Figure 9 The process is 900 out of 930.

[0241] At point 1160, a determination is made regarding whether the file is malicious. In some embodiments, 1160 corresponds to or is similar to... Figure 10 The process is 1000 to 1050.

[0242] In response to determining at 1160 that the file is malicious, process 1100 continues to 1170, where one or more security policies are applied relative to the file. In some embodiments, 1170 corresponds to or is similar to Figure 8 The process is 860 from 800. After that, process 1100 continues to 1190.

[0243] In response to determining at 1160 that the file is not malicious, process 1100 continues to 1180, where the file is processed as non-malicious. In some embodiments, 1180 corresponds to or is similar to Figure 8 The process is 800 out of 850.

[0244] At 1190, a determination is made regarding whether process 1100 is complete. In some embodiments, in response to determining that no further sample analysis is needed (e.g., no further prediction of documents is required), no further business analysis is performed, the administrator instructs to pause or stop process 1100, etc., to determine whether process 1000 is complete. In response to determining that process 1100 is complete, process 1100 ends. In response to determining that process 1100 is not complete, process 1100 returns to 1110.

[0245] Various examples of embodiments described herein are illustrated in conjunction with flowcharts. Although examples may include certain steps performed in a particular order, according to various embodiments, various steps may be performed in various orders and / or various steps may be combined into a single step or in parallel.

[0246] Although the foregoing embodiments have been described in detail for clarity of understanding, the invention is not limited to the details provided. Many alternative ways of implementing the invention are possible. The disclosed embodiments are illustrative and not restrictive.< / mappingflags> < / libraryname> < / functionname> < / mappingflags> < / functionname> < / libraryname> < / mappingflags> < / functionname> < / libraryname> < / libraryname> < / functionname> < / functionname> < / libraryname> < / functionname> < / libraryname> < / crtimplementationdetails> < / crtimplementationdetails> < / crtimplementationdetails> < / crtimplementationdetails> < / crtimplementationdetails> < / crtimplementationdetails> < / functionname> < / libraryname> < / functionname> < / libraryname>

Claims

1. A system comprising: One or more processors are configured as follows: Receive samples including .NET files; The imported API function names are obtained at least in part from the .NET header of the .NET file; Determine a hash of a list of unmanaged import API function names, wherein the list of unmanaged import API function names includes a pair for each unmanaged import API function in the list, the pair including (a) an indication of the unmanaged API function name and (b) an indication of the corresponding library name. as well as Whether a sample is malware is determined at least in part based on a hash of a list of unmanaged import API function names; as well as The memory is coupled to the one or more processors and configured to provide instructions to the one or more processors.

2. The system according to claim 1, wherein obtaining the imported API function name includes: Parse the .NET header of a .NET file; as well as Extract the imported API function names from the parsed .NET header.

3. The system of claim 2, wherein the imported API function name is extracted from the parsed .NET header at least in part based on an index table.

4. The system of claim 3, wherein the index table includes an ImplMap table that indicates a set of unmanaged methods imported in conjunction with the execution of a .NET file.

5. The system of claim 2, wherein the imported API function name is extracted from the parsed .NET header based at least in part on the MethodDef table.

6. The system of claim 1, wherein the hash for determining the list of unmanaged imported API function names comprises: The set of unmanaged imported API functions is determined based at least in part on the imported API function names obtained from the .NET header; as well as The hash of the list of unmanaged imported API function names is generated, at least in part, based on a predefined hash function.

7. The system according to claim 6, wherein the predetermined hash function includes at least one of the SHA-256 hash algorithm, the MD5 hash algorithm, and the SHA-1 hash algorithm.

8. The system of claim 1, wherein the one or more processors are further configured to: Sending samples to a secure entity is a malicious instruction.

9. The system of claim 8, wherein sending an indication that the sample is malicious to a security entity comprises: The blacklist of files deemed malicious is updated to include identifiers corresponding to the samples.

10. The system of claim 1, wherein the security entity corresponds to a firewall.

11. The system of claim 1, wherein determining whether a sample is malware is executed in a sandbox environment based at least in part on a hash of a list of unmanaged import API function names.

12. The system of claim 1, wherein the determination of whether a sample is malware is based at least in part on a hash of a list of unmanaged import API function names is executed at a secure entity.

13. The system of claim 1, wherein the import API function name is obtained at least in part based on the value corresponding to the ImportName field.

14. The system of claim 13, wherein in response to determining that the value corresponding to the ImportName field is not equal to 0, obtaining the import API function name includes determining a set of one or more library names based at least in part on the #Strings stream of the .NET file.

15. The system of claim 13, wherein obtaining the import API function name in response to determining that the value corresponding to the ImportName field is equal to 0 includes: Get the value from the row of the MethodDef table; The function name is obtained from the #Strings stream, at least in part, based on the values ​​of the rows from the MethodDef table; Determine whether the PE header of the .NET file includes an import table; as well as In response to the fact that the PE header of a .NET file does not include an import table, the library name corresponding to the function name is determined at least in part based on the import table.

16. The system of claim 15, wherein determining the library name corresponding to the function name based at least in part on the import table includes resolving the import table to obtain the library name from the corresponding function name.

17. The system of claim 1, wherein the one or more processors are further configured to: Ensure that the library name corresponding to the unmanaged function does not have a file extension; The string is determined at least in part based on the library name and unmanaged functions; and Add the string to the list of unmanaged imported API function names.

18. The system of claim 17, wherein the string is determined by ensuring that the library name and the name of the unmanaged function do not contain uppercase letters and by appending the name of the unmanaged function to the library name using a predefined separator included between the name of the unmanaged function and the library name.

19. The system of claim 1, wherein hashing the list of unmanaged imported API function names comprises: Based at least in part on the indication of the unmanaged API function name and the indication of the corresponding library name of the multiple unmanaged API functions, a string representing a list of unmanaged imported API function names is generated; and The hash is calculated based on a list of strings representing the names of unmanaged imported API functions.

20. A method comprising: Receive samples including .NET files; The imported API function names are obtained at least in part from the .NET header of the .NET file; Determine a hash of a list of unmanaged import API function names, wherein the list of unmanaged import API function names includes a pair for each unmanaged import API function in the list, the pair including (a) an indication of the unmanaged API function name and (b) an indication of the corresponding library name. as well as Whether a sample is malware is determined at least in part based on a hash of a list of unmanaged import API function names.

21. A computer program product embodied in a non-transitory computer-readable medium and comprising computer instructions for: Receive samples including .NET files; The imported API function name is obtained at least in part from the .NET header of the .NET file; Determine a hash of a list of unmanaged import API function names, wherein the list of unmanaged import API function names includes a pair for each unmanaged import API function in the list, the pair including (a) an indication of the unmanaged API function name and (b) an indication of the corresponding library name. as well as Whether a sample is malware is determined at least in part based on a hash of a list of unmanaged import API function names.