Deep reinforcement learning techniques for detecting malware

By combining deep reinforcement learning neural networks with event classifiers and file classifiers, and dynamically adjusting file execution decisions, the low detection efficiency caused by fixed event sequences in existing technologies is solved, achieving more efficient malware detection.

CN111954881BActive Publication Date: 2025-12-12MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980024816.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-04-10
Filing Date
2019-03-26
Publication Date
2025-12-12
Estimated Expiration
2039-03-26

AI Technical Summary

Technical Problem

Existing malware detection methods rely on fixed-length event sequences, making it difficult to flexibly decide whether to pause or continue file execution, resulting in low detection efficiency.

Method used

By employing a deep reinforcement learning (DRL) neural network combined with an event classifier and a file classifier, the pause point for file execution is dynamically adjusted, and decisions are made based on the unique characteristics of each file.

Benefits of technology

It significantly improved the accuracy of classifying unknown files, reduced the false positive rate, and increased the true positive detection rate by 30.6%, achieving more flexible and efficient malware detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111954881B_ABST
    Figure CN111954881B_ABST
Patent Text Reader

Abstract

Techniques for detecting malware based on a reinforcement learning model, for detecting whether a file is malicious or benign, and for determining the optimal time to pause execution of a file in such detection process. The reinforcement learning model is combined with an event classifier and a file classifier, learning is whether to pause execution or continue execution after enough state information has been observed, if more events are needed to make a high-confidence determination. The disclosed algorithm allows the system to decide when to stop on a per-file basis.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] Despite decades of research in computer security and tools to eliminate security threats, users and organizations continue to rely on commercial malware products that use some primary means to attempt to detect malware. First, static analysis based on malware "signatures" is used to search for malicious code sequences in files or processes. Next, dynamic analysis is used to emulate execution of a file, often in an isolated space. This emulation can not involve a full virtual machine ("VM"). Instead, an emulator can mimic the responses of a typical operating system. If the system can detect malicious behavior while emulating the file, the system can prevent execution on the native operating system and identify the file as a malicious file. As a result, infection of the computer can be avoided. If the system cannot detect malicious behavior during emulation, the file can be installed and / or executed on the computer. After installation, the malware system typically continues to monitor the dynamic behavior of the file whenever it is executed on the computer. If the malware system detects the malware file on the computer, it typically takes one or more actions to protect the computer from the file. SUMMARY

[0002] The summary provided in this section summarizes one or more portions or the entirety of example embodiments of the technology described herein in order to provide the reader with a basic high-level understanding of the technology. The summary is not an extensive description of the technology, and it can not identify key elements or aspects of the technology or explain the scope of the technology. Its sole purpose is to present various aspects of the technology in a simplified form as a preamble to the detailed description provided below. In general, the technology should not be limited to any particular embodiment(s) or example(s) provided herein or combination(s) thereof.

[0003] The computer-related technology disclosed herein is primarily directed to a novel invention that detects the optimal time to pause file execution in order to determine whether a file is malicious or benign based on deep reinforcement learning ("DRL"). The resulting DRL neural network ("NN") learns in conjunction with an event classifier and a file classifier: whether to pause emulation after enough state information has been observed or to continue execution if more events are needed to make a high-confidence determination. Unlike previously proposed solutions, the DRL algorithm disclosed herein allows the system to decide when to stop execution on a per-file basis. By doing so, the invention is a step towards using artificial intelligence in a critical area of cybersecurity.

[0004] For example, the results of the analysis of a set of malware and benign files by the deep reinforcement learning system indicate a significant improvement in the overall classification of unknown files. The proposed deep reinforcement learning system improves the true positive detection rate by 30.6% with a false positive rate of 1.0%.

[0005] One of the weaknesses of these early systems is that they use fixed length sequences of events to make the decision to stop or pause file execution. In the present invention, a new deep reinforcement learning method is used to decide better execution pause points with good confidence, which helps the anti-malware system to learn to be more flexible in the required sequence length of events.

[0006] Reinforcement learning is a special type of machine learning method that uses the concept of stochastic optimization. It aims to solve an optimization problem such that an agent will take actions in a stochastic environment to maximize some notion of cumulative reward. In one example of the present invention, the environment is defined as a malware file to be screened, the agent is defined as an anti-malware system, and the reward is defined in a way that can train the agent to be as smart as possible when choosing between two actions: continue executing the file (because the file is determined to be benign) or pause file execution (because the file is determined to be malicious) by maximizing its expected reward. BRIEF DESCRIPTION OF DRAWINGS

[0007] The detailed description provided below in connection with the appended drawings is intended as a description of the present technology and is not intended to represent that the present technology can not be practiced with the claimed technology. The detailed description set forth below in connection with the appended drawings is intended as a description of the present technology and is not intended to represent that the present technology can not be practiced with the claimed technology.

[0008] Figure 1 is a block diagram illustrating an example computing environment 100 in which the technology described herein can be implemented.

[0009] Figure 2 is a block diagram illustrating an example malware detection system 200 based on the disclosed technology.

[0010] Figure 3 is a diagram illustrating various data structures used in detecting malware.

[0011] Figure 4 is a block diagram illustrating an example method 400 for determining whether a file being executed is malicious or benign.

[0012] Figure 5 is a block diagram illustrating an example execution control module 510.

[0013] Figure 6 is a block diagram illustrating an example method 600 for determining event scores and making execution decisions to continue or pause file execution.

[0014] Figure 7 is a block diagram illustrating an example inference model 720.

[0015] Figure 8 is a block diagram illustrating an example method 800 for determining an improvement score, which indicates a likelihood that a file being executed is malicious or benign.

[0016] Figure 9 is a block diagram illustrating an example classifier 920, which can be used to implement the event classifier 512 and / or the file classifier 722.

[0017] Like-numbered elements in different figures are used to represent similar or identical elements or steps. DETAILED DESCRIPTION

[0018] The detailed description provided in this section describes one or more embodiments of the disclosed technology in connection with the appended figures, but is not intended to describe all possible embodiments of the technology. This detailed description sets forth various examples of at least some of the systems and / or methods of the disclosed technology. However, similar or equivalent technologies, systems, and / or methods can also be implemented according to other examples.

[0019] Computing Environment

[0020] Although the examples provided herein are described and illustrated as being implemented in a computing environment, the described environment is provided as an example only and is not limiting. As those skilled in the art will appreciate, the disclosed examples are suitable for use in a variety of different computing environments.

[0021] Figure 1 is a block diagram illustrating an example computing environment 100 in which the technology described herein can be implemented. Any of a variety of general purpose or special purpose computing devices can be used to implement a suitable computing environment. Examples of such devices include, but are not limited to, personal digital assistants (“PDAs”), personal computers (“PCs”), hand-held or laptop devices, microprocessor-based systems, multi-processor systems, system-on-a-chip (“SOCs”), servers, internet appliances, workstations, consumer electronics devices, cellular telephones, set-top boxes, and the like. In all cases, such systems are strictly limited to articles of manufacture such as those that fall within the purview of 35 U.S.C. § 101.

[0022] The computing environment 100 generally includes at least one computing device 101 coupled with various components, such as peripheral devices, including a display 102, input / output devices 103, and the like. These can include components such as input / output devices 103 that can operate via one or more input / output ("I / O") interfaces 112, input / output devices 103 such as voice recognition technology, touch pads, buttons, keyboards, and / or pointing devices such as mice or trackballs. The components of the computing device 101 can include one or more processing units 107 (including central processing units ("CPUs"), graphics processing units ("GPUs"), microprocessors ("μΡs"), and the like), system memory 109, and a system bus 108 that generally couples the various components. The processing unit(s) 107 generally process or execute various computer-executable instructions and, based on those instructions, control the operation of the computing device 101. This can include the computing device 101 communicating with other electronic and / or computing devices, systems, or environments (not shown) via various communication technologies, such as network connections 114, and the like. The system bus 108 represents any number of bus structures including a memory bus or memory controller, a peripheral bus, a serial bus, an accelerated graphics port, a processor or local bus using any of a variety of bus architectures, and the like.

[0023] The system memory 109 can include computer-readable media in the form of volatile memory, such as random-access memory ("RAM"), and / or non-volatile memory, such as read-only memory ("ROM") or flash memory ("FLASH"). A basic input / output system ("BIOS") can be stored in non-volatile memory, among other things. The system memory 109 generally stores data, computer-executable instructions, and / or program modules, including computer-executable instructions, that are immediately accessible and / or currently being operated on by the processing unit(s) 107. The term "system memory" as used herein strictly refers to physical articles of manufacture or similar.

[0024] Storage devices 104 and mass storage 110 can be coupled to the computing device 101 or incorporated into the computing device 101 via a coupling to the system bus. Such storage devices 104 and mass storage 110 can include non-volatile RAM, a disk drive, a floppy disk drive with a removable storage disk (e.g., a "floppy" or "soft" disk), and / or an optical disk drive with a removable storage disk (e.g., a CD ROM, DVD ROM 106). Alternatively, mass storage devices can include non-removable storage media such as hard disk or other like storage devices. Other mass storage devices can include storage tanks, storage cartridges, tape storage devices, etc. As used herein, the term "mass storage device" refers strictly to one or more physical articles of manufacture or the like.

[0025] Any number of computer programs, files, data structures, etc. can be stored in the mass storage 110, other storage devices 104, 105, 106, and system memory 109 (typically subject to available space), including, by way of example and not limitation, operating systems, application programs, data files, directory structures, computer-executable instructions, etc.

[0026] Output components or devices, such as a display 102, can be coupled to the computing device 101 via an interface, such as a display adapter 111. The display 102 as an output device can be a liquid crystal display ("LCD"). Other example output devices can include a printer, audio output, voice output, a cathode ray tube ("CRT") display, a haptic device or other sensory output mechanism, etc. The output devices can enable the computing device 101 to interact with an operator or other machines, systems, computing environments, etc. A user can interface with the computing environment 100 via any number of different input / output devices 103, such as a touchpad, buttons, a keyboard, a mouse, a joystick, a gamepad, a data port, etc. These and other I / O devices can be coupled to the processing unit(s) 107 via an I / O interface 112, which can be coupled to the system bus 108, and / or can be coupled through other interface and bus structures, such as a parallel port, a game port, a universal serial bus ("USB"), a Firewire, an infrared ("IR") port, etc.

[0027] Computing device 101 can operate in a networked environment via a communication connection to one or more remote computing devices through one or more cellular networks, wireless networks, local area networks (“LANs”), wide area networks (“WANs”), storage area networks (“SANs”), the Internet, radio links, optical links, etc. Computing device 101 can be coupled to the network via network adapter 113, or alternatively via modem, digital subscriber line (“DSL”) link, integrated services digital network (“ISDN”) link, Internet link, wireless link, etc.

[0028] Communication connections, such as network connections, typically provide coupling to communication media, such as networks. Communication media typically use modulated data signals, such as carrier waves, or other transmission mechanisms to provide computer-readable and computer-executable instructions, data structures, files, program modules, and other data. The term "modulated data signal" generally refers to a signal that sets or alters one or more characteristics of information in a manner that encodes it as a signal. By way of example and not limitation, communication media can include wired media such as wired networks or direct wired connections, and wireless media such as acoustic, radio frequency, infrared, or other wireless communication mechanisms.

[0029] A power source 190, such as a battery or power supply, typically provides some or all of the power to the computing environment 100. In the case where the computing environment 100 is a mobile device or portable device, the power source 190 may be a battery. Alternatively, in the case where the computing environment 100 is a desktop computer or server, the power source 190 may be a power supply designed to be connected to an AC power source, such as via a wall-mounted power outlet.

[0030] Some mobile devices may only include a combination Figure 1 Some of the components described. For example, an electronic badge may consist of a coil, etc., together with a simple processing unit 107, etc., the coil being configured to act as a power source 190 when near a card reader device, etc. This coil may also be configured to act as an antenna coupled to the processing unit 107, etc., capable of radiating / receiving communication between the electronic badge and another device such as a card reader device. This communication may not involve networking, but may alternatively be general or dedicated communication via telemetry, point-to-point, RF, IR, audio, or other means. The electronic card may not include a display 102, input / output device 103, or a combination thereof. Figure 1 Many other components are described. As an example and not a limitation, combinations may be omitted. Figure 1 Other mobile devices that include many of the components described include electronic wristbands, electronic tags, implantable devices, etc.

[0031] Those skilled in the art will realize that storage devices used to provide computer-readable and computer-executable instructions and data can be distributed over a network. For example, a remote computer or storage device can store software applications and data that can be accessed by a local computer or storage device via the network. A local computer or storage device can download portions or all of the software applications and data from the remote computer or storage device via the network, and execute the computer-executable instructions. Alternatively, a local computer or storage device can download pieces of the software or data as needed, or distribute portions of the software or data across multiple computers or storage devices. Those skilled in the art will also realize that at least a portion of the software can be converted into one or more hardware equivalents, such as digital signal processors ("DSPs"), application specific integrated circuits ("ASICs"), field programmable gate arrays ("FPGAs"), etc. Those skilled in the art will recognize that the depicted example is not intended to limit the scope of this disclosure to one or more particular computer systems or computing devices.

[0032] Those skilled in the art will further realize that at least a portion of the software can be converted into one or more computer programs comprising one or more computer program objects. Those skilled in the art will further realize that at least a portion of the software can be converted into one or more computer programs comprising one or more computer program objects. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0033] As used herein, the term "firmware" generally includes and refers to executable instructions, code, data, applications, programs, program modules, etc. maintained in electronic devices such as ROMs. As used herein, the term "software" generally includes and refers to computer-executable instructions, code, data, applications, programs, program modules, firmware, etc. maintained in any form or type of computer-readable media configured to store computer-executable instructions, etc. in a manner that is accessible to computing devices.

[0034] As used herein and in the claims, the term "computer-readable medium" "computer-readable media" and the like are limited to one or more tangible, lawfully manufactured devices, machines, articles of manufacture, etc. that are not signals or carriers. Thus, as used herein, the term "computer-readable medium" is intended to and should be interpreted to cover lawfully manufactured devices, machines, articles of manufacture, etc.

[0035] As used herein and in the claims, the term "computing device" is limited to one or more lawfully manufactured devices, articles of manufacture, etc. that are not signals or carriers, such as computing device 101, which encompasses client devices, mobile devices, one or more servers, network services such as one or more computer-based Internet services or corporate network services, etc., and / or any combination thereof. Thus, as used herein, the term computing device is also intended to and should be interpreted to cover lawfully manufactured devices, machines, articles of manufacture, etc.

[0036] System Overview

[0037] Figure 2is a block diagram illustrating an example malware detection system ("MDS") 200 based on the disclosed technology. The MDS 200 generally includes three main components: an execution control module ("ECM") 210, an inference module ("IM") 220, and an event monitor ("EM") 230. Each of these components can be implemented in hardware or software or any combination thereof. Further, in other embodiments, these components can alternatively be combined in any combination. Generally, the MDS 200 accepts input 250 and produces output 260 and / or 270. Further, in some embodiments, the IM 220 is optional.

[0038] Generally, the input 250 is in the form of a file. The term "file" (including in the claim language) as used herein refers to any conventional executable file, as well as any process, program, code, firmware, function, software, script (including non-executable scripts), object, data (e.g., email attachments, web pages, digital images, videos, files, and any other form of container of digital information), and the like (all referred to herein for simplicity as "files"). Further, the term "execution" (including in the claim language) as used herein refers to conventional execution as well as emulation, interpretation (as in interpreting non-executable scripts), and the like (all referred to herein for simplicity as "execution"). Such "execution" can be performed in any one of system memory of a computer, a virtual machine, any isolation space, an emulator or simulator, an operating system, and the like.

[0039] In the context of monitoring by the EM 230, such files can be executed in a VM (or in other isolation space where execution of malware cannot harm the host), or directly on the host itself. The EM 230 will generally monitor the executing file (the executing file) for particular types of operations or events that it performs. For example, monitored events can include performance of file input / output ("I / O") operations, as well as calls by the executing file to registry application programming interfaces (APIs), networking APIs, thread / process creation / control APIs, interprocess communication APIs, and debugging APIs. This list is non-limiting and can also include any other events performed by the executing file that are determined to be relevant to detecting malware now and in the future. Generally, the term "monitored event" as used herein, and in particular in the claim language, refers to operations or events performed by the executing file that are relevant to detecting malware and generally include, but are not limited to, the example operations and events listed above.

[0040] Further, in one embodiment, each type of event monitored by EM 230 is indicated by an event identifier (“ID”), which uniquely identifies that event type from all other monitored event types. For example, an event of type “file open” can be indicated by event ID 54 (some unique identifier), while an event of type “file close” can be indicated by event ID 55 (some other unique identifier). Such unique event IDs can take any suitable form, whether numerical or otherwise. Typically, the output of EM 230 includes event IDs that identify the monitored events performed by the executing file. In one example, EM 230 provides the ID for each event e t in sequence to ECM 210, where e t indicates a monitored event at step t in the sequence of monitored events as they are performed by the executing file. In another example, EM 230 provides the sequence of event IDs to ECM 210 and IM 220 one step at a time.

[0041] In one example, an event ID can include parameters for the corresponding event e t . For example, if the event is a “file open” event, it can include a final name and path(s) parameter, among others. In this example, any or all such parameters can be referenced by or included in the event ID in any suitable form. Note that such events typically represent a regular operating system or other interface, each of which has zero or more various parameters. In many cases, such interfaces and parameters are documented by their provider.

[0042] ECM 210 typically includes two main components: an event classifier (“EC”) 212 and a reinforcement learning model 214, in one embodiment a deep reinforcement learning model (“DRL”). ECM 210 produces control decisions, such as h t , for continuing or pausing the execution of a file. For example, if MDS 200 detects a malicious sequence of events, these control decisions can be used to decide to pause the execution of a file. In another example, these decisions 260 are provided to IM 220.

[0043] IM 220 typically includes a file classifier (“FC”) 222 that employs a file classification model to help determine the improved likelihood that a file is malicious or benign. This likelihood y RL,t is typically provided as output 270, and is typically used to classify an executing file as malware (malicious) or benign. In conjunction with Figure 9 FC 222 and its operation are described in more detail.

[0044] Figure 3 is a diagram illustrating various data structures used in detecting malware. Sequence log 310 represents a sequence log of event IDs, which generally correspond to monitored events in the order in which they were executed by the execution file. Further, an event state s t 320 is typically generated for each new monitored event. Generally, event state s t corresponds to the event e t .

[0045] In one example of event state s t 320, each instance includes three fields: (1) an event ID field 322, which generally includes the event ID of the monitored event e t at step t in the sequence of monitored events in the order in which they were executed by the execution file; (2) an event position number or "step" field 324, which generally includes the position number or step t of the monitored event in the sequence of monitored events executed by the execution file since the beginning of the file execution; and (3) an event histogram field 326, which generally includes a histogram of event IDs.

[0046] In one embodiment, the event histogram takes the form of an ordered array representing all monitored event types. For example, given 100 different monitored event types, the first position in the ordered array represents event ID 1, the second position represents event ID 2, and so on, up to the one-hundredth position in the ordered array representing event ID 100. The histogram is updated at each step t in the sequence of monitored events in the order in which the monitored events were executed by the execution file. In one example, all positions in the histogram are initially set to zero. Then, as Figure 3If the monitored event at step 1 is of type 12 (e.g., event ID = 12), as illustrated in the middle, the 12th position in the ordered array (representing event ID 12) is incremented by 1, indicating that a first instance of monitored event type 12 occurred in the sequence of monitored events. Next, if the monitored event at step 2 is of type 45 (e.g., event ID = 45), the 45th position in the ordered array (representing event ID 45) is incremented by 1, indicating that a first instance of monitored event type 45 occurred in the sequence of monitored events. Finally, if the monitored event at step 19 is of type 23 (e.g., event ID = 23), the 23rd position in the ordered array (representing event ID 23) will be incremented by 1, indicating that an instance of monitored event type 23 occurred in the sequence of monitored events. Note that the example illustrated by histogram 326 indicates that, as of step 19, one instance of event ID 1 has been executed, 0 instances of event types 2 and 3 have been executed, 3 instances of event type 99 have been executed, and 1 instance of event ID 100 has been executed.

[0047] Generally, the sequence log 310 and event status s t 320. In one example, the sequence log and / or event status are created and updated in real-time as the monitored events are executed by the execution file. Additionally or alternatively, the sequence log can be created in real-time as the execution file executes the monitored events, can be saved once execution is complete, and the event status can be created using the saved sequence log at any time after the file is executed.

[0048] The exact format and / or structure of the sequence log 310 and / or event status 320 is not critical to the present invention; any form and / or structure suitable for a particular implementation is acceptable.

[0049] Figure 4 is a block diagram illustrating an example method 400 for determining whether an execution file is malicious or benign. In one embodiment, the method 400 is performed by the MDS 200, or the like. In one example, the method 400 is performed as follows.

[0050] Block 410 generally indicates detecting execution of a monitored event by the execution file. In one example, the monitored events are detected as described in connection with the EM 230. Further, as described above, these monitored events are among the types of operations and events monitored by the EM 230. Each monitored event e tAn event identifier ("ID") typically identifies the event, which uniquely identifies the event type from all other monitored event types. In one example, each event ID is provided in real-time as the monitored event is executed by the executing file. In another example, a sequence of event IDs is provided in the form of a sequence log 310 or the like. The sequence of event IDs typically corresponds to the monitored events in the order they are executed by the executing file. After providing the event ID of the corresponding monitored event at step t in the sequence, the method 400 typically continues at block 412.

[0051] Block 412 typically indicates that, based on the provided event ID for the latest event e t at step t in the sequence of monitored events in the order they are executed by the executing file, a corresponding event state s t is constructed. In one example, a particular event state is constructed as described in connection with Figure 3 Such an event state can be constructed in real-time as the file is executed. Alternatively, such an event state can be constructed from an event sequence log such as the sequence log 310. Typically, only a single instance of the event state is needed. This instance s t is typically updated at each step t to correspond to the latest event e t In this way, memory requirements are minimized. Once the event state s t is constructed or updated at step t in the sequence of monitored events, the method 400 typically continues at step 414.

[0052] Block 414 typically indicates that, in response to the provided event ID for the latest event e t at step t in the sequence of monitored events in the order they are executed by the executing file, a likelihood of a malicious event sequence is determined. This likelihood is typically determined by the EC 212 and is referred to herein as a score y t for the monitored event e e,t at step t in the sequence of monitored events. Once the event score y e,t is provided to the DRL model 214 at step t in the sequence of monitored events, the method 400 typically continues at block 416. Here, the term "event score" as used in particular in the claim language refers to a likelihood that the latest event history indicates a malicious event sequence, where the likelihood can optionally represent a probability. The EC 212 and its operation are described in more detail in connection with Figure 9

[0053] As described in more detail below, block 416 typically indicates that, in response to the event state s t and the event score y e,t ​, producing an execution decision to continue or pause execution of the file. In one example, this decision is provided by the MDS 200 as output 260. Once an execution control decision is produced at step t in the sequence of monitored events, the method 400 optionally continues at block 418.

[0054] Block 418 generally indicates that, in response to the execution control decision, a refined score is determined that indicates a likelihood that the file being executed is malicious or benign. This determination is typically performed by the IM 220 if such classification of the executing file is desired, otherwise this step can be excluded. Once a refined score is determined for step t, the method 400 generally restarts for step t+1.

[0055] Figure 5 is a block diagram illustrating an example execution control module (“ECM”) 510. The ECM 510 is generally the same as the ECM 210 and performs the same functions as the ECM 210, but illustrates additional details in connection with the ECM 510. In addition to two main components, an event classifier (“EC”) 512 (the same as the EC 212) and a deep reinforcement learning (“DRL”) model 514 (the same as the DRL model 214), the ECM 510 also includes a sliding event window (“SEW”) 516, an event state 518 (e.g., the event state s t 320), and an action state module (“ASM”) 520. The input e t 580 is generally in the form of a sequence of event IDs, e.g., one event e t at each step t in the sequence of monitored events in the order in which they are executed by the executing file, such as from the EM 230. And as with the ECM 210, the output 560 is generally the same as the execution control decision h t 260 to continue or pause execution of the file.

[0056] SEW 516 is typically a sliding window structure, in one example a first-in-first-out ("FIFO") queue, maintained by ECM 510 and typically retains or indicates E latest event IDs corresponding to a sequence of monitored events in the order they were executed. In one example, SEW 516 retains or indicates approximately 200 latest event IDs. In other examples, SEW 516 retains or indicates some other number of latest event IDs. In one embodiment, the number E can be determined based on hyperparameter tuning that yields the best performance of EC 512 in predicting malicious event activity. If E is too small (i.e., the latest event history in SEW 516 is too short), EC 512 can not have enough events to process to make a reliable decision. Likewise, if E is too large, then malicious activity can be too brief to be detected by EC 512. The term "latest event history" as used herein (including in the claim language) refers to a list of event IDs of the n latest monitored events in a sequence of monitored events in the order they were executed, where n is some integer. Here, SEW 516 lists the latest event history in the form of E latest event IDs in a sequence of monitored events in the order they were executed.

[0057] EC 512 is typically the same as EC 212 and performs the same functions as EC 212. In one example, EC 512 is a two-stage neural network structure where the first stage is a recurrent neural language model that generates a feature vector, which is then input to a second classifier stage. The recurrent neural language model can be a recurrent neural network ("RNN") model. Alternatively, the recurrent neural language model can be a long short-term memory ("LSTM") model, gated recurrent unit (GRU), or any suitable recurrent neural model. In another embodiment, the recurrent neural language model can be replaced with a sequential convolutional neural network (CNN). The classifier stage can be any supervised classifier such as a logistic regression-based classifier, support vector machine, neural network, or deep neural network.

[0058] EC 512 typically evaluates the latest event history indicated by SEW 516 to determine an event score y e,t , the event score indicating a likelihood that the latest event history indicates malicious activity, where e t indicates an event at step t in a sequence of monitored events in the order they were executed. In training system 200, score y e,t is typically provided to a reward function of DRL model 514 via path 530 to determine an output of at least one Q-value thereof. When system 200 has been trained and is being used to detect malware, path 530 is typically not used and DRL model 514 determines the output of at least one Q-value based on event state st To determine the output with at least one Q value. In one example, the event score y e,t It is also provided as output 562. Combined with... Figure 9 The EC512 and its operation are described in more detail. As described in conjunction with block 412 of method 400, and as at least... Figure 3 As illustrated in the diagram, ES 518 is typically created step-by-step by ECM 510 based on the sequence of event IDs generated by the executable file.

[0059] DRL model 514 is generally identical to DRL model 214 and performs the same function. In one example, DRL model 514 is implemented as a nonlinear approximator, such as a deep neural network. In alternative examples, DRL model 514 can be implemented as a linear approximator or a quantum computer. In one example, the output of DRL model 514 can be a value given the input event state s. t The Q-values ​​are in the form of pairs, where one Q-value is for a continuing action and the other Q-value is for a pausing action. Alternatively, a single Q-value can be generated by the DRL model 514. The term "Q-value" as used herein (including in the language of the claims) refers to a given action a t In step t, the given state s t The expected utility at that time.

[0060] Typically, the DRL model 514 must be trained before it can be used for malware detection. Specifically, during training, the DRL model 514 operates based on event states, actions, rewards, and policies. For example, event states... t Event states like 518 combined Figure 3 Described. The action is defined as an action performed on the corresponding monitored event e. t Given input event state s t The terms "continue" and "pause" are used. Rewards are typically constructed during training and used internally by the DRL model to determine its output, which has at least one Q-value. The policy generally refers to the mapping function from event states to actions, and is discussed further below.

[0061] In one embodiment, the reward function of the DRL model 514 used during the training of system 200 is defined as:

[0062] r t =0.5-|y e,t -L|×e -βt

[0063] Where r tis a reward at step t, a label L e {0, 1} is defined as the true label of the training file, where 0 indicates that the known file is benign and 1 indicates that the known file is malware. The decay factor β is typically selected through experimentation and in one example is 0.01. The reward value r t is used by the DRL model 514 to determine its output of at least one Q-value. In the context of training, the Q-value follows the optimal policy π and in one example is defined as:

[0064]

[0065] where R t includes both the reward value r t at the state s t t, and the cumulative reward to be obtained in the future considering the policy from the current state s t to its neighboring states s t +1, etc. by taking a specific action a t at step t. The action here corresponds to the execution control decision of the ECM 510 being provided as output 560: i.e., to continue or to pause the execution of the file. The output of the DRL model 514 is in the form of at least one Q-value.

[0066] In one example, the following algorithm describes the training process with example starting values for training the DRL model 514. Other variables for training are also possible.

[0067]

[0068]

[0069]

[0070] Once trained, the system 200 can be used to detect malware. Once trained, the output of the DRL model 514 can be in the form of a pair of Q-values based on a given input event state s t , where one Q-value of the pair is for the continue action and the other Q-value of the pair is for the pause action. Alternatively, the DRL model 514 can produce a single Q-value.

[0071] The ASM 520 generally filters the Q-values output from the DRL model 514 to produce an execution control signal or decision 560 for the file being executed. In one embodiment, the ASM 520 filters the Q-values based on a majority vote of the K most recent Q-values to determine whether file execution should continue or be paused. In one example, the ASM 520 filters approximately 200 Q-values or Q-value pairs to make the decision. In other examples, the ASM 520 filters other numbers of Q-values or Q-value pairs. In one embodiment, the number K can be determined based on hyperparameter tuning. In one embodiment, the output 560 is provided as input to the IM 220.

[0072] Figure 6 is a block diagram illustrating an example method 600 for determining event scores and producing an execution decision of whether to continue or pause file execution. The method 600 generally aligns with the method 400, but includes more detail. In one embodiment, the method 600 is performed by the ECM 510, or the like. In one example, the method 600 is performed as follows.

[0073] Block 610 generally indicates constructing a most recent event history in response to the most recent monitored event. The most recent event history is generally relative to the most recent monitored event e t and takes the form of a sliding window structure, in one example a first-in-first-out (“FIFO”) queue, such as the structure of the SEW 516. The most recent event history is generally constructed to retain or indicate E most recent event IDs that correspond to a sequence of monitored events in the order in which they were executed by the file execution. When the most recent monitored event e t is received and added to the full history, the oldest event e t-E in the history is removed so that E most recent event IDs are always maintained in the history. As such, the most recent event history is constructed or reconstructed upon receipt of each new monitored event e t via the input 580. In one example, the most recent event history is initially populated with padding events. In some examples, the most recent event history can not be needed in the method 600; otherwise, block 610 can not be needed in the method 600, only when training the system 200 or using the event score histograms for training the system 200 or detecting malware. Once the most recent event history is constructed for the most recent monitored event e t , the method 600 generally continues at block 612.

[0074] Block 612 generally indicates that, in response to the most recent event e tThe event ID is used, and based on the latest event history in SEW 516, the likelihood that the latest event history indicates malicious activity is determined. In one example, this is done by evaluating the latest monitored event e relative to the most recent monitored event via EC 512. t The latest event history of SEW 516 determines the event relative to event e. t The latest event history indicates the event score of malicious activity. e,t This is accomplished by event e. t The event y indicates the event at step t in the sequence of monitored events executed in the order they were executed by the executable file. In some examples, the event score y is only needed when training system 200, or when using an event score histogram to train system 200, or when detecting malware. e,t (Provided via path 532 to construct the event state); otherwise, box 610 may not be required in method 600. Once the event score y corresponding to the most recent monitored event is determined... e,t Method 600 usually continues at box 614.

[0075] Box 614 typically indicates: based on the latest event e at step t in the sequence of monitored events in the order in which they are executed by the executable file. t The provided event ID is used to construct the corresponding event state s. t In one example, such as combining Figure 4 As described in box 412, specific event states are constructed. Furthermore, event states s t This can be constructed by including additional event score histograms for each event type. For example, given 100 different monitored event types, the event states s t 320 can include a set of 100 event score histograms, one event score histogram for each event type. Figure 3 (Not shown in the image). These event score histograms may be included in event state 320 or otherwise indicated by event state 320 so that they can be used by DRL model 514.

[0076] In one example, each event score histogram in this set of event score histograms is presented as an ordered array of partitions (buckets). For instance, suppose the event score indicates a value between 0 and 1 relative to event e. t The latest event history indicates the probability of malicious activity. Four partitions divide this probability into four equal parts (e.g., [0-0.24], [0.25-0.49], [0.5-0.74], [0.75-1]), while ten partitions divide the probability into ten equal parts. Any number of partitions can be used, although more partitions tend to require more memory. Typically, all partitions are initialized to zero.

[0077] For example, suppose the latest event e t It is type 1, and for example, assuming the corresponding event score y e,t If the probability is 0.29, then for event type 1, the second partition of the four partitions in the four-part event score histogram is incremented by 1 to indicate event e. t Event score y e,t It is type 1 and falls between 0.25 and 0.49. Considering a ten-zone histogram, the third zone (e.g., indicating 0.20-0.29) would be incremented. On the other hand, if event e t If the type is type 87 instead of type 1, then the score histogram for the 87th event will be the modified histogram. Thus, when using a bisected score histogram as the event state s... t When a portion of 320 is reached, these histograms will be based on the event score y as described above. e,t Updated, the event score is y e,t The latest event e corresponds to step t in the sequence of monitored events in the order in which they are executed by the executable file. t In some examples, multiple event score histograms can be combined into a single histogram.

[0078] Box 616 typically indicates: based on the most recent monitored event e t event states t This generates Q-values ​​or Q-value pairs for continuing and / or pausing file execution. In one example, this is achieved by evaluating the most recent monitored event e using DRL model 514. t event states t (such as event state 320) and its corresponding score y e,t This is accomplished. In this example, the event state includes at least an event histogram, such as that described in conjunction with event state 320. The event state may additionally or alternatively include an event score histogram as described above. Once generated with the most recent monitored event e... t The corresponding Q-value or Q-value pair, method 600 typically continues at box 618. When training system 200, as in detecting malware after training, this generation is typically based on the most recent monitored event. t event states t and its corresponding score y e,t .

[0079] Box 618 typically indicates that an execution decision should be made regarding whether the file should continue execution or be suspended, based on the K most recent Q values ​​or Q value pairs. In one embodiment, this is done via ASM 520 based on a majority voting pair relative to the most recent monitored event e. tK most recent Q values or Q value pairs are filtered to produce a decision h whether to continue or suspend execution of the file t Once the decision h is produced t , the method 600 can continue at block 610 with the next monitored event e t+1 , or can further process the decision h as described in Figure 8 t .

[0080] Figure 7 is a block diagram showing an example inference model (“IM”) 720, which is generally the same as and performs the same functions as the IM 220, although additional details are shown in connection with the IM 720. In addition to the main component, a file classifier (“FC”) 722 (which is the same as the FC 222), the IM 720 includes an event buffer (“EB”) 724, a sliding event probability window (“SEPW”) 726, and a score module (“SM”) 728. In some embodiments, the EB 724 and the FC 722 are optional and can not be included.

[0081] The input 780e t (which is the same as 580) is generally from the EM 230 and is generally in the form of a sequence of event IDs, where e t indicates the event at step t in the sequence of monitored events in the order in which they were executed by the execution file. The input 782p e,t is generally from the output 562 of the ECM 510 and is generally an event score y t indicating malicious activity relative to the event e e,t at step t. The input 784h t is generally from the output 560 of the ECM 510 and is generally a suspend or continue decision relative to the event e t at step t. And as with the MDS 200, the output 770 is generally the same as the output 270, which is generally used to classify the execution file as malware (malicious) or benign.

[0082] ​EB 724 is typically an event buffer structure, in one example a queue, that is typically maintained by IM 720 and that typically retains or indicates a first event history comprising V first event IDs received at input 780 from a sequence of event IDs corresponding to monitored events in the order they were executed by the file. In one example, EB 724 retains or indicates the first 200 event IDs of the first 200 monitored events in the order they were executed by the file (as opposed to the most recent monitored events). In other examples, EB 724 retains or indicates other numbers of event IDs. In one embodiment, the number V can be determined based on hyperparameter tuning. The term “first event history” as used herein (including in the claim language) refers to a list of the first V monitored events in the sequence of monitored events in the order they were executed by the file.

[0083] FC 722 is typically the same as FC 222 and performs the same functions as FC 222. In one example, FC 722 is a two-stage neural network structure in which the first stage is a recurrent neural language model that generates a feature vector, which is then input to a second classifier stage. The recurrent neural language model can be a recurrent neural network (“RNN”) model. Alternatively, the recurrent neural language model can be a long short-term memory (“LSTM”) model, gated recurrent unit (GRU), or any suitable recurrent neural model. In another embodiment, the recurrent neural language model can be replaced with a sequential convolutional neural network (CNN). The classifier stage can be any supervised classifier, such as a logistic regression-based classifier, support vector machine, neural network, or deep neural network.

[0084] FC 722 typically evaluates the first event history comprising the first V monitored event IDs of EB 724 in order to determine a file score y f,t that the file being executed is a malicious or benign file in the sequence of monitored events in the order they were executed by the file. f,t is typically provided to SM 728. In conjunction with Figure 9 FC 722 and its operation are described in more detail.

[0085] SEPW 726 is typically a sliding window structure, in one example a first-in-first-out (“FIFO”) queue, which is typically maintained by IM 720 and typically retains or indicates the latest event score history, including W latest event scores received from input 782, which correspond to W latest event IDs from an event ID sequence, which correspond to monitored events in the order they were executed by the executable file. In one example, SEPW 726 retains or indicates approximately 200 event scores. In other examples, SEPW 726 retains or indicates other numbers of event scores. In one embodiment, this number may be determined based on hyperparameter tuning. As used herein, the term “latest event score history” (included in the language of the claims) refers to a list of event scores corresponding to W latest monitored events in the sequence of monitored events executed by the executable file.

[0086] SM 728 typically responds to a decision instructing a suspension of execution. t Input 784 to calculate the final improved file classifier score y for the file being executed. RL,t In one example, this score is an improved score of whether the executable is malicious or benign, and is relative to step t in the sequence of monitored events in the order in which they are executed by the executable, and is based on three inputs: (1) the score y from FC 722. f (2) The W most recent event scores relative to step t from SEPW 726, and (3) the decisions h from input 784 that may be considered too noisy to be used directly. t In one example, the computation is performed as follows: in response to a decision instructing to pause execution, h... t If y f If the value is greater than 0.5, the executable file is more likely to be malicious, therefore y RL,t Set to the maximum value y from the W most recent event scores e,t Otherwise, if y f If the value is ≤0.5, the executable file is more likely to be benign, therefore y RL,t Set to the minimum value y from the scores of the W most recent events e,t Improve score y RL,t It is typically provided as an output of 770 and is an improved score indicating that the executable is malware (malicious).

[0087] Figure 8 This is a block diagram illustrating an example method 800 for determining an improvement score, which indicates the likelihood that the executable file is malicious or benign. In one embodiment, method 800 is executed by an IM 720, etc. In one example, method 800 is executed as follows.

[0088] Box 810 typically indicates that a first event history is constructed based on the first V monitored events in the order they were executed from the start of file execution. In one example, the first event history takes the form of a queue that holds or indicates the first V monitored events, such as EB724. Typically, the first event history is constructed to hold or indicate the first V monitored events in the order they were executed from the start of file execution. For example, the first event history typically consists of event IDs 1 through V. In one example, the queue is initially filled with fill events. Once the first event history is constructed, it typically remains unchanged, and method 800 typically continues at box 812.

[0089] Box 812 typically indicates: in response to receiving the most recent monitored event e at step t. t Furthermore, based on the first event history of EB 724, a document score y was determined to indicate the likelihood that the instruction execution document was malicious. f In one example, this is a file score y used to determine if the executable file is malicious by evaluating the first event history of EB 724 through FC 722. f This is achieved by [method 800]. Once the file score is determined, method 800 typically continues at box 814. The steps in boxes 810 and 812 are optional and may not be included in all embodiments.

[0090] Box 814 typically indicates: based on the most recent monitored event e t Corresponding score y e,t This constructs the latest score history. The latest event score history is typically relative to the most recent monitored event. t And it takes the form of a sliding window structure, in one example a first-in-first-out (“FIFO”) queue, such as the SEPW 726. Typically, an event score history is constructed to retain or indicate the W most recent event scores received from input 782, corresponding to the W most recent event IDs from the event ID sequence, which correspond to the monitored events in the order they were executed by the executable file. With the most recent score y... e,t Add it to the complete history, the oldest score in history y e,t-W This will be removed so that W most recent event scores are always kept in the history. Thus, when each new monitored event is received via input 780... t At that time, the latest event history is built or reconstructed. In one example, the latest event history is initially populated with filling events. Once the most recent score y is applied... e,t Once the latest score history has been constructed, method 800 typically continues at box 816.

[0091] Box 816 typically indicates: an improved score for determining the likelihood that the executable file is malicious or benign. In one example, this determination is performed by SM 728 based on the inputs and calculations described in conjunction with SM 728. Once the improved score is determined, method 800 is typically complete.

[0092] Figure 9 This is a block diagram illustrating an example classifier 920 that can be used to implement event classifier 512 and / or document classifier 722. History 912 is typically considered a separate input to classifier 920 and indicates the latest event history, such as that provided by SEW 516 in the case of event classifier 512, or the latest event score history, such as that provided by SEPW 726 in the case of document classifier 726. In one embodiment, the appropriate history provided at each step t is provided as input to embedding layer 921, and the result is then provided as input to recursive layer 922. In one example, recursive layer 922 is implemented as a recursive neural network (“RNN”). The hidden states of the recursive layer are then provided as input to max-pooling layer 923, which is generally better at detecting malicious activity in the history.

[0093] Next, feature vector 926 is formed, comprising: (1) a bag-of-words (“BOW”) representing history; (2) the final hidden state of recursive layer 922, which is recursive layer embedding 924; and (3) the output of max pooling layer 923, which is max pooling embedding 925. In various examples, feature vector 926 can be either a sparse binary feature vector or a dense binary feature vector. In one example, the BOW 924 of the feature vector consists of 114 features, the recursive layer embedding 924 consists of 1500 features, and the max pooling embedding 925 consists of 1500 features, resulting in a feature vector 926 of size 3114×1. In other examples, other numbers of features can be used. In other examples, the sparse binary features may contain only the max pooling embedding 925.

[0094] Finally, the feature vector 926 is provided as the output of classifier layer 927. Layer 927 can typically be any supervised classifier, such as a logistic regression-based classifier, support vector machine, neural network, shallow neural network, or deep neural network. The output of classifier layer 927 is typically a sigmoid function. Specifically, as an event classifier 512, the output is the event score y. e,t Event score y e,t The latest event history indicates the likelihood of malicious activity, where e t The event at step t in the sequence of monitored events according to the order in which they were executed by the executed file. Alternatively, in the file classifier 722, the output is a file score y. fwhich indicates a likelihood that the executing file is malicious. This score is provided as a classifier output 990.

[0095] CONCLUSION

[0096] In a first example, a method performed on at least one computing device including at least one processor and memory, the method comprising: executing, by the at least one computing device, at least a portion of a file; monitoring, by the at least one computing device, execution of the file, the execution of the file being sufficient to identify a sequence of monitored events performed by the executing file; constructing, by the at least one computing device, an event state including an event histogram based on the monitored events; generating, by the at least one computing device, at least one Q-value based on the event state; producing, by the at least one computing device, a decision to continue executing the file or to pause execution of the file based on the at least one Q-value; pausing, by the at least one computing device, execution of the at least a portion of the file responsive at least to the decision to pause execution of the file.

[0097] In a second example, there is at least one computing device comprising: at least one processor and memory coupled to the at least one processor and comprising computer-executable instructions that, based on execution by the at least one processor, configure the at least one computing device to perform acts comprising: executing, by the at least one computing device, at least a portion of a file; monitoring, by the at least one computing device, execution of the file, the execution of the file being sufficient to identify a sequence of monitored events performed by the executing file; constructing, by the at least one computing device, an event state including an event histogram based on the monitored events; generating, by the at least one computing device, at least one Q-value based on the event state; producing, by the at least one computing device, a decision to continue executing the file or to pause execution of the file based on the at least one Q-value; pausing, by the at least one computing device, execution of the at least a portion of the file responsive at least to the decision to pause execution of the file.

[0098] In a third example, at least one computer-readable medium comprising computer-executable instructions that, based on execution by at least one computing device, configure the at least one computing device to perform acts comprising: executing, by the at least one computing device, at least a portion of a file; monitoring, by the at least one computing device, execution of the file, the execution of the file being sufficient to identify a sequence of monitored events performed by the executing file; constructing, by the at least one computing device, an event state including an event histogram based on the monitored events; generating, by the at least one computing device, at least one Q-value based on the event state; producing, by the at least one computing device, a decision to continue executing the file or to pause execution of the file based on the at least one Q-value; pausing, by the at least one computing device, execution of the at least a portion of the file responsive at least to the decision to pause execution of the file.

[0099] In the first, second, and third examples: the generating is further based on an event histogram; the generating is further based on a set of event score histograms corresponding to the monitored events; the generating is further based on a step number of the monitored events corresponding to the event state; the generating is further based on an identifier of the monitored events corresponding to the event state; the reinforcement learning model is a deep reinforcement learning model; and / or the method and actions further comprise: constructing, by the at least one computing device, an event score history based on the event scores; and determining, by the at least one computing device, an improved score indicating a likelihood that the executing file is malicious or benign based on the event score history and the decision, wherein the pausing is further based on the improved score indicating that the executing file is malicious.

Claims

1. A method performed on at least one computing device comprising at least one processor and memory, the method comprising: executing at least a portion of a file; identifying a sequence of monitored events during execution of the at least a portion of the file, the sequence of monitored events having a sequence length that varies according to content of the file, the sequence of monitored events comprising any type of event; based on the monitored events, constructing an event state comprising an event histogram, wherein the event histogram is an ordered array representing a plurality of monitored event types, and wherein each position in the event histogram is initially set to zero and is associated with a different event identifier (ID) representing a monitored event type, at each time step, a position in the ordered array is incremented according to an event ID of the monitored events; generating, from a reinforcement learning model, at least one value based at least on the event state comprising the event histogram, the at least one value representing an expected utility of pausing the execution of the at least a portion of the file; and pausing the execution of the at least a portion of the file based on the at least one value.

2. The method of claim 1, wherein the generating is further based on a set of event score histograms corresponding to the monitored events.

3. The method of claim 1, wherein the generating is further based on a step number of the monitored events corresponding to the event state.

4. The method of claim 1, wherein the generating is further based on an identifier of the monitored events corresponding to the event state.

5. The method of claim 1, wherein the reinforcement learning model is a deep reinforcement learning model, the method further comprising: prior to generating the at least one value, determining an event score associated with a most recent event in the sequence of monitored events, the event score representing that a most recent event history indicates a malicious event sequence; providing the event score to the deep reinforcement learning model; wherein generating the at least one value further comprises generating the at least one value based on the event score; and wherein the at least one value comprises at least one Q-value for pausing the execution.

6. The method of claim 1, further comprising: constructing, by the at least one computing device, an event score history based on event scores; and determining, by the at least one computing device, an improvement score based at least on the event score history, the improvement score indicating a probability that the file is malicious, wherein the pausing is further based on the improvement score indicating that the file is malicious.

7. At least one computing device comprising: at least one processor and memory coupled to the at least one processor and comprising computer-executable instructions that, based on execution by the at least one processor, configure the at least one computing device to perform acts comprising: executing at least a portion of a file; identifying a sequence of monitored events during execution of the at least a portion of the file, the sequence of monitored events having a sequence length that varies according to content of the file, the sequence of monitored events comprising any type of event; based on the monitored events, constructing an event state comprising an event histogram, wherein the event histogram is an ordered array representing a plurality of monitored event types, and wherein each position in the event histogram is initially set to zero and is associated with a different event identifier (ID) representing a monitored event type, at each time step, a position in the ordered array is incremented according to an event ID of the monitored events; generating, from a reinforcement learning model, at least one value based on at least the event state comprising the event histogram, the at least one value representing an expected utility of pausing the execution of the at least one portion of the file; and pausing the execution of the at least one portion of the file based on the at least one value.

8. The at least one computing device of claim 7, wherein the generating is further based on a set of event score histograms corresponding to the monitored events.

9. The at least one computing device of claim 7, wherein the generating is further based on a step number of a monitored event corresponding to the event state.

10. The at least one computing device of claim 7, wherein the generating is further based on an identifier of a monitored event corresponding to the event state.

11. The at least one computing device of claim 7, wherein the reinforcement learning model is a deep reinforcement learning model.

12. The at least one computing device of claim 7, the actions further comprising: constructing, by the at least one computing device, an event score history based on event scores; and determining, by the at least one computing device, an improvement score based on at least the event score history, the improvement score indicating a probability that the file is malicious, wherein the pausing is further based on the improvement score indicating that the file is malicious.

13. At least one computer-readable medium comprising computer-executable instructions to configure at least one computing device, based on execution of the computer- executable instructions by the at least one computing device, to perform actions comprising: executing at least one portion of a file; identifying, during execution of the at least one portion of the file, a sequence of monitored events, the sequence of monitored events having a sequence length that varies according to content of the file, the sequence of monitored events comprising any type of event; based on the monitored events, constructing an event state comprising an event histogram, wherein the event histogram is an ordered array representing a plurality of monitored event types, and wherein each position in the event histogram is initially set to zero and is associated with a different event identifier (ID) representing a monitored event type, at each time step, a position in the ordered array is incremented according to an event ID of the monitored events; generating, from a reinforcement learning model, at least one value based on at least the event state comprising the event histogram, the at least one value representing an expected utility of pausing the execution of the at least one portion of the file; and and pausing the execution of the at least one portion of the file based on the at least one value.

14. The at least one computer readable medium of claim 13, wherein the generating is further based on a set of event score histograms corresponding to the monitored events.

15. The at least one computer readable medium of claim 13, wherein the generating is further based on a step number or identifier of a monitored event corresponding to the event state.

16. The at least one computer readable medium of claim 13, wherein the reinforcement learning model is a deep reinforcement learning model.

17. The at least one computer readable medium of claim 13, the acts further comprising: building, by the at least one computing device, an event score history based on event scores; and determining, by the at least one computing device, an improved score based on the event score history, the improved score indicating a probability that the file is malicious, wherein the pausing is further based on the improved score indicating that the file is malicious.

Citation Information

Patent Citations

  • System and Method for Distributed Denial of Service Identification and Prevention

    US20100082513A1

  • System and method for automated machine-learning, zero-day malware detection

    US20140090061A1