Detection of Malicious Behavior of Computer Programs

By extracting subsets from system class calls intercepted by computer programs on Android system and converting them into vector representations, combining statistical information, using pre-trained machine learning models for real-time detection, the problem of high computing and storage requirements in the existing technology is solved, and efficient real-time malicious behavior detection is achieved.

CN113196268BActive Publication Date: 2025-05-30HUAWEI TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN201980083156.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-01-10
Publication Date
2025-05-30
Estimated Expiration
2039-01-10

AI Technical Summary

Technical Problem

The prior art has high computing and storage requirements when detecting real-time malicious behavior of computer programs on Android systems, especially on mobile devices.

Method used

By extracting a subset of system class calls from a large sequence of intercepting system class calls generated by a computer program and converting them into vector representations, combined with statistical information, a pre-trained machine learning model is used for real-time detection.

Benefits of technology

It reduces the computing and storage requirements on mobile devices, improves detection accuracy, can effectively monitor the behavior of computer programs in real time, and reduces the damage to device resources and data by malware.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113196268B_ABST
    Figure CN113196268B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for determining real-time malicious behaviors of a computer program, such as on an Android system. Store and convert a first API sequence in the total sequence of intercepted APIs generated by the computer program into a vector representation, where the first API sequence includes inputs and statistical information of APIs in the first sequence and APIs in the total sequence, for determining whether the behaviors of the computer program are abnormal. The determination is made through pre-trained data sets and models in different types of machine learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to computer science, and in particular to a real-time detection of whether a computer program behavior is abnormal. The judgment is based on the following inputs: 1. Select a most recent subset of system class calls from a large set of intercepted system class calls of a computer program. 2. Statistical information related to the system class calls in the subset and the large set. A pre-trained judgment module uses machine learning to judge whether the input indicates abnormal behavior. In order to reduce the computing and storage requirements of devices such as mobile devices, the input is converted into a vector representation, and only the most recent subset is stored, and the statistical information includes relatively old information. Background Art

[0002] According to a report by information security company G Data (C. Lueck, "8,400 New Android Malware Samples Every Day"), 8,400 Android malware were discovered every day in 2017, that is, a new Android malware was discovered every 10 seconds. In order to cope with the ever-changing Android malware, in recent years, application-audit systems based on machine learning have been developed to automatically detect Android malware. Although static or dynamic analysis systems deployed on the market side or the device side can detect most malware, they are restricted by, for example, code obfuscation techniques, anti-emulator techniques, etc.

[0003] The Drebin system was published by D. Arp, H. Gascon, M. Huebner, K. Rieck, and M. Spreitzenbarth in the proceedings of the 21st Annual Network and Distributed System Security Symposium (NDSS'14) held by the Internet Society in San Diego, California, USA in February 2014, "Drebin: An Efficient and Explainable Portable System for Detecting Android Malware". The Drebin system is a lightweight system that directly detects Android malware on a smartphone during software installation. Drebin performs static analysis on Android applications, and collects 8 feature sets from the manifest file and disassembled code. Using one-hot vector representation, these features are mapped to a 545,000-dimensional vector space. Then, a linear support vector machine model is trained based on malware classes and benign classes, and is deployed on the smartphone to detect malicious applications.

[0004] N. McLaughlin, J.M. Deneline, B. Kang, S. Yerima, P. Miller, S. Shas, Y. Safai, E. Chikker, Z. Zhao, A. Dup, and G.J. Ann published "Deep Android Malware Detection" in the proceedings of the 7th ACM Conference on Data and Application Security and Privacy (CODASPY’17), held at 301 - 308, New York, NY, USA. A deep Android malware detection based on a deep convolutional neural network (CNN) takes the disassembled opcode sequence as input. These disassembled opcodes are encoded as one - hot 218 - dimensional vectors.

[0005] The focus of the above two systems is on the static malware analysis of Android applications. Therefore, these two systems have the disadvantages of static analysis, such as difficulty in overcoming code obfuscation techniques, native code, etc.

[0006] The inventors found that there are real - time detection methods based on third - party learning, but the public information in this regard is relatively limited. In addition, the inventors' analysis shows that these methods require limited statistical analysis of the monitoring units and have high storage and computational requirements due to the size of the monitoring unit sequences. Summary of the Invention

[0007] The embodiment provides a method for determining real - time malicious behavior of a computer program, such as on an Android system. The system call first sequence (such as API) in a large sequence (such as the total sequence) of intercepted system calls generated by the computer program is stored and converted into a vector representation. The vector representation includes inputs that combine the statistics of the APIs in the first sequence and the APIs in the large sequence to determine whether the behavior of the computer program is abnormal. In this embodiment, the determination is made through a pre - trained data set and model in different types of machine learning. Other embodiments provide an apparatus and a computer program for implementing the method of the present invention.

[0008] To achieve the above object, the present embodiment adopts the following technical solutions:

[0009] The first aspect of the present embodiment provides a method for monitoring a computer program, including:

[0010] - At time t0, obtain a second number N2 of monitoring units called by the computer program from a first number N1 of monitoring units called by the computer program, where the second number N2 of monitoring units is pre - configured to be less than the first number N1 of monitoring units, and one monitoring unit corresponds to one action performed by the computer program;

[0011] - Obtain the statistical information corresponding to the first quantity N1 of monitoring units;

[0012] - Convert the monitoring units of the second quantity N2 of monitoring units into vector representations of the second quantity N2 of monitoring units, such that the monitoring units of the second quantity N2 of monitoring units have their respective corresponding vector representations; and

[0013] - Determine whether the behavior of the computer program is abnormal based on the statistical information and the vector representations of the second quantity N2 of monitoring units.

[0014] This method can be used in any computing environment where malware may occur, such as third-party applications. Preferably, the method can be used in a computing environment using an operating system such as Android.

[0015] The "preferably" used in the present invention can be understood as illustrative but not limiting to one possible implementation manner, and it is not necessarily the only implementation manner, that is, there are also other possible implementation manners.

[0016] Furthermore, it can be understood that although the present invention mainly relates to the Android system, it is not limited thereto.

[0017] In addition, the term "and / or" means that there are three possible cases, namely existing alone or existing simultaneously.

[0018] Preferably, the method can be used in end-user devices, such as portable devices with relatively limited processing power and memory, such as terminals or telephones.

[0019] The method can be part of a detection system.

[0020] The computer program can be pre-installed or installed subsequently. The computer program can be a third-party application, including hacked or malware versions of what are apparently official computer programs. As is well known, hacked or malware versions can cause harm, such as causing instability, data loss or theft, stealing processing cycles, etc.

[0021] The computer program is preferably run when obtaining the monitoring units, so that the method can monitor the computer program in real time according to the actions generated by the computer program.

[0022] The obtaining of the monitoring units can include intercepting, recording, or staging the monitoring units. For example, sensitive APIs (such as sendSMS()) can be suspended, so as to be intercepted by modifying the OS framework. It may not be possible to obtain all the monitoring units called by the computer program.

[0023] The monitoring unit may include system calls at the kernel level or application programming interface calls at the Android system framework level, enabling the computer program to access system resources and services. Therefore, each monitoring unit corresponds to an action of the computer program during runtime. When the computer program runs, the monitoring unit can be called or generated. When the computer program runs, a series of monitoring units are usually generated, and the series of monitoring units increases as the computer program runs (i.e., the sequence increases in length or size). The computer program may call millions of monitoring units.

[0024] The first quantity N1 of monitoring units may include any number of monitoring units. It may include the total number of monitoring units called by the computer program during runtime. For example, when the monitoring units called by the computer program at time t include API1, API2, API3, and API4, the first quantity N1 of monitoring units is 4. Preferably, the first quantity N1 of monitoring units starts from the first monitoring unit called by the computer program. Preferably, the first quantity N1 of monitoring units is obtained and arranged according to the time series corresponding to the call order of the computer program. For example, if the monitoring unit is an API, the total sequence may be API1, API2, API3, and API4 arranged in order at time t.

[0025] The second quantity N2 of monitoring units may include any number of monitoring units less than the first quantity N1 of monitoring units. Preferably, the second quantity N2 of monitoring units may include a subset selected in chronological order from the first quantity N1 of monitoring units. The second quantity N2 of monitoring units can be pre-configured (fixed), for example, according to the requirements and capabilities of the operating system. The pre-configured size needs to consider detection accuracy and performance cost. The larger the pre-configuration, the higher the detection accuracy, but the higher the storage cost. For example, if the monitoring unit is an API, the second quantity N2 of monitoring units may be API3 and API4 arranged in the call order at time t. Therefore, the second quantity N2 of monitoring units can represent a relatively more recent API sequence, which is a subset of the first quantity N1.

[0026] Preferably, the statistical information corresponding to the first number N1 of monitoring units may be related to all or (optionally) a part of the first number N1 of monitoring units. For example, the statistical information may correspond to the sampling data of the first number N1 of monitoring units. The statistical information may correspond to the remaining monitoring units among the first number N1 of monitoring units that are not in the second number N2 of monitoring units. The statistical information corresponding to the first number N1 of monitoring units may also correspond to the monitoring units in the second number N2 of monitoring units and their mutual relationships. The statistical information may be related to the occurrence times (proportions), counts, or deviation ratios of each API in the second number N2 of monitoring units corresponding to the first number N1 of monitoring units. The statistical information may be related to the total API sequence, and after the total API sequence is converted into a vector representation, it can be regarded as a compressed vector representing the total API sequence.

[0027] Since the statistical information of the monitoring units not in the second number N2 of monitoring units is considered, the detection accuracy is improved. Those skilled in the art should understand several measurement methods for malware detection accuracy, including accuracy, recall, precision, and F1-score. Preferably, the present invention uses accuracy to measure binary classification. The quantization formula for binary accuracy is: accuracy = (TP + TN) / (TP + TN + FP + FN), where TP is True positive, FP is False positive, TN is True negative, and FN is False negative.

[0028] During the judgment process, all representations of each monitoring unit not in the second number N2 of monitoring units are not stored or included, which can reduce the computational requirements and improve the speed.

[0029] Preferably, each monitoring unit in the second number N2 of monitoring units can be converted into its corresponding vector representation for convenient analysis or input into a machine learning model. The vector representation may be a numerical vector representation. The first number N1 of monitoring units can also be converted into their corresponding vector representations, but they are not stored separately.

[0030] Based on the input of the statistical information and the vector representations of the second number N2 of monitoring units, it is determined whether the behavior of the computer program is abnormal. Preferably, it is determined by detection models such as support vector machines, random forests, and deep neural networks. The detection models may have been pre-trained with appropriate data sets.

[0031] The method can take further actions according to the judgment result. For example, when the judgment result is abnormal behavior, the method can notify the operating system or the end user to delete or isolate the computer program. By detecting in real time and enabling subsequent actions, the destruction of device resources or data can be minimized. If the judgment result is not abnormal behavior, the method can continue to monitor by repeating the steps of the method, for example, at a predetermined time interval, such as every 2 minutes.

[0032] The order of the steps of the method is not fixed and can be executed in any order, including the shown order. For example, the order of obtaining the statistical information corresponding to the first number N1 of monitoring units and the order of converting the second number N2 of monitoring units into vector representations can be interchanged.

[0033] In a possible design, obtaining the second number N2 of monitoring units called by the computer program from the first number N1 of monitoring units called by the computer program further includes:

[0034] The monitoring units in the second number N2 of monitoring units and / or the first number N1 of monitoring units are arranged in a sequence in chronological order, such as the chronological order of calls of the computer program.

[0035] In a possible design, the monitoring units in the second number N2 are the most recently obtained monitoring units compared with the monitoring units in the first number.

[0036] Whether a monitoring unit is the most recently obtained can be determined with reference to the call time of the monitoring unit.

[0037] Only including the most recent monitoring units in the second number N2 of monitoring units, the storage and processing requirements are relatively low in this way. For example, if the monitoring units are APIs, the total sequence may include 4 APIs, namely API1, API2, API3, and API4 arranged in order at time t. API4 is closer than API3, API3 is closer than API2, and API2 is closer than API1. For example, if the second number N2 of monitoring units is preconfigured to be 2, less than 4, then the second number N2 of monitoring units may include API3 and API4. At time t+1, for example, another most recent API5 is generated, the total sequence is API1, API2, API3, API4, and API5, and the second number N2 of monitoring units includes API4 and API5.

[0038] In a possible design, at time t0, obtaining the second number N2 of monitoring units called by the computer program from the first number N1 of monitoring units called by the computer program includes:

[0039] Obtain a second quantity N2 of monitoring units, where the end monitoring unit among the second quantity N2 of monitoring units is one of the nearest monitoring units in a set including a first quantity N1 of monitoring units, and the other end monitoring unit among the second quantity N2 of monitoring units is a nearest monitoring unit compared to the monitoring units in the first quantity N1 of monitoring units.

[0040] In one possible design, the obtaining the statistical information corresponding to the first quantity N1 of monitoring units includes:

[0041] By representing the obtained statistical information as a vector representation of the statistical information, convert it into the statistical information used in the judgment step.

[0042] Preferably, the statistical information used in the judgment step is in vector representation format, which is convenient for machine learning processing, especially in the case of a large number of monitoring units. Preferably, for simplicity of calculation, before judgment, the vector representation of the statistical information is combined with the vector representation of the second quantity N2 of monitoring units.

[0043] In one possible design, the obtaining the second quantity N2 of monitoring units called by the computer program from the first quantity N1 of monitoring units called by the computer program includes:

[0044] Obtain a second quantity N2 of monitoring units from the first quantity N1 of monitoring units through a sliding window of a preconfigured size, such that at any specific time t, the monitoring units in the second quantity N2 of monitoring units include the nearest monitoring units in the time series of the first quantity N1 of monitoring units.

[0045] Each computer program can have its own corresponding sliding window of a preconfigured size. The sliding window is a sampling tool of a fixed size during the operation of the computer program. The size of the sliding window can be adjusted, but it cannot be adjusted during the operation of the computer program. Preferably, the sliding window moves or filters the input sequence of monitoring units, such that the monitoring unit recently added to the second quantity N2 of monitoring units is the latest (i.e., the nearest) monitoring unit in the input sequence. The last monitoring unit in the order within the sliding window or covered by the sliding window is the oldest monitoring unit covered by the sliding window. When the sliding window moves from one monitoring unit to the nearest monitoring unit, the last monitoring unit in the order will be deleted from the second quantity N2 of monitoring units.

[0046] In one possible design, the method for monitoring a computer program includes:

[0047] A sliding window of the preconfigured size is obtained by subtracting the index value of the sliding window from the number of the first quantity N1 of monitoring units, where the index value corresponds to the earliest monitoring unit among the second quantity N2 of monitoring units.

[0048] In the stage of API sampling using a sliding window, a sliding window of a fixed size can be used, so that only the APIs in the sliding window need to be stored during runtime. Let n be the size of the API sequence and m be the index of the head of the sliding window. The API sequence can be represented as S = {x i , i ∈ [1, n]}, where x i is an API intercepted by the system under detection. Preferably, the size of the set S is greater than 1 and can include millions corresponding to the runtime APIs. The sliding window can be represented as W = {x j , j ∈ [m, n]}, where x j is an API belonging to the set S. Therefore, the size of the sliding window is n–m, and the size is preconfigured. When designing the OS, the size can be fixed and determined according to empirical values. Preferably, in the present invention, the size range of the sliding window is [200, 500].

[0049] In a possible design, obtaining the second quantity N2 of monitoring units called by the computer program from the first quantity N1 of monitoring units called by the computer program includes:

[0050] Storing the second quantity N2 of monitoring units.

[0051] This can save storage space.

[0052] In a possible design, obtaining the second quantity N2 of monitoring units called by the computer program from the first quantity N1 of monitoring units called by the computer program includes:

[0053] Not storing the monitoring units that are not in the second quantity N2 of monitoring units but in the first quantity N1 of monitoring units.

[0054] This can save storage space and only retain the relevant statistical information.

[0055] Different types of statistical information can be used. The three types of statistical information disclosed in the present invention are only examples, and the present invention is not limited thereto.

[0056] In a possible design, obtaining the statistical information corresponding to the first quantity N1 of monitoring units includes:

[0057] Obtain the occurrence ratio of the monitoring unit according to the occurrence times of one monitoring unit among the second quantity N2 monitoring units relative to the occurrence times of the monitoring unit among the first quantity N1 monitoring units.

[0058] In this possible design, the ratio can be defined as: for API x j ∈W, its ratio is rate_x j = the size of the set {x i , x i = x j , i ∈ [1, j]} / the size of the set {x i , i ∈ [1, j]}.

[0059] For example, assume that the size of the sliding window is 2, and the total sequence is <API_1, API_2, API_3, API_2, API_1>. The API sequence in the sliding window is <API_2, API_1>. For API_2 in the sliding window, its ratio is 2 / 4 = 0.5. For API_1 in the sliding window, its ratio is 2 / 5 = 0.4.

[0060] Store the API and its calculated ratio, such as storing in a data storage unit. The data storage unit is used to store data, and may include a list of the most recent API sequences, whose size is equal to the size of the sliding window, and a list of the common or all API ratios. The data can be used as an input for judging the behavior of a computer program. Key values can be used to retrieve each API ratio. The data storage unit can use different types of memories, such as a HashMap structure and shared memory, to store this data.

[0061] In a possible design, the obtaining of the statistical information corresponding to the first quantity N1 monitoring units includes:

[0062] Obtain the count value of the monitoring unit.

[0063] In this possible design, the count value of the monitoring unit refers to the number of times the API is called. The count value represents the frequency ratio of the API. For example, assume that the size of the sliding window is 2, and the total sequence is <API_1, API_2, API_3, API_2, API_1>. The API sequence in the sliding window is <API_2, API_1>. For API_2 in the sliding window, its API count value is 2. For API_1 in the sliding window, its API count value is also 2.

[0064] The method of the present invention can be repeatedly executed from time to time, for example, every 2 minutes. Each time point is a detection stage. The number of detections is a value representing the number of times the detection stage is called. For example, in the above example, at 2 minutes, the number of detections is 2; at 4 minutes, the number of detections is 3. The count value of each API increases as the computer program runs. Optionally, the statistical information obtained by dividing the count value by the number of detection stages can indicate whether the API is frequently called.

[0065] In a possible design, the obtaining of the statistical information corresponding to the first number N1 of monitoring units includes:

[0066] Obtaining the deviation of the ratio of the monitoring units in the second number N2 of monitoring units relative to the average ratio of the monitoring units in the sliding window.

[0067] In this possible design, the statistical information can be defined as: for API x j ∈W, the deviation value is rate_x j –Σrate_x k / (n–m), where rate_x is the ratio of the API, and x k ∈W. Several levels of deviation values can be predefined, such as very large, relatively large, equal, relatively small, very small. Very large / relatively large means the deviation value is a large / small positive number, and relatively small / very small means the deviation value is a large / small negative number. Preferably, three levels of deviation values can also be predefined as large, equal, and small. Large means the ratio of the API in the sliding window is greater than the average ratio of the API in the sliding window. Small means the ratio of the API in the sliding window is less than the average ratio of the API in the sliding window. Equal means the ratio is the same as the average value. The deviation value can represent the frequency ratio of API calls.

[0068] The determination of whether the behavior of the computer program is abnormal is further made based on the information of one or more sets of the fourth number N4 of monitoring units called at time t - 1, where time t - 1 is earlier than time t0, and the size of each set of the fourth number N4 is equal to the size of the second number N2 of monitoring units.

[0069] For example, the stored information can include a list of the most recent API sequences, the size of which is equal to the size of the sliding window, and / or a list of API ratios as described above.

[0070] The determination of whether the behavior of the computer program is abnormal is further made based on the stored statistical information of the monitoring units.

[0071] For example, the stored information can include a list of the most recent API sequences, the size of which is equal to the size of the sliding window, and / or a list of API ratios as described above.

[0072] In a possible design, the monitoring unit is an Application Programming Interface (API) or a system call.

[0073] In a possible design, the conversion to the vector representation and / or statistical information of the monitoring unit includes conversion to one-hot representation or conversion based on word embedding technology.

[0074] For example, there are 3 different APIs (i.e., API_1, API_2, and API_3) for one-hot representation. These 3 APIs are represented in one-hot as:

[0075] API_1 = [1, 0, 0]

[0076] API_2 = [0, 1, 0]

[0077] API_3 = [0, 0, 1]

[0078] Preferably, an efficient word2vec technology, a kind of word embedding technology, is adopted.

[0079] In a possible design, the determination of whether the behavior of the computer program is abnormal includes:

[0080] - Inputting the vector representation and the statistical information of the second number N2 of monitoring units into a pre-trained judgment module; and

[0081] - The judgment module determines whether the behavior of the computer program is abnormal.

[0082] The judgment module can be pre-trained according to a suitable data set, and the data set can be obtained from various sources, such as from the APKs in the Google Play store. The judgment module can implement any type of machine learning.

[0083] In a possible design, the method further includes:

[0084] - At a moment t1 greater than the moment t0, obtaining a third number N3 of monitoring units called by the computer program from the first number N1 of monitoring units called by the computer program, where the first number N1 of monitoring units includes the second number N2 of monitoring units, and the third number N3 of monitoring units is pre-configured to be less than the first number N1 of monitoring units;

[0085] - Convert each of the third number N3 monitoring units into a vector representation of the third number N3 monitoring units, such that each of the third number N3 monitoring units has its corresponding vector representation;

[0086] - Obtain statistical information corresponding to the first number N1 monitoring units and the third number N3 monitoring units; and

[0087] - Determine whether the behavior of the computer program is abnormal based on the statistical information and the vector representation of the third number N3 monitoring units.

[0088] The method can be iterated such that at each subsequent time point, a set of closer monitoring units is obtained from a larger sequence of monitoring units. Various possible designs disclosed in the present invention can be adapted and applied at each subsequent time point. The time point t1 can be a subsequent time, such as the fixed time interval of 2 minutes as described above, i.e., the time points t0, t1, t1 + n, etc.

[0089] The second aspect of this embodiment provides a device for monitoring a computer program based on a computer program call of a monitoring unit, including:

[0090] - A statistical information acquisition module for obtaining statistical information corresponding to the first number N1 monitoring units;

[0091] - A vector representation conversion unit for converting the monitoring units in the second number N2 monitoring units into their corresponding vector representations, where the second number N2 monitoring units are a subset of the set of the first number N1 monitoring units including the computer program call, and the second number N2 monitoring units are preconfigured to be less than the first number N1 monitoring units; and

[0092] - A data storage unit for storing the statistical information obtained by the statistical information acquisition module.

[0093] The vector representation conversion unit is responsible for mapping the APIs and API ratios in the current sliding window into a numerical vector. The vector representation conversion unit can retrieve the APIs in the sliding window from memory, for example, it can retrieve from the data storage unit storing the APIs. Optionally, the data storage unit can also be preconfigured, or the data storage unit can also store additional information, such as a list of the most recent API sequences and / or a list of API ratios with a size equal to the size of the sliding window.

[0094] In a possible design, the device further includes:

[0095] An acquisition module, configured to acquire a second number N2 of monitoring units called by the computer program from a first number N1 of monitoring units called by the computer program; wherein

[0096] The acquisition module is configured to select the second number N2 of monitoring units from the first number N1 of monitoring units through a sliding window of a preconfigured size, and is configured to select monitoring units that are closer than other monitoring units in the set of the first number N1 of monitoring units from the second number N2 of monitoring units.

[0097] The acquisition module may acquire the second number N2 of monitoring units, or may acquire the second number N2 of monitoring units by other means and provide them to the statistical information acquisition module.

[0098] In a possible design, the apparatus further includes:

[0099] A detection module, configured to determine whether the behavior of the computer program is abnormal according to an input including a vector representation of the second number N2 of monitoring units and the statistical information.

[0100] In a possible design, the apparatus further includes:

[0101] A notification module, when the computer behavior is determined to be abnormal, is configured to send a notification message about the abnormal behavior. Similarly, if it is determined to be normal, a notification may be sent.

[0102] The apparatus or its related corresponding functional module can be used to execute any method or method step of the present invention.

[0103] In the present invention, the method may also be implemented in computer program instructions such as computer program code. The computer program code may be stored in a storage device connected to a processor. The storage device may be a memory device such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), an optical disc, a magneto-optical disc, or a semiconductor memory, etc. The storage device may also be a non-transitory computer-readable storage medium storing a computer program for controlling the operation of a mobile device.

[0104] The processor may read the computer program stored in the storage device, store the computer program in the RAM, and control the operation of the mobile device 10 according to the computer program read from the RAM. It should be noted that the computer program for controlling the operation of the mobile device may be pre-stored in the ROM, or may be downloaded from the network through the communication capability of the mobile device.

[0105] For the sake of clarity, any of the above embodiments can be combined with any one or more of the other above embodiments to create a new embodiment within the scope of the present invention. These features and other features will be described in detail below with reference to the accompanying drawings and the claims.

[0106] The embodiments of the present invention have at least the following advantages:

[0107] Low performance cost: By storing the APIs in a sliding window of a fixed size instead of storing the entire API sequence at runtime, the performance cost (such as storage) is reduced;

[0108] High detection accuracy: The vector representation in the embodiments of the present invention takes into account the statistical information of the APIs, rather than ignoring the early APIs that are not in the sliding window. The statistical information of the APIs is retained, that is, the information of those APIs that are not in the sliding window is not discarded. This improves the detection accuracy. Experiments conducted using the method of the present invention show that the true positive rate has increased by approximately 5% compared to traditional solutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0109] Figure 1 is an overview diagram of the method steps for monitoring a computer program provided by an embodiment of the present invention;

[0110] Figure 2 is an example block diagram of the configuration and workflow of functional units in a device provided by an embodiment of the present invention;

[0111] Figure 3 is a schematic diagram of the functional modules of a device provided on a mobile device according to an embodiment of the present invention;

[0112] Figure 4 is an example block diagram of the hardware configuration of a mobile device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0113] The technical solutions of the embodiments will be described below with reference to the accompanying drawings. It can be understood that the embodiments described below are not all embodiments, but only some embodiments related to the present invention. It should be noted that other embodiments obtained by those skilled in the art without creative efforts based on the embodiments described below are within the protection scope of the present invention.

[0114] The following will be combined with Figure 1 to introduce the method of monitoring a computer program on an operating system.

[0115] When a computer program runs, an API sequence is generated, and the API sequence continuously increases as the computer program runs, as shown by the arrow on the right side of the API sequence in Figure 1 As shown in Figure 1The intercepted API sequence 10 shown, i.e., the first number N1 of monitoring units, includes an API sequence at a specific moment t during detection. The API sequence includes 8 APIs, namely 1, 2, 3, 4, 5, 6, 7, and 8.

[0116] Among the 8 APIs in the intercepted API sequence, the 4 most recent APIs 5, 6, 7, and 8 are within a sliding window 20 of a fixed size of 4 APIs. These 4 most recent APIs are sampled to obtain the 4 most recent APIs. These 4 most recent APIs are the second number N2 of monitoring units in this embodiment and form a sequence in chronological order, with the larger numbered APIs being closer.

[0117] These 4 most recent APIs are converted 200 into vector representations 50, 60, and 70. These 4 most recent APIs have their respective corresponding vector representations. For simplicity, Figure 1 only 3 vector representations are shown. Vector representation 50 corresponds to API5, vector representation 60 corresponds to API6, and vector representation 70 corresponds to API7.

[0118] Statistical information of the APIs in the intercepted sequence is obtained and converted 300 into vector representations. Figure 1 Below the dotted line of each vector representation of the APIs are shown the compact vectors 55, 65, and 75. In this embodiment, the statistical information corresponds to the APIs in both the first number N1 and the second number N2. Although Figure 1 only 3 vector representations are shown, each API in these 4 most recent APIs has corresponding statistical information. The vector representation 55 of the statistical information corresponds to vector representation 55, the vector representation 65 of the statistical information corresponds to vector representation 60, and the vector representation 75 of the statistical information corresponds to vector representation 70.

[0119] This embodiment combines the vector representations of the statistical information with the corresponding vector representations.

[0120] Each vector in the combined vector includes the vector representation of the APIs in the second number N2 of monitoring units. The statistical information corresponding to the vector representation of the APIs is input 400 into a learning-based, pre-trained model 30. The model 30 determines 500 whether the behavior of the computer program is abnormal or normal, i.e., whether there is malware or it is benign.

[0121] Figure 2 Shows the functional modules corresponding to the Figure 1 actions therein and their interconnections. The functional modules can be implemented by hardware, software, or a combination of both. For example, the functional modules can be logical entities or library functions for performing specific functions.

[0122] This embodiment provides an acquisition module 2000, a data processing module 3000, a detection module 4000, and a notification module 5000. The data processing module includes the following sub-modules: a statistical information acquisition module 3005, a vector representation conversion unit 3015, and a data storage unit 3010. It can be understood that these sub-modules can also be provided outside the data processing module.

[0123] The acquisition module 2000 is responsible for acquiring the entire framework API (i.e., all APIs generated by the computer program during operation) or a specific API set and its information, including parameters, PID, and UID, etc. The acquisition module 2000 intercepts the API. The intercepted APIs form an API sequence 10. The acquisition module is also responsible for obtaining the APIs in the second number N2 of monitoring units through a selection function such as Figure 1 a fixed-size sliding window as shown, so as to form another sequence.

[0124] After an API is intercepted by the acquisition module, the statistical information acquisition module obtains the statistical information of the sequence or each sequence. In this embodiment, the statistical information is the API ratio. The API sequence can be expressed as S = {x i , i ∈ [1, n]}, where x i is an intercepted API. The sliding window can be expressed as W = {x j , j ∈ [m, n]}, where x j is an API belonging to the set S. The API ratio can be defined as: for the API x j ∈ W, its ratio is rate_x j = the size of the set {x i , x i = x j , i ∈ [1, j]} / the size of the set {x i , i ∈ [1, j]}.

[0125] Then, the data storage unit 3010 stores the API and the API ratio. The data storage unit 3010 stores a list of the most recent API sequences, whose size is equal to the size of the sliding window, and a list of API ratios. The data storage unit 3010 can use different types of memories, such as HashMap and shared memory, to store this data.

[0126] The vector representation conversion unit 3015 is responsible for mapping the APIs and API ratios in the current sliding window into numerical vectors.

[0127] Then, these vectors are input into the pre-trained model provided by the detection module 4000.

[0128] Finally, the notification module 5000 is responsible for notifying the end user of the details of whether malicious behavior has been detected.

[0129] In another embodiment (not shown), the statistical information obtained by the statistical information acquisition unit is not the API ratio, but the API count value disclosed in the present invention. Other functions in this embodiment remain unchanged and are implemented in the same manner as other embodiments.

[0130] In another embodiment (not shown), the statistical information obtained by the statistical information acquisition unit is not the API ratio, but the API ratio deviation relative to the average value of the API ratios in the sliding window. The statistical information can be defined as: for API x j ∈W, the deviation value is rate_x j -Σrate_x k / (n–m), where rate_x is the API ratio and x k ∈W. Several levels disclosed in the present invention can be defined to improve the detection effect. Other functions in this embodiment remain unchanged and are implemented in the same manner as other embodiments.

[0131] Figure 3 and Figure 4 relate to the hardware configuration of the mobile device.

[0132] In Figure 3 , the mobile device 6000 is configured with the data processing module 3000 disclosed in the present invention for monitoring the computer program 8000, and the computer program 8000 can run in the memory, such as RAM 7400 (as Figure 4 shown).

[0133] Optionally, in the mobile device 6000 as Figure 4 shown, the functions of the data processing module and other modules disclosed in the present invention are implemented by the programming code stored in the ROM 7200 and / or RAM 7400. When the programming code runs, the CPU 7000 performs corresponding actions. The memory and the CPU are connected by a bus.

[0134] The API can be intercepted by the suspension device 7600, and the suspension device is used to obtain information about the intercepted API.

[0135] The notification module can provide a warning output displayed on the screen of the mobile device 6000.

[0136] Furthermore, optionally, this embodiment provides a storage device 8000 for storing historical information of API sequences and general API ratios. The storage device can also store or obtain information about the training data set in the detection module.

[0137] The above description only discloses the preferred embodiments and is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that the above embodiments and all or part of other embodiments and modifications obtained according to the scope of the claims of the present invention are within the scope of the present invention.

Claims

1. A method for monitoring a computer program, characterized in that, comprising: - At time t0, obtain a second number N2 of monitoring units called by the computer program from a first number N1 of monitoring units called by the computer program, wherein the second number N2 of monitoring units is pre-configured to be less than the first number N1 of monitoring units, and one monitoring unit corresponds to one action phase executed by the computer program; the monitoring units in the second number N2 are the most recently obtained monitoring units compared with the monitoring units in the first number; - Obtain statistical information corresponding to the first number N1 of monitoring units; the statistical information corresponds to the remaining monitoring units in the first number N1 of monitoring units that are not in the second number N2 of monitoring units; - Convert the monitoring units in the second number N2 of monitoring units into a vector representation of the second number N2 of monitoring units, such that the monitoring units in the second number N2 of monitoring units have their respective corresponding vector representations; and - Determine whether the behavior of the computer program is abnormal according to the statistical information and the vector representation of the second number N2 of monitoring units; the determination of whether the behavior of the computer program is abnormal is further determined according to information of one or more sets of a fourth number N4 of monitoring units called at time t-1, wherein time t-1 is earlier than time t0, and the size of each set of the fourth number N4 is equal to the size of the second number N2 of monitoring units.

2. The method for monitoring a computer program according to claim 1, characterized in that, the obtaining of the second number N2 of monitoring units called by the computer program from the first number N1 of monitoring units called by the computer program further comprises: The monitoring units in the second number N2 of monitoring units and / or the monitoring units in the first number N1 of monitoring units are arranged in a sequence in the time order of being called by the computer program.

3. The method for monitoring a computer program according to claim 1 or 2, characterized in that, the obtaining of the second number N2 of monitoring units called by the computer program from the first number N1 of monitoring units called by the computer program at time t0 comprises: Obtain the second number N2 of monitoring units, wherein the end monitoring unit in the second number N2 of monitoring units is one of the most recent monitoring units in the set including the first number N1 of monitoring units, and the other end monitoring unit in the second number N2 of monitoring units is one of the most recent monitoring units compared with the monitoring units in the first number N1 of monitoring units.

4. The method for monitoring a computer program according to claim 1 or 2, characterized in that, the obtaining of the statistical information corresponding to the first number N1 of monitoring units comprises: Convert the obtained statistical information into the statistical information used in the determination step by representing the obtained statistical information as a vector representation of the statistical information.

5. The method for monitoring a computer program according to claim 1, characterized in that, Obtaining the second number N2 of monitoring units called by the computer program from the first number N1 of monitoring units called by the computer program includes: Obtaining the second number N2 of monitoring units from the first number N1 of monitoring units through a sliding window of a preconfigured size, such that at any specific time t, the monitoring units in the second number N2 of monitoring units include the most recent monitoring units in the time series of the first number N1 of monitoring units.

6. The method for monitoring a computer program according to claim 5, wherein, it includes: Obtaining the sliding window of the preconfigured size by subtracting the index value of the sliding window from the number of the first number N1 of monitoring units, wherein the index value corresponds to the earliest monitoring unit among the second number N2 of monitoring units.

7. The method for monitoring a computer program according to claim 1, wherein, Obtaining the second number N2 of monitoring units called by the computer program from the first number N1 of monitoring units called by the computer program includes: Storing the second number N2 of monitoring units.

8. The method for monitoring a computer program according to claim 1, wherein, Obtaining the second number N2 of monitoring units called by the computer program from the first number N1 of monitoring units called by the computer program includes: Not storing the monitoring units that are not in the second number N2 of monitoring units among the first number N1 of monitoring units.

9. The method for monitoring a computer program according to any one of claims 1, 2, 5 - 7, wherein, Obtaining the statistical information corresponding to the first number N1 of monitoring units includes: Obtaining the occurrence ratio of a monitoring unit according to the number of occurrences of the monitoring unit in the second number N2 of monitoring units relative to the number of occurrences of the monitoring unit in the first number N1 of monitoring units.

10. The method for monitoring a computer program according to any one of claims 1, 2, 5 - 7, wherein, Obtaining the statistical information corresponding to the first number N1 of monitoring units includes: Obtaining the count value of the monitoring unit.

11. The method for monitoring a computer program according to claim 5, wherein, Obtaining the statistical information corresponding to the first number N1 of monitoring units includes: Obtaining the deviation of the ratio of the monitoring unit in the second number N2 of monitoring units relative to the average ratio of the monitoring unit in the sliding window.

12. The method for monitoring a computer program according to any one of claims 1, 2, 5 - 7, 11, wherein, Judging whether the behavior of the computer program is abnormal is further judged according to the stored statistical information of the monitoring units.

13. The method for monitoring a computer program according to any one of claims 1, 2, 5 - 7, 11, wherein, The monitoring unit is an application programming interface or a system call.

14. The method for monitoring a computer program according to any one of claims 1, 2, 5 - 7, 11, characterized in that, the conversion to one - hot representation or the conversion based on word embedding technology.

15. The method for monitoring a computer program according to any one of claims 1, 2, 5 - 7, 11, characterized in that, judging whether the behavior of the computer program is abnormal includes: - inputting the vector representations of the second number N2 monitoring units and the statistical information into a pre - trained judgment module; and - the judgment module judges whether the behavior of the computer program is abnormal.

16. The method for monitoring a computer program according to any one of claims 1, 2, 5 - 7, 11, characterized in that, the method further includes: - at a time t1 greater than the time t0, obtaining a third number N3 of monitoring units called by the computer program from the first number N1 of monitoring units called by the computer program, wherein the first number N1 of monitoring units includes the second number N2 of monitoring units, and the third number N3 of monitoring units is pre - configured to be less than the first number N1 of monitoring units; - converting each of the third number N3 of monitoring units into a vector representation of the third number N3 of monitoring units, such that each of the third number N3 of monitoring units has its corresponding vector representation; - obtaining the statistical information corresponding to the first number N1 of monitoring units and the third number N3 of monitoring units; and - judging whether the behavior of the computer program is abnormal according to the statistical information and the vector representations of the third number N3 of monitoring units.

17. A device for monitoring a computer program based on monitoring units called by the computer program, characterized in that, comprising: a statistical information acquisition module for acquiring statistical information corresponding to a first number N1 of monitoring units; the statistical information corresponds to the remaining monitoring units among the first number N1 of monitoring units that are not in the second number N2 of monitoring units; a vector representation conversion unit for converting the monitoring units in the second number N2 of monitoring units into their respective corresponding vector representations, wherein the second number N2 of monitoring units is a subset of the set of the first number N1 of monitoring units called by the computer program, the second number N2 of monitoring units is pre - configured to be less than the first number N1 of monitoring units; the monitoring units in the second number N2 are the most recently acquired monitoring units compared with the monitoring units in the first number; and a data storage unit for storing the statistical information acquired by the statistical information acquisition module. A detection module, configured to determine whether the behavior of the computer program is abnormal according to an input including a vector representation of the second number N2 of monitoring units and the statistical information; the determination of whether the behavior of the computer program is abnormal is further made according to information of one or more sets of the fourth number N4 of monitoring units called at time t - 1, where time t - 1 is earlier than time t0, and the size of each set of the fourth number N4 is equal to the size of the second number N2 of monitoring units.

18. The apparatus according to claim 17, wherein, further comprising: an obtaining module, configured to obtain the second number N2 of monitoring units called by the computer program from the first number N1 of monitoring units called by the computer program; wherein the obtaining module is configured to select the second number N2 of monitoring units from the first number N1 of monitoring units through a sliding window with a preconfigured size, and to select monitoring units closer to other monitoring units in the set of the first number N1 of monitoring units among the second number N2 of monitoring units.

19. The apparatus according to claim 17 or 18, wherein, further comprising: a notification module, configured to send a notification message about abnormal behavior when the behavior of the computer program is determined to be abnormal.