A method for constructing a malicious program behavior rule set

The malicious program behavior rule set is constructed by ontology-based API call frequency description method and algorithm, which solves the single problem of scalability and feature analysis in malicious program detection, and realizes efficient and accurate detection of unknown samples.

CN115310086BActive Publication Date: 2025-07-29GUILIN UNIV OF ELECTRONIC TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210935122.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-05
Publication Date
2025-07-29
Estimated Expiration
2042-08-05

AI Technical Summary

Technical Problem

The prior art has problems such as scalability limitations and single feature analysis in malicious program detection, and traditional methods are difficult to effectively detect the variability and polymorphism of new malicious programs.

Method used

The ontology-based API call frequency formalized extension description method is used to build a fine-grained behavior rule set of malicious programs with association rule algorithm and decision tree algorithm, and a coarse-grained behavior rule set is used to build a behavior rule set that can detect program categories.

Benefits of technology

The inference detection rate of unknown program samples is improved. The generated rule set is not limited by scalability in terms of generation efficiency, has high detection accuracy and low false alarm rate, and can fully cover unknown samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115310086B_ABST
    Figure CN115310086B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for constructing a malicious program behavior rule set, which includes steps of describing malicious program behaviors using an ontology; extracting malicious program behavior features; constructing a fine-grained malicious program behavior rule set; constructing a coarse-grained program behavior rule set; and constructing a behavior rule set for detectable program categories. The construction method of the present invention proposes a formalized extended description method based on the API call frequency of malicious programs, and realizes a complete description of program behaviors based on the ontology. The behavior rule set for detectable program categories jointly constructed by the fine-grained malicious program behavior rule set and the coarse-grained program behavior rule set is complete. By inputting the inference rules into an inference engine and binding the inference engine to the ontology, the behavior rule set can be used to perform inference detection on unknown samples, mark the categories of unknown programs, and comprehensively cover the inference detection of unknown samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet security technology, and particularly relates to a method for constructing a behavior rule set for detecting program categories based on the ontology of malicious program behaviors. Background Art

[0002] In recent years, the number of malicious programs has shown explosive growth. Malicious programs pose a huge threat to network and information security. Therefore, the detection of malicious programs is becoming increasingly important. Using ontology formal description language can not only reason and analyze malicious programs, but also more comprehensively display the behavior characteristics of malicious programs.

[0003] With the rapid development of network information, new malicious programs often have variability, polymorphism and family characteristics. Traditional signature-based methods cannot detect them well. Existing malicious program detection methods mainly include static detection, dynamic detection, and a combination of static and dynamic methods. First, in terms of the behavior analysis of malicious programs in static detection, it directly analyzes the semantics, syntax, and call sequences of the code, and obtains the malicious behaviors of the program based on the static characteristics such as the binary files, assembly instructions, functions, and API calls of the program. Secondly, in terms of the behavior of malicious programs in dynamic detection, sandboxes and related software can be used to monitor the true intentions of malicious programs, and further analyze the API sequences and opcode sequences. Thereafter, related research combines static and dynamic features in a hybrid method for malicious program behavior classification.

[0004] Ontology is a form of knowledge expression that describes things in the real world and can be applied to the field of malicious program behavior knowledge, with a clearer and more explicit expression form. In the research on the domain ontology of malicious programs, most define the inference rule set manually, which is often limited by scalability, and the time for extracting features in defining inference rules is relatively long, unable to meet the application requirements for malicious detection and analysis of application programs; a few studies use algorithms to construct the inference rule set, but the feature analysis of malicious programs is relatively single. Summary of the Invention

[0005] The present invention provides a method for constructing a malicious program behavior rule set based on ontology, which improves the inference detection rate of unknown program samples, solves the problems of being limited by scalability and long time for extracting features in defining malicious program ontology rules manually, and at the same time solves the shortcoming of single feature analysis in constructing the inference rule set by traditional algorithms.

[0006] The technical solutions for achieving the objectives of the present invention are as follows:

[0007] A method for constructing a malicious program behavior rule set includes the following steps:

[0008] 1) Describe the behavior of malicious programs using an ontology:

[0009] Through dynamic analysis of malicious programs, monitor the running process of malicious programs in real time to obtain the attribute description of the behavior of malicious programs; obtain a malicious behavior report based on the operation of the Cuckoo Sandbox; since the formal description method of traditional triples will cause certain information omission, according to the malicious behavior report of unknown samples, this technical solution adopts a formal extended description method based on the API call frequency of malicious programs in the ontology software to enhance the description of program behavior by the ontology;

[0010] 2) Extract the behavior characteristics of malicious programs:

[0011] The malicious behavior report mainly includes the call sequence of API functions and the specific records of operation objects such as files, processes, networks, and registries in the operating system; each piece of information in the malicious program behavior report consists of an API function and an operation object, and each piece of information can completely describe an operation behavior of the malicious program. Extract each complete operation behavior as a characteristic attribute; at the same time, since the frequency information of API functions is an important characteristic of malicious programs, feature extraction is performed on the frequency information of API functions.

[0012] 3) Construct a fine-grained behavior rule set for malicious programs:

[0013] Use the association rule algorithm to mine the characteristic attributes of the behavior of malicious programs in the training samples, and select the decision tree algorithm to analyze the frequency information of API function calls of malicious programs. The results obtained jointly construct a fine-grained behavior rule set for malicious programs;

[0014] 4) Construct a coarse-grained behavior rule set for programs:

[0015] Use the Fisher linear discriminant algorithm to analyze the frequency information of API function calls of 34 API functions that can best describe the behavior of application programs selected. Combine the obtained linear discriminant function with the inference rules to obtain a coarse-grained behavior rule set for programs that can distinguish normal or malicious.

[0016] 5) Construct a behavior rule set that can detect program categories: jointly constructed by the fine-grained behavior rule set of malicious programs and the coarse-grained behavior rule set of programs.

[0017] Furthermore, the formal extended description method based on the API call frequency of malicious programs described in step 1) is defined as:

[0018]

[0019] On the basis that the subject, predicate, and object are consistent with the meanings of the malicious program behavior triple, add objectproperty as the connection property between the subject and the predicate; add the connection property hasvalue between the predicate and the object, indicating that the operation object value of the predicate is the object; add the hasnum property between the predicate and the predicate occurrence frequency predicate_num, indicating that the call frequency value of the predicate is the frequency predicate_num.

[0020] Furthermore, for the extraction of feature attributes described in step 2), the feature attributes are program behavior actions represented by API functions <predicate>The operation object of the API function <object>Composition, where malicious program behavior actions <predicate>And the operation object <object>They are connected by the symbol "*". The part before "*" represents the malicious program behavior action, and the part after "*" represents the operation object; at the end of the behavioral feature attributes of each training sample malicious program, add the malicious program category, thus constituting an item set of this malicious program. The item sets of multiple malicious programs are constructed into the entire dataset Dataset1;

[0021] For the frequency information of the API functions described in step 2), feature extraction is to count the frequency information of API functions for multiple training samples, forming a dataset Dataset2 of the frequency information of malicious program Windows API function calls.

[0022] Furthermore, for the construction of the fine-grained behavior rule set of malicious programs described in step 3), the specific method includes the following steps:

[0023] 3.1) Use the association rule algorithm to mine the feature attributes of the malicious program behavior of the training samples:

[0024] The association rule algorithm is an algorithm for classification by mining the associations between various data items in the dataset, that is, using the apriori method in the efficient_apriori module; the operation process of the apriori method is to first identify all frequent item sets in the dataset Dataset1, and then further construct relevant rules;

[0025] The set of all items in the dataset Dataset1 is I = {I1, I2, I3,......, I m}, through association rule mining, find the behavioral feature attribute item set A and the malicious program category label item set B, and satisfy the following conditions:

[0026] ① The item set A is often a multi-item set or a single-item set, and the item set B is often a single-item set;

[0027] ② Satisfy A ∩ B = Φ;

[0028] ③ Minimum support (min_support) condition:

[0029]

[0030] ④ Minimum confidence (min_confidence) condition:

[0031]

[0032] At this time, the result of the association rule can be obtained, in the form of: A → B;

[0033] ⑤ Filter the result of the association rule:

[0034] Since item set B is a single-item set in set I, item set B may be either a behavioral characteristic attribute item of a malicious program or a malicious program category item. Therefore, the results of the association rules are filtered, and only the rules with item set B being a malicious program category item are retained;

[0035] 3.2) Relationship mapping between the results of association rules and the rules:

[0036] The results obtained by the association rules are themselves a set of rules. The behavioral characteristic attribute item set A and the malicious program category label item set B are associated by the symbol "→". The behavioral characteristic attribute item set A is the prerequisite for the association, and the malicious program category label item set B is the result of the association;

[0037] 3.3) Use the decision tree algorithm to analyze the frequency information of malicious program API function calls:

[0038] The classification decision tree model is a tree structure that uses descriptions to distinguish the categories of samples. The decision tree is composed of nodes and directed edges. The nodes are divided into two categories: internal nodes and leaf nodes. Each internal node represents the discrimination of the characteristic attributes of the frequency information of malicious program API function calls, and the leaf nodes represent that the samples are determined to be a certain category after being discriminated by multiple internal nodes; call the DecisionTreeClassifier classifier in the sklearn.tree module, and use the DecisionTree() function to use information entropy or Gini coefficient as the basis for node feature selection. To avoid overfitting on the training data set, the adjustment of the decision tree depth can be increased;

[0039] 3.4) Relationship mapping between the results of the decision tree algorithm and the rules:

[0040] The mining rules can be naturally converted from the model established by the decision tree. Each rule created corresponds to each path from the root node to the leaf node of the decision tree; the conditions of the rules also correspond to the characteristic attribute discrimination of the internal nodes of each path, and the results of the rules correspond to the categories of the leaf nodes. Finally, a set of rules equivalent to the decision tree can be obtained;

[0041] 3.5) Use the Merge() function to merge the rule sets obtained by the association rule algorithm and the decision tree algorithm according to the malicious program category, and output the fine-grained rule set of malicious program behaviors;

[0042] SWRL presents rules in a semantic form and can be directly described using ontology. The general form of SWRL rules can be expressed as:

[0043] Rule:M(?x)∧N(?x,?y)→P(?x),

[0044] Among them, the variables x and y represent instances of OWL and data values of OWL respectively, M and P represent class descriptions, and the class attribute is represented by N; if the left part of the rule holds as a premise, the conclusion on the right is true;

[0045] 3.6) Use the SWRL rule language to perform the transformation of formal semantics on the rule set:

[0046] Map the preconditions and results in the rule set to concepts and attributes in the OWL ontology language, so as to associate the ontology SWRL rules with the ontology, and combine the inference engine to detect and classify unknown program samples.

[0047] Furthermore, for step 4) mentioned in step 4), the method for constructing the program coarse-grained behavior rule set specifically includes the following steps:

[0048] 4.1) Select 34 API functions that can best describe the behavior of the application program, which are: ntCreateFile, CreateDirectory, CopyFile, FindFirstFile / FindNextFile, GetFileAttribute, GetFileSize, GetFileType, ntOpenFile, ntReadFile, ntWriteFile, DeleteFile, ntCreateProcess, ntOpenProcess, ntTerminateProcess, regOpenKey, ntOpenKey, regQueryInfoKey, regQueryValue, regSetValue, regCloseKey, CreateThread, ntOpenThread, ntResumeThread, GetSystemInfo, GetSystemDirectory, GetSystemMetrics, GetSystemTimeAsFileTime, GetSystemWindowsDirectory, ntCreateMutant, ntOpenMutant, ntClose, ntDelayExecution, GetCursorpos, wsaStartUp;

[0049] 4.2) Normalize the frequency information of the above 34 API function calls:

[0050] Use the normalization function f(x) = round(10 * lgx) to process the frequency information;

[0051] 4.3) Import the data into the Fisher linear discriminant algorithm to obtain the optimal linear discriminant function:

[0052] Fisher linear discriminant uses the projection method to project the m-dimensional space samples onto a straight line in one-dimensional space At the same time, based on one-way analysis of variance, following the principle of maximizing the between-group mean square deviation and minimizing the within-group mean square deviation, the samples in the one-dimensional space direction are best dispersed, and the coefficients for obtaining the linear discriminant function are determined to create a linear discriminant equation. The specific steps are as follows:

[0053] ① Call the getAvg() function to calculate the mean vectors of the normal and malicious training samples respectively;

[0054] ② Call the getSi() function to calculate the within-class scatter matrix S of the normal and malicious training samples respectively i and the total between-class scatter matrix S between the normal and malicious training samples w = S0 + S1;

[0055] ③ Calculate the between-class scatter matrix S of the samples b ;

[0056] ④ Call the getW() function to calculate the discriminant index weight vector ω * ;

[0057] ⑤ Project all the samples in the training set and obtain the projection points through the dot() function;

[0058] ⑥ Calculate the means of the projected normal and malicious training samples respectively;

[0059] ⑦ Calculate the classification threshold;

[0060] ⑧ Create the optimal linear discriminant function with the classification threshold and the discriminant index weight vector;

[0061] 4.4) Combining the linear discriminant function with the inference rules can obtain the program coarse-grained behavior rule set.

[0062] Beneficial effects produced by the construction method of the present invention:

[0063] 1. The present invention proposes a formal extended description method based on the API call frequency of malicious programs to achieve a complete description of program behavior based on ontology.

[0064] 2. The behavior rule set for detecting program categories generated by the present invention has a significant improvement in generation efficiency, is not limited by scalability, and has a high accuracy in inferring and detecting categories for unknown samples and a low false alarm rate.

[0065] 3. The behavior rule set for detectable program categories jointly constructed by the fine-grained behavior rule set of malicious programs and the coarse-grained behavior rule set of programs is complete and can comprehensively cover the inference detection of unknown samples. Description of the Drawings

[0066] Figure 1 It is a flowchart of the construction method of the malicious program behavior rule set in the embodiment;

[0067] Figure 2 It is an example of a partial behavior analysis report of a certain malicious program in the embodiment;

[0068] Figure 3 It is an example diagram of the formalized extension description method based on the API call frequency of malicious programs in the embodiment. Embodiment

[0069] In order to make the purpose, technical solutions and advantages of the implementation of the present invention clearer, the technical solutions of the present invention will be elaborated in detail through specific embodiments below.

[0070] Embodiment:

[0071] A method for constructing a malicious program behavior rule set, as Figure 1 shown, includes the following steps:

[0072] 1) Use ontology to describe the behavior of malicious programs;

[0073] 2) Extract the behavior characteristics of malicious programs;

[0074] 3) Construct a fine-grained behavior rule set for malicious programs;

[0075] 4) Construct a coarse-grained behavior rule set for programs;

[0076] 5) Construct a behavior rule set for detectable program categories.

[0077] In step 1), using ontology to describe the behavior of malicious programs, the specific method is:

[0078] Through the dynamic analysis of malicious programs, the running process of malicious programs is monitored in real time to obtain the attribute description of malicious program behavior; a malicious behavior report is obtained based on the operation of the Cuckoo Sandbox; according to the malicious behavior report of unknown samples, in the ontology software, a formalized extension description method based on the API call frequency of malicious programs is adopted to enhance the description of program behavior by the ontology;

[0079] The formalized extension description method based on the API call frequency of malicious programs is defined as:

[0080]

[0081] When formally describing the behavior of malicious programs in the traditional form, a triple <subject; predicate; object> is used, where subject represents the type of malicious program or a specific instance individual within the type, predicate represents the action, i.e., the behavior operated by the malicious program, and object represents the type or instance individual of the operating system object;

[0082] This embodiment is based on a formal extended description method of the API call frequency of malicious programs, as Figure 3 shown: On the basis that subject, predicate, and object are consistent with the meaning of the malicious program behavior triple, objectproperty is added as the connection property between the subject and the predicate; a connection property of hasvalue is added between the predicate and the object, indicating that the operation object value of the predicate is the object; a hasnum property is added between the predicate and the predicate occurrence frequency predicate_num, indicating that the call frequency value of the predicate is the frequency predicate_num.

[0083] Step 2) Malicious program behavior feature extraction, including:

[0084] 2.1) Extracting feature attributes: Each complete operation behavior obtained from the malicious behavior report of the training sample is extracted as a feature attribute, and the feature attribute is the program behavior action represented by the API function <predicate>The operation object of the API function <object>Composition, to Figure 2 Taking the example of the partial behavior analysis report of a certain malicious program in Figure 2 as an example, it is in the form of: ntCreateFile*C:\WINDOWS\system32\sample.exe. Among them, the malicious program behavior <predicate>and the operating object <object>They are connected by the symbol "*". The part before "*" represents the malicious program behavior action, and the part after "*" represents the operation object. Taking Figure 2 as an example of the behavior report, the behavior characteristic attributes of sample.exe are expressed as:

[0085] Sample.exe: ntCreateFile*C:\WINDOWS\system32\sample.exe, ntReadFile*C:\sample.exe, ntWriteFile*C:\WINDOWS\sample.exe, ntWriteFile*C:\WINDOWS\system32\kernel32.dll, regOpenKey*HKLM\Software\Microsoft\Windows NT\CurrentVersion\Windows……

[0086] Add the malicious program category at the end of the behavior characteristic attributes of each training sample malicious program to form an item set of this malicious program. The item sets of multiple malicious programs are constructed into the entire dataset Dataset1.

[0087] 2.2) Feature extraction is performed on the frequency information of API functions: Taking Figure 2 as an example of the behavior analysis report, the frequency information of each function call in the API function column is counted, and finally it is stored in the form of a dictionary as: Sample.exe: ’ntCreateFile’: 1, ’ntReadFile’: 1, ’ntWriteFile’: 3, ’regOpenKey’: 3, ’regSetValue’: 1…… Similarly, the frequency information of API functions of multiple training samples is counted to form the frequency information dataset Dataset2 of malicious program Windows API function calls.

[0088] Step 3) Construct a fine-grained behavior rule set for malicious programs:

[0089] The fine-grained behavior rule set for malicious programs is applicable to detecting whether the unknown sample has the corresponding characteristic attributes of the malicious program categories in the dataset or the frequency information of API function calls, so as to classify the malicious program categories of the unknown sample;

[0090] 3.1) Mining the characteristic attributes of malicious program behaviors in the training samples using the association rule algorithm: The association rule algorithm is an algorithm for classification by mining the associations between various data items in a dataset, that is, using the apriori method in the efficient_apriori module; the operation process of the apriori method is to first identify all frequent item sets in the dataset Dataset1, and then further construct relevant rules; the set of all items in the dataset Dataset1 is I = {I1, I2, I3,......, I m}, through association rule mining, find the behavioral characteristic attribute item set A and the malicious program category label item set B, and satisfy the following conditions:

[0091] ① The item set A is often a multi-item set or a single-item set, and the item set B is often a single-item set;

[0092] ② Satisfy A ∩ B = Φ;

[0093] ③ Minimum support (min_support) condition:

[0094]

[0095] ④ Minimum confidence (min_confidence) condition:

[0096]

[0097] At this time, the results of the association rules can be obtained, in the form of: A → B;

[0098] ⑤ Filter the results of the association rules:

[0099] Since the item set B is a single-item set in the set I, the item set B may be either a behavioral characteristic attribute item of a malicious program or a malicious program category item. Therefore, filter the results of the association rules and only retain the rules where the item set B is a malicious program category item;

[0100] 3.2) Relationship mapping between the results of association rules and rules: The results obtained by the association rules are themselves a set of rules. The behavioral characteristic attribute item set A and the malicious program category label item set B are associated by the symbol "→". The behavioral characteristic attribute item set A is the prerequisite for the association, and the malicious program category label item set B is the result of the association;

[0101] 3.3) Analyze the frequency information of malicious program API function calls using the decision tree algorithm: The classification decision tree model is a tree structure that uses descriptions to distinguish different types of samples; a decision tree is composed of nodes and directed edges. Nodes are divided into two categories: internal nodes and leaf nodes. Each internal node represents the discrimination of the characteristic attributes of the frequency information of malicious program API function calls, and the leaf node represents that the sample is determined to be a certain category after being discriminated by multiple internal nodes; call the DecisionTreeClassifier classifier in the sklearn.tree module, and use the DecisionTree() function to use information entropy or Gini coefficient as the basis for node feature selection. To avoid overfitting on the training data set, the adjustment of the decision tree depth can be increased;

[0102] 3.4) Relationship mapping between the results of the decision tree algorithm and the rules: Mining rules can be naturally converted from the model established by the decision tree. Each created rule corresponds to each path from the root node to the leaf node of the decision tree; the conditions of the rules also correspond to the discrimination of the characteristic attributes of the internal nodes of each path, and the results of the rules correspond to the categories of the leaf nodes. Finally, a rule set equivalent to the decision tree can be obtained;

[0103] 3.5) Use the Merge() function to merge the rule sets obtained by the association rule algorithm and the decision tree algorithm according to the malicious program category, and output the fine-grained rule set of malicious program behaviors;

[0104] SWRL (Semantic Web Rule Language) presents rules in a semantic form and can directly use ontological descriptions. The general form of SWRL rules can be expressed as:

[0105] Rule: M(?x) ∧ N(?x,?y) → P(?x),

[0106] Among them, the variables x and y represent OWL instances and OWL data values respectively, M and P represent class descriptions, and the class attribute is represented by N; if the left part of the rule holds as a premise, the conclusion on the right is true.

[0107] 3.6) Use the SWRL rule language to perform formal semantic transformation on the rule set: Map the preconditions and results in the rule set to concepts and attributes in the OWL ontology language to associate the ontology SWRL rules with the ontology, and combine the inference engine to detect and classify unknown program samples.

[0108] Step 4) Construct the coarse-grained behavior rule set of the program

[0109] If the unknown sample does not have the corresponding features of the malicious program categories in the dataset, inference detection cannot be performed. To improve the comprehensive inference detection of unknown samples, the Fisher linear discriminant algorithm is used to coarsely classify programs into malicious programs and normal programs, and a coarse-grained behavior rule set of programs is constructed. When detecting an unknown sample, if it is not a malicious program category in the rule set, basic malicious or normal inference can also be performed.

[0110] 4.1) The 34 API functions that best describe the behavior of the application are selected, which are: ntCreateFile, CreateDirectory, CopyFile, FindFirstFile / FindNextFile, GetFileAttribute, GetFileSize, GetFileType, ntOpenFile, ntReadFile, ntWriteFile, DeleteFile, ntCreateProcess, ntOpenProcess, ntTerminateProcess, regOpenKey, ntOpenKey, regQueryInfoKey, regQueryValue, regSetValue, regCloseKey, CreateThread, ntOpenThread, ntResumeThread, GetSystemInfo, GetSystemDirectory, GetSystemMetrics, GetSystemTimeAsFileTime, GetSystemWindowsDirectory, ntCreateMutant, ntOpenMutant, ntClose, ntDelayExecution, GetCursorpos, wsaStartUp...

[0111] 4.2) Normalize the frequency information of the above 34 API function calls: To solve the deviation influence of the detection result caused by too large or too small data values, the normalization function f(x) = round(10 * lgx) is used to process the frequency information.

[0112] 4.3) Import the data into the Fisher linear discriminant algorithm to obtain the best linear discriminant function: The Fisher linear discriminant uses the projection method to project the samples in the m-dimensional space onto a straight line in the one-dimensional space. At the same time, based on one-way analysis of variance, according to the principle of the largest between-group mean square deviation and the smallest within-group mean square deviation, the samples in the one-dimensional space direction are best dispersed, and the coefficients of the linear discriminant function are determined to create a linear discriminant equation. The specific steps are as follows:

[0113] ① Call the getAvg() function to calculate the mean vectors of normal and malicious training samples respectively.

[0114] ② Call the getSi() function to calculate the intra-class scatter matrices S of normal and malicious training samples i and the total inter-class scatter matrix S between normal and malicious training samples w = S0 + S1;

[0115] ③ Calculate the inter-class scatter matrix S of the samples b ;

[0116] ④ Call the getW() function to calculate the discriminant index weight vector ω * ;

[0117] ⑤ Project all samples in the training set and obtain the projection points through the dot() function;

[0118] ⑥ Calculate the means of the two types of samples after projection respectively;

[0119] ⑦ Calculate the classification threshold;

[0120] ⑧ Create the optimal linear discriminant function with the classification threshold and the discriminant index weight vector;

[0121] 4.4) Combining the linear discriminant function with the inference rules can obtain the program coarse-grained behavior rule set

[0122] Step 5) Construct the behavior rule set that can detect program categories: jointly constructed by the malicious program fine-grained behavior rule set and the program coarse-grained behavior rule set.

[0123] The construction method of the present invention, the behavior rule set that can detect program categories jointly constructed by the malicious program fine-grained behavior rule set and the program coarse-grained behavior rule set is complete. Input the inference rules into the inference engine, bind the inference engine with the ontology, and then the behavior rule set can be used to perform inference detection on unknown samples, mark the categories of unknown programs, and the inference detection of unknown samples can be comprehensively covered.< / object> < / predicate> < / object> < / predicate> < / object> < / predicate> < / object> < / predicate>

Claims

1. A method for constructing a malicious program behavior rule set, characterized in that, It includes the following steps: 1) Describe the behavior of malicious programs using an ontology: Through dynamic analysis of malicious programs, monitor the running process of malicious programs in real time to obtain the attribute description of malicious program behavior; obtain a malicious program behavior report based on the operation of the Cuckoo Sandbox. According to the malicious behavior report of unknown samples, in the ontology software, adopt a formal extension description method based on the API call frequency of malicious programs to enhance the description of program behavior by the ontology. 2) Extract malicious program behavior features: Each piece of information in the malicious program behavior report consists of an API function and an operation object. Each piece of information can completely describe an operation behavior of the malicious program. Extract each complete operation behavior as a feature attribute. At the same time, since the frequency information of API functions is an important feature of malicious programs, feature extraction is performed on the frequency information of API functions. 3) Construct a fine-grained behavior rule set for malicious programs: Use the association rule algorithm to mine the feature attributes of the malicious program behavior of the training samples, and select the decision tree algorithm to analyze the frequency information of the malicious program API function calls. The obtained results jointly construct a fine-grained behavior rule set for malicious programs. 4) Construct a coarse-grained behavior rule set for programs: Use the Fisher linear discriminant algorithm to analyze the frequency information of the API function calls of the 34 most representative API functions that can describe the behavior of application programs. Combine the obtained linear discriminant function with the inference rule to obtain a coarse-grained behavior rule set for discriminating normal or malicious programs. 5) Construct a behavior rule set for detecting program categories: It is jointly constructed by the fine-grained behavior rule set of malicious programs and the coarse-grained behavior rule set of programs.

2. The method for constructing a malicious program behavior rule set according to claim 1, characterized in that: The formal extension description method based on the API call frequency of malicious programs described in step 1) is defined as: On the basis that the subject, predicate, and object are consistent with the meaning of the malicious program behavior triple, add objectproperty as the connection property between the subject and the predicate; add the connection property hasvalue between the predicate and the object, indicating that the operation object value of the predicate is the object; add the hasnum property between the predicate and the predicate occurrence frequency predicate_num, indicating that the call frequency value of the predicate is the frequency predicate_num.

3. The method for constructing a malicious program behavior rule set according to claim 2, characterized in that: The feature attribute extraction described in step 2), the feature attribute Program behavior actions represented by API functions <predicate>The operation object of the API function <object>Composition, where malicious program behavior actions <predicate>And the operation object <object>They are connected by the symbol "*". The part before "*" represents the malicious program behavior action, and the part after "*" represents the operation object; add the malicious program category at the end of the behavioral feature attributes of each training sample malicious program, thus forming an item set of this malicious program, and the item sets of multiple malicious programs are constructed into the entire dataset Dataset1; The feature extraction of the frequency information of the API functions described in step 2) is to statistically analyze the frequency information of the API functions of multiple training samples to form a frequency information dataset Dataset2 of malicious program Windows API function calls.

4. The method for constructing a malicious program behavior rule set according to claim 3, wherein: The specific method for constructing the fine-grained behavior rule set of malicious programs in step 3) includes the following steps: 3.1) Use the association rule algorithm to mine the characteristic attributes of the malicious program behavior in the training samples: The association rule algorithm is an algorithm for classification by mining the associations between various data items in the data set, that is, using the apriori method in the efficient_apriori module; the operation process of the apriori method is to first identify all frequent item sets in the data set Dataset1, and then further construct relevant rules; The set of all items in Dataset1 is \(I = \{I1, I2, I3,\cdots, I\}\). m Through association rule mining, the behavioral feature attribute item set \(A\) and the malicious program category label item set \(B\) are found and satisfy the following conditions: ① The item set A is often a multi-item set or a single-item set, and the item set B is often a single-item set; ② Satisfy A ∩ B = Φ; ③ Minimum support min_support condition: ④ Minimum confidence min_confidence condition: At this time, the result of the association rule can be obtained, in the form of: A → B; ⑤ Filter the result of the association rule: Since the item set B is a single-item set in the set I, the item set B may be either a behavior characteristic attribute item of the malicious program or a malicious program category item. Therefore, filter the result of the association rule and only retain the rules where the item set B is a malicious program category item; 3.2) Relationship mapping between the association rule result and the rule: The result obtained by the association rule is itself a set of rules. The behavior characteristic attribute item set A and the malicious program category label item set B are associated by the symbol "→". The behavior characteristic attribute item set A is the prerequisite for the association, and the malicious program category label item set B is the result of the association; 3.3) Use the decision tree algorithm to analyze the frequency information of malicious program API function calls: The classification decision tree model is a tree structure that uses descriptions to distinguish samples into categories. The decision tree is composed of nodes and directed edges. The nodes are divided into two categories: internal nodes and leaf nodes. Each internal node represents the discrimination of the characteristic attributes of the frequency information of malicious program API function calls, and the leaf node represents that the sample is determined to be a certain category after being discriminated by multiple internal nodes; call the DecisionTreeClassifier classifier in the sklearn.tree module, and use the DecisionTree() function to use information entropy or Gini coefficient as the basis for node feature selection. To avoid overfitting on the training data set, the adjustment of the decision tree depth can be increased; 3.4) Relationship mapping between the decision tree algorithm result and the rule: The mining rules can be naturally converted from the model established by the decision tree. Each rule created corresponds to each path from the root node to the leaf node of the decision tree; the conditions of the rules also correspond to the characteristic attribute discriminations of the internal nodes of each path, and the results of the rules correspond to the categories of the leaf nodes. Finally, a rule set equivalent to the decision tree can be obtained; 3.5) Use the Merge() function to merge the rule sets obtained by the two methods, the association rule algorithm and the decision tree algorithm, according to the malicious program category, and output the fine-grained rule set of malicious program behavior; SWRL presents rules in a semantic form and can directly use ontology descriptions. The general form of an SWRL rule can be expressed as: Rule: M(?x) ∧ N(?x,?y) → P(?x), where the variables x and y represent OWL instances and OWL data values respectively, M and P represent class descriptions, and the class attribute is represented by N; if the left part of the rule holds as a premise, then the conclusion on the right is true; 3.6) Use the SWRL rule language to perform formal semantic transformation on the rule set: Map the preconditions and results in the rule set to concepts and properties in the OWL ontology language to associate the ontology SWRL rules with the ontology, and combine an inference engine to detect and classify unknown program samples.

5. The construction method of the malicious program behavior rule set according to claim 4, characterized in that: The specific method for constructing the coarse-grained behavior rule set of the program in step 4) includes the following steps: 4.1) The 34 API functions that can best describe the behavior of the application are selected, which are: ntCreateFile, CreateDirectory, CopyFile, FindFirstFile / FindNextFile, GetFileAttribute, GetFileSize, GetFileType, ntOpenFile, ntReadFile, ntWriteFile, DeleteFile, ntCreateProcess, ntOpenProcess, ntTerminateProcess, regOpenKey, ntOpenKey, regQueryInfoKey, regQueryValue, regSetValue, regCloseKey, CreateThread, ntOpenThread, ntResumeThread, GetSystemInfo, GetSystemDirectory, GetSystemMetrics, GetSystemTimeAsFileTime, GetSystemWindowsDirectory, ntCreateMutant, ntOpenMutant, ntClose, ntDelayExecution, GetCursorpos, wsaStartUp; 4.2) Normalize the frequency information of the above 34 API function calls: Use the normalization function f(x) = round(10 * lgx) to process the frequency information; 4.3) Import the data into the Fisher linear discriminant algorithm to obtain the optimal linear discriminant function: Fisher linear discriminant uses the projection method to project the samples in the m-dimensional space onto a straight line in the one-dimensional space At the same time, based on one-way analysis of variance, according to the principle of maximizing the between-group mean square deviation and minimizing the within-group mean square deviation, the samples in the one-dimensional space direction are best dispersed, the coefficients of the linear discriminant function are determined, and the linear discriminant equation is created. The specific steps are as follows: ① Call the getAvg() function to calculate the mean vectors of the normal and malicious training samples respectively; ② Call the getSi() function to calculate the within-class scatter matrices S of normal and malicious training samples respectively i and the total between-class scatter matrix S between normal and malicious training samples w = S0 + S1; ③ Calculate the between-class scatter matrix S of the samples b ; ④ Call the getW() function to calculate the discriminant index weight vector ω * ; ⑤ Project all samples in the training set and obtain the projection points through the dot() function; ⑥ Calculate the means of the projected normal and malicious training samples respectively; ⑦ Calculate the classification threshold; ⑧ Create the optimal linear discriminant function with the classification threshold and the discriminant index weight vector; 4.4) Combining the linear discriminant function with the inference rules can obtain the program coarse-grained behavior rule set. < / object> < / predicate> < / object> < / predicate>

Citation Information

Patent Citations

  • Android malicious application detection method and system integrating frequent item set and random forest algorithm

    CN109753800A

  • Android malicious application detection method

    CN112395615A