Malicious Program Behavior Characterization Method Based on Morpheme Word Vector Model

Through the malicious program behavior representation method based on the morpheme word vector model, the problem of difficulty in understanding the behavior of malicious program with encryption and confusing malicious program behavior is solved, and the rapid perception and vectorized representation is achieved, which improves the efficiency of malicious program detection and analysis.

CN115587361BActive Publication Date: 2025-07-11CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211264125.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-14
Publication Date
2025-07-11
Estimated Expiration
2042-10-14

AI Technical Summary

Technical Problem

Traditional methods are difficult to quickly perceive and understand malicious program behaviors with encryption and confusion, and cannot efficiently mine, correlate analysis, and reason and judgment, which affects auxiliary decision-making for malicious file detection, analysis, disposal and prevention.

Method used

The malicious program behavior representation method based on the morpheme word vector model, by sorting function calls in order of execution, using the GSP sequence mining algorithm to de-redundancy and de-obfuscation, combining stemming extraction and word vector model to calculate the TF-IDF of the function, vectorized representation of malicious program behavior is realized.

Benefits of technology

It realizes rapid perception and understanding of malicious program behavior, provides an effective vectorized representation method, laying a solid foundation for the analysis, detection and classification of malicious programs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115587361B_ABST
    Figure CN115587361B_ABST
Patent Text Reader

Abstract

The present invention discloses a malicious program behavior characterization method based on a morpheme word vector model, which includes sorting and abstracting the captured malicious program function call information, extracting high-frequency sequences, and setting segmentation points; performing redundancy removal and de-obfuscation to obtain a new function sequence S'; obtaining the morphemes in the function name f to obtain the corresponding morpheme list M of the function, filling the marker Mask for the morpheme list with a non-maximum length, numbering the morphemes and function names respectively, and applying one-hot encoding to the numbers of the function names, morphemes, and placeholders to train the word vector model; calculating the feature vector of the function and calculating the TF-IDF of each function. It effectively solves the encryption and obfuscation problems existing in dynamic function calls and can quickly perceive and understand the behavior of malicious programs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of network security and representation learning, and particularly relates to a malicious program behavior representation method based on a morpheme word vector model. Background Art

[0002] A malicious program refers to a type of aggressive program. An attacker uses the malicious program to complete some specific functions, such as virus propagation, remote control, data theft, and destruction. During the code writing process, the attacker hides the code segments with attack intentions through means such as encryption and obfuscation, which increases the difficulty of detecting malicious code based on static code.

[0003] Applying natural language processing methods to vectorize and represent the function calls during the dynamic execution of malicious programs is an effective malicious program behavior representation method. However, when facing malicious program segments with encryption and obfuscation, traditional methods usually cannot quickly perceive and understand the malicious behaviors and typical features of samples, and it is difficult to efficiently mine, perform correlation analysis, and make reasoning judgments on effective information, and cannot provide auxiliary decision-making references for refined malicious file detection, analysis, disposal, response, and prevention. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide a malicious program behavior representation method based on a morpheme word vector model, which effectively solves the problems of encryption and obfuscation existing in dynamic function calls and can quickly perceive and understand the behaviors of malicious programs.

[0005] To solve the above technical problems, the technical solution adopted by the present invention is that a malicious program behavior representation method based on a morpheme word vector model includes the following steps:

[0006] Step 1: Sort the captured malicious program function call information in the order of instruction execution and abstract it into a function sequence S;

[0007] Step 2: Use the GSP sequence mining algorithm to extract the high-frequency sequences of the function sequence S in Step 1. The high-frequency sequences correspond to the hot code segments in the malicious program. Set breakpoints between the high-frequency subsequences to cut the function sequence S into several sequence segments s i ;

[0008] Step 3: According to the coincidence degree of adjacent sequence segments, remove redundancy and de-obfuscate the function sequence S to obtain a new function sequence S';

[0009] Step 4: For the functions in the function sequence S' in Step 3, use the stemming algorithm to obtain the morphemes in the function name f, and get the corresponding morpheme list M of the function. Fill the tag Mask for the morpheme list with a non-maximum length until the length of each morpheme list is the same. The Mask is used to fill the morpheme list with a length less than the maximum length.

[0010] Step 5: Number the morphemes and function names respectively, and apply one-hot encoding to the numbers of the function names, morphemes, and placeholders. In the encoding layer of the model, the final vector w of the function vec has the following relationship with the morpheme list M and the function name f:

[0011] where encode(·) is the numbering and one-hot encoding processing function, and m i is the i-th morpheme in the morpheme list M, n is the length of the morpheme list, W is the trainable weight matrix in the word vector model, and λ is the hyperparameter. The obtained w vec is a function vector that combines the morphological information of the function name and the window context information in the function call.

[0012] Step 6: Train the word vector model. The word vector model uses the CBOW and Bert models. Set the training task of the upper and lower sentences in the Bert model. The two models adjust the logic of the function vectorization in the encoding layer, that is, the vector encoding of the function is obtained from w in Step 5). vec Obtained;

[0013] Step 7: Use the trained morpheme word vector model to calculate the feature vector of the function according to formula (1), and calculate the TF-IDF of each function. The vector representation s of the function sequence vec is as follows:

[0014]

[0015] s vec That is, the final vectorized representation of the malicious program behavior. tf_idf(f) represents the TF-IDF value of the function f in the function sequence, and w vec (f,M) represents the feature vector of the function calculated according to formula (1), and len represents the number of functions in the current function sequence.

[0016] The beneficial effects of the present invention are as follows: First, sort the function call information of malicious programs according to the execution order and serialize the function information; then cut the function sequence into several sequence segments through a sequence mining algorithm, treeify each sequence segment, and identify and delete the common subtree parts of adjacent treeified sequence segments; also, use a stemming method to identify and extract the morphemes in the function names, and calculate the feature vectors of each function using a morpheme word vector model; finally, calculate the TF-IDF of each function, and weight the feature vectors according to the TF-IDF to obtain a vectorized representation of the behavior of malicious programs. The present invention provides an effective vectorized representation method for describing the dynamic behavior of malicious programs, realizes the application of the word vector model with morphemes in malicious program representation, effectively solves the encryption and obfuscation problems existing in dynamic function calls, and lays a solid foundation for the analysis, detection, and classification of downstream malicious programs. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0018] Figure 1 It is a schematic flowchart of the method for representing the behavior of malicious programs based on the morpheme word vector model of the present invention;

[0019] Figure 2 It is the network structure of the morpheme word vector;

[0020] Figure 3 It is a schematic diagram of the common subtree of adjacent subtrees. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0022] The method for representing the behavior of malicious programs based on the morpheme word vector model has a process as Figure 1 shown, including the following steps:

[0023] Step 1: Sort the captured function call information of malicious programs according to the instruction execution order and abstract it into a function sequence S;

[0024] Among them, the function call information contains multiple pieces of function call data. Each piece of function call data contains a calling function, the corresponding called function, and the relative execution time. The relative execution time can determine the order before and after each function call. Serialize the calling function and the called function according to the relative execution time. Among the same function call, the calling function is sorted before the called function. The function sequence is the function call data serialized according to the instruction execution order, marked as S = (a1, b1), (a2, b2), …(a n , b n ). Among them, (a i , b i ) represents the calling function and the called function in the i-th function call data, where i = 1, 2…n.

[0025] Step 2: Use the GSP sequence mining algorithm to extract the high-frequency sequences of the function sequence S in Step 1. The high-frequency sequences correspond to the hot code segments in the malicious program. Set split points between the high-frequency subsequences to cut the function sequence S into several sequence segments s i ;

[0026] Among them, identify the frequent sequences in the function call sequence in Step 1 based on the sequence mining algorithm and set split points. Specifically, set split points in any of the following situations when traversing S:

[0027] Situation 1): Set a split point before a1; that is, set a split point at the start of the sequence to ensure that there is at least one split point in the sequence S.

[0028] Situation 2): If the current function call group (a i , b i ) appears between the previous split point and the previous function group, then set a split point before a i ; The appearance of a repeated function call indicates that there may be redundancy between adjacent sequence segments.

[0029] Situation 3): If the current function call group (a i , b i ) and the previous function group between the previous split point are not in the same connected component at the graph structure level, then set a split point before a i ; Thus, to ensure that Step 3 can generate a normal hierarchical call tree during the treeification of the sequence, it is necessary to ensure that the sequence segment can be treeified into a connected component.

[0030] Situation 4): If the function call group (a i , b i ) from the previous split point to the current one is a frequent sequence identified by the sequence mining algorithm, then set a split point after b i .

[0031] There are generally multiple loop and conditional control blocks in the code written in high-level programming languages. The number of loops has a great impact on the behavior characterization of malicious programs. The high-frequency sequences in the dynamic function call sequence are manifestations of the loop blocks in the static code. Therefore, it is necessary to set breakpoints between high-frequency sequences.

[0032] Finally, according to the positions of the breakpoints, the function sequence is segmented into several sequence segments s i , and each sequence segment s i corresponds to a functional block in the malicious program.

[0033] Step 3: According to the degree of overlap between adjacent sequence segments, remove redundancy and confusion from the function sequence S to obtain a new function sequence S′;

[0034] Specifically, as Figure 3 shown, determine a unique spanning tree according to the function order and the inter-function call relationship of the sequence segment s i , and build a tree for each sequence segment, marked as t i , identify the common subtrees of adjacent t i according to the tree isomorphism algorithm, and delete the common subtree part. Pre-order traverse t i after deleting the common subtree to obtain a new sequence segment s i , and reorganize the sequence segments in the original order to obtain a new function sequence S′;

[0035] More specifically, it includes the following steps:

[0036] Step 3.1: Connect a i , b i in each function call group with an edge. Setting breakpoints in case 3) of step 2 can ensure that all nodes belong to the same connected component, and the function call characteristics can ensure that the connected component is a tree structure. After connecting a i , b i in all call groups with an edge, a unique spanning tree is established, that is, the call tree t i corresponding to each sequence segment is obtained. The call tree t i contains the behavior and structure information in the functional block;

[0037] Step 3.2: For adjacent call trees (t i , t i+1 ), i ∈ [0, p - 1], where p is the number of sequence segments, apply t i , t i+1 to the tree isomorphism algorithm to identify the common subtrees of adjacent call trees t i , t i+1 , and delete the common subtree part of t i , t i+1 . Pre-order traverse t′ i, obtain the traversal sequence as the updated sequence segment s'. i , the new sequence segment s'. i eliminates redundant and confusing information in the neighbor sequence segments;

[0038] Step 3.3. Recombine the updated sequence segment s' in Step 3.2 in the order of the sequence segments in S i to obtain a new de-confused function sequence S'.

[0039] Step 4. For the functions in the function sequence S' in Step 3, use the stemming algorithm morfessor to obtain the morphemes in the function name f, and obtain the corresponding morpheme list M of the function. Fill the marker Mask for the morpheme list that is not the maximum length. Mask is used to fill the morpheme list with a length less than the maximum length until the length of each morpheme list is the same;

[0040] Among them, use the stemming algorithm to obtain the morphemes in the function name. A morpheme is the smallest grammatical unit of a word. A word generally consists of several morphemes. Since the function name may contain more than one word, before using the stemming algorithm, first use the word segmentation algorithm to split the function name into several words, and then apply the stemming algorithm to each word to cut it into several morphemes, which are marked as the morpheme list M. Fill the marker Mask for the morpheme list that is not the maximum length until the length of each morpheme list is the same. In particular, no word segmentation and stemming processing are performed on meaningless function names.

[0041] Step 5. Number the morphemes and function names respectively, and apply one-hot encoding to the numbers of the function names, morphemes, and placeholders. In the encoding layer of the model, the final vector w of the function vec has the following relationship with the morpheme list M and the function name f:

[0042]

[0043] where encode(·) is the numbering and one-hot encoding processing function, m i is the i-th morpheme in the morpheme list M, n is the length of the morpheme list, W is the trainable weight matrix in the word vector model, λ is a hyperparameter, and the calculated w vec is a function vector that combines the morphological information of the function name and the window context information in the function call. In malicious file detection, built-in functions and user-defined functions with the same morphemes generally have similar behavioral semantics. Therefore, adding the morphological information of the function name helps in the detection of malicious files.

[0044] In Step 5, the numbering of the morphemes and function names is independent of each other to avoid conflicts in the numbering of the morphemes and function names. Under the principle of ensuring no conflicts, the numbering of the morphemes and function names is random. The final vector w vecIn the relationship with the morpheme list M and the function name f, the λ hyperparameter is used to adjust the weight between the morpheme encoding and the function name encoding.

[0045] Step 6: Train the word vector model. In this embodiment, the word vector model adopts two models for word embedding, namely CBOW (continuous bag of words) and Bert (bidirectional Encoder Representation from Transformers), without changing the network structure of the encoding layer in the model. Set the training task of the upper and lower sentences in the Bert model, and adjust the logic of function vectorization in the encoding layer for both models, that is, the vector encoding of the function is obtained from w in step 5). vec obtained.

[0046] Specifically, it includes:

[0047] Set the window size of the sequence, apply S′ to the word vector model, and apply the formula (1) in step 5 to the encoding layer of the word vector model to update the parameters of the word vector neural network model until the training end condition is met;

[0048] More specifically, it includes:

[0049] Step 6.1: Set the sliding window size of the sequence. The central word is set as the function in the middle of the sliding window, and the remaining functions in the window are used as context. Then input: the one-hot encoding of the morpheme and the function name number of each function in the window as the model input, that is, encode(m i ) and encode(f);

[0050] Concatenate all function sequences into a long sequence. The window slides forward 1 step each time to get a new sample until the sliding window moves to the end of the long sequence, obtaining a sample set, and not distinguishing between the training set and the test set.

[0051] Step 6.2, as Figure 2 shown, take the sample set as the input of the morpheme word vector model, and then encode: After multiplying encode(m i ) and encode(f) with the weight matrix W in the encoding layer, obtain the corresponding fixed-length vectors, sum and average the vectors corresponding to each non-Mask morpheme in the morpheme list, and finally add the λ-weighted function name vector.

[0052] Step 6.3: Calculate the loss: The difference relationship between the central word encoding and the context within the window is used as the model loss.

[0053] Calculate the loss function value according to step 6.3 and optimize the network parameters.

[0054] Step 7: Use the trained morpheme word vector model to calculate the feature vector of the function according to formula (1), and calculate the TF-IDF (term frequency–inverse document frequency) of each function. The vector representation s of the function sequence is as follows: vec as follows:

[0055]

[0056] s vec That is, the final vectorized representation of the malicious program behavior. tf_idf(f) represents the TF-IDF value of function f in the function sequence, w vec (f, M) represents the feature vector of the function calculated according to formula (1), and len represents the number of functions in the current function sequence.

[0057] Among them, TF-IDF is a statistical method for evaluating the importance of words. Calculating the TF-IDF value of functions in the sequence can evaluate the importance of functions to the function sequence, and thus assign reasonable weights to each function to vectorize the function sequence.

[0058] Example 1:

[0059] Step 1: Sort the background malicious program function call information in the four datasets DS1, DS2, DS3, and DS4 respectively. The sizes of the four datasets are 183, 264, 518, and 1061 respectively. The function call information of each background malicious program contains multiple execution records, and each execution record entry has a calling function name, a called function name, and an execution relative time, which is abstracted into a function sequence S.

[0060] Step 2: Use the GSP sequence mining algorithm to identify the high-frequency sequences in S. The shortest length and the lowest occurrence frequency of the high-frequency sequences are set to 3 and 4 respectively, and a split point is set in any of the following situations when traversing S: (1) Set a split point before a1; (2) If the current function call group (a i , b i ) appears in the previous split point to the previous function group, then set a split point before a i . (3) If the current function call group (a i , b i ) is not in the same connected component as the previous split point to the previous function group at the graph structure level, then set a split point before a i . (4) If the previous split point to the current function call group (a i , b i ) is a frequent sequence identified by the sequence mining algorithm, then set a split point before b iSet the splitting point later. Finally, split the sequence into several sequence segments s according to the position of the splitting point i 。

[0061] Step 3: According to the overlapping degree of adjacent sequence segments, perform redundancy removal and confusion removal on the function S to obtain a new function sequence S′, which specifically includes the following steps:

[0062] Step 3.1: Connect the calling function and the called function in each function call group with an edge to obtain the call tree t corresponding to each sequence segment i 。Since the connected components are guaranteed to be tree structures when setting the splitting point in Step 2, a unique spanning tree is established after connecting the calling function and the called function in the call group with an edge

[0063] Step 3.2: For adjacent call trees (t i , t i+1 ), i ∈ [0, p - 1], where p is the number of sequence segments, apply t i , t i+1 to the tree isomorphism algorithm to identify the common subtree of adjacent call trees t i , t i+1 , and delete the common subtree part of t i , t i+1 . Pre-order traverse t′ after deleting the common subtree i to obtain the traversal sequence as the updated sequence segment s′ i ;

[0064] Step 3.3: Recombine the updated sequence segments s′ in Step 3.2) according to the order of the sequence segments in S i to obtain the confusion-removed function sequence S′.

[0065] Step 4: For the functions in the S′ sequence in Step 3, use the stemming algorithm morfessor to obtain the morphemes in the function name f, and obtain the corresponding morpheme list M of the function. Fill the non-maximum length morpheme list with the marker Mask until the length of each morpheme list is the same.

[0066] Step 5: Number the function names in the sequence S′ starting from 1, number the morphemes starting from the maximum function number + 1, and number the Mask placeholders as 0. And apply one-hot encoding to the numbers of the function names, morphemes, and placeholders. In the encoding layer of the model, the vector w of the function vec has the following relationship with the morpheme list M and the function name f: λ is a hyperparameter, and in this embodiment, λ = 0.5.

[0067] Step 6: Train the word vector model. In this embodiment, the word vector model adopts two models for word embedding, namely CBOW and Bert, without changing the network structure of the encoding layer in the models. The logic of vectorizing functions in the encoding layer is adjusted for both models, that is, the vector encoding of the function is w in step 5) vec obtained. The specific steps are as follows:

[0068] Step 6.1): Set the sliding window size of the sequence to 5, set the central word as the third function in the sliding window, and the remaining functions in the window as the context. Only set the training task of predicting the central word in the CBOW and Bert models;

[0069] Step 6.2): The one-hot encoding of the central word and the corresponding morpheme list in each window, as well as the context and the corresponding morpheme list number, is used as a sample, that is, encode(m i ). Concatenate all function sequences into a long sequence. Each time the window slides forward 1 step to obtain a new sample until the sliding window moves to the end of the long sequence, a sample set is obtained, and the training set and the test set are not distinguished;

[0070] Step 6.3): Use the sample set in step 6.2) as the input of the morpheme word vector model. The model loss is calculated by the mean square error between the central word and the context of the window. The learning rate is taken as 0.001, and the number of training iterations is set to 1000. In the encoding layer, after encode(m i ) and encode(f) are multiplied by the weight matrix W in the encoding layer, the corresponding fixed-length vectors are obtained. The vectors corresponding to each non-Mask morpheme in the morpheme list are summed and averaged, and finally added to the function name vector weighted by λ to optimize the network parameters.

[0071] Step 7: Use the trained morpheme word vector model to calculate the feature vector of the function according to formula (1), and calculate the TF-IDF of each function. The vector representation s of the function sequence vec is as follows: s vec That is, the final vectorized representation of the malicious program behavior.

[0072] To prove the effectiveness of the present invention, the s obtained from step 7 vec will be used as the feature vector for Webshell classification and used as the input of the Kmeans clustering method. Set the parameter k of Kmeans to be the same as the number of categories in the real data, and use the 4 pieces of Webshell function call information obtained from the real world in step 1) as the experimental data.

[0073] The effects of Experimental Example 1 are as follows:

[0074]

[0075] The numbers in the table represent the accuracy of clustering, which is calculated by the Hungarian maximum matching algorithm.

[0076] To further prove the effectiveness of the present invention, the following two auxiliary embodiments are set up for ablation comparison experiments.

[0077] Embodiment 2:

[0078] This embodiment includes steps 1), 4), 5), 6), and 7) in Embodiment 1, does not include steps 2) and 3), and the parameters in steps 1), 4), 5), 6), and 7) remain the same as those in Embodiment 1. This embodiment is a simplified and de-confused version of Embodiment 1. The effects of Experimental Example 2 are as follows:

[0079]

[0080]

[0081] Embodiment 3:

[0082] This embodiment includes steps 1), 2), 3), 6), and 7) in Embodiment 1, does not include steps 4) and 5), and the parameters in steps 1), 4), 5), 6), and 7) remain the same as those in Embodiment 1. In step 6, the morpheme word vector model is weakened to a word vector model. This embodiment is a word vector version of Embodiment 1 without morphemes. The effects of Experimental Example 3 are as follows:

[0083]

[0084] Embodiment 1 has a higher clustering accuracy compared to Embodiments 2 and 3, proving the effectiveness of the present invention.

[0085] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the corresponding part of the method embodiment for the related content.

[0086] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A malicious program behavior characterization method based on a morpheme word vector model, characterized in that, It includes the following steps: Step 1: Sort the captured malicious program function call information according to the instruction execution order and abstract it into a function sequence S; Step 2: Use the GSP sequence mining algorithm to extract the high-frequency sequences of the function sequence S in Step 1. The high-frequency sequences correspond to the hot code segments in the malicious program. Set breakpoints between the high-frequency subsequences to cut the function sequence S into several sequence segments s i ; Step 3: According to the coincidence degree of adjacent sequence segments, remove redundancy and de-obfuscate the function sequence S to obtain a new function sequence S'; Step 4: For the functions in the function sequence S' in Step 3, use the stemming algorithm to obtain the morphemes in the function name f, and obtain the corresponding morpheme list M of the function. Fill the morpheme list with a non-maximum length with a marker Mask. Mask is used to fill the morpheme list with a length less than the maximum length until the length of each morpheme list is the same; Step 5: Number the morphemes and function names respectively, and apply one-hot encoding to the numbers of the function names, morphemes, and placeholders. In the encoding layer of the model, the final vector w of the function vec has the following relationship with the morpheme list M and the function name f: where encode(·) is the numbering and one-hot encoding processing function, m i is the i-th morpheme in the morpheme list M, n is the length of the morpheme list, W is the trainable weight matrix in the word vector model, and λ is the hyperparameter; the obtained w vec is a function vector that combines the morphological information of the function name and the window context information in the function call; Step 6: Train the word vector model. The word vector model adopts the CBOW and Bert models. Set the training task of the upper and lower sentences in the Bert model. Adjust the logic of function vectorization in the encoding layer for both models, that is, the vector encoding of the function is obtained from w in step 5) vec obtained; Step 7: Calculate the feature vector of the function according to formula (1) by using the trained morpheme word vector model, and calculate the TF-IDF of each function, and the vector representation s of the function sequence vec is as follows: s vec That is, the final malicious program behavior vectorized representation, tf_idf(f) represents the TF-IDF value of function f in the function sequence, w vec (f, M) represents the feature vector of the function calculated according to formula (1), and len represents the number of functions in the current function sequence.

2. The malicious program behavior characterization method based on the morpheme word vector model according to claim 1, characterized in that In the said step 1, the function call information includes multiple function call data. Each piece of function call data includes a calling function, the corresponding called function, and the relative execution time. The calling function and the called function are serialized according to the relative execution time. In the same function call, the calling function is sorted before the called function. The function sequence is the function call data serialized according to the instruction execution order, marked as S=(a1,b1),(a2,b2),…(a n ,b n ), where (a i ,b i ) represents the calling function and the called function in the i-th function call data, and i = 1, 2…n.

3. The malicious program behavior characterization method based on the morpheme word vector model according to claim 1, wherein In Step 2, based on the sequence mining algorithm, identify the frequent sequences in the function call sequence in Step 1 and set splitting points. When traversing S, set splitting points in any of the following situations: Situation 1): Set a splitting point before a1; that is, set a splitting point at the start of the sequence; Case 2): If the current function call group (a i , b i ) appears in the previous function group from the previous split point, then set a split point before a i . Case 3): If the current function call group (a i , b i ) and the previous function group from the previous split point are not in the same connected component at the graph structure level, then set a split point before a i ; Case 4): If the previous split point to the current function call group (a i , b i ) is a frequent sequence recognized by the sequence mining algorithm, then set a split point after b i ; Finally, the function sequence is segmented into several sequence segments s according to the positions of the segmentation points. i , and each sequence segment s i corresponds to a functional block in the malicious program.

4. The malicious program behavior characterization method based on the morpheme word vector model according to claim 1 or 3, characterized in that, Step 3 includes: Step 3.

1. Connect a i , b i in each function call group with an edge. After connecting all a i , b i in the call groups with edges, a unique spanning tree is established, that is, the call tree t corresponding to each sequence segment is obtained i . The call tree t i contains the behavior and structure information in the functional block; Step 3.

2. For adjacent call trees (t i , t i+1 ), where i ∈ [0, p - 1] and p is the number of sequence segments, apply t i , t i+1 to the tree isomorphism algorithm to identify the common subtrees of adjacent call trees t i , t i+1 , and delete the common subtree parts of t i , t i+1 . Perform a preorder traversal of t′ i after deleting the common subtrees to obtain a traversal sequence as the updated sequence segment s′ i . The new sequence segment s′ i eliminates redundant and confusing information in the neighbor sequence segments; Step 3.

3. Reorganize the updated sequence segment a′ in Step 3.2 according to the order of the sequence segments in S i to obtain a new function sequence S′ after de - obfuscation.

5. The malicious program behavior characterization method based on the morpheme word vector model according to claim 1 or 3, characterized in that, Step 6 includes: Step 6.1, set the sliding window size of the sequence, set the central word as the function in the middle of the sliding window, and use the remaining functions in the window as context, and then input: the one-hot encoding of the morpheme and function name number of each function in the window as the model input, that is, encode(m i ) and encode(f); Concatenate all function sequences into a long sequence. The window slides forward 1 step each time to obtain a new sample until the sliding window moves to the end of the long sequence, obtaining a sample set, and not distinguishing between the training set and the test set; Step 6.2, take the sample set as the input of the morpheme word vector model, and then encode: encode(m i ) and encode(f) are multiplied by the weight matrix W in the encoding layer to obtain corresponding fixed-length vectors. The vectors corresponding to each non-Mask morpheme in the morpheme list are summed and averaged, and finally added to the function name vector weighted by λ; Step 6.3, calculate the loss: The difference relationship between the central word encoding and the context within the window is used as the model loss to optimize the network parameters.

Citation Information

Patent Citations

  • Parameter learning device, similarity calculation device and method, and program

    JP2016197289A

  • Semantic representation model-based text classification method and apparatus, and computer device

    WO2021051503A1