Malicious software detection and classification method based on multi-dimensional feature fusion

By establishing a system call-traffic map and using multimodal detection methods, the error judgment problem caused by single-modal detection in the prior art is solved, and more accurate malware detection is achieved.

CN120068074AActive Publication Date: 2025-05-30GUANGDONG UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510216070.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

Existing malware detection methods mostly use single-modal detection, which can easily lead to missed judgments and ignore the dependence of system call order and noise interference of background traffic.

Method used

A malware detection and classification method based on multi-dimensional feature fusion is adopted. By obtaining system call files and traffic files, a system call-traffic map is established, edge feature matrix and node feature matrix are generated, and multimodal detection is performed using the BERT model and classifier.

Benefits of technology

It realizes complete retention of system call order dependence and traffic characteristics, reduces the possibility of missed judgments and improves detection quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068074A_ABST
    Figure CN120068074A_ABST
Patent Text Reader

Abstract

The invention relates to a malicious software detection and classification method based on multi-dimensional feature fusion. The method comprises the following steps: acquiring a system call file and a traffic file generated in a dynamic operation environment of a to-be-detected software sample; establishing a node graph based on system calling through the system calling file; the node graph takes input and output of system calling as edges of the image and takes system calling parameters as nodes of the image; extracting a data load in the traffic file, matching the data load with a system call parameter, adding a traffic node and an edge of the matched system call and traffic node on a node graph based on a matching result, and obtaining a system call-traffic graph; according to the system call-flow diagram, generating an edge feature matrix, and inputting the system call-flow diagram into a BERT model to obtain a node feature matrix; and inputting the edge feature matrix and the node feature matrix into a classifier to obtain a malicious software detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of malware detection, and particularly to a malware detection and classification method based on multi-dimensional feature fusion. Background Art

[0002] In the prior art (malware detection), in most of the current malware detection methods, a single modality is used for detection. The traditional defense model is based on a single dimension of data representation, while the attack behavior data of application software is rich, heterogeneous, and multimodal. Therefore, using a single modality for application software detection may lead to the possibility of misjudgment and missed judgment, affecting the detection quality.

[0003] In the existing solution "A container intrusion detection method and system based on the threat level of system calls", the system call sequence is filtered by a method based on word frequency statistics weighting to generate a new system call sequence for detection, and this method ignores the dependence on the call order.

[0004] In the existing solution "A malware detection method, device, and electronic device based on traffic", it uses traffic features for malware detection. In the case where the malware does not generate any traffic or the noise generated by background traffic is too large, the detection result will be poor.

[0005] In the existing solution "An attack detection system based on Linux system calls", it removes redundant system calls, generates a new system call sequence, and then matches it with the system call sequence in the dataset to determine whether it is malicious. This method causes chaos in the system call order. Summary of the Invention

[0006] In order to solve the problems existing in the above prior art, the purpose of the present invention is to provide a malware detection and classification method based on multi-dimensional feature fusion. The present invention accurately extracts system calls, filters out redundant system calls, maps system calls to traffic data packets one by one, establishes a stable data relationship, and uses machine learning for multimodal malware detection.

[0007] To achieve the above object, the present invention provides the following solution:

[0008] A malware detection and classification method based on multi-dimensional feature fusion, comprising:

[0009] Obtaining a system call file and a traffic file generated in the dynamic running environment of a software sample to be detected;

[0010] Call a file through the system to establish a node graph based on system calls; the node graph includes: edges with the inputs and outputs of system calls as images, and nodes with system call parameters as images;

[0011] Extract the data payload in the traffic file, match the data payload with the system call parameters, and based on the matching result, add traffic nodes and edges between the matched system calls and the traffic nodes on the node graph to obtain a system call-traffic graph;

[0012] Generate an edge feature matrix according to the system call-traffic graph, and input the system call-traffic graph into the BERT model to obtain a node feature matrix;

[0013] Input the edge feature matrix and the node feature matrix into a classifier to obtain a malware detection result; the classifier is trained using a training set and adjusts model parameters by minimizing the cross-entropy loss function; the training set includes: an original node feature matrix and an original edge feature matrix; a LabelEmbedding layer is added on the basis of the input end of the classifier for adding an embedded expression of node labels.

[0014] Optionally, establishing the node graph includes:

[0015] Map the system calls to file descriptors through the system call file, and according to the relationship between the system calls and the file descriptors, that is, when the file descriptors generated by the system calls are applied in subsequent system calls, connect the system calls with the subsequent system calls to form a system call chain, thereby further generating the node graph.

[0016] Optionally, extracting the data payload in the traffic file includes:

[0017] Extract the Burst traffic in the traffic file, where the Burst traffic is a sequence of consecutive data packets in the same transmission direction in the network flow, and capture the data packets in the Burst traffic;

[0018] Associate the data packets belonging to the same network flow according to the five-tuples of the data packets to obtain a data packet set;

[0019] Set extraction target conditions, and based on the extraction target conditions, extract consecutive subsequences of the data packet set, and use the consecutive subsequences as individual Burst traffic, thereby extracting the data payloads of all data packets in the individual Burst traffic:

[0020]

[0021] Among them, Payload(p) represents the data payload corresponding to a single packet, and B k represents a continuous subsequence. Payload(B k ) represents the data payload corresponding to the continuous subsequence.

[0022] Optionally, the extraction target conditions include:

[0023] Continuity constraint: B k ={p i , p i+1 ,..., p j};

[0024] Direction consistency constraint:

[0025] Among them, p i and p j are both data packets of continuous subsequences, d i is the direction identifier, represents the direction consistency constraint.

[0026] Optionally, obtaining the system call - traffic graph includes:

[0027] Adding traffic nodes to the node graph;

[0028] Adding an edge between the matched system call and the traffic node to the node graph:

[0029] If then it proves the system call triggered by the out - direction or in - direction of a single Burst traffic;

[0030] Determining the edge direction between the matched system call and the traffic node according to the out - direction and the in - direction:

[0031]

[0032] Among them, is the out - direction, is the in - direction, Payload(p) represents the data payload corresponding to a single packet, Parameters(s i ) represents the parameters of the system call;

[0033] Obtaining the system call - traffic graph by adding traffic nodes and the edge between the matched system call and the traffic node to the node graph.

[0034] 6. The malware detection and classification method based on multi-dimensional feature fusion according to claim 1, wherein the edge feature matrix is a three-dimensional feature, and the three-dimensional feature includes: the frequency of edge occurrence, the interval time, and the transmitted data packet; the frequency of edge occurrence is the probability of the edge occurrence between system calls in the system call-traffic graph, the interval time is the interval time feature between system calls, and the transmitted data packet is the total size of Burst network traffic data packets.

[0035] Optionally, obtaining the node feature matrix includes:

[0036] Decompose the target content in the system call parameters in the system call-traffic graph into first words, represent the traffic nodes with hexadecimal sequences for data packets, and perform double-byte encoding, with every two bytes represented as a second word;

[0037] Embed each of the first words and each of the second words:

[0038] E i = TokenEmbedding(t i )

[0039] P i = PositonEmbedding(t i )

[0040] wherein, E i is the token embedding, P i is the position embedding, TokenEmbedding(t i ) is the word vector embedding of the i-th word, TokenEmbedding(.) is the embedding layer, and PositionEmbedding(t i ) is the position information embedding position embedding layer of this word;

[0041] Input the embedding result into the BERT model, and use the stacked Transformer encoders in the BERT model to capture long-range dependencies while paying attention to different regions of the input sequence:

[0042] {H' 1 ,H' 2 ,...,H' n} = TransformerEncoder({H 1 ,H 2 ,...,H n})

[0043] wherein, H' nIs the output of H after passing through the Transformer encoder n Output of TransformerEncoder({H 1 , H 2 ,..., H n} encodes the input sequence using the Transformer encoder architecture, and H n is the embedding result of the nth character;

[0044] Based on the long-range dependencies and different regions of the input sequence, generate embeddings for the entire system call nodes or traffic nodes, as the feature vectors corresponding to each row generated by the expression attributes of the nodes, thereby generating the node feature matrix.

[0045] Optionally, obtaining the malware detection result includes:

[0046] Input the edge feature matrix and the node feature matrix into the classifier. A Label Embedding layer is added at the input end of the classifier for adding the embedding expression of the node labels:

[0047] Input = H CLS + W 1 * Laplacian matrix + W 2 * Node Label

[0048]

[0049] where Input is the model input, H' CLS is the embedding feature of the node, Laplacian matrix is the Laplacian matrix, Node Label is the node label, W 1 is a learnable weight matrix, w 2 is a learnable weight matrix, syscall is the system call node type, and burst is the network traffic node type;

[0050] Perform self-attention calculation between system call nodes with edges in the input data and between system call nodes with edges and traffic nodes, and combine with the global average pooling layer to obtain the graph-level embedding result:

[0051]

[0052] where H G is the graph-level embedding result, is the representation of the feature of the ith node at the Lth layer (the last layer), and v iis the i-th node in the graph, V is the set of nodes in the graph, and L is the output of the last graph transformation layer;

[0053] Input the graph-level embedding result into the first fully connected layer to map the global features to the target dimension, and perform random dropout through the dropout layer:

[0054] H D = Dropout(ReLU(W 1 H G + b 1 ))

[0055] where H D is the intermediate feature representation after the dropout layer, W 1 is the weight of the first fully connected layer, and b 1 is the bias of the first fully connected layer;

[0056] Input the embedding result after random dropout into the second fully connected layer to generate the malware detection result.

[0057] Optionally, the self-attention calculation between system call nodes with edges in the input data includes:

[0058] Set the update of node features:

[0059]

[0060] where b Si and h Sj are the node features of system call nodes S i and S j respectively, e ij is the edge feature, W O is a different learnable parameter matrix, is the weight matrix for updating node features (a weighted combination of node features h Si , h Sj , and edge feature e ij ), is the output of the node features of the h-th attention head, is the vector after concatenating all attention heads, h Si is the result of normalizing the sum of the vector obtained by concatenating the outputs of all attention heads and the original node vector, is the output of the last attention head, d K is the dimension of the node feature vector, and h Si is the feature vector of the original system call node S i ;

[0061] Set the update of edge features:

[0062]

[0063] Among them, is the result of concatenating the weight matrices for updating node features of all attention heads, is a learnable parameter matrix, is the weight matrix of the last attention head, e' ij is the result of normalizing the sum of the concatenation of the weight matrices of all attention heads and the original edge feature matrix.

[0064] Optionally, performing self-attention calculation between system call nodes and traffic nodes with edges in the input data includes:

[0065] Setting the query vector, key vector, and value vector of the traffic node to be multiplied by a first excitation factor:

[0066]

[0067] where α is the first excitation factor, is the weight matrix for updating node features (a weighted combination of node features h Bi 、h Si , and edge feature e ij ), is the node feature output of the h-th attention head, is a learnable parameter matrix, h Si is the system call node feature matrix, W O is a learnable parameter matrix, is the traffic node feature output by the last attention head, h′ Bj is the result of normalizing the sum of the vector obtained by concatenating the network traffic node features output by all attention heads and the original traffic node features, h Bj is the original traffic node feature;

[0068] Setting the key vector and value vector of the traffic node to be multiplied by a second excitation factor:

[0069]

[0070] where β is the second excitation factor, is the weight matrix for updating node features, is a learnable weight matrix, is a learnable parameter matrix, is a learnable parameter matrix, is the system call node feature output of the h-th attention head, h Bj is the feature of the traffic node, is the result of concatenating the node features output by all attention heads, W O is a learnable parameter matrix, is the system call node feature output by the last attention head, h′ Si is the result of normalizing the sum of the vector obtained by concatenating the outputs of all attention heads and the original node vector, h Si is the original system call node feature.

[0071] The beneficial effects of the present invention are:

[0072] The present invention constructs the edges of the graph according to the correlation between the output and input of system calls, and completely retains the dependencies in the call order.

[0073] The present invention maps the traffic generated by the application through the parameters of system calls, and there is no interference such as background traffic.

[0074] Since the present invention establishes a relationship according to the file descriptors of system calls and the parameter situations of system calls, the disorder of system call order caused by CPU scheduling can be solved.

[0075] Currently, the more commonly used method is single-modal detection. When using the sequence method to statistically extract features during the feature extraction process, too much meaningless data will be introduced, thus generating noise. To address the above problems, the present invention accurately extracts system calls and filters out redundant system calls, and performs a one-to-one mapping between system calls and traffic data packets, establishing a stable data relationship and using machine learning for multi-modal malware detection.

[0076] In summary, the present invention realizes accurately capturing specific traffic from a complex and massive network traffic environment according to the system calls of a single process or application by fusing system calls and traffic features. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0078] Figure 1 is a flowchart of a malware detection and classification method based on multi-dimensional feature fusion according to an embodiment of the present invention;

[0079] Figure 2 is a schematic diagram of establishing a graph based on system calls according to an embodiment of the present invention;

[0080] Figure 3 The figure within the frame shows a schematic diagram of a Burst traffic for an embodiment of the present invention;

[0081] Figure 4 The figure shows a schematic diagram of the data payload and system call mapping of the Burst for an embodiment of the present invention;

[0082] Figure 5 The figure shows a schematic diagram of the process of embedding traffic nodes for an embodiment of the present invention;

[0083] Figure 6 The figure shows a schematic diagram of the classifier structure for an embodiment of the present invention. Detailed implementation manners

[0084] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0085] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0086] This embodiment discloses a malicious software detection and classification method based on multi-dimensional feature fusion, including: obtaining a system call file and a traffic file generated in the dynamic running environment of a software sample to be detected; establishing a node graph based on system calls through the system call file; the node graph includes: using the input and output of system calls as the edges of the image and the system call parameters as the nodes of the image; extracting the data payload in the traffic file, matching the data payload with the system call parameters, and based on the matching result, adding traffic nodes and the edges between the matched system calls and traffic nodes on the node graph to obtain a system call-traffic graph; generating an edge feature matrix according to the system call-traffic graph, and inputting the system call-traffic graph into a BERT model to obtain a node feature matrix; inputting the edge feature matrix and the node feature matrix into a classifier to obtain a malicious software detection result; the classifier is obtained by training with a training set and adjusting the model parameters by minimizing the cross-entropy loss function; the training set includes: an original node feature matrix and an original edge feature matrix; a Label Embedding layer is added on the basis of the input end of the classifier for adding the embedded expression of node labels.

[0087] Specifically: This embodiment discloses a malware detection and classification method based on multi-dimensional feature fusion, which uses a graph structure to fuse traffic features and system call features, so that their modal structures are in the same dimension. The present invention also proposes a new graph feature with system call parameter embedding and traffic embedding in units of Burst as node features. The background and flowchart of the present invention are as Figure 1 shown.

[0088] Step1 In a dynamic running environment, capture the system call file and traffic file generated by the software application.

[0089] Step2 Establish a graph based on system calls: Establish edges according to the inputs and outputs of each system call. The edges represent the data flow relationship, and the nodes represent system calls.

[0090] Step3 Extract Burst (burst) traffic and the data payload of the packets contained in the Burst from the traffic file.

[0091] Step4 In subsequent work, the present invention matches the data payload of the Burst with the parameters of the system call. If the match is successful, a traffic node is added to the graph, and an edge between the matched system call and the traffic node is added. The edge represents the causal relationship.

[0092] Step5 Use the obtained system call-traffic graph to complete node embedding through the BERT model to obtain the node feature matrix and edge matrix of the graph.

[0093] Step6 Input the obtained feature matrix and edge matrix into the classifier to obtain the classification result.

[0094] Further, establishing the node graph includes: mapping the system call to a file descriptor through the system call file. According to the relationship between the system call and the file descriptor, that is, when the file descriptor generated by the system call is applied in subsequent system calls, the system call is connected to the subsequent system call to form a system call chain, thereby further generating the node graph.

[0095] Specifically:

[0096] Capture the system call and use the obtained system call to construct a system call set, denoted as S = {s 1 , s 2 , …, s n}, where n is the number of system calls.

[0097] Establish a behavior graph according to the captured system call set and packet set.

[0098] Establish a graph based on system calls. The system call s iis an element in the set S = {s 1 , s 2 , …, s n}. The present invention constructs a system call base graph by mapping system calls to file descriptors (fd) and establishing edges according to the relationships between them.

[0099] For where i < j < k, when s i generates a file descriptor fd 1 , in subsequent system calls, s j , s k uses fd 1 as the first parameter (a system can call multiple parameters). Then connect them into a system call chain: s i →s j →s k . If s L uses fd 1 as the second parameter, directly connect s i and s L . Finally, use the system call parameters to represent the node.

[0100] For example: File descriptor 62 is generated. The subsequent [B][C][D][E] all use this descriptor 62, forming a system call chain. And [L] uses two file descriptors 73, 62. Among them, 62 is the second parameter, so [A] can be directly connected to [L]. And use the system call parameters corresponding to this node as its expression. For example, the [D] node is expressed as its parameters (62, prot(443), ip(100.8.8.8)), rather than getsockopot.

[0101] As Figure 2 shown, the input of this algorithm is: a system call sequence; the output is: a graph based on system calls constructed through file descriptors.

[0102] Furthermore, extracting the data payload in the traffic file includes: extracting the Burst traffic in the traffic file, where the Burst traffic is a sequence of consecutive data packets in the same transmission direction in the network flow, capturing the data packets in the Burst traffic; associating the data packets belonging to the same network flow according to the five-tuples of the data packets to obtain a data packet set; setting extraction target conditions, and based on the extraction target conditions, extracting consecutive subsequences of the data packet set, and taking the consecutive subsequences as individual Burst traffic, so as to extract the data payload of all data packets in the individual Burst traffic:

[0103]

[0104] Among them, Payload(p) represents the data payload corresponding to a single packet, and B k represents a continuous subsequence. Payload(B k ) represents the data payload corresponding to the continuous subsequence.

[0105] Furthermore, the extraction target conditions include: Continuity constraint: B k ={p i , p i+1 ,..., p j};

[0106] Direction consistency constraint: Among them, p i and p j are both data packets of continuous subsequences, d i is the direction identifier, represents the direction consistency constraint.

[0107] Specifically:

[0108] Capture the data packets Packet (hereinafter referred to as "packets") transmitted in the network (TCP / IP protocol communication transmission), and use the obtained packets to construct a packet set, denoted as p={p 1 , p 2 ,…, p m}, where m is the number of packets in the packet set.

[0109] Extract the mapping between the data payload in the data packet and Burst.

[0110] In network traffic analysis, Burst (burst) is defined as a continuous data packet sequence in the same transmission direction in the network flow, and this sequence has strict time and direction consistency. This sequence usually represents a short-term high-density data transmission behavior in network communication. As Figure 3 shown.

[0111] Associate the packets belonging to the same flow (distinguished according to the five-tuple of the packet - source address, source port, destination address, destination port, transport layer protocol). Define it as a network flow F. Then F can be represented as a set of data packets arranged in chronological order:

[0112] F={p 1 , p 2 , p 3 ,..., p n};

[0113] Among them, each data packet has an attribute: direction identifier d i ∈(in, out), indicating the data packet direction.

[0114] If a continuous subsequence B in F k , satisfies the following conditions:

[0115] Continuity constraint: B k ={p i , p i+1 ,..., p j};

[0116] Direction consistency constraint (same direction):

[0117] Then the subsequence B k is called a single Burst. It can be represented as (in direction) or (out direction).

[0118] Payload(p) represents the data payload corresponding to a single packet. Therefore, the extraction of the data payloads of all packets in a certain B k can be expressed as:

[0119]

[0120] Furthermore, obtaining the system call - traffic graph includes: adding traffic nodes to the node graph;

[0121] Adding edges between the matched system calls and traffic nodes to the node graph:

[0122] If then it proves the system calls triggered by the out direction or in direction of a single Burst traffic; where is an empty set;

[0123] Determining the edge directions between the matched system calls and traffic nodes according to the out direction and in direction:

[0124]

[0125] where is the out direction, is the in direction, Payload(p) represents the data payload corresponding to a single packet, and Parameters(s i ) represents the parameters of the system call;

[0126] Obtain the system call - traffic graph by adding traffic nodes and edges between the matched system calls and traffic nodes to the node graph.

[0127] Specifically:

[0128] For example Figure 4As shown, map system call parameters to a single Burst to add traffic nodes. Based on the data payload in the established Burst, establish a connection with the system call node. Specifically, the parameters of the system call are represented as Parameters(s i ), where s i ∈S, if:

[0129]

[0130] Then it indicates that the system call s i caused (outbound direction), or the system call triggered by this (inbound direction).

[0131] Determine the direction of the edge to be established according to :

[0132]

[0133] For example:

[0134] sendto(62<TCPv6:[[::ffff:10.0.2.15]:53672->[::ffff:104.16.160.145]:443]>,"\26\3\1\2\0\1\0\1\374\3\3\336V\272%tR\225\202l\207\364\3036\317\17\257\355u\240M\242\315\313\267=\2152f\235Z'\231\\\f\211\344\336\305\346\332ar\21\202\323\303\263\17\340Z&\307\342\305\360\250On}3\0\251\266\"\0\"\23\1\23\2\23\3\300+\300,\314\251\300 / \3000\314\250\300\t\300\n\300\23\300\24\0\234\0\235\0 / \0005\1\0\1\221\0\0\0\26\0\24\0\0\21api"...)=517;

[0135] Subsequently, filter the flow F to find this Burst according to the five-tuple in the file descriptor 62<TCPv6:[[::ffff:10.0.2.15]:53672->[::ffff:104.16.160.145]:443]>, as Figure 2As shown, the hexadecimal application layer data payload, after being converted by ASCII code and trimming a section of the string, corresponds to the above Burst content.

[0136] The Burst content is:

[0137] "\"\26\3\1\2\0\1\0\1\374\3\3\336V\272%tR\225\202l\207\364\3036\317\17\257\355u\240M\242\315\313\267=\2152f\235Z'\231\\\f\211\344\336\305\346\332ar\21\202\323\303\263\17\340Z&\307\342\305\360\250On}3\0\251\266\"\0\"\23\1\23\2\23\3\300+\300,\314\251\300 / \3000\314\250\300\t\300\n\300\23\300\24\0\234\0\235\0 / \0005\1\0\1\221\0\0\0\26\0\24\0\0\21api"”.

[0138] Subsequently, the hexadecimal sequences of these two packets are used to represent the Burst. And it is the Burst in the outgoing direction.

[0139] The Burst is represented as: the hexadecimal sequence of the first data packet (64 bytes) + the hexadecimal sequence of the second data packet (64 bytes).

[0140] 64 bytes: “52 55 0a 00 02 02 5a 36 c0 42 4a f0 08 00 45 00 00 28 a2 3d40 00 40 06 83 e2 0a 00 02 0f 68 10 a0 91 d1 a8 01 bb df aa 3a ff 8c c3 46 0250 10 ff ff 14 cb 00 00...”

[0141] Finally, establish the edge between the system call node and the traffic node:

[0142] The constructed graph G=(V, E), where V represents the nodes in the graph, represented as system calls or traffic. Among them, the edge (v i , v j ) indicates the data flow relationship from node v i to node v j .

[0143] For example:

[0144] The current expression matrix of this graph is: And an adjacency matrix is used to represent the connection of edges, as shown in Table 1, and Table 1 is the representation of symbols.

[0145] Table 1

[0146]

[0147] Furthermore, the edge feature matrix is a three-dimensional feature, and the three-dimensional feature includes: the frequency of edge occurrence, the interval time, and the transmitted data packet; the frequency of edge occurrence is the probability of the edge occurrence between system calls in the system call - traffic graph and subsequent system calls, the interval time is the interval time feature between system calls and subsequent system calls, and the transmitted data packet is the sum of the edges between system calls and traffic nodes.

[0148] Specifically:

[0149] Generate the edge feature matrix of graph G:

[0150] For the edge feature matrix of graph G, there are a total of three-dimensional features: [frequency of edge occurrence, interval time, size of transmitted data packet].

[0151] For the frequency of edge occurrence, count the frequency of the edges of Si and Sj appearing in this graph. That is, count the frequency of the edge E(Si, Sj) to represent the concentrated place of application behavior.

[0152] For the feature of the interval time, since the present invention constructs a new system call sequence through file descriptors, the present invention needs to count the interval time feature of system calls Si and Sj to confirm the dependency strength between this system call pair. That is, count the number of system calls separated between Si and Sj.

[0153] For example:

[0154] The original system call sequence was: Si, S1, S2, S3, Sj, and the constructed sequence is: Si, Sj. Then the interval time feature of the edge E(Si, Sj) is 3.

[0155] For the feature of the size of the transmitted data packet: For the edge B between system call Si and the traffic node j the total size E(S i , B j ). By calculating the transmitted Burst data packet B j = {p i , p i+1 ,..., p k}(Total size in bytes) to represent the interaction between the system call and the traffic. That is:

[0156] Finally, the feature matrix of the edge [x1, x2, x3] is obtained.

[0157] Furthermore, obtaining the node feature matrix includes:

[0158] Decompose the target content in the system call parameters in the system call - traffic graph into the first words, represent the traffic nodes with hexadecimal sequences for data packets, and perform double - byte encoding, with every two bytes represented as a second word;

[0159] Embed each first word and each second word:

[0160] E i = TokenEmbedding(t i );

[0161] P i = PositionEmbedding(t i );

[0162] Among them, E i is the token embedding, and P i is the position embedding;

[0163] Input the embedding result into the BERT model, and use the stacked Transformer encoders in the BERT model to capture long - range dependencies while paying attention to different regions of the input sequence:

[0164] {H' 1 ,H' 2 ,...,H' n}} = TransformerEncoder({H 1 ,H 2 ,...,H n});

[0165] Among them, H' n is the output of H n after passing through the Transformer encoder, TransformerEncoder({H 1 ,H 2 ,...,H n}} is to encode the input sequence by the Transformer encoder architecture, and H n is the embedding result of the nth character;

[0166] Generate the embeddings of the entire system call nodes or traffic nodes based on long-range dependencies and different regions of the input sequence, as the feature vectors corresponding to each line generated by the expression attributes of the nodes, that is, a node feature matrix is generated.

[0167] Specifically:

[0168] Generate the feature matrix of graph G: For generating the feature matrix of a graph, use the expression attributes of the nodes for vectorization.

[0169] Illustrate with an example:

[0170] Input: This matrix is input into the BERT model to generate the vectors corresponding to each line, that is, the node feature matrix of the graph is generated, and the output matrix X:

[0171]

[0172] The following will explain this step in detail:

[0173] For the system call nodes and traffic nodes in graph G, the present invention uses the BERT model to generate feature vectors, thereby generating rich context embeddings. In BERT, there are a maximum position embedding of 512, a hidden size of 768, 12 attention heads, and 6 hidden layers. Therefore, each system call node and traffic node is represented by a 768-dimensional vector.

[0174] Generate system call node Tokens. For the system call node s i =[a 1 , a 2 ,..., a n , where a i represents a certain part of its parameters. Further decompose each part a i into sub-word Tokens:

[0175]

[0176] Therefore, for the system call node s i is finally expressed as: s i =[t 1 , t 2 ,..., t k , where k is the total number of Tokens after all word segmentations.

[0177] Generate traffic node Tokens. For the traffic node B k ={p 1 , p 2 ,..., p j}, represent the data packet with a hexadecimal sequence, and then use double-byte encoding to represent every two bytes as a Token. Therefore, for traffic node B k Finally expressed as: B k =[t 1 ,t 2 ,...,t j , where j is the total number of Tokens.

[0178] For example: Suppose a certain B i The hexadecimal sequences of the two packets it contains are [0x00, 0x04...] and [...0xdc, 0xf9, 0x4a].

[0179] Form Tokens [0004, 0400,..., dc9f, f94a] through double-byte encoding. Subsequently, input the Tokens into BERT to obtain the embeddings. As Figure 5 shown.

[0180] For each Tokent i , before being used as the input to BERT, the following embeddings need to be calculated:

[0181] E i =TokenEmbedding(t i );

[0182] P i =PositionEmbedding(t i );

[0183] Then, the resulting embedding H i is obtained by adding the token embedding E i and the position embedding P i , and it is fed into a stacked Transformer encoder, which consists of 6 self-attention layers. This enables the model to focus on different regions of the input sequence while capturing long-range dependencies.

[0184] H i =E i +P i

[0185] {H' 1 ,H' 2 ,...,H' n}=TransformerEncoder({H 1 ,H 2 ,...,H n});

[0186] The output is a series of context-aware embeddings, and the output using the [CLS] token represents the entire input sequence:

[0187] H′ CLS = CLS;

[0188] where H′ CLS represents the embedding of the entire s i or B k embedding.

[0189] BERT can effectively learn the relationship patterns among the parts in the parameters and the relationship patterns among the parts in the traffic. Finally, the formed H′ CLS is used as the feature vector of this node.

[0190] By associating the relationship between the parameters in the system call and the data payload of the network traffic, a new graph structure is established for aggregating the system calls and traffic generated by the application. And the present invention innovatively uses the embedding of the system call parameters and the traffic embedding in units of Burst as the features of the graph.

[0191] This method can accurately capture the traffic characteristics according to the system calls generated by the application in a complex and massive network traffic environment, so as to realize the detection of the software at the system level and network level.

[0192] Furthermore, obtaining the malware detection result includes:

[0193] Input the edge feature matrix and the node feature matrix into the classifier, and a Label Embedding layer is added at the input end of the classifier for adding the embedding expression of the node label:

[0194] Input = H′ CLS + W 1 * Laplacian matrix + W 2 * Node Label;

[0195]

[0196] where Input is the model input;

[0197] Perform self-attention calculation between the system call nodes with edges in the input data and between the system call nodes with edges and the traffic nodes, and combine with the global average pooling layer to obtain the graph-level embedding result:

[0198]

[0199] where H G is the graph-level embedding result;

[0200] Input the embedding results at the figure level into the first fully connected layer to map the global features to the target dimension, and perform random dropout through the dropout layer:

[0201] H D = Dropout(ReLU(W 1 H G + b 1 ));

[0202] Among them, H D is the intermediate feature representation after the dropout layer, W 1 is the weight of the first fully connected layer, and b 1 is the bias of the first fully connected layer;

[0203] Input the embedding results after random dropout into the second fully connected layer to generate malware detection results.

[0204] Furthermore, the self-attention calculation between system call nodes with edges in the input data includes:

[0205] Set the update of node features:

[0206]

[0207] Among them, h Si , h Sj are the node features of system call nodes Si and Sj respectively, e ij is the edge feature, W O is a different learnable parameter matrix;

[0208] Set the update of edge features:

[0209]

[0210] Among them, is the result of concatenating the weight matrices for updating node features of all attention heads, is a learnable parameter matrix, is the weight matrix of the last attention head, and e' ij is the result of normalizing the sum of the concatenation of the weight matrices of all attention heads and the original edge feature matrix.

[0211] Furthermore, the self-attention calculation between system call nodes with edges in the input data and traffic nodes includes:

[0212] Set the query vector, key vector, and value vector of the traffic node to be multiplied by the first excitation factor:

[0213]

[0214] Among them, α is the first excitation factor, is the weight matrix for updating node features (a weighted combination of node features h Bi 、h Si , and edge feature e ij );

[0215] Multiply the key vector and value vector of the traffic node by the second excitation factor:

[0216]

[0217] Among them, β is the second excitation factor.

[0218] S6 performs system call - traffic graph classification:

[0219] S61 The obtained represents the node feature matrix of graph G, where D = 768 represents the attribute dimension. N represents the number of nodes |V| of graph G. represents the edge feature matrix of graph G, where M is the number of edges, and the edge features have three dimensions [frequency of edge appearance, interval time, size of sent data packets]. Put X F 、X E into the set Graph Transformer.

[0220] Specifically: The input of the original Graph Transformer is:

[0221] Input = Node Embedding + Position Embedding;

[0222] The detailed input is: Input = H′ CLS +W 1 *Laplacian matrix;

[0223] Among them, H′ CLS is the embedded feature of the node, Laplacian matrix = A - D, where is the adjacency matrix of the graph, is the degree matrix of the graph nodes.

[0224] The present invention changes the input to:

[0225] Input = Node Embedding + Position Embedding + Label Embedding

[0226] That is, the embedded expression of the node label is added to enhance the input expression of the node.

[0227] Specifically, as Figure 6 added by the box, the input is changed to:

[0228] Input = H CLS + W 1 * Laplacian matrix + W 2 * Node Label;

[0229]

[0230] For the self - attention calculation between system call nodes with an edge E:

[0231] First, for system call nodes Si, Sj with an edge E. Their node features are h Si , h Sj , and the edge feature is e ij .

[0232] For the update of node features h Si , h Sj :

[0233]

[0234] Among them, h = [1, H], where H represents the number of attention heads;

[0235] For the update of edge feature e ij :

[0236]

[0237] For the self - attention calculation between system call nodes with an edge E and network traffic nodes:

[0238] Since the number of system call nodes and network traffic nodes is unbalanced, the number of system call nodes > the number of network traffic nodes.

[0239] Set the query vector key vector and value vector for network traffic nodes multiplied by an excitation factor:

[0240] For the node features corresponding to system call node Si and network traffic Bi node: h Fi , h Bj , and the corresponding edge feature is e ij . For the query vector Q of traffic nodes, multiply by an excitation factor α > 1:

[0241]

[0242] For the key vector K and value vector V of the traffic node, multiply by the excitation factor β > 1:

[0243]

[0244] To perform graph classification, the present invention applies a global average pooling layer to obtain a graph-level embedding H G . It can be represented by the equation:

[0245] where L represents the output of the last graph transformation layer.

[0246] In the classification prediction stage, the first fully connected layer maps the global features to a 32-dimensional intermediate representation and randomly drops neurons through a dropout layer:

[0247] H D = Dropout(ReLU(W 1 H G + b 1 ));

[0248] W 1 and b 1 represent the weights and biases of the first fully connected layer. The ReLU activation function is used to introduce non-linearity. Dropout is used to randomly drop some neurons to prevent overfitting. H D is the intermediate feature representation after the dropout layer.

[0249] The second fully connected layer generates the final classification result to predict whether the graph G represents malicious behavior:

[0250]

[0251] where W c and b c are the weights and biases of the classifier respectively. After training, the model adjusts the parameters by minimizing the cross-entropy loss function to optimize the detection accuracy and improve the training efficiency. Here, N is the total number of samples in the training set.

[0252]

[0253] where, y i is the true label of sample i, is the predicted value of sample i.

[0254] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A malware detection and classification method based on multi-dimensional feature fusion, characterized in that: include: Obtaining system call files and flow files generated in the dynamic running environment of the software sample to be detected; A node graph based on the system call is established through the system call file; the node graph includes: inputs and outputs of the system call as edges of the graph and system call parameters as nodes of the graph; Extracting the data load in the traffic file, matching the data load with the system call parameter, and adding a traffic node and an edge between the matched system call and the traffic node on the node graph based on the matching result to obtain a system call-traffic graph; Generate an edge feature matrix according to the system call-flow graph, and input the system call-flow graph into a BERT model to obtain a node feature matrix; The edge feature matrix and the node feature matrix are input into the classifier to obtain malware detection results; the classifier is trained by using a training set and the model parameters are adjusted by minimizing the cross entropy loss function; the training set includes: an original node feature matrix and an original edge feature matrix; a LabelEmbedding layer is added on the basis of the classifier input end to add an embedded expression of the node label.

2. The malware detection and classification method based on multi-dimensional feature fusion according to claim 1 is characterized in that: Establishing the node graph includes: Through the system call file, the system call is mapped to a file descriptor. According to the relationship between the system call and the file descriptor, that is, when the file descriptor generated by the system call is applied in a subsequent system call, the system call is connected with the subsequent system call to form a system call chain, thereby further generating the node graph.

3. The malware detection and classification method based on multi-dimensional feature fusion according to claim 1 is characterized in that: Extracting the data payload in the traffic file includes: Extracting the Burst traffic in the traffic file, the Burst traffic is a sequence of continuous data packets in the same transmission direction in the network flow, and capturing the data packets in the Burst traffic; According to the quintuple of the data packet, the data packets belonging to the same network flow are associated to obtain a data packet set; An extraction target condition is set, and based on the extraction target condition, a continuous subsequence of the data packet set is extracted, and the continuous subsequence is used as a single Burst flow, thereby extracting the data payload of all data packets in the single Burst flow: Where Payload(p) represents the data load corresponding to a single packet, B k Indicates a continuous subsequence, Payload(B k ) represents the data load corresponding to the continuous subsequence.

4. The malware detection and classification method based on multi-dimensional feature fusion according to claim 3 is characterized in that: The extraction target conditions include: Continuity constraint: B k ={p i , p i+1 , ..., p j }; Directional consistency constraints: Among them, p i and p j are all continuous subsequence packets, d i is the direction identifier, Represents a directional consistency constraint.

5. The malware detection and classification method based on multi-dimensional feature fusion according to claim 1 is characterized in that: Obtaining the system call-flow graph includes: Adding a flow node to the node graph; Add an edge between the matched system call and the traffic node on the node graph: like This proves the system calls caused by the outbound or inbound direction of a single Burst flow; According to the outgoing direction and the incoming direction, determine the edge direction of the matched system call and the traffic node: in, For the outgoing direction, In the inbound direction, Payload (p) indicates the data load corresponding to a single packet, Parameters (s i ) represents the parameters of the system call; The system call-traffic graph is obtained by adding a traffic node and an edge between the matched system call and the traffic node to the node graph.

6. The malware detection and classification method based on multi-dimensional feature fusion according to claim 1 is characterized in that: The edge feature matrix is ​​a three-dimensional feature, and the three-dimensional features include: the frequency of edge occurrence, the interval time and the sent data packet; the frequency of edge occurrence is the probability of edge occurrence between the system call and the subsequent system call in the system call-flow graph, the interval time is the interval time feature between the system call and the subsequent system call, and the sent data packet is the sum of the sizes of the Burst network traffic data packets.

7. The malware detection and classification method based on multi-dimensional feature fusion according to claim 1 is characterized in that: Acquiring the node feature matrix includes: Decomposing the target content in the system call parameter in the system call-traffic graph into first words, representing the data packet by the traffic node using a hexadecimal sequence, and performing double-byte encoding, representing each two bytes as a second word; Embed each of the first words and each of the second words: E i N TokenEmbedding ( t i ) P i =PositionEmbedding(t i ) Among them, E i For token embedding, P i For position embedding, TokenEmbedding(t i ) is the word vector embedding of the i-th word, TokenEmbedding(.) is the embedding layer, PositionEmbedding(t i ) is the position embedding layer for embedding the position information of the word; The embedding result is input into the BERT model, and the stacked Transformer encoder in the BERT model is used to capture long-range dependencies while focusing on different areas of the input sequence: {H'1,H'2,...,H' n }=TransformerEncoder({H1,H2,...,H n }) Among them, H' n is H after the Transformer encoder n The output of TransformerEncoder({H1,H2,...,H n } is to encode the input sequence for the Transformer encoder architecture, H n is the embedding result of the nth character; Based on the long-range dependency and different regions of the input sequence, an embedding of the entire system call node or traffic node is generated, and a feature vector corresponding to each row generated as the expression attribute of the node is generated, that is, the node feature matrix is ​​generated.

8. The malware detection and classification method based on multi-dimensional feature fusion according to claim 1 is characterized in that: Obtaining the malware detection result includes: The edge feature matrix and the node feature matrix are input into the classifier, and a LabelEmbedding layer is added on the basis of the classifier input to add the embedding expression of the node label: Input=H′ CLS +W1*Laplacian matrix+W2*Node Label Among them, Input is the model input, H′ CLS is the embedding feature of the node, Laplacian matrix is ​​the Laplacian matrix, Node Label is the node label, W1 is the learnable weight matrix, W2 is the learnable weight matrix, syscall is the system call node type, and burst is the network traffic node type; Self-attention calculations are performed between system call nodes with edges in the input data, and between system call nodes and traffic nodes with edges, and combined with the global average pooling layer to obtain graph-level embedding results: Among them, H G is the graph-level embedding result, is the representation of the feature of the i-th node in the L-th layer, v i is the i-th node in the graph, V is the set of nodes in the graph, and L is the output of the last graph conversion layer; The graph-level embedding result is input into the first fully connected layer to map the global features to the target dimension, and then randomly discarded through the discard layer: H D =Dropout(ReLU(W1H G +b1)) Among them, H D is the intermediate feature representation after the discard layer, W1 is the weight of the first fully connected layer, and b1 is the bias of the first fully connected layer; The embedding result after random discarding is input into the second fully connected layer to generate the malware detection result.

9. The malware detection and classification method based on multi-dimensional feature fusion according to claim 8 is characterized in that: The self-attention calculation between the system call nodes with edges in the input data includes: To set the update of node characteristics: Among them, h Si 、h Sj They are system call nodes S i , S j The node characteristics, e ij is the edge feature, W O are different learnable parameter matrices, is the weight matrix used to update node features, is the node feature output of the h-th attention head, is the concatenated vector of all attention heads, h′ Si It is the result of normalizing the concatenated vector of the outputs of all attention heads and the original node vector. is the output of the last attention head, d K is the dimension of the node feature vector, h Si The original system call node S i The eigenvector of To set the update of edge features: in, is the concatenation of the weight matrices of all attention heads used to update node features. is the learnable parameter matrix, is the weight matrix of the last attention head, e' ij It is the result of concatenating the weight matrices of all attention heads and adding them to the original edge feature matrix and then normalizing them.

10. The malware detection and classification method based on multi-dimensional feature fusion according to claim 8, characterized in that: The self-attention calculation between the system call node and the traffic node with an edge in the input data includes: The query vector, key vector and value vector of the traffic node are set to be multiplied by the first incentive factor: Among them, α is the first excitation factor, is the weight matrix used to update node features, is the node feature output of the h-th attention head, is the learnable parameter matrix, h Si is the system call node feature matrix, W O is the learnable parameter matrix, is the traffic node feature output by the last attention head, h′ Bj h is the result of normalizing the concatenated vector of network traffic node features output by all attention heads and adding the original traffic node features. Bj is the original traffic node feature; Set the key vector and value vector of the traffic node to be multiplied by the second incentive factor: Among them, β is the second excitation factor, is the weight matrix used to update node features, is the learnable weight matrix, is the learnable parameter matrix, is the learnable parameter matrix, is the system call node feature output of the h-th attention head output, h Bj is the characteristic of the traffic node, is the result of concatenating the node features output by all attention heads, W O is the learnable parameter matrix, is the system call node feature output by the last attention head, h′ Sj h is the normalized result of adding the concatenated output vector of all attention heads to the original node vector. Si It is the original system call node feature.

Citation Information

Patent Citations

  • Application privacy leakage detection method and system, terminal and medium

    CN113158251A

  • Malicious application detection method and system and electronic equipment

    CN117763547A

  • Reinforced Android malicious application robust detection method based on mixed features

    CN117852032A

  • Method and system for detecting malicious software based on convolutional neural network

    CN118332551A