A Fast Malware Detection Method Based on Graph Neural Networks

By constructing a heterogeneous graph network and a graph neural network with self-attention mechanism, the problems of low accuracy and slow speed in malware detection are solved, and fast and accurate detection of unknown software is achieved.

CN115344863BActive Publication Date: 2025-07-08GUANGXI POWER GRID CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210996905.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-19
Publication Date
2025-07-08
Estimated Expiration
2042-08-19

AI Technical Summary

Technical Problem

Existing malware detection methods have insufficient detection accuracy and speed, especially the detection speed of unknown software is too slow, and existing methods fail to effectively utilize the implicit relationships and higher-order neighbor attributes between malware samples.

Method used

Using a graph neural network-based method, a high-level leader node message aggregation is used to use the HG2vec algorithm to perform semantic fusion with self-attention mechanism, and a classification detection is used to construct a rapid detection model of incremental aggregation.

Benefits of technology

It improves the accuracy and speed of malware detection, can better mine the semantic relationships between software nodes, and improves the speed and accuracy of detection of unknown malware.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115344863B_ABST
    Figure CN115344863B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of network and information security technology, and specifically relates to a method for rapid malware detection based on graph neural networks. The method includes: constructing a malware detection model, using different meta-structures to mine hidden information in different entities of software nodes; capturing the correlation based on high-order content between nodes, using an attention mechanism to perform semantic fusion on meta-paths; adopting the Sim2vec algorithm based on meta-structure similarity matching to perform incremental aggregation on the embeddings of unknown software nodes and similar known software nodes, so as to improve the detection speed; considering the problem of inaccurate detection accuracy caused by the diversity of different malware entities and the complexity of semantic relationships, the present invention constructs a model using a heterogeneous information network, utilizes a high-order graph neural network to mine high-order feature information of malware, and then uses a similarity algorithm for matching, which can effectively perform rapid detection of malware.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network and information security, and particularly relates to a method for quickly detecting malicious software based on a graph neural network. Background Art

[0002] With the continuous development of science and technology, network security has attracted more and more attention, especially the security of software. Among them, malicious software poses a serious threat to network security and information security. How to accurately detect malicious software and how to reduce the harm caused by malicious software have become urgent problems to be solved. Existing detections can be divided into two major types: static (signature-based) detection and dynamic (behavior-based) detection. In static detection, the detection will fail due to simple code obfuscation techniques, while in dynamic detection, it depends on the version of the operating system and the running state of the device, which requires a high detection cost. Therefore, although the malicious software data contains a large number of content features in the process of detecting software by the existing technology, these detection methods all ignore the implicit relationship between malicious software sample nodes; and most heterogeneous graph embedding methods only consider the influence of directly adjacent neighbor nodes, while ignoring the attributes of higher-order neighbors, resulting in incomplete acquisition of the content relevance of nodes, and there will be a semantic confusion phenomenon in the semantic propagation of deep neural networks; and when detecting unknown software, because it is necessary to re-perform heterogeneous graph embedding on the unknown software and re-train the downstream classifier at the same time, the detection speed of the model for unknown software is too slow. How to effectively learn the representation of unknown software in a heterogeneous graph is still a great difficulty. Summary of the Invention

[0003] Aiming at the problems in the prior art that the accuracy of the detection result is low due to the diversity of malicious software samples, the complexity of the relationship between samples, and the content-based relevance between malicious sample nodes, and the detection speed of unknown malicious software is slow, the present invention proposes a method for quickly detecting malicious software based on a graph neural network. The method includes: constructing a malicious software detection model; obtaining application data of the software to be detected, and inputting the application data of the software to be detected into the malicious software detection model to obtain a detection result;

[0004] The process of using the malicious software detection model to process the application data of the software to be detected includes:

[0005] S1. Using different meta-paths and meta-graphs to extract features from the input application program data of the software to be detected, and constructing a heterogeneous graph network according to the extracted features;

[0006] S2. Using the HG2vec algorithm to aggregate messages of higher-order neighbor nodes of software entity nodes in the heterogeneous graph network to obtain a higher-order enhanced embedding representation of the software node to be detected;

[0007] S3. Use the self-attention mechanism to perform semantic fusion on the high-order enhanced embedding representations passing through different meta-structures to obtain the final embedding vector representation of the fused nodes;

[0008] S4. Use the Sim2vec algorithm to classify and detect the final embedding vector representation of the software node to be tested to obtain the detection result of the software to be tested.

[0009] Preferably, the process of using different meta-paths and meta-graphs to extract features from the input application data to be detected includes:

[0010] S11. Obtain the data information of the software to be tested, and the data information includes code files, configuration files, and signature information; extract five types of entities in the data information as the key features of the software, and the five types of entities include APIs, permissions, components, signatures, and services;

[0011] S12. Construct different meta-structures according to the five types of entities of the software to be tested, obtain the adjacency matrices of each meta-structure, and perform dot multiplication on each adjacency matrix to obtain a heterogeneous graph network.

[0012] Preferably, the process of using the HG2vec algorithm to extract features from the hidden information includes:

[0013] S21. Use the mapping function composed of a single-layer neural network to project the hidden information of the high-order neighbor nodes in different meta-paths into the semantic space, and use the aggregation function to aggregate all the neighbor information projected into the semantic space to obtain node embeddings;

[0014] S22. Use the GAT method to calculate the weights between each node in the node embeddings, and perform Lipschitz normalization on the calculated weights to obtain weight coefficients;

[0015] S23. Aggregate the adjacent nodes in each node embedding according to the weight coefficients and node embeddings to obtain the aggregated node-level semantic features;

[0016] S24: Use the stacked node aggregation method to enhance the high-order attributes of the nodes to obtain the high-order embedding representation of the nodes based on a single meta-structure.

[0017] Preferably, the process of using the self-attention mechanism to fuse the high-order features passing through different meta-structures includes:

[0018] S31. Obtain the learning weights of each meta-structure and construct them into a weight matrix;

[0019] S32. Calculate the importance of each meta-structure according to the self-attention mechanism through the weight matrix;

[0020] S33. Fuse the high-order features of different meta-structures according to the importance of each meta-structure after that to obtain the fused node embedding vector representation.

[0021] Preferably, the process of classifying and detecting the fused known node embedding vector representation using the Sim2vec algorithm includes:

[0022] S41. Connect the relationship matrix of the software node to be detected with the relationship matrix of the known software node to obtain a new path from the software node to be detected to the known software node, and construct an incremental adjacency matrix with a new adjacency relationship according to the new path;

[0023] S42. Calculate the meta-structure similarity between the unknown node and the existing known nodes, and extract the top N most similar known node embedding representations;

[0024] S43. Aggregate the corresponding known nodes according to the weight of the similarity to obtain the node-level embedding representation of a single node;

[0025] S44. Perform semantic fusion on the node-level embedding representations of the top N single nodes according to the meta-structure weights of the existing model to obtain the node embedding representation of the unknown software, and the node embedding representation of the unknown software is the classification result.

[0026] Advantages of the present invention:

[0027] 1. This method applies the heterogeneous information network to the analysis of malicious code, providing a new method for feature construction and semantic description of malicious code analysis. And aiming at the problem that most methods designed for heterogeneous graphs cannot learn complex semantic representations, it proposes to use the embedding representation methods of meta-paths and meta-graphs to improve the expression of malicious software features, so as to more accurately explore the interaction relationships of content-based nodes.

[0028] 2. By integrating multiple node aggregation layers into the attention network, this method proposes a new method for high-order enhanced node embedding, so that the initial input node features can be gradually enhanced layer by layer based on node interactions in heterogeneous graph embedding. Especially by integrating multiple node layers, the model can obtain the attributes of higher-order neighbors, and in the node-level semantic propagation mechanism, a novel normalization layer is introduced for the attention-based neural network. The application of this normalization layer enhances the continuity of the self-attention layer, enabling the deep graph neural network to capture the features of each node and learn distinguishable node embeddings through a deeper neural network structure.

[0029] 3. This method constructs a fast detection model with incremental aggregation. When an unknown software sample node is first input into the model, the embedding representations of the unknown software node and similar known software sample nodes can be incrementally aggregated, and at the same time, the model parameters are fine-tuned, so as to improve the detection speed of unknown software.

[0030] 4. The present invention can be applied to the security protection of major software platforms or terminals, which helps to protect data security and personal information, and can also explore the impact of software node behavior data and relationship structures in the software network on detection. It can also enable the regulatory authorities to more accurately grasp the outbreak of malicious software and guide and control it. Brief Description of the Drawings

[0031] Figure 1 is a schematic diagram of a method for quickly detecting malicious software in an embodiment of the present invention;

[0032] Figure 2 is a flowchart of a method for quickly detecting malicious software in a preferred embodiment of the present invention;

[0033] Figure 3 is a schematic diagram of the meta-structure adopted in an embodiment of the present invention;

[0034] Figure 4 is a schematic diagram of node fusion adopted in an embodiment of the present invention;

[0035] Figure 5 is a schematic diagram of high-order node attribute enhancement in an embodiment of the present invention. Detailed Embodiments

[0036] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0037] In view of the problems of low detection accuracy and high detection cost in existing malware detection methods, the present invention proposes a fast malware detection model based on graph attention network considering high-order neighbor attributes and unknown application detection. First, taking advantage of the fact that heterogeneous information networks can discover hidden information by retaining comprehensive semantics and structures, a heterogeneous information network is used to construct the model. At the same time, a self-attention mechanism and multiple meta-structures are added to the model to obtain the implicit relationships between malware samples. Subsequently, multiple denoised node layers with attention mechanisms are stacked to obtain node embeddings after fusing high-order semantics. Finally, the node embeddings of unknown software and known software are incrementally aggregated to improve the detection speed of the proposed model. Based on this method, not only can the semantic relationships between software nodes be better mined, but also the detection speed of unknown malware can be improved.

[0038] A fast malware detection method based on graph neural network, the method includes: constructing a malware detection model; obtaining application data of the software to be detected, and inputting the application data of the software to be detected into the malware detection model to obtain a detection result;

[0039] As Figure 2 shown, the process of using the malware detection model to process the application data of the software to be detected includes:

[0040] S1. Extract features from the input application program data of the software to be detected using different meta-paths and meta-graphs, and construct a heterogeneous graph network according to the extracted features;

[0041] S2. Use the HG2vec algorithm to aggregate messages of high-order neighbor nodes of software entity nodes in the heterogeneous graph network to obtain a high-order enhanced embedding representation of the software node to be detected;

[0042] S3. Use the self-attention mechanism to perform semantic fusion on the high-order enhanced embedding representations passing through different meta-structures to obtain the final embedding vector representation of the fused nodes;

[0043] S4. Use the Sim2vec algorithm to perform classification detection on the final embedding vector representation of the software node to be tested to obtain the detection result of the software to be tested.

[0044] A specific implementation manner of a fast malware detection method based on graph neural network, such as Figure 1As shown, the rapid detection method of the malware requires inputting the software features of the malware; then, different meta-structures are used to extract the homogeneous graph feature space of the malware. In this feature space, the HG2vec algorithm is used to extract the high-order features of the software nodes; the self-attention mechanism is added to each meta-structure to obtain the node embedding vector representation after the meta-structure fusion; these vectors are input into the Sim2vec algorithm to obtain the binary classification result of the unknown nodes.

[0045] The specific implementation of a rapid detection method for malware includes: malware feature extraction; extracting the high-order feature attributes of software nodes according to the obtained software node behavior data; fusing the embedding vector representations of different meta-structures by using the self-attention mechanism; and using the Sim2vec algorithm to obtain the incremental learning embedding representation of unknown nodes.

[0046] In this embodiment, the process of malware feature extraction includes: in order to input the node network into the subsequent algorithm, it is necessary to convert the heterogeneous graph composed of various entities into a homogeneous graph containing only App nodes. The key operation is to merge the relationships between App entities and other entities into the combined connections between Apps. Specifically, given a meta-structure, the heterogeneous graph can be converted into an exclusive homogeneous graph, where each node has neighbor nodes specific to the meta-structure. In fact, the meta-path connects a pair of nodes with semantic relationships.

[0047] In this embodiment, to further enrich the meta-structure, the meta-graph is incorporated into the meta-path. The meta-graph can be used as an extended template to capture the combination of any existing semantic relationships between a pair of nodes. In fact, the meta-structure can serve as a semantic bridge for converting a heterogeneous graph into a homogeneous graph, where all nodes satisfy specific complex semantics. It can be said that according to different meta-structures, nodes will have different structural relationships in different graphs. To some extent, each graph can be regarded as a sub-graph of the whole under a specific view - each sub-graph satisfies the semantic constraints given by the meta-structure.

[0048] A series of matrix operations are used to calculate the adjacency degree of nodes in the graph. For a given meta-path MP(A1,..., An), the adjacency matrix can be calculated by the following formula:

[0049]

[0050] where, is the relationship matrix between entity A j and A j+1 (for example, in the meta-path PID1: the adjacency matrix of the graph under A-API-A) ψ ij >0 indicates that App i and App jAre interconnected, i.e., they are neighbors based on the meta-path PID1. Specifically, this value represents the count of meta-path instances between nodes i and j, i.e., the number of paths. Similarly, for a given meta-graph MG, which represents a combination of several meta-paths, i.e., (MP1,..., MP m ), the node adjacency matrix is:

[0051]

[0052] where ⊙ represents the Hadamard product. For example, the meta-graph PID6 can be calculated as By graph modeling each meta-structure, the original App heterogeneous graph is transformed into multiple App homogeneous graphs, each graph belonging to an adjacency matrix given a K - meta-structure, i.e., a set of K adjacency relation matrices, i.e.,

[0053] This embodiment provides a method for extracting high - order feature attributes of software nodes according to the obtained software node behavior data, including: Given a meta - path Φ, first, a mapping function f Φ is given to project the App nodes into the semantic space. Then, an aggregation function g Φ aggregates the neighbor information of the App nodes based on the meta - path and meta - graph to learn specific node embeddings:

[0054] Z Φ = g Φ (f Φ (X))

[0055] where X represents the initial feature matrix, and Z Φ represents the semantic - specific node embedding. To process the heterogeneous graph, the semantic projection function f Φ projects the nodes into the semantic space as follows:

[0056] f Φ (X)= σ(X · W Φ + b Φ )

[0057] Use GAT to determine the weights between various nodes. Given a pair of nodes (i, j) connected by the meta - path Φ, the node - level attention can understand the importance of node j to node i. The importance of the node pair (i, j) based on the meta - path can be expressed as follows:

[0058]

[0059] where attnode represents the deep neural network that performs node - level attention.

[0060] Calculate the node of wherein denotes the meta-path-based neighbors (including itself) of node i. After obtaining the importance between node pairs based on the meta-path, it is normalized by the softmax function to obtain the weight coefficients

[0061]

[0062] where σ represents the activation function, || represents the concatenation operation, and a Φ is the node-level attention vector of the meta-path Φ. The weight coefficients of (i, j) depend on their features. The meta-path-based embedding of node i can be aggregated through the projected features of its neighbors, and the corresponding coefficients are as follows:

[0063]

[0064] where is the learned embedding of node i of the meta-path Φ. Each node embedding is aggregated by its adjacent nodes. Since the attention weights are generated for a single meta-path, it is semantic-specific and can extract a kind of semantic information. Due to the scale-free property of the heterogeneous graph, the variance of the graph data is large.

[0065] In this embodiment, the node-level attention is extended to multi-head attention, so that the training process is more stable. Specifically, in order to enhance the attention model, multiple independent attention models are concatenated. The multi-head model can thus simultaneously focus on multiple directions of the input space and is usually more powerful in practice. The standard process includes first projecting the input vector into multiple low-dimensional spaces, and then combining the results of all attention layers using a linear function. Let d I (or d O ) be the input (or output) dimension, h be the number of attention heads, and are matrices of dimension h + 1; all the matrices are concatenated row by row, that is:

[0066] MultiAtt(X) = W o (Att(W1X) ||... || Att(W h X))

[0067] In this embodiment, the Lipschitz constant of the multi-head attention will be used as the constant bound for each attention head, so as to calculate the coefficients of each attention head respectively. Specifically, if each attention head is Lipschitz continuous, then the multi-head attention is also Lipschitz continuous and can be defined as:

[0068]

[0069] The process of processing data by LipschitzNorm includes: given an attention model M with a scoring function g, define M-Lip as the updated attention model to which LipschitzNorm is applied. The specific steps include:

[0070] Step 1: The Frobenius norm of the query; that is, calculate the Frobenius norm of Q:

[0071]

[0072] Step 2: Calculate the maximum value of the 2-norm (Euclidean distance) of the input vector based on the Frobenius norm of Q: v = max i ||x i ||2;

[0073] Step 3: Scale the function according to the calculated data, that is, divide the score function by the product uv.

[0074] Each attention head is processed separately, so each head adopts all the norms and maximum values. In addition, in the case of graph attention, the norms and maximum values are calculated for neighbors, that is, for each node, calculate the maximum value of the 2-norm of its neighbors.

[0075] Given the meta-path sets Φ0, Φ1,..., Φ P , after inputting the node features into the node-level attention, P groups of semantics-specific node embeddings are obtained, denoted as

[0076] In this embodiment, the HG2vec framework stacks multiple node-level aggregations together. It constructs a high-order attribute enhancement architecture to enhance the input node features in a hierarchical manner through content-based node interactions. Take the 4th-order HG2vec as a specific example to illustrate how the high-order architecture works. By stacking three NALs, HG2vec (4) can capture the interactions between the first-order content-based and the high-order semantic structure-based neighborhoods.

[0077] NAL enhances the embedding of each App by embedding its first-order semantic structure based on the neighborhood. Therefore, as Figure 5As shown, the first NAL enhances App1 using App2 and App6, and enhances App6 using App1 and App7. Then, the second NAL repeats this process and enhances App1 with App3. At the same time, App3 is enhanced with App2 and App8 through the first NAL. Therefore, the second NAL indirectly enhances App1 through App3. Similarly, the third NAL enhances App1 with App4. In this way, HGat2vec (3) combines the embedding of App1 with the embeddings of its first-order (App2 and App6), second-order (App3), and third-order (App4) neighbors. As Figure 3 shown, HGat2vec (3) model takes the initial feature matrix of the target type nodes as input and outputs the enhanced node embedding representation.

[0078] This embodiment discloses a method for fusing the embedding vector representations of different meta-structures using the self-attention mechanism, including: Each node in the heterogeneous graph contains multiple types of semantic information, and the node embedding of a specific semantics can only reflect the node from one aspect. To learn more comprehensive node embeddings, the present invention reveals multiple semantics by fusing multiple meta-structures. To address the challenges of meta-path and meta-graph selection and semantic fusion in heterogeneous graphs, the self-attention mechanism is used to automatically learn the importance of different meta-structures and fuse them into specific tasks. Taking k groups of semantics-specific node embeddings learned from node-level aggregation as input, each meta-structure has the following learning weights:

[0079]

[0080] where att sem represents the semantic-level deep neural network with the self-attention mechanism added. To understand the importance of each meta-path, the semantic embedding is first transformed through a non-linear transformation (e.g., one layer of MLP), and the semantic-level attention vector q is used to measure the importance of the semantics-specific embedding, that is, the similarity of the transformed embedding. When averaging the importance of all semantics-specific node embeddings, it can be interpreted as the importance of each meta-path; the formula for the importance of each meta-path is:

[0081]

[0082] where W is the weight matrix, b is the bias vector, and q is the attention vector at the semantic level. Note that all the above parameters are shared among all meta-structures and specific semantics embeddings. After obtaining the importance of each meta-structure, they are normalized through the softmax function. The weight of the meta-path Φ i is denoted as , which can be obtained by normalizing the importance of all meta-structures using the softmax function:

[0083]

[0084] According to the above formula, it can be known that the higher, the more important the meta-path Φ i For different tasks, the meta-path Φ i may have different weights. Using the learned weights as coefficients, the semantic-specific embeddings are fused to obtain the final embedding Z, and its expression is:

[0085]

[0086] To better understand the aggregation process of the semantic layer, as Figure 4 shown, the final embedding representation is aggregated from all semantic-specific embeddings. According to the final embedding, it is applied to a specific task, and different loss functions are designed. For semi-supervised node classification, the cross-entropy of all labeled nodes between the ground truth and the prediction can be minimized:

[0087]

[0088] where C is the parameter of the classifier, and Y l is the set of node indices with labels, Y l and Z l are the labels and embeddings of the labeled nodes. Under the guidance of the labeled data, the proposed model can be optimized by backpropagation, and the embeddings of the nodes can be learned.

[0089] In this embodiment, the process of using the Sim2vec algorithm to obtain the incremental learning embedding representation of unknown nodes includes: To better represent the unknown Apps not included in the training process, the Sim2vec algorithm is adopted, and an incremental learning mechanism that uses the in-sample App embedding representation learned from HG2vec to quickly represent those out-of-sample Apps. For clarity, use App out to represent any out-of-sample node in the HIN.

[0090] Accurately finding the potential connections between unknown nodes and existing nodes in the HIN plays a key role in the fast embedding representation for malware detection. For this purpose, it is necessary to calculate and accumulate the similarity between App out and existing nodes. Given the node similarity between node v i and node v j under a meta-path is defined as:

[0091]

[0092] Among them, represents the number of meta-structures between two connected nodes. Therefore, a higher similarity indicates a closer association between these two nodes. Thus, for node v under the meta-graph MG i and node v j the node similarity between them is:

[0093]

[0094] The main task is to capture incremental relationships and construct graph information. In the given meta-structure, only the adjacency matrix is updated, which quantifies the connections between out-of-sample nodes and existing in-sample App nodes. To reduce the training cost, it is carried out in an incremental manner. The relationship matrix of the new App node is concatenated with the relationship matrix of the existing App nodes to form an incremental segment of the node adjacency -- the path from in-sample App nodes to the new node. Taking PID1 as an example; first obtain the relationship matrix A out of all new nodes, and then generate the matrix through . This design ensures that the incremental adjacency matrix can function independently of the established adjacency matrix , while together they serve as an overall abstraction of the connectivity between all nodes.

[0095] Using Sim2vec can calibrate the existing node representations while also being able to embed new nodes. Similar to HG2vec, this model consists of two steps: node-level aggregation and semantic-level fusion. Given the semantic meta-structure M k , substitute into the similarity function formula to calculate Sim Mk (v j , v out ), that is, the similarity between the new node v out and any example application node v j . Repeating this operation for all out-of-sample nodes and in-sample application nodes will form a similarity matrix X Mk , where larger values essentially indicate that the two nodes are closer. Therefore, a set of similarity matrices for all meta-structures {X M1 ,..., X Mk} can be obtained.

[0096] To better represent the new node in the numerical vector, the existing embedding results of the existing nodes that are very close to the new node should be fully aggregated. For this purpose, based on the similarity matrix select the top-σ in the sample App nodes and aggregate their vectors to embed the new node:

[0097]

[0098] Among them represents a node in the meta-structure M k of the weight, represents the incremental embedding information of the out-of-sample node. The weight can be calculated as follows:

[0099]

[0100] By internally aggregating the K individual representations under all meta-structures to recalibrate the embedding:

[0101]

[0102] wherein, can be obtained from the attention formula.

[0103] The above-described embodiments further illustrate in detail the object, technical solution and advantages of the present invention. It should be understood that the above-described embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for quickly detecting malware based on a graph neural network, characterized in that, The method includes: constructing a malware detection model; obtaining the application data of the software to be detected, inputting the application data of the software to be detected into the malware detection model, and obtaining a detection result; The process of using the malware detection model to process the application data of the software to be detected includes: S1. Extract features from the input application program data of the software to be detected using different meta-paths and meta-graphs, and construct a heterogeneous graph network based on the extracted features; specifically including: S11. Obtain the data information of the software to be detected, where the data information includes code files, configuration files, and signature information; extract five types of entities in the data information as the key features of the software, and the five types of entities include APIs, permissions, components, signatures, and services; S12. Construct different meta-structures based on the five types of entities of the software to be detected, obtain the adjacency matrix of each meta-structure, and perform dot multiplication on each adjacency matrix to obtain a heterogeneous graph network; S2. Use the HG2vec algorithm to aggregate the high-order neighbor nodes of the software entity nodes in the heterogeneous graph network to obtain the high-order enhanced embedding representation of the software nodes to be detected; specifically including: S21. Use a mapping function composed of a single-layer neural network to project the hidden information of the high-order neighbor nodes in different meta-paths into the semantic space, and use an aggregation function to aggregate all the neighbor information projected into the semantic space to obtain node embeddings; S22. Use the GAT method to calculate the weights between each node in the node embeddings, and perform Lipschitz normalization on the calculated weights to obtain weight coefficients; S23. Aggregate the adjacent nodes in each node embedding according to the weight coefficients and node embeddings to obtain the aggregated node-level semantic features; S24: Use the stacked node aggregation method to enhance the high-order attributes of the nodes to obtain the high-order embedding representation of the nodes based on a single meta-structure; S3. Use the self-attention mechanism to perform semantic fusion on the high-order enhanced embedding representations passing through different meta-structures to obtain the final embedding vector representation of the fused nodes; S4. Use the Sim2vec algorithm to perform classification detection on the final embedding vector representation of the software node to be tested to obtain the detection result of the software to be tested.

2. The malware rapid detection method based on a graph neural network according to claim 1, characterized in that The formula for obtaining node embeddings is: Among them, Z Φ represents the node embedding with specific semantics, g Φ represents the aggregation function, f Φ represents the mapping function, X represents the initial feature matrix, σ represents the activation function, W Φ represents the weight matrix of the meta-path Φ, b Φ represents the bias vector of the meta-path Φ.

3. The malware rapid detection method based on graph neural network according to claim 1, characterized in that The process of using the stacked node aggregation method to enhance the high-order attributes of the nodes includes: S241. Input the initial feature matrix of the target type nodes into the HG2vec framework, perform enhanced representation on its first-order neighbor nodes, and use the enhanced node embedding representation as the output; S242. Use the output of the second node attention layer NAL as the input of the next NAL, and stack N layers in this way to obtain the enhanced embedding representation of the high-order neighbor nodes of a single node.

4. A rapid malware detection method based on graph neural network according to claim 1, characterized in that, The expression of the weight coefficient is: Among them, represents the importance between the i-th node and the j-th node, and σ represents the activation function, represents the node-level attention vector of the meta-path Φ, h′ i represents the projected feature of the i node, || represents the concatenation operation, h′ j represents the projected feature of the j node, represents the meta-structure-based neighbor of node i.

5. A method for quickly detecting malware based on a graph neural network according to claim 1, characterized in that, The process of using the self-attention mechanism to fuse the high-order features passing through different meta-structures includes: S31. Obtain the learning weights of each meta-structure and construct them into a weight matrix; S32. Calculate the importance of each meta-structure according to the self-attention mechanism through the weight matrix; S33. Fuse the high-order features of different meta-structures according to the importance weights of each calculated meta-structure to obtain the fused node embedding vector representation.

6. The malware rapid detection method based on graph neural network according to claim 5, characterized in that The importance formula of the meta-path is: Among them, V represents all specific semantic nodes, q T represents the attention vector at the semantic level, tanh represents the activation function, W is the weight matrix, z i Φ represents the high-order enhanced representation vector of the node, and b is the bias vector.

7. A method for rapid malware detection based on graph neural network according to claim 1, characterized in that The process of using the Sim2vec algorithm to quickly classify and detect unknown nodes includes: S41. Connect the relationship matrix of the software node to be detected with the relationship matrix of the known software node to obtain a new path from the software node to be detected to the known software node, and construct an incremental adjacency matrix with a new adjacency relationship according to the new path; S42. Calculate the meta-structure similarity between the unknown node and the existing known nodes, and extract the embedding representations of the top N most similar known nodes; S43. Aggregate the corresponding known nodes according to the weight of the similarity to obtain the node-level embedding representation of a single node; S44. Perform semantic fusion on the node-level embedding representations of the top N single nodes according to the meta-structure weights of the existing model to obtain the node embedding representation of the unknown software, and the node embedding representation of the unknown software is the classification result.

Citation Information

Patent Citations

  • Method of detecting android malware based on heterogeneous graph and apparatus thereof

    US20240241954A1