Malicious software security detection method based on graph neural network
By combining graph neural networks and quasi-Newton methods, a software graph structure is constructed and jointly optimized, which solves the problem of insufficient generalization ability of malware detection methods under diverse and mutated conditions, and realizes efficient and intelligent malware detection, adapting to complex security detection scenarios.
Patent Information
- Application Number
- CN202511146070.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-28
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing malware detection methods based on graph neural networks lack generalization ability and robustness when facing diverse and mutated malware, making it difficult to achieve efficient fusion and adaptive adjustment of multi-source heterogeneous features, resulting in a tradeoff between detection accuracy and real-time performance.
By employing graph neural networks combined with the quasi-Newton method, a software graph structure is constructed to extract multi-dimensional features. The weight parameters and structural hyperparameters are jointly optimized, and multi-scale neighborhood aggregation, heterogeneous edge type weight allocation, and node-edge interaction gating are used to achieve intelligent identification and adaptive detection of malware.
It significantly improves the accuracy and efficiency of malware detection, can efficiently identify unknown, mutated and obfuscated samples, has good generalization ability and robustness, adapts to large-scale malware detection scenarios, and meets the high security requirements of practical applications.
Smart Images

Figure CN121030741A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network and information security technology, and in particular to a method for security detection of malware based on graph neural networks. Background Technology
[0002] Currently, the number and types of malware are growing rapidly, posing a significant threat to global information security. Traditional malware detection methods primarily include signature-based static detection and behavior-based dynamic detection. Static detection methods typically identify malicious samples by extracting features from binary files, instruction sequences, API call lists, or signature information, and comparing them with known malware databases. While static detection methods are highly efficient, their effectiveness decreases significantly for obfuscated, packed, or mutated malware samples, making them susceptible to evasion by adversarial samples. Dynamic detection, on the other hand, captures malicious behavior patterns by monitoring software execution behavior, system calls, file operations, and network communications. However, dynamic detection methods are resource-intensive, have long detection cycles, and struggle to balance accuracy and real-time performance when facing highly concealed or delayed-triggered malicious samples.
[0003] With the development of artificial intelligence and deep learning technologies, more and more research is attempting to leverage neural network models to improve the intelligence and automation of malware detection. Among these, graph neural networks (GNNs) have become an important research direction in the field of intelligent detection due to their ability to effectively model complex call relationships, dependencies, and behavioral chains within software samples. However, existing GNN-based detection methods still suffer from several prominent problems. On the one hand, most current models rely on fixed network structures and hyperparameters set by human experience, making it difficult to adapt to the structural differences and behavioral variations of different types of malicious samples, resulting in limited generalization ability and robustness. On the other hand, traditional optimization methods such as first-order gradient descent are prone to getting trapped in local optima and have slow convergence speeds in large-scale, high-dimensional parameter spaces, making it difficult to achieve efficient synergistic optimization of weight parameters and structural hyperparameters, thus affecting overall detection performance.
[0004] The diversification of malware and the increasing sophistication of attack techniques place higher demands on the adaptability and self-learning capabilities of detection systems. Existing methods generally lack the ability to intelligently identify and adaptively adjust to malware variants, unknown families, and obfuscated samples, and cannot fully explore and integrate multi-source heterogeneous features such as static and dynamic data, thus limiting the detection effectiveness and application scope in practical scenarios.
[0005] Therefore, how to provide a malware security detection method based on graph neural networks is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0006] One objective of this invention is to propose a malware security detection method based on graph neural networks. This invention fully utilizes the modeling ability of graph neural networks to model the complex structure of static and dynamic behavior of software, and integrates the quasi-Newton method for the joint optimization mechanism of network weight parameters and structural hyperparameters. It describes in detail the method for adaptive extraction and efficient classification of multi-dimensional software sample features, which has the advantages of high detection accuracy, strong ability to identify unknown and mutated malware, high optimization efficiency, and good model generalization ability.
[0007] A malware security detection method based on graph neural networks according to an embodiment of the present invention includes the following steps:
[0008] S1. Collect static and dynamic behavior data of the software to be tested, and construct a software graph structure. The nodes of the software graph structure represent functions, APIs, files or behavior events, and the edges represent calling, dependency or time sequence relationships.
[0009] S2. Standardize and feature-encode the software graph structure, extract node attributes, edge attributes and global structural features to form multi-dimensional graph data input samples;
[0010] S3. Based on the multi-dimensional graph data input samples, initialize the graph neural network model, set the structural hyperparameters of the graph neural network model, and initialize the weight parameters;
[0011] S4. Input the multi-dimensional graph data input samples into the initialized graph neural network model, perform forward propagation, and obtain the embedding features and preliminary classification results of each software sample;
[0012] S5. The quasi-Newton method is used to jointly optimize the weight parameters and structural hyperparameters of the graph neural network model. During the training process of the graph neural network model, the structure of the graph neural network model is dynamically adjusted according to the performance index and the parameters are continuously optimized.
[0013] S6. Process the new sample data of the software to be detected according to the process of steps S1 to S4, input it into the optimized graph neural network model, output the malware detection results, and realize the intelligent identification of unknown, mutated and confused samples.
[0014] Optionally, static behavioral data specifically includes the binary code, file structure, API call list, instruction sequence, and resource dependency information of the software under test, while dynamic behavioral data specifically includes system calls, process behavior, network communication, memory operations, and file read / write activities of the software under test during operation.
[0015] Optionally, S2 specifically includes:
[0016] S21. For each node v in the software graph structure iExtract node attribute vector x i The node attribute vector includes function type, API call frequency, file type, and behavior event type;
[0017] S22. For each edge e in the software graph structure ij Extracting edge attribute vector e ij The edge attribute vector includes the number of calls, call weight, and the order of call times;
[0018] S23. Based on the adjacency matrix A∈{0,1} N×N This represents the connection relationships in a software graph structure, where N is the total number of nodes. If node v i With node v j If there is an edge between them, then A ij =1, otherwise A ij =0;
[0019] S24. For the node attribute matrix X = [x1, x2, ..., x...] N ] T And edge attribute tensor E = [e ij Normalization is performed to ensure that the features of each dimension of the node attribute vector and edge attribute vector have a uniform scale;
[0020] S25. Extract the global structural features of the software graph structure. g This includes obtaining the number of connections for each node by calculating the degree matrix D, where D is a diagonal matrix, and the element in the i-th row and i-th column is D. ii , representing node v in the software graph structure i The number of directly connected edges, combined with node attribute vectors, edge attribute vectors, and adjacency matrix information, are used to construct a multi-dimensional graph data input sample containing node attributes, edge attributes, and global structural features.
[0021] Optionally, S3 specifically includes:
[0022] S31. Receive the obtained multi-dimensional graph data input sample. Where A is the adjacency matrix, X is the node attribute matrix, E is the edge attribute tensor, and S... g For global structural features;
[0023] S32. Construct the basic structure of the graph neural network model. The graph neural network model includes an input layer, several graph neural network hidden layers and an output layer. The input layer receives multi-dimensional graph data input samples. The hidden layers perform aggregation and transformation of node and edge features in sequence. The output layer generates malware detection and classification results.
[0024] S33. Set an aggregation function for each hidden layer. The aggregation function performs joint aggregation of node attributes and edge attributes to generate node feature vectors.
[0025] S34. Set the number of layers L of the graph neural network model, and configure the activation function, number of hidden layer units, and number of node samples for each hidden layer.
[0026] S35. Introduce a multi-scale neighborhood aggregation radius hyperparameter r in each hidden layer. (l) The neighborhood range of node information aggregation at each layer can be dynamically controlled. The aggregation radius of different layers can be set separately, and the initial value can be randomly generated or given by preset rules.
[0027] S36. Set the heterogeneous edge type weight allocation coefficient α for each edge type in the software graph structure. t In the process of node aggregation, the information contribution of different types of edges is weighted, and the initial value can be set according to prior experience or uniform distribution.
[0028] S37. Introduce a node-edge interaction gating weight parameter β in each hidden layer. (l) The fusion ratio of node features and edge features in each layer to the final node representation is adaptively adjusted, and the initial value can be randomly set in the range of 0 to 1.
[0029] S38. Initialize the weight parameters W of each layer in the graph neural network model. (l) W (l) Linear transformations are used for node and edge features, with initial weight parameters randomly initialized using a Gaussian distribution;
[0030] S39. The hyperparameter r of the multi-scale neighborhood aggregation radius (l) Heterogeneous edge type weight allocation coefficient α t Gating weight parameter β for node-edge interaction (l) Together they are used as structural hyperparameters and weight parameters;
[0031] S310. Initialize the graph neural network model, which includes the multi-scale neighborhood aggregation radius hyperparameter, heterogeneous edge type weight allocation coefficients, node-edge interaction gating weight parameters, and weight parameters. Then, initialize the node feature vectors output from the last layer. It serves as the input basis for forward propagation and malware sample embedding features and classification.
[0032] Optionally, S4 specifically includes:
[0033] S41. Receive the initialized graph neural network model, which includes the multi-scale neighborhood aggregation radius hyperparameter r. (l) Heterogeneous edge type weight allocation coefficient α tNode-edge interaction gating weight parameter β (l) and the weight parameters W of each layer (l) ;
[0034] S42. Input the obtained multidimensional graph data into the sample. The input is fed into the input layer of the initialized graph neural network model, where A is the adjacency matrix, X is the node attribute matrix, E is the edge attribute tensor, and S... g For global structural features;
[0035] S43. In each hidden layer l, perform node neighborhood information feature fusion for each node v according to the aggregation function, and update the node feature vector.
[0036] S44. Perform the aggregation operation and feature transformation described in S43 sequentially on all hidden layers to finally obtain the feature vector of each node in the last layer. Where L is the number of layers in the graph neural network model;
[0037] S45. Employ graph-level readout operations to extract the last layer feature vector of each node. Global pooling is used to obtain the embedding features z of each software sample. G ;
[0038] S46. Input the embedded features of each software sample into the output layer, perform a linear transformation on the embedded features through a fully connected layer to obtain the score of each category, and then input the score of each category into the Softmax function to normalize all category scores and convert them into a probability distribution. The probability value represents the likelihood of each software sample belonging to each category. Finally, output the preliminary classification result of each software sample according to the category corresponding to the highest probability.
[0039] Optionally, S5 specifically includes:
[0040] S51. Define the loss function The loss function combines the cross-entropy loss and regularization term between the initial classification result of each software sample and the true class label;
[0041] S52, the weight parameters W of the graph neural network model (l) Multi-scale neighborhood aggregation radius hyperparameter r (l) Heterogeneous edge type weight allocation coefficient α t and node-edge interaction gating weight parameter β (l) The parameters to be optimized are uniformly composed into a vector θ.
[0042] S53. Employing a sparse approximation and low-rank update mechanism, only the Hessian matrix approximation B is retained in each iteration. k Principal components or significantly nonzero elements;
[0043] S54. Employ a block gradient collaborative mechanism to adjust the current parameter vector θ. k gradient Grouping calculations are performed, dividing all parameters into several sub-blocks. Each sub-block independently calculates its local gradient and local sparse low-rank Hessian matrix approximation. At the global level, the optimization directions of all sub-blocks are merged through weighted summation or principal component aggregation to obtain the global optimization direction d. k ;
[0044] S55. Adopt an adaptive line search and dynamic step size strategy. Based on the current rate of decline of the loss function, the change of the gradient, and the optimization results of the previous round of graph neural network model, dynamically evaluate and select the step size of the parameter update in this round. When the loss function decreases significantly, the step size is increased appropriately. When the loss function fluctuates or oscillates, the step size is decreased appropriately. The step size adjustment process combines threshold judgment and historical step size reference.
[0045] S56. Based on the obtained global optimization direction and the determined step size, adjust the parameter vector to be optimized item by item. For each weight parameter and structural hyperparameter, according to the increasing or decreasing trend indicated by the optimization direction, add the corresponding step size value in sequence. After all parameters are updated, immediately use them for the next round of graph neural network forward propagation and performance evaluation.
[0046] S57. Apply the updated parameters to the graph neural network model, perform forward propagation, obtain the embedding features and preliminary classification results for each software sample, and recalculate the loss function and graph neural network model performance metrics.
[0047] S58. Based on the performance metrics of the graph neural network model, namely accuracy, loss value, and recall, dynamically adjust the structure of the graph neural network model, including adjusting the number of network layers, multi-scale neighborhood aggregation radius, number of node samples, number of hidden layer units, and activation function.
[0048] S59. Parallel and distributed quasi-Newton optimization mechanisms are adopted to divide the parameter space or training data into multiple sub-blocks. The gradients and approximations of the Hessian matrix of each sub-block are calculated synchronously or asynchronously in a multi-core or distributed computing environment, and the results are updated through distributed synchronous fusion.
[0049] S510 continuously executes steps such as sparse approximation and low-rank update, adaptive line search and dynamic step size, block gradient collaboration and parallel distributed optimization, and iteratively optimizes parameters and structure until the loss function converges or reaches the preset number of training rounds, and finally obtains the optimal graph neural network model parameters and structure configuration.
[0050] Optionally, S6 specifically includes:
[0051] S61. Collect new sample data of the software to be tested, including static behavior data and dynamic behavior data, and construct the corresponding software graph structure;
[0052] S62. Standardize and feature-encode the constructed software graph structure, extract node attributes, edge attributes and global structural features to form multi-dimensional graph data input samples;
[0053] S63. Input the obtained multi-dimensional graph data input samples into the graph neural network model that has been optimized in terms of parameters and adjusted in terms of structure;
[0054] S64. Using the optimized graph neural network model, perform forward propagation on the input multi-dimensional graph data input samples to obtain the embedding features and preliminary classification results of each new software sample;
[0055] S65. Based on the embedding features and preliminary classification results of each new software sample, output the malware detection results to achieve intelligent identification of unknown, mutated and confused samples.
[0056] The beneficial effects of this invention are:
[0057] The malware security detection method proposed in this invention, based on a combination of graph neural networks and quasi-Newton methods, significantly improves the overall performance of malware detection systems. When modeling the complex static and dynamic behavioral relationships of malware samples, this invention fully utilizes the deep fusion capability of graph neural networks for multi-source heterogeneous features, achieving accurate representation of call relationships, dependency structures, and behavioral links. By introducing innovative structural hyperparameters such as multi-scale neighborhood aggregation, heterogeneous edge type weight allocation, and node-edge interaction gating, this invention effectively enhances the model's adaptability to diverse malware structural and behavioral variations. The joint optimization of the weight parameters and structural hyperparameters of the graph neural network using quasi-Newton methods not only improves the model's convergence speed and global optimum search capability but also achieves adaptive dynamic adjustment of network structure and parameters, avoiding the limitations of traditional optimization methods that easily fall into local optima.
[0058] In actual detection processes, this invention achieves high-accuracy intelligent identification of unknown, mutated, and obfuscated samples, significantly improving the detection and defense capabilities against novel malware. The optimized graph neural network model possesses excellent generalization ability and robustness, not only compatible with diverse sample structures but also adaptable to large-scale malware detection scenarios, meeting the high efficiency and high security requirements of practical applications. This invention provides a novel malware detection technology solution for the field of network and information security that is efficient, intelligent, and possesses self-learning capabilities, demonstrating outstanding technological advancement and application promotion value. Attached Figure Description
[0059] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0060] Figure 1 This is a flowchart of a malware security detection method based on graph neural networks proposed in this invention;
[0061] Figure 2 This is a schematic diagram of the graph neural network model structure and parameter optimization process of the malware security detection method based on graph neural networks proposed in this invention. Detailed Implementation
[0062] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0063] refer to Figure 1 and Figure 2 A malware security detection method based on graph neural networks includes the following steps:
[0064] S1. Collect static and dynamic behavior data of the software to be tested, and construct a software graph structure. The nodes of the software graph structure represent functions, APIs, files or behavior events, and the edges represent calling, dependency or time sequence relationships.
[0065] S2. Standardize and feature-encode the software graph structure, extract node attributes, edge attributes and global structural features to form multi-dimensional graph data input samples;
[0066] S3. Based on the multi-dimensional graph data input samples, initialize the graph neural network model, set the structural hyperparameters of the graph neural network model, and initialize the weight parameters;
[0067] S4. Input the multi-dimensional graph data input samples into the initialized graph neural network model, perform forward propagation, and obtain the embedding features and preliminary classification results of each software sample;
[0068] S5. The quasi-Newton method is used to jointly optimize the weight parameters and structural hyperparameters of the graph neural network model. During the training process of the graph neural network model, the structure of the graph neural network model is dynamically adjusted according to the performance index and the parameters are continuously optimized.
[0069] S6. Process the new sample data of the software to be detected according to the process of steps S1 to S4, input it into the optimized graph neural network model, output the malware detection results, and realize the intelligent identification of unknown, mutated and confused samples.
[0070] In this embodiment, static behavioral data specifically includes the binary code, file structure, API call list, instruction sequence, and resource dependency information of the software under test, while dynamic behavioral data specifically includes system calls, process behavior, network communication, memory operations, and file read / write activity information of the software under test during operation.
[0071] In this embodiment, S2 specifically includes:
[0072] S21. For each node v in the software graph structure i Extract node attribute vector x i The node attribute vector includes function type, API call frequency, file type, and behavior event type;
[0073] S22. For each edge e in the software graph structure ij Extracting edge attribute vector e ij The edge attribute vector includes the number of calls, call weight, and the order of call times;
[0074] S23. Based on the adjacency matrix A∈{0,1} N×N This represents the connection relationships in a software graph structure, where N is the total number of nodes. If node v i With node v j If there is an edge between them, then A ij =1, otherwise A ij =0;
[0075] S24. For the node attribute matrix X = [x1, x2, ..., x...] N ] T And edge attribute tensor E = [e ij Normalization is performed to ensure that the features of each dimension of the node attribute vector and edge attribute vector have a uniform scale;
[0076] S25. Extract the global structural features of the software graph structure. g This includes obtaining the number of connections for each node by calculating the degree matrix D, where D is a diagonal matrix, and the element in the i-th row and i-th column is D. ii , representing node v in the software graph structure i The number of directly connected edges, combined with node attribute vectors, edge attribute vectors, and adjacency matrix information, are used to construct a multi-dimensional graph data input sample containing node attributes, edge attributes, and global structural features.
[0077] In this embodiment, S3 specifically includes:
[0078] S31. Receive the obtained multi-dimensional graph data input sample. Where A is the adjacency matrix, X is the node attribute matrix, E is the edge attribute tensor, and S...g For global structural features;
[0079] S32. Construct the basic structure of the graph neural network model. The graph neural network model includes an input layer, several graph neural network hidden layers and an output layer. The input layer receives multi-dimensional graph data input samples. The hidden layers perform aggregation and transformation of node and edge features in sequence. The output layer generates malware detection and classification results.
[0080] S33. Set an aggregation function for each hidden layer. The aggregation function performs joint aggregation of node attributes and edge attributes to generate node feature vectors.
[0081]
[0082] Where σ is the activation function, W (l) For the weight parameters of the l-th layer, AGGREGATE (l) () represents the aggregation operation, α t β is the weighting coefficient for heterogeneous edge types, where t is the edge type. (l) For node-edge interaction gating weight parameters, e represents the feature representation of neighbor node u in the previous layer. uv Let be the edge attribute vector between node u and node v. With v as the center and r as the radius (l) The set of neighboring nodes, r (l) The hyperparameter for the multi-scale neighborhood aggregation radius;
[0083] S34. Set the number of layers L of the graph neural network model, and configure the activation function, number of hidden layer units, and number of node samples for each hidden layer.
[0084] S35. Introduce a multi-scale neighborhood aggregation radius hyperparameter r in each hidden layer. (l) The neighborhood range of node information aggregation at each layer can be dynamically controlled. The aggregation radius of different layers can be set separately, and the initial value can be randomly generated or given by preset rules.
[0085] S36. Set the heterogeneous edge type weight allocation coefficient α for each edge type in the software graph structure. t In the process of node aggregation, the information contribution of different types of edges is weighted, and the initial value can be set according to prior experience or uniform distribution.
[0086] S37. Introduce a node-edge interaction gating weight parameter β in each hidden layer. (l) The fusion ratio of node features and edge features in each layer to the final node representation is adaptively adjusted, and the initial value can be randomly set in the range of 0 to 1.
[0087] S38. Initialize the weight parameters W of each layer in the graph neural network model. (l) W (l) Linear transformations are used for node and edge features, with initial weight parameters randomly initialized using a Gaussian distribution;
[0088] S39. The hyperparameter r of the multi-scale neighborhood aggregation radius (l) Heterogeneous edge type weight allocation coefficient α t Gating weight parameter β for node-edge interaction (l) Together they are used as structural hyperparameters and weight parameters;
[0089] S310. Initialize the graph neural network model, which includes the multi-scale neighborhood aggregation radius hyperparameter, heterogeneous edge type weight allocation coefficients, node-edge interaction gating weight parameters, and weight parameters. Then, initialize the node feature vectors output from the last layer. It serves as the input basis for forward propagation and malware sample embedding features and classification.
[0090] In this embodiment, S4 specifically includes:
[0091] S41. Receive the initialized graph neural network model, which includes the multi-scale neighborhood aggregation radius hyperparameter r. (l) Heterogeneous edge type weight allocation coefficient α t Node-edge interaction gating weight parameter β (l) and the weight parameters W of each layer (l) ;
[0092] S42. Input the obtained multidimensional graph data into the sample. The input is fed into the input layer of the initialized graph neural network model, where A is the adjacency matrix, X is the node attribute matrix, E is the edge attribute tensor, and S... g For global structural features;
[0093] S43. In each hidden layer l, perform node neighborhood information feature fusion for each node v according to the aggregation function, and update the node feature vector.
[0094] S44. Perform the aggregation operation and feature transformation described in S43 sequentially on all hidden layers to finally obtain the feature vector of each node in the last layer. Where L is the number of layers in the graph neural network model;
[0095] S45. Employ graph-level readout operations to extract the last layer feature vector of each node. Global pooling is used to obtain the embedding features z of each software sample. G ;
[0096] S46. Input the embedded features of each software sample into the output layer, perform a linear transformation on the embedded features through a fully connected layer to obtain the score of each category, and then input the score of each category into the Softmax function to normalize all category scores and convert them into a probability distribution. The probability value represents the likelihood of each software sample belonging to each category. Finally, output the preliminary classification result of each software sample according to the category corresponding to the highest probability.
[0097] In this embodiment, S5 specifically includes:
[0098] S51. Define the loss function The loss function combines the cross-entropy loss between the initial classification result of each software sample and the true class label with a regularization term:
[0099]
[0100] Where M is the number of samples, C is the number of categories, and y ic For real labels, To predict the probability, λ is the regularization coefficient. θ is the regularization term, and θ is the vector of parameters to be optimized.
[0101] S52, the weight parameters W of the graph neural network model (l) Multi-scale neighborhood aggregation radius hyperparameter r (l) Heterogeneous edge type weight allocation coefficient α t and node-edge interaction gating weight parameter β (l) The parameters to be optimized are uniformly composed into a vector θ.
[0102] S53. Employing a sparse approximation and low-rank update mechanism, only the Hessian matrix approximation B is retained in each iteration. k Principal components or significantly nonzero elements:
[0103]
[0104] Among them, B k+1 Let y be the approximate value of the Hessian matrix after the (k+1)th iteration. k Let s be the difference between the gradients of the loss function with respect to the parameter vector θ at the k-th and k+1-th iterations. k The difference between the parameter vector θ to be optimized in the k-th and k+1-th iterations;
[0105] S54. Employ a block gradient collaborative mechanism to adjust the current parameter vector θ. k gradient Grouping calculations are performed, dividing all parameters into several sub-blocks. Each sub-block independently calculates its local gradient and local sparse low-rank Hessian matrix approximation. At the global level, the optimization directions of all sub-blocks are merged through weighted summation or principal component aggregation to obtain the global optimization direction d. k ;
[0106] S55. Adopt an adaptive line search and dynamic step size strategy. Based on the current rate of decline of the loss function, the change of the gradient, and the optimization results of the previous round of graph neural network model, dynamically evaluate and select the step size of the parameter update in this round. When the loss function decreases significantly, the step size is increased appropriately. When the loss function fluctuates or oscillates, the step size is decreased appropriately. The step size adjustment process combines threshold judgment and historical step size reference.
[0107] S56. Based on the obtained global optimization direction and the determined step size, adjust the parameter vector to be optimized item by item. For each weight parameter and structural hyperparameter, according to the increasing or decreasing trend indicated by the optimization direction, add the corresponding step size value in sequence. After all parameters are updated, immediately use them for the next round of graph neural network forward propagation and performance evaluation.
[0108] S57. Apply the updated parameters to the graph neural network model, perform forward propagation, obtain the embedding features and preliminary classification results for each software sample, and recalculate the loss function and graph neural network model performance metrics.
[0109] S58. Based on the performance metrics of the graph neural network model, namely accuracy, loss value, and recall, dynamically adjust the structure of the graph neural network model, including adjusting the number of network layers, multi-scale neighborhood aggregation radius, number of node samples, number of hidden layer units, and activation function.
[0110] S59. Parallel and distributed quasi-Newton optimization mechanisms are adopted to divide the parameter space or training data into multiple sub-blocks. The gradients and approximations of the Hessian matrix of each sub-block are calculated synchronously or asynchronously in a multi-core or distributed computing environment, and the results are updated through distributed synchronous fusion.
[0111] S510 continuously executes steps such as sparse approximation and low-rank update, adaptive line search and dynamic step size, block gradient collaboration and parallel distributed optimization, and iteratively optimizes parameters and structure until the loss function converges or reaches the preset number of training rounds, and finally obtains the optimal graph neural network model parameters and structure configuration.
[0112] In this embodiment, S6 specifically includes:
[0113] S61. Collect new sample data of the software to be tested, including static behavior data and dynamic behavior data, and construct the corresponding software graph structure;
[0114] S62. Standardize and feature-encode the constructed software graph structure, extract node attributes, edge attributes and global structural features to form multi-dimensional graph data input samples;
[0115] S63. Input the obtained multi-dimensional graph data input samples into the graph neural network model that has been optimized in terms of parameters and adjusted in terms of structure;
[0116] S64. Using the optimized graph neural network model, perform forward propagation on the input multi-dimensional graph data input samples to obtain the embedding features and preliminary classification results of each new software sample;
[0117] S65. Based on the embedding features and preliminary classification results of each new software sample, output the malware detection results to achieve intelligent identification of unknown, mutated and confused samples.
[0118] Example 1:
[0119] To verify the feasibility of this invention in practice, it was applied to the data center of a large financial institution. This data center hosts several key business operations of the institution, including mobile payment, e-banking, risk control systems, and customer information systems, processing millions of sensitive customer data entries daily. In recent years, with the frequent occurrence of new ransomware, remote control Trojans, cryptocurrency mining viruses, and advanced persistent threat attacks targeting financial institutions, traditional methods relying on virus signature database comparison and behavioral rule detection are no longer sufficient to meet real-time defense requirements. Especially when facing obfuscated, mutated, and unknown families of malware, traditional methods suffer from high false positive rates and frequent false negatives, posing a significant risk to the secure operation of business.
[0120] The financial institution's technical team deployed and implemented the detection method proposed in this invention, performing real-time sampling of the data center's business servers and terminals. Static behavioral data included the binary file structure, API call details, resource dependencies, and program execution instruction sequences of the software under test. Dynamic behavioral data covered network communication records, system call logs, memory access records, file read / write behavior, and process operation information during software runtime. This collected multi-dimensional data was standardized to construct a multi-dimensional software graph structure with functions, APIs, files, and key events as nodes, and calls, dependencies, and timing as edges.
[0121] The system uses an optimized graph neural network model to automatically analyze and classify the constructed software graph structure. By setting multi-scale neighborhood aggregation radius, heterogeneous edge type weight allocation coefficients, and node-edge interaction gating weight parameters, the graph neural network achieves accurate capture of complex features and structural information of malicious software. The system uses an improved quasi-Newton method to jointly optimize network parameters and structural hyperparameters, achieving rapid convergence of model parameters and adaptive adjustment of model structure.
[0122] During its actual operation, the system performed daily real-time detection and classification of new samples on financial institution servers and terminals. Over the 30-day monitoring period (November 1, 2024 - November 30, 2024), a total of 36,250 new software samples were processed. The system promptly detected and accurately identified 517 malicious samples, including new ransomware, advanced remote control Trojans, cryptocurrency mining virus variants, and highly obfuscated malicious plugins. Among these, 32 new mutated malware samples were successfully detected, preventing serious data leaks and service interruptions, and receiving high praise from operations personnel and security managers.
[0123] Table 1 Comparison of the actual application performance of the detection system of the present invention and traditional detection platforms.
[0124]
[0125] As shown in Table 1, the malware security detection method based on graph neural networks proposed in this invention exhibits significant advantages in practical application scenarios. The detection system of this invention detected a total of 36,250 samples, achieving an accuracy rate of 99.15%, a recall rate of 98.45%, and a false positive rate of only 0.42%. The average detection time per sample was only 0.95 seconds. This invention not only accurately and efficiently completes real-time detection of malware but also effectively avoids the problems of missed and false positives, ensuring the security and stability of the business system.
[0126] In comparison, traditional detection platforms A, B, and C, when testing the same number of samples, achieved accuracies of 94.83%, 93.56%, and 95.44%, respectively, with recall rates of only 85.79%, 83.12%, and 88.67%. Not only were their overall accuracy and recall relatively low, but their false positive rates were also high, reaching 2.05%, 2.42%, and 1.92%, respectively. The average time for testing a single sample was also significantly longer, at 2.15 seconds, 2.34 seconds, and 1.98 seconds, respectively. This traditional method has significant limitations when dealing with novel variants and confounding samples.
[0127] While manual review and sampling improved the accuracy (97.60%) and recall (92.15%) of the tests, the false alarm rate remained high (1.75%), and the sampling process was time-consuming, averaging 65 minutes per batch, which was difficult to meet the needs of rapid response and full coverage testing.
[0128] In summary, the detection system proposed in this invention outperforms traditional detection methods and manual verification methods in all aspects, demonstrating significant advantages such as high accuracy, high recall, low false alarm rate, and rapid detection. It is also more adaptable to complex security detection scenarios in real-world environments, characterized by a wide variety of malware, high variability, and high response requirements, and has great potential for widespread application.
[0129] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A malware security detection method based on graph neural networks, characterized in that, Includes the following steps: S1. Collect static and dynamic behavior data of the software to be tested, and construct a software graph structure. The nodes of the software graph structure represent functions, APIs, files or behavior events, and the edges represent calling, dependency or time sequence relationships. S2. Standardize and feature-encode the software graph structure, extract node attributes, edge attributes and global structural features to form multi-dimensional graph data input samples; S3. Based on the multi-dimensional graph data input samples, initialize the graph neural network model, set the structural hyperparameters of the graph neural network model, and initialize the weight parameters; S4. Input the multi-dimensional graph data input samples into the initialized graph neural network model, perform forward propagation, and obtain the embedding features and preliminary classification results of each software sample; S5. The quasi-Newton method is used to jointly optimize the weight parameters and structural hyperparameters of the graph neural network model. During the training process of the graph neural network model, the structure of the graph neural network model is dynamically adjusted according to the performance index and the parameters are continuously optimized. S6. Process the new sample data of the software to be detected according to the process of steps S1 to S4, input it into the optimized graph neural network model, output the malware detection results, and realize the intelligent identification of unknown, mutated and confused samples.
2. The malware security detection method based on graph neural networks according to claim 1, characterized in that, Static behavioral data specifically includes the binary code, file structure, API call list, instruction sequence, and resource dependency information of the software under test. Dynamic behavioral data specifically includes system calls, process behavior, network communication, memory operations, and file read / write activities of the software under test during operation.
3. The malware security detection method based on graph neural networks according to claim 1, characterized in that, S2 specifically includes: S21. For each node v in the software graph structure i Extract node attribute vector x i The node attribute vector includes function type, API call frequency, file type, and behavior event type; S22. For each edge e in the software graph structure ij Extracting edge attribute vector e ij The edge attribute vector includes the number of calls, call weight, and the order of call times; S23. Based on the adjacency matrix A∈{0,1} N×N This represents the connection relationships in a software graph structure, where N is the total number of nodes. If node v i With node v j If there is an edge between them, then A ij =1, otherwise A ij =0; S24. For the node attribute matrix X = [x1, x2, ..., x...] N ] T And edge attribute tensor E = [e ij Normalization is performed to ensure that the features of each dimension of the node attribute vector and edge attribute vector have a uniform scale; S25. Extract the global structural features of the software graph structure. g This includes obtaining the number of connections for each node by calculating the degree matrix D, where D is a diagonal matrix, and the element in the i-th row and i-th column is D. ii , representing node v in the software graph structure i The number of directly connected edges, combined with node attribute vectors, edge attribute vectors, and adjacency matrix information, are used to construct a multi-dimensional graph data input sample containing node attributes, edge attributes, and global structural features.
4. The malware security detection method based on graph neural networks according to claim 1, characterized in that, S3 specifically includes: S31. Receive the obtained multi-dimensional graph data input sample. Where A is the adjacency matrix, X is the node attribute matrix, E is the edge attribute tensor, and S... g For global structural features; S32. Construct the basic structure of the graph neural network model. The graph neural network model includes an input layer, several graph neural network hidden layers and an output layer. The input layer receives multi-dimensional graph data input samples. The hidden layers perform aggregation and transformation of node and edge features in sequence. The output layer generates malware detection and classification results. S33. Set an aggregation function for each hidden layer. The aggregation function performs joint aggregation of node attributes and edge attributes to generate node feature vectors. S34. Set the number of layers L of the graph neural network model, and configure the activation function, number of hidden layer units, and number of node samples for each hidden layer. S35. Introduce a multi-scale neighborhood aggregation radius hyperparameter r in each hidden layer. (l) The neighborhood range of node information aggregation at each layer can be dynamically controlled. The aggregation radius of different layers can be set separately, and the initial value can be randomly generated or given by preset rules. S36. Set the heterogeneous edge type weight allocation coefficient α for each edge type in the software graph structure. t In the process of node aggregation, the information contribution of different types of edges is weighted, and the initial value can be set according to prior experience or uniform distribution. S37. Introduce a node-edge interaction gating weight parameter β in each hidden layer. (l) The fusion ratio of node features and edge features in each layer to the final node representation is adaptively adjusted, and the initial value can be randomly set in the range of 0 to 1. S38. Initialize the weight parameters W of each layer in the graph neural network model. (l) W (l) Linear transformations are used for node and edge features, with initial weight parameters randomly initialized using a Gaussian distribution; S39. The hyperparameter r of the multi-scale neighborhood aggregation radius (l) Heterogeneous edge type weight allocation coefficient α t Gating weight parameter β for node-edge interaction (l) Together they are used as structural hyperparameters and weight parameters; S310. Initialize the graph neural network model, which includes the multi-scale neighborhood aggregation radius hyperparameter, heterogeneous edge type weight allocation coefficients, node-edge interaction gating weight parameters, and weight parameters. Then, initialize the node feature vectors output from the last layer. It serves as the input basis for forward propagation and malware sample embedding features and classification.
5. The malware security detection method based on graph neural networks according to claim 1, characterized in that, S4 specifically includes: S41. Receive the initialized graph neural network model, which includes the multi-scale neighborhood aggregation radius hyperparameter r. (l) Heterogeneous edge type weight allocation coefficient α t Node-edge interaction gating weight parameter β (l) and the weight parameters W of each layer (l) ; S42. Input the obtained multidimensional graph data into the sample. The input is fed into the input layer of the initialized graph neural network model, where A is the adjacency matrix, X is the node attribute matrix, E is the edge attribute tensor, and S... g For global structural features; S43. In each hidden layer l, perform node neighborhood information feature fusion for each node v according to the aggregation function, and update the node feature vector. S44. Perform the aggregation operation and feature transformation described in S43 sequentially on all hidden layers to finally obtain the feature vector of each node in the last layer. Where L is the number of layers in the graph neural network model; S45. Employ graph-level readout operations to extract the last layer feature vector of each node. Global pooling is used to obtain the embedding features z of each software sample. G ; S46. Input the embedded features of each software sample into the output layer, perform a linear transformation on the embedded features through a fully connected layer to obtain the score of each category, and then input the score of each category into the Softmax function to normalize all category scores and convert them into a probability distribution. The probability value represents the likelihood of each software sample belonging to each category. Finally, output the preliminary classification result of each software sample according to the category corresponding to the highest probability.
6. The malware security detection method based on graph neural networks according to claim 1, characterized in that, S5 specifically includes: S51. Define the loss function The loss function combines the cross-entropy loss and regularization term between the initial classification result of each software sample and the true class label; S52, the weight parameters W of the graph neural network model (l) Multi-scale neighborhood aggregation radius hyperparameter r (l) Heterogeneous edge type weight allocation coefficient α t and node-edge interaction gating weight parameter β (l) The parameters to be optimized are uniformly composed into a vector θ. S53. Employing a sparse approximation and low-rank update mechanism, only the Hessian matrix approximation B is retained in each iteration. k Principal components or significantly nonzero elements; S54. Employ a block gradient collaborative mechanism to adjust the current parameter vector θ. k gradient Grouping calculations are performed, dividing all parameters into several sub-blocks. Each sub-block independently calculates its local gradient and local sparse low-rank Hessian matrix approximation. At the global level, the optimization directions of all sub-blocks are merged through weighted summation or principal component aggregation to obtain the global optimization direction d. k ; S55. Adopt an adaptive line search and dynamic step size strategy. Based on the current rate of decline of the loss function, the change of the gradient, and the optimization results of the previous round of graph neural network model, dynamically evaluate and select the step size of the parameter update in this round. When the loss function decreases significantly, the step size is increased appropriately. When the loss function fluctuates or oscillates, the step size is decreased appropriately. The step size adjustment process combines threshold judgment and historical step size reference. S56. Based on the obtained global optimization direction and the determined step size, adjust the parameter vector to be optimized item by item. For each weight parameter and structural hyperparameter, according to the increasing or decreasing trend indicated by the optimization direction, add the corresponding step size value in sequence. After all parameters are updated, immediately use them for the next round of graph neural network forward propagation and performance evaluation. S57. Apply the updated parameters to the graph neural network model, perform forward propagation, obtain the embedding features and preliminary classification results for each software sample, and recalculate the loss function and graph neural network model performance metrics. S58. Based on the performance metrics of the graph neural network model, namely accuracy, loss value, and recall, dynamically adjust the structure of the graph neural network model, including adjusting the number of network layers, multi-scale neighborhood aggregation radius, number of node samples, number of hidden layer units, and activation function. S59. Parallel and distributed quasi-Newton optimization mechanisms are adopted to divide the parameter space or training data into multiple sub-blocks. The gradients and approximations of the Hessian matrix of each sub-block are calculated synchronously or asynchronously in a multi-core or distributed computing environment, and the results are updated through distributed synchronous fusion. S510 continuously executes steps such as sparse approximation and low-rank update, adaptive line search and dynamic step size, block gradient collaboration and parallel distributed optimization, and iteratively optimizes parameters and structure until the loss function converges or reaches the preset number of training rounds, and finally obtains the optimal graph neural network model parameters and structure configuration.
7. The malware security detection method based on graph neural networks according to claim 1, characterized in that, S6 specifically includes: S61. Collect new sample data of the software to be tested, including static behavior data and dynamic behavior data, and construct the corresponding software graph structure; S62. Standardize and feature-encode the constructed software graph structure, extract node attributes, edge attributes and global structural features to form multi-dimensional graph data input samples; S63. Input the obtained multi-dimensional graph data input samples into the graph neural network model that has been optimized in terms of parameters and adjusted in terms of structure; S64. Using the optimized graph neural network model, perform forward propagation on the input multi-dimensional graph data input samples to obtain the embedding features and preliminary classification results of each new software sample; S65. Based on the embedding features and preliminary classification results of each new software sample, output the malware detection results to achieve intelligent identification of unknown, mutated and confused samples.