A Fine-Grained Malicious Behavior Location Method and System Based on Semantic Enhancement

Through semantic enhancement-based methods, the features of malicious samples are extracted and fused, and the graph attention autoencoder model is used for reconstruction, which solves the problems of insufficient fine-grained and insufficient semantic refinement of malicious behavior in the prior art, and achieves high-precision and reliable malicious behavior positioning.

CN119830288BActive Publication Date: 2025-06-10NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510309189.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-10
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

The existing malicious behavior positioning methods are insufficient in fine-grained positioning, which cannot meet the needs of virus analysts for specific malicious function positioning, and insufficient refinement of API semantics and context semantics, so optimized matching cannot be achieved.

Method used

Using a semantic enhancement method, by obtaining benign samples and malicious samples, extracting the feature set of semantic embedded API sequences, constructing attribute control flow graphs, optimizing graph structures, and reconstructing them using bidirectional propagation and multiple masks, to achieve fine-grained malicious behavior positioning at the basic block level.

Benefits of technology

It significantly improves the accuracy and accuracy of malicious behavior positioning, and can locate malicious behavior in the absence of labeled data sets, improving the reliability and long-term availability of malicious behavior positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119830288B_ABST
    Figure CN119830288B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of digital information application security technology, and particularly relates to a fine-grained malicious behavior localization method and system based on semantic enhancement. The method includes: extracting features from benign samples and malicious samples to be detected, designing a feature set containing semantic embedded API sequences, encoding through weights to generate node embedding expressions, fusing numerical features and sequence features to construct an attribute control flow graph; optimizing the attribute control flow graph; training a graph attention autoencoder model containing bidirectional propagation and multiple masks based on benign samples; inputting the malicious samples into the trained autoencoder model for reconstruction, calculating the reconstruction error, and realizing fine-grained malicious behavior localization at the basic block level; that is, highly condensing the semantics related to malicious behaviors in malicious samples, fusing feature encoding technology and graph structure optimization strategies, and enabling the trained model to accurately locate the basic blocks of malicious behaviors through a customized loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital information application security, and particularly relates to a fine-grained malicious behavior positioning method and system based on semantic enhancement. Background Art

[0002] With the increasingly severe threat of malware, efficient malware sample analysis has become an important topic in the field of network security. Traditional manual analysis methods are inadequate in malware sample analysis, and their time-consuming and inefficient characteristics limit the rapid response to malware threats. The development of automated analysis technology is crucial for improving the efficiency of malware sample analysis, which can significantly shorten the analysis time and improve the response speed.

[0003] However, the existing automated analysis technology is in its infancy in the application of malware behavior positioning, and due to the limitations of the environment and dataset, the progress in this field is relatively slow. In addition, the existing methods do not sink to vertical applications, fail to fully combine the typical characteristics of malicious behaviors and the business requirements of reverse analysis, and cannot achieve optimized matching for malicious behaviors.

[0004] Therefore, the current malicious behavior analysis methods have the following problems: (1) There are few fine-grained malicious behavior localization methods. Existing methods mainly focus on coarsely distinguishing benign and malicious samples, or classifying malicious samples. Usually, malicious behavior detection is carried out on a file-by-file basis, which cannot meet the needs of virus analysts for locating specific malicious functions. (2) The existing malicious behavior localization methods lack sufficient refinement of API semantics. Methods based on API statistical data have a high false negative rate. For example, there are two basic block control flow graphs (a) and (b) with similar structures, which call the CreateDirectory and AddUsersToEncryptedFile APIs respectively. CreateDirectory is used to create a new directory, and AddUsersToEncryptedFile is used to add users to an encrypted file to allow these users to access the file. Both are APIs related to file operations, but obviously AddUsersToEncryptedFile is more sensitive. Methods based on API statistical data will represent the situations of (a) and (b) as the same feature during the feature representation process, ignoring the semantic differences between AddUsersToEncryptedFile and CreateDirectory. While methods based on API call sequences can identify malicious behaviors more accurately, they cannot effectively understand synonymous APIs. For example, CreateSymbolicLink is an infrequently used API that is used to create a symbolic link pointing to an existing file or directory. If this API is not involved in the training dataset, even a benign behavior will be misreported as malicious. This is because the model cannot understand the actual meaning of the API, resulting in difficulty in dealing with rapidly mutating malware and obvious model aging problems. (3) The existing malicious behavior localization methods lack sufficient refinement of context semantics. Existing deep learning-based malicious behavior localization methods transplant mature model frameworks without adapting and optimizing them for the malicious behavior localization scenario, resulting in errors in context semantic refinement. In terms of preprocessing, a rough truncation method is used to process variable-length inputs, without considering that the basic blocks at the truncation positions lose context information; in terms of the model, traditional rule-based machine learning methods cannot cope with rapidly mutating malware, and classic neural networks are based on sequential data or undirected graph data, which do not conform to the context association characteristics of program data, and there are biases in the information aggregation process. Summary of the Invention

[0005] In order to solve the technical problems that the existing deep learning analysis methods fail to fully combine the typical characteristics of malicious behaviors and reverse analysis, lack of refinement of API semantics and context semantics, lack of a malicious behavior localization method of "optimized matching", unable to meet the needs of analysts for fine-grained localization methods, and the localization method lacks refinement of API semantics and context semantics and has poor understanding ability, the purpose of the present invention is to provide a fine-grained malicious behavior localization method based on semantic enhancement, and the specific technical solutions adopted are as follows:

[0006] Obtain benign samples and malicious samples to be detected respectively, and perform feature extraction to design a feature set including semantic-embedded API sequences, encode through weights to generate node embedding expressions, fuse numerical features and sequence features, and construct an attribute control flow graph;

[0007] Optimize the attribute control flow graph and preprocess the malicious samples to be detected;

[0008] Train a graph attention autoencoder model including bidirectional propagation and multiple masks based on benign samples;

[0009] Input the preprocessed malicious samples into the trained autoencoder model for reconstruction, calculate the reconstruction error, and realize fine-grained malicious behavior localization at the basic block level.

[0010] Preferably, obtain benign samples and malicious samples to be detected respectively, perform feature extraction based on the two samples to obtain a feature set including semantic-embedded API sequences, encode through weights to generate node embedding expressions, fuse numerical features and sequence features, and construct an attribute control flow graph, including:

[0011] Extract numerical features and sequence features based on benign samples and malicious samples to be detected, perform semantic embedding on the sequence features to obtain variable-length API call sequences, denoted as sequence feature vectors;

[0012] Calculate the difference value between any two nodes, denoted as the distance between nodes, count the frequency of repetition of each feature performance as the weight, cluster all nodes through the weight to obtain several clusters, generate a codebook according to the nodes closest to the corresponding cluster centroids in each cluster, and perform encoding;

[0013] Construct an attribute control flow graph, denoted as G=(V,A,X), where G represents the attribute control flow graph; V represents the feature set; A represents the adjacency matrix of the attribute control flow graph; X represents the node feature matrix.

[0014] Preferably, the numerical features include the number of basic arithmetic instructions, the number of logical operation instructions, the number of shift instructions, the number of stack operation instructions, the number of register operation instructions, and the number of port operation instructions; the sequence features include API call sequences.

[0015] Preferably, the difference value of the computing nodes is calculated, and the corresponding calculation formula is:

[0016]

[0017] wherein, v i and v j represent any two nodes; N represents the number of numerical features; f i a and respectively represent the numerical feature vectors corresponding to the nodes v i and the nodes v j ; f i b and respectively represent the sequence feature vectors corresponding to the nodes v i and the nodes v j ; represents the edit distance between the two sequence feature vectors f i b and ; len(f i b ) and respectively represent the lengths of the sequences f i b and the sequence .

[0018] Preferably, the attribute control flow graph is optimized, and the malicious samples to be detected are preprocessed, including:

[0019] Pruning the unimportant nodes with all features being 0, removing the corresponding nodes and the out-edges and in-edges representing dependencies, and supplementing new edges to connect all the predecessor nodes and successor nodes of the pruned nodes;

[0020] Fusing the consecutive basic block nodes of the sequential structure;

[0021] Obtaining subgraphs of equal size through subgraph sampling.

[0022] Preferably, training a graph attention autoencoder model including bidirectional propagation and multiple masks based on the benign samples, including:

[0023] Transmitting the benign samples into the graph attention autoencoder model including the bidirectional propagation mechanism and the multiple re-masking mechanisms, and designing the initial loss function, and the corresponding calculation formula is:

[0024]

[0025] wherein, represents the initial loss function after the j-th re-masking; x i and x mrespectively represent the feature vectors corresponding to the \(i\)-th node \(v\) i and the \(m\)-th node \(v\) m ; respectively represent the reconstructed feature vectors output after the \(j\)-th re-masking of node \(v\) i and node \(v\) m and being operated on by the decoder; \(n\) represents the number of nodes in the attribute control flow graph input to the model; \(\gamma\) and \(\lambda\) both represent hyperparameters; \(L\) 2 represents the regularization term; \(W\) q represents the weight matrix of the \(q\)-th layer in the graph attention autoencoder model containing a bidirectional propagation mechanism and multiple re-masking mechanisms;

[0026] The average value is obtained through the number of times of multiple masking, and the initial loss function is corrected to obtain the total loss function. The corresponding calculation formula is:

[0027]

[0028] where represents the total loss function; \(k\) represents the number of times of multiple masking; represents the initial loss function after the \(j\)-th re-masking.

[0029] Preferably, the preprocessed malicious samples are input into the trained autoencoder model for reconstruction, and the reconstruction error is calculated to achieve fine-grained malicious behavior localization at the basic block level, including:

[0030] The preprocessed malicious samples are input into the trained autoencoder model for reconstruction, and the reconstruction error is calculated. The corresponding calculation formula is:

[0031]

[0032] where \(r\) represents the reconstruction error; \(k\) represents the number of times of multiple masking; \(x\) i and \(x\) m respectively represent the feature vectors corresponding to the \(i\)-th node \(v\) i and the \(m\)-th node \(v\) m ; respectively represent the reconstructed feature vectors output after the \(j\)-th re-masking of node \(v\) i and node \(v\) m and being operated on by the decoder; \(n\) represents the number of nodes in the attribute control flow graph input to the model; \(\gamma\) and \(\lambda\) both represent hyperparameters; \(W\) q represents the weight matrix of the \(q\)-th layer in the graph attention autoencoder model containing a bidirectional propagation mechanism and multiple re-masking mechanisms;

[0033] Set a judgment threshold. When the reconstruction error is greater than the judgment threshold, it indicates that the basic block corresponding to the current analysis node is a malicious basic block; when the reconstruction error is less than the judgment threshold, it indicates that the basic block corresponding to the current analysis node is a benign basic block.

[0034] To solve the above technical problems, the present invention provides another technical solution as follows: A fine-grained malicious behavior localization system based on semantic enhancement, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of a fine-grained malicious behavior localization method based on semantic enhancement as described in any one of the foregoing.

[0035] The present invention has the following beneficial effects:

[0036] 1. The present invention adopts weight-based clustering feature encoding, fuses numerical features and sequence features of different scales, and effectively solves the problems of single feature processing and inability to fully utilize API semantic information in traditional methods; analyzes the basic blocks in malicious samples for feature analysis and modeling, realizes the basic block-level localization of malicious behaviors, significantly improves the accuracy and precision of malicious behavior localization, solves the deficiencies of existing methods in fine-grained localization, and provides a more accurate malicious function localization tool for virus analysts; trains a graph attention autoencoder model containing bidirectional propagation and multiple masks with benign samples, enables the model to learn the rules of benign samples, can effectively restore the code information flow, enhances the understanding of malicious behaviors, and can locate malicious behaviors according to the difference between good and evil in the absence of a labeled dataset, improving the reliability of malicious behavior localization; at the same time, through variable-length output processing technology, quickly narrows the localization range, adapts to the fixed-length input requirement, and improves the efficiency of malicious sample analysis; in addition, by performing semantic embedding on the API call sequence, combined with pruning and subgraph sampling, it can improve the ability to cope with newly emerging malicious behaviors and variants, reduce the dependence on frequently updated models, and improve the long-term usability and stability of the malicious behavior localization method.

[0037] 2. The present invention also provides a fine-grained malicious behavior localization system based on semantic enhancement for implementing the fine-grained malicious behavior localization method provided above. This system has the same beneficial effects as the above-mentioned fine-grained malicious behavior localization method based on semantic enhancement and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0039] Figure 1 Schematic flowchart of a fine-grained malicious behavior localization method based on semantic enhancement provided by an embodiment of the present invention;

[0040] Figure 2 Schematic flowchart of feature extraction and encoding of a fine-grained malicious behavior localization method based on semantic enhancement provided by an embodiment of the present invention;

[0041] Figure 3 Schematic structural diagram of a graph attention autoencoder model with bidirectional propagation and multiple masks in a fine-grained malicious behavior localization method based on semantic enhancement provided by an embodiment of the present invention. Detailed implementation manners

[0042] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, describe in detail a fine-grained malicious behavior localization method and system based on semantic enhancement proposed according to the present invention, including its specific implementation manners, structures, features, and effects. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0044] The following will specifically describe the specific solutions of a fine-grained malicious behavior localization method and system based on semantic enhancement provided by the present invention with reference to the accompanying drawings.

[0045] Due to the rapid mutation ability and evasion means of malware, traditional static analysis methods rely on rich prior knowledge, are limited by fixed rules, have rapidly aging models, and fail to fully combine the typical characteristics of malicious behaviors and the business requirements of reverse analysis. They lack sufficient refinement of API semantics and context semantics and cannot achieve "optimized matching" for malicious behavior localization. The first embodiment of the present invention provides a fine-grained malicious behavior localization method based on semantic enhancement. This method, through feature encoding and graph structure optimization, uses a graph attention autoencoder model with a bidirectional propagation mechanism and multiple re-masking mechanisms trained on benign samples to highly condense the semantics related to malicious behaviors in malicious samples, enabling the model to effectively distinguish the benign part and the malicious part in malicious samples, accurately locate the basic blocks of malicious behaviors, and achieve rapid diagnosis of malicious samples, thereby improving the automation and accuracy of static malicious sample analysis. To solve the fine-grained malicious behavior localization method based on semantic enhancement, the second embodiment of the present invention provides a fine-grained malicious behavior localization system based on semantic enhancement. This system is essentially a software system composed of units that implement corresponding functions. Now, the specific steps in this method will be introduced in detail.

[0046] Please refer to Figure 1 , which shows a schematic flowchart of a fine-grained malicious behavior localization method based on semantic enhancement provided by an embodiment of the present invention. The method includes:

[0047] Step S1: Obtain benign samples and malicious samples to be detected respectively, and perform feature extraction to design a feature set including semantic-embedded API sequences, encode them through weights to generate node embedding expressions, fuse numerical features and sequence features, and construct an attribute control flow graph;

[0048] Step S2: Optimize the attribute control flow graph and preprocess the malicious samples to be detected;

[0049] Step S3: Train a graph attention autoencoder model with bidirectional propagation and multiple masks based on benign samples;

[0050] Step S4: Input the preprocessed malicious samples into the trained autoencoder model for reconstruction, calculate the reconstruction error, and achieve fine-grained malicious behavior localization at the basic block level.

[0051] It can be noted that with the development of the mobile Internet and the popularization of intelligent terminal devices, malicious behaviors are becoming increasingly common during program operation, and the user group affected by malicious behaviors on terminal devices is gradually expanding. This not only disrupts the normal order of the network market but also poses a threat to users. That is, most malicious software implants a large number of advertisements, affecting the normal use of users. Some malicious software even implants some hidden malicious codes, affecting the information security of users; or uses application vulnerabilities to write malicious scripts for the purpose of making money or stealing privacy, threatening the daily life and property security of users. The methods proposed in the market have certain limitations, resulting in low accuracy in locating malicious behaviors in program codes, errors in semantic refinement, and inability to meet the application requirements of "optimized matching". Therefore, the technical solution of this embodiment is proposed to improve the accuracy and precision of malicious behavior location and provide a more efficient and reliable solution for malicious behavior location.

[0052] For better illustration, a basic block refers to an instruction sequence in a program that has a single entry and exit, without internal jumps. As the basic unit for malicious behavior analysis, it provides basic data support for subsequent graph structure construction and feature extraction; the API (Application Programming Interface) call sequence refers to the interface call sequence for software to interact with the operating system in each basic block, which is used to capture the high-level semantics of software behavior and can reflect the operation intention of the software.

[0053] As an alternative implementation, in this embodiment, a benign sample refers to a normal behavior or harmless data sample that conforms to the normal use and expected behavior of the system and does not contain any malicious intent or potential attack behavior, such as normal program operation, legal user behavior, or normal network communication, etc.; a malicious sample refers to a data sample of potential malicious behavior or known malicious behavior, usually an operation initiated by an attacker or malicious user of the system or network, such as abnormal system calls, unknown code execution paths, suspicious network traffic, or privacy-infringing data tampering, etc.

[0054] Please refer to Figure 2 , which shows a schematic flowchart of feature extraction and encoding of a fine-grained malicious behavior location method based on semantic enhancement provided by an embodiment of the present invention.

[0055] Furthermore, in step S1, it includes:

[0056] Step S11: Extract numerical features and sequence features based on benign samples and malicious samples to be detected, and perform semantic embedding based on the sequence features to obtain an indefinite-length API call sequence, denoted as a sequence feature vector.

[0057] Furthermore, the numerical features include the number of basic arithmetic instructions, the number of logical operation instructions, the number of shift instructions, the number of stack operation instructions, the number of register operation instructions, and the number of port operation instructions; the sequence feature includes the API call sequence.

[0058] Explanation is made that the numerical features refer to the statistics and analysis of the number of specific instructions during the computer program design and execution process, including the number of basic arithmetic instructions such as addition, subtraction, multiplication, and division; the number of logical operation instructions such as AND, OR, NOT, etc.; the number of shift instructions for left or right shifting of bits; the number of stack operation instructions referring to pushing and popping in the stack structure; the number of register operation instructions for directly reading and writing to the internal registers of the processor; the number of port operation instructions referring to data exchange with the input / output ports of computer hardware; the sequence feature, that is, the API call sequence can serve as a key feature to capture high-level system operations required for performing malicious behavior operations, such as network communication, file system interaction, registry access, and process management, etc., thus directly mapping the specific behavior patterns of malware in the underlying operating system or host environment.

[0059] In real life, malware often uses APIs with similar functions but different ones to avoid detection. Therefore, an API semantic knowledge graph is constructed to perform semantic embedding on all the APIs provided by the system to reduce the detection error. That is, the API semantic knowledge graph is constructed using API documentation to strengthen the solution to the technical problem of insufficient API semantic refinement, so as to enhance the understanding of semantics in model training through semantic embedding, ensuring that each API can accurately understand and execute the tasks assigned to it; then all the APIs in the system are clustered to obtain different cluster categories, and the APIs in the same cluster category are regarded as semantically similar, and the specific API names are replaced with the numbers of the cluster categories where the APIs are located to simplify the calculation, and then the variable-length API call sequences after semantic embedding corresponding to any node are determined, denoted as the sequence feature vector of the basic block corresponding to the node.

[0060] Step S12: Calculate the difference value between any two nodes, denoted as the distance between the nodes, count the frequency of repetition of each feature performance as the weight, cluster all the nodes through the weight to obtain several clusters, generate a codebook according to the node closest to the centroid of the corresponding cluster in each cluster, and perform encoding.

[0061] Furthermore, the formula for calculating the difference value of the nodes is as follows:

[0062]

[0063] where, v i 、v j represent any two nodes; N represents the number of numerical features; fi a , respectively represent the numerical feature vectors corresponding to node v i and node v j ; f i b , respectively represent the sequence feature vectors corresponding to node v i and node v j ; represents the edit distance between f i b , two sequence feature vectors; len(f i b ) and respectively represent the lengths of sequence f i b and sequence .

[0064] It is explained that N represents the number of numerical features, and specifically, the value range of i is [1, N]; represents the edit distance between f i b , two sequence feature vectors, which refers to how many insertion, deletion, or replacement operations are required for the sequence feature vector f i b to become the sequence feature vector That is, it measures the minimum number of operations required to convert one string to another; len(f i b ) and respectively represent the lengths of sequence f i b and sequence . In the worst case, that is, when the number of operations between the two sequences reaches the maximum possible value, it means that each element has to be modified through insertion, deletion, or replacement until the two sequences are completely different. At this time, the edit distance between the two sequence feature vectors is the sum of the lengths of the two sequence feature vectors.

[0065] Specifically, the frequency of repetition of each feature manifestation is recorded as the weight, and the DBSCAN (Density-Based Spatial Clustering of Applications with Noise, that is, density-based spatial clustering) algorithm based on the weight is used to perform clustering analysis on the nodes to obtain several clusters. In this embodiment, k clusters are obtained according to the sample situation, and the node closest to the centroid of each cluster is taken, that is, v s , s = 1, 2,..., k, to generate the codebook C, ci ∈ C, i = 1, 2, ..., k; Encode the feature of node v as The corresponding logical formula is:

[0066] f i = dis(v, c i ), i = 1, 2, ..., k

[0067] where f i ∈ F, i = 1, 2, …, k represents the difference value dis(v, c i ) between node v and the corresponding node c in the password book i ).

[0068] It can be explained that assuming a specific node v a and its most similar password book node c i , use a k-dimensional vector x a to represent node v a , and this vector is 0 in most dimensions, and only has non-zero values at the positions corresponding to the password book node c i , that is, if node v a is most similar to the password book node c i , then its k-dimensional vector x a contains k - 1 zeros, but for the value of the password book node c i at the i-th dimension is the difference value f a between node v i and the password book node c i ; In addition, according to the actual application scenario, it is also possible to select the nodes in the top three password books most similar to node v a to construct the vector x a . At this time, there are non-zero values in the dimensions corresponding to these three password book nodes in the vector x a , and the remaining dimensions are still 0; In this way, the data structure can be simplified, the encoding processing volume can be reduced, and the most similar password book node to the node can be quickly located through the representation of non-zero values, which speeds up the information retrieval speed and enables quick recognition and classification of similar nodes during the deep learning process.

[0069] Step S13: Construct an attribute control flow graph, denoted as G = (V, A, X), where G represents the attribute control flow graph; V represents the feature set; A represents the adjacency matrix of the attribute control flow graph; X represents the node feature matrix.

[0070] Make an explanation, v i ∈ V, i ∈ (1, n), n represents the number of nodes; represents the adjacency matrix of the attribute control flow graph; ​Represents the node feature matrix, where n represents the dimension of the input node feature, that is, the length of the codebook.

[0071] Furthermore, step S2 includes:

[0072] Step S21: Prune unimportant nodes whose features are all 0, remove the corresponding nodes and the outgoing and incoming edges representing the dependencies, and add new edges to connect all the predecessor nodes and successor nodes of the pruned nodes.

[0073] To clarify, pruning refers to reducing the complexity of data by removing redundant, irrelevant or unimportant parts of the data, eliminating unnecessary nodes or edges, and highlighting the key behaviors of the program; specifically, first remove unimportant nodes whose features are all 0, and their outgoing and incoming edges that represent dependencies, that is, according to the program dependencies, remove nodes and redundant paths that are irrelevant as a whole or irrelevant to the target behavior to retain valuable information; then supplement the pruned nodes to maintain the integrity and connectivity of the network structure.

[0074] Step S22: Merge consecutive basic block nodes of the sequential structure.

[0075] It is explained that the sequential structure refers to the nodes in the program that are executed sequentially from top to bottom, and the continuous basic block nodes represent a continuous code sequence that does not contain jump instructions in the attribute control flow graph; specifically, multiple adjacent basic blocks that do not contain branch or jump instructions are merged into a single basic block, and the merger does not destroy the original program logic, thereby reducing the number of nodes in the attribute control flow graph and improving the efficiency of model optimization to simplify the attribute control flow graph.

[0076] Step S23: obtaining sub-images of equal size through sub-image sampling.

[0077] Preferably, in this embodiment, the shaDow-GNN (Shadow Graph Neural Network) method is used for subgraph sampling to process graphs of different sizes of the attribute control flow graph into subgraphs of the same size, solve the problem of variable-length input data, and provide a more compact and efficient graph structure for model training and reasoning; the subgraph contains sufficient information to support accurate learning and avoid the technical problem of insufficient semantic extraction; that is, subgraph sampling is to extract a subset from the entire graph in the attribute control flow graph, and the subset can be analyzed as an independent graph.

[0078] It is explained that shaDow-GNN is a graph neural network algorithm used to solve the problems of neighborhood explosion and global graph over-smoothing in the process of learning large-scale graph data. First, an IID (Independent and Identically Distributed) node sampler randomly selects nodes from the entire graph to ensure the independent and identical distribution of samples, reduce bias, and enhance the generalization ability of the model. Second, a k-hop sampler expands from the selected central node to its k-hop neighbors to prevent neighborhood explosion by controlling the receptive field size of each node, where k represents the maximum number of hops from the central node. In addition, a PPR (Personalized PageRank) sampler uses a personalized version based on PageRank to assign probability weights according to the importance of specific nodes, making the sampling process tend to select neighbor nodes more important to the central node and retaining key information flow paths. Finally, a randomized version of the PPR sampler, as a variant of the PPR sampler, further enhances the diversity of sampling by introducing randomness, which helps to avoid overfitting and improve the robustness of the model. By jointly applying the above components, shaDow-GNN can ensure that the subgraph contains sufficient information while avoiding the problems of global graph over-smoothing and neighborhood explosion.

[0079] Please refer to Figure 3 , which shows a schematic structural diagram of a graph attention autoencoder model with two-way propagation and multiple masks for a fine-grained malicious behavior localization method provided by an embodiment of the present invention.

[0080] It can be understood that the two-way propagation mechanism is an optimization based on the traditional propagation mechanism, that is, inside the original encoder and decoder, it is divided into two branches, forward propagation and backward propagation, to independently aggregate information until the last layer of the encoder and decoder, and then the two branches of the two-way propagation are re-aggregated. In the branch structure of the attribute control flow graph, different branches cannot be executed simultaneously at runtime. Among them, loops or other situations that cause the reuse of a certain code block are regarded as multiple runs. Therefore, it is not in line with the actual requirements for different branches under the same node to collect each other's information. Therefore, the adjacency matrix of the directed graph and the transpose of the adjacency matrix are used, that is, the two-way propagation mechanism is used as two independent information aggregation routes respectively. Among them, the forward propagation and the backward propagation share weights. The forward propagation uses the real adjacency matrix of the directed graph, and the backward propagation uses the transpose of the adjacency matrix to achieve the reverse control flow propagation of information, which can not only solve the problem of context connection but also effectively avoid the information blending of different branches.

[0081] Specifically, in the two-way propagation mechanism, it is possible to specifically collect the node information in the same control flow path, and the corresponding logical formula is:

[0082]

[0083] Among them, represents the output of the forward propagation of the l-th layer, represents the output of the backward propagation of the l-th layer; A represents the adjacency matrix; A T represents the transpose of the adjacency matrix; W (l) represents the weight of the l-th layer; then, the last layers of the encoder and decoder merge the information independently aggregated on each node, and the corresponding logical formula is:

[0084] H = H f + H b

[0085] Among them, H represents the encoding, that is, the final output of the graph attention autoencoder model containing the bidirectional propagation mechanism and the multiple re-masking mechanism, which corresponds to the input of the second masking in the multiple masking mechanism.

[0086] It can be explained that the multiple re-masking mechanism is a technology for dealing with the situation where the target is occluded. When the target is partially or completely occluded by non-target parts, the multiple re-masking mechanism can help the system identify, track, and reconstruct the object.

[0087] Specifically, this mechanism performs two maskings on the data when inputting the encoder and when the encoded H is input to the decoder. Among them, the first masking is before the data is input to the encoder, and is randomly selected for masking. When masking, the adjacency matrix A is not changed, and the masking is performed on the node feature matrix X. After masking, is obtained, so that the decoder cannot obtain complete data information, so as to prompt it to better learn the features of benign samples and master the change rules of benign samples; the second masking is before the encoded H is input to the decoder, and k nodes are randomly selected for re-masking on the encoded H j ∈ {1, k} represents the masked node selected for the j-th time. After masking, is obtained, that is, the multiple mask technology is used to simulate the situation where some nodes or edges are missing, and the model's ability to capture local and global information is enhanced.

[0088] Then, the graph attention mechanism is used to enable the graph attention autoencoder model containing bidirectional propagation and multiple masks to automatically learn the importance weights assigned to different neighbor nodes, so as to be more efficient and accurate when aggregating neighbor information.

[0089] Optionally, in this embodiment, the calculation method of additive attention coefficients is used; specifically, the node pair (i, j) is defined, where i represents the target node and j represents the neighbor node corresponding to the target node, and the importance between the two nodes of the node pair is calculated. The corresponding calculation formula is:

[0090] eij = atten([Wh i || Wh j )

[0091] where h i and h j represent the feature vectors of nodes i and j respectively; W represents the shared weight matrix for linearly transforming the node feature vectors; the atten function is a single-layer feed-forward neural network that takes the input [Wh i || Wh j to perform a concatenation operation, concatenating the features of nodes i and j after linear transformation to form a new vector, i.e., implementing the similarity calculation in additive attention and obtaining the importance of all neighbor nodes of the target node being currently analyzed. The softmax function is used to convert all importances into a probability distribution to obtain the attention weight α ij , and the corresponding calculation formula is:

[0092]

[0093] where α ij represents the attention weight, i.e., the attention coefficient of the final target node i to the neighbor node j, reflecting the importance of the neighbor node to the target node; represents the set of neighbor nodes of the currently analyzed node i. In the forward propagation branch In the backward propagation branch The softmax function can ensure that the sum of the attention coefficients of all neighbor nodes of each target node i is 1, and regard the attention coefficients as a probability distribution, which is conducive to screening out the target information from the input information, reducing data overload, decreasing the attention to irrelevant information, and even filtering out irrelevant information, improving the efficiency and accuracy of task processing.

[0094] It can be explained that after training the graph attention autoencoder model with bidirectional propagation and multiple masks using benign samples, the autoencoder model can better master the change rules of benign samples, that is, through learning benign samples, the model can identify the features and patterns of benign behaviors, and then more quickly identify the features and patterns of malicious behaviors, improving the reliability of malicious behavior localization; it can also effectively restore the code information flow, capture the complex relationships between program codes, enhance the model's understanding of malicious behaviors, and solve the problem of fine-grained localization.

[0095] Specifically, then the obtained after multiple re-masking is feature reconstructed through the decoder to obtain the corresponding The initial loss function between The value, which is the average of the loss function values taken k times is used as the total loss function of the attribute control flow graph.

[0096] Furthermore, in step S3, it includes:

[0097] Step S31: Transmit the benign samples to the graph attention autoencoder model with a two-way propagation mechanism and a multiple re-masking mechanism, and design an initial loss function. The corresponding calculation formula is:

[0098]

[0099] where represents the initial loss function after the j-th re-masking; x i and x m respectively represent the feature vectors corresponding to the i-th node v i and the m-th node v m ; respectively represent the reconstructed feature vectors of the node v i and the node v m after the j-th re-masking and calculated by the decoder; n represents the number of nodes in the attribute control flow graph input to the model; γ and λ both represent hyperparameters; L 2 represents the regularization term; W q represents the weight matrix of the q-th layer in the graph attention autoencoder model with a two-way propagation mechanism and a multiple re-masking mechanism.

[0100] It should be noted that represents the mean square error, which is used to evaluate the difference between the predicted value and the true value of the model; n represents the number of nodes in the attribute control flow graph input to the model, that is, the number of rows of the input matrix; represents the range constraint part, which is used to calculate the maximum difference between the actual value x and the reconstructed value x j .

[0101] Step S32: Obtain the average value through the number of multiple masks, and correct the initial loss function to obtain the total loss function. The corresponding calculation formula is:

[0102]

[0103] where represents the total loss function; k represents the number of multiple masks; represents the initial loss function after the j-th re-masking.

[0104] It is explained that by using the mask multiple times to process the loss of nodes and taking the average value, the influence brought by single randomness can be reduced, making the model more stable and improving the robustness and generalization ability of the model; it can also effectively prevent the model from overfitting the data in the benign samples and improve the interpretability of the model.

[0105] Further, in step S4, it includes:

[0106] Input the preprocessed malicious samples into the trained autoencoder model for reconstruction, and calculate the reconstruction error. The corresponding calculation formula is:

[0107]

[0108] where r represents the reconstruction error; k represents the number of multiple masks; x i and x m respectively represent the feature vectors corresponding to the i-th node v i and the m-th node v m ; respectively represent the reconstructed feature vectors output after the j-th re-masking of the node v i and the node v m and calculated by the decoder; n represents the number of nodes in the attribute control flow graph input to the model; γ and λ both represent hyperparameters; W q represents the weight matrix of the q-th layer in the graph attention autoencoder model including the bidirectional propagation mechanism and the multiple re-masking mechanism;

[0109] Set a judgment threshold. When the reconstruction error is greater than the judgment threshold, it indicates that the basic block corresponding to the current analysis node is a malicious basic block; when the reconstruction error is less than the judgment threshold, it indicates that the basic block corresponding to the current analysis node is a benign basic block.

[0110] Optionally, denote the judgment threshold as δ, which represents the boundary error value for distinguishing benign samples and malicious samples and is determined according to the actual situation of the sample data. For example, in a certain detection of malicious behavior, the reconstruction error of any basic block is 0.735. At this time, the corresponding judgment threshold δ = 0.4, that is, this basic block is a malicious basic block, and the location of this malicious basic block can be obtained through its corresponding node. Conversely, if the reconstruction error of this basic block is less than δ, it indicates that this basic block is a benign basic block, indicating that the current program is running well.

[0111] Understandably, the present invention adopts weight-based clustering feature encoding, which integrates numerical features and sequence features of different scales, effectively solving the problems of single feature processing and inability to fully utilize API semantic information in traditional methods; analyzes basic blocks in malicious samples for feature analysis and modeling, realizes basic block-level positioning of malicious behaviors, significantly improves the accuracy and precision of malicious behavior positioning, solves the deficiencies of existing methods in fine-grained positioning, and provides a more accurate malicious function positioning tool for virus analysts; trains a graph attention autoencoder model with bidirectional propagation and multiple masks using benign samples, enabling the model to learn the rules of benign samples, effectively restoring the code information flow, enhancing the understanding of malicious behaviors, and being able to locate malicious behaviors based on the differences between benign and malicious samples in the absence of a labeled dataset, improving the reliability of malicious behavior positioning; at the same time, through variable-length output processing technology, quickly narrows the positioning range, adapts to the fixed-length input requirement, and improves the efficiency of malicious sample analysis; in addition, by performing semantic embedding on the API call sequence, combining pruning and subgraph sampling, the ability to handle newly emerging malicious behaviors and variants can be improved, reducing the dependence on frequently updated models, and improving the long-term usability and stability of the malicious behavior positioning method.

[0112] An embodiment of the present invention also proposes a fine-grained malicious behavior positioning system based on semantic enhancement, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of a fine-grained malicious behavior positioning method based on semantic enhancement as described in any one of the foregoing embodiments.

[0113] It should be noted that a fine-grained malicious behavior positioning system based on semantic enhancement has the same beneficial effects as the foregoing provided fine-grained malicious behavior positioning method based on semantic enhancement, and will not be elaborated here.

[0114] It should be noted that: the above-mentioned sequence of embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0115] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments.

Claims

1. A fine-grained malicious behavior location method based on semantic enhancement, characterized in that: The method comprises: We obtain benign samples and malicious samples to be detected respectively, and perform feature extraction to design a feature set containing semantically embedded API sequences. We encode them through weights, generate node embedding expressions, fuse numerical features and sequence features, and construct an attribute control flow graph, including: Based on benign samples and malicious samples to be detected, numerical features and sequence features are extracted, and semantic embedding is performed based on the sequence features to obtain an indefinite length API call sequence, which is recorded as a sequence feature vector; Calculate the difference between any two nodes and record it as the distance between the nodes. Count the frequency of repetition of each feature as the weight. Cluster all nodes by weight to get several clusters. Generate a codebook based on the node closest to the corresponding cluster centroid in each cluster and encode it. Construct an attribute control flow graph, denoted as G = (V, A, X), where G represents the attribute control flow graph; V represents the feature set; A represents the adjacency matrix of the attribute control flow graph; X represents the node feature matrix; Optimize the attribute control flow graph and pre-process the malicious samples to be detected, including: Prune unimportant nodes whose features are all 0, remove the corresponding nodes and the outgoing and incoming edges that represent dependencies, and add new edges to connect all predecessor nodes and successor nodes of the pruned nodes; Fusing consecutive basic block nodes of sequential structures; Obtain subgraphs of equal size through subgraph sampling; Based on benign samples, we train a graph attention autoencoder model with bidirectional propagation and multiple masks, including: The benign samples are transferred to the graph attention autoencoder model including the bidirectional propagation mechanism and the multiple re-masking mechanism, and the initial loss function is designed. The corresponding calculation formula is: in, represents the initial loss function after the jth re-masking; x i 、x m Respectively represent the i-th node v i and the mth node v m The corresponding eigenvector; Respectively represent the node v i and node v m The reconstructed feature vector output after the jth re-masking and decoder operation; n represents the number of nodes in the attribute control flow graph in the input model; γ and λ both represent hyperparameters; L2 represents the regularization term; W q represents the weight matrix of the qth layer in the graph attention autoencoder model that includes a bidirectional propagation mechanism and a multiple re-masking mechanism; The average value is obtained by the number of multiple masks, and the initial loss function is corrected to obtain the total loss function. The corresponding calculation formula is: in, represents the total loss function; k represents the number of multiple masking; represents the initial loss function after the jth re-masking; The preprocessed malicious samples are input into the trained autoencoder model for reconstruction, and the reconstruction error is calculated to achieve fine-grained malicious behavior positioning at the basic block level, including: The preprocessed malicious samples are input into the trained autoencoder model for reconstruction, and the reconstruction error is calculated. The corresponding calculation formula is: Where r represents the reconstruction error; k represents the number of multiple masking; x i 、x m Respectively represent the i-th node v i and the mth node v m The corresponding eigenvector; Respectively represent the node v i and node v m The reconstructed feature vector output after the jth re-masking and decoder operation; n represents the number of nodes in the attribute control flow graph in the input model; γ and λ both represent hyperparameters; W q represents the weight matrix of the qth layer in the graph attention autoencoder model that includes a bidirectional propagation mechanism and a multiple re-masking mechanism; A judgment threshold is set. When the reconstruction error is greater than the judgment threshold, it indicates that the basic block corresponding to the current analysis node is a malicious basic block; when the reconstruction error is less than the judgment threshold, it indicates that the basic block corresponding to the current analysis node is a benign basic block.

2. According to the semantic enhancement-based fine-grained malicious behavior positioning method of claim 1, it is characterized in that: The numerical features include the number of basic arithmetic instructions, the number of logical operation instructions, the number of displacement instructions, the number of stack operation instructions, the number of register operation instructions and the number of port operation instructions; the sequence features include the API call sequence.

3. According to the semantic enhancement-based fine-grained malicious behavior positioning method of claim 2, it is characterized in that: Calculate the difference value of the node, the corresponding calculation formula is: Among them, v i 、v j represents any two nodes; N represents the number of numerical features; f i a 、f j a Respectively represent the node v i and node v j The corresponding numerical eigenvector; f i b 、f j b Respectively represent the node v i and node v j The corresponding sequence feature vector; ld(f i b ,f j b ) indicates f i b 、f j b The edit distance between two sequence feature vectors; len(f i b )、len(f j b ) represent the sequence f i b and the sequence f j b Length.

4. A semantically enhanced fine-grained malicious behavior positioning system, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of a fine-grained malicious behavior positioning method based on semantic enhancement as described in any one of claims 1 to 3 are implemented.

Citation Information

Patent Citations

  • Binary malicious code feature analysis method, system, equipment and medium

    CN117725579A

  • Dynamic malicious software detection method based on enhanced semantic API sequence features

    CN118656827A